An image generation method, apparatus, electronic device, and computer-readable medium
By correcting the contour description data of the first object to match the pose data of the second object, the effect of simultaneously adjusting the pose of the object and maintaining the object features in image generation is achieved, which solves the problem of poor pose and feature consistency in the prior art and improves the image generation quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2023-11-21
- Publication Date
- 2026-06-23
AI Technical Summary
Existing image generation techniques are not very effective in maintaining the consistency of object pose and features, especially in face image generation, where face shape and hairstyle are poorly preserved and pose adjustment is not ideal.
By acquiring the first object contour description data and pose description data, the first object contour description data is corrected using the second object pose description data to generate corrected contour description data, and the image is generated by combining the second object pose data to ensure that the generated image maintains the consistency of object pose and other features.
It improves the adjustment effect of object pose during image generation and maintains the consistency of object features such as face shape and hairstyle, thereby enhancing the image generation effect.
Smart Images

Figure CN120031989B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to an image generation method, apparatus, electronic device, and computer-readable medium. Background Technology
[0002] In some image generation scenarios, there may be a need to generate a new image based on certain information, ensuring that the new image remains consistent with that information. For ease of understanding, examples are provided below.
[0003] As an example, for some face image generation scenarios, there may be the following requirement: to generate a new image based on some face state description information so that the face state represented by the new image is consistent with the face state described by the face state description information. Summary of the Invention
[0004] This application provides an image generation method, apparatus, electronic device, and computer-readable medium that can achieve the requirements described above.
[0005] To achieve the above objectives, the technical solution provided in this application is as follows:
[0006] This application provides an image generation method, the method comprising:
[0007] Acquire first object contour description data, first object pose description data, and second object pose description data; the first object contour description data and the first object pose description data are both determined based on the first object description image;
[0008] Based on the difference representation data between the second object pose description data and the first object pose description data, the first object contour description data is corrected to obtain corrected contour description data.
[0009] Image generation processing is performed based on the corrected contour description data and the second object pose description data to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature other than the object pose.
[0010] In one possible implementation, the difference between the object pose of the contour described by the corrected contour description data and the object pose described by the second object pose description data is smaller than the difference between the object pose of the contour described by the first object contour description data and the object pose described by the second object pose description data.
[0011] In one possible implementation, the first object contour description data is an image;
[0012] The process of determining the corrected contour description data includes:
[0013] The position offset is determined based on the difference representation data between the second object pose description data and the first object pose description data;
[0014] Based on the position offset, pixel position correction processing is performed on the contour description pixels in the first object contour description data to obtain corrected contour description data.
[0015] In one possible implementation, the first object pose description data is used to describe the position of at least one first key point; the second object pose description data is used to describe the position of at least one second key point; if there is a correspondence between a first target point among the at least one first key point and a second target point among the at least one second key point, then the difference characterization data includes the position difference between the first target point and the second target point; the position offset is determined based on the position difference.
[0016] In one possible implementation, the step of performing pixel position correction processing on the contour description pixels in the first object contour description data based on the position offset to obtain corrected contour description data includes:
[0017] According to the stated position offset, the contour description pixels in the first object contour description data are shifted to obtain the corrected contour description data.
[0018] In one possible implementation, the second object pose description data is determined based on a second object description image, the object pose described by the second object description image being different from the object pose described by the first object description image.
[0019] In one possible implementation, the second object pose description data is obtained by perturbing the first object pose description data so that there is a difference between the object pose described by the second object pose description data and the object pose described by the first object pose description data.
[0020] In one possible implementation, the first object pose description data is obtained by performing keypoint detection processing on the first object description image using a first keypoint detection algorithm; the second object pose description data is obtained by performing keypoint detection processing on the first object description image using a second keypoint detection algorithm; the second keypoint detection algorithm is different from the first keypoint detection algorithm.
[0021] In one possible implementation, the first object description image includes at least one image to be processed; different images to be processed are used to provide information to be retained under different object features; the first object contour description data includes object contour description data corresponding to each of the images to be processed; the first object pose description data includes object pose description data corresponding to each of the images to be processed.
[0022] The step of correcting the first object contour description data based on the difference representation data between the second object pose description data and the first object pose description data to obtain corrected contour description data includes:
[0023] For any of the images to be processed, based on the difference representation data between the second object pose description data and the object pose description data corresponding to the image to be processed, the object contour description data corresponding to the image to be processed is corrected to obtain the corrected contour description data corresponding to the image to be processed.
[0024] The step of performing image generation processing based on the corrected contour description data and the second object pose description data to obtain a generated image includes:
[0025] Image generation processing is performed based on the pose description data of the second object and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image.
[0026] In one possible implementation, the at least one image to be processed includes a face shape providing image and a hairstyle providing image, wherein the face shape providing image is used to provide information to be retained under the object's face shape, and the hairstyle providing image is used to provide information to be retained under the object's hairstyle.
[0027] In one possible implementation, for any of the images to be processed, the object pose description data corresponding to the image to be processed is determined based on the image to be processed, and the object pose description data corresponding to the image to be processed is used to describe the position of at least one key point in the facial region.
[0028] In one possible implementation, the object outline description data corresponding to the hairstyle image is provided as the hairstyle outline description data corresponding to the hairstyle image.
[0029] In one possible implementation, the object contour description data corresponding to the hairstyle image is the head contour description data corresponding to the hairstyle image; the head contour description data carries face shape contour description information and hairstyle contour description information.
[0030] In one possible implementation, the process of determining the corrected contour description data corresponding to the hairstyle image includes:
[0031] Based on the image segmentation result of the provided hairstyle image, information removal processing is performed on the facial contour description information in the head contour description data corresponding to the provided hairstyle image to obtain the hairstyle contour description data corresponding to the provided hairstyle image.
[0032] Based on the difference representation data between the second object pose description data and the object pose description data corresponding to the hairstyle image, the hairstyle contour description data corresponding to the hairstyle image is corrected to obtain the corrected contour description data corresponding to the hairstyle image.
[0033] In one possible implementation, before performing image generation processing based on the second object pose description data and the corrected contour description data corresponding to the at least one image to be processed to obtain the generated image, the method further includes:
[0034] Based on the image segmentation result of the provided image of the hairstyle, information removal processing is performed on the facial contour description information in the corrected contour description data corresponding to the provided image of the hairstyle to obtain the removed contour description data corresponding to the provided image of the hairstyle.
[0035] The step of performing image generation processing based on the second object pose description data and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image includes:
[0036] The generated image is obtained by performing image generation processing based on the pose description data of the second object, the corrected contour description data corresponding to the face shape image, and the removed contour description data corresponding to the hairstyle image.
[0037] In one possible implementation, the image generation process based on the corrected contour description data and the second object pose description data to obtain a generated image includes:
[0038] The first object description image, the corrected contour description data, and the second object pose description data are input into a pre-constructed graph model to obtain the generated image output by the graph model; the graph model is used to process the first object description image using the corrected contour description data and the second object pose description data as conditions to obtain the generated image.
[0039] In one possible implementation, the image generation process based on the corrected contour description data and the second object pose description data to obtain a generated image includes:
[0040] The noise data, prompt text, the corrected contour description data, and the second object pose description data are input into a pre-constructed text image model to obtain the generated image output by the text image model; the text image model is used to process the noise data using the prompt text, the corrected contour description data, and the second object pose description data as conditions to obtain the generated image.
[0041] In one possible implementation, the at least one object feature includes at least one of the object's face shape and the object's hairstyle.
[0042] This application provides an image generation apparatus, comprising:
[0043] The acquisition unit is used to acquire first object contour description data, first object pose description data, and second object pose description data; the first object contour description data and the first object pose description data are both determined based on the first object description image;
[0044] The correction unit is used to perform correction processing on the first object contour description data based on the difference representation data between the second object posture description data and the first object posture description data to obtain corrected contour description data.
[0045] The generation unit is used to perform image generation processing based on the corrected contour description data and the second object pose description data to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature other than the object pose.
[0046] This application provides an electronic device, the device comprising: a processor and a memory;
[0047] The memory is used to store instructions or computer programs;
[0048] The processor is configured to execute the instructions or computer program in the memory, so that the electronic device performs the image generation method provided in this application.
[0049] This application provides a computer-readable medium storing instructions or a computer program that, when executed on a device, causes the device to perform the image generation method provided in this application.
[0050] This application provides a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the image generation method provided in this application.
[0051] Compared with related technologies, this application has at least the following advantages:
[0052] The technical solution provided in this application first obtains first object contour description data (e.g., a contour image extracted from image 1), first object pose description data (e.g., a facial key point description image extracted from image 1), and second object pose description data (e.g., from image 1). Figure 2 The first object contour description data and the first object pose description data are both determined based on the first object description image (e.g., Image 1), so that the first object contour description data can describe the contours of some object features (e.g., face shape, hairstyle, etc.) that need to be preserved in the first object description image, so that the first object pose description data can describe the object pose corresponding to these object features in the first object description image (e.g., the positions of the five key points used to describe the face pose, etc.), and so that the second object pose description data can describe the object pose required to be presented in the final generated image; then, based on the difference representation data between the second object pose description data and the first object pose description data, the object is... The first object contour description data is corrected to obtain corrected contour description data, so that the object pose presented in the corrected contour description data is closer to the object pose described by the second object pose description data. Then, image generation processing is performed based on the corrected contour description data and the second object pose description data to obtain a generated image, so that the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature (e.g., face shape, hairstyle, etc.) other than the object pose. In this way, the purpose of adjusting the object pose and maintaining some object features can be taken into account during the image generation process, which is beneficial to improving the image generation effect. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 A flowchart illustrating an image generation method provided in this application embodiment;
[0055] Figure 2 A schematic diagram of a first object contour description data provided in an embodiment of this application;
[0056] Figure 3 A schematic diagram of corrected contour description data provided in an embodiment of this application;
[0057] Figure 4 This is a schematic diagram of the structure of an image generation device provided in an embodiment of this application;
[0058] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0059] Research has revealed that for certain application scenarios (e.g., image-to-image generation), the following solution addresses the aforementioned needs: Input the original input image, a contour image extracted from the original input image (e.g., a Canny image of a face), and a keypoint description image extracted from a pose-providing image into a diffusion model (e.g., a ControlNet model). The diffusion model then processes the original input image using the contour image and the keypoint description image as conditions to obtain a new image. The contour image provides the contour features (e.g., face shape) of certain object parts (e.g., face) within the original input image, enabling contour control during image generation. The keypoint description image provides the object's pose (e.g., facial feature positions), enabling pose control during image generation. However, this solution suffers from defects in the generated image (e.g., poor face shape retention, poor hairstyle retention, and unsatisfactory pose adjustment results), leading to poor image generation performance.
[0060] Based on the above research, in order to better improve the image generation effect, this application provides an image generation method, which includes: firstly acquiring first object contour description data (e.g., a contour image extracted from image 1), first object pose description data (e.g., a facial key point description image extracted from image 1), and second object pose description data (e.g., from image 1). Figure 2 The first object contour description data and the first object pose description data are both determined based on the first object description image (e.g., Image 1), so that the first object contour description data can describe the contours of some object features (e.g., face shape, hairstyle, etc.) that need to be preserved in the first object description image, so that the first object pose description data can describe the object pose corresponding to these object features in the first object description image (e.g., the positions of the five key points used to describe the face pose, etc.), and so that the second object pose description data can describe the object pose required to be presented in the final generated image; then, based on the difference representation data between the second object pose description data and the first object pose description data, the object is... The first object contour description data is corrected to obtain corrected contour description data, so that the object pose presented in the corrected contour description data is closer to the object pose described by the second object pose description data. Then, image generation processing is performed based on the corrected contour description data and the second object pose description data to obtain a generated image, so that the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature (e.g., face shape, hairstyle, etc.) other than the object pose. In this way, the purpose of adjusting the object pose and maintaining some object features can be taken into account during the image generation process, which is beneficial to improving the image generation effect.
[0061] Furthermore, this application does not limit the executing entity of the image generation method provided in the embodiments of this application. For example, the image generation method provided in the embodiments of this application can be applied to a terminal device or a server. Alternatively, the image generation method provided in the embodiments of this application can also be implemented through a data interaction process between a terminal device and a server. The terminal device can be a smartphone, computer, personal digital assistant (PDA), tablet computer, etc. The server can be a standalone server, a cluster server, or a cloud server.
[0062] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0063] To better understand the technical solution provided in this application, the image generation method provided in this application will be explained below with reference to some accompanying drawings. For example... Figure 1 As shown, the image generation method provided in this application includes S1-S3 as described below. Wherein, the... Figure 1 This is a flowchart of an image generation method provided in an embodiment of this application.
[0064] S1: Obtain first object contour description data, first object pose description data, and second object pose description data; the first object contour description data and the first object pose description data are both determined based on the first object description image.
[0065] The first object description image is primarily used to describe the state of an object under at least one object feature other than its pose, so that these object features can be preserved during subsequent image generation. The object refers to the foreground information described by pixels other than background pixels in the first object description image; however, this application does not limit the object, for example, the object can be a person, animal, virtual robot, or object. The object pose is used to describe the pose of some or all parts of the object in the first object description image; however, this application does not limit the object pose, for example, in some application scenarios, if the first object description image is used to describe a person or animal, then the object pose can refer to the object's facial pose (e.g., the location of some facial key points). The at least one object feature refers to other feature dimensions of the object presented in the first object description image besides its pose; however, this application embodiment does not limit the at least one object feature, for example, in some application scenarios, if the first object description image is used to describe a person or animal, and the object pose refers to the object's facial pose, then the at least one object feature can include the object's face shape and / or hairstyle. The object's face shape describes the face shape of the object as presented in the first object description image. The object's hairstyle describes the hairstyle of the object as presented in the first object description image.
[0066] The first object contour description data refers to data (e.g., an image) extracted from the aforementioned first object description image, used to describe the contours of some or all parts of the object presented by the first object description image, so that the first object contour description data can represent some contour-related characteristics of the part or all parts presented in the first object description image (e.g., the position of the part or all parts in the first object description image, the shape of the part or all parts in the first object description image, the size of the part or all parts in the first object description image, etc.). For example, the first object contour description data can be... Figure 2 The outline image shown is 1.
[0067] Furthermore, this application does not limit the representation method of the first object contour description data described above. For example, in some application scenarios, the first object contour description data can be represented by coordinate data that can indicate the position of the contour of some or all parts of an object. As another example, in some application scenarios, the first object contour description data can be represented by an image (e.g., Figure 2 The outline image 1 shown is used for representation. It can be seen that, in one possible implementation, the first object outline description data can be an outline image obtained by performing outline acquisition processing on the first object description image above, so that the outline image can represent the outline of part or all of the object (e.g., head) presented by the first object description image in image form.
[0068] Furthermore, this application does not limit the method of obtaining the first object contour description data mentioned above. For example, it can be implemented using any existing or future method capable of contour acquisition processing of a foreground target in an image (e.g., the Canny method). Therefore, in one possible implementation, the process of obtaining the first object contour description data can specifically be as follows: using a preset contour acquisition method, the first object description image mentioned above is processed to obtain the first object contour description data (e.g., a contour image such as a Canny image), so that the first object contour description data can accurately represent the contours of some or all parts of the object presented by the first object description image.
[0069] The first object pose description data refers to the data (e.g., an image) extracted from the first object description image above, used to describe the pose of some or all parts of the object presented by the first object description image, so that the first object pose description data can represent some pose-related features of that part or all parts presented in the first object description image (e.g., the location of some facial key points). For example, the first object pose description data can adopt a similar... Figure 2The key point description image of the target pose diagram shown is implemented. This key point description image is used to describe the distribution of key points in this part or all of the region.
[0070] Furthermore, this application does not limit the representation method of the first object pose description data described above. For example, in some application scenarios, the first object pose description data can be represented using key point coordinate data. Also, in some application scenarios, the first object pose description data can be represented using an image (e.g., similar to...). Figure 2 The target pose image shown is represented by a keypoint description image. Therefore, in one possible implementation, the first object pose description data can be a keypoint description image obtained by performing keypoint acquisition processing on the first object description image described above (e.g., similar to...). Figure 2 The key point description image of the target pose diagram shown is used to represent, in image form, the pose of some or all parts (e.g., the head) of the object presented by the first object description image.
[0071] Furthermore, this application does not limit the method of obtaining the first object pose description data mentioned above. For example, it can be implemented using any existing or future method capable of performing pose acquisition processing on a foreground target in an image (e.g., a key point extraction method). Therefore, in one possible implementation, the process of obtaining the first object pose description data can specifically be as follows: using a preset key point extraction method to process the first object description image mentioned above to obtain the first object pose description data (e.g., a key point description image), so that the first object pose description data can accurately represent the pose of some or all parts of the object presented by the first object description image.
[0072] Based on the aforementioned content regarding the first object contour description data and the first object pose description data, it is clear that since both the first object contour description data and the first object pose description data are determined based on the aforementioned first object description image, they describe the state presented in the same image. Therefore, the first object pose description data can represent the object pose of the contour described by the first object contour description data, and consequently, the first object pose description data can also be used to describe the object pose presented in the first object contour description data. It is evident that the object pose presented by the first object contour description data, the object pose described by the first object pose description data, and the object pose presented by the first object description image are all the same.
[0073] The second object pose description data is used to describe the pose of an object so that the object pose can be maintained during subsequent image generation. For example, the second object pose description data can be... Figure 2 The target pose diagram is shown. It should be noted that the implementation method of this second object pose description data is similar to the implementation method of the first object pose description data described above, and will not be repeated here for the sake of brevity.
[0074] Furthermore, the second object pose description data mentioned above differs from the first object pose description data mentioned above, so that the object pose described by the second object pose description data is different from the object pose described by the first object pose description data. This is so that the second object pose description data can be used as pose control information to influence the pose generation process involved in subsequent image generation, thereby ensuring that the object pose presented in the final generated image is consistent with the object pose described by the second object pose description data.
[0075] Furthermore, this application does not limit the process of obtaining the pose description data of the second object mentioned above. For ease of understanding, the following explanation is provided in conjunction with some other situations.
[0076] Scenario 1: In some application scenarios, the object pose required for the image generation process comes from one image, but other object features besides the object pose come from other images. This allows for the generation of new images by combining the different features of different images.
[0077] Based on the above situation 1, in one possible implementation, when both the first object contour description data and the first object pose description data are determined based on the first object description image, the second object pose description data can be determined based on the second object description image. Furthermore, the object pose described by the second object description image is different from the object pose described by the first object description image, so that the object pose described by the second object pose description data is different from the object pose described by the first object pose description data, thereby making the object pose described by the second object pose description data different from the object pose of the contour described by the first object contour description data. The second object description image is mainly used to provide an object pose so that the object pose can be maintained during subsequent image generation.
[0078] It should be noted that this application does not limit the process of acquiring the second object pose description data in the preceding paragraph. For example, the implementation method of acquiring the second object pose description data is similar to the implementation method of acquiring the first object pose description data described above. Therefore, in one possible implementation, the process of acquiring the second object pose description data can specifically be: processing the second object description image described above using a preset key point extraction method to obtain the second object pose description data (e.g., a key point description image), so that the second object pose description data can accurately represent the pose of some or all parts of the object presented by the second object description image.
[0079] Scenario 2: In some application scenarios (such as sample image augmentation), the second object pose description data can be obtained by perturbing the first object pose description data mentioned above, so that there is a difference between the second object pose description data and the first object pose description data, so that a new image different from the first object description image can be generated based on the second object pose description data.
[0080] Based on the above situation 2, in one possible implementation, when both the first object contour description data and the first object pose description data are determined based on the first object description image, the second object pose description data can be obtained by perturbing the first object pose description data, so that there is a difference between the object pose described by the second object pose description data and the object pose described by the first object pose description data. This makes the object pose described by the second object pose description data different from the object pose of the contour described by the first object contour description data, and thus makes the new image generated based on the second object pose description data different from the first object description image in terms of object pose. In this way, it is possible to obtain a new image different from the first object description image using the first object description image.
[0081] It should be noted that this application does not limit the implementation method of the perturbation processing in the above paragraph. For example, it can be implemented using any existing or future method that can perform perturbation processing on a piece of data (e.g., image data).
[0082] Scenario 3: For some application scenarios, in order to better meet certain requirements, different pose extraction methods (such as pose extraction methods with different precision, pose extraction methods with different working principles, etc.) can be used to obtain pose description data of the same image, so that the contour image extracted from the image can be corrected based on the difference between the two pose description data.
[0083] Based on the above situation 3, in one possible implementation, when the first object pose description data is obtained by performing keypoint detection processing on the first object description image using the first keypoint detection algorithm, the second object pose description data can be obtained by performing keypoint detection processing on the first object description image using the second keypoint detection algorithm. Because the second keypoint detection algorithm differs from the first keypoint detection algorithm, there is a difference between the second object pose description data and the first object pose description data, resulting in a certain difference between the object pose described by the second object pose description data and the object pose described by the first object pose description data.
[0084] It should be noted that, for the first and second object pose description data shown above, in order to better ensure the normal progress of the subsequent contour correction process, the following constraint must be met between these two data: there must be an intersection between the key points described by the first object pose description data and the key points described by the second object pose description data, so that the positional differences of the key points in the intersection can be used to complete the subsequent contour correction process. For example, if the first object pose description data is used to describe the positions of 106 facial key points, and the second object pose description data is used to describe 5 facial key points (e.g., ... Figure 2 If the positions of points 1 to 5 shown are given, then the 106 facial key points include some or all of the 5 facial key points.
[0085] It should also be noted that this application does not limit the implementation of the first key point detection algorithm and the second key point detection algorithm mentioned above. For example, these two can be pre-set according to the actual application scenario.
[0086] It should be further noted that this application does not limit the process of determining the first object pose description data and the second object pose description data under the above-described situation 3. For example, it can specifically be as follows: First, multiple keypoint detection algorithms are used to perform keypoint detection processing on the first object description image mentioned above, and the keypoint detection results corresponding to each keypoint detection algorithm are obtained; then, the keypoint similarity between the keypoint detection results corresponding to each keypoint detection algorithm and the keypoints in the first object contour description data mentioned above is calculated; then, the keypoint detection algorithm with the highest keypoint similarity is selected as the first keypoint detection algorithm, and the keypoint detection results corresponding to the keypoint detection algorithm with the highest keypoint similarity are selected as the first keypoint detection algorithm. The result is used as the first object pose description data; and any other keypoint detection algorithm besides the one with the highest keypoint similarity is used as the second keypoint detection algorithm, and the keypoint detection result corresponding to the other keypoint detection algorithm is used as the second object pose description data, so that the object pose presented by the new image generated based on the second object pose description data presents some subtle differences with the object pose presented by the first object description image, so that the tuple (the new image, the first object description image) can be used to complete other tasks (such as recognition processing for subtle differences).
[0087] Based on the relevant content in S1 above, in some application scenarios (such as head image generation), if it is desired to maintain the object pose described by the second object pose description data (e.g., face pose) and other object features described by the first object description image besides the object pose (e.g., face shape, hairstyle, etc.) during image generation, the second object pose description data can be obtained, and the first object contour description data and the first object pose description data can be extracted from the first object description image. This allows for subsequent adjustments to the first object contour description data based on the differences in object pose between the first and second object pose description data, in order to minimize (or even eliminate) the gap between the object pose described by the first and second object pose description data. This helps avoid defects caused by a large gap between the object pose described by the first and second object pose description data (e.g., inability to maintain face shape or hairstyle, poor pose adjustment effect, etc.).
[0088] S2: Based on the difference representation data between the second object pose description data and the first object pose description data, the first object contour description data is corrected to obtain the corrected contour description data.
[0089] The difference representation data between the second object pose description data and the first object pose description data is used to represent the differences between the object pose described by the second object pose description data and the object pose described by the first object pose description data.
[0090] Furthermore, this application does not limit the implementation method of the difference representation data in the preceding paragraph. For example, in some application scenarios, when the first object pose description data is used to describe the position of at least one first key point, and the second object pose description data is used to describe the position of at least one second key point, if there is a correspondence between the first target point among the at least one first key point and the second target point among the at least one second key point, then the difference representation data may include the positional difference between the first target point and the second target point. Here, the first key point refers to the key point described by the first object pose description data. The second key point refers to the key point described by the second object pose description data. The first target point refers to the key point among the key points described by the first object pose description data that corresponds to the second target point among the at least one second key point. The second target point refers to the key point among the key points described by the second object pose description data that corresponds to the first target point among the at least one first key point. The correspondence between the first target point and the second target point is used to indicate that the first target point and the second target point represent the same target (e.g., the left eye). The positional difference between the first target point and the second target point is used to represent the difference between the position of the first target point described by the first object attitude description data and the position of the second target point described by the second object attitude description data.
[0091] Based on the above, for some application scenarios, after obtaining the first object pose description data and the second object pose description data, if there is a correspondence between the i-th first target point described by the first object pose description data and the i-th second target point described by the first object pose description data (for example, the i-th first target point and the i-th second target point both represent the same target), then the difference between the position of the i-th first target point recorded by the first object pose description data and the position of the i-th second target point recorded by the second object pose description data can be calculated first, and this difference can be used as the value between the i-th first target point and the i-th second target point. The positional difference between the punctuation marks, where i is a positive integer, i≤I, and I represents the number of the first target point or the number of the second target point; then, the positional difference between the first first target point and the first second target point, the positional difference between the second first target point and the second second target point, ... (and so on), and the positional difference between the I-th first target point and the I-th second target point are collected to obtain the difference characterization data between the pose description data of the second object and the pose description data of the first object, so that the difference characterization data includes these positional differences, so that the contour description data of the first object above can be corrected based on the difference characterization data to obtain the corrected contour description data.
[0092] The corrected contour description data refers to the result of correcting the first object contour description data mentioned above. This corrected data ensures that the difference between the object pose described by the corrected contour description data and the object pose described by the second object pose description data is smaller than the difference between the object pose described by the first and second object pose description data. This makes the object pose presented by the corrected contour description data as close as possible (or even completely equal to) the object pose described by the second object pose description data. For example, if the first object contour description data is... Figure 2 The outline image 1 shown is shown, and the pose description data of the second object is... Figure 2 The target attitude diagram shown can be used as the corrected contour description data. Figure 3 The outline image shown is 2.
[0093] Furthermore, this application does not limit the process of obtaining the corrected contour description data (that is, the implementation of S2). For example, it can be implemented using any existing or future correction method.
[0094] In addition, in order to better improve the correction effect, this application also provides a possible implementation of S2 above. In this implementation, if the first object contour description data above is an image, then S2 may include steps 11-12 below.
[0095] Step 11: Determine the position offset based on the difference representation data between the second object pose description data and the first object pose description data.
[0096] Here, the position offset refers to the offset required when correcting the position of the contour description pixels in the first object contour description data mentioned above. The contour description pixels refer to the pixels in the first object contour description data that carry contour information (e.g., ...). Figure 2 (The pixels used to represent the color black in the outline image 1 shown). For example, the positional offset may include an offset in at least one coordinate dimension (e.g., horizontal and vertical coordinates). The at least one coordinate dimension refers to the coordinate dimension involved in the image coordinate system in the first object outline description data.
[0097] Furthermore, this application does not limit the determination process of the above-mentioned position offset. For example, when the difference representation data between the second object posture description data and the first object posture description data includes the position difference between the first first target point and the first second target point, the position difference between the second first target point and the second second target point, ... (and so on), and the position difference between the I-th first target point and the I-th second target point, if the position offset can include the offset under at least one coordinate dimension (e.g., horizontal and vertical coordinates), then the position offset can specifically be: calculate the average value of the values presented by all position differences in the difference representation data under the j-th coordinate dimension to obtain the offset under the j-th coordinate dimension, where j is a positive integer, j≤J, and J represents the number of dimensions in the at least one coordinate dimension.
[0098] Step 12: Based on the position offset mentioned above, perform pixel position correction processing on the contour description pixels in the first object contour description data to obtain the corrected contour description data.
[0099] It should be noted that this application does not limit the implementation of step 12 above. For example, step 12 can specifically be: after obtaining the position offset mentioned above, the pixel position of the kth contour description pixel in the first object contour description data can be summed with the position offset to obtain the offset position of the kth contour description pixel (for example, the coordinate value in the j-th coordinate dimension of the pixel position of the kth contour description pixel is summed with the offset in the j-th coordinate dimension of the position offset to obtain the coordinate value in the j-th coordinate dimension of the offset position of the kth contour description pixel, where j is a positive integer, j≤J, etc.), k is a positive integer, k≤K, K is a positive integer, K This represents the number of contour description pixels in the first object contour description data. Then, the pixel value at the offset position of the k-th contour description pixel in the first object contour description data is set as the pixel value of the k-th contour description pixel, where k is a positive integer, k≤K, and K is a positive integer. The pixel values at other pixel positions in the first object contour description data other than the offset positions of these K contour description pixels are set as background pixel values. This yields the corrected contour description data corresponding to the first object contour description data, so that the object pose described by the corrected contour description data is closer to (or even equal to) the object pose described by the second object pose description data mentioned above.
[0100] For example, in some application scenarios, step 12 above can specifically be: according to the position offset mentioned above, the contour description pixels in the first object contour description data are shifted to obtain corrected contour description data, so that the position difference between the position of each contour description pixel in the corrected contour description data and the position of the corresponding contour description pixel in the first object contour description data is the position offset. It should be noted that this application does not limit the implementation method of this position shifting process. For example, it can be implemented using any existing or future method capable of shifting the position of foreground pixels in an image.
[0101] Based on the relevant content of S2 above, for some application scenarios, after obtaining the second object pose description data mentioned above and extracting the first object pose description data and the first object contour description data from the first object description image mentioned above, the difference representation data between the second object pose description data and the first object pose description data can be calculated first. This difference representation data can represent the difference between the object pose described by the second object pose description data and the object pose described by the first object pose description data, thereby enabling the difference representation data to represent the difference between the object pose described by the second object pose description data and the object pose presented by the first object contour description data. Then, the first object contour description data is corrected based on the difference representation data to obtain corrected contour description data, so that the object pose presented by the corrected contour description data is closer to the object pose described by the second object pose description data. This allows the corrected contour description data to be used as contour control information to influence the generation process of other object features besides the object pose in the subsequent image generation process. In particular, because the difference between the corrected contour description data and the second object pose description data in terms of object pose is small, the corrected contour description data and the second object pose description data have a high degree of consistency in terms of object pose. This makes it easier to avoid defects caused by the large difference in object pose between the two data when generating images based on these two data, thereby improving the image generation effect.
[0102] S3: Perform image generation processing based on the corrected contour description data and the second object pose description data to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image is consistent with the first object description image in at least one object feature other than the object pose.
[0103] The generated image refers to a new image generated based on the corrected contour description data and the second object pose description data, so that the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the contour features (e.g., face shape, hairstyle, etc.) presented by the generated image are consistent with the contour features described by the corrected contour description data, thereby making the generated image consistent with the first object description image above in at least one object feature (e.g., face shape, hairstyle, etc.) other than the object pose.
[0104] Furthermore, this application does not limit the implementation method of S3 above. For ease of understanding, some situations will be described below.
[0105] In some application scenarios (such as image-to-image scenarios), a new image can be obtained by processing the image described in the first object above (such as adding noise and removing noise).
[0106] Based on the above situation one, in one possible implementation, S3 above can specifically be: inputting the above-mentioned first object description image, the above-mentioned corrected contour description data, and the above-mentioned second object pose description data into a pre-constructed graph-to-image model to obtain a generated image output by the graph-to-image model, so that the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and that the generated image and the first object description image are consistent in at least one object feature other than the object pose. This enables the adjustment of the object pose while preserving certain object features (such as hairstyle or face shape) in the first object description image, thereby improving the image generation effect. The graph-to-image model is used to process the first object description image using the corrected contour description data and the second object pose description data as conditions to obtain the generated image; moreover, this application does not limit the implementation of the graph-to-image model. For example, it can be implemented using any existing or future model capable of generating a new image based on a graph and some conditions (such as diffusion models similar to ControlNet).
[0107] In some application scenarios (such as text-to-image scenarios), the prompt text, the corrected outline description data from the preceding text, and the pose description data of the second object from the preceding text can all be used as conditions for image generation processing.
[0108] Based on the above-described second scenario, in one possible implementation, S3 can specifically be: inputting noise data, prompt text, the corrected contour description data, and the second object pose description data into a pre-constructed raw image model to obtain a generated image output by the raw image model, so that the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image is consistent with the first object description image in at least one object feature other than the object pose. Here, noise data refers to the noise data used during image generation processing; and this application does not limit the method of generating the noise data. For example, it can be implemented using any existing or future noise data generation method (e.g., a method of randomly generating noise data or a method of adding noise to certain images, etc.). The prompt text refers to the text referenced during image processing so that the prompt text can describe certain constraints in string form; and this application does not limit the prompt text. The text-to-image model is used to process the noise data using the prompt text, the corrected contour description data, and the second object pose description data as conditions to obtain a generated image. Moreover, this application does not limit the implementation of the text-to-image model. For example, it can be implemented using any existing or future model that can generate a new image based on noise data, prompt text, and some additional conditions (e.g., a diffusion model).
[0109] Based on the relevant content of S1 to S3 above, it can be seen that for the image generation method provided in this application embodiment, the first object contour description data (e.g., the contour image extracted from image 1), the first object pose description data (e.g., the facial key point description image extracted from image 1), and the second object pose description data (e.g., the facial key point description image extracted from image 1) are first obtained. Figure 2The first object contour description data and the first object pose description data are both determined based on the first object description image (e.g., Image 1), so that the first object contour description data can describe the contours of some object features (e.g., face shape, hairstyle, etc.) that need to be preserved in the first object description image, so that the first object pose description data can describe the object pose corresponding to these object features in the first object description image (e.g., the positions of the five key points used to describe the face pose, etc.), and so that the second object pose description data can describe the object pose required to be presented in the final generated image; then, based on the difference representation data between the second object pose description data and the first object pose description data, the object is... The first object contour description data is corrected to obtain corrected contour description data, so that the object pose presented in the corrected contour description data is closer to the object pose described by the second object pose description data. Then, image generation processing is performed based on the corrected contour description data and the second object pose description data to obtain a generated image, so that the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature (e.g., face shape, hairstyle, etc.) other than the object pose. In this way, the purpose of adjusting the object pose and maintaining some object features can be taken into account during the image generation process, which is beneficial to improving the image generation effect.
[0110] Furthermore, based on the relevant content of the first object description image involved in the image generation method above, it can be known that in some application scenarios, the first object description image can refer to an image, and the image can simultaneously provide multiple object features that need to be preserved (such as object features such as object hairstyle and object face shape), so that a new image can be generated subsequently using the process shown in S1-S3 above. In this way, image generation processing can be achieved while preserving multiple object features provided by an image.
[0111] In addition, for some application scenarios, there may be the following requirements: at least one object feature (such as hairstyle and face shape) mentioned above comes from different images, so as to maintain the different object features provided by different images during the image generation process (such as the object face shape provided by one image and the object hairstyle provided by another image).
[0112] To better fulfill the requirements of the above paragraph, this application also provides a possible implementation of the image generation method, in which the image generation method includes steps 21-23 below.
[0113] Step 21: Obtain first object contour description data, first object pose description data, and second object pose description data; the first object contour description data and the first object pose description data are both determined based on the first object description image; the first object description image includes at least one image to be processed; different images to be processed are used to provide information to be preserved under different object features; the first object contour description data includes object contour description data corresponding to each image to be processed; the first object pose description data includes object pose description data corresponding to each image to be processed.
[0114] In this application, in some application scenarios, if different images are needed to provide different object features that need to be preserved, N images to be processed can be obtained first, so that different images to be processed can be used to provide information to be preserved under different object features, where N is a positive integer. Then, for the nth image to be processed to provide information to be preserved under the nth object feature, contour extraction processing is performed on the nth image to obtain the object contour description data corresponding to the nth image, and pose extraction processing (e.g., key point extraction processing) is performed on the nth image to obtain the object pose description data corresponding to the nth image, so that the object contour description data corresponding to the nth image can be corrected based on the difference representation data between the second object pose description data mentioned above and the object pose description data corresponding to the nth image, where n is a positive integer, n≤N, and N is a positive integer.
[0115] The nth image to be processed describes the state of an object under its nth feature (excluding its pose), providing information to be preserved under that nth feature so that it can be retained during subsequent image generation. n is a positive integer, n≤N, where N is a positive integer. This information to be preserved refers to the pixel information in the nth image that characterizes the nth object feature.
[0116] It should be noted that this application does not limit the implementation method of the information to be retained in the above paragraph. For example, if the nth image to be processed above is used to retain information to be retained under the object's face shape, then the information to be retained may refer to the pixel information existing in the nth image to be processed that is used to characterize the object's face shape. Similarly, if the nth image to be processed above is used to retain information to be retained under the object's hairstyle, then the information to be retained may refer to the pixel information existing in the nth image to be processed that is used to characterize the object's hairstyle.
[0117] It should also be noted that this application does not limit the implementation of the above-mentioned at least one image to be processed. For example, the at least one image to be processed may include a face shape providing image and a hairstyle providing image. The face shape providing image is used to provide information to be retained under the object's face shape. The hairstyle providing image is used to provide information to be retained under the object's hairstyle.
[0118] The object contour description data corresponding to the nth image to be processed mentioned above refers to the data (e.g., image) extracted from the nth image to be processed to describe the contours of some or all parts of the object presented by the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent some contour-related features of the part or all parts presented in the nth image to be processed (e.g., the position of the part or all parts in the nth image to be processed, the shape of the part or all parts presented in the nth image to be processed, the size of the part or all parts presented in the nth image to be processed, etc.).
[0119] Furthermore, this application does not limit the implementation method of the object contour description data corresponding to the nth image to be processed mentioned above. For ease of understanding, some examples are provided below.
[0120] Example 1: If the nth image to be processed is used to preserve information about the face shape of an object (e.g., the nth image to be processed provides an image of the face shape), then the object contour description data corresponding to the nth image to be processed can be implemented using the face shape contour description data corresponding to the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent the contour-related features of the object's face shape presented in the nth image to be processed. Here, the face shape contour description data corresponding to the nth image to be processed refers to the data (e.g., an image) extracted from the nth image to be processed to describe the contour of the face in the object presented by the nth image to be processed.
[0121] Based on Example 1 above, in one possible implementation, when the nth image to be processed provides an image of a face shape, the object contour description data corresponding to the nth image to be processed can be the face shape contour description data corresponding to the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent the contour-related features of the object face shape presented in the nth image to be processed.
[0122] Example 2: If the nth image to be processed is used to preserve information about the hairstyle of an object (e.g., the nth image to be processed provides an image of the hairstyle), then the object contour description data corresponding to the nth image to be processed can be implemented using the hairstyle contour description data corresponding to the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent the contour-related features of the hairstyle presented in the nth image to be processed. Here, the hairstyle contour description data corresponding to the nth image to be processed refers to the data (e.g., an image) extracted from the nth image to be processed to describe the contour of the hair in the object presented by the nth image to be processed.
[0123] Based on Example 2 above, in one possible implementation, when the nth image to be processed provides an image of a hairstyle, the object contour description data corresponding to the nth image to be processed can be the hairstyle contour description data corresponding to the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent the contour-related features of the hairstyle in the nth image to be processed.
[0124] The object pose description data corresponding to the nth image to be processed mentioned above refers to the data (e.g., images) extracted from the nth image to be processed to describe the pose of some or all parts of the object presented by the nth image to be processed, so that the object pose description data corresponding to the nth image to be processed can represent some pose-related features of the part or all parts presented in the nth image to be processed (e.g., the location of some facial key points).
[0125] Furthermore, this application does not limit the implementation method of the object pose description data corresponding to the nth image to be processed mentioned above. For example, when the nth image to be processed is a face shape image or a hairstyle image, the object pose description data corresponding to the nth image to be processed can adopt the facial key point description data corresponding to the nth image to be processed (e.g., Figure 2 The implementation uses the target pose map shown to enable the object pose description data corresponding to the nth image to be processed to represent the facial pose presented by the nth image to be processed. The facial key point description data corresponding to the nth image to be processed is used to describe the positions of some facial key points in the nth image to be processed; moreover, this application does not limit the representation method of the facial key point description data. For example, in some application scenarios, the facial key point description data can be in the form of an image (e.g., Figure 2 The implementation will proceed as shown in the target attitude diagram.
[0126] Based on the relevant content of step 21 above, in some application scenarios (e.g., different object features other than object pose come from different image scenes), if it is desired to maintain the object pose described by the second object pose description data and retain some object features described by some images to be processed during image generation, the second object pose description data can be obtained, and the object contour description data and object pose description data corresponding to each image to be processed can be extracted from each image to be processed. This allows for subsequent adjustments to the object contour description data corresponding to the corresponding image to be processed based on the differences in object pose between the object pose description data corresponding to each image to be processed and the second object pose description data, so as to minimize (or even eliminate) the gap between the object pose described by the object contour description data and the object pose described by the second object pose description data. This helps to avoid defects caused by a large gap between the object pose described by the object contour description data and the object pose described by the second object pose description data.
[0127] Step 22: For any image to be processed, based on the difference representation data between the second object pose description data mentioned above and the object pose description data corresponding to the image to be processed, the object contour description data corresponding to the image to be processed is corrected to obtain the corrected contour description data corresponding to the image to be processed.
[0128] In this application, for the nth image to be processed, after obtaining the object contour description data and object pose description data corresponding to the nth image to be processed, the difference representation data between the second object pose description data mentioned above and the object pose description data corresponding to the nth image to be processed can be calculated first. This difference representation data can represent the difference between the object pose described by the second object pose description data and the object pose described by the nth image to be processed, thereby enabling the difference representation data to represent the difference between the object pose described by the second object pose description data and the object pose presented by the object contour description data corresponding to the nth image to be processed. Then, based on the difference representation data, the object contour description data corresponding to the nth image to be processed is corrected to obtain the corrected contour description data corresponding to the nth image to be processed, so that the difference between the object pose described by the corrected contour description data and the object pose described by the second object pose description data is small (or even non-existent). Wherein, n is a positive integer, n≤N.
[0129] Step 23: Perform image generation processing based on the second object pose description data and the corrected contour description data corresponding to at least one image to be processed above to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the nth object feature described by the generated image is consistent with the nth object feature described by the nth image to be processed, where n is a positive integer, n≤N, and N represents the number of images in the at least one image to be processed.
[0130] It should be noted that the implementation method of step 23 is similar to the implementation method of S3 above, and will not be repeated here for the sake of brevity.
[0131] Based on the relevant content of steps 21 to 23 above, it can be seen that in some application scenarios, if it is necessary to perform image generation processing based on different object features provided by different images, the object contour description data corresponding to the corresponding image can be corrected based on the difference between the object pose description data corresponding to each image and the second object pose description data mentioned above; then, a new image can be generated based on these correction results and the second object pose description data. In this way, the object pose can be adjusted while maintaining the object features provided by multiple images, which is beneficial to improving the image generation effect.
[0132] In fact, in order to better improve the image generation effect, this application also provides some possible implementations of the image generation method described above. The following description uses some implementations of the image generation method as examples.
[0133] In one possible implementation, the image generation method described above may include steps 31-35 below.
[0134] Step 31: Obtain the object contour description data corresponding to the face shape image, the object pose description data corresponding to the face shape image, the object contour description data corresponding to the hairstyle image, the object pose description data corresponding to the hairstyle image, and the second object pose description data.
[0135] The object contour description data corresponding to the face shape image refers to the data (e.g., an image) extracted from the face shape image to describe the contour of the face in the object presented by the face shape image, so that the object contour description data corresponding to the face shape image can represent the contour features of the object's face presented in the face shape image; moreover, this application does not limit the implementation method of the object contour description data corresponding to the face shape image. For example, the object contour description data corresponding to the face shape image can be implemented using the face shape contour description data corresponding to the face shape image.
[0136] The object pose description data corresponding to the face shape image refers to the data (e.g., an image) extracted from the face shape image to describe the pose of the face in the object presented by the face shape image, so that the object pose description data corresponding to the face shape image can represent the pose characteristics of the object's face in the face shape image; moreover, this application does not limit the implementation method of the object pose description data corresponding to the face shape image. For example, the object pose description data corresponding to the face shape image can adopt the facial key point description data corresponding to the face shape image (e.g., similar to...). Figure 2 The implementation is carried out using the key point description image of the target attitude diagram shown.
[0137] The object contour description data corresponding to the hairstyle image refers to the data (e.g., image) extracted from the hairstyle image to describe the contour of the hair in the object presented by the hairstyle image, so that the object contour description data corresponding to the hairstyle image can represent the contour features of the hair of the object presented in the hairstyle image.
[0138] In addition, for the object contour description data corresponding to the above hairstyle image, in order to better improve the subsequent pose correction effect, the object contour description data corresponding to the hairstyle image can be aligned with the object pose description data corresponding to the hairstyle image in some aspects.
[0139] Based on the foregoing content, in order to better improve the correction effect, this application also provides a possible implementation of the object contour description data corresponding to the hairstyle image mentioned above. In this implementation, the object contour description data corresponding to the hairstyle image can be the head contour description data corresponding to the hairstyle image (e.g., Figure 2 The outline image 1 shown is implemented such that the object outline description data corresponding to the hairstyle image can represent the outline features (e.g., facial outline features and hair outline features) of the object's head. Here, the head outline description data corresponding to the hairstyle image refers to data (e.g., an image) extracted from the hairstyle image to describe the outline of the head in the object presented by the hairstyle image, so that the object outline description data corresponding to the hairstyle image can represent the outline features of the object's head in the hairstyle image, thereby making the head outline description data corresponding to the hairstyle image carry facial outline description information and hairstyle outline description information. The facial outline description information is used to represent the outline features of the facial region in the object's head in the hairstyle image. The hairstyle outline description information is used to represent the outline features of the hair region in the object's head in the hairstyle image.
[0140] The object pose description data corresponding to the hairstyle image refers to the data (e.g., an image) extracted from the hairstyle image to describe the pose of the face in the object presented by the hairstyle image, so that the object pose description data corresponding to the hairstyle image can represent the pose characteristics of the object's face in the hairstyle image; moreover, this application does not limit the implementation method of the object pose description data corresponding to the hairstyle image. For example, the object pose description data corresponding to the hairstyle image can adopt the facial key point description data corresponding to the hairstyle image (e.g., similar to...). Figure 2 The implementation is carried out using the key point description image of the target attitude diagram shown.
[0141] Based on the relevant content of step 31 above, in some application scenarios, if it is desired to maintain the object pose described by the second object pose description data, maintain the object face shape described by the face shape image, and maintain the object hairstyle described by the hairstyle image during the image generation process, the second object pose description data can be obtained, and the object contour description data and object pose description data corresponding to the face shape image can be extracted from the face shape image, and the object contour description data and object pose description data corresponding to the hairstyle image can be extracted from the hairstyle image.
[0142] Step 32: Based on the difference representation data between the second object pose description data above and the object pose description data corresponding to the face shape image above, perform correction processing on the object contour description data corresponding to the face shape image to obtain the corrected contour description data corresponding to the face shape image.
[0143] It should be noted that the implementation method of step 32 is similar to that of step 22 above, and will not be repeated here for the sake of brevity.
[0144] Step 33: Based on the image segmentation results of the hairstyle image provided above, perform information removal processing on the facial contour description information in the object contour description data (i.e., head contour description data) corresponding to the hairstyle image to obtain the hairstyle contour description data corresponding to the hairstyle image.
[0145] The image segmentation result of the hairstyle image is used to represent the positions of different parts of an object within the hairstyle image. Furthermore, this application does not limit the implementation of the image segmentation result of the hairstyle image; for example, the image segmentation result of the hairstyle image may include at least face position representation data and hair position representation data. The face position representation data is used to describe the position of the face in the hairstyle image. The hair position representation data is used to describe the position of the hair in the hairstyle image.
[0146] Furthermore, this application does not limit the implementation of step 33 above. For example, when the head contour description data corresponding to the hairstyle image above is an image, step 33 can specifically be: first, obtain facial position representation data from the image segmentation result of the hairstyle image; then, adjust the pixel values at the pixel positions described by the facial position representation data in the head contour description data corresponding to the hairstyle image to background pixel values to obtain the hairstyle contour description data corresponding to the hairstyle image, so that only the hairstyle contour description information is retained in the hairstyle contour description data, thereby enabling the hairstyle contour description data to represent some contour features of the object hairstyle presented in the hairstyle image.
[0147] Step 34: Based on the difference representation data between the second object pose description data above and the object pose description data corresponding to the hairstyle image above, perform correction processing on the hairstyle contour description data corresponding to the hairstyle image above to obtain the corrected contour description data corresponding to the hairstyle image above.
[0148] It should be noted that the implementation method of step 34 is similar to the implementation method of step 22 above, and will not be repeated here for the sake of brevity.
[0149] Step 35: Based on the second object pose description data, the corrected contour description data corresponding to the face shape image, and the corrected contour description data corresponding to the hairstyle image, perform image generation processing to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, the object face shape described by the generated image is consistent with the object face shape described by the face shape image, and the object hairstyle described by the generated image is consistent with the object hairstyle described by the hairstyle image.
[0150] It should be noted that the implementation method of step 35 is similar to the implementation method of S3 above, and will not be repeated here for the sake of brevity.
[0151] Based on the relevant content of steps 31 to 35 above, in some application scenarios, if it is desired to maintain the feature of the hairstyle described by the hairstyle image, the head contour description data and object pose description data corresponding to the hairstyle image can be extracted first. Then, the facial contour description information in the head contour description data corresponding to the hairstyle image is processed to remove information, resulting in the hairstyle contour description data corresponding to the hairstyle image. Then, based on the difference representation data between the second object pose description data and the object pose description data corresponding to the hairstyle image, the hairstyle contour description data corresponding to the hairstyle image is corrected to obtain the corrected contour description data corresponding to the hairstyle image. This allows subsequent image generation processing to be performed based on the second object pose description data and the corrected contour description data corresponding to the hairstyle image, thus improving the image generation effect.
[0152] In one possible implementation, the image generation method described above may include steps 41-45 below.
[0153] Step 41: Obtain the object contour description data corresponding to the face shape image, the object pose description data corresponding to the face shape image, the object contour description data corresponding to the hairstyle image, the object pose description data corresponding to the hairstyle image, and the second object pose description data.
[0154] It should be noted that the relevant content of step 41 can be found in step 31 above.
[0155] Step 42: Based on the difference representation data between the second object pose description data above and the object pose description data corresponding to the face shape image above, perform correction processing on the object contour description data corresponding to the face shape image to obtain the corrected contour description data corresponding to the face shape image.
[0156] It should be noted that the implementation method of step 42 is similar to that of step 22 above, and will not be repeated here for the sake of brevity.
[0157] Step 43: Based on the difference representation data between the second object pose description data and the object pose description data corresponding to the hairstyle image above, the object contour description data (that is, head contour description data) corresponding to the hairstyle image is corrected to obtain the corrected contour description data corresponding to the hairstyle image, so that the corrected contour description data carries face contour description information and hairstyle contour description information, thereby enabling the corrected contour description data to represent the contour features of the head in the pose-corrected object.
[0158] It should be noted that the implementation method of step 43 is similar to that of step 22 above, and will not be repeated here for the sake of brevity.
[0159] Step 44: Based on the image segmentation results of the hairstyle image provided above, perform information removal processing on the facial contour description information in the corrected contour description data corresponding to the hairstyle image to obtain the removed contour description data corresponding to the hairstyle image.
[0160] It should be noted that the implementation of step 44 is similar to the implementation of step 33 above. Therefore, in one possible implementation, step 44 can specifically be: first, obtain facial position representation data from the image segmentation result of the provided image of the hairstyle; then, adjust the pixel values at the pixel positions described by the facial position representation data in the corrected contour description data corresponding to the provided image of the hairstyle to background pixel values, thereby obtaining the removed contour description data corresponding to the provided image of the hairstyle, so that only the hairstyle contour description information is retained in the removed contour description data, thus enabling the removed contour description data to represent the contour characteristics presented by the hairstyle in the pose-adjusted object.
[0161] Step 45: Based on the second object pose description data, the corrected contour description data corresponding to the face shape image provided above, and the removed contour description data corresponding to the hairstyle image provided above, perform image generation processing to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, the object face shape described by the generated image is consistent with the object face shape described by the face shape image provided above, and the object hairstyle described by the generated image is consistent with the object hairstyle described by the hairstyle image provided above.
[0162] It should be noted that the implementation method of step 45 is similar to the implementation method of S3 above, and will not be repeated here for the sake of brevity.
[0163] Based on the relevant content of steps 41 to 45 above, in some application scenarios, if it is desired to maintain the feature of the hairstyle described by the hairstyle image, the head contour description data and object pose description data corresponding to the hairstyle image can be extracted first. Then, based on the difference representation data between the second object pose description data and the object pose description data corresponding to the hairstyle image, the head contour description data corresponding to the hairstyle image is corrected to obtain the corrected contour description data corresponding to the hairstyle image, so that the corrected contour description data can represent the contour features of the object head (e.g., face shape + hairstyle) after pose adjustment. Then, the face shape contour description information in the corrected contour description data corresponding to the hairstyle image is removed to obtain the removed contour description data corresponding to the hairstyle image, so that subsequent image generation processing can be performed based on the second object pose description data and the removed contour description data corresponding to the hairstyle image, which helps to improve the image generation effect.
[0164] Based on the image generation method provided in the embodiments of this application, the embodiments of this application also provide an image generation apparatus, which will be discussed below. Figure 4 Explanation and clarification will be provided. Among them, Figure 4 This is a schematic diagram of an image generation apparatus provided in an embodiment of this application. It should be noted that for technical details of the image generation apparatus provided in this application embodiment, please refer to the relevant content of the image generation method above.
[0165] like Figure 4 As shown, the image generation apparatus 400 provided in this application embodiment includes:
[0166] The acquisition unit 401 is used to acquire first object contour description data, first object pose description data, and second object pose description data; the first object contour description data and the first object pose description data are both determined based on the first object description image;
[0167] The correction unit 402 is used to perform correction processing on the first object contour description data based on the difference characterization data between the second object posture description data and the first object posture description data to obtain corrected contour description data.
[0168] The generation unit 403 is used to perform image generation processing based on the corrected contour description data and the second object pose description data to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature other than the object pose.
[0169] In one possible implementation, the difference between the object pose of the contour described by the corrected contour description data and the object pose described by the second object pose description data is smaller than the difference between the object pose of the contour described by the first object contour description data and the object pose described by the second object pose description data.
[0170] In one possible implementation, the first object contour description data is an image;
[0171] The correction unit 402 includes:
[0172] The offset determination subunit is used to determine the position offset based on the difference representation data between the second object posture description data and the first object posture description data;
[0173] The position correction subunit is used to perform pixel position correction processing on the contour description pixels in the first object contour description data according to the position offset, so as to obtain the corrected contour description data.
[0174] In one possible implementation, the first object pose description data is used to describe the position of at least one first key point; the second object pose description data is used to describe the position of at least one second key point; if there is a correspondence between a first target point among the at least one first key point and a second target point among the at least one second key point, then the difference characterization data includes the position difference between the first target point and the second target point; the position offset is determined based on the position difference.
[0175] In one possible implementation, the position correction subunit is specifically used to: perform position translation processing on the contour description pixels in the first object contour description data according to the position offset, so as to obtain corrected contour description data.
[0176] In one possible implementation, the second object pose description data is determined based on a second object description image, the object pose described by the second object description image being different from the object pose described by the first object description image.
[0177] In one possible implementation, the second object pose description data is obtained by perturbing the first object pose description data so that there is a difference between the object pose described by the second object pose description data and the object pose described by the first object pose description data.
[0178] In one possible implementation, the first object pose description data is obtained by performing keypoint detection processing on the first object description image using a first keypoint detection algorithm; the second object pose description data is obtained by performing keypoint detection processing on the first object description image using a second keypoint detection algorithm; the second keypoint detection algorithm is different from the first keypoint detection algorithm.
[0179] In one possible implementation, the first object description image includes at least one image to be processed; different images to be processed are used to provide information to be retained under different object features; the first object contour description data includes object contour description data corresponding to each of the images to be processed; the first object pose description data includes object pose description data corresponding to each of the images to be processed.
[0180] The correction unit 402 is specifically used to: for any image to be processed, based on the difference characterization data between the second object pose description data and the object pose description data corresponding to the image to be processed, perform correction processing on the object contour description data corresponding to the image to be processed to obtain the corrected contour description data corresponding to the image to be processed.
[0181] The generation unit 403 is specifically used to: perform image generation processing based on the second object pose description data and the corrected contour description data corresponding to the at least one image to be processed, to obtain a generated image.
[0182] In one possible implementation, the at least one image to be processed includes a face shape providing image and a hairstyle providing image, wherein the face shape providing image is used to provide information to be retained under the object's face shape, and the hairstyle providing image is used to provide information to be retained under the object's hairstyle.
[0183] In one possible implementation, for any of the images to be processed, the object pose description data corresponding to the image to be processed is determined based on the image to be processed, and the object pose description data corresponding to the image to be processed is used to describe the position of at least one key point in the facial region.
[0184] In one possible implementation, the object outline description data corresponding to the hairstyle image is provided as the hairstyle outline description data corresponding to the hairstyle image.
[0185] In one possible implementation, the object contour description data corresponding to the hairstyle image is the head contour description data corresponding to the hairstyle image; the head contour description data carries face shape contour description information and hairstyle contour description information.
[0186] In one possible implementation, the correction unit 402 is specifically configured to: perform information removal processing on the facial contour description information in the head contour description data corresponding to the hairstyle image based on the image segmentation result of the hairstyle image, to obtain hairstyle contour description data corresponding to the hairstyle image; and perform correction processing on the hairstyle contour description data corresponding to the hairstyle image based on the difference representation data between the second object pose description data and the object pose description data corresponding to the hairstyle image, to obtain corrected contour description data corresponding to the hairstyle image.
[0187] In one possible implementation, the image generation apparatus 400 further includes:
[0188] The removal unit is used to perform information removal processing on the facial contour description information in the corrected contour description data corresponding to the hairstyle image based on the image segmentation result of the hairstyle image, so as to obtain the removed contour description data corresponding to the hairstyle image.
[0189] The generation unit 403 is specifically used to: perform image generation processing based on the second object pose description data, the corrected contour description data corresponding to the face shape image, and the removed contour description data corresponding to the hairstyle image, to obtain a generated image.
[0190] In one possible implementation, the generation unit 403 is specifically configured to: input the first object description image, the corrected contour description data, and the second object pose description data into a pre-constructed graph model to obtain the generated image output by the graph model; the graph model is configured to process the first object description image using the corrected contour description data and the second object pose description data as conditions to obtain the generated image.
[0191] In one possible implementation, the generation unit 403 is specifically used to: input noise data, prompt text, the corrected contour description data, and the second object pose description data into a pre-constructed text image model to obtain the generated image output by the text image model; the text image model is used to process the noise data using the prompt text, the corrected contour description data, and the second object pose description data as conditions to obtain the generated image.
[0192] In one possible implementation, the at least one object feature includes at least one of the object's face shape and the object's hairstyle.
[0193] Based on the aforementioned content regarding the image generation apparatus 400, it is understood that the image generation apparatus 400 provided in this application first acquires first object contour description data (e.g., a contour image extracted from image 1), first object pose description data (e.g., a facial key point description image extracted from image 1), and second object pose description data (e.g., a facial key point description image extracted from image 1). Figure 2 The first object contour description data and the first object pose description data are both determined based on the first object description image (e.g., Image 1), so that the first object contour description data can describe the contours of some object features (e.g., face shape, hairstyle, etc.) that need to be preserved in the first object description image, so that the first object pose description data can describe the object pose corresponding to these object features in the first object description image (e.g., the positions of the five key points used to describe the face pose, etc.), and so that the second object pose description data can describe the object pose required to be presented in the final generated image; then, based on the difference representation data between the second object pose description data and the first object pose description data, the object is... The first object contour description data is corrected to obtain corrected contour description data, so that the object pose presented in the corrected contour description data is closer to the object pose described by the second object pose description data. Then, image generation processing is performed based on the corrected contour description data and the second object pose description data to obtain a generated image, so that the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature (e.g., face shape, hairstyle, etc.) other than the object pose. In this way, the purpose of adjusting the object pose and maintaining some object features can be taken into account during the image generation process, which is beneficial to improving the image generation effect.
[0194] In addition, this application embodiment also provides an electronic device, the device including a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device performs any implementation of the image generation method provided in this application embodiment.
[0195] See Figure 5The diagram illustrates a structural schematic of an electronic device 500 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0196] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0197] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0198] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0199] The electronic device provided in this embodiment belongs to the same inventive concept as the method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0200] This application also provides a computer-readable medium storing instructions or a computer program that, when executed on a device, causes the device to perform any implementation of the image generation method provided in this application.
[0201] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0202] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0203] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0204] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, enable the electronic device to perform the aforementioned methods.
[0205] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0206] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0207] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units / modules do not necessarily limit the specific unit itself.
[0208] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0209] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0210] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
[0211] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0212] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0213] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0214] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image generation method, characterized in that, The method includes: Acquire first object contour description data, first object pose description data, and second object pose description data; the first object contour description data and the first object pose description data are both determined based on the first object description image, and the first object contour description data is an image; Based on the difference representation data between the second object pose description data and the first object pose description data, the first object contour description data is corrected to obtain corrected contour description data; the process of determining the corrected contour description data includes: determining the position offset based on the difference representation data between the second object pose description data and the first object pose description data; and performing pixel position correction processing on the contour description pixels in the first object contour description data based on the position offset to obtain corrected contour description data. Image generation processing is performed based on the corrected contour description data and the second object pose description data to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature other than the object pose.
2. The method according to claim 1, characterized in that, The difference between the object pose of the contour described by the corrected contour description data and the object pose described by the second object pose description data is smaller than the difference between the object pose of the contour described by the first object contour description data and the object pose described by the second object pose description data.
3. The method according to claim 1, characterized in that, The first object pose description data is used to describe the position of at least one first key point; The second object pose description data is used to describe the position of at least one second key point; If there is a correspondence between a first target point among the at least one first key point and a second target point among the at least one second key point, then the difference characterization data includes the positional difference between the first target point and the second target point; the positional offset is determined based on the positional difference.
4. The method according to claim 1, characterized in that, The step of performing pixel position correction processing on the contour description pixels in the first object contour description data based on the position offset to obtain corrected contour description data includes: According to the stated position offset, the contour description pixels in the first object contour description data are shifted to obtain the corrected contour description data.
5. The method according to claim 1, characterized in that, The second object pose description data is determined based on the second object description image, and the object pose described by the second object description image is different from the object pose described by the first object description image.
6. The method according to claim 1, characterized in that, The second object pose description data is obtained by perturbing the first object pose description data so that there is a difference between the object pose described by the second object pose description data and the object pose described by the first object pose description data.
7. The method according to claim 1, characterized in that, The first object pose description data is obtained by performing key point detection processing on the first object description image using the first key point detection algorithm; The second object pose description data is obtained by performing key point detection processing on the first object description image using the second key point detection algorithm; The second keypoint detection algorithm is different from the first keypoint detection algorithm.
8. The method according to claim 1, characterized in that, The first object description image includes at least one image to be processed; different images to be processed are used to provide information to be preserved under different object features; the first object contour description data includes object contour description data corresponding to each of the images to be processed; the first object pose description data includes object pose description data corresponding to each of the images to be processed. The step of correcting the first object contour description data based on the difference representation data between the second object pose description data and the first object pose description data to obtain corrected contour description data includes: For any of the images to be processed, based on the difference representation data between the second object pose description data and the object pose description data corresponding to the image to be processed, the object contour description data corresponding to the image to be processed is corrected to obtain the corrected contour description data corresponding to the image to be processed. The step of performing image generation processing based on the corrected contour description data and the second object pose description data to obtain a generated image includes: Image generation processing is performed based on the pose description data of the second object and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image.
9. The method according to claim 8, characterized in that, The at least one image to be processed includes a face shape image and a hairstyle image, wherein the face shape image is used to provide information to be retained under the object's face shape, and the hairstyle image is used to provide information to be retained under the object's hairstyle.
10. The method according to claim 9, characterized in that, For any of the images to be processed, the object pose description data corresponding to the image to be processed is determined based on the image to be processed, and the object pose description data corresponding to the image to be processed is used to describe the position of at least one key point in the face region.
11. The method according to claim 9, characterized in that, The hairstyle provides object outline description data corresponding to the image, which is the hairstyle outline description data corresponding to the image.
12. The method according to claim 9, characterized in that, The image provided by the hairstyle provides object contour description data corresponding to the head contour description data corresponding to the hairstyle image; the head contour description data carries face shape contour description information and hairstyle contour description information.
13. The method according to claim 12, characterized in that, The process of determining the corrected contour description data corresponding to the hairstyle image includes: Based on the image segmentation result of the provided hairstyle image, information removal processing is performed on the facial contour description information in the head contour description data corresponding to the provided hairstyle image to obtain the hairstyle contour description data corresponding to the provided hairstyle image. Based on the difference representation data between the second object pose description data and the object pose description data corresponding to the hairstyle image, the hairstyle contour description data corresponding to the hairstyle image is corrected to obtain the corrected contour description data corresponding to the hairstyle image.
14. The method according to claim 12, characterized in that, Before performing image generation processing based on the second object pose description data and the corrected contour description data corresponding to the at least one image to be processed, the method further includes: Based on the image segmentation result of the provided image of the hairstyle, information removal processing is performed on the facial contour description information in the corrected contour description data corresponding to the provided image of the hairstyle to obtain the removed contour description data corresponding to the provided image of the hairstyle. The step of performing image generation processing based on the second object pose description data and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image includes: The generated image is obtained by performing image generation processing based on the pose description data of the second object, the corrected contour description data corresponding to the face shape image, and the removed contour description data corresponding to the hairstyle image.
15. The method according to claim 1, characterized in that, The step of performing image generation processing based on the corrected contour description data and the second object pose description data to obtain a generated image includes: The first object description image, the corrected contour description data, and the second object pose description data are input into a pre-constructed graph model to obtain the generated image output by the graph model; the graph model is used to process the first object description image using the corrected contour description data and the second object pose description data as conditions to obtain the generated image.
16. The method according to claim 1, characterized in that, The step of performing image generation processing based on the corrected contour description data and the second object pose description data to obtain a generated image includes: The noise data, prompt text, the corrected contour description data, and the second object pose description data are input into a pre-constructed text image model to obtain the generated image output by the text image model; the text image model is used to process the noise data using the prompt text, the corrected contour description data, and the second object pose description data as conditions to obtain the generated image.
17. The method according to claim 1, characterized in that, The at least one object feature includes at least one of the object's face shape and the object's hairstyle.
18. An image generation apparatus, characterized in that, include: The acquisition unit is used to acquire first object contour description data, first object pose description data, and second object pose description data; the first object contour description data and the first object pose description data are both determined based on the first object description image, and the first object contour description data is an image; The correction unit is used to correct the first object contour description data based on the difference representation data between the second object pose description data and the first object pose description data to obtain corrected contour description data. The process of determining the corrected contour description data includes: determining a position offset based on the difference representation data between the second object pose description data and the first object pose description data; and performing pixel position correction processing on the contour description pixels in the first object contour description data based on the position offset to obtain corrected contour description data. The generation unit is used to perform image generation processing based on the corrected contour description data and the second object pose description data to obtain a generated image; the object pose described by the generated image is consistent with the object pose described by the second object pose description data, and the generated image and the first object description image are consistent in at least one object feature other than the object pose.
19. An electronic device, characterized in that, The electronic device includes: a processor and a memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory to cause the electronic device to perform the method according to any one of claims 1-17.
20. A computer-readable medium, characterized in that, The computer-readable medium stores instructions or computer programs that, when executed on the device, cause the device to perform the method according to any one of claims 1-17.