Image generation method and apparatus, electronic device, and computer readable medium
By acquiring and correcting the object outline and pose description data in the image, and generating images with the second object pose, the problem of object pose adjustment and feature maintenance in image generation is solved, and efficient image generation effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/133475
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively maintain the consistency between images and information in image generation scenarios, especially when adjusting the object posture, which can easily lead to poor maintenance of features such as face shape and hairstyle.
By obtaining the first object outline description data, the first object posture description data and the second object posture description data, the first object outline description data is used to correct the first object outline description data, and the corrected outline description data is generated, and image generation processing is performed in combination with the second object posture description data to ensure that the object posture of the generated image is consistent with the second object posture description data, and at the same time, it is consistent with the original image in terms of object characteristics such as face shape, hairstyle, etc.
The purpose of adjusting the object posture and maintaining the object characteristics during the image generation process is achieved, which improves the image generation effect and ensures the consistency between the generated image and the original image in key features.
Smart Images

Figure CN2024133475_30052025_PF_FP_ABST
Abstract
Description
Image generation method, device, electronic device, and computer-readable medium
[0001] This application claims priority to Chinese patent application No. 202311559362.X filed on November 21, 2023. The contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The present disclosure relates to an image generation method, an apparatus, an electronic device, and a computer-readable medium. Background Art
[0003] For some image generation scenarios, there may be a requirement to generate a new image based on some information so that the new image is consistent with the information. For ease of understanding, the following is an example.
[0004] As an example, for some facial image generation scenarios, these facial image generation scenarios may have the following requirements: generating a new image based on some facial state description information so that the facial state represented by the new image is consistent with the facial state described by the facial state description information. Summary of the Invention
[0005] The present disclosure provides an image generation method, device, electronic device, and computer-readable medium, which can achieve the requirements shown above.
[0006] In order to achieve the above objectives, the technical solutions provided by the present disclosure are as follows:
[0007] The present disclosure provides an image generation method, the method comprising:
[0008] Acquire first object contour description data, first object posture description data, and second object posture description data; the first object contour description data and the first object posture description data are both determined based on the first object description image;
[0009] Correcting the first object contour description data according to difference representation data between the second object posture description data and the first object posture description data to obtain corrected contour description data;
[0010] Image generation processing is performed based on the corrected contour description data and the second object posture description data to obtain a generated image; the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature other than the object posture.
[0011] In one possible implementation, the difference between the object posture of the contour described by the corrected contour description data and the object posture described by the second object posture description data is smaller than the difference between the object posture of the contour described by the first object contour description data and the object posture described by the second object posture description data.
[0012] In a possible implementation manner, the first object contour description data is an image;
[0013] The process of determining the corrected contour description data includes:
[0014] determining a position offset based on difference representation data between the second object posture description data and the first object posture description data;
[0015] According to the position offset, pixel position correction processing is performed on the contour description pixel points in the first object contour description data to obtain corrected contour description data.
[0016] In one possible implementation, the first object posture description data is used to describe the position of at least one first key point; the second object posture description data is used to describe the position of at least one second key point; if there is a corresponding relationship between a first target point in the at least one first key point and a second target point in the at least one second key point, the difference characterization data includes the position difference between the first target point and the second target point; the position offset is determined based on the position difference.
[0017] In a possible implementation manner, performing pixel position correction processing on the contour description pixel points in the first object contour description data based on the position offset to obtain corrected contour description data includes:
[0018] According to the position offset, the contour description pixel points in the first object contour description data are subjected to position translation processing to obtain corrected contour description data.
[0019] In a possible implementation manner, the second object posture description data is determined based on a second object description image, and the object posture described by the second object description image is different from the object posture described by the first object description image.
[0020] In a possible implementation, the second object posture description data is obtained by performing perturbation processing on the first object posture description data, so that there is a difference between the object posture described by the second object posture description data and the object posture described by the first object posture description data.
[0021] In one possible implementation, the first object posture description data is obtained by performing key point detection processing on the first object description image using a first key point detection algorithm; the second object posture description data is obtained by performing key point detection processing on the first object description image using a second key point detection algorithm; and the second key point detection algorithm is different from the first key point detection algorithm.
[0022] In one possible implementation, the first object description image includes at least one image to be processed; different images to be processed are used to provide information to be retained under different object features; the first object contour description data includes object contour description data corresponding to each of the images to be processed; and the first object posture description data includes object posture description data corresponding to each of the images to be processed;
[0023] The step of correcting the first object contour description data based on the difference representation data between the second object posture description data and the first object posture description data to obtain corrected contour description data includes:
[0024] For any of the images to be processed, correcting the object contour description data corresponding to the image to be processed based on the difference representation data between the second object posture description data and the object posture description data corresponding to the image to be processed, to obtain corrected contour description data corresponding to the image to be processed;
[0025] The performing image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image includes:
[0026] Image generation processing is performed based on the second object posture description data and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image.
[0027] In one possible implementation, the at least one image to be processed includes a face shape providing image and a hairstyle providing image, wherein the face shape providing image is used to provide information to be retained under the object's face shape, and the hairstyle providing image is used to provide information to be retained under the object's hairstyle.
[0028] In one possible implementation, for any of the images to be processed, the object posture description data corresponding to the image to be processed is determined based on the image to be processed, and the object posture description data corresponding to the image to be processed is used to describe the position of at least one key point in the facial area.
[0029] In a possible implementation manner, the object contour description data corresponding to the hairstyle-provided image is hairstyle contour description data corresponding to the hairstyle-provided image.
[0030] In a possible implementation manner, the object contour description data corresponding to the hairstyle-provided image is head contour description data corresponding to the hairstyle-provided image; the head contour description data carries facial contour description information and hairstyle contour description information.
[0031] In one possible implementation, the process of determining the corrected contour description data corresponding to the hairstyle image includes:
[0032] performing information removal processing on face contour description information in head contour description data corresponding to the hairstyle provided image based on the image segmentation result of the hairstyle provided image, to obtain hairstyle contour description data corresponding to the hairstyle provided image;
[0033] Correction processing is performed on the hairstyle contour description data corresponding to the hairstyle provided image based on the difference representation data between the second object posture description data and the object posture description data corresponding to the hairstyle provided image to obtain corrected contour description data corresponding to the hairstyle provided image.
[0034] In one possible implementation manner, before performing image generation processing based on the second object posture description data and the corrected contour description data corresponding to the at least one image to be processed to obtain the generated image, the method further includes:
[0035] performing information removal processing on the face contour description information in the corrected contour description data corresponding to the hairstyle provided image based on the image segmentation result of the hairstyle provided image, to obtain the removed contour description data corresponding to the hairstyle provided image;
[0036] The performing image generation processing based on the second object posture description data and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image includes:
[0037] Image generation processing is performed based on the second object posture description data, the corrected contour description data corresponding to the face shape image, and the removed contour description data corresponding to the hairstyle image to obtain a generated image.
[0038] In one possible implementation, performing image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image includes:
[0039] The first object description image, the corrected contour description data and the second object posture description data are input into a pre-constructed image-based model to obtain the generated image output by the image-based model; the image-based model is used to process the first object description image with the corrected contour description data and the second object posture description data as conditions to obtain the generated image.
[0040] In one possible implementation, performing image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image includes:
[0041] The noise data, the prompt text, the corrected contour description data, and the second object posture description data are input into a pre-built Wensheng graph model to obtain the generated image output by the Wensheng graph model; the Wensheng graph model is used to process the noise data with the prompt text, the corrected contour description data, and the second object posture description data as conditions to obtain the generated image.
[0042] In a possible implementation manner, the at least one object feature includes at least one of the object's face shape and the object's hairstyle.
[0043] The present disclosure provides an image generating device, comprising:
[0044] an acquisition unit, configured to acquire first object contour description data, first object posture description data, and second object posture description data; the first object contour description data and the first object posture description data are both determined based on the first object description image;
[0045] a correction unit, configured to correct the first object contour description data according to difference representation data between the second object posture description data and the first object posture description data, to obtain corrected contour description data;
[0046] a generating unit configured to perform image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image; the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature other than the object posture.
[0047] The present disclosure provides an electronic device, the device comprising: a processor and a memory;
[0048] The memory is used to store instructions or computer programs;
[0049] The processor is configured to execute the instructions or computer program in the memory so that the electronic device executes the image generation method provided by the present disclosure.
[0050] The present disclosure provides a computer-readable medium having instructions or a computer program stored therein. When the instructions or the computer program are executed on a device, the device executes the image generation method provided by the present disclosure.
[0051] The present disclosure provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, wherein the computer program contains program code for executing the image generation method provided by the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0053] FIG1 is a flow chart of an image generation method provided by an embodiment of the present disclosure;
[0054] FIG2 is a schematic diagram of first object outline description data provided by an embodiment of the present disclosure;
[0055] FIG3 is a schematic diagram of corrected contour description data provided by an embodiment of the present disclosure;
[0056] FIG4 is a schematic structural diagram of an image generating device provided by an embodiment of the present disclosure;
[0057] FIG5 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0058] Research has found that for some application scenarios (e.g., image-to-image), the following solution is provided to address the above requirements: an original input image, a contour image (e.g., a facial canny map) extracted from the original input image, and a keypoint description image extracted from a pose-provided image are input into a diffusion model (e.g., a ControlNet model). The diffusion model processes the original input image using the contour image and the keypoint description image as conditions to generate a new image. The contour image is used to provide the contour characteristics (e.g., face shape) of certain object parts (e.g., face) in the original input image, enabling the contour image to perform contour control during image generation. The keypoint description image is used to provide the object's pose (e.g., the location of facial features), enabling the keypoint description image to perform pose control during image generation. However, the new images generated using this solution have defects (e.g., poor facial shape preservation, poor hairstyle preservation, and unsatisfactory pose adjustment), resulting in poor image generation performance.
[0059] Based on the above research, in order to better improve the image generation effect, the present disclosure provides an image generation method, which includes: first obtaining first object contour description data (for example, the contour image extracted from image 1, etc.), first object posture description data (for example, the facial key point description image extracted from image 1, etc.), and second object posture description data (for example, the facial key point description image extracted from Figure 2, etc.), and the first object contour description data and the first object posture description data are both determined based on the first object description image (for example, image 1, etc.), so that the first object contour description data can describe the contours of some object features (for example, face shape, hairstyle, etc.) that need to be maintained in the first object description image, so that the first object posture description data can describe the object posture corresponding to these object features in the first object description image (for example, the positions of five key points used to describe facial posture, etc.), and the second object posture description data can be used to describe the facial key point description image. The description data can describe the object posture required to be presented in the final generated image; then, based on the difference representation data between the second object posture description data and the first object posture description data, the first object contour description data is corrected to obtain corrected contour description data, so that the object posture presented in the corrected contour description data is closer to the object posture described by the second object posture description data; then, image generation processing is performed based on the corrected contour description data and the second object posture description data to obtain a generated image, so that the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature (for example, face shape, hairstyle, etc.) other than the object posture. In this way, the purpose of simultaneously adjusting the object posture and maintaining some object features in the image generation process can be achieved, thereby helping to improve the image generation effect.
[0060] In addition, the present disclosure does not limit the execution subject of the image generation method provided in the embodiments of the present disclosure. For example, the image generation method provided in the embodiments of the present disclosure can be applied to a terminal device or a server. For another example, the image generation method provided in the embodiments of the present disclosure can also be implemented with the help of a data interaction process between a terminal device and a server. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a standalone server, a cluster server, or a cloud server.
[0061] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0062] To better understand the technical solutions provided by the present disclosure, the image generation method provided by the present disclosure is described below with reference to some accompanying figures. As shown in Figure 1, the image generation method provided by an embodiment of the present disclosure includes the following steps S1-S3. Figure 1 is a flowchart of an image generation method provided by an embodiment of the present disclosure.
[0063] S1: Acquire first object contour description data, first object posture description data, and second object posture description data; the first object contour description data and the first object posture description data are both determined based on a first object description image.
[0064] The first object description image is primarily used to describe the state of an object presented by at least one object feature other than the object's pose, so that these object features can be maintained during subsequent image generation. The object refers to the foreground information described by pixels other than background pixels in the first object description image. This disclosure does not limit this object; for example, the object can be a person, an animal, a virtual robot, or an object. The object pose describes the pose of part or all of the object's parts as presented in the first object description image. This disclosure does not limit this object pose. For example, in some application scenarios, if the first object description image is used to describe a person or an animal, the object pose can refer to the object's facial pose (e.g., the location of certain facial landmarks). The at least one object feature refers to a characteristic dimension of the object presented in the first object description image other than the object pose. This disclosure does not limit this at least one object feature. For example, in some application scenarios, if the first object description image is used to describe a person or an animal, and the object pose refers to the object's facial pose, the at least one object feature can include the object's facial shape and / or hairstyle. The object face shape is used to describe the face shape of the object in the first object description image. The object hairstyle is used to describe the hairstyle of the object in the first object description image.
[0065] The first object contour description data refers to data (e.g., an image, etc.) extracted from the first object description image and used to describe the contour of part or all of the object presented by the first object description image, so that the first object contour description data can represent some contour-related characteristics of the part or all of the object presented in the first object description image (e.g., the position of the part or all of the object in the first object description image, the shape of the part or all of the object in the first object description image, the size of the part or all of the object in the first object description image, etc.). For example, the first object contour description data may be contour image 1 shown in FIG.
[0066] In addition, the present disclosure does not limit the representation method of the first object contour description data mentioned above. For example, in some application scenarios, the first object contour description data can be represented by some coordinate data that can represent the location of the contour of part or all parts of an object. For another example, in some application scenarios, the first object contour description data can be represented by an image (for example, the contour image 1 shown in Figure 2). It can be seen that in one possible implementation, the first object contour description data can be a contour image obtained by performing contour acquisition processing on the first object description image mentioned above, so that the contour image can represent the contour of part or all parts (for example, the head) of the object presented by the first object description image in the form of an image.
[0067] Furthermore, the present disclosure does not limit the method for acquiring the first object contour description data. For example, it can be implemented using any existing or future method capable of performing contour acquisition processing on a foreground object in an image (e.g., the Canny method). Thus, in one possible implementation, the process for acquiring the first object contour description data can specifically include: processing the first object description image using a preset contour acquisition method to obtain the first object contour description data (e.g., a contour image such as a Canny image), so that the first object contour description data can accurately represent the contours of some or all parts of the object represented by the first object description image.
[0068] The first object posture description data refers to data (e.g., an image, etc.) extracted from the first object description image and used to describe the posture of part or all of the object presented by the first object description image, so that the first object posture description data can represent some posture-related characteristics of the part or all of the object presented by the first object description image (e.g., the positions of some facial key points, etc.). For example, the first object posture description data can be implemented using a key point description image similar to the target posture graph shown in Figure 2. The key point description image is used to describe the distribution of key points in the part or all of the object.
[0069] In addition, the present disclosure does not limit the representation method of the above-mentioned first object posture description data. For example, in some application scenarios, the first object posture description data can be represented by some key point coordinate data. For another example, in some application scenarios, the first object posture description data can be represented by an image (for example, a key point description image similar to the target posture diagram shown in Figure 2). It can be seen that in one possible implementation method, the first object posture description data can be a key point description image obtained by performing key point acquisition processing on the above-mentioned first object description image (for example, a key point description image similar to the target posture diagram shown in Figure 2), so that the key point description image can represent the posture of part or all parts (for example, the head) of the object presented by the first object description image in the form of an image.
[0070] Furthermore, the present disclosure does not limit the method for acquiring the first object pose description data. For example, it may be implemented using any existing or future method capable of acquiring the pose of a foreground object in an image (e.g., a key point extraction method). Thus, in one possible implementation, the acquisition process of the first object pose description data may specifically include: processing the first object description image using a preset key point extraction method to obtain the first object pose description data (e.g., a key point description image), such that the first object pose description data accurately represents the pose of some or all parts of the object represented by the first object description image.
[0071] Based on the relevant content of the first object contour description data and the first object posture description data described above, it can be seen that because the first object contour description data and the first object posture description data are both determined based on the first object description image described above, so that the first object contour description data and the first object posture description data both describe the state presented in the same image, the first object posture description data can represent the object posture of the contour described by the first object contour description data, and thus the first object posture description data can also be used to describe the object posture presented in the first object contour description data. It can be seen that the object posture presented by the first object contour description data, the object posture described by the first object posture description data, and the object posture presented by the first object description image are all the same.
[0072] The second object pose description data is used to describe the pose of an object so that the pose can be maintained during subsequent image generation. For example, the second object pose description data may be the target pose graph shown in FIG2 . It should be noted that the implementation of the second object pose description data is similar to the implementation of the first object pose description data described above and, for the sake of brevity, will not be repeated here.
[0073] In addition, for the second object posture description data above, there is a difference between the second object posture description data and the first object posture description data above, so that the object posture described by the second object posture description data is different from the object posture described by the first object posture description data, so that the second object posture description data can be used as posture control information to affect the posture generation process involved in the subsequent image generation process, thereby ensuring that the object posture presented in the final generated image is consistent with the object posture described by the second object posture description data.
[0074] In addition, the present disclosure does not limit the above-mentioned process of obtaining the second object posture description data. For ease of understanding, the following description is given in conjunction with some situations.
[0075] Case 1: For some application scenarios, the object pose required for the image generation process comes from one image, but other object features except the object pose come from other images. In this way, it is possible to generate a new image by combining the different characteristics of different images.
[0076] Based on the above situation 1, it can be seen that under one possible implementation, when the above first object contour description data and the above first object posture description data are both determined based on the first object description image, the above second object posture description data can be determined based on the second object description image, and the object posture described by the second object description image is different from the object posture described by the first object description image, so that the object posture described by the second object posture description data is different from the object posture described by the first object posture description data, and thus the object posture described by the second object posture description data is different from the object posture of the contour described by the first object contour description data. The second object description image is mainly used to provide an object posture so that the object posture can be maintained in the subsequent image generation process.
[0077] It should be noted that the present disclosure does not limit the process for acquiring the second object pose description data described in the preceding paragraph. For example, the implementation of the second object pose description data acquisition process is similar to the implementation of the first object pose description data acquisition process described above. Therefore, in one possible implementation, the process for acquiring the second object pose description data may specifically include: processing the second object description image described above using a preset key point extraction method to obtain the second object pose description data (e.g., a key point description image), so that the second object pose description data can accurately represent the pose of some or all parts of the object represented by the second object description image.
[0078] Case 2: In some application scenarios (for example, sample image augmentation scenarios, etc.), the second object posture description data can be obtained by perturbing the first object posture description data above, so that there is a difference between the second object posture description data and the first object posture description data, so that a new image different from the first object description image above can be generated based on the second object posture description data.
[0079] Based on the above situation 2, it can be seen that in a possible implementation manner, when the above-mentioned first object contour description data and the above-mentioned first object posture description data are both determined based on the first object description image, the above-mentioned second object posture description data can be obtained by perturbation processing on the first object posture description data, so that there is a difference between the object posture described by the second object posture description data and the object posture described by the first object posture description data, so that the object posture described by the second object posture description data is different from the object posture of the contour described by the first object contour description data, and thus there is a difference in object posture between the new image subsequently generated based on the second object posture description data and the first object description image, so that it is possible to obtain a new image different from the first object description image using the first object description image.
[0080] It should be noted that the present disclosure does not limit the implementation method of the disturbance processing in the above paragraph. For example, it can be implemented by adopting any existing or future method that can perform disturbance processing on a piece of data (for example, image data).
[0081] Case 3: For some application scenarios, in order to better meet certain requirements in these application scenarios, different posture extraction methods (for example, posture extraction methods with different accuracies, posture extraction methods with different working principles, etc.) can be used to obtain posture description data of the same image, so that the contour image extracted from the image can be corrected based on the difference between the two posture description data.
[0082] Based on the above scenario 3, it can be seen that in one possible implementation, when the first object pose description data is obtained by performing key point detection processing on the first object description image using a first key point detection algorithm, the second object pose description data may be obtained by performing key point detection processing on the first object description image using a second key point detection algorithm. Because the second key point detection algorithm is different from the first key point detection algorithm, there is a difference between the second object pose description data and the first object pose description data, and thus there is a certain difference between the object pose described by the second object pose description data and the object pose described by the first object pose description data.
[0083] It should be noted that for the first and second object pose description data shown in the previous paragraph, in order to better ensure the normal progress of the subsequent contour correction process, the two data must also meet the following constraints: there must be an intersection between the key points described by the first object pose description data and the key points described by the second object pose description data, so that the subsequent contour correction process can be completed by using the position differences of the key points in the intersection. For example, if the first object pose description data is used to describe the positions of 106 facial key points, and the second object pose description data is used to describe the positions of 5 facial key points (e.g., points 1 to 5 shown in Figure 2), then the 106 facial key points include some or all of the 5 facial key points.
[0084] It should also be noted that the present disclosure does not limit the implementation methods of the first key point detection algorithm and the second key point detection algorithm. For example, the two can be set in advance according to the actual application scenario.
[0085] It should be noted that the present disclosure does not limit the determination process of the first object posture description data and the second object posture description data under the above-mentioned situation 3. For example, it can be specifically as follows: first, use multiple key point detection algorithms to perform key point detection processing on the above-mentioned first object description image respectively, and obtain the key point detection results corresponding to each key point detection algorithm; then calculate the key point similarity between the key point detection results corresponding to each key point detection algorithm and the key points in the above-mentioned first object contour description data; then, use the key point detection algorithm with the highest key point similarity as the first key point detection algorithm, and use the key point detection algorithm corresponding to the key point detection algorithm with the highest key point similarity as the first key point detection algorithm. The result is used as the first object posture description data; and any other key point detection algorithm except the key point detection algorithm with the highest key point similarity is used as the second key point detection algorithm, and the key point detection result corresponding to any other key point detection algorithm is used as the second object posture description data, so that there are some subtle differences between the object posture presented by the new image subsequently generated based on the second object posture description data and the object posture presented by the first object description image, so that the binary group (the new image, the first object description image) can be used subsequently to complete other tasks (for example, tasks such as recognition processing with subtle differences).
[0086] Based on the relevant content of S1 above, it can be known that in some application scenarios (for example, head image generation scenarios), if it is desired to maintain the object posture (for example, facial posture) described by the second object posture description data and maintain some other object features (for example, face shape, hairstyle, etc.) described by the first object description image in addition to the object posture during the image generation process, the second object posture description data can be obtained, and the first object contour description data and the first object posture description data can be extracted from the first object description image, so that the first object contour description data can be subsequently adjusted based on the difference in object posture between the first object posture description data and the second object posture description data, so as to reduce (or even eliminate) the gap between the object posture described by the first object contour description data and the object posture described by the second object posture description data as much as possible. This is conducive to avoiding defects caused by a large gap between the object posture described by the first object contour description data and the object posture described by the second object posture description data (for example, defects such as inability to maintain face shape or hairstyle, poor posture adjustment effect, etc.).
[0087] S2: Correcting the first object contour description data based on the difference representation data between the second object posture description data and the first object posture description data to obtain corrected contour description data.
[0088] The difference representation data between the second object posture description data and the first object posture description data is used to represent the difference between the object posture described by the second object posture description data and the object posture described by the first object posture description data.
[0089] In addition, the present disclosure does not limit the implementation method of the difference representation data in the above paragraph. For example, in some application scenarios, when the first object posture description data above is used to describe the position of at least one first key point, and the second object posture description data above is used to describe the position of at least one second key point, if there is a corresponding relationship between the first target point in the at least one first key point and the second target point in the at least one second key point, then the difference representation data may include the position difference between the first target point and the second target point. The first key point refers to the key point described by the first object posture description data. The second key point refers to the key point described by the second object posture description data. The first target point refers to the key point that exists in the key points described by the first object posture description data and corresponds to the second target point in the at least one second key point. The second target point refers to the key point that exists in the key points described by the second object posture description data and corresponds to the first target point in the at least one first key point. The corresponding relationship between the first target point and the second target point is used to indicate that the first target point and the second target point represent the same target (e.g., the left eye). The position difference between the first target point and the second target point is used to represent a difference between a position of the first target point described by the first object posture description data and a position of the second target point described by the second object posture description data.
[0090] Based on the content of the above paragraph, it can be known that for some application scenarios, after obtaining the above-mentioned first object posture description data and the above-mentioned second object posture description data, if there is a corresponding relationship between the i-th first target point described by the first object posture description data and the i-th second target point described by the first object posture description data (for example, the i-th first target point and the i-th second target point are both used to represent the same target, etc.), then the difference between the position of the i-th first target point recorded by the first object posture description data and the position of the i-th second target point recorded by the second object posture description data can be calculated first, as the difference between the position of the i-th first target point and the i-th second target point. The position difference between the punctuation points, i is a positive integer, i≤I, I represents the number of the first target points or the number of the second target points; then, the position difference between the first first target point and the first second target point, the position difference between the second first target point and the second second target point, ... (and so on), and the position difference between the I first target point and the I second target point are collected to obtain the difference representation data between the second object posture description data and the first object posture description data, so that the difference representation data includes these position differences, so that the first object contour description data can be corrected based on the difference representation data to obtain the corrected contour description data.
[0091] The corrected contour description data refers to the correction processing result of the first object contour description data, so that the difference between the object posture of the contour described by the corrected contour description data and the object posture described by the second object posture description data is smaller than the difference between the object posture described by the first object posture description data and the object posture described by the second object posture description data, so that the difference between the object posture of the contour described by the corrected contour description data and the object posture described by the second object posture description data is smaller than the difference between the object posture of the contour described by the first object contour description data and the object posture described by the second object posture description data, so that the object posture presented by the corrected contour description data is as close as possible to (or even completely equal to) the object posture described by the second object posture description data. For example, if the first object contour description data is the contour image 1 shown in Figure 2, and the second object posture description data is the target posture diagram shown in Figure 2, then the corrected contour description data can be the contour image 2 shown in Figure 3.
[0092] In addition, the present disclosure does not limit the above-mentioned process of acquiring the corrected contour description data (ie, the implementation of S2 ); for example, it may be implemented using any existing or future correction method.
[0093] In addition, in order to better improve the correction effect, the present disclosure also provides a possible implementation of the above S2. In this implementation, if the above first object contour description data is an image, then S2 may include the following steps 11 and 12.
[0094] Step 11: Determine a position offset based on the difference representation data between the second object posture description data and the first object posture description data.
[0095] The position offset refers to the offset required to correct the position of the outline description pixel in the first object outline description data. The outline description pixel refers to the pixel carrying outline information in the first object outline description data (for example, the pixel used to represent the color black in the outline image 1 shown in Figure 2). For example, the position offset may include an offset in at least one coordinate dimension (for example, a horizontal coordinate and a vertical coordinate). The at least one coordinate dimension refers to the coordinate dimension involved in the image coordinate system in the first object outline description data.
[0096] In addition, the present disclosure does not limit the determination process of the above position offset. For example, when the difference representation data between the second object posture description data and the first object posture description data include the position difference between the first first target point and the first second target point, the position difference between the second first target point and the second second target point, ... (and so on), and the position difference between the Ith first target point and the Ith second target point, if the position offset can include the offset in at least one coordinate dimension (for example, dimensions such as the horizontal coordinate and the vertical coordinate), then the position offset can be specifically: calculate the average value of the values presented by all position differences in the difference representation data in the jth coordinate dimension to obtain the offset in the jth coordinate dimension, where j is a positive integer, j≤J, and J represents the number of dimensions in the at least one coordinate dimension.
[0097] Step 12: Based on the above position offset, pixel position correction processing is performed on the contour description pixel points in the first object contour description data to obtain corrected contour description data.
[0098] It should be noted that the present disclosure does not limit the implementation method of the above step 12. For example, the specific implementation method of step 12 may be: after obtaining the above position offset, the pixel position of the k-th contour description pixel point in the first object contour description data may be added to the position offset to obtain the offset position of the k-th contour description pixel point (for example, the coordinate value of the pixel position of the k-th contour description pixel point in the j-th coordinate dimension is added to the offset of the position offset in the j-th coordinate dimension to obtain the coordinate value of the offset position of the k-th contour description pixel point in the j-th coordinate dimension, where j is a positive integer, j≤J, etc.), k is a positive integer, k≤K, K is a positive integer, and K Represents the number of contour description pixels in the first object contour description data; then, the pixel value at the offset position of the k-th contour description pixel in the first object contour description data is set as the pixel value of the k-th contour description pixel, k is a positive integer, k≤K, K is a positive integer, and the pixel values at other pixel positions in the first object contour description data except the offset positions of the K contour description pixels are set as background pixel values, to obtain corrected contour description data corresponding to the first object contour description data, so that the object posture described by the corrected contour description data is closer to (or even equal to) the object posture described by the second object posture description data above.
[0099] For example, in some application scenarios, step 12 may specifically be performed by performing positional translation processing on the contour description pixels in the first object contour description data according to the positional offset, to obtain corrected contour description data, such that the positional difference between the position of each contour description pixel in the corrected contour description data and the position of the corresponding contour description pixel in the first object contour description data is the positional offset. It should be noted that the present disclosure does not limit the implementation method of the positional translation processing; for example, it may be implemented using any existing or future method capable of positionally translating foreground pixels in an image.
[0100] Based on the relevant content of S2 above, it can be known that for some application scenarios, after obtaining the above second object posture description data and extracting the first object posture description data and the first object contour description data from the above first object description image, the difference representation data between the second object posture description data and the first object posture description data can be calculated first, so that the difference representation data can represent the difference between the object posture described by the second object posture description data and the object posture described by the first object posture description data, thereby enabling the difference representation data to represent the difference between the object posture described by the second object posture description data and the object posture presented by the first object contour description data; and then the first object contour description data is corrected based on the difference representation data to obtain corrected contour description data, so that the object posture presented by the corrected contour description data is closer to the object posture described by the second object posture description data, so that the corrected contour description data can be used as contour control information to affect the generation process of other object features other than the object posture involved in the subsequent image generation process. Among them, since the difference in object posture between the corrected contour description data and the second object posture description data is small, the corrected contour description data and the second object posture description data have a high consistency in the object posture, so that when generating an image based on the two data, the defects caused by the large difference in object posture between the two data can be better avoided, which is beneficial to improving the image generation effect.
[0101] S3: Perform image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image; the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature except the object posture.
[0102] Among them, the generated image refers to a new image generated based on the above-mentioned corrected contour description data and the above-mentioned second object posture description data, so that the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the contour features (such as face shape, hairstyle, etc.) presented by the generated image are consistent with the contour features described by the corrected contour description data, so that the generated image is consistent with the above-mentioned first object description image in at least one object feature (such as face shape, hairstyle, etc.) other than the object posture.
[0103] In addition, the present disclosure does not limit the implementation of the above S3. For ease of understanding, some situations are described below.
[0104] Case 1: In some application scenarios (eg, image-to-image scenarios, etc.), a new image can be obtained by processing the image described above for the first object (eg, noise addition + denoising, etc.).
[0105] Based on the above situation, it can be seen that in one possible implementation, S3 above can be specifically: inputting the above first object description image, the above corrected contour description data, and the above second object posture description data into a pre-built graph-based graph model to obtain a generated image output by the graph-based graph model, so that the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature other than the object posture, so that the object posture can be adjusted while maintaining certain object features (such as hairstyle or face shape) in the first object description image, thereby facilitating the image generation effect. In which, the graph-based graph model is used to process the first object description image with the corrected contour description data and the second object posture description data as conditions to obtain the generated image; and the present disclosure does not limit the implementation of the graph-based graph model. For example, it can be implemented using any existing or future model that can generate a new graph based on a graph and some conditions (for example, a diffusion model similar to Controlnet).
[0106] Case 2: In some application scenarios (eg, text-image scenarios, etc.), the prompt text, the above corrected contour description data, and the above second object posture description data can all be used as conditions for image generation processing.
[0107] Based on the above situation 2, it can be seen that in one possible implementation, the above S3 can be specifically: inputting the noise data, prompt text, the above corrected contour description data and the above second object posture description data into a pre-built text-based graph model to obtain a generated image output by the text-based graph model, so that the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the above first object description image in at least one object feature other than the object posture. Among them, noise data refers to the noise data required for image generation processing; and the present disclosure does not limit the method for generating the noise data. For example, it can be implemented using any existing or future noise data generation method (for example, a method for randomly generating noise data or a method for adding noise to certain images, etc.). The prompt text refers to the text required to be referenced when performing image processing, so that the prompt text can describe a certain constraint condition in the form of a character string; and the present disclosure does not limit the prompt text. The text-generated graph model is used to process the noise data using the prompt text, the corrected contour description data, and the second object posture description data as conditions to generate an image. Moreover, the present disclosure does not limit the implementation method of the text-generated graph model. For example, it can be implemented using any existing or future model (e.g., a diffusion model) that can generate a new graph based on noise data, prompt text, and some additional conditions.
[0108] Based on the relevant contents of S1 to S3 above, it can be known that for the image generation method provided by the embodiment of the present disclosure, firstly obtain the first object contour description data (for example, the contour image extracted from image 1, etc.), the first object posture description data (for example, the facial key point description image extracted from image 1, etc.), and the second object posture description data (for example, the facial key point description image extracted from FIG. 2, etc.), and the first object contour description data and the first object posture description data are both determined based on the first object description image (for example, image 1, etc.), so that the first object contour description data can describe the contours of some object features (for example, face shape, hairstyle, etc.) that need to be maintained in the first object description image, so that the first object posture description data can describe the object posture corresponding to these object features in the first object description image (for example, the positions of the five key points used to describe the facial posture, etc.), and the second object posture description data can be used to describe the object posture corresponding to these object features in the first object description image (for example, the positions of the five key points used to describe the facial posture, etc.), and the second object posture description data can be used to describe the object posture corresponding to these object features in the first object description image. The data can describe the object posture required to be presented in the final generated image; then, based on the difference representation data between the second object posture description data and the first object posture description data, the first object contour description data is corrected to obtain corrected contour description data, so that the object posture presented in the corrected contour description data is closer to the object posture described by the second object posture description data; then, image generation processing is performed based on the corrected contour description data and the second object posture description data to obtain a generated image, so that the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature (for example, face shape, hairstyle, etc.) other than the object posture. In this way, the purpose of simultaneously adjusting the object posture and maintaining some object features in the image generation process can be achieved, thereby helping to improve the image generation effect.
[0109] In addition, based on the relevant content of the first object description image involved in the above image generation method, it can be known that in some application scenarios, the first object description image may refer to an image, and the image can simultaneously provide multiple object features that need to be maintained (for example, object features such as object hairstyle and object face shape), so that a new image can be generated subsequently with the help of the process shown in S1-S3 above, thereby realizing image generation processing while maintaining multiple object features provided by an image.
[0110] In addition, for some application scenarios, these application scenarios may have the following requirements: different object features in at least one of the above object features (for example, features such as hairstyle and face shape) come from different images, so as to achieve maintaining different object features provided by different images during the image generation process (for example, the object face shape provided by one image, the object hairstyle provided by another image, etc.).
[0111] In order to better meet the above requirements, the present disclosure also provides a possible implementation of the image generation method. In this implementation, the image generation method includes the following steps 21 to 23.
[0112] Step 21: Acquire first object contour description data, first object posture description data, and second object posture description data; the first object contour description data and the first object posture description data are both determined based on a first object description image; the first object description image includes at least one image to be processed; different images to be processed are used to provide information to be retained under different object features; the first object contour description data includes object contour description data corresponding to each image to be processed; the first object posture description data includes object posture description data corresponding to each image to be processed.
[0113] In the present disclosure, in some application scenarios, if different images are required to provide different object features that need to be maintained, N images to be processed can be obtained first, so that different images to be processed are used to provide information to be maintained under different object features, where N is a positive integer; then, for the nth image to be processed used to provide information to be maintained under the nth object feature, contour extraction processing is performed on the nth image to be processed to obtain object contour description data corresponding to the nth image to be processed, and posture extraction processing (for example, key point extraction processing) is performed on the nth image to be processed to obtain object posture description data corresponding to the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can be corrected based on the difference representation data between the second object posture description data above and the object posture description data corresponding to the nth image to be processed, where n is a positive integer, n≤N, and N is a positive integer.
[0114] The nth image to be processed is used to describe the state of an object under the nth object feature in addition to the object's posture, so that the nth image to be processed is used to provide the information to be retained under the nth object feature, so that the information to be retained under the nth object feature can be retained in the subsequent image generation process. n is a positive integer, n≤N, and N is a positive integer. The information to be retained refers to the pixel information present in the nth image to be processed that is used to represent the nth object feature.
[0115] It should be noted that the present disclosure does not limit the implementation of the information to be retained in the above paragraph. For example, if the nth image to be processed is used to retain the information to be retained for the subject's facial shape, then the information to be retained may refer to the pixel information present in the nth image to be processed that represents the subject's facial shape. For another example, if the nth image to be processed is used to retain the information to be retained for the subject's hairstyle, then the information to be retained may refer to the pixel information present in the nth image to be processed that represents the subject's hairstyle.
[0116] It should also be noted that the present disclosure is not limited to the aforementioned embodiments of the at least one image to be processed. For example, the at least one image to be processed may include a face shape image and a hairstyle image. The face shape image is used to provide information to be retained for the subject's face shape. The hairstyle image is used to provide information to be retained for the subject's hairstyle.
[0117] The object contour description data corresponding to the nth image to be processed above refers to data (for example, an image, etc.) extracted from the nth image to be processed and used to describe the contours of part or all of the parts of the object presented by the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent some contour-related characteristics of the part or all of the parts presented in the nth image to be processed (for example, the position of the part or all of the parts in the nth image to be processed, the shape presented by the part or all of the parts in the nth image to be processed, the size presented by the part or all of the parts in the nth image to be processed, etc.).
[0118] In addition, the present disclosure does not limit the implementation method of the object contour description data corresponding to the nth image to be processed. For ease of understanding, some examples are provided below for illustration.
[0119] For example, if the nth image to be processed is used to maintain the information to be maintained regarding the face of an object (e.g., the nth image to be processed provides an image of the face), then the object contour description data corresponding to the nth image to be processed can be implemented using the facial contour description data corresponding to the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent the contour-related features of the subject's face as presented in the nth image to be processed. The facial contour description data corresponding to the nth image to be processed refers to data (e.g., an image, etc.) extracted from the nth image to be processed and used to describe the contour of the face of the object presented by the nth image to be processed.
[0120] Based on the above, it can be seen that in one possible implementation, when the nth image to be processed above provides an image of a face shape, the object contour description data corresponding to the nth image to be processed can be the face shape contour description data corresponding to the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent the contour-related characteristics of the object face shape presented in the nth image to be processed.
[0121] For example, if the nth image to be processed is used to preserve information related to a subject's hairstyle (e.g., the nth image to be processed provides an image of the hairstyle), the subject contour description data corresponding to the nth image to be processed can be implemented using the hairstyle contour description data corresponding to the nth image to be processed, so that the subject contour description data corresponding to the nth image to be processed can represent the contour-related features of the subject's hairstyle as presented in the nth image to be processed. The hairstyle contour description data corresponding to the nth image to be processed refers to data (e.g., an image, etc.) extracted from the nth image to be processed and used to describe the contour of the subject's hair as presented by the nth image to be processed.
[0122] Based on the above, it can be seen that in one possible implementation, when the nth image to be processed above provides an image of a hairstyle, the object contour description data corresponding to the nth image to be processed can be the hairstyle contour description data corresponding to the nth image to be processed, so that the object contour description data corresponding to the nth image to be processed can represent the contour-related characteristics of the object hairstyle presented in the nth image to be processed.
[0123] The object posture description data corresponding to the nth image to be processed above refers to data (for example, an image, etc.) extracted from the nth image to be processed and used to describe the posture of part or all parts of the object presented by the nth image to be processed, so that the object posture description data corresponding to the nth image to be processed can represent some posture-related characteristics of the part or all parts presented in the nth image to be processed (for example, the positions of some facial key points, etc.).
[0124] In addition, the present disclosure does not limit the implementation method of the object posture description data corresponding to the nth image to be processed above. For example, when the nth image to be processed provides an image of a face shape or a hairstyle, the object posture description data corresponding to the nth image to be processed can be implemented using the facial key point description data corresponding to the nth image to be processed (for example, the target posture diagram shown in FIG2 ), so that the object posture description data corresponding to the nth image to be processed can represent the facial posture presented by the nth image to be processed. The facial key point description data corresponding to the nth image to be processed is used to describe the positions of some facial key points in the nth image to be processed; and the present disclosure does not limit the representation method of the facial key point description data. For example, in some application scenarios, the facial key point description data can be implemented using an image (for example, the target posture diagram shown in FIG2 ).
[0125] Based on the relevant content of step 21 above, it can be known that in some application scenarios (for example, other different object features other than the object posture come from different image scenes, etc.), if you want to maintain the object posture described by the second object posture description data and maintain some object features described by some images to be processed during the image generation process, you can obtain the second object posture description data, and extract the object contour description data and object posture description data corresponding to each image to be processed from each image to be processed, so that the object contour description data corresponding to the corresponding image to be processed can be adjusted based on the difference in object posture between the object posture description data corresponding to each image to be processed and the second object posture description data, so as to reduce (or even eliminate) the gap between the object posture described by the object contour description data and the object posture described by the second object posture description data as much as possible. This is conducive to avoiding defects caused by a large gap between the object posture described by the object contour description data and the object posture described by the second object posture description data.
[0126] Step 22: For any image to be processed, based on the difference representation data between the second object posture description data and the object posture description data corresponding to the image to be processed, the object contour description data corresponding to the image to be processed is corrected to obtain the corrected contour description data corresponding to the image to be processed.
[0127] In the present disclosure, for the nth image to be processed, after obtaining the object contour description data and the object posture description data corresponding to the nth image to be processed, the difference representation data between the second object posture description data and the object posture description data corresponding to the nth image to be processed can be calculated first, so that the difference representation data can represent the difference between the object posture described by the second object posture description data and the object posture described by the nth image to be processed, so that the difference representation data can represent the difference between the object posture described by the second object posture description data and the object posture presented by the object contour description data corresponding to the nth image to be processed; then, based on the difference representation data, the object contour description data corresponding to the nth image to be processed is corrected to obtain the corrected contour description data corresponding to the nth image to be processed, so that the difference between the object posture described by the corrected contour description data and the object posture described by the second object posture description data is smaller (or even has no difference). Wherein, n is a positive integer, n≤N.
[0128] Step 23: Perform image generation processing based on the second object posture description data and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image; the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the nth object feature described by the generated image is consistent with the nth object feature described by the nth image to be processed, where n is a positive integer, n≤N, and N represents the number of images in the at least one image to be processed.
[0129] It should be noted that the implementation of step 23 is similar to the implementation of S3 above, and for the sake of brevity, it will not be repeated here.
[0130] Based on the relevant content of steps 21 to 23 above, it can be seen that in some application scenarios, if it is necessary to perform image generation processing based on different object features provided by different images, the object contour description data corresponding to the corresponding image can be corrected based on the difference between the object posture description data corresponding to each image and the second object posture description data above; and then a new image is generated based on these correction results and the second object posture description data. In this way, the object posture can be adjusted while maintaining the object features provided by multiple images, which is beneficial to improving the image generation effect.
[0131] In fact, in order to better improve the image generation effect, the present disclosure also provides some possible implementation methods of the above image generation method, and some implementation methods of the image generation method are used as examples for description below.
[0132] In a possible implementation, the above image generation method may include the following steps 31 to 35 .
[0133] Step 31: Obtain object contour description data corresponding to the face shape provided image, object posture description data corresponding to the face shape provided image, object contour description data corresponding to the hairstyle provided image, object posture description data corresponding to the hairstyle provided image, and second object posture description data.
[0134] Among them, the object contour description data corresponding to the face shape provided image refers to data (for example, an image) extracted from the face shape provided image and used to describe the contour of the face of the object presented by the face shape provided image, so that the object contour description data corresponding to the face shape provided image can represent the contour characteristics of the object's face presented in the face shape provided image; and the present disclosure does not limit the implementation method of the object contour description data corresponding to the face shape provided image. For example, the object contour description data corresponding to the face shape provided image can be implemented using the face contour description data corresponding to the face shape provided image.
[0135] The object posture description data corresponding to the face shape provided image refers to data (for example, an image) extracted from the face shape provided image and used to describe the posture of the face of the object presented by the face shape provided image, so that the object posture description data corresponding to the face shape provided image can represent the posture characteristics of the object's face presented in the face shape provided image; and the present disclosure does not limit the implementation method of the object posture description data corresponding to the face shape provided image. For example, the object posture description data corresponding to the face shape provided image can be implemented using the facial key point description data corresponding to the face shape provided image (for example, a key point description image similar to the target posture diagram shown in Figure 2).
[0136] The object contour description data corresponding to the hairstyle provided image refers to data (for example, an image) extracted from the hairstyle provided image and used to describe the contour of the hair in the object presented by the hairstyle provided image, so that the object contour description data corresponding to the hairstyle provided image can represent the contour characteristics of the object's hair presented in the hairstyle provided image.
[0137] In addition, for the object contour description data corresponding to the above hairstyle-provided image, in order to better improve the subsequent posture correction effect, the object contour description data corresponding to the hairstyle-provided image and the object posture description data corresponding to the hairstyle-provided image can be aligned in some aspects.
[0138] Based on the above, to further improve the correction effect, the present disclosure also provides a possible implementation of the object contour description data corresponding to the hairstyle-provided image. In this implementation, the object contour description data corresponding to the hairstyle-provided image can be implemented using the head contour description data corresponding to the hairstyle-provided image (e.g., contour image 1 shown in FIG2 ), so that the object contour description data corresponding to the hairstyle-provided image can represent the contour features (e.g., facial contour features and hair contour features) presented by the subject's head. The head contour description data corresponding to the hairstyle-provided image refers to data (e.g., an image) extracted from the hairstyle-provided image that describes the contour of the subject's head presented by the hairstyle-provided image. This allows the object contour description data corresponding to the hairstyle-provided image to represent the contour features presented by the subject's head in the hairstyle-provided image, thereby causing the head contour description data corresponding to the hairstyle-provided image to carry facial contour description information and hairstyle contour description information. The facial contour description information is used to represent the contour features of the facial region of the subject's head as presented in the hairstyle-provided image. The hairstyle contour description information is used to represent the contour features of the hair region of the subject's head as presented in the hairstyle-provided image.
[0139] The object posture description data corresponding to the hairstyle-provided image refers to data (e.g., an image) extracted from the hairstyle-provided image and used to describe the posture of the face of the object presented by the hairstyle-provided image, so that the object posture description data corresponding to the hairstyle-provided image can represent the posture characteristics of the object's face presented in the hairstyle-provided image; and the present disclosure does not limit the implementation method of the object posture description data corresponding to the hairstyle-provided image. For example, the object posture description data corresponding to the hairstyle-provided image can be implemented using the facial key point description data corresponding to the hairstyle-provided image (e.g., a key point description image similar to the target posture diagram shown in Figure 2).
[0140] Based on the relevant content of step 31 above, it can be known that in some application scenarios, if you want to maintain the object posture described by the second object posture description data, maintain the object face shape described by the face shape provided image, and maintain the object hairstyle described by the hairstyle provided image during the image generation process, then you can obtain the second object posture description data, extract the object contour description data and object posture description data corresponding to the face shape provided image from the face shape provided image, and extract the object contour description data and object posture description data corresponding to the hairstyle provided image from the hairstyle provided image.
[0141] Step 32: Based on the difference representation data between the second object posture description data and the object posture description data corresponding to the face-shaped image, the object contour description data corresponding to the face-shaped image is corrected to obtain corrected contour description data corresponding to the face-shaped image.
[0142] It should be noted that the implementation of step 32 is similar to the implementation of step 22 above, and for the sake of brevity, it will not be repeated here.
[0143] Step 33: Based on the image segmentation result of the hairstyle image provided above, the facial contour description information in the object contour description data (i.e., the head contour description data) corresponding to the hairstyle image is removed to obtain the hairstyle contour description data corresponding to the hairstyle image provided.
[0144] The image segmentation result of the hairstyle-provided image is used to represent the locations of different parts of an object in the hairstyle-provided image. Furthermore, the present disclosure does not limit the implementation of the image segmentation result of the hairstyle-provided image. For example, the image segmentation result of the hairstyle-provided image may include at least facial position representation data and hair position representation data. The facial position representation data is used to describe the location of the face of the object in the hairstyle-provided image. The hair position representation data is used to describe the location of the hair of the object in the hairstyle-provided image.
[0145] In addition, the present disclosure does not limit the implementation method of the above step 33. For example, when the head contour description data corresponding to the above hairstyle-provided image is an image, the step 33 may specifically be: first, obtain facial position representation data from the image segmentation result of the hairstyle-provided image; then, adjust the pixel values at the pixel positions described by the facial position representation data in the head contour description data corresponding to the hairstyle-provided image to background pixel values, and obtain the hairstyle contour description data corresponding to the hairstyle-provided image, so that only the hairstyle contour description information is retained in the hairstyle contour description data, so that the hairstyle contour description data can represent some contour features of the object hairstyle presented in the hairstyle-provided image.
[0146] Step 34: Based on the difference representation data between the second object posture description data and the object posture description data corresponding to the hairstyle provided image, the hairstyle contour description data corresponding to the hairstyle provided image is corrected to obtain corrected contour description data corresponding to the hairstyle provided image.
[0147] It should be noted that the implementation of step 34 is similar to the implementation of step 22 above, and for the sake of brevity, it will not be repeated here.
[0148] Step 35: Perform image generation processing based on the second object posture description data, the corrected contour description data corresponding to the face shape provided image, and the corrected contour description data corresponding to the hairstyle provided image to obtain a generated image; the object posture described by the generated image is consistent with the object posture described by the second object posture description data, the object face shape described by the generated image is consistent with the object face shape described by the face shape provided image, and the object hairstyle described by the generated image is consistent with the object hairstyle described by the hairstyle provided image.
[0149] It should be noted that the implementation of step 35 is similar to the implementation of S3 above, and for the sake of brevity, it will not be repeated here.
[0150] Based on the relevant contents of steps 31 to 35 above, it can be seen that in some application scenarios, if it is desired to maintain the feature of the hairstyle of the object described by the hairstyle-provided image, the head contour description data and the object posture description data corresponding to the hairstyle-provided image can be first extracted from the hairstyle-provided image; then, the face contour description information in the head contour description data corresponding to the hairstyle-provided image is subjected to information removal processing to obtain the hairstyle contour description data corresponding to the hairstyle-provided image; then, the hairstyle contour description data corresponding to the hairstyle-provided image is directly corrected based on the difference representation data between the second object posture description data and the object posture description data corresponding to the hairstyle-provided image to obtain the corrected contour description data corresponding to the hairstyle-provided image, so that image generation processing can be performed subsequently based on the second object posture description data and the corrected contour description data corresponding to the hairstyle-provided image, which is conducive to improving the image generation effect.
[0151] In a possible implementation, the above image generation method may include the following steps 41 to 45 .
[0152] Step 41: Obtain object contour description data corresponding to the face shape provided image, object posture description data corresponding to the face shape provided image, object contour description data corresponding to the hairstyle provided image, object posture description data corresponding to the hairstyle provided image, and second object posture description data.
[0153] It should be noted that for the relevant content of step 41, please refer to step 31 above.
[0154] Step 42: Based on the difference representation data between the second object posture description data and the object posture description data corresponding to the face-shaped image, the object contour description data corresponding to the face-shaped image is corrected to obtain corrected contour description data corresponding to the face-shaped image.
[0155] It should be noted that the implementation of step 42 is similar to the implementation of step 22 above, and for the sake of brevity, it will not be repeated here.
[0156] Step 43: Based on the difference representation data between the second object posture description data and the object posture description data corresponding to the hairstyle-provided image, the object contour description data corresponding to the hairstyle-provided image (that is, the head contour description data) is corrected to obtain corrected contour description data corresponding to the hairstyle-provided image, so that the corrected contour description data carries facial contour description information and hairstyle contour description information, thereby enabling the corrected contour description data to represent the contour characteristics of the head of the object after posture correction.
[0157] It should be noted that the implementation of step 43 is similar to the implementation of step 22 above, and for the sake of brevity, it will not be repeated here.
[0158] Step 44: Based on the image segmentation result of the hairstyle provided image, perform information removal processing on the face contour description information in the corrected contour description data corresponding to the hairstyle provided image to obtain the removed contour description data corresponding to the hairstyle provided image.
[0159] It should be noted that the implementation of step 44 is similar to the implementation of step 33 above. Thus, in one possible implementation, step 44 may specifically include: first, obtaining facial position representation data from the image segmentation result of the hairstyle provided image; then, adjusting the pixel values at the pixel positions described by the facial position representation data in the corrected contour description data corresponding to the hairstyle provided image to background pixel values, thereby obtaining post-removal contour description data corresponding to the hairstyle provided image, so that only the hairstyle contour description information is retained in the post-removal contour description data, thereby enabling the post-removal contour description data to represent the contour characteristics of the hairstyle of the object after posture adjustment.
[0160] Step 45: Perform image generation processing based on the second object posture description data, the corrected contour description data corresponding to the face shape provided image, and the removed contour description data corresponding to the hairstyle provided image to obtain a generated image; the object posture described by the generated image is consistent with the object posture described by the second object posture description data, the object face shape described by the generated image is consistent with the object face shape described by the face shape provided image, and the object hairstyle described by the generated image is consistent with the object hairstyle described by the hairstyle provided image.
[0161] It should be noted that the implementation of step 45 is similar to the implementation of S3 above, and for the sake of brevity, it will not be repeated here.
[0162] Based on the relevant contents of steps 41 to 45 above, it can be seen that in some application scenarios, if it is desired to maintain the feature of the object's hairstyle described by the hairstyle-provided image, the head contour description data and the object posture description data corresponding to the hairstyle-provided image can be first extracted from the hairstyle-provided image; then, based on the difference representation data between the second object posture description data and the object posture description data corresponding to the hairstyle-provided image, the head contour description data corresponding to the hairstyle-provided image is corrected to obtain the corrected contour description data corresponding to the hairstyle-provided image, so that the corrected contour description data can represent the contour characteristics of the object's head (for example, face shape + hairstyle, etc.) after the posture is adjusted; then, the face contour description information in the corrected contour description data corresponding to the hairstyle-provided image is removed to obtain the removed contour description data corresponding to the hairstyle-provided image, so that image generation processing can be performed subsequently based on the second object posture description data and the removed contour description data corresponding to the hairstyle-provided image, which is conducive to improving the image generation effect.
[0163] Based on the image generation method provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide an image generation device, which will be explained and illustrated below in conjunction with Figure 4. Figure 4 is a schematic structural diagram of the image generation device provided in the embodiments of the present disclosure. It should be noted that for the technical details of the image generation device provided in the embodiments of the present disclosure, please refer to the relevant content of the image generation method above.
[0164] As shown in FIG4 , an image generating apparatus 400 provided in an embodiment of the present disclosure includes:
[0165] An acquisition unit 401 is configured to acquire first object contour description data, first object posture description data, and second object posture description data; the first object contour description data and the first object posture description data are both determined based on the first object description image;
[0166] a correction unit 402 configured to correct the first object contour description data based on difference representation data between the second object posture description data and the first object posture description data to obtain corrected contour description data;
[0167] A generation unit 403 is configured to perform image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image; the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature other than the object posture.
[0168] In one possible implementation, the difference between the object posture of the contour described by the corrected contour description data and the object posture described by the second object posture description data is smaller than the difference between the object posture of the contour described by the first object contour description data and the object posture described by the second object posture description data.
[0169] In a possible implementation manner, the first object contour description data is an image;
[0170] The correction unit 402 includes:
[0171] an offset determination subunit, configured to determine a position offset based on difference representation data between the second object posture description data and the first object posture description data;
[0172] The position correction subunit is configured to perform pixel position correction processing on the contour description pixel points in the first object contour description data according to the position offset to obtain corrected contour description data.
[0173] In one possible implementation, the first object posture description data is used to describe the position of at least one first key point; the second object posture description data is used to describe the position of at least one second key point; if there is a corresponding relationship between a first target point in the at least one first key point and a second target point in the at least one second key point, the difference characterization data includes the position difference between the first target point and the second target point; the position offset is determined based on the position difference.
[0174] In a possible implementation manner, the position correction subunit is specifically configured to perform position translation processing on the contour description pixel points in the first object contour description data according to the position offset to obtain corrected contour description data.
[0175] In a possible implementation manner, the second object posture description data is determined based on a second object description image, and the object posture described by the second object description image is different from the object posture described by the first object description image.
[0176] In a possible implementation, the second object posture description data is obtained by performing perturbation processing on the first object posture description data, so that there is a difference between the object posture described by the second object posture description data and the object posture described by the first object posture description data.
[0177] In one possible implementation, the first object posture description data is obtained by performing key point detection processing on the first object description image using a first key point detection algorithm; the second object posture description data is obtained by performing key point detection processing on the first object description image using a second key point detection algorithm; and the second key point detection algorithm is different from the first key point detection algorithm.
[0178] In one possible implementation, the first object description image includes at least one image to be processed; different images to be processed are used to provide information to be retained under different object features; the first object contour description data includes object contour description data corresponding to each of the images to be processed; and the first object posture description data includes object posture description data corresponding to each of the images to be processed;
[0179] The correction unit 402 is specifically configured to: for any of the images to be processed, perform correction processing on the object contour description data corresponding to the image to be processed based on the difference representation data between the second object posture description data and the object posture description data corresponding to the image to be processed, to obtain corrected contour description data corresponding to the image to be processed;
[0180] The generating unit 403 is specifically configured to perform image generation processing according to the second object posture description data and the corrected contour description data corresponding to the at least one image to be processed, to obtain a generated image.
[0181] In one possible implementation, the at least one image to be processed includes a face shape providing image and a hairstyle providing image, wherein the face shape providing image is used to provide information to be retained under the object's face shape, and the hairstyle providing image is used to provide information to be retained under the object's hairstyle.
[0182] In one possible implementation, for any of the images to be processed, the object posture description data corresponding to the image to be processed is determined based on the image to be processed, and the object posture description data corresponding to the image to be processed is used to describe the position of at least one key point in the facial area.
[0183] In a possible implementation manner, the object contour description data corresponding to the hairstyle-provided image is hairstyle contour description data corresponding to the hairstyle-provided image.
[0184] In a possible implementation manner, the object contour description data corresponding to the hairstyle-provided image is head contour description data corresponding to the hairstyle-provided image; the head contour description data carries facial contour description information and hairstyle contour description information.
[0185] In one possible implementation, the correction unit 402 is specifically used to: perform information removal processing on the face contour description information in the head contour description data corresponding to the hairstyle-provided image based on the image segmentation result of the hairstyle-provided image, to obtain the hairstyle contour description data corresponding to the hairstyle-provided image; perform correction processing on the hairstyle contour description data corresponding to the hairstyle-provided image based on the difference representation data between the second object posture description data and the object posture description data corresponding to the hairstyle-provided image, to obtain the corrected contour description data corresponding to the hairstyle-provided image.
[0186] In a possible implementation manner, the image generating device 400 further includes:
[0187] a removal unit configured to perform information removal processing on the face contour description information in the corrected contour description data corresponding to the hairstyle provided image based on the image segmentation result of the hairstyle provided image, to obtain the removed contour description data corresponding to the hairstyle provided image;
[0188] The generating unit 403 is specifically configured to perform image generation processing based on the second object posture description data, the corrected contour description data corresponding to the face shape image, and the removed contour description data corresponding to the hairstyle image to obtain a generated image.
[0189] In one possible implementation, the generation unit 403 is specifically used to: input the first object description image, the corrected contour description data, and the second object posture description data into a pre-constructed image-based model to obtain the generated image output by the image-based model; the image-based model is used to process the first object description image with the corrected contour description data and the second object posture description data as conditions to obtain the generated image.
[0190] In one possible implementation, the generation unit 403 is specifically configured to: input the noise data, the prompt text, the corrected contour description data, and the second object posture description data into a pre-constructed Wensheng graph model to obtain the generated image output by the Wensheng graph model; and the Wensheng graph model is configured to process the noise data using the prompt text, the corrected contour description data, and the second object posture description data as conditions to obtain the generated image.
[0191] In a possible implementation manner, the at least one object feature includes at least one of the object's face shape and the object's hairstyle.
[0192] Based on the relevant content of the above-mentioned image generating device 400, it can be known that for the image generating device 400 provided by the present disclosure, first object contour description data (for example, the contour image extracted from image 1, etc.), first object posture description data (for example, the facial key point description image extracted from image 1, etc.), and second object posture description data (for example, the facial key point description image extracted from Figure 2, etc.) are first obtained, and the first object contour description data and the first object posture description data are both determined based on the first object description image (for example, image 1, etc.), so that the first object contour description data can describe the contours of some object features (for example, face shape, hairstyle, etc.) that need to be maintained in the first object description image, so that the first object posture description data can describe the object posture corresponding to these object features in the first object description image (for example, the positions of the five key points used to describe the facial posture, etc.), and the second object posture description data can be used to describe the facial key point description image. The description data can describe the object posture required to be presented in the final generated image; then, based on the difference representation data between the second object posture description data and the first object posture description data, the first object contour description data is corrected to obtain corrected contour description data, so that the object posture presented in the corrected contour description data is closer to the object posture described by the second object posture description data; then, image generation processing is performed based on the corrected contour description data and the second object posture description data to obtain a generated image, so that the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature (for example, face shape, hairstyle, etc.) other than the object posture. In this way, the purpose of simultaneously adjusting the object posture and maintaining some object features in the image generation process can be achieved, thereby helping to improve the image generation effect.
[0193] In addition, an embodiment of the present disclosure also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the image generation method provided by the embodiment of the present disclosure.
[0194] Referring to FIG5 , a schematic diagram of the structure of an electronic device 500 suitable for implementing embodiments of the present disclosure is shown. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG5 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.
[0195] As shown in Figure 5, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0196] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG5 shows the electronic device 500 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0197] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0198] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0199] The embodiments of the present disclosure further provide a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the image generation method provided by the embodiments of the present disclosure.
[0200] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0201] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0202] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0203] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.
[0204] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0205] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0206] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.
[0207] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0208] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0209] It should be noted that the various embodiments of this disclosure are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the descriptions of the systems or devices disclosed in the embodiments for similarities and differences between them. Since the systems or devices disclosed in the embodiments correspond to the methods disclosed in the embodiments, their descriptions are relatively simple, and reference can be made to the descriptions of the methods for any related details.
[0210] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0211] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0212] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0213] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present disclosure. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein, but is intended to be construed in the widest manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating an image, comprising: Acquire first object contour description data, first object posture description data, and second object posture description data; The first object contour description data and the first object posture description data are both determined based on the first object description image; According to the difference representation data between the second object posture description data and the first object posture description data, the first object contour description data is corrected to obtain corrected contour description data; Image generation processing is performed based on the corrected contour description data and the second object posture description data to obtain a generated image; the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature other than the object posture.
2. The method according to claim 1, wherein: The difference between the object pose of the contour described by the corrected contour description data and the object pose described by the second object pose description data is smaller than the difference between the object pose of the contour described by the first object contour description data and the object pose described by the second object pose description data.
3. The method according to claim 1, wherein: The first object contour description data is an image; The process of determining the corrected contour description data includes: Determining a position offset according to difference representation data between the second object posture description data and the first object posture description data; According to the position offset, pixel position correction processing is performed on the contour description pixel points in the first object contour description data to obtain the corrected contour description data.
4. The method according to claim 3, wherein: The first object posture description data is used to describe the position of at least one first key point; The second object posture description data is used to describe the position of at least one second key point; If there is a corresponding relationship between a first target point in the at least one first key point and a second target point in the at least one second key point, the difference characterization data includes a position difference between the first target point and the second target point; and the position offset is determined based on the position difference.
5. The method according to claim 3 or 4, wherein: The step of performing pixel position correction processing on the contour description pixel points in the first object contour description data according to the position offset to obtain the corrected contour description data includes: According to the position offset, the contour description pixel points in the first object contour description data are subjected to position translation processing to obtain the corrected contour description data.
6. The method according to any one of claims 1 to 5, wherein: The second object posture description data is determined based on a second object description image, and the object posture described by the second object description image is different from the object posture described by the first object description image.
7. The method according to any one of claims 1 to 5, wherein: The second object posture description data is obtained by performing a perturbation process on the first object posture description data, so that there is a difference between the object posture described by the second object posture description data and the object posture described by the first object posture description data.
8. The method according to any one of claims 1 to 5, wherein: The first object posture description data is obtained by performing key point detection processing on the first object description image using a first key point detection algorithm; The second object posture description data is obtained by performing key point detection processing on the first object description image using a second key point detection algorithm; The second key point detection algorithm is different from the first key point detection algorithm.
9. The method according to claim 1, wherein: The first object description image includes at least one image to be processed; different images to be processed are used to provide information to be retained under different object features; the first object contour description data includes object contour description data corresponding to each of the images to be processed; the first object posture description data includes object posture description data corresponding to each of the images to be processed; The step of correcting the first object contour description data based on the difference characterization data between the second object posture description data and the first object posture description data to obtain corrected contour description data includes: For any of the images to be processed, according to the difference representation data between the second object posture description data and the object posture description data corresponding to the image to be processed, correct the object contour description data corresponding to the image to be processed to obtain the corrected contour description data corresponding to the image to be processed; The step of performing image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image includes: Image generation processing is performed according to the second object posture description data and the corrected contour description data corresponding to the at least one image to be processed to obtain the generated image.
10. The method according to claim 9, wherein: The at least one image to be processed includes a face shape providing image and a hairstyle providing image, wherein the face shape providing image is used to provide information to be kept under the face shape of the object, and the hairstyle providing image is used to provide information to be kept under the hairstyle of the object.
11. The method according to claim 10, wherein: For any of the images to be processed, the object posture description data corresponding to the image to be processed is determined based on the image to be processed, and the object posture description data corresponding to the image to be processed is used to describe the position of at least one key point in the facial region.
12. The method according to claim 10, wherein: The object contour description data corresponding to the hairstyle providing image is the hairstyle contour description data corresponding to the hairstyle providing image.
13. The method according to claim 10, wherein: The object contour description data corresponding to the hairstyle-provided image is head contour description data corresponding to the hairstyle-provided image; the head contour description data carries face shape contour description information and hairstyle contour description information.
14. The method according to claim 13, wherein: The process of determining the corrected contour description data corresponding to the hairstyle provided image includes: According to the image segmentation result of the hairstyle provided image, performing information removal processing on the face contour description information in the head contour description data corresponding to the hairstyle provided image, so as to obtain the hairstyle contour description data corresponding to the hairstyle provided image; According to the difference representation data between the second object posture description data and the object posture description data corresponding to the hairstyle provided image, the hairstyle contour description data corresponding to the hairstyle provided image is corrected to obtain the corrected contour description data corresponding to the hairstyle provided image.
15. The method according to claim 13, wherein: Before performing image generation processing based on the second object posture description data and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image, the method further includes: According to the image segmentation result of the image provided by the hairstyle, the face contour description information in the corrected contour description data corresponding to the image provided by the hairstyle is subjected to information removal processing to obtain the removed contour description data corresponding to the image provided by the hairstyle; The step of performing image generation processing based on the second object posture description data and the corrected contour description data corresponding to the at least one image to be processed to obtain a generated image includes: Image generation processing is performed based on the second object posture description data, the corrected contour description data corresponding to the face shape image, and the removed contour description data corresponding to the hairstyle image to obtain the generated image.
16. The method according to any one of claims 1 to 8, wherein: The step of performing image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image includes: The first object description image, the corrected contour description data and the second object posture description data are input into a pre-constructed image-generated image model to obtain the generated image output by the image-generated image model; the image-generated image model is used to process the first object description image with the corrected contour description data and the second object posture description data as conditions to obtain the generated image.
17. The method according to any one of claims 1 to 8, wherein: The step of performing image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image includes: The noise data, the prompt text, the corrected contour description data and the second object posture description data are input into a pre-constructed Wensheng graph model to obtain the generated image output by the Wensheng graph model; the Wensheng graph model is used to process the noise data with the prompt text, the corrected contour description data and the second object posture description data as conditions to obtain the generated image.
18. The method according to any one of claims 1 to 17, wherein: The at least one subject feature includes at least one of a face shape of the subject and a hairstyle of the subject.
19. An image generating device, comprising: an acquisition unit, configured to acquire first object contour description data, first object posture description data, and second object posture description data, wherein the first object contour description data and the first object posture description data are both determined based on the first object description image; a correction unit configured to perform correction processing on the first object contour description data according to difference representation data between the second object posture description data and the first object posture description data to obtain corrected contour description data; A generating unit is configured to perform image generation processing based on the corrected contour description data and the second object posture description data to obtain a generated image, wherein the object posture described by the generated image is consistent with the object posture described by the second object posture description data, and the generated image is consistent with the first object description image in at least one object feature other than the object posture.
20. An electronic device comprising a processor and a memory, wherein: The memory is configured to store instructions or computer programs; The processor is configured to execute the instructions or computer programs in the memory so that the electronic device performs the method according to any one of claims 1 to 18.
21. A computer-readable medium storing instructions or computer programs, which, when executed on a device, enable the device to execute the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Figure hair style replacement method and device based on neural network, equipment and medium
CN112102149A
Image and data processing method and device
CN113327190A
Image processing method and device, processor, electronic equipment and storage medium
CN113569790A
Image generation method and device, electronic equipment and computer readable medium
CN114627529A
Controllable image generation method and device and electronic equipment
CN116363249A