Method and apparatus for generating image, electronic device, and computer readable medium

By acquiring and processing the three-dimensional pose representation data and face texture map of the image, and using the image generation model to generate new images with consistent poses, it solves the problem that pose consistency is difficult to achieve in the image generation process in the prior art, and improves the accuracy and effect of image generation.

WO2025113657A1PCT designated stage expired Publication Date: 2025-06-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/135748
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-11-29
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The prior art is difficult to maintain posture consistency between the original image and the generated image during image generation, especially in terms of hairstyle and face contour.

Method used

By acquiring the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data and the reference face texture map, the image generation model is used to determine the pose motion representation characteristics, and a new image maintaining pose consistency is generated by combining the encoding characteristics of the original image and the encoding characteristics of the reference face texture map.

Benefits of technology

It realizes the pose consistency between the original image and the generated image during the image generation process, and improves the accuracy and effect of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135748_05062025_PF_FP_ABST
    Figure CN2024135748_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method and apparatus for generating an image, an electronic device, and a computer readable medium. The method comprises: first, obtaining an original image, original three-dimensional pose representation data, reference three-dimensional pose representation data, and a reference facial texture map; then, on the basis of the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data, determining a pose motion representation feature; and then, on the basis of an encoding feature of the original image, an encoding feature of the reference facial texture map, and the pose motion representation feature, obtaining a generated image, so that the generated image carries facial state information described by the original image and the reference facial texture map, and also carries three-dimensional space information described by the original three-dimensional pose representation data and the reference three-dimensional pose representation data. As a result, the generated image can depict a three-dimensional facial state, so that the generated image can better describe the facial state, thereby facilitating improving an image generation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method, device, electronic device, and computer-readable medium

[0001] This application claims priority to Chinese Patent Application No. 202311629107.8 filed on November 30, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] The present disclosure relates to an image generation method, an apparatus, an electronic device, and a computer-readable medium. Background Art

[0003] For some image-related scenarios, these scenarios may have the following requirements: generating a new image based on an original image provided by a user, so that the new image is consistent with the original image in some aspects, such as hairstyle, facial contour, etc. Summary of the Invention

[0004] The present disclosure provides an image generation method, an apparatus, an electronic device, and a computer-readable medium.

[0005] In order to achieve the above objectives, the technical solutions provided by the present disclosure are as follows:

[0006] The present disclosure provides an image generation method, the method comprising:

[0007] Acquiring an original image, original three-dimensional pose representation data, reference three-dimensional pose representation data, and a reference facial texture map, wherein the original three-dimensional pose representation data is determined based on the original image; and the reference three-dimensional pose representation data and the reference facial texture map are determined based on reference information corresponding to the original image;

[0008] Determining a gesture motion representation feature based on the original image, the original three-dimensional gesture representation data, and the reference three-dimensional gesture representation data;

[0009] A generated image is obtained based on the coding features of the original image, the coding features of the reference facial texture map, and the posture and motion representation features.

[0010] In one possible implementation, the generated image is determined by an image generation model based on the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference facial texture map;

[0011] After obtaining the generated image, the method further includes:

[0012] The image generation model is updated using the generated image and the supervision information corresponding to the generated image.

[0013] In one possible implementation, the generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image; the supervision information corresponding to the generated image includes supervision information corresponding to the two-dimensional plane image, supervision information corresponding to the depth map, and supervision information corresponding to the background segmentation map.

[0014] In a possible implementation manner, the reference information includes supervision information corresponding to the two-dimensional plane image.

[0015] In one possible implementation, the supervisory information corresponding to the depth map is obtained by performing depth map generation processing on the supervisory information corresponding to the two-dimensional plane image; the supervisory information corresponding to the background segmentation map is obtained by performing background segmentation processing on the supervisory information corresponding to the two-dimensional plane image.

[0016] In one possible implementation, obtaining the generated image based on the encoding features of the original image, the encoding features of the reference facial texture map, and the posture and motion representation features includes:

[0017] Processing the encoded features of the original image according to the posture motion representation features to obtain processed features;

[0018] splicing the processed features and the coded features of the reference facial texture map to obtain spliced ​​features;

[0019] The generated image is obtained according to the splicing features.

[0020] In a possible implementation manner, the generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image;

[0021] Obtaining the generated image based on the splicing features includes:

[0022] The splicing features are decoded to obtain the two-dimensional plane image, the depth map, and the background segmentation map.

[0023] In one possible implementation, the generated image is determined by an image generation model based on the original image, the original three-dimensional posture representation data, the reference three-dimensional posture representation data, and the reference facial texture map; the image generation model includes a motion estimation network, an encoder, and a decoder; the motion estimation network is used to determine the posture motion representation features based on the original image, the original three-dimensional posture representation data, and the reference three-dimensional posture representation data; the encoder is used to determine the encoding features of the original image and the encoding features of the reference facial texture map; the decoder is used to decode the splicing features to obtain the two-dimensional plane image, the depth map, and the background segmentation map.

[0024] In one possible implementation, the generated image includes a three-dimensional image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image;

[0025] Obtaining the generated image based on the splicing features includes:

[0026] Decoding the splicing features to obtain the two-dimensional plane image, the depth map, and the background segmentation map;

[0027] The three-dimensional stereoscopic image is generated according to the two-dimensional planar image and the depth map.

[0028] In one possible implementation, the three-dimensional stereoscopic image includes a first left-eye image and a first right-eye image; the first left-eye image and the first right-eye image are both obtained by stereoscopically transposing the two-dimensional plane image using the depth map.

[0029] In a possible implementation manner, the three-dimensional stereoscopic image includes a second left-eye image and a second right-eye image;

[0030] Generating the three-dimensional stereoscopic image according to the two-dimensional planar image and the depth map includes:

[0031] Performing stereo transposition processing on the two-dimensional plane image using the depth map to obtain a first left-eye image and a first right-eye image;

[0032] Using the background segmentation map, performing background removal processing on the first left-eye image to obtain the second left-eye image;

[0033] The background segmentation map is used to perform background removal processing on the first right-eye image to obtain the second right-eye image.

[0034] In one possible implementation, the process of determining the posture motion characterization feature includes: splicing the original image, the original three-dimensional posture characterization data, and the reference three-dimensional posture characterization data to obtain a splicing result; and performing motion estimation processing on the splicing result to obtain the posture motion characterization feature.

[0035] In one possible implementation, the original three-dimensional posture representation data is obtained by mapping the three-dimensional facial mesh corresponding to the original image to a two-dimensional image space; the three-dimensional facial mesh corresponding to the original image is constructed based on the original image; the reference three-dimensional posture representation data is obtained by mapping the three-dimensional facial mesh corresponding to the reference information to a two-dimensional image space; and the three-dimensional facial mesh corresponding to the reference information is constructed based on the reference information.

[0036] In one possible implementation, the original three-dimensional posture representation data is obtained by performing feature extraction processing on the three-dimensional facial mesh corresponding to the original image; the three-dimensional facial mesh corresponding to the original image is constructed based on the original image; the reference three-dimensional posture representation data is obtained by performing feature extraction processing on the three-dimensional facial mesh corresponding to the reference information; and the three-dimensional facial mesh corresponding to the reference information is constructed based on the reference information.

[0037] In one possible implementation, the generated image includes a two-dimensional plane image; the reference information is supervisory information corresponding to the two-dimensional plane image; and the three-dimensional facial mesh corresponding to the reference information is obtained by performing three-dimensional facial reconstruction processing on the supervisory information corresponding to the two-dimensional plane image.

[0038] In one possible implementation, determining the three-dimensional facial mesh corresponding to the reference information includes: performing three-dimensional facial reconstruction on the supervisory information corresponding to the two-dimensional plane image to obtain a first facial mesh; and performing expression adjustment on the first facial mesh according to preset expressionless parameters to obtain the three-dimensional facial mesh corresponding to the reference information.

[0039] In a possible implementation, the reference information includes a head posture; and the three-dimensional face mesh corresponding to the reference information is obtained by performing posture adjustment processing on the three-dimensional face mesh corresponding to the original image using the head posture.

[0040] In one possible implementation, determining the three-dimensional facial mesh corresponding to the original image includes: performing three-dimensional facial reconstruction on the original image to obtain a second facial mesh; and performing expression adjustment on the second facial mesh according to preset expressionless parameters to obtain the three-dimensional facial mesh corresponding to the original image.

[0041] In one possible implementation, the generated image includes a two-dimensional plane image; the reference information is the supervisory information corresponding to the two-dimensional plane image; and the reference facial texture map refers to a facial texture map obtained by performing three-dimensional facial reconstruction processing on the supervisory information corresponding to the two-dimensional plane image.

[0042] In one possible implementation, the reference information includes a facial expression coefficient; the reference facial texture map is obtained by performing expression adjustment processing on the facial texture map corresponding to the original image using the facial expression coefficient; and the facial texture map corresponding to the original image is constructed based on the original image.

[0043] In one possible implementation, the method is applied to a virtual reality (VR) device.

[0044] In one possible implementation, the reference information includes facial expression coefficients and head posture; the generated image includes a three-dimensional stereo image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image.

[0045] In one possible implementation, the generated image includes a three-dimensional stereoscopic image; the VR device and a target device are in video communication, and the target device is a three-dimensional image display device;

[0046] After obtaining the generated image, the method further includes:

[0047] The three-dimensional stereoscopic image is sent to the target device, and the target device is used to display the three-dimensional stereoscopic image.

[0048] In one possible implementation, the generated image includes a two-dimensional plane image and a three-dimensional stereoscopic image, and the three-dimensional stereoscopic image includes at least two channel images; the VR device is in video communication with a target device, and the target device is a two-dimensional image display device;

[0049] After obtaining the generated image, the method further includes:

[0050] The two-dimensional plane image or any one of the channel images in the three-dimensional stereoscopic image is sent to the target device, and the target device is used to display the two-dimensional plane image or any one of the channel images.

[0051] In one possible implementation, the reference information corresponding to the original image includes a facial expression coefficient and a head posture acquired by the VR device for the target object;

[0052] The obtaining of the original image, the original three-dimensional posture representation data, the reference three-dimensional posture representation data, and the reference facial texture map includes:

[0053] Receiving the original image, the three-dimensional face mesh corresponding to the original image, and the face texture map corresponding to the original image sent by the server;

[0054] The original three-dimensional posture representation data, the reference three-dimensional posture representation data, and the reference facial texture map are obtained based on the three-dimensional facial mesh corresponding to the original image, the facial texture map corresponding to the original image, the facial expression coefficient, and the head posture.

[0055] In a possible implementation manner, the server is configured to generate the original image based on the facial representation image and style description information specified by the target object.

[0056] The present disclosure provides an image generating device, comprising:

[0057] an acquisition unit, configured to acquire an original image, original 3D posture representation data, reference 3D posture representation data, and a reference facial texture map, wherein the original 3D posture representation data is determined based on the original image; and the reference 3D posture representation data and the reference facial texture map are determined based on reference information corresponding to the original image;

[0058] a determining unit, configured to determine a gesture motion representation feature based on the original image, the original three-dimensional gesture representation data, and the reference three-dimensional gesture representation data;

[0059] A generating unit is configured to obtain a generated image based on the coding features of the original image, the coding features of the reference facial texture map, and the gesture motion representation features.

[0060] The present disclosure provides an electronic device, the device comprising: a processor and a memory;

[0061] The memory is used to store instructions or computer programs;

[0062] The processor is configured to execute the instructions or computer program in the memory so that the electronic device executes the image generation method provided by the present disclosure.

[0063] The present disclosure provides a computer-readable medium, characterized in that instructions or computer programs are stored in the computer-readable medium, and when the instructions or computer programs are executed on a device, the device executes the image generation method provided by the present disclosure.

[0064] The present disclosure provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, wherein the computer program contains program code for executing the image generation method provided by the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0066] FIG1 is a flow chart of an image generation method provided by an embodiment of the present disclosure;

[0067] FIG2 is a schematic diagram of an image generation scenario provided by an embodiment of the present disclosure;

[0068] FIG3 is a schematic diagram of a process for constructing an original image according to an embodiment of the present disclosure;

[0069] FIG4 is a schematic diagram of an image generation process in a model training process provided by an embodiment of the present disclosure;

[0070] FIG5 is a schematic diagram of an image generation process in an image generation scenario provided by an embodiment of the present disclosure;

[0071] FIG6 is a schematic structural diagram of an image generating device provided by an embodiment of the present disclosure;

[0072] FIG7 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0073] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0074] To better understand the technical solutions provided by the present disclosure, the image generation method provided by the present disclosure is described below with reference to some accompanying figures. As shown in Figure 1, the image generation method provided by an embodiment of the present disclosure includes the following steps S1-S3. Figure 1 is a flowchart of an image generation method provided by an embodiment of the present disclosure.

[0075] S1: Acquire an original image, original three-dimensional posture representation data, reference three-dimensional posture representation data, and a reference facial texture map, wherein the original three-dimensional posture representation data is determined based on the original image; the reference three-dimensional posture representation data and the reference facial texture map are determined based on reference information corresponding to the original image.

[0076] The original image refers to a two-dimensional image used in the image generation process to provide partial facial constraints. This disclosure does not limit the constraints provided by this original image; for example, this original image can be used to provide constraints on static facial features. Static facial features refer to facial features that barely change or only change slightly over a short period of time. This disclosure does not limit the implementation of these static facial features; for example, they can include facial contours, hairstyle, etc.

[0077] In addition, the present disclosure does not limit the implementation of the above original image. For example, the original image can be implemented using the original image shown in Figure 2, the original image shown in Figure 3, the original image 1 shown in Figure 4, or the original image 2 shown in Figure 5.

[0078] In addition, the present disclosure does not limit the above-mentioned process of obtaining the original image. For ease of understanding, the following description is given in combination with two scenarios.

[0079] In scenario 1, when the image generation method provided by this disclosure is applied to a model training scenario, such as the one shown in Figure 4, the original image acquisition process described above can be: a frame of video image randomly extracted from a sample video is used as the original image. The sample video refers to the video required for model training; this disclosure does not limit the method for acquiring the sample video.

[0080] Scenario 2: When the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in FIG2 or the binocular image generation scenario shown in FIG5 , if the image generation method is applied to a virtual reality (VR) device, the process for the VR device to obtain the above-mentioned original image may be: first, the server generates the original image based on the facial representation image and style description information specified by the target object; then, the server sends the original image to the VR device so that the VR device can subsequently perform image generation processing based on the original image, such as the image generation processing shown in FIG5 . The target object refers to the user of the VR device, so that the VR device can be used to obtain the target object's facial state in real time. The server and the VR device can communicate data so that the server can provide the VR device with certain information, such as the original image and related information shown in FIG2 . The facial representation image refers to an image specified by the target object and used to describe the facial state; and the present disclosure does not limit the implementation of the facial representation image. For example, it can refer to an image provided by the target object through some means, such as an image captured by an image capture device. The style description information refers to information specified by the target object that describes the style characteristics; this disclosure does not limit the style description information; for example, the style description information may include style description text. The style description text is used to describe the image style required for the generation of the original image. It should be noted that the use of the facial representation image and style description information is performed with the authorization of the target object.

[0081] In addition, the present disclosure does not limit the implementation method of the step of "generating the original image based on the facial representation image and style description information specified by the target object" in the previous paragraph. For example, this step can be implemented using the process shown in Figure 3. Based on this, it can be seen that the present disclosure provides a possible implementation method for the above-mentioned server to generate the original image. Under this implementation method, when a large number of pre-constructed style images are stored in the server, such as the base maps of various styles shown in Figure 3, and the styles of different style images are different, the generation process of the original image may include: after the server receives the style description text provided by the target object, the server searches for the style image that best matches the style description text from these stored style images, so that the best-matching style image can represent an image that meets the style requirements of the target object; then, the server redraws the best-matching style image to obtain multiple candidate images, such as the multiple candidate base maps shown in Figure 3, so that the style of any two candidate images is consistent with the style of the best-matching style image, but there are some differences between the any two candidate images, such as different hair lengths, different hair tip colors, etc., which is conducive to improving image diversity; secondly, the server determines the final selection of the target object based on the selection operation triggered by the target object for these candidate images. The server then performs face-swapping processing on the finally selected candidate image using the facial representation image specified by the target object to obtain a face-swapping image, so that the facial description information carried by the face-swapping image is consistent with the facial description information carried by the facial representation image, and the other information carried by the face-swapping image other than the facial description information is consistent with the other information carried by the finally selected candidate image other than the facial description information; finally, the server performs a secondary redrawing processing, such as stylization processing, on the face-swapping image to obtain the original image, so that the stylization degree of the original image is higher than the stylization degree of the face-swapping image, and the similarity between the facial description information carried by the original image and the facial description information carried by the facial representation image is lower than the similarity between the facial description information carried by the face-swapping image and the facial description information carried by the facial representation image, so as to achieve a balance between the facial similarity and the stylization degree, so that the original image can take into account both the facial state requirements and the stylization requirements of the target object, thereby improving the image generation effect.

[0082] In addition, the present disclosure does not limit the implementation method of the step of "redrawing the best-matching style image to obtain multiple candidate images" in the previous paragraph. For example, it can be specifically: performing image generation processing based on random noise and the best-matching style image to obtain a candidate image, so that there are a small number of differences between the candidate image and the best-matching style image. For another example, it can be specifically: performing image generation processing based on the adjustment description text and the best-matching style image to obtain a candidate image, so that the candidate image meets the facial state constraint described by the adjustment description text, so that adjustment processing can be performed on a certain part of the best-matching style image. The adjustment description text is used to describe which adjustment processing is performed on which part of the best-matching style image; and the present disclosure does not limit the implementation method of the adjustment description text. For example, the adjustment description text can be implemented using the character string "hair blue".

[0083] Furthermore, the present disclosure does not limit the implementation of the step of "performing a secondary redrawing process on the face-swapped image to obtain the original image" described above. For example, the step may specifically include: performing image generation processing on the face-swapped image using a pre-built stylized image generation model to obtain the original image, thereby increasing the stylization of the original image. The stylized image generation model refers to a machine learning model pre-built for the style described by the style description text, such that the stylized image generation model can be used to generate images that conform to the style described by the style description text.

[0084] Based on the above four paragraphs and the content shown in Figure 3, it can be seen that if the image generation method provided by the present disclosure is applied to a VR device, such as VR device 1 or VR device 2 shown in Figure 2, and the VR device can communicate data with a server, then the original image required for the VR device to perform image generation processing can be provided by the server; and the server can be used to generate the original image based on the facial representation image and style description information specified by the target object. The facial representation image refers to an image provided by the target object to the server in a certain manner to describe the facial state, so that the facial representation image is used to provide facial state constraints for the generation process of the original image. The style description information refers to the information provided by the target object to the server in a certain manner for describing the style of the image, so that the style description information is used to provide style constraints for the generation process of the original image; and the present disclosure does not limit the implementation method of the style description information. For example, the style description information may include the above-mentioned style description text and the above-mentioned "selection operation triggered for these candidate images" so that the style description information can more accurately represent the style requirements of the target object, thereby making the original image ultimately generated more consistent with the style requirements, which is conducive to improving the image generation effect. It should be noted that the present disclosure does not limit the implementation method of the target object providing the facial representation image and style description information to the server. For example, the target object may provide the facial representation image and style description information to the server with the help of a terminal device, such as the VR device.

[0085] The original 3D posture representation data is used to represent the state of the facial posture described by the original image in three-dimensional space, so that the original 3D posture representation data can better represent the facial posture described by the original image. For example, when the original image is original image 1 shown in Figure 4, the original 3D posture representation data can be implemented using the displayed 3D key point map of original image 1 shown in Figure 4. For another example, when the original image is original image 2 shown in Figure 5, the original 3D posture representation data can be implemented using the displayed 3D key point map of original image 2 shown in Figure 5.

[0086] In addition, for the above-mentioned original three-dimensional posture representation data, the original three-dimensional posture representation data is determined based on the above-mentioned original image, so that the original three-dimensional posture representation data can represent the state of the facial posture described by the original image in the three-dimensional space; and the present disclosure does not limit the determination process of the original three-dimensional posture representation data, for example, it can specifically include the following steps 11-12.

[0087] Step 11: Construct a three-dimensional facial mesh corresponding to the original image, so that the three-dimensional facial mesh can represent the state of the facial posture described by the original image in three-dimensional space.

[0088] The 3D facial mesh corresponding to the original image refers to a 3D facial mesh constructed based on the original image, so that the 3D facial mesh can represent the state of the facial posture described by the original image in the 3D space.

[0089] In addition, the present disclosure does not limit the implementation method of the above step 11. For example, it can specifically be: performing three-dimensional facial reconstruction processing on the above original image to obtain a three-dimensional facial mesh corresponding to the original image. It should be noted that the present disclosure does not limit the implementation method of the three-dimensional facial reconstruction processing. For example, the three-dimensional facial reconstruction processing can be implemented by using any existing or future method that can achieve three-dimensional facial reconstruction processing, such as using a pre-built three-dimensional facial reconstruction model. Among them, the three-dimensional facial reconstruction model refers to a pre-built model with good three-dimensional facial reconstruction function; and the present disclosure does not limit the implementation method of the three-dimensional facial reconstruction model. For example, it can be implemented by using any existing or future model with three-dimensional facial reconstruction function, such as the three-dimensional facial reconstruction model 2 shown in Figure 5.

[0090] In addition, in order to prevent facial expressions from interfering with facial postures, the present disclosure also provides a possible implementation of the above step 11. In this implementation, the step 11 may specifically include the following steps 111-112.

[0091] Step 111: Perform three-dimensional facial reconstruction on the original image to obtain a second facial mesh, so that the second facial mesh can represent the state of the facial posture and facial expression described by the original image in three-dimensional space.

[0092] The second facial mesh refers to a three-dimensional facial mesh obtained by performing three-dimensional facial reconstruction processing on the original image, so that the second facial mesh can represent the state of the facial posture and facial expression described by the original image in three-dimensional space.

[0093] Step 112: Expression adjustment processing is performed on the second facial mesh according to preset expressionless parameters to obtain a three-dimensional facial mesh corresponding to the original image, so that the three-dimensional facial mesh can represent the state of the facial posture described by the original image in three-dimensional space, but the three-dimensional facial mesh cannot represent the state of the facial expression described by the original image in three-dimensional space.

[0094] The preset neutral expression parameters refer to the parameters used to adjust a 3D facial mesh to a neutral expression state. The present disclosure does not limit the preset neutral expression parameters; for example, the preset neutral expression parameters can be determined based on the actual scenario. The neutral expression state refers to a pre-set standard expression state, such as not smiling, not crying, or not speaking.

[0095] In addition, the present disclosure does not limit the implementation method of the expression adjustment process in the above step 112. For example, it can be implemented using any existing or future method that can adjust the expression of the three-dimensional facial mesh according to a certain expression.

[0096] Based on the relevant contents of steps 111 to 112 above, it can be seen that for the original image above, the original image is first subjected to three-dimensional facial reconstruction processing to obtain a three-dimensional facial mesh with expression and posture, so that the three-dimensional facial mesh can represent the facial posture described by the original image and the state of the facial expression in three-dimensional space; then, according to the preset expressionless parameters, the three-dimensional facial mesh is subjected to expression adjustment processing to obtain a three-dimensional facial mesh with expressionless posture as the three-dimensional facial mesh corresponding to the original image. This can effectively overcome the interference caused by facial expression on facial posture, so that the three-dimensional facial mesh can better represent the state of the facial posture described by the original image in three-dimensional space, which is conducive to improving the image generation effect.

[0097] Furthermore, the present disclosure does not limit the implementation of step 11 above. For example, in some application scenarios, such as the holographic video generation scenario shown in Figure 2, to better reduce the computing pressure of the VR device, step 11 can be performed by the server. Based on this, it can be seen that step 11 can specifically be: the server constructs a 3D facial mesh corresponding to the original image, so that the server can subsequently send the 3D facial mesh corresponding to the original image to the VR device, so that the VR device can complete the image generation task based on the 3D facial mesh corresponding to the original image.

[0098] Step 12: Determine the original 3D posture representation data based on the 3D facial mesh corresponding to the original image above.

[0099] It should be noted that the present disclosure does not limit the implementation of the above step 12. For example, the step 12 may specifically be: directly determining the three-dimensional facial mesh corresponding to the original image as the original three-dimensional posture representation data.

[0100] In addition, in some application scenarios, in order to better avoid defects caused by the excessive amount of data carried by the three-dimensional facial mesh, the present disclosure also provides a possible implementation method of the above step 12. In this implementation method, the step 12 can specifically be: performing feature extraction processing on the three-dimensional facial mesh corresponding to the above original image to obtain original three-dimensional posture representation data, so that the original three-dimensional posture representation data can represent the features extracted from the three-dimensional facial mesh, and the amount of data carried by the original three-dimensional posture representation data is less than the amount of data carried by the three-dimensional facial mesh, so that the original three-dimensional posture representation data can represent the three-dimensional information described by the three-dimensional facial mesh with a smaller amount of data, thereby making the image generation process based on the original three-dimensional posture representation data have less resource consumption.

[0101] In addition, in some application scenarios, in order to better improve the representation effect of facial gestures while avoiding defects caused by the excessive amount of data carried by the three-dimensional facial mesh, the facial gestures can be represented in a display manner. Based on this, the present disclosure also provides a possible implementation of the above step 12. In this implementation, the step 12 can specifically be: mapping the three-dimensional facial mesh corresponding to the above original image to a two-dimensional image space to obtain original three-dimensional gesture representation data, such as the explicit three-dimensional key point map of the original image 1 shown in Figure 4 or the explicit three-dimensional key point map of the original image 2 shown in Figure 5, so that the original three-dimensional gesture representation data is a two-dimensional image, so that the original three-dimensional gesture representation data can visually represent the state of the facial gesture described by the original image in three-dimensional space with the help of the two-dimensional image, and the amount of data carried by the original three-dimensional gesture representation data is less than the amount of data carried by the three-dimensional facial mesh, thereby enabling the original three-dimensional gesture representation data to visually represent the three-dimensional information described by the three-dimensional facial mesh with a smaller amount of data. In this way, the facial gesture representation effect can be improved while minimizing resource consumption, thereby improving the image generation effect. The two-dimensional image space is used to describe the space in which the two-dimensional image resides. Furthermore, this disclosure does not limit the implementation of this two-dimensional image space; for example, the two-dimensional image space may refer to a blank two-dimensional image. It should be noted that this disclosure does not limit the implementation of this mapping; for example, it may be implemented using any existing or future method capable of mapping a three-dimensional facial mesh into a two-dimensional image.

[0102] In addition, for the original three-dimensional posture representation data shown in the above two paragraphs, in order to better perform subsequent splicing processing, such as the splicing processing shown in Figure 4 or the splicing processing shown in Figure 5, the original three-dimensional posture representation data can have the following characteristics: the size of the original three-dimensional posture representation data in the target dimension is consistent with the size of the original image in the target dimension, so that the original three-dimensional posture representation data and the original image can be spliced ​​based on the target dimension. The target dimension refers to a pre-set dimension that needs to be aligned during the splicing process; and the target dimension can be set according to the actual application scenario, for example, the target dimension can be width or height.

[0103] Furthermore, the present disclosure does not limit the implementation of step 12 above. For example, in some application scenarios, such as the holographic video generation scenario shown in FIG2 , step 12 can be performed by a VR device. Based on this, it can be seen that step 12 can specifically be: the VR device determines the original 3D posture representation data based on the 3D facial mesh corresponding to the original image above.

[0104] Based on the relevant content of steps 11 to 12 above, it can be seen that for some application scenarios, after obtaining the original image, a three-dimensional facial mesh corresponding to the original image can be first constructed so that the three-dimensional facial mesh can represent the state of the facial posture described by the original image in three-dimensional space; then, the original three-dimensional posture representation data is determined based on the three-dimensional facial mesh so that the original three-dimensional posture representation data can also represent the state of the facial posture described by the original image in three-dimensional space.

[0105] For the original image described above, the reference information corresponding to the original image refers to information required for providing facial constraints other than those described by the original image when performing image generation processing based on the original image. Furthermore, this disclosure does not limit the constraints provided by this reference information. For example, this reference information can be used to provide constraints on aspects such as dynamic facial features. These dynamic facial features refer to facial features that can change, such as significantly, within a short period of time. Furthermore, this disclosure does not limit the implementation of these dynamic facial features. For example, these dynamic facial features can include facial expressions, facial postures, etc.

[0106] In addition, the present disclosure does not limit the implementation method of the reference information corresponding to the original image. For ease of understanding, the following description is combined with two scenarios.

[0107] Scenario 1: When the image generation method provided by the present disclosure is applied to a model training scenario, such as the model training scenario shown in FIG4 , the reference information corresponding to the original image mentioned above can be implemented using a reference image corresponding to the original image, such as the reference image shown in FIG4 . The reference image is used to provide facial expression constraints and facial posture constraints required when performing image generation processing based on the original image, so that the model performance evaluation process can be performed subsequently using the similarity between the reference image and the two-dimensional image generated based on the original image. The “two-dimensional image generated based on the original image” refers to an image generated based on the original image and the reference image, such as the predicted image 1 shown in FIG4 . In addition, the present disclosure does not limit the method for obtaining the reference image corresponding to the original image. For example, when the original image is a frame of video image randomly extracted from a sample video, the reference image can be any other frame of video image in the sample video except the original image.

[0108] Scenario 2: When the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in FIG2 or the binocular image generation scenario shown in FIG5 , if the image generation method is applied to a VR device, the reference information corresponding to the above-mentioned original image may include the facial expression coefficient and head posture obtained by the VR device for the target object. The head posture is used to describe the facial posture of the target object. Moreover, the present disclosure does not limit the method for obtaining the head posture. For example, the head posture can be collected in real time by a posture sensor in the VR device, such as an inertial sensor similar to a DOF sensor, for the target object. The facial expression coefficient is used to represent the facial expression of the target object. Moreover, the present disclosure does not limit the process for obtaining the facial expression coefficient. For example, the facial expression coefficient can be a 52-dimensional expression coefficient determined in real time for the target object by a 52-dimensional expression coefficient acquisition module in the VR device. The 52-dimensional expression coefficient acquisition module refers to a module already deployed in the VR device that has the function of acquiring 52-dimensional expression coefficients. Furthermore, this disclosure does not limit the operating principle of the 52-dimensional expression coefficient acquisition module. For example, it may specifically be: after the 52-dimensional expression coefficient acquisition module acquires the facial image captured for the target object, the 52-dimensional expression coefficient acquisition module performs 52-dimensional expression coefficient extraction processing on the facial image to obtain the 52-dimensional expression coefficients corresponding to the facial image. It should be noted that the use of the facial image is performed under the premise of having obtained the authorization of the target object.

[0109] In addition, to further improve the accuracy of facial expressions, the present disclosure also provides a process for acquiring the facial expression coefficients mentioned in the previous paragraph. Specifically, after the 52-dimensional expression coefficient acquisition module in the VR device acquires the 52-dimensional expression coefficients from the facial image of the target subject, and the lip shape acquisition module in the VR device acquires the 52-dimensional expression coefficients from the voice data of the target subject, these two 52-dimensional expression coefficients are subjected to certain processing, such as weighted summation, to obtain the facial expression coefficients. It should be noted that the use of the voice data is performed on the premise that the target subject has obtained authorization.

[0110] For the reference information corresponding to the original image above, the reference information can be used to provide facial posture constraints and facial expression constraints, so that after obtaining the reference information corresponding to the original image, reference three-dimensional posture representation data and a reference facial texture map can be determined based on the reference information, so that the reference three-dimensional posture representation data is used to represent the facial posture constraints, and the reference facial texture map is used to represent the facial expression constraints, so that subsequent image generation processing can be performed based on the reference three-dimensional posture representation data and the reference facial texture map to generate an image that satisfies the facial posture constraints and facial expression constraints.

[0111] The reference three-dimensional posture representation data is used to represent the state of the facial posture described by the reference information corresponding to the original image in three-dimensional space, so that the reference three-dimensional posture representation data can better represent the facial posture described by the reference information. For example, when the reference information is the reference image shown in Figure 4, the reference three-dimensional posture representation data can be implemented using the displayed three-dimensional key point map of the reference image shown in Figure 4. For another example, when the reference information is the reference information provided by the VR device shown in Figure 5, the reference three-dimensional posture representation data can be implemented using the displayed three-dimensional key point map of the reference information shown in Figure 5.

[0112] In addition, the present disclosure does not limit the determination process of the above reference three-dimensional posture representation data. For example, it may include the following steps 21 and 22.

[0113] Step 21: Construct a three-dimensional facial mesh corresponding to the reference information, so that the three-dimensional facial mesh can represent the state of the facial posture described by the reference information in the three-dimensional space.

[0114] The three-dimensional facial mesh corresponding to the above reference information refers to a three-dimensional facial mesh constructed based on the reference information, so that the three-dimensional facial mesh can represent the state of the facial posture described by the reference information in three-dimensional space.

[0115] In addition, the present disclosure does not limit the implementation of the above step 21. For ease of understanding, the following description is combined with two scenarios.

[0116] Scenario 1: When the image generation method provided by the present disclosure is applied to a model training scenario, such as the model training scenario shown in FIG4 , if the reference information corresponding to the original image above is an image, such as the reference image shown in FIG4 , or the supervisory information corresponding to the two-dimensional plane image below, then the implementation of step 21 above is similar to the implementation of step 11 above. Based on this, it can be seen that in one possible implementation, step 21 can specifically be: performing three-dimensional facial reconstruction processing on the reference information to obtain a three-dimensional facial mesh corresponding to the reference information. In another possible implementation, to better avoid the interference of facial expressions on facial posture, step 21 can specifically be: performing three-dimensional facial reconstruction processing on the reference information to obtain a first facial mesh, so that the first facial mesh can represent the facial posture and facial expression described by the reference information in three-dimensional space; then, according to preset expressionless parameters, performing expression adjustment processing on the first facial mesh to obtain a three-dimensional facial mesh corresponding to the reference information. The first facial mesh refers to a three-dimensional facial mesh obtained by performing three-dimensional facial reconstruction processing on the reference information, so that the first facial mesh can represent the state of the facial posture and facial expression described by the reference information in three-dimensional space.

[0117] Scenario 2: When the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in FIG2 or the binocular image generation scenario shown in FIG5 , if the image generation method is applied to a VR device, and the reference information corresponding to the original image includes the facial expression coefficient and head posture obtained by the VR device for the target object, then step 21 above may specifically be: using the head posture to perform posture adjustment processing on the three-dimensional facial mesh corresponding to the original image above, to obtain a three-dimensional facial mesh corresponding to the reference information, so that the three-dimensional facial mesh can describe the state of the head posture in three-dimensional space. It should be noted that the present disclosure does not limit the implementation method of the posture adjustment processing. For example, it can be implemented using any existing or future method that can perform posture adjustment processing on a three-dimensional facial mesh based on a certain posture.

[0118] Furthermore, the present disclosure does not limit the implementation of step 21 above. For example, in some application scenarios, such as the holographic video generation scenario shown in FIG2 , step 21 is performed by a VR device. Based on this, it can be seen that step 21 can specifically be as follows: after the VR device obtains reference information for the target object, the VR device constructs a three-dimensional facial mesh corresponding to the reference information, so that the three-dimensional facial mesh can represent the state of the facial posture described by the reference information in three-dimensional space.

[0119] Step 22: Determine reference 3D posture representation data based on the 3D facial mesh corresponding to the above reference information.

[0120] It should be noted that the implementation of step 22 is similar to the implementation of step 12. For ease of understanding, the following description is given with reference to examples.

[0121] As an example, in one possible implementation, in order to better avoid defects caused by the excessive amount of data carried by the three-dimensional facial mesh, the above step 22 may specifically be: performing feature extraction processing on the three-dimensional facial mesh corresponding to the above reference information to obtain the reference three-dimensional posture representation data, so that the reference three-dimensional posture representation data can represent the features extracted from the three-dimensional facial mesh, and the amount of data carried by the reference three-dimensional posture representation data is less than the amount of data carried by the three-dimensional facial mesh, so that the reference three-dimensional posture representation data can represent the three-dimensional information described by the three-dimensional facial mesh with a smaller amount of data, thereby making the image generation process implemented based on the reference three-dimensional posture representation data have less resource consumption.

[0122] As an example, in one possible implementation, in order to better improve the facial gesture representation effect while avoiding defects caused by the excessive amount of data carried by the three-dimensional facial mesh, the above step 22 can specifically be: mapping the three-dimensional facial mesh corresponding to the above reference information to a two-dimensional image space to obtain reference three-dimensional gesture representation data, such as the explicit three-dimensional key point map of the reference image shown in Figure 4 or the explicit three-dimensional key point map of the reference information shown in Figure 5, so that the reference three-dimensional gesture representation data is a two-dimensional image, so that the reference three-dimensional gesture representation data can visually represent the state of the facial gesture described by the reference information in the three-dimensional space with the help of the two-dimensional image, and the amount of data carried by the reference three-dimensional gesture representation data is less than the amount of data carried by the three-dimensional facial mesh, so that the reference three-dimensional gesture representation data can visually represent the three-dimensional information described by the three-dimensional facial mesh with a smaller amount of data. In this way, the facial gesture representation effect can be improved while minimizing resource consumption, thereby improving the image generation effect.

[0123] It should be noted that, for the reference three-dimensional posture representation data shown in the above two paragraphs, in order to better carry out subsequent stitching processing, such as the stitching processing shown in Figure 4 or the stitching processing shown in Figure 5, the reference three-dimensional posture representation data has the following characteristics: the size of the reference three-dimensional posture representation data in the target dimension is consistent with the size of the original image in the target dimension, so that the reference three-dimensional posture representation data and the original image can be stitched together based on the target dimension.

[0124] Furthermore, the present disclosure does not limit the implementation of step 22 above. For example, in some application scenarios, such as the holographic video generation scenario shown in FIG2 , step 22 may be performed by a VR device. Based on this, it can be seen that step 22 may specifically be: the VR device determines reference 3D posture representation data based on the 3D facial mesh corresponding to the reference information above.

[0125] Based on the relevant content of steps 21 to 22 above, it can be seen that for some application scenarios, after obtaining the reference information corresponding to the original image above, a three-dimensional facial mesh corresponding to the reference information can be first constructed so that the three-dimensional facial mesh can represent the state of the facial posture described by the reference information in the three-dimensional space; then, reference three-dimensional posture representation data can be determined based on the three-dimensional facial mesh so that the reference three-dimensional posture representation data can also represent the state of the facial posture described by the reference information in the three-dimensional space.

[0126] The reference facial texture map is used to describe the facial expression described by the reference information corresponding to the original image. For example, when the reference information is the reference image shown in FIG4 , the reference facial texture map can be implemented using the facial texture map of the reference image shown in FIG4 . For another example, when the reference information is the reference information shown in FIG5 , the reference facial texture map can be implemented using the facial texture map of the reference information shown in FIG5 .

[0127] In addition, the present disclosure is not limited to the implementation method of referring to the facial texture map. For ease of understanding, the following description is combined with two scenarios.

[0128] Scenario 1. When the image generation method provided by the present disclosure is applied to a model training scenario, such as the model training scenario shown in Figure 4, if the reference information corresponding to the original image above is an image, such as the reference image shown in Figure 4 or the supervision information corresponding to the two-dimensional plane image below, then the reference facial texture map above may refer to a facial texture map (Texture) obtained by performing three-dimensional facial reconstruction processing on the reference information.

[0129] Scenario 2: When the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in FIG2 or the binocular image generation scenario shown in FIG5, if the image generation method is applied to a VR device, and the reference information corresponding to the original image includes the facial expression coefficient and head posture obtained by the VR device for the target object, then the reference facial texture map can be obtained by using the facial expression coefficient to perform expression adjustment processing on the facial texture map corresponding to the original image, so that the facial expression described by the reference facial texture map is consistent with the facial expression described by the facial expression coefficient, and the facial states other than the facial expression described by the reference facial texture map are consistent with the facial states other than the facial expression described by the facial texture map corresponding to the original image. The facial texture map corresponding to the original image is constructed based on the original image; and the present disclosure does not limit the construction process of the facial texture map corresponding to the original image. For example, it can be specifically: performing three-dimensional facial reconstruction processing on the original image to obtain the facial texture map corresponding to the original image.

[0130] In addition, the present disclosure is not limited to the implementation of S1 above. For ease of understanding, the following description is combined with two scenarios.

[0131] Scenario 1. When the image generation method provided by the present disclosure is applied to a model training scenario, such as the model training scenario shown in Figure 4, the above S1 can be specifically as follows: first, two frames of video images are randomly selected from the sample video as the original image and the reference information corresponding to the original image; then, the original three-dimensional posture representation data is determined based on the original image, and the reference three-dimensional posture representation data and the reference facial texture map are determined based on the reference information.

[0132] Scenario 2: When the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in Figure 2 or the binocular image generation scenario shown in Figure 5, if the image generation method is applied to a VR device, and the reference information corresponding to the original image above includes the facial expression coefficient and head posture obtained by the VR device for the target object, then the above S1 can specifically include the following steps 31-32.

[0133] Step 31: The VR device receives the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image sent by the server.

[0134] In the present disclosure, for the server, after the server generates an original image based on the facial representation image and style description information specified by the target object, the server can construct a three-dimensional facial mesh corresponding to the original image and a facial texture map corresponding to the original image based on the original image, so that the three-dimensional facial mesh can represent the facial posture described by the original image, and the facial texture map can represent the facial expression described by the original image. The server then sends the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image to the VR device, so that the VR device can complete the image generation task based on this information, such as the image generation task shown in Figure 5.

[0135] Step 32: The VR device obtains original 3D posture representation data, reference 3D posture representation data, and a reference facial texture map based on the 3D facial mesh corresponding to the original image, the facial texture map corresponding to the original image, and the facial expression coefficients and head posture obtained by the VR device for the target object.

[0136] In the present disclosure, for a VR device, after the VR device receives the original image sent by the server, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image, the VR device determines the original three-dimensional posture representation data based on the three-dimensional facial mesh corresponding to the original image, and after the VR device obtains the facial expression coefficient and the head posture for the target object, the VR device uses the facial expression coefficient to perform expression adjustment processing on the facial texture map to obtain a reference facial texture map, and the VR device uses the head posture to perform posture adjustment processing on the three-dimensional facial mesh, and the VR device determines the reference three-dimensional posture representation data based on the three-dimensional facial mesh after posture adjustment, so that the VR device can subsequently complete the image generation task based on this information, such as the image generation task shown in Figure 5.

[0137] Based on the relevant content of steps 31 to 32 above, it can be seen that in some application scenarios, for the above VR device, when the VR device performs an image generation task, part of the information required for the image generation task can come from the server, and the other part of the information is determined by the VR device based on the information provided by the server. In this way, the server shares part of the computing pressure of the VR device, which is beneficial to reduce the computing pressure of the VR device, and further beneficial to ensure the rapid execution of the image generation task.

[0138] In addition, for some video generation scenarios, such as the holographic video generation scenario shown in Figure 2, since the original image required to generate each frame of video image is the same image, in order to better improve the video generation efficiency, after the VR device receives the original image sent by the server, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image, the VR device can directly store this information so that when the generation task of each frame of video image is triggered, this information can be directly read from the storage space for use. This can effectively avoid the consumption of computing resources caused by repeatedly obtaining the original image and its related information, thereby helping to improve the video generation efficiency.

[0139] Based on the content of the above paragraph, it can be seen that in one possible implementation, for the above VR device, the working principle of the VR device can be: after the VR device receives the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image sent by the server, the VR device stores the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image, so that when the VR device performs the task of generating the n-th frame of video image, the VR device first reads the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image from the storage space; then, the VR device obtains the original three-dimensional posture representation data, the reference three-dimensional posture representation data, and the reference facial texture map based on the three-dimensional facial mesh corresponding to the original image, the facial texture map corresponding to the original image, and the facial expression coefficient and head posture obtained by the VR device for the target object, so that the VR device can subsequently complete the image generation task based on this information, such as the image generation task shown in Figure 5.

[0140] Based on the relevant content of S1 above, it can be known that for some scenarios, such as the model training scenario shown in Figure 4 or the three-dimensional stereo image generation scenario shown in Figure 5, the original image, original three-dimensional posture representation data, reference three-dimensional posture representation data and reference facial texture map are obtained so that image generation processing can be performed based on this information later.

[0141] S2: Determine the posture motion representation feature based on the original image, the original three-dimensional posture representation data, and the reference three-dimensional posture representation data.

[0142] Among them, the posture motion representation feature is used to represent the posture motion information carried by the above reference three-dimensional posture representation data, so that the posture motion representation feature can represent the posture motion information required when changing from the posture represented by the above original three-dimensional posture representation data to the posture represented by the reference three-dimensional posture representation data, so that the posture motion representation feature can, to a certain extent, represent the difference between the facial posture represented by the original three-dimensional posture representation data and the facial posture represented by the reference three-dimensional posture representation data.

[0143] In addition, the present disclosure does not limit the implementation of the above S2. For example, it may specifically include the following steps 41 and 42.

[0144] Step 41: splicing the original image, the original 3D posture representation data, and the reference 3D posture representation data to obtain a splicing result.

[0145] The stitching result refers to the result obtained by stitching the above original image, the above original 3D posture representation data, and the above reference 3D posture representation data.

[0146] In addition, the present disclosure does not limit the implementation method of the splicing in the above step 41. For example, it can be implemented by using any existing or future method that can splice multiple data into one data, such as the splicing process shown in Figure 4 or the splicing process shown in Figure 5.

[0147] Step 42: Perform motion estimation processing on the above splicing results to obtain posture motion representation features.

[0148] It should be noted that the present disclosure does not limit the implementation method of the motion estimation processing in the above step 42. For example, it can be implemented using any existing or future method that can perform motion estimation processing, such as the motion estimation network shown in Figure 4 or the motion estimation network shown in Figure 5.

[0149] Based on the relevant content of steps 41 to 42 above, it can be seen that for some application scenarios, after obtaining the above original image, the above original three-dimensional posture representation data and the above reference three-dimensional posture representation data, these three data can be spliced ​​first; then the splicing result can be motion estimated to obtain posture motion representation features, so that the posture motion representation features are used to represent the posture motion information carried by the reference three-dimensional posture representation data.

[0150] In addition, the present disclosure does not limit the implementation of S2 above. For example, when the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in Figure 2 or the binocular image generation scenario shown in Figure 5, if the image generation method is applied to a VR device, then S2 can be executed by the VR device. Based on this, it can be seen that in one possible implementation, S2 can specifically be: the VR device determines the posture motion representation feature based on the original image, the original three-dimensional posture representation data, and the reference three-dimensional posture representation data.

[0151] Based on the relevant content of S2 above, it can be known that after obtaining the above original image, the above original three-dimensional posture representation data and the above reference three-dimensional posture representation data, the posture motion representation feature can be determined based on these three data, so that the posture motion representation feature is used to represent the posture motion information carried by the reference three-dimensional posture representation data, so that subsequent image generation processing can be performed based on the posture motion representation feature.

[0152] S3: Obtain a generated image based on the encoding features of the original image, the encoding features of the reference facial texture map, and the posture and motion representation features.

[0153] The generated image refers to an image generated based on the coding features of the original image, the coding features of the reference facial texture map, and the posture and motion representation features.

[0154] In addition, the present disclosure does not limit the above-mentioned implementation method of generating images. For ease of understanding, the following description is combined with two scenarios.

[0155] Scenario 1: When the image generation method provided by the present disclosure is applied to a model training scenario, such as the model training scenario shown in FIG4 , the generated image may include a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image. The two-dimensional plane image refers to a two-dimensional image generated based on the encoding features of the original image, the encoding features of the reference facial texture map, and the posture and motion representation features, so that the two-dimensional plane image can represent the state of the final generated face in the two-dimensional image space. Furthermore, the present disclosure does not limit the implementation of the two-dimensional plane image; for example, the two-dimensional plane image can be implemented using the predicted image 1 shown in FIG4 . The depth map is used to describe the three-dimensional information corresponding to the two-dimensional plane image. Furthermore, the present disclosure does not limit the implementation of the depth map; for example, the depth map can be implemented using the depth map of the predicted image 1 shown in FIG4 . The background segmentation map is used to describe the location of the background in the two-dimensional plane image. Furthermore, the present disclosure does not limit the implementation of the background segmentation map; for example, the background segmentation map can be implemented using the background segmentation map of the predicted image 1 shown in FIG4 . For another example, the background segmentation map can be implemented using a foreground mask map of the two-dimensional plane image.

[0156] Scenario 2: When the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in FIG2 or the binocular image generation scenario shown in FIG5, if the image generation method is applied to a VR device, the generated image may include a three-dimensional stereo image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image. The three-dimensional stereo image refers to a three-dimensional image generated based on the encoding features of the original image, the encoding features of the reference facial texture map, and the posture and motion representation features, such as a binocular image, so that the three-dimensional stereo image can represent the state of the final generated face in the three-dimensional space.

[0157] In addition, the present disclosure does not limit the implementation method of the above three-dimensional stereoscopic image. For example, the three-dimensional stereoscopic image may include at least two channel images, and different channel images are used to describe facial states under different perspectives in the two-dimensional image space, so that each channel image belongs to a two-dimensional image. It can be seen that under one possible implementation method, the three-dimensional stereoscopic image may include a left-eye channel image and a right-eye channel image. Among them, the left-eye channel image is used to describe the facial state under the perspective of the left eye in the two-dimensional image space. The right-eye channel image is used to describe the facial state under the perspective of the right eye in the two-dimensional image space.

[0158] In addition, for the above-mentioned three-dimensional stereoscopic image, the three-dimensional stereoscopic image is determined based on the above-mentioned two-dimensional plane image and the depth map corresponding to the two-dimensional plane image, so that the three-dimensional stereoscopic image carries the facial state description information in the two-dimensional plane image, and the three-dimensional stereoscopic image carries the three-dimensional information described by the depth map.

[0159] In addition, the present disclosure does not limit the above-mentioned process of determining a three-dimensional stereoscopic image. For ease of understanding, the process is described below with reference to examples.

[0160] As an example, in some application scenarios, it may be necessary to retain the background described in the above two-dimensional plane image when constructing the above three-dimensional stereo image. Based on this, it can be seen that when the three-dimensional stereo image includes a first left-eye image and a first right-eye image, the process of determining the three-dimensional stereo image is: using the depth map corresponding to the two-dimensional plane image to perform stereo transposition processing on the two-dimensional plane image to obtain the first left-eye image and the first right-eye image. Among them, the first left-eye image refers to the left-eye channel image carrying the background, such as the left-eye image shown in Figure 5. The first right-eye image refers to the right-eye channel image carrying the background, such as the right-eye image shown in Figure 5. It should be noted that the present disclosure does not limit the implementation method of the stereo transposition processing. For example, it can be implemented using any existing or future method that can stereo transpose a two-dimensional image using a depth map to obtain a binocular image.

[0161] As an example, in some application scenarios, in order to better avoid interference caused by the background, the present disclosure also provides a possible implementation method of the above-mentioned three-dimensional stereoscopic image determination process. In this implementation method, when the three-dimensional stereoscopic image includes a second left-eye image and a second right-eye image, the three-dimensional stereoscopic image determination process may specifically include the following steps 51 to 53.

[0162] Step 51: Perform stereo transposition processing on the two-dimensional plane image using the depth map corresponding to the two-dimensional plane image to obtain a first left-eye image and a first right-eye image.

[0163] Step 52: Using the background segmentation map corresponding to the above two-dimensional plane image, background removal processing is performed on the first left-eye image to obtain a second left-eye image, so that the foreground described by the second left-eye image is consistent with the foreground described by the first left-eye image, but the background described by the first left-eye image does not exist in the second left-eye image.

[0164] The second left-eye image refers to a left-eye channel image without background, so that the second left-eye image can describe the state of the foreground from the perspective of the left eye.

[0165] In addition, the present disclosure does not limit the implementation of the above step 52. For example, it can be implemented by using any existing or future method that can perform background removal based on the background segmentation map.

[0166] Step 53: Using the background segmentation map corresponding to the above two-dimensional plane image, perform background removal processing on the first right eye image to obtain a second right eye image, so that the foreground described by the second right eye image is consistent with the foreground described by the first right eye image, but the background described by the first right eye image does not exist in the second right eye image.

[0167] The second right-eye image refers to a right-eye channel image without background, so that the second right-eye image can describe the state of the foreground from the perspective of the right eye.

[0168] In addition, the present disclosure does not limit the implementation of the above step 53. For example, the implementation of the above step 53 is similar to the implementation of the above step 52.

[0169] In addition, the present disclosure does not limit the correlation between the execution time of step 53 and the execution time of step 52. For example, the execution time of step 53 and step 52 may be the same. For example, the execution time of step 53 may be earlier than the execution time of step 52. For example, the execution time of step 52 may be later than the execution time of step 53.

[0170] Based on the relevant content of steps 51 to 53 above, it can be seen that after generating a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image, a three-dimensional stereoscopic image without a background can be generated based on the two-dimensional plane image, the depth map, and the background segmentation map, so that the three-dimensional stereoscopic image can be adapted to various backgrounds, so that subsequent VR devices can configure any background for the three-dimensional stereoscopic image. This can effectively avoid the interference caused by the background automatically generated during the image generation process, thereby helping to increase the scope of use of the three-dimensional stereoscopic image.

[0171] In addition, the present disclosure does not limit the acquisition process of the above generated image. For example, it may specifically include the following steps 61 to 63.

[0172] Step 61: Process the encoding features of the original image according to the above posture and motion representation features to obtain processed features.

[0173] The coding feature of the original image refers to a feature obtained by encoding the original image, so that the coding feature can represent the image information carried by the original image.

[0174] In addition, the present disclosure does not limit the method for obtaining the coding features of the above original image. For example, it can be implemented by using any existing or future method for encoding images, such as a method using an encoder.

[0175] The processed feature refers to the feature obtained by processing the coding feature of the original image according to the above-mentioned posture and motion characterization feature, so that the processed feature carries the image information described by the coding feature and the posture and motion information carried by the posture and motion characterization feature, so that the processed feature can represent the result obtained by adjusting the posture of the coding feature according to the posture and motion characterization feature, and thus make the facial posture described by the processed feature as close as possible to the facial posture described by the above-mentioned reference three-dimensional posture characterization data. It should be noted that the present disclosure does not limit the implementation method of the step of "adjusting the posture of the coding feature according to the posture and motion characterization feature". For example, it can be specifically: the coding feature of the original image is deformed using the posture and motion characterization feature to obtain the processed feature.

[0176] Step 62: Concatenate the processed features and the encoded features of the reference facial texture map to obtain concatenated features.

[0177] The coding feature of the reference facial texture image refers to a feature obtained by encoding the reference facial texture image, so that the coding feature can represent the image information carried by the reference facial texture image.

[0178] In addition, the method for obtaining the coding features of the reference facial texture map is similar to the method for obtaining the coding features of the original image.

[0179] The spliced ​​feature refers to the result obtained by splicing the processed feature above and the encoded feature of the reference facial texture map above.

[0180] Step 63: Based on the above splicing features, a generated image is obtained.

[0181] It should be noted that the present disclosure does not limit the implementation of the above step 63. For ease of understanding, the following description is given in combination with two scenarios.

[0182] Scenario 1: When the image generation method provided by the present disclosure is applied to a model training scenario, such as the model training scenario shown in FIG4 , if the above-mentioned generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image, then the above-mentioned step 63 can specifically be: decoding the above-mentioned splicing features to obtain the two-dimensional plane image, the depth map, and the background segmentation map. It should be noted that the present disclosure does not limit the implementation method of the decoding process. For example, it can be implemented using any existing or future decoding method, such as the decoder shown in FIG4 .

[0183] Scenario 2. When the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in Figure 2 or the binocular image generation scenario shown in Figure 5, if the image generation method is applied to a VR device, and the above-generated image includes a three-dimensional stereo image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image, then the above step 63 can be specifically as follows: first decode the above splicing features to obtain the two-dimensional plane image, the depth map and the background segmentation map; and then generate the three-dimensional stereo image based on the two-dimensional plane image and the depth map.

[0184] Based on the relevant content of steps 61 to 63 above, it can be seen that for some application scenarios, after obtaining the coding features of the above original image, the coding features of the above reference facial texture map, and the above posture motion representation features, the coding features of the original image are first deformed using the posture motion representation features to obtain processed features, so that the facial posture described by the processed features is as close as possible to the facial posture described by the above reference three-dimensional posture representation data, such as the facial posture described by the processed features is consistent with the facial posture described by the reference three-dimensional posture representation data; then the processed features are spliced ​​with the coding features of the reference facial texture map to obtain spliced ​​features; finally, based on the decoding results of the spliced ​​features, a generated image is determined, so that the facial posture described by the generated image is consistent with the facial posture described by the reference three-dimensional posture representation data, and the facial texture similar to facial expression described by the generated image is consistent with the corresponding facial texture described by the reference facial texture map.

[0185] Based on the relevant contents of S1 to S3 above, it can be known that for the image generation method provided by the embodiment of the present disclosure, the original image, original three-dimensional posture representation data, reference three-dimensional posture representation data and reference facial texture map are first obtained, and the original three-dimensional posture representation data is determined based on the original image; the reference three-dimensional posture representation data and the reference facial texture map are both determined based on the reference information corresponding to the original image; and then based on the original image, the original three-dimensional posture representation data and the reference three-dimensional posture representation data, the posture motion representation feature is determined so that the posture motion representation feature can represent the posture motion information carried by the reference three-dimensional posture representation data, so that the posture motion representation feature can The method expresses the posture motion information required when changing from the posture represented by the original three-dimensional posture representation data to the posture represented by the reference three-dimensional posture representation data; then, based on the coding features of the original image, the coding features of the reference facial texture map and the posture motion representation features, a generated image is obtained, so that the generated image not only carries the facial state information described by the original image and the reference facial texture map, but also carries the three-dimensional spatial information described by the original three-dimensional posture representation data and the reference three-dimensional posture representation data, so that the generated image can describe the three-dimensional facial state, and then the generated image can better describe the facial state, which is conducive to improving the image generation effect.

[0186] In order to better understand the image generation method provided by the present disclosure, two scenarios are described below.

[0187] Scenario 1: When the image generation method provided by the present disclosure is applied to a model training scenario, such as the model training scenario shown in FIG4 , the model construction process provided by the present disclosure may include the following steps 71 to 74 .

[0188] Step 71: Acquire the original image, original 3D posture representation data, reference 3D posture representation data, and reference facial texture map.

[0189] It should be noted that the relevant content of step 71 can be found in S1 above, and for the sake of brevity, it will not be repeated here.

[0190] Step 72: Determine the gesture motion representation feature based on the original image, the original 3D gesture representation data, and the reference 3D gesture representation data.

[0191] It should be noted that the relevant content of step 72 can be found in S2 above, and for the sake of brevity, it will not be repeated here.

[0192] In addition, in some application scenarios, step 72 above can be implemented by a module in the image generation model. Based on this, it can be seen that when the generated image above is determined by the image generation model based on the original image, original 3D pose representation data, reference 3D pose representation data, and a reference facial texture map, and the image generation model includes a motion estimation network, the motion estimation network can be used to implement step 72. Based on this, it can be seen that in one possible implementation, the motion estimation network can be used to determine the pose motion representation features based on the original image, original 3D pose representation data, and reference 3D pose representation data.

[0193] In addition, the present disclosure does not limit the working principle of the above motion estimation network. For example, as shown in Figure 4, the working principle of the motion estimation network can be specifically as follows: after the original image, the original three-dimensional posture representation data and the reference three-dimensional posture representation data are spliced ​​to obtain the splicing result, the motion estimation network performs motion estimation processing on the splicing result to obtain the posture motion representation feature.

[0194] Based on the content of the previous paragraph, it can be seen that in one possible implementation, the above step 72 can specifically include: first splicing the original image, the original three-dimensional posture representation data and the reference three-dimensional posture representation data to obtain a splicing result; then the motion estimation network in the image generation model performs motion estimation processing on the splicing result to obtain posture motion representation features.

[0195] Step 73: Based on the coding features of the original image, the coding features of the reference facial texture map, and the posture and motion representation features, a generated image is obtained, where the generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image.

[0196] It should be noted that the relevant content of step 73 can be found in S3 above, and for the sake of brevity, it will not be repeated here.

[0197] In addition, in some application scenarios, step 73 above can be implemented by a module in the image generation model. Based on this, it can be seen that when the above-mentioned generated image is determined by the image generation model based on the original image, the original three-dimensional posture representation data, the reference three-dimensional posture representation data and the reference facial texture map, and the image generation model includes the above-mentioned motion estimation network, the encoder and the decoder, the encoder can be used to determine the encoding features of the above-mentioned original image and the encoding features of the above-mentioned reference facial texture map, and the decoder is used to decode the above-mentioned splicing features to obtain a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image. Among them, the relevant content of the splicing feature can be found above.

[0198] Based on the content of the above paragraph, it can be seen that in one possible implementation, the above step 73 can specifically include: after the encoder in the above image generation model determines the coding features of the above original image and the coding features of the above reference facial texture map, and the motion estimation network in the image generation model outputs the posture motion representation features, the coding features of the original image can be first subjected to posture adjustment processing according to the posture motion representation features to obtain processed features; then the processed features and the coding features of the reference facial texture map are spliced ​​to obtain spliced ​​features; finally, the decoder in the image generation model decodes the spliced ​​features to obtain a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image.

[0199] Step 74: Use the above generated image and the supervision information corresponding to the generated image to update the image generation model, and return to continue executing the above step 71 and subsequent steps until the preset stopping condition is reached.

[0200] The supervisory information corresponding to the generated image refers to guidance information pre-set for the generated image, for example, image generation guidance information pre-configured for the original image mentioned above.

[0201] In addition, the present disclosure does not limit the implementation method of the supervision information corresponding to the generated image above. For example, when the generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image, the supervision information corresponding to the generated image includes the supervision information corresponding to the two-dimensional plane image, the supervision information corresponding to the depth map, and the supervision information corresponding to the background segmentation map.

[0202] For the two-dimensional image described above, the supervisory information corresponding to the two-dimensional image refers to guidance information pre-set for the two-dimensional image, such as two-dimensional image generation guidance information pre-configured for the original image described above, so that the supervisory information can guide the generation process of the two-dimensional image. Furthermore, this disclosure does not limit the implementation of the supervisory information corresponding to the two-dimensional image. For example, the supervisory information corresponding to the two-dimensional image can be implemented using a reference image corresponding to the original image, such as the reference image shown in FIG4 .

[0203] Based on the content of the previous paragraph, it can be seen that when the image generation method provided by the present disclosure is applied to a model training scenario, such as the model training scenario shown in Figure 4, in one possible implementation method, the reference information corresponding to the original image above can be the supervision information corresponding to the two-dimensional plane image, so that the content of "reference information" that appears in each step that needs to be performed for the reference information in the model training scenario provided by the present disclosure can be replaced with the content of "supervision information corresponding to the two-dimensional plane image", so as to achieve some processing processes for the reference information with the help of the supervision information corresponding to the two-dimensional plane image.

[0204] For the depth map corresponding to the above two-dimensional plane image, the supervisory information corresponding to the depth map refers to the guidance information set in advance for the depth map, such as the depth map generation guidance information pre-configured for the original image above, so that the supervisory information can guide the generation process of the depth map. In addition, the present disclosure does not limit the implementation method of the supervisory information corresponding to the depth map. For example, the supervisory information corresponding to the depth map may refer to a manually provided depth map. For another example, the supervisory information corresponding to the depth map may be obtained by performing a depth map generation process on the supervisory information corresponding to the two-dimensional plane image. In this way, the supervisory information corresponding to the depth map can be automatically constructed, thereby effectively reducing the difficulty of obtaining training data. It should be noted that the present disclosure does not limit the implementation method of the depth map generation process. For example, it can adopt any existing or future method that can perform depth map generation processing on an image, such as with the help of a pre-built depth map generation model and other methods. Among them, the depth map generation model is used to perform depth map generation processing on the input data of the depth map generation model.

[0205] For the background segmentation map corresponding to the above two-dimensional plane image, the supervisory information corresponding to the background segmentation map refers to the guidance information set in advance for the background segmentation map, such as the guidance information generated in advance for the background segmentation map configured for the above original image, so that the supervisory information can guide the generation process of the background segmentation map. In addition, the present disclosure does not limit the implementation method of the supervisory information corresponding to the background segmentation map. For example, the supervisory information corresponding to the background segmentation map may refer to a manually provided background segmentation map. For another example, the supervisory information corresponding to the background segmentation map may be obtained by performing background segmentation processing on the supervisory information corresponding to the two-dimensional plane image. In this way, the supervisory information corresponding to the background segmentation map can be automatically constructed, thereby effectively reducing the difficulty of obtaining training data. It should be noted that the present disclosure does not limit the implementation method of the background segmentation processing. For example, it can adopt any existing or future method that can perform background segmentation processing on an image, such as with the help of a pre-constructed background segmentation model and other methods. Among them, the background segmentation model is used to perform background segmentation processing on the input data of the background segmentation model.

[0206] The preset stopping condition refers to the condition that needs to be met when model training stops; and the present disclosure does not limit the implementation method of the preset stopping condition. For example, the preset stopping condition may include: the model loss of the above-mentioned image generation model is lower than a preset loss threshold. For another example, the preset stopping condition may include: the rate of change of the model loss of the image generation model is lower than a preset loss change rate threshold. For another example, the preset stopping condition may include: the number of updates of the image generation model reaches a preset number threshold. The model loss of the image generation model is used to represent the performance of the image generation model; and the model loss of the image generation model is determined based on the above-mentioned generated image and the supervisory information corresponding to the generated image, such as the model loss is determined based on the difference representation data between the two-dimensional plane image and the supervisory information corresponding to the two-dimensional plane image, the difference representation data between the depth map corresponding to the two-dimensional plane image and the supervisory information corresponding to the depth map, and the difference representation data between the background segmentation map corresponding to the two-dimensional plane image and the supervisory information corresponding to the background segmentation map. It should be noted that the present disclosure does not limit the calculation method of the model loss; and the present disclosure does not limit the implementation method of the difference characterization data. For example, the difference characterization data can be determined using any distance calculation formula.

[0207] Based on the relevant content of steps 71 to 74 above, it can be seen that in some application scenarios, the training process of the above image generation model can be: first, randomly select two frames of video images from the sample video as the original image and the reference information corresponding to the original image; then determine the original three-dimensional posture representation data based on the original image, and determine the reference three-dimensional posture representation data and the reference facial texture map based on the reference information; then the image generation model performs image generation processing based on the original image, the original three-dimensional posture representation data, the reference three-dimensional posture representation data and the reference facial texture map to obtain a generated image, such as the predicted image 1, the background segmentation map of the predicted image 1 and the depth map of the predicted image 1 shown in Figure 4; then, based on the difference between the generated image and the supervisory information corresponding to the generated image, update the image generation model so that the updated image generation model has a better image generation effect, so that the next round of training process can be performed based on the updated image generation model, and the cycle is repeated until the preset stop condition is reached.

[0208] In addition, the present disclosure does not limit the execution entity of the above-mentioned image generation model training process. For example, in some application scenarios, the execution entity of the image generation model training process is a VR device. For another example, in some application scenarios, in order to better improve the model training effect, the execution entity of the image generation model training process is a server, so that the VR device can subsequently use the image generation model sent by the server to complete the image generation task, such as the image generation task shown in Figure 5.

[0209] Scenario 2: When the image generation method provided by the present disclosure is applied to a certain image generation scenario, such as the holographic video generation scenario shown in Figure 2 or the binocular image generation scenario shown in Figure 5, if the image generation method is applied to a VR device, the working principle of the VR device may include some or all of the steps in steps 81 to 84 below.

[0210] Step 81: The VR device obtains the original image, the original 3D posture representation data, the reference 3D posture representation data, and the reference facial texture map.

[0211] It should be noted that the relevant content of step 81 can be found in S1 above, and for the sake of brevity, it will not be repeated here.

[0212] Step 82: The VR device determines the gesture motion representation feature based on the original image, the original three-dimensional gesture representation data, and the reference three-dimensional gesture representation data.

[0213] It should be noted that the relevant contents of step 82 can be found in S2 above, and for the sake of brevity, they will not be described here in detail.

[0214] As an example, in one possible implementation, step 82 above may specifically be: the VR device first splices the original image, the original three-dimensional posture representation data, and the reference three-dimensional posture representation data to obtain a splicing result; and then the motion estimation network in the image generation model deployed on the VR device performs motion estimation processing on the splicing result to obtain posture motion representation features.

[0215] Step 83: The VR device obtains a generated image based on the coding features of the original image, the coding features of the reference facial texture map, and the posture and motion representation features. The generated image includes a three-dimensional stereo image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image.

[0216] It should be noted that the relevant contents of step 83 can be found in S3 above, and for the sake of brevity, they will not be described here in detail.

[0217] As an example, in one possible implementation, when the image generation model deployed on the above-mentioned VR device includes the above-mentioned motion estimation network, encoder and decoder, the above-mentioned step 83 may be specifically as follows: after the encoder in the image generation model determines the coding features of the above-mentioned original image and the coding features of the above-mentioned reference facial texture map, and the motion estimation network in the image generation model outputs the posture motion representation features, the VR device may first use the posture motion representation features to perform posture adjustment processing on the coding features of the original image to obtain processed features; then, the VR device splices the processed features with the coding features of the reference facial texture map to obtain spliced ​​features; then, the decoder in the image generation model decodes the spliced ​​features to obtain a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image, such as the predicted image 2, the background segmentation map of the predicted image 2, and the depth map of the predicted image 2 shown in Figure 5; finally, the VR device generates a three-dimensional stereo image based on the two-dimensional plane image and the depth map, such as the left-eye image and the right-eye image shown in Figure 5.

[0218] Step 84: The VR device displays the above three-dimensional image.

[0219] Based on the relevant contents of steps 81 to 84 above, it can be known that in some application scenarios, for the above VR device, after the VR device obtains the facial expression coefficient and head posture for the target object, the VR device can first perform image generation processing based on the facial expression coefficient, head posture and relevant information of the original image obtained from the server to obtain a three-dimensional stereo image, such as the left-eye image and the right-eye image shown in Figure 5; the VR device then displays the three-dimensional stereo image so that the user of the VR device can see the three-dimensional stereo image on the VR device. In this way, the VR device can perform image generation processing based on the facial expression coefficient and head posture obtained in real time, thereby enabling the VR device to generate a holographic video.

[0220] In addition, in some application scenarios, such as video calls, the VR device described above also needs to send the generated images to other VR devices for display. Based on this, the present disclosure also provides a possible implementation of the operating principle of the VR device. In this implementation, when the VR device is in video communication with a target device, the target device is a three-dimensional image display device, and the generated images include three-dimensional stereo images, the operating principle of the VR device may include at least the following step 85. The execution time of step 85 is later than the execution time of step 83 described above.

[0221] Step 85: The VR device sends the above three-dimensional stereoscopic image to the target device, and the target device is used to display the three-dimensional stereoscopic image.

[0222] The target device refers to a device that performs video communication with the VR device mentioned above. For example, when the VR device is the VR device 1 shown in FIG2 , the target device may be the VR device 2 shown in FIG2 .

[0223] In addition, for the target device mentioned above, the target device can be a three-dimensional image display device, so that the target device can be used to display three-dimensional stereoscopic images. It should be noted that the present disclosure does not limit the implementation of the three-dimensional image display device. For example, it can be implemented using any existing or future device capable of displaying three-dimensional stereoscopic images, such as a VR device.

[0224] Based on the relevant content of step 85 above, it can be known that when the above VR device and the target device are in a video communication state, and the target device is a three-dimensional image display device, the VR device performs image generation processing based on the facial expression coefficient and head posture obtained in real time to obtain a three-dimensional stereo image. After that, the VR device can send the three-dimensional stereo image to the target device for display, so that the user of the target device can view the three-dimensional stereo image sent by the VR device on the target device, thereby enabling the user of the target device to view the holographic video including multiple three-dimensional stereo images sent by the VR device on the target device.

[0225] In addition, in some application scenarios, such as video calls, the VR device described above also needs to send its generated images to other non-VR devices, such as the mobile phone shown in Figure 2. Based on this, the present disclosure also provides a possible implementation of the operating principle of the VR device. In this implementation, when the VR device is in video communication with a target device, the target device is a two-dimensional image display device, and the generated image includes a two-dimensional plane image and a three-dimensional stereo image, and the three-dimensional stereo image includes at least two channel images, the operating principle of the VR device may include at least the following step 86. The execution time of step 86 is later than the execution time of step 83 described above.

[0226] Step 86: The VR device sends the above two-dimensional plane image or any channel image in the above three-dimensional stereoscopic image to the target device, and the target device is used to display the two-dimensional plane image or the channel image.

[0227] Among them, the two-dimensional image display device is used to display two-dimensional images; and the present disclosure does not limit the implementation method of the two-dimensional image display device. For example, it can be implemented using any existing or future device that can display two-dimensional images, such as a mobile phone or other terminal device as shown in Figure 2.

[0228] Based on the relevant content of step 86 above, it can be known that when the above VR device and the target device are in a video communication state, the target device is a two-dimensional image display device, the above generated image includes a two-dimensional plane image and a three-dimensional stereo image, and the three-dimensional stereo image includes a left-eye channel image and a right-eye channel image, the VR device performs image generation processing based on the facial expression coefficient and head posture acquired in real time, and obtains the above two-dimensional plane image, the left-eye channel image and the right-eye channel image. After that, the VR device can send the two-dimensional plane image, the left-eye channel image or the right-eye channel image to the target device for display, so that the user of the target device can view the two-dimensional image sent by the VR device on the target device, thereby enabling the user of the target device to view the video including multiple two-dimensional images sent by the VR device on the target device.

[0229] Based on the image generation method provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide an image generation device, which will be explained and illustrated below in conjunction with Figure 6. Figure 6 is a schematic structural diagram of the image generation device provided in the embodiments of the present disclosure. It should be noted that for the technical details of the image generation device provided in the embodiments of the present disclosure, please refer to the relevant content of the image generation method above.

[0230] As shown in FIG6 , an image generating apparatus 600 provided by an embodiment of the present disclosure includes:

[0231] An acquisition unit 601 is configured to acquire an original image, original 3D posture representation data, reference 3D posture representation data, and a reference facial texture map, wherein the original 3D posture representation data is determined based on the original image; and the reference 3D posture representation data and the reference facial texture map are determined based on reference information corresponding to the original image.

[0232] A determining unit 602 is configured to determine a gesture motion representation feature based on the original image, the original 3D gesture representation data, and the reference 3D gesture representation data;

[0233] The generating unit 603 is configured to obtain a generated image based on the coding features of the original image, the coding features of the reference facial texture map, and the gesture motion representation features.

[0234] In one possible implementation, the generated image is determined by an image generation model based on the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference facial texture map;

[0235] The image generating device 600 further includes:

[0236] An updating unit is used to update the image generation model using the generated image and the supervision information corresponding to the generated image.

[0237] In one possible implementation, the generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image; the supervision information corresponding to the generated image includes supervision information corresponding to the two-dimensional plane image, supervision information corresponding to the depth map, and supervision information corresponding to the background segmentation map.

[0238] In a possible implementation manner, the reference information includes supervision information corresponding to the two-dimensional plane image.

[0239] In one possible implementation, the supervisory information corresponding to the depth map is obtained by performing depth map generation processing on the supervisory information corresponding to the two-dimensional plane image; the supervisory information corresponding to the background segmentation map is obtained by performing background segmentation processing on the supervisory information corresponding to the two-dimensional plane image.

[0240] In a possible implementation, the generating unit 603 includes:

[0241] a feature processing subunit, configured to process the coded features of the original image according to the gesture motion representation features to obtain processed features;

[0242] a feature splicing subunit, configured to splice the processed features with the coded features of the reference facial texture map to obtain a spliced ​​feature;

[0243] The first determining subunit is configured to obtain the generated image based on the splicing feature.

[0244] In a possible implementation manner, the generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image;

[0245] The first determining subunit is specifically configured to decode the splicing features to obtain the two-dimensional plane image, the depth map, and the background segmentation map.

[0246] In one possible implementation, the generated image is determined by an image generation model based on the original image, the original three-dimensional posture representation data, the reference three-dimensional posture representation data, and the reference facial texture map; the image generation model includes a motion estimation network, an encoder, and a decoder; the motion estimation network is used to determine the posture motion representation features based on the original image, the original three-dimensional posture representation data, and the reference three-dimensional posture representation data; the encoder is used to determine the encoding features of the original image and the encoding features of the reference facial texture map; the decoder is used to decode the splicing features to obtain the two-dimensional plane image, the depth map, and the background segmentation map.

[0247] In one possible implementation, the generated image includes a three-dimensional image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image;

[0248] The first determining subunit includes:

[0249] A decoding processing subunit, configured to decode the splicing features to obtain the two-dimensional plane image, the depth map, and the background segmentation map;

[0250] The image generation subunit is configured to generate the three-dimensional stereoscopic image based on the two-dimensional planar image and the depth map.

[0251] In one possible implementation, the three-dimensional stereoscopic image includes a first left-eye image and a first right-eye image; the first left-eye image and the first right-eye image are both obtained by stereoscopically transposing the two-dimensional plane image using the depth map.

[0252] In a possible implementation manner, the three-dimensional stereoscopic image includes a second left-eye image and a second right-eye image;

[0253] The image generation subunit is specifically used to: use the depth map to perform stereo transposition processing on the two-dimensional plane image to obtain a first left-eye image and a first right-eye image; use the background segmentation map to perform background removal processing on the first left-eye image to obtain the second left-eye image; use the background segmentation map to perform background removal processing on the first right-eye image to obtain the second right-eye image.

[0254] In one possible implementation, the determination unit 602 is specifically used to: splice the original image, the original three-dimensional posture representation data and the reference three-dimensional posture representation data to obtain a splicing result; and perform motion estimation processing on the splicing result to obtain the posture motion representation feature.

[0255] In one possible implementation, the original three-dimensional posture representation data is obtained by mapping the three-dimensional facial mesh corresponding to the original image to a two-dimensional image space; the three-dimensional facial mesh corresponding to the original image is constructed based on the original image; the reference three-dimensional posture representation data is obtained by mapping the three-dimensional facial mesh corresponding to the reference information to a two-dimensional image space; and the three-dimensional facial mesh corresponding to the reference information is constructed based on the reference information.

[0256] In one possible implementation, the original three-dimensional posture representation data is obtained by performing feature extraction processing on the three-dimensional facial mesh corresponding to the original image; the three-dimensional facial mesh corresponding to the original image is constructed based on the original image; the reference three-dimensional posture representation data is obtained by performing feature extraction processing on the three-dimensional facial mesh corresponding to the reference information; and the three-dimensional facial mesh corresponding to the reference information is constructed based on the reference information.

[0257] In one possible implementation, the generated image includes a two-dimensional plane image; the reference information is supervisory information corresponding to the two-dimensional plane image; and the three-dimensional facial mesh corresponding to the reference information is obtained by performing three-dimensional facial reconstruction processing on the supervisory information corresponding to the two-dimensional plane image.

[0258] In one possible implementation, determining the three-dimensional facial mesh corresponding to the reference information includes: performing three-dimensional facial reconstruction on the supervisory information corresponding to the two-dimensional plane image to obtain a first facial mesh; and performing expression adjustment on the first facial mesh according to preset expressionless parameters to obtain the three-dimensional facial mesh corresponding to the reference information.

[0259] In a possible implementation, the reference information includes a head posture; and the three-dimensional face mesh corresponding to the reference information is obtained by performing posture adjustment processing on the three-dimensional face mesh corresponding to the original image using the head posture.

[0260] In one possible implementation, determining the three-dimensional facial mesh corresponding to the original image includes: performing three-dimensional facial reconstruction on the original image to obtain a second facial mesh; and performing expression adjustment on the second facial mesh according to preset expressionless parameters to obtain the three-dimensional facial mesh corresponding to the original image.

[0261] In one possible implementation, the generated image includes a two-dimensional plane image; the reference information is the supervisory information corresponding to the two-dimensional plane image; and the reference facial texture map refers to a facial texture map obtained by performing three-dimensional facial reconstruction processing on the supervisory information corresponding to the two-dimensional plane image.

[0262] In one possible implementation, the reference information includes a facial expression coefficient; the reference facial texture map is obtained by performing expression adjustment processing on the facial texture map corresponding to the original image using the facial expression coefficient; and the facial texture map corresponding to the original image is constructed based on the original image.

[0263] In a possible implementation, the image generating device 600 is deployed on a virtual reality (VR) device.

[0264] In one possible implementation, the reference information includes facial expression coefficients and head posture; the generated image includes a three-dimensional stereo image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image.

[0265] In one possible implementation, the generated image includes a three-dimensional stereoscopic image; the VR device and a target device are in video communication, and the target device is a three-dimensional image display device;

[0266] The image generating device 600 further includes:

[0267] The sending unit is configured to send the three-dimensional stereoscopic image to the target device, and the target device is configured to display the three-dimensional stereoscopic image.

[0268] In one possible implementation, the generated image includes a two-dimensional plane image and a three-dimensional stereoscopic image, and the three-dimensional stereoscopic image includes at least two channel images; the VR device is in video communication with a target device, and the target device is a two-dimensional image display device;

[0269] The image generating device 600 further includes:

[0270] A sending unit is used to send the two-dimensional plane image or any one of the channel images in the three-dimensional stereoscopic image to the target device, and the target device is used to display the two-dimensional plane image or any one of the channel images.

[0271] In one possible implementation, the reference information corresponding to the original image includes a facial expression coefficient and a head posture acquired by the VR device for the target object;

[0272] The acquisition unit 601 is specifically configured to receive the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image, sent by the server; and obtain the original three-dimensional posture representation data, the reference three-dimensional posture representation data, and the reference facial texture map based on the three-dimensional facial mesh corresponding to the original image, the facial texture map corresponding to the original image, the facial expression coefficient, and the head posture.

[0273] In a possible implementation manner, the server is configured to generate the original image based on the facial representation image and style description information specified by the target object.

[0274] Based on the relevant content of the above-mentioned image generating device 600, it can be known that for the image generating device 600 provided by the embodiment of the present disclosure, the original image, original three-dimensional posture representation data, reference three-dimensional posture representation data and reference facial texture map are first obtained, and the original three-dimensional posture representation data is determined based on the original image; the reference three-dimensional posture representation data and the reference facial texture map are both determined based on the reference information corresponding to the original image; and then based on the original image, the original three-dimensional posture representation data and the reference three-dimensional posture representation data, the posture motion representation feature is determined so that the posture motion representation feature can represent the posture motion information carried by the reference three-dimensional posture representation data, so that the posture motion representation feature can represent the posture motion information carried by the reference three-dimensional posture representation data, thereby making the posture motion representation feature The feature can represent the posture motion information required when changing from the posture represented by the original three-dimensional posture representation data to the posture represented by the reference three-dimensional posture representation data; then, based on the encoding features of the original image, the encoding features of the reference facial texture map and the posture motion representation features, a generated image is obtained, so that the generated image not only carries the facial state information described by the original image and the reference facial texture map, but also carries the three-dimensional spatial information described by the original three-dimensional posture representation data and the reference three-dimensional posture representation data, so that the generated image can describe the three-dimensional facial state, and then the generated image can better describe the facial state, which is conducive to improving the image generation effect.

[0275] In addition, an embodiment of the present disclosure also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the image generation method provided by the embodiment of the present disclosure.

[0276] Referring to FIG7 , a schematic diagram of the structure of an electronic device 700 suitable for implementing embodiments of the present disclosure is shown. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG7 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.

[0277] As shown in Figure 7, electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of electronic device 700 are also stored in RAM 703. Processing device 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0278] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 7 shows the electronic device 700 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.

[0279] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0280] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0281] The embodiments of the present disclosure further provide a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the image generation method provided by the embodiments of the present disclosure.

[0282] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0283] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0284] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0285] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.

[0286] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0287] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0288] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.

[0289] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0290] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0291] It should be noted that the various embodiments of this disclosure are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the descriptions of the systems or devices disclosed in the embodiments for similarities and differences between them. Since the systems or devices disclosed in the embodiments correspond to the methods disclosed in the embodiments, their descriptions are relatively simple, and reference can be made to the descriptions of the methods for any related details.

[0292] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0293] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0294] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0295] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present disclosure. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein, but is intended to be construed in the widest manner consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating an image, comprising: Acquire an original image, original three-dimensional posture representation data, reference three-dimensional posture representation data, and a reference facial texture map, wherein the original three-dimensional posture representation data is determined based on the original image; The reference three-dimensional posture representation data and the reference facial texture map are determined based on reference information corresponding to the original image; Determining a gesture motion representation feature based on the original image, the original three-dimensional gesture representation data, and the reference three-dimensional gesture representation data; A generated image is obtained based on the encoding features of the original image, the encoding features of the reference facial texture map, and the gesture motion representation features.

2. The method according to claim 1, wherein: The generated image is determined by an image generation model based on the original image, the original three-dimensional posture representation data, the reference three-dimensional posture representation data and the reference facial texture map; After obtaining the generated image, the method further includes: The image generation model is updated using the generated image and the supervisory information corresponding to the generated image.

3. The method according to claim 2, wherein: The generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image; The supervisory information corresponding to the generated image includes supervisory information corresponding to the two-dimensional plane image, supervisory information corresponding to the depth map, and supervisory information corresponding to the background segmentation map.

4. The method according to claim 3, wherein: The reference information includes supervision information corresponding to the two-dimensional plane image.

5. The method according to claim 3 or 4, wherein: The supervisory information corresponding to the depth map is obtained by performing a depth map generation process on the supervisory information corresponding to the two-dimensional plane image; The supervisory information corresponding to the background segmentation map is obtained by performing background segmentation processing on the supervisory information corresponding to the two-dimensional plane image.

6. The method according to claim 1, wherein: The step of obtaining a generated image based on the coding features of the original image, the coding features of the reference facial texture map, and the gesture motion representation features includes: Processing the encoded features of the original image according to the gesture motion representation features to obtain processed features; splicing the processed features and the coded features of the reference facial texture map to obtain spliced ​​features; The generated image is obtained according to the splicing features.

7. The method according to claim 6, wherein: The generated image includes a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image; The step of obtaining the generated image based on the splicing features includes: The splicing features are decoded to obtain the two-dimensional plane image, the depth map and the background segmentation map.

8. The method according to claim 7, wherein: The generated image is determined by an image generation model based on the original image, the original three-dimensional posture representation data, the reference three-dimensional posture representation data and the reference facial texture map; The image generation model includes a motion estimation network, an encoder and a decoder; The motion estimation network is used to determine the gesture motion representation feature based on the original image, the original three-dimensional gesture representation data and the reference three-dimensional gesture representation data; The encoder is used to determine the encoding features of the original image and the encoding features of the reference facial texture map; The decoder is used to decode the splicing features to obtain the two-dimensional plane image, the depth map and the background segmentation map.

9. The method according to claim 6, wherein: The generated image includes a three-dimensional stereo image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image; The step of obtaining the generated image based on the splicing features includes: Decoding the splicing features to obtain the two-dimensional plane image, the depth map and the background segmentation map; The three-dimensional image is generated according to the two-dimensional planar image and the depth map.

10. The method according to claim 9, wherein: The three-dimensional stereoscopic image includes a first left-eye image and a first right-eye image; The first left-eye image and the first right-eye image are both obtained by performing a stereoscopic transposition process on the two-dimensional plane image using the depth map.

11. The method according to claim 9, wherein: The three-dimensional stereoscopic image includes a second left-eye image and a second right-eye image; The step of generating the three-dimensional image based on the two-dimensional plane image and the depth map includes: Performing stereoscopic transposition processing on the two-dimensional plane image using the depth map to obtain a first left-eye image and a first right-eye image; Using the background segmentation map, performing background removal processing on the first left-eye image to obtain the second left-eye image; The background segmentation map is used to perform background removal processing on the first right eye image to obtain the second right eye image.

12. The method according to any one of claims 1 to 11, wherein: The process of determining the gesture motion representation feature includes: splicing the original image, the original three-dimensional posture representation data, and the reference three-dimensional posture representation data to obtain a splicing result; Perform motion estimation processing on the splicing result to obtain the posture motion representation feature.

13. The method according to claim 1, wherein: The original three-dimensional posture representation data is obtained by mapping the three-dimensional face mesh corresponding to the original image to the two-dimensional image space; the three-dimensional face mesh corresponding to the original image is constructed based on the original image; The reference three-dimensional posture representation data is obtained by mapping the three-dimensional face mesh corresponding to the reference information to the two-dimensional image space; the three-dimensional face mesh corresponding to the reference information is constructed based on the reference information.

14. The method according to claim 1, wherein: The original three-dimensional posture representation data is obtained by performing feature extraction processing on the three-dimensional face mesh corresponding to the original image; the three-dimensional face mesh corresponding to the original image is constructed based on the original image; The reference three-dimensional posture representation data is obtained by performing feature extraction processing on the three-dimensional face mesh corresponding to the reference information; the three-dimensional face mesh corresponding to the reference information is constructed based on the reference information.

15. The method according to claim 13 or 14, wherein: The generated image includes a two-dimensional plane image; The reference information is supervision information corresponding to the two-dimensional plane image; The three-dimensional face mesh corresponding to the reference information is obtained by performing three-dimensional face reconstruction processing on the supervision information corresponding to the two-dimensional plane image.

16. The method according to claim 15, wherein: The process of determining the three-dimensional facial mesh corresponding to the reference information includes: Performing three-dimensional facial reconstruction processing on the supervision information corresponding to the two-dimensional plane image to obtain a first facial mesh; According to the preset expressionless parameters, expression adjustment processing is performed on the first facial mesh to obtain a three-dimensional facial mesh corresponding to the reference information.

17. The method according to claim 13 or 14, wherein: The reference information includes head posture; The three-dimensional face mesh corresponding to the reference information is obtained by performing posture adjustment processing on the three-dimensional face mesh corresponding to the original image using the head posture.

18. The method according to claim 13 or 14, wherein: The process of determining the three-dimensional face mesh corresponding to the original image includes: Performing three-dimensional facial reconstruction processing on the original image to obtain a second facial mesh; According to the preset expressionless parameters, expression adjustment processing is performed on the second facial mesh to obtain a three-dimensional facial mesh corresponding to the original image.

19. The method according to claim 1, wherein: The generated image includes a two-dimensional plane image; the reference information is the supervisory information corresponding to the two-dimensional plane image; the reference facial texture map refers to a facial texture map obtained by performing three-dimensional facial reconstruction processing on the supervisory information corresponding to the two-dimensional plane image; or, The reference information includes facial expression coefficients; the reference facial texture map is obtained by using the facial expression coefficients to perform expression adjustment processing on the facial texture map corresponding to the original image; the facial texture map corresponding to the original image is constructed based on the original image.

20. The method according to claim 1, wherein: The method is applied to virtual reality VR equipment.

21. The method according to claim 20, wherein: The reference information includes facial expression coefficients and head posture; The generated image includes a three-dimensional stereo image, a two-dimensional plane image, a depth map corresponding to the two-dimensional plane image, and a background segmentation map corresponding to the two-dimensional plane image.

22. The method according to claim 20 or 21, wherein: The generated image includes a three-dimensional stereo image; The VR device is in video communication with a target device, and the target device is a three-dimensional image display device; After obtaining the generated image, the method further includes: The three-dimensional stereoscopic image is sent to the target device, and the target device is used to display the three-dimensional stereoscopic image.

23. The method according to claim 20 or 21, wherein: The generated image includes a two-dimensional plane image and a three-dimensional stereoscopic image, and the three-dimensional stereoscopic image includes at least two channel images; The VR device and the target device are in video communication state, and the target device is a two-dimensional image display device; After obtaining the generated image, the method further includes: The two-dimensional plane image or any one of the channel images in the three-dimensional stereoscopic image is sent to the target device, and the target device is used to display the two-dimensional plane image or any one of the channel images.

24. The method according to any one of claims 20 to 23, wherein: The reference information corresponding to the original image includes a facial expression coefficient and a head posture acquired by the VR device for the target object; The obtaining of the original image, the original three-dimensional posture representation data, the reference three-dimensional posture representation data and the reference facial texture map includes: Receiving the original image, the three-dimensional face mesh corresponding to the original image, and the face texture map corresponding to the original image sent by the server; The original three-dimensional posture representation data, the reference three-dimensional posture representation data and the reference facial texture map are obtained according to the three-dimensional facial mesh corresponding to the original image, the facial texture map corresponding to the original image, the facial expression coefficient and the head posture.

25. The method according to claim 24, wherein: The server is used to generate the original image according to the facial representation image and style description information specified by the target object.

26. An image generating device, comprising: an acquisition unit, configured to acquire an original image, original three-dimensional posture representation data, reference three-dimensional posture representation data, and a reference facial texture map, wherein the original three-dimensional posture representation data is determined based on the original image; the reference three-dimensional posture representation data and the reference facial texture map are determined based on reference information corresponding to the original image; a determination unit configured to determine a gesture motion representation feature based on the original image, the original three-dimensional gesture representation data, and the reference three-dimensional gesture representation data; The generating unit is configured to obtain a generated image according to the encoding features of the original image, the encoding features of the reference facial texture map and the posture motion representation features.

27. An electronic device comprising a processor and a memory, in, The memory is configured to store instructions or computer programs; The processor is configured to execute the instructions or computer programs in the memory so that the electronic device performs the method according to any one of claims 1 to 25.

28. A computer readable medium storing instructions or a computer program, wherein: When the instructions or computer programs are executed on a device, the device is caused to execute the method according to any one of claims 1 to 25.

Citation Information

Patent Citations

  • Image processing method, device and apparatus, medium and electronic equipment

    CN111583399A

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN113221847A

  • Face image replaying method and system, electronic equipment and storage medium

    CN116310146A

  • Video generation method and device, storage medium and computer equipment

    CN117036583A

  • System and method for self-supervised depth and ego-motion overfitting

    US20210350222A1