Image generation method and device, electronic equipment and computer readable medium
By acquiring and processing the three-dimensional posture and texture features of the image, generating new images that maintain consistency of certain features is solved, which solves the problem of difficulty in generating consistent images in the prior art and improves the effect of image generation.
Patent Information
- Application Number
- CN202311629107.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively generate new images that maintain consistent certain features, such as hairstyles and facial contours.
By acquiring the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map, the pose motion representation features are determined and new images are generated using these features.
It realizes the generation of new images that are consistent with the original image in certain features, improving the effect and consistency of image generation.
Smart Images

Figure CN120070723A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and particularly to an image generation method, apparatus, electronic device, and computer-readable medium. Background Art
[0002] For some image-related scenarios, these scenarios may have the following requirements: generating a new image based on the original image provided by the user, so that the new image is consistent with the original image in some aspects, such as hairstyle, facial contour, etc. Summary of the Invention
[0003] The present application provides an image generation method, apparatus, electronic device, and computer-readable medium, which are beneficial to better meeting the image generation requirements.
[0004] To achieve the above object, the technical solutions provided by the present application are as follows:
[0005] The present application provides an image generation method, and the method includes:
[0006] Obtaining an original image, original three-dimensional pose representation data, reference three-dimensional pose representation data, and a reference facial texture map, where the original three-dimensional pose representation data is determined based on the original image; the reference three-dimensional pose representation data and the reference facial texture map are determined based on the reference information corresponding to the original image;
[0007] Determining a pose motion representation feature based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data;
[0008] Obtaining a generated image based on the encoded feature of the original image, the encoded feature of the reference facial texture map, and the pose motion representation feature.
[0009] In a possible implementation manner, the generated image is determined by an image generation model based on the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference facial texture map;
[0010] After obtaining the generated image, the method further includes:
[0011] Updating the image generation model by using the generated image and the supervision information corresponding to the generated image.
[0012] In a possible implementation, the generated image includes a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image; the supervision information corresponding to the generated image includes the supervision information corresponding to the two-dimensional planar image, the supervision information corresponding to the depth map, and the supervision information corresponding to the background segmentation map.
[0013] In a possible implementation, the reference information includes the supervision information corresponding to the two-dimensional planar image.
[0014] In a possible implementation, the supervision information corresponding to the depth map is obtained by performing depth map generation processing on the supervision information corresponding to the two-dimensional planar image; the supervision information corresponding to the background segmentation map is obtained by performing background segmentation processing on the supervision information corresponding to the two-dimensional planar image.
[0015] In a possible implementation, obtaining the generated image according to the encoded features of the original image, the encoded features of the reference facial texture map, and the pose motion representation features includes:
[0016] Processing the encoded features of the original image according to the pose motion representation features to obtain processed features;
[0017] Concatenating the processed features and the encoded features of the reference facial texture map to obtain concatenated features;
[0018] Obtaining the generated image according to the concatenated features.
[0019] In a possible implementation, the generated image includes a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image;
[0020] Obtaining the generated image according to the concatenated features includes:
[0021] Performing decoding processing on the concatenated features to obtain the two-dimensional planar image, the depth map, and the background segmentation map.
[0022] In a possible implementation manner, the generated image is determined by an image generation model based on the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map; the image generation model includes a motion estimation network, an encoder, and a decoder; the motion estimation network is used to determine a pose motion representation feature based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data; the encoder is used to determine an encoded feature of the original image and an encoded feature of the reference face texture map; the decoder is used to perform a decoding process on the spliced feature to obtain the two-dimensional planar image, the depth map, and the background segmentation map.
[0023] In a possible implementation manner, the generated image includes a three-dimensional stereoscopic image, a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image;
[0024] Obtaining the generated image according to the spliced feature includes:
[0025] Performing a decoding process on the spliced feature to obtain the two-dimensional planar image, the depth map, and the background segmentation map;
[0026] Generating the three-dimensional stereoscopic image based on the two-dimensional planar image and the depth map.
[0027] In a possible implementation manner, the three-dimensional stereoscopic image includes a first left-eye image and a first right-eye image; both the first left-eye image and the first right-eye image are obtained by performing a stereoscopic transposition process on the two-dimensional planar image using the depth map.
[0028] In a possible implementation manner, the three-dimensional stereoscopic image includes a second left-eye image and a second right-eye image;
[0029] Generating the three-dimensional stereoscopic image based on the two-dimensional planar image and the depth map includes:
[0030] Performing a stereoscopic transposition process on the two-dimensional planar image using the depth map to obtain a first left-eye image and a first right-eye image;
[0031] Performing a background removal process on the first left-eye image using the background segmentation map to obtain the second left-eye image;
[0032] Performing a background removal process on the first right-eye image using the background segmentation map to obtain the second right-eye image.
[0033] In a possible implementation, the process of determining the pose motion representation feature includes: splicing the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data to obtain a splicing result; performing motion estimation processing on the splicing result to obtain the pose motion representation feature.
[0034] In a possible implementation, the original three-dimensional pose representation data is obtained by mapping the three-dimensional face mesh corresponding to the original image to a two-dimensional image space; the three-dimensional face mesh corresponding to the original image is constructed based on the original image; the reference three-dimensional pose representation data is obtained by mapping the three-dimensional face mesh corresponding to the reference information to a two-dimensional image space; the three-dimensional face mesh corresponding to the reference information is constructed based on the reference information.
[0035] In a possible implementation, the original three-dimensional pose representation data is obtained by performing feature extraction processing on the three-dimensional face mesh corresponding to the original image; the three-dimensional face mesh corresponding to the original image is constructed based on the original image; the reference three-dimensional pose representation data is obtained by performing feature extraction processing on the three-dimensional face mesh corresponding to the reference information; the three-dimensional face mesh corresponding to the reference information is constructed based on the reference information.
[0036] In a possible implementation, the generated image includes a two-dimensional planar image; the reference information is the supervision information corresponding to the two-dimensional planar image; the three-dimensional face mesh corresponding to the reference information is obtained by performing three-dimensional face reconstruction processing on the supervision information corresponding to the two-dimensional planar image.
[0037] In a possible implementation, the process of determining the three-dimensional face mesh corresponding to the reference information includes: performing three-dimensional face reconstruction processing on the supervision information corresponding to the two-dimensional planar image to obtain a first face mesh; performing expression adjustment processing on the first face mesh according to a preset non-expression parameter to obtain the three-dimensional face mesh corresponding to the reference information.
[0038] In a possible implementation, the reference information includes a head pose; the three-dimensional face mesh corresponding to the reference information is obtained by performing pose adjustment processing on the three-dimensional face mesh corresponding to the original image by using the head pose.
[0039] In a possible implementation, the process of determining the three-dimensional face mesh corresponding to the original image includes: performing three-dimensional face reconstruction processing on the original image to obtain a second face mesh; performing expression adjustment processing on the second face mesh according to a preset non-expression parameter to obtain the three-dimensional face mesh corresponding to the original image.
[0040] In a possible implementation, the generated image includes a two-dimensional planar image; the reference information is the supervision information corresponding to the two-dimensional planar image; the reference face texture map refers to the face texture map obtained by performing three-dimensional face reconstruction processing on the supervision information corresponding to the two-dimensional planar image.
[0041] In a possible implementation, the reference information includes face expression coefficients; the reference face texture map is obtained by performing expression adjustment processing on the face texture map corresponding to the original image by using the face expression coefficients; the face texture map corresponding to the original image is constructed based on the original image.
[0042] In a possible implementation, the method is applied to a virtual reality (VR) device.
[0043] In a possible implementation, the reference information includes face expression coefficients and head pose; the generated image includes a three-dimensional stereoscopic image, a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image.
[0044] In a possible implementation, the generated image includes a three-dimensional stereoscopic image; the VR device and the target device are in a video communication state, and the target device is a three-dimensional image display device;
[0045] After obtaining the generated image, the method further includes:
[0046] Sending the three-dimensional stereoscopic image to the target device, and the target device is used to display the three-dimensional stereoscopic image.
[0047] In a possible implementation, the generated image includes a two-dimensional planar image and a three-dimensional stereoscopic image, and the three-dimensional stereoscopic image includes at least two channel images; the VR device and the target device are in a video communication state, and the target device is a two-dimensional image display device;
[0048] After obtaining the generated image, the method further includes:
[0049] Sending either the two-dimensional planar image or any one of the channel images of the three-dimensional stereoscopic image to the target device, and the target device is used to display the two-dimensional planar image or any one of the channel images.
[0050] In a possible implementation, the reference information corresponding to the original image includes face expression coefficients and head pose acquired by the VR device for the target object;
[0051] The obtaining of the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference facial texture map includes:
[0052] Receiving the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image sent by the server;
[0053] Based on the three-dimensional facial mesh corresponding to the original image, the facial texture map corresponding to the original image, the facial expression coefficients, and the head pose, obtaining the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference facial texture map.
[0054] In a possible implementation manner, the server is used to generate the original image according to the facial representation image and the style description information specified by the target object.
[0055] This application provides an image generation device, including:
[0056] An acquisition unit, configured to acquire an original image, original three-dimensional pose representation data, reference three-dimensional pose representation data, and a reference facial texture map, where the original three-dimensional pose representation data is determined according to the original image; the reference three-dimensional pose representation data and the reference facial texture map are determined according to the reference information corresponding to the original image;
[0057] A determination unit, configured to determine pose motion representation features according to the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data;
[0058] A generation unit, configured to obtain a generated image according to the encoded features of the original image, the encoded features of the reference facial texture map, and the pose motion representation features.
[0059] This application provides an electronic device, where the device includes: a processor and a memory;
[0060] The memory is used to store instructions or computer programs;
[0061] The processor is configured to execute the instructions or computer programs in the memory, so that the electronic device executes the image generation method provided by this application.
[0062] This application provides a computer-readable medium, characterized in that instructions or computer programs are stored in the computer-readable medium, and when the instructions or computer programs run on a device, the device is caused to execute the image generation method provided by this application.
[0063] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium. The computer program contains program codes for executing the image generation method provided by the present application.
[0064] Compared with the related art, the present application has at least the following advantages:
[0065] In the technical solution provided by the present application, first, an original image, original three-dimensional pose representation data, reference three-dimensional pose representation data, and a reference face texture map are obtained. The original three-dimensional pose representation data is determined based on the original image; both the reference three-dimensional pose representation data and the reference face texture map are determined based on the reference information corresponding to the original image. Then, based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data, a pose motion representation feature is determined, so that the pose motion representation feature can represent the pose motion information carried by the reference three-dimensional pose representation data, thereby enabling the pose motion representation feature to represent the pose motion information required when changing from the pose represented by the original three-dimensional pose representation data to the pose represented by the reference three-dimensional pose representation data. Then, based on the encoding feature of the original image, the encoding feature of the reference face texture map, and the pose motion representation feature, a generated image is obtained, so that the generated image not only carries the face state information described by the original image and the reference face texture map, but also carries the three-dimensional space information described by the original three-dimensional pose representation data and the reference three-dimensional pose representation data, thereby enabling the generated image to describe a three-dimensional face state, and further enabling the generated image to better describe the face state. This is beneficial to improving the image generation effect. Description of the Drawings
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0067] Figure 1 It is a flowchart of an image generation method provided by an embodiment of the present application;
[0068] Figure 2 It is a schematic diagram of an image generation scenario provided by an embodiment of the present application;
[0069] Figure 3 It is a schematic diagram of a construction process of an original image provided by an embodiment of the present application;
[0070] Figure 4Schematic diagram of an image generation process during model training provided by an embodiment of the present application;
[0071] Figure 5 Schematic diagram of an image generation process in an image generation scenario provided by an embodiment of the present application;
[0072] Figure 6 Schematic structural diagram of an image generation device provided by an embodiment of the present application;
[0073] Figure 7 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0074] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0075] To better understand the technical solution provided by the present application, the image generation method provided by the present application will be described below with reference to some accompanying drawings first. As Figure 1 shown, the image generation method provided by the embodiment of the present application includes S1 - S3 below. Among them, the Figure 1 is a flowchart of an image generation method provided by an embodiment of the present application.
[0076] S1: Obtain an original image, original three - dimensional pose characterization data, reference three - dimensional pose characterization data, and a reference face texture map. The original three - dimensional pose characterization data is determined based on the original image; the reference three - dimensional pose characterization data and the reference face texture map are determined based on the reference information corresponding to the original image.
[0077] Among them, the original image refers to a two - dimensional image required for use in the image generation process and used to provide partial face constraints. Moreover, the present application does not limit the constraints provided by the original image. For example, the original image can be used to provide constraints in aspects such as facial static features. The facial static features refer to facial features that hardly change or change slightly within a short period of time. Moreover, the present application does not limit the implementation manner of the facial static features. For example, the facial static features can include facial contours, hairstyles, etc.
[0078] In addition, the present application does not limit the implementation manner of the above - mentioned original image. For example, the original image can adopt the Figure 2 original image shown in Figure 3The original image shown, Figure 4 the original image 1 shown, or Figure 5 the original image 2 shown is implemented.
[0079] In addition, this application does not limit the acquisition process of the above original image. For the sake of easy understanding, the following will be described in combination with two scenarios.
[0080] Scenario 1, when the image generation method provided by this application is applied to a model training scenario, such as Figure 4 in the model training scenario shown, the acquisition process of the above original image can be: a video image randomly extracted from a sample video is used as the original image. Among them, the sample video refers to the video required during the model training process; and this application does not limit the acquisition method of the sample video.
[0081] Scenario 2, when the image generation method provided by this application is applied to a certain image generation scenario, such as Figure 2 the holographic video generation scenario shown or Figure 5 the binocular image generation scenario shown, if the image generation method is applied to a Virtual Reality (VR) device, the process for the VR device to obtain the above original image can be: first, the server generates the original image according to the facial representation image and style description information specified by the target object; then the server sends the original image to the VR device so that the VR device can perform image generation processing based on the original image later, such as Figure 5 the image generation processing shown. Among them, the target object refers to the user of the VR device, so that the VR device can be used to obtain the facial state of the target object in real time. Data communication can be carried out between the server and the VR device so that the server can provide some information to the VR device, such as Figure 2 the original image shown and its related information, etc. The facial representation image refers to the image specified by the target object and used to describe the facial state; and this application does not limit the implementation manner of the facial representation image. For example, it can be the image provided by the target object through some means, such as the image captured by an image capture device, etc. The style description information refers to the information specified by the target object and used to describe the style characteristics; and this application does not limit the style description information. For example, the style description information can include a style description text. Among them, the style description text is used to describe the image style required during the generation process of the original image. It should be noted that the use processes of the facial representation image and the style description information are all carried out on the premise of obtaining the authorization of the target object.
[0082] In addition, the present application does not limit the implementation manner of the step of "generating the original image according to the face representation image and the style description information specified by the target object" in the above paragraph. For example, this step can be implemented by using Figure 3 the process shown. Based on this, it can be known that the present application provides a possible implementation manner for the server to generate the original image above. In this implementation manner, when a large number of pre-constructed style images are stored in the server, such as Figure 3 the base maps under various styles shown, when the styles of different style images are different, the generation process of the original image may include: after the server receives the style description text provided by the target object, the server searches for the style image that best matches the style description text from these stored style images, so that the best-matched style image can represent an image that meets the style requirements of the target object; then, the server performs redrawing processing on the best-matched style image to obtain a plurality of candidate images, such as Figure 3 the multiple candidate base maps shown, so that the styles of any two candidate images are the same as the style of the best-matched style image, but there are some differences between any two candidate images, such as different hair lengths, different hair tip colors, etc., which is beneficial to improving image diversity; secondly, the server determines the candidate image finally selected by the target object according to the selection operation triggered by the target object for these candidate images; subsequently, the server performs face replacement processing on the finally selected candidate image by using the face representation image specified by the target object to obtain a face-replaced image, so that the face description information carried by the face-replaced image is the same as the face description information carried by the face representation image, and the other information carried by the face-replaced image except the face description information is the same as the other information carried by the finally selected candidate image except the face description information; finally, the server performs secondary redrawing processing on the face-replaced image, such as stylization processing, etc., to obtain the original image, so that the stylization degree of the original image is higher than the stylization degree of the face-replaced image, and the similarity degree between the face description information carried by the original image and the face description information carried by the face representation image is lower than the similarity degree between the face description information carried by the face-replaced image and the face description information carried by the face representation image, so as to balance the face similarity degree and the stylization degree, so that the original image can take into account the face state requirements and stylization requirements of the target object, thereby being beneficial to improving the image generation effect.
[0083] In addition, the present application does not limit the implementation manner of the step of "performing redrawing processing on the most matching style image to obtain a plurality of candidate images" in the above paragraph. For example, specifically, it may be: performing image generation processing on the most matching style image based on random noise to obtain a candidate image, so that there are a small number of difference points between the candidate image and the most matching style image. Another example is that specifically, it may be: performing image generation processing on the most matching style image based on an adjusted description text to obtain a candidate image, so that the candidate image meets the facial state constraints described in the adjusted description text, so as to realize the adjustment processing for a certain part of the most matching style image. Among them, the adjusted description text is used to describe which part of the most matching style image is to be adjusted and how; and the present application does not limit the implementation manner of the adjusted description text. For example, the adjusted description text may be implemented using a string such as "hair is blue".
[0084] Furthermore, the present application does not limit the implementation manner of the step of "performing secondary redrawing processing on the face-swapped image to obtain the original image" above. For example, specifically, it may be: performing image generation processing on the face-swapped image using a pre-constructed stylized image generation model to obtain the original image, so as to increase the stylization degree of the original image. Among them, the stylized image generation model refers to a machine learning model pre-constructed for the style described by the above style description text, so that the stylized image generation model can be used to generate images conforming to the style described by the above style description text.
[0085] Based on the above four paragraphs and Figure 3 the content shown, it can be seen that if the image generation method provided by the present application is applied to a VR device, such as Figure 2The VR device 1 or VR device 2 shown, etc., and this VR device can communicate with the server for data. Then, the original image required for the VR device to perform image generation processing can be provided by the server; and the server can be used to generate the original image according to the facial representation image and style description information specified by the target object. Among them, the facial representation image refers to an image provided by the target object to the server in a certain way for describing the facial state, so that the facial representation image is used to provide facial state constraints for the generation process of the original image. The style description information refers to the information provided by the target object to the server in a certain way for describing the image style, so that the style description information is used to provide style constraints for the generation process of the original image; and this application does not limit the implementation manner of the style description information. For example, the style description information can include the above style description text and the above "selection operation triggered for these candidate images", so that the style description information can more accurately represent the style requirements of the target object, so that the finally generated original image better conforms to the style requirements, which is beneficial to improving the image generation effect. It should be noted that this application does not limit the implementation manner of the target object providing the facial representation image and style description information to the server. For example, the target object can use a certain terminal device, such as the VR device, etc., to provide the facial representation image and style description information to the server.
[0086] The original three-dimensional pose representation data is used to represent the state of the facial pose described by the above original image in three-dimensional space, so that the original three-dimensional pose representation data can better represent the facial pose described by the original image. For example, when the original image is Figure 4 the original image 1 shown, the original three-dimensional pose representation data can be implemented using the display three-dimensional key point map of the original image 1 shown Figure 4 For another example, when the original image is Figure 5 the original image 2 shown, the original three-dimensional pose representation data can be implemented using the display three-dimensional key point map of the original image 2 shown Figure 5 For another example, when the original image is
[0087] In addition, for the above original three-dimensional pose representation data, the original three-dimensional pose representation data is determined according to the above original image, so that the original three-dimensional pose representation data can represent the state of the facial pose described by the original image in three-dimensional space; and this application does not limit the determination process of the original three-dimensional pose representation data. For example, it can specifically include step 11-step 12 below.
[0088] Step 11: Construct a three-dimensional facial mesh corresponding to the original image, so that the three-dimensional facial mesh can represent the state of the facial pose described by the original image in three-dimensional space.
[0089] Among them, the three-dimensional face mesh corresponding to the original image refers to the three-dimensional face mesh (Mesh) constructed based on the original image, so that the three-dimensional face mesh can represent the state of the face pose described by the original image in three-dimensional space.
[0090] In addition, the present application does not limit the implementation manner of step 11 above. For example, specifically, it can be: performing three-dimensional face reconstruction processing on the original image above to obtain the three-dimensional face mesh corresponding to the original image. It should be noted that the present application does not limit the implementation manner of the three-dimensional face reconstruction processing. For example, the three-dimensional face reconstruction processing can adopt any existing or future method capable of realizing three-dimensional face reconstruction processing, such as methods using a pre-constructed three-dimensional face reconstruction model, etc., for implementation. Among them, the three-dimensional face reconstruction model refers to a pre-constructed model with good three-dimensional face reconstruction function; and the present application does not limit the implementation manner of the three-dimensional face reconstruction model. For example, it can adopt any existing or future model with three-dimensional face reconstruction function, such as Figure 5 the three-dimensional face reconstruction model 2 shown, etc., for implementation.
[0091] In addition, in order to avoid the interference of facial expressions on the face pose, the present application also provides a possible implementation manner of step 11 above. In this implementation manner, step 11 specifically may include the following steps 111-step 112.
[0092] Step 111: Perform three-dimensional face reconstruction processing on the original image to obtain a second face mesh, so that the second face mesh can represent the state of the face pose and facial expression described by the original image in three-dimensional space.
[0093] Among them, the second face mesh refers to the three-dimensional face mesh obtained by performing three-dimensional face reconstruction processing on the original image, so that the second face mesh can represent the state of the face pose and facial expression described by the original image in three-dimensional space.
[0094] Step 112: Perform expression adjustment processing on the second face mesh according to the preset non-expression parameter to obtain the three-dimensional face mesh corresponding to the original image, so that the three-dimensional face mesh can represent the state of the face pose described by the original image in three-dimensional space, but the three-dimensional face mesh cannot represent the state of the facial expression described by the original image in three-dimensional space.
[0095] Among them, the preset expressionless parameter refers to the parameter required when adjusting a three-dimensional face mesh to an expressionless state; moreover, the present application does not limit this preset expressionless parameter. For example, this preset expressionless parameter can be determined according to the actual scenario. The expressionless state refers to a preset standard expression state, such as the states of not smiling, not crying, not opening the mouth, etc.
[0096] In addition, the present application does not limit the implementation manner of the expression adjustment process in step 112 above. For example, it can be implemented by using any existing or future method that can perform expression adjustment processing on a three-dimensional face mesh according to a certain expression.
[0097] Based on the relevant content of steps 111 to 112 above, for the original image above, first perform three-dimensional face reconstruction processing on the original image to obtain a three-dimensional face mesh with expression and pose, so that the three-dimensional face mesh can represent the face pose and the state of the face expression in the three-dimensional space described by the original image; then, according to the preset expressionless parameter, perform expression adjustment processing on the three-dimensional face mesh to obtain a three-dimensional face mesh with expressionless and pose, as the three-dimensional face mesh corresponding to the original image. In this way, it can effectively overcome the interference caused by the face expression to the face pose, so that the three-dimensional face mesh can better represent the state of the face pose described by the original image in the three-dimensional space, and thus is beneficial to improving the image generation effect.
[0098] Furthermore, the present application does not limit the implementation manner of step 11 above. For example, in some application scenarios, such as Figure 2 in the holographic video generation scenario shown, in order to better reduce the computing pressure of the VR device, this step 11 can be executed by the server. Based on this, it can be known that this step 11 can specifically be: the server constructs a three-dimensional face mesh corresponding to the original image, so that the server can subsequently send the three-dimensional face mesh corresponding to the original image to the VR device, so that the VR device can complete the image generation task according to the three-dimensional face mesh corresponding to the original image.
[0099] Step 12: Determine the original three-dimensional pose characterization data according to the three-dimensional face mesh corresponding to the original image above.
[0100] It should be noted that the present application does not limit the implementation manner of step 12 above. For example, this step 12 can specifically be: directly determine the three-dimensional face mesh corresponding to the original image as the original three-dimensional pose characterization data.
[0101] In addition, in some application scenarios, in order to better avoid the defects caused by the excessive amount of data carried by the three-dimensional face mesh, the present application also provides a possible implementation manner of step 12 above. In this implementation manner, step 12 may specifically be: performing feature extraction processing on the three-dimensional face mesh corresponding to the original image above to obtain original three-dimensional pose representation data, so that the original three-dimensional pose representation data can represent the features extracted for the three-dimensional face mesh, and making the amount of data carried by the original three-dimensional pose representation data less than the amount of data carried by the three-dimensional face mesh, thereby enabling the original three-dimensional pose representation data to represent the three-dimensional information described by the three-dimensional face mesh with less data, and further enabling the image generation process implemented based on the original three-dimensional pose representation data to have less resource consumption.
[0102] In addition, in some application scenarios, in order to better improve the representation effect of the face pose on the premise of avoiding the defects caused by the excessive amount of data carried by the three-dimensional face mesh, a visualization method can be used to represent the face pose. Based on this, the present application also provides a possible implementation manner of step 12 above. In this implementation manner, step 12 may specifically be: mapping the three-dimensional face mesh corresponding to the original image above to the two-dimensional image space to obtain original three-dimensional pose representation data, such as Figure 4 the explicit three-dimensional key point map of the original image 1 shown or Figure 5 the explicit three-dimensional key point map of the original image 2 shown, etc., so that the original three-dimensional pose representation data is a two-dimensional image, thereby enabling the original three-dimensional pose representation data to visually represent the state of the face pose described by the original image in the three-dimensional space by means of the two-dimensional image, and making the amount of data carried by the original three-dimensional pose representation data less than the amount of data carried by the three-dimensional face mesh, and further enabling the original three-dimensional pose representation data to visually represent the three-dimensional information described by the three-dimensional face mesh with less data. In this way, it is possible to improve the representation effect of the face pose on the premise of minimizing resource consumption as much as possible, which is beneficial to improving the image generation effect. Among them, the two-dimensional image space is used to describe the space where the two-dimensional image is located; and the present application does not limit the implementation manner of the two-dimensional image space. For example, the two-dimensional image space may refer to a blank two-dimensional image. It should be noted that the present application does not limit the implementation manner of the mapping. For example, it may be implemented by using any existing or future method that can map a three-dimensional face mesh into a two-dimensional image.
[0103] Furthermore, for the original three-dimensional pose representation data shown in the above two paragraphs, in order to better perform subsequent splicing processing, such as Figure 4 the splicing processing shown or Figure 5For the splicing process shown, etc., the original three-dimensional pose representation data may have the following characteristics: the size of the original three-dimensional pose representation data in the target dimension is the same as the size of the original image in the target dimension, so that subsequent splicing processing of the original three-dimensional pose representation data and the original image can be performed based on the target dimension. Herein, the target dimension refers to a preset dimension that needs to be aligned during the splicing process; moreover, the target dimension can be set according to the actual application scenario. For example, the target dimension can be the width or the height.
[0104] In addition, the present application does not limit the implementation manner of step 12 above. For example, in some application scenarios, such as Figure 2 in the holographic video generation scenario shown, step 12 can be executed by a VR device. Based on this, it can be known that step 12 can specifically be: the VR device determines the original three-dimensional pose representation data based on the three-dimensional face mesh corresponding to the original image above.
[0105] Based on the relevant content of steps 11 to 12 above, for some application scenarios, after obtaining the original image, the three-dimensional face mesh corresponding to the original image can be constructed first, so that the three-dimensional face mesh can represent the state of the face pose described by the original image in the three-dimensional space; then the original three-dimensional pose representation data can be determined based on the three-dimensional face mesh, so that the original three-dimensional pose representation data can also represent the state of the face pose described by the original image in the three-dimensional space.
[0106] For the original image above, the reference information corresponding to the original image refers to the information required for image generation processing based on the original image and used to provide other face constraints in addition to the face constraints described by the original image; moreover, the present application does not limit the constraints provided by the reference information. For example, the reference information can be used to provide constraints in aspects such as face dynamic characteristics. Herein, the face dynamic characteristics refer to the face characteristics that change within a short period of time, such as significant changes; moreover, the present application does not limit the implementation manner of the face dynamic characteristics. For example, the face dynamic characteristics can include facial expressions, face poses, etc.
[0107] In addition, the present application does not limit the implementation manner of the reference information corresponding to the original image above. For the convenience of understanding, the following is described in combination with two scenarios.
[0108] Scenario 1, when the image generation method provided by the present application is applied to a model training scenario, such as Figure 4 in the model training scenario shown, the reference information corresponding to the original image above can adopt the reference image corresponding to the original image, such as Figure 4The reference image and the like shown are implemented. Among them, the reference image is used to provide facial expression constraints, facial pose constraints, etc. required when performing image generation processing based on the original image, so that the similarity between the reference image and the two-dimensional image generated based on the original image can be used for model performance evaluation processing subsequently. The "two-dimensional image generated based on the original image" refers to the image generated based on the original image and the reference image. For example, Figure 4 The predicted image 1 shown. In addition, the present application does not limit the acquisition method of the reference image corresponding to the original image. For example, when the original image is a video image randomly extracted from a sample video, the reference image can be any other video image in the sample video except the original image.
[0109] Scenario 2, when the image generation method provided by the present application is applied to a certain image generation scenario, such as Figure 2 The holographic video generation scenario shown or Figure 5 When the binocular image generation scenario shown is used, if the image generation method is applied to a VR device, the reference information corresponding to the original image above may include the facial expression coefficient and the head pose obtained by the VR device for the target object. Among them, the head pose is used to describe the facial pose of the target object; and the present application does not limit the acquisition method of the head pose. For example, the head pose can be collected in real time by an attitude sensor in the VR device, such as an inertial sensor similar to a DOF sensor, for the target object. The facial expression coefficient is used to represent the facial expression of the target object; and the present application does not limit the acquisition process of the facial expression coefficient. For example, the facial expression coefficient can be the 52-dimensional expression coefficient determined in real time by the 52-dimensional expression coefficient acquisition module in the VR device for the target object. Among them, the 52-dimensional expression coefficient acquisition module refers to a module deployed in the VR device that has the function of acquiring 52-dimensional expression coefficients; and the present application does not limit the working principle of the 52-dimensional expression coefficient acquisition module. For example, specifically, after the 52-dimensional expression coefficient acquisition module acquires the facial image collected for the target object, the 52-dimensional expression coefficient acquisition module performs 52-dimensional expression coefficient extraction processing on the facial image to obtain the 52-dimensional expression coefficient corresponding to the facial image. It should be noted that the use process of the facial image is carried out on the premise that the authorization of the target object has been obtained.
[0110] In addition, to better improve the expression accuracy, the present application also provides a process for obtaining the facial expression coefficient in the above paragraph. Specifically, it can be as follows: After the 52-dimensional expression coefficient acquisition module in the VR device obtains the 52-dimensional expression coefficient from the facial image of the target object, and the mouth shape acquisition module in the VR device obtains the 52-dimensional expression coefficient from the voice data of the target object, these two 52-dimensional expression coefficients are subjected to a certain process, such as weighted summation processing, etc., to obtain the facial expression coefficient. It should be noted that the use process of the voice data is carried out on the premise that the authorization of the target object has been obtained.
[0111] For the reference information corresponding to the original image above, the reference information can be used to provide facial pose constraints and facial expression constraints. After obtaining the reference information corresponding to the original image, the reference three-dimensional pose representation data and the reference facial texture map can be determined based on the reference information, so that the reference three-dimensional pose representation data is used to represent the facial pose constraint, and the reference facial texture map is used to represent the facial expression constraint, so that subsequent image generation processing can be carried out based on the reference three-dimensional pose representation data and the reference facial texture map to generate an image that meets the facial pose constraint and the facial expression constraint.
[0112] The reference three-dimensional pose representation data is used to represent the state of the facial pose described by the reference information corresponding to the original image above in the three-dimensional space, so that the reference three-dimensional pose representation data can better represent the facial pose described by the reference information. For example, when the reference information is Figure 4 the reference image shown, the reference three-dimensional pose representation data can be implemented using the display three-dimensional key point map of the reference image shown. Another example is when the reference information is Figure 4 the reference information provided by the VR device shown, the reference three-dimensional pose representation data can be implemented using the display three-dimensional key point map of the reference information shown. Figure 5 shown. Figure 5 shown.
[0113] In addition, the present application does not limit the determination process of the above reference three-dimensional pose representation data. For example, it may include step 21-step 22 below.
[0114] Step 21: Construct a three-dimensional facial mesh corresponding to the above reference information, so that the three-dimensional facial mesh can represent the state of the facial pose described by the reference information in the three-dimensional space.
[0115] Among them, the three-dimensional facial mesh corresponding to the above reference information refers to the three-dimensional facial mesh (Mesh) constructed based on the reference information, so that the three-dimensional facial mesh can represent the state of the facial pose described by the reference information in the three-dimensional space.
[0116] In addition, the present application does not limit the implementation manner of step 21 above. For ease of understanding, the following will be described in conjunction with two scenarios.
[0117] Scenario 1: When the image generation method provided by the present application is applied to a model training scenario, such as Figure 4 the model training scenario shown, if the reference information corresponding to the original image above is an image, such as Figure 4 the reference image shown or the supervision information corresponding to the two-dimensional planar image below, etc., then the implementation manner of step 21 above is similar to the implementation manner of step 11 above. Based on this, it can be known that in a possible implementation manner, this step 21 may specifically be: performing three-dimensional face reconstruction processing on the reference information to obtain a three-dimensional face mesh corresponding to the reference information. In another possible implementation manner, in order to better avoid the interference of facial expressions on facial poses, this step 21 may specifically be: performing three-dimensional face reconstruction processing on the reference information to obtain a first face mesh, so that the first face mesh can represent the facial pose and the state of the facial expression in three-dimensional space described by the reference information; then performing expression adjustment processing on the first face mesh according to a preset non-expression parameter to obtain a three-dimensional face mesh corresponding to the reference information. Among them, the first face mesh refers to the three-dimensional face mesh obtained by performing three-dimensional face reconstruction processing on the reference information, so that the first face mesh can represent the facial pose and the state of the facial expression in three-dimensional space described by the reference information.
[0118] Scenario 2: When the image generation method provided by the present application is applied to a certain image generation scenario, such as Figure 2 the holographic video generation scenario shown or Figure 5 the binocular image generation scenario shown, if the image generation method is applied to a VR device, and the reference information corresponding to the original image above includes the facial expression coefficient and the head pose obtained by the VR device for the target object, then this step 21 may specifically be: using the head pose to perform pose adjustment processing on the three-dimensional face mesh corresponding to the original image above to obtain a three-dimensional face mesh corresponding to the reference information, so that the three-dimensional face mesh can describe the state of the head pose in three-dimensional space. It should be noted that the present application does not limit the implementation manner of this pose adjustment processing. For example, it can be implemented by using any existing or future method that can perform pose adjustment processing on a three-dimensional face mesh according to a certain pose.
[0119] Furthermore, the present application does not limit the implementation manner of step 21 above. For example, in some application scenarios, such as Figure 2In the holographic video generation scenario shown, this step 21 is performed by the VR device. Based on this, it can be known that this step 21 can specifically be: after the VR device obtains the reference information for the target object, the VR device constructs a three-dimensional face mesh corresponding to the reference information, so that the three-dimensional face mesh can represent the state of the face pose described by the reference information in three-dimensional space.
[0120] Step 22: Determine the reference three-dimensional pose representation data based on the three-dimensional face mesh corresponding to the above-mentioned reference information.
[0121] It should be noted that the implementation manner of the above step 22 is similar to the implementation manner of the above step 12. For the convenience of understanding, the following will be described with three examples.
[0122] Example 1, in a possible implementation manner, in order to better avoid the defects caused by the excessive amount of data carried by the three-dimensional face mesh, the above step 22 can specifically be: perform feature extraction processing on the three-dimensional face mesh corresponding to the above-mentioned reference information to obtain the reference three-dimensional pose representation data, so that the reference three-dimensional pose representation data can represent the features extracted for the three-dimensional face mesh, and make the amount of data carried by the reference three-dimensional pose representation data less than the amount of data carried by the three-dimensional face mesh, so that the reference three-dimensional pose representation data can represent the three-dimensional information described by the three-dimensional face mesh with less data, and further make the image generation process implemented based on the reference three-dimensional pose representation data have less resource consumption.
[0123] Example 2, in a possible implementation manner, in order to better improve the representation effect of the face pose on the premise of avoiding the defects caused by the excessive amount of data carried by the three-dimensional face mesh, the above step 22 can specifically be: map the three-dimensional face mesh corresponding to the above-mentioned reference information to the two-dimensional image space to obtain the reference three-dimensional pose representation data, such as Figure 4 the explicit three-dimensional key point map of the reference image shown or Figure 5 the explicit three-dimensional key point map of the reference information shown, etc., so that the reference three-dimensional pose representation data is a two-dimensional image, so that the reference three-dimensional pose representation data can visually represent the state of the face pose described by the reference information in three-dimensional space through the two-dimensional image, and make the amount of data carried by the reference three-dimensional pose representation data less than the amount of data carried by the three-dimensional face mesh, and further make the reference three-dimensional pose representation data visually represent the three-dimensional information described by the three-dimensional face mesh with less data, so that it is possible to improve the representation effect of the face pose on the premise of minimizing resource consumption as much as possible, which is beneficial to improving the image generation effect.
[0124] It should be noted that for the reference three-dimensional pose representation data shown in the above two paragraphs, in order to better perform subsequent stitching processing, such as the stitching processing shown in Figure 4 or the stitching processing shown in Figure 5 and so on, the reference three-dimensional pose representation data has the following characteristics: the size of the reference three-dimensional pose representation data in the target dimension is the same as the size of the above-mentioned original image in the target dimension, so that subsequent stitching processing of the reference three-dimensional pose representation data and the original image can be performed based on this target dimension.
[0125] In addition, the present application does not limit the implementation manner of step 22 above. For example, in some application scenarios, such as the holographic video generation scenario shown in Figure 2 step 22 may be executed by a VR device. Based on this, it can be known that step 22 may specifically be: the VR device determines the reference three-dimensional pose representation data according to the three-dimensional face mesh corresponding to the above-mentioned reference information.
[0126] Based on the relevant content of steps 21 to 22 above, for some application scenarios, after obtaining the reference information corresponding to the above-mentioned original image, a three-dimensional face mesh corresponding to the reference information may be constructed first, so that the three-dimensional face mesh can represent the state of the face pose described by the reference information in three-dimensional space; then the reference three-dimensional pose representation data is determined according to the three-dimensional face mesh, so that the reference three-dimensional pose representation data can also represent the state of the face pose described by the reference information in three-dimensional space.
[0127] The reference face texture map is used to describe the expression state described by the reference information corresponding to the above-mentioned original image. For example, when the reference information is the reference image shown in Figure 4 , the reference face texture map may be implemented using the face texture map of the reference image shown in Figure 4 . Another example is that when the reference information is the reference information shown in Figure 5 , the reference face texture map may be implemented using the face texture map of the reference information shown in Figure 5 .
[0128] In addition, the present application does not limit the implementation manner of the above-mentioned reference face texture map. For the sake of easy understanding, two scenarios are described below.
[0129] Scenario 1, when the image generation method provided by the present application is applied to a model training scenario, such as the model training scenario shown in Figure 4 , if the reference information corresponding to the above-mentioned original image is an image, such as the image shown in Figure 4The reference image shown or the supervision information corresponding to the two-dimensional planar image below, etc. Then, the above-mentioned reference face texture map may refer to the face texture map (Texture) obtained by performing three-dimensional face reconstruction processing on the reference information.
[0130] Scenario 2, when the image generation method provided in this application is applied to a certain image generation scenario, such as Figure 2 the holographic video generation scenario shown or Figure 5 the binocular image generation scenario shown. When the image generation method is applied to a VR device, and the reference information corresponding to the above-mentioned original image includes the facial expression coefficients and head postures obtained by the VR device for the target object, the above-mentioned reference face texture map may be obtained by performing expression adjustment processing on the face texture map corresponding to the above-mentioned original image using the facial expression coefficients, so that the facial expression described by the reference face texture map is consistent with the facial expression described by the facial expression coefficients, and other facial states except the facial expression described by the reference face texture map are consistent with other facial states except the facial expression described by the face texture map corresponding to the original image. Among them, the face texture map corresponding to the original image is constructed based on the original image; and this application does not limit the construction process of the face texture map corresponding to the original image. For example, specifically, it may be: performing three-dimensional face reconstruction processing on the original image to obtain the face texture map corresponding to the original image.
[0131] In addition, this application does not limit the implementation manner of S1 above. For the convenience of understanding, the following is described in conjunction with two scenarios.
[0132] Scenario 1, when the image generation method provided in this application is applied to a model training scenario, such as Figure 4 the model training scenario shown. Specifically, S1 above may be: randomly extracting two video images from the sample video as the original image and the reference information corresponding to the original image respectively; then determining the original three-dimensional pose representation data based on the original image, and determining the reference three-dimensional pose representation data and the reference face texture map based on the reference information.
[0133] Scenario 2, when the image generation method provided in this application is applied to a certain image generation scenario, such as Figure 2 the holographic video generation scenario shown or Figure 5 the binocular image generation scenario shown. When the image generation method is applied to a VR device, and the reference information corresponding to the above-mentioned original image includes the facial expression coefficients and head postures obtained by the VR device for the target object, then S1 above may specifically include the following steps 31 - step 32.
[0134] Step 31: The VR device receives the original image sent by the server, the three-dimensional face mesh corresponding to the original image, and the face texture map corresponding to the original image.
[0135] In the present application, for the server, after the server generates the original image according to the facial representation image and style description information specified by the target object, the server can construct a three-dimensional facial mesh corresponding to the original image and a facial texture map corresponding to the original image according to the original image, so that the three-dimensional facial mesh can represent the facial posture described by the original image, and the facial texture map can represent the facial expression described by the original image. The server then sends the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image to the VR device, so that the VR device can complete the image generation task based on this information, such as Figure 5 The image generation task shown.
[0136] Step 32: The VR device obtains original three-dimensional posture representation data, reference three-dimensional posture representation data, and a reference facial texture map based on the three-dimensional facial mesh corresponding to the original image, the facial texture map corresponding to the original image, and the facial expression coefficient and head posture obtained by the VR device for the target object.
[0137] In the present application, for a VR device, after the VR device receives an original image sent by a server, a three-dimensional face mesh corresponding to the original image, and a face texture map corresponding to the original image, the VR device determines the original three-dimensional posture representation data based on the three-dimensional face mesh corresponding to the original image, and after the VR device obtains a facial expression coefficient and a head posture for a target object, the VR device uses the facial expression coefficient to perform expression adjustment processing on the face texture map to obtain a reference facial texture map, and the VR device uses the head posture to perform posture adjustment processing on the three-dimensional face mesh, and the VR device determines the reference three-dimensional posture representation data based on the three-dimensional face mesh after posture adjustment, so that the VR device can subsequently complete an image generation task based on this information, such as Figure 5 The image generation task shown.
[0138] Based on the relevant contents of steps 31 to 32 above, it can be known that in some application scenarios, for the VR device above, when the VR device performs an image generation task, part of the information required for the image generation task can come from the server, and the other part of the information is determined by the VR device based on the information provided by the server. In this way, the server shares part of the computing pressure of the VR device, which is beneficial to reduce the computing pressure of the VR device, and further beneficial to ensure the rapid execution of the image generation task.
[0139] In addition, for some video generation scenarios, such as Figure 2 the holographic video generation scenario shown, since the original image required for generating each frame of the video image is the same image, in order to better improve the video generation efficiency, after the VR device receives the original image sent by the server, the three-dimensional face mesh corresponding to the original image, and the face texture map corresponding to the original image, the VR device can directly store this information, so that when the generation task of each frame of the video image is triggered, these information can be directly read from the storage space for use. In this way, it can effectively avoid the consumption of computing resources caused by repeatedly obtaining the original image and its related information, which is beneficial to improving the video generation efficiency.
[0140] Based on the above content, in a possible implementation manner, for the above VR device, the working principle of the VR device can be: after the VR device receives the original image sent by the server, the three-dimensional face mesh corresponding to the original image, and the face texture map corresponding to the original image, the VR device stores the original image, the three-dimensional face mesh corresponding to the original image, and the face texture map corresponding to the original image, so that when the VR device executes the generation task of the nth frame of the video image, first the VR device reads the original image, the three-dimensional face mesh corresponding to the original image, and the face texture map corresponding to the original image from the storage space; then the VR device obtains the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map based on the three-dimensional face mesh corresponding to the original image, the face texture map corresponding to the original image, and the face expression coefficient and head pose obtained by the VR device for the target object, so that the VR device can complete the image generation task based on this information later, such as Figure 5 the image generation task shown.
[0141] Based on the relevant content of S1 above, for some scenarios, such as Figure 4 the model training scenario shown or Figure 5 the three-dimensional stereoscopic image generation scenario shown, obtain the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map, so that subsequent image generation processing can be performed based on this information.
[0142] S2: Determine the pose motion representation features based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data.
[0143] Among them, the pose motion representation feature is used to represent the pose motion information carried by the above-mentioned reference three-dimensional pose representation data, so that the pose motion representation feature can represent the pose motion information required when the pose represented by the above-mentioned original three-dimensional pose representation data changes to the pose represented by the reference three-dimensional pose representation data, so that the pose motion representation feature can to a certain extent represent the difference between the facial pose represented by the original three-dimensional pose representation data and the facial pose represented by the reference three-dimensional pose representation data.
[0144] In addition, the present application does not limit the implementation manner of S2 above. For example, it may specifically include the following steps 41-step 42.
[0145] Step 41: Concatenate the above-mentioned original image, the above-mentioned original three-dimensional pose representation data, and the above-mentioned reference three-dimensional pose representation data to obtain a concatenation result.
[0146] Among them, the concatenation result refers to the result obtained by concatenating the above-mentioned original image, the above-mentioned original three-dimensional pose representation data, and the above-mentioned reference three-dimensional pose representation data.
[0147] In addition, the present application does not limit the implementation manner of the concatenation in step 41 above. For example, it may adopt any existing or future method that can concatenate multiple data into one data, such as Figure 4 the concatenation processing shown or Figure 5 the concatenation processing shown, etc., for implementation.
[0148] Step 42: Perform motion estimation processing on the above-mentioned concatenation result to obtain a pose motion representation feature.
[0149] It should be noted that the present application does not limit the implementation manner of the motion estimation processing in step 42 above. For example, it may adopt any existing or future method that can perform motion estimation processing, such as Figure 4 the motion estimation network shown or Figure 5 the motion estimation network shown, etc., for implementation.
[0150] Based on the relevant content of the above steps 41 to 42, for some application scenarios, after obtaining the above-mentioned original image, the above-mentioned original three-dimensional pose representation data, and the above-mentioned reference three-dimensional pose representation data, the three data can be concatenated first; then motion estimation processing is performed on the concatenation result to obtain a pose motion representation feature, so that the pose motion representation feature is used to represent the pose motion information carried by the reference three-dimensional pose representation data.
[0151] In addition, the present application does not limit the implementation manner of S2 above. For example, when the image generation method provided by the present application is applied to a certain image generation scenario, such asFigure 2 the holographic video generation scenario shown or Figure 5 in the binocular image generation scenario shown, if the image generation method is applied to a VR device, then this S2 can be executed by the VR device. Based on this, it can be known that in a possible implementation manner, this S2 can specifically be: The VR device determines the pose motion representation feature based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data.
[0152] Based on the relevant content of S2 above, after obtaining the above-mentioned original image, the above-mentioned original three-dimensional pose representation data, and the above-mentioned reference three-dimensional pose representation data, the pose motion representation feature can be determined based on these three data, so that the pose motion representation feature is used to represent the pose motion information carried by the reference three-dimensional pose representation data, so that subsequent image generation processing can be performed based on the pose motion representation feature.
[0153] S3: Obtain a generated image based on the encoding feature of the original image, the encoding feature of the reference face texture map, and the pose motion representation feature.
[0154] Among them, the generated image refers to the image generated based on the encoding feature of the above-mentioned original image, the encoding feature of the above-mentioned reference face texture map, and the above-mentioned pose motion representation feature.
[0155] In addition, the present application does not limit the implementation manner of the above-mentioned generated image. For the sake of understanding, the following will be described in conjunction with two scenarios.
[0156] Scenario 1, when the image generation method provided by the present application is applied to a model training scenario, such as Figure 4 in the model training scenario shown, the above-mentioned generated image may include a two-dimensional plane image, the depth map corresponding to the two-dimensional plane image, and the background segmentation map corresponding to the two-dimensional plane image. Among them, the two-dimensional plane image refers to the two-dimensional image generated based on the encoding feature of the above-mentioned original image, the encoding feature of the above-mentioned reference face texture map, and the above-mentioned pose motion representation feature, so that the two-dimensional plane image can represent the state of the finally generated face in the two-dimensional image space; and the present application does not limit the implementation manner of the two-dimensional plane image. For example, the two-dimensional plane image can be implemented using Figure 4 the prediction image 1 shown. The depth map is used to describe the three-dimensional information corresponding to the two-dimensional plane image; and the present application does not limit the implementation manner of the depth map. For example, the depth map can be implemented using Figure 4 the depth map of the prediction image 1 shown. The background segmentation map is used to describe the position of the background in the two-dimensional plane image; and the present application does not limit the implementation manner of the background segmentation map. For example, the background segmentation map can be implemented using Figure 4Implement the background segmentation map of the predicted image 1 shown. For another example, the background segmentation map can be implemented using the foreground mask map (mask) of the two-dimensional plane image.
[0157] Scenario 2, when the image generation method provided in this application is applied to a certain image generation scenario, such as Figure 2 the holographic video generation scenario shown or Figure 5 the binocular image generation scenario shown, if the image generation method is applied to a VR device, the generated image above can include a three-dimensional stereoscopic image, a two-dimensional plane image, the depth map corresponding to the two-dimensional plane image, and the background segmentation map corresponding to the two-dimensional plane image. Among them, the three-dimensional stereoscopic image refers to a three-dimensional image generated based on the coding features of the original image above, the coding features of the reference face texture map above, and the pose motion representation features above, such as binocular images, so that the three-dimensional stereoscopic image can represent the state of the finally generated face in the three-dimensional space.
[0158] In addition, this application does not limit the implementation manner of the three-dimensional stereoscopic image above. For example, the three-dimensional stereoscopic image can include at least two channel images, and different channel images are used to describe the face state from different perspectives in the two-dimensional image space, so that each channel image belongs to a two-dimensional image. It can be seen that in a possible implementation manner, the three-dimensional stereoscopic image can include a left-eye channel image and a right-eye channel image. Among them, the left-eye channel image is used to describe the face state from the perspective of the left eye in the two-dimensional image space. The right-eye channel image is used to describe the face state from the perspective of the right eye in the two-dimensional image space.
[0159] In addition, for the three-dimensional stereoscopic image above, the three-dimensional stereoscopic image is determined based on the two-dimensional plane image above and the depth map corresponding to the two-dimensional plane image, so that the three-dimensional stereoscopic image carries the face state description information in the two-dimensional plane image, and the three-dimensional stereoscopic image carries the three-dimensional information described by the depth map.
[0160] Furthermore, this application does not limit the determination process of the three-dimensional stereoscopic image above. For the convenience of understanding, two examples are described below.
[0161] Example 1, in some application scenarios, it may be necessary to retain the background described in the two-dimensional plane image above when constructing the three-dimensional stereoscopic image above. Based on this, it can be known that when the three-dimensional stereoscopic image includes a first left-eye image and a first right-eye image, the determination process of the three-dimensional stereoscopic image is: performing a stereoscopic transposition process on the two-dimensional plane image using the depth map corresponding to the two-dimensional plane image to obtain the first left-eye image and the first right-eye image. Among them, the first left-eye image refers to the left-eye channel image carrying the background, such as Figure 5The left-eye image shown. This first right-eye image refers to the right-eye channel image with the background, such as Figure 5 The right-eye image shown. It should be noted that the present application does not limit the implementation manner of this stereo transposition processing. For example, it can be implemented by any existing or future method that can perform stereo transposition processing on a two-dimensional image using a depth map to obtain a binocular image.
[0162] Example 2, in some application scenarios, in order to better avoid the interference caused by the background, the present application also provides a possible implementation manner of the above-mentioned determination process of the three-dimensional stereo image. In this implementation manner, when the three-dimensional stereo image includes a second left-eye image and a second right-eye image, the determination process of the three-dimensional stereo image may specifically include the following steps 51-step 53.
[0163] Step 51: Perform stereo transposition processing on the two-dimensional plane image using the depth map corresponding to the above two-dimensional plane image to obtain a first left-eye image and a first right-eye image.
[0164] Step 52: Use the background segmentation map corresponding to the above two-dimensional plane image to perform background removal processing on the first left-eye image to obtain a second left-eye image, so that the foreground described by the second left-eye image is consistent with the foreground described by the first left-eye image, but there is no background described by the first left-eye image in the second left-eye image.
[0165] Among them, the second left-eye image refers to the left-eye channel image without the background, so that the second left-eye image can describe the state of the foreground from the perspective of the left eye.
[0166] In addition, the present application does not limit the implementation manner of the above step 52. For example, it can be implemented by any existing or future method that can perform background removal processing based on the background segmentation map.
[0167] Step 53: Use the background segmentation map corresponding to the above two-dimensional plane image to perform background removal processing on the first right-eye image to obtain a second right-eye image, so that the foreground described by the second right-eye image is consistent with the foreground described by the first right-eye image, but there is no background described by the first right-eye image in the second right-eye image.
[0168] Among them, the second right-eye image refers to the right-eye channel image without the background, so that the second right-eye image can describe the state of the foreground from the perspective of the right eye.
[0169] In addition, the present application does not limit the implementation manner of the above step 53. For example, the implementation manner of this step 53 is similar to the implementation manner of the above step 52.
[0170] In addition, this application does not limit the correlation between the execution time of step 53 above and the execution time of step 52 above. For example, the two can be the same. Another example is that the former is earlier than the latter. Still another example is that the former is later than the latter.
[0171] Based on the relevant content of steps 51 to 53 above, after generating a two-dimensional planar image, the depth map corresponding to the two-dimensional planar image, and the background segmentation map corresponding to the two-dimensional planar image, a three-dimensional stereoscopic image without a background can be generated based on the two-dimensional planar image, the depth map, and the background segmentation map, so that the three-dimensional stereoscopic image can be adapted to various backgrounds, so that the subsequent VR device can configure any background for the three-dimensional stereoscopic image, which can effectively avoid the interference caused by the automatically generated background during the image generation process, thus facilitating the improvement of the usage range of the three-dimensional stereoscopic image.
[0172] In addition, this application does not limit the acquisition process of the above-generated image. For example, it can specifically include steps 61 - 63 below.
[0173] Step 61: Process the encoded feature of the original image according to the above-mentioned pose motion representation feature to obtain a processed feature.
[0174] Among them, the encoded feature of the original image refers to the feature obtained by encoding the original image, so that the encoded feature can represent the image information carried by the original image.
[0175] In addition, this application does not limit the acquisition method of the encoded feature of the above-mentioned original image. For example, it can be implemented by using any existing or future method for encoding an image, such as the method of using an encoder, etc.
[0176] The processed feature refers to the feature obtained by processing the encoded feature of the original image according to the above-mentioned pose motion representation feature, so that the processed feature carries the image information described by the encoded feature and the pose motion information carried by the pose motion representation feature, so that the processed feature can represent the result obtained by performing pose adjustment on the encoded feature according to the pose motion representation feature, and further make the facial pose described by the processed feature as close as possible to the facial pose described by the above-mentioned reference three-dimensional pose representation data. It should be noted that this application does not limit the implementation manner of the step of "performing pose adjustment on the encoded feature according to the pose motion representation feature". For example, specifically, it can be: deforming the encoded feature of the original image using the pose motion representation feature to obtain the processed feature.
[0177] Step 62: Stitch the above-mentioned processed feature and the encoded feature of the above-mentioned reference facial texture map to obtain a stitched feature.
[0178] Among them, the encoded feature of the reference facial texture map refers to the feature obtained by encoding the reference facial texture map, so that the encoded feature can represent the image information carried by the reference facial texture map.
[0179] In addition, the acquisition method of the encoded feature of the above-mentioned reference facial texture map is similar to the acquisition method of the encoded feature of the above-mentioned original image.
[0180] The splicing feature refers to the result obtained by splicing the above-mentioned processed feature and the encoded feature of the above-mentioned reference facial texture map.
[0181] Step 63: Obtain a generated image according to the above-mentioned splicing feature.
[0182] It should be noted that the present application does not limit the implementation manner of the above-mentioned step 63. For the sake of easy understanding, the following will be described in conjunction with two scenarios.
[0183] Scenario 1, when the image generation method provided by the present application is applied to a model training scenario, such as Figure 4 the model training scenario shown, if the above-mentioned generated image includes a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image, then the above-mentioned step 63 may specifically be: performing a decoding process on the above-mentioned splicing feature to obtain the two-dimensional planar image, the depth map, and the background segmentation map. It should be noted that the present application does not limit the implementation manner of the decoding process. For example, it may adopt any existing or future decoding method, such as Figure 4 the decoder shown, for implementation.
[0184] Scenario 2, when the image generation method provided by the present application is applied to a certain image generation scenario, such as Figure 2 the holographic video generation scenario shown or Figure 5 the binocular image generation scenario shown, if the image generation method is applied to a VR device, and the above-mentioned generated image includes a three-dimensional stereoscopic image, a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image, then the above-mentioned step 63 may specifically be: first performing a decoding process on the above-mentioned splicing feature to obtain the two-dimensional planar image, the depth map, and the background segmentation map; then generating the three-dimensional stereoscopic image according to the two-dimensional planar image and the depth map.
[0185] Based on the relevant content of steps 61 to 63 above, for some application scenarios, after obtaining the encoded features of the original image above, the encoded features of the reference facial texture map above, and the pose motion representation features above, first deform the encoded features of the original image using the pose motion representation features to obtain processed features, so that the facial pose described by the processed features is as close as possible to the facial pose described by the reference 3D pose representation data above, such as the facial pose described by the processed features being consistent with the facial pose described by the reference 3D pose representation data above; then splice the processed features with the encoded features of the reference facial texture map to obtain spliced features; finally, determine the generated image according to the decoding result of the spliced features, so that the facial pose described by the generated image is consistent with the facial pose described by the reference 3D pose representation data above, and the facial texture such as the facial expression described by the generated image is consistent with the corresponding facial texture described by the reference facial texture map above.
[0186] Based on the relevant content of S1 to S3 above, for the image generation method provided in the embodiments of the present application, first obtain an original image, original 3D pose representation data, reference 3D pose representation data, and a reference facial texture map. The original 3D pose representation data is determined according to the original image; both the reference 3D pose representation data and the reference facial texture map are determined according to the reference information corresponding to the original image; then, according to the original image, the original 3D pose representation data, and the reference 3D pose representation data, determine pose motion representation features, so that the pose motion representation features can represent the pose motion information carried by the reference 3D pose representation data, so that the pose motion representation features can represent the pose motion information required when changing from the pose represented by the original 3D pose representation data to the pose represented by the reference 3D pose representation data; then, according to the encoded features of the original image, the encoded features of the reference facial texture map, and the pose motion representation features, obtain a generated image, so that the generated image not only carries the facial state information described by the original image and the reference facial texture map, but also carries the 3D space information described by the original 3D pose representation data and the reference 3D pose representation data, so that the generated image can describe a three-dimensional facial state, and further the generated image can better describe the facial state, which is beneficial to improving the image generation effect.
[0187] To better understand the image generation method provided in the present application, the following will be described in combination with two scenarios.
[0188] Scenario 1, when the image generation method provided in the present application is applied to a model training scenario, such as Figure 4When the model training scenario shown is considered, the model construction process provided by this application may include the following steps 71 - step 74.
[0189] Step 71: Obtain the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map.
[0190] It should be noted that for the relevant content of step 71, please refer to S1 above. For the sake of brevity, it will not be elaborated here.
[0191] Step 72: Determine the pose motion representation features based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data.
[0192] It should be noted that for the relevant content of step 72, please refer to S2 above. For the sake of brevity, it will not be elaborated here.
[0193] In addition, in some application scenarios, step 72 above may be implemented by a certain module in the image generation model. Based on this, it can be known that when the generated image above is determined by the image generation model based on the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map, and the image generation model includes a motion estimation network, the motion estimation network can be used to implement step 72. Based on this, it can be known that in a possible implementation manner, the motion estimation network can be used to determine the pose motion representation features based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data.
[0194] In addition, this application does not limit the working principle of the motion estimation network above. For example, as Figure 4 shown, the working principle of the motion estimation network can specifically be: after splicing the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data to obtain a splicing result, the motion estimation network performs motion estimation processing on the splicing result to obtain the pose motion representation features.
[0195] Based on the above content, it can be known that in a possible implementation manner, step 72 above may specifically include: first splicing the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data to obtain a splicing result; then the motion estimation network in the image generation model performs motion estimation processing on the splicing result to obtain the pose motion representation features.
[0196] Step 73: Obtain a generated image based on the encoded features of the original image, the encoded features of the reference face texture map, and the pose motion representation features. The generated image includes a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image.
[0197] It should be noted that for the relevant content of step 73, please refer to S3 above. For the sake of brevity, it will not be elaborated here.
[0198] In addition, in some application scenarios, step 73 above can be implemented by a certain module in the image generation model. Based on this, it can be known that when the generated image above is determined by the image generation model according to the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map, and the image generation model includes the motion estimation network, the encoder, and the decoder above, the encoder can be used to determine the encoded features of the original image above and the encoded features of the reference face texture map above, and the decoder is used to decode the above-mentioned splicing features to obtain a two-dimensional plane image, the depth map corresponding to the two-dimensional plane image, and the background segmentation map corresponding to the two-dimensional plane image. Among them, for the relevant content of the splicing features, please refer to the above.
[0199] Based on the above paragraph, it can be known that in a possible implementation manner, step 73 above can specifically include: after the encoder in the above-mentioned image generation model determines the encoded features of the original image above and the encoded features of the reference face texture map above, and the motion estimation network in the image generation model outputs the pose motion representation features, the pose adjustment process can be first performed on the encoded features of the original image according to the pose motion representation features to obtain the processed features; then, the processed features and the encoded features of the reference face texture map are spliced to obtain the splicing features; finally, the decoder in the image generation model decodes the splicing features to obtain a two-dimensional plane image, the depth map corresponding to the two-dimensional plane image, and the background segmentation map corresponding to the two-dimensional plane image.
[0200] Step 74: Use the generated image above and the supervision information corresponding to the generated image to update the image generation model, and return to continue to execute step 71 and its subsequent steps above until a preset stop condition is reached.
[0201] Among them, the supervision information corresponding to the generated image refers to the guiding information preset for the generated image. For example, the image generation guiding information configured in advance for the original image above.
[0202] In addition, the present application does not limit the implementation manner of the supervision information corresponding to the generated image above. For example, when the generated image includes a two-dimensional plane image, the depth map corresponding to the two-dimensional plane image, and the background segmentation map corresponding to the two-dimensional plane image, the supervision information corresponding to the generated image includes the supervision information corresponding to the two-dimensional plane image, the supervision information corresponding to the depth map, and the supervision information corresponding to the background segmentation map.
[0203] For the above two-dimensional planar image, the supervision information corresponding to the two-dimensional planar image refers to the guiding information set in advance for the two-dimensional planar image, such as the two-dimensional image generation guiding information configured in advance for the above original image, etc., so that the supervision information can guide the generation process of the two-dimensional planar image. In addition, the present application does not limit the implementation manner of the supervision information corresponding to the two-dimensional planar image. For example, the supervision information corresponding to the two-dimensional planar image can adopt the reference image corresponding to the original image, such as Figure 4 the reference image shown, etc., for implementation.
[0204] Based on the above content, it can be known that when the image generation method provided by the present application is applied to a model training scenario, such as Figure 4 the model training scenario shown, in a possible implementation manner, the reference information corresponding to the above original image can be the supervision information corresponding to the two-dimensional planar image, so that the content of "reference information" that appears in each step that needs to be performed on the reference information in the model training scenario provided by the present application can be replaced with the content of "supervision information corresponding to the two-dimensional planar image", so as to complete some processing processes for the reference information by means of the supervision information corresponding to the two-dimensional planar image.
[0205] For the depth map corresponding to the above two-dimensional planar image, the supervision information corresponding to the depth map refers to the guiding information set in advance for the depth map, such as the depth map generation guiding information configured in advance for the above original image, etc., so that the supervision information can guide the generation process of the depth map. In addition, the present application does not limit the implementation manner of the supervision information corresponding to the depth map. For example, the supervision information corresponding to the depth map can refer to the depth map provided manually. Another example is that the supervision information corresponding to the depth map can be obtained by performing depth map generation processing on the supervision information corresponding to the two-dimensional planar image, so that the supervision information corresponding to the depth map can be automatically constructed, thereby effectively reducing the difficulty of obtaining training data. It should be noted that the present application does not limit the implementation manner of the depth map generation processing. For example, it can adopt any existing or future method capable of performing depth map generation processing on an image, such as by means of a pre-constructed depth map generation model, etc., for implementation. Among them, the depth map generation model is used to perform depth map generation processing on the input data of the depth map generation model.
[0206] For the background segmentation map corresponding to the above two-dimensional planar image, the supervision information corresponding to the background segmentation map refers to the guiding information preset for the background segmentation map, such as the guiding information for generating the background segmentation map pre-configured for the above original image, etc., so that the supervision information can guide the generation process of the background segmentation map. In addition, the present application does not limit the implementation manner of the supervision information corresponding to the background segmentation map. For example, the supervision information corresponding to the background segmentation map may refer to a manually provided background segmentation map. Another example is that the supervision information corresponding to the background segmentation map can be obtained by performing background segmentation processing on the supervision information corresponding to the two-dimensional planar image, so that the supervision information corresponding to the background segmentation map can be automatically constructed, thereby effectively reducing the difficulty of obtaining training data. It should be noted that the present application does not limit the implementation manner of the background segmentation processing. For example, it can adopt any existing or future method capable of performing background segmentation processing on an image, such as methods using a pre-constructed background segmentation model, etc., for implementation. Among them, the background segmentation model is used to perform background segmentation processing on the input data of the background segmentation model.
[0207] The preset stop condition refers to the condition required when the model training stops; and the present application does not limit the implementation manner of the preset stop condition. For example, the preset stop condition may include: the model loss of the above image generation model is lower than a preset loss threshold. Another example is that the preset stop condition may include: the change rate of the model loss of the image generation model is lower than a preset loss change rate threshold. Still another example is that the preset stop condition may include: the number of updates of the image generation model reaches a preset number threshold. Among them, the model loss of the image generation model is used to represent the performance of the image generation model; and the model loss of the image generation model is determined based on the above generated image and the supervision information corresponding to the generated image. For example, the model loss is determined based on the difference representation data between the two-dimensional planar image and the supervision information corresponding to the two-dimensional planar image, the difference representation data between the depth map corresponding to the two-dimensional planar image and the supervision information corresponding to the depth map, and the difference representation data between the background segmentation map corresponding to the two-dimensional planar image and the supervision information corresponding to the background segmentation map. It should be noted that the present application does not limit the calculation method of the model loss; and the present application also does not limit the implementation manner of the difference representation data. For example, the difference representation data can be determined using any distance calculation formula.
[0208] Based on the relevant content of the above steps 71 to 74, in some application scenarios, the training process of the above image generation model can be as follows: First, randomly select two video images from the sample video as the original image and the reference information corresponding to the original image respectively; then determine the original three-dimensional pose representation data based on the original image, and determine the reference three-dimensional pose representation data and the reference facial texture map based on the reference information; then the image generation model performs image generation processing based on the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference facial texture map to obtain a generated image, such as Figure 4 the predicted image 1 shown, the background segmentation map of the predicted image 1, and the depth map of the predicted image 1; then, according to the difference between the generated image and the supervision information corresponding to the generated image, update the image generation model so that the updated image generation model has a better image generation effect, so as to perform the next round of training process based on the updated image generation model subsequently, and iterate in this way until the preset stop condition is reached.
[0209] In addition, the execution subject of the training process of the above image generation model is not limited in this application. For example, in some application scenarios, the execution subject of the training process of the image generation model is a VR device. Another example is that in some application scenarios, in order to better improve the model training effect, the execution subject of the training process of the image generation model is a server, so that the VR device can use the image generation model sent by the server to complete the image generation task subsequently, such as Figure 5 the image generation task shown.
[0210] Scenario 2, when the image generation method provided in this application is applied to a certain image generation scenario, such as Figure 2 the holographic video generation scenario shown or Figure 5 the binocular image generation scenario shown, if the image generation method is applied to a VR device, the working principle of the VR device may include some or all of the following steps 81 - step 84.
[0211] Step 81: The VR device acquires the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference facial texture map.
[0212] It should be noted that for the relevant content of step 81, please refer to S1 above. For the sake of brevity, it will not be elaborated here.
[0213] Step 82: The VR device determines the pose motion representation feature based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data.
[0214] It should be noted that for the relevant content of step 82, please refer to S2 above. For the sake of brevity, it will not be elaborated here. For ease of understanding, it will be described below with examples.
[0215] As an example, in a possible implementation, step 82 above may specifically be: first, the VR device splices the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data to obtain a splicing result; then, the motion estimation network in the image generation model deployed on the VR device performs motion estimation processing on the splicing result to obtain pose motion representation features.
[0216] Step 83: The VR device obtains a generated image based on the encoding features of the original image, the encoding features of the reference face texture map, and the pose motion representation features. The generated image includes a three-dimensional stereoscopic image, a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image.
[0217] It should be noted that for the relevant content of step 83, please refer to S3 above. For the sake of brevity, it will not be elaborated here. For ease of understanding, it will be described below with examples.
[0218] As an example, in a possible implementation, when the image generation model deployed on the above VR device includes the above motion estimation network, encoder, and decoder, step 83 above may specifically be: after the encoder in the image generation model determines the encoding features of the above original image and the encoding features of the above reference face texture map, and the motion estimation network in the image generation model outputs the pose motion representation features, the VR device may first perform pose adjustment processing on the encoding features of the original image using the pose motion representation features to obtain processed features; then, the VR device splices the processed features and the encoding features of the reference face texture map to obtain splicing features; then, the decoder in the image generation model performs decoding processing on the splicing features to obtain a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image, such as Figure 5 the predicted image 2 shown, the background segmentation map of the predicted image 2, and the depth map of the predicted image 2; finally, the VR device generates a three-dimensional stereoscopic image based on the two-dimensional planar image and the depth map, such as Figure 5 the left-eye image and the right-eye image shown.
[0219] Step 84: The VR device displays the above three-dimensional stereoscopic image.
[0220] Based on the relevant content of steps 81 to 84 above, it can be seen that in some application scenarios, for the VR device above, after the VR device obtains the facial expression coefficient and head posture for the target object, the VR device can first perform image generation processing based on the facial expression coefficient, head posture, and relevant information of the original image obtained from the server to obtain a three-dimensional stereoscopic image, such as Figure 5 the left-eye image and right-eye image shown, etc.; then the VR device displays the three-dimensional stereoscopic image, so that the user of the VR device can see the three-dimensional stereoscopic image on the VR device. In this way, it can be realized that the VR device performs image generation processing based on the facial expression coefficient and head posture obtained in real time, so that the VR device can generate a holographic video.
[0221] In addition, in some application scenarios, such as video call scenarios, the VR device above also needs to send the generated image to other VR devices for display. Based on this, the present application also provides a possible implementation manner of the working principle of the VR device. In this implementation manner, when the VR device and the target device are in a video communication state, the target device is a three-dimensional image display device, and the generated image above includes a three-dimensional stereoscopic image, the working principle of the VR device can at least include step 85 below. Among them, the execution time of step 85 is later than the execution time of step 83 above.
[0222] Step 85: The VR device sends the three-dimensional stereoscopic image above to the target device, and the target device is used to display the three-dimensional stereoscopic image.
[0223] Among them, the target device refers to the device that conducts video communication with the VR device above. For example, when the VR device is Figure 2 the VR device 1 shown, the target device can be Figure 2 the VR device 2 shown.
[0224] In addition, for the target device above, the target device can be a three-dimensional image display device, so that the target device can be used to display a three-dimensional stereoscopic image. It should be noted that the present application does not limit the implementation manner of the three-dimensional image display device. For example, it can be implemented using any existing or future device capable of displaying a three-dimensional stereoscopic image, such as a VR device, etc.
[0225] Based on the relevant content of step 85 above, when the VR device above is in a video communication state with the target device, and the target device is a three-dimensional image display device, after the VR device performs image generation processing based on the face expression coefficient and head pose obtained in real time to obtain a three-dimensional stereoscopic image, the VR device can send the three-dimensional stereoscopic image to the target device for display, so that the user of the target device can view the three-dimensional stereoscopic image sent by the VR device on the target device, thereby enabling the user of the target device to view the holographic video including multiple three-dimensional stereoscopic images sent by the VR device on the target device.
[0226] In addition, in some application scenarios, such as video call scenarios, the VR device above also needs to send the images it generates to other non-VR devices, such as Figure 2 the mobile phone shown. Based on this, the present application also provides a possible implementation manner of the working principle of the VR device. In this implementation manner, when the VR device is in a video communication state with the target device, the target device is a two-dimensional image display device, the generated images above include two-dimensional planar images and three-dimensional stereoscopic images, and the three-dimensional stereoscopic image includes at least two channel images, the working principle of the VR device can at least include step 86 below. Among them, the execution time of step 86 is later than the execution time of step 83 above.
[0227] Step 86: The VR device sends either the two-dimensional planar image above or any one of the channel images in the three-dimensional stereoscopic image above to the target device, and the target device is used to display the two-dimensional planar image or the channel image.
[0228] Among them, the two-dimensional image display device is used to display two-dimensional images; and the present application does not limit the implementation manner of the two-dimensional image display device. For example, it can adopt any existing or future device capable of displaying two-dimensional images, such as Figure 2 the mobile phone shown and other terminal devices for implementation.
[0229] Based on the relevant content of step 86 above, when the VR device above is in a video communication state with the target device, and the target device is a two-dimensional image display device, and the generated image above includes a two-dimensional planar image and a three-dimensional stereoscopic image, and the three-dimensional stereoscopic image includes a left-eye channel image and a right-eye channel image, after the VR device performs image generation processing based on the face expression coefficient and head pose obtained in real time to obtain the two-dimensional planar image, the left-eye channel image, and the right-eye channel image above, the VR device can send the two-dimensional planar image, the left-eye channel image, or the right-eye channel image to the target device for display, so that the user of the target device can view the two-dimensional image sent by the VR device on the target device, thereby enabling the user of the target device to view a video including multiple two-dimensional images sent by the VR device on the target device.
[0230] Based on the image generation method provided in the embodiments of the present application, the embodiments of the present application further provide an image generation device, which will be explained and described below in conjunction with Figure 6 This is for explanation and illustration. Among them, Figure 6 This is a schematic structural diagram of an image generation device provided in the embodiments of the present application. It should be noted that for the technical details of the image generation device provided in the embodiments of the present application, please refer to the relevant content of the image generation method above.
[0231] As Figure 6 shown, the image generation device 600 provided in the embodiments of the present application includes:
[0232] An acquisition unit 601, configured to acquire an original image, original three-dimensional pose characterization data, reference three-dimensional pose characterization data, and a reference face texture map, where the original three-dimensional pose characterization data is determined based on the original image; the reference three-dimensional pose characterization data and the reference face texture map are determined based on the reference information corresponding to the original image;
[0233] A determination unit 602, configured to determine a pose motion characterization feature based on the original image, the original three-dimensional pose characterization data, and the reference three-dimensional pose characterization data;
[0234] A generation unit 603, configured to obtain a generated image based on the encoded feature of the original image, the encoded feature of the reference face texture map, and the pose motion characterization feature.
[0235] In a possible implementation manner, the generated image is determined by an image generation model based on the original image, the original three-dimensional pose characterization data, the reference three-dimensional pose characterization data, and the reference face texture map;
[0236] The image generation device 600 further includes:
[0237] An updating unit, configured to update the image generation model by using the generated image and the supervision information corresponding to the generated image.
[0238] In a possible implementation manner, the generated image includes a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image; the supervision information corresponding to the generated image includes the supervision information corresponding to the two-dimensional planar image, the supervision information corresponding to the depth map, and the supervision information corresponding to the background segmentation map.
[0239] In a possible implementation manner, the reference information includes the supervision information corresponding to the two-dimensional planar image.
[0240] In a possible implementation manner, the supervision information corresponding to the depth map is obtained by performing depth map generation processing on the supervision information corresponding to the two-dimensional planar image; the supervision information corresponding to the background segmentation map is obtained by performing background segmentation processing on the supervision information corresponding to the two-dimensional planar image.
[0241] In a possible implementation manner, the generating unit 603 includes:
[0242] A feature processing subunit, configured to process the encoded feature of the original image according to the pose motion representation feature to obtain a processed feature;
[0243] A feature splicing subunit, configured to splice the processed feature and the encoded feature of the reference face texture map to obtain a spliced feature;
[0244] A first determination subunit, configured to obtain the generated image according to the spliced feature.
[0245] In a possible implementation manner, the generated image includes a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image;
[0246] The first determination subunit is specifically configured to: perform decoding processing on the spliced feature to obtain the two-dimensional planar image, the depth map, and the background segmentation map.
[0247] In a possible implementation manner, the generated image is determined by an image generation model based on the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map; the image generation model includes a motion estimation network, an encoder, and a decoder; the motion estimation network is used to determine a pose motion representation feature based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data; the encoder is used to determine an encoded feature of the original image and an encoded feature of the reference face texture map; the decoder is used to perform decoding processing on the spliced feature to obtain the two-dimensional planar image, the depth map, and the background segmentation map.
[0248] In a possible implementation manner, the generated image includes a three-dimensional stereoscopic image, a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image;
[0249] The first determination sub-unit includes:
[0250] A decoding processing sub-unit, configured to perform decoding processing on the spliced feature to obtain the two-dimensional planar image, the depth map, and the background segmentation map;
[0251] An image generation sub-unit, configured to generate the three-dimensional stereoscopic image based on the two-dimensional planar image and the depth map.
[0252] In a possible implementation manner, the three-dimensional stereoscopic image includes a first left-eye image and a first right-eye image; both the first left-eye image and the first right-eye image are obtained by performing a stereoscopic transposition process on the two-dimensional planar image by using the depth map.
[0253] In a possible implementation manner, the three-dimensional stereoscopic image includes a second left-eye image and a second right-eye image;
[0254] The image generation sub-unit is specifically configured to: perform a stereoscopic transposition process on the two-dimensional planar image by using the depth map to obtain a first left-eye image and a first right-eye image; perform background removal processing on the first left-eye image by using the background segmentation map to obtain the second left-eye image; perform background removal processing on the first right-eye image by using the background segmentation map to obtain the second right-eye image.
[0255] In a possible implementation manner, the determination unit 602 is specifically configured to: splice the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data to obtain a splicing result; perform motion estimation processing on the splicing result to obtain the pose motion representation feature.
[0256] In a possible implementation manner, the original three-dimensional pose characterization data is obtained by mapping the three-dimensional face mesh corresponding to the original image into a two-dimensional image space; the three-dimensional face mesh corresponding to the original image is constructed based on the original image; the reference three-dimensional pose characterization data is obtained by mapping the three-dimensional face mesh corresponding to the reference information into a two-dimensional image space; the three-dimensional face mesh corresponding to the reference information is constructed based on the reference information.
[0257] In a possible implementation manner, the original three-dimensional pose characterization data is obtained by performing feature extraction processing on the three-dimensional face mesh corresponding to the original image; the three-dimensional face mesh corresponding to the original image is constructed based on the original image; the reference three-dimensional pose characterization data is obtained by performing feature extraction processing on the three-dimensional face mesh corresponding to the reference information; the three-dimensional face mesh corresponding to the reference information is constructed based on the reference information.
[0258] In a possible implementation manner, the generated image includes a two-dimensional planar image; the reference information is the supervision information corresponding to the two-dimensional planar image; the three-dimensional face mesh corresponding to the reference information is obtained by performing three-dimensional face reconstruction processing on the supervision information corresponding to the two-dimensional planar image.
[0259] In a possible implementation manner, the process of determining the three-dimensional face mesh corresponding to the reference information includes: performing three-dimensional face reconstruction processing on the supervision information corresponding to the two-dimensional planar image to obtain a first face mesh; performing expression adjustment processing on the first face mesh according to a preset non-expression parameter to obtain the three-dimensional face mesh corresponding to the reference information.
[0260] In a possible implementation manner, the reference information includes a head pose; the three-dimensional face mesh corresponding to the reference information is obtained by performing pose adjustment processing on the three-dimensional face mesh corresponding to the original image by using the head pose.
[0261] In a possible implementation manner, the process of determining the three-dimensional face mesh corresponding to the original image includes: performing three-dimensional face reconstruction processing on the original image to obtain a second face mesh; performing expression adjustment processing on the second face mesh according to a preset non-expression parameter to obtain the three-dimensional face mesh corresponding to the original image.
[0262] In a possible implementation manner, the generated image includes a two-dimensional planar image; the reference information is the supervision information corresponding to the two-dimensional planar image; the reference face texture map refers to the face texture map obtained by performing three-dimensional face reconstruction processing on the supervision information corresponding to the two-dimensional planar image.
[0263] In a possible implementation manner, the reference information includes a facial expression coefficient; the reference facial texture map is obtained by performing an expression adjustment process on the facial texture map corresponding to the original image by using the facial expression coefficient; the facial texture map corresponding to the original image is constructed based on the original image.
[0264] In a possible implementation manner, the image generation device 600 is deployed on a virtual reality (VR) device.
[0265] In a possible implementation manner, the reference information includes a facial expression coefficient and a head pose; the generated image includes a three-dimensional (3D) stereoscopic image, a two-dimensional (2D) planar image, a depth map corresponding to the 2D planar image, and a background segmentation map corresponding to the 2D planar image.
[0266] In a possible implementation manner, the generated image includes a 3D stereoscopic image; the VR device and a target device are in a video communication state, and the target device is a 3D image display device.
[0267] The image generation device 600 further includes:
[0268] A sending unit, configured to send the 3D stereoscopic image to the target device, and the target device is configured to display the 3D stereoscopic image.
[0269] In a possible implementation manner, the generated image includes a 2D planar image and a 3D stereoscopic image, and the 3D stereoscopic image includes at least two channel images; the VR device and a target device are in a video communication state, and the target device is a 2D image display device.
[0270] The image generation device 600 further includes:
[0271] A sending unit, configured to send either the 2D planar image or any one of the channel images of the 3D stereoscopic image to the target device, and the target device is configured to display the 2D planar image or the any one of the channel images.
[0272] In a possible implementation manner, the reference information corresponding to the original image includes a facial expression coefficient and a head pose acquired by the VR device for a target object.
[0273] The obtaining unit 601 is specifically configured to: receive the original image, the three-dimensional face mesh corresponding to the original image, and the face texture map corresponding to the original image sent by the server; obtain the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map according to the three-dimensional face mesh corresponding to the original image, the face texture map corresponding to the original image, the face expression coefficient, and the head pose.
[0274] In a possible implementation manner, the server is configured to generate the original image according to the face representation image and the style description information specified by the target object.
[0275] Based on the related content of the above image generation device 600, for the image generation device 600 provided in the embodiment of the present application, first obtain the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map. The original three-dimensional pose representation data is determined according to the original image; the reference three-dimensional pose representation data and the reference face texture map are both determined according to the reference information corresponding to the original image; then, determine the pose motion representation feature according to the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data, so that the pose motion representation feature can represent the pose motion information carried by the reference three-dimensional pose representation data, so that the pose motion representation feature can represent the pose motion information required when changing from the pose represented by the original three-dimensional pose representation data to the pose represented by the reference three-dimensional pose representation data; then, obtain the generated image according to the encoding feature of the original image, the encoding feature of the reference face texture map, and the pose motion representation feature, so that the generated image not only carries the face state information described by the original image and the reference face texture map, but also carries the three-dimensional space information described by the original three-dimensional pose representation data and the reference three-dimensional pose representation data, so that the generated image can describe the three-dimensional face state, and further make the generated image better describe the face state, which is beneficial to improving the image generation effect.
[0276] In addition, the embodiment of the present application further provides an electronic device, where the device includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation manner of the image generation method provided in the embodiment of the present application.
[0277] See Figure 7, which shows a schematic structural diagram of an electronic device 700 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The shown electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0278] As Figure 7 shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage device 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0279] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 the shown electronic device 700 has various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0280] Particularly, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.
[0281] The electronic device provided by the embodiments of the present disclosure and the method provided by the above embodiments belong to the same inventive concept. For technical details not described in detail in this embodiment, reference may be made to the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0282] An embodiment of the present application further provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs run on a device, the device is caused to execute any implementation manner of the image generation method provided by the embodiments of the present application.
[0283] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0284] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0285] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.
[0286] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the above method.
[0287] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0288] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0289] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the name of the unit / module does not, in some cases, constitute a limitation on the unit itself.
[0290] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0291] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash Memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0292] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0293] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one)" or its similar expression below refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0294] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0295] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0296] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image generation method, characterized in that, the method includes: obtaining an original image, original three-dimensional pose representation data, reference three-dimensional pose representation data, and a reference face texture map, where the original three-dimensional pose representation data is determined based on the original image; the reference three-dimensional pose representation data and the reference face texture map are determined based on reference information corresponding to the original image; determining pose motion representation features based on the original image, the original three-dimensional pose representation data, and the reference three-dimensional pose representation data; obtaining a generated image based on the encoded features of the original image, the encoded features of the reference face texture map, and the pose motion representation features.
2. The method according to claim 1, characterized in that, the generated image is determined by an image generation model based on the original image, the original three-dimensional pose representation data, the reference three-dimensional pose representation data, and the reference face texture map; after obtaining the generated image, the method further includes: updating the image generation model using the generated image and the supervision information corresponding to the generated image.
3. The method according to claim 2, characterized in that, the generated image includes a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image; the supervision information corresponding to the generated image includes the supervision information corresponding to the two-dimensional planar image, the supervision information corresponding to the depth map, and the supervision information corresponding to the background segmentation map.
4. The method according to claim 3, characterized in that, the reference information includes the supervision information corresponding to the two-dimensional planar image.
5. The method according to claim 3, characterized in that, the supervision information corresponding to the depth map is obtained by performing depth map generation processing on the supervision information corresponding to the two-dimensional planar image; the supervision information corresponding to the background segmentation map is obtained by performing background segmentation processing on the supervision information corresponding to the two-dimensional planar image.
6. The method according to claim 1, characterized in that, obtaining the generated image based on the encoded features of the original image, the encoded features of the reference face texture map, and the pose motion representation features includes: processing the encoded features of the original image according to the pose motion representation features to obtain processed features; concatenating the processed features and the encoded features of the reference face texture map to obtain concatenated features; obtaining the generated image based on the concatenated features.
7. The method according to claim 6, characterized in that, the generated image includes a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image; obtaining the generated image based on the concatenated features includes: performing decoding processing on the concatenated features to obtain the two-dimensional planar image, the depth map, and the background segmentation map.
8. The method according to claim 7, characterized in that, The generated image is determined by an image generation model based on the original image, the original three-dimensional pose characterization data, the reference three-dimensional pose characterization data, and the reference facial texture map; The image generation model includes a motion estimation network, an encoder, and a decoder; The motion estimation network is used to determine pose motion characterization features based on the original image, the original three-dimensional pose characterization data, and the reference three-dimensional pose characterization data; The encoder is used to determine the encoded features of the original image and the encoded features of the reference facial texture map; The decoder is used to perform decoding processing on the spliced features to obtain the two-dimensional planar image, the depth map, and the background segmentation map.
9. The method according to claim 6, wherein, The generated image includes a three-dimensional stereoscopic image, a two-dimensional planar image, a depth map corresponding to the two-dimensional planar image, and a background segmentation map corresponding to the two-dimensional planar image; The obtaining of the generated image based on the spliced features includes: Performing decoding processing on the spliced features to obtain the two-dimensional planar image, the depth map, and the background segmentation map; Generating the three-dimensional stereoscopic image based on the two-dimensional planar image and the depth map.
10. The method according to claim 9, wherein, The three-dimensional stereoscopic image includes a first left-eye image and a first right-eye image; Both the first left-eye image and the first right-eye image are obtained by performing stereoscopic transposition processing on the two-dimensional planar image using the depth map.
11. The method according to claim 9, wherein, The three-dimensional stereoscopic image includes a second left-eye image and a second right-eye image; The generating of the three-dimensional stereoscopic image based on the two-dimensional planar image and the depth map includes: Performing stereoscopic transposition processing on the two-dimensional planar image using the depth map to obtain a first left-eye image and a first right-eye image; Performing background removal processing on the first left-eye image using the background segmentation map to obtain the second left-eye image; Performing background removal processing on the first right-eye image using the background segmentation map to obtain the second right-eye image.
12. The method according to claim 1, wherein, The process of determining the pose motion characterization features includes: Splicing the original image, the original three-dimensional pose characterization data, and the reference three-dimensional pose characterization data to obtain a splicing result; Performing motion estimation processing on the splicing result to obtain the pose motion characterization features.
13. The method according to claim 1, wherein, The original three-dimensional pose characterization data is obtained by mapping the three-dimensional facial mesh corresponding to the original image into the two-dimensional image space; the three-dimensional facial mesh corresponding to the original image is constructed based on the original image; The reference three-dimensional pose characterization data is obtained by mapping the three-dimensional facial mesh corresponding to the reference information into the two-dimensional image space; the three-dimensional facial mesh corresponding to the reference information is constructed based on the reference information.
14. The method according to claim 1, wherein, The original 3D pose representation data is obtained by performing feature extraction processing on the 3D face mesh corresponding to the original image; the 3D face mesh corresponding to the original image is constructed based on the original image; The reference 3D pose representation data is obtained by performing feature extraction processing on the 3D face mesh corresponding to the reference information; the 3D face mesh corresponding to the reference information is constructed based on the reference information.
15. The method according to claim 13 or 14, wherein, the generated image includes a 2D planar image; the reference information is the supervision information corresponding to the 2D planar image; the 3D face mesh corresponding to the reference information is obtained by performing 3D face reconstruction processing on the supervision information corresponding to the 2D planar image.
16. The method according to claim 15, wherein, the process of determining the 3D face mesh corresponding to the reference information includes: performing 3D face reconstruction processing on the supervision information corresponding to the 2D planar image to obtain a first face mesh; performing expression adjustment processing on the first face mesh according to a preset non-expression parameter to obtain the 3D face mesh corresponding to the reference information.
17. The method according to claim 13 or 14, wherein, the reference information includes a head pose; the 3D face mesh corresponding to the reference information is obtained by performing pose adjustment processing on the 3D face mesh corresponding to the original image by using the head pose.
18. The method according to claim 13 or 14, wherein, the process of determining the 3D face mesh corresponding to the original image includes: performing 3D face reconstruction processing on the original image to obtain a second face mesh; performing expression adjustment processing on the second face mesh according to a preset non-expression parameter to obtain the 3D face mesh corresponding to the original image.
19. The method according to claim 1, wherein, the generated image includes a 2D planar image; the reference information is the supervision information corresponding to the 2D planar image; the reference face texture map refers to the face texture map obtained by performing 3D face reconstruction processing on the supervision information corresponding to the 2D planar image; or, the reference information includes a face expression coefficient; the reference face texture map is obtained by performing expression adjustment processing on the face texture map corresponding to the original image by using the face expression coefficient; the face texture map corresponding to the original image is constructed based on the original image.
20. The method according to claim 1, wherein, the method is applied to a virtual reality (VR) device.
21. The method according to claim 20, wherein, the reference information includes a face expression coefficient and a head pose; the generated image includes a 3D stereoscopic image, a 2D planar image, the depth map corresponding to the 2D planar image, and the background segmentation map corresponding to the 2D planar image.
22. The method according to claim 20, wherein, the generated image includes a 3D stereoscopic image; The VR device is in a video communication state with a target device, and the target device is a three-dimensional image display device; After obtaining the generated image, the method further includes: Sending the three-dimensional stereoscopic image to the target device, and the target device is used to display the three-dimensional stereoscopic image.
23. The method according to claim 20, wherein, The generated image includes a two-dimensional planar image and a three-dimensional stereoscopic image, and the three-dimensional stereoscopic image includes at least two channel images; The VR device is in a video communication state with a target device, and the target device is a two-dimensional image display device; After obtaining the generated image, the method further includes: Sending either the two-dimensional planar image or any one of the channel images in the three-dimensional stereoscopic image to the target device, and the target device is used to display the two-dimensional planar image or any one of the channel images.
24. The method according to claim 20, wherein, The reference information corresponding to the original image includes the facial expression coefficient and the head pose obtained by the VR device for the target object; The obtaining of the original image, the original three-dimensional pose characterization data, the reference three-dimensional pose characterization data, and the reference facial texture map includes: Receiving the original image, the three-dimensional facial mesh corresponding to the original image, and the facial texture map corresponding to the original image sent by the server; Based on the three-dimensional facial mesh corresponding to the original image, the facial texture map corresponding to the original image, the facial expression coefficient, and the head pose, obtaining the original three-dimensional pose characterization data, the reference three-dimensional pose characterization data, and the reference facial texture map.
25. The method according to claim 24, wherein, The server is used to generate the original image according to the facial representation image and the style description information specified by the target object.
26. An image generation device, wherein, including: An acquisition unit, configured to acquire an original image, original three-dimensional pose characterization data, reference three-dimensional pose characterization data, and a reference facial texture map, and the original three-dimensional pose characterization data is determined according to the original image; The reference three-dimensional pose characterization data and the reference facial texture map are determined according to the reference information corresponding to the original image; A determination unit, configured to determine pose motion characterization features according to the original image, the original three-dimensional pose characterization data, and the reference three-dimensional pose characterization data; A generation unit, configured to obtain a generated image according to the coding features of the original image, the coding features of the reference facial texture map, and the pose motion characterization features.
27. An electronic device, wherein, The device includes: a processor and a memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer programs in the memory, so that the electronic device executes the method according to any one of claims 1-25.
28. A computer-readable medium, wherein, The computer-readable medium stores instructions or a computer program which, when run on a device, cause the device to perform the method according to any one of claims 1-25.