Method, device, equipment and computer program product for generating virtual image
By extracting and establishing three-dimensional feature mapping relationships, posture correction and fitting are performed, the problem of low quality of virtual image images is solved, and a higher quality and fit virtual image images are achieved.
Patent Information
- Application Number
- CN202510100245.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to improve the quality of virtual image images, resulting in problems such as face marginal distortion, invisible expressions or movements, inconsistent facial features, insufficient resolution or insufficient image details.
By extracting three-dimensional features from the original image and the generated image, establishing a three-dimensional mapping relationship, performing posture correction, and applying the corrected virtual image to the original image to generate high-quality virtual image images.
The quality of virtual image images is improved, making it more suitable for the original image, reducing unnatural seams at the edges of virtual image, and improving the user experience.
Smart Images

Figure CN120107462A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of virtual image technology, and in particular to a method, device, equipment and computer program product for generating a virtual image. Background Art
[0002] With the development of avatar technology, more and more people are creating personalized avatars in media, games or social platforms to enhance interactivity and participation. In this process, it is necessary to generate a avatar based on the original image of the generated object, and then fit the avatar back to the original image to obtain the avatar image for playback.
[0003] Therefore, how to improve the quality of virtual image images and thus improve people's user experience has become a technical problem that needs to be urgently solved by those skilled in the art. Summary of the invention
[0004] In view of this, the present application proposes a method, apparatus, device and computer program product for generating a virtual avatar image, which method can improve the quality of the virtual avatar image.
[0005] The technical solutions proposed in this application are as follows:
[0006] In a first aspect, an embodiment of the present application provides a method for generating a virtual image, comprising:
[0007] Extracting a first three-dimensional feature from an original image of the generated object, extracting a second three-dimensional feature from the generated image, and establishing a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature; wherein the generated image is generated based on the original image, and the generated image includes a virtual image of the generated object;
[0008] Using the three-dimensional mapping relationship, performing posture correction on the virtual image in the generated image to obtain a virtual image after posture correction;
[0009] The virtual image after posture correction is fitted into the original image to obtain the virtual image image of the generated object.
[0010] Furthermore, in the above method, extracting the first three-dimensional feature from the original image of the generated object and extracting the second three-dimensional feature from the generated image include:
[0011] Mapping a first image feature extracted from an original image of the generated object into a three-dimensional space to obtain a first three-dimensional feature of the original image;
[0012] The second image feature extracted from the generated image is mapped into a three-dimensional space to obtain a second three-dimensional feature of the generated image.
[0013] Furthermore, in the above method, the establishing of the three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature includes:
[0014] A three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature is established according to the feature representation of each feature point in the first three-dimensional feature and the second three-dimensional feature.
[0015] Furthermore, in the above method, when the feature point includes a lip region feature point of the generated object, establishing a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature according to the feature representation of the same feature point in the first three-dimensional feature and the second three-dimensional feature includes:
[0016] A three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature in the lip region is established according to feature representation of the specific lip region feature point in the first three-dimensional feature and the second three-dimensional feature.
[0017] Furthermore, in the above method, using the three-dimensional mapping relationship to perform posture correction on the virtual image in the generated image to obtain the virtual image after posture correction includes:
[0018] According to the three-dimensional mapping relationship, feature transformation is performed on the second image feature extracted from the generated image to obtain a third image feature;
[0019] The third image feature is used to perform posture correction on the virtual image in the generated image to obtain a virtual image with corrected posture.
[0020] Furthermore, in the above method, using the third image feature to perform posture correction on the virtual image in the generated image to obtain the virtual image after posture correction includes:
[0021] fusing the third image feature, the first image feature extracted from the original image of the generated object, and the three-dimensional mapping relationship to obtain a fused feature;
[0022] Performing posture correction on the second image feature by using the fusion feature to obtain a virtual image feature;
[0023] A virtual image with posture correction is generated according to the virtual image features.
[0024] Furthermore, in the above method, generating a virtual image with posture correction according to the virtual image features includes:
[0025] Performing upsampling processing on the virtual image features to obtain upsampled features;
[0026] A virtual image whose pixels are consistent with those of the original image is generated according to the up-sampling features.
[0027] Furthermore, in the above method, the step of fitting the virtual image after posture correction to the original image to obtain the virtual image of the generated object includes:
[0028] Mask the image area of the generated object in the original image;
[0029] The virtual image after posture correction is fitted to the mask area to obtain the virtual image image of the generated object.
[0030] In a second aspect, an embodiment of the present application provides a device for generating a virtual image, comprising:
[0031] an extraction unit, configured to extract a first three-dimensional feature from an original image of a generated object, extract a second three-dimensional feature from a generated image, and establish a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature; wherein the generated image is generated based on the original image, and the generated image includes a virtual image of the generated object;
[0032] A correction unit, configured to perform posture correction on the virtual image in the generated image by using the three-dimensional mapping relationship to obtain a virtual image after posture correction;
[0033] The fitting unit is used to fit the virtual image after posture correction to the original image to obtain the virtual image image of the generated object.
[0034] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0035] A memory and a processor; wherein the memory is used to store programs; and the processor is used to implement any of the methods described above by running the programs in the memory.
[0036] In a fourth aspect, an embodiment of the present application provides a computer program product, the computer program product comprising a computer program, and when the computer program is executed by a processor, the computer program implements any one of the above methods. Optionally, the computer program can be stored in a readable storage medium of a computer device or in the cloud; the processor of the computer device reads the computer program from the readable storage medium or the cloud.
[0037] The method for generating a virtual image proposed in the present application can extract a first three-dimensional feature from an original image of a generated object, extract a second three-dimensional feature from a generated image, and establish a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature. The generated image is generated based on the original image, and the generated image contains a virtual image of the generated object. Then, the three-dimensional mapping relationship is used to perform posture correction on the virtual image in the generated image to obtain a posture-corrected virtual image, so as to fit the posture-corrected virtual image to the original image to obtain a virtual image of the generated object. In this way, the virtual image after posture correction using the three-dimensional mapping relationship has a higher consistency in posture with the generated object in the original image, and the virtual image is more closely aligned with the original image, which can reduce unnatural seams at the edges of the virtual image and improve the quality of the virtual image image. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0039] Figure 1 It is a schematic diagram of a feasible application scenario of the method for generating a virtual image provided in an embodiment of the present application.
[0040] Figure 2 It is a flowchart of a method for generating a virtual image provided in an embodiment of the present application.
[0041] Figure 3 It is a structural schematic diagram of a device for generating a virtual image provided in an embodiment of the present application.
[0042] Figure 4 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0044] Early virtual images originated from animation and image processing technology. With the development of digital entertainment and computer graphics, virtual images have gradually become an important part of various media, games and social platforms. With the rapid development of artificial intelligence (AI), deep learning and image processing technology, virtual images are no longer limited to simple animation images. Modern technology can make virtual images more realistic and dynamic through deep learning and three-dimensional modeling of facial features. Through technologies such as facial recognition and expression capture, virtual characters can respond to user input in real time and even show natural movements and expressions in virtual environments.
[0045] In this process, it is necessary to generate a virtual image based on the original image of the generated object, and then fit the virtual image back to the original image to obtain a virtual image image for playback. However, for different usage scenarios, the generated virtual image images still have some problems, such as face edge distortion, expressions or movements that are not as vivid as real people, inconsistent facial features, insufficient resolution or insufficient image details, etc. Therefore, how to improve the quality of virtual image images and thus improve people's user experience has become a technical problem that needs to be solved urgently by those skilled in the art.
[0046] Based on this, the present application proposes a method, device, equipment and computer program product for generating a virtual image. The technical solution utilizes the three-dimensional mapping relationship between the original image and the generated image to correct the posture of the virtual image in the generated image, and fits the corrected virtual image back to the original image, thereby reducing the unnatural seams at the edges of the virtual image and improving the quality of the virtual image.
[0047] Figure 1 A feasible application scenario of the method for generating a virtual image is shown, such as Figure 1 In the scenario shown, a client and a server are set up.
[0048] The client can be an electronic device with network access capability. Specifically, for example, the client can be a desktop computer, a tablet computer, a laptop computer, a smart phone, a digital assistant, a smart wearable device, a shopping guide terminal, a television, etc. Among them, the smart wearable device includes but is not limited to a smart bracelet, a smart watch, a smart glasses, a smart helmet, a smart necklace, etc. Alternatively, the client can also be software that can run in an electronic device.
[0049] The server can be an electronic device with certain computing and processing capabilities. It can have a network communication module, a processor, and a memory, etc. Of course, the server can also refer to software running in an electronic device. The server can also be a distributed server, which can be a system with multiple processors, memories, network communication modules, etc. operating in collaboration. Alternatively, the server can also be a server cluster formed by several servers. Alternatively, with the development of science and technology, the server can also be a new technical means that can realize the corresponding functions of the implementation method of the specification. For example, it can be a new form of "server" based on quantum computing.
[0050] The client and the server can communicate through the target network. The target network can be any type of network. For example, the target network can be a network that can be subdivided into multiple subnetworks. The target network or the multiple subnetworks contained in the target network can be at least one of a cellular mobile network (such as 2G, 3G, 4G or 5G), ZIGBEE, Wi-Fi, Bluetooth, or any combination of at least one of these networks and other networks.
[0051] In the above feasible application scenario, the client can obtain the original image of the generated object and send the original image of the generated object to the server. The server processes the original image to obtain a generated image, wherein the generated image contains a virtual image of the generated object, and then extracts the first three-dimensional feature from the original image of the generated object, extracts the second three-dimensional feature from the generated image, and establishes a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature. The virtual image in the generated image is posture-corrected using the three-dimensional mapping relationship to obtain a posture-corrected virtual image, and the posture-corrected virtual image is fitted to the original image to obtain a virtual image image of the generated object. The server sends the virtual image image of the generated object to the client so that the client can output the virtual image of the generated object.
[0052] In this way, the virtual image after posture correction using the three-dimensional mapping relationship has a higher posture consistency with the object generated in the original image, and the virtual image is more closely aligned with the original image, which can reduce unnatural seams at the edges of the virtual image and improve the quality of the virtual image image.
[0053] Furthermore, the present application embodiment provides a method for generating a virtual image, which can be executed by an electronic device, which can be any device with data and instruction processing functions, such as a notebook computer, a tablet computer, a desktop computer, a mobile device (e.g., a mobile phone, a personal digital assistant, a dedicated messaging device) and other types of user terminals, or a combination of any two or more of these electronic devices, or a server, such as the server in the above embodiment. Figure 2 As shown, the method includes:
[0054] S101, extracting a first three-dimensional feature from an original image of a generated object, extracting a second three-dimensional feature from a generated image, and establishing a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature.
[0055] The generated object refers to the object for which the virtual image is generated, and the generated object is generally a living being, such as a person or an animal, etc., which is not limited in this embodiment. The original image of the generated object refers to the real image of the generated object, and the original image of the generated object generally includes the facial image of the generated object, which can be obtained in real time by an optical device or imported by other devices, which is not limited in this embodiment.
[0056] The generated image is generated based on the original image, and the generated image includes a virtual image of the generated object. In the embodiment of the present application, the original image of the generated object is first processed to obtain a generated image including the virtual image of the generated object. In some embodiments, the original image can be processed using an image generation algorithm to obtain a corresponding generated image. Among them, the image generation algorithm can adopt a mature algorithm in the prior art, such as Generative Adversarial Networks (GAN), etc., which is not limited in this embodiment.
[0057] In order to make the virtual image more consistent with the posture of the generated object in the original image, the virtual image and the original image more closely fit, and reduce unnatural seams at the edges of the virtual image, the present application requires further processing of the generated image.
[0058] Extracting the first 3D feature of the original image and the second 3D feature of the generated image. In some embodiments, the original image and the generated image are processed to obtain the first 3D feature and the second 3D feature by training a 3D feature extraction model capable of performing 3D feature extraction tasks.
[0059] Exemplarily, a data set containing a large number of two-dimensional face images and corresponding three-dimensional face models can be prepared first. The two-dimensional face images can be used as training samples, and the corresponding three-dimensional face models can be used as training labels. The training samples are input into the three-dimensional feature extraction model to obtain the output results of the three-dimensional feature extraction model. By comparing the output results of the three-dimensional feature extraction model with the training labels, the loss value of the three-dimensional feature extraction model is determined. The parameters of the three-dimensional feature extraction model are adjusted with the goal of reducing the loss of the three-dimensional feature extraction model, and then the above training process is repeated until the parameters of the three-dimensional feature extraction model meet the training requirements. Then, the last fully connected layer of the three-dimensional feature extraction model is modified so that the output of the three-dimensional feature extraction model is a three-dimensional feature with depth, including the posture, expression, geometric structure, etc. of the face.
[0060] Among them, the three-dimensional feature extraction model can use a convolutional neural network model as a basic model, which is not limited in this embodiment.
[0061] The original image is input into the three-dimensional feature extraction model so that the three-dimensional feature extraction model processes the original image to obtain the first three-dimensional feature output by the three-dimensional feature extraction model; the generated image is input into the three-dimensional feature extraction model so that the three-dimensional feature extraction model processes the generated image to obtain the second three-dimensional feature output by the three-dimensional feature extraction model.
[0062] Furthermore, a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature is established. Specifically, the three-dimensional mapping relationship in the first three-dimensional feature and the second three-dimensional feature can be established according to the feature representation of each feature point in the first three-dimensional feature and the second three-dimensional feature.
[0063] The positions and number of the feature points can be set according to actual conditions, and are not limited in this embodiment. For example, 68 feature points or 106 feature points can be selected from the face, and are not limited in this embodiment.
[0064] For any feature point, determine the feature representation of the feature point in the first three-dimensional feature and the feature representation of the feature point in the second three-dimensional feature, and then establish a three-dimensional mapping relationship between the two feature representations, that is, define the conversion rule of the feature point from one feature representation to another feature representation. This mapping can be implemented based on a variety of mathematical functions or logical operations, and can be set according to actual conditions. This embodiment does not limit it. For each feature point, a three-dimensional mapping relationship is established according to the same conversion rule, and then a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature can be obtained.
[0065] S102, using the three-dimensional mapping relationship, performing posture correction on the virtual image in the generated image to obtain a virtual image with posture correction.
[0066] Specifically, in this embodiment, the three-dimensional mapping relationship can be used to perform posture correction on the virtual image in the generated image to obtain the virtual image with corrected posture.
[0067] In some embodiments, the feature points in the virtual image are spatially mapped and adjusted using a three-dimensional mapping relationship to achieve posture correction of the virtual image. The three-dimensional mapping relationship is essentially a spatial transformation matrix that includes the correspondence between two three-dimensional feature spaces. Through this spatial transformation matrix, the virtual image in the generated image can better match the posture and spatial position of the original image.
[0068] For example, the grid_sample function can be used to perform posture correction on the virtual image in the generated image. The three-dimensional mapping relationship and the generated image are input into the grid_sample function so that the grid_sample function processes the input generated image according to the given three-dimensional mapping relationship and outputs the virtual image with posture correction.
[0069] In this way, the posture of the virtual image is corrected by using the three-dimensional mapping relationship, so that the posture of the virtual image in the three-dimensional space is closer to the posture of the object generated in the original image, thereby reducing the problem of visually inconsistent edge distortion during fitting.
[0070] S103, fitting the virtual image after posture correction to the original image to obtain a virtual image of the generated object.
[0071] In this embodiment, the virtual image after posture correction is fused with the original image.
[0072] The above-mentioned fusion operation may include first using a mask to mask the image area of the generated object in the original image, and then fitting the virtual image after posture correction to the mask position, so as to obtain the virtual image image of the generated object.
[0073] Furthermore, in order to further improve the naturalness of the virtual image, a weighted average method can be used for fusion. This method can smoothly transition the boundary between the two images and reduce the abruptness.
[0074] In order to ensure a natural fusion effect, a weight function can be designed, which determines the contribution ratio of the virtual image after posture correction and the original image at each position. Linear gradient, Gaussian kernel or Poisson editing can be used as the above weight function, which is not limited in this embodiment.
[0075] The value of each pixel in the avatar image is calculated as:
[0076] I out (x,y)=(1-w(x,y))·I orig (x,y)+w(x,y)·I gen (x,y)
[0077] Among them, w(x,y) represents the weight function, I orig represents the original image, I gen represents the virtual image after posture correction, I out Indicates an avatar image.
[0078] It should be noted that w(x,y) needs to be multiplied by the mask M to ensure that only genA weighted average is applied within the affected area. Furthermore, if I orig and I gen There is a clear color difference between the two, which can be used to identify the gen Perform color matching or white balance adjustments.
[0079] In this way, the virtual image with posture correction can be seamlessly fitted into the original image, and the overall effect is realistic with no obvious distortion at the edges.
[0080] In the above embodiment, a first three-dimensional feature can be extracted from the original image of the generated object, a second three-dimensional feature can be extracted from the generated image, and a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature can be established. The generated image is generated based on the original image, and the generated image contains a virtual image of the generated object. Then, the virtual image in the generated image is corrected in posture using the three-dimensional mapping relationship to obtain a virtual image after posture correction, so as to fit the virtual image after posture correction to the original image to obtain a virtual image image of the generated object. In this way, the virtual image after posture correction using the three-dimensional mapping relationship has a higher consistency with the posture of the generated object in the original image, and the virtual image is more closely fitted to the original image, which can reduce the unnatural seams at the edge of the virtual image and improve the quality of the virtual image image.
[0081] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment extract the first three-dimensional feature from the original image of the generated object and extract the second three-dimensional feature from the generated image, which may specifically include the following steps:
[0082] The first image feature extracted from the original image of the generated object is mapped into the three-dimensional space to obtain the first three-dimensional feature of the original image; the second image feature extracted from the generated image is mapped into the three-dimensional space to obtain the second three-dimensional feature of the generated image.
[0083] Specifically, the original image may be first feature extracted and encoded to obtain the first image feature, and the generated image may be feature extracted and encoded to obtain the second image feature. In some embodiments, a convolutional neural network (CNN) deep learning model is used to extract and encode the original image to obtain the first image feature, and to extract and encode the generated image to obtain the second image feature.
[0084] The first image feature is the representation of the original image at different levels, capturing information such as texture, edge, and posture in the original image; the second image feature is the representation of the generated image at different levels, capturing information such as texture, edge, and posture in the generated image.
[0085] Then, the first image feature is mapped into the three-dimensional space to obtain the first three-dimensional feature of the original image, and the second image feature extracted from the generated image is mapped into the three-dimensional space to obtain the second three-dimensional feature of the generated image.
[0086] In some embodiments, a three-dimensional mapping model capable of performing three-dimensional space mapping tasks is trained, and the first image feature and the second image feature are processed respectively to obtain a first three-dimensional feature and a second three-dimensional feature.
[0087] Exemplarily, a data set containing a large number of two-dimensional face images and corresponding three-dimensional face models can be prepared first. Image features are extracted from the two-dimensional face images as training samples, and the corresponding three-dimensional face models are used as training labels. The training samples are input into the three-dimensional mapping model to obtain the output results of the three-dimensional mapping model. By comparing the output results of the three-dimensional mapping model with the training labels, the loss value of the three-dimensional mapping model is determined. The parameters of the three-dimensional mapping model are adjusted with the goal of reducing the loss of the three-dimensional mapping model, and then the above training process is repeated until the parameters of the three-dimensional mapping model meet the training requirements. Then, the last fully connected layer of the three-dimensional mapping model is modified so that the output of the three-dimensional mapping model is a three-dimensional feature with depth, including the posture, expression, geometric structure, etc. of the face.
[0088] The 3D mapping model may use a 3D Morphable Model (3DMM) as a basic model, which is not limited in this embodiment. 3DMM is a model commonly used for 3D face reconstruction, which can extract 3D feature information from a 2D image, including the posture, expression, and geometric structure of the face.
[0089] In addition to the supervised training method described above, an unsupervised training method may also be used to train the three-dimensional mapping model, which is not limited in this embodiment.
[0090] The first image feature is input into the three-dimensional mapping model so that the three-dimensional mapping model processes the first image feature to obtain the first three-dimensional feature output by the three-dimensional mapping model; the second image feature is input into the three-dimensional mapping model so that the three-dimensional mapping model processes the second image feature to obtain the second three-dimensional feature output by the three-dimensional mapping model.
[0091] In the above embodiment, the two-dimensional image features can be converted into three-dimensional features with depth information, and the posture of the virtual image can be corrected using the three-dimensional features with depth information. The virtual image after posture correction can be more consistent with the original image.
[0092] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment establish a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature, which may specifically include the following steps:
[0093] A three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature is established according to the feature representation of each feature point in the first three-dimensional feature and the second three-dimensional feature.
[0094] Among them, feature points refer to fixed positions with specific anatomical significance in a face image. These points are usually located in key parts of facial structures such as eyes, eyebrows, nose, mouth, chin, etc., and can accurately describe the geometric shape and appearance characteristics of the face. The position and number of feature points can be set according to actual conditions, and this embodiment does not limit it. Exemplarily, 68 feature points or 106 feature points can be selected from the face, and this embodiment does not limit it. For example, the above 68 feature points include 17 point contour feature points, 10 eyebrow feature points, 12 eye feature points, 9 nose feature points and 20 mouth feature points.
[0095] For any feature point, determine the feature representation of the feature point in the first three-dimensional feature and the feature representation of the feature point in the second three-dimensional feature, and then establish a three-dimensional mapping relationship between the two feature representations, that is, define the conversion rule of the feature point from one feature representation to another feature representation. This mapping can be implemented based on a variety of mathematical functions or logical operations, and can be set according to actual conditions, and this embodiment does not limit it.
[0096] Furthermore, a three-dimensional mapping relationship is established for each feature point according to the same conversion rule, and then a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature can be obtained. The three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature is essentially a spatial transformation matrix, which includes the correspondence between the two three-dimensional feature spaces. Through this spatial transformation matrix, the generated image can better match the posture and spatial position of the original image.
[0097] As an optional implementation, another embodiment of the present application discloses that, in the case where the feature points include feature points of the lip region of the generated object, the steps of the above embodiment establish a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature according to the feature representation of the same feature point in the first three-dimensional feature and the second three-dimensional feature, and specifically may include the following steps:
[0098] A three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature in the lip region is established according to feature representation of the specific lip region feature point in the first three-dimensional feature and the second three-dimensional feature.
[0099] Specifically, the movements of the lip region are consistent, and in order to keep the driven lip shape accurate, in this embodiment, the feature points of the lip region are fixed to have the same three-dimensional mapping relationship. In other words, the three-dimensional mapping relationship of the feature points of the lip region is the same.
[0100] A specific lip region feature point can be selected from the lip region as a representative point of the lip region, and a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature of all feature points in the lip region is established based on the feature representation of the specific lip region feature point in the first three-dimensional feature and the second three-dimensional feature. It should be noted that the specific lip region feature point can be selected according to actual needs, for example, the feature point of the lip peak region is selected as the specific lip region feature point, which is not limited in this embodiment.
[0101] In a specific embodiment, the lip region includes 20 feature points, namely feature point 1, feature point 2, feature point 3, ..., feature point 19, and feature point 20. Feature point 1 is selected as the feature point of the specific lip region. According to the feature representation of feature point 1 in the first three-dimensional feature and the feature representation in the second three-dimensional space, a three-dimensional mapping relationship corresponding to feature point 1 is established. It is determined that the three-dimensional mapping relationships corresponding to feature point 2-feature point 20 are all the three-dimensional mapping relationships corresponding to feature point 1.
[0102] This arrangement makes the three-dimensional mapping relationship of the lip area consistent, so that when the posture of the virtual image in the generated image is corrected, the movement of the lip area is consistent and more realistic.
[0103] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment use a three-dimensional mapping relationship to perform posture correction on the virtual image in the generated image to obtain a virtual image after posture correction, which may specifically include the following steps:
[0104] According to the three-dimensional mapping relationship, the second image feature extracted from the generated image is transformed to obtain the third image feature; the third image feature is used to perform posture correction on the virtual image in the generated image to obtain the virtual image after posture correction.
[0105] Specifically, the third image feature is obtained by performing feature transformation on the second image feature using the three-dimensional mapping relationship. The three-dimensional mapping relationship is essentially a spatial transformation matrix that contains the correspondence between two three-dimensional feature spaces. Through this spatial transformation matrix, the third image feature can better match the posture and spatial position of the original image.
[0106] Exemplarily, the grid_sample function can be used to perform feature transformation on the second image feature, and the three-dimensional mapping relationship and the second image feature can be input into the grid_sample function, so that the grid_sample function processes the input second image feature according to the given three-dimensional mapping relationship and outputs the third image feature.
[0107] Then, the third image feature is used to perform posture correction on the virtual image in the generated image to obtain the virtual image after posture correction. In some embodiments, the third image feature and the second image feature can be spliced to obtain a spliced feature, and then the spliced feature is decoded to obtain the virtual image after posture correction.
[0108] In the above embodiment, the feature transformation of the second image feature is performed using a three-dimensional mapping relationship, so that the obtained third image feature is closer to the posture of the object generated in the original image in the three-dimensional space. The posture of the virtual image is corrected using the third image feature, and the obtained virtual image is closer to the posture of the object generated in the original image, thereby reducing the problem of visually inharmonious edge distortion during fitting.
[0109] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment use the third image feature to perform posture correction on the virtual image in the generated image to obtain the virtual image after posture correction, which may specifically include the following steps:
[0110] The third image feature, the first image feature extracted from the original image of the generated object, and the three-dimensional mapping relationship are fused to obtain a fused feature; the second image feature is posture-corrected using the fused feature to obtain a virtual image feature; and a virtual image with posture correction is generated according to the virtual image feature.
[0111] Specifically, the third image feature, the first image feature extracted from the original image of the generated object, and the three-dimensional mapping relationship can be fused to obtain a fused feature. A simple concatenation operation can be used to splice the third image feature, the first image feature, and the three-dimensional mapping relationship together to obtain a fused feature; an attention mechanism can also be used to splice the third image feature, the first image feature, and the three-dimensional mapping relationship together to obtain a fused feature, which is not limited in this embodiment.
[0112] The fusion feature is a combination of features based on the third image feature, the first image feature, and the three-dimensional mapping relationship, which can better express the overall information of the image.
[0113] Then, the fused feature is used to perform posture correction on the second image feature to obtain a virtual image feature. Specifically, the fused feature and the second image feature can be fused together again in the manner described in the above embodiment to obtain a second fused feature. Then, a deep neural network, such as CNN or GAN, is used to perform feature extraction again based on the second fused feature to generate a higher quality image feature as a virtual image feature.
[0114] Then, a virtual image with posture correction is generated based on the virtual image features. In some embodiments, the virtual image features are decoded by a decoder, and the high-dimensional features are mapped back to the pixel space to obtain the virtual image with posture correction.
[0115] With such a setting, the obtained virtual image features are closer to the posture of the object generated in the original image, and the virtual image decoded from the virtual image features is also closer to the posture of the object generated in the original image, thereby reducing the problem of visually inharmonious edge distortion during fitting.
[0116] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment generate a virtual image after posture correction according to the virtual image features, which may specifically include the following steps:
[0117] The virtual image features are upsampled to obtain upsampled features; and a virtual image whose pixels are consistent with those of the original image is generated based on the upsampled features.
[0118] Specifically, the decoder can map the high-dimensional virtual image features back to the pixel space; the virtual image features are upsampled through deconvolution or upsampling operations to obtain upsampled features; based on the upsampled features, a virtual image whose pixels are consistent with those of the original image is generated to restore the resolution of the image. In other words, the decoder maps the high-dimensional features back to the pixel space, and the deconvolution or upsampling operation is used to restore the resolution of the image.
[0119] Such a setting can improve the resolution of the generated virtual image to keep it consistent with the original image, avoid the problem of insufficient resolution or insufficient image details, and improve the quality of the virtual image image.
[0120] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment fit the virtual image after posture correction to the original image to obtain the virtual image of the generated object, which may specifically include the following steps:
[0121] The image area of the generated object in the original image is masked; the virtual image after posture correction is fitted to the masked area to obtain the virtual image image of the generated object.
[0122] Specifically, the image area of the generated object in the original image can be masked by a mask to obtain a masked area, and then the virtual image after posture correction is attached to the masked area to obtain a virtual image image of the generated object.
[0123] For example, the virtual image after posture correction can be fitted to the mask area by aligning the boundary of the mask area with the boundary of the virtual image after posture correction. It is also possible to extract feature points of the generated object in the original image and feature points in the virtual image after posture correction, and fit the virtual image after posture correction to the mask area by aligning the same feature point of the generated object in the original image with the virtual image after posture correction.
[0124] The fitted boundary area can also be smoothed through Gaussian or Poisson fusion to ensure that the details of the virtual image after posture correction are naturally integrated with the original image and reduce the visual difference at the edge.
[0125] This setting ensures that the virtual image after posture correction can be seamlessly fitted into the original image, ensuring that the overall effect is more realistic and there is no obvious distortion at the edges.
[0126] In the above embodiment, feature extraction, 3DMM three-dimensional modeling, feature fusion, upsampling and image decoding technology in deep learning are used to generate high-quality images, and the generated virtual image is seamlessly connected with the original image through image fusion technology. This method can effectively reduce the inconsistency between the generated virtual image and the original image in posture, resolution and edge details, and improve the naturalness and quality of the virtual image.
[0127] Corresponding to the above-mentioned method for generating a virtual image, the present application embodiment also discloses a device for generating a virtual image, see Figure 3 As shown, the device comprises:
[0128] The extraction unit 100 is used to extract a first three-dimensional feature from an original image of a generated object, extract a second three-dimensional feature from a generated image, and establish a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature; wherein the generated image is generated based on the original image, and the generated image includes a virtual image of the generated object;
[0129] The correction unit 110 is used to perform posture correction on the virtual image in the generated image by using the three-dimensional mapping relationship to obtain the virtual image after posture correction;
[0130] The fitting unit 120 is used to fit the virtual image after posture correction to the original image to obtain a virtual image image of the generated object.
[0131] As an optional implementation, another embodiment of the present application discloses that the extraction unit 100 of the above embodiment, when extracting the first three-dimensional feature from the original image of the generated object and extracting the second three-dimensional feature from the generated image, is specifically used to:
[0132] The first image feature extracted from the original image of the generated object is mapped into the three-dimensional space to obtain the first three-dimensional feature of the original image; the second image feature extracted from the generated image is mapped into the three-dimensional space to obtain the second three-dimensional feature of the generated image.
[0133] As an optional implementation, another embodiment of the present application discloses that the extraction unit 100 in the above embodiment, when establishing a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature, is specifically used to:
[0134] A three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature is established according to the feature representation of each feature point in the first three-dimensional feature and the second three-dimensional feature.
[0135] As an optional implementation, another embodiment of the present application discloses that, when the feature points include feature points of the lip region of the generated object, the extraction unit 100 of the above embodiment, when establishing a three-dimensional mapping relationship in the first three-dimensional feature and the second three-dimensional feature according to the feature representation of each feature point in the first three-dimensional feature and the second three-dimensional feature, is specifically used to:
[0136] A three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature in the lip region is established according to feature representation of the specific lip region feature point in the first three-dimensional feature and the second three-dimensional feature.
[0137] As an optional implementation, another embodiment of the present application discloses that the correction unit 110 of the above embodiment, when using the three-dimensional mapping relationship to perform posture correction on the virtual image in the generated image to obtain the virtual image after posture correction, is specifically used to:
[0138] According to the three-dimensional mapping relationship, the second image feature extracted from the generated image is transformed to obtain the third image feature; the third image feature is used to perform posture correction on the virtual image in the generated image to obtain the virtual image after posture correction.
[0139] As an optional implementation, another embodiment of the present application discloses that the correction unit 110 of the above embodiment, when performing posture correction on the virtual image in the generated image using the third image feature to obtain the virtual image after posture correction, is specifically used to:
[0140] The third image feature, the first image feature extracted from the original image of the generated object, and the three-dimensional mapping relationship are fused to obtain a fused feature; the second image feature is posture-corrected using the fused feature to obtain a virtual image feature; and a virtual image with posture correction is generated according to the virtual image feature.
[0141] As an optional implementation, another embodiment of the present application discloses that the correction unit 110 of the above embodiment, when generating a virtual image after posture correction according to the virtual image features, is specifically used to:
[0142] The virtual image features are upsampled to obtain upsampled features; and a virtual image whose pixels are consistent with those of the original image is generated based on the upsampled features.
[0143] Specifically, the device provided in this embodiment belongs to the same application concept as the method provided in the above embodiment of this application, can execute the method provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not fully described in this embodiment, please refer to the specific processing content of the method provided in the above embodiment of this application, which will not be repeated here.
[0144] The functions implemented by the above units can be implemented by the same or different processors respectively, and the embodiments of the present application are not limited thereto.
[0145] It should be understood that the units in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits, and the functions of some or all units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by the configuration file, so as to realize the functions of some or all of the above units. All units of the above devices can be implemented in the form of a processor calling software, or in the form of hardware circuits, or in part by a processor calling software, and the remaining part is implemented in the form of hardware circuits.
[0146] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.
[0147] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0148] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0149] An embodiment of the present application further provides a control device, which includes a processor and an interface circuit. The processor in the control device is connected to an input-output component through the interface circuit of the control device.
[0150] The input-output component specifically refers to a hardware component that enables a user to input information and output information to the user, such as a microphone, keyboard, handwriting tablet, touch screen, display, speaker, printer, etc.
[0151] The above-mentioned interface circuit can be any interface circuit that can realize the data communication function, for example, it can be a USB interface circuit, a Type-C interface circuit, a serial port circuit, a PCIE circuit, etc.
[0152] The processor in the control device is a circuit with signal processing capability, which improves the quality of the virtual image by executing any one of the methods for generating a virtual image described in the above embodiments. The specific implementation of the processor can refer to the above processor implementation, and the embodiments of this application are not strictly limited.
[0153] When the control device is applied to a device with a human-computer interaction function, the input and output components of the control device may be input components and output components on the device, such as a microphone, a keyboard, a handwriting tablet, a touch screen, a display, an audio player, etc. At the same time, the processor of the control device may be a CPU or GPU, etc. provided by the device, and the interface circuit of the control device may be an interface circuit between the information input component of the device and a processor such as a CPU or GPU.
[0154] Corresponding to the above method for generating a virtual image, the present application embodiment further discloses an electronic device, see Figure 4 As shown, the electronic device includes:
[0155] Memory 200 and processor 210;
[0156] The memory 200 is connected to the processor 210 and is used to store programs;
[0157] The processor 210 is used to implement the method for generating a virtual image disclosed in any of the above embodiments by running the program stored in the memory 200.
[0158] Specifically, the electronic device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .
[0159] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected to each other via a bus.
[0160] A bus may include a pathway that transfers information between components of a computer system.
[0161] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0162] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0163] The memory 200 stores a program for executing the technical solution of the present application, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.
[0164] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0165] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0166] The communication interface 220 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0167] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of the method for generating a virtual image provided in the above embodiment of the present application.
[0168] In addition to the above methods and devices, the embodiments of the present application may also be a computer program product, which includes a computer program. When the computer program is executed by a processor, the method for generating a virtual image provided by any of the above embodiments of the present application may be executed. Optionally, the computer program may be stored in a readable storage medium or in the cloud of a computer device; the processor of the computer device reads the computer program from the readable storage medium or the cloud.
[0169] The computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, etc., and also conventional procedural programming languages such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0170] The computer program product may be implemented in hardware, software or a combination thereof. In one optional embodiment, the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0171] In addition, an embodiment of the present application may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes each step of the method for generating a virtual image provided in the above embodiment.
[0172] The computer readable storage medium may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0173] Specifically, the specific working contents of each part of the above-mentioned electronic device, computer program product and storage medium, as well as the specific processing contents when the computer program product or the computer program on the above-mentioned storage medium is executed by the processor, can all be found in the contents of the various embodiments of the above-mentioned method for generating a virtual image, and will not be repeated here.
[0174] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.
[0175] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0176] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0177] The modules and sub-modules in the devices and terminals in the various embodiments of the present application can be combined, divided and deleted according to actual needs.
[0178] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, such as multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0179] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.
[0181] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0182] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0183] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including a..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0184] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating a virtual image, characterized in that: include: Extracting a first three-dimensional feature from an original image of the generated object, extracting a second three-dimensional feature from the generated image, and establishing a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature; wherein the generated image is generated based on the original image, and the generated image includes a virtual image of the generated object; Using the three-dimensional mapping relationship, performing posture correction on the virtual image in the generated image to obtain a virtual image after posture correction; The virtual image after posture correction is fitted into the original image to obtain the virtual image image of the generated object.
2. The method according to claim 1, characterized in that: The step of extracting the first three-dimensional feature from the original image of the generated object and extracting the second three-dimensional feature from the generated image comprises: Mapping a first image feature extracted from an original image of the generated object into a three-dimensional space to obtain a first three-dimensional feature of the original image; The second image feature extracted from the generated image is mapped into a three-dimensional space to obtain a second three-dimensional feature of the generated image.
3. The method according to claim 1, characterized in that The establishing of the three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature includes: A three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature is established according to the feature representation of each feature point in the first three-dimensional feature and the second three-dimensional feature.
4. The method according to claim 3, characterized in that In a case where the feature point includes a lip region feature point of the generated object, establishing a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature according to a feature representation of the same feature point in the first three-dimensional feature and the second three-dimensional feature includes: A three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature in the lip region is established according to feature representation of the specific lip region feature point in the first three-dimensional feature and the second three-dimensional feature.
5. The method according to claim 1, characterized in that: The method of using the three-dimensional mapping relationship to perform posture correction on the virtual image in the generated image to obtain the virtual image after posture correction includes: According to the three-dimensional mapping relationship, feature transformation is performed on the second image feature extracted from the generated image to obtain a third image feature; The third image feature is used to perform posture correction on the virtual image in the generated image to obtain a virtual image with corrected posture.
6. The method according to claim 5, characterized in that The step of using the third image feature to perform posture correction on the virtual image in the generated image to obtain a virtual image after posture correction includes: fusing the third image feature, the first image feature extracted from the original image of the generated object, and the three-dimensional mapping relationship to obtain a fused feature; Performing posture correction on the second image feature by using the fusion feature to obtain a virtual image feature; A virtual image with posture correction is generated according to the virtual image features.
7. The method according to claim 6, characterized in that The step of generating a posture-corrected virtual image according to the virtual image features comprises: Performing upsampling processing on the virtual image features to obtain upsampled features; A virtual image whose pixels are consistent with those of the original image is generated according to the up-sampling features.
8. The method according to claim 1, characterized in that The step of fitting the posture-corrected virtual image to the original image to obtain the virtual image of the generated object includes: Mask the image area of the generated object in the original image; The virtual image after posture correction is fitted to the mask area to obtain the virtual image image of the generated object.
9. A device for generating a virtual image, characterized in that: include: an extraction unit, configured to extract a first three-dimensional feature from an original image of a generated object, extract a second three-dimensional feature from a generated image, and establish a three-dimensional mapping relationship between the first three-dimensional feature and the second three-dimensional feature; wherein the generated image is generated based on the original image, and the generated image includes a virtual image of the generated object; A correction unit, configured to perform posture correction on the virtual image in the generated image by using the three-dimensional mapping relationship to obtain a virtual image after posture correction; The fitting unit is used to fit the virtual image after posture correction to the original image to obtain the virtual image image of the generated object.
10. An electronic device, characterized in that: include: Memory and processor; Wherein, the memory is used to store programs; The processor is used to implement the method according to any one of claims 1 to 8 by running the program in the memory.
11. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.