Image rendering method, apparatus, device, and storage medium
By optimizing rendering parameters through 3D reconstruction and training models based on shape and pose parameters, the problem of mismatch between 3D model rendering parameters is solved, thereby improving rendering effects and image quality.
Patent Information
- Application Number
- CN202110721851.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-28
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2041-09-16
AI Technical Summary
In existing technologies, the rendering parameters of the 3D model do not match the model, resulting in poor rendering effects.
3D reconstruction is performed by obtaining shape and pose parameters based on sample video frames. Rendering parameters are determined using an image rendering model, and the rendering effect is optimized by training the model. The influence of virtual camera and sample objects is considered to match the rendering parameters.
It improves rendering effects, enhances the capabilities of the image rendering model, makes the rendered images closer to the sample video frames, and improves image quality.
Smart Images

Figure CN113822977B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and particularly relates to an image rendering method and device, equipment and a storage medium. BACKGROUND
[0002] With the development of computer technology, human body reconstruction technology is applied more and more widely. Human body reconstruction technology refers to a technology of obtaining a three-dimensional model of a human body according to a two-dimensional video or photo of the human body. The human body reconstruction technology is widely applied in live broadcast and animation production scenes. For example, in a live broadcast scene, a host can convert his / her own image into an animal through the human body reconstruction technology.
[0003] In the related art, a depth camera is used to collect point cloud data of a human body, and human body reconstruction is performed based on the collected point cloud data. However, when a three-dimensional model is rendered, rendering parameters are usually manually set by a technician. The rendering parameters can not match the three-dimensional model, resulting in poor rendering effect of the three-dimensional model. SUMMARY
[0004] Embodiments of the present application provide an image rendering method, device, equipment and storage medium, which can improve the rendering effect. The technical solution is as follows:
[0005] In one aspect, an image rendering method is provided, and the method comprises:
[0006] Based on a sample video frame, shape parameters and pose parameters of a sample object are obtained, and the sample video frame comprises the sample object;
[0007] Based on the shape parameters and the pose parameters, three-dimensional reconstruction is performed on the sample object to obtain a three-dimensional model of the sample object;
[0008] Through an image rendering model, a plurality of first rendering parameters are determined based on camera parameters of a virtual camera, the shape parameters and the pose parameters, the camera parameters of the virtual camera being the same as camera parameters of an actual camera that captures the sample video frame; based on the plurality of first rendering parameters, a first target image is rendered to output a first rendered image, the first target image being an image captured by the virtual camera on the three-dimensional model;
[0009] Based on difference information between the sample video frame and the first rendered image, the image rendering model is trained, and the image rendering model is used to render an image captured by the virtual camera.
[0010] In one possible implementation, before the plurality of first rendering parameters are determined based on the camera parameters, the shape parameters and the pose parameters, the method further comprises:
[0011] Based on the sample video frame, the camera parameter is acquired.
[0012] In a possible implementation, the first rendering parameter includes a color parameter and a density parameter, and the determining the color and the opacity of the pixel point based on the virtual ray between the pixel point and the virtual camera and the first rendering parameter corresponding to the pixel point includes:
[0013] integrating first relationship data on the virtual ray to obtain the color, the first relationship data being associated with the color parameter and the density parameter;
[0014] integrating second relationship data on the virtual ray to obtain the opacity, the second relationship data being associated with the density parameter.
[0015] In a possible implementation, the method further includes:
[0016] performing shape adjustment on the three-dimensional model of the sample object based on a target shape parameter to obtain a shape-adjusted three-dimensional model;
[0017] inputting the camera parameter, the pose parameter and the target shape parameter into the trained image rendering model, and determining a plurality of third rendering parameters based on the camera parameter, the pose parameter and the target shape parameter;
[0018] performing rendering on the third target image based on the plurality of third rendering parameters to output a third rendering image, the third target image being an image obtained by the virtual camera capturing the shape-adjusted three-dimensional model.
[0019] In a possible implementation, the method further includes:
[0020] inputting a target camera parameter, the pose parameter and the target shape parameter into the trained image rendering model, and determining a plurality of fourth rendering parameters based on the pose parameter, the shape parameter and the target camera parameter;
[0021] performing rendering on a fourth target image based on the plurality of fourth rendering parameters to output a fourth rendering image, the fourth target image being an image obtained by the virtual camera capturing the three-dimensional model under the target camera parameter.
[0022] In an aspect, an image rendering method is provided, and the method includes:
[0023] displaying a target video frame, the target video frame including a target object;
[0024] In response to a three-dimensional reconstruction operation on the target object, a three-dimensional model of the target object is displayed, the three-dimensional model being generated based on a shape parameter and a pose parameter of the target object, the shape parameter and the pose parameter being determined based on the target video frame;
[0025] In response to a shooting operation on the three-dimensional model, a first target image is displayed, the first target image being an image obtained by a virtual camera shooting the three-dimensional model;
[0026] In response to a rendering operation on the first target image, a first rendered image is displayed, the first rendered image being obtained by a trained image rendering model rendering the first target image based on a plurality of first rendering parameters, the plurality of first rendering parameters being determined by the image rendering model based on camera parameters of the virtual camera, the shape parameter, and the pose parameter, the image rendering model being used to render images shot by the virtual camera.
[0027] In one aspect, an image rendering apparatus is provided, and the apparatus comprises:
[0028] A parameter acquisition module is configured to acquire a shape parameter and a pose parameter of the sample object based on the sample video frame, the sample video frame comprising a sample object.
[0029] A three-dimensional reconstruction module is configured to perform three-dimensional reconstruction on the sample object based on the shape parameter and the pose parameter, to obtain a three-dimensional model of the sample object.
[0030] A rendering module is configured to determine a plurality of first rendering parameters by an image rendering model based on camera parameters of a virtual camera, the shape parameter, and the pose parameter, the camera parameters of the virtual camera being identical to camera parameters of a real camera that shot the sample video frame; render a first target image based on the plurality of first rendering parameters, and output a first rendered image, the first target image being an image obtained by the virtual camera shooting the three-dimensional model.
[0031] A training module is configured to train the image rendering model based on difference information between the sample video frame and the first rendered image, the image rendering model being used to render images shot by the virtual camera.
[0032] In one possible implementation, the three-dimensional reconstruction module is configured to adjust a shape of a reference three-dimensional model based on the shape parameter, to adjust a pose of the reference three-dimensional model based on the pose parameter, and to obtain the three-dimensional model of the sample object, the reference three-dimensional model being obtained based on shape parameters and pose parameters of a plurality of objects.
[0033] In a possible implementation, the apparatus further includes:
[0034] a region determination module, configured to perform image segmentation on the sample video frame to obtain a target region, the target region being a region where the sample object is located;
[0035] The parameter acquisition module is configured to acquire the shape parameter and the pose parameter of the sample object based on the target region.
[0036] In a possible implementation, the parameter acquisition module is configured to perform pose estimation on the sample object based on the sample video frame to obtain a pose parameter of the sample object; perform shape estimation on the sample object based on a plurality of video frames in a sample video to obtain a plurality of reference shape parameters of the sample object, one reference shape parameter corresponding to one video frame, the sample video including the sample video frame; and determine the shape parameter of the sample object based on the plurality of reference shape parameters.
[0037] In a possible implementation, the camera parameter includes a position parameter of the virtual camera in the first virtual space, and the rendering module is configured to determine at least one virtual ray in the first virtual space based on a perspective of the virtual camera on the three-dimensional model and the position parameter, the virtual ray being a line connecting the virtual camera and a pixel point on the first target image, the first virtual space being a virtual space established based on the camera parameter; determine the plurality of first rendering parameters based on coordinates of a plurality of first sampling points on the at least one virtual ray, the shape parameter, and the pose parameter, the coordinates of the first sampling point being coordinates of the first sampling point in the first virtual space.
[0038] In a possible implementation, the rendering module is configured to transform the plurality of first sampling points into a second virtual space based on the coordinates of the plurality of first sampling points, the pose parameter, and a reference pose parameter to obtain a plurality of second sampling points, one first sampling point corresponding to one second sampling point, the reference pose parameter being a pose parameter corresponding to the second virtual space, the coordinates of the second sampling point being coordinates of the second sampling point in the second virtual space; and determine the plurality of first rendering parameters based on the coordinates of the plurality of second sampling points in the second virtual space, the shape parameter, and the pose parameter.
[0039] In a possible implementation, the rendering module is configured to, for one of the first sampling points, obtain a first pose transformation matrix and a second pose transformation matrix of the first sampling point, the first pose transformation matrix being a transformation matrix of a first vertex from a first pose to a second pose, the second pose transformation matrix being a transformation matrix of the first vertex from the first pose to a third pose, the first pose being a reference pose, the second pose being a pose corresponding to the pose parameter, the third pose being a pose corresponding to the reference pose parameter, the first vertex being a vertex of the three-dimensional model that meets a target condition in terms of distance to the first sampling point; and obtain a second sampling point corresponding to the first sampling point based on a skinning weight corresponding to the first vertex, the first pose transformation matrix, and the second pose transformation matrix.
[0040] In a possible implementation, the rendering module is configured to, for one of the second sampling points, splice a coordinate of the second sampling point in the second virtual space, the shape parameter, and the pose parameter to obtain a first parameter set; and perform full connection processing on the first parameter set to obtain the first rendering parameter.
[0041] In a possible implementation, the rendering module is configured to, for one of the pixel points on the first target image, determine a color and an opacity of the pixel point based on a virtual ray between the pixel point and the virtual camera and the first rendering parameter corresponding to the pixel point; and render the pixel point based on the color and the opacity, and output the pixel point after rendering.
[0042] In a possible implementation, the first rendering parameter includes a color parameter and a density parameter, and the rendering module is configured to integrate first relationship data on the virtual ray to obtain the color, the first relationship data being associated with the color parameter and the density parameter; and integrate second relationship data on the virtual ray to obtain the opacity, the second relationship data being associated with the density parameter.
[0043] In a possible implementation, the rendering module is further configured to perform pose adjustment on the three-dimensional model of the sample object based on a target pose parameter to obtain a three-dimensional model after pose adjustment; input the camera parameter, the shape parameter, and the target pose parameter into the trained image rendering model, determine a plurality of second rendering parameters based on the camera parameter, the shape parameter, and the target pose parameter; and render a second target image based on the plurality of second rendering parameters to output a second rendering image, the second target image being an image obtained by the virtual camera capturing the three-dimensional model after pose adjustment.
[0044] In a possible implementation, the apparatus further includes:
[0045] The camera parameter acquisition module is configured to acquire the camera parameter based on the sample video frame.
[0046] In a possible implementation, the rendering module is further configured to perform shape adjustment on the three-dimensional model of the sample object based on target shape parameters to obtain a shape-adjusted three-dimensional model; input the camera parameter, the pose parameter and the target shape parameter into the trained image rendering model, determine a plurality of third rendering parameters based on the camera parameter, the pose parameter and the target shape parameter; and perform rendering on a third target image based on the plurality of third rendering parameters to output a third rendering image, the third target image being an image obtained by photographing the shape-adjusted three-dimensional model by the virtual camera.
[0047] In a possible implementation, the rendering module is further configured to input target camera parameters, the pose parameter and the target shape parameter into the trained image rendering model, determine a plurality of fourth rendering parameters based on the pose parameter, the shape parameter and the target camera parameter; and perform rendering on a fourth target image based on the plurality of fourth rendering parameters to output a fourth rendering image, the fourth target image being an image obtained by photographing the three-dimensional model by the virtual camera under the target camera parameter.
[0048] In an aspect, an image rendering apparatus is provided, and the apparatus includes:
[0049] The video frame display module is configured to display a target video frame, the target video frame including a target object.
[0050] The three-dimensional model display module is configured to display a three-dimensional model of the target object in response to a three-dimensional reconstruction operation on the target object, the three-dimensional model being generated based on shape parameters and pose parameters of the target object, the shape parameters and the pose parameters being determined based on the target video frame.
[0051] The target image display module is configured to display a first target image in response to a photographing operation on the three-dimensional model, the first target image being an image obtained by photographing the three-dimensional model by a virtual camera.
[0052] The rendering image display module is configured to display a first rendering image in response to a rendering operation on the first target image, the first rendering image being obtained by rendering the first target image based on a plurality of first rendering parameters determined by the image rendering model based on the camera parameters, the shape parameters and the pose parameters of the virtual camera, the image rendering model being configured to render an image captured by the virtual camera.
[0053] In an aspect, a computer device is provided, which includes one or more processors and one or more memories having stored therein at least one computer program, which is loaded and executed by the one or more processors to implement the image rendering method.
[0054] In an aspect, a computer readable storage medium is provided, which has stored therein at least one computer program, which is loaded and executed by a processor to implement the image rendering method.
[0055] In an aspect, a computer program product or computer program is provided, which includes program code stored in a computer readable storage medium, the program code being read by a processor of a computer device from the computer readable storage medium, and the processor executes the program code to cause the computer device to perform the image rendering method.
[0056] By the technical solutions provided in the embodiments of the present application, when training the image rendering model, the sample object is reconstructed in three dimensions based on the shape parameters and the pose parameters of the sample object, the rendering parameters are determined based on the camera parameters, the shape parameters and the pose parameters, the influence of the virtual camera and the sample object is considered when determining the rendering parameters, the rendering parameters are more matched with the three-dimensional model of the sample object, the first target image is rendered based on the rendering parameters to obtain the first rendering image, and the image rendering model is trained based on the difference information between the first rendering image and the sample video frame, so that the image rendering model obtained in this way has stronger image rendering capability. BRIEF DESCRIPTION OF DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0058] Figure 1 FIG. 1 is a schematic diagram of an implementation environment of an image rendering method provided by an embodiment of the present application;
[0059] Figure 2 is a flowchart of an image rendering method provided by an embodiment of the present application;
[0060] Figure 3 is a flowchart of an image rendering method provided by an embodiment of the present application;
[0061] Figure 4 is a flowchart of an image rendering method provided by an embodiment of the present application;
[0062] Figure 5 is a flowchart of an image rendering method provided by an embodiment of the present application;
[0063] Figure 6 is a posture transformation schematic diagram provided by an embodiment of the present application;
[0064] Figure 7 is a flowchart of an image rendering method provided by an embodiment of the present application;
[0065] Figure 8 is a flowchart of an image rendering method provided by an embodiment of the present application;
[0066] Figure 9 is a view transformation schematic diagram provided by an embodiment of the present application;
[0067] Figure 10 is a view transformation schematic diagram provided by an embodiment of the present application;
[0068] Figure 11 is a flowchart of an image rendering method provided by an embodiment of the present application;
[0069] Figure 12 is an interface schematic diagram provided by an embodiment of the present application;
[0070] Figure 13 is an interface schematic diagram provided by an embodiment of the present application;
[0071] Figure 14 is a structure schematic diagram of an image rendering device provided by an embodiment of the present application;
[0072] Figure 15 is a structure schematic diagram of an image rendering device provided by an embodiment of the present application;
[0073] Figure 16 is a structure schematic diagram of a terminal provided by an embodiment of the present application;
[0074] Figure 17 is a structure schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION
[0075] In order to make the purposes, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0076] The terms "first", "second", and the like are used in the present application to distinguish between the same or similar items with substantially the same function, and it should be understood that there is no logical or chronological dependency between "first", "second", and "nth", and the number and execution order are not limited.
[0077] The term "at least one" in the present application means one or more, and the meaning of "multiple" is two or more, for example, multiple reference face images refer to two or more reference face images.
[0078] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0079] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0080] Computer vision (CV) is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to identify, track and measure targets, and further process graphics so that the computer processing is more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision researches related theories and technologies, trying to establish artificial intelligence systems that can obtain information from images or multidimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, etc.
[0081] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory, etc. It is a specialized study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge sub-models, and continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.
[0082] Normalization: mapping a number series with different value ranges to the interval (0, 1) for easy data processing. In some cases, the normalized value can be directly implemented as a probability.
[0083] Figure 1 is a schematic diagram of an implementation environment of an image rendering method provided by an embodiment of the present application, referring to Figure 1 The implementation environment can include a terminal 110 and a server 140.
[0084] The terminal 110 is connected to the server 140 through a wireless network or a wired network. Optionally, the terminal 110 is a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal 110 is installed and runs an application program supporting image rendering.
[0085] Optionally, the server is a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0086] Optionally, the terminal 110 generally refers to one of multiple terminals, and the embodiments of the present application are only exemplified by the terminal 110.
[0087] Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminal is only one, or the above terminal is several tens or several hundreds, or more, at this time the above implementation environment also includes other terminals. The number and type of the terminal are not limited in the embodiments of the present application.
[0088] After introducing the implementation environment of the technical solutions provided by the embodiments of the present application, the application scenarios of the technical solutions provided by the embodiments of the present application will be described in combination with the above implementation environment. In the following description, the terminal is the terminal 110 in the above implementation environment, and the server is the terminal 140 in the above implementation environment.
[0089] The technical solutions provided by the embodiments of the present application can be applied in the scenarios of multi-view image synthesis and new action image synthesis. The multi-view image synthesis refers to that, given an image of a target object taken at a first view, the image of the target object at other views can be obtained through the technical solutions provided by the embodiments of the present application. For example, the terminal trains an image rendering model based on the images of the target object taken at different views, and after the training is completed, the image of the target object taken at any angle can be obtained. The new action image synthesis refers to that, given an image of a target object performing a first action, the image of the target object performing other actions can be obtained through the technical solutions provided by the embodiments of the present application. For example, the terminal trains an image rendering model based on the images of the target object taken at different views, and after the training is completed, the image of the target object performing other actions can be obtained.
[0090] After introducing the implementation environment and the application scenarios of the embodiments of the present application, the image rendering method provided by the embodiments of the present application will be described.
[0091] It should be noted that, in the following description of the technical solutions provided by the present application, the server is taken as an example of the execution subject. In other possible implementation manners, the terminal can also be taken as the execution subject to execute the technical solutions provided by the present application, and the embodiments of the present application do not limit the type of the execution subject.
[0092] Figure 2 is a flowchart of an image rendering method provided by an embodiment of the present application, referring to Figure 2 , the method comprises the following steps.
[0093] 201. The server obtains shape parameters and pose parameters of a sample object based on a sample video frame, wherein the sample video frame comprises the sample object.
[0094] The sample object can be a human body, an animal, a plant, a building, a vehicle, or the like, and the embodiments of the present application do not limit the sample object. The sample video frame belongs to a sample video, and the sample video frame comprising the sample object means that the sample object is displayed in the sample video frame, and accordingly, the sample video is a video of the sample object.
[0095] 202. The server performs three-dimensional reconstruction on the sample object based on the shape parameters and the pose parameters, to obtain a three-dimensional model of the sample object.
[0096] The shape parameter is used to describe the shape of the sample object, and the pose parameter is used to describe the pose of the sample object. If the sample object is a human body, the shape refers to the body shape of the human body, for example, the shape parameter is used to describe the height, weight, and other body shapes of the human body. The pose refers to the action of the human body, for example, the pose parameter is used to describe the action of the human body, such as opening arms, bending over, or lifting legs.
[0097] 203. The server determines a plurality of first rendering parameters based on the camera parameters, the shape parameters, and the pose parameters of the virtual camera by using the image rendering model, wherein the camera parameters of the virtual camera are the same as the camera parameters of the real camera that captures the sample video frame.
[0098] The camera parameters are used to describe the related attributes of the virtual camera, for example, the camera parameters include the focal length parameter of the virtual camera, the size parameter of the captured image, and the position parameter of the camera, and the like. The focal length parameter is used to describe the focal length of the virtual camera. The size parameter of the captured image is used to describe the height and width of the image captured by the virtual camera. The position parameter is used to describe the position of the virtual camera. In some embodiments, the focal length parameter and the size parameter of the captured image of the virtual camera are also referred to as the intrinsic parameters of the virtual camera, and the position parameter of the virtual camera is also referred to as the extrinsic parameters of the virtual camera. The first rendering parameter is a parameter used for rendering an image.
[0099] 204. The server renders the first target image based on the plurality of first rendering parameters by using the image rendering model, and outputs a first rendered image, wherein the first target image is an image captured by the virtual camera from the three-dimensional model.
[0100] The first target image captured by the virtual camera from the three-dimensional model means that the first target image is an image captured by the virtual camera from the three-dimensional model at a specific position and angle, wherein the specific position and angle are determined by the camera parameters of the virtual camera. In some embodiments, the position of the virtual camera is equivalent to the position of the human eye, and the capturing of the three-dimensional model by the virtual camera is equivalent to the observation of the three-dimensional model by the human eye.
[0101] 205. The server trains the image rendering model based on the difference information between the sample video frame and the first rendered image, wherein the image rendering model is used to render the image captured by the virtual camera.
[0102] The first rendered image is an image obtained by the server through the image rendering model, and the sample video frame is an original image. Therefore, when training the image rendering model, the image rendering model is trained based on the sample video frame as supervision. The purpose of the training is to make the image rendered by the image rendering model as close as possible to the corresponding sample video frame.
[0103] By means of the technical solutions provided in the embodiments of the present application, when the image rendering model is trained, the sample object is reconstructed in three dimensions based on the shape parameters and the pose parameters of the sample object, the rendering parameters are determined based on the camera parameters, the shape parameters and the pose parameters, the influence of the virtual camera and the sample object is considered when the rendering parameters are determined, the rendering parameters are more matched with the three-dimensional model of the sample object, the first target image is rendered based on the rendering parameters, the first rendered image is obtained, and the image rendering model is trained based on the difference information between the first rendered image and the sample video frame. The image rendering model obtained in this way has stronger image rendering capability.
[0104] The above steps 201-205 are a simple description of the image rendering method provided in the embodiments of the present application. In the following, the image rendering method provided in the embodiments of the present application will be described in more detail in combination with some examples. Figure 3 FIG. 1 is a flowchart of an image rendering method provided in an embodiment of the present application. Referring to FIG. 1, Figure 3 Taking a server as an example, the method comprises the following steps.
[0105] 301. The server acquires a sample video, the sample video comprising a plurality of sample video frames, and the sample video frames comprising a sample object.
[0106] The sample object can be a human body, an animal, a plant, a building or a vehicle, and the embodiments of the present application do not limit the sample object. The sample video frame belongs to the sample video, and the sample video frame comprising the sample object means that the sample object is displayed in the sample video frame. Correspondingly, the sample video is a video of the sample object.
[0107] In a possible implementation, a camera is used to capture the sample object to obtain a sample video. The camera uploads the captured sample video to the server, and the server acquires the sample video. In some embodiments, the camera is a monocular camera, and correspondingly, the sample video is a monocular video, and correspondingly, the sample video frame is a monocular video frame.
[0108] For example, a camera is used to capture the sample object from all sides to obtain video frames of the sample object at different angles, the plurality of video frames constitute a sample video, and the camera uploads the sample video to the server, and the server acquires the sample video. By capturing the sample object from all sides, video frames of the sample object at different angles can be obtained, which helps to improve the modeling effect of the server on the sample object.
[0109] The camera capturing the sample object from all sides includes the following two ways:
[0110] In a manner 1, taking a human body as an example of a sample object, a user sets up a camera at a shooting position, a human body to be photographed moves into a shooting range of the camera, the camera photographs, the human body rotates 1-2 circles in the shooting range, and the camera photographs to obtain a sample video. In some embodiments, the human body can adopt a T-pose (T-shaped pose) or an A-pose (A-shaped pose) when rotating in the shooting range. The T-pose means that the human body horizontally spreads arms, and the two feet are close together to form a "T" shape. The A-pose means that the human body spreads arms downward at a certain angle with the body to form an "A" shape. After the photographing is completed, the camera uploads the sample video to a server, and the server obtains the sample video. In this manner, the position of the camera can be kept unchanged when the sample video is photographed, so that a more stable and clear sample video is obtained.
[0111] In a manner 2, taking a human body as an example of a sample object, a human body to be photographed moves to a specified position, a user holds a camera to surround the human body 1-2 circles, and simultaneously photographs the human body to obtain a sample video. In some embodiments, the human body can adopt a T-pose (T-shaped pose) or an A-pose (A-shaped pose) at the specified position. After the photographing is completed, the camera uploads the sample video to a server, and the server obtains the sample video. In this manner, the human body can be kept unchanged when the sample video is photographed, so that the posture of the human body in the sample video is more stable.
[0112] In the above description, a camera is taken as an ordinary camera for description. In other possible embodiments, the camera can also be a depth camera, such as a depth camera based on a structured light principle or a depth camera based on a time of flight (TOF) principle, and the like. The embodiments of the present application do not limit this. If the camera is a depth camera, when a sample object is photographed, not only image information of the sample object, that is, a video frame of a sample image, but also depth information of the sample object, that is, distances between different positions of the sample object and the camera, can be obtained. The server can subsequently model based on the depth information of the sample object.
[0113] In a possible embodiment, a server obtains a sample video from a network. The sample video is a video obtained by photographing a sample object by other users. In this case, a user does not need to re-photograph a sample video and then upload it to a server. The user can directly obtain the sample video from the network, which is more efficient.
[0114] In a possible implementation, the sample object is a sample three-dimensional model, and the sample video is a video obtained by using a virtual camera to shoot the sample three-dimensional model. For example, the server creates a three-dimensional model of a human body, uses a virtual camera to shoot the three-dimensional model of the human body from all directions, obtains video frames of the three-dimensional model of the human body from different angles, and the multiple video frames constitute the sample video.
[0115] In this implementation, when actual shooting is inconvenient, the server can also obtain the sample video by using a virtual camera to shoot the three-dimensional model, thereby improving the acquisition approach of the sample video and reducing the acquisition difficulty of the sample video.
[0116] 302. The server obtains shape parameters and pose parameters of the sample object based on the sample video frame.
[0117] In a possible implementation, the server performs shape estimation and pose estimation on the sample object based on the sample video frame, to obtain the shape parameters and the pose parameters of the sample object.
[0118] The above implementation is described below through two examples.
[0119] Example 1. The server extracts features from the sample video frame to obtain sample video frame features corresponding to the sample video frame. The server performs regression processing on the sample video frame features corresponding to the sample video frame to obtain the shape parameters and the pose parameters of the sample object. In this way, the server can directly obtain the shape parameters and the pose parameters of the sample object by performing feature extraction and regression processing on the sample video frame, which is efficient.
[0120] For example, the server inputs the sample video frame into a first parameter extraction model, performs convolution processing on the sample video frame through a convolutional layer of the first parameter extraction model to obtain sample video frame features corresponding to the sample video frame. The server performs full connection processing on the sample video frame features corresponding to the sample video frame through a regression layer of the first parameter extraction model, and maps the sample video frame features to the shape parameters and the pose parameters of the sample object. In some embodiments, the first parameter extraction model is an HMR (Human Mesh Recovery) model.
[0121] In an embodiment, the server inputs the sample video into a second parameter extraction model, and performs convolution processing on the sample video through a convolutional layer of the second parameter extraction model to obtain sample video frame features corresponding to the sample video frames respectively. The server performs time coding on the sample video frame features based on the arrangement order of the sample video frames in the sample video through a time coding layer of the second parameter extraction model to obtain time coding features of each sample video frame. In some embodiments, the time coding layer is a GRU (Gate Recurrent Unit). The server performs full connection processing on the time coding features corresponding to the sample video frames through a regression layer of the second parameter extraction model, and maps the time coding features to the shape parameters and the pose parameters of the sample object.
[0122] For example, the server inputs the sample video into a second parameter extraction model, and performs convolution processing on the sample video through a convolutional layer of the second parameter extraction model to obtain sample video frame features corresponding to the sample video frames respectively. The server performs time coding on the sample video frame features based on the arrangement order of the sample video frames in the sample video through a time coding layer of the second parameter extraction model to obtain time coding features of each sample video frame. In some embodiments, the time coding layer is a GRU (Gate Recurrent Unit). The server performs full connection processing on the time coding features corresponding to the sample video frames through a regression layer of the second parameter extraction model, and maps the time coding features to the shape parameters and the pose parameters of the sample object. In some embodiments, the second parameter extraction model is a VIBE (Video Inference for Human Body Pose and Shape Estimation) model.
[0123] In a possible implementation, the server performs image segmentation on the sample video frame to obtain a target region, and the target region is a region where the sample object is located. The server performs shape estimation and pose estimation on the target region to obtain the shape parameters and the pose parameters of the sample object.
[0124] In this embodiment, the server can perform image segmentation on the sample video frame before performing shape estimation and pose estimation on the sample video frame, determine the region where the sample object is located from the sample video frame, and perform shape estimation and pose estimation on the region where the sample object is located. This can avoid the influence of the background in the sample video frame on the shape estimation and the pose estimation, and improve the accuracy of the shape parameters and the pose parameters.
[0125] The above embodiments are described below by two examples.
[0126] Example 1, the server determines a plurality of candidate boxes (Region Proposals) on the sample video frame, respectively extracts features of the plurality of candidate boxes to obtain a plurality of candidate box features, and one candidate box feature corresponds to one candidate box. The server performs full connection processing and normalization processing on the plurality of candidate box features to obtain the probability of whether the sample object is contained in each candidate box. The server splices the candidate boxes whose probabilities meet the target probability condition to obtain the target region. The server extracts features of the target region to obtain the region feature corresponding to the target region. The server performs regression processing on the region feature corresponding to the target region to obtain the shape parameter and the pose parameter of the sample object.
[0127] For example, the server inputs the sample video frame into the first image segmentation model, generates a plurality of candidate boxes on the sample video frame through the candidate box generation layer of the first image segmentation model. The server performs convolution processing on the images in the plurality of candidate boxes through the convolution layer of the first image segmentation model to obtain the candidate box features corresponding to each candidate box. The server classifies the candidate box features corresponding to each candidate box through the classification layer of the first image segmentation model to obtain the probability of whether the sample object is contained in each candidate box. The server splices the candidate boxes whose probabilities are greater than or equal to the probability threshold to obtain the target region, and the target region is also the region including the target object. The server inputs the target region into the first parameter extraction model, performs convolution processing on the target region through the convolution layer of the first parameter extraction model to obtain the target region feature corresponding to the target region. The server performs full connection processing on the target region feature corresponding to the target region through the regression layer of the first parameter extraction model, and maps the target region feature to the shape parameter and the pose parameter of the sample object. In some embodiments, the first segmentation model is R-CNN (Region-CNN, Region Convolutional Neural Network).
[0128] In addition, the server can also perform image segmentation on the plurality of sample video frames in the sample video to obtain a target region in each sample video frame, where the method for the server to perform image segmentation on the plurality of sample video frames belongs to the same inventive concept as described above, and the implementation process is described previously and will not be repeated here. After obtaining the plurality of target regions, the server can input the plurality of target regions into the second parameter extraction model, perform convolution processing on the sample video through the convolution layer of the second parameter extraction model, and obtain sample video frame features corresponding to the plurality of sample video frames respectively. The server performs time coding on the plurality of sample video frame features based on the arrangement order of the plurality of sample video frames in the sample video through the time coding layer of the second parameter extraction model, and obtains time coding features of each sample video frame. In some embodiments, the time coding layer is a GRU.
[0129] In example 2, the server performs feature extraction on the sample video frame to obtain a feature map of the sample video frame. The server determines a plurality of candidate boxes (Region Proposals) on the feature map. The server performs full connection processing and normalization processing on the feature map corresponding to the plurality of candidate boxes to obtain a probability of whether a sample object is contained in each candidate box. The server splices the candidate boxes whose probabilities meet a target probability condition to obtain a target region. The server performs feature extraction on the target region to obtain a region feature corresponding to the target region. The server performs regression processing on the region feature corresponding to the target region to obtain a shape parameter and a pose parameter of the sample object.
[0130] For example, the server inputs the sample video frame into the second image segmentation model, performs convolution processing on the sample video frame through the convolution layer of the second image segmentation model, and obtains a feature map corresponding to the sample video frame. The server generates a plurality of candidate boxes on the feature map corresponding to the sample video frame through the candidate box generation layer of the second image segmentation model. The server classifies the feature map corresponding to each candidate box through the classification layer of the second image segmentation model to obtain a probability of whether a sample object is contained in each candidate box. The server splices the candidate boxes whose probabilities are greater than or equal to a probability threshold to obtain a target region, which is a region including a target object. The server inputs the target region into the first parameter extraction model, performs convolution processing on the target region through the convolution layer of the first parameter extraction model, and obtains a target region feature corresponding to the target region. The server performs full connection processing on the target region feature corresponding to the target region through the regression layer of the first parameter extraction model, and maps the target region feature to a shape parameter and a pose parameter of the sample object. In some embodiments, the second segmentation model is a FastR-CNN (Fast Region-CNN).
[0131] In addition, the server can also perform image segmentation on the plurality of sample video frames in the sample video to obtain a target region in each sample video frame, where the method for the server to perform image segmentation on the plurality of sample video frames belongs to the same inventive concept as described above, and the implementation process is described previously and will not be repeated here. After obtaining the plurality of target regions, the server can input the plurality of target regions into the second parameter extraction model respectively, perform convolution processing on the sample video through the convolution layer of the second parameter extraction model, and obtain sample video frame features corresponding to the plurality of sample video frames respectively. The server performs time coding on the plurality of sample video frame features based on the arrangement order of the plurality of sample video frames in the sample video through the time coding layer of the second parameter extraction model, and obtains time coding features of each sample video frame. In some embodiments, the time coding layer is a GRU. The server performs full connection processing on the time coding features corresponding to the sample video frame through the regression layer of the second parameter extraction model, and maps the time coding features to the shape parameter and the pose parameter of the sample object.
[0132] In a possible implementation, the server performs pose estimation on the sample object based on the sample video frame to obtain a pose parameter of the sample object. The server performs shape estimation on the sample object based on a plurality of video frames in the sample video to obtain a plurality of reference shape parameters of the sample object, one reference shape parameter corresponding to one video frame. The sample video includes the sample video frame. The server determines a shape parameter of the sample object based on the plurality of reference shape parameters.
[0133] In this case, the shape parameter of the same sample object should be the same. The server combines the shape parameters determined through the plurality of sample video frames when obtaining the shape parameter of the sample object, so that the determined shape parameter is more accurate.
[0134] The method for the server to perform pose estimation on the sample object based on the sample video frame to obtain a pose parameter of the sample object, and the method for the server to obtain a plurality of reference shape parameters belong to the same inventive concept as described previously, and the implementation process is described previously and will not be repeated here. For the server to determine a shape parameter of the sample object based on a plurality of reference shape parameters, the server obtains an average value of the plurality of reference shape parameters, and takes the average value as the shape parameter of the sample object. For example, the server determines the shape parameter of the sample object based on the plurality of reference shape parameters through the following formula (1).
[0135]
[0136] wherein β is the shape parameter of the sample object, is a reference shape parameter numbered t, and n is the number of reference shape parameters, t and n are both positive integers.
[0137] The principle behind the above implementation is as follows: Experiments revealed that the estimated shape parameters did not effectively align the sample object and the 3D model. Inaccurate shape parameters have a very negative impact on image rendering, often resulting in blurry images. To avoid this problem, the shape parameters were fine-tuned during training; that is, the average of multiple reference shape parameters was used as the shape parameters of the sample object. Experiments showed that the fine-tuned shape parameters better aligned the sample object and the 3D model, leading to clearer results.
[0138] 303. The server performs 3D reconstruction of the sample object based on shape and pose parameters to obtain a 3D model of the sample object.
[0139] Among them, shape parameters describe the shape of the sample object, and pose parameters describe the pose of the sample object. If the sample object is a human body, then shape refers to the body shape, such as the shape parameter describing the height, weight, and other body types. Pose refers to the human body's movements, such as the pose parameter describing movements like arms outstretched, bending over, or raising a leg.
[0140] In one possible implementation, the server inputs shape and pose parameters into a baseline 3D model, which is trained based on the shape and pose parameters of multiple objects. The server adjusts the shape of the baseline 3D model using the shape parameters and adjusts its pose using the pose parameters to obtain a 3D model of the sample object.
[0141] To provide a clearer explanation of the above embodiments, the reference three-dimensional model in the above embodiments will be introduced below.
[0142] In some embodiments, the reference three-dimensional model, also referred to as a standard SMPL (Skinned Multi-Person Linear) model, corresponds to a standard-shaped human body, and includes 6890 vertices and 23 joints, one joint having a binding relationship with multiple vertices, i.e., the movement of one joint can drive the movement of multiple vertices. In some embodiments, since the distance between the vertices having a binding relationship with the joint and the joint is different, or one vertex has a binding relationship with multiple joints, the binding relationship between the joint and the vertex can be embodied in the form of a skinning weight, for example, the higher the weight, the stronger the binding relationship between the joint and the vertex, and the lower the weight, the weaker the binding relationship between the joint and the vertex. In some embodiments, the server uses a matrix to record the vertex coordinates and joint coordinates of the reference three-dimensional model, for example, uses matrix T to record the coordinates of 6890 vertices, since the coordinates of the vertices are three-dimensional coordinates, the size of matrix T is 6890x3. For example, the server uses matrix J to record the coordinates of the joints, since the coordinates of the joints are three-dimensional coordinates, the size of matrix T is (23+1)x3, wherein 1 refers to recording the root node of the reference three-dimensional model. For example, the server uses matrix W to record the skinning weight between the joints and different vertices, and the size of matrix W is 24x6890.
[0143] In some embodiments, the reference three-dimensional model is trained based on the shape parameters and pose parameters of multiple objects, or in other words, the reference three-dimensional model can embody the average shape and average pose of multiple objects. After the training of the reference three-dimensional model is completed, the reference three-dimensional model can automatically adjust the shape or pose after inputting the shape parameters or pose parameters, i.e., adjusting matrix J or matrix T. In the above implementation, the server inputs the shape parameters of the sample object and the pose parameters of the sample object into the reference three-dimensional model. The server adjusts the shape of the reference three-dimensional model through the shape parameters, i.e., adjusts the vertex coordinates recorded in matrix T; adjusts the pose of the reference three-dimensional model through the pose parameters, i.e., adjusts the joint coordinates recorded in matrix J, to obtain the three-dimensional model of the sample object.
[0144] It should be noted that in the above description, the reference three-dimensional model is taken as the SMPL model as an example for description, and in other possible embodiments, the reference three-dimensional model can also be implemented as other types of models, such as an SMPLH (Skinned Multi-Person Linear Hand, multi-person linear hand), an SMPLX (Skinned Multi-Person Linear X, multi-person linear multi-part), and a STAR (A Sparse Trained Articulated Human Body Regressor, a sparse trained articulated human body regressor) model, and the embodiments of the present application do not limit this.
[0145] 304、The server obtains camera parameters of the virtual camera based on the sample video frame, and the camera parameters of the virtual camera are the same as camera parameters of a real camera that captures the sample video frame.
[0146] The camera parameters are used to describe the related attributes of the virtual camera, such as the camera parameters including a focal length parameter of the virtual camera, a size parameter of a captured image, and a position parameter of the camera, and the like, wherein the focal length parameter is used to describe the focal length of the virtual camera, the size parameter of the captured image is used to describe the height and width of the image captured by the virtual camera, and the position parameter is used to describe the position of the virtual camera. In some embodiments, the focal length parameter and the size parameter of the captured image of the virtual camera are also referred to as the intrinsic parameters of the virtual camera, and the position parameter of the virtual camera is also referred to as the extrinsic parameters of the virtual camera.
[0147] In some embodiments, the position of the virtual camera is equivalent to the position of the human eye, and the virtual camera capturing the three-dimensional model is equivalent to the human eye observing the three-dimensional model. The camera parameters of the virtual camera being the same as the camera parameters of the real camera that captures the sample video frame means that, for example, the real camera captures the sample video frame at a position 3 meters in front of the sample object, and then for the virtual camera, if the three-dimensional model is placed in a first virtual space, the virtual camera is also located at a position 3 meters in front of the three-dimensional model. The first virtual space is a virtual space associated with the camera parameters of the virtual camera, and in some embodiments, the first virtual space is also referred to as an observation space (Observation Space).
[0148] In one possible implementation, the server extracts features of the sample video frame to obtain sample video frame features corresponding to the sample video frame. The server performs regression processing on the sample video frame features corresponding to the sample video frame to obtain camera parameters corresponding to the sample video frame.
[0149] For example, the server inputs the sample video frame into the third parameter extraction model, and performs convolution processing on the sample video frame through the convolution layer of the third parameter extraction model to obtain the sample video frame feature corresponding to the sample video frame. The server performs full connection processing on the sample video frame feature corresponding to the sample video frame through the regression layer of the third parameter extraction model, and maps the sample video frame feature to the camera parameter corresponding to the sample video frame. In some embodiments, the camera parameter can also be obtained in step 302, that is, the server can obtain the camera parameter while obtaining the shape parameter and the pose parameter of the sample object based on the sample video frame. In this implementation, the third parameter extraction model is also the first parameter extraction model in step 302, that is, the server extracts the features of the sample video frame to obtain the sample video frame feature corresponding to the sample video frame. The server performs regression processing on the sample video frame feature corresponding to the sample video frame to obtain the shape parameter, the pose parameter of the sample object, and the camera parameter corresponding to the sample video frame. In this way, the server can directly obtain the shape parameter, the pose parameter of the sample object, and the camera parameter by performing feature extraction and regression processing on the sample video frame, which is more efficient.
[0150] In a possible implementation, the server extracts features of the sample video to obtain a plurality of sample video frame features, one sample video frame feature corresponding to one sample video frame in the sample video. The server performs time coding on the plurality of sample video frame features based on the arrangement order of the plurality of sample video frames in the sample video to obtain the time coding feature of each sample video frame, the time coding feature of one sample video frame fusing the sample video frame feature of the sample video frame and the sample video frame features of the plurality of sample video frames before the sample video frame. The server performs regression processing on the sample video frame feature of each sample video frame to obtain the camera parameter corresponding to the sample video frame. In this implementation, the server combines the camera parameters corresponding to the sample video frames before each sample video frame when obtaining the camera parameter corresponding to the sample video frame, which can improve the accuracy of determining the camera parameter by the server.
[0151] For example, the server inputs the sample video into the fourth parameter extraction model, performs convolution processing on the sample video through the convolution layer of the fourth parameter extraction model, and obtains sample video frame features corresponding to multiple sample video frames respectively. The server performs time coding on the multiple sample video frame features based on the arrangement order of the multiple sample video frames in the sample video through the time coding layer of the fourth parameter extraction model, and obtains time coding features of each sample video frame. In some embodiments, the time coding layer is a GRU. The server performs full connection processing on the time coding features corresponding to the sample video frame through the Regression layer of the fourth parameter extraction model, and maps the time coding features to the camera parameters corresponding to the sample video frame. In some embodiments, the camera parameters can also be obtained in step 302, that is, the server can obtain the camera parameters while obtaining the shape parameters and the pose parameters of the sample object based on the sample video frame. In this implementation, the fourth parameter extraction model is also the second parameter extraction model in step 302, that is, the server performs feature extraction on the sample video to obtain multiple sample video frame features, and each sample video frame feature corresponds to a sample video frame in the sample video. The server performs time coding on the multiple sample video frame features based on the arrangement order of the multiple sample video frames in the sample video, and obtains time coding features of each sample video frame. The time coding features of a sample video frame fuse the sample video frame features of the sample video frame and the sample video frame features of multiple sample video frames before the sample video frame. The server performs Regression processing on the sample video frame features of each sample video frame to obtain the shape parameters, the pose parameters of the sample object, and the camera parameters corresponding to the sample video frame.
[0152] 305. The server inputs the camera parameters, the shape parameters, and the pose parameters into an image rendering model, and the image rendering model is used to render an image captured by a virtual camera.
[0153] The image rendering model is used to determine rendering parameters based on the camera parameters, the shape parameters, and the pose parameters, render the image captured by the virtual camera based on the rendering parameters, and generate a rendered image.
[0154] 306. The server determines multiple first rendering parameters based on the camera parameters, the shape parameters, and the pose parameters of the virtual camera through the image rendering model.
[0155] The first rendering parameters are parameters used to render the image.
[0156] In a possible implementation, the camera parameter of the virtual camera includes a position parameter of the virtual camera, the server determines at least one virtual ray in the first virtual space based on the position parameter and a perspective of the virtual camera on the three-dimensional model, the virtual ray is a line connecting the virtual camera and a pixel point on the first target image, and the first virtual space is a virtual space established based on the camera parameter. The server determines the plurality of first rendering parameters based on coordinates of a plurality of first sampling points on the at least one virtual ray, the shape parameter, and the pose parameter, the coordinates of the first sampling point being coordinates of the first sampling point in the first virtual space.
[0157] In some embodiments, the continuous integration can be obtained by sampling a plurality of sampling points between the near plane and the far plane along the camera ray.
[0158] To make the above-mentioned embodiments clearer, the following will be divided into two parts to explain the above-mentioned embodiments.
[0159] The first part is to explain the method for determining, by the server, at least one virtual ray in the first virtual space based on the position parameter and the perspective of the virtual camera on the three-dimensional model.
[0160] In a possible implementation, the position parameter is a coordinate of the virtual camera in the first virtual space, the server determines the first target image in the first virtual space based on the coordinate of the virtual camera in the first virtual space and the perspective of the virtual camera on the three-dimensional model, the first target image being an image of the three-dimensional model at a perspective. The server connects the optical center of the virtual camera with at least one pixel point on the first target image to obtain at least one virtual ray. In some embodiments, the first target image includes a plurality of pixel points, and the server can obtain a plurality of virtual rays, each virtual ray corresponding to a perspective of observing the three-dimensional model.
[0161] The second part is to explain the method for determining, by the server, the plurality of first rendering parameters based on the coordinates of the plurality of first sampling points on the at least one virtual ray, the shape parameter, and the pose parameter.
[0162] In a possible implementation, the server transforms the plurality of first sampling points into a second virtual space based on the coordinates of the plurality of first sampling points, the pose parameter, and a reference pose parameter to obtain a plurality of second sampling points, one first sampling point corresponding to one second sampling point, the reference pose parameter being a pose parameter corresponding to the second virtual space, and the coordinates of the second sampling point being coordinates of the second sampling point in the second virtual space. The server determines the plurality of first rendering parameters based on the coordinates of the plurality of second sampling points in the second virtual space, the shape parameter, and the pose parameter.
[0163] The first sampling point refers to a sampling point on a virtual ray in the first virtual space, and the number of the first sampling points is set by a technician according to an actual situation, and embodiments of the present application do not limit this. In some embodiments, in order to guarantee the representativeness of the plurality of first sampling points to the virtual ray, the server can divide the virtual ray into a plurality of parts on average, and randomly determine a first sampling point on each part to obtain a plurality of first sampling points on the virtual ray. Since the first sampling point is determined by the server on the virtual ray, the server can determine the coordinates of the plurality of first sampling points based on the function of the virtual ray, that is, the function of the straight line between the optical center of the virtual camera and the corresponding pixel point on the first target image.
[0164] The method of the server for transforming the plurality of first sampling points into a plurality of second sampling points in the second virtual space will be described first.
[0165] In a possible implementation, for a first sampling point, the server obtains a first pose transformation matrix and a second pose transformation matrix of the first sampling point, the first pose transformation matrix is a transformation matrix of a first vertex from a first pose to a second pose, the second pose transformation matrix is a transformation matrix of the first vertex from the first pose to a third pose, the first pose is a reference pose, the second pose is a pose corresponding to a pose parameter, the third pose is a pose corresponding to a reference pose parameter, and the first vertex is a vertex on the three-dimensional model that meets a target condition with the first sampling point. The server obtains a second sampling point corresponding to the first sampling point based on a skin weight corresponding to the first vertex, the first pose transformation matrix and the second pose transformation matrix.
[0166] The reference pose is the pose of the reference three-dimensional model described in step 304, the second pose is the pose of the sample object, the third pose is the pose corresponding to the reference pose parameter, and the reference pose parameter is the pose parameter corresponding to the second virtual space. Therefore, the third pose parameter is the pose parameter corresponding to the standard three-dimensional model of the second virtual space. In some embodiments, the second virtual space is also referred to as a candidate space (Canonical Space).
[0167] For example, the server transforms the first sampling point in the first virtual space into the second sampling point in the second virtual space by the following formulas (2)-(5).
[0168]
[0169] x is the coordinate of the first sampling point, x 0 is the coordinate of the second sampling point, is a transformation function, θ t is the second pose, θ0 is the third pose, M j (θt ) refers to the first attitude change matrix, M j (θ0) refers to the second attitude change matrix, v j These are the coordinates of the first vertex, numbered j. This refers to the skinning weight of the first vertex that is closest to the first sampling point, b j This refers to the skinning weight ω of the first vertex with the number j corresponding to the first sampling point. j ω refers to the change weight corresponding to the first vertex numbered j, ω refers to the sum of the change weights corresponding to multiple first vertices corresponding to the first sampling point, and N(x) refers to the set of multiple first vertices corresponding to the first sampling point. In some embodiments, Considering that a vertex may be affected by different body parts at the same time, resulting in ambiguous or meaningless movement, skinning weights are used to distinguish different body parts and to emphasize the influence of the nearest vertex.
[0170] The following describes the method by which the server determines multiple first rendering parameters based on the coordinates, shape parameters, and pose parameters of multiple second sampling points in the second virtual space.
[0171] In one possible implementation, for a second sampling point, the server concatenates the coordinates, shape parameters, and pose parameters of the second sampling point in a second virtual space to obtain a first parameter set. The server then performs a fully connected operation on the first parameter set to obtain the first rendering parameters.
[0172] For example, the server obtains the first rendering parameters using the following formula (6).
[0173] F(D(x,β t θ t ))=(c,σ) (6)
[0174] Where F is a fully connected function, and D(x, β) t θ t To transform the coordinates of the first sampling point x to the coordinates of the second sampling point x. 0 The transformation function of the coordinates, (c, σ) is the first rendering parameter, c is the color parameter, c = (r, g, b), and σ is the density parameter.
[0175] In some embodiments, the method for determining the first rendering parameter described above is also referred to as the method of Neural Radiance Fields (NeRF).
[0176] 307. The server renders the first target image based on multiple first rendering parameters using an image rendering model, and outputs the first rendered image. The first target image is an image obtained by a virtual camera capturing a 3D model.
[0177] In a possible implementation, for a pixel point on the first target image, the server determines a color and an opacity of the pixel point based on a virtual ray between the pixel point and the virtual camera and the first rendering parameter corresponding to the pixel point. The server renders the pixel point based on the color and the opacity, and outputs the rendered pixel point.
[0178] For example, the server integrates the first relationship data on the virtual ray to obtain the color, the first relationship data being associated with a color parameter and a density parameter. The server integrates the second relationship data on the virtual ray to obtain the opacity, the second relationship data being associated with the density parameter. For example, the server obtains the color by using formula (7) and formula (9), and obtains the opacity by using formula (8) and formula (9).
[0179]
[0180] wherein, is the color, is the opacity, x k is the second sampling point, c t is the color parameter, σ t is the density, T k is the probability that the ray reaches x k , η t (x k ) is a prior mask used to provide geometric priors and handle ambiguities that can arise during the deformation process, δ k =‖x k+1 -x k ‖, is the distance between adjacent sampling points.
[0181] In some embodiments, the server can employ a coarse network and a fine network to determine the color of the pixel points respectively, wherein the coarse network and the fine network are determined according to the number of the first sampling points, for the same virtual ray, the server can determine different numbers of the first sampling points on the virtual ray, if the server determines a larger number of the first sampling points on the virtual ray, the larger number of the first sampling points can be used to construct the fine network, if the server determines a smaller number of the first sampling points on the virtual ray, the smaller number of the first sampling points can be used to construct the coarse network. Subsequently, the server can train the image rendering model through the coarse network and the fine network. That is, the sample object is expressed using the coarse network and the fine network. In the experiment, 64 points are uniformly sampled for each ray of the coarse network, and 64+16 points are sampled again according to the density distribution of the results of the coarse network for the fine network. For each first target image, 1024 rays are randomly sampled, 80% of the rays are sampled in the foreground part, and the remaining 20% of the rays are sampled in the background part. In the experiment, the other hyperparameters are set as follows: |N(i)| = 4, δ = 0.2, λ1 = 0.001, λ2 = 0.01, λ d = 0.1, and the resolution of the picture in all experiments is 512x512.
[0182] In some embodiments, the sampling points far away from the three-dimensional model have little influence on the image rendering, in this case, a priori geometry for accelerating the training of the model is introduced: the density of the points far away from the surface of the body should be zero, which is determined by the following formula (10) and formula (11).
[0183] η t (x k )=d(x k )≤δ (10)
[0184]
[0185] Wherein, d(x k ) is the distance between the first sampling point x k and the first vertex, and δ is the distance threshold for limiting the first sampling point to the first vertex.
[0186] 308、The server trains the image rendering model based on the difference information between the sample video frame and the first rendered image.
[0187] In a possible implementation, the server constructs a first loss function based on the color difference information between the sample video frame and the first rendered image, and trains the image rendering model using the first loss function.
[0188] For example, the server constructs a first loss function based on color difference information between the color determined by the coarse network and the color of the sample video frame, and color difference information between the color determined by the fine network and the color of the sample video frame. For example, the server constructs the first loss function by the following formula (12).
[0189]
[0190] wherein L c is the first loss function, C t (r) is the color of the sample video frame, is the color determined by the coarse network, is the color determined by the fine network.
[0191] In a possible implementation, the server performs regularization constraint on the pose parameters of adjacent sample video frames to ensure that the pose parameters between adjacent sample video frames are as close as possible, and to ensure that the optimized pose parameters and the pre-optimized pose parameters are not too different. That is, in order to obtain stable and smooth pose parameters, a second loss function is added in the image rendering process to optimize the pose parameters and the initial parameters not to be too different, and the pose parameters between adjacent frames should be as close as possible, wherein the optimized pose parameters refer to the average pose parameters of multiple sample video frames in the sample video. For example, the server performs regularization constraint by the following formula (13).
[0192]
[0193] wherein L p is the second loss function, λ1 and λ2 are weights, is the optimized pose parameter of the sample video frame with the serial number t, θ t is the pre-optimized pose parameter of the sample video frame with the serial number t, is the optimized pose parameter of the sample video frame with the serial number t+1.
[0194] In a possible implementation, the server constructs a third loss function based on opacity difference information between the opacity of the sample video frame and the opacity of the first rendered image, and adopts the third loss function to train the image rendering model. In some embodiments, this implementation is also referred to as background regularization.
[0195] For example, the server constructs the third loss function based on opacity difference information between the opacity determined by the coarse network and the opacity of the sample video frame, and opacity difference information between the opacity determined by the fine network and the opacity of the sample video frame. For example, the server constructs the third loss function by the following formula (14).
[0196]
[0197] wherein L c is a third loss function, D t (r) is the opacity of the sample video frame, is the opacity determined by the coarse network, is the opacity determined by the fine network.
[0198] In a possible implementation, the server can employ at least two of the above three loss functions to train the image rendering model, such as the server employs all the three loss functions to train the image rendering model, that is, the image rendering model is trained by the following formula (15).
[0199] L = L c + L p + λ d *L d (15)
[0200] wherein L is a joint loss function, λ d is a weight.
[0201] During the experiment, experiments were conducted on two different data sets: the People-Snapshot data set and the iPER data set contain monocular videos of target human bodies in real scenes keeping A-pose turning around. In the experiment, 200 images were uniformly selected from each video for training, and the target person turned about 1-2 circles.
[0202] The above steps 301-308 will be described below in conjunction with Figure 4 .
[0203] Referring to Figure 4 , the server obtains a sample video frame 401, obtains pose parameters, shape parameters of a sample object, and camera parameters corresponding to the sample video frame based on the sample video frame 401. The server places a three-dimensional model of the sample object in an observation space 402, and obtains coordinates of a first sampling point on a plurality of virtual rays. The server transforms the first sampling point into a candidate space 403, inputs the pose parameters, the shape parameters of the sample object, and the coordinates of the second sampling point into the image rendering model, determines rendering parameters based on the pose parameters, the shape parameters of the sample object, and the coordinates of the second sampling point, and in some embodiments, the server determines the rendering parameters based on a neural radiance field method, that is, determines the rendering parameters based on a multi-layer perception (MLP). The server renders a first target image based on the rendering parameters to obtain a first rendered image 404. The server trains the image rendering model based on the difference information between the sample video frame 401 and the first rendered image 404.
[0204] All the optional technical solutions described above can be combined to form optional embodiments of the present application, which will not be described again.
[0205] Through the technical solutions provided by the embodiments of the present application, when training the image rendering model, the sample object is reconstructed in three dimensions based on the shape parameters and the pose parameters of the sample object, the rendering parameters are determined based on the camera parameters, the shape parameters and the pose parameters, the influence of the virtual camera and the sample object is considered when determining the rendering parameters, the rendering parameters are more matched with the three-dimensional model of the sample object, the first target image is rendered based on the rendering parameters to obtain a first rendered image, and the image rendering model is trained based on the difference information between the first rendered image and the sample video frame. The image rendering model obtained in this way has stronger image rendering capability.
[0206] After the server trains the image rendering model through the above steps 301-308, the method of using the image rendering model will be described below, referring to Figure 5 , the method comprises:
[0207] 501. The server adjusts the pose of the three-dimensional model of the sample object based on the target pose parameters to obtain a three-dimensional model after pose adjustment.
[0208] The target pose parameters are the pose parameters input by the user, and the target pose parameters are used to change the pose of the sample object.
[0209] In one possible implementation, the server obtains a reference image, and the reference image includes a target object. The server estimates the pose of the target object based on the reference image to obtain target pose parameters of the target object. The server adjusts the pose of the three-dimensional model of the sample object based on the target pose parameters to obtain a three-dimensional model after pose adjustment. In some embodiments, the reference image is an image uploaded by the user to the server through the terminal. In this case, when the user wants to adjust the pose of the sample object to the pose of the target object, the user only needs to upload the reference image containing the target object to the server, and the server estimates the pose of the reference image, and adjusts the three-dimensional model of the sample object based on the target pose parameters of the target object.
[0210] For example, the server inputs the reference image into the first parameter extraction model, convolves the reference image through the convolution layer of the first parameter extraction model to obtain reference image features corresponding to the reference image. The server performs full connection processing on the reference image features through the regression layer of the first parameter extraction model, and maps the reference image features to the pose parameters of the target object. The server inputs the target pose parameters into the three-dimensional model of the sample object, adjusts the three-dimensional model of the sample object through the target pose parameters, and obtains a three-dimensional model after pose adjustment.
[0211] In a possible implementation, the terminal uploads the target pose parameter to the server, and the server acquires the target pose parameter. The server adjusts the pose of the three-dimensional model of the sample object based on the target pose parameter, and obtains a three-dimensional model after the pose adjustment. In some embodiments, the reference image is an image uploaded by the user to the server through the terminal. In this case, when the user wants to adjust the pose of the sample object to the pose corresponding to the target pose parameter, the user only needs to upload the target pose parameter to the server, and the server performs subsequent processing based on the target pose parameter.
[0212] 502. The server inputs the camera parameter, the shape parameter, and the target pose parameter into the trained image rendering model, and determines a plurality of second rendering parameters based on the camera parameter, the shape parameter, and the target pose parameter.
[0213] The method for determining the second rendering parameter by the server belongs to the same inventive concept as the method for determining the first rendering parameter by the server in step 306, and the implementation process is described with reference to the related description of step 306, which will not be described here again.
[0214] 503. The server renders the second target image based on the plurality of second rendering parameters, and outputs a second rendering image, the second target image being an image obtained by photographing the three-dimensional model after the pose adjustment by a virtual camera.
[0215] The method for rendering the second target image by the server based on the second rendering parameter belongs to the same inventive concept as the method for rendering the second target image by the server based on the first rendering parameter in step 307, and the implementation process is described with reference to the related description of step 307, which will not be described here again.
[0216] Referring to Figure 6 , the above steps 501-503 are used to change the pose of the sample object, and a second rendering image is obtained. In Figure 6 , 601-603 are three sample video frames, and each sample video frame is followed by a plurality of second rendering images, each second rendering image including a sample object after the pose adjustment.
[0217] Through the above steps 501-503, when the user wants to adjust the pose of the sample object to the pose of the target object, the user only needs to upload the reference image including the target object to the server, and the server performs pose estimation on the reference image to obtain the target pose parameter of the target object. The server adjusts the pose of the three-dimensional model of the sample object based on the target pose parameter, and obtains a three-dimensional model after the pose adjustment. The server can output a second rendering image based on the camera parameter, the target pose parameter, and the shape parameter.
[0218] The steps 501-503 are a method for adjusting a three-dimensional model of a sample object based on a target pose parameter and outputting a second rendered image. The following will describe a method for adjusting a three-dimensional model of a sample object based on a target shape parameter and outputting a third rendered image.
[0219] 701. The server adjusts the three-dimensional model of the sample object based on the target shape parameter to obtain a shape-adjusted three-dimensional model.
[0220] The target shape parameter is a shape parameter input by the user, and the target shape parameter is used to change the shape of the sample object.
[0221] In a possible implementation, the server obtains a reference image including a target object. The server performs shape estimation on the target object based on the reference image to obtain a target shape parameter of the target object. The server adjusts the three-dimensional model of the sample object based on the target shape parameter to obtain a shape-adjusted three-dimensional model. In some embodiments, the reference image is an image uploaded by the user to the server through the terminal. In this case, when the user wants to adjust the shape of the sample object to the shape of the target object, the user only needs to upload the reference image including the target object to the server, and the server performs shape estimation on the reference image and adjusts the three-dimensional model of the sample object based on the target shape parameter of the target object.
[0222] For example, the server inputs the reference image into the first parameter extraction model, performs convolution processing on the reference image through the convolution layer of the first parameter extraction model to obtain reference image features corresponding to the reference image. The server performs full connection processing on the reference image features through the regression layer of the first parameter extraction model, and maps the reference image features to the shape parameter of the target object. The server inputs the target shape parameter into the three-dimensional model of the sample object, adjusts the three-dimensional model of the sample object through the target shape parameter, and obtains a shape-adjusted three-dimensional model.
[0223] In a possible implementation, the terminal uploads a target shape parameter to the server, and the server obtains the target shape parameter. The server adjusts the three-dimensional model of the sample object based on the target shape parameter to obtain a shape-adjusted three-dimensional model. In some embodiments, the reference image is an image uploaded by the user to the server through the terminal. In this case, when the user wants to adjust the shape of the sample object to the shape corresponding to the target shape parameter, the user only needs to upload the target shape parameter to the server, and the server performs subsequent processing based on the target shape parameter.
[0224] 702. The server inputs the camera parameters, target shape parameters, and pose parameters into the trained image rendering model, and determines multiple third rendering parameters based on the camera parameters, target shape parameters, and pose parameters.
[0225] The method by which the server determines the third rendering parameter is the same inventive concept as the method by which the server determines the first rendering parameter in step 306 above. The implementation process is described in the relevant description of step 306 above and will not be repeated here.
[0226] 703. The server renders the third target image based on multiple third rendering parameters and outputs the third rendered image. The third target image is an image obtained by the virtual camera taking a picture of the 3D model after the pose has been adjusted.
[0227] The method by which the server renders the third target image based on the third rendering parameters is the same inventive concept as the method by which the server renders the third target image based on the first rendering parameters in step 307 above. The implementation process is described in the relevant description of step 307 above, and will not be repeated here.
[0228] Through steps 701-703 above, when a user wants to adjust the shape of a sample object to match the shape of a target object, they only need to upload a reference image containing the target object to the server. The server then performs shape estimation on the reference image to obtain the target shape parameters of the target object. Based on the target shape parameters, the server adjusts the shape of the 3D model of the sample object to obtain the 3D model with the adjusted shape. Based on the camera parameters, target shape parameters, and pose parameters, the server can output a third-party rendered image.
[0229] This application proposes a SMPL-based pose-guided deformation strategy for reconstructing neural radiation fields from dynamic monocular video scenes. This method relaxes the requirements for static scenes while preserving details such as clothing and hair. Considering the impact of inaccurate SMPL parameters, a strategy for jointly optimizing SMPL parameters and neural radiation fields is proposed, significantly improving the quality of 3D reconstruction. High-quality monocular video 3D human reconstruction is achieved, and high-quality images can be rendered from any viewpoint. Due to the controllable geometric deformation based on SMPL, this method can synthesize new actions for animation-driven applications.
[0230] Steps 701-703 above describe the method by which the server adjusts the pose of the 3D model of the sample object based on the target pose parameters and outputs the second rendered image. The following will explain the method by which the server outputs sample objects at different angles based on the target camera parameters.
[0231] 801、The server inputs the target camera parameter, the pose parameter and the target shape parameter into the trained image rendering model, determines a plurality of fourth rendering parameters based on the pose parameter, the shape parameter and the target camera parameter.
[0232] The method for determining the fourth rendering parameter by the server belongs to the same inventive concept as the method for determining the first rendering parameter by the server in step 306, and the implementation process is described in the related description of step 306, which will not be described here.
[0233] 802、The server renders the fourth target image based on the plurality of fourth rendering parameters, outputs a fourth rendered image, and the fourth target image is an image of the three-dimensional model taken by the virtual camera under the target camera parameter.
[0234] The method for rendering the first target image based on the third rendering parameter by the server belongs to the same inventive concept as the method for rendering the first target image based on the first rendering parameter by the server in step 307, and the implementation process is described in the related description of step 307, which will not be described here.
[0235] In some embodiments, after changing the camera parameter, the server can extract the three-dimensional model at different angles based on the Marching Cubes (MC) algorithm, as shown in Figure 9 901-903 are sample video frames, and the images behind the sample video frames are three-dimensional models of the sample object at different angles.
[0236] As shown in Figure 10 , the fourth rendered images generated based on the three-dimensional models at different angles are shown, wherein 1001-1006 are sample video frames, and the images behind the sample video frames are the fourth rendered images. That is, by using the trained image rendering model, the images rendered from two different perspectives while keeping the pose of the target human body (sample object) unchanged. From the results, the image rendering method provided in the embodiments of the application successfully reconstructs a high-quality static human body scene from a dynamic scene.
[0237] Through the above steps 801-802, when a user wants to obtain rendered images of a sample object at different angles, the camera parameter only needs to be changed, which is efficient.
[0238] Another image rendering method is provided in the embodiments of the application, as shown in Figure 11 Taking a terminal as an example, the method comprises the following steps.
[0239] 1101、The terminal displays a target video frame, and the target video frame comprises a target object.
[0240] The definitions of the target video frame and the target object are the same as those of the sample video frame and the sample object, and details are described above in relation to step 301, which will not be repeated here.
[0241] In a possible implementation, the terminal displays an image rendering interface, and the image rendering interface displays a target video selection control. In response to a click operation on the target video selection control, the terminal displays a target video selection interface, and the target video selection interface displays identifiers of a plurality of target videos. In response to an identifier corresponding to a target video being selected, the terminal displays a target video frame corresponding to the target video frame on the image rendering interface. In some embodiments, the target video frame is the first video frame of the target video frame.
[0242] For example, referring to Figure 12 , the terminal displays an image rendering interface 1201, and the image rendering interface displays a target video selection control 1202. In response to a click operation on the target video selection control 1202, the terminal displays a target video selection interface 1203, and the target video selection interface displays identifiers of a plurality of target videos. In response to an identifier 1204 corresponding to a target video being selected, the terminal displays a target video frame 1205 corresponding to the target video frame on the image rendering interface 1201.
[0243] 1102. In response to a three-dimensional reconstruction operation on the target object, the terminal displays a three-dimensional model of the target object, which is generated based on a shape parameter and a pose parameter of the target object, and the shape parameter and the pose parameter are determined based on the target video frame.
[0244] The generation method of the three-dimensional model of the target object belongs to the same inventive concept as the generation method of the sample object, and the implementation process is described above in relation to steps 302 and 303, which will not be repeated here.
[0245] In a possible implementation, in response to a click operation on a three-dimensional reconstruction control displayed on the image rendering interface, the terminal sends a three-dimensional model acquisition request to the server, and the three-dimensional model acquisition request carries a target video corresponding to the target video frame. In response to receiving the three-dimensional model acquisition request, the target video is acquired from the three-dimensional model acquisition request. The server performs three-dimensional reconstruction on the target object based on the target video to obtain a three-dimensional model of the target object. The server sends the three-dimensional model of the target object to the terminal, and the terminal displays the three-dimensional model of the target object on the image rendering interface.
[0246] For example, referring to Figure 13 , in response to a click operation on a three-dimensional reconstruction control 1302 displayed on an image rendering interface 1301, the terminal displays a three-dimensional model 1303 of a target object on the image rendering interface 1301.
[0247] In some embodiments, the terminal can also adjust the shape parameters and the pose parameters of the three-dimensional model when displaying the three-dimensional model of the target object, so as to change the shape and the pose of the three-dimensional model. If the shape parameters and the pose parameters of the three-dimensional model are adjusted, the terminal can also display the three-dimensional model generated based on the adjusted shape parameters and the adjusted pose parameters on the image rendering interface, and can subsequently perform rendering based on the adjusted shape parameters and the adjusted pose parameters.
[0248] 1103. In response to the shooting operation on the three-dimensional model, the terminal displays a first target image, which is an image obtained by shooting the three-dimensional model by a virtual camera.
[0249] In a possible implementation, in response to a click operation on a shooting control displayed on the image rendering interface, the terminal controls the virtual camera to shoot the three-dimensional model to obtain a first target image. The terminal displays the first target image on the image rendering interface.
[0250] For example, referring to Figure 13 In response to a click operation on a shooting control 1304 displayed on the image rendering interface 1301, the terminal displays a first target image 1305 on the image rendering interface 1301.
[0251] In some embodiments, when the terminal controls the virtual camera to shoot the three-dimensional model, the terminal can adjust the angle at which the virtual camera shoots the three-dimensional model, so as to obtain a first target image of the target object at different angles. Subsequent rendering of different first target images can obtain rendering images of the target object at different angles.
[0252] 1104. In response to a rendering operation on the first target image, the terminal displays a first rendering image, which is obtained by rendering the first target image based on a plurality of first rendering parameters determined by an image rendering model based on camera parameters, shape parameters and pose parameters of the virtual camera, the image rendering model being used to render images shot by the virtual camera.
[0253] In a possible implementation, in response to a click operation on a rendering control displayed on the image rendering interface, the terminal displays a first rendering image on the image rendering interface.
[0254] For example, referring to Figure 13 In response to a click operation on a rendering control 1306 displayed on the image rendering interface 1301, the terminal displays a first rendering image 1307 on the image rendering interface 1301.
[0255] By means of the technical scheme provided in the embodiments of the present application, when the first rendered image is generated, the target object is three-dimensionally reconstructed based on the shape parameter and the pose parameter of the target object, the rendering parameter is determined based on the camera parameter, the shape parameter and the pose parameter, the influence of the virtual camera and the target object is considered when the rendering parameter is determined, the rendering parameter is more matched with the three-dimensional model of the target object, the first target image is rendered based on the rendering parameter, and the first rendered image is obtained, and the effect of image rendering is better.
[0256] Figure 14 is a structural schematic diagram of an image rendering device provided by the embodiments of the present application, referring to Figure 14 The device comprises a parameter acquisition module 1401, a three-dimensional reconstruction module 1402, a rendering module 1403 and a training module 1404.
[0257] The parameter acquisition module 1401 is configured to acquire the shape parameter and the pose parameter of a sample object based on a sample video frame, wherein the sample video frame comprises the sample object.
[0258] The three-dimensional reconstruction module 1402 is configured to perform three-dimensional reconstruction on the sample object based on the shape parameter and the pose parameter, and obtain a three-dimensional model of the sample object.
[0259] The rendering module 1403 is configured to determine a plurality of first rendering parameters based on the camera parameter, the shape parameter and the pose parameter of a virtual camera by using an image rendering model, wherein the camera parameter of the virtual camera is the same as the camera parameter of an actual camera used for shooting the sample video frame; and render a first target image based on the plurality of first rendering parameters, and output a first rendered image, wherein the first target image is an image obtained by shooting the three-dimensional model by using the virtual camera.
[0260] The training module 1404 is configured to train the image rendering model based on difference information between the sample video frame and the first rendered image, wherein the image rendering model is used for rendering an image shot by the virtual camera.
[0261] In a possible implementation, the three-dimensional reconstruction module 1402 is configured to adjust the shape of a reference three-dimensional model by using the shape parameter, adjust the pose of the reference three-dimensional model by using the pose parameter, and obtain the three-dimensional model of the sample object, wherein the reference three-dimensional model is obtained by training based on the shape parameter and the pose parameter of a plurality of objects.
[0262] In a possible implementation, the parameter acquisition module 1401 is configured to perform shape estimation and pose estimation on the sample object based on the sample video frame, and obtain the shape parameter and the pose parameter of the sample object.
[0263] In a possible implementation, the device further comprises:
[0264] The region determination module is configured to perform image segmentation on the sample video frame to obtain a target region, the target region being a region in which the sample object is located.
[0265] The parameter acquisition module 1401 is configured to perform shape estimation and pose estimation on the target region to obtain a shape parameter and a pose parameter of the sample object.
[0266] In a possible implementation, the parameter acquisition module 1401 is configured to perform pose estimation on the sample object based on the sample video frame to obtain the pose parameter of the sample object. Shape estimation is performed on the sample object based on a plurality of video frames in the sample video to obtain a plurality of reference shape parameters of the sample object, one reference shape parameter corresponding to one video frame, the sample video including the sample video frame. The shape parameter of the sample object is determined based on the plurality of reference shape parameters.
[0267] In a possible implementation, the camera parameter includes a position parameter of the virtual camera in the first virtual space, and the rendering module 1403 is configured to determine at least one virtual ray in the first virtual space based on the position parameter and a perspective of the virtual camera on the three-dimensional model, the virtual ray being a line connecting the virtual camera and a pixel point on the first target image, the first virtual space being a virtual space established based on the camera parameter. A plurality of first rendering parameters are determined based on coordinates of a plurality of first sampling points on the at least one virtual ray, the shape parameter, and the pose parameter, the coordinates of the first sampling point being coordinates of the first sampling point in the first virtual space.
[0268] In a possible implementation, the rendering module 1403 is configured to transform the plurality of first sampling points into a second virtual space based on the coordinates of the plurality of first sampling points, the pose parameter, and a reference pose parameter to obtain a plurality of second sampling points, one first sampling point corresponding to one second sampling point, the reference pose parameter being a pose parameter corresponding to the second virtual space, the coordinates of the second sampling point being coordinates of the second sampling point in the second virtual space. The plurality of first rendering parameters are determined based on the coordinates of the plurality of second sampling points in the second virtual space, the shape parameter, and the pose parameter.
[0269] In a possible implementation, the rendering module 1403 is configured to, for a first sampling point, obtain a first pose transformation matrix and a second pose transformation matrix of the first sampling point, the first pose transformation matrix being a transformation matrix of a first vertex from a first pose to a second pose, the second pose transformation matrix being a transformation matrix of the first vertex from the first pose to a third pose, the first pose being a reference pose, the second pose being a pose corresponding to the pose parameter, and the third pose being a pose corresponding to a reference pose parameter, the first vertex being a vertex of the three-dimensional model that meets a target condition with the first sampling point; and obtain a second sampling point corresponding to the first sampling point based on a skinning weight corresponding to the first vertex, the first pose transformation matrix, and the second pose transformation matrix.
[0270] In a possible implementation, the rendering module 1403 is configured to, for a second sampling point, splice a coordinate of the second sampling point in the second virtual space, the shape parameter, and the pose parameter to obtain a first parameter set; and perform full connection processing on the first parameter set to obtain the first rendering parameter.
[0271] In a possible implementation, the rendering module 1403 is configured to, for a pixel point on the first target image, determine a color and an opacity of the pixel point based on a virtual ray between the pixel point and the virtual camera and the first rendering parameter corresponding to the pixel point; and perform rendering on the pixel point based on the color and the opacity, to output a rendered pixel point.
[0272] In a possible implementation, the first rendering parameter includes a color parameter and a density parameter, and the rendering module 1403 is configured to integrate first relationship data on the virtual ray to obtain the color, the first relationship data being associated with the color parameter and the density parameter; and integrate second relationship data on the virtual ray to obtain the opacity, the second relationship data being associated with the density parameter.
[0273] In a possible implementation, the rendering module 1403 is further configured to perform pose adjustment on the three-dimensional model of the sample object based on a target pose parameter, to obtain a three-dimensional model after pose adjustment; input the camera parameter, the shape parameter, and the target pose parameter into the trained image rendering model, and determine a plurality of second rendering parameters based on the camera parameter, the shape parameter, and the target pose parameter; and perform rendering on a second target image based on the plurality of second rendering parameters, to output a second rendering image, the second target image being an image obtained by photographing the three-dimensional model after pose adjustment by the virtual camera.
[0274] In a possible implementation, the apparatus further includes:
[0275] The camera parameter acquisition module is configured to obtain the camera parameter based on the sample video frame.
[0276] In a possible implementation, the rendering module 1403 is further configured to perform shape adjustment on the three-dimensional model of the sample object based on the target shape parameter to obtain a shape-adjusted three-dimensional model. The camera parameter, the pose parameter and the target shape parameter are input into the trained image rendering model, and a plurality of third rendering parameters are determined based on the camera parameter, the pose parameter and the target shape parameter. The third target image is rendered based on the plurality of third rendering parameters, and a third rendered image is output, where the third target image is an image obtained by photographing the shape-adjusted three-dimensional model by the virtual camera.
[0277] In a possible implementation, the rendering module 1403 is further configured to input the target camera parameter, the pose parameter and the target shape parameter into the trained image rendering model, determine a plurality of fourth rendering parameters based on the pose parameter, the shape parameter and the target camera parameter, and render a fourth target image based on the plurality of fourth rendering parameters to output a fourth rendered image, where the fourth target image is an image obtained by photographing the three-dimensional model by the virtual camera under the target camera parameter.
[0278] By the technical scheme provided in the embodiments of the present application, when the image rendering model is trained, the sample object is reconstructed in three dimensions based on the shape parameter and the pose parameter of the sample object, and the rendering parameters are determined based on the camera parameter, the shape parameter and the pose parameter, so that the influence of the virtual camera and the sample object is considered when the rendering parameters are determined, the rendering parameters are more matched with the three-dimensional model of the sample object, the first target image is rendered based on the rendering parameters to obtain a first rendered image, and the image rendering model is trained based on the difference information between the first rendered image and the sample video frame, so that the obtained image rendering model has stronger image rendering capability.
[0279] Figure 15 FIG. 1 is a structural schematic diagram of an image rendering device provided by an embodiment of the present application, referring to FIG. 1, Figure 15 The device includes a video frame display module 1501, a three-dimensional model display module 1502, a target image display module 1503 and a rendered image display module 1504.
[0280] The video frame display module 1501 is configured to display a target video frame, where the target video frame includes a target object.
[0281] The three-dimensional model display module 1502 is configured to display a three-dimensional model of the target object in response to a three-dimensional reconstruction operation on the target object, where the three-dimensional model is generated based on a shape parameter and a pose parameter of the target object, and the shape parameter and the pose parameter are determined based on the target video frame.
[0282] The target image display module 1503 is configured to display a first target image in response to a shooting operation on the three-dimensional model, the first target image being an image obtained by shooting the three-dimensional model by using a virtual camera.
[0283] The rendered image display module 1504 is configured to display a first rendered image in response to a rendering operation on the first target image, the first rendered image being obtained by rendering the first target image based on a plurality of first rendering parameters determined by using the image rendering model, the plurality of first rendering parameters being determined by using the image rendering model based on camera parameters of the virtual camera, the shape parameters, and the pose parameters, the image rendering model being used to render an image shot by a virtual camera.
[0284] By the technical solutions provided in the embodiments of the present application, when the first rendered image is generated, the target object is three-dimensionally reconstructed based on the shape parameters and the pose parameters of the target object, the rendering parameters are determined based on the camera parameters, the shape parameters, and the pose parameters, the influence of the virtual camera and the target object is considered when the rendering parameters are determined, the rendering parameters are more matched with the three-dimensional model of the target object, the first target image is rendered based on the rendering parameters, and the first rendered image is obtained, and the effect of image rendering is better.
[0285] The embodiments of the present application provide a computer device for executing the above method, and the computer device can be implemented as a terminal or a server. Here, the structure of the terminal is introduced first.
[0286] Figure 16 FIG. 1 is a structural diagram of a terminal according to an embodiment of the present application. The terminal 1600 can be a smart phone, a tablet computer, a notebook computer, or a desktop computer. The terminal 1600 can also be referred to as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, or other names.
[0287] Generally, the terminal 1600 includes one or more processors 1601 and one or more memories 1602.
[0288] The processor 1601 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 1601 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 1601 can also include a main processor and a coprocessor, the main processor being a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit), and the coprocessor being a low-power processor for processing data in a standby state. In some embodiments, the processor 1601 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing of content to be displayed by the display screen. In some embodiments, the processor 1601 can further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0289] The memory 1602 can include one or more computer-readable storage media that can be non-transitory. The memory 1602 can also include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1602 is used to store at least one computer program for being executed by the processor 1601 to implement the image rendering method provided by the method embodiments in the present application.
[0290] In some embodiments, the terminal 1600 can also optionally include a peripheral device interface 1603 and at least one peripheral device. The processor 1601, the memory 1602, and the peripheral device interface 1603 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 1603 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1604, a display screen 1605, a camera assembly 1606, an audio circuit 1607, and a power supply 1609.
[0291] The peripheral interface 1603 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1601 and the memory 1602. In some embodiments, the processor 1601, the memory 1602 and the peripheral interface 1603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1601, the memory 1602 and the peripheral interface 1603 can be implemented on a separate chip or circuit board, for which the present embodiments are not limited.
[0292] The radio frequency circuit 1604 is used to receive and send RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1604 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 1604 converts electrical signals into electromagnetic signals for transmission, or converts electromagnetic signals received into electrical signals. Optionally, the radio frequency circuit 1604 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like.
[0293] The display screen 1605 is used to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 1605 is a touch display screen, the display screen 1605 also has the ability to collect touch signals on or above the surface of the display screen 1605. The touch signals can be input as control signals to the processor 1601 for processing. At this time, the display screen 1605 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards.
[0294] The camera assembly 1606 is used to capture images or videos. Optionally, the camera assembly 1606 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is disposed on the front panel of the terminal, and the rear-facing camera is disposed on the back of the terminal.
[0295] The audio circuit 1607 can include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals input to the processor 1601 for processing, or input to the radio frequency circuit 1604 to realize voice communication.
[0296] The power supply 1609 is used to supply power to each component in the terminal 1600. The power supply 1609 can be alternating current, direct current, disposable batteries or rechargeable batteries.
[0297] In some embodiments, the terminal 1600 further comprises one or more sensors 1610. The one or more sensors 1610 include, but are not limited to, an acceleration sensor 1611, a gyroscope sensor 1612, a pressure sensor 1613, an optical sensor 1615, and a proximity sensor 1616.
[0298] The acceleration sensor 1611 can detect the acceleration magnitude in three coordinate axes of a coordinate system established by the terminal 1600.
[0299] The gyroscope sensor 1612 can detect the body direction and rotation angle of the terminal 1600. The gyroscope sensor 1612 can cooperate with the acceleration sensor 1611 to collect the 3D action of the user on the terminal 1600.
[0300] The pressure sensor 1613 can be arranged on the side frame of the terminal 1600 and / or the lower layer of the display screen 1605. When the pressure sensor 1613 is arranged on the side frame of the terminal 1600, the holding signal of the user on the terminal 1600 can be detected, and the left-hand or right-hand recognition or shortcut operation can be performed by the processor 1601 according to the holding signal collected by the pressure sensor 1613. When the pressure sensor 1613 is arranged on the lower layer of the display screen 1605, the controllable control on the UI interface can be controlled by the processor 1601 according to the pressure operation of the user on the display screen 1605.
[0301] The optical sensor 1615 is used to collect the ambient light intensity. In one embodiment, the processor 1601 can control the display brightness of the display screen 1605 according to the ambient light intensity collected by the optical sensor 1615.
[0302] The proximity sensor 1616 is used to collect the distance between the user and the front of the terminal 1600.
[0303] Those skilled in the art can understand that, Figure 16 The structure shown in the above figure does not constitute a limitation on the terminal 1600, and can include more or fewer components than the figure, or combine certain components, or use different component arrangements.
[0304] The above computer device can also be implemented as a server, and the structure of the server will be introduced as follows:
[0305] Figure 17Fig. 7 is a schematic diagram of a structure of a server according to an embodiment of the present application. The server 700 can have a great difference due to different configurations or performances, and can include one or more processors (Central Processing Units, CPUs) 701 and one or more memories 702. The one or more memories 702 store at least one computer program, which is loaded and executed by the one or more processors 701 to implement the method provided by any of the above method embodiments. Of course, the server 700 can also have a wired or wireless network interface, a keyboard, an input and output interface, and other components for implementing device functions, and the like, so as to perform input and output. The server 700 can also include other components for implementing device functions, which are not described herein.
[0306] In an example embodiment, a computer-readable storage medium, such as a memory including a computer program, is also provided. The computer program can be executed by a processor to complete the image rendering method in the above embodiments. For example, the computer-readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0307] In an example embodiment, a computer program product or computer program is also provided. The computer program product or computer program includes program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium, and executes the program code to cause the computer device to perform the above image rendering method.
[0308] In some embodiments, the computer program related to the embodiments of the present application can be deployed to execute on one computer device, or on multiple computer devices located in one place, or on multiple computer devices distributed in multiple places and interconnected through a communication network, which can constitute a blockchain system.
[0309] Those of ordinary skill in the art can understand that all or part of the steps of the above embodiments can be completed by hardware, or by a program instructing relevant hardware, which can be stored in a computer-readable storage medium. The storage medium mentioned above can be a Read-Only Memory, a magnetic disk or an optical disk, etc.
[0310] The above merely is the optional embodiment of the present application, and does not limit the present application, and any modification, equivalent replacement, improvement and the like made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. An image rendering method, characterized by, The method comprises: based on the sample video frame, obtain the shape parameter and the pose parameter of the sample object, the sample video frame includes the sample object; based on the shape parameter and the pose parameter, three-dimensional reconstruction is carried out on the sample object, and a three-dimensional model of the sample object is obtained; by an image rendering model, based on the camera parameter of the virtual camera, the shape parameter and the pose parameter, a plurality of first rendering parameters are determined, the camera parameter of the virtual camera is the same as the camera parameter of the real camera for shooting the sample video frame; based on the plurality of first rendering parameters, a first target image is rendered, and a first rendering image is output, the first target image is an image obtained by shooting the three-dimensional model by the virtual camera; based on the difference information between the sample video frame and the first rendering image, the image rendering model is trained; wherein, the image rendering model is used to determine rendering parameters based on input camera parameters, input shape parameters and input pose parameters, and then render an image shot by the virtual camera based on the determined rendering parameters; the image shot by the virtual camera is an image shot by the virtual camera under the input camera parameter or an image shot by the virtual camera on the adjusted three-dimensional model, and the adjusted three-dimensional model is obtained by adjusting the three-dimensional model of the sample object based on the input shape parameter or the input pose parameter.
2. The method of claim 1, wherein, based on the shape parameter and the pose parameter, three-dimensional reconstruction is carried out on the sample object, and a three-dimensional model of the sample object is obtained; by adjusting the shape of the reference three-dimensional model by the shape parameter and adjusting the pose of the reference three-dimensional model by the pose parameter, the three-dimensional model of the sample object is obtained, and the reference three-dimensional model is obtained based on the shape parameter and the pose parameter of a plurality of objects.
3. The method of claim 1, wherein, Before the method based on the sample video frame, the shape parameter and the pose parameter of the sample object are obtained, the method further comprises: image segmentation is performed on the sample video frame to obtain a target region, and the target region is a region where the sample object is located; based on the sample video frame, the shape parameter and the pose parameter of the sample object are obtained. based on the sample video frame, the shape parameter and the pose parameter of the sample object are obtained.
4. The method of claim 1, wherein, based on the sample video frame, the shape parameter and the pose parameter of the sample object are obtained. based on the sample video frame, the shape parameter and the pose parameter of the sample object are obtained. based on the plurality of reference shape parameters, the shape parameter of the sample object is determined. the camera parameter includes the position parameter of the virtual camera in the first virtual space, and the camera parameter of the virtual camera, the shape parameter and the pose parameter are used to determine a plurality of first rendering parameters.
5. The method of claim 1, wherein, determining at least one virtual ray in the first virtual space based on the position parameter and a perspective of the virtual camera on the three-dimensional model, the virtual ray being a line connecting the virtual camera and a pixel point on the first target image, the first virtual space being a virtual space established based on the camera parameter; determining the plurality of first rendering parameters based on coordinates of a plurality of first sampling points on the at least one virtual ray, the shape parameter, and the pose parameter, the coordinates of the first sampling point being coordinates of the first sampling point in the first virtual space.
6. The method of claim 5, wherein, The determining the plurality of first rendering parameters based on the coordinates of the plurality of first sampling points on the at least one virtual ray, the shape parameter, and the pose parameter comprises: transforming the plurality of first sampling points into a second virtual space based on the coordinates of the plurality of first sampling points, the pose parameter, and a reference pose parameter to obtain a plurality of second sampling points, one of the first sampling points corresponding to one of the second sampling points, the reference pose parameter being a pose parameter corresponding to the second virtual space, the coordinates of the second sampling point being coordinates of the second sampling point in the second virtual space; determining the plurality of first rendering parameters based on the coordinates of the plurality of second sampling points in the second virtual space, the shape parameter, and the pose parameter.
7. The method of claim 6, wherein, The transforming the plurality of first sampling points into a second virtual space based on the coordinates of the plurality of first sampling points, the pose parameter, and a reference pose parameter to obtain a plurality of second sampling points comprises: for one of the first sampling points, obtaining a first pose transformation matrix and a second pose transformation matrix of the first sampling point, the first pose transformation matrix being a transformation matrix of a first vertex from a first pose to a second pose, the second pose transformation matrix being a transformation matrix of the first vertex from the first pose to a third pose, the first pose being a reference pose, the second pose being a pose corresponding to the pose parameter, the third pose being a pose corresponding to the reference pose parameter, the first vertex being a vertex on the three-dimensional model and having a distance to the first sampling point meeting a target condition; obtaining a second sampling point corresponding to the first sampling point based on a skinning weight corresponding to the first vertex, the first pose transformation matrix, and the second pose transformation matrix.
8. The method of claim 6, wherein, The determining the plurality of first rendering parameters based on the coordinates of the plurality of second sampling points in the second virtual space, the shape parameter, and the pose parameter comprises: for one of the second sampling points, concatenating the coordinates of the second sampling point in the second virtual space, the shape parameter, and the pose parameter to obtain a first parameter set; performing full connection processing on the first parameter set to obtain the first rendering parameter.
9. The method of claim 1, wherein, The rendering the first target image based on the plurality of first rendering parameters to output a first rendered image comprises: For a pixel point on the first target image, based on a virtual ray between the pixel point and the virtual camera and a first rendering parameter corresponding to the pixel point, a color and an opacity of the pixel point are determined; Based on the color and the opacity, the pixel point is rendered, and the rendered pixel point is output.
10. The method of claim 9, wherein, The first rendering parameter includes a color parameter and a density parameter, and the determination of the color and the opacity of the pixel point based on the virtual ray between the pixel point and the virtual camera and the first rendering parameter corresponding to the pixel point includes: On the virtual ray, first relationship data associated with the color parameter and the density parameter is integrated to obtain the color; On the virtual ray, second relationship data associated with the density parameter is integrated to obtain the opacity.
11. The method of claim 1, wherein, Before the determination of the plurality of first rendering parameters based on the camera parameter of the virtual camera, the shape parameter and the pose parameter, the method further includes: Based on the sample video frame, the camera parameter is obtained.
12. The method of claim 1, wherein, The method further includes: Based on a target pose parameter, a pose of the three-dimensional model of the sample object is adjusted to obtain a three-dimensional model after pose adjustment; The camera parameter, the shape parameter and the target pose parameter are input into the trained image rendering model, and a plurality of second rendering parameters are determined based on the camera parameter, the shape parameter and the target pose parameter; Based on the plurality of second rendering parameters, a second target image is rendered to output a second rendering image, and the second target image is an image obtained by the virtual camera shooting the three-dimensional model after pose adjustment.
13. The method of claim 1, wherein, The method further includes: Based on a target shape parameter, a shape of the three-dimensional model of the sample object is adjusted to obtain a three-dimensional model after shape adjustment; The camera parameter, the pose parameter and the target shape parameter are input into the trained image rendering model, and a plurality of third rendering parameters are determined based on the camera parameter, the pose parameter and the target shape parameter; based on the plurality of third rendering parameters, a third target image is rendered to output a third rendering image, and the third target image is an image obtained by the virtual camera shooting the three-dimensional model after shape adjustment.
14. The method of claim 1, wherein, The method further includes: The target camera parameter, the pose parameter and the shape parameter are input into the trained image rendering model, and a plurality of fourth rendering parameters are determined based on the pose parameter, the shape parameter and the target camera parameter; based on the plurality of fourth rendering parameters, a fourth target image is rendered to output a fourth rendering image, and the fourth target image is an image obtained by the virtual camera shooting the three-dimensional model under the target camera parameter.
15. An image rendering method, characterized by, The method includes: A target video frame is displayed, and the target video frame includes a target object; In response to a three-dimensional reconstruction operation on the target object, a three-dimensional model of the target object is displayed, the three-dimensional model being generated based on a shape parameter and a pose parameter of the target object, the shape parameter and the pose parameter being determined based on the target video frame; In response to a shooting operation on the three-dimensional model, a first target image is displayed, the first target image being an image obtained by a virtual camera shooting the three-dimensional model; In response to a rendering operation on the first target image, a first rendered image is displayed, the first rendered image being obtained by an image rendering model trained, rendering the first target image based on a plurality of first rendering parameters, the plurality of first rendering parameters being determined by the image rendering model based on camera parameters of the virtual camera, the shape parameter, and the pose parameter, the image rendering model being configured to determine rendering parameters based on input camera parameters, input shape parameters, and input pose parameters, and further configured to render an image obtained by the virtual camera based on the determined rendering parameters, the image obtained by the virtual camera being an image obtained by the virtual camera shooting under the input camera parameters or an image obtained by the virtual camera shooting an adjusted three-dimensional model, the adjusted three-dimensional model being obtained by adjusting the three-dimensional model based on the input shape parameters or the input pose parameters.
16. An image rendering apparatus, characterized by comprising: The apparatus comprises: A parameter acquisition module configured to acquire a shape parameter and a pose parameter of a sample object based on a sample video frame, the sample video frame comprising a sample object; A three-dimensional reconstruction module configured to perform three-dimensional reconstruction on the sample object based on the shape parameter and the pose parameter to obtain a three-dimensional model of the sample object; A rendering module configured to determine a plurality of first rendering parameters based on camera parameters of a virtual camera, the shape parameter, and the pose parameter by an image rendering model, the camera parameters of the virtual camera being identical to camera parameters of a real camera shooting the sample video frame; render a first target image based on the plurality of first rendering parameters to output a first rendered image, the first target image being an image obtained by the virtual camera shooting the three-dimensional model; A training module configured to train the image rendering model based on difference information between the sample video frame and the first rendered image; The image rendering model is configured to determine rendering parameters based on input camera parameters, input shape parameters, and input pose parameters, and further configured to render an image obtained by the virtual camera based on the determined rendering parameters, the image obtained by the virtual camera being an image obtained by the virtual camera shooting under the input camera parameters or an image obtained by the virtual camera shooting an adjusted three-dimensional model, the adjusted three-dimensional model being obtained by adjusting the three-dimensional model of the sample object based on the input shape parameters or the input pose parameters.
17. An image rendering apparatus, characterized by comprising: The apparatus comprises: A video frame display module configured to display a target video frame, the target video frame comprising a target object; The three-dimensional model display module is configured to display a three-dimensional model of the target object in response to a three-dimensional reconstruction operation on the target object, the three-dimensional model being generated based on a shape parameter and a pose parameter of the target object, the shape parameter and the pose parameter being determined based on the target video frames; The target image display module is configured to display a first target image in response to a shooting operation on the three-dimensional model, the first target image being an image obtained by shooting the three-dimensional model by a virtual camera; The rendering image display module is configured to display a first rendering image in response to a rendering operation on the first target image, the first rendering image being obtained by rendering the first target image based on a plurality of first rendering parameters determined by the image rendering model based on camera parameters of the virtual camera, the shape parameter, and the pose parameter, the image rendering model being configured to determine rendering parameters based on input camera parameters, input shape parameters, and input pose parameters, and to render an image obtained by the virtual camera based on the determined rendering parameters, the image obtained by the virtual camera being an image obtained by the virtual camera under the input camera parameters or an image obtained by shooting an adjusted three-dimensional model by the virtual camera, the adjusted three-dimensional model being obtained by adjusting the three-dimensional model based on the input shape parameters or the input pose parameters.
18. A computer device, comprising: The computer device includes one or more processors and one or more memories, and the one or more memories store at least one computer program, which is loaded and executed by the one or more processors to implement the image rendering method according to any one of claims 1 to 15.
19. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one computer program, which is loaded and executed by the processor to implement the image rendering method according to any one of claims 1 to 15.
20. A computer program product, comprising a program code stored in a computer readable storage medium, a processor of a computer device reading the program code from the computer readable storage medium, and the processor executing the program code to implement the image rendering method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Data synthesis method and device, storage medium and electronic device
CN111105489A
Image processing method and device, equipment and computer readable storage medium
CN112037320A