Human body free viewpoint rendering method and system guided by local information

Through the local information-guided free viewpoint rendering method of human body, using the SMPL model and three-dimensional Gaussian distribution, combined with multi-layer perceptron and adaptive density control, the problems of high professional equipment, low rendering efficiency, and insufficient detail fitting in the existing technology are solved, and efficient and rich in detail are achieved.

CN120070696AActive Publication Date: 2025-05-30SUN YAT SEN UNIV

Patent Information

Application Number
CN202510008785.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-30
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The prior art has problems in the free viewpoint rendering of human body, with high equipment professionalism, low rendering efficiency, and insufficient detail fitting.

Method used

A local information-guided free viewpoint rendering method is proposed. The human body is fitted with SMPL model, a three-dimensional Gaussian distribution is constructed, and a multi-layer perceptron is used for deformation and transformation. Combined with the local perceptual adaptive density control strategy, the Gaussian density is dynamically adjusted to improve rendering quality.

Benefits of technology

It realizes efficient rendering of free viewpoint images of human bodies without the need for professional RGBD cameras, improving rendering efficiency and detail fitting effect, and reducing storage needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070696A_ABST
    Figure CN120070696A_ABST
Patent Text Reader

Abstract

The invention discloses a local information guided human body free viewpoint rendering method, which comprises the following steps: inputting a single-view human body video data set, fitting SMPL attitude and shape parameters and carrying out random sampling to obtain three-dimensional point cloud data; gaussian distribution is constructed based on the point cloud, Gaussian distribution attributes are initialized, parameter offset is trained by using a multilayer perceptron and fitting is carried out, a deformed three-dimensional Gaussian position is obtained, LBS transformation is carried out in combination with a joint point rotation matrix, and Gaussian distribution of an observation frame attitude is obtained; decomposing and predicting the Gaussian color to obtain a final color; and projecting the three-dimensional gauss to a two-dimensional plane according to the visual angle parameters, performing rasterization, rendering a human body two-dimensional image, dynamically calculating a densification threshold value according to a local perception adaptive density control strategy, and performing density control on the gauss. The invention further discloses a local information guided human body free viewpoint rendering system. The detail fitting degree and the color prediction accuracy can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image rendering, and particularly to a method and system for rendering a free viewpoint of a human body guided by local information. Background Art

[0002] In recent years, with the continuous development of computer technology and the continuous improvement of device performance, people's demand for the understanding and reconstruction of three-dimensional scenes and objects has been continuously increasing. Compared with text and images, three-dimensional data has a richer and more three-dimensional presentation and is more in line with our real world. Among the vast amounts of data, human body data is the most common. With the continuous development of artificial intelligence technology, virtual digital humans have gradually entered our daily lives and have extensive applications in fields and industries such as the virtual reality industry, game and movie production, and digital e-commerce. High-quality free viewpoint rendering has broad application prospects and potential commercial value in fields such as film and television entertainment and holographic communication.

[0003] The human body free viewpoint image rendering technology is a process that aims to use computer graphics and computer vision technologies to restore the geometric shape and appearance of the human body from data sources such as images, videos, and depth data, model the human body, and render realistic human body images from a specific perspective according to camera parameters. How to use portable devices to achieve accurate, efficient, and realistic digital human modeling and free viewpoint rendering is a major difficulty in this field.

[0004] One of the current existing technologies is a new perspective rendering technology of the human body based on depth information by Yu T et al. in "Function4d: Real-time human volumetric capture from very sparse consumer rgbd sensors". This technology uses a sparse multi-view RGBD sensor to collect RGBD images from multiple views, uses a dynamic sliding fusion technology to fuse the RGBD images of adjacent frames to obtain a voxelized fusion result, and then re-renders the RGBD images with the same perspective as the original. The multi-view RGBD images obtained by re-rendering are finally used to completely reconstruct the human body after passing through a multi-view image encoder, feature aggregation, and geometric and color decoders. In the scenario where a new perspective needs to be rendered, a traditional rendering model such as BRDF is used to obtain the human body image at any viewpoint. The disadvantages of this technology are: the RGBD sensor required by this technology has not been widely popularized at present, and at the same time, multi-view image acquisition requires strict calibration of the camera, and its professionalism is also beyond the reach of the general public. In addition, the three-dimensional mesh of the human body generated by this technology is a displayed one, rather than an end-to-end new perspective human body image, which consumes additional storage space without considering physical interaction.

[0005] The second current prior art is a novel view synthesis technology for human bodies based on neural radiance fields in "Instantavatar: Learning avatars from monocular video in 60 seconds" by Jiang T et al. This technology represents a 3D human body as an implicit neural radiance field, performs ray casting and point sampling according to the camera view at the observed pose, and maps the sampled points to the standard pose. For the sampled points outside the human body, this technology designs a blank space skipping strategy, maintains a common occupancy grid in the standard space for all observed frames, and updates this common occupancy grid by integrating the medium probability density of the sampled points in different observed frames. During inference, according to this occupancy grid, the sampled points far from the human body surface are filtered out. Finally, according to the coordinates of the sampled points in the standard pose, the point features at different resolutions are retrieved through hashing, and after splicing, they are input into the neural radiance field to calculate the medium probability density and color of the points, and render the human body image at any viewing point. The disadvantages of this technology are as follows: This technology uses a multi-layer perceptron as the storage medium in the neural radiance field. When rendering images, dense point sampling is required, and the time cost required for forward calculation during training and inference is relatively large, making it impossible to achieve real-time performance. In addition, this technology uses a continuous implicit neural radiance field to render images, and for edge regions with drastic color changes, blurring often occurs.

[0006] The third current prior art is a novel view synthesis technology for human bodies based on 3D mixture Gaussian and splat rendering in "Hugs: Human gaussian splats" by Kocabas M et al. This technology samples 3D point clouds on a parameterized template of a human body in the standard pose, establishes a learnable 3D Gaussian distribution for each point cloud, and stores attributes such as the mean position, covariance matrix, scaling vector, and spherical harmonic function coefficients of the Gaussian distribution. The Gaussian position passes through a feature triplane, samples to obtain features, and after passing through a multi-layer perceptron, outputs 3D Gaussian attributes and linear skinning weights. The 3D Gaussian in the standard pose is transformed by the linear skinning weights to obtain the 3D Gaussian distribution in the observed pose. Given any view, a differentiable GPU parallel rasterization strategy is executed to render the human body image. The disadvantages of this technology are as follows: This technology does not consider the combination of 3D Gaussian with other geometric information. As a special type of explicit point cloud, 3D Gaussian can be associated with other attributes with explicit features. In addition, this technology does not consider that the detail granularity of different regions of the human body is different, resulting in a lack of high-frequency details. Summary of the Invention

[0007] The purpose of the present invention is to overcome the shortcomings of existing methods and propose a method and system for human body free viewpoint rendering guided by local information. The main problems solved by the present invention are: 1) How to solve the problem that the human body new perspective rendering technology based on depth information requires a professional RGBD camera, which is difficult to obtain, and the generated human body three-dimensional mesh has high requirements for storage conditions, which is not conducive to large-scale popularization and application; 2) How to solve the problem that the human body free perspective synthesis technology based on neural radiation field uses a multi-layer perceptron as a medium for storing radiation field, and dense point sampling is required when rendering images, and the training and reasoning speed is slow and the time cost is high; 3) How to solve the problem that the human body new perspective synthesis technology based on three-dimensional mixed Gaussian and splash rendering does not consider the connection and combination of three-dimensional Gaussian and other geometric attributes, and the fitting effect of high-frequency details is not ideal.

[0008] In order to solve the above problems, the present invention proposes a human body free viewpoint rendering method guided by local information, the method comprising:

[0009] Input a single-view human video dataset, and fit the corresponding Skinned Multi-Person Linear (SMPL) model for each frame in the video in the dataset. At the same time, remove the background in the video frame according to the human mask provided by the dataset, and randomly sample the surface of the SMPL model in a standard posture to obtain 3D point cloud data;

[0010] Based on the three-dimensional point cloud data, a three-dimensional Gaussian distribution is constructed for each point, and the attributes of the Gaussian distribution are initialized. The position of the three-dimensional Gaussian is passed to the multi-layer perceptron MLP_1, and the offset of the storage parameter in the Gaussian distribution is obtained by training, and the offset is applied to the initial three-dimensional Gaussian distribution to obtain the deformed three-dimensional Gaussian position;

[0011] The deformed three-dimensional Gaussian position and the joint point rotation matrix of the SMPL model are used to predict the linear blend skinning (LBS) transformation weights from the standard posture to the observation frame posture through a multi-layer perceptron MLP_2, and the LBS transformation is performed to obtain the Gaussian distribution of the observation frame posture;

[0012] Decompose the Gaussian color into specular reflection component and diffuse reflection component, estimate them separately based on the normal vector, and combine them to get the final predicted color;

[0013] According to the viewing angle parameters, the 3D Gaussian is projected onto a 2D plane, and the 2D Gaussian is rasterized to render a 2D image of the human body. Based on the local perception adaptive density control strategy, the densification threshold is dynamically calculated by the neighborhood variance of the Gaussian's normal vector to control the density of the Gaussian.

[0014] Preferably, the SMPL model is specifically:

[0015] The SMPL model is a vertex - based skinning model that can accurately represent various body shapes in natural human postures. The parameters of this model include shape and pose parameters, a standard pose template, and predefined blend skinning weights. Among them, the shape parameters depict the general shape of the human body. The standard pose refers to the pose where the human body spreads its arms and legs in a "big" character shape. The pose parameters in the SMPL model depict the rotation and translation degrees of the human body's joint points relative to the standard pose. The standard pose template is the human SMPL model in the standard pose, which consists of several points and triangular patches formed by connecting points.

[0016] Preferably, based on the three - dimensional point cloud data, a three - dimensional Gaussian distribution is constructed for each point, and the attributes of the Gaussian distribution are initialized. The position of the three - dimensional Gaussian is passed into the multi - layer perceptron MLP_1, and the offset of the stored parameters in the Gaussian distribution is obtained through training and applied to the initial three - dimensional Gaussian distribution to obtain the deformed three - dimensional Gaussian position. Specifically:

[0017] The mathematical definition of the three - dimensional Gaussian distribution is as follows:

[0018]

[0019] where \(x\) is the point cloud point, \(\mu\) is the center of the three - dimensional Gaussian distribution, that is, the position mean, \(\Sigma\) is the covariance matrix, which is used to represent the size, shape, and direction of the three - dimensional Gaussian, and is calculated from the scaling vector \(s\) and the rotation matrix \(R\). In the three - dimensional Gaussian distribution, the following parameters are stored: position \(\mu\), rotation quaternion \(q\), scaling vector \(s\), normal vector \(n\), opacity parameter \(\alpha\), \(f\) s is the specular reflection color feature, \(f\) d is the diffuse reflection color feature. The rotation quaternion is initialized as \(q = [1,0,0,0]\), which is equivalent to the identity matrix in the rotation matrix, that is, no rotation is performed; the scaling vector \(s\) is initialized as the logarithm of the distance from the point to the three nearest neighbor points; the opacity parameter \(\alpha\) is initialized as 1, \(f\) s and \(f\) d is initialized as a vector of all 0s. Among them, the rotation quaternion \(q=[w,x,y,z]\) is converted into the rotation matrix \(R\) through the following formula:

[0020]

[0021] According to the three - dimensional point cloud data, the three - dimensional Gaussian is initialized. Among them, the position \(\mu\) is initialized as the coordinate of the point cloud, and the normal vector \(n\) is initialized as the surface normal vector of the SMPL triangular patch where the sampling point is located;

[0022] Train a multi-layer perceptron MLP_1 to learn the offsets of three-dimensional Gaussian attributes, specifically:

[0023] (Δμ, Δq, Δs, Δn) = MLP_1(μ)

[0024] The output is the offsets of the position μ, rotation quaternion q, scaling vector s, and normal vector n, that is, the offsets of the stored parameters, used to fit the offset of the clothing relative to the human body surface. MLP_1 consists of an input layer, two hidden layers, and an output layer. The width of the hidden layers is 128. The predicted offsets act on the initialized three-dimensional Gaussian distribution according to the following formula:

[0025]

[0026] Among them, the · operation between q and Δq is equivalent to converting them into matrices and then performing matrix multiplication. e is the exponential function. The predicted offset of the normal vector needs to be converted into a rotation matrix and then multiplied by the original normal vector n.

[0027] Preferably, the deformed three-dimensional Gaussian position and the joint rotation matrix of the SMPL model pass through a multi-layer perceptron MLP_2 to predict the linear blend skinning (LBS) transformation weights from the standard pose to the observed frame pose, and perform the LBS transformation to obtain the Gaussian distribution of the observed frame pose, specifically:

[0028] The initialization of the three-dimensional Gaussian and the deformation fitting of the clothing are both completed in the standard pose space. It is necessary to perform a linear blend skinning LBS transformation to achieve the transformation from the standard pose space to the observed pose space. The LBS transformation regards the changes of the human body surface vertices caused by actions as being driven by the skeleton. The trajectory changes of the points are calculated by weighted averaging the transformations of the joints on the skeleton, that is, the position of each vertex is the weighted average of the vertex positions after rigid skinning of the joints that affect it;

[0029] The SMPL model predicted from the observed frame contains pose parameters. Calculate the human joint bone rotation matrix from the pose parameters. Input the deformed three-dimensional Gaussian position and the SMPL joint rotation matrix into the multi-layer perceptron MLP_2 to predict the weights of the LBS transformation:

[0030]

[0031] Among them, is the bone rotation matrix, k represents the kth joint point, K represents the total number of joint points, w k represents the transformation of the kth joint point relative to μ dThe weight of MLP_2 consists of an input layer, three hidden layers, and an output layer. The width of the hidden layer is 128. After obtaining the LBS weight of the three-dimensional Gaussian through MLP_2, the LBS transformation is implemented according to the following formula:

[0032]

[0033] where T is the transformation matrix, obtained by multiplying the weight of the bone predicted by MLP_2 by the bone transformation matrix calculated from SMPL, and R d is the rotation matrix calculated from the rotation quaternion q d in it.

[0034] Preferably, decomposing the Gaussian color into specular and diffuse components, estimating them separately based on the normal vector, and combining them to obtain the final predicted color, specifically:

[0035] Decompose the Gaussian color into specular and diffuse components, and use the normal vector as one of the bases for predicting the two color components, specifically as follows:

[0036]

[0037] where f s is the specular color feature stored in the Gaussian, f d is the diffuse color feature stored in the Gaussian, d is the camera view angle, c s is the specular component, c d is the diffuse component. MLP_3 and MLP_4 have the same structure, both consisting of an input layer, a hidden layer, and an output layer. The width of the hidden layer is 64. Add the two components:

[0038] c = c s + c d

[0039] Thus, the final predicted color c is obtained.

[0040] Preferably, according to the view angle parameters, project the three-dimensional Gaussian onto a two-dimensional plane, rasterize the two-dimensional Gaussian, render the two-dimensional human image, and dynamically calculate the densification threshold based on the normal vector neighborhood variance of the Gaussian according to the local perception adaptive density control strategy to control the density of the Gaussian, specifically:

[0041] According to the view angle parameters, project the three-dimensional Gaussian onto a two-dimensional plane to obtain the corresponding two-dimensional Gaussian, specifically as follows:

[0042]

[0043] where i represents the i-th Gaussian distribution; αi represents the opacity parameter of the i-th Gaussian, μ i represents the two-dimensional mean coordinate obtained after projecting the i-th three-dimensional Gaussian, Σ′ i =(JWΣ i W T J T ) 1:2,1:2 is the two-dimensional covariance matrix obtained after projecting the i-th three-dimensional Gaussian, where W is the transformation matrix of the camera from the world coordinate system to the camera coordinate system, and J is the Jacobian matrix of the perspective projection transformation; p is the coordinate of each pixel point, and e is the exponential function;

[0044] The final RGB color value of each pixel in the rendered image is obtained by the α-blending method, and its calculation formula is as follows:

[0045]

[0046] where c i is the i-th final predicted color, f i is the probability value of the position of pixel p after projection for the i-th two-dimensional Gaussian distribution, and the result C obtained after final α-blending is the RGB color value of pixel p;

[0047] Execute the local perception adaptive density control strategy for the three-dimensional Gaussian to dynamically adjust the number and distribution of the three-dimensional Gaussian;

[0048] Regard the normal vector as a hint for the details of the human body surface, and regard the local variance of the normal vector as a measure of the local similarity of the Gaussian. During the splitting and cloning process, dynamically calculate the densification threshold of the Gaussian, as follows:

[0049]

[0050] where var is the variance operation. According to the knnk nearest neighbor algorithm, find the k Gaussians closest to the position of the three-dimensional Gaussian, and calculate the variance of their normal vectors var(n d ), a, β, and b are all constants. According to the local perception adaptive density strategy, the calculated splitting and cloning thresholds are lower, and densification is more likely to occur, making the number and distribution of the Gaussian more reasonable, and further making the finally rendered image have richer high-frequency details;

[0051] The loss function of the training process is as follows:

[0052]

[0053] where, is the mean error between each pixel of the rendered image and the corresponding pixel of the real image, It is the structural similarity loss between the rendered image and the real image, which measures the similarity between the two images. Using ECON, the corresponding normal map is estimated from the observed frame as a supervisory signal. At the same time, the RGB color value C of the pixel p is replaced by the Gaussian normal vector to render the normal map. The same loss as the image is calculated for the two normal maps. and λ 1 and λ 2 is the loss weight.

[0054] Accordingly, the present invention also provides a human body free viewpoint rendering system guided by local information, comprising:

[0055] The data acquisition unit is used to input a single-view human video dataset, and for each frame in the video in the dataset, the corresponding Skinned Multi-Person Linear (SMPL) model is fitted. At the same time, the background in the video frame is removed according to the human mask provided by the dataset, and the surface of the SMPL model is randomly sampled under a standard posture to obtain three-dimensional point cloud data;

[0056] A parameter deformation unit is used to construct a three-dimensional Gaussian distribution for each point based on the three-dimensional point cloud data, initialize the properties of the Gaussian distribution, pass the position of the three-dimensional Gaussian into the multi-layer perceptron MLP_1, train to obtain the offset of the storage parameter in the Gaussian distribution, and act on the initial three-dimensional Gaussian distribution to obtain the deformed three-dimensional Gaussian position;

[0057] A weight transformation unit is used to predict the linear blend skinning (LBS) transformation weights from the standard posture to the observation frame posture through the multi-layer perceptron MLP_2 of the deformed three-dimensional Gaussian position and the joint point rotation matrix of the SMPL model, and implement LBS transformation to obtain the Gaussian distribution of the observation frame posture;

[0058] A color prediction unit is used to decompose the Gaussian color into specular reflection component and diffuse reflection component, and estimate them separately based on the normal vector, and combine them to obtain the final predicted color;

[0059] The image rendering unit is used to project the three-dimensional Gaussian onto a two-dimensional plane according to the viewing angle parameters, rasterize the two-dimensional Gaussian, render the two-dimensional image of the human body, and dynamically calculate the densification threshold from the neighborhood variance of the normal vector of the Gaussian according to the local perception adaptive density control strategy to control the density of the Gaussian.

[0060] The implementation of the present invention has the following beneficial effects:

[0061] This scheme adds a normal vector to the attributes of the three-dimensional Gaussian. As a kind of geometric information, it can guide the three-dimensional Gaussian to pay more attention to the surface geometry of the human body during the network learning process, thereby assisting the learning of other attributes of the three-dimensional Gaussian.

[0062] This scheme takes into account that different parts of the human body have different detail granularity. Simply using a fixed threshold to control the Gaussian density ignores local information. Therefore, a locally aware adaptive density control strategy is used to dynamically calculate the densification threshold of each Gaussian, so that the Gaussian can fit more human surface details.

[0063] This scheme combines traditional graphics knowledge to divide color into specular reflection component and diffuse reflection component, and considers the different factors affecting the two, and predicts them separately. At the same time, the normal vector also plays a prompting role in color calculation, so it is also used as the input of the color prediction model. In addition, the color decomposition strategy of this scheme can help the three-dimensional Gaussian better fit the human body surface color in areas with drastic color changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a flow chart of a method for human body free viewpoint rendering guided by local information according to an embodiment of the present invention;

[0065] Figure 2 It is a structural diagram of a local information guided human body free viewpoint rendering system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0066] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0067] Figure 1 is a flow chart of a method for rendering a human body free viewpoint guided by local information according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0068] S1, the input is a single-view human video. For each frame in the video, the corresponding SMPL model is fitted, and the background in the video frame is removed according to the human mask provided by the dataset;

[0069] S2, in the standard posture space, the SMPL model, i.e., the standard posture template surface, is sampled to obtain three-dimensional point cloud data;

[0070] S3. Based on the point cloud data in step S2, construct a three-dimensional Gaussian distribution for each point and initialize the attributes of the Gaussian distribution;

[0071] S4. Input the position of the three-dimensional Gaussian into the multi-layer perceptron MLP_1, train to obtain the offset of the stored parameters in the Gaussian distribution, and apply it to the initial three-dimensional Gaussian distribution in step S3;

[0072] S5. Pass the deformed three-dimensional Gaussian position and the SMPL joint point rotation matrix through the multi-layer perceptron MLP_2 to predict the LBS transformation weight from the standard pose space to the observed frame pose space, and perform the LBS transformation;

[0073] S6. For the estimation of Gaussian color, decompose the color into specular reflection component and diffuse reflection component, estimate them respectively based on the normal vector, and combine them to obtain the final color;

[0074] S7. Project the three-dimensional Gaussian onto a two-dimensional plane according to the camera parameters of a certain perspective;

[0075] S8. Rasterize the two-dimensional Gaussian to render the two-dimensional human body image;

[0076] S9. According to the local perception adaptive density control strategy, dynamically calculate the densification threshold from the normal vector neighborhood variance of the Gaussian to control the density of the Gaussian.

[0077] Step S1 is as follows:

[0078] S1-1: The Skinned Multi-Person Linear (SMPL) model is a vertex skinning-based model that can accurately represent various body shapes in natural human postures; the parameters of this model include shape and pose parameters, standard pose templates, predefined blend skinning weights, etc. Among them, the shape parameters depict the general shape of the human body, the standard pose refers to the pose where the human body spreads its arms and legs in a "big" character shape, and the pose parameters in SMPL depict the rotation and translation degree of the human body's joint points relative to the standard pose. The standard pose template is the human SMPL model in the standard pose, which consists of several points and triangular patches connecting the points. There are already mature methods such as HMR2.0 that can fit the SMPL model from two-dimensional images.

[0079] Step S2 is as follows:

[0080] S2-1: The sampling method is random sampling to ensure that the sampled points evenly cover the human body surface.

[0081] Step S3 is as follows:

[0082] S3-1: The mathematical definition of the three-dimensional Gaussian distribution is as follows:

[0083]

[0084] Among them, μ is the center of the three-dimensional Gaussian distribution, that is, the position mean value, and Σ is the covariance matrix, which is used to represent the size, shape and direction of the three-dimensional Gaussian, and is calculated from the scaling vector s and the rotation matrix R. In the three-dimensional Gaussian distribution, the following parameters are stored: position μ, rotation quaternion q, scaling vector s, normal vector n, opacity parameter α, and color feature f s and f d . The rotation quaternion is initialized to q = [1, 0, 0, 0], which is equivalent to the identity matrix in the rotation matrix, that is, no rotation is performed; the scaling vector s is initialized to the logarithm of the distance from the point to the three nearest neighbor points; the opacity α is initialized to 1, f s and f d is initialized to a vector of all 0s. Among them, the rotation quaternion q = [w, x, y, z] can be converted into the rotation matrix R through the following formula:

[0085]

[0086] S3-2: According to the point cloud in S2, initialize the three-dimensional Gaussian. Among them, μ is initialized to the coordinates of the point cloud, and n is initialized to the surface normal vector of the SMPL triangular patch where the sampling point is located.

[0087] Step S4 is as follows:

[0088] S4-1: The SMPL model is without human clothing, and there is a deviation in directly using the three-dimensional Gaussian obtained from SMPL to fit a single-view human body. Therefore, train a multi-layer perceptron MLP_1 to learn the offset of the three-dimensional Gaussian attributes:

[0089] (Δμ, Δq, Δs, Δn) = MLP_1(μ)

[0090] The output is the offsets of the position, rotation quaternion, scaling vector, and normal vector, which are used to fit the offset of the clothing relative to the human body surface. MLP_1 consists of an input layer, two hidden layers, and an output layer, where the width of the hidden layers is 128. The predicted offsets act on the initialized three-dimensional Gaussian distribution according to the following formula:

[0091]

[0092] Among them, the · operation between q and Δq is equivalent to converting the two into matrices and then performing matrix multiplication. The e exponential function. The predicted normal vector offset needs to be converted into a rotation matrix and then multiplied by the original normal vector n.

[0093] Step S5 is as follows:

[0094] S5-1: The initialization of the 3D Gaussian and the fitting of the clothing deformation are both completed in the standard pose space. It is necessary to perform a transformation from the standard pose space to the observed pose space through Linear Blend Skinning (LBS). The LBS transformation regards the changes in the vertices on the human body surface due to actions as being driven by the skeleton. The trajectory changes of the points can be calculated by weighted calculation of the transformations of the joint points on the skeleton, that is, the position of each vertex is the weighted average of the vertex positions after rigid skinning of the joints that affect it.

[0095] S5-2: The SMPL model predicted from the observed frame contains pose parameters, and the rotation matrix of the human joint point bones can be calculated from the pose parameters. The position of the deformed 3D Gaussian and the SMPL joint point rotation matrix are input into the multi-layer perceptron MLP_2 to predict the weights of the LBS transformation:

[0096]

[0097] Among them, is the bone transformation matrix, k represents the k-th joint point, K represents the total number of joint points, and w k represents the weight of the transformation of the k-th joint point relative to μ d . MLP_2 consists of an input layer, three hidden layers, and an output layer, where the width of the hidden layer is 128. After obtaining the LBS weights of the 3D Gaussian through MLP_2, the LBS transformation is implemented according to the following formula:

[0098]

[0099] Among them, R d is the rotation matrix calculated from the rotation quaternion q d .

[0100] Step S6 is as follows:

[0101] S6-1: In traditional computer graphics, color is formed by light reflection, and light reflection can be further divided into specular reflection and diffuse reflection. Combining this, this solution also decomposes color into specular reflection components and diffuse reflection components. In the color calculation model of traditional computer graphics, the normal vector is also involved in the color calculation. Therefore, the normal vector also contributes to the color calculation and is also used as one of the bases for predicting the two color components:

[0102]

[0103] Among them, f s is the specular reflection color feature stored in the Gaussian, and f dis the diffuse color feature stored in Gauss, d is the camera view angle, and c s is the specular color component, and c d is the diffuse color component. Since the specular reflection is related to the view angle and the diffuse reflection is not related to the view angle, d is used as the input of MLP_3 instead of MLP_4. The structures of MLP_3 and MLP_4 are the same, both consisting of an input layer, a hidden layer, and an output layer, where the width of the hidden layer is 64. After adding the two components, the final predicted color is obtained:

[0104] c = c s + c d .

[0105] Step S7 is as follows:

[0106] S7-1: The calculation of the two-dimensional Gaussian is as follows:

[0107]

[0108] where i represents the i-th Gaussian distribution; α i represents the opacity parameter of the i-th Gaussian; μ i represents the two-dimensional mean coordinate obtained after projecting the i-th three-dimensional Gaussian; Σ′ i =(JWΣ i W T J T ) 1:2,1:2 is the two-dimensional covariance matrix obtained after projecting the i-th three-dimensional Gaussian, where W is the transformation matrix of the camera from the world coordinate system to the camera coordinate system, and J is the Jacobian matrix of the perspective projection transformation; p is the coordinate of each pixel point; e is the exponential function.

[0109] Step S8 is as follows:

[0110] S8-1: The final RGB color value of each pixel in the rendered image is obtained by the α-blending method, and the calculation formula is as follows:

[0111]

[0112] c i is the color of the i-th Gaussian calculated according to S6, and f i is the probability value of the position of the pixel p in the projected position for the i-th two-dimensional Gaussian distribution calculated in S7. The final result C obtained after α-blending is the RGB color value of the pixel p.

[0113] Step S9 is as follows:

[0114] S9-1: Execute a local perception adaptive density control strategy on the three-dimensional Gaussian to dynamically adjust the number and distribution of the three-dimensional Gaussian.

[0115] During the optimization of the Gaussian, every 100 iterations, the Gaussian will be split, cloned, or pruned to adjust the number of Gaussians. The detail granularity of each part of the human body is different, so each part also requires a different number of three-dimensional Gaussians to represent it. This process requires the help of local information. The normal vector can be regarded as a hint for the details of the human body surface, and the local variance of the normal vector can be regarded as a measure of the local similarity of the Gaussian. Combining this point, during the splitting and cloning process, the densification threshold of the Gaussian is dynamically calculated:

[0116]

[0117] Among them, var is the variance operation. According to the knnk nearest neighbor algorithm, find the k Gaussians that are closest to the position of the three-dimensional Gaussian, and calculate the variance of their normal vectors var(n d ). a, β, and b are all constants. a takes 1.0, b takes 0.0002, and b takes 0.00003. For a Gaussian with a relatively large variance of the neighborhood normal vector, it can be considered that its similarity with the neighborhood Gaussian is relatively low, and the current number of Gaussians is not sufficient to fit the human body part at that place. Therefore, according to the local perception adaptive density strategy, the calculated splitting and cloning thresholds are also relatively low, and it is more likely to be densified, making the number and distribution of Gaussians more reasonable, and further making the finally rendered image have richer high-frequency details;

[0118] S9-2: The loss function of the training process is as follows:

[0119]

[0120] Among them, is the mean error between each pixel of the rendered image and the corresponding pixel of the real image; is the structural similarity loss between the rendered image and the real image, which measures the similarity between the two images. Using ECON, estimate the corresponding normal map from the observed frame as the supervision signal. At the same time, replace the color in S8 with the normal vector of the Gaussian to render the normal map, and calculate the same loss as the image for the two and λ 1 and λ 2 are the loss weights.

[0121] Correspondingly, the present invention also provides a local information-guided human free-viewpoint rendering system, as Figure 2 shown, including:

[0122] The data acquisition unit 1 is used to input a single-view human video dataset, and for each frame in the video in the dataset, a corresponding Skinned Multi-Person Linear (SMPL) model is fitted, and the background in the video frame is removed according to the human mask provided by the dataset, and the surface of the SMPL model is randomly sampled under a standard posture to obtain three-dimensional point cloud data;

[0123] Parameter deformation unit 2, used for constructing a three-dimensional Gaussian distribution for each point based on the three-dimensional point cloud data, initializing the properties of the Gaussian distribution, passing the position of the three-dimensional Gaussian into the multi-layer perceptron MLP_1, training to obtain the offset of the storage parameter in the Gaussian distribution, and acting on the initial three-dimensional Gaussian distribution to obtain the deformed three-dimensional Gaussian position;

[0124] The weight transformation unit 3 is used to predict the linear blend skinning (LBS) transformation weight from the standard posture to the observation frame posture through the multi-layer perceptron MLP_2 of the deformed three-dimensional Gaussian position and the joint point rotation matrix of the SMPL model, and implement the LBS transformation to obtain the Gaussian distribution of the observation frame posture;

[0125] A color prediction unit 4 is used to decompose the Gaussian color into a specular reflection component and a diffuse reflection component, and estimate them respectively based on the normal vector, and combine them to obtain the final predicted color;

[0126] The image rendering unit 5 is used to project the three-dimensional Gaussian onto a two-dimensional plane according to the viewing angle parameters, rasterize the two-dimensional Gaussian, render the two-dimensional image of the human body, and dynamically calculate the densification threshold from the neighborhood variance of the normal vector of the Gaussian according to the local perception adaptive density control strategy to control the density of the Gaussian.

[0127] Therefore, this scheme adds the normal vector to the attributes of the 3D Gaussian. As a kind of geometric information, it can guide the 3D Gaussian to pay more attention to the surface geometry of the human body during the network learning process, thereby assisting the learning of other attributes of the 3D Gaussian. This scheme takes into account that different parts of the human body have different detail granularity. Simply controlling the Gaussian density with a fixed threshold ignores local information. Therefore, a local-aware adaptive density control strategy is used to dynamically calculate the densification threshold of each Gaussian, so that the Gaussian can fit more human surface details. This scheme combines traditional graphics knowledge, divides the color into specular reflection component and diffuse reflection component, and considers the different factors affecting the two, and predicts them separately. At the same time, the normal vector also plays a prompting role in color calculation, so it is also used as the input of the color prediction model. In addition, the color decomposition strategy of this scheme can help the 3D Gaussian better fit the human surface color in areas with drastic color changes.

[0128] The above has introduced in detail the method and system for rendering a human free viewpoint guided by local information provided by the embodiments of the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A human body free viewpoint rendering method guided by local information, characterized in that: The method comprises: Input a single-view human video dataset. For each frame in the video in the dataset, fit the corresponding skin multi-human linear SMPL model. At the same time, remove the background in the video frame according to the human mask provided by the dataset. In the standard posture, randomly sample the surface of the SMPL model to obtain 3D point cloud data. Based on the three-dimensional point cloud data, a three-dimensional Gaussian distribution is constructed for each point, and the attributes of the Gaussian distribution are initialized. The position of the three-dimensional Gaussian is passed to the multi-layer perceptron MLP_1, and the offset of the storage parameter in the Gaussian distribution is obtained by training, and the offset is applied to the initial three-dimensional Gaussian distribution to obtain the deformed three-dimensional Gaussian position; The deformed three-dimensional Gaussian position and the joint point rotation matrix of the SMPL model are used to predict the linear mixed skinning (LBS) transformation weights from the standard posture to the observation frame posture through a multi-layer perceptron MLP_2, and the LBS transformation is performed to obtain the Gaussian distribution of the observation frame posture; Decompose the Gaussian color into specular reflection component and diffuse reflection component, estimate them separately based on the normal vector, and combine them to get the final predicted color; According to the viewing angle parameters, the 3D Gaussian is projected onto a 2D plane, and the 2D Gaussian is rasterized to render a 2D image of the human body. Based on the local perception adaptive density control strategy, the densification threshold is dynamically calculated by the neighborhood variance of the Gaussian's normal vector to control the density of the Gaussian.

2. A human body free viewpoint rendering method guided by local information as claimed in claim 1, characterized in that: The SMPL model is specifically: The SMPL model is a vertex skinning-based model that can accurately represent various body shapes in natural human body postures. The parameters of the model include shape and posture parameters, a standard posture template, and predefined hybrid skinning weights. The shape parameters describe the approximate shape of the human body. The standard posture refers to a posture in which the human body's arms and legs are spread out in a "big" shape. The posture parameters in the SMPL model describe the degree of rotation and translation of the human body's joints relative to the standard posture. The standard posture template is a human body SMPL model in a standard posture, which is composed of a number of points and triangular patches formed by connecting the points.

3. The method for rendering a human body free viewpoint guided by local information as claimed in claim 1, characterized in that: Based on the three-dimensional point cloud data, a three-dimensional Gaussian distribution is constructed for each point, and the properties of the Gaussian distribution are initialized. The position of the three-dimensional Gaussian is passed to the multi-layer perceptron MLP_1, and the offset of the storage parameter in the Gaussian distribution is obtained by training, and applied to the initial three-dimensional Gaussian distribution to obtain the deformed three-dimensional Gaussian position, which is specifically: The mathematical definition of the three-dimensional Gaussian distribution is as follows: Among them, x is a point cloud point, μ is the center of the three-dimensional Gaussian distribution, that is, the position mean, Σ is the covariance matrix, which is used to represent the size, shape and direction of the three-dimensional Gaussian, which is calculated by the scaling vector s and the rotation matrix R. In the three-dimensional Gaussian distribution, the following parameters are stored: position μ, rotation quaternion q, scaling vector s, normal vector n, opacity parameter α, f s is the specular color feature, f d is the diffuse color feature. The rotation quaternion is initialized to q = [1, 0, 0, 0], which is equivalent to the unit matrix in the rotation matrix, that is, no rotation is performed; the scaling vector s is initialized to the logarithm of the distance from the point to the three nearest neighboring points; the opacity parameter α is initialized to 1, f s and f d Initialized to a full zero vector, where the rotation quaternion q = [w, x, y, z] is converted into a rotation matrix R using the following formula: Initialize the three-dimensional Gaussian according to the three-dimensional point cloud data, wherein the position μ is initialized to the coordinates of the point cloud, and the normal vector n is initialized to the surface normal vector of the SMPL triangle patch where the sampling point is located; Train a multi-layer perceptron MLP_1 to learn the offset of three-dimensional Gaussian attributes, specifically: (Δμ,Δq,Δs,Δn)=MLP_1(μ) The output is the offset of position μ, rotation quaternion q, scaling vector s, and normal vector n, that is, the offset of the storage parameters, which is used to fit the offset of clothing relative to the human body surface. MLP_1 consists of an input layer, two hidden layers, and an output layer. The width of the hidden layer is 128. The predicted offset is applied to the initialized three-dimensional Gaussian distribution using the following formula: The · operation between q and Δq is equivalent to converting the two into matrices and then performing matrix multiplication. e is an exponential function. The predicted normal vector offset needs to be converted into a rotation matrix and then multiplied by the original normal vector n.

4. The method for rendering a human body free viewpoint guided by local information as claimed in claim 3, characterized in that: The deformed three-dimensional Gaussian position and the joint rotation matrix of the SMPL model are used to predict the linear mixed skinning LBS transformation weights from the standard posture to the observation frame posture through the multi-layer perceptron MLP_2, and the LBS transformation is implemented to obtain the Gaussian distribution of the observation frame posture, specifically: The initialization of the three-dimensional Gaussian and the deformation of the fitted clothing are completed in the standard posture space. The linear blending skinning (LBS) transformation is required to achieve the transformation from the standard posture space to the observed posture space. The LBS transformation regards the changes of the vertices on the human body surface due to the action as driven by the skeleton. The trajectory changes of the points are obtained by weighted calculation of the transformation of the joint points on the skeleton, that is, the position of each vertex is the weighted average of the vertex positions after the rigid skinning of the joints that affect it. The SMPL model predicted from the observation frame contains the posture parameters. The human joint bone rotation matrix is ​​calculated from the posture parameters. The deformed three-dimensional Gaussian position and the SMPL joint rotation matrix are passed to the multi-layer perceptron MLP_2 to predict the weight of the LBS transformation: in, is the bone rotation matrix, k represents the kth joint point, K represents the total number of joint points, and w k Represents the transformation of the kth joint point relative to μ d The weight of MLP_2 is composed of an input layer, three hidden layers, and an output layer. The width of the hidden layer is 128. After obtaining the LBS weight of the three-dimensional Gaussian through MLP_2, the LBS transformation is implemented according to the following formula: Among them, T is the transformation matrix, which is obtained by multiplying the weight of the bone predicted by MLP_2 by the bone transformation matrix calculated from SMPL, and R d is the rotation quaternion q d The rotation matrix calculated in .

5. The method for rendering a human body free viewpoint guided by local information as claimed in claim 3, characterized in that: The Gaussian color is decomposed into a specular reflection component and a diffuse reflection component, and they are estimated based on the normal vector, and the final predicted color is obtained by combining them, specifically: The Gaussian color is decomposed into a specular reflection component and a diffuse reflection component, and the normal vector is used as one of the bases for predicting the two color components, as follows: Among them, f s is the specular color feature stored in Gaussian, f d is the diffuse color feature stored in the Gaussian, d is the camera angle, c s is the specular component, c d It is the diffuse reflection component. The structures of MLP_3 and MLP_4 are the same, both of which consist of an input layer, a hidden layer and an output layer. The width of the hidden layer is 64. The two components are added: c=c s +c d Thus, the final predicted color c is obtained.

6. The method for rendering a human body free viewpoint guided by local information as claimed in claim 5, characterized in that: According to the viewing angle parameter, the three-dimensional Gaussian is projected onto a two-dimensional plane, the two-dimensional Gaussian is rasterized, and a two-dimensional image of a human body is rendered. According to the local perception adaptive density control strategy, the densification threshold is dynamically calculated by the neighborhood variance of the normal vector of the Gaussian, and the density of the Gaussian is controlled, specifically: According to the viewing angle parameter, the three-dimensional Gaussian is projected onto a two-dimensional plane to obtain the corresponding two-dimensional Gaussian, as follows: Among them, i represents the i-th Gaussian distribution; α i represents the opacity parameter of the i-th Gaussian, μ i represents the two-dimensional mean coordinates of the i-th three-dimensional Gaussian after projection, Σ′ i =(JWΣ i W T J T ) 1:2,1:2 is the two-dimensional covariance matrix obtained after the projection of the i-th three-dimensional Gaussian, where W is the transformation matrix of the camera from the world coordinate system to the camera coordinate system, J is the Jacobian matrix of the perspective projection transformation; p is the coordinate of each pixel point, and e is the exponential function; The final RGB color value of each pixel in the rendered image is obtained by the alpha blending method, and the calculation formula is as follows: Among them, c i is the final predicted color of the ith i is the probability value of the position of pixel p after projection for the i-th two-dimensional Gaussian distribution, and the result C obtained after α mixing is the RGB color value of pixel p; Implement a local-aware adaptive density control strategy on 3D Gaussians to dynamically adjust the number and distribution of 3D Gaussians; The normal vector is regarded as a hint of the surface details of the human body, and the local variance of the normal vector is regarded as a measure of the local similarity of the Gaussian. During the splitting and cloning process, the Gaussian densification threshold is dynamically calculated as follows: Among them, var is the variance operation. According to the knnk nearest neighbor algorithm, find the k Gaussians closest to the three-dimensional Gaussian position and calculate their normal vector variance var(n d ), a, β and b are all constants. According to the local-aware adaptive density strategy, the calculated splitting and cloning thresholds are lower, and densification is more likely to be performed, so that the number and distribution of Gaussians are more reasonable, and further the final rendered image has richer high-frequency details; The loss function of the training process is as follows: in, is the mean error between each pixel of the rendered image and the corresponding pixel of the real image, It is the structural similarity loss between the rendered image and the real image, which measures the similarity between the two images. Using ECON, the corresponding normal map is estimated from the observed frame as a supervisory signal. At the same time, the RGB color value c of the pixel p is replaced by the Gaussian normal vector to render the normal map. The same loss as the image is calculated for the two normal maps. and λ1 and λ2 are loss weights.

7. A human body free viewpoint rendering system guided by local information, characterized in that: The system comprises: The data acquisition unit is used to input a single-view human video dataset, and for each frame in the video in the dataset, the corresponding skin multi-human linear SMPL model is fitted, and the background in the video frame is removed according to the human mask provided by the dataset, and the surface of the SMPL model is randomly sampled under a standard posture to obtain three-dimensional point cloud data; A parameter deformation unit is used to construct a three-dimensional Gaussian distribution for each point based on the three-dimensional point cloud data, initialize the properties of the Gaussian distribution, pass the position of the three-dimensional Gaussian into the multi-layer perceptron MLP_1, train to obtain the offset of the storage parameter in the Gaussian distribution, and act on the initial three-dimensional Gaussian distribution to obtain the deformed three-dimensional Gaussian position; A weight transformation unit is used to predict the linear blending skinning (LBS) transformation weights from the standard posture to the observation frame posture through the multi-layer perceptron MLP_2 of the deformed three-dimensional Gaussian position and the joint point rotation matrix of the SMPL model, and implement LBS transformation to obtain the Gaussian distribution of the observation frame posture; A color prediction unit is used to decompose the Gaussian color into specular reflection component and diffuse reflection component, and estimate them separately based on the normal vector, and combine them to obtain the final predicted color; The image rendering unit is used to project the three-dimensional Gaussian onto a two-dimensional plane according to the viewing angle parameters, rasterize the two-dimensional Gaussian, render the two-dimensional image of the human body, and dynamically calculate the densification threshold from the neighborhood variance of the normal vector of the Gaussian according to the local perception adaptive density control strategy to control the density of the Gaussian.

8. The human body free viewpoint rendering system guided by local information as claimed in claim 7, characterized in that: The SMPL model is specifically: The SMPL model is a vertex skinning-based model that can accurately represent various body shapes in natural human body postures. The parameters of the model include shape and posture parameters, a standard posture template, and predefined hybrid skinning weights. The shape parameters describe the approximate shape of the human body. The standard posture refers to a posture in which the human body's arms and legs are spread out in a "big" shape. The posture parameters in the SMPL model describe the degree of rotation and translation of the human body's joints relative to the standard posture. The standard posture template is a human body SMPL model in a standard posture, which is composed of a number of points and triangular patches formed by connecting the points.

9. The human body free viewpoint rendering system guided by local information as claimed in claim 7, characterized in that: The parameter deformation unit is used to construct a three-dimensional Gaussian distribution for each point based on the three-dimensional point cloud data, initialize the properties of the Gaussian distribution, pass the position of the three-dimensional Gaussian into the multi-layer perceptron MLP_1, train to obtain the offset of the storage parameter in the Gaussian distribution, and act on the initial three-dimensional Gaussian distribution to obtain the deformed three-dimensional Gaussian position, specifically: The mathematical definition of the three-dimensional Gaussian distribution is as follows: Among them, x is a point cloud point, μ is the center of the three-dimensional Gaussian distribution, that is, the position mean, Σ is the covariance matrix, which is used to represent the size, shape and direction of the three-dimensional Gaussian, which is calculated by the scaling vector s and the rotation matrix R. In the three-dimensional Gaussian distribution, the following parameters are stored: position μ, rotation quaternion q, scaling vector s, normal vector n, opacity parameter α, f s is the specular color feature, f d is the diffuse color feature. The rotation quaternion is initialized to q = [1, 0, 0, 0], which is equivalent to the unit matrix in the rotation matrix, that is, no rotation is performed; the scaling vector s is initialized to the logarithm of the distance from the point to the three nearest neighboring points; the opacity parameter α is initialized to 1, f s and f d Initialized to a full zero vector, where the rotation quaternion q = [w, x, y, z] is converted into a rotation matrix R using the following formula: Initialize the three-dimensional Gaussian according to the three-dimensional point cloud data, wherein the position μ is initialized to the coordinates of the point cloud, and the normal vector n is initialized to the surface normal vector of the SMPL triangle patch where the sampling point is located; Train a multi-layer perceptron MLP_1 to learn the offset of three-dimensional Gaussian attributes, specifically: (Δμ,Δq,Δs,Δn)=MLP_1(μ) The output is the offset of position μ, rotation quaternion q, scaling vector s, and normal vector n, that is, the offset of the storage parameters, which is used to fit the offset of clothing relative to the human body surface. MLP_1 consists of an input layer, two hidden layers, and an output layer. The width of the hidden layer is 128. The predicted offset is applied to the initialized three-dimensional Gaussian distribution using the following formula: The · operation between q and Δq is equivalent to converting the two into matrices and then performing matrix multiplication. e is an exponential function. The predicted normal vector offset needs to be converted into a rotation matrix and then multiplied by the original normal vector n.

10. The human body free viewpoint rendering system guided by local information as claimed in claim 9, characterized in that: The weight transformation unit is used to predict the linear blending skinning LBS transformation weights from the standard posture to the observation frame posture through the deformed three-dimensional Gaussian position and the SMPL joint point rotation matrix through the multi-layer perceptron MLP_2, and implement the LBS transformation to obtain the Gaussian distribution of the observation frame posture, specifically: The initialization of the three-dimensional Gaussian and the deformation of the fitted clothing are completed in the standard posture space. The linear blending skinning (LBS) transformation is required to achieve the transformation from the standard posture space to the observed posture space. The LBS transformation regards the changes of the vertices on the human body surface due to the action as driven by the skeleton. The trajectory changes of the points are obtained by weighted calculation of the transformation of the joint points on the skeleton, that is, the position of each vertex is the weighted average of the vertex positions after the rigid skinning of the joints that affect it. The SMPL model predicted from the observation frame contains the posture parameters. The human joint bone rotation matrix is ​​calculated from the posture parameters. The deformed three-dimensional Gaussian position and the SMPL joint rotation matrix are passed to the multi-layer perceptron MLP_2 to predict the weight of the LBS transformation: in, is the bone rotation matrix, k represents the kth joint point, K represents the total number of joint points, and w k Represents the transformation of the kth joint point relative to μ d The weight of MLP_2 is composed of an input layer, three hidden layers, and an output layer. The width of the hidden layer is 128. After obtaining the LBS weight of the three-dimensional Gaussian through MLP_2, the LBS transformation is implemented according to the following formula: Among them, T is the transformation matrix, which is obtained by multiplying the weight of the bone predicted by MLP_2 by the bone transformation matrix calculated from SMPL, and R d is the rotation quaternion q d The rotation matrix calculated in .

11. The human body free viewpoint rendering system guided by local information as claimed in claim 9, characterized in that: The color prediction unit is used to decompose the Gaussian color into a specular reflection component and a diffuse reflection component, and estimate them respectively based on the normal vector, and combine them to obtain the final predicted color, specifically: The Gaussian color is decomposed into a specular reflection component and a diffuse reflection component, and the normal vector is used as one of the bases for predicting the two color components, as follows: Among them, f s is the specular color feature stored in Gaussian, f d is the diffuse color feature stored in the Gaussian, d is the camera angle, c s is the specular component, c d It is the diffuse reflection component. The structures of MLP_3 and MLP_4 are the same, both of which consist of an input layer, a hidden layer and an output layer. The width of the hidden layer is 64. The two components are added: c=c s +c d Thus, the final predicted color c is obtained.

12. The human body free viewpoint rendering system guided by local information according to claim 11, characterized in that: The image rendering unit is used to project the three-dimensional Gaussian onto a two-dimensional plane according to the viewing angle parameter, rasterize the two-dimensional Gaussian, render the two-dimensional image of the human body, and dynamically calculate the densification threshold from the neighborhood variance of the normal vector of the Gaussian according to the local perception adaptive density control strategy, and perform density control on the Gaussian, specifically: According to the viewing angle parameter, the three-dimensional Gaussian is projected onto a two-dimensional plane to obtain the corresponding two-dimensional Gaussian, as follows: Among them, i represents the i-th Gaussian distribution; α i represents the opacity parameter of the i-th Gaussian, μ i represents the two-dimensional mean coordinates of the i-th three-dimensional Gaussian after projection, Σ′ i =(JWΣ i W T J T ) 1:2,1:2 is the two-dimensional covariance matrix obtained after the projection of the i-th three-dimensional Gaussian, where W is the transformation matrix of the camera from the world coordinate system to the camera coordinate system, J is the Jacobian matrix of the perspective projection transformation; p is the coordinate of each pixel point, and e is the exponential function; The final RGB color value of each pixel in the rendered image is obtained by the alpha blending method, and the calculation formula is as follows: Among them, c i is the final predicted color of the ith i is the probability value of the position of pixel p after projection for the i-th two-dimensional Gaussian distribution, and the result C obtained after α mixing is the RGB color value of pixel p; Implement a local-aware adaptive density control strategy on 3D Gaussians to dynamically adjust the number and distribution of 3D Gaussians; The normal vector is regarded as a hint of the surface details of the human body, and the local variance of the normal vector is regarded as a measure of the local similarity of the Gaussian. During the splitting and cloning process, the Gaussian densification threshold is dynamically calculated as follows: Among them, var is the variance operation. According to the knnk nearest neighbor algorithm, find the k Gaussians closest to the three-dimensional Gaussian position and calculate their normal vector variance var(n d ), a, β and b are all constants. According to the local-aware adaptive density strategy, the calculated splitting and cloning thresholds are lower, and densification is more likely to be performed, so that the number and distribution of Gaussians are more reasonable, and further the final rendered image has richer high-frequency details; The loss function of the training process is as follows: in, is the mean error between each pixel of the rendered image and the corresponding pixel of the real image, It is the structural similarity loss between the rendered image and the real image, which measures the similarity between the two images. Using ECON, the corresponding normal map is estimated from the observed frame as a supervisory signal. At the same time, the RGB color value C of the pixel p is replaced by the Gaussian normal vector to render the normal map. The same loss as the image is calculated for the two normal maps. and λ1 and λ2 are loss weights.

Citation Information

Patent Citations

  • Three-dimensional virtual image expression animation generation method based on reality rendering technology

    CN115937387A

  • Dynamic human body modeling method based on three-dimensional Gaussian

    CN117671108A

  • Virtual human arbitrary view angle rendering method and system based on three-dimensional Gaussian spattering

    CN118736092A

  • Three dimensional gaussian splatting initialization based on trained neural radiance field representations

    US20240355047A1

Cited By

  • Three-order training method and system for dynamic Gaussian role video

    CN121170138A

  • Third-order training method and system for dynamic gaussian role video

    CN121170138B

  • Virtual digital human generation method and electronic equipment

    CN121353562A

  • Method and system for displaying human body image at any visual angle of virtual digital human

    CN121921423A

  • Virtual digital person arbitrary view human image display method and system

    CN121921423B