A human body reconstruction and rendering method and device cooperating with light field and occupancy field

CN117315153BActive Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311273506.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2026-09-25
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

由此,提出一种可从用户级设备稀疏视图输入中进行高质量人体自由视角渲染的方法仍然是一个挑战

Benefits of technology

[0086](1)提出了一种新型的人体捕捉方法,可以在100ms左右从稀疏RGB输入视图中创建1K分辨率的自由视角视频。该方法可以泛化到未见过的表演者上,而无需经过进一步优化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315153B_ABST
    Figure CN117315153B_ABST
Patent Text Reader

Abstract

The application discloses a human body reconstruction and rendering method and device cooperating with light fields and occupancy fields, which can be directly applied to new shooting targets. The method integrates PIFu and NeRF expression, and makes the two expressions work cooperatively by predicting human body occupancy fields and light fields respectively. PIFu depends on a denoised depth image, so that the model introduces human body priori, and assists the new view rendering task. The application introduces a network model SRONet to process geometry and human body parts simultaneously, and uses the occupancy field to assist the light field drawing. In the training process, the geometry and color supervision signals are applied to the network model simultaneously, so as to enhance the ability of the network to capture high-quality texture details. The application also introduces a light ray up-sampling method based on neural network fusion, a depth image denoising model, and a two-layer tree data structure based on the denoised depth image, so as to efficiently up-sample a low-resolution image to a target resolution and efficiently sample a light ray rendering point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer three-dimensional visual reconstruction and computer graphics rendering, and in particular to a method and apparatus for human body reconstruction and rendering based on cooperative light field and occupancy field. Background Technology

[0002] Human 3D reconstruction and the creation of free-viewpoint videos of the human body are crucial components in many applications, such as virtual reality and augmented reality, distance education, and virtual conferencing. To provide users with an immersive experience, these applications require the ability to capture high-quality human models as close to real-time as possible using consumer-grade capture devices and render them from a free-viewpoint perspective. Recently, neural implicit field representations have been widely used in human performance capture systems. Pixel-aligned implicit functions (PIFu) can efficiently reconstruct the mesh and texture of dynamic human 3D models. The surface model is extracted through a trained implicit occupancy field, while the texture is obtained by predicting the RGB values ​​of surface points using a trained network. Neural network-based light field (NeRF) is a coordinate-based implicit network model that encodes volume density and color fields. NeRF is widely popular because it can render photorealistic images with dense point sampling. However, both representations have certain limitations in the field of 3D human reconstruction. First, PIFu-based rendering often produces blurry results and cannot produce viewpoint-dependent rendering effects, nor can it handle materials such as translucent hair. Second, NeRF rendering speed is often insufficient for real-time scenes, and its generalization ability is poor. Even the latest generalizable NeRF cannot effectively reconstruct target objects not seen in the training set from sparse views. For new target objects, high rendering quality is usually only achieved after online optimization. Therefore, proposing a method for high-quality free-view rendering of human bodies from sparse view input from user-level devices remains a challenge. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention proposes a method for human body reconstruction and rendering based on collaborative light field and occupancy field, which can be directly generalized to rendering unseen target human bodies and has low rendering latency.

[0004] When PIFu and NeRF, both based on depth data, work together, NeRF can generate high-quality texture results using global light field information. Geometric ambiguity or noise can be further reduced by the occupancy field of PIFu (due to the introduction of geometric constraints), which helps improve the quality of surface reconstruction. Furthermore, when PIFu and NeRF work together, since surface points contribute the most to light color, the rendering quality requires high depth accuracy. When the input of PIFu depends only on the input RGB image features, the geometric quality reconstructed from a sparse perspective cannot be guaranteed. By introducing accurate depth information, the geometric surface field constructed by PIFu can be constrained, allowing the light field model image features to better align with the surface model, thus benefiting NeRF learning.

[0005] Based on the above observations, the method of this invention first introduces a novel depth image denoising model, which references the architecture of UNet, outputting an optimized depth image to reduce depth noise and fill in holes in the original depth image. Then, a novel network model, SRONet, is introduced to model the human body by combining occupancy field and light field. Specifically, SRONet reconstructs the human body by predicting the human occupancy field based on pixel-aligned denoised depth map features, and predicts the light field based on geometric features and pixel-aligned RGB features to render viewpoint-dependent human body textures. During training, geometric and color supervision signals work together to enhance SRONet's ability to capture high-quality details during reconstruction and rendering. Furthermore, this invention constructs a novel two-layer tree structure from the denoised depth map for efficient geometric storage, fast ray-voxel intersection, and 3D point sampling. Additionally, this invention constructs a light line sampling method based on neural network fusion, enabling 1K resolution rendering from new viewpoints with relatively low computational overhead.

[0006] To achieve the above, the technical solution of the present invention is as follows: a method for human body reconstruction and rendering based on coordinated light field and occupancy field, the method comprising the following steps:

[0007] S1: Construct a deep image denoising model F based on a convolutional neural network structure d The model's input includes a human RGB image I and an original depth image D, and its output is a denoised depth image D. rf The model is trained using depth, normal, and 3D consistency loss functions, and then used for deep denoising tasks.

[0008] S2: Based on the denoised multi-view depth image and the camera parameters of each viewpoint, a global human point cloud is fused. A two-layer tree data structure is constructed from the point cloud, and the parent and child nodes of the tree are stored as a global list. During the inference process, voxel post-processing operations are introduced to denoise the tree structure, retaining only voxels near the human body surface.

[0009] S3: Construct a collaborative network model SRONet for human body mesh reconstruction and new perspective image rendering. This network model comprises two sub-network models: the occupancy field network OCCNet and the light field network ColorNet. For a given input: a multi-view RGBD image... And the sampling point x in three-dimensional space, the viewing direction d, where N is the number of views, and the sub-network models respectively implement the following tasks:

[0010] Given a multi-view depth image as input And for the sampling point x, OCCNet predicts the voxel occupancy o for point x based on the pixel-aligned implicit function (PIFu). x ∈[0,1], representing the probability that the point is located inside the human body mesh model;

[0011] Given an input multi-view RGB image {I i} i=1,...N ColorNet predicts the RGB three-channel color vector c for a point x along the viewing direction d based on PIFu. x ;

[0012] Occupancy based on point x prediction o x The human body mesh model is calculated from the estimated occupancy field using the marching cubes algorithm for 3D human body reconstruction. SRONet is trained using a multi-view human body dataset, combining the geometry-color collaborative loss function and the depth error loss function. After the model is trained, it is used to predict the occupancy field and the light field.

[0013] S4: Calculate the light rays for each pixel in the new perspective image based on the corresponding camera parameters, and perform the voxel intersection process according to the voxel traversal algorithm to determine the voxels that intersect with the light rays in the tree data structure; record the depth values ​​of the near and far intersection points of all intersecting voxels, and design sampling weights along the light rays according to the voxel size within the voxels. Calculate the occupancy value and color of the sampling points according to S3; use the volume rendering formula to calculate the color fusion weights based on the occupancy value to fuse the colors of each sampling point on the light rays and calculate the final color of the light rays;

[0014] S5: Upsample each ray to improve the resolution and quality of the rendered image; construct a feature fusion network, the input of which is information shared by each sub-ray, including the original color, depth, ray features, and the input RGB images of two adjacent viewpoints; the output is the final color of each sub-ray; train the feature fusion network by ray-by-ray color error, structural similarity, and feature loss function, and use it for ray upsampling operations after training to obtain the target resolution rendered image.

[0015] Furthermore, the specific design of the depth denoising process is as follows:

[0016] (1) The BackgroundMatting-v2 algorithm is used to extract the human figure region from the original input image as the human figure region mask image Ψ, and the RGBD images I and D of the human figure region are obtained at the same time; the input RGBD images are normalized to the interval [-1, 1], and the depth maximum and minimum values ​​are recorded at the same time; the normalized RGB and depth images are merged as the depth image denoising model F. d Input;

[0017] (2) Constructing model F based on UNet structure d Two independent feature extraction networks, HRNetV2-W18-Small-v2, are used to encode RGB and depth images respectively. The dilated spatial convolutional pooling pyramid module ASPP and the residual attention module ResCBAM are used to fuse RGB and depth features, upsample the fused features, and feed them back to the feature extraction network.

[0018] (3) Model F d Output the image in the range [-1, 1]. Use the depth extrema in (1) to inverse normalize the output image back to the original value range. Use Ψ to post-process the image, retaining only the depth value of the human figure region, denoted as D. rf ;

[0019] (4) Used for model F d The training loss function is as follows:

[0020] Deep consistency loss L D : Used to penalize the true depth image at viewpoint i With the predicted depth image The per-pixel deviation is defined as:

[0021]

[0022] Normal consistency loss L N Used to penalize depth images and The deviation between the calculated normals is defined as:

[0023]

[0024] in, Represents the predicted depth image and its calculated normal diagram The depth value and normal vector at the i-th viewpoint pixel p. This represents the corresponding true depth value and normal vector; Describes the L1 loss function. The L2 loss function is represented by <·, and ·> represents the vector dot product operation;

[0025] 3D consistency loss L P Used to further constrain passage The consistency between the fused point cloud and the real point cloud, in order to reduce depth fusion noise in the construction of a two-layer tree structure, is defined as:

[0026]

[0027] in, Indicates from The fused point cloud, F is the truncated symbolic range field fusion algorithm TSDF-Fusion, K i RT i These are the camera's intrinsic and extrinsic parameters, P. gt The point cloud is sampled from a real 3D human body mesh model. The chamfer function represents the chamfer distance loss function.

[0028] The loss function used for deep denoising is expressed as: L = L D +λ N L N +λ P L P , where λ N , λ P The loss function is optimized using the ADAM algorithm, with weights assigned to it.

[0029] Furthermore, the specific design of the two-layer tree data structure is as follows:

[0030] (1) The truncated symbolic distance field fusion algorithm TSDF-Fusion is used to fuse the denoised depth images from all viewpoints. Simultaneously, the global human point cloud P is obtained. rf With Fusion Cube V tsdf ; For V tsdf After binarization, let it be denoted as V. occ The resolution can be set to 128. 3 However, it is not limited to this, the occupancy value (occ) V at spatial location x. occ (x) is:

[0031]

[0032] Among them, s v V represents the size of the voxel, α is the truncation sign distance threshold, and V tsdf (x) represents V at spatial location x. tsdf The truncation symbol distance;

[0033] (2) Voxel post-processing: Using OCCNet in S3, a cube with a set resolution is quickly constructed based on the Realtime-PIFu real-time human reconstruction algorithm, denoted as . After binarization, V is constructed using (1). occ Fusion to eliminate V occ The floating noise voxels are used to obtain the denoised cube. Then the fusion occupancy value at spatial location x for:

[0034]

[0035] in, It is a binary function based on a threshold β, where γ is the threshold used to filter floating voxels, and | represents the OR operation;

[0036] (3) The denoised cube middle The voxels are denoted as effective voxels. Parent voxels are fused according to a preset ratio (for example, if the ratio is set to 64:1, then each effective voxel within a cube with dimensions 4 corresponds to one effective parent voxel). All effective voxels are then stored in a global list L. v Each node corresponds to a voxel, and each node records the index (list position), spatial position, and size information of its parent or child node; construct an index cube V. idx This stores the index of each valid parent node in the global list; for invalid parent nodes, their index is set to -1. This is achieved by constructing V... idx To facilitate efficient calculation of ray-voxel intersection during the S4 process.

[0037] Furthermore, the specific design of the SRONet is as follows:

[0038] (1) OCCNet: Based on depth information, it uses the feature encoder HRNetV2-W18 to encode depth images; for a sampling point x, OCCNet predicts the occupancy value o corresponding to that point by pooling the pixel-aligned depth features from each viewpoint. x Define the occupancy field as a function.

[0039]

[0040] Among them, W i W represents the depth feature map of viewpoint i after encoding. i (x) represents the projection of point x onto viewpoint i from W. i The deep features obtained from c i(x) represents the depth value of point x projected onto the camera coordinate system at viewpoint i and the distance to the truncation sign; the implicit function f1, represented by a fully connected network, is used to obtain the geometric features of each viewpoint, and the fused global features are obtained through the average pooling operation Avg. The global features are then fed into the second implicit function f2 to calculate the occupancy value o. x ;

[0041] (2) ColorNet: Light field based on color and geometric features. It encodes RGB images using the same feature encoder as OCCNet to obtain color features; for sampling point x, ColorNet uses additional inputs: view direction d and geometric features. To aggregate color features from each viewpoint to predict viewpoint-related color vectors Among them, geometric features are represented as Where f3 is the implicit function used for encoding; the light field is defined as a function :

[0042]

[0043] Among them, M i M represents the color feature map of viewpoint i after encoding. i (x) and rgb i This indicates that the projection of point x onto viewpoint i is from M. i The color features and pixel colors obtained from the data; f4 and f5 represent the implicit functions used for further feature processing; It is a feature fusion function implemented using a transformer, with the basic feature fusion unit implemented using Hydra Attention; the view direction d in the camera coordinate system. i =R i d, where R i This represents the rotation matrix in the extrinsic parameters of the camera at viewpoint i;

[0044] SRONet predicts the occupancy value of the sample point x and the view-dependent color.

[0045] Furthermore, the specific design of ray-voxel intersection and sampling in S4 is as follows:

[0046] (1) For a emitted ray l, use a voxel traversal algorithm to detect the parent nodes that intersect along ray l, and record the depth value of the intersection point of the intersecting parent voxels from the viewpoint of l. The depth values ​​of the near and far intersection points are denoted as D. far With D near For each intersecting parent voxel, continue using the voxel traversal algorithm to detect intersecting child voxels, and record the near and far depth values, denoted as D′. far With D′ nearThe sum of the depth values ​​for all records is as follows: and

[0047] (2) Assign the number of sampling points between the near and far intersections of each voxel, and calculate the sampling weight w of the i-th voxel. i as follows:

[0048]

[0049] Where, d far (i) and d near (i) respectively represent and The near and far depths of voxel i, N v s represents the number of all intersecting voxels. i The scale of voxel i can be represented by setting the parent and child voxels to 1 and 4 respectively, so that more points can be allocated to the child voxel; specifically, the number of sampling points m within each voxel i. i for:

[0050]

[0051] Where M is the total number of sampling points, This indicates rounding down; during sampling, sampling points are assigned to each voxel i along the direction of light emission to ensure that the voxel closest to the camera on the light ray is always sampled; if ∑ i m i If M < M, then the remaining M-∑ will be redistributed according to the sampling order. i m i The depth value d of the j-th sampling point within the i-th voxel. i (j) The calculation formula is as follows:

[0052]

[0053] Where j starts from 0; sampling point x passes through P cam +d i P is calculated using (j)·d. cam d represents the camera position, and d represents the viewing direction.

[0054] Furthermore, the volume rendering described in S4 and the loss function of SRONet used for training the model described in S3 are specifically designed as follows:

[0055] (1) To calculate the final color of ray l, a normalized surface (Uni Surf) and volume rendering technique are used, based on the occupancy value o of each sampling point on l. x Calculate the color blending weights and use the blending point color cx to calculate the color of the light rays. as follows:

[0056]

[0057] Wherein, the color fusion weight ω of the i-th sampling point x (i)=o x (i)Π j<i (1-o x (j)); At the same time, by weighting the depth value d of the i-th sampling point i To calculate the depth value at the intersection of ray l and the human body surface.

[0058] (2) Color based on light estimation With depth value The following loss function is designed to train and optimize the reconstruction and rendering parts of SRONet:

[0059] Geometric-color collaborative loss L syn : Based on PIFu-based supervision, OCCNet is trained by sampling points y in space, and the estimated occupancy value o is penalized. y Compared with actual occupancy value Errors between them; and penalties for each ray's true color C * (l) and estimated color The error is due to the two loss functions working together:

[0060]

[0061] Where S and R represent the set of sampling points and the set of rays, respectively. Represents the cross-entropy loss function. Let μ represent the L1 loss function. o μ c These are the weights of the occupancy value loss term and the color loss term, respectively.

[0062] Depth error loss L D′ Used to penalize estimated ray depth values Compared with the true depth value D * (l) error, to further improve reconstruction and rendering details:

[0063]

[0064] in, Represents the L2 loss function;

[0065] The loss function used for SRONet is expressed as: L syn +λ D′ L D′ , where λ D′ It is a balancing term, and the loss function is optimized using the ADAM algorithm.

[0066] Furthermore, the specific design of ray upsampling in S5 is as follows:

[0067] (1) Ray upsampling: For a ray l passing through pixel position (x, y), its color is upsampled. Depth value Light fusion features The values ​​are distributed to four sub-pixels, corresponding to the positions: (x, y), (x+0.5, y), (x, y+0.5), (x+0.5, y+0.5); where ft color Feature fusion function The output features; thus generating a coarse upsampling result;

[0068] (2) Feature fusion operation: The original resolution RGB images of two adjacent viewpoints n0 and n1 of the target viewpoint are used to enhance the result in (1); specifically, UNet is used to encode two adjacent RGB images, each sub-ray corresponds to a sub-pixel, and the surface position of the sub-ray is calculated by the sub-ray depth value. Projecting point p onto neighboring images and features to obtain color. and feature and Computational visibility and The formula for calculating visibility is: z i This represents the depth value of p projected at viewpoint i. Represents the denoised depth image from viewpoint i. The depth value obtained from σ v It is a preset visibility weight coefficient; through a feature fusion network Calculate fusion weights The three colors of the merged sub-ray: The final color of the sub-ray is obtained; defined as follows:

[0069]

[0070] Where f6 represents the implicit function for processing features, The original resolution RGB images of adjacent viewpoints n0 and n1;

[0071] (3) Training the feature fusion network Loss function:

[0072] Ray-by-ray color error and structural similarity loss L B Used to penalize sampled color blocks With real color blocks The color error between them is defined as follows:

[0073]

[0074] Where R represents the set of light rays, It is of size S patch ·S patch color blocks, It is the final estimated color of the ray r located at (i, j); Let μ1 denote the L1 loss function, and SSIM denote the structural similarity function; μ1 and μ2 are the balance terms of the loss function.

[0075] Feature loss function L ft Used to punish color blocks and The feature error between the two is used to further enhance the quality of the rendered image; the loss function is calculated using a pre-trained VGG-16 network, and the loss function is defined as follows:

[0076]

[0077] in, The L1 loss function represents the loss between VGG features. Specifically, the loss function is calculated using the three features fed into the first three max pooling layers, MaxPool2d.

[0078] For feature fusion networks The loss function is expressed as: L B +μ vgg ·L ft , where μ vgg It is a balancing term, trained independently using the ADAM algorithm. The parameters.

[0079] Furthermore, the specific design for the parallel accelerated rendering process is as follows:

[0080] (1) Use two GPUs to accelerate the rendering process; each GPU processes half of the data (image and light) and synchronizes the two batches of data on the CPU via memory. The rendering process is accelerated by pipeline.

[0081] Specifically, the pipeline is divided into three parts, with each GPU accelerated by three independent data streams: 1. I / O process (CPU to GPU) and depth denoising process; 2. Constructing a two-layer tree data structure, multi-view RGB image encoding and depth image encoding in SRONet, and image encoding in ray upsampling; 3. Ray sampling, calculating point occupancy values ​​and colors through SRONet, and ray upsampling based on a feature fusion network; finally, all calculated rays are converted into images for display.

[0082] (2) Use TensorRT technology for half-precision quantization and acceleration of the depth image denoising model, multi-view RGB image encoding and depth image encoding in SRONet, and image encoding in ray upsampling; use the fully fused scheme to accelerate all implicit functions and feature fusion functions implemented by transformer through GPU shared memory; the accelerated model can render a new perspective image of 1K resolution in about 100ms.

[0083] The present invention also provides a human body reconstruction and rendering apparatus based on coordinated light field and occupancy field, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement the above-mentioned human body reconstruction and rendering method based on coordinated light field and occupancy field.

[0084] The present invention also provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the above-described method for human body reconstruction and rendering based on the cooperative light field and occupancy field.

[0085] The applicant trained the designed model on a synthetic human dataset and tested it on a public dataset. Compared with other methods, this invention achieved better results. In summary, the main beneficial effects of this invention are as follows:

[0086] (1) A novel human capture method is proposed, which can create 1K resolution free-view video from a sparse RGB input view in about 100ms. The method can be generalized to unseen performers without further optimization.

[0087] (2) A hybrid representation and a novel network model, SRONet, are proposed. The model is based on the results of deep denoising, utilizes pixel-aligned RGBD features, and coordinates occupancy field and light field to perform accurate human body reconstruction and rendering.

[0088] (3) A two-layer tree-based data structure is proposed for effective point sampling, a neural fusion-based optical line sampling technique is proposed for rapidly improving the resolution of rendered images, and a parallel computing pipeline is proposed for rendering acceleration. Attached Figure Description

[0089] Figure 1This is an overview diagram of the human body reconstruction and rendering method based on the coordinated light field and occupancy field in this embodiment of the invention. Given an RGBD stream image captured by a sparse Azure Kinect sensor as input, (1) a depth image denoising model removes depth noise based on the input RGBD image and fills in the original depth holes; (2) given a denoised depth image, a novel two-layer tree structure is reconstructed for discretizing and storing global geometric information; (3) efficient ray-voxel intersection and ray point sampling for rendering are performed; (4) a novel network model SRONet is proposed to coordinate the light field and occupancy field for human body reconstruction and free-view rendering; (5) the output of SRONet is improved to the target resolution by a neural fusion-based light line sampling technique.

[0090] Figure 2 This is the specific process of constructing a two-layer tree structure in this embodiment of the invention. It includes (1) converting the denoised depth image into a full-body point cloud; (2) constructing a cube based on the point cloud, where voxels are stored in the GPU; (3) merging smaller voxels at a ratio of 64:1 as parent voxels to construct a two-layer tree structure, and converting the voxels into a node list for storage; (4) during inference, using OCCNet, the cube is quickly constructed based on the Realtime-PIFu real-time human reconstruction algorithm. After binarization, it is compared with the first-level cube V of the tree structure. occ The floating noise voxels are fused together to obtain a denoised cube.

[0091] Figure 3 This describes the specific process of ray-voxel intersection and ray point sampling in this embodiment of the invention. A voxel traversal method is used to determine the parent and child voxels that intersect with the rays in the tree structure.

[0092] Figure 4 This is the structure diagram of the proposed SRONet and the supervision signal. SRONet consists of two parts, OCCNet and ColorNet, which predict the input point occupancy value and color, respectively. During the supervision process, the geometric loss function and the color loss function are trained together.

[0093] Figure 5 This describes the proposed ray upsampling and neural fusion process. Sub-rays share the color, depth, and features of the main emitted ray. For each sub-pixel, neural fusion takes features from neighboring viewpoints, viewpoint direction, and visibility as input, predicting weights to fuse the sub-pixel color with the colors from neighboring viewpoints. Detailed Implementation

[0094] The main module embodiments of the present invention will be further described below with reference to the accompanying drawings, and the effectiveness verification experiments of the present invention will be explained in order to enable those skilled in the art to better understand the present invention.

[0095] like Figure 1 (2) and Figure 2 As shown, the two-layer tree data structure in this invention is constructed based on the denoised depth point cloud. Constructing a two-layer tree structure can effectively limit the rendering range to the vicinity of the three-dimensional surface of the human body, thus playing a role in geometric constraint. In addition, storing the tree structure in the GPU facilitates fast voxel search for ray voxel intersection and efficient voxel intra-point sampling. The tree is designed as two layers mainly because: (1) the invisible areas of the point cloud cannot be covered by the voxels of the leaf layer, which will lead to missing data during rendering; (2) the two-layer structure uses large voxels to cover the whole body area, and higher layers require more overhead without performance improvement. During storage, the parent and child voxels of the tree structure are converted into a global list L. v In a GPU, each node stores information such as index, size, center position, and parent (child) voxel index.

[0096] like Figure 1 (2) and Figure 3 As shown, given the emitted ray l, the voxel traversal algorithm is first used ( Figure 3 (c) and (d) detect valid parent voxels that intersect with l and record the depth value of the ray-voxel intersection point at the target viewpoint t. Figure 3 In example (c), a total of 4 valid voxels intersect with ray l, and the recorded near and far depth values ​​are: D near =d({v i} i=0,1,2,4,5 RT t ), D far =d({v i} i=1,2,3,5,6 RT t ). Where d(·, RT t ) is a projection function used to obtain the depth value of a 3D point at a viewpoint t. RT t Let be the camera extrinsic parameters at viewpoint t. Then, continue using the voxel traversal algorithm to detect all valid child voxels and record their near and far depth values ​​as a list: D′ near , D′ far ,like Figure 3 As shown in (d). Finally, all near and far depths are merged into a global list, recorded as follows. and The final ray-voxel intersection result determines the point sampling range. The point sampling process is performed on all recorded near and far depth data pairs, for example, sampling points in {v0, v1} and {q0, q1}. Ultimately, all sampled points on ray l are used to calculate the ray color.

[0097] The training dataset involved in this invention consists of a human body rendering dataset.

[0098] In this invention, we used the THuman2.0 human model dataset introduced in the paper "Function4D: Real-time Human Volumetric Capture from Very Sparse Consumer RGBD Sensors" to generate the training dataset, which contains 500 high-quality 3D human scanning models. Along the y-axis, we used CUDA acceleration to render realistic human RGBD images at 6-degree intervals. We split this dataset into training and testing sets in a 4:1 ratio. For the input raw depth image, we referenced the method in the paper "Kinect v2 for mobile robotnavigation: Evaluation and modeling" to render a realistic depth map D. gt Add sensor noise, including noise related to the pixel depth value z: 1.5z 2 -1.5z +1.375, Gaussian noise (average 1.5cm), holes (average width 3 pixels) to cover as many possible noise conditions as possible.

[0099] We implemented the model involved in this invention using PyTorch and CUDA. During the actual model training process, we used two NVIDIA RTX 3090 graphics cards to train the depth image denoising model, SRONet, and ray upsampling network respectively, and used the ADAM optimizer to optimize the model parameters. The parameters β1 and β2 in ADAM were set to 0.5 and 0.99, respectively. The batch size of the depth image denoising model was set to 8, and the learning rate was set to 1e. -3 The training run consisted of 10 epochs. The SRONet batch size was set to 4, and the learning rate was set to 1e. -4 The learning rate is halved every 5 epochs, for a total of 20 epochs. The ray upsampling network has the same learning rate as SRONet, a batch size of 2, and is trained for 10 epochs. For all loss function balancing terms, we adjust the λ of the depth image denoising model. N , λ P Set them to 0.5 and 0.01 respectively; adjust the μ values ​​in SRONet. o μ c , λ D′ Set them to 0.5, 1.0, and 1.0 respectively; adjust the μ1, μ2, and μ values ​​of the ray upsampling network. vgg Set them to 0.4, 0.6, and 0.01 respectively; calculate the visibility weight σ. vSet to 200. In addition, set the hyperparameters α and γ in the two-layer tree structure to 40 and 0.01 respectively. Set the number of sampling points M during training and inference to 48. Set the final resolution of the rendered image to 1K (1024*1024).

[0100] To verify the effectiveness of the method of this invention, we conducted rendering quality tests on the Thuman2.0 test set and real-world data. The verification experiments were compared with six mainstream deep learning-based methods on the aforementioned test set: PixelNeRF (Neural Radiance Fields from One or Few Images), IBRNet (IBRNet: Learning Multi-View Image-Based Rendering), MPSNeRF (Generalizable 3D Human Rendering from Multiview Images), NHP (Learning generalizable radiance fields for human performance rendering), NPBG++ (Accelerating Neural Point-Based Graphics), and PIFu(RGBD) (Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization). Evaluations were also conducted on four image similarity metrics: PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity), LPIPS (Learned Perceptual Patch Similarity), and MAE (Mean Absolute Error). The comparison results are shown in Table 1 (Test Set) and Table 2 (Real-world Data).

[0101] Table 1

[0102]

[0103] Table 2

[0104]

[0105] Higher PSNR and SSIM values ​​indicate better results, while lower LPIPS and MAE values ​​indicate even better results. The best result in each row is represented by a bold number, and the second best result by an underscore. The results show that the proposed method (Ours) is superior to other methods overall, and significantly outperforms them in average rendering time (Avg Time). This comparative experiment fully demonstrates the superiority of the proposed method.

[0106] In addition, we conducted some decomposition experiments on the test set (Thuman2.0 Dataset) and real-world data (Our Real Captured Dataset) to analyze the effectiveness of SRONet's introduced loss functions (geometric loss function, depth loss function), depth denoising process, collaborative geometric and light field representation, feature fusion operation, and ray upsampling operation. We removed or modified the above content and retrained the entire network model. The experimental results are shown in Table 3.

[0107] Table 3

[0108]

[0109] Wherein, w / o GT Depth indicates that the depth consistency loss function was removed during SRONet training; w / o GT Occ indicates that geometric supervision was removed during SRONet training; w / o Denoised Depth indicates that depth denoising was removed; Soft Occ.→Density indicates that the collaborative occupancy field and light field representation were replaced with the original light field representation of NeRF, that is, OCCNet predicts the volume density of sampling points instead of the occupancy value, and the rendering method used in this invention is replaced with NeRF's volume rendering method. OccMLP→DbMLP indicates that OCCNet simultaneously predicts the volume density and occupancy value of sampling points, retains the geometric supervision loss function, and uses NeRF's volume rendering method for rendering. Hydra Att.→Self Att. indicates that the hydra attention module of the feature fusion function H is replaced with the self attention module of the transformer model. w / o Upsampling indicates that the upsampling module based on neural fusion is removed. The above decomposition experiments show that the loss function or module introduced in this invention improves the overall rendering quality to a certain extent.

[0110] This invention also provides a human body reconstruction and rendering apparatus based on coordinated light field and occupancy field, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement the aforementioned human body reconstruction and rendering method based on coordinated light field and occupancy field.

[0111] This invention also provides a computer-readable storage medium storing a program that, when executed by a processor, implements the aforementioned method for human body reconstruction and rendering based on cooperative light field and occupancy field.

[0112] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit the scope of one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the protection scope of one or more embodiments of this specification.

Claims

1. A method for human body reconstruction and rendering based on coordinated light field and occupancy field, characterized in that, Includes the following steps: S1: Construct a deep image denoising model F based on a convolutional neural network structure d The model's input includes a human RGB image I and an original depth image D, and its output is a denoised depth image D. rf The model is trained using depth, normal, and 3D consistency loss functions, and then used for deep denoising tasks. S2: Based on the denoised multi-view depth image and the camera parameters of each viewpoint, a global human point cloud is fused. A two-layer tree data structure is constructed from the point cloud, and the parent and child nodes of the tree are stored as a global list. During the inference process, voxel post-processing operations are introduced to denoise the tree structure, retaining only voxels near the human body surface. S3: Construct a collaborative network model SRONet for human body mesh reconstruction and new perspective image rendering. This network model comprises two sub-network models: the occupancy field network OCCNet and the light field network ColorNet. For a given input: a multi-view RGBD image... And the sampling point x in three-dimensional space, the viewing direction d, where N is the number of views, and the sub-network models respectively implement the following tasks: Given a multi-view depth image as input And for the sampling point x, OCCNet predicts the voxel occupancy o for point x using the pixel-aligned implicit function PIFu. x ∈[0,1], representing the probability that the point is located inside the human body mesh model; Given an input multi-view RGB image {I i } i=1,…N ColorNet predicts the RGB three-channel color vector c for a point x along the viewing direction d based on PIFu. x ; Occupancy based on point x prediction o x The equivalent cube search algorithm is used to calculate the human body mesh model from the estimated occupancy field for 3D human body reconstruction. Using a multi-view human body dataset, SRONet is trained by combining the geometry-color collaborative loss function and the depth error loss function. After the model is trained, it is used to predict the occupancy field and the light field. S4: Calculate the light rays for each pixel in the new perspective image based on the corresponding camera parameters, and perform the light ray voxel intersection process according to the voxel traversal algorithm to determine the voxels that intersect with the light rays in the tree data structure; Record the depth values ​​of the near and far intersection points of all intersecting voxels, and design sampling weights along the light ray within the voxel based on the voxel size. Calculate the occupancy value and color of the sampling points according to S3. Using the volume rendering formula, calculate the color blending weight based on the occupancy value to blend the color of each sampling point on the light ray and calculate the final color of the light ray. S5: Upsample each ray to improve the resolution and quality of the rendered image; A feature fusion network is constructed. The input of the network is information shared by each sub-ray, including the original color, depth, ray features, and input RGB images from two adjacent viewpoints. The output is the final color of each sub-ray. The feature fusion network is trained by ray-by-ray color error, structural similarity, and feature loss function. After training, it is used for ray upsampling operations to obtain a target resolution rendering image.

2. The method for human body reconstruction and rendering based on coordinated light field and occupancy field according to claim 1, characterized in that, The specific design of the deep denoising process is as follows: (1) The BackgroundMatting-v2 algorithm is used to extract the human figure region from the original input image as the human figure region mask image Ψ, and the RGBD images I and D of the human figure region are obtained at the same time; the input RGBD images are normalized to the interval [-1,1], and the depth maximum and minimum values ​​are recorded at the same time; the normalized RGB and depth images are merged as the depth image denoising model F. d Input; (2) Constructing model F based on UNet structure d Two independent feature extraction networks, HRNetV2-W18-Small-v2, are used to encode RGB and depth images respectively. A dilated spatial convolutional pooling pyramid module and a residual attention module are used to fuse RGB and depth features. The fused features are upsampled and fed back to the feature extraction network. (3) Model F d Output the image in the range [-1, 1]. Use the depth extrema in (1) to inverse normalize the output image back to the original value range. Use Ψ to post-process the image, retaining only the depth value of the human figure region, denoted as D. rf ; (4) Used for model F d The training loss function is as follows: Deep consistency loss L D : Used to penalize the true depth image at viewpoint i With the predicted depth image The per-pixel deviation is defined as: Normal consistency loss L N Used to penalize depth images and The deviation between the calculated normals is defined as: in, Represents the predicted depth image and its calculated normal diagram The depth value and normal vector at the i-th viewpoint pixel p. This represents the corresponding true depth value and normal vector; Describes the L1 loss function. The L2 loss function is represented by <·, ·>, which represent the vector dot product operation. 3D consistency loss L P Used to further constrain passage The consistency between the fused point cloud and the real point cloud, in order to reduce depth fusion noise in the construction of a two-layer tree structure, is defined as: in, Indicates from The fused point cloud, F is the truncated symbolic range field fusion algorithm, K i ,RT i These are the camera's intrinsic and extrinsic parameters, P. gt The point cloud is sampled from a real 3D human body mesh model. This represents the chamfer distance loss function; The loss function used for deep denoising is expressed as: L = L D +λ N L N +λ P L P , where λ N ,λ P The loss function is optimized using the ADAM algorithm, with weights assigned to it.

3. The method for human body reconstruction and rendering based on coordinated light field and occupancy field according to claim 1, characterized in that, The specific design of the two-level tree data structure is as follows: (1) Use the truncated symbolic distance field fusion algorithm to fuse the denoised depth images from all viewpoints. Simultaneously, the global human point cloud P is obtained. rf With Fusion Cube V tsdf ; For V tsdf After binarization, let it be denoted as V. occ Then the occupancy value V at spatial location x occ (x) is: Among them, s v V represents the size of the voxel, α is the truncation sign distance threshold, and V tsdf (x) represents V at spatial location x. tsdf The truncation distance; (2) Voxel post-processing: Using OCCNet in S3, a cube with a set resolution is quickly constructed based on the real-time human reconstruction algorithm, denoted as . After binarization, V is constructed using (1). occ Fusion to eliminate V occ The floating noise voxels are used to obtain the denoised cube. Then the fusion occupancy value at spatial location x for: in, It is a binarization function based on the threshold β, where γ is the threshold used to filter floating voxels, and | represents the OR operation; (3) The denoised cube middle The voxels are denoted as effective voxels, and parent voxels are fused together according to a preset ratio. All effective voxels are then stored in a global list L. v Each node corresponds to a voxel, and each node records the index, spatial location, and size information of its parent or child node; construct an indexed cube V. idx This is used to store the index value of each valid parent node in the global list. For invalid parent nodes, their index value is set to -1.

4. The method for human body reconstruction and rendering based on coordinated light field and occupancy field according to claim 1, characterized in that, The specific design of SRONet is as follows: (1) OCCNet: Based on depth information, it uses the feature encoder HRNetV2-W18 to encode depth images; for a sampling point x, OCCNet predicts the occupancy value o corresponding to that point by pooling the pixel-aligned depth features from each viewpoint. x Define the occupancy field as a function. Among them, W i W represents the depth feature map of viewpoint i after encoding. i (x) represents the projection of point x onto viewpoint i from W. i The deep features obtained from c i (x) represents the depth value of point x projected onto the camera coordinate system at viewpoint i and the distance to the truncation sign; the implicit function f1, represented by a fully connected network, is used to obtain the geometric features of each viewpoint, and the fused global features are obtained through the average pooling operation Avh. The global features are then fed into the second implicit function f2 to calculate the occupancy value o. x ; (2) ColorNet: Light field based on color and geometric features. It encodes RGB images using the same feature encoder as OCCNet to obtain color features; for sampling point x, ColorNet uses additional inputs: view direction d and geometric features. To aggregate color features from each viewpoint to predict viewpoint-related color vectors Among them, geometric features are represented as Where f3 is the implicit function used for encoding; the light field is defined as a function Among them, M i M represents the color feature map of viewpoint i after encoding. i (x) and rgb i This indicates that the projection of point x onto viewpoint i is from M. i The color features and pixel colors obtained from the data; f4 and f5 represent the implicit functions used for further feature processing; The feature fusion function is implemented using a transformer; the view direction d in the camera coordinate system. i =R i d, where R i This represents the rotation matrix in the extrinsic parameters of the camera at viewpoint i; SRONet predicts the occupancy value of the sample point x and the view-dependent color.

5. The method for human body reconstruction and rendering based on coordinated light field and occupancy field according to claim 1, characterized in that, The specific design of ray-voxel intersection and point sampling in S4 is as follows: (1) For a emitted ray l, use a voxel traversal algorithm to detect the parent nodes that intersect along ray l, and record the depth value of the intersection point of the intersecting parent voxels from the viewpoint of l. The depth values ​​of the near and far intersection points are denoted as D. far With D near For each intersecting parent voxel, continue using the voxel traversal algorithm to detect intersecting child voxels, and record the near and far depth values, denoted as D′. far With D′ near The sum of the depth values ​​for all records is as follows: and (2) Assign the number of sampling points between the near and far intersections of each voxel, and calculate the sampling weight w of the i-th voxel. i as follows: Where, d far (i) and d near (i) respectively represent and The near and far depths of voxel i, N v s represents the number of all intersecting voxels. i The scale of voxel i is represented; specifically, the number of sampling points m within each voxel i. i for: Where M is the total number of sampling points, This indicates rounding down; during sampling, sampling points are assigned to each voxel i along the direction of light emission to ensure that the voxel closest to the camera on the light ray is always sampled; if ∑ i m i If M < M, then the remaining M-∑ will be redistributed according to the sampling order. i m i The depth value d of the j-th sampling point within the i-th voxel. i (j) The calculation formula is as follows: Where j starts from 0; sampling point x passes through p cam +d i (j)·d is calculated to obtain p cam d represents the camera position, and d represents the viewing direction.

6. The method for human body reconstruction and rendering based on coordinated light field and occupancy field according to claim 1, characterized in that, The specific design of the volume rendering described in S4 and the loss function of SRONet used for training the model described in S3 is as follows: (1) To calculate the final color of ray l, normalized surface and volume rendering techniques are used, based on the occupancy value o of each sampling point on l. x Calculate the color blending weights and blend the point colors c. x Calculate the color of light as follows: Wherein, the color fusion weight ω of the i-th sampling point x (i)=o x (i)∏ j<i (1-o x (j)); At the same time, by weighting the depth value d of the i-th sampling point i To calculate the depth value at the intersection of ray l and the human body surface. (2) Color based on light estimation With depth value The following loss function is designed to train and optimize the reconstruction and rendering parts of SRONet: Geometric-color collaborative loss L syn : Based on PIFu-based supervision, OCCNet is trained by sampling points y in space, and the estimated occupancy value o is penalized. y Compared with actual occupancy value Errors between them; and penalties for each ray's true color C * (l) and estimated color The error is due to the two loss functions working together: Where S and R represent the set of sampling points and the set of rays, respectively. Represents the cross-entropy loss function. Let μ represent the L1 loss function. o ,μ c These are the weights of the occupancy value loss term and the color loss term, respectively. Depth error loss L D′ Used to penalize estimated ray depth values Compared with the true depth value D * (l) error, to further improve reconstruction and rendering details: in, Represents the L2 loss function; The loss function used for SRONet is expressed as: L syn +λ D′ L D′ , where λ D′ It is a balancing term, and the loss function is optimized using the ADAM algorithm.

7. The method for human body reconstruction and rendering based on coordinated light field and occupancy field according to claim 1, characterized in that, The specific design of ray upsampling in S5 is as follows: (1) Ray upsampling: For the ray l passing through pixel position (x,y), its color is upsampled. Depth value Light fusion features The values ​​are distributed to four sub-pixels, corresponding to the positions: (x, y), (x+0.5, y), (x, y+0.5), (x+0.5, y+0.5); where ft color The output features of the feature fusion function H are used to generate a coarse upsampling result. (2) Feature fusion operation: The original resolution RGB images of two adjacent viewpoints n0 and n1 of the target viewpoint are used to enhance the result in (1); specifically, UNet is used to encode two adjacent RGB images, each sub-ray corresponds to a sub-pixel, and the surface position of the sub-ray is calculated by the depth value of the sub-ray. Projecting point p onto neighboring images and features to obtain color. and feature and Computational visibility and The formula for calculating visibility is: z i This represents the depth value of p projected at viewpoint i. Represents the denoised depth image from viewpoint i. The depth value obtained from σ v It is a preset visibility weight coefficient; through a feature fusion network Calculate fusion weights The three colors of the merged sub-ray: The final color of the sub-ray is obtained; defined as follows: Where f6 represents the implicit function for processing features, The original resolution RGB images of adjacent viewpoints n0 and n1; (3) Training the feature fusion network Loss function: Ray-by-ray color error and structural similarity loss L B Used to penalize sampled color blocks With real color blocks The color error between them is defined as follows: Where R represents the set of light rays, It is of size S patch ·S patch color blocks, It is the final estimated color of the ray r located at (i,j); Let μ1 denote the L1 loss function, and SSIM denote the structural similarity function; μ1 and μ2 are the balance terms of the loss function. Feature loss function L ft Used to punish color blocks and The feature error between the two is used to further enhance the quality of the rendered image; the loss function is calculated using a pre-trained VGG-16 network, and the loss function is defined as follows: in, The L1 loss function represents the relationship between VGG features; For feature fusion networks The loss function is expressed as: L B +μ vgg ·L ft , where μ vgg It is a balancing term, trained independently using the ADAM algorithm. The parameters.

8. The method for human body reconstruction and rendering based on coordinated light field and occupancy field according to claim 1, characterized in that, The specific design for the parallel accelerated rendering process is as follows: (1) Use two GPUs to accelerate the rendering process; each GPU processes half of the data and synchronizes two batches of data on the CPU via memory. The rendering process is accelerated by the pipeline; specifically, the pipeline is divided into 3 parts, and each GPU is accelerated by 3 independent data streams:

1. I / O process and depth denoising process; 2. Construct a two-layer tree data structure, multi-view RGB image encoding and depth image encoding in SRONet, and image encoding in ray upsampling; 3. Ray sampling: The point occupancy value and color are calculated using SRONet, and ray sampling is performed on the light rays based on the feature fusion network; finally, all the calculated rays are converted into images for display. (2) Use TensorRT technology for half-precision quantization and acceleration of the depth image denoising model, multi-view RGB image encoding and depth image encoding in SRONet, and image encoding in ray upsampling; accelerate all implicit functions and feature fusion functions implemented by transformer through GPU shared memory.

9. A human body reconstruction and rendering apparatus based on coordinated light field and occupancy field, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that, When the processor executes the executable code, it is used to implement the human body reconstruction and rendering method based on the cooperative light field and occupancy field as described in any one of claims 1-8.

10. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the human body reconstruction and rendering method based on the cooperative light field and occupancy field as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Structured observation and sparse representation collaborative optimization method for optical field camera

    CN108492239A

  • Three-Dimensional Object Localization Using a Lookup Table

    US20180286073A1