A multi-view based human body surface reconstruction method and system

Through the multi-view-based human surface reconstruction method, the multi-view depth estimation neural network and the maching-cubes algorithm are used to solve the problems of high cost, time-consuming and limited accuracy in the existing technology, and achieve fast and accurate three-dimensional reconstruction of the human body, especially fidelity in clothing details.

CN114663599BActive Publication Date: 2025-05-27NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210398618.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2025-05-27
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction technology of human body has problems such as high cost, long algorithms and limited accuracy, making it difficult to achieve fast and accurate human body reconstruction.

Method used

Using a multi-view-based human surface reconstruction method, by obtaining multi-view RGB images, using the trained human UV estimation network and multi-view depth estimation neural network, the optical flow and initial depth map are calculated, the human body point cloud is extracted, and the marching-cubes algorithm is fused into a three-dimensional surface.

Benefits of technology

It realizes fast and accurate three-dimensional reconstruction of the human body, which can finely retain the details of the deformation of clothes without the need for a high-precision depth camera, reducing equipment costs and algorithm execution time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114663599B_ABST
    Figure CN114663599B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-view based human body surface reconstruction method and system. The method includes obtaining indoor RGB pictures from multiple perspectives, and using a trained human body UV estimation network to obtain the UV coordinate map for each perspective; calculating the optical flow between two adjacent perspectives according to the UV coordinate map of each perspective; and obtaining an initial depth map according to the optical flow and the camera extrinsic parameter matrix; determining the depth map for each perspective by using a trained multi-view depth estimation neural network according to the indoor RGB pictures from multiple perspectives and the initial depth map; extracting the human body point cloud according to the depth map of each perspective, and fusing the extracted human body point cloud into a three-dimensional human body surface based on the marching-cubes algorithm; and performing rendering visualization at any perspective according to the three-dimensional human body surface. The present invention realizes an end-to-end three-dimensional human body surface reconstruction system, enabling users to quickly and accurately obtain the human body reconstruction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a multi-view based human body surface reconstruction method and system. Background Art

[0002] Indoor human body surface reconstruction technology mainly involves recovering the three-dimensional shape of the human body from information such as images and radars, and then applying it to virtual reality (VR), digital human body, game graphics and image applications, etc.

[0003] Dynamic three-dimensional reconstruction of the human body is also a very important issue in the fields of computer vision, computer graphics, and virtual reality. High-precision and high-efficiency reconstruction is still one of the relatively challenging tasks in the current academic and industrial circles.

[0004] Currently, the methods for indoor three-dimensional reconstruction of the human body mainly include those based on depth cameras, multi-view stereo matching, and deep learning-based strategies. Although the related algorithms based on depth cameras and multi-view stereo matching under multiple views can currently obtain relatively high-quality human body reconstruction results under limited conditions, there are still many obstacles to wide applications in terms of equipment cost and algorithm execution efficiency. For example, for the reconstruction algorithm based on the Kinect depth camera, in order to achieve the effect of full-body reconstruction, multiple cameras are often required to collect depth information, so the cost is relatively high. The method based on multi-view stereo matching does not require depth information, but requires more high-definition cameras to capture partial images of the human body. At the same time, the dynamic programming and optimization processes involved in the stereo matching algorithm are time-consuming, and it is difficult to achieve fast human body reconstruction.

[0005] In recent years, with the gradual deepening of the research on deep learning in the field of vision, researchers have begun to explore deep learning-based human body reconstruction algorithms, such as three-dimensional parameter estimation of the human body based on CNN, single-view human body reconstruction based on implicit representation of neural networks, etc. The former generally adopts the idea of a parameterized human body model, and infers the motion parameters and body shape parameters of the human body from a single image by training a neural network. The latter generally extracts image features through a neural network and predicts the signed distance function (SDF) of the human body. However, these methods either cannot recover high-frequency information such as the wrinkles on the surface of the human body clothes, or cannot achieve high-precision human body reconstruction due to the ambiguity of a single view.

[0006] In summary, the existing methods mainly have the following problems: high cost, long algorithm running time, and limited accuracy. Therefore, there is an urgent need for a new method or system to solve the above problems. Summary of the Invention

[0007] The objective of the present invention is to provide a multi-view-based human body surface reconstruction method and system, to implement an end-to-end three-dimensional human body surface reconstruction system, enabling users to quickly and accurately obtain the human body reconstruction results.

[0008] To achieve the above objective, the present invention provides the following solutions:

[0009] A multi-view-based human body surface reconstruction method, comprising:

[0010] Obtain indoor RGB pictures from multiple perspectives, and use a trained human body UV estimation network to obtain the UV coordinate map for each perspective;

[0011] Calculate the optical flow between two adjacent perspectives according to the UV coordinate map of each perspective; and obtain the initial depth map according to the optical flow and the camera extrinsic parameter matrix;

[0012] Determine the depth map for each perspective according to the indoor RGB pictures from multiple perspectives and the initial depth map by using a trained multi-view depth estimation neural network;

[0013] Extract the human body point cloud according to the depth map of each perspective, and fuse the extracted human body point cloud into a three-dimensional human body surface based on the marching-cubes algorithm;

[0014] Perform rendering visualization at any perspective according to the three-dimensional human body surface.

[0015] Optionally, before the step of obtaining indoor RGB pictures from multiple perspectives and using a trained human body UV estimation network to obtain the UV coordinate map for each perspective, it further includes:

[0016] Set up an indoor environment with multiple cameras for shooting, and use a calibration board for calibration to obtain the camera extrinsic parameter matrix.

[0017] Optionally, the step of obtaining indoor RGB pictures from multiple perspectives and using a trained human body UV estimation network to obtain the UV coordinate map for each perspective specifically includes:

[0018] Use a trained segmentation network to obtain a two-dimensional mask for the human body part;

[0019] Crop the indoor RGB pictures of the perspective according to the position of the two-dimensional mask for the human body part.

[0020] Optionally, the trained multi-view depth estimation neural network includes: a Feature Pyramid Network (FPN) for encoding the information of a single perspective picture, a Feature Cross-Correlation Module for fusing perspective features at different depths of the viewing rays in the source perspective, a depth search mechanism at multiple scales, an attention module, a feature fusion module, and a depth probability estimation module.

[0021] A multi-view based human body surface reconstruction system, comprising:

[0022] A UV coordinate map determination unit, configured to obtain indoor RGB pictures of multiple perspectives, and use a trained human body UV estimation network to obtain the UV coordinate map of each perspective;

[0023] An initial depth map determination unit, configured to calculate the optical flow between two adjacent perspectives according to the UV coordinate map of each perspective; and obtain an initial depth map according to the optical flow and the camera extrinsic parameter matrix;

[0024] A depth map determination unit, configured to use a trained multi-view depth estimation neural network to determine the depth map of each perspective according to the indoor RGB pictures of multiple perspectives and the initial depth map;

[0025] A human body three-dimensional surface determination unit, configured to extract the human body point cloud according to the depth map of each perspective, and fuse the extracted human body point cloud into a human body three-dimensional surface based on the marching-cubes algorithm;

[0026] A rendering visualization unit, configured to perform rendering visualization at any perspective according to the human body three-dimensional surface.

[0027] Optionally, it further comprises:

[0028] A camera extrinsic parameter matrix acquisition unit, configured to set up an indoor environment with multiple cameras for shooting, and use a calibration board for calibration to obtain the camera extrinsic parameter matrix.

[0029] Optionally, the UV coordinate map determination unit specifically comprises:

[0030] A human body part two-dimensional mask determination subunit, configured to obtain a human body part two-dimensional mask using a trained segmentation network;

[0031] A cropping subunit, configured to crop the indoor RGB picture of the perspective according to the position of the human body part two-dimensional mask.

[0032] Optionally, the trained multi-view depth estimation neural network includes: a feature pyramid network FPN for encoding the information of a single perspective picture, a feature cross-correlation module for fusing perspective features at different depths of the source perspective observation ray, a depth search mechanism at multiple scales, an attention module, a feature fusion module, and a depth probability estimation module.

[0033] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0034] A method and system for human body surface reconstruction based on multi-views provided by the present invention calculates the optical flow between two adjacent views according to the UV coordinate maps of each view; and obtains an initial depth map based on the optical flow and the external camera parameter matrix; determines the depth map of each view by using the trained multi-view depth estimation neural network according to the indoor RGB pictures of multi-views and the initial depth map; extracts the human body point cloud according to the depth map of each view, and fuses the extracted human body point clouds into a three-dimensional human body surface based on the marching-cubes algorithm; by using the trained multi-view depth estimation neural network to determine the depth map of each view, that is, a more refined depth map, without the need for a high-precision depth camera; using RGB pictures with normal resolution, the three-dimensional reconstruction of the human body can be realized while finely retaining the deformation details of clothes. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0036] Figure 1 It is a schematic flow chart of a method for human body surface reconstruction based on multi-views provided by the present invention;

[0037] Figure 2 It is a flow chart of a multi-view human body depth estimation network provided by the present invention;

[0038] Figure 3 It is a schematic structural diagram of the trained multi-view depth estimation neural network and the schematic structural diagram of the attention module provided by the present invention;

[0039] Figure 4 It is a schematic structural diagram of a system for human body surface reconstruction based on multi-views provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0041] The purpose of the present invention is to provide a method and system for human body surface reconstruction based on multi-views, to realize an end-to-end three-dimensional human body surface reconstruction system, so that users can quickly and accurately obtain the human body reconstruction results.

[0042] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] Figure 1 It is a schematic flowchart of a multi-view based human body surface reconstruction method provided by the present invention. Figure 2 It is a flowchart of a multi-view human body depth estimation network provided by the present invention. As Figure 1 and Figure 2 shown, a multi-view based human body surface reconstruction method provided by the present invention includes:

[0044] S101, Obtain indoor RGB pictures of multiple views, and use the trained human body UV estimation network to obtain the UV coordinate map of each view;

[0045] Use cameras to capture indoor RGB pictures of the human body from multiple views I1, I2,... In. After preprocessing, use the trained human body UV estimation network to obtain the UV coordinate map of each view; Calculate the pixel position of each pixel corresponding to the pixel at the nearest UV coordinate in another view according to the human body UV coordinates and the label of the human body part where each pixel is located. Here, the kd-tree algorithm can be used to quickly calculate the approximate matching position. A certain matching error is allowed in this step; After completion, the optical flow of adjacent views is obtained. According to the obtained optical flow map and the internal and external parameters of the camera, calculate the rough depth result of each view using solid geometry. And downsample it to a scale of 1 / 8 of the original resolution.

[0046] Before S101, it also includes:

[0047] Build an indoor environment with multiple cameras for shooting, and use a calibration board for calibration to obtain the external parameter matrix of the camera.

[0048] S101 specifically includes:

[0049] Use the trained segmentation network to obtain the two-dimensional mask mask of the human body part;

[0050] Crop the indoor RGB pictures of the view according to the position of the two-dimensional mask mask of the human body part.

[0051] S102, Calculate the optical flow between two adjacent views according to the UV coordinate map of each view; And obtain the initial depth map according to the optical flow and the external parameter matrix of the camera;

[0052] According to the UV coordinate maps of the current view and other views, use the kd-tree search algorithm to find the pixel coordinates of each pixel in the current view corresponding to other views, so as to calculate the optical flow map.

[0053] Calculate the initial depth of the current view according to the internal and external parameters of the corresponding two views and the optical flow between them.

[0054] S103. Based on the indoor RGB images from multiple perspectives and the initial depth map, use the trained multi-view depth estimation neural network to determine the depth map for each perspective; the trained multi-view depth estimation neural network includes: a Feature Pyramid Network (FPN) for encoding the information of a single perspective image, a Feature Cross-Correlation Module for fusing perspective features at different depths along the viewing rays of the source perspective, a depth search mechanism at multiple scales, an attention module, a feature fusion module, and a depth probability estimation module.

[0055] The trained multi-view depth estimation neural network takes each perspective as the source perspective one by one, and regards all the other perspectives as the target perspectives. The result of each network inference is the human body depth map under the current source perspective. After multiple cycles, the depths of all perspectives are obtained. Integrate the depths of all perspectives for human body point cloud fusion and surface reconstruction.

[0056] The network structure consists of depth estimation modules at four resolutions from low to high. The input of each module is the RGB image at the current resolution and the depth estimated by the previous layer module as the initial depth of the current module, and the output is the refined depth of each pixel at the current resolution. The coarse-to-fine idea is adopted. Among them, the initial depth of the lowest resolution is provided by the initial depth obtained from UV.

[0057] When the network performs depth optimization, it uses a strategy of searching in the depth direction. That is, whenever the depth estimation of the previous low resolution is completed, the initial depth at the current resolution is obtained by doubling the previous one. Within an equidistant range before and after the three-dimensional position represented by the current initial depth, a number of points are uniformly sampled. The current module estimates the probability that each point is at the position of the real surface point under the action of the 2D image features pointed to by these different depth points.

[0058] As Figure 3 shown, the Feature Pyramid Network (FPN) consists of four layers of convolution and four layers of deconvolution, and cross-layer connections are made between the features obtained by convolution of the same size and the features obtained by deconvolution. The output is multi-channel features of the same size as the original image.

[0059] The Feature Cross-Correlation Module obtains the cost volume in the depth search direction by performing cross-correlation between the features at the current pixel position of the source perspective projected to a certain position of the target perspective at different depths.

[0060] The network predicts depth maps at 4 different resolutions from low to high. And the prediction module at the higher resolution conducts a uniform search within a certain range near the depth predicted at the previous low resolution until the depth estimation at the highest resolution is completed.

[0061] The attention module takes the features of each target view related to the source view as input, predicts the correlation weight values of each view at different pixel positions, and is used to characterize the attention of different views to the current position.

[0062] The attention module first encodes the features of each view image, and based on the initial depth obtained from the previous layer, calculates the matching positions of other views for each pixel of the source view, deforms the features of the target view to the source view, and after splicing the two features, inputs them into the decoder of the attention module to output an attention heat map with 1 output channel. The heat maps of multiple views are used to perform a softmax operation on each position to obtain the final view weights.

[0063] After being calculated by the attention module, at the current initial depth, cross-correlations are performed on the feature vectors at the positions of 2*r + 1 points searched at a certain interval in the vicinity with the projected positions of other views.

[0064] Assume that the current pixel position is (u1, v1), the depth is d, and the internal and external camera parameters are K1, P1. After calculation, the original three-dimensional position is p = ∏ -1 (u1, v1, d; K1, P1); the position projected onto the second view is (u2, v2) = ∏(p; K2, P2); thus, the feature at the current position can be obtained as f2 = BIL(FPN(I2); u2, v2); the cross-correlation calculation is Corr(f1, f2) = <f1, f2>.

[0065] Assume that the current initial depth is d, the depth interval is dur, and the search radius is r. Then, the search cost volume V at this level can be obtained, with a shape of H*W*(2*r + 1).

[0066] Subsequently, each target view is searched and correlated with the source view to obtain cost volumes V1, V2, … Vn.

[0067] After the images of all target views and the source view are used to extract features by another independent feature pyramid network, the attention weights of each view, W1, W2 … Wn, are obtained through the Decoder and softmax operations. Extract and fuse the relevant cost volumes of all target views:

[0068] V_fuse and the features of the source view are sent to the decoder together to output heat maps at different depth positions. A softmax calculation is performed in the depth dimension to obtain the probability values at each depth, denoted as P1, P2, … Pn. The further refined depth value predicted at this level is d_pred =

[0069]

[0070] The network consists of four layers of depth estimation modules with different resolutions, namely M0, M1, M2, and M3. The depth estimated by the i-th layer is \(d_i = M_i(I1, I2, up(d_{i - 1}))\), which means that each layer uses the upsampled result of the depth value estimated by the previous layer. Among them, the initial depth of the first module M0 is calculated from UV. Therefore, the final depth estimate is \(d = M3(M2(M1(M0(d_{init}, I1, I2, … In))))\).

[0071] The feature fusion module fuses multi-view information based on the perspective weight values and the features of each target perspective.

[0072] The depth probability estimation module takes the fused multi-view features and the source perspective features as inputs, and outputs the probability values for each pixel point in the source perspective at each position within the searched depth range. The final estimated depth value is calculated based on the probability distribution.

[0073] In the training stage, first, a dataset acquisition environment needs to be set up. We completed the acquisition of the dataset based on three Kinect cameras and twelve high-definition single-lens reflex cameras. After synchronization, the Kinect cameras simultaneously capture videos of human actions, and at the same time, the infrared depth cameras capture the depth information of the human body. The single-lens reflex cameras only need to take 12 pictures in one frame, and then use the structure from motion method to extract a human body reconstruction model with higher accuracy.

[0074] The loss of training the network is to calculate the absolute error loss of depth estimation at multiple resolutions, and calculate the total loss through weights. That is where 4 represents the number of different resolutions, D i represents the depth map predicted at a certain layer resolution, D i ’ represents the true depth value at this resolution. \(\lambda_i\) represents the loss weight of this layer resolution. L all represents the overall loss target.

[0075] Use the above loss to train the entire depth estimation network.

[0076] S104, extract the human body point cloud from the depth map of each perspective, and fuse the extracted human body point clouds into a human body three-dimensional surface based on the marching-cubes algorithm;

[0077] S105, perform rendering visualization at any perspective based on the human body three-dimensional surface.

[0078] Figure 4 This is a schematic structural diagram of a multi-view based human body surface reconstruction system provided by the present invention. As Figure 4 shown, a multi-view based human body surface reconstruction system provided by the present invention includes:

[0079] A UV coordinate map determination unit 401, configured to obtain indoor RGB pictures of multiple perspectives and use a trained human body UV estimation network to obtain a UV coordinate map for each perspective;

[0080] An initial depth map determination unit 402, configured to calculate the optical flow between two adjacent perspectives according to the UV coordinate map of each perspective; and obtain an initial depth map according to the optical flow and the camera extrinsic parameter matrix;

[0081] A depth map determination unit 403, configured to determine the depth map of each perspective by using a trained multi-view depth estimation neural network according to the indoor RGB pictures of multiple perspectives and the initial depth map;

[0082] A human body three-dimensional surface determination unit 404, configured to extract a human body point cloud according to the depth map of each perspective and fuse the extracted human body point clouds into a human body three-dimensional surface based on the marching-cubes algorithm;

[0083] A rendering visualization unit 405, configured to perform rendering visualization at any perspective according to the human body three-dimensional surface.

[0084] A multi-view based human body surface reconstruction system provided by the present invention further includes:

[0085] A camera extrinsic parameter matrix acquisition unit, configured to build an indoor environment photographed by multiple cameras and perform calibration using a calibration board to obtain a camera extrinsic parameter matrix.

[0086] The UV coordinate map determination unit 401 specifically includes:

[0087] A human body part two-dimensional mask determination subunit, configured to obtain a human body part two-dimensional mask using a trained segmentation network;

[0088] A cropping subunit, configured to crop the indoor RGB picture of the perspective according to the position of the human body part two-dimensional mask.

[0089] The trained multi-view depth estimation neural network includes: a feature pyramid network FPN for encoding the information of a single perspective picture, a feature cross-correlation module for fusing perspective features at different depths of the source perspective observation ray, a depth search mechanism at multiple scales, an attention module, a feature fusion module, and a depth probability estimation module.

[0090] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0091] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A multi-view based human body surface reconstruction method, characterized in that, it includes: Obtain indoor RGB pictures from multiple perspectives, and use the trained human body UV estimation network to obtain the UV coordinate map of each perspective; Calculate the optical flow between two adjacent perspectives according to the UV coordinate map of each perspective; And obtain the initial depth map according to the optical flow and the camera extrinsic parameter matrix; According to the indoor RGB pictures from multiple perspectives and the initial depth map, use the trained multi-view depth estimation neural network to determine the depth map of each perspective; The trained multi-view depth estimation neural network includes: a feature pyramid network FPN for encoding the information of a single perspective picture, a feature cross-correlation module for fusing perspective features at different depths of the source perspective observation ray, a depth search mechanism at multiple scales, an attention module, a feature fusion module, and a depth probability estimation module; Extract the human body point cloud according to the depth map of each perspective, and fuse the extracted human body point cloud into a three-dimensional human body surface based on the marching-cubes algorithm; Perform rendering visualization at any perspective according to the three-dimensional human body surface.

2. The multi-view based human body surface reconstruction method according to claim 1, characterized in that, Before the step of obtaining indoor RGB pictures from multiple perspectives and using the trained human body UV estimation network to obtain the UV coordinate map of each perspective, it further includes: Set up an indoor environment with multiple cameras for shooting, and use a calibration board for calibration to obtain the camera extrinsic parameter matrix.

3. The multi-view based human body surface reconstruction method according to claim 1, characterized in that, The step of obtaining indoor RGB pictures from multiple perspectives and using the trained human body UV estimation network to obtain the UV coordinate map of each perspective specifically includes: Use the trained segmentation network to obtain the two-dimensional mask of the human body part; Crop the indoor RGB picture of the perspective according to the position of the two-dimensional mask of the human body part.

4. A multi-view based human body surface reconstruction system, characterized in that, it includes: A UV coordinate map determination unit for obtaining indoor RGB pictures from multiple perspectives and using the trained human body UV estimation network to obtain the UV coordinate map of each perspective; An initial depth map determination unit for calculating the optical flow between two adjacent perspectives according to the UV coordinate map of each perspective; And obtaining the initial depth map according to the optical flow and the camera extrinsic parameter matrix; A depth map determination unit for using the trained multi-view depth estimation neural network to determine the depth map of each perspective according to the indoor RGB pictures from multiple perspectives and the initial depth map; The trained multi-view depth estimation neural network includes: a feature pyramid network FPN for encoding the information of a single perspective picture, a feature cross-correlation module for fusing perspective features at different depths of the source perspective observation ray, a depth search mechanism at multiple scales, an attention module, a feature fusion module, and a depth probability estimation module; A three-dimensional human body surface determination unit for extracting the human body point cloud according to the depth map of each perspective and fusing the extracted human body point cloud into a three-dimensional human body surface based on the marching-cubes algorithm; A rendering visualization unit for performing rendering visualization from any perspective based on the three-dimensional surface of the human body.

5. A multi-view based human body surface reconstruction system according to claim 4, characterized in that, it further comprises: An external camera parameter matrix acquisition unit for setting up an indoor environment with multiple cameras for shooting and using a calibration board for calibration to obtain the external camera parameter matrix.

6. A multi-view based human body surface reconstruction system according to claim 4, characterized in that, the UV coordinate map determination unit specifically includes: A human body part two-dimensional mask determination subunit for obtaining a human body part two-dimensional mask using a trained segmentation network; A cropping subunit for cropping the indoor RGB image of the perspective according to the position of the human body part two-dimensional mask.

Citation Information

Patent Citations

  • Multi-visual angle video image depth detecting method and depth estimating method

    CN101231754A

  • Light stream optimization based three-dimensional reconstruction method and device

    CN102800127A