Double-view point cloud reconstruction method based on rotation invariant region consistency
By introducing a method based on rotational invariant region consistency in multi-view three-dimensional reconstruction, using technical means such as PINet and cross attention mechanism, the problem of difficulty in aggregation of different views is solved, and high-quality point cloud reconstruction is achieved.
Patent Information
- Application Number
- CN202510195575.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing multi-view 3D reconstruction methods are difficult to accurately aggregate different view features, resulting in poor reconstruction performance, especially in exploring object area rotation invariance and cross-view object area consistency.
A two-view point cloud reconstruction method based on rotation invariant region consistency is proposed. By introducing a point cloud initialization network (PINet) and a region-level rotation invariant feature extraction network, combining the cross attention mechanism and AoA model, it realizes accurate aggregation and matching of regional features of point clouds in different views.
By learning the rotation invariance and regional consistency of the object area, the accuracy of multi-view feature aggregation is significantly improved. The generated point cloud has fine-grained details and smooth surfaces, which significantly improves reconstruction performance.
Smart Images

Figure CN120047622A_ABST
Abstract
Description
Technical Field:
[0001] A proposed dual-viewpoint cloud reconstruction method based on rotational invariant region consistency belongs to the field of multi-view Figure 3 dimensional reconstruction, involving exploring the rotational invariance of object regions, learning the regional consistency of cross-view objects for multi-view feature aggregation to enhance the reconstruction performance. Background Art:
[0002] The mainstream strategy of current multi-view Figure 3 dimensional reconstruction methods is to aggregate features of different views from objects with the help of different methods, and then reconstruct the three-dimensional shape of the object by decoding the aggregated features. Therefore, feature aggregation plays an important role in multi-view Figure 3 dimensional reconstruction.
[0003] Early multi-view Figure 3 dimensional reconstruction methods used recurrent neural networks (RNNs) to aggregate features from multiple views and decoded the aggregated features into three-dimensional shapes. For example, 3D-R2N2 first used a three-dimensional convolutional neural network (3D-CNN) to extract features of multiple input views and aggregated these features through a three-dimensional long short-term memory neural network (3D-LSTM), and then introduced a deconvolutional neural network to decode the aggregated features and generate a voxel representation of the two-dimensional image. Kar et al. proposed a multi-view Figure 3 dimensional reconstruction method called LSM. This architecture first used a U-Net network to capture grid features of different views, further calculated the relationships between these grid features through a recurrent neural network to guide feature aggregation, and finally used a 3D U-Net network to decode the aggregated features into a voxel representation. R-MVSNet extracted features of multiple views through a convolutional neural network, and at the same time used a gated recurrent unit to selectively retain the extracted features and further aggregate them to reconstruct the point cloud representation of the input object. Due to the lack of long-term memory in recurrent neural networks, the information of the input images is easily gradually forgotten, and multi-view feature aggregation cannot be accurately performed.
[0004] In addition, some works achieve aggregating different view features through pooling techniques or bilinear interpolation algorithms. For example, DeepMVS and 3DensiNet simply use max pooling to aggregate multiple features from a set of images. 3D-RETR extracts features of each input image using ViT and then aggregates these features through adaptive pooling to generate a point cloud representation of the input object. The Pixel2Mesh++ series of methods fuse features of different views through bilinear interpolation and use the fused features to guide the deformation of the template mesh to generate a three-dimensional shape describing the object surface. Although the above methods avoid the limitations of recurrent neural networks, they cannot model the relationships between multiple views and have relatively poor reconstruction performance.
[0005] Therefore, some studies attempt to use the attention mechanism to mine the associations between different views to guide multi-view feature aggregation. DVPC predicts a rough point cloud for each view and captures the mutual information between the rough point clouds from different perspectives through the attention mechanism, thereby performing feature aggregation to reconstruct the point cloud representation. UMIFormer develops a Transformer network for multi-view Figure 3 dimensional reconstruction. It divides each image into sub-image patches of a fixed size and then mines the correlations between the sub-image patches through the attention mechanism to aggregate different view features to reconstruct the three-dimensional shape of the object. Yang et al. propose an attention mechanism based on remote grouping for multi-view Figure 3 dimensional reconstruction. It alternately processes the associations within a single view and between different views to achieve feature aggregation, and then uses deconvolution and self-attention mechanisms to decode the aggregated features to generate a voxel representation of the object. In fact, an object usually has structurally similar regions under different views, and there is a certain degree of rotation between these regions (as Figure 1 shown). Therefore, on the basis of exploring the rotation invariance of object regions, learning the regional consistency of cross-view objects should contribute to multi-view feature aggregation. However, existing methods have not explored this idea, resulting in inaccurate aggregation of different view features and affecting the reconstruction performance.
[0006] Therefore, this paper proposes a dual-view point cloud reconstruction method based on rotation-invariant regional consistency (Dual-view Point cloud reconstruction based on Rotation-invariant Regional consistency, DPR2), which consists of a view encoder and a point cloud decoder, takes a pair of RGB images of different views as input, and gradually restores a high-quality point cloud representation of the given object. Summary of the Invention:
[0007] A dual-viewpoint point cloud reconstruction method based on rotational invariant region consistency, comprising step 1: introducing a point cloud initialization network PINet, and using the PINet network to generate a point cloud x corresponding to the input RGB images of different viewpoints 1 and the point cloud x 2 ;
[0008] Step 2: R2Net divides the point cloud generated by PINet into N regions by farthest point sampling and K-nearest neighbor;
[0009] Step 3: R2Net extracts N region-level rotational invariant features of the point cloud, making the region features have rotational and translational invariance. The region-level rotational invariant features of the point cloud x 1 and the point cloud x 2 are respectively and where represents the rotational invariant feature of the t-th region of the point cloud x 1 , represents the rotational invariant feature of the t-th region from the point cloud x 2 ;
[0010] Step 4: DCM matches the features and through the original cross-attention mechanism to obtain a matching result E = [e 1 ,..., e N , represents the matching result of the feature of each region of x 2 and the feature of the t-th region of x 1 ; Through region feature matching, the regional consistency of different viewpoint point clouds (x 1 and x 2 ) can be initially obtained, that is and e n are a pair of features with regional consistency;
[0011] Step 5: The AoA model is introduced into the original cross-attention mechanism to eliminate the error information in E and obtain an optimized feature matching result
[0012] Step 6: Concatenate the features and the matching result to achieve accurate aggregation of the rotationally invariant features of different regions of the point cloud x 1 and the point cloud x 2 . The aggregated region-level features are denoted as F, and the F is used for the feature aggregation result of the point cloud x 1 and the point cloud x 2 ;
[0013] Step 7: Use PRNet to establish the connection between the aggregated feature F and the point cloud x 1 to generate the second point cloud data with fine structure, denoted as point cloud O;
[0014] Step 8: Feed the output O of the point cloud deformation module into the local folding network to obtain the third point cloud Y with high smoothness. The third point cloud contains the structural information from different views and is used to capture the complex topological structure of the input object.
[0015] Furthermore, the PINet network includes a PSG unit and a convolutional gated recurrent neural network unit. The PINet network repeatedly performs convolutional operations and deconvolutional operations on the input RGB image using an hourglass network to obtain the shape attributes and local details of the RGB image, models the relationship between the shape attributes and the local details using the convolutional gated recurrent neural network unit, selectively updates the shape attributes and the local details in each iteration, and merges the updated shape attributes and local details.
[0016] Furthermore, the step 3 includes: S301: Calculate the Euclidean distance between the centroid point in the region and the neighborhood points, and regard this distance as the "edge" of the region; S302: Use a multi-layer perceptron to extract all the "edges" in the region, and combine all the "edges" in the region to describe the basic structural information of the "edges"; S303: Improve the combination of the "edges" to clearly depict the structural information of the point cloud region.
[0017] Furthermore, according to the original cross-attention, when performing regional feature matching on x1 and x2, is regarded as the query Q, is regarded as the key K and value V; DCM calculates the attention score between Q and K.
[0018] Furthermore, to optimize the matching result E, the AoA model is introduced into the original cross-attention mechanism to only retain the relevant information between Q and K and eliminate the error information in E.
[0019] Furthermore, the point cloud deformation module is pre-constructed. The point cloud deformation model is used to improve the fineness of the point cloud. The point cloud deformation module uses an MLP and a reshaping operation to decode F into a guidance matrix T of the same size as the point cloud x 1 to guide the deformation of each point in the point cloud x 1 to generate the point cloud O.
[0020] This application takes RGB images of two different views as input, promotes multi-view feature aggregation by learning the rotational invariance and regional consistency of object regions to reconstruct high-quality point clouds. It uses a regional-level rotational invariant feature extraction network in the view encoder to obtain the rotational invariance of object regions, uses a two-stage cross-attention mechanism in the point cloud decoder to initially model the regional consistency between two rough point clouds, and uses a point cloud refiner in the point cloud decoder to explore the regional consistency between point clouds of different views, thereby achieving accurate feature aggregation. In addition, the point cloud decoder uses the aggregated features to reconstruct the fine point cloud of the input object, and can generate a point cloud with fine-grained details and a smooth surface, significantly improving the quality of the reconstructed point cloud. And a quantitative and qualitative comparison with existing advanced methods will be carried out below. Experiments on the ShapeNet and Pix3D datasets show that the method proposed in this paper is superior to existing advanced methods in terms of numerical comparison and visualization results. Description of the Drawings:
[0021] Figure 1 There is a certain degree of rotation between regions
[0022] Figure 2 Overall structure of the model method
[0023] Figure 3 Flow of rotational invariant feature extraction for the t-th region of the point cloud
[0024] Figure 4 Structural diagram of the two-stage attention mechanism
[0025] Figure 5 For the ShapeNet dataset, compared with existing multi-view Figure 3 Qualitative comparison of dimensional reconstruction methods
[0026] Figure 6 For the Pix3D dataset, compared with existing multi-view Figure 3 Qualitative comparison of dimensional reconstruction methods Detailed implementation method:
[0027] The following uses an example combined with an example diagram to illustrate this method in detail.
[0028] The present invention includes:
[0029] Step 1: Introduce a point cloud initialization network, called PINet. It follows the structure of the point cloud generation network in the DN-Net method and can process RGB images of different views in parallel. Without camera parameters, it can initialize a point cloud for each image (see Figure 2 point cloud x in 1 and point cloud x 2), the output point cloud is oriented according to the perspective of the given object. Specifically, PINet consists of a PSG unit and a convolutional gated recurrent neural network unit. It uses a hourglass network to repeatedly perform convolutional operations and deconvolutional operations on the input RGB image to obtain its shape attributes and local details respectively. The convolutional gated recurrent neural network unit models the relationship between the shape attributes and local details and selectively updates them in each iteration to more effectively merge these two sets of information. Therefore, PINet has high flexibility in capturing the point cloud structure.
[0030] Step 2: R2Net uses farthest point sampling and K-nearest neighbors to divide the rough point cloud generated by PINet into N regions, where N is set to 4.
[0031] Step 3: R2NeT extracts the features of N regions of the point cloud while ensuring that these features have rotational and translational invariance. Figure 3 Shows the rotational invariance feature extraction process of the t-th region of the point cloud.
[0032] Specifically, it includes: S301: Calculate the Euclidean distance between the centroid point and the neighborhood points in the region and regard this distance as the "edge" of the region. S302: Combine all the "edges" in the region to describe the basic structural information of the "edges". S303: Further improve the combination of the "edges" to clearly depict the structural information of the point cloud region.
[0033] S301 includes: Assume that the centroid point in the t-th region of the point cloud is C t and the l-th neighborhood point of point C t and point B tl form the "edge" J tl The calculation is as follows:
[0034]
[0035] S302 includes: Use a multilayer perceptron (MLP) to extract the features of all the edges in the t-th region:
[0036] P t =[p t1 , p t2 ,..., p tL (2)
[0037] p tl =MLP(J tl ) (3)
[0038] where Represents the feature of the l-th edge in the t-th region of the point cloud Represents the features of all edges in the t-th region of the point cloud. L represents that there are L "edges" in this region, and d is the dimension of the feature. The MLP consists of three linear transformation layers
[0039] In addition, inspired by the Transformer model, R2Net adopts the self-attention mechanism to learn the correlation between "edges" within the region. According to this correlation, the "edges" within the region are further combined to describe the basic geometric structure of the region. Technically, P t is mapped to the query Q = P t W Q , the key K = P t W K , and the value V = P t W V , where is the linear transformation matrix. The correlation w t between all "edges" in the t-th region can be calculated as follows
[0040]
[0041] where d k represents a constant scaling factor
[0042] Subsequently, w t is used to calculate the weighting of each "edge" with other "edges", and the combination of all "edges" is calculated through reshaping operations and fully connected operations
[0043] r t = reshape(MLP(w t V)) (5)
[0044] In the formula, MLP consists of three linear transformation layers; reshape(·) represents the reshaping operation is the rotation-invariant feature of the t-th region of the point cloud and can describe the basic geometric structure of this region
[0045] S303 includes: Based on research findings, regardless of whether the query and the key have no relevant elements, self-attention always outputs a weighted average, which leads to incorrect information in the combination of "edges". Therefore, step 5: Introduce the AoA model to measure the correlation between the combined result r t and the query Q, and filter the incorrect information in r t . The AoA model consists of an information matrix and an attention gate matrix. In the t-th region, r t and Q are linearly transformed to obtain the information matrix I t , r tPerform additional linear transformation on Q and sigmoid activation to construct the attention gate matrix B t .
[0046]
[0047] Among them, represents the mapping matrix for the t-th region; d is the dimension of r t and Q; σ represents the sigmoid activation function.
[0048] B t The value of each channel in can be regarded as the correlation of the value of I t on the corresponding channel. Therefore, multiplying B t with I t by Hadamard Product can improve the error information in r t .
[0049]
[0050] Among them, ⊙ represents the Hadamard Product. represents the improved feature of the t-th region of the point cloud, which can clearly depict the geometric structure of this region.
[0051] Based on Step 3, the region-level rotation invariant features of the point cloud x 1 and the point cloud x 2 are respectively extracted as and Among them represents the rotation invariant feature of the t-th region of the point cloud x 1 , represents the rotation invariant feature of the t-th region from the point cloud x 2 , and t ranges from 1 to N.
[0052] Step 4: DCM matches features through the original cross-attention mechanism and to initially model the regional consistency between two rough point clouds.
[0053] The feature matching process in Step 4 is as follows: According to the original cross-attention, when performing regional feature matching on x1 and x2, is regarded as the query Q, is regarded as the key K and value V. DCM calculates the attention score between Q and the key K.
[0054]
[0055] Among them,[[]] represents the linear transformation matrix, ant Describe x 1 The correlation between the nth region and x 2 The correlation between the tth region
[0056] Matrix Represents the correlation between all regions of two point clouds, which is used to weight V to achieve feature matching between different regions of two point clouds
[0057]
[0058] Among them, Represents the feature matching results of N regions Represents x 2 The feature of each region and the feature matching results of the tth region of x 1
[0059] Step 5, after the original cross-attention mechanism, introduce the AoA model
[22] Optimize the matching results to obtain high-quality regional consistency
[0060] Specifically, through regional feature matching, the regional consistency of rough point clouds in different views can be initially obtained, that is And e n Are a pair of features with regional consistency. Further observe that regardless of whether the query Q and the key K are relevant, the original cross-attention mechanism always calculates the attention score of Q and K. When Q and K are completely irrelevant, there will be error information in the matching result E. To further optimize E, the AoA model is introduced into the original cross-attention mechanism, only retaining the relevant information between Q and K and eliminating the error information in E
[0061] First, use E and Q to construct the information matrix Z and the attention gate matrix Gate, as shown in formulas (11) and (12). Then, perform the Hadamard product of Z and Gate to eliminate the misleading information in E, as shown in formula (13)
[0062]
[0063] Among them, Represents the optimized feature matching results. Z represents the information matrix, and Gate represents the attention gate matrix Is the linear transformation matrix, d represents the dimensions of E and Q. σ represents the sigmoid activation function, and ⊙ represents the Hadamard product
[0064] Through the above operations, DCM can learn the high-quality regional consistency between cross-view rough point clouds, that is And Are a pair of features with high-quality regional consistency.
[0065] Step 6, by splicing and Can accurately aggregate the rotation-invariant features of different regions of the rough point cloud.
[0066]
[0067] Among them, Represents the feature aggregation result of x 1 and x 2
[0068] Step 7: After DCM, PRNet is proposed. It can establish the connection between the aggregated feature F and the rough point cloud to generate a point cloud with a fine structure and a smooth surface. PRNet includes a point cloud deformation module and a local folding network.
[0069] First, a point cloud deformation module is constructed to improve the fineness of the rough point cloud. Specifically, the point cloud deformation module uses MLP and reshaping operations to decode F into a guidance matrix T of the same size as x 1 For guiding the deformation of each point in x 1
[0070] T = reshape(MLP(F)) (15)
[0071] Among them, T represents the guidance signal, which is a matrix of the same size as the point cloud x 1 . reshape(·) represents the reshaping operation. MLP consists of two linear transformation layers.
[0072] To achieve the stable deformation of x 1 , this paper hopes to find a gating matrix c related to and For controlling the information flow of x 1 . A direct way to define c is to splice and And then map it to a weight matrix of the same size as x 1 . The calculation process of the gating matrix c is shown in formula (16), and the calculation process of the point cloud deformation is shown in formula (17).
[0073]
[0074] O = T + c ⊙ x 1 (17)
[0075] Among them, the values in the matrix c consist of real numbers between 0 and 1, σ represents the sigmoid activation function, and O represents the result of the point cloud deformation module.
[0076] Step 8. Previous studies have found that local folding networks prove to be good at approximating smooth surfaces. Therefore, the output O of the point cloud deformation module is fed into this network to further improve its smoothness and generate a point cloud Y of size 8192×3. This point cloud contains structural information from different views, so it can capture the complex topological structure of the input object.
[0077] Quantitative and qualitative comparison with existing advanced methods:
[0078] To evaluate the reconstruction performance of DPR2, we made quantitative and qualitative comparisons of the proposed method with existing advanced methods 3D-R2N2, LSM, DV-Net, DVPC, P2M++, and MVP2M++ on the ShapeNet dataset.
[0079] Quantitative comparison. For fair comparison, the proposed method used the same training and test data as all other methods. All results are from the papers published by methods P2M++, DVPC, and MVP2M++. As shown in Tables 1 and 2, for the average CD metric, DPR2 improved by 80.00%, 57.92%, 30.22%, 23.02%, 23.62%, and 9.06% compared with methods 3D-R2N2, LSM, DV-Net, DVPC, P2M++, and MVP2M++ respectively. For the average F-Score(λ), DPR2 increased by 62.84%, 51.84%, 19.48%, 16.56%, 12.32%, and 7.72% compared with 3D-R2N2, LSM, DV-Net, DVPC, P2M++, and MVP2M++ respectively. For the average F-Score(2λ), DPR2 increased by 36.67%, 31.66%, 7.59%, 5.31%, 5.22%, and 1.42% compared with 3D-R2N2, LSM, DV-Net, DVPC, P2M++, and MVP2M++ respectively. Although the above methods have proposed some relatively advanced feature aggregation strategies, they did not explore object rotation invariance and ignored exploring the regional consistency of cross-view objects, resulting in difficult accurate aggregation of features from multiple views and poor reconstruction results.
[0080] Qualitative comparison. To more clearly illustrate the superiority of the proposed method, we made a qualitative comparison with existing advanced methods. As Figure 5As shown, the voxels generated by 3D-R2N2 have a low resolution and an incomplete structure. The point cloud generated by DV-Net has many outliers and lacks details. Although DVPC can reconstruct a relatively accurate point cloud, it is still difficult to reconstruct objects with complex topological structures, such as the hollow structure of the chair in the second row of the figure. MVP2M++ can recover the 3D mesh of a given object, but the reconstruction of this method lacks reliability. For example, in the fifth row of the figure, the reconstruction of the speaker is not smooth enough. Different from the above methods, the reconstruction of DPR2 is more visually expressive. It can clearly recover the slender structure of the chair legs and the hollow structure of the armrests in the first, second, and fifth rows of the figure. For the car in the fourth row, DPR2 can reconstruct the fine-grained details of the wing part. In addition, it can also effectively capture the smooth surfaces of the sofa and the speaker.
[0081] To verify the generalization ability of DPR2, we made a qualitative comparison of the proposed method with DV-Net, DVPC, and MVP2M++ on the real-world dataset Pix3D. For each input image, we used the mask provided by Pix3D to remove the background. For fair comparison, none of the methods were pre-trained on Pix3D.
[0082] As Figure 6 shown, the proposed method captures the fine-grained details of the object more clearly. Existing advanced methods are difficult to accurately aggregate multi-view features, resulting in insufficient generalization ability. DPR2 explores the regional consistency of cross-view objects from the perspective of the rotation invariance of object regions and can accurately achieve multi-view feature aggregation. Even if the real-world images are not pre-trained, the proposed method can still reconstruct high-quality point clouds. For example, the proposed method captures the slender structure of the chair legs and recovers the smooth surfaces of the bed and the cabinet. Experiments show that DPR2 has good generalization ability for real-world objects.
[0083] Table 1 Quantitative comparison results with existing advanced methods on the CD metric for the ShapeNet dataset
[0084]
[0085] Table 2 Quantitative comparison results with existing advanced methods on the F-Score metric for the ShapeNet dataset
[0086]
Claims
1. A dual-view point cloud reconstruction method based on rotation-invariant region consistency, characterized by: Step 1: Introduce the point cloud initialization network PINet, and use the PINet network to generate point clouds x1 and point clouds x2 corresponding to the input RGB images of different perspectives; Step 2: R2Net uses farthest point sampling and K nearest neighbors to divide the point cloud generated by PINet into N regions; Step 3: R2Net extracts N region-level rotation invariant features of the point cloud, so that the region features have rotation and translation invariance. The region-level rotation invariant features of the point cloud x1 and point cloud x2 are respectively and in represents the rotation invariant feature of the t-th region of the point cloud x1, represents the rotation invariant feature of the t-th region from the point cloud x2, where t ranges from 1 to N; Step 4: DCM matches features via the original cross-attention mechanism and Get the matching result E = [e1,...,e N ], represents the matching result of the features of each region of x2 and the features of the tth region of x1; through regional feature matching, the regional consistency of point clouds x1 and point clouds x2 from different views can be preliminarily obtained, that is, and e n is a pair of features with regional consistency; Step 5: The AoA model is introduced into the original cross-attention mechanism to eliminate the error information in E and obtain the optimized feature matching result. Step 6: The features and matching results Splicing is performed to achieve accurate aggregation of rotation-invariant features of different regions of point cloud x1 and point cloud x2. The aggregated region-level features are recorded as F, and F is used for the feature aggregation results of point cloud x1 and point cloud x2; Step 7: Use PRNet to establish a connection between the aggregated feature F and the point cloud x1 to generate a second point cloud data with a fine structure, recorded as point cloud O; Step 8: Feed the output O of the point cloud deformation module into the local folding network to obtain a third point cloud Y with high smoothness, which contains structural information from different views and is used to capture the complex topological structure of the input object.
2. The method according to claim 1, characterized in that: The PINet network includes a PSG unit and a convolutional gated recurrent neural network unit. The PINet network uses an hourglass network to repeatedly perform convolution operations and deconvolution operations on the input RGB image to obtain shape attributes and local details of the RGB image, uses a convolutional gated recurrent neural network unit to model the relationship between the shape attributes and the local details, and selectively updates the shape attributes and the local details in each iteration, and merges the updated shape attributes and the local details.
3. The method according to claim 1, characterized in that: The step 3 includes: S301: calculating the Euclidean distance between the centroid point and the neighboring points in the region, and considering the distance as the "edge" of the region; S302: extracting all the "edges" in the region using a multilayer perceptron, and combining all the "edges" in the region to describe the basic structural information of the "edges"; S303: improving the combination of "edges" to clearly characterize the structural information of the point cloud region.
4. The method according to claim 1, characterized in that: According to the original cross attention, when performing regional feature matching on x1 and x2, is considered as a query Q, is regarded as a key K and a value V; DCM computes the attention score between Q and the key K.
5. The method according to claim 1, characterized in that: In order to optimize the matching result E, the AoA model is introduced into the original cross-attention mechanism, retaining only the relevant information between Q and K and eliminating the error information in E.
6. The method according to claim 1, characterized in that: The point cloud deformation module is pre-constructed. The point cloud deformation model is used to improve the fineness of the point cloud. The point cloud deformation module uses MLP and reshaping operations to decode F into a guidance matrix T of the same size as the point cloud x1, which is used to guide the deformation of each point in the point cloud x1 to generate a point cloud O.
Citation Information
Patent Citations
Multi-view three-dimensional reconstruction method and system based on uncalibrated image
CN118470219A
Point cloud reconstruction method and apparatus based on pyramid transformer, device, and medium
US11488283B1
Digital image calculation method and system for RGB-d camera multi-view matching based on variable template
US20240428430A1
Data processing method, apparatus, system and storage media
WO2019079766A1
Digital image calculation method and system for deformable template-based RGB-d camera multi-view matching
WO2025000574A1