Dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower hoisting model
Through the dynamic 3D scene reconstruction method of multi-dimensional point cloud and hybrid tower crane model, the problem of high-fidelity real-time rendering of dynamic scenes in wind power projects is solved, and high-quality real-time display is achieved, which is suitable for a variety of dynamic scene applications.
Patent Information
- Application Number
- CN202510865165.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-14
AI Technical Summary
Existing technologies lack effective dynamic three-dimensional scene representation and rendering methods during the mixed tower crane installation process in wind power projects, resulting in insufficient dynamic interference prediction and delayed spatial collision detection, making it impossible to achieve high-fidelity real-time dynamic display, seriously restricting the improvement of intelligent construction level.
A dynamic 3D scene reconstruction method based on multi-dimensional point cloud and hybrid tower crane model is adopted. By collecting multi-view RGB videos, constructing a multi-dimensional feature grid, combining discrete image fusion network and multi-frequency domain material transfer model, a differential layered rendering pipeline is designed under heterogeneous computing architecture to achieve high-fidelity real-time rendering.
It improves the real-time dynamic display quality of the mixed tower crane model and reduces the computational burden of the inference stage. It is suitable for fields such as interactive virtual reality, motion capture reconstruction and virtual human synthesis.
Smart Images

Figure CN120782991A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graphic image processing, and in particular to a dynamic three-dimensional scene reconstruction method based on a multi-dimensional point cloud and a mixed tower crane model. Background Art
[0002] With the rapid development of the global renewable energy industry, the scale of large-scale infrastructure construction, particularly wind power projects, continues to expand. Concrete tower hoisting, a critical step in the construction of core support structures for wind turbines, involves complex spatial pose transformations and multi-physics field coupling. Traditional construction plan planning relies primarily on two-dimensional drawings and discrete construction simulations. These factors, such as insufficient dynamic interference prediction and delayed spatial collision detection, can easily lead to irreversible damage to multi-million-dollar concrete tower components during the hoisting process. Existing 3D rendering technology faces three technical bottlenecks in engineering applications: First, conventional laser scanning point cloud data lacks a temporal dimension, making it difficult to characterize the dynamic relationship between the hoisting robot's motion trajectory and the deformation of concrete tower components. Second, commercial modeling software often uses static LOD (level of detail) models, which cannot respond in real time to cable sway and tower deflection changes caused by wind loads. Third, data gaps exist between construction simulation systems and BIM platforms, resulting in a lack of a unified analytical framework for visualizing material stress distribution and verifying hoisting plans. Although the field of computer graphics has made progress in real-time volume rendering and point cloud compression and transmission, there is still a lack of spatiotemporal continuous modeling methods for heavy equipment collaborative operation scenarios. In particular, there is a lack of effective visualization methods for the multi-scale deformation mechanism of steel-concrete composite structures during non-rigid lifting, which seriously restricts the improvement of the intelligent construction level of wind power projects.
[0003] Therefore, there is an urgent need for a new type of dynamic three-dimensional scene representation and rendering mechanism to achieve high-speed real-time rendering while maintaining rendering quality, and to improve the high-fidelity real-time dynamic display of the mixed tower crane model. Summary of the Invention
[0004] The purpose of the present invention is to provide a dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model to solve the problems existing in the background technology.
[0005] To achieve the above objectives, the present invention provides a dynamic three-dimensional scene reconstruction method based on a multi-dimensional point cloud and a mixed tower crane model, comprising the following steps:
[0006] S1. Collect RGB video of the multi-view mixed tower crane installation site, covering multiple viewpoints of the target dynamic scene, and obtain data of the target dynamic 3D scene;
[0007] S2. Use time-series-aware point cloud incremental reconstruction to generate point clouds for each frame, obtain a 3D point set for each frame, and construct a point cloud sequence.
[0008] S3, construct a multidimensional feature grid based on the K-Planes method, and encode points into feature vectors through six feature planes;
[0009] S4, using MLP to predict the radius and density of the points and construct the volume structure of the semi-transparent point cloud;
[0010] S5. Combining discrete image fusion network and multi-frequency domain material transfer model to build a mixed tower crane model;
[0011] S6. Use the depth peeling algorithm for differentiable rendering, design a differentiable layered rendering pipeline under heterogeneous computing architecture, extract multiple forward points through K-fold rasterization, and calculate pixel color by combining the volume rendering formula;
[0012] S7. Use the pixel loss MSE and perceptual loss between the rendered image and the real image to train the mixed tower crane model;
[0013] S8. Pre-calculate point attributes during the inference phase and perform real-time rendering.
[0014] Preferably, the multidimensional feature grid in S3 is formed by six feature planes θ xy ,θ xz ,θ yz ,θ tx ,θ ty ,θ tz Composition, integrating spatial and temporal dimension features;
[0015] Given a rough point cloud of the target scene (tower hoisting), a neural network and feature mesh are used to represent its dynamic geometry and appearance. First, six feature planes θ are defined. xy ,θ xz ,θ yz ,θ tx ,θ ty , and θ tz , in order to assign the feature vector f to any point x at frame t, the K-Planes strategy is adopted to model the multidimensional feature field Θ(x, t) using six planes:
[0016]
[0017] Among them, x = (x, y, z) is the input point, Represents the concatenation operator.
[0018] Preferably, the content of S4 is as follows:
[0019] Based on the coarse point cloud, the dynamic scene geometry is represented by learning three entries at each point: position p∈R 3, radius r∈R and density σ∈R, using the point entry, calculate the volume density of the spatial point x relative to the image pixel u for volume rendering.
[0020] Preferably, the contents of the S5 medium-mixed tower crane model are as follows:
[0021] c(x,t,d)=c ibr (x,t,d)+c sh (s,d) (2)
[0022] where c ibr represents the discrete view-dependent appearance; c sh represents the continuous view-dependent appearance; t represents the t-th frame; x represents the point at the t-th frame; d represents the viewing direction; and s represents the SH coefficient at point x.
[0023] Preferably, the content of S6 is as follows:
[0024] Develop a custom shader to implement the depth peeling algorithm consisting of K rendering passes. Consider a specific image pixel u. In the first pass, the point cloud is first rendered on the image using the hardware rasterizer. The point x0 closest to the camera is assigned to pixel u. The depth of point x0 is denoted as t0. Subsequently, in the kth rendering pass, all depth values t k Less than the depth t recorded in the previous rendering process k-1 All points are discarded, resulting in the kth pixel u being closest to the camera point x k ; Since discarding closer points is implemented in a custom shader, it still supports hardware rasterization. After K passes, pixel u has a set of sorted points {x k |1,…,K}; based on the point {x k |k=1,…,K}, use volume rendering to synthesize the color of pixel u; the point {x k The density of |k=1,…,K} is defined based on the distance between the projected point and the pixel u on the 2D image:
[0025]
[0026] Point x k The density is denoted as α k , the color formula of pixel u in volume rendering is:
[0027]
[0028] where c k It is point x k color.
[0029] Preferably, in S4, the radius r and density σ are predicted by feeding the feature vector f in equation (1) to the MLP network.
[0030] Preferably, the discrete image fusion network in S5 is used for view-dependent color prediction and supports pre-computation;
[0031] The multi-frequency domain material transmission model provides support for continuous perspective changes and compensates for the shortcomings of the discrete model.
[0032] Preferably, the training process in S7 is as follows:
[0033] Given a rendered pixel color C(u), compare it to the ground truth pixel color C gt (u) Comparison is performed to optimize the mixed tower crane model in an end-to-end manner using the following loss function:
[0034]
[0035] in is a set of image pixels; except for the MSE loss L img , and also applied the perceptual loss L lpips :
[0036] L lpips =||Φ(I)-Φ(I gt )||1 (6)
[0037] Where Φ is the perception function (VGG16 network), I, I gt They are rendered images and real images respectively.
[0038] Preferably, the mixed tower crane model includes:
[0039] Multi-source data input layer, 4D spatiotemporal point cloud (including XYZ coordinates and timestamp sequence) captured in real time by the lidar array, data stream of the hoisting manipulator posture sensor, and wind speed vector field transmitted by the meteorological monitoring module;
[0040] Image fusion module, based on a real-time texture mapping pipeline with multi-view image geometric correction, uses a feature point-driven non-rigid registration algorithm to align point clouds with BIM models, and an embedded point cloud voxelization engine converts disordered scan data into a spatiotemporally continuous four-dimensional grid;
[0041] A multi-frequency domain material transfer model and a frequency domain illumination transfer system use low-order spherical harmonic basis functions to decompose the bidirectional reflectance distribution function (BRDF) of the mixed tower surface, dynamically integrating the shadow casting of the slings with the scattering characteristics of the concrete material. The radiosity calculation unit and the image fusion module share the GPU memory pool.
[0042] Dynamic coupling interface, cross-module data exchange mechanism, the pixel-by-pixel depth map output by the discrete image fusion network is injected into the illumination pre-integration cache of the multi-frequency domain material transfer model through the CUDA kernel. At the same time, the irradiance field generated by the multi-frequency domain material transfer model reversely drives the material detail enhancement algorithm of the image fusion module;
[0043] The output layer presents a holographic view of the hoisting process at the bottom, integrating physical simulation and visual perception. It includes a thermal map of the tower deflection, a predicted envelope of the cable swing trajectory, and an interference area warning sign. The status panel on the right synchronously displays the GPU rendering frame rate and deformation calculation residual.
[0044] Therefore, the present invention adopts the above-mentioned dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model, which has the following beneficial effects:
[0045] (1) The hybrid appearance modeling strategy that integrates image features and spherical harmonic coefficients can effectively improve the fidelity of texture details and occlusion edges;
[0046] (2) The image fusion part does not depend on the viewing angle and can be pre-calculated after training, reducing the computational burden in the inference stage;
[0047] (3) Compatible with a variety of dynamic scenes, suitable for interactive virtual reality, motion capture reconstruction, virtual human synthesis and other fields.
[0048] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Schematic diagram of the process of the dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model of the present invention;
[0050] Figure 2 Schematic diagram of point cloud feature extraction and four-dimensional feature grid allocation for the dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model of the present invention, wherein (a) is a schematic diagram of the mixed tower crane site; (b) is a schematic diagram of the four-dimensional feature grid allocation;
[0051] Figure 3 It is a structural schematic diagram of the hybrid tower crane model of the dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and hybrid tower crane model of the present invention. DETAILED DESCRIPTION
[0052] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort shall fall within the scope of protection of the present invention.
[0053] See also Figure 1 ,A dynamic 3D scene reconstruction method based on multi-dimensional point cloud and ,mixed tower crane model includes the following steps:
[0054] S1, collect RGB video of the multi-view mixed tower crane installation site, cover multiple viewpoints of the target dynamic scene, and obtain data of the target dynamic 3D scene. Figure 2 , demonstrated point cloud generation and feature projection of dynamic objects from multiple perspectives.
[0055] S2. Use time-series-aware point cloud incremental reconstruction to generate point clouds for each frame, obtain a three-dimensional point set for each frame, and construct a point cloud sequence.
[0056] S3. Construct a multidimensional feature grid based on the K-Planes method and encode points into feature vectors through six feature planes.
[0057] S4. Predict the radius and density of the points through MLP to construct the volume structure of the semi-transparent point cloud. The radius r and density σ are predicted by feeding the feature vector f in formula (1) into the MLP network.
[0058] S5. Combining the discrete image fusion network with the multi-frequency domain material transfer model, a mixed tower crane model is constructed.
[0059] Discrete image fusion network for view-dependent color prediction with pre-computation support;
[0060] The multi-frequency domain material transmission model provides support for continuous perspective changes and compensates for the shortcomings of the discrete model.
[0061] S6. Use the depth peeling algorithm for differentiable rendering, design a differential layered rendering pipeline under heterogeneous computing architecture, extract multiple forward points through K-times rasterization, and calculate pixel color based on the volume rendering formula.
[0062] The multidimensional feature grid in S3 is divided into six feature planes θ xy ,θ xz ,θ yz ,θ tx ,θ ty ,θ tz Composition, integrating spatial and temporal dimension features;
[0063] Given a rough point cloud of the target scene (tower hoisting), a neural network and feature mesh are used to represent its dynamic geometry and appearance. First, six feature planes θ are defined. xy ,θ xz ,θ yz ,θ tx ,θ ty , and θ tz, in order to assign the feature vector f to any point x at frame t, the K-Planes strategy is adopted to model the multidimensional feature field Θ(x, t) using six planes:
[0064]
[0065] Among them, x = (x, y, z) is the input point, Represents the concatenation operator.
[0066] Based on the coarse point cloud, the dynamic scene geometry is represented by learning three entries at each point: position p∈R 3 , radius r∈R and density σ∈R, using the point entry, calculate the volume density of the spatial point x relative to the image pixel u for volume rendering.
[0067] The contents of the S5 medium-sized concrete tower crane model are as follows:
[0068] c(x,t,d)=c ibr (x,t,d)+c sh (s,d) (2)
[0069] where c ibr represents the discrete view-dependent appearance; c sh represents the continuous view-dependent appearance; t represents the t-th frame; x represents the point at the t-th frame; d represents the viewing direction; and s represents the SH coefficient at point x.
[0070] Develop a custom shader to implement the depth peeling algorithm consisting of K rendering passes. Consider a specific image pixel u. In the first pass, the point cloud is first rendered on the image using the hardware rasterizer. The point x0 closest to the camera is assigned to pixel u. The depth of point x0 is denoted as t0. Subsequently, in the kth rendering pass, all depth values t k Less than the depth t recorded in the previous rendering process k-1 All points are discarded, resulting in the kth pixel u being closest to the camera point x k ; Since discarding closer points is implemented in a custom shader, it still supports hardware rasterization. After K passes, pixel u has a set of sorted points {x k |1,…,K}; based on the point {x k |k=1,…,K}, use volume rendering to synthesize the color of pixel u; the point {x k The density of |k=1,…,K} is defined based on the distance between the projected point and the pixel u on the 2D image:
[0071]
[0072] Point x k The density is denoted as αk , the color formula of pixel u in volume rendering is:
[0073]
[0074] where c k It is point x k color.
[0075] S7. Use the pixel loss MSE and perceptual loss between the rendered image and the real image to train the mixed tower crane model.
[0076] Given a rendered pixel color C(u), compare it to the ground truth pixel color C gt (u) Comparison is performed to optimize the mixed tower crane model in an end-to-end manner using the following loss function:
[0077]
[0078] in is a set of image pixels; except for the MSE loss L img , and also applies the perceptual loss L lpips :
[0079] L lpips =||Φ(I)-Φ(I gt )||1 (6)
[0080] Where Φ is the perception function (VGG16 network), I, I gt They are rendered images and real images respectively.
[0081] S8. Pre-calculate point attributes during the inference phase and perform real-time rendering.
[0082] like Figure 3 , mixed tower crane models include:
[0083] Multi-source data input layer, 4D spatiotemporal point cloud (including XYZ coordinates and timestamp sequence) captured in real time by the lidar array, data stream of the hoisting manipulator posture sensor, and wind speed vector field transmitted by the meteorological monitoring module;
[0084] Image fusion module, based on a real-time texture mapping pipeline with multi-view image geometric correction, uses a feature point-driven non-rigid registration algorithm to align point clouds with BIM models, and an embedded point cloud voxelization engine converts disordered scan data into a spatiotemporally continuous four-dimensional grid;
[0085] A multi-frequency domain material transfer model and a frequency domain illumination transfer system use low-order spherical harmonic basis functions to decompose the bidirectional reflectance distribution function (BRDF) of the mixed tower surface, dynamically integrating the shadow casting of the slings with the scattering characteristics of the concrete material. The radiosity calculation unit and the image fusion module share the GPU memory pool.
[0086] Dynamic coupling interface, cross-module data exchange mechanism, the pixel-by-pixel depth map output by the discrete image fusion network is injected into the illumination pre-integration cache of the multi-frequency domain material transfer model through the CUDA kernel. At the same time, the irradiance field generated by the multi-frequency domain material transfer model reversely drives the material detail enhancement algorithm of the image fusion module;
[0087] The output layer presents a holographic view of the hoisting process at the bottom, integrating physical simulation and visual perception. It includes a thermal map of the tower deflection, a predicted envelope of the cable swing trajectory, and an interference area warning sign. The status panel on the right synchronously displays the GPU rendering frame rate and deformation calculation residual.
[0088] Therefore, the present invention adopts the above-mentioned dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model to improve image fidelity and significantly accelerate the rendering process, combining real-time and high quality, and has broad application prospects.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A dynamic 3D scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model, characterized in that: The following steps are involved: S1. Collect RGB video of the multi-view mixed tower crane installation site, covering multiple viewpoints of the target dynamic scene, and obtain data of the target dynamic 3D scene; S2. Use time-series-aware point cloud incremental reconstruction to generate point clouds for each frame, obtain a 3D point set for each frame, and construct a point cloud sequence. S3, construct a multidimensional feature grid based on the K-Planes method, and encode points into feature vectors through six feature planes; S4, using MLP to predict the radius and density of the points and construct the volume structure of the semi-transparent point cloud; S5. Combining discrete image fusion network and multi-frequency domain material transfer model to build a mixed tower crane model; S6. Use the depth peeling algorithm for differentiable rendering, design a differentiable layered rendering pipeline under heterogeneous computing architecture, extract multiple forward points through K-fold rasterization, and calculate pixel color by combining the volume rendering formula; S7. Use the pixel loss MSE and perceptual loss between the rendered image and the real image to train the mixed tower crane model; S8. Pre-calculate point attributes during the inference phase and perform real-time rendering.
2. The dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model according to claim 1 is characterized by: The multidimensional feature grid in S3 is divided into six feature planes θ xy ,θ xz ,θ yz ,θ tx ,θ ty ,θ tz Composition, integrating spatial and temporal dimension features; Given a rough point cloud of the target scene, we use neural networks and feature meshes to represent its dynamic geometry and appearance. We first define six feature planes θ xy ,θ xz ,θ yz ,θ tx ,θ ty , and θ tz , the K-Planes strategy is used to assign the feature vector f to any point x at frame t, and six planes are used to model the multidimensional feature field Θ(x, t): Among them, x = (x, y, z) is the input point, Represents the concatenation operator.
3. The dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model according to claim 2 is characterized in that: S4 content is as follows: Based on the coarse point cloud, the dynamic scene geometry is represented by learning three entries at each point: the position p∈R 3 , radius r∈R and density σ∈R, using the point entry, calculate the volume density of the spatial point x relative to the image pixel u for volume rendering.
4. The dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model according to claim 3 is characterized in that: The contents of the S5 medium-sized concrete tower crane model are as follows: c(x,t,d)=c ibr (x,t,d)+c sh (s,d) (2) where c ibr represents the discrete view-dependent appearance; c sh represents the continuous view-dependent appearance; t represents the t-th frame; x represents the point at the t-th frame; d represents the viewing direction; and s represents the SH coefficient at point x.
5. The dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model according to claim 4 is characterized in that: S6 content is as follows: Develop a custom shader to implement a depth peeling algorithm consisting of K rendering passes. Consider a specific image pixel u. In the first pass, first use the hardware rasterizer to render the point cloud on the image, assign the point x0 closest to the camera to pixel u, and denote the depth of point x0 as t0. Subsequently, in the kth rendering pass, discard all depth values t k Less than the depth t recorded in the previous rendering process k-1 The point that causes pixel u to be the kth closest to the camera point x k ; After K renderings, pixel u has a set of sorted points {x k |1,…,K}; based on the point {x k |k=1,…,K}, use volume rendering to synthesize the color of pixel u; the point {x k The density of |k=1,…,K} is defined based on the distance between the projected point and the pixel u on the 2D image: Point x k The density is denoted as α k , the color formula of pixel u in volume rendering is: where c k It is point x k color.
6. The dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and hybrid tower crane model according to claim 2 is characterized by: In S4, the radius r and density σ are predicted by feeding the feature vector f in Equation (1) into the MLP network.
7. The dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and hybrid tower crane model according to claim 1 is characterized by: Discrete image fusion network in S5 for view-dependent color prediction, supporting pre-computation; The multi-frequency domain material transmission model provides support for continuous perspective changes and compensates for the shortcomings of the discrete model.
8. The dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model according to claim 5 is characterized in that: The training process in S7 is as follows: Given a rendered pixel color C(u), compare it to the ground truth pixel color C gt (u) For comparison, the mixed tower crane model is optimized in an end-to-end manner using the following loss function: in is the set of image pixels; the perceptual loss L lpips The formula is as follows: L lpips =||Φ(I)-Φ(I gt )||1 (6) Where Φ is the perception function; I, I gt They are rendered images and real images respectively.
9. The dynamic three-dimensional scene reconstruction method based on multi-dimensional point cloud and mixed tower crane model according to claim 5 is characterized in that: The mixed tower crane model includes: Multi-source data input layer: 4D spatiotemporal point cloud captured in real time by the lidar array, data stream from the hoisting manipulator's position and posture sensor, and wind speed vector field transmitted by the meteorological monitoring module; Image Fusion Module: A real-time texture mapping pipeline based on multi-view image geometric correction uses a feature point-driven non-rigid registration algorithm to align point clouds with BIM models. An embedded point cloud voxelization engine converts unordered scan data into a spatiotemporally continuous four-dimensional grid. Multi-frequency domain material transfer model: A frequency domain illumination transfer system decomposes the bidirectional reflectance distribution function (BRDF) of the mixed tower surface through low-order spherical harmonic basis functions, dynamically integrating the shadow casting of the slings with the scattering characteristics of the concrete material. Its radiosity calculation unit and image fusion module share the GPU memory pool. Dynamic Coupling Interface: A cross-module data exchange mechanism. The pixel-by-pixel depth map output by the discrete image fusion network is injected into the illumination pre-integration cache of the multi-frequency domain material transfer model through the CUDA kernel. At the same time, the irradiance field generated by the multi-frequency domain material transfer model reversely drives the material detail enhancement algorithm of the image fusion module. Output layer: The bottom layer presents a holographic view of the hoisting process that integrates physical simulation and visual perception. It includes a heat map of tower deflection, a predicted envelope of cable swing trajectories, and an interference area warning sign. The status panel on the right simultaneously displays the GPU rendering frame rate and deformation calculation residual.
Citation Information
Patent Citations
Scene reconstruction and rendering method based on point cloud, storage medium and electronic equipment
CN114898028A
Border-free scene new view angle synthesis method based on mixed neural radiation field
CN116977536A
Efficient volume video representing and rendering method and device
CN117252969A
Indoor three-dimensional object reconstruction method and apparatus, computer device and storage medium
WO2024230151A1
Cited By
Complex scene three-dimensional modeling method and system based on monocular vision
CN120997409A
Camera and illumination combined controllable 4D video generation method, device and equipment
CN121567935A