High-definition three-dimensional scene reconstruction method and device, electronic equipment and storage medium
By performing spatial benchmark alignment and standardized preprocessing on the original scene data, combined with point cloud registration fusion and network optimization, a high-definition 3D target model is generated, which solves the problem of insufficient accuracy in multi-source data fusion and reconstruction in existing technologies, and realizes efficient 3D scene reconstruction and industry application adaptation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN FUTURE QINGYAN INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-05-26
- Publication Date
- 2026-07-31
AI Technical Summary
Existing 3D scene reconstruction methods have shortcomings in multi-source data fusion, spatial alignment, and reconstruction accuracy, and lack the ability to be integrated with industry applications, resulting in low automation.
By performing spatial benchmark alignment and standardization preprocessing on the synchronously acquired raw scene data, point cloud registration fusion and joint reconstruction are carried out, network optimization and spatial semantic information enhancement are performed, and feature extraction is performed according to the task scene type to generate a high-definition 3D target model.
It improves the consistency and reconstruction accuracy of multi-source data, takes into account the integrity and stability of static/dynamic scenarios, and realizes full-process automation and multi-scenario adaptation from reconstruction to delivery.
Smart Images

Figure CN122492937A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of 3D construction technology, and in particular to a method, apparatus, electronic device, and storage medium for high-definition 3D scene reconstruction. Background Technology
[0002] In fields such as industrial quality inspection, intelligent security and criminal investigation, cultural heritage protection, and smart city digital twins, the demand for 3D scene reconstruction is constantly increasing. This requires not only high-precision reconstruction of complex real-world scenes but also support for dynamic scene modeling, multi-source data fusion, and subsequent industry application analysis, thereby achieving end-to-end digitalization and intelligentization from data acquisition to deliverables. For example, industrial inspection requires sub-millimeter-level dimensional analysis, criminal investigation requires millimeter-level scene replication and trajectory extrapolation, and cultural heritage preservation requires micrometer-level detail restoration. These applications all place higher demands on the accuracy, completeness, efficiency, and semantic understanding capabilities of 3D reconstruction.
[0003] Existing 3D scene reconstruction methods are mainly divided into two categories: one is high-precision reconstruction methods based on professional equipment, which rely on LiDAR or multi-camera systems to obtain high-precision models through point cloud registration, structured light or photogrammetry, but usually require manual completion of complex operations such as calibration, noise reduction and model repair, resulting in low automation; the other is lightweight methods based on consumer-grade devices, which lower the barrier to entry, but have significant shortcomings in multi-source data fusion, spatial alignment and reconstruction accuracy. At the same time, existing methods usually take the generation of 3D models as the end point and lack the ability to be combined with industry applications. Summary of the Invention
[0004] In view of this, this application provides a high-definition 3D scene reconstruction method, apparatus, electronic device, and storage medium, which can solve the problem of low adaptability of 3D models to business scenarios in the prior art.
[0005] In a first aspect, embodiments of this application provide a high-definition three-dimensional scene reconstruction method, the method comprising: Spatial reference alignment and standardization preprocessing are performed on the synchronously acquired raw scene data to obtain standard scene detection data; The standard scene detection data is subjected to point cloud registration, fusion, and joint reconstruction to generate an initial 3D mesh model; The initial 3D mesh model is subjected to network optimization and spatial semantic information enhancement processing to obtain a high-definition 3D target model; Based on the received task scenario type, scene features are extracted from the high-definition 3D target model to generate 3D scene application data.
[0006] In some embodiments, the step of performing spatial reference alignment and standardization preprocessing on the synchronously acquired raw scene data to obtain standard scene detection data includes: The original scene data is obtained by synchronously collecting data from preset targets using preset acquisition devices. Based on the device parameters corresponding to each acquisition device, the original scene data is calibrated and transformed to obtain calibrated scene data with unified spatial reference. The calibrated scene data is then cleaned and the neighborhood weighted filling of the void regions is performed to obtain the first optimized scene data. Global spatial registration is performed between different types of data in the first optimized scenario data to obtain the second optimized scenario data. According to the preset configuration confidence weight, the conflict data in the spatially overlapping areas between the second optimized scenario data are weighted averaged and fused to obtain the third optimized scenario data. Calculate the quality index corresponding to the third optimized scenario data, and then filter out standard scenario detection data from the third optimized scenario data based on the preset qualified threshold and the quality index.
[0007] In some embodiments, the step of performing point cloud registration, fusion, and joint reconstruction on the standard scene detection data to generate an initial 3D mesh model includes: Global alignment and registration processing is performed on multiple frames of point clouds in the standard scene detection data to obtain global point cloud data; Based on the global point cloud data, voxel fusion is performed to generate a continuous voxel field, and the isosurface of the continuous voxel field is extracted to obtain an initial fused mesh model. Camera pose estimation and minimization of reprojection error calculation are performed on the viewpoint images in the standard scene detection data to obtain sparse 3D point cloud data and camera pose of each image. Based on the camera pose, the sparse 3D point cloud data is subjected to patch diffusion processing to generate dense depth information. The dense depth information and the initial fused mesh model are then geometrically superimposed and vertex-fused to generate an initial 3D mesh model.
[0008] In some embodiments, the method further includes: Gray-scale change analysis is performed on the viewpoint images of adjacent frames in the standard scene detection data to mark dynamic regions; The dynamic region is segmented into instances and the dynamic objects are tracked to obtain the trajectory information of the dynamic objects. The trajectory information of the dynamic objects is then processed for missing regions and occlusion completion to obtain dynamic geometric information. The dynamic geometric information is spatially aligned and the mesh is overlaid and fused with the initial 3D mesh model to update the initial 3D mesh model.
[0009] In some embodiments, the step of performing point cloud registration, fusion, and joint reconstruction on the standard scene detection data to generate an initial 3D mesh model includes: Global alignment and registration processing is performed on multiple frames of point clouds in the standard scene detection data to obtain global point cloud data; Calculate the average point spacing of the global point cloud data, and extract regional features from the global point cloud data based on the average point spacing to obtain a word feature sequence; Geometric feature enhancement processing is performed on each feature in the word feature sequence according to the preset attention weight coefficients to generate a structured grid model; Gray-scale change analysis is performed on the viewpoint images of adjacent frames in the standard scene detection data to mark dynamic regions and extract Gaussian distribution parameters within the dynamic regions; The Gaussian distribution parameters and the structured mesh model are fused to generate an initial three-dimensional mesh model.
[0010] In some embodiments, performing network optimization and spatial semantic information enhancement processing on the initial 3D mesh model to obtain a high-resolution 3D target model includes: The initial 3D mesh model is subjected to mesh topology repair and mesh simplification processing to obtain an optimized mesh model; Based on the standard scene detection data, the optimized mesh model is texture mapped and rendered to generate a high-fidelity mesh model. Based on the preset semantic feature types, local feature extraction and neighborhood feature correlation analysis are performed on the high-fidelity mesh model to obtain instance objects with the same semantic labels and spatial adjacency. Based on the received interaction instructions, spatial association analysis and semantic information enhancement of the target object are performed on the instance object to obtain a high-definition three-dimensional target model.
[0011] In some embodiments, the step of extracting scene features from the high-definition 3D target model according to the received task scene type to generate 3D scene application data includes: When the task scenario type is a preset industrial quality inspection scenario, the high-definition three-dimensional target model is subjected to size deviation calculation and defect feature extraction to generate industrial quality inspection data. When the task scenario type is a preset security and criminal investigation scenario, physical evidence annotation and trajectory backtracking and deduction are performed on the high-definition three-dimensional target model to generate security monitoring data; When the task scenario type is a preset repair and protection scenario, the damaged areas of the high-definition 3D target model are repaired and the texture is restored to generate cultural heritage display data. When the task scenario type is a preset digital twin scenario, scene change monitoring is performed on the high-definition 3D target model to generate digital twin scenario data; The three-dimensional scene application data includes the industrial quality inspection data, the security monitoring data, the cultural heritage display data, and the digital twin scene data.
[0012] Secondly, this application also provides a high-definition three-dimensional scene reconstruction device, the device comprising: The scene detection module is used to perform spatial reference alignment and standardization preprocessing on the synchronously acquired raw scene data to obtain standard scene detection data. The joint reconstruction module is used to perform point cloud registration and fusion and joint reconstruction on the standard scene detection data to generate an initial three-dimensional mesh model. The optimization and enhancement module is used to perform network optimization and spatial semantic information enhancement processing on the initial three-dimensional mesh model to obtain a high-definition three-dimensional target model. The feature extraction module is used to extract scene features from the high-definition 3D target model according to the received task scene type, and generate 3D scene application data.
[0013] Thirdly, this application also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the high-definition three-dimensional scene reconstruction method as described above.
[0014] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the high-definition three-dimensional scene reconstruction method as described above.
[0015] In summary, this application includes at least the following beneficial technical effects: 1. Improve the consistency and reconstruction accuracy of multi-source data through time synchronization, spatial calibration, cleaning and registration, and quality screening.
[0016] 2. By employing classic reconstruction, dynamic scene compensation, mesh optimization, and semantic enhancement, the integrity and stability of both static and dynamic scenes are ensured.
[0017] 3. By generating application data according to industry scenarios, the entire process from reconstruction to delivery is automated and adapted to multiple scenarios. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of an embodiment of a high-definition three-dimensional scene reconstruction method provided in this application; Figure 2 This is a schematic diagram of an embodiment of an electronic device provided in this application; Figure 3 This is a structural block diagram of a high-definition three-dimensional scene reconstruction device provided in this application. Detailed Implementation
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0021] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0022] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0023] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), unless otherwise expressly and specifically defined.
[0024] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0025] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0026] Firstly, please refer to Figure 1 , Figure 1 This is a schematic flowchart of an embodiment of a high-definition 3D scene reconstruction method provided in this application. The high-definition 3D scene reconstruction method provided in this application includes the following steps.
[0027] Step S1: Perform spatial reference alignment and standardization preprocessing on the synchronously acquired raw scene data to obtain standard scene detection data.
[0028] The raw scene data refers to the set of raw data with spatiotemporal reference markers obtained by unifying the timing and trigger control of acquisition devices such as LiDAR, multi-view camera arrays, RGBD depth cameras, inertial measurement units (IMUs), and UAV aerial survey cameras under the same time reference, after the FPGA hardware triggering center and cross-time synchronization mechanism are used to capture data of the same preset acquisition target (such as an industrial part, a cultural relic site, or a crime scene). Specifically, a 64-bit nanosecond-level precision system clock counter is built inside the FPGA chip. This counter simultaneously receives the GPS PPS second pulse from the UAV flight control and the synchronization message from the local area network PTP master clock. The rising edge of the GPS PPS is taken as the absolute time zero point, and the deviation is corrected every 100 milliseconds of the PTP message to generate a unified system clock reference. Under this reference, the FPGA triggering center outputs a global shutter synchronization trigger signal to the camera array, a point cloud frame synchronization trigger signal to the LiDAR, and a 1kHz sampling synchronization trigger signal to the IMU, and records the corresponding system clock count value at each trigger moment. After acquiring data based on the trigger signal, each acquisition device encapsulates its own recorded local timestamp along with the system clock count into a data frame. The system then executes a linear clock drift compensation algorithm: for each device, a linear model between the local timestamp and the system reference time is established. The drift rate was fitted using the least squares method with paired data from 10 consecutive synchronization cycles. With fixed offset The system detects deviations every 500 milliseconds, and refits parameters when the deviation exceeds 2 milliseconds. After compensation, each data record contains a data frame type identifier, device number, unified timestamp, and original data content (point cloud coordinate array, image pixel matrix, IMU measurement vector, etc.). These records together constitute the spatiotemporally aligned original scene dataset. The purpose of this original scene data is to provide multi-source input with accurate time synchronization for subsequent spatial benchmark unification, avoiding spatial misalignment or ghosting after reconstruction due to differences in timing systems of different sensors (such as the mismatch between UAV GPS timing and local area network PTP timing).
[0029] After obtaining the raw scene data, the system immediately begins spatial reference alignment. Since the internal optical parameters and installation positions relative to the world coordinate system of different sensors vary, multi-sensor joint calibration must be performed first to transform the data collected by each sensor in the raw scene data to a unified spatial reference. For camera intrinsic parameters, the Zhang Zhengyou calibration method is used: a checkerboard calibration board is placed in the scene, 20 to 30 images from different angles are acquired, the sub-pixel coordinates of the checkerboard corner points are automatically detected, and the intrinsic parameter matrix is solved. With distortion coefficient vector ,in , Focal length The coordinates of the main point are used. For the extrinsic parameters between multiple sensors, a hand-eye calibration method based on a calibration board is adopted: the calibration board is placed within the common field of view of all sensors, and LiDAR point clouds and camera images are acquired simultaneously. The 3D coordinates of the corner points of the calibration board are extracted from the point cloud, and the 2D pixel coordinates of the corresponding corner points are extracted from the image. The transformation matrix from the camera to the calibration board is obtained by solving the PnP problem. The transformation matrix of the laser radar to the calibration plate is obtained by fitting the point cloud plane. Then, the extrinsic parameter matrix of the laser radar to the camera is calculated. For the IMU, zero-bias values are acquired by idling for 30 seconds, and then subtracted from the corresponding zero-bias value at each sampling time to compensate for attitude drift. After calibration, all point clouds, images, and depth maps in the original scene data are projected onto a unified world coordinate system with the central camera or LiDAR as the origin through their respective extrinsic parameter matrices, thus obtaining calibrated scene data with a unified spatial reference. Uncalibrated multi-source data may have spatial deviations at the millimeter or even centimeter level due to sensor installation deviations or manufacturing tolerances. If directly used for reconstruction, it will lead to double contours or structural distortions in the model. This deviation is unacceptable, especially in scenarios with sub-millimeter accuracy requirements such as industrial quality inspection and cultural heritage sites.
[0030] After obtaining the calibration scene data, the system performs data cleaning and neighborhood-weighted filling of empty regions. Data cleaning includes three sub-operations: outlier removal, duplicate data removal, and missing value filling. Outlier removal uses the DBSCAN clustering algorithm: each point cloud point in the 3D space is considered as a sample, the k-distance of each sample is calculated, and the core distance threshold is determined based on the k-distance distribution of all samples. The minimum number of neighborhood points will fall on the core point. Points outside the neighborhood are marked as noise points and deleted. Deduplication of duplicate data uses a spatial coordinate hashing method: the space is divided into a grid with sides equal to 0.5 times the average point spacing; for multiple points falling into the same grid, only the point closest to the grid center is retained, and the rest are deleted. Missing value imputation uses the KNN nearest neighbor algorithm: for hollow regions in the point cloud (e.g., areas where the LiDAR does not receive echoes due to specular reflection from object surfaces), the nearest neighbor is searched around the hollow region. points ( The values are adaptively adjusted according to the hole size (usually between 5 and 10). The average coordinates and color of these points are calculated, and the average value is used as the position and color of the filling point to insert into the hole. For bad pixels in the image data (such as overexposed or dead pixels), they are replaced using interpolation of adjacent normal pixels. After cleaning and filling, the point cloud data no longer contains isolated outliers or redundant duplicate points, and the hole areas are reasonably filled, thus obtaining the first optimized scene data. It should be noted that in criminal scene reconstruction, small pieces of evidence such as bloodstains and bullet casings may have missing data due to occlusion or reflection. If filling is not performed, subsequent measurements and annotations cannot be accurately completed. In the scene of cultural relic digitization, the filling of damaged areas must be based on the neighborhood geometric features for a smooth transition; otherwise, the virtually restored model will appear abrupt.
[0031] After obtaining the first optimized scene data, the system performs global spatial registration between data from different types of sensors to eliminate minor offsets caused by calibration residual errors or acquisition motion, generating the second optimized scene data. Registration uses the LiDAR point cloud as the reference set, and the sparse point cloud generated by the camera using the MVS method and the depth point cloud generated by the RGBD depth camera are used as the registration sets. Coarse registration is performed using the NDT algorithm: the space containing the reference point cloud is divided into a voxel grid, and the normal distribution parameters (mean vector) of the point cloud are calculated within each voxel. Covariance Matrix For each point in the point cloud to be registered, calculate its probability density falling within the target voxel, and solve for the initial transformation matrix by maximizing the probability density. Then, using this initial transformation as the initial value, perform ICP fine registration from point to surface: minimize the sum of squared distances from each point in the point cloud to the local plane containing the nearest neighbor of the reference point cloud, and iteratively solve for the rotation matrix. With translation vector until the change update amount is less than The process involves either measuring meters or iterating 100 times. After registration, the LiDAR point cloud, camera depth map, and RGBD data are unified into the same global coordinate system, thus obtaining the second optimized scene data. The significance of this step is that even after calibration, there may still be slight spatial misalignments between different sensors (e.g., due to thermal expansion and contraction of the camera mount caused by temperature changes). Global registration can control the misalignment to the sub-millimeter level, ensuring the geometric consistency of multi-source data in subsequent reconstruction.
[0032] Subsequently, the system performs a weighted average fusion calculation on the conflicting data in the spatially overlapping areas of the second optimized scene data according to preset confidence weights to obtain the third optimized scene data. Since LiDAR, cameras, and RGBD sensors may provide different measurement values in spatially overlapping areas (e.g., LiDAR measures a distance of 1.000 meters, while the RGBD depth camera measures 1.005 meters at the same location), fusion needs to be performed based on the reliability of each sensor. The system presets weights for the LiDAR. Preset weights for the MVS depth map of the industrial camera. Preset weights for RGBD depth cameras Regarding spatial location Suppose there is Each sensor has data at this location, and each sensor provides coordinate values. The weight is Then the merged coordinate values For color information, a weighted average is also used, with the color weight of the RGBD camera set to 0.5 and the color weight of the visible light camera set to 0.9, because the latter has higher color fidelity. After conflict resolution, the previously divergent areas are unified into a continuous and consistent geometric and color value, thus obtaining the third optimized scene data. In large-scale smart city scenarios, UAV aerial survey data and ground-based LiDAR data often conflict on the same building facade. Weighted averaging can combine the advantages of large-scale high-altitude coverage and high-precision ground-based data, avoiding layered or double-skin phenomena in the reconstruction results.
[0033] Finally, the system calculates the quality indicators corresponding to the third optimized scene data and filters out standard scene detection data from the data according to the preset qualified threshold. In this embodiment, the system automatically calculates six quality indicators: data coverage (the ratio of the spatial volume occupied by the preprocessed point cloud to the total volume of the collected scene), noise standard deviation (the standard deviation of the distance from each point in the point cloud to its fitted local plane), average point spacing (the average distance from each point in the point cloud to its nearest neighbor), registration error (the average point-to-point distance in the final iteration of the ICP algorithm), color consistency error (the root mean square of the color difference of the same physical point captured by different cameras in the overlapping area), and texture sharpness (the energy proportion of high-frequency components in the image). Each indicator is compared with the qualified threshold, for example, requiring a data coverage of no less than 95%, a noise standard deviation of less than 0.1 mm, an average point spacing of no more than 1.2 times the preset sampling density, a registration error of less than 0.3 mm, a color consistency error of less than 5 gray levels, and a texture sharpness of more than 0.7. If all indicators meet the thresholds, the current third optimized scenario data is directly output as standard scenario detection data, along with an attached metadata file (recording calibration parameters, cleaning parameters, registration transformation matrix, and values of various quality indicators). If any indicator fails to meet the requirements, the system automatically triggers a reprocessing procedure: based on the type of the failed indicator, the corresponding parameter in the cleaning or registration stage is located (e.g., if the noise standard deviation is too high, the DBSCAN algorithm's...). Threshold; if the registration error is too large, the number of ICP iterations is increased. The system automatically adjusts the parameters according to the preset step size and re-executes the corresponding processing steps until all indicators are qualified or the maximum number of retries (usually 3) is reached. Through this quality screening mechanism, the system ensures that the standard scene detection data that finally enters the 3D reconstruction engine has the characteristics of being noise-free, highly consistent, and having high coverage, meeting the sub-millimeter accuracy requirements of industrial quality inspection and the need for micron-level detail restoration of cultural heritage sites.
[0034] Step S2: Perform point cloud registration, fusion, and joint reconstruction on the standard scene detection data to generate an initial 3D mesh model.
[0035] It should be understood that standard scene detection data refers to the multi-source fusion data obtained in step S1, which has passed quality index screening and has been supplemented with metadata files. This data includes LiDAR point cloud frame sequences, multi-view synchronized image sequences, RGBD depth maps, IMU pose data, and calibration parameters and registration transformation matrices between various sensors. The purpose of this data is to provide highly consistent, noise-free, and spatiotemporally unified input for 3D reconstruction, ensuring that subsequent registration, fusion, and mesh generation can be performed with sub-millimeter accuracy.
[0036] For the first implementation of point cloud registration fusion and joint reconstruction, the system first performs global alignment registration on multiple frames of point clouds in the standard scene detection data to obtain global point cloud data. Since the acquisition device (such as a handheld LiDAR or drone) moves during the acquisition process, there are rotation and translation deviations between adjacent frame point clouds. These frame point clouds must be unified to the same world coordinate system. The system uses the NDT algorithm for coarse registration: the space where the reference frame point cloud is located is divided into a voxel grid, and the normal distribution parameters (mean vector and covariance matrix) of the point cloud are calculated within each voxel. Then, each point in the frame to be registered is mapped to a voxel, and its probability density is calculated. The initial transformation matrix is solved by maximizing the probability density. Next, point-to-surface ICP fine registration is performed: the sum of squared distances from each point in the point cloud to be registered to the local plane containing the nearest neighbor point in the reference point cloud is minimized, and the rotation matrix and translation vector are iteratively solved until the transformation update amount is less than a preset threshold. After the above registration, all point cloud frames are transformed to the same global coordinate system, resulting in global point cloud data. In a smart city scenario spanning square kilometers, if the hundreds of point cloud frames collected by drones are not globally aligned, multiple layers of ghosting will appear on the same building, making subsequent voxel fusion impossible.
[0037] After obtaining global point cloud data, the system performs voxel fusion to generate a continuous voxel field and extracts the isosurfaces of this continuous voxel field to obtain an initial fused mesh model. Voxel fusion employs the TSDF (Truncate Signed Distance Function) algorithm: the global space is divided into a uniform voxel grid, and each voxel stores a truncated signed distance value and a fusion weight. For the center point of each voxel, the signed distance to the nearest surface in each point cloud frame is calculated (negative inside the surface, positive outside), and this distance is truncated within a preset interval. Then, a weighted average is used to fuse the distance values of all frames to obtain the TSDF field. Finally, the zero isosurface is extracted from the TSDF field using the moving cube algorithm to generate a triangular mesh model (i.e., the initial fused mesh model). This initial fused mesh model mainly preserves the large-scale geometry of the scene, such as the walls of buildings, the ground, and the outer contours of large equipment, but lacks fine textures and minute details. In industrial quality inspection scenarios, this mesh model can be used to quickly locate the overall dimensional deviations of parts, but for minor defects such as surface scratches and pores, further reconstruction with higher precision is still required.
[0038] Simultaneously, the system performs camera pose estimation and minimizes reprojection error calculation on the viewpoint images in the standard scene detection data to obtain sparse 3D point cloud data and the camera pose of each image. This operation extracts a multi-view synchronized image sequence from the standard scene detection data, with each image containing RGB three-channel color information. The system extracts AKAZE feature points from each image and generates binary feature descriptors. By matching feature point pairs between adjacent viewpoint images, the correspondence between images is established. For each successful matching pair, the RANSAC algorithm is used to estimate the fundamental matrix and remove mismatched points. Then, the relative rotation matrix and translation vector between the two views are obtained through essential matrix factorization. The incremental SFM framework starts with two-view reconstruction and adds new images sequentially: for each new image, the reprojection matching of existing 3D points in the image is used, and the rotation matrix and translation vector of the current camera are solved using the PnP algorithm; then, triangulation is used to generate new 3D points. After adding a predetermined number of images, the system performs a global bundle adjustment to minimize the sum of squared errors between the pixel coordinates of all observed 3D points reprojected onto the image and the original feature point coordinates, optimizing the objective function as follows: ,in For projection function, The pixel coordinates of the observed feature points. This is the Huber robust kernel function. After this optimization, the system outputs sparse 3D point cloud data and the camera poses (including rotation matrices and translation vectors) for each image. In the cultural heritage scene, the sparse point cloud can quickly outline the overall contour of the Buddha statue, providing a benchmark for subsequent dense reconstruction, while the camera poses are used to guide the direction of texture mapping.
[0039] Based on the camera pose, the sparse 3D point cloud data is processed by patch diffusion to generate dense depth information. This dense depth information is then geometrically superimposed and vertex-fused with the initial fused mesh model to generate the initial 3D mesh model. Patch diffusion employs the PMVS (Patch-based Multi-View Stereo) algorithm: For each image, it is divided into a pixel grid, and a patch is initialized within each grid cell (the center point is the 3D point corresponding to that pixel, and the normal vector is aligned with the direction of the line connecting the camera's optical center to that point). Then, through the patch diffusion propagation step, if there are no patches at the adjacent pixel positions of the current patch, and the current patch has high consistency with the projected color of its neighboring images, a new patch is generated at that adjacent pixel position. During diffusion, the system optimizes the geometric parameters of the patches (center point position and normal vector) to minimize the reprojection color error of the patches across multiple visible images. After diffusion, patch filtering is performed to remove patches visible in fewer images and patches in areas of depth discontinuity. The center points of all patches constitute a dense point cloud. The system reconstructs the dense point cloud into a dense mesh model using Poisson surface reconstruction, and then performs geometric overlay and vertex fusion with the initial fused mesh model obtained through TSDF: for the same location in space, the vertex closer to the original point cloud in both meshes is selected as the final vertex; for topological structures, detailed triangular faces in the dense mesh are retained first, while the initial fused mesh is used to fill any possible holes. This fusion operation can comprehensively utilize the wide-range accuracy of LiDAR point clouds and the rich detail of camera images to generate an initial 3D mesh model that has both macroscopic geometric accuracy and microscopic texture detail. In intelligent security and criminal investigation scenarios, this mesh model can simultaneously present the overall layout of the room (from LiDAR) and the fine structure of small evidence such as bloodstains and bullet casings (from camera MVS), providing a reliable 3D base map for subsequent evidence annotation.
[0040] For standard scene detection data containing moving objects, the system further performs dynamic region labeling and processing. The system analyzes grayscale changes in the viewpoint images of adjacent frames in the standard scene detection data to label dynamic regions. Specifically, it calculates the grayscale difference between corresponding pixels in two adjacent frames to generate a difference image; after Gaussian smoothing of the difference image, a grayscale change threshold (usually 30 grayscale levels) is set, and consecutive pixel regions exceeding the threshold are labeled as candidate dynamic regions. Combined with motion information measured by the IMU (if the device itself is moving, the entire image will have a large range of grayscale changes; in this case, optical flow must be used to distinguish between background and foreground motion), the system finally determines the dynamic regions truly generated by the movement of independent objects in the scene. Because the reconstruction of static backgrounds can reuse information from multiple frames to improve accuracy, while indiscriminately participating in the reconstruction of dynamic objects can lead to object trailing, blurring, or ghostly semi-transparent artifacts. For example, in traffic monitoring scenarios, if moving vehicles participate in TSDF fusion simultaneously with static roads, they will leave afterimages at the locations where the vehicles pass.
[0041] After labeling dynamic regions, the system performs instance segmentation and tracking of dynamic objects within these regions to obtain their trajectory information. It then performs missing region detection and occlusion completion processing on this trajectory information to obtain dynamic geometric information. The system employs a Mask R-CNN instance segmentation network to perform pixel-level segmentation of dynamic regions in each frame, generating a unique instance identifier and contour mask for each dynamic object (e.g., pedestrian, vehicle). In the point cloud domain, the pixel regions corresponding to the masks are back-projected into 3D space to extract a subset of the point cloud belonging to each dynamic object. For the same instance in consecutive frames, a Kalman filter is used to predict its position in the next frame, and inter-frame matching is performed using a Hungarian algorithm to form the object's motion trajectory. When a dynamic object is partially occluded by other objects in the scene (e.g., pillars, walls), its corresponding point cloud will show missing regions. The system uses a generative adversarial mesh-based occlusion completion algorithm: taking the 3D shape and motion trend of the known parts of the dynamic object as input, it generates the shape and point cloud distribution of the occluded part. Each completed dynamic object possesses full 3D geometric information, along with pose and trajectory data for each frame. This information collectively constitutes dynamic geometric information. In industrial automated production line scenarios, workpieces on conveyor belts are dynamic objects. Through instance segmentation, the position, orientation, and speed of each workpiece can be tracked individually, providing data support for subsequent robot grasping.
[0042] Finally, the system aligns the dynamic geometric information with the initial 3D mesh model using spatial coordinates and performs mesh overlay fusion to update the initial 3D mesh model. Since each dynamic object mesh model in the dynamic geometric information already carries pose information for each frame, the system first transforms the current dynamic object mesh into the world coordinate system using its pose, and then merges it with the initial 3D mesh model (which mainly contains the static background). During merging, for the spatial location occupied by the dynamic object, the corresponding vertices in the static mesh are replaced with the mesh vertices of the dynamic object; for the boundary area between the dynamic object and the static background, local mesh smoothing and vertex welding are performed to eliminate seams. In the updated initial 3D mesh model, static scenes (such as buildings, roads, and fixed equipment) have high-precision geometry and texture, while dynamic objects (such as pedestrians, vehicles, and moving objects) are overlaid as independent meshes at the correct timestamp positions, which can be extracted or replayed in animations as needed. In the smart city digital twin scenario, this capability allows the platform to simultaneously display static urban landscapes and dynamic vehicles in real-time traffic flow, achieving a virtual-real integrated situational visualization.
[0043] In another optional implementation, the system employs Transformer-based end-to-end mesh generation instead of the classic SFM-MVS pipeline, while incorporating 3D Gaussian splashing to process dynamic objects. First, global alignment and registration are performed on multiple frames of point clouds in the standard scene detection data (same as the previous implementation), obtaining global point cloud data. Then, the system calculates the average point spacing of this global point cloud data to extract regional features based on this average point spacing, obtaining a token feature sequence. Specifically, the bounding box of the point cloud space is divided into a uniform voxel grid at twice the average point spacing, and tokens are generated only for valid voxels containing at least three point clouds. Each token has an initial 128-dimensional feature vector, containing the average coordinates, average normal vector, principal curvature value, color mean and variance, voxel size, and point cloud density of all points within the voxel. For flat, featureless regions (where the absolute value of the principal curvature is less than 0.5 times the global average curvature), 2×2×2 voxel merging downsampling is performed; for regions with high detail (where the absolute value of the principal curvature is greater than twice the global average curvature), the original voxel tokens are retained. Subsequently, a 16-dimensional spherical harmonic position code is generated for each token, which is concatenated with 128-dimensional features to form a 144-dimensional word feature sequence. This operation transforms the unstructured point cloud into a structured Transformer-processable sequence, laying the foundation for subsequent neural network-generated meshes.
[0044] Based on preset attention weight coefficients, the system performs geometric feature enhancement processing on each feature in the token feature sequence to generate a structured mesh model. The system inputs this token feature sequence into a network consisting of a 6-layer Transformer encoder and a 3-layer Transformer decoder. The encoder employs sparse local geometric enhancement attention: each token calculates its attention weight only with the tokens corresponding to its spatially adjacent 27 voxels (3×3×3 neighborhood), and the attention weight is multiplied by a geometric feature enhancement matrix M, where M=2 for high-feature regions and M=1 for flat regions. The decoder outputs the mesh vertex coordinate matrix, triangular face topological index matrix, and vertex normal vector matrix. Through the trained network, the system directly generates a structured mesh model without the need for traditional SFM and MVS intermediate steps. In static high-precision scenarios (such as cultural relic digitization), this end-to-end method can complete reconstruction within 3 minutes, improving efficiency by more than 5 times while maintaining micrometer-level geometric accuracy.
[0045] Meanwhile, the system performs grayscale change analysis on the viewpoint images of adjacent frames in the standard scene detection data to mark dynamic regions and extract Gaussian distribution parameters within these regions. An optimized 3D Gaussian splashing algorithm is used: the dynamic scene is represented as a set of three-dimensional Gaussian ellipsoids, each defined by its center position, covariance matrix, color spherical harmonic coefficients, and opacity. The system first determines the dynamic regions through grayscale change analysis (same as the previous implementation), then generates Gaussian ellipsoids only by backprojecting each pixel within these regions and calculates their parameters. For each frame, frustum clipping is performed to remove Gaussians outside the frustum; color variance, ellipsoid radius, and contribution are calculated, and low-value Gaussians with color variance less than 0.01, radius less than 0.1 mm, and contribution ranking in the bottom 20% are removed. A CUDA fusion rasterization operator is used to fuse Gaussian sorting, alpha mixing, and gradient calculation into two composite operators, and FP16 half-precision floating-point calculations are used. After identifying dynamic regions using optical flow, the Gaussian parameters of these regions are updated in real-time every frame, while the Gaussian parameters of static regions remain fixed. The extracted Gaussian distribution parameters within the dynamic regions include the center position, covariance, spherical harmonic coefficients, and opacity of each Gaussian. These parameters can be used to render a high-fidelity appearance of dynamic objects in real time.
[0046] Finally, the system fuses the Gaussian distribution parameters and the structured mesh model to generate an initial 3D mesh model. The fusion process is divided into two levels: for static background areas, the structured mesh model generated by Transformer is directly used as the geometric base; for dynamic areas, the Gaussian ellipsoids of the 3D Gaussian splash are projected onto the mesh surface, and the original mesh texture is replaced or superimposed with the color and transparency rendered by Gaussian. In practice, the center position of each Gaussian ellipsoid is associated with the nearest triangle in the structured mesh. For mesh vertices whose distance from the center of the Gaussian ellipsoid is less than the radius of the Gaussian ellipsoid, their color is replaced with the RGB value calculated by the Gaussian spherical harmonic function. For semi-transparent Gaussian ellipsoids (such as glass and water reflections), the Gaussian color is weighted and blended with the original mesh color by opacity through alpha blending. After fusion, the generated initial 3D mesh model has both the complete topological and geometric accuracy of the structured mesh and the dynamic, high-fidelity realistic appearance provided by Gaussian splashing. It is especially suitable for complex scenes containing reflections, transparent materials, or fast-moving objects, such as the instantaneous glass breaking in smart security or the dynamic water flow effect in cultural heritage sites.
[0047] Step S3: Perform network optimization and spatial semantic information enhancement processing on the initial three-dimensional mesh model to obtain a high-definition three-dimensional target model.
[0048] The initial 3D mesh model refers to the triangular mesh data formed after point cloud registration and fusion, joint reconstruction, and dynamic object overlay. It contains the geometric shape, topological structure, and preliminary color information of the scene. However, the model may have topological defects such as holes, non-manifold edges, and degenerate triangular faces. At the same time, its texture resolution is low and it lacks semantic tags, which cannot meet the requirements of industrial quality inspection, cultural heritage, or intelligent security scenarios for model quality and intelligent analysis capabilities.
[0049] First, the system performs mesh topology repair and mesh simplification on the initial 3D mesh model to obtain an optimized mesh model. Topology repair consists of three parallel sub-operations. The first sub-operation is automatic hole filling: the system traverses all boundary edges of the mesh, constructs boundary loops based on the adjacency relationships between boundary edges, and for each loop, uses an implicit surface fitting method based on radial basis functions, with the vertex positions and normal vectors on the loop boundary as constraints, to generate new triangular patches inside the hole region, ensuring that the normal vectors of the newly added patches smoothly transition with the surrounding mesh. The second sub-operation is non-manifold structure repair: when an edge is detected to be shared by more than two triangular faces (i.e., a non-manifold edge), the system re-triangulates the patches around that edge locally; when a vertex is detected to connect multiple geometrically unconnected sets of patches (i.e., a non-manifold vertex), the system splits that vertex into multiple independent vertices, each independent vertex corresponding to a connected set of patches. The third sub-operation is normal unification: Starting from any triangle, the system traverses all adjacent triangles using a depth-first search, adjusting the normal direction of adjacent faces to match the current face (by determining whether the vertex order of shared edges between adjacent faces needs to be reversed), ultimately ensuring that the normals of all triangles in the entire mesh point outwards. The reason for this repair operation is that in industrial quality inspection scenarios, if holes in the model are not filled, subsequent 3D measurements may misjudge the hole boundaries as edge defects; in the digitization of cultural heritage, non-manifold structures can cause texture mapping algorithms to fail to determine the continuity of UV coordinates, resulting in texture tearing.
[0050] After topology repair, the system performs feature-preserving mesh simplification. An edge-folding algorithm based on quadratic error metrics is employed. This algorithm maintains a 4×4 quadratic error matrix for each vertex. This matrix encodes a metric for the sum of squared distances from a vertex to its adjacent triangles. For each connected vertex... and The algorithm calculates the new vertex after folding the edge. The optimal position is found that minimizes the quadratic error between the new vertex and the adjacent face of the original edge. This minimum error value is the cost of folding the edge, calculated using the following formula: The system sorts all edges by cost from smallest to largest, then iteratively folds the edge with the lowest cost. After each fold, the cost of the affected edges is updated and the order is reordered. During the folding process, the system simultaneously calculates the angle between the normal vectors of the triangular faces on both sides of the folded edge. If the angle is greater than 15 degrees, the edge is considered to be on a feature edge, and its folding cost is multiplied by a large penalty coefficient (e.g., 10 times), thus reducing its folding priority. The simplified target face count ratio is set by the user according to downstream application requirements (e.g., simplification to 5% of the original face count for mobile display). After simplification, the total number of faces in the mesh is significantly reduced, but core geometric features (such as edges, sharp corners, and areas with drastic changes in surface curvature) are fully preserved. The feature loss rate is automatically calculated by the system and controlled to within 1%, thus obtaining an optimized mesh model. In large-scale smart city scenarios, the simplified optimized mesh model can compress the original tens of gigabytes of data to hundreds of megabytes, thereby supporting real-time browsing on the web.
[0051] Subsequently, the system performs texture mapping and rendering on the optimized mesh model based on standard scene detection data, generating a high-fidelity mesh model. The standard scene detection data retains the original multi-view high-resolution images and their corresponding camera pose parameters, providing the pixel source for texture mapping. The system first performs intelligent UV unwrapping on the optimized mesh model: using a parameterized algorithm based on minimizing stretching (Least Squares Conformal Maps), each vertex of the triangular mesh is mapped to a point on the 2D UV plane, while minimizing the deformation stretching energy of each triangular face from 3D to 2D. The stretching energy is measured using... ,in and Here are the singular values of the Jacobian matrix. After obtaining the UV coordinates of each vertex by solving the sparse linear system, the system places all the triangular faces of the mesh into a 2D UV texture image and automatically performs multi-quadrant arrangement to maximize texture utilization and reduce seam length. The core of texture mapping is color baking: for each triangular face on the mesh, the system selects the image with the largest projected area of the triangular face and without occlusion by other objects from all camera views as the texture source. Then, based on the UV coordinates of the three vertices of the triangular face, the color of the corresponding pixel in the image is assigned to the texture pixel through a barycentric coordinate interpolation method. Since color discontinuity often occurs at the seams between different texture blocks, the system adopts a U-based... The Net's generative adversarial network performs automatic seam repair: This network takes an 8-pixel-wide strip region on each side of the seam as input and outputs a smoothly transitioned pixel value, ensuring the color difference between the two sides of the seam is less than one gray level. To improve the visual resolution of the texture, the system further performs AI texture super-resolution: using an ESRGAN (Enhanced Generative Adversarial Network) model, the original 4K resolution texture map is enlarged to 16K resolution, while multiple residual dense blocks in the generator are used to recover high-frequency details, such as fabric textures, wood rings, or scratches on metal surfaces. The system also generates PBR materials based on multi-angle lighting images from standard scene detection data: the albedo, normal map, roughness, and metallicity of each pixel are calculated using photometric stereo; the ambient occlusion value is calculated by emitting multiple rays from the mesh surface towards the upper hemisphere and counting the proportion of occlusion. All texture maps (diffuse map, normal map, roughness map, metallicity map, ambient occlusion map) and their corresponding UV coordinates are packaged and mapped onto the optimized mesh model to obtain a high-fidelity mesh model. In the context of cultural heritage, this operation can realistically reproduce the fine details such as gilded patterns and stone inscriptions on the surface of Buddha statues at 16K resolution, making the virtual exhibition achieve a visual effect close to that of the real object.
[0052] After obtaining the high-fidelity mesh model, the system performs local feature extraction and neighborhood feature correlation analysis on the model according to the preset semantic feature types to obtain instance objects with the same semantic labels and spatial adjacency. The system loads a pre-trained PointNet++ 3D semantic segmentation model, whose hierarchical structure includes multiple ensemble abstraction layers (for downsampling and local feature extraction) and feature propagation layers (for upsampling and point-by-point feature recovery). The coordinates of each vertex of the high-fidelity mesh model are then... and optional color information As input, the model outputs a probability vector for each vertex belonging to a preset semantic category (e.g., wall, ground, ceiling, beam, column, vehicle, pedestrian, pipeline, cultural relic, vegetation, etc.). The system takes the category corresponding to the maximum probability as the semantic label of that vertex. After completing vertex-by-vertex semantic annotation, the system performs connected component analysis: constructing a graph structure where each vertex is a node and each grid edge is a connection, clustering all vertices with the same semantic label that are directly connected by edges or indirectly connected by vertices with the same label into a connected component. Each connected component is an instance object; for example, multiple columns in a building will be distinguished as different instance objects, even if they have the same "column" semantic label. The system generates a unique identifier for each instance object and records its bounding box, vertex index range, surface area, volume, and other geometric attributes. The significance of this operation is that, in intelligent security and criminal investigation scenarios, physical evidence such as bloodstains, bullet casings, and murder weapons at the scene need to be distinguished as independent instance objects so that they can be individually labeled and tracked later; in the digital twin of smart cities, every building, every street lamp, and every roadside tree needs to be identified as an independent instance so that its business attributes (such as building name, year of construction, maintenance records, etc.) can be associated.
[0053] Finally, based on the received interaction commands, the system performs spatial association analysis on the instance objects and enhances the semantic information of the target objects to obtain a high-definition 3D target model. Interaction commands are spatial analysis operation requests issued by users through desktop clients, web clients, or mobile clients, such as distance measurement between two points, vertical distance measurement from a point to a surface, area calculation of a specified region, volume calculation of a closed space, extraction of profile lines along a given plane, visibility analysis (determining whether there is occlusion between the observation point and the target point), and collision detection (determining whether two instance objects will intersect during movement). After receiving the interaction command, the system first parses the target instance object identifier and analysis parameters contained in the command (such as the spatial coordinates of the measurement start and end points, the equation coefficients of the cutting plane, the coordinates of the observation point, etc.). For distance measurement, the system calculates the Euclidean distance between the two points. If the user requests a curved distance along the mesh surface, Dijkstra's shortest path algorithm is used to search for the shortest path along the mesh edges and the edge lengths are accumulated. For profile generation, the system calculates the intersection line between the specified cutting plane and each triangular face (by judging the signs of the three vertices of the triangle relative to the plane), connects and sorts all intersection segments, and outputs a closed 2D contour line, which can be exported in DXF format. For visibility analysis, the system emits rays from the viewpoint to the target point, and checks whether the ray intersects any triangular face in the mesh and whether the distance from the intersection point to the viewpoint is less than the distance from the target point to the viewpoint. If so, it is considered invisible; otherwise, it is considered visible. For collision detection, the system constructs an axis-aligned bounding box hierarchy for each instance object, quickly filters out potentially intersecting object pairs, and then performs precise triangular face intersection tests on the candidate pairs. After completing the spatial correlation analysis, the system enhances the analysis results (such as measured values, profile graphics, visibility judgment conclusions, and collision detection reports) on the high-fidelity mesh model in the form of text, highlighted colors, or overlaid lines. Simultaneously, users can manually add custom attributes to instance objects (such as the name of the cultural relic, its age, and material; or, in security scenarios, the evidence number and the person who extracted it). This additional information is stored as metadata for the instance objects and bound to the corresponding geometry. The model processed in the above way is a high-definition 3D target model. This model includes high-precision geometric meshes and 16K-level high-fidelity textures, and also carries semantic tags, instance partitioning, and queryable spatial analysis capabilities. It can be directly used for generating dimensional deviation reports in industrial quality inspection, associating hotspot annotations with voice narration in cultural heritage exhibitions, deducing spatial relationships of evidence in security and criminal investigation, and visualizing traffic conditions in smart cities.
[0054] Step S4: Extract scene features from the high-definition 3D target model according to the received task scene type to generate 3D scene application data.
[0055] In step S4, the system performs differentiated scene feature extraction and industry application data generation operations on the high-definition 3D target model generated in step S3, based on the task scene type identifier carried in the pre-received 3D reconstruction task instruction. The high-definition 3D target model refers to a 3D data volume obtained after mesh optimization, high-fidelity texture mapping, and semantic information enhancement. This model not only includes sub-millimeter precision geometric meshes and 16K resolution texture maps, but also semantic tags, instance segmentation objects, and intermediate spatial analysis results. The task scene type identifier is specified by the user when submitting the reconstruction task from four options: industrial quality inspection, security and criminal investigation, restoration and protection (cultural heritage), and digital twin. Based on this identifier, the system loads the corresponding industry rule engine from a preset process template and performs post-processing on the high-definition 3D target model according to the rule set, ultimately generating the corresponding 3D scene application data.
[0056] When the task scenario is industrial quality inspection, the system calculates dimensional deviations and extracts defect features from the high-definition 3D target model to generate industrial quality inspection data. The core requirement of industrial quality inspection is to compare the reconstructed 3D workpiece model with a standard CAD design model to identify dimensional deviation areas and surface defects. The system first imports the user-provided STEP or IGS format CAD design model and discretizes it into a triangular mesh. Then, an improved ICP algorithm with geometric feature constraints is used to automatically register the high-definition 3D target model with the CAD model. This improved algorithm adds feature point constraints to the traditional point-to-point ICP: the system extracts edge feature points (by calculating the principal curvature of each vertex, points with a curvature change rate exceeding a set threshold are marked as edge points) and regular geometric primitive features (such as cylindrical surfaces, spheres, and planes, extracted through RANSAC algorithm fitting) from both the reconstructed model and the CAD model. During the matching process, feature point pairs are assigned a weight coefficient five times higher than that of ordinary points, ensuring that the registration results do not deviate in geometrically critical areas. After registration, the system calculates the Euclidean distance from each vertex of the high-resolution 3D target model to the nearest corresponding point in the CAD model. And the distance value is calculated according to the following... arrive The range is mapped as a color gradient from blue to red (green at zero deviation), generating a color deviation heatmap. For continuous regions where the absolute value of the deviation exceeds the user-defined tolerance threshold (e.g., ±0.1 mm), the system automatically marks them as dimensional deviation defects. For surface defect extraction, the system further analyzes the normal deviation between the reconstructed model and the CAD model: calculating the normal vector at the vertices of the reconstructed model. Normal vector at the corresponding point in the CAD model The included angle When the included angle is greater than 15 degrees and the distance deviation exceeds 0.05 mm, the area is marked as a surface depression or protrusion defect; when holes or isolated small patches are detected in the reconstructed model, they are marked as surface damage or porosity defects. The location coordinates, deviation values, defect types, and corresponding color heatmaps of all defects are summarized. The system automatically fills in a preset inspection report template (including workpiece name, inspection time, inspector, calibration information, etc.), affixes an electronic signature, and generates PDF format industrial quality inspection data. The significance of this process is that in precision manufacturing, such as the inspection of aero-engine blades, manual visual inspection is unlikely to detect contour deviations as small as 0.1 mm. This method can automatically quantify and visualize the error at each location, providing an objective basis for quality judgment.
[0057] When the task scenario is a security and criminal investigation scenario, the system performs evidence annotation, trajectory backtracking, and deduction on the high-definition 3D target model to generate security monitoring data. Security and criminal investigation scenarios require rapid (within 10 minutes) completion of millimeter-level 3D replication of the crime scene, supporting spatial positioning of evidence and reconstruction of the crime process. The system first extracts the scene's geometric data (including ground plane equations, wall positions, furniture outlines, etc.) from the high-definition 3D target model. The ground plane equation is then fitted using the RANSAC algorithm. The system determines the vertical direction of the world coordinate system. It provides an evidence annotation tool: users can drag and drop icons such as bloodstains, bullet casings, shoe prints, and murder weapons from a pre-made icon library and place them at any spatial coordinate position in the 3D model. The system automatically records the 3D coordinates of the icon. The system includes placement orientation angles and corresponding physical evidence description text. For the suspect's height reference line, the system automatically detects the footprint location (where the user clicks on the footprint area), draws a ray perpendicular to the ground from that location, and subtracts the footprint's Z-coordinate from the Z-coordinate of the highest intersection point of the ray and the humanoid mesh in the model to obtain the estimated height. For trajectory backtracking and deduction, the user can select a series of path points (corresponding to the suspect's walking path or vehicle driving path) on or above the ground in the 3D scene. The system uses a cubic B-spline curve interpolation algorithm to generate a smooth, continuous trajectory curve; then, it calls a proxy character model (humanoid or vehicle model), moves the character model along the trajectory curve, and records the character model's pose at each keyframe, generating a dynamic deduction video of the crime scene through offline rendering. The system also supports packaging physical evidence icons, trajectory curves, and dynamic deduction videos, encrypting the entire data package using the AES256 algorithm to generate an encrypted scene archive model; simultaneously, it exports FBX or OBJ format files required for court presentation, along with digital watermarks and access control information, forming security monitoring data. The value of this processing method in criminal investigation practice lies in the fact that investigators can intuitively display the perpetrator's movement path, crime actions, and spatial relationships between physical evidence in a three-dimensional model, replacing the traditional two-dimensional sketch, making courtroom evidence presentation more intuitive and credible.
[0058] When the task scenario is a restoration and protection scenario (i.e., a cultural heritage protection scenario), the system repairs damaged areas and restores textures on a high-definition 3D target model to generate cultural heritage display data. This scenario requires reverse modeling of cultural relics with micron-level precision and supports virtual restoration and interactive exhibition of damaged parts of ancient sites. The system first detects damaged areas on the high-definition 3D target model: by calculating the rate of curvature change and color gradient of the model surface, areas with abnormal curvature abrupt changes and areas where the color texture differs from the surrounding area by more than a preset threshold are marked as suspected damaged areas; at the same time, users can manually select damaged areas on the model (such as missing facial features of Buddha statues or worn text on stone tablets). For each marked damaged area, the system performs topology restoration: extracting the boundary loops of the damaged area; if the artifact has symmetry (e.g., bronzes and porcelain often have rotational or mirror symmetry), a symmetry completion algorithm is used to mirror the geometric structure of the corresponding position on the symmetrical side to the damaged side, and then the boundary is smoothly blended; if the artifact lacks symmetry, a complete model stored in a database of similar artifacts (e.g., the complete rim of a similar porcelain bowl) is referenced, and a non-rigid registration algorithm is used to deform and fit the corresponding part of the reference model to the damaged area. The restored geometric mesh needs further texture virtual restoration: the system trains a deep learning-based texture generation model, which uses the texture pattern around the damaged area as input to predict the texture pixels of the damaged area. Specifically, a partial convolutional network is used to process irregular damaged areas, gradually filling the holes through multiple dilation convolutions. After restoration, the system generates the data required for a virtual exhibition hall tour: users can add interactive hotspots to the artifact model, each hotspot associated with a three-dimensional spatial coordinate. The system binds text descriptions (such as "Tang Dynasty Celadon Double Dragon Vase, 38cm high") or pre-recorded audio narration files. With one click, the system generates a standalone web-based virtual exhibition package, containing HTML pages, JavaScript code, a WebGL rendering engine based on Three.js, artifact model files (glTF format), and hotspot configuration files. This exhibition package can be directly deployed on any web server, allowing viewers to zoom and rotate 360 degrees to observe the artifacts through a browser, and click on hotspots to view descriptions or listen to narration. The final packaged cultural heritage display data includes both restored high-precision artifact models and complete interactive exhibition content, conforming to the relevant standards for digital preservation of cultural relics issued by the State Administration of Cultural Heritage and is compatible with digital archive systems for cultural relics.
[0059] When the task scenario type is a digital twin scenario, the system monitors scene changes on the high-definition 3D target model and generates digital twin scene data. Digital twin scenarios typically cover large-scale cities or industrial parks covering square kilometers, requiring support for comparative analysis of multiple models and dynamic data fusion and visualization. The system first packages the high-definition 3D target model into 3D Tiles or I3S format according to OGC standards for loading on Cesium or ArcGIS platforms. For scene change monitoring, the system receives two high-definition 3D target models from the same area but different periods (e.g., January 2024 and January 2025) as input. Coarse registration followed by fine registration aligns the two models: the coarse registration stage uses rigid transformation based on building corner points or GPS control point coordinates; the fine registration stage uses the ICP algorithm to further eliminate residual biases. After alignment, the system calculates the 3D displacement vector for each building's corner points (extracted through plane segmentation and edge detection). , displacement modulus This refers to the deformation at that point. All deformations are overlaid on the model as a heatmap (0 mm in green, 5 mm in yellow, and over 10 mm in red), and a list of buildings whose deformation exceeds a preset threshold (e.g., 5 mm) along with their displacement direction and magnitude is output. The system can also integrate real-time dynamic data: for example, obtaining the real-time location, speed, and direction of each ride-hailing vehicle or bus from the city traffic management center (subscribed via the MQTT protocol), mapping this data onto the road grid of the 3D model, and using different colored particle flows to represent the degree of traffic congestion (red for congestion, green for smooth flow). Furthermore, for smart park scenarios, the system can access IoT sensor data from devices (such as temperature, humidity, and vibration), displaying sensor values as floating tags next to the corresponding device models. Finally, digital twin scene data conforming to OGC standards is generated. This data includes basic geographic information, building reality models, deformation monitoring heatmap layers, and real-time traffic flow or sensor data layers, enabling direct connection to the city's intelligent system or park management platform for visualized command and dispatch. The significance of this approach is that city managers can visually see which buildings have experienced uneven settlement and which road sections are severely congested through a three-dimensional digital twin model, thus enabling them to take timely engineering or management measures.
[0060] In summary, step S4 extracts differentiated scene features from the high-definition 3D target model based on the task scenario type. The resulting four types of 3D scene application data serve industrial manufacturing quality inspection, public security criminal investigation scene reconstruction, cultural heritage digital protection, and urban / park digital twin monitoring, respectively, realizing a closed-loop delivery from general 3D models to industry-specific solutions.
[0061] On the other hand, please see Figure 2 , Figure 2This is a schematic diagram of an embodiment of an electronic device provided in this application.
[0062] like Figure 2 As shown, the electronic device 2 in this embodiment includes: at least one processor 21 ( Figure 2 Only one is shown in the diagram), memory 22, and computer program 23 stored in the memory 22 and executable on the at least one processor 21, wherein the processor 21 executes the computer program 23 to implement the steps in the embodiments of the high-definition three-dimensional scene reconstruction method of this application.
[0063] Figure 2 The illustrated electronic device 2 may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art will understand that... Figure 2 This is merely an example of electronic device 2 and does not constitute a limitation on electronic devices. It may have more or fewer components than shown in the figure, or combine certain components, or have different components. For example, it may also include input / output devices, network access devices, etc.
[0064] The processor 21 may be a central processing unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0065] In some embodiments, the memory 22 may be an internal storage unit of the electronic device, such as a hard drive or memory. In other embodiments, the memory 22 may be an external storage device of the electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital card (SD), flash card, etc. Furthermore, the memory 22 may include both internal and external storage units of the electronic device. The memory 22 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 22 can also be used to temporarily store data that has been output or will be output.
[0066] Furthermore, in one embodiment, as Figure 3 As shown, this application also provides a high-definition three-dimensional scene reconstruction device, including: The scene detection module 1100 is used to perform spatial reference alignment and standardization preprocessing on the synchronously acquired raw scene data to obtain standard scene detection data; The joint reconstruction module 1200 is used to perform point cloud registration and fusion and joint reconstruction on the standard scene detection data to generate an initial three-dimensional mesh model. The optimization and enhancement module 1300 is used to perform network optimization and spatial semantic information enhancement processing on the initial three-dimensional mesh model to obtain a high-definition three-dimensional target model. The feature extraction module 1400 is used to extract scene features from the high-definition 3D target model according to the received task scene type, and generate 3D scene application data.
[0067] In some embodiments, the scene detection module 1100 is used for: The original scene data is obtained by synchronously collecting data from preset targets using preset acquisition devices. Based on the device parameters corresponding to each acquisition device, the original scene data is calibrated and transformed to obtain calibrated scene data with unified spatial reference. The calibrated scene data is then cleaned and the neighborhood weighted filling of the void regions is performed to obtain the first optimized scene data. Global spatial registration is performed between different types of data in the first optimized scenario data to obtain the second optimized scenario data. According to the preset configuration confidence weight, the conflict data in the spatially overlapping areas between the second optimized scenario data are weighted averaged and fused to obtain the third optimized scenario data. Calculate the quality index corresponding to the third optimized scenario data, and then filter out standard scenario detection data from the third optimized scenario data based on the preset qualified threshold and the quality index.
[0068] In some embodiments, the joint reconstruction module 1200 is used for: Global alignment and registration processing is performed on multiple frames of point clouds in the standard scene detection data to obtain global point cloud data; Based on the global point cloud data, voxel fusion is performed to generate a continuous voxel field, and the isosurface of the continuous voxel field is extracted to obtain an initial fused mesh model. Camera pose estimation and minimization of reprojection error calculation are performed on the viewpoint images in the standard scene detection data to obtain sparse 3D point cloud data and camera pose of each image. Based on the camera pose, the sparse 3D point cloud data is subjected to patch diffusion processing to generate dense depth information. The dense depth information and the initial fused mesh model are then geometrically superimposed and vertex-fused to generate an initial 3D mesh model.
[0069] In some embodiments, the joint reconstruction module 1200 is further configured to: Gray-scale change analysis is performed on the viewpoint images of adjacent frames in the standard scene detection data to mark dynamic regions; The dynamic region is segmented into instances and the dynamic objects are tracked to obtain the trajectory information of the dynamic objects. The trajectory information of the dynamic objects is then processed for missing regions and occlusion completion to obtain dynamic geometric information. The dynamic geometric information is spatially aligned and the mesh is overlaid and fused with the initial 3D mesh model to update the initial 3D mesh model.
[0070] In some embodiments, the joint reconstruction module 1200 is further configured to: Global alignment and registration processing is performed on multiple frames of point clouds in the standard scene detection data to obtain global point cloud data; Calculate the average point spacing of the global point cloud data, and extract regional features from the global point cloud data based on the average point spacing to obtain a word feature sequence; Geometric feature enhancement processing is performed on each feature in the word feature sequence according to the preset attention weight coefficients to generate a structured grid model; Gray-scale change analysis is performed on the viewpoint images of adjacent frames in the standard scene detection data to mark dynamic regions and extract Gaussian distribution parameters within the dynamic regions; The Gaussian distribution parameters and the structured mesh model are fused to generate an initial three-dimensional mesh model.
[0071] In some embodiments, the optimization and enhancement module 1300 is used for: The initial 3D mesh model is subjected to mesh topology repair and mesh simplification processing to obtain an optimized mesh model; Based on the standard scene detection data, the optimized mesh model is texture mapped and rendered to generate a high-fidelity mesh model. Based on the preset semantic feature types, local feature extraction and neighborhood feature correlation analysis are performed on the high-fidelity mesh model to obtain instance objects with the same semantic labels and spatial adjacency. Based on the received interaction instructions, spatial association analysis and semantic information enhancement of the target object are performed on the instance object to obtain a high-definition three-dimensional target model.
[0072] In some embodiments, the feature extraction module 1400 is used for: When the task scenario type is a preset industrial quality inspection scenario, the high-definition three-dimensional target model is subjected to size deviation calculation and defect feature extraction to generate industrial quality inspection data. When the task scenario type is a preset security and criminal investigation scenario, physical evidence annotation and trajectory backtracking and deduction are performed on the high-definition three-dimensional target model to generate security monitoring data; When the task scenario type is a preset repair and protection scenario, the damaged areas of the high-definition 3D target model are repaired and the texture is restored to generate cultural heritage display data. When the task scenario type is a preset digital twin scenario, scene change monitoring is performed on the high-definition 3D target model to generate digital twin scenario data; It should be noted that the high-definition 3D scene reconstruction device can be understood as a virtual device that can be installed in the electronic device in the aforementioned embodiments. The electronic device calls the high-definition 3D scene reconstruction device through a processor, thereby running the specific implementation scheme in the above-mentioned high-definition 3D scene reconstruction method embodiments. The information interaction and execution process between the aforementioned devices / units are based on the same concept as the method embodiments of this application; their specific functions and technical effects can be found in the method embodiments section.
[0073] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0074] This application also provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps in the above-described method embodiments.
[0075] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0076] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0077] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0078] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for high-definition 3D scene reconstruction, characterized in that, The method includes: Spatial reference alignment and standardization preprocessing are performed on the synchronously acquired raw scene data to obtain standard scene detection data; The standard scene detection data is subjected to point cloud registration, fusion, and joint reconstruction to generate an initial 3D mesh model; The initial 3D mesh model is subjected to network optimization and spatial semantic information enhancement processing to obtain a high-definition 3D target model; Based on the received task scenario type, scene features are extracted from the high-definition 3D target model to generate 3D scene application data.
2. The method according to claim 1, characterized in that, The process of performing spatial reference alignment and standardization preprocessing on the synchronously acquired raw scene data to obtain standard scene detection data includes: The original scene data is obtained by synchronously collecting data from preset targets using preset acquisition devices. Based on the device parameters corresponding to each acquisition device, the original scene data is calibrated and transformed to obtain calibrated scene data with unified spatial reference. The calibrated scene data is then cleaned and the neighborhood weighted filling of empty areas is performed to obtain the first optimized scene data. Global spatial registration is performed between different types of data in the first optimized scenario data to obtain the second optimized scenario data. According to the preset configuration confidence weight, the conflict data in the spatially overlapping areas between the second optimized scenario data are weighted averaged and fused to obtain the third optimized scenario data. Calculate the quality index corresponding to the third optimized scenario data, and then filter out standard scenario detection data from the third optimized scenario data based on the preset qualified threshold and the quality index.
3. The method according to claim 1, characterized in that, The step of performing point cloud registration, fusion, and joint reconstruction on the standard scene detection data to generate an initial 3D mesh model includes: Global alignment and registration processing is performed on multiple frames of point clouds in the standard scene detection data to obtain global point cloud data; Based on the global point cloud data, voxel fusion is performed to generate a continuous voxel field, and the isosurface of the continuous voxel field is extracted to obtain an initial fused mesh model. Camera pose estimation and minimization of reprojection error calculation are performed on the viewpoint images in the standard scene detection data to obtain sparse 3D point cloud data and camera pose of each image. Based on the camera pose, the sparse 3D point cloud data is subjected to patch diffusion processing to generate dense depth information. The dense depth information and the initial fused mesh model are then geometrically superimposed and vertex-fused to generate an initial 3D mesh model.
4. The method according to claim 3, characterized in that, The method further includes: Gray-scale change analysis is performed on the viewpoint images of adjacent frames in the standard scene detection data to mark dynamic regions; The dynamic region is segmented into instances and the dynamic objects are tracked to obtain the trajectory information of the dynamic objects. The trajectory information of the dynamic objects is then processed for missing regions and occlusion completion to obtain dynamic geometric information. The dynamic geometric information is spatially aligned and the mesh is overlaid and fused with the initial 3D mesh model to update the initial 3D mesh model.
5. The method according to claim 1, characterized in that, The step of performing point cloud registration, fusion, and joint reconstruction on the standard scene detection data to generate an initial 3D mesh model includes: Global alignment and registration processing is performed on multiple frames of point clouds in the standard scene detection data to obtain global point cloud data; Calculate the average point spacing of the global point cloud data, and extract regional features from the global point cloud data based on the average point spacing to obtain a word feature sequence; Geometric feature enhancement processing is performed on each feature in the word feature sequence according to the preset attention weight coefficients to generate a structured grid model; Gray-scale change analysis is performed on the viewpoint images of adjacent frames in the standard scene detection data to mark dynamic regions and extract Gaussian distribution parameters within the dynamic regions; The Gaussian distribution parameters and the structured mesh model are fused to generate an initial three-dimensional mesh model.
6. The method according to claim 1, characterized in that, The step of performing network optimization and spatial semantic information enhancement processing on the initial 3D mesh model to obtain a high-definition 3D target model includes: The initial 3D mesh model is subjected to mesh topology repair and mesh simplification processing to obtain an optimized mesh model; Based on the standard scene detection data, the optimized mesh model is texture mapped and rendered to generate a high-fidelity mesh model. Based on the preset semantic feature types, local feature extraction and neighborhood feature correlation analysis are performed on the high-fidelity mesh model to obtain instance objects with the same semantic labels and spatial adjacency. Based on the received interaction instructions, spatial association analysis and semantic information enhancement of the target object are performed on the instance object to obtain a high-definition three-dimensional target model.
7. The method according to claim 1, characterized in that, The step of extracting scene features from the high-definition 3D target model based on the received task scene type to generate 3D scene application data includes: When the task scenario type is a preset industrial quality inspection scenario, the high-definition three-dimensional target model is subjected to size deviation calculation and defect feature extraction to generate industrial quality inspection data. When the task scenario type is a preset security and criminal investigation scenario, physical evidence annotation and trajectory backtracking and deduction are performed on the high-definition three-dimensional target model to generate security monitoring data; When the task scenario type is a preset repair and protection scenario, the damaged areas of the high-definition 3D target model are repaired and the texture is restored to generate cultural heritage display data. When the task scenario type is a preset digital twin scenario, scene change monitoring is performed on the high-definition 3D target model to generate digital twin scenario data; The three-dimensional scene application data includes the industrial quality inspection data, the security monitoring data, the cultural heritage display data, and the digital twin scene data.
8. A high-definition three-dimensional scene reconstruction device, applied to the high-definition three-dimensional scene reconstruction method of claim 1, characterized in that, The device includes: The scene detection module is used to perform spatial reference alignment and standardization preprocessing on the synchronously acquired raw scene data to obtain standard scene detection data. The joint reconstruction module is used to perform point cloud registration and fusion and joint reconstruction on the standard scene detection data to generate an initial three-dimensional mesh model. The optimization and enhancement module is used to perform network optimization and spatial semantic information enhancement processing on the initial three-dimensional mesh model to obtain a high-definition three-dimensional target model. The feature extraction module is used to extract scene features from the high-definition 3D target model according to the received task scene type, and generate 3D scene application data.
9. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the high-definition three-dimensional scene reconstruction method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the high-definition three-dimensional scene reconstruction method according to any one of claims 1 to 7.