Three-dimensional reconstruction method and system for urban element universe
By constructing multi-scale scene models and performing cross-scale fusion, the problem of balancing modeling efficiency and accuracy in urban-level 3D reconstruction is solved, and the natural transition and efficient rendering of multi-scale scenes in the urban metaverse are achieved, meeting users' real-time interaction needs.
Patent Information
- Application Number
- CN202510763852.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies in city-level 3D reconstruction and rendering have problems such as difficulty in balancing modeling efficiency and accuracy, cross-scale connection relying on manual adjustment, insufficient expression of dynamic radiation characteristics, and bottlenecks in neural radiation field solutions in scene scale and multi-scale details, which cannot meet the needs of real-time interaction.
Construct a multi-scale scene model by acquiring basic three-dimensional data at different scales, including macro, meso and micro data, and construct corresponding scene models respectively. Then, perform cross-scale fusion through real-time interactive data matching activation strategy to eliminate the sense of visual fragmentation and improve operational efficiency.
It achieves natural transition and efficient rendering of multi-scale scenes in the urban metaverse, meets users' needs for different perspectives, and improves the operating efficiency and visual smoothness of three-dimensional reconstruction.
Smart Images

Figure CN120807767A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of three-dimensional reconstruction, in particular to a three-dimensional reconstruction method, system, electronic device, computer readable storage medium and computer program product of urban meta-universe. BACKGROUND
[0002] Three-dimensional reconstruction refers to establishing a mathematical model suitable for computer representation and processing of a three-dimensional object, which is the basis for processing, operating and analyzing its properties in a computer environment. The key to three-dimensional reconstruction is how to convert the collected three-dimensional data into a model expression in a computer.
[0003] Currently, urban-level three-dimensional reconstruction and rendering have significant characteristics such as large scale, multi-scale, high detail, etc., which makes the existing technology have obvious limitations, such as: (1) The traditional three-dimensional modeling method faces the problem of difficulty in balancing efficiency and accuracy: macro-scale (block-level) modeling takes a long time; micro-scale (indoor-level) texture resolution is limited and data volume grows exponentially; cross-scale connection relies on manual adjustment and lacks dynamic radiation characteristic expression.
[0004] (2) The traditional radiation field modeling has the problem of dimension disaster in global illumination model based on voxelized radiation transfer equation in urban scenes, which cannot meet the real-time interaction demand.
[0005] (3) The traditional neural radiation field (NeRF) scheme has bottlenecks in scene scale, multi-scale detail and training efficiency.
[0006] (4) The 3D Gaussian Splatting scheme improves rendering speed, but still has deficiencies in single Gaussian kernel scale, cross-scale feature confusion and dynamic reflective material modeling.
[0007] Therefore, a three-dimensional reconstruction method, system, electronic device, computer readable storage medium and computer program product of urban meta-universe are needed. SUMMARY
[0008] The present application provides a three-dimensional reconstruction method, system, electronic device, computer readable storage medium and computer program product of urban meta-universe, which can meet the needs of different perspectives of users by constructing a multi-scale scene model and displaying a preset scale scene. According to the real-time interaction data matching activation strategy, the user's focus of attention can be accurately adapted, the system resources can be reasonably allocated, and the running efficiency can be improved. When adjacent scale fusion is involved, the cross-scale fusion technology effectively eliminates the visual fragmentation, making the scene transition natural and smooth.
[0009] The three-dimensional reconstruction method of urban meta-universe provided by the present application adopts the following technical scheme, which includes: constructing a multi-scale scene model; the scene model comprising a plurality of reconstructed objects; presenting a preset-scale scene model to the user; obtaining real-time interaction data between the user and the reconstructed objects; and matching an activation strategy for a region block where the reconstructed objects are located according to the real-time interaction data; when the activation strategy involves a scene model of an adjacent scale, performing cross-scale fusion on the scene model of the adjacent scale; after the cross-scale fusion, presenting a rendered virtual scene to the user.
[0010] Optionally, the constructing a multi-scale scene model comprises: obtaining basic three-dimensional data of a target region; the basic three-dimensional data comprising one or more of macro-scale data, meso-scale data, and micro-scale data; inputting the basic three-dimensional data into a three-dimensional reconstruction model to construct a plurality of original scene models; wherein, a macro-scale scene model is constructed based on the macro-scale data; a meso-scale scene model is constructed based on the meso-scale data; and a micro-scale scene model is constructed based on the micro-scale data.
[0011] Optionally, the constructing a macro-scale scene model based on the macro-scale data comprises: performing macro-reconstructed object extraction on the target region according to satellite remote sensing data; performing block segmentation on the macro-reconstructed objects in combination with LiDAR point cloud data and unmanned aerial vehicle oblique photography data to obtain a plurality of macro blocks; performing local optimization on each of the macro blocks to obtain locally optimized macro blocks; performing global joint optimization on all of the macro-reconstructed objects to obtain the macro-scale scene model.
[0012] Optionally, the constructing a meso-scale scene model based on the meso-scale data comprises: preprocessing the meso-scale data to obtain preprocessed meso-scale data; extracting macro features and detail features of meso-reconstructed objects from the preprocessed meso-scale data; performing cross-scale alignment on the macro features and the detail features to obtain the meso-scale scene model; performing detail enhancement on the meso-scale scene model.
[0013] Optionally, the constructing a micro-scale scene model based on the micro-scale data comprises: determining micro-reconstructed objects corresponding to the micro-scale; super-resolution reconstruction is performed on the micro-reconstruction object based on the micro-scale data to obtain radiation field information; micro-surface modeling is performed on the micro-reconstruction object to obtain surface modeling data; micro-surface detection is performed on the micro-reconstruction object through a multi-spectral defect map to obtain micro-surface detection data; The surface modeling data and the micro-surface detection data are summarized to obtain the micro scene model.
[0014] Optionally, when the activation strategy involves a scene model of a neighboring scale, cross-scale fusion is performed on the scene model of the neighboring scale, including: A scale transition region of the scene model of the neighboring scale is searched; Curved surface optimization is performed on the scale transition region according to a differentiable distance field; The scale transition region is forced to be smooth in terms of radiation field based on a radiation gradient consistency constraint.
[0015] The three-dimensional reconstruction system of the urban meta-universe provided in the application adopts the following technical solution, including: A scene construction module is configured to construct a multi-scale scene model; the scene model includes a plurality of reconstruction objects; An initialization module is configured to show a user a scene model of a preset scale; A strategy matching module is configured to obtain real-time interaction data between a user and the reconstruction objects, and match an activation strategy for a region block where the reconstruction objects are located according to the real-time interaction data; A scale fusion module is configured to perform cross-scale fusion on a scene model of a neighboring scale when the activation strategy involves the scene model of the neighboring scale; A scene rendering module is configured to show a user a rendered virtual scene after the cross-scale fusion.
[0016] Optionally, the scene construction module includes: An acquisition sub-module is configured to acquire basic three-dimensional data of a target region; the basic three-dimensional data includes one or more of macro-scale data, meso-scale data, and micro-scale data; A reconstruction sub-module is configured to input the basic three-dimensional data into a three-dimensional reconstruction model to construct a plurality of original scene models; The reconstruction sub-module includes: A macro-reconstruction unit is configured to construct a macro scene model based on the macro-scale data; A meso-reconstruction unit is configured to construct a meso scene model based on the meso-scale data; A micro-reconstruction unit is configured to construct a micro scene model based on the micro-scale data.
[0017] Optionally, the macro reconstruction unit comprises: a data extraction subunit configured to extract macro reconstruction objects from the target area according to satellite remote sensing data; a block segmentation subunit configured to segment the macro reconstruction objects into a plurality of macro blocks in combination with LiDAR point cloud data and unmanned aerial vehicle oblique photography data; a local optimization subunit configured to locally optimize each of the macro blocks to obtain locally optimized macro blocks; a joint optimization subunit configured to globally and jointly optimize all the macro reconstruction objects to obtain the macro scene model.
[0018] Optionally, the meso reconstruction unit comprises: a preprocessing subunit configured to preprocess the meso scale data to obtain preprocessed meso scale data; a detail extraction subunit configured to extract macro features and detail features of meso reconstruction objects from the preprocessed meso scale data; an alignment subunit configured to cross-scale align the macro features and the detail features to obtain the meso scene model; a detail enhancement subunit configured to enhance the meso scene model in detail.
[0019] Optionally, the micro reconstruction unit comprises: an object confirmation subunit configured to determine micro reconstruction objects corresponding to the micro scale; a super-resolution reconstruction subunit configured to perform super-resolution reconstruction on the micro reconstruction objects based on the micro scale data to obtain radiation field information; a surface modeling subunit configured to perform micro surface modeling on the micro reconstruction objects to obtain surface modeling data; a surface detection subunit configured to perform micro surface detection on the micro reconstruction objects through multi-spectral defect maps to obtain micro surface detection data; a summarizing subunit configured to summarize the surface modeling data and the micro surface detection data to obtain the micro scene model.
[0020] Optionally, the scale fusion module comprises: a region searching sub-module configured to search for a scale transition region of the scene model of the adjacent scale; a first optimization sub-module configured to perform surface optimization on the scale transition region according to a differentiable distance field; a second optimization sub-module configured to perform forced radiation field smoothing on the scale transition region based on a radiation gradient consistency constraint.
[0021] The specification also provides an electronic device, wherein the electronic device includes: a processor; and a memory storing computer-executable instructions that, when executed, cause the processor to perform any of the above methods.
[0022] The specification also provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs that, when executed by a processor, implement any of the above methods.
[0023] The specification also provides a computer program product, wherein the computer program product includes: computer programs / instructions that, when executed by a processor, implement any of the above methods.
[0024] In the present application, a multi-scale scene model is constructed; the scene model includes a plurality of reconstructed objects; a scene model of a preset scale is presented to the user; real-time interaction data between the user and the reconstructed objects is obtained; an activation strategy is matched for a region block where the reconstructed object is located according to the real-time interaction data; when the activation strategy involves a scene model of an adjacent scale, the scene model of the adjacent scale is cross-scale fused; after cross-scale fusion, a rendered virtual scene is presented to the user, eliminating visual fragmentation and improving scene transition fluency. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 A principle diagram of a three-dimensional reconstruction method of an urban meta-universe provided by an embodiment of the specification; Figure 2 A flowchart of a three-dimensional reconstruction method of an urban meta-universe provided by an embodiment of the specification; Figure 3 A structure diagram of a three-dimensional reconstruction system of an urban meta-universe provided by an embodiment of the specification; Figure 4 A structure diagram of an electronic device provided by an embodiment of the specification; Figure 5 A principle diagram of a computer-readable medium provided by an embodiment of the specification. DETAILED DESCRIPTION
[0026] The following description is presented to enable any person skilled in the art to practice the application as claimed. Various modifications to the preferred embodiment will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the present application. Thus, the present application is not intended to be limited to the preferred embodiments described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0027] Example embodiments of the present application will now be described in greater detail from a consideration of the following description in conjunction with the accompanying drawings. The features, structures, characteristics or other details described in the description of certain embodiments can be combined in a suitable manner in one or more other embodiments without departing from the spirit and scope of the present application.
[0028] The term "and / or" or "and / or" includes all combinations of one or more of the associated listed items.
[0029] If the technical solution of the present application involves personal information, the product applying the technical solution of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solution of the present application involves sensitive personal information, the product applying the technical solution of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and the personal information will be collected. If the person voluntarily enters the collection range, it is considered to agree to collect the personal information. Or, on the device for processing personal information, the personal information processing rules are informed through obvious signs / information, and the personal authorization is obtained through pop-up information or by asking the person to upload his / her personal information. The personal information processing rules can include personal information processor, personal information processing purpose, processing method, and personal information type.
[0030] Figure 1 A principle schematic diagram of a three-dimensional reconstruction method of an urban meta-universe provided by an embodiment of the present application, the method comprising: S1, constructing a multi-scale scene model; the scene model comprising a plurality of reconstruction objects; S2, showing a preset scale scene model to the user; S3, acquiring real-time interaction data between the user and the reconstruction object; and matching an activation strategy of a region block where the reconstruction object is located according to the real-time interaction data; S4, when the activation strategy involves a scene model of an adjacent scale, performing cross-scale fusion on the scene model of the adjacent scale; S5, after the cross-scale fusion, showing a rendered virtual scene to the user.
[0031] Three-dimensional reconstruction refers to the establishment of a mathematical model suitable for computer representation and processing of three-dimensional objects. It is the basis for processing, operating and analyzing their properties in a computer environment. How to convert the collected three-dimensional data into a model expression in a computer is the key to three-dimensional reconstruction.
[0032] Currently, city-level 3D reconstruction and rendering have significant characteristics such as large scale, multi-scale, and high detail, which leads to obvious limitations of existing technologies, such as: (1) Traditional 3D modeling solutions use laser scanning and polygon mesh modeling technology (such as CityGML), which presents a contradiction between modeling efficiency and detail accuracy. Macro-scale (block-level) modeling requires manual topology optimization, which takes more than 300 man-hours per km². When the texture resolution of micro-scale (indoor level) is less than 0.1mm / pixel, the data volume increases exponentially (approximately 3.2PB / km²). Cross-scale connection relies on manual LOD adjustment, and the expression of dynamic radiation characteristics is lacking.
[0033] (2) Traditional radiation field modeling solutions, based on the voxelized radiative transfer equation (VRE) global illumination model, face the dimensionality curse in urban scenes. A single 20-story building requires processing approximately 10^15 voxel parameters, which cannot meet the requirements of real-time interaction (>30FPS). The error rate of dynamic lighting simulation is >37% (ISO / IEC 23894 standard test).
[0034] (3) In traditional NeRF solutions, Neural Radiance Fields (NRFs) face three challenges in urban applications: ① scene scale limitations (<1 km²); ② loss of multi-scale detail (MSE > 0.25 @ 4K resolution); and ③ excessive training time (> 72 hours @ NVIDIA A100×8). Existing improved solutions, such as Block-NeRF, produce radiation discontinuities at cross-block splicing (PSNR drops ≥ 6.2 dB).
[0035] (4) Although the traditional 3D Gaussian sputtering solution improves rendering speed (>60FPS@1080P), it faces the following problems in urban scenes: ① Gaussian kernel scale is simplistic (σ∈[0.5m,2m]); ② Cross-scale feature confusion (KL divergence>1.2); ③ Lack of dynamic reflective material modeling (metal material BRDF error rate>42%).
[0036] Overall, there are four major technical gaps in existing 3D reconstruction technology for urban scenes: (1) Lack The unified framework of three-dimensional representation leads to a break in scale continuity.
[0037] (2) The lack of high dynamic range (HDR>20EV) illumination modeling limits the dynamics of radiation.
[0038] (3) Multi-source sensor (LiDAR / tilt photography / IoT sensor) data fusion efficiency is low, and there is a data heterogeneity barrier.
[0039] (4) Distributed neural radiation field training has a super-linear relationship (O(n^1.8)) with the size of the scene, causing a computing scalability barrier.
[0040] To solve the above problems, the present application provides a three-dimensional reconstruction method of urban meta-universe, as shown in the figure, the method comprises: Figure 2 S1 constructing a multi-scale scene model; S11 obtaining basic three-dimensional data of the target area; S111 obtaining modeling requirements of the target area; The modeling requirements include at least one scale dimension.
[0041] The scale dimension includes at least one of macro scale (block level, modeling area > 500 m²), meso scale (building level, 100 m²≤ modeling area≤ 500 m²), and micro scale (indoor level, modeling area < 100 m²). Among them, the macro scale and the meso scale are adjacent scales, and the meso scale and the micro scale are adjacent scales.
[0042] When the modeling requirements include one of the multiple scale dimensions (such as macro scale; meso scale; micro scale), the modeling requirements are single-scale modeling, that is, a single scale dimension.
[0043] When the modeling requirements include a combination of two adjacent scale dimensions (macro scale + meso scale; meso scale + micro scale), or when the modeling requirements include a combination of three adjacent scale dimensions (macro scale + meso scale + micro scale), the modeling requirements are multi-scale modeling.
[0044] When the modeling requirements are multi-scale modeling, to avoid geometric feature distortion, scale continuity constraints can also be performed by prohibiting cross-non-adjacent scale combinations (such as macro scale + micro scale) to avoid non-physical cross-scale data fusion.
[0045] In one embodiment of the present application, when the target area is urban meta-universe, to provide a better experience to users, the modeling requirements include: macro scale + meso scale + micro scale.
[0046] S112 obtaining basic three-dimensional data of the target area based on the modeling requirements; The basic three-dimensional data is used to represent the geometric structure, physical properties and texture features of the target area in three-dimensional space. The basic three-dimensional data includes one or more of macro scale data, meso scale data and micro scale data.
[0047] It is easy to understand that the type of basic three-dimensional data is related to modeling requirements, wherein single-scale modeling calls a single data source of the corresponding scale dimension. Multi-scale modeling calls data sources of adjacent scale dimensions.
[0048] Specifically, when the scale dimension in the modeling requirement includes a macro scale, macro scale data is called; when the scale dimension in the modeling requirement includes a meso scale, meso scale data is called; when the scale dimension in the modeling requirement includes a micro scale, micro scale data is called.
[0049] S112-1 when the scale dimension includes a macro scale, macro scale data is called; The macro scale data is macro block level image data, and the macro scale data includes satellite remote sensing data, LiDAR point cloud data, and unmanned aerial vehicle oblique photography data.
[0050] In an embodiment of the present specification, the ground sampling interval (GSD) of the satellite remote sensing data is 0.31 meters. The ground sampling interval (GSD) of the unmanned aerial vehicle oblique photography data is 2 centimeters, and the overlap rate is > 85%. The point density of the LiDAR point cloud data is > 500 pt / m².
[0051] GSD = flight height / (camera focal length * sensor size / image resolution).
[0052] S112-2 when the scale dimension includes a meso scale, meso scale data is called; The meso scale data is meso building level image data, and the meso scale data includes unmanned aerial vehicle close-range photography data, local point cloud data, and illumination intensity data.
[0053] In an embodiment of the present specification, the unmanned aerial vehicle close-range photography data is 2mm / pixel level texture data obtained using a 100mm fixed focus lens. The local point cloud data is collected by a ground laser scanner. The illumination intensity data is collected by an IOT sensor deployed with an illumination intensity meter.
[0054] S112-3 when the scale dimension includes a micro scale, micro scale data is called; The micro scale data is micro indoor level image data, and the micro scale data includes 4K resolution photos.
[0055] The 4K resolution photos include micro feature images and / or multispectral defect maps. The micro feature images are 4k RGB images, wherein for a small scene, the number of 4K resolution photos is at least 50-100, and for a multi-scene, the number of 4K resolution photos is increased according to coverage requirements. When shooting, multi-directional and wrap-around shooting is performed to reduce blind areas.
[0056] If there is external calibration (such as laser radar or photogrammetry precise external parameters), 4K resolution photos can be directly imported; if there is no external calibration: to support the COLMAP process, high overlap and clear feature points are needed; inaccurate camera external parameters will cause distortion of the NeRF model, and calibration quality needs to be ensured as much as possible.
[0057] When shooting, the same scene is shot under similar lighting as much as possible to avoid day and night or strong light / weak light mixed shooting to maintain light consistency; for static environment, reduce moving objects (people, cars), or use image mask to remove dynamic interference; in order to avoid serious blur, inaccurate focusing, and properly increase the shutter to prevent trailing.
[0058] The data format can be in.PNG format, and the pose is in COLMAP file (.txt) or custom JSON / YAML; optional depth map: PNG(16-bit) / EXR.
[0059] Specifically, the microscopic feature image includes: a synchronous phase projection sequence and a multi-view stereo image pair.
[0060] When acquiring the microscopic feature image, in an embodiment of the present specification, a customized multi-spectral structured light projector (450nm / 650nm / 850nm / 1050nm) and a multi-view stereo camera array are deployed; the camera-projector internal and external parameters are calibrated (the accuracy of the calibration board is ≤0.01mm); the synchronous control module ensures that the light field projection and the camera exposure timing are aligned (error <1ms) to collect microscopic scale data, and obtain a synchronous phase projection sequence and a multi-view stereo image pair.
[0061] As preferred, the macro / meso scale data acquisition accuracy: decimeter level (5-10 cm). Microscopic scale data acquisition accuracy: centimeter level (1-3 cm or better).
[0062] When collecting macro / meso / micro scale data, multi-sensor combination can be used, such as LiDAR + high-resolution industrial camera + IMU + GNSS, with positioning accuracy better than 10 cm.
[0063] For important target areas (such as complex buildings, key objects), multi-angle high overlap collection is performed, with an overlap of ≥50%-70%.
[0064] For a single target, not less than 3-5 times of scanning is needed, and high-resolution images are combined to ensure full-azimuth texture.
[0065] Use HDR panoramic cameras or multispectral sensors; collect lighting information in multiple time periods and weather conditions; and use photometers / radiometers to ensure that the accuracy of subsequent radiation field modeling is ≥90%.
[0066] During data collection, in order to generate high-resolution and high-frame-rate data, the acquisition device must support 4K / 30fps or above; this lays a good data foundation for meeting rendering indicators of ≥4K, ≥30fps, and PSNR≥30dB.
[0067] During the acquisition process, quality control is carried out: specifically, point cloud density, image clarity, IMU / GNSS drift, etc. are monitored in real time; after the acquisition is completed, rapid pre-processing is carried out on site to check data integrity and timely supplement any missing data.
[0068] S12 inputs the basic three-dimensional data into a three-dimensional reconstruction model to construct an original scene model; The type of original scene model is related to the scale dimension involved in the modeling requirements.
[0069] When the modeling requirements include macro scale, a macro scene model will be constructed based on the macro scale data; when the modeling requirements include meso scale, a meso scene model will be constructed based on the meso scale data; when the modeling requirements include micro scale, a micro scene model will be constructed based on the micro scale data.
[0070] Therefore, the original scene model includes one or more of a macroscopic scene model, a mesoscopic scene model, and a microscopic scene model.
[0071] S121 constructs a macroscopic scene model based on the macroscopic scale data; S121-1 extracts macroscopic reconstruction objects from the target area based on satellite remote sensing data; Satellite remote sensing data (such as Sentinel-2 and Landsat) can cover urban and even regional scenes (such as 10km×10km), providing a global perspective for dynamic segmentation and avoiding the blindness of local segmentation.
[0072] S121-2 combines the LiDAR point cloud data and the UAV oblique photography data to perform block segmentation on the macro-reconstructed object to obtain a plurality of macro blocks; S121-2-1 combines LiDAR point cloud data and UAV oblique photography data to perform complexity evaluation on the macro-reconstructed object and generate regional features; Regional characteristics include: building density, building materials, and geometric coefficient of variation.
[0073] Most of the current dynamic block methods rely on traditional clustering (such as K-means) or image similarity-based segmentation, which lack explicit complexity modeling. Therefore, the present application constructs a complexity evaluation model based on a graph neural network for the hierarchical block neural radiance field dynamic block division. Specifically: (1) Determine the building density of the macro reconstruction object according to the LiDAR point cloud data; The building density is used to represent the edge number density of the structural elements (walls, beams and columns) in a unit area, reflecting the geometric complexity of the scene.
[0074] Specifically, according to the LiDAR point cloud data, the building outline is extracted, and a three-dimensional structure line segment is generated; the number of edge points in a unit area (such as a 1m x 1m grid) is counted; and the building density is determined based on the edge number statistics.
[0075] Wherein; the building density includes high complexity and low complexity, and the building density is determined based on the edge number per unit area. When the edge number per unit area is , the building density is high complexity, otherwise the building density is low complexity.
[0076] (2) Determine the building material of the macro reconstruction object according to the unmanned aerial vehicle oblique photography data; The building material is calculated by RGB color, texture features (such as normal vector, roughness) and reflectivity.
[0077] Specifically, the surface texture features of the macro reconstruction object are extracted according to the unmanned aerial vehicle oblique photography data; the surface texture features are classified by a ResNet model, and the building material of each part of the macro reconstruction object is determined based on the texture entropy In an embodiment of the present application, ResNet-50 is used to extract the texture entropy , and when the texture entropy , the building material is marked.
[0078] (3) Determine the geometric coefficient of variation of the macro reconstruction object according to the LiDAR point cloud data; The geometric coefficient of variation is used to represent the mutation degree of local geometric shape. Specifically, the LiDAR point cloud data is down-sampled and the normal vector is estimated; the local curvature (such as principal curvature, Gaussian curvature) is calculated; the standard deviation and mean of the regional curvature are calculated to obtain the geometric coefficient of variation. In an embodiment of the present application, when the geometric coefficient of variation , the geometric coefficient of variation is high, otherwise the geometric coefficient of variation is low.
[0079] S121-2-2 determines the complexity distribution of the macro reconstruction object according to the region features; (1) Uniformly mesh the macro reconstruction object, and take each mesh as a node (the number of nodes is ); Node features ; wherein, is the building density; is the building material; is the geometric variation coefficient. The edge is the adjacent mesh (8-neighborhood), and the edge weight is the material similarity.
[0080] (2) Construct a complexity evaluation model; The complexity evaluation model is preferably a graph neural network.
[0081] (3) Take the node feature matrix (dimension ) as input; in the graph convolution layer, the material similarity and geometric features of the neighbor nodes are aggregated based on message passing, and the node hidden representation is output based on the MLP encoder; the complexity score of each node is output by using the multilayer perceptron MLP, as the complexity distribution.
[0082] S121-2-3 hierarchically models the macro reconstruction object based on the complexity distribution to obtain a plurality of hierarchical macro blocks; The macro block is a dynamic granularity block.
[0083] S121-3 locally optimizes each macro block to obtain a locally optimized macro block; S121-3-1 locally trains each macro block to obtain a trained macro block; (1) Adjust the granularity of the macro block in combination with the scene complexity category of the target region; The boundary of the adjacent block needs to reserve a certain overlap area (such as 5%-10%), to avoid the reconstructed scene from appearing “cracks” or rendering discontinuously due to the hard segmentation. Specifically: according to the scene complexity category of the target region, the boundary overlap degree is determined; The scene complexity category of the target region includes: a first complexity category, a second complexity category, and a third complexity category. The complexity degree of each scene complexity category is: the first complexity category > the second complexity category > the third complexity category.
[0084] Each scene complexity category corresponds to an adjustment strategy for adjusting the default block granularity and the default boundary overlap rate.
[0085] In an embodiment of the present specification, the granularity of the default macro block is 200m x 200m, and the default boundary overlap rate is 5%.
[0086] When the category of the target area is the first complex category, the granularity of the macro block is reduced and the boundary overlap rate is increased according to the first adjustment strategy.
[0087] For example, the granularity of the macro block is adjusted to 50m x 50m, ensuring that the details within each block can be fully captured by the NeRF model; the parameter sharing rate is reduced (e.g., from 72% to a lower value), avoiding the blurring of details in adjacent blocks due to parameter sharing. The corresponding boundary overlap rate is determined to be 10%, and through the cross-training of details in the overlapping area, the connection between adjacent fine blocks is ensured to be natural. Among them, the first complex category is a commercial area, which generally involves a dense building group, contains complex geometric structures, and contains a large number of details (such as carvings, special-shaped buildings), which require higher modeling accuracy, and is of high complexity, which requires higher modeling accuracy.
[0088] When the category of the target area is the second complex category, the granularity of the macro block is reduced and the boundary overlap rate is increased according to the second adjustment strategy.
[0089] For example, the granularity of the macro block is adjusted to 100m x 100m; the corresponding boundary overlap rate is determined to be 8%, and the cross-block consistency of 0.1mm-level carving details is preserved. Among them, the second complex category is a historical block, which is of medium complexity and high detail.
[0090] When the category of the target area is the third complex category, the granularity of the macro block is maintained or expanded and the boundary overlap rate is maintained or reduced according to the third adjustment strategy.
[0091] For example, the granularity of the macro block is maintained at 200m x 200m or continues to be expanded (e.g., adjusted to 500m x 500m), reducing the number of blocks to improve training efficiency; the parameter sharing rate is increased, and the calculation cost is reduced by sharing basic parameters across blocks; the corresponding boundary overlap rate is determined to be 5%, reducing redundant calculations, and using the smoothing characteristics of low complexity to mask slight segmentation marks to facilitate modeling of large-area low-detail areas. Among them, the third complex category is a suburban open space, which generally involves an open square, mainly containing simple geometric bodies, and is of low complexity. Due to the single structure and low detail requirement, the computing resources can be appropriately compressed. The block segmentation loss function is optimized through gradient backpropagation to automatically coordinate the weights of "rendering quality" (detail accuracy) and "training efficiency" (computing resource consumption): The block segmentation loss function is optimized through gradient backpropagation .
[0092] When the rendering quality is insufficient (e.g., the boundary sawtooth is obvious), the overlap rate is automatically increased or the granularity is refined; When the training efficiency is low (e.g., the calculation time is too long), the granularity is automatically expanded or the overlap rate is reduced.
[0093] Through the dynamic balance of the scene complexity corresponding to the category of the target area and the block connection efficiency, the precision and computational efficiency optimization of three-dimensional modeling are realized.
[0094] Through the above method, the complex area is "refined on demand", avoiding resource waste or detail loss caused by global uniform granularity; dynamic control of overlap rate ensures smooth transition between blocks, improving the visual consistency of the overall scene; through scene complexity driven parameter sharing and block size optimization, the modeling quality and training speed are balanced.
[0095] (2) Multi-resolution grid construction is performed for each macro block, and each macro block is divided into coarse and fine layers; Multi-resolution grid construction is performed for each macro block, and coarse and fine layers are constructed; (3) Local block NeRF training is performed for each macro block according to the coarse and fine division results; Local block NeRF training is performed for each macro block according to the coarse and fine division results, and coarse and fine layers are optimized in stages; specifically: For the coarse layer of each macro block, a low-resolution voxel grid is used, multi-view images are input according to the UAV oblique photography data, and coarse-grained color and density are output; For the fine layer of each macro block, a high-resolution grid is used, and the output of the coarse layer is used as initialization, and local details are refined through upsampling or super-resolution network.
[0096] High similarity blocks are identified by SCAC algorithm; high similarity blocks include but are not limited to: continuous commercial building groups, continuous glass curtain wall building groups (material similarity of building materials ) and the like.
[0097] High similarity blocks share base network weights, and the sharing rate is preferably 72%; cross-block radiation differences are suppressed through contrast learning loss function ).
[0098] Among them, 8-node NVIDIA DGX A100 cluster is used for training with mixed precision (FP16+FP32). Asynchronous parameter server (APS) realizes gradient synchronization delay <8ms.
[0099] S121-3-2 fuses and aligns the macro blocks to obtain locally optimized macro blocks; The boundary overlap area of adjacent macro blocks is determined based on the boundary overlap rate; For the boundary overlap area, Poisson fusion or weighted average (weight decays with distance) is used; The relative pose between macro blocks is fine-tuned through feature matching, and the macro blocks are aligned.
[0100] S121-4 globally jointly optimizes all the macro blocks to obtain the macro scene model; Joint all macro blocks, optimize the global consistency and high frequency details of the macro reconstruction object. Specifically, progressive hash encoding (PHE) and radiance transport field consistency constraint (RTFC) are used for distributed radiance field fusion to solve the cross-block radiation discontinuity problem (Delta PSNR<0.3dB).
[0101] The application significantly improves the modeling efficiency while maintaining sub-millimeter geometric accuracy (RMSE=0.07m).
[0102] S122 constructs a mesoscopic scene model based on the mesoscopic scale data; S122-1 pre-processes the mesoscopic scale data to obtain pre-processed mesoscopic scale data; The unmanned aerial vehicle close-range photography data includes: RGB images and multispectral images, which contain the color, texture and geometric structure of the object surface.
[0103] The light intensity data includes: HDRi environment data; The HDRi environment data is used to record the dynamic lighting information of the scene.
[0104] S122-1-1 determines the mesoscopic reconstruction object in the macro reconstruction object according to the modeling requirements; Specifically, the relative position of the mesoscopic reconstruction object in the macro scene model is determined to obtain the correspondence between the mesoscopic scene model and the macro scene model. In an embodiment of the present application, the key feature points of the mesoscopic reconstruction object in the macro scene model are extracted, such as edges, junctions, etc., and the corresponding position of the mesoscopic reconstruction object in the macro scene model is found through a feature matching algorithm (such as SIFT, SURF, etc.).
[0105] S122-1-2 performs feature point matching and defect labeling on the mesoscopic reconstruction object by local point cloud data and unmanned aerial vehicle close-range photography data; Align the calibration parameters of the RGB image and the multispectral image to ensure spatial consistency.
[0106] The multispectral image includes: infrared imaging data and ultraviolet fluorescence imaging data; The infrared imaging data is collected based on the near-infrared band (900-1700nm) to detect the internal structural defects of the building facade; The ultraviolet fluorescence imaging data is collected based on ultraviolet fluorescence imaging (365nm) to identify the surface coating aging area.
[0107] S122-1-3 determining HDRi environment data corresponding to the photographing data of the UAV; Segmenting the building facade (windows, walls, decorative components) using a U-Net network.
[0108] S122-2 extracting macro features and detail features of the meso reconstruction object from the preprocessed meso-scale data; S122-2-1 retrieving macro features of the meso reconstruction object; In an embodiment of the present specification, a HRNet network is used to extract geometric priors (plane, edge distribution) of LiDAR point cloud to generate a macro feature map.
[0109] In another embodiment of the present specification, based on the correspondence between the macro-scale data and the meso-scale data, a macro reconstruction object corresponding to the meso reconstruction object is determined; a macro scene model of the macro reconstruction object is retrieved; Based on the mapping relationship between the meso reconstruction object and the macro reconstruction object, the region of the meso reconstruction object in the macro scene model is determined as the macro geometric skeleton of the meso reconstruction object. Specifically, local point cloud data is obtained, and feature point matching is performed to determine the distribution of the meso reconstruction object in the macro scene model.
[0110] S122-2-2 extracting detail features from the meso-scale data; Performing ESRGAN super-resolution reconstruction based on the close-range aerial photography data of the UAV to extract texture details.
[0111] S122-3 aligning the macro features and the detail features across scales to obtain the meso scene model; Aligning the macro features and the texture details according to a hybrid-scale feature pyramid (HFP), constraining geometric consistency, and converting the fused features into the meso scene model.
[0112] S122-4 enhancing details of the meso scene model; In an embodiment of the present specification, PBR material maps (Albedo / Roughness / Metallic) are generated based on HDRi environment data; Driving the bidirectional reflectance distribution function (BRDF) of complex materials such as metal / glass based on a physically-based rendering (PBR) material library and real-time high-dynamic HDRi environment data for dynamic fitting.
[0113] Specifically, a physical parameter library (such as metal degree, roughness, and base reflectivity) containing materials such as metal and glass is constructed, and synchronous real-time acquisition of ambient light data is performed through a high dynamic range imaging device to generate a multi-level HDR cube map.
[0114] In the preprocessing stage, the optical properties of the material are decomposed into distribution functions, geometric occlusion and Fresnel reflection components in combination with the micro-surface theory, for example, the GGX model is used to describe the influence of surface roughness on light scattering, the Fresnel term calculation is simplified through Schlick approximation, and the radiance distribution in the real-time HDRi environment data is used to pre-integrate the specular reflection intensity, and the Mipmap sampling level in the cube map of the material with different roughness is dynamically adjusted to match the blurred reflection effect.
[0115] To realize dynamic fitting, the system takes the HDRi light source direction and the viewing angle geometric relationship as input, uses the weighted least squares method to iteratively optimize the BRDF kernel function coefficients, preferentially calculates the light contribution weight similar to the current observation angle, simultaneously simulates complex light interaction (such as the attenuation of the transmission path of the glass material in the volume fog) by combining real-time ray tracing technology, and balances the calculation efficiency under the condition of multiple light sources through Monte Carlo integration. On the performance optimization level, the BRDF parameter fitting task is decomposed to Compute Shader asynchronous execution through GPU parallel computing, and a dynamic LOD strategy (such as reducing the HDRi sampling resolution or switching to a simplified Fresnel model) is enabled for distant or low importance objects. Finally, based on the spectral angle mapping or the root mean square error, the consistency of the fitting result and the measured material reflection characteristics is verified, so as to realize high-precision dynamic material performance while ensuring real-time rendering efficiency.
[0116] In an embodiment of the present application, in order to realize cross-domain feature migration, CycleGAN can also be used to map multispectral features to the visible light domain to generate PBR material maps with defect labels; the BRDF fitting network is used to correct the specular reflection error of the metal curtain wall; wherein a spectral consistency constraint term is added in the BRDF fitting, that is, , the specular reflection error rate of the metal material is reduced.
[0117] The cross-scale features of the unmanned aerial vehicle close-range photography data and the local point cloud data are aligned through the CFAL loss function.
[0118] The present application optimizes the facade reconstruction accuracy through the cross-scale feature alignment loss (Cross-scale Feature Alignment Loss, CFAL), so that the mean square error MSE is less than 0.01.
[0119] S123 constructs a micro-scale scene model based on the micro-scale data; S123-1 determines a micro-reconstruction object corresponding to the micro-scale; According to the modeling requirements, the micro-reconstruction object in the meso-reconstruction object is determined. In an embodiment of the present specification, the micro-reconstruction object can be a key element at the micro level, such as fine carvings on the surface of a building, unique decorations in a small landscape, etc.
[0120] The relative position of the micro-reconstruction object in the meso scene model is determined, and the correspondence between the micro scene model and the meso scene model is obtained. In an embodiment of the present specification, key feature points of the micro-reconstruction object in the meso scene model are extracted, such as edges, junctions, etc., and the corresponding position of the micro-reconstruction object in the meso scene model is found through a feature matching algorithm (such as SIFT, SURF, etc.).
[0121] S123-2 performs super-resolution reconstruction on the micro-reconstruction object based on the micro-scale data to obtain radiation field information; In an embodiment of the present specification, a cascaded generative adversarial network (Cascaded GAN) and a micro-surface radiance estimator (Micro-surface Radiance Estimator, MSRE) are used to improve the input image resolution from 4K to 16K (peak signal-to-noise ratio PSNR> 42dB). 16K super-resolution radiance field (SR ratio x 4) is generated.
[0122] In another embodiment of the present specification, a 4K resolution photo is processed by DASR-Net to generate a 16K super-resolution radiance field (SR ratio x 4).
[0123] Specifically, S123-2-1 calculates the initial phase map according to the phase projection sequence; ; wherein, is a four-step phase shift image; is a fringe order correction term.
[0124] S123-2-2 noise suppression is performed on the initial phase map to obtain a denoised phase map; Noise separation is performed based on a variational autoencoder (VAE), ; wherein, is a hyperparameter for balancing the weight of the reconstruction loss and the latent distribution regularization term; and are the encoder and decoder parameters of the variational autoencoder (VAE), respectively, which are learned through end-to-end training; is the initial phase map without denoising processing.
[0125] S123-2-3 combines the denoised phase map with the multi-view stereo image pair to generate a normal map; Estimating micro-surface normal map by MSRE , obtaining a normal map with confidence Mask).
[0126] S123-2-4 obtains a NeRF radiation field optimized by geometric constraints based on the normal map and a geometric loss function; Constructing a geometric loss function: ; wherein, is the normal vector predicted by the NeRF model, generated by the radiation field parameters (such as density, color network weight) of NeRF; is the output of PLFA (micro-surface normal map), sharing parameters with MSRE; PLFA confidence mask (regions with >0.8 take effect).
[0127] The micro-surface normal map output by PLFA is used as the geometric constraint of NeRF, and a NeRF radiation field optimized by geometric constraints is obtained.
[0128] S123-2-5 super-resolution, refinement and detail enhancement of the NeRF radiation field to generate 16K super-resolution radiation field information.
[0129] In an embodiment of the present specification, the resolution of the radiation field is improved using a super-resolution algorithm, the detail features of the micro-object are highlighted using a detail enhancement algorithm, and the geometric structure of the radiation field is optimized through a refinement algorithm. Finally, 16K super-resolution radiation field information is generated, which provides more abundant data support for the construction of micro-scene models.
[0130] S123-3 micro-surface modeling of the micro-reconstructed object; Phase offset light field analysis (PLFA) is used to capture surface fluctuations of 0.05mm; MSRE module generates anisotropic micro-surface normal map.
[0131] S123-4 micro-surface detection of the micro-reconstructed object by multi-spectral defect map; Internal cracks and repair marks are detected by multi-spectral defect map. The multi-spectral defect map has the same imaging method as the aforementioned multi-spectral image.
[0132] S123-5 summarizes the surface modeling data and micro-surface detection data to obtain a micro-scene model; Microsurface modeling data (such as anisotropic microsurface normal maps and surface geometry) and microsurface inspection data (such as internal crack locations and repair trace information) are aggregated and integrated to construct a complete microscopic scene model. This model accurately describes the surface features and internal defects of the microscopically reconstructed object.
[0133] The present invention uses a macroscopic scene model (such as a block) as the base, uses the intersection of the mesoscale model in the macroscale model as the determination point of the positional relationship, and then establishes the positional relationship between the mesoscale model and the microscale model through the intersection of the mesoscale model and the microscale model.
[0134] After the model is built, its macroscale model, mesoscale model, and microscale model must meet the following conditions: (1) All components of the external structure of the macro-scale model and the meso-scale model must be accurate in shape, proportion, orientation, and position, with a precision of decimeters (5-10 cm); the soft and hard edges of the smooth group must be correct; the accuracy of the micro-scale model must reach the centimeter level (1-3 cm), with no specific limit on the number of faces.
[0135] (2) The UV layout of the macro scene model, the meso scene model, and the micro scene model is reasonable, the UV directions are consistent, and the proportions are correct; the UV seams of the exposed models have consistent accuracy at the UV seams; (3) Each scene model corresponds to 6 ultra-clear maps, namely: diffuse map BaseColor, normal map Normal, transparent map Opacity (provided by translucent objects), etc. The size is temporarily set to 2048x2048 and will be changed according to the actual restoration effect. The map resolution is .
[0136] (4) The number of triangles in each scene model depends on the actual restoration effect and must be able to achieve the accuracy level required by the project.
[0137] (5) The same type of materials in each scene model are merged into the same material channel, and the number of material channels for each model is controlled at around 15.
[0138] (6) When collecting data in key areas, try to eliminate the influence of light and shadow effects and collect data on cloudy days or in environments with no obvious light source.
[0139] (7) Any scene model of the full-scale scene needs to be processed individually.
[0140] (8) The objects of 3D scene reconstruction include: 3D model reconstruction of roads, buildings, rivers, green areas, street lights, traffic lights, docks, bus stations, vehicles, ships, etc.; and the objects of reconstruction shall be of no less than 20 categories.
[0141] S2 presents the preset scale scene model to the user; The multi-scale scene model of the urban meta-universe includes a macro scene model, a meso scene model, and a micro scene model.
[0142] The scene model includes a plurality of reconstruction objects. The number of facets of the macro scene model is < the number of facets of the meso scene model is ≈ the number of facets of the micro scene model is >. .
[0143] The preset scale scene model is one of a macro scene model, a meso scene model, and a micro scene model.
[0144] When the scene model is loaded, the preset scale scene model is displayed.
[0145] S3 obtains real-time interaction data between the user and the reconstruction object, and matches an activation strategy of a region block where the reconstruction object is located according to the real-time interaction data. S31 obtains real-time interaction data between the user and the reconstruction object. S311 collects a current viewpoint of the user. In an embodiment of the present specification, a binocular disparity depth estimation network (DPEN) is used; the DPEN is a network model based on deep learning, which can accurately estimate the depth of an object by analyzing the disparity information of binocular images, and then determine the current viewpoint of the user. In another embodiment of the present specification, the current viewpoint is collected according to the 6DoF pose data (120Hz sampling) of the head-mounted device; the 6DoF pose data of the head-mounted device can accurately reflect the position and direction of the user's head, providing a reliable basis for viewpoint determination.
[0146] S312 calculates the Euclidean distance d between the current viewpoint and the surface of each reconstruction object in real time as real-time interaction data.
[0147] The collected current viewpoint information is matched with the reconstruction objects in the scene model in terms of spatial position, the Euclidean distance d between the current viewpoint and the surface of each reconstruction object is calculated in real time, and this distance will be used as subsequent real-time interaction data for matching the activation strategy.
[0148] S32 determines the activation strategy of the region block where the reconstruction object is located according to the real-time interaction data and multi-level loading conditions.
[0149] The multi-level loading conditions include: When the real-time interactive data d>100m, the macro scene model containing the area block where the reconstructed object is located is activated. The threshold of 100m here is set based on a comprehensive consideration of actual application scenarios and user experience. When the user is far away from the reconstructed object, the macro scene model can provide an overall scene overview to meet the user's basic browsing needs.
[0150] When 10m≤real-time interactive data d≤100m, the mesoscopic scene model containing the area block where the reconstructed object is located is activated; within this distance range, the user's demand for scene details increases, and the mesoscopic scene model can display more local details and enhance the user's visual experience.
[0151] When the real-time interactive data d≤10m, the microscopic scene model containing the area block where the reconstructed object is located is activated; when the user approaches the reconstructed object, the microscopic scene model can present the finest details, meeting the user's needs for in-depth exploration of the scene.
[0152] In another embodiment of the present specification, user interaction intention is predicted based on the handle operation trajectory: Specifically, the handle operation trajectory (3D coordinate sequence) is analyzed by the convolutional LSTM network. ); define attention weight ,when , forces the activation of the microscopic scene model containing the reconstructed object.
[0153] S33 adjusts the activation policy based on the system resource status; The trigger mechanism is not only based on user interaction behaviors (such as changes in viewpoint distance or handle operations), but also includes system resource status monitoring and scene dynamic event response.
[0154] In one embodiment of this specification, when the real-time viewpoint approaches a reconstructed object (real-time interactive data d ≤ 10m), the microscopic scene model is automatically loaded and integrated with the surrounding mesoscopic scene models to provide richer scene details and transition effects. Simultaneously, the distant macroscopic model is unloaded to free up system resources and ensure smooth system operation.
[0155] In another embodiment of the present invention, when video memory or memory pressure is detected, the accuracy of the scene model of secondary areas is dynamically reduced or redundant data layers are merged. For example, for areas farther from the user's viewpoint and less worthy of user attention, the texture resolution of the scene model can be lowered or the number of polygons in the model can be reduced, or multiple similar reconstructed objects can be merged to reduce system resource usage.
[0156] S4: when the activation strategy involves scene models of adjacent scales, performing cross-scale fusion on the scene models of adjacent scales; The triggering and implementation of cross-scale fusion can be based on the dynamic coordination and resource scheduling of multi-scale scene models.
[0157] S41 searches for scale transition regions of scene models at adjacent scales; The reconstructed objects in a macroscale model are typically groups of objects, such as an office building. The macro-mesoscale transition region is the boundary between group objects and individual objects, specifically the relationship between a local region composed of multiple reconstructed objects and a single reconstructed object. For example, an office building in a macroscale model takes the form of a rectangular block, with the hollow interior representing the interior area and the outer shell representing the building facade. In this case, the macro-mesoscale transition can be understood as the relationship between a local region composed of multiple rectangular blocks and a single rectangular block, i.e., the process of transitioning from displaying the entire office building area to focusing on a single office building.
[0158] The meso-micro scale transition zone is the connection between the building facade and the interior. While the mesoscopic scene model displays the building's exterior details, the microscopic scene model needs to show the building's interior fine structure. The connection between the two is the scale transition zone.
[0159] S42 performs surface optimization on the scale transition region according to the differentiable distance field; Differentiable distance field , where S is the surface of the scale transition region and σ is the Sigmoid function. The Sigmoid function can map the output of MLP(x) to the interval [0,1], thus normalizing it.
[0160] By optimizing the DDF parameters through gradient descent, the surface in the scale transition area can be made smoother and more natural, and the splicing traces between scene models of different scales can be reduced.
[0161] S43 performs forced radiation field smoothing on the scale transition region based on a radiation gradient consistency constraint; Construct a radiation gradient consistency constraint: ; In the scale transition region The internal forced radiation field is smooth. 、 It is the radiation field of the scene model at adjacent scales (such as the mesoscopic and microscopic scene models).
[0162] is the gradient of the radiation field (usually a three-dimensional vector containing the direction and magnitude of RGB or brightness changes).
[0163] It is the transition area (such as the adjacent area at the junction of the mesoscopic and microscopic models).
[0164] Through the constraint, the change of the radiation field of the adjacent scale scene model in the transition region can be made more smooth, and obvious light and shadow mutation can be avoided.
[0165] The application realizes smooth transition through a differentiable distance field (DDF) and a radiance gradient consistency constraint (RGC) in the scale transition region, and eliminates the visual fragmentation of a traditional scheme.
[0166] In order to improve the consistency of cross-scale light and shadow, the application further comprises: adopting a hierarchical feature transfer strategy to map microscopic physical properties to a macroscopic behavior model, balancing the contribution weights of different scales through a dynamic weighting algorithm, and realizing cross-scale light and shadow consistency by using a ray tracing technology.
[0167] The application structures a target region into independent blocks based on spatial grid division, takes a block where a reconstructed object is located as a current block, loads a microscopic scene model for the current block, and combines with background asynchronous preloading of adjacent blocks of the current block to ensure switching fluency.
[0168] S5 shows the rendered virtual scene to the user after cross-scale fusion.
[0169] S51 replaces the scene model in combination with resource dynamic scheduling; In order to avoid lag and optimize system resources, the application further comprises: dynamically allocating computing resources through a priority queue: Specifically, according to real-time state data, a high-frequency interaction region is determined; the high-frequency interaction region mainly includes a line-of-sight focus region, that is, a region that a current viewpoint of the user focuses on.
[0170] Full-detail rendering is reserved for the high-frequency interaction region to ensure that the user can see the clearest scene details; and a simplified proxy model is used for a low-frequency interaction region (an edge region and a static background region) to reduce the occupation of computing resources.
[0171] Time series analysis is introduced to predict user behavior trends and predict a medium-frequency interaction region; model resources that can be called are scheduled in advance, and finally efficient rendering and physical simulation of a global scene are realized in multi-scale dynamic balance.
[0172] In order to improve rendering effect and user experience, the application further comprises: S52 real-time dynamic weather simulation.
[0173] S521 collects real-time meteorological data and pre-processes the real-time meteorological data to obtain a weather parameter set; S521-1 acquires real-time meteorological data; Specifically, the original data of precipitation type, intensity, wind speed, humidity, etc. are collected through weather API or sensors; S521-2 After extracting key parameters through JSON parsing, a sliding window algorithm is used to filter transient abnormal values (such as sudden wind speed noise); S521-3 Missing data caused by sensor packet loss is completed through linear interpolation or Kalman filtering, and finally a standardized weather parameter set is output.
[0174] S522 Weather parameter configuration based on the weather parameter set; S522-1 Through the stream processing engine (such as Flink), the data stream is processed by frame rate slicing, and the precipitation intensity is mapped to the particle emission rate (for example, the number of raindrops per second = intensity × 1000), and the wind speed and direction are converted to particle motion vector (for example, the number of raindrops per second = intensity × 1000), and the wind speed and direction are converted to particle motion vector (for example, the number of raindrops per second = intensity × 1000). ); S522-2 Calculate the fog concentration parameter according to the visibility and humidity, and generate a dynamic configuration containing particle system parameters, physical field vector and fog effect density.
[0175] S523 Real-time dynamic weather simulation based on the weather parameter; The particle system dynamically adjusts the emission area according to the camera frustum range, uses GPU Instancing to generate raindrop or snowflake particles in batches, simulates the superposition effect of gravity and wind resistance through Newtonian mechanics, and combines ray detection to trigger ground splashing.
[0176] Based on the screen space volume fog technology, the basic fog concentration is calculated by an exponential decay formula, the diffusion effect of wind field on fog is simulated by a simplified CFD model, and the calculation efficiency is optimized by using adaptive ray step (0.1m for close-up / 1m for long shot) and early termination strategy.
[0177] Through the post-processing pipeline, the particle layer and the fog effect layer are superimposed to the scene, motion blur is added for high-speed particles, and LOD grading (such as reducing the number of particles in the distance by 75%) or turning off secondary special effects (such as raindrop splashing) is dynamically enabled according to the real-time frame time.
[0178] Through the double buffering mechanism, the data analysis and rendering threads are isolated, and the physical simulation is executed asynchronously by Compute Shader, ensuring that the end-to-end delay is stable and controlled within 40ms, that is, the rendering delay of rain and fog effect is <40ms.
[0179] For environmental mutations (such as weather conversion events), the system synchronously updates the model parameters of the associated scale (such as the linkage of macroscopic terrain hydrology and microscopic particle effect), to ensure the coordination and consistency of different scale scene models when the weather changes.
[0180] Figure 3A structural schematic diagram of a three-dimensional reconstruction system of an urban meta-universe is provided for an embodiment of the present specification, comprising: a scene construction module 310, configured to construct a multi-scale scene model; the scene model comprises a plurality of reconstruction objects; an initialization module 320, configured to show the user a preset scale scene model; a strategy matching module 330, configured to obtain real-time interaction data between the user and the reconstruction object; and match an activation strategy for a region block where the reconstruction object is located according to the real-time interaction data; a scale fusion module 340, configured to perform cross-scale fusion on the scene model of the adjacent scale when the activation strategy involves the scene model of the adjacent scale; a scene rendering module 350, configured to show the user a rendered virtual scene after cross-scale fusion.
[0181] Optionally, the scene construction module 310 comprises: an acquisition sub-module, configured to acquire basic three-dimensional data of a target region; the basic three-dimensional data comprises one or more of macro-scale data, meso-scale data and micro-scale data; a reconstruction sub-module, configured to input the basic three-dimensional data into a three-dimensional reconstruction model to construct a plurality of original scene models; The reconstruction sub-module comprises: a macro-reconstruction unit, configured to construct a macro scene model based on the macro-scale data; a meso-reconstruction unit, configured to construct a meso scene model based on the meso-scale data; a micro-reconstruction unit, configured to construct a micro scene model based on the micro-scale data.
[0182] Optionally, the macro-reconstruction unit comprises: a data extraction sub-unit, configured to extract macro-reconstruction objects from the target region according to satellite remote sensing data; a block segmentation sub-unit, configured to segment the macro-reconstruction objects into a plurality of macro blocks in combination with LiDAR point cloud data and unmanned aerial vehicle oblique photography data; a local optimization sub-unit, configured to locally optimize each macro block to obtain a locally optimized macro block; a joint optimization sub-unit, configured to globally jointly optimize all the macro-reconstruction objects to obtain the macro scene model.
[0183] Optionally, the meso-reconstruction unit comprises: a preprocessing subunit, configured to preprocess the mesoscale data to obtain preprocessed mesoscale data; a detail extraction subunit, configured to extract macro features and detail features of a mesoscopic reconstructed object from the preprocessed mesoscopic scale data; an alignment subunit, configured to align the macro features with the detail features across scales to obtain the mesoscopic scene model; The detail enhancement subunit is used to perform detail enhancement on the mesoscopic scene model.
[0184] Optionally, the microscopic reconstruction unit includes: an object confirmation subunit, configured to determine the microscopic reconstruction object corresponding to the microscopic scale; a super-resolution reconstruction subunit, configured to perform super-resolution reconstruction on the microscopic reconstruction object based on the microscopic scale data to obtain radiation field information; a surface modeling subunit, configured to perform micro-surface modeling on the microscopic reconstructed object to obtain surface modeling data; a surface detection subunit, configured to perform micro-surface detection on the microscopic reconstructed object using a multi-spectral defect map to obtain micro-surface detection data; The summarizing subunit is used to summarize the surface modeling data and the micro-surface detection data to obtain the microscopic scene model.
[0185] Optionally, the scale fusion module 340 includes: A region search submodule, configured to search for scale transition regions of scene models of adjacent scales; A first optimization submodule is configured to perform surface optimization on the scale transition region according to a differentiable distance field; The second optimization submodule is used to force radiation field smoothing in the scale transition region based on radiation gradient consistency constraints.
[0186] The functions of the system of the embodiment of the present invention have been described in the above method embodiment. Therefore, for details not fully described in this embodiment, please refer to the relevant descriptions in the above embodiment and will not be repeated here.
[0187] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of this specification, which includes: a memory 401 and a processor 402, the memory 401 is used to store computer-executable instructions, and when the computer-executable instructions are executed by the processor 402, they can implement the steps of the above method embodiment.
[0188] Figure 5A structural diagram of a computer readable storage medium provided by an embodiment of the present specification is shown in FIG. 5. The computer readable storage medium 500 stores one or more computer programs, which, when executed by a processor, can implement the steps of the above method embodiments.
[0189] An embodiment of the present specification also provides a computer program product, which includes computer programs / computer executable instructions, which, when executed by a processor, can implement the steps of the above method embodiments.
[0190] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by a computer program instructing related hardware, and the computer program, when executed, can include the processes of the above embodiments of each method.
[0191] Obviously, those of ordinary skill in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the claims of the present application and their equivalents, the present application also intends to include these modifications and variations.
Claims
1. A three-dimensional reconstruction method of an urban metaverse, characterized in that: include: Constructing a multi-scale scene model; the scene model includes a plurality of reconstructed objects; Displaying a scene model of a preset scale to the user; Acquire real-time interaction data between the user and the reconstructed object; match an activation strategy for the area block where the reconstructed object is located according to the real-time interaction data; When the activation strategy involves scene models of adjacent scales, cross-scale fusion is performed on the scene models of adjacent scales; After cross-scale fusion, the rendered virtual scene is presented to the user.
2. The three-dimensional reconstruction method of the urban metaverse according to claim 1, characterized in that: The multi-scale scene model is constructed, comprising: Acquire basic three-dimensional data of the target area; the basic three-dimensional data includes: one or more of macro-scale data, meso-scale data, and micro-scale data; The basic three-dimensional data is input into a three-dimensional reconstruction model to construct several original scene models; wherein, a macroscopic scene model is constructed based on the macroscopic scale data; a mesoscopic scene model is constructed based on the mesoscopic scale data; and a microscopic scene model is constructed based on the microscopic scale data.
3. The three-dimensional reconstruction method of the urban metaverse according to claim 2, characterized in that: The constructing of a macroscopic scene model based on the macroscopic scale data includes: Performing macro-reconstruction object extraction on the target area according to satellite remote sensing data; Combining LiDAR point cloud data and UAV oblique photography data to segment the macro-reconstructed object into blocks to obtain a plurality of macro blocks; Performing local optimization on each of the macro blocks to obtain a locally optimized macro block; All the macroscopic reconstruction objects are globally jointly optimized to obtain the macroscopic scene model.
4. The three-dimensional reconstruction method of the urban metaverse according to claim 2, characterized in that: The constructing of a mesoscopic scene model based on the mesoscopic-scale data includes: Preprocessing the mesoscale data to obtain preprocessed mesoscale data; extracting macro features and detail features of a mesoscopic reconstructed object from the preprocessed mesoscopic-scale data; Aligning the macro features with the detail features across scales to obtain the mesoscopic scene model; The mesoscopic scene model is enhanced in detail.
5. The three-dimensional reconstruction method of the urban metaverse according to claim 2, characterized in that: The constructing of a microscopic scene model based on the microscopic scale data includes: determining a microscopic reconstruction object corresponding to the microscopic scale; Performing super-resolution reconstruction on the microscopic reconstruction object based on the microscopic scale data to obtain radiation field information; Performing micro-surface modeling on the microscopic reconstructed object to obtain surface modeling data; Performing micro-surface detection on the microscopic reconstructed object using a multispectral defect image to obtain micro-surface detection data; The surface modeling data and the micro-surface detection data are aggregated to obtain the microscopic scene model.
6. The three-dimensional reconstruction method of the urban metaverse according to claim 1, characterized in that: When the activation strategy involves scene models of adjacent scales, cross-scale fusion of the scene models of adjacent scales includes: Finding scale transition regions of scene models at adjacent scales; performing surface optimization on the scale transition region according to a differentiable distance field; The radiation field in the scale transition region is forced to be smoothed based on the radiation gradient consistency constraint.
7. A three-dimensional reconstruction system for urban metaverse, characterized in that: include: A scene construction module, configured to construct a multi-scale scene model comprising a plurality of reconstructed objects; An initialization module, configured to display a scene model of a preset scale to the user; A strategy matching module is used to obtain real-time interaction data between the user and the reconstructed object; and match an activation strategy for the area block where the reconstructed object is located according to the real-time interaction data; a scale fusion module, configured to perform cross-scale fusion on the scene models of adjacent scales when the activation strategy involves scene models of adjacent scales; The scene rendering module is used to display the rendered virtual scene to the user after cross-scale fusion.
8. A computer device, characterized in that: The computer equipment includes: processor; and, A memory storing computer executable instructions, which, when executed, cause the processor to perform the method of any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs / instructions, and when the one or more programs / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Cited By
Remote sensing change detection method and system based on element cosmic space and pellet driving
CN121746950A
Method and system for generating constructable landscape based on three-dimensional modeling
CN122197353A