A rapid and non-damage digital reconstruction method and device for cultural relics and ancient buildings

By improving the quality of image sets and depth data of cultural relics and ancient buildings and identifying feature regions, and combining the reconstruction model for three-dimensional reconstruction and scale calibration, the problem of difficulty in quickly generating high-precision three-dimensional models in existing technologies has been solved, and efficient and accurate digital reconstruction has been achieved.

CN122289571APending Publication Date: 2026-06-26NINGBO DIGITAL TWIN (EASTERN UNIV OF TECH) RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610752032.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In the existing technology, it is difficult to quickly generate high-precision three-dimensional digital models of cultural relics and ancient buildings without damaging them, which makes it difficult to optimize the reconstruction results in terms of speed, accuracy and direct usability.

Method used

By acquiring image sets and depth data, data quality is improved and pose is calculated. Different feature regions are identified and fusion weights are assigned. Combined with the reconstruction model, 3D reconstruction and scale calibration are performed, including denoising, hidden point removal, depth completion, motion recovery of structure, feature extraction and matching, bundle adjustment optimization and other processing steps.

Benefits of technology

It achieves a balance between processing efficiency, reconstruction accuracy, and model size availability with limited data input, improving the overall effect of digital reconstruction results and ensuring high integrity and accuracy of complex cultural relic surfaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289571A_ABST
    Figure CN122289571A_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for rapid and non-destructive digital reconstruction of cultural relics and ancient buildings. The reconstruction method includes: acquiring an image set of the target object and corresponding depth data; performing quality improvement processing on the depth data and performing pose calculation on the image set to obtain pose data; identifying different feature regions on the surface of the target object based on the image features of the image set, and assigning corresponding fusion weights to the processed depth data according to the characteristics of each feature region; using the image set, the weighted depth data, and the pose data as input data, performing three-dimensional reconstruction of the target object through a reconstruction model, and completing scale calibration during the reconstruction process; and outputting a scale-calibrated three-dimensional model of the target object. The technical problem solved by this invention is that limitations in the performance indicators of object acquisition and processing make it difficult to coordinate and optimize the speed, accuracy, and direct usability of the overall reconstruction method, resulting in less than ideal digital reconstruction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional reconstruction technology, and more specifically, to a method and apparatus for rapid and non-destructive digital reconstruction of cultural relics and ancient buildings. Background Technology

[0002] Digital reconstruction of cultural relics and ancient buildings is a key means of emergency protection of cultural heritage. It is necessary to quickly generate high-precision three-dimensional digital models without damaging the relics or disturbing the surrounding environment, for use in scenarios such as cultural relic restoration, disease monitoring, emergency protection, and academic research.

[0003] However, the relevant technologies have at least one of the following problems: due to limitations in the performance indicators of the collected and processed objects, it is difficult to optimize the speed, accuracy and direct usability of the overall reconstruction method, resulting in less than ideal digital reconstruction results. Summary of the Invention

[0004] The technical problem solved by this invention is that due to limitations in the performance indicators of the objects collected and processed, it is difficult to optimize the speed, accuracy and direct usability of the overall reconstruction method in a coordinated manner, resulting in less than ideal digital reconstruction results.

[0005] To address the aforementioned problems, this invention provides a method for rapid and non-destructive digital reconstruction of cultural relics and ancient buildings, comprising: acquiring an image set of the target object and corresponding depth data; performing quality enhancement processing on the depth data and performing pose calculation on the image set to obtain pose data; identifying different feature regions on the surface of the target object based on the image features of the image set, and assigning corresponding fusion weights to the processed depth data according to the characteristics of each feature region; using the image set, the weighted depth data, and the pose data as input data, performing three-dimensional reconstruction of the target object through a reconstruction model, and completing scale calibration during the reconstruction process; and outputting a scale-calibrated three-dimensional model of the target object.

[0006] Compared with existing technologies, the technical effects achieved by adopting this technical solution are as follows: This application constructs a processing flow from data acquisition, data processing, weight allocation, model reconstruction and calibration to result output, ensuring that with limited data input, it can balance processing efficiency, reconstruction accuracy, and the direct usability of the output model size, ultimately improving the overall effect of digital reconstruction results.

[0007] In one embodiment of the present invention, the process of improving the quality of depth data and calculating the pose of an image set to obtain pose data includes: performing a first processing on the depth data; the first processing includes at least denoising and hidden point removal; performing depth completion on the depth data after the first processing; and performing feature extraction and matching based on the image set using a motion reconstruction method to solve for and obtain pose data.

[0008] Compared with existing technologies, the technical effects achieved by this solution are as follows: improving the cleanliness of depth data through denoising and hidden point removal, enhancing data integrity through depth completion, and obtaining accurate camera pose through motion reconstruction methods, thereby providing a high-quality and fusionable input data foundation for all subsequent processing stages.

[0009] In one embodiment of the present invention, feature extraction and matching based on an image set using a motion reconstruction structure method are performed to obtain pose data. This includes: extracting and matching image features from the image set and extracting local features; removing unmatched features from the feature matching results to retain reliable inliers; based on the inliers, performing initial pose calculation and relative pose calculation of the image acquisition device using multi-view pose extension and a preset algorithm; and optimizing the calculated pose data using bundle adjustment. During the pose data calculation and optimization process, closed-loop detection and epipolar constraints are used to enhance the pose.

[0010] Compared with existing technologies, the technical effects achieved by this solution are as follows: the accuracy and stability of pose calculation are improved through mismatch elimination and global optimization using the bundle adjustment method, and the pose is further corrected and enhanced by closed-loop detection and epipolar constraints, thereby constructing a high-precision spatial geometric framework for 3D reconstruction and effectively avoiding structural distortion of the model.

[0011] In one embodiment of the present invention, after depth completion of the depth data after the first processing, the reconstruction method further includes: acquiring a depth image after depth completion; the depth image contains the depth value of each pixel; for each valid depth value in the depth image, a preset sampling depth interval is determined; during three-dimensional reconstruction, the distribution of spatial sampling points in the three-dimensional space is constrained according to the sampling depth interval, so that the spatial sampling points are restricted to the sampling depth interval corresponding to each depth value, and sampling data located outside the sampling depth interval is excluded.

[0012] Compared with existing technologies, the technical effects achieved by this solution are as follows: by using reliable depth data as prior knowledge, the calculated sampling points are focused on the narrow spatial ranges most likely to appear on the surface of the artifact, thereby significantly reducing redundant calculations while ensuring visual quality. This is a key efficiency optimization method for achieving the goal of rapid reconstruction.

[0013] In one embodiment of the present invention, different feature regions on the surface of a target object are identified based on the image features of an image set, including: extracting texture features, edge features, and grayscale statistical features from the image set; and segmenting the surface of the target object into decorative areas, diseased areas, and weak texture areas based on the texture features, edge features, and grayscale statistical features.

[0014] Compared with existing technologies, the technical effects achieved by this solution are as follows: by integrating three types of visual features—texture, edge, and grayscale statistics—it enables automated and quantitative classification of surface areas, providing a scientific and repeatable objective basis for the subsequent implementation of differentiated data processing strategies.

[0015] In one embodiment of the present invention, according to the characteristics of each feature region, corresponding fusion weights are assigned to the processed depth data, including: increasing the fusion weight of the depth data for the identified weak texture areas; and decreasing the fusion weight of the depth data for the identified decorative areas and disease areas.

[0016] Compared with existing technologies, the technical effect achieved by this solution is as follows: through the adaptive fusion strategy of weak texture area information depth and detail area information image, the complementarity and enhancement of multi-source data at the feature level are realized, thereby improving the overall integrity and accuracy of the reconstruction of complex cultural relic surfaces and areas containing both flat and fine features.

[0017] In one embodiment of the present invention, an image set, weighted depth data, and pose data are used as input data. A reconstruction model is used to reconstruct the target object in three dimensions, and scale calibration is performed during the reconstruction process. This includes: setting independently learnable scale parameters in the reconstruction model; optimizing the scale parameters by combining the historical process size patterns of the target object; and controlling the scale parameters to converge within a predetermined training round during the training process of the reconstruction model, so as to achieve scale alignment between the three-dimensional model and the real target object.

[0018] Compared with existing technologies, the technical effects achieved by this solution are as follows: by embedding individually learnable scale parameters into the model and optimizing it using prior knowledge such as historical process size patterns, the reconstructed model is quickly aligned with the actual physical size.

[0019] In one embodiment of the present invention, 3D reconstruction of a target object is performed using a reconstruction model, comprising: adopting a corresponding training strategy according to the training stage of the reconstruction model; if the reconstruction model is in the early stage, a first training strategy is executed; the first training strategy includes simultaneously optimizing image reconstruction loss and depth perception loss, and optimizing the scale parameters in the reconstruction model; if the reconstruction model is in the middle stage, a second training strategy is executed; the second training strategy includes initiating a sampling strategy and focusing on model fitting to the core region of the target object; if the reconstruction model is in the late stage, a third training strategy is executed; the third training strategy includes fixing the scale parameters and adjusting the reconstruction model until convergence.

[0020] Compared with existing technologies, the technical effects achieved by this solution are as follows: Through a progressive strategy of jointly optimizing the foundation and scale in the early stage, focusing on sampling to accelerate fitting in the middle stage, and fine-tuning with a fixed scale in the later stage, the training process is scientifically managed in stages and resources are allocated. This guides the model to converge to a better state more efficiently within a limited time, thus systematically improving training efficiency and final model quality.

[0021] In one embodiment of the present invention, the reconstruction method further includes: constructing a loss function for optimizing the reconstruction model, wherein the loss function includes at least image color reconstruction loss and depth perception loss; and iteratively optimizing the reconstruction model by minimizing the loss function.

[0022] Compared with existing technologies, the technical effects achieved by this solution are as follows: by clearly specifying that optimization must simultaneously take into account both image color reconstruction loss and depth perception loss, it provides a balanced and clear supervision signal for model learning, ensuring the unity of color fidelity and geometric accuracy of the final 3D model from the root of the optimization algorithm.

[0023] In one embodiment of the present invention, a rapid and non-destructive digital reconstruction device for cultural relics and ancient buildings is also provided, capable of implementing any of the reconstruction methods described above. The reconstruction device includes: a data acquisition module for acquiring an image set of the target object and corresponding depth data; a data processing module for improving the quality of the depth data and performing pose calculation on the image set to obtain pose data; a feature fusion module for identifying different feature regions on the surface of the target object based on the image features of the image set, and assigning corresponding fusion weights to the processed depth data according to the characteristics of each feature region; a reconstruction and calibration module for using the image set, weighted depth data, and pose data as inputs, performing three-dimensional reconstruction of the target object through a reconstruction model, and completing scale calibration during the reconstruction process; and an output module for outputting a scale-calibrated three-dimensional model of the target object.

[0024] Compared with existing technologies, the technical effects achieved by adopting this technical solution are: to achieve any of the technical effects described above, which will not be elaborated further here.

[0025] By adopting the technical solution of the present invention, the following technical effects can be achieved: (1) This application constructs a processing flow from data acquisition, data processing, weight allocation, model reconstruction and calibration to result output, which ensures that under limited data input, it can take into account processing efficiency, reconstruction accuracy and direct usability of output model size, and ultimately improve the overall effect of digital reconstruction results. (2) By using the adaptive fusion strategy of weak texture area information depth and detail area information image, the complementarity and enhancement of multi-source data at the feature level are realized, thereby improving the overall integrity and accuracy of the reconstruction of complex cultural relics surface and areas containing flat and fine areas. (3) Through a progressive strategy of joint optimization of the foundation and scale in the early stage, focused sampling to accelerate fitting in the middle stage, and fine-tuning of fixed scale in the later stage, the training process was scientifically managed in stages and resources were allocated. This guided the model to converge to a better state more efficiently within a limited time, and systematically improved the training efficiency and the quality of the final model. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings to be used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 A flowchart illustrating a method for rapid and non-destructive digital reconstruction of cultural relics and ancient buildings provided in this embodiment of the invention; Figure 2 A module diagram of a rapid and non-destructive digital reconstruction device for cultural relics and ancient buildings provided in an embodiment of the present invention.

[0027] Explanation of reference numerals in the attached figures: 100. Reconstruction device; 10. Data acquisition module; 20. Data processing module; 30. Feature fusion module; 40. Reconstruction and calibration module; 50. Output module. Detailed Implementation

[0028] Embodiments of the present invention will now be described in detail. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0029] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a link, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0030] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0031] See Figure 1 , Figure 1 A flowchart of a method for rapid and non-destructive digital reconstruction of cultural relics and ancient buildings provided in this invention; a method for rapid and non-destructive digital reconstruction of cultural relics and ancient buildings includes: S1: Obtain the image set of the target object and its corresponding depth data; S2: Perform quality enhancement processing on the depth data and perform pose calculation on the image set to obtain pose data; S3: Based on the image features of the image set, identify different feature regions on the surface of the target object, and assign corresponding fusion weights to the processed depth data according to the characteristics of each feature region; S4: Using image sets, weighted depth data, and pose data as input data, the target object is reconstructed in three dimensions through a reconstruction model, and scale calibration is completed during the reconstruction process. S5: Outputs the scale-calibrated 3D model of the target object.

[0032] Preferably, a lightweight, non-contact acquisition device is used to acquire the image set and depth data of the target object.

[0033] Alternatively, lightweight non-contact data acquisition devices include: high-definition cameras, LiDAR, etc.

[0034] Specifically, operators use non-contact acquisition devices to collect images of the core structure of the target object and its corresponding depth data while maintaining a safe distance from the artifact. Subsequently, a data processing module performs quality enhancement processing on the depth data, including denoising, hidden point removal, and depth completion. Simultaneously, feature extraction and matching are performed on the image set to obtain accurate pose data. Next, a precise depth fusion stage is entered. Based on the texture, edge, and other features of the image set, different feature regions are automatically identified, and corresponding fusion weights are assigned to the processed depth data according to the characteristics of each region. Then, using the aforementioned image set, weighted depth data, and pose data as input, a reconstruction model is used for 3D reconstruction. During training, this model optimizes an internally learnable scale parameter to achieve scale calibration with the real artifact. Finally, a 3D digital model with realistic scale is output.

[0035] Furthermore, the depth data undergoes quality enhancement processing, and the image set is used to calculate pose data, including: The depth data undergoes a first processing step; this first processing step includes at least denoising and hidden point removal. Depth completion is performed on the depth data after the first processing; and feature extraction and matching are performed on the image set using the motion reconstruction method to obtain pose data.

[0036] Preferably, denoising involves using a statistical outlier removal filter to remove floating objects and sparse outliers in the depth point cloud.

[0037] Preferably, hidden point removal involves projecting the point cloud based on the camera pose and arranging it in ascending order of depth, removing the more distant points around each closer point, thereby generating an initial depth map that avoids perspective errors.

[0038] Preferably, the motion recovery structure method is the SfM (Structure from Motion) method.

[0039] To reduce the impact of environmental features such as lighting and shadows on depth data, filtering, hidden point removal, and depth completion algorithms were added to improve the quality of the acquired depth data. These algorithms filter out environmental interference points such as tourists or pets, as well as moving objects. Specifically, denoising is used to filter out environmental interference points and moving objects in the depth data. Since the point cloud acquired by LiDAR contains a large number of noise points and holes, the depth data is denoised by using a statistical outlier removal filter and a low-pass filter to remove floating objects and sparse outliers from the original point cloud. This process can effectively reduce erroneous values ​​in the generated depth map.

[0040] Furthermore, hidden point removal is used to address perspective errors during point cloud projection. If a point cloud is directly projected into a depth map, perspective distortion occurs, revealing occluded areas that shouldn't be visible. This is because, although the point cloud has a high density, when projected onto the image, points in front cannot completely obscure points behind. Therefore, before generating the depth map, hidden point removal is performed on the point cloud based on the corresponding camera pose before projection; that is, hidden point removal is performed on the depth data. Specifically, based on the acquired point cloud and camera pose information, projection yields a depth map corresponding to the pose. Points behind the camera and points outside the image range are filtered out. The points are sorted in ascending order of depth; for each pixel location, only the point with the smallest depth value is retained, and points with larger depth values ​​surrounding each smaller point are removed, using closer points to cover farther points, thus obtaining a correct initial depth map.

[0041] Building upon this, depth completion is used to fill holes in the depth map, improving data integrity. Since the depth map processed above is relatively sparse and many areas fail to display the true depth values, a depth completion method is introduced and improved to complete the generated depth map. Specifically: distance transformation interpolation is used to fill small holes, morphological closing operations are used to fill large holes, and guided filtering is used for smoothing. Crucially, the depth completion result is applied only to points with a depth value of 0 in the depth map after hidden point removal, ensuring that the true depth values ​​acquired by the LiDAR remain unchanged.

[0042] It should be noted that depth completion only covers pixels marked as holes (such as those with a depth value of 0) in the depth map after hidden point removal, while retaining the true and effective depth values ​​collected by the LiDAR, thus ensuring the accuracy of the original measurement data.

[0043] The purpose of using the motion recovery structure method to extract and match features from the image set to obtain pose data is to automatically recover the spatial position and pose parameters of each image at the time of shooting from the unordered image set, providing a set framework for the 3D reconstruction of the target object.

[0044] Furthermore, based on the image set, feature extraction and matching are performed using the structure-of-motion method to obtain pose data, including: Image feature extraction and matching are performed on the image set, and local features are extracted; Unmatched features are removed from the feature matching results to retain reliable inliers; Based on interior points, the initial solution of the pose of the image acquisition device and the calculation of the relative pose are performed through multi-view pose extension and preset algorithms. The calculated pose data were optimized using the bundle adjustment method; In the process of solving and optimizing pose data, closed-loop detection and epipolar constraints are used to enhance the pose.

[0045] Specifically, using the SfM method, image feature extraction and matching are first performed on the image set. Local features are extracted primarily using the SIFT (Scale-Invariant Feature Transform) method, and feature point matching is performed using the FLANN (Fast Library for Approximate Nearest Neighbors) algorithm. Subsequently, mismatches are removed from the feature matching results, specifically using the RANSAC (Random Sample Consensus) algorithm to eliminate incorrect matching pairs and retain reliable inliers. Then, based on these inliers, the initial pose of the image acquisition device and the relative pose are calculated using multi-view pose extension and the PnP (Perspective-n-Point) algorithm. Next, the calculated pose data is globally optimized using bundle adjustment to improve overall accuracy. Furthermore, throughout the entire pose data solution and optimization process, loop closure detection and epipolar constraints are employed to further enhance the pose and ensure spatial consistency.

[0046] The image feature extraction and matching process involves using the SIFT method to extract key points from each artifact image. Subsequently, the FLANN algorithm is used to match the key points of all images, establishing a preliminary correspondence between feature points in image pairs.

[0047] Next, mismatch removal is performed to obtain reliable interior points. Since the initial matching contains a large number of errors, the RANSAC algorithm is used, combined with the epipolar geometric constraint model for iterative estimation, thereby robustly removing erroneous matching pairs and filtering out the geometrically consistent and reliable set of feature point correspondences (i.e., interior points).

[0048] Then, the initial camera pose is solved. Based on the aforementioned interior points, the relative camera pose between adjacent images is solved using the PnP algorithm in multi-view geometry. An incremental reconstruction strategy is adopted, starting from the initial image pairs, and new images are gradually registered to the already generated parts of the scene.

[0049] Furthermore, global optimization is performed; using bundle adjustment, all recovered camera pose parameters and 3D point cloud coordinates are optimized globally and nonlinearly. By minimizing the reprojection error of feature points, the overall consistency and accuracy of the camera motion trajectory and scene geometry are improved.

[0050] Finally, pose enhancement is implemented. During the optimization process, closed loops are detected in the image sequence, i.e., the same scene is seen again in the first, last, or middle images, introducing loop closure constraints to effectively correct accumulated errors. Simultaneously, epipolar geometry is used as an additional constraint to verify and enhance the consistency of feature matching and pose estimation results, ensuring high accuracy and robustness of the pose data in complex cultural relic scenes.

[0051] Furthermore, after performing depth completion on the depth data after the first processing, the reconstruction method also includes: Obtain the depth image after depth completion; the depth image contains the depth value of each pixel. For each valid depth value in the depth image, a preset sampling depth range is determined; During 3D reconstruction, the distribution of spatial sampling points in 3D space is constrained based on the sampling depth range, so that the spatial sampling points are limited to the sampling depth range corresponding to each depth value, and sampling data located outside the sampling depth range are excluded.

[0052] Specifically, to handle redundant data and improve efficiency, after obtaining the processed depth image, depth information is used to constrain the sampling depth interval, limiting it to a range near the depth value. Data far from the depth information is discarded. This significantly reduces the number of invalid sampling points while ensuring rendering quality, effectively accelerating the convergence speed of the network model, improving the quality of the predicted image and depth map, and alleviating the problem of floating objects in the rendered scene. Specifically, a depth-completed depth image is obtained, containing the depth values ​​of each pixel. For each valid depth value in the depth image, a preset sampling depth interval is determined centered on that value. This interval is defined by the near-plane distance and the far-plane distance; Formula 1 is used to calculate the near-plane distance (Near), and Formula 2 is used to calculate the far-plane distance (Far). Formula 1: ; Formula 2: ; in, θ is the depth value of the input depth image; θ is a configurable sampling range constraint value.

[0053] During 3D reconstruction, the distribution of sampling points in 3D space is constrained based on this sampling depth range, so that the sampling points are restricted to the range corresponding to each depth value. This actively excludes sampling data that is outside the range and is considered redundant, thereby achieving focused calculation and accelerating convergence.

[0054] Furthermore, based on the image features of the image set, different feature regions on the surface of the target object are identified, including: Extract texture features, edge features, and grayscale statistical features from the image set; Based on texture features, edge features, and grayscale statistical features, the surface of the target object is divided into decorative areas, diseased areas, and weak texture areas.

[0055] Among them, the carved and painted decorative areas have complex and highly repetitive patterns on their surfaces; therefore, in the image, they are characterized by high texture complexity (high entropy value), dense edges and strong regularity of arrangement. At the same time, due to the concavity or color changes, their local gray-scale variance and gradient amplitude are also relatively large.

[0056] Among them, diseased areas such as cracks and weathering are characterized by abnormal changes in local structure. In images, this manifests as irregular textures, abrupt changes, and although the edges may be sparse, their direction is chaotic. Furthermore, significant local gray-scale abrupt changes or high gradient variance will appear at the diseased areas.

[0057] Among them, areas with weak texture, such as smooth brick and stone walls and floors, have a uniform surface material and lack significant details. Therefore, in images, they appear as having a simple and uniform texture (low entropy), very little or no edge information, and very low overall gray-level variance and gradient magnitude.

[0058] Based on the above, when performing region segmentation, the LBP texture entropy value, Canny edge density and orientation consistency, and gray-level statistical variance or gradient magnitude mean of each local region in the image are calculated, and thresholds matching the above characteristics are set for these three types of features respectively. Then, based on these extracted features, the surface of the target object is automatically segmented into different categories by setting multi-dimensional thresholds for texture, edge density and gray-level variance for comparison and judgment, namely, fine decorative areas (such as carving and painting), diseased areas (such as cracks and weathering) and weak texture areas (such as flat walls and floors).

[0059] Furthermore, based on the characteristics of each feature region, corresponding fusion weights are assigned to the processed depth data, including: For the identified weak texture areas, increase the fusion weight of depth data; For the identified decorative and diseased areas, the fusion weight of the depth data is reduced.

[0060] After identifying different feature regions, a differentiated strategy is adopted when assigning different weights to depth data. For identified weak texture regions, due to their weak image feature information, the fusion weight of the depth data assigned to these regions is increased, making the reconstruction process rely more on the geometric constraints of depth information in these regions. Conversely, for identified decorative and diseased areas, which have rich image features, the fusion weight of the depth data assigned to them is reduced, making the reconstruction process rely more on the color and texture information of the image to restore details in these regions, thereby achieving adaptive fusion.

[0061] For example, the Canny operator can be used for edge feature extraction, extracting feature images from visible light images. For regions rich in feature information, the depth loss weight is reduced; while for regions with less obvious feature information, the depth loss weight is increased. By weighting the depth loss based on feature values, color information is given more weight in areas with strong texture, while depth information is given more weight in areas with weak texture.

[0062] Furthermore, feature weighting can be understood as calculating the feature values ​​extracted from the image; the formula for calculating feature weighting satisfies Formula 3. Formula 3: ; Among them, W feat For feature weights; f is the result of taking the cube root of the feature value extracted from the image; max(f) and min(f) are the maximum and minimum values ​​of f; the result of [f-min(f)] / [max(f)-min(f)] is used to determine the position of the feature of this point in the whole image.

[0063] Furthermore, using image sets, weighted depth data, and pose data as input data, a reconstruction model is used to perform 3D reconstruction of the target object, and scale calibration is completed during the reconstruction process, including: Set scale parameters that can be learned independently in the reconstruction model; The dimensional parameters are optimized based on the historical dimensional patterns of the target object. During the training of the reconstruction model, the scale parameter is controlled to converge within a predetermined training round to achieve scale alignment between the 3D model and the real target object.

[0064] Specifically, during 3D reconstruction using a reconstruction model, scale calibration is achieved by setting a separately learnable scale parameter within the model. The optimization of this scale parameter incorporates historical dimensional patterns of the target object as prior knowledge or constraints, guiding its optimization direction. During the training of the reconstruction model, this scale parameter is optimized independently, and its optimization process is controlled to ensure rapid convergence and stabilization within a predetermined training cycle, ultimately achieving scale alignment between the generated 3D model and the real target object.

[0065] Furthermore, the target object is reconstructed in three dimensions using a reconstruction model, including: The corresponding training strategy is adopted according to the training stage of the reconstruction model; If the reconstruction model is in the early stage, the first training strategy is executed; the first training strategy includes simultaneously optimizing the image reconstruction loss and the depth perception loss, and optimizing the scale parameters in the reconstruction model. If the reconstruction model is in the intermediate stage, a second training strategy is executed; the second training strategy includes initiating a sampling strategy and focusing on model fitting to the core region of the target object. If the reconstruction model is in the later stage, a third training strategy is executed; the third training strategy includes fixing the scale parameters and adjusting the reconstruction model until convergence.

[0066] Specifically, firstly, a corresponding training strategy is adopted according to the training stage of the reconstruction model. When the model is in the early training stage, the first training strategy is executed, which includes simultaneously optimizing the image reconstruction loss and depth perception loss, and optimizing the scale parameters in the model. When the model enters the mid-training stage, the second training strategy is executed, which includes activating a sampling strategy and focusing sampling and computation on the core region of the target object to accelerate model fitting. When the model is in the late training stage, the third training strategy is executed, which includes fixing the previously optimized scale parameters and only fine-tuning other parts of the reconstruction model until the model converges as a whole.

[0067] For example, the first 30 minutes are the early training phase, during which the adopted strategy is frozen, mainly to optimize the image reconstruction loss, depth perception loss, and scale calibration parameters simultaneously. The next 30 minutes after the early phase is the mid-training phase, during which a targeted sampling strategy is launched to focus on the core region of the target object to accelerate model fitting. The next 30 minutes after the mid-training phase is the late training phase, during which the optimization of scale calibration parameters is stopped, and the reconstruction model is adjusted until the reconstruction model converges.

[0068] Experiments revealed that the depth values ​​predicted by deep learning methods do not match the scale of real-world depth data; instead, they are often proportional, resulting in depth maps generated by the model failing to accurately represent the depth data in the real scene. Therefore, this application introduces the concept of depth scale. A separate optimizer is used, trained with a high learning rate (0.01) on the depth loss function for the first 5000 training iterations, a lower learning rate (0.001) for iterations 5000-10000, and then frozen after 10000 iterations. This allows for rapid fitting of the scale of depth data in the model to the real-world depth data.

[0069] Regarding dynamic sampling, by limiting the sampling range to the vicinity of the depth value, the sampling points can be concentrated near the surface of the object that is most likely to be there. This greatly reduces the number of invalid sampling points while ensuring rendering quality, which can effectively speed up the convergence of the network model, improve the quality of the predicted image and depth map, and alleviate the problem of floating objects in the rendering scene.

[0070] Furthermore, in accordance with the corresponding training strategies adopted during the training phase of the reconstructed model, the reconstruction method also includes: Construct a loss function for optimizing the reconstruction model, which includes at least image color reconstruction loss and depth perception loss; The reconstruction model is iteratively optimized by minimizing the loss function.

[0071] Specifically, during training, a loss function needs to be constructed as the training objective to optimize the reconstruction model. This loss function contains at least two core components: image color reconstruction loss and depth perception loss; the loss function is shown in Equation 4. Formula 4: ; Where L is the loss function, Lc is the image color reconstruction loss, Ld represents the depth perception loss, and λ is a hyperparameter used to balance depth and color supervision.

[0072] Furthermore, noting that using only depth loss might be too simplistic, we introduced disparity loss. This is because depth loss represents global depth, focusing more on overall representation capabilities, while disparity directly reflects the distance between pixels, helping the network better learn the details and structure of local regions. Therefore, our depth loss is: ; Where R is the set of three-dimensional rays that originate from the optical center of the camera and pass through a pixel on the image plane; μ is used to balance depth loss and parallax loss, and is generally set to μ=0.01. The input depth image contains depth values. According to The obtained disparity value is... Dep and Disp are the depth and disparity values ​​obtained by the model calculation. The essence of model training is to minimize the loss function through continuous forward calculation and backpropagation, thereby driving the reconstruction model parameters to iteratively optimize and finally obtain a model that can accurately restore the three-dimensional shape and appearance of the target object.

[0073] Furthermore, in NeRF and its derivative methods, r is a three-dimensional ray originating from the camera's optical center and passing through a pixel on the image plane. Mathematically, it is usually parameterized as: r(t) = o + td; where: o: the three-dimensional coordinates of the camera's optical center (the origin of the ray); d: the unit direction vector of the ray (calculated from the pixel position and the camera's intrinsic / extrinsic parameters); t: the distance parameter along the ray (t≥0), used to locate the three-dimensional point on the ray; r(t): the coordinates of the three-dimensional point on the ray r at a distance t from the origin.

[0074] Among these features, the loss function can be adjusted based on the different characteristic information of various types of cultural relics.

[0075] For further information, please refer to [link / reference]. Figure 2 The present invention also provides a rapid and non-destructive digital reconstruction device 100 for cultural relics and ancient buildings, capable of implementing any of the reconstruction methods described above. The reconstruction device 100 includes: a data acquisition module 10, a data processing module 20, a feature fusion module 30, a reconstruction and calibration module 40, and an output module 50. The data acquisition module 10 is used to acquire an image set of the target object and the corresponding depth data. The data processing module 20 is used to perform quality improvement processing on the depth data and to perform pose calculation on the image set to obtain pose data. The feature fusion module 30 is used to identify different feature regions on the surface of the target object based on the image features of the image set, and to assign corresponding fusion weights to the processed depth data according to the characteristics of each feature region. The reconstruction and calibration module 40 is used to perform three-dimensional reconstruction of the target object through a reconstruction model using the image set, the weighted depth data, and the pose data as input, and to complete scale calibration during the reconstruction process. The output module 50 is used to output the three-dimensional model of the target object after scale calibration.

[0076] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.

Claims

1. A method for rapid and non-destructive digital reconstruction of cultural relics and ancient buildings, characterized in that, include: Obtain the image set and corresponding depth data of the target object; The depth data is subjected to quality enhancement processing, and the image set is subjected to pose calculation to obtain pose data; Based on the image features of the image set, different feature regions on the surface of the target object are identified, and corresponding fusion weights are assigned to the processed depth data according to the characteristics of each feature region. Using the image set, the weighted depth data, and the pose data as input data, the target object is reconstructed in three dimensions through a reconstruction model, and scale calibration is completed during the reconstruction process. Output the scale-calibrated 3D model of the target object.

2. The reconstruction method according to claim 1, characterized in that, The process of performing quality enhancement processing on the depth data and performing pose calculation on the image set to obtain pose data includes: The depth data undergoes a first process; the first process includes at least denoising and hidden point removal. The depth data after the first processing is depth-completed; and based on the image set, feature extraction and matching are performed using the structure-of-motion method to solve for and obtain the pose data.

3. The reconstruction method according to claim 2, characterized in that, The step of extracting and matching features from the image set using the structure-of-motion method to obtain the pose data includes: Image feature extraction and matching are performed on the image set, and local features are extracted; Unmatched features are removed from the feature matching results to retain reliable inliers; Based on the inlier, the initial solution of the pose of the image acquisition device and the calculation of the relative pose are performed through multi-view pose extension and preset algorithm. The calculated pose data were optimized using the bundle adjustment method; In the process of solving and optimizing the pose data, closed-loop detection and epipolar constraints are used to enhance the pose.

4. The reconstruction method according to claim 2, characterized in that, After performing depth completion on the depth data after the first processing, the reconstruction method further includes: Obtain a depth image after depth completion; the depth image contains the depth value of each pixel. For each valid depth value in the depth image, a preset sampling depth range is determined; During 3D reconstruction, the distribution of spatial sampling points in 3D space is constrained according to the sampling depth range, so that the spatial sampling points are restricted to the sampling depth range corresponding to each depth value, and sampling data located outside the sampling depth range are excluded.

5. The reconstruction method according to claim 1, characterized in that, The process of identifying different feature regions on the surface of the target object based on the image features of the image set includes: Extract the texture features, edge features, and grayscale statistical features of the image set; Based on the texture features, edge features, and grayscale statistical features, the surface of the target object is divided into a decorative area, a diseased area, and a weak texture area.

6. The reconstruction method according to claim 5, characterized in that, The step of assigning corresponding fusion weights to the processed depth data based on the characteristics of each feature region includes: For the identified weak texture areas, increase the fusion weight of the depth data; For the identified decorative area and the diseased area, the fusion weight of the depth data is reduced.

7. The reconstruction method according to claim 1, characterized in that, The process of using the image set, the weighted depth data, and the pose data as input data to perform 3D reconstruction of the target object through a reconstruction model, and completing scale calibration during the reconstruction process, includes: In the reconstruction model, a scale parameter that can be learned independently is set; The dimensional parameters are optimized based on the historical dimensional patterns of the target object. During the training of the reconstruction model, the scale parameter is controlled to converge within a predetermined training round to achieve scale alignment between the 3D model and the real target object.

8. The reconstruction method according to claim 1, characterized in that, The step of performing three-dimensional reconstruction of the target object using a reconstruction model includes: The corresponding training strategy is adopted according to the training phase of the reconstruction model; If the reconstruction model is in the early stage, the first training strategy is executed; the first training strategy includes simultaneously optimizing the image reconstruction loss and the depth perception loss, and optimizing the scale parameters in the reconstruction model. If the reconstruction model is in the intermediate stage, a second training strategy is executed; the second training strategy includes initiating a sampling strategy and focusing on model fitting to the core region of the target object; If the reconstruction model is in a later stage, a third training strategy is executed; the third training strategy includes fixing the scale parameter and adjusting the reconstruction model until convergence.

9. The reconstruction method according to claim 8, characterized in that, The reconstruction method further includes: (1) Adopting a corresponding training strategy based on the training phase of the reconstructed model; (2) Specific training strategies are used for each phase of the reconstructed model. Construct a loss function for optimizing the reconstruction model, the loss function including at least image color reconstruction loss and depth perception loss; The reconstruction model is iteratively optimized by minimizing the loss function.

10. A device for rapid and non-destructive digital reconstruction of cultural relics and ancient buildings, characterized in that, The reconstruction apparatus, capable of implementing the reconstruction method as described in any one of claims 1 to 9, comprises: The data acquisition module is used to acquire the image set of the target object and the corresponding depth data; The data processing module is used to perform quality improvement processing on the depth data and to perform pose calculation on the image set to obtain the pose data; The feature fusion module is used to identify different feature regions on the surface of the target object based on the image features of the image set, and to assign corresponding fusion weights to the processed depth data according to the characteristics of each feature region. The reconstruction and calibration module is used to take the image set, the weighted depth data, and the pose data as input, perform three-dimensional reconstruction of the target object through the reconstruction model, and complete scale calibration during the reconstruction process. The output module is used to output the three-dimensional model of the target object after scale calibration.