Three-dimensional reconstruction system for immovable cultural relics

By using a 3D reconstruction system for immovable cultural relics and employing deep learning and SLAM technologies to construct virtual restoration models, the problems of low data acquisition accuracy and insufficient automation in the digital restoration of antiquities have been solved. This has enabled the automation and intelligentization of antiquities restoration, improved restoration efficiency and accuracy, and promoted the inheritance and standardization of techniques.

CN121053296APending Publication Date: 2025-12-02ZHEJIANG COLLEGE OF CONSTR
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511166270.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing digital restoration methods for antiquities suffer from low data acquisition accuracy and insufficient automation, while traditional restoration techniques rely on skill and are inefficient.

Method used

A three-dimensional reconstruction system for immovable cultural relics is adopted, including a data acquisition module, a feature extraction module, a restoration simulation module, and an optimization and decision-making module. It utilizes deep learning algorithms and SLAM technology, combined with NeRF technology, to construct a virtual restoration model, thereby achieving automation and precise positioning of the restoration process.

Benefits of technology

It has significantly improved the efficiency and precision of restoration, solved the problem of the interruption of the inheritance of skills, promoted industry standardization, realized the automation and intelligence of antiquities restoration, and shortened the restoration cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BSA0000301531710000031
    Figure BSA0000301531710000031
  • Figure BSA0000301531710000101
    Figure BSA0000301531710000101
  • Figure BSA0000301531710000121
    Figure BSA0000301531710000121
Patent Text Reader

Abstract

The invention discloses an immovable cultural relic three-dimensional reconstruction system which comprises a data acquisition module, a feature extraction module, a restoration simulation module and an optimization and decision module. The data acquisition module is used for acquiring three-dimensional data of an antique; the feature extraction module is used for extracting key features of antiques by using a deep learning algorithm; the restoration simulation module is used for constructing a virtual restoration model based on the extracted features, simulating an antique restoration process and evaluating a restoration scheme; and the optimization and decision module is used for realizing precise positioning and navigation of a repair site in combination with the SLAM technology, optimizing a repair path and assisting a repairer in making a scientific decision. According to the method, the repairing efficiency and precision are remarkably improved, automation and intelligentization of the ancient object repairing process are achieved by introducing advanced technologies such as deep learning and NeRF, the repairing efficiency and precision are greatly improved, and the repairing period is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional reconstruction systems for immovable cultural relics, and specifically to a three-dimensional reconstruction system for immovable cultural relics. Background Technology

[0002] Currently, antiquities restoration techniques are mainly divided into traditional manual restoration and digitally assisted restoration. Traditional techniques rely on the skills of restorers, which suffers from problems such as a break in the transmission of skills, inconsistent standards, and low efficiency. Although digital technology has introduced auxiliary means such as 3D scanning and VR / AR, it still faces shortcomings such as low data acquisition accuracy and insufficient automation. Summary of the Invention

[0003] To address the aforementioned problems, this invention proposes a three-dimensional reconstruction system for immovable cultural relics, which solves the shortcomings of low data acquisition accuracy and insufficient automation in existing digital-assisted restoration of antiquities.

[0004] The technical solution adopted in this invention is as follows:

[0005] A three-dimensional reconstruction system for immovable cultural relics includes:

[0006] A data acquisition module; the data acquisition module is used to acquire the three-dimensional data of the artifact;

[0007] A feature extraction module; the feature extraction module is used to extract key features of the artifact using deep learning algorithms;

[0008] A restoration simulation module; the restoration simulation module is used to construct a virtual restoration model based on extracted features, simulate the restoration process of antiquities, and evaluate restoration plans; and

[0009] An optimization and decision-making module; the optimization and decision-making module is used to combine SLAM technology to achieve accurate positioning and navigation at the repair site, optimize the repair path, and assist repair technicians in making scientific decisions.

[0010] Optionally, the three-dimensional reconstruction system for immovable cultural relics of the present invention further includes a preprocessing module, which is used to perform denoising, registration, and segmentation preprocessing operations on the acquired three-dimensional data.

[0011] Optionally, the denoising process includes statistical filtering and bilateral filtering; the registration operation includes using SAC-IA or FPFH feature matching to obtain an initial transformation matrix.

[0012] Optionally, the registration operation employs the ICP algorithm, which includes iterative optimization, convergence conditions, and global optimization. The iterative optimization involves accelerating the nearest point search using a KD-tree and solving for the optimal transformation using SVD. The convergence conditions specifically involve setting an error threshold or a maximum number of iterations. Global optimization involves using Pose Graph optimization to eliminate accumulated errors.

[0013] Optionally, the three-dimensional reconstruction system for immovable cultural relics of the present invention further includes an output and display module, which is used to output the restoration results in the form of a three-dimensional model and a restoration report.

[0014] Optionally, the optimization and decision-making module incorporates biological visual mechanisms, physical modeling, and cross-scale feature fusion, wherein the cross-scale feature fusion strategy includes a pyramid attention module, a frequency domain enhancement method, and a dynamic environment adaptation scheme.

[0015] Optionally, the specific operations of constructing the virtual restoration model include using NeRF technology to construct a neural radiation field model of the original form of the artifact and using virtual restoration to simulate the effects of different restoration schemes in the NeRF model.

[0016] Optionally, the key features of the antiquities include geometric features, material properties, and texture features.

[0017] Optionally, the geometric features are the curvature and depth variations of the damaged area; the material features are the spectral reflectance; and the texture features are the texture direction and density of historical traces.

[0018] The present invention discloses a method for using a three-dimensional reconstruction system for immovable cultural relics, comprising the following steps:

[0019] 1) Use a 3D scanner to obtain 3D data of the artifacts;

[0020] 2) Perform denoising, registration, and segmentation preprocessing operations on the acquired 3D data;

[0021] 3) Use deep learning algorithms to extract key features of antiquities.

[0022] 4) Based on the extracted features, a virtual restoration model is constructed using NeRF technology to simulate the restoration process of antiquities and evaluate restoration plans;

[0023] 5) Combining SLAM technology to achieve precise positioning and navigation at the restoration site, optimize restoration paths, and assist restorers in making scientific decisions;

[0024] 6) Output the repair results in the form of 3D models and repair reports, and combine AR / VR technology to provide an immersive display and interactive experience.

[0025] The beneficial effects of the present invention include at least the following:

[0026] 1. This invention significantly improves restoration efficiency and accuracy, and by introducing advanced technologies such as deep learning and NeRF, it realizes the automation and intelligence of the restoration process of antiquities, greatly improving restoration efficiency and accuracy and shortening the restoration cycle.

[0027] 2. This invention promotes the inheritance and standardization of skills: digitally recording and disseminating restoration techniques, combined with a VR teaching system, effectively solves the problem of the interruption of skill inheritance, and establishes a standardized digital restoration process to ensure restoration quality and consistency.

[0028] 3. This invention promotes interdisciplinary integration and innovation: This invention integrates technologies from multiple fields such as computer vision, deep learning, and 3D reconstruction, promoting interdisciplinary integration and innovation, and bringing new technical ideas and solutions to the field of antiquities preservation.

[0029] 4. This invention enhances the standardization and normalization of the industry: It promotes the establishment of industry standards and acceptance specifications in the field of digital restoration, which helps to improve the standardization and normalization of the entire industry and promotes the healthy development of antiquities restoration technology.

[0030] 5. This invention employs a combination strategy of SAC-IA and FPFH, which can accurately estimate the initial transformation matrix even when the overlap area is less than 50%, significantly improving the registration success rate. Detailed Implementation

[0031] The specific embodiments of the present invention will be described in further detail below with reference to examples. These examples are used to illustrate the present invention, but are not intended to limit the scope of the invention.

[0032] The high-precision 3D laser scanner used in this invention can be the FARO Focus3D series.

[0033] The structured light scanner in this invention can be an Artec Eva. The high-definition camera in this invention can be a Canon EOS R5.

[0034] In this invention, the ICP algorithm (Iterative Closest Point Algorithm) is one of the core methods for point cloud registration. It solves the optimal spatial transformation parameters (rotation matrix and translation vector) between two sets of point clouds through iterative optimization.

[0035] The goal is to minimize the difference between the source point cloud P5 and the target point cloud P. t Sum of squared Euclidean distances between corresponding points:

[0036]

[0037] Where R is the rotation matrix, t is the translation vector, and pi and qi are the matching point pairs.

[0038] During this process, rotation and translation are separated and solved by centroidalization (calculating the geometric center of the point cloud); the rotation matrix R is obtained by using SVD decomposition of the covariance matrix or the quaternion method to achieve decoupling calculation.

[0039] In this invention, SAC-1A (Sample Consensus Initial Alignment) is a coarse registration algorithm for point clouds based on feature matching and random sampling consistency, mainly used to solve the problem of aligning 3D point clouds with large initial pose differences. Specifically, it establishes similarity relationships between points by extracting local feature descriptors of the point cloud (such as FPFH (Fast Point Feature Histogram) or 3DSC (3D Shape Context)), and then combines the sampling consistency criterion to select reliable matching point pairs.

[0040] In this invention, the KD-tree (K-dimensional tree) is a binary tree data structure used for efficiently organizing and retrieving data points in k-dimensional space. It is primarily applied to range searches (e.g., finding points within a specified region) and nearest neighbor searches (e.g., finding the point closest to the query point). Its basic structure and construction principles include: Spatial recursive partitioning: Each non-leaf node represents a hyperplane, dividing the entire k-dimensional space into two mutually exclusive subspaces (the left subtree corresponds to one side of the hyperplane, and the right subtree to the other side). Partition dimension selection: During construction, dimensions are selected alternately for partitioning (e.g., the x-axis is selected in the first layer, the y-axis in the second layer, and so on). Alternatively, the dimension with the largest variance can be dynamically selected based on the data distribution to improve partition uniformity. Partition point selection: On the selected dimension, the median of all data points is taken as the partition point to ensure a balance of data volume between the left and right subtrees. Recursive termination: When a subspace has only one data point or no data, a leaf node is generated.

[0041] In this invention, Pose Graph is an efficient backend optimization method for optimizing the motion trajectory of robots or cameras. It simplifies the computational complexity of traditional Bundle Adjustment (BA) and focuses on the constraint relationships between pose nodes to achieve global consistency.

[0042] In this invention, CNN (Convolutional Neural Network) is a deep learning model specifically designed for processing grid-structured data (such as images and audio), and it performs exceptionally well in the field of computer vision. Its core lies in simulating the hierarchical feature extraction mechanism of biological visual systems, efficiently capturing spatial features through designs such as local connectivity and weight sharing.

[0043] The NeRF model (Neural Radiance Fields) in this invention is a technique that uses deep learning for implicit reconstruction and rendering of 3D scenes.

[0044] In this invention, the FPGA implementation uses the Xilinx Zynq series, and the throughput of the heat conduction equation module reaches 230 FPS.

[0045] In this invention, embedded optimization achieves 1080p real-time processing (30FPS) on NVIDIA Jetson AGX Xavier.

[0046] The Short-Time Fourier Transform (STFT) in this invention is a time-frequency analysis method for analyzing non-stationary signals. It achieves localized frequency analysis by sliding a fixed window function across the signal, thereby overcoming the limitation of traditional Fourier Transform in handling time-varying signals.

[0047] Example 1

[0048] The technical solution adopted in this invention is as follows:

[0049] This invention discloses a three-dimensional reconstruction system for immovable cultural relics, comprising:

[0050] The system comprises a data acquisition module, a preprocessing module, a feature extraction module, a restoration simulation module, an optimization and decision-making module, and an output and display module. The data acquisition module acquires the 3D data of the artifact; the preprocessing module performs denoising, registration, and segmentation preprocessing on the acquired 3D data. The feature extraction module uses deep learning algorithms to extract key features of the artifact; the restoration simulation module constructs a virtual restoration model based on the extracted features, simulates the restoration process, and evaluates restoration plans; the optimization and decision-making module combines SLAM technology to achieve precise positioning and navigation at the restoration site, optimize restoration paths, and assist restorers in making scientific decisions.

[0051] This invention proposes a systematic solution to the core challenges in the reconstruction of immovable cultural relics through a modular collaborative architecture. The data acquisition module provides high-precision foundational information for subsequent processing through 3D data acquisition, directly improving the accuracy of the data source. The feature extraction module uses deep learning algorithms to replace traditional manual feature annotation, enhancing the automated recognition of complex features of ancient artifacts (such as damage patterns and historical traces). The restoration simulation module constructs a virtual restoration model based on extracted features, using digital simulation technology to pre-evaluate restoration plans and avoid secondary damage to cultural relics caused by traditional trial-and-error restoration methods. The optimization and decision-making module introduces SLAM technology, combining spatial positioning and path planning to provide dynamic navigation support for operators during the physical restoration phase. Simultaneously, it optimizes restoration paths through algorithms, improving restoration efficiency and the scientific nature of decision-making. These modules form a closed-loop process, systematically improving data acquisition accuracy, the automation level of feature processing, the scientific nature of restoration plans, and the precision of on-site operations.

[0052] In this embodiment, during the registration stage, a set of feature points is first randomly selected from the point cloud using the SAC-IA algorithm. For example, four corresponding points are selected in each iteration to calculate the rigid body transformation. The RANSAC mechanism is then used to select the transformation matrix with the highest proportion of interior points as the initial registration result. Furthermore, a local geometric feature histogram is constructed using the FPFH feature descriptor. For example, the angle between normal vectors is divided into 11 intervals, and the curvature value is divided into 5 intervals. Corresponding point pairs are optimized through feature matching, and finally, accurate initial transformation parameters are output for subsequent ICP algorithm iteration optimization.

[0053] In the registration operation, the SAC-IA algorithm achieves robust feature matching through the sampling consistency mechanism, while the FPFH feature descriptor enhances feature discriminability by calculating the histogram of local geometric properties. The two methods work together to obtain a high-precision initial transformation matrix, providing reliable initial values ​​for subsequent iterative optimization, thereby reducing the model misalignment problem caused by registration errors.

[0054] In this embodiment, a high-precision 3D laser scanner is used to ensure the acquisition of point cloud data with millimeter-level accuracy.

[0055] In this embodiment, the denoising process includes statistical filtering and bilateral filtering; the registration operation includes using SAC-IA or FPFH feature matching to obtain an initial transformation matrix. Statistical filtering refers to the operation of removing discrete noise points based on the statistical characteristics of the point cloud neighborhood. Specifically, it can be implemented by calculating the point cloud density distribution and setting a standard deviation threshold to eliminate outliers caused by scanning equipment errors.

[0056] Bilateral filtering refers to a smoothing algorithm that combines spatial and color domain weights. Specifically, it can be implemented by using a Gaussian kernel function to calculate the weighted average of neighboring points, thereby eliminating surface noise while preserving the edge sharpness of the engraved texture on the artifact. In this invention, denoising improves registration accuracy, coarse registration provides a good initial position for ICP, and global optimization ensures consistency of point clouds from multiple perspectives.

[0057] Segmentation preprocessing refers to dividing 3D data into local regions with independent semantics. This can be achieved using clustering algorithms based on curvature or color thresholds. For example, the surface of a cultural relic can be divided into complete regions, damaged regions, or regions with attachments, providing a structured data foundation for subsequent restoration simulation.

[0058] Specifically, the preprocessing module eliminates scattered noise points in the point cloud through denoising operations, such as isolated points caused by scanner jitter, to prevent noise from being misjudged as real structures in the subsequent feature extraction process; it aligns multi-view scan data through registration operations, such as fusing laser scan data with photogrammetric data to form a complete 3D model; and it decomposes the complex 3D model into semantic units through segmentation operations.

[0059] This invention integrates a multi-stage preprocessing workflow to eliminate noise interference while maintaining the integrity of geometric details. It achieves efficient data alignment through automated feature matching and provides structured input for repair simulations through semantic segmentation, significantly reducing manual operations and improving data quality. Conventional ICP algorithms heavily rely on the initial pose and are prone to getting trapped in local optima when the point cloud overlap is less than 60%. This invention employs a combined strategy of SAC-IA and FPFH, which can accurately estimate the initial transformation matrix even when the overlap area is less than 50%, significantly improving the registration success rate.

[0060] In another embodiment, the data acquisition module further includes auxiliary devices, including a structured light scanner and a high-definition camera. This invention combines a structured light scanner and a high-definition camera to acquire texture and color information.

[0061] In the scanning strategy of the data acquisition module, the resolution is adjusted according to the size of the artifact (e.g., 0.1mm-1mm) when setting the resolution. During multi-angle scanning, the scanning path is planned to ensure unobstructed areas and cover the entire surface of the artifact.

[0062] When setting up environmental records, it is necessary to record light intensity, temperature, and humidity simultaneously for subsequent data correction.

[0063] In this embodiment, high-precision point cloud provides the foundation for subsequent registration and reconstruction, while environmental data assists in denoising and error compensation.

[0064] The output and display module is used to output the repair results in the form of a 3D model and a repair report.

[0065] The registration operation employs the ICP algorithm, which includes iterative optimization, convergence conditions, and global optimization. The iterative optimization accelerates the nearest point search using a KD-tree and solves for the optimal transformation using SVD. The convergence conditions specifically involve setting an error threshold (e.g., 0.01 mm) or a maximum number of iterations. Global optimization uses Pose Graph optimization to eliminate accumulated errors.

[0066] In this embodiment, KD-tree acceleration of nearest neighbor search refers to hierarchical partitioning of point cloud data by constructing a spatial index structure. Specifically, a binary tree structure can be used to recursively divide the 3D space into sub-regions, with each node storing a subset of the point cloud corresponding to that region, thereby reducing computational complexity when searching for nearest neighbors. SVD solution for optimal transformation refers to calculating the rotation matrix and translation vector between point clouds through singular value decomposition. Specifically, the optimal rigid body transformation parameters can be obtained using the eigenvalue decomposition of the covariance matrix, ensuring the orthogonality constraint of the transformation matrix. Error threshold or maximum number of iterations as convergence conditions refers to setting iteration termination rules. Specifically, the calculation can be stopped when the mean square error is below a preset threshold or when a specified number of iterations is reached, preventing invalid loops and balancing accuracy and efficiency. Pose Graph global optimization refers to jointly optimizing the multi-view registration results by establishing a pose graph model. Specifically, a nonlinear least squares method can be used to adjust the pose nodes of each frame, eliminating the cumulative error caused by adjacent registrations.

[0067] In this embodiment, the deep learning algorithms include CNN algorithm and PointNet++.

[0068] The specific operations for constructing the virtual restoration model include using NeRF technology to construct a neural radiation field model of the original form of the artifact and using virtual restoration to simulate the effects of different restoration schemes in the NeRF model.

[0069] The output and display module is also used to provide immersive display and interactive experiences by combining AR / VR technology.

[0070] In this embodiment, for example, the output of the 3D model generates a surface topology structure through point cloud reconstruction and mesh optimization algorithms, and uses texture mapping technology to restore the color information of the artifact. Users can adjust the viewing angle through touch screen or mouse operation. The restoration report automatically generates a standardized document containing charts and text descriptions by calling the timestamps, material parameters, and evaluation indicators of the restoration process from the database. Immersive display uses SLAM technology to locate the user's viewpoint in real time, superimposing the 3D model onto the real environment. Users can simulate restoration actions by operating virtual tools with gestures.

[0071] AR technology, or Augmented Reality technology, can be implemented using SLAM (Simultaneous Localization and Image Recognition) technology. It captures real-world scenes using a camera and overlays the 3D information of a virtual restoration model onto a display device. This technology allows for spatial alignment of the restoration effect with the actual artifact, achieving a visual verification that blends the virtual and real worlds.

[0072] VR technology refers to virtual reality technology, which can be implemented using head-mounted display devices and spatial positioning systems to present restored 3D models in a holographic virtual space. This technology allows observers to freely adjust their viewing angle and distance within the virtual environment, enabling multi-dimensional exploration of details.

[0073] Gesture recognition technology refers to capturing a user's hand movements using a depth camera or inertial sensor. Specifically, it can employ convolutional neural networks to identify gesture types and map preset gestures to model rotation, scaling, and other operation commands. This technology can replace traditional input devices and achieve natural interaction.

[0074] Spatial positioning technology refers to pose tracking based on UWB or infrared optical markers, specifically using multi-sensor fusion algorithms to calculate the user's position and orientation. This technology ensures that the user's relative position to the virtual model remains accurately aligned during movement.

[0075] In one embodiment, the 3D model of this invention is output in OBJ / PLY format, with a before-and-after comparison. The repair report records the repair steps, materials, timelines, and quality assessment.

[0076] During the restoration plan demonstration phase, AR technology overlays a virtual restoration model onto the damaged area of ​​the artifact through real-time rendering, allowing restorers to visually compare the morphological differences before and after restoration. When a user wears a VR device, the system generates a holographic environment containing the complete restoration model, allowing the user to observe details of the model from various angles by rotating their head and moving their position. During the interaction, the gesture recognition module captures the user's hand movements, such as pinching to trigger model scaling and waving to switch restoration plan versions. The spatial positioning module synchronously updates the user's coordinates in the virtual space, ensuring real-time matching between operation commands and model responses. This closed-loop interactive system allows the restoration effect evaluation process to be completed through multi-dimensional observation and dynamic operation. In this embodiment, the artifact restoration process can be displayed via mobile phone / tablet. A virtual exhibition hall is provided, allowing users to interactively view restoration details.

[0077] In this embodiment, the key features of the antiquities include geometric features, material properties, and texture features.

[0078] In this embodiment, the geometric features are the curvature and depth variations of the damaged area; the material features are spectral reflectance; and the texture features are the texture direction and density of historical traces. The virtual repair model is compared with actual data to assess the feasibility of repair. During this process, quantitative indicators are used for comparison, specifically shape similarity and texture consistency.

[0079] Among them, geometric features refer to the calculation of the surface curvature and depth changes of the damaged area through three-dimensional point cloud data. Specifically, curvature estimation algorithms can be combined with depth sensor measurements to quantify the degree of structural deformation of the artifact.

[0080] Curvature refers to a quantitative indicator of the degree of local bending of a three-dimensional surface. It can be calculated using a curvature estimation algorithm after acquiring point cloud data with a 3D scanner, and is used to identify the geometric deformation characteristics of damaged areas on the surface of artifacts. Depth variation refers to the vertical displacement of concave or convex areas on the surface, which can be measured using structured light or lidar, and is used to determine the repair filling thickness of damaged areas. Spectral reflectance refers to the material's ability to reflect light of different wavelengths, which can be acquired using multispectral imaging equipment, and is used to distinguish the physical properties of different materials on the surface of artifacts. Texture direction refers to the arrangement direction of surface patterns or traces, which can be achieved by extracting gradient direction histograms using image processing algorithms, and is used to preserve the spatial distribution characteristics of historical traces. Texture density refers to the number of texture elements distributed per unit area, which can be achieved using local binary mode analysis, and is used to quantify the density of historical traces.

[0081] Specifically, curvature calculation and depth change measurement can accurately identify the spatial deformation parameters of damaged areas of artifacts, providing geometric constraints for restoration path planning. Spectral reflectance data, acquired through multispectral imaging equipment, establishes a mapping relationship between material characteristics and restoration materials, avoiding deviations in restoration plans due to material misjudgment. Texture direction and density, extracted using image processing algorithms, form two-dimensional distribution parameters of historical traces, ensuring the integrity of cultural information during restoration. The combination of these three elements constitutes a multidimensional feature dataset, serving as input to a deep learning model to improve the comprehensiveness and accuracy of feature extraction.

[0082] Among them, material properties refer to the spectral reflectance parameters of materials obtained through spectral imaging equipment. Specifically, this can be achieved by using a multispectral scanner to collect data in the visible to near-infrared bands, which is used to establish a physical property database of material composition.

[0083] Among them, texture features refer to the analysis of the texture direction and distribution density of historical traces through high-resolution image analysis. Specifically, it can be achieved by combining the directional gradient histogram algorithm with the texture segmentation model, and is used to identify surface wear patterns and historical processing traces.

[0084] In one embodiment, after the 3D scanning device acquires point cloud data of the artifact's surface, a curvature estimation algorithm extracts local geometric features from the damaged area to generate a curvature distribution map to locate the structural deformation area; a multispectral scanner collects material reflectance spectral data, and a spectral matching algorithm compares it with material parameters in the database to determine the original material composition of the artifact; after a high-resolution camera captures surface images, an directional gradient histogram algorithm calculates the principal direction angle of the texture, and a texture segmentation model statistically analyzes the line density per unit area to identify the distribution differences between artificial carving marks and natural weathering marks. These three types of feature data are then standardized and input into the restoration simulation system to form a multi-dimensional restoration benchmark that includes structural deformation, material properties, and historical traces.

[0085] Based on the results of feature extraction, the NeRF model provides a visual repair solution to assist repair technicians in making decisions.

[0086] In this embodiment, the optimization and decision-making module incorporates biological visual mechanisms, physical modeling, and cross-scale feature fusion. Specifically, a bio-inspired dual-channel dynamic thresholding mechanism is used to simulate retinal ganglion cells, constructing an ON / OFF dual-pathway processing. The ON channel generates a positive activation map R through brightness enhancement response. + (x, y) OFF channel: generates a negative activation map R through brightness attenuation response. - (x, y). In image processing formulas, x and y represent the spatial coordinates of a pixel in the image. x: represents the horizontal coordinate (column index) of the pixel, with a value range of [0, image width - 1], and y: represents the vertical coordinate (row index) of the pixel, with a value range of [0, image height - 1].

[0087] The dynamic threshold formula is:

[0088]

[0089] Where T(x,y): the threshold or target value at coordinates (x,y); α, β, γ are weighting coefficients; Median(R+) is the region R. + The median of pixel values ​​within the target / positive sample region; Median(R-) is the median of the region R. - Median pixel value within the (background / negative sample region); The gradient of image l; σ I denoted as the standard deviation of image I; ∈: a minimum constant (to avoid zero in the denominator).

[0090] In one embodiment of the industrial inspection scenario, ∈ = 0.5 (structural detail preservation) and the number of iterations = 15 (convergence stability).

[0091] In another embodiment, in an outdoor monitoring scenario, ε = 1.2 (resistance to changes in illumination) and the number of iterations = 25 (adaptation to complex textures).

[0092] Regarding dual-channel coordinated control, a competition-cooperation mechanism can be adopted, specifically by weighted fusion of the outputs of the ON / OFF channels to avoid signal conflicts.

[0093] Alternatively, pulse coding can be used, specifically to simulate the pulse firing pattern of biological neurons, and event-driven computation (such as SNN models) can be employed to reduce power consumption.

[0094] There is a close synergistic relationship between the spiking neural network integration and the bio-inspired dual-channel dynamic thresholding mechanism. By simulating the dynamic characteristics and information processing patterns of biological nervous systems, both enhance the network's spatiotemporal information representation capabilities, computational efficiency, and adaptability. This includes implementing event-driven threshold updates. Specifically, event-driven threshold updates involve converting the image into a pulse time series, with each pixel independently generating a trigger time t. fire .

[0095] Membrane potential dynamic equation:

[0096] V(t)=V(t-1)+I(x,y)-T(t) if t=t fire

[0097] Where V(t) is the state quantity at time t (such as voltage, energy, etc.), V(t-1): the value of the state quantity at time t-1 before the previous moment; I(x,y) is the input quantity related to the spatial location (x,y) (such as image pixel intensity, current, etc.); T(t) is the loss quantity or threshold quantity at time t (such as temperature, energy loss, etc.); t fire The trigger moment or ignition moment (the specific point in time when the formula takes effect).

[0098] The dynamic equation of membrane potential and the biological visual mechanism together reveal how biological systems encode and process information through dynamic electrical activity. The membrane potential equation describes the dynamic changes in neuronal membrane potential with ion channel current, while the biological visual mechanism relies on the light-dependent changes in membrane potential in photoreceptor cells to transmit visual information.

[0099] Threshold decay model:

[0100] T(t) = T0·e- λt withλ=0.05

[0101] Where T(t) is the physical quantity value at time t; T0 is the physical quantity value at the initial time (t=0); e is the natural constant (approximately 2.71828); λ is the decay coefficient (0.05 here); and t is time. Temperature here refers to ambient temperature. The physical values ​​in this invention specifically include temperature and humidity.

[0102] This model accurately reproduces the response characteristics of photoreceptor cells to light stimuli in biological visual systems, as well as the ability of visual neurons to synchronously process temporal information, by simulating the dynamic changes in the excitation threshold of neurons.

[0103] In this embodiment, the physically constrained diffusion model includes anisotropic heat conduction equations and threshold propagation equations. Specifically, the anisotropic heat conduction equations include:

[0104] Diffusion coefficient design

[0105]

[0106] Where λ1 and λ2 are the eigenvalues ​​of the structure tensor.

[0107] The specific threshold propagation equation is as follows:

[0108]

[0109] Boundary conditions: Neumann boundary (zero flux)

[0110] in, The partial derivative of temperature T with respect to time t represents the rate of change of temperature over time. T is temperature. t is time. is the gradient operator, used to represent the spatial derivative. D is the thermal diffusivity (or diffusion coefficient). The temperature gradient represents the rate of change of temperature in space. For the divergence operator to act on Describes the diffusion process of heat conduction; LBP(I) stands for Local Binary Pattern, which is usually related to the local texture features of image I, and here it is a part of the source term. I represents the input image.

[0111] In this embodiment, the physical constraints also include implementation using the phase-field method, specifically the free energy function implemented using the phase-field method:

[0112]

[0113] Where G represents the energy functional (e.g., the free energy functional). dx is the spatial volume element. ε 2 Small parameters that are positive (related to material properties, length, or diffusion coefficient). This represents the square of the temperature gradient (reflecting the spatial variation of the temperature field). T represents temperature (the core variable, describing the spatial temperature distribution). 。 (T 2 -1) 2 This is a temperature-dependent nonlinear barrier term (it takes a minimum value of 0 when T = ±1, corresponding to the equilibrium state).

[0114] The specific numerical solution scheme adopts the semi-implicit Fourier spectrum method, with a time step of Δt = 0.1 and a spatial discretization of 256 × 256.

[0115] In this embodiment, the cross-scale feature fusion strategy includes a pyramid attention module, a frequency domain enhancement method, and a dynamic environment adaptation scheme;

[0116] In the pyramid attention module, four layers of Gaussian pyramid feature extraction are performed, and LBP+Gabor joint features are calculated for each layer:

[0117]

[0118] Gaussian pyramids are a multi-scale image decomposition technique that uses Gaussian blur and downsampling to construct image layers of different resolutions for extracting features at different scales. The formula above is a step in the Local Binary Pattern (LBP) feature extraction process.

[0119] F l This refers to the LBP feature extracted at a specific scale or location *l*. It is a vector or value describing the local texture features of an image. LBP, or Local Binary Pattern Operator, is a method used for image texture analysis. The LBP operator generates a binary pattern by comparing the center pixel with its neighboring pixels, thus obtaining a feature value describing the local texture. l A region or pixel value of an image at a specific scale or location. It is usually a local window of the image or the grayscale value of a specific pixel.

[0120] Gabor filters are frequency-domain texture analysis tools that extract texture features using filters of different frequencies and directions, simulating the "direction-frequency" sensitivity of biological vision. By fusing the LBP and Gabor features from the four-layer pyramid according to scale, frequency, and direction, a multi-dimensional description of the texture is formed.

[0121] This formula represents the application of Gabor filters in image processing; specifically, it describes the process of applying Gabor filters in different directions. Where I... l This represents the grayscale value or feature map of the input image at time t. In image processing, I... l It can be the original image or an image that has undergone some kind of preprocessing.

[0122] θ represents the direction parameter of the Gabor filter. In this formula, θ is set to a series of discrete direction values, namely [0, π / 4, π / 2, 3π / 4]. These direction values ​​represent the direction in which the Gabor filter is applied to the image, and are typically used to capture texture information in different directions within the image.

[0123] Cross-scale weight allocation:

[0124]

[0125] Among them, w l Z represents the weight of the l-th feature map. It indicates the importance of that feature map in the final output. l The scalar value of the l-th feature map is used to calculate the unnormalized weights. W2 is a weight matrix used to map intermediate features to the final scalar value Z. l δ is an activation function, typically used to introduce non-linearity. W1 is a weight matrix used to map the average pooled features to intermediate features. AvgPool(F l For the l-th feature map F l Perform average pooling. Average pooling reduces the spatial dimension of the feature map while preserving its main features.

[0126] F l This is the l-th feature map, which is usually extracted from a layer of a convolutional neural network (CNN).

[0127] In the frequency domain enhancement method of this embodiment, a short-time Fourier transform is used. The specific parameters of the short-time Fourier transform are: the window function is a Hanning window (length 64), and the overlap rate is 75%.

[0128] Phase consistency weighting:

[0129] T freq =T spatial ·(1+0.5·PC(x,y))

[0130] T freq This represents the temperature value after frequency or location-related correction. T spatial This represents the base space temperature, i.e., the temperature value without any corrections. 1 + 0.5·PC(x,y) is the correction factor used to adjust the base temperature T based on the location (x,y). spatial .

[0131] Both phase consistency weighting and cross-scale weighting achieve accurate capture of information saliency and cross-scale complementarity in multi-scale feature fusion through dynamic weight allocation mechanisms. Phase consistency weighting focuses on local feature weighting driven by signal phase similarity, while cross-scale weight allocation focuses on dynamic modeling of the importance differences between cross-scale features.

[0132] Specifically, the Kovesi algorithm (8-direction Gabor filter bank) is used for PC calculations.

[0133] This embodiment also adopts a dynamic environment adaptation scheme, in which the specific online parameter optimization adopts dual time scale updates and multi-sensor fusion.

[0134] The specific dual-timescale updates include fast adaptation (per frame) and slow learning (per 100 frames).

[0135] Fast adaptation (per frame) specifically means:

[0136]

[0137] Where, φ t φ is the parameter value at time t. t+1 The parameter value at time t+1 is 0.01, which is the learning rate (parameter update step size). Let L be the gradient of the loss function L with respect to the parameter φ. Loss function (measures the difference between the prediction and the true value); Tφ( / t): the model's prediction result for the input / t under parameter φ. The true label (target value) at time t.

[0138] Slow learning (per 100 frames) specifically refers to:

[0139]

[0140] Here, θ represents the updated values ​​of the model parameters after the (t+1)th iteration. In machine learning, θ is the model's weights / parameters (such as the weight matrix of a neural network, the coefficients of linear regression, etc.). t+1 It is through "current parameter θ" t The new parameters are obtained by "+gradient descent update amount". θ t This represents the current value of the model parameters at iteration t. 0.001 is the learning rate. This is the gradient operator with respect to θ, used to calculate the partial derivative of the loss function with respect to each parameter (if θ is a multidimensional vector, the gradient is a vector composed of the partial derivatives of each dimension). E[L] is the loss function, which measures the difference between the model's prediction and the true value (such as mean squared error MSE, cross-entropy CE, etc.).

[0141] The multi-sensor fusion in this embodiment adopts a joint illumination-distance model and a threshold compensation formula.

[0142] Specifically:

[0143] P(L|d)=N(500e -0.1d 50 2 )

[0144] In this embodiment, under condition d (representing distance), the random variable L has a distribution with a mean of 500e. -0.1d The variance is 50. 2 It follows a normal distribution.

[0145] The threshold compensation formula is as follows:

[0146]

[0147] Among them, T comp This is the corrected temperature (e.g., compression correction temperature);

[0148] Where T is the base temperature; P is the pressure; L is the length-related parameter; and d is the diameter or characteristic dimension.

[0149] The illumination-distance joint model and the threshold compensation formula work together to influence the system's perception or response mechanism. By combining illumination and distance information, the system's response threshold is dynamically adjusted to optimize performance and adapt to different environmental conditions.

[0150] This invention compensates for visual odometry drift through point cloud registration, and realizes three major strategies for the fusion of vision and lidar: visual-assisted lidar uses visual information to improve the closed-loop detection accuracy of lidar SLAM or assist in relocalization.

[0151] Visual features (such as ORB and SIFT) and laser point cloud features (such as FPFH and SHOT) are jointly optimized to enhance the laser features. The specific error function is as follows:

[0152] E=ω I E visual +ω2E LiDAR .

[0153] Where ω1 and ω2 are weighting coefficients used to adjust the contribution ratio of the two items; E visual E represents the energy or cost term related to vision. LiDAR This refers to the energy or cost associated with lidar.

[0154] By visually recognizing environmental semantics (such as lane lines and buildings), a semantic map is constructed, improving the accuracy of laser matching and the performance of laser SLAM in structured environments.

[0155] In this embodiment, laser-assisted vision utilizes laser point clouds to provide scale information or motion priors for visual features.

[0156] Specifically, the LIM0 scheme is adopted, which projects laser point clouds onto the image plane, estimates the scale of visual features, and constructs a joint optimization problem:

[0157]

[0158] This formula represents the objective function of an optimization problem, commonly used in tasks within machine learning or computer vision, such as pose estimation or point cloud registration. Below is an explanation of the symbols in the formula:

[0159] Where T and λ are the variables to be optimized, and π(λ) i X i ) is a scaling factor or weight π, a projection function that projects a 3D point X. i Projected onto a two-dimensional plane. λ i For point X i Related scaling factor. p i This refers to the projected target point or observation point. LiDAR T represents the transformation matrix obtained from LiDAR (Light Detection and Ranging) data, and indicates the coordinate system of the LiDAR point cloud. pred The transformation matrix representing the prediction is the result obtained through the optimization process.

[0160] This embodiment utilizes high-frequency pose estimation from LiDAR to correct point cloud distortion, improving the accuracy of subsequent visual matching. It also addresses the scale blur issue in monocular vision, enhancing robustness in dynamic environments.

[0161] This embodiment employs a tightly coupled fusion approach to jointly optimize observation data from both visual and lidar systems, constructing a globally consistent nonlinear optimization problem. It adopts the V-LOAM approach, specifically by setting a high-frequency visual odometry (e.g., 10-30Hz) to estimate camera pose at the front end. In the middle stage, visual pose is used to correct motion distortion in the lidar point cloud. At the back end, ICP (Iterative Closest Point) matching is used to match the corrected point cloud, estimating a low-frequency (1-10Hz) but high-precision lidar pose.

[0162] This embodiment also employs a closed-loop optimization operation, specifically combining the bag-of-words (BoW) model and laser point cloud features to perform loop closure detection and correct the global trajectory.

[0163] This embodiment adopts the DVL0 scheme and proposes a bidirectional structure-aligned local-global fusion network, which treats image pixels as pseudo-points and laser points to achieve fine-grained feature fusion.

[0164] On the KITTI dataset, the mean relative pose error (RPE) is reduced by 62% compared to a purely vision-based approach, fully leveraging the complementarity of the two sensors to achieve centimeter-level positioning accuracy.

[0165] In this embodiment, the hierarchical registration strategy employs a combination of coarse and fine registration. The coarse registration uses FPFH feature matching, while the fine registration combines ICP-RANSAC with NDT and anisotropic smoothing constraints.

[0166] In this embodiment, path planning includes repair path formulation, which specifically includes multi-scenario simulation and contingency planning. Multi-scenario simulation: testing different repair paths in a 3D model.

[0167] Emergency response plan: Develop alternative plans to address the risks.

[0168] Example 2

[0169] This invention also discloses a method for using a three-dimensional reconstruction system for immovable cultural relics, comprising the following steps:

[0170] 1) Use a 3D scanner to obtain 3D data of the artifacts;

[0171] 2) Perform denoising, registration, and segmentation preprocessing operations on the acquired 3D data;

[0172] 3) Use deep learning algorithms to extract key features of antiquities.

[0173] 4) Based on the extracted features, a virtual restoration model is constructed using NeRF technology to simulate the restoration process of antiquities and evaluate restoration plans;

[0174] 5) Combining SLAM technology to achieve precise positioning and navigation at the restoration site, optimize restoration paths, and assist restorers in making scientific decisions;

[0175] 6) Output the repair results in the form of 3D models and repair reports, and combine AR / VR technology to provide an immersive display and interactive experience.

[0176] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0177] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.

[0178] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0179] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0180] The above description is only a preferred embodiment of the present invention and does not limit the scope of patent protection of the present invention. Any equivalent structural transformations made using the present invention specification, whether directly or indirectly applied to other related technical fields, are similarly included within the scope of protection of the present invention.

Claims

1. A three-dimensional reconstruction system for immovable cultural relics, characterized in that, include: A data acquisition module; the data acquisition module is used to acquire the three-dimensional data of the artifact; Feature extraction module; The feature extraction module is used to extract key features of the artifacts using deep learning algorithms. A restoration simulation module; the restoration simulation module is used to construct a virtual restoration model based on extracted features, simulate the restoration process of antiquities, and evaluate restoration plans; and An optimization and decision-making module; the optimization and decision-making module is used to combine SLAM technology to achieve accurate positioning and navigation at the repair site, optimize the repair path, and assist repair technicians in making scientific decisions.

2. The three-dimensional reconstruction system for immovable cultural relics as described in claim 1, characterized in that, It also includes a preprocessing module, which is used to perform denoising, registration, and segmentation preprocessing operations on the acquired 3D data.

3. The three-dimensional reconstruction system for immovable cultural relics as described in claim 2, characterized in that, The denoising process includes statistical filtering and bilateral filtering; the registration operation includes using SAC-IA or FPFH feature matching to obtain an initial transformation matrix.

4. The three-dimensional reconstruction system for immovable cultural relics as described in claim 2, characterized in that, The registration operation employs the ICP algorithm, which includes iterative optimization, convergence conditions, and global optimization. The iterative optimization accelerates the nearest point search using a KD-tree and solves for the optimal transformation using SVD. The convergence conditions specifically involve setting an error threshold or a maximum number of iterations. Global optimization uses Pose Graph optimization to eliminate accumulated errors.

5. A three-dimensional reconstruction system for immovable cultural relics as described in claim 1, 2, 3, or 4, characterized in that, It also includes an output and display module, which is used to output the repair results in the form of a 3D model and a repair report.

6. A three-dimensional reconstruction system for immovable cultural relics as described in claim 1, 2, 3, or 4, characterized in that, The optimization and decision-making module incorporates biological vision mechanisms, physical modeling, and cross-scale feature fusion. The cross-scale feature fusion strategy includes a pyramid attention module, a frequency domain enhancement method, and a dynamic environment adaptation scheme.

7. A three-dimensional reconstruction system for immovable cultural relics as described in claim 1, 2, 3, or 4, characterized in that, The specific operations for constructing the virtual restoration model include using NeRF technology to construct a neural radiation field model of the original form of the artifact and using virtual restoration to simulate the effects of different restoration schemes in the NeRF model.

8. A three-dimensional reconstruction system for immovable cultural relics as described in claim 1, 2, 3, or 4, characterized in that, The key features of the artifacts include geometric features, material properties, and texture features.

9. A three-dimensional reconstruction system for immovable cultural relics as described in claim 8, characterized in that, The geometric features are the curvature and depth variations of the damaged area; the material features are the spectral reflectance; and the texture features are the texture direction and density of historical traces.

10. A method for using a three-dimensional reconstruction system for immovable cultural relics, characterized in that, The immovable cultural relic three-dimensional reconstruction system is the immovable cultural relic three-dimensional reconstruction system according to any one of claims 1 to 9, and includes the following steps: 1) Use a 3D scanner to obtain 3D data of the artifacts; 2) Perform denoising, registration, and segmentation preprocessing operations on the acquired 3D data; 3) Use deep learning algorithms to extract key features of antiquities. 4) Based on the extracted features, a virtual restoration model is constructed using NeRF technology to simulate the restoration process of antiquities and evaluate restoration plans; 5) Combining SLAM technology to achieve precise positioning and navigation at the restoration site, optimize restoration paths, and assist restorers in making scientific decisions; 6) Output the repair results in the form of 3D models and repair reports, and combine AR / VR technology to provide an immersive display and interactive experience.

Citation Information

Cited By

  • Cultural relic restoration decision-making method and system based on computer three-dimensional modeling

    CN121256880A

  • Computer three-dimensional modeling-based cultural relic restoration decision method and system

    CN121256880B

  • Method for reasoning and reconstructing original appearance of decorative component of damaged ancient building

    CN121746887A

  • Method for reasoning and rebuilding original appearance of damaged decorative components of ancient buildings

    CN121746887B

  • Weathering monitoring and demonstration method and device for immovable cultural relics and storage medium

    CN122156503A