Automatic disease detection and repair guiding system for multi-spectral three-dimensional reconstruction of cultural relics

By establishing a unified fusion model of multispectral data and three-dimensional geometry, the problem of inaccurate correspondence between disease detection results and three-dimensional geometric surfaces was solved, enabling accurate identification of disease boundaries and automatic planning of repair paths, thus improving the automation and safety of cultural relic protection.

CN121366266BActive Publication Date: 2026-03-31CHONGQING UNIV ARCHITECTURAL PLANNING & DESIGN RES INST CO LTD +4
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, there is a lack of a unified fusion mechanism between multispectral data and three-dimensional geometric data, which makes it impossible to accurately correspond the disease detection results with the three-dimensional geometric surface, and it is difficult to accurately distinguish between different types of diseases. In particular, misjudgment is prone to occur in complex backgrounds, affecting the accuracy and safety of repair operations.

Method used

By acquiring visible light, near-infrared light, ultraviolet-induced fluorescence, and structured light depth data, joint calibration and registration are performed to establish an apparent geometric fusion model of reflectivity and geometric consistency, generating a homogeneous appearance map of the area. The disease boundary is corrected using geometrically normal abrupt edges and curvature ridges, and a disease heat map is generated by combining a depth network to plan the repair pose guidance path.

Benefits of technology

It achieves deep fusion of multimodal data, improves the accuracy and boundary fit of disease detection, provides direct guidance for repair paths, reduces human error, and enhances the automation level and safety of cultural relic disease detection and repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366266B_ABST
    Figure CN121366266B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of image processing, and particularly relates to a disease automatic detection and repair guiding system for cultural relic multispectral three-dimensional reconstruction, which comprises a spectrum acquisition and model fusion unit, an automatic detection unit and a repair pose guiding unit; wherein the spectrum acquisition and model fusion unit is used for forming an apparent geometric fusion model with consistent reflectivity and geometry; the automatic detection unit is used for mapping multispectral images in a unified coordinate system to a three-dimensional surface, and outputting a disease heat map attached to the three-dimensional surface; and the repair pose guiding unit is used for generating a pose guiding map according to the disease heat map, planning a path for avoiding a fragile area on the three-dimensional curved surface by using an A-star search algorithm, and outputting a repair pose guiding sequence. The application can realize deep fusion of multi-modal data, and improve the accuracy and boundary fitting of disease detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, specifically relating to an automatic detection and repair guidance system for multispectral three-dimensional reconstruction of cultural relics. Background Technology

[0002] In the field of cultural relic conservation, 3D digitization and multispectral imaging have become important tools for research and restoration. Current technologies commonly employ 3D laser scanning, structured light scanning, or photogrammetry to model the shape of cultural relics and then use visible, infrared, or ultraviolet light imaging to reveal surface features that are difficult to observe with the naked eye. For example, visible light images can reflect the surface color and pigment residue of cultural relics, near-infrared light can penetrate part of the pigment layer to reveal underlying cracks, and ultraviolet-induced fluorescence can distinguish organic residues, repair adhesives, or paint layers. These methods can provide valuable information to researchers at individual stages and have become standard application procedures in some museums, archaeological excavation sites, and laboratories.

[0003] However, existing technologies still have several significant problems. First, there is often a lack of a unified fusion mechanism between multispectral data and 3D geometric data. Most systems only analyze damage information on a 2D image plane or attach texture maps separately to 3D models, lacking rigorous geometric calibration and registration. This results in a lack of precise correspondence between damage detection results and 3D geometric surfaces, making it difficult to use directly in actual repair operations, especially when the construction location and scope need to be clearly defined, leading to significant deviations. Second, most damage detection methods rely on traditional image processing techniques. For example, crack identification is usually accomplished through edge detection or threshold segmentation, while stain detection relies on grayscale differences or texture analysis. These methods are prone to misjudgment in complex backgrounds, with drastic color changes, or in the presence of lighting interference. Furthermore, damage types often overlap; for example, pigment chalking areas may exhibit reduced fluorescence under ultraviolet light while displaying bright granular features under visible light. Such complex situations are difficult to accurately distinguish using a single algorithm. Summary of the Invention

[0004] Therefore, the main objective of this invention is to provide an automatic detection and restoration guidance system for multispectral 3D reconstruction of cultural relics. By acquiring visible light, near-infrared light, ultraviolet-induced fluorescence, and structured light depth data, and through joint calibration and registration, an apparent geometric fusion model with reflectivity and geometric consistency is established. Based on this, a homogeneous apparent map of the curved surface area is generated as the basic unit. A consistent map of the area is obtained through reverse calculation and majority consistency rules. Then, boundary traction and correction are performed using geometrically normal abrupt edges and curvature ridges to obtain accurate boundary information for the affected area. Finally, a heat map of the affected area is generated using decision rules and a depth network for verification. Finally, the system constructs a pose guidance map based on the heat map of the affected area and uses an A* search algorithm based on mesh edge counting on the 3D surface to plan a minimum-intrusion construction path that avoids vulnerable areas, outputting a restoration pose guidance sequence. This invention enables deep fusion of multimodal data, improving the accuracy and boundary fit of disease detection. At the same time, it provides directly executable paths and pose guidance for restoration operations, avoiding human error and reducing interference with non-disease areas, thereby significantly improving the automation, repeatability and safety of cultural relic disease detection and restoration.

[0005] The technical solution adopted in this invention is as follows:

[0006] An automated detection and restoration guidance system for multispectral 3D reconstruction of cultural relics includes: a spectral acquisition and model fusion unit, an automatic detection unit, and a restoration pose guidance unit. The spectral acquisition and model fusion unit acquires visible light, near-infrared light, ultraviolet-induced fluorescence, and structured light depth data. After joint calibration, coarse and fine registration, and global consistency processing, it forms an appearance geometric fusion model with reflectivity and geometric consistency. The automatic detection unit, using curved surface regions as basic processing units, maps multi-view, multispectral images onto a 3D surface in a unified coordinate system, generating a homogeneous appearance map of the region. It then calculates the pixel values ​​of each region in reverse order. The system identifies the hit locations on the 3D surface and counts them in the evidence stack, forming a region consensus map according to the majority consensus rule. Using geometrical abrupt changes and curvature ridges as traction, the boundary of this consensus map is mapped onto the 3D geometric surface and its orientation is corrected to obtain a region disease boundary that closely adheres to the 3D geometry. Based on this correction result, disease candidate regions are generated according to the judgment rules of cracks, flaking, stains, and powdering, and the intersection is verified using U-Net or SegFormer networks to output a disease heatmap attached to the 3D surface. A repair pose guidance unit is used to generate a pose guidance map based on the disease heatmap, plan a path to avoid vulnerable areas on the 3D curved surface using the A* search algorithm, and output a repair pose guidance sequence.

[0007] Furthermore, the specific processing flow of the spectral acquisition and model fusion unit includes: completing the joint calibration of the imaging unit and the structured light projection unit to determine the camera's intrinsic and extrinsic parameters and the projection geometry; extracting and matching SIFT or SURF features on multi-view images, using a random consistency algorithm to eliminate mismatches, and obtaining the initial poses of each viewpoint; using a point-to-surface iterative nearest-point algorithm to eliminate residual alignment errors, and using a factor graph optimization algorithm to globally constrain the poses of all viewpoints to obtain a consistent 3D surface; performing occlusion culling in a unified coordinate system, mapping the multispectral image pixel by pixel onto the 3D surface to form an appearance atlas corresponding one-to-one with the surface elements; performing region growth on the 3D surface according to normal variation and curvature continuity to obtain a set of surface regions, and generating a unique region identifier for each surface region; the set of surface regions and the appearance atlas together constitute the appearance geometry fusion model.

[0008] Furthermore, the automatic detection unit generates a response map for each viewpoint and each spectral channel in the following order: edge-preserving smoothing is performed on the input image to retain boundary details and reduce isolated noise; a set of strong boundaries and a set of weak boundaries are extracted from the processed image, where the strong boundaries are used for crack and flaking detection, and the weak boundaries are used for stain and chalking detection; connected regions are labeled based on a joint strategy of four-connectivity and eight-connectivity to generate a connected component index; and local fragments that stand out only in a single channel are removed by utilizing the intersection of stable regions between different spectral channels under the same viewpoint to obtain the response map.

[0009] Furthermore, the process of the automatic detection unit generating a homogeneous appearance map of a region includes: for each curved region, its surface elements are transformed to images from various viewpoints according to a unified coordinate system to obtain the region's visible field of view; within the region's visible field of view, the stable pixel set of each spectral channel is counted, where a stable pixel is defined as a pixel that is simultaneously located in the neighborhood of a strong boundary or a weak boundary and appears consecutively in three adjacent viewpoints; the area of ​​the connected component formed by the stable pixel, the boundary sharpness, and the degree of conformity with the region boundary are compared and sorted in sequence to obtain a channel validity sequence; starting from the beginning of the channel validity sequence, mutual cropping consistency checks are performed one by one; if two channels have conflicting boundary directions near the region boundary, only the channel that is more consistent with the region boundary is retained until there is no conflict, forming a region-selected channel set; in this region-selected channel set, the cumulative histogram alignment is first performed on the response map of each channel, and then the homogeneous appearance map of the region is generated by stitching them together inside the region according to the seamless stitching rule of connected components.

[0010] Furthermore, the process of the automatic detection unit forming a region consistency map specifically includes: for each pixel in the homogeneous appearance map of the region, using the camera geometry in a unified coordinate system, calculating the corresponding facet on the three-dimensional surface, recording the hit facet number, and forming a pixel-facet hit list; establishing an evidence stack based on the curved surface region, incrementing the count value of the corresponding facet for each pixel-facet hit record, and simultaneously recording the source viewpoint and spectral channel of the hit; for the same facet, if the number of hit records from different viewpoints exceeds half of the total number of hits and their source channels are all in the region gating channel set, then the facet is marked as cross-viewpoint consistent; the facets marked as cross-viewpoint consistent form the region consistency map.

[0011] Furthermore, the process of obtaining the disease boundary of the automatic detection unit includes: extracting normal abrupt edges and curvature ridges on the three-dimensional surface to form a set of feature edges; starting from the boundary point of the area consistency map, growing along the nearest feature edge, and preferentially selecting edge segments consistent with the direction of the area boundary loop at the grid boundary to generate a boundary traction polyline; mapping the boundary traction polyline from three-dimensional space to a two-dimensional homogeneous appearance map of the area to form a boundary correction field; if the boundary correction field and the area consistency map overlap but have inconsistent directions in a certain grid cell, then aligning the boundary of the cell to the feature edge direction along the path with the fewest grid edges; thus obtaining the disease boundary of the area that strictly conforms to the geometric features.

[0012] Furthermore, the process of automatically generating candidate regions for defects by the detection unit includes: For cracks, the determination rule is as follows: within the homogeneous appearance image of the area, the anchored boundary is refined to obtain a skeleton. If the skeleton is elongated and the difference in grayscale values ​​of the pixels on both sides in the visible and near-infrared channels is greater than a preset first threshold, and the skeleton is roughly parallel to the normal abrupt change edge, then it is marked as a crack candidate; For scale buildup, the determination rule is as follows: in the stable response of the ultraviolet-induced fluorescence channel, closed or semi-closed boundary loops are searched. If the change in surface normal in the area inside and outside the loop is greater than a preset second threshold and accompanies the curvature ridge line, then it is marked as a scale buildup candidate; For stains, the determination rule is as follows: The rules are as follows: A blurred image is generated from the homogeneous appearance image of the area and the difference is made with the original image. If the gray value of the obtained low-frequency area in the near-infrared channel is lower than the preset third threshold and the boundary does not significantly overlap with the feature edge, it is marked as a stain candidate. The judgment rule for pigment chalking is: the intersection of the bright fine-grained area in the visible light channel and the area where the response value of ultraviolet-induced fluorescence is lower than the preset fourth threshold is calculated. If the intersection is distributed in spots within the area and does not extend along the feature edge, it is marked as a chalking candidate. Then, opening and closing operations and hole filling are performed on each candidate area to remove isolated pieces that only contact the boundary loop of the area with the corner point, thus obtaining the disease candidate area.

[0013] Furthermore, the automatic detection unit uses the candidate disease region as a priori mask to input into the inference process of the U-Net or SegFormer network, and takes the intersection of the priori mask and the network output mask as the final disease pixel set; it summarizes the number of surface element hits of the final disease pixel set according to the surface area and normalizes it to a preset level range to obtain a disease heat map attached to the three-dimensional surface.

[0014] Furthermore, the execution process of the repair pose guidance unit specifically includes: calculating the median direction of the normal set of each surface region as the tool approach direction, and using the tangent of the region boundary loop as the tool sliding direction, together forming a pose guidance map; on the three-dimensional surface, the lowest level region of the disease heat map is selected as the preferred channel, and the A-star search algorithm based on grid edge counting is used to select paths in the following lexicographical priority order: the first priority is to avoid high-level disease regions, the second priority is to travel along grid edges with smaller curvature, and the third priority is to keep the direction change of adjacent path segments to a minimum; the obtained path is discretized into a path point sequence according to the grid vertex order, and the approach direction and sliding direction in the pose guidance map are added to each path point to form a repair pose guidance sequence.

[0015] By employing the above technical solutions, this invention achieves the following beneficial effects: By constructing an apparent geometric fusion model with reflectivity and geometric consistency in a unified coordinate system, this invention realizes the association of multi-view, multispectral information with three-dimensional geometric information on the same surface, avoiding positional drift and misindication caused by switching back and forth between two-dimensional images and three-dimensional models in traditional methods. Using curved surface regions as basic processing units, a homogeneous apparent image of the region is first generated. Then, the hit position of each pixel on the three-dimensional surface is calculated in reverse, and evidence stack counting is performed. A region consistency image is formed according to the majority consensus rule, suppressing occasional errors caused by occlusion, specular reflection, and local blurring at the acquisition level, and improving cross-view consistency. Subsequently, using geometrically abrupt changes in the normal direction and curvature ridges as traction, the boundary of the region consistency image is mapped to the three-dimensional geometric surface and its orientation is corrected, resulting in a region defect boundary that closely adheres to the real structure. This achieves strong constraint coupling between appearance and geometry, reducing the diffusion impact of missegmentation on subsequent operations. Based on this, candidate regions for defects are generated according to deterministic judgment rules for cracks, flaking, stains, and powdering. The intersection of these regions is then verified using U-Net or SegFormer, outputting a defect heatmap attached to the 3D surface. This allows for a natural transition of defect representation from the image domain to the construction domain. Finally, a pose guidance map is generated based on the defect heatmap. An A* search algorithm based on mesh edge counting is used to plan paths that avoid vulnerable areas on the 3D surface, outputting a repair pose guidance sequence that achieves a balance between minimal intrusion and sufficient coverage. The entire process does not rely on weights or historical data; decisions are executed in a fixed order, facilitating auditing and review. By combining stable processing of local areas with global consistency, it can depict minute defects with fine granularity while maintaining overall coordinate uniformity on large-scale objects. This reduces human error, minimizes the risk of accidentally touching non-defective areas, and improves the repeatability, traceability, and on-site adaptability of the repair work. Attached Figure Description

[0016] Figure 1 A schematic diagram of the system structure of the automatic detection and repair guidance system for multispectral three-dimensional reconstruction of cultural relics provided in an embodiment of the present invention;

[0017] Figure 2 This is a schematic diagram illustrating the core mechanism by which the automatic detection unit forms a region consistency map according to an embodiment of the present invention.

[0018] Figure 3 This is a schematic diagram of the face element hit record provided in an embodiment of the present invention. Detailed Implementation

[0019] All features disclosed in this specification, or all steps in all disclosed methods or processes, may be combined in any way, except for mutually exclusive features and / or steps.

[0020] Any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.

[0021] refer to Figure 1 1. An automatic detection and repair guidance system for multispectral three-dimensional reconstruction of cultural relics, characterized in that the system includes: a spectral acquisition and model fusion unit, an automatic detection unit, and a repair pose guidance unit.

[0022] The spectral acquisition and model fusion unit acquires visible light, near-infrared light, ultraviolet induced fluorescence and structured light depth data. After joint calibration, coarse and fine registration and global consistency processing, it forms an apparent geometric fusion model with reflectivity and geometric consistency.

[0023] In one embodiment, the imaging unit and the structured light projection unit are fixed on a rigid support, with their optical axes forming an angle of approximately 20° and a baseline distance of approximately 0.25 meters. This angle ensures sufficient parallax for the structured light fringes, facilitating the calculation of structured light depth data, while avoiding excessive tilt angles that could cause shadow expansion. Visible light illumination, near-infrared light illumination, and ultraviolet-induced fluorescence illumination are applied in front of the artifact surface. The color temperature of the visible light illumination is controlled between 5000 and 6500 Kelvin, the center wavelength of the near-infrared light illumination is 850 nm, and the center wavelength of the ultraviolet-induced fluorescence illumination is 365 nm. The three types of illumination are applied at different times to avoid mutual interference.

[0024] The imaging unit acquires images in the following order: visible light → near-infrared light → ultraviolet-induced fluorescence → structured light fringe sequence. Visible and near-infrared light are acquired first because these two types of illumination provide relatively stable surface reflection from pigments and stains. Ultraviolet-induced fluorescence is acquired later to reduce fluorescence decay caused by prolonged exposure. Structured light fringe is acquired last to avoid strong stripes leaving texture residue in previous images. At each viewing angle, after acquiring the three types of spectral images and the structured light fringe sequence, the support is moved to the next viewing angle. The rotation step between adjacent viewing angles is controlled between 8 and 12 degrees, covering at least 12 global viewing angles. This step size ensures sufficient overlap of feature points in adjacent viewing angles, thereby improving the stability of subsequent coarse registration. During acquisition, a standard white board and a standard black board are placed at the edge area of ​​the artifact. The white board has a reflectance close to 1, and the black board has a reflectance close to 0. These are used for subsequent reflectance normalization and dark level correction, ensuring the reflectance of the apparent geometric fusion model is comparable.

[0025] One set each of high-contrast checkerboard and dotted targets, measuring 300 mm x 300 mm and 200 mm x 200 mm respectively, were used. The checkerboard target was used for visible and near-infrared calibration; fluorescent strips were affixed to the black and white boundaries of the checkerboard to ensure visibility under UV-induced fluorescence illumination for UV-induced fluorescence calibration. Three types of illumination were sequentially applied, allowing the calibration targets to occupy different positions and angles within the imaging area, and at least 20 calibration images were acquired for each. Target coordinates were obtained through corner and center point detection, and the focal length, principal point position, and distortion parameters of the imaging unit were calculated to ensure that the reprojection error of each target point was less than 0.3 pixels. Controlling the reprojection error within this range significantly reduces geometric drift when projecting back onto the 3D surface. A multi-frequency structured light sequence composed of Gray code and phase-shifted fringes was played, with the fringe period decreasing from coarse to fine, and the sequence length set to 48 to 72 frames. The fringe index and phase of each pixel were obtained through fringe decoding, thus establishing a dense correspondence between the imaging unit pixels and the projection coordinates. Combining the aforementioned intrinsic parameters, the relative pose between the imaging unit and the structured light projection unit is obtained. Using multi-frequency combination effectively suppresses ambiguity caused by fringe repetition: after the coarse period determines the approximate position, the fine period is selected only within a local range, significantly reducing the probability of misselection of the fringe index. The calibration target is observed in at least six different poses, and the extrinsic parameters of each viewpoint relative to the reference viewpoint are calculated using 3D intersection, ensuring that the average baseline orientation angle error is less than 0.2 degrees and the average translation error is less than 0.5 millimeters. This accuracy ensures that the initial convergence region for feature matching during subsequent coarse registration is sufficiently large.

[0026] With all illumination and structured light projection turned off, 10 dark-field images were acquired and averaged to create a dark level map. The purpose of acquiring dark-field images was to eliminate sensor-fixed pattern noise and prevent the misinterpretation of non-artifact signals as true reflections during reflectivity recovery. A standard diffuse reflective surface was uniformly illuminated using visible light, near-infrared light, and ultraviolet-induced fluorescence, and flat-field images were acquired and normalized to 200 gray levels. Each acquired image was first subtracted from the dark level map and then divided by the corresponding flat-field image to correct for illumination non-uniformity. 200 gray levels were chosen as the target value to balance the dynamic range of highlights and shadows and reduce saturation and clipping. A standard white board area was selected, its average gray level was calculated, and the entire image was scaled proportionally to map the average gray level of the white board area to 200 levels. The zero point was verified to be close to levels 0 to 5 using a standard black board area; if it deviated, both the zero point and gain were adjusted simultaneously. Through dual-point constraints on the white board and black board, imaging from various viewpoints and spectra can be unified to the same reflectivity scale without introducing any formulas.

[0027] SIFT or SURF features are extracted from visible light images view-by-view, with at least 2000 feature points obtained for each image. Bidirectional nearest neighbor matching is performed on features between adjacent viewpoints, and a random consensus algorithm is used to remove erroneous matches, typically between 50% and 70%. The retained matches are used to estimate the initial pose relationship between adjacent viewpoints. The poses of adjacent viewpoints are concatenated in the order of acquisition to form a global initial pose map. The image with the most features is selected as the reference viewpoint to reduce the cumulative error caused by reference drift.

[0028] Point clouds are reconstructed from structured light depth data at each viewpoint. Using one viewpoint's point cloud as a reference, the point clouds from other views are sequentially compared to the reference point cloud using an iterative nearest-neighbor algorithm. The nearest neighbor search uses triangular mesh normals for surface constraint matching to avoid slippage in sparse regions during point-to-point matching. The iteration terminates when the average distance change is less than 0.02 mm or the number of iterations reaches 30. Point-to-surface matching is chosen over point-to-point matching because artifact surfaces often contain large areas of weak texture; the normal provides stable geometric constraints, allowing registration to converge even in the absence of local details. A sparse factor map is constructed using the poses of each viewpoint as nodes and the residuals between the relative poses of adjacent viewpoints and the output of the iterative nearest-neighbor algorithm as constraints. A robust kernel is used to suppress the influence of a small number of distorted viewpoints, ensuring that the global error converges along both the baseline direction and the normal direction. The termination condition is when the pose change of any node is less than 0.001 degrees and 0.02 mm, or the number of iterations reaches 50. By using global constraints in the factor graph, local alignment errors can be evenly distributed across all viewpoints, eliminating the cumulative error of a single link.

[0029] All structured light depth data from all perspectives were unified to a reference coordinate system and fused using a voxel stack, with voxel side lengths set to 0.2 mm to 0.5 mm. The number of observations for each voxel was counted, and the depth distribution along the camera ray direction was recorded. In the presence of translucent enamel layers or highly reflective regions, the same ray may correspond to multiple depth peaks. For the depth histogram distribution of each ray, the dominant peak closest to the imaging unit was selected as the target depth range. This approach prioritizes preserving the first effective reflection from the truly visible surface, avoiding the misrepresentation of distant peaks generated by multiple reflections or subsurface scattering as surface depth, thereby reducing the virtual thickness and double-layer effect of the reconstructed surface. For each triangular facet, an occlusion test was performed along the line of sight of the imaging unit. If a closer facet existed along this line of sight, the occluded facet was marked as invisible at that perspective and would not participate in the fusion of that perspective during subsequent appearance mapping. This rule avoids appearance contamination caused by over-boundary smearing at edge locations.

[0030] For each triangular facet, its observation angle and distance at various viewpoints are statistically analyzed, prioritizing viewpoints with observation angles less than 60 degrees and distances between 0.8 meters and 1.2 meters. The order of selecting angle first and then distance is to reduce cross-facet aliasing caused by squint. Within the set of qualified viewpoints, the apparent values ​​for that facet are selected in the following order: first priority is the smallest observation angle; if there are ties, the second priority is the viewpoint most frequently observed when the line of sight passes through voxel fusion; if still ties, the third priority is the viewpoint with the highest local sharpness score. Using a deterministic order of progressive priority, rather than a weighted average, avoids introducing cross-viewpoint blur bands at boundaries. The selected pixel values ​​are written back to the facet attributes using reflectance scales for each of the three spectral types, forming a reflectance description corresponding to each facet. A random inspection of standard white and blackboard areas is performed; if the average value of the whiteboard deviates from 200 levels by more than 5 levels or the average value of the blackboard is higher than 5 levels, the gain and zero point of that batch of images are recalibrated. By traversing all surface elements, reflectance values ​​for the three spectral types are assigned, generating an appearance atlas. This appearance atlas, together with the unified 3D surface, constitutes the appearance component of the appearance geometry fusion model.

[0031] For each triangular facet, the changes in the included angle of the facet's normal and the number of changes in its interior angles are statistically analyzed within a ring of its neighborhood to obtain two geometric metrics for region growing. A seed facet is randomly selected; if the included angle of the normal of an adjacent facet to the seed facet is less than 15 degrees and the number of changes in its interior angles is less than 2, it is included in the same region. Growth in that direction is stopped when adjacent facets do not meet these conditions. This threshold range ensures continuous curvature within the region and that the boundaries are located near geometric polygons, facilitating subsequent mapping and alignment using curved surface regions as basic processing units. The boundary edge set of each region is traversed clockwise to form closed or semi-closed boundary loops for subsequent region visibility calculation and occlusion removal. Each curved surface region is sorted according to three attributes: acquisition order, region area, and boundary length, generating a unique region identifier code. The identifier code length is set to 16 to 32 bits to ensure no conflicts occur even in large-scale artifact scenes. A unified 3D surface, appearance atlas, and curved surface region set together constitute the appearance geometry fusion model. The model has consistency in reflectivity and geometry, and can be directly used by subsequent automatic detection units.

[0032] In another embodiment, in glazed or metal-inlaid areas with strong specular reflection, a linear polarizer is added and cross-polarized illumination is used to reduce the specular component, thereby improving the stability of reflectivity estimation.

[0033] In areas with very weak texture or high reflectivity, the length of the stripe sequence is increased to 96 stripes, while the finest period is shortened to improve stripe uniqueness and reduce the probability of aliasing when stripes cross boundaries.

[0034] For wall artifacts with an area exceeding 3 square meters, an electric sliding table is used to collect data in sections with a translation step of 100 mm. The above steps are repeated for each section, and cross-section constraints are introduced in the factor graph optimization algorithm stage to integrate the panoramic 3D surface.

[0035] To protect cultural relics, the illuminance of ultraviolet-induced fluorescence illumination is limited to no more than 1 watt per square meter, and the continuous irradiation time at a single point is limited to no more than 10 seconds. Time-division scanning is used to cover the entire area to avoid localized over-irradiation. In coated areas with subsurface scattering, the main depth peak closest to the imaging unit is preferentially selected as the target depth range; if multiple peaks have similar intensities, the peak consistent with most viewing angles is selected as the target. This selection ensures a stable representation of the true outer surface, avoiding the incorporation of spurious signals from internal layers into the unified three-dimensional surface.

[0036] The automatic detection unit, based on the apparent geometry fusion model, uses curved surface regions as the basic processing unit to map multi-view, multi-spectral images onto a three-dimensional surface in a unified coordinate system, generating a homogeneous appearance map of the region. It then calculates the hit position of each pixel on the three-dimensional surface and counts them in the evidence stack, forming a region consistency map according to the majority consensus rule. Using geometrical abrupt changes and curvature ridges as traction, it maps the boundary of this consistency map onto the three-dimensional geometric surface and corrects its orientation, obtaining a region disease boundary that closely adheres to the three-dimensional geometry. Based on this correction result, it generates disease candidate regions according to the judgment rules for cracks, scale, stains, and powdering, and uses a U-Net or SegFormer network to verify and take the intersection, outputting a disease heatmap attached to the three-dimensional surface.

[0037] In one implementation, the automatic detection unit reads the set of facet elements and boundary loops of the surface patch in the appearance geometry fusion model. A best-fit plane is fitted using the vertex set of this facet element set, and a two-dimensional sampling mesh is created on this plane for the surface patch, with a mesh resolution of 0.2 mm per pixel to 0.5 mm per pixel. The best-fit plane is used because the normal variation within the surface patch is limited to a small range; projecting onto this plane keeps the surface stretching within an acceptable range, facilitating the formation of a regular two-dimensional homogeneous appearance map of the patch. For each viewpoint, the triangular facet elements of the surface patch are projected onto the multispectral image of that viewpoint using a unified coordinate system to obtain the patch projection field of view. Occlusion determination is performed for each viewpoint: if there is an intersecting facet element closer to the current facet element at that viewpoint, the facet element is marked as invisible at that viewpoint, and subsequent mapping does not use pixels from that viewpoint. Occlusion determination can prevent edge-boundary overlay, thereby ensuring that the homogeneous appearance map of the patch is not contaminated by erroneous pixels. The observation angle and distance of the surface patch are calculated for each viewpoint. Prioritize retaining viewing angles no greater than 60 degrees and distances between 0.8 meters and 1.2 meters. Controlling the viewing angle first, then the distance, can significantly reduce cross-surface aliasing caused by squint, ensuring that the retained pixels mainly come from the target surface itself.

[0038] In one implementation, to ensure that the judgment rules for cracks, flaking, stains, and powdering have a directly implementable threshold setting method, a preset first threshold, a preset second threshold, a preset third threshold, and a preset fourth threshold are set before the start of each detection task according to the same deterministic process and remain unchanged throughout the entire detection task. First, three types of reference regions are selected within the same surface area in the apparent geometry fusion model: the first type of reference region is a passable area reference region with a disease heatmap level of 1 and stable boundaries; the second type of reference region is a cautious area reference region reference region with a disease heatmap level of 2 and stable connected domain area; and the third type of reference region is a vulnerable area reference region reference region with a disease heatmap level of 3 and continuous boundaries. Each type of reference region contains at least one connected domain, and the number of projection views for each connected domain is not less than three. Then, for each type of reference region, the pixel grayscale distribution is statistically analyzed on the homogeneous appearance map of the region under visible light, near-infrared light, and ultraviolet induced fluorescence. The 10th percentile, 50th percentile, and 90th percentile of the grayscale distribution are taken as the grayscale feature values ​​of that type of reference region, and the threshold determination rules are given as follows.

[0039] A preset first threshold is used to determine the "difference in grayscale values ​​between pixels on both sides in the visible and near-infrared channels" in the crack detection. It is set as the upper bound of the difference in grayscale feature values ​​between the visible and near-infrared channels in the passable reference area, used to exclude false cracks caused by normal texture fluctuations. Specifically, within the passable reference area, for each boundary pixel, three pixels are taken from the left and right sides along a direction perpendicular to the skeleton to form two side windows. The average grayscale difference between the two side windows is calculated to obtain the grayscale difference sets for the visible and near-infrared channels. The larger of the 90th percentile of the visible and near-infrared grayscale difference sets is used as the initial value of the preset first threshold, and this initial value is limited to between 12 and 30. When the texture of the passable reference area is strong enough that the initial value is higher than 30, the preset first threshold is fixed at 30; when the initial value is lower than 12, the preset first threshold is fixed at 12. Through these settings, crack detection needs to achieve a contrast difference higher than normal texture fluctuations, which can reduce misjudgments caused by brush strokes and texture shadows.

[0040] A preset second threshold is used to determine the "surface normal change" in the surface roughness assessment. It is set as the boundary between the normal change distribution in the cautious and vulnerable reference regions, distinguishing normal surface undulations from raised edges. Specifically, within each reference region, taking a concentric ring of the surface patch as the range, the median value of the angle between each surface element and the normal of its neighboring surface elements is calculated as the normal change of that surface element. These are then aggregated to obtain the set of normal changes for that reference region. The 90th percentile of the set of normal changes in the cautious reference region is denoted as value A, and the 10th percentile of the set of normal changes in the vulnerable reference region is denoted as value B. The smaller of these two values ​​is used as the initial value of the preset second threshold, which is limited to between 6 and 18. When the initial value is less than 6, the preset second threshold is fixed at 6; when the initial value is greater than 18, the preset second threshold is fixed at 18. This setting makes roughness candidates more likely to appear in locations with significant geometric changes and local raised edges, thereby reducing false detections caused by slow undulations.

[0041] A preset third threshold is used in stain detection to determine if the "grayscale value of the near-infrared channel is below the threshold." It is set as the low grayscale boundary of the passable area reference region in the near-infrared light channel, used to distinguish low-frequency dark areas from normal material differences. Specifically, the near-infrared light grayscale distribution is statistically analyzed within the passable area reference region. The 10th percentile is denoted as value C, and the 50th percentile is denoted as value D. The average of these two values ​​is used as the initial value of the preset third threshold, which is limited to between 70 and 120. When the initial value is less than 70, the preset third threshold is fixed at 70; when the initial value is greater than 120, the preset third threshold is fixed at 120. By using the lower grayscale boundary of the passable area as a reference, stain detection prioritizes capturing low-frequency areas that are significantly darker than the normal substrate, thereby reducing misjudgments caused by shadows or dark local textures.

[0042] A preset fourth threshold is used in the chalking determination process to judge whether the "UV-induced fluorescence response value is lower than the threshold". It is set as the boundary between the weak response in the UV-induced fluorescence channel of the accessible reference area and the vulnerable reference area, distinguishing between normal fluorescence differences and weak responses caused by chalking. Specifically, the 10th percentile of the UV-induced fluorescence grayscale distribution is statistically analyzed within the accessible reference area and recorded as value E. The 50th percentile of the UV-induced fluorescence grayscale distribution is statistically analyzed within the vulnerable reference area and recorded as value F. The smaller of the two values ​​is used as the initial value of the preset fourth threshold, which is limited to between 40 and 110. When the initial value is less than 40, the preset fourth threshold is fixed at 40; when the initial value is greater than 110, the preset fourth threshold is fixed at 110. This setting makes chalking candidates more likely to fall in areas where UV-induced fluorescence is significantly weakened, while avoiding misjudging natural fluorescence differences caused by material background as chalking.

[0043] In an optional implementation, if the lighting conditions collected on-site vary significantly, resulting in a distinct bimodal distribution of grayscale within the same curved surface area, a deterministic brightness uniformity step can be added before the threshold determination process described above. This involves linearly normalizing the homogeneous appearance of the entire area using the target grayscale values ​​of a standard whiteboard in visible light, near-infrared light, and ultraviolet-induced fluorescence, ensuring that the average grayscale of the whiteboard area falls around 200 with a deviation of no more than 5. Then, the preset first threshold, the preset third threshold, and the preset fourth threshold are calculated according to the aforementioned rules. The preset second threshold remains unaffected by the normal variation set calculation. In another optional implementation, when the vulnerable area reference region is too small, resulting in an insufficient normal variation set, a cautionary area reference region can be used to replace the vulnerable area reference region in the preset second threshold calculation. The preset second threshold is then fixed between 10 and 15 to ensure that the initial assessment can still be performed and the assessment boundary is clear.

[0044] For each candidate viewpoint, the 90th percentile of the local gradient magnitude is calculated as the sharpness score for the multispectral image. The grayscale values ​​of the standard white and black areas are randomly checked to ensure they are close to level 200 and level 0 to 5, respectively. If the sharpness score is lower than 80% of the adjacent viewpoints, or the white / black deviation exceeds 5 levels, the viewpoint is discarded to avoid introducing defocusing and brightness drift into the homogeneous appearance of the area. For each retained viewpoint, strong and weak boundary sets are extracted from the visible, near-infrared, and ultraviolet-induced fluorescence channels. Pixels that appear consecutively in the neighborhood of a strong or weak boundary and in three adjacent viewpoints are grouped into a stable pixel set. Consecutive appearance is required to filter out occasional noise and single flicker, ensuring that the pixels participating in the stitching maintain consistent appearance characteristics across different viewpoints.

[0045] The automatic detection unit uses the boundary loop of the curved surface region as a reference and compares three deterministic indicators for each channel in turn: those with larger connected region area, clearer boundaries, and boundaries more consistent with the curved surface region boundary are ranked first, resulting in a channel validity sequence. Starting from the beginning of the sequence, pairwise mutual clipping checks are performed in the boundary neighborhood of the curved surface region: if the boundary directions of two channels have obvious conflicts near the same point, only the channel more consistent with the curved surface region boundary is retained, until there are no conflicts, forming a set of selected channels for the region. Using the sorting-checking-retention order, rather than weighted merging, can avoid generating fuzzy bands at the boundaries.

[0046] For each channel in the selected channel set of the surface region, the cumulative histogram is aligned using the gray-level percentile anchor point of the surface region to ensure consistent brightness structure across different viewpoints within the same channel. Then, seamless stitching is performed according to the seam direction of the connected components: first, a transition zone of 5 to 10 pixels is selected on each side of the seam, and pixels with continuous texture are selected as the master pixels along the seam direction; on the other side, points are only added when the master pixels are missing. This master-then-fill stitching method eliminates seam steps without weighted averaging. The stitched channel pixels are then filled pixel by pixel on the 2D sampling grid of the surface region, yielding three homogeneous appearance images of the region for visible light, near-infrared light, and ultraviolet-induced fluorescence. Each image corresponds one-to-one with the grid of the surface region and can be directly used for subsequent inverse calculations of its hit position on the 3D surface.

[0047] The automatic detection unit reads the corresponding two-dimensional sampling coordinates of each pixel in the homogeneous appearance image of the area in a unified coordinate system, and calculates its hit position on the three-dimensional surface in reverse along the line of sight of the imaging unit. Specifically, starting from the pixel's position on the sampling plane, it searches for the nearest triangular intersection point along the normal direction from the sampling plane to the curved surface area, recording the hit triangular cell number and the centroid position of the intersection point on that cell. Choosing the nearest intersection point prioritizes the correspondence to the real visible outer surface, avoiding mistaking distant intersection points caused by subsurface scattering or multiple reflections for the real surface, thus reducing surface false thickness and double-layer effects. An evidence stack is established for each curved surface area. For each pixel hit record, the count of the corresponding triangular cell is incremented by 1, and the source viewpoint and source channel markers are recorded. Count tables are established for the three types of channels, with counts only accumulated as integers, without any weighting or historical accumulation, ensuring that each hit has the same impact on the evidence stack.

[0048] Then, for the same surface element, the number of hits from different viewpoints is counted. If the number of hits from different viewpoints exceeds half of the total number of hits for that surface element, and all the source channels of these hits belong to the region's gating channel set, then the surface element is marked as cross-viewpoint consistent. Majority consistency means that the same surface location is repeatedly confirmed by multiple viewpoints with similar appearances, which can effectively suppress occasional misjudgments caused by single-viewpoint occlusion, specular reflection, or local blurring. All surface elements marked as cross-viewpoint consistent are placed as foreground on the 2D sampling mesh of the curved surface region, and the rest are placed as background, resulting in a region consistency map. Isolated points with an area less than 3 pixels and small holes with a diameter of no more than 5 pixels are removed from the region consistency map to improve connectivity and facilitate subsequent boundary traction and orientation correction.

[0049] In the edge band (3 to 7 pixels wide) of the region consistency map, count the number of foreground and background alternations. If the number of alternations per unit length exceeds 4, it indicates possible texture aliasing or occlusion misjudgment. Increase the mutual clipping test range to twice the edge band and regenerate the region homogeneity appearance map. For each pixel marked as cross-viewpoint consistent, record the number of views participating in the majority consensus decision. If the average number of views for a surface region is less than 3, query for supplementary views in adjacent surface regions to avoid region fragmentation due to insufficient views. Output three types of region homogeneity appearance maps, the corresponding pixel-to-pixel hit list, the evidence stack count results, and the region consistency map. These results will be used in subsequent steps for geometric feature guidance and orientation correction.

[0050] In another implementation, when the curvature of the curved surface region varies significantly, the automatic detection unit can reduce the resolution of the two-dimensional sampling grid to 0.5 mm per pixel to 1.0 mm per pixel, and change the sampling direction of the homogeneous appearance image of the region to sequentially spread along the geodesic direction of the curved surface region. Coarser sampling and geodesic spreading can reduce stretching distortion at extreme curvatures. In addition to occlusion detection, a reverse penetration check is added: the pixels of the homogeneous appearance image of the region are slightly shifted outward along the normal direction by 1 to 2 pixels, and their hit positions on the three-dimensional surface are calculated again in reverse. If the two hits fall on different facets, the pixel is removed from the stacked data to further reduce misjudgments caused by edge interleaving.

[0051] In another implementation, the automatic detection unit replaces the 90th percentile of the local gradient magnitude with the 80th percentile of the local Laplacian response, while retaining the rule of discarding images with a sharpness score below 80% of adjacent viewpoints. Both scoring methods rely on deterministic sorting and thresholds, without involving weighted fusion. When shooting conditions are limited, the lower limit for majority consistent viewpoints can be fixed at 3, meaning that cross-viewpoint consistency is only marked when at least 3 different viewpoints provide hit records for the same facet element and all come from the area gating channel set. Fixing the lower limit can maintain stability when the number of viewpoints is small, avoiding fluctuations in judgment due to changes in the total number of times. The selection of the principal pixel in seamless stitching is replaced with a priority of connected components, starting with the longest and ending with the shortest: the lengths of continuous boundary segments are compared on both sides of the seam, and the side with the longer length is selected as the source of the principal pixel, with the other side only used to fill in missing points.

[0052] Next, the automatic detection unit calculates the change in normal and the number of changes in interior angles of the ring neighborhood for each triangular element on the 3D surface, marking mesh edges with a normal angle greater than 15 degrees as abrupt changes in normal. This threshold can stably capture polylines, gaps, and peeling edges, avoiding mistaking slow undulations for abrupt changes. Within one to two ring neighborhoods, the concave and convex transition positions of adjacent elements along the principal curvature direction are statistically analyzed and connected to form continuous polylines, resulting in curvature ridges. When the distance between adjacent transitions is less than 1 mm, they are merged into a single polyline to avoid jagged edges caused by high-frequency noise. The abrupt changes in normal and the curvature ridges are combined into a feature edge set. If the distance between them is less than 0.5 mm, they are merged into a single geometric polyline segment, preserving the spatial orientation and endpoint order of the polyline segment for subsequent traction along the existing structure. This concentrates high-confidence geometric boundaries on a small number of clear polyline segments, reducing ambiguity in subsequent traction.

[0053] The automatic detection unit extracts the outer boundary of the homogeneous appearance map of each surface region, obtaining a boundary loop composed of pixel-level polylines. Based on the pixel-to-surface correspondence in a unified coordinate system, each pixel in the boundary loop is assigned to its corresponding triangular facet, and the centroid position and normal orientation of the pixel on that facet are recorded, forming a preliminary draft of the 3D polyline. Visibility verification is performed on line segments between adjacent mapping points: if a line segment crosses an occlusion area or enters a facet outside the current surface region, it is segmented at the crossing point, and the segment crossing the boundary is discarded, retaining only the segment within the current surface region. This step ensures that subsequent traction is always constrained to the truly visible 3D geometric surface.

[0054] The automatic detection unit uses each vertex of the initial 3D polyline draft as its center, setting a spatial search radius of 1 mm to 3 mm. It searches the feature edge set for the polyline segment closest to that vertex, selecting it as a candidate traction segment. This radius covers the actual geometric polylines near the boundary while avoiding attachment to distant, irrelevant structures. For each candidate traction segment, the angle between its tangent and the tangent of the current polyline segment is calculated. If the angle is no greater than 20 degrees, it is considered to have the same orientation. If multiple candidate traction segments have the same orientation, they are selected according to the following deterministic priority: first, the closest segment; second, the segment with the longest continuous polyline length; and third, the segment whose endpoint is more consistent with the loop direction of the area boundary. This order of consistency, distance, continuity, and loop direction uniquely determines the traction target without any weighting. The candidate traction segment with the same orientation replaces the line segment of the initial 3D polyline draft within the corresponding interval, retaining 1 to 2 original vertices at each end of the replaced segment as transition points to avoid sudden reversals. The process proceeds segment by segment along the boundary loop until the entire loop is processed, resulting in a coherent boundary traction polyline. If there is a gap between adjacent replacement segments with a gap length of no more than 2 mm, the gap is connected along the feature edge set using the minimum number of edges. If the gap exceeds 2 mm, it is retained and used as a discontinuity indicator in subsequent judgments to prevent over-extrapolation. The boundary traction polyline is mapped from 3D space back to the homogeneous appearance map of the surface region using a unified coordinate system. A strip-shaped region one to two pixels wide is generated at the pixel location covered by the polyline, serving as the boundary correction field. The boundary correction field is superimposed on the region consistency map. If the two overlap in a certain grid cell but their orientations differ significantly, the boundary of the region consistency map is rewritten as the orientation of the boundary correction field along the path with the fewest grid edges within that cell. This process is repeated until conflict-free cells are reached, resulting in a region defect boundary that conforms to the geometric structure.

[0055] Within the same curved surface area, candidate disease regions are generated for different types using the following steps, with all gray levels ranging from 0 to 255 after reflectance normalization: The disease boundary of the area is refined pixel-level to obtain a skeleton; if the skeleton is elongated and the gray level difference between its two sides in visible light and near-infrared light is not less than 20 and 15 levels respectively, and the skeleton's direction is basically parallel to the normal abrupt change edge with a length not less than 30 pixels and an average width not greater than 3 pixels, then this connected region is marked as a crack candidate. The combination of elongated shape and contrast can stably correspond to narrow cracks without misreporting brushstrokes and shadows as cracks. Closed or semi-closed brightness difference rings are searched in the UV-induced fluorescence channel; if the average normal difference between the inside and outside of the ring is not less than 10 degrees, and the inner side of the ring is distributed along the curvature ridge, while the outer side shows edge deepening in the visible light channel, then this connected region is marked as a flaking candidate. The combination of fluorescence enhancement with normal abrupt change and curvature ridge can accurately indicate the flaking edge. A blurred image is generated in the visible light channel of the homogeneous surface image of the affected area and subtracted from the original image. If the average gray level of the resulting low-frequency region in the near-infrared light channel is no higher than 90, and the overlap length between its boundary and the feature edge set is less than 20% of the perimeter of the region, it is marked as a stain candidate. The near-infrared light dimming and non-overlapping with the geometric boundary can distinguish surface adhesions from structural edges. High-brightness fine-grained regions are detected in the visible light channel, and weak response regions are detected in the ultraviolet-induced fluorescence channel. If the intersection of the two is a patchy distribution, with a single spot diameter between 1 and 3 pixels, and the overall distribution does not extend along the feature edge set and the density is no less than 5 spots per 100 square pixels, it is marked as a powdering candidate. The combination of fine-grained brightness and fluorescence reduction can characterize the weak binding state of powdering. Opening and closing operations of 3 to 5 pixels and hole filling with a diameter no greater than 5 pixels are performed on each candidate, and isolated small patches that only touch the boundary of the lesion area at the corner are removed to obtain connected and clearly defined lesion candidate areas.

[0056] The automatic detection unit combines the visible light, near-infrared light, and ultraviolet-induced fluorescence images of the homogeneous appearance of the region into a three-channel input in channel order. The input is then divided into 512 pixel x 512 pixel blocks with a 32-pixel overlap between blocks and zero-padding around the blocks. This block division ensures consistent network input resolution across large curved surfaces. Semantic segmentation maps are obtained using either U-Net or SegFormer inference. The outputs of adjacent blocks are seamlessly combined into a complete segmentation map in a fixed order, first covering the center and then the edges. The intersection of the complete segmentation map with candidate lesion regions is calculated for each category, and only pixels within the intersection are considered as final lesion pixels. Fragments with an intersection area less than 20 pixels are directly discarded, and sharp corners at the intersection boundaries are trimmed with a 3-pixel radius to maintain stability in subsequent projections. Verifying the intersection ensures that the rules and network output mutually constrain each other, eliminating noise that might be generated by either alone.

[0057] Finally, the automatic detection unit maps the final defect pixels back to the 3D geometric surface using a unified coordinate system, marking the hit pixels as defect pixels and counting the number of hits from different viewpoints. For each defect pixel, it is divided into three levels based on the number of hits: Level 3 for at least 5 hits, Level 2 for at least 3 hits and less than 5 hits, and Level 1 for at least 1 hit and less than 3 hits. This classification relies solely on integer counting to avoid introducing weighting. A majority vote is performed on the defect level within the same surface region: if a pixel has the same level as at least half of its neighbors, it remains unchanged; otherwise, it is changed to the level with the highest frequency among its neighbors. This eliminates isolated jumps without altering the judgment of the majority of pixels. The defect heatmap is rendered on the 3D geometric surface by mapping the level to three fixed color intervals, resulting in a defect heatmap attached to the 3D surface. Simultaneously, the defect coverage ratio and average level of each surface region are output for use in the repair process.

[0058] If all candidate traction segments of a boundary traction polyline fail to meet the consistency condition, the original 3D polyline segment is retained and marked with a dashed line in the boundary correction field, indicating that subsequent rules should not consider this segment as a valid source of the fracture skeleton. If a boundary traction polyline crosses the boundary of a curved area multiple times within a continuous 10 mm range, it is considered a false adsorption, and the line is reverted to the previous stable segment, with the traction radius reduced from 3 mm to 1 mm before re-traction. When the center-to-center distance of candidate regions of the same type of disease is less than 5 pixels and their connection does not affect the continuity of the skeleton, they are merged to reduce redundant annotations.

[0059] In another implementation, the automatic detection unit replaces the fixed 1 mm to 3 mm radius with an adaptive radius equivalent to the average side length of the grid, with a minimum of no less than twice the grid side length and a maximum of no more than six times, to maintain the same geometric adhesion on grids of varying thicknesses. The 20-degree radius is replaced with an optional range of 15 to 25 degrees to accommodate differences in detail intensity on different material surfaces, making it easier to adhere to the main geometric direction on surfaces with weak textures. In addition to the fluorescence brightness difference ring, a dark ring extending 3 pixels outward is searched in the visible light channel. If both rings exist simultaneously and the distance between the inner and outer rings is between 2 and 6 pixels, they are preferentially marked as candidates for edge lifting, improving the ability to identify micro-raised edges.

[0060] The repair pose guidance unit generates a pose guidance map based on the heat map of the defect, plans a path to avoid the vulnerable area on the three-dimensional surface using the A-star search algorithm, and outputs the repair pose guidance sequence.

[0061] In one implementation, the repair pose guidance unit reads the disease heatmap and the set of surface regions, marking the third-level surface elements of the disease heatmap as vulnerable areas, the second-level as cautious areas, and the first-level as passable areas. One to two buffer zones, 3 to 6 millimeters wide, are generated along the three-dimensional surface outside the vulnerable areas; areas within these buffer zones are considered impassable. This absorbs actual tool contact and positioning errors within the buffer zones, preventing the path from closely adhering to high-risk boundaries. Within each surface region, the connectivity of passable areas is checked. If there is no connection between the same disease area and an external passable area, a single transition zone, 1 to 2 pixels wide and no longer than 30 millimeters, is created within the cautious area. This transition zone is used only for entry and exit, preventing prolonged stays within the cautious area.

[0062] Then, the pose guidance unit repairs the tool approach direction within each surface patch, using the median direction of the patch's normal set as the tool approach direction. The median direction is chosen because it is insensitive to isolated spikes and stably represents the patch's main orientation, preventing sudden impacts during tool entry and exit. The tangential direction of the surface patch's boundary loop is used as the tool sliding direction. Sliding along the tangential direction of the boundary loop allows the tool's minor operations near the boundary to unfold along existing textures or seams, reducing disturbance to the entire area. The tool approach direction and tool sliding direction are written into each mesh vertex of the 3D surface to form a pose guidance map. If the angle between the tool approach directions of adjacent vertices exceeds 15 degrees, a transition direction is inserted between the two vertices, ensuring that the angle between two consecutive transition directions does not exceed 10 degrees, guaranteeing smooth posture changes and reducing instantaneous acceleration of the handheld device or robotic arm.

[0063] For each disease-affected connected domain, the nearest accessible grid vertex to the domain's boundary is selected as the entry point, and the second nearest accessible grid vertex to the domain's boundary that does not coincide with the entry point is selected as the exit point. Distance measurements are taken using the shortest grid path length on the 3D surface to avoid traversing invisible internal regions. For cracks and flaking-type diseases, service points are set every 10 mm along the skeleton line of the connected domain; for stains and pigment powdering diseases, service points are set every 15 mm along the outer boundary of the connected domain. Service points define critical locations that the path must pass through, ensuring the integrity of the repair coverage. The entry point, all service points, and exit point are arranged in the order of proximity. If equidistant points are adjacent, the point with the smaller height difference is prioritized, followed by the point with the smaller curvature, forming a visit sequence. Fixed adjacency rules eliminate the randomness of path generation.

[0064] The A-Star search algorithm based on grid edge counting (using the lowest-level area of ​​the disease heatmap as the preferred channel) includes the following process: Constructing a search graph on the 3D surface with grid vertices as nodes and grid edges as connections, the traversal cost of each grid edge is uniformly counted as one grid edge. This setup directly equates the path length to the number of grid edges, naturally corresponding to the geodesic orientation of the 3D surface. The heuristic step count is estimated deterministically: calculating the 3D straight-line distance between the current node and the target node, dividing by the average grid edge length, and rounding up to obtain the estimated step count. Straight-line distance is used because it is always no shorter than the actual geodesic direction step count, ensuring stable compression of the search range without missing solutions. The expansion order uses lexicographical priority, comparing the following three items one by one: First priority: Whether the target adjacent edge leads to a passable area. Edges leading to passable areas are prioritized; if all adjacent edges are not passable areas, only one edge is selected within the transition zone for further expansion. Second priority: Whether the included angle between the normals of the two triangular facets crossed by the adjacent edge is no greater than 10 degrees. Prioritize edges with an angle no greater than 10 degrees; if tied, the edge with the smaller angle comes first. Third priority: Does the change in direction between the added edge and the previous path edge not exceed 15 degrees? Edges with a change in direction no greater than 15 degrees are prioritized; if tied, the edge with the smaller change in direction comes first. If still tied, select the edge with the smaller heuristic step count; if still tied, select the edge with the smaller 3D straight-line distance to the target node; if still tied, select the edge with the smaller mesh vertex index. This fixed order avoids any weighted aggregation and ensures that each step's selection is unique and reproducible.

[0065] For each pair of adjacent targets in the access sequence (entry point to the first service point, service point to the next service point, last service point to exit point), the search is performed sequentially to obtain segmented paths; the beginning and end of each segment are then concatenated to form a complete path. If the search fails, the transition band width is first widened to 2 to 3 pixels within the caution zone; if it still fails, the included angle threshold for the second priority is widened from 10 degrees to 15 degrees; if it fails again, sliding on the boundary between the traversable zone and the caution zone is allowed to be no more than 20 millimeters. This progressive widening ensures that a path can always be found, while always limiting the exposure to risk.

[0066] Polyline simplification is performed on the obtained complete path: when the maximum distance between the middle vertex of three adjacent polylines and the preceding and following straight lines does not exceed 1 mm and does not intrude into the buffer zone, the middle vertex is deleted. This reduces unnecessary backtracking without changing the macroscopic direction. Minimum spacing check: the three-dimensional nearest distance from any vertex on the path to the vulnerable area must not be less than 3 mm; if this is violated, the path is translated in the opposite direction of the normal to the nearest passable vertex, and the local segment is searched again. Attitude smoothing: an attitude point is sampled every 2 mm along the path, and the tool approach direction and tool sliding direction in the pose guidance diagram are read. If the angle between the tool approach directions of two consecutive attitude points exceeds 15 degrees, a transition attitude point is inserted between them until the adjacent angle is no greater than 10 degrees. This approach avoids tool slippage at abrupt changes in curvature.

[0067] The path is discretized at equal intervals of 2 mm to 5 mm to obtain a path point sequence. Each path point includes its position, tool approach direction, tool sliding direction, and action marker. An approach action marker is added to the first path point at the entry point, processing action markers are added to path points within a 5 mm radius of each service point, and a departure action marker is added to the last path point at the exit point. Action markers are used to switch the speed and pressure settings of the construction equipment. For cracks and flaking defects, when the skeleton orientation and tool sliding direction are inconsistent, the skeleton orientation covers the tool sliding direction to ensure that construction proceeds along the main direction of the defect, reducing lateral disturbance. The output repair pose guidance sequence includes the path point sequence, the tool approach direction, tool sliding direction, and action marker for each point, and is displayed in conjunction with the defect heatmap for easy visual verification by the operator on the 3D surface.

[0068] For path points marked with machining actions, a coverage band is projected onto the 3D surface with an influence radius of 5 mm. The coverage band should completely cover the corresponding connected domain of the defect. If there are any uncovered areas, a new service point is added at the densest defect pixels in that area, and the local search is restarted and stitched together. If there are points on the path where the angle between the normal and the tool approach direction exceeds 30 degrees, a local transition path is inserted 5 mm on each side of the point to ensure that the angle does not exceed 30 degrees. When there are two or more paths with identical lexicographical priority, the one with the lower total mesh edge count is selected; if the counts are the same, the one with the smaller cumulative change in the tool sliding direction is selected to ensure that the execution time and attitude adjustment are minimized simultaneously.

[0069] Figure 2 This demonstrates the core mechanism by which the automatic detection unit in an automated detection and restoration guidance system for multispectral 3D reconstruction of cultural relics generates a consistent map of a specific area. For example... Figure 2As shown, multiple facets, including facets 001 to 009, are distributed within the curved surface area of ​​the artifact's three-dimensional surface. Each facet serves as the target location for reverse calculation of the homogeneous appearance image pixels within the area. The system establishes an evidence stack based on the curved surface area and counts and statistically analyzes the hit records from different perspectives. Hit records from perspective 1 are represented by short dashed lines, primarily hitting facets 001, 002, 003, and 006. Hit records from perspective 2 are represented by medium dashed lines, primarily hitting facets 003, 004, 005, and 007. Hit records from perspective 3 are represented by long dashed lines, covering facets 002, 003, 004, and 008. Hit records from perspective 4 are represented by dotted lines, primarily affecting facets 003, 004, 005, and 009. During the evidence stack counting process, each pixel-to-cell hit record increments the corresponding cell's count by one, while simultaneously recording the source viewpoint and spectral channel information. Cell 001 receives a count of 1, cell 002 receives a count of 2, cell 003 receives the highest count of 4, cells 004 and 005 each receive a count of 3, cell 006 receives a count of 1, and cells 007, 008, and 009 each receive a count of 2. The system uses a majority consensus rule for cross-viewpoint consistency determination; that is, for the same cell, if the number of hit records from different viewpoints exceeds half of the total hit count and their source channels are all within the area gating channel set, then the cell is marked as cross-viewpoint consistent.

[0070] refer to Figure 3 According to this rule, elements 003, 004, and 005 were marked as cross-view consistent because their hit records were 4, 3, and 3 respectively, all exceeding the majority threshold. The source viewpoint and spectral channel records detail the specific hit situation for each consistent element. The hit records for element 003 include viewpoint 1 (visible light channel), viewpoint 2 (near-infrared channel), viewpoint 3 (ultraviolet fluorescence channel), and viewpoint 4 (visible light channel), demonstrating good multispectral consistency. The hit sources for element 004 are viewpoint 2 (near-infrared channel), viewpoint 3 (visible light channel), and viewpoint 4 (ultraviolet fluorescence channel). The records for element 005 show the hit situation for viewpoint 1 (near-infrared channel), viewpoint 2 (visible light channel), and viewpoint 4 (ultraviolet fluorescence channel). The elements marked as cross-view consistent ultimately form a region consistency map. This consistency map, through rigorous multi-view verification and spectral channel cross-confirmation, ensures the reliability and consistency of data during the disease detection process, providing a stable basic data structure for subsequent boundary correction and disease candidate region generation.

[0071] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative. Those skilled in the art can omit, substitute, and modify the details of the above methods and systems in various ways without departing from the principles and essence of the present invention. For example, combining the above method steps to perform substantially the same function and achieve substantially the same result according to substantially the same method falls within the scope of the present invention. Therefore, the scope of the present invention is defined only by the appended claims.

Claims

1. An automatic disease detection and repair guidance system for multispectral three-dimensional reconstruction of cultural heritage artifacts, characterized by, The system comprises a spectrum acquisition and model fusion unit, an automatic detection unit and a repair pose guiding unit; wherein the spectrum acquisition and model fusion unit is used to acquire visible light, near-infrared light, ultraviolet-induced fluorescence and structure light depth data, and after joint calibration, coarse and fine registration and global consistency processing, an apparent geometric fusion model with reflectivity and geometric consistency is formed; the automatic detection unit is used to take a curved surface area as a basic processing unit on the apparent geometric fusion model, map multi-view and multi-spectral images to a three-dimensional surface in a unified coordinate system to generate a homogeneous apparent image of the area; the hit position of each pixel in the area on the three-dimensional surface is reversely calculated and counted in the evidence stack, and a consistent image of the area is formed according to the majority consistent rule; the boundary of the consistent image is mapped to the three-dimensional geometric surface and corrected in the direction of the geometric normal mutation edge and the curvature ridge line to obtain a disease boundary of the area closely attached to the three-dimensional geometry; the corrected result is used to generate a disease candidate area according to the judgment rules of cracks, starting shell, stains and powdering, and the U-Net or SegFormer network is used to review and take the intersection to output a disease heat map attached to the three-dimensional surface; the repair pose guiding unit is used to generate a pose guiding map according to the disease heat map, plan a path to avoid the fragile area on the three-dimensional curved surface by using the A-star search algorithm, and output a repair pose guiding sequence; The process of the automatic detection unit generating the homogeneous apparent image of the area comprises: for each curved surface area, the surface elements thereof are transformed to each view image according to the unified coordinate system to obtain an area visible domain; in the area visible domain, the stable pixel sets of each spectral channel are respectively counted, wherein the stable pixel is defined as the pixel located in the strong boundary or weak boundary neighborhood and continuously appearing in three adjacent view angles; the area, boundary definition and boundary coincidence of the connected domain formed by the stable pixels are compared and sorted in turn to obtain a channel effectiveness sequence; starting from the front part of the channel effectiveness sequence, the mutual clipping consistency test is performed one by one, if the two channels appear mutually conflicting boundary directions near the area boundary, only the channel more consistent with the area boundary is retained until there is no conflict, and an area selected channel set is formed; in the area selected channel set, the response graph of each channel is first executed to perform cumulative histogram alignment, and then the area internal splicing is performed according to the connected domain seamless splicing rule to generate the homogeneous apparent image of the area; The process of the automatic detection unit forming the consistent image of the area comprises: for each pixel in the homogeneous apparent image of the area, the corresponding surface element of the pixel on the three-dimensional surface is reversely calculated by using the camera geometric relationship in the unified coordinate system, the hit surface element number is recorded to form a pixel-surface element hit list; an evidence stack is established in units of curved surface areas, for each pixel-surface element hit record, the count value of the corresponding surface element is increased by one, and the source view angle and spectral channel of the hit are recorded; for the same surface element, if the number of hit records from different views exceeds half of the total number of hits and the source channels are all in the area selected channel set, the surface element is marked as cross-view consistent; the surface elements marked as cross-view consistent are combined to form the consistent image of the area; The automatic detection unit obtains the disease boundary of the patch area, including: extracting normal mutation edges and curvature ridges on the three-dimensional surface to form a set of feature edges; taking the boundary points of the patch consistent map as the starting point, growing along the nearest feature edge, and preferentially selecting the edge segment consistent with the loop direction of the patch boundary at the grid boundary to generate a boundary traction polyline; mapping the boundary traction polyline from the three-dimensional space to the two-dimensional patch homogenous appearance map to form a boundary correction field; if the boundary correction field and the patch consistent map overlap each other in a certain grid cell but have inconsistent directions, aligning the boundary of the cell to the feature edge direction along the path with the least number of grid edges; obtaining the disease boundary of the patch area strictly adhering to the geometric features.

2. The automatic detection and repair guidance system for the disease of the cultural heritage multispectral three-dimensional reconstruction according to claim 1, characterized in that, The specific processing procedure of the spectrum acquisition and model fusion unit includes: completing joint calibration of the imaging unit and the structured light projection unit, determining the camera internal and external parameters and the projection geometric relationship; extracting and matching SIFT or SURF features on the multi-view images, and using a random consistency algorithm to remove false matches to obtain initial poses of each view; using a point-to-plane iterative closest point algorithm to align and eliminate residual alignment errors, and using a factor graph optimization algorithm to globally constrain all view poses to obtain a consistent three-dimensional surface; performing occlusion removal in a unified coordinate system, and mapping the multispectral images to the three-dimensional surface pixel by pixel to form an appearance map set corresponding to each surface element; performing region growing on the three-dimensional surface according to the normal variation and curvature continuity to obtain a set of surface patches, and generating a unique patch identification code for each surface patch; and the set of surface patches and the appearance map set jointly constitute the appearance-geometry fusion model.

3. The automatic detection and repair guidance system for the disease of the cultural heritage multispectral three-dimensional reconstruction according to claim 2, characterized in that, The automatic detection unit generates a response map for each view and each spectral channel in the following order: performing edge-preserving smoothing on the input image to retain boundary details and weaken isolated noise points; extracting a strong boundary set and a weak boundary set on the processed image, wherein the strong boundary is used for crack and incipient blister determination, and the weak boundary is used for stain and chalking determination; labeling connected regions based on a joint strategy of four-connected and eight-connected to generate a connected domain index; and using the stable region intersection between different spectral channels under the same view to remove local fragments that only stand out in a single channel to obtain the response map.

4. The automatic detection and repair guidance system for the disease of the cultural property multispectral three-dimensional reconstruction according to claim 3, characterized in that, The process of generating the disease candidate region by the automatic detection unit includes: the determination rule for cracks is that the skeleton is obtained by refining the anchor boundary in the homogeneous appearance map of the patch area, if the skeleton is in the form of a long strip and the difference between the pixel grayscale values of the two sides of the skeleton in the visible light and near-infrared channels is greater than a preset first threshold value, and the skeleton is approximately parallel to the normal mutation edge, then it is marked as a crack candidate; the determination rule for the starting cap is that a closed or semi-closed boundary ring is searched in the stable response of the ultraviolet-induced fluorescence channel, if the surface normal variation of the inside and outside regions of the ring is greater than a preset second threshold value and is accompanied by a curvature ridge line, then it is marked as a starting cap candidate; the determination rule for stains is that a fuzzy map is generated for the homogeneous appearance map of the patch area and is subtracted from the original map, if the low-frequency region obtained has a grayscale value in the near-infrared channel that is lower than a preset third threshold value and the boundary does not significantly coincide with the characteristic edge, then it is marked as a stain candidate; the determination rule for pigment powdering is that the intersection of the high-brightness fine-grained region in the visible light channel and the region with a response value in the ultraviolet-induced fluorescence that is lower than a preset fourth threshold value is calculated, if the intersection is distributed in patches in the form of spots in the patch area and does not extend along the characteristic edge, then it is marked as a powdering candidate; then, open-close operation and small hole filling are performed on each candidate region, and isolated patches that only touch the patch boundary loop at corner points are removed, to obtain the disease candidate region.

5. The automatic detection and repair guidance system for the disease of the cultural heritage multispectral three-dimensional reconstruction according to claim 4, characterized in that, The automatic detection unit inputs the disease candidate region as a prior mask into the inference process of the U-Net or SegFormer network, and takes the intersection of the prior mask and the network output mask as the final disease pixel set; the face element hit frequency of the final disease pixel set is summarized according to the curved patch area and is normalized to a preset grade interval, to obtain a disease heat map attached to the three-dimensional surface.

6. The automatic detection and repair guidance system for the disease of the cultural property multispectral three-dimensional reconstruction according to claim 5, wherein, The execution process of the repair pose guidance unit specifically includes: calculating the median direction of the patch normal set as the tool approach direction for each curved patch area, and taking the tangent of the patch boundary loop as the tool sliding direction, to jointly constitute a pose guidance map; on the three-dimensional surface, the lowest grade region of the disease heat map is used as the preferred channel, and the A-star search algorithm based on grid edge counting is used to select paths in turn according to the lexicographic priority.

7. The automatic detection and repair guidance system for the disease of the cultural property multispectral three-dimensional reconstruction according to claim 6, wherein, The lexicographic priority includes: the first priority is to avoid high-grade disease areas, the second priority is to travel along grid edges with smaller curvature, and the third priority is to keep the direction change amount of adjacent path segments to be the smallest; the obtained path is discretized into a path point sequence in the order of grid vertices, and the approach direction and the sliding direction in the pose guidance map are attached to each path point, to form a repair pose guidance sequence.

Citation Information

Patent Citations

  • Fresco scaling damage assessment method based on near-infrared hyperspectrum

    CN104677853A

  • Inscription character image enhancement and restoration system based on multi-modal feature fusion

    CN120976065A