Inscription character image enhancement and restoration system based on multi-modal feature fusion

The image enhancement and restoration system for inscription text, which integrates multimodal feature fusion, solves the problems of low efficiency and secondary damage in traditional inscription restoration, and achieves efficient and reproducible restoration of inscription text, restoring the integrity of the original information.

CN120976065AInactive Publication Date: 2025-11-18HEZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510989263.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current methods for restoring inscriptions rely on manual intervention, which is inefficient, cannot be replicated, and is prone to causing secondary damage. This makes it difficult to promote on a large scale, and the integrity of the information is also limited.

Method used

A system for enhancing and restoring inscription images based on multimodal feature fusion is adopted, including data acquisition and preprocessing, adaptive image enhancement and restoration, multispectral fusion enhancement analysis and restoration output unit. By utilizing three-dimensional geometric, multispectral and texture features, and through the complementarity of infrared and visible light and multimodal data, the restoration strategy is dynamically adjusted to restore the continuity and details of strokes and enhance the contrast of text.

Benefits of technology

It significantly improves the restoration effect of complex weathered inscriptions, provides an efficient and replicable digital restoration solution, reduces the risk of secondary damage, restores the integrity of the original information, and provides an innovative technical approach for cultural relic protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976065A_ABST
    Figure CN120976065A_ABST
Patent Text Reader

Abstract

The invention provides a multimodal feature fusion-based inscription character image enhancement and restoration system, which belongs to the technical field of inscription character restoration and comprises a data acquisition and preprocessing unit, a self-adaptive image enhancement and restoration unit, a multispectral fusion enhancement analysis unit and a restoration output unit, the data acquisition and preprocessing unit is connected with the adaptive image enhancement and restoration unit, the adaptive image enhancement and restoration unit is connected with the multispectral fusion enhancement analysis unit, and the multispectral fusion enhancement analysis unit is connected with the restoration output unit. According to the method, three-dimensional geometry, multispectral and textural features are combined, the information limitation of single-mode restoration is broken through, and the problem that the restoration effect of a traditional GAN in a complex weathered area is unstable is solved through a weathering label and multi-mode feature dynamic adjustment generation strategy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of epigraphic character restoration, and particularly relates to an epigraphic character image enhancement and restoration system based on multi-modal feature fusion. BACKGROUND

[0002] As an important carrier of human civilization, epigraphs carry multi-dimensional cultural heritage information such as historical events, religious beliefs, literature and art, and social systems, and are the core physical data for studying ancient politics, economy, culture, and art. For example, Chinese Han Wei epigraphs, Tang and Song cliff inscriptions, and Western medieval religious inscriptions are all listed as key cultural heritage protection objects due to their irreplaceable historical evidence value. However, due to natural erosion (such as wind and sand erosion, acid rain corrosion, and biological attachment) and human damage (such as carving, material removal, and improper repair), a large number of epigraphs have shown surface erosion, blurred characters, and structural fracture, resulting in the loss or damage of their original information, which seriously threatens the inheritance of cultural heritage. Therefore, efficient and accurate digital enhancement and restoration of epigraphic characters has become one of the core tasks in the field of cultural heritage protection.

[0003] Traditional epigraph restoration mainly relies on manual intervention, and the main methods include physical cleaning (removing surface contaminants), chemical reinforcement (suppressing further weathering), and manual completion (experts copying missing characters based on experience). Although these methods have certain effects in local restoration, they have significant limitations: low efficiency and reproducibility: manual repair is time-consuming and labor-intensive, and highly dependent on the professional experience of the restorer, making it difficult to be widely promoted; risk of secondary damage: improper chemical reinforcement or physical cleaning operations may cause irreversible damage to the epigraph; limited information integrity: manually completed characters are easily affected by subjective factors, making it difficult to restore the original appearance and lacking quantitative basis. Therefore, it is necessary to design an epigraphic character image enhancement and restoration system based on multi-modal feature fusion. SUMMARY

[0004] The purpose of the present application is to provide an epigraphic character image enhancement and restoration system based on multi-modal feature fusion, which solves the technical problems of low efficiency, non-reproducibility, and risk of secondary damage in existing epigraphic character manual restoration.

[0005] In order to achieve the above-mentioned purpose, the technical solution adopted by the present application is as follows:

[0006] The epigraphic character image enhancement and restoration system based on multi-modal feature fusion comprises a data acquisition and preprocessing unit, an adaptive image enhancement and restoration unit, a multi-spectral fusion enhancement analysis unit, and a restoration output unit. The data acquisition and preprocessing unit is connected to the adaptive image enhancement and restoration unit, the adaptive image enhancement and restoration unit is connected to the multi-spectral fusion enhancement analysis unit, and the multi-spectral fusion enhancement analysis unit is connected to the restoration output unit.

[0007] The data acquisition and preprocessing unit acquires the information of the stele surface, and then pre-processes the information of the stele surface; the adaptive image enhancement and repair unit is used for dynamically adjusting the repair strategy for different weathering degree areas of the stele, restoring the stroke continuity and details, and the complementarity of infrared and visible light; the multi-spectral fusion enhancement analysis unit is used for enhancing the contrast between the text and the background, restoring the underlying text covered by pigments, and the repair output unit is used for repairing and outputting the stele.

[0008] Further, the data acquisition and preprocessing unit includes a laser scanning module, a spectral imaging module, an image calibration module, and a data preprocessing module. The laser scanning module uses a line scanning TOF laser radar to cover the overall surface of the stele, outputs point cloud data and a depth map, and the point cloud data includes XYZ coordinates and reflection intensity. The spectral imaging module is used to configure visible light and infrared cameras, synchronously trigger shooting, eliminate the influence of environmental light changes, and record the reflection spectrum curve under different light conditions by matching adjustable LED light sources. The LED light source is incident from several angles, and the incident angle is 0°-80°. The image calibration module is used to calibrate the three-dimensional point cloud graph of the laser and the two-dimensional image of the camera by using coded marker points and a structured light calibration board. The data preprocessing module pre-processes the images of the stele.

[0009] Further, in the laser scanning module, the three-dimensional spatial position of each sampling point on the surface of the stele is recorded, a complete three-dimensional model of the stele is constructed, the geometric features and surface degradation areas of the text are quantitatively analyzed by point cloud, and a structural reference is provided for repair. The geometric features include stroke height, spacing, and inclination angle. The surface degradation areas include the depth of the erosion pit and the three-dimensional trend of the crack. The reflection intensity reflects the surface roughness, material composition, and pollution coverage related information of the stele surface. The original stone surface that is not weathered usually has a reflection intensity higher than a set value due to the dense surface and uniform diffuse reflection. The weathered or polluted area has a reflection intensity lower than the set value due to the rough surface or enhanced light absorption. The reflection intensity can be used as a semantic tag to assist in distinguishing the original text and the degradation area.

[0010] Further, in the spectral imaging module, the visible light and the camera are triggered by a TTL level trigger to ensure that the two images reflect the true state of the stele surface under the same lighting conditions at the same time. The synchronously triggered infrared image can record the reflection signal at the same instant, avoiding the introduction of artifacts due to time difference. The incident angle is the angle between the light source and the normal of the stele surface. By collecting the reflection spectrum curve under different incident angles, the surface pollution interference is suppressed. By analyzing the angle dependence of the spectrum curve, the degradation degree of the stele can be quantified.

[0011] Furthermore, the data preprocessing module includes a point cloud optimization submodule, a modal registration submodule, and an illumination normalization submodule. The 3D point cloud data of the inscribed text is susceptible to environmental vibrations, scanning noise, and surface cracks and pits, leading to uneven point cloud distribution, local voids, or outliers. The point cloud optimization submodule achieves noise suppression and structural fidelity through bilateral filtering denoising and normal vector consistency restoration, providing a high-precision 3D benchmark for subsequent registration and restoration. By calculating the local density and normal vector variance of the point cloud, it identifies void regions caused by cracks or erosion, performs neighborhood searches on the void boundary points, and filters continuous point sets with consistent normal vector directions. It then generates the coordinates and normal vectors of missing points through linear interpolation or surface-based fitting. For linear voids caused by cracks, point clouds are generated by interpolation along the crack direction. For local depressions caused by erosion, the point cloud of the depression area is completed by surface fitting of the surrounding point cloud. The modal registration submodule achieves cross-modal feature association by extracting local features of the image and point cloud. For visible light image SIFT feature extraction, the point cloud is projected onto the imaging plane of the visible light image, and key points of the point cloud corresponding to key points of the image are extracted and their three-dimensional coordinates and normal vector directions are calculated. The key points of the image and point cloud are matched by descriptor similarity to establish an initial correspondence. The illumination normalization submodule corrects the illumination difference of multispectral images based on Lambertian reflection and extracts diffuse reflection components to eliminate specular reflection interference.

[0012] Furthermore, the adaptive image enhancement and restoration unit includes a feature extraction module, a weathering grading module, a region segmentation module, and an adversarial network module. The feature extraction module is used to fuse three-dimensional depth gradient, two-dimensional texture gradient, and spectral reflectance as weathering features. The weathering grading module is used to train a weathering degree classifier through K-means clustering and output pixel-level weathering labels. The region segmentation module is used to separate isolated weathered regions from continuous healthy regions by combining morphological operations and connected component analysis. The adversarial network module adopts a U-Net architecture, takes multimodal feature maps as input, preserves high-frequency details through skip connections, introduces an attention mechanism to focus on the edges and textures of weathered regions, and uses a PatchGAN structure to distinguish whether the input is a real image or a generated image. The adversarial network module is used to set a loss function to constrain the difference between the generated image and the real image in the VGG feature space, and implements an adaptive weathering loss.

[0013] Furthermore, the multispectral fusion enhancement analysis unit includes a spectral feature-level fusion module and a contrast adaptive adjustment module. The spectral feature-level fusion module is used to extract the thermal response features of text in the infrared image. The difference in heat absorption of the underlying stone is more obvious in the pigment peeling area. The edge is preserved by guided filtering, and the texture and color information of the visible light image is extracted. The noise is removed by bilateral filtering, and the Laplacian pyramid fusion is used. The infrared features are used for the global structure, and the visible light features are used for local details. The enhanced image after fusion is output. The contrast adaptive adjustment module is used to introduce spectral weights in the adaptive histogram equalization. The contrast is enhanced in the infrared area, and the natural color tone is preserved in the visible light area. Combined with the three-dimensional depth map, the brightness of the concave area that is prone to dust accumulation is increased, and the edge contrast of the convex area that is prone to wear is enhanced.

[0014] Furthermore, the repair output unit includes an optimization repair module and a visualization output module. The optimization repair module is used to first repair the fracture using cGAN, then enhance the details using spectral fusion, and finally eliminate artifacts using nonlocal mean filtering. The visualization output module is used to compare the original image, depth map, infrared image and repaired image, and provides an OCR interface to verify the readability of the text.

[0015] The present invention, by adopting the above-described technical solution, has the following beneficial effects:

[0016] This invention combines three-dimensional geometry, multispectral and texture features to overcome the information limitations of single-modal restoration. By dynamically adjusting the generation strategy through weathering tags and multimodal features, it solves the problem of unstable restoration effects of traditional GANs in complex weathered areas. By utilizing the material response differences between infrared and visible light, it achieves perspective enhancement of the underlying text in areas where pigments have peeled off. Through multimodal data complementarity and adaptive generative restoration, it significantly improves the text enhancement effect of complex weathered inscriptions, providing an innovative technical path for the field of cultural relic protection. Attached Figure Description

[0017] Figure 1 This is a system unit block diagram of the present invention;

[0018] Figure 2 This is a block diagram of the data acquisition and preprocessing unit module of the present invention;

[0019] Figure 3 This is a block diagram of the adaptive image enhancement and restoration unit module of the present invention;

[0020] Figure 4 This is a block diagram of the multispectral fusion enhancement analysis unit module of the present invention;

[0021] Figure 5 This is a block diagram of the repair output unit module of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and preferred embodiments. However, it should be noted that many details listed in the specification are merely to provide the reader with a thorough understanding of one or more aspects of the present invention, and these aspects of the invention can be implemented even without these specific details.

[0023] like Figure 1 As shown, the image enhancement and restoration system for inscription text based on multimodal feature fusion includes a data acquisition and preprocessing unit, an adaptive image enhancement and restoration unit, a multispectral fusion enhancement and analysis unit, and a restoration output unit. The data acquisition and preprocessing unit is connected to the adaptive image enhancement and restoration unit, the adaptive image enhancement and restoration unit is connected to the multispectral fusion enhancement and analysis unit, and the multispectral fusion enhancement and analysis unit is connected to the restoration output unit.

[0024] The data acquisition and preprocessing unit acquires surface information of the stele and then preprocesses the surface information; the adaptive image enhancement and restoration unit dynamically adjusts the restoration strategy for areas of different weathering degrees of the stele to restore the continuity and details of the strokes, the complementarity of infrared and visible light, and the multispectral fusion enhancement analysis unit to enhance the contrast between the text and the background and restore the underlying text covered by pigment; the restoration output unit is used to restore the stele and output the finished product.

[0025] In embodiments of the present invention, such as Figure 2 As shown, the data acquisition and preprocessing unit includes a laser scanning module, a spectral imaging module, an image calibration module, and a data preprocessing module. The laser scanning module uses a line-scanning TOF lidar to cover the entire surface of the monument, outputting point cloud data and a depth map. The point cloud data includes XYZ coordinates and reflection intensity. The spectral imaging module is used to configure visible light and infrared cameras, synchronously triggering shooting to eliminate the influence of changes in ambient light. It is equipped with an adjustable LED light source to record the reflection spectrum curves under different lighting conditions. The LED light source is incident from several angles, with incident angles ranging from 0° to 80°. The image calibration module is used to register the three-dimensional point cloud map of the laser with the two-dimensional image of the camera using coded marker points and a structural cursor calibration board. The data preprocessing module preprocesses the image of the monument.

[0026] In the laser scanning module, the three-dimensional spatial position of each sampling point on the surface of the inscription is recorded to construct a complete three-dimensional model of the inscription. The point cloud can be used to quantitatively analyze the geometric features of the text and the surface degradation areas, providing a structural benchmark for restoration. The geometric features include stroke height, spacing, and tilt angle. The surface degradation areas include the depth of the erosion pits and the three-dimensional direction of the cracks. The reflection intensity reflects the surface roughness, material composition, and contaminant coverage of the inscription surface. The original stone surface that has not been weathered usually has a higher reflection intensity than the set value because the surface is dense and the diffuse reflection is uniform. The reflection intensity of weathered or contaminant-covered areas is lower than the set value because the surface is rough or the light absorption is enhanced. The reflection intensity can be used as a semantic label of the surface condition to help distinguish the original text from the degradation areas.

[0027] In the spectral imaging module, visible light and the camera are triggered by a TTL level trigger and exposed at the same time to ensure that the two images reflect the true state of the surface of the stele under the same lighting conditions. The synchronously triggered infrared image can record the reflection signal at the same instant, avoiding artifacts introduced by time difference. The incident angle is the angle between the light source and the normal of the stele surface. By collecting the reflection spectrum curves under different incident angles, the interference of surface pollutants is suppressed. By analyzing the angle dependence of the spectrum curve, the degree of degradation of the stele can be quantified.

[0028] A line-scanning TOF (Time-of-Flight) lidar (such as Riegl VZ-4000) with a scanning accuracy of ≤50μm is employed, covering the entire surface of the inscription and outputting point cloud data (XYZ coordinates + reflection intensity) and a depth map (based on point cloud projection). Visible light (400-700nm) and short-wave infrared (700-1700nm) cameras (such as the XiIMEA xiC-64) are configured to support synchronous triggering and eliminate the influence of ambient light variations. An adjustable LED light source (multi-angle incidence, simulating 0°-80° incident angles) is used to record the reflectance spectrum curves under different lighting conditions.

[0029] The data preprocessing module includes a point cloud optimization submodule, a modal registration submodule, and an illumination normalization submodule. The 3D point cloud data of the inscribed text is susceptible to environmental vibrations, scanning noise, and surface cracks and pits, leading to uneven point cloud distribution, local voids, or outliers. The point cloud optimization submodule uses bilateral filtering for noise reduction and normal vector consistency restoration to achieve noise suppression and structural fidelity, providing a high-precision 3D benchmark for subsequent registration and restoration. By calculating the local density and normal vector variance of the point cloud, it identifies void regions caused by cracks or erosion, performs neighborhood searches on the void boundary points, and filters continuous point sets with consistent normal vector directions. Finally, it generates the coordinates and normal vectors of missing points through linear interpolation or surface-based fitting. For linear voids caused by cracks, point clouds are generated by interpolation along the crack direction. For local depressions caused by erosion, the point cloud of the depression area is completed by surface fitting of the surrounding point cloud. The modal registration submodule achieves cross-modal feature association by extracting local features of the image and point cloud. For visible light image SIFT feature extraction, the point cloud is projected onto the imaging plane of the visible light image, and key points of the point cloud corresponding to key points of the image are extracted and their three-dimensional coordinates and normal vector directions are calculated. The key points of the image and point cloud are matched by descriptor similarity to establish an initial correspondence. The illumination normalization submodule corrects the illumination difference of multispectral images based on Lambertian reflection and extracts diffuse reflection components to eliminate specular reflection interference.

[0030] In embodiments of the present invention, such as Figure 3 As shown, the adaptive image enhancement and restoration unit includes a feature extraction module, a weathering grading module, a region segmentation module, and an adversarial network module. The feature extraction module is used to fuse three-dimensional depth gradient, two-dimensional texture gradient, and spectral reflectance as weathering features. The weathering grading module is used to train a weathering degree classifier through K-means clustering and output pixel-level weathering labels. The region segmentation module is used to separate isolated weathered regions from continuous healthy regions by combining morphological operations and connected component analysis. The adversarial network module adopts a U-Net architecture, takes multimodal feature maps as input, preserves high-frequency details through skip connections, introduces an attention mechanism to focus on the edges and textures of weathered regions, and uses a PatchGAN structure to distinguish whether the input is a real image or a generated image. The adversarial network module is used to set a loss function to constrain the difference between the generated image and the real image in the VGG feature space, and implements an adaptive weathering loss.

[0031] In embodiments of the present invention, such as Figure 4As shown, the multispectral fusion enhancement analysis unit includes a spectral feature-level fusion module and a contrast adaptive adjustment module. The spectral feature-level fusion module is used to extract the thermal response features of text in the infrared image. The difference in heat absorption of the underlying stone is more obvious in the pigment peeling area. The edge is preserved by guided filtering, and the texture and color information of the visible light image is extracted. The noise is removed by bilateral filtering, and the Laplacian pyramid fusion is used. The infrared features are used for the global structure, and the visible light features are used for local details. The enhanced image after fusion is output. The contrast adaptive adjustment module is used to introduce spectral weights in the adaptive histogram equalization. The contrast is enhanced in the infrared area, and the natural color tone is preserved in the visible light area. Combined with the three-dimensional depth map, the brightness of the concave area that is easy to accumulate dust is increased, and the edge contrast of the convex area that is easy to wear is enhanced.

[0032] In embodiments of the present invention, such as Figure 5 As shown: The repair output unit includes an optimization repair module and a visualization output module. The optimization repair module is used to repair the fracture first through cGAN, then enhance the details through spectral fusion, and finally eliminate artifacts through nonlocal mean filtering. The visualization output module is used to compare the original image, depth map, infrared image and repaired image, and provides an OCR interface to verify the readability of the text.

[0033] This system provides rapid digital restoration for outdoor stone inscriptions, assists in the creation of high-precision digital archives, restores blurred text content, helps interpret historical information (such as the date and author of the inscription), provides highly realistic 3D models of stone inscriptions for virtual museums, and supports interactive displays.

[0034] Matters not covered in this invention are common knowledge.

[0035] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A system for enhancing and restoring images of inscribed characters based on multimodal feature fusion, characterized in that: It includes a data acquisition and preprocessing unit, an adaptive image enhancement and restoration unit, a multispectral fusion enhancement and analysis unit, and a restoration output unit. The data acquisition and preprocessing unit is connected to the adaptive image enhancement and restoration unit, the adaptive image enhancement and restoration unit is connected to the multispectral fusion enhancement and analysis unit, and the multispectral fusion enhancement and analysis unit is connected to the restoration output unit. The data acquisition and preprocessing unit acquires the surface information of the stele and then preprocesses the surface information of the stele. The adaptive image enhancement and restoration unit is used to dynamically adjust the restoration strategy for areas of different weathering degrees on the stele, restore the continuity and details of the strokes, and utilize the complementarity of infrared and visible light. The multispectral fusion enhancement and analysis unit is used to enhance the contrast between the text and the background and restore the underlying text covered by pigment. The restoration output unit is used to restore the stele and output the finished product.

2. The system for enhancing and restoring inscribed text images based on multimodal feature fusion according to claim 1, characterized in that: The data acquisition and preprocessing unit includes a laser scanning module, a spectral imaging module, an image calibration module, and a data preprocessing module. The laser scanning module uses a line-scanning TOF lidar to cover the entire surface of the monument, outputting point cloud data and a depth map. The point cloud data includes XYZ coordinates and reflection intensity. The spectral imaging module is used to configure visible light and infrared cameras, synchronously triggering shooting to eliminate the influence of changes in ambient light. It is equipped with an adjustable LED light source to record the reflection spectrum curves under different lighting conditions. The LED light source is incident from several angles, with incident angles ranging from 0° to 80°. The image calibration module is used to register the three-dimensional point cloud map of the laser with the two-dimensional image of the camera using coded marker points and a structural cursor calibration board. The data preprocessing module preprocesses the images of the monument.

3. The system for enhancing and restoring inscribed text images based on multimodal feature fusion according to claim 2, characterized in that: In the laser scanning module, the three-dimensional spatial position of each sampling point on the surface of the inscription is recorded to construct a complete three-dimensional model of the inscription. The point cloud can be used to quantitatively analyze the geometric features of the text and the surface degradation areas, providing a structural benchmark for restoration. The geometric features include stroke height, spacing, and tilt angle. The surface degradation areas include the depth of the erosion pits and the three-dimensional direction of the cracks. The reflection intensity reflects the surface roughness, material composition, and contaminant coverage of the inscription surface. The original stone surface that has not been weathered usually has a higher reflection intensity than the set value because the surface is dense and the diffuse reflection is uniform. The reflection intensity of weathered or contaminant-covered areas is lower than the set value because the surface is rough or the light absorption is enhanced. The reflection intensity can be used as a semantic label of the surface condition to help distinguish the original text from the degradation areas.

4. The system for enhancing and restoring inscribed text images based on multimodal feature fusion according to claim 2, characterized in that: In the spectral imaging module, visible light and the camera are triggered by a TTL level trigger and exposed at the same time to ensure that the two images reflect the true state of the surface of the stele under the same lighting conditions. The synchronously triggered infrared image can record the reflection signal at the same instant, avoiding artifacts introduced by time difference. The incident angle is the angle between the light source and the normal of the stele surface. By collecting the reflection spectrum curves under different incident angles, the interference of surface pollutants is suppressed. By analyzing the angle dependence of the spectrum curve, the degree of degradation of the stele can be quantified.

5. The system for enhancing and restoring inscribed text images based on multimodal feature fusion according to claim 2, characterized in that: The data preprocessing module includes a point cloud optimization submodule, a modal registration submodule, and an illumination normalization submodule. The 3D point cloud data of the inscribed text is susceptible to environmental vibrations, scanning noise, and surface cracks and pits, leading to uneven point cloud distribution, local voids, or outliers. The point cloud optimization submodule uses bilateral filtering for noise reduction and normal vector consistency restoration to achieve noise suppression and structural fidelity, providing a high-precision 3D benchmark for subsequent registration and restoration. By calculating the local density and normal vector variance of the point cloud, it identifies void regions caused by cracks or erosion, performs neighborhood searches on the void boundary points, and filters continuous point sets with consistent normal vector directions. Finally, it generates the coordinates and normal vectors of missing points through linear interpolation or surface-based fitting. For linear voids caused by cracks, point clouds are generated by interpolation along the crack direction. For local depressions caused by erosion, the point cloud of the depression area is completed by surface fitting of the surrounding point cloud. The modal registration submodule achieves cross-modal feature association by extracting local features of the image and point cloud. For visible light image SIFT feature extraction, the point cloud is projected onto the imaging plane of the visible light image, and key points of the point cloud corresponding to key points of the image are extracted and their three-dimensional coordinates and normal vector directions are calculated. The key points of the image and point cloud are matched by descriptor similarity to establish an initial correspondence. The illumination normalization submodule corrects the illumination difference of multispectral images based on Lambertian reflection and extracts diffuse reflection components to eliminate specular reflection interference.

6. The system for enhancing and restoring inscribed text images based on multimodal feature fusion according to claim 1, characterized in that: The adaptive image enhancement and restoration unit includes a feature extraction module, a weathering grading module, a region segmentation module, and an adversarial network module. The feature extraction module fuses 3D depth gradient, 2D texture gradient, and spectral reflectance as weathering features. The weathering grading module trains a weathering degree classifier using K-means clustering and outputs pixel-level weathering labels. The region segmentation module combines morphological operations and connected component analysis to separate isolated weathered regions from continuous healthy regions. The adversarial network module uses a U-Net architecture, taking multimodal feature maps as input, preserving high-frequency details through skip connections, introducing an attention mechanism to focus on the edges and textures of weathered regions, and employing a PatchGAN structure to distinguish between real and generated images. The adversarial network module sets a loss function to constrain the difference between generated and real images in the VGG feature space, using an adaptive weathering loss.

7. The system for enhancing and restoring inscribed text images based on multimodal feature fusion according to claim 1, characterized in that: The multispectral fusion enhancement analysis unit includes a spectral feature-level fusion module and a contrast adaptive adjustment module. The spectral feature-level fusion module is used to extract the thermal response features of text in infrared images. The difference in heat absorption of the underlying stone is more obvious in areas where pigment has peeled off. The edge is preserved by guided filtering, and the texture and color information of the visible light image is extracted. Noise is removed by bilateral filtering, and Laplacian pyramid fusion is used. Infrared features are used for the global structure, and visible light features are used for local details. The fused enhanced image is output. The contrast adaptive adjustment module is used to introduce spectral weights in adaptive histogram equalization. The contrast is enhanced in the infrared region, and the natural color tone is preserved in the visible light region. Combined with the three-dimensional depth map, the brightness of the concave areas that are prone to dust accumulation is increased, and the edge contrast of the convex areas that are prone to wear is enhanced.

8. The system for enhancing and restoring inscribed text images based on multimodal feature fusion according to claim 1, characterized in that: The repair output unit includes an optimization repair module and a visualization output module. The optimization repair module first repairs the fracture using cGAN, then enhances the details using spectral fusion, and finally eliminates artifacts using nonlocal mean filtering. The visualization output module is used to compare the original image, depth map, infrared image, and repaired image, and provides an OCR interface to verify text readability.

Citation Information

Cited By

  • Disease automatic detection and repair guide system for cultural relic multispectral three-dimensional reconstruction

    CN121366266A

  • Automatic disease detection and repair guiding system for multi-spectral three-dimensional reconstruction of cultural relics

    CN121366266B