Stele culture digital intelligence display interaction system

By generating high-precision 3D models through laser scanning and multispectral imaging technology, and combining virtual and real fusion rendering with SLAM and AR engines, along with infrared gesture and voice noise reduction technologies, the problem of insufficient modeling accuracy and poor interactive experience in traditional stone inscription culture display is solved. This achieves high-precision fusion of virtual reality and real environment and highly robust user interaction.

CN121764327APending Publication Date: 2026-03-31XIAN UNVERSITY OF ARTS & SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional methods of displaying epigraphic culture suffer from insufficient precision in 3D modeling, significant visual differences between virtual and real environments, poor interactive experience in complex environments, and low accuracy in recognizing user intent.

Method used

Laser scanning and multispectral imaging fusion are used to acquire point cloud and hyperspectral data of the tombstone. A high-precision 3D model is generated by Poisson surface reconstruction and GAN texture restoration. Combined with SLAM spatial localization and AR engine, high-precision spatial registration and visual consistency fusion of virtual content and real environment are achieved. In addition, infrared gesture recognition, end-to-end voice noise reduction and context-aware interactive scheduling mechanism are integrated.

Benefits of technology

It achieves multi-dimensional digital preservation of the geometric shape, material characteristics and cultural semantics of the stone tablet, provides an immersive, intelligent and environmentally adaptable interactive experience, and improves the robustness of user intent recognition and the naturalness of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121764327A_ABST
    Figure CN121764327A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of virtual reality, in particular to an inscription culture digital intelligence display interaction system which comprises a data acquisition and three-dimensional modeling module, an augmented reality display module, a multi-modal interaction control module and a cross-platform adaptation module. Through an SLAM space positioning technology and an AR engine, in combination with an illumination perception virtual-real fusion rendering and multi-sensor fusion positioning system, a three-dimensional model and multimedia information are superposed to a real stele position in real time according to a 6DoF coordinate, and illumination and positioning are adaptively adjusted; high-precision spatial registration and visual consistency fusion of virtual content and a real environment are achieved, an interactive scheduling mechanism of infrared gesture recognition, end-to-end voice noise reduction and context perception is integrated, and high-robustness user intention recognition and multi-modal input fusion control is achieved in a low-light and noise environment. And immersive, intelligent and high-environment-adaptability interaction experience is provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality technology, specifically to a digital display and interactive system for stone inscription culture. Background Technology

[0002] Stele inscriptions are stone carriers inscribed with text, patterns, or symbols. They are typically used to record historical events, deeds of figures, cultural heritage, and religious beliefs, and possess significant historical, cultural, artistic, and scientific value. With the development of digital technology, the digital display of stele culture has become an important way to inherit civilization. However, traditional display methods have significant shortcomings: Insufficient accuracy in 3D modeling: Traditional laser scanning technology is unable to fully capture details such as weathering texture and fine carvings on the surface of the stone tablet, and lacks material spectral information, resulting in geometric and visual deviations between the 3D model and the real stone tablet. The visual differences between virtual and real integration are obvious: the lighting and shadow effects of virtual models and real environments do not match in AR displays, and SLAM positioning is prone to drift in complex scenes, resulting in a stiff virtual-real superposition effect and insufficient immersion in cultural displays. Poor interactive experience in complex environments: In noisy museum environments or low-light outdoor scenes, traditional interaction methods (such as touch and ordinary voice recognition) are easily affected by noise and light, resulting in low accuracy in recognizing user intent and failing to achieve natural and efficient human-computer interaction.

[0003] Based on this, the present invention provides a digital interactive display system for stele culture to solve the aforementioned technical problems. Summary of the Invention

[0004] The purpose of this invention is to provide a digital and intelligent interactive system for displaying stele culture. This invention uses laser scanning and multispectral imaging fusion to acquire point cloud and hyperspectral data of the stele. After Poisson surface reconstruction and GAN texture restoration, a high-precision 3D model is generated. Based on LLM, a structured knowledge graph is constructed by automatically parsing the inscription content. This achieves multi-dimensional digital preservation and knowledge association of the stele's geometric shape, material characteristics, and cultural semantics. Based on SLAM spatial positioning technology and an AR engine, combined with a light-sensing virtual-real fusion rendering and a multi-sensor fusion positioning system, the 3D model and multimedia information are superimposed on the real stele position in real time according to 6DoF coordinates, and the lighting and positioning are adaptively adjusted. This achieves high-precision spatial registration and visual consistency fusion of virtual content and real environment. It also integrates infrared gesture recognition, end-to-end voice noise reduction, and context-aware interactive scheduling mechanisms. In low-light and noisy environments, it achieves highly robust user intent recognition and multimodal input fusion control, providing users with an immersive, intelligent, and environmentally adaptable interactive experience.

[0005] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a digital interactive system for displaying and showcasing stele culture, comprising a data acquisition and 3D modeling module, an augmented reality display module, a multimodal interactive control module, and a cross-platform adaptation module, wherein: The data acquisition and 3D modeling module is used to obtain point cloud and hyperspectral data of the stele by laser scanning and multispectral imaging fusion, generate a 3D model by Poisson surface reconstruction and GAN texture restoration, and construct a knowledge graph based on LLM automatic parsing of the inscription content. The augmented reality display module: Based on SLAM spatial positioning technology, it uses an AR engine to overlay 3D models and multimedia information onto the real location of the monument in the user's field of vision in real time according to 6DoF coordinates, and integrates virtual and real fusion rendering technology with light perception and multi-sensor fusion positioning system to adaptively adjust the lighting effects and spatial positioning of virtual content. The multimodal interaction control module is used to integrate infrared gesture recognition, end-to-end voice noise reduction and context-aware interaction scheduling mechanism to perform robust user intent recognition and multimodal input fusion control in low light and noise environments. The cross-platform adaptation module is used to balance the performance and image quality of different devices by dynamically scaling the resolution and optimizing the rendering pipeline to ensure compatibility with iOS / Android mobile terminals and AR glasses.

[0006] The data acquisition and 3D modeling module includes a laser scanning unit, a multispectral imaging unit, a 3D model generation unit, and a knowledge graph construction unit, wherein: The laser scanning unit is used to acquire point cloud data of the stele through a laser scanner, and accurately capture the geometric shape and spatial coordinates of the stele. The multispectral imaging unit is used to acquire spectral information of the surface material of the tombstone using a multispectral camera. The three-dimensional model generation unit: based on the Poisson surface reconstruction algorithm, it converts point cloud data into a high-precision three-dimensional mesh model and repairs incomplete textures using GAN technology; The knowledge graph construction unit is used to automatically construct a structured cultural knowledge graph containing historical background, relationships between figures, etc., by parsing the semantics of the inscription through LLM.

[0007] The 3D model generation unit is based on the Poisson surface reconstruction algorithm, which converts point cloud data into a high-precision 3D mesh model and repairs incomplete textures using GAN technology. The specific operations are as follows: A1: Point cloud preprocessing: The Statistical Outlier Removal algorithm is used to denoise the point cloud data obtained by laser scanning. The distance threshold is set to the mean + 2 times the standard deviation, and more than 95% of the valid points are retained. A2: Poisson Surface Reconstruction: Import the preprocessed point cloud into the Mitsuba renderer, construct the Signed Distance Field by solving the Poisson equation, set the reconstruction depth parameter to 8, and generate a triangular mesh model; A3: GAN Texture Restoration: ① Construct a training set of stone inscription textures: Collect high-definition texture images of 1000+ complete stone tablets, with a uniform size of 2048×2048 pixels; ②Design the generator network:Use the U-Net architecture with residual blocks, the input is the incomplete texture block, and the output is the repaired complete texture; ③ Training parameter settings: Use the Adam optimizer, with an adversarial loss to perceptual loss weight ratio of 1:5, and train until the FID score is ≤30.

[0008] The augmented reality display module includes a SLAM positioning unit, an AR rendering engine unit, a lighting perception unit, and a multi-sensor fusion unit, wherein: The SLAM localization unit is used to register the spatial coordinates of the real scene and the virtual model through simultaneous localization and mapping (SLAM) technology. The AR rendering engine unit: Based on the AR engine, it overlays 3D models and multimedia information onto the user's field of view in real time according to 6DoF coordinates; The illumination sensing unit is used to analyze ambient illumination parameters using a deep learning illumination probe and dynamically adjust the shadow and reflection effects of the virtual model. The multi-sensor fusion unit is used to integrate sensor data from SLAM, GPS, and inertial navigation.

[0009] The illumination sensing unit uses a deep learning illumination probe to analyze ambient illumination parameters and dynamically adjust the shadow and reflection effects of the virtual model. The specific operation is as follows: B1: Acquire environmental images of the current scene using a multispectral camera; B2: Analyze the image using a pre-trained deep learning illumination probe model to extract ambient light intensity, direction, color temperature, and shadow distribution parameters; B3: Convert the extracted illumination parameters into a spherical harmonic function representation; B4: During AR rendering, the material reflection, shadow projection, and surface specular effects of the virtual 3D model are dynamically adjusted based on the real-time lighting information. B5: Adaptive lighting fusion to achieve visual consistency between virtual content and the real environment.

[0010] The illumination parameters are represented by second-order spherical harmonic functions, and the illumination coefficient... The specific formula is as follows: , in, For spherical harmonic basis functions, For sampling weight function, For ambient light in direction Irradiance.

[0011] The multimodal interaction control module includes an infrared gesture recognition unit, a voice noise reduction processing unit, an interaction scheduling unit, and an intent recognition unit, wherein: The infrared gesture recognition unit is used to accurately capture and analyze hand gestures in low-light environments using infrared sensing technology. The speech noise reduction processing unit is used to filter environmental noise and enhance the recognition accuracy of speech commands through an end-to-end speech noise reduction algorithm. The interactive scheduling unit dynamically allocates response priorities for multimodal inputs such as gestures and voice based on a context-aware mechanism. The intent recognition unit is used to fuse multimodal input data, understand user operation intents such as rotation, scaling, query explanation, and trigger corresponding interactions.

[0012] The speech noise reduction processing unit uses an end-to-end speech noise reduction algorithm to filter environmental noise and enhance the recognition accuracy of speech commands. The specific operation is as follows: C1: Data Preprocessing and Feature Extraction The input noisy speech signal is framed; Applying the Hanning window function to reduce spectral leakage and extracting log-Mel spectral features; C2: End-to-end noise reduction model construction: It adopts the Wave-U-Net architecture and includes: ① Encoder: 12 layers of one-dimensional convolution, each layer has a kernel size of 3, a stride of 2, and the number of channels increases from 16 to 512; ② Decoder: 12 symmetrical deconvolution layers, fusing multi-scale features through skip connections; ③Loss function: Combined SI-SDR loss and multi-resolution STFT loss, with a weighting ratio of 3:1; C3: Model Training and Optimization Training dataset: contains a mixture of 500 hours of museum ambient noise and 1000 hours of clean speech; Data augmentation: Randomly add noise with a signal-to-noise ratio ranging from -5dB to 20dB to simulate different environmental intensities; Optimizer: Adam, trained until SI-SDR on the validation set ≥ 15dB; C4: Real-time noise reduction inference: It adopts a streaming processing mechanism, processing 200ms audio segments at a time with a latency of ≤30ms; The speech recognition accuracy remains at ≥92% under 85dB ambient noise.

[0013] The specific formula for the multi-resolution STFT loss function is as follows: , in, For pure speech spectrum, To predict the speech spectrum, and They represent the first Amplitude and phase spectra at a resolution of [resolution value] These are the weighting coefficients.

[0014] The cross-platform adaptation module includes a dynamic resolution adaptation unit, a rendering pipeline optimization unit, a device compatibility unit, and a performance monitoring unit, wherein: The dynamic resolution adaptation unit is used to automatically adjust the rendering resolution according to the device screen parameters. The rendering pipeline optimization unit is used to balance the model rendering performance and image quality of various devices through GPU shader optimization and LOD level management. The device compatibility unit is used to provide differentiated drivers and interface adaptations for the hardware characteristics of iOS / Android mobile terminals and AR glasses. The performance monitoring unit is used to monitor the device's operating status in real time and dynamically adjust rendering parameters to ensure system smoothness.

[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention acquires point cloud and hyperspectral data of a stele through laser scanning and multispectral imaging fusion. A high-precision 3D model is generated through Poisson surface reconstruction and GAN texture restoration. A structured knowledge graph is constructed based on LLM automatic parsing of the inscription content, achieving multi-dimensional digital preservation and knowledge association of the stele's geometric shape, material characteristics, and cultural semantics. Based on SLAM spatial positioning technology and an AR engine, combined with a light-aware virtual-real fusion rendering and multi-sensor fusion positioning system, the 3D model and multimedia information are superimposed onto the real stele location in real time according to 6DoF coordinates, with adaptive adjustments to lighting and positioning. This achieves high-precision spatial registration and visual consistency fusion of virtual content and the real environment. Furthermore, it integrates infrared gesture recognition, end-to-end voice noise reduction, and context-aware interactive scheduling mechanisms, enabling highly robust user intent recognition and multimodal input fusion control in low-light and noisy environments. This provides users with an immersive, intelligent, and environmentally adaptable interactive experience. Attached Figure Description

[0016] Figure 1 This is a system diagram of a digital interactive display system for stone inscription culture according to the present invention.

[0017] Figure 2 This is a system architecture diagram of a digital display and interactive system for stone inscription culture according to the present invention.

[0018] Figure 3 This is a flowchart of AR rendering and lighting adaptation in a digital display and interactive system for stone inscription culture according to the present invention.

[0019] Explanation of icon numbers: 1. Data Acquisition and 3D Modeling Module; 11. Laser Scanning Unit; 12. Multispectral Imaging Unit; 13. 3D Model Generation Unit; 14. Knowledge Graph Construction Unit; 2. Augmented Reality Display Module; 21. SLAM Localization Unit; 22. AR Rendering Engine Unit; 23. Light Perception Unit; 24. Multi-Sensor Fusion Unit; 3. Multimodal Interaction Control Module; 31. Infrared Gesture Recognition Unit; 32. Speech Noise Reduction Processing Unit; 33. Interaction Scheduling Unit; 34. Intent Recognition Unit; 4. Cross-Platform Adaptation Module; 41. Dynamic Resolution Adaptation Unit; 42. Rendering Pipeline Optimization Unit; 43. Device Compatibility Unit; 44. Performance Monitoring Unit. Detailed Implementation

[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention. Example:

[0021] like Figures 1-3As shown, this embodiment provides a digital interactive system for displaying and showcasing stele culture, including a data acquisition and 3D modeling module 1, an augmented reality display module 2, a multimodal interactive control module 3, and a cross-platform adaptation module 4. Specifically: the data acquisition and 3D modeling module 1 is used to acquire point cloud and hyperspectral data of the stele through laser scanning and multispectral imaging fusion, generate a 3D model through Poisson surface reconstruction and GAN texture restoration, and construct a knowledge graph based on LLM automatic parsing of the inscription content; the augmented reality display module 2, based on SLAM spatial positioning technology, uses an AR engine to integrate the 3D model and multimedia information in real time according to 6DoF coordinates. The system overlays the real location of the monument onto the user's field of vision and integrates virtual-real fusion rendering technology with light perception and a multi-sensor fusion positioning system to adaptively adjust the lighting effects and spatial positioning of the virtual content; Multimodal interaction control module 3: used to integrate infrared gesture recognition, end-to-end voice noise reduction and context-aware interaction scheduling mechanism to perform highly robust user intent recognition and multimodal input fusion control in low light and noise environments; Cross-platform adaptation module 4: used to achieve compatibility with iOS / Android mobile terminals and AR glasses through dynamic resolution scaling and rendering pipeline optimization, balancing the performance and image quality of different devices.

[0022] It should be noted that the data acquisition and 3D modeling module 1 constructs a digital model and knowledge base of the monument through multi-source perception; the augmented reality display module 2 realizes virtual-real fusion rendering based on precise spatial positioning; the multimodal interaction control module 3 parses user intent and provides feedback on operation commands; and the cross-platform adaptation module 4 dynamically optimizes system performance.

[0023] In this embodiment, it should also be noted that the data acquisition and 3D modeling module 1 includes a laser scanning unit 11, a multispectral imaging unit 12, a 3D model generation unit 13, and a knowledge graph construction unit 14, wherein: the laser scanning unit 11 is used to acquire point cloud data of the stele through a laser scanner, accurately capturing the geometric shape and spatial coordinates of the stele; the multispectral imaging unit 12 is used to acquire spectral information of the surface material of the stele using a multispectral camera; the 3D model generation unit 13 is based on the Poisson surface reconstruction algorithm, converting the point cloud data into a high-precision 3D mesh model, and repairing the missing texture through GAN technology; the specific operations are as follows: A1: Point cloud preprocessing: The Statistical Outlier Removal algorithm is used to denoise the point cloud data acquired by laser scanning, setting the distance threshold to mean + 2 times standard deviation, and retaining more than 95% of the valid points; A2: Poisson surface reconstruction: The preprocessed point cloud is imported into the Mitsuba renderer, and the SignedDistance is constructed by solving the Poisson equation. Field: Set the reconstruction depth parameter to 8 to generate a triangular mesh model; A3: GAN texture restoration: ① Construct a training set for inscription textures: Collect high-resolution texture images of 1000+ complete stone tablets, with a uniform size of 2048×2048 pixels; ② Design the generator network: Adopt a U-Net architecture with residual blocks, where the input is the incomplete texture block and the output is the restored complete texture; ③ Training parameter settings: Use the Adam optimizer, with a weight ratio of 1:5 for adversarial loss and perceptual loss, and train until the FID score is ≤30. Knowledge graph construction unit 14: Used to automatically construct a structured cultural knowledge graph containing historical background, relationships between figures, etc., by parsing the semantics of the inscriptions using LLM.

[0024] It should be noted that the laser scanning unit 11 and the multispectral imaging unit 12 simultaneously acquire the geometric and material data of the stele. The 3D model generation unit 13 generates a high-fidelity 3D model through Poisson reconstruction and GAN repair. The knowledge graph construction unit 14 parses the semantics of the inscription based on LLM and constructs a structured knowledge base.

[0025] Furthermore, it should be noted that the laser scanning unit 11 uses a FARO Focus S150 laser scanner with a scanning distance of 0.6-150 meters and a point cloud acquisition accuracy of 0.05mm. It achieves comprehensive data acquisition of the monument through multi-point point cloud stitching technology. During the scanning process, the system automatically generates point cloud color information, providing a foundation for subsequent texture mapping. The multispectral imaging unit 12 is equipped with a Pika L multispectral camera, covering a spectral range of 400-2500nm (visible to near-infrared), supporting simultaneous imaging in 8 bands. The spectral data can be used to analyze the mineral composition and weathering degree of the monument surface, providing a scientific basis for texture restoration. The reconstruction depth is set to 8 in the Mitsuba renderer, corresponding to a mesh resolution of 1 / 2. 8Meters (about 0.39 mm), and the generated triangular mesh model is smoothed by the Loop subdivision algorithm to enhance surface continuity. The knowledge graph construction unit 14 constructs an inscription understanding engine based on the GPT-4 large language model. After identifying the inscription text through OCR technology, it combines with a historical literature knowledge base (such as the "Chinese Stone Inscriptions Database") to achieve: ① Sentence segmentation processing: Using the BiLSTM+CRF model, the sentence segmentation accuracy rate is ≥98%; ② Semantic parsing: Extract entities such as time, place, and person in the inscription, and construct triple relationships (such as "Ode of the Sacred Spring at the Nine-Iron Palace - author - Ouyang Xun"); ③ Knowledge association: Construct an association graph of inscriptions - historical events - artistic styles through the Neo4j graph database to support multi-dimensional retrieval.

[0026] In this embodiment, it should also be noted that the augmented reality display module 2 includes a SLAM positioning unit 21, an AR rendering engine unit 22, a light perception unit 23, and a multi-sensor fusion unit 24, where: SLAM positioning unit 21: Used to register the spatial coordinates of the real scene and the virtual model through simultaneous localization and mapping technology; AR rendering engine unit 22: Based on the AR engine, the three-dimensional model and multimedia information are superimposed on the user's field of view in real time according to 6DoF coordinates; Light perception unit 23: Used to analyze the environmental light parameters using a deep learning light probe and dynamically adjust the shadow and reflection effects of the virtual model; The specific operations are as follows: B1: Collect the environmental image of the current scene through a multi-spectral camera; B2: Use a pre-trained deep learning light probe model to analyze the image and extract environmental light intensity, direction, color temperature, and shadow distribution parameters; B3: Convert the extracted light parameters into a spherical harmonic function representation; The light parameters are represented by second-order spherical harmonic functions, and the light coefficients The specific formula is as follows: [ , Among them, is the spherical harmonic basis function, is the sampling weight function, is the irradiance of the ambient light in the direction . B4: During the AR rendering process, dynamically adjust the material reflection, shadow projection, and surface highlight effects of the virtual three-dimensional model according to the real-time acquired light information; B5: Achieve adaptive light fusion for visual consistency between virtual content and the real environment. Multi-sensor fusion unit 24: Used to integrate sensor data of SLAM, GPS, and inertial navigation.

[0027] Among them, it should be noted that the SLAM positioning unit 21 establishes a centimeter-level spatial coordinate system through multi-source sensor fusion. The multi-sensor fusion unit 24 continuously optimizes the positioning accuracy. The light perception unit 23 real-time analyzes the environmental light parameters and converts them into spherical harmonic function coefficients. The AR rendering engine unit 22 implements 6DoF virtual-real fusion rendering based on the above data.

[0028] Furthermore, it should be noted that the SLAM positioning unit 21 is developed based on the ORB-SLAM3 framework, combining feature point detection and keyframe management to achieve: indoor scene positioning error ≤1cm, and outdoor scene SLAM fine positioning error ≤5cm after initial GPS positioning; in dynamic environments, a repositioning mechanism solves the tracking loss problem when the user's viewpoint moves rapidly. The AR rendering engine unit 22 is developed using Unity AR Foundation and supports: 6DoF (six degrees of freedom) model control, allowing users to view the virtual monument from any angle through head movements or gestures; multimedia information overlay: registering the inscription animation and historical scene restoration video (such as the monument carving process animation) to the corresponding positions of the virtual model according to spatial coordinates. Environmental image acquisition in B1 uses an Intel RealSense D455 RGB-D camera with a resolution of 1280×720 and a frame rate of 30fps, simultaneously acquiring RGB images and depth information. The multi-sensor fusion unit 24 fuses the following data through a Kalman filter: SLAM visual positioning data (high frequency, high accuracy but susceptible to texture); BeiDou satellite positioning data (low frequency, outdoor accuracy 10m); and IMU inertial navigation data (high frequency, short-term motion prediction). In areas with weak GPS signals (such as indoors or under the shade of trees), it automatically switches to SLAM+AprilTag visual marker positioning to ensure positioning continuity.

[0029] In this embodiment, it should also be noted that the multimodal interaction control module 3 includes an infrared gesture recognition unit 31, a voice noise reduction processing unit 32, an interaction scheduling unit 33, and an intent recognition unit 34. Specifically: the infrared gesture recognition unit 31 is used to accurately capture and analyze gestures in low-light environments using infrared sensing technology; the voice noise reduction processing unit 32 is used to filter environmental noise and enhance the recognition accuracy of voice commands through an end-to-end voice noise reduction algorithm; the specific operation is as follows: C1: Data preprocessing and feature extraction: The input noisy speech... Audio signal framing; applying the Hanning window function to reduce spectral leakage and extracting log-Mel spectral features; C2: End-to-end noise reduction model construction: adopting the Wave-U-Net architecture, including: ① Encoder: 12 layers of one-dimensional convolutions, each layer with a kernel size of 3, a stride of 2, and the number of channels increasing from 16 to 512; ② Decoder: 12 symmetrical deconvolutions, fusing multi-scale features through skip connections; ③ Loss function: combining SI-SDR loss and multi-resolution STFT loss, with a weight ratio of 3:1; the specific formula of the multi-resolution STFT loss function is as follows: , in, For pure speech spectrum, To predict the speech spectrum, and They represent the first Amplitude and phase spectra at a resolution of [resolution value] For weighting coefficients. C3: Model Training and Optimization: Training dataset: containing a mixture of 500 hours of museum environmental noise and 1000 hours of clean speech; Data augmentation: randomly adding noise with a signal-to-noise ratio of -5dB to 20dB to simulate different environmental intensities; Optimizer: Adam, trained until the validation set SI-SDR ≥ 15dB; C4: Real-time noise reduction inference: adopting a streaming processing mechanism, processing 200ms audio segments each time, with a latency ≤ 30ms; under 85dB environmental noise, the speech recognition accuracy remains ≥ 92%. Interaction scheduling unit 33: dynamically allocates the response priority of multimodal inputs such as gestures and speech based on a context-aware mechanism; Intent recognition unit 34: used to fuse multimodal input data, understand user operation intent such as rotation, zoom, query explanation, and trigger corresponding interactions.

[0030] It should be noted that the infrared gesture recognition unit 31 and the voice noise reduction processing unit 32 collect user input signals in parallel, and achieve robust environmental perception through high-precision sensing and deep learning noise reduction; the interaction scheduling unit 33 dynamically allocates the priority of multimodal instructions based on the context; the intent recognition unit 34 fuses the processed multi-source input data, accurately analyzes the user's operation intent, and triggers the system response.

[0031] Furthermore, it should be noted that the infrared gesture recognition unit 31 uses an Intel RealSense D415 depth camera, equipped with an infrared emitter and a binocular infrared camera, to achieve: real-time tracking of 22 hand joints within a distance of 0.5-5 meters, with a recognition latency of ≤80ms; it supports 6 basic gestures (click, pinch, swipe, rotate, etc.), and uses the DTW (Dynamic Time Warping) algorithm to match gesture patterns, achieving a recognition accuracy of ≥95%. C1 data preprocessing: input speech sampling rate 16kHz, frame length 25ms, frame shift 10ms, extracting 80-dimensional log-Mel spectrum features. The Wave-U-Net network contains 24 layers (12 encoder layers + 12 decoder layers), with each layer having a convolutional kernel size of 3, and feature transfer is enhanced through residual connections. Adam optimizer parameters... , The weights decay by 1e-6, and training is conducted for 300 rounds until the validation set SI-SDR ≥ 15dB. The interaction scheduling unit 33 constructs a user behavior model based on a time window: when a user zooms in and out of the same area of ​​the inscription three times consecutively, a high-definition texture loading and academic annotation pop-up are automatically triggered; when a user's voice query for "inscription history" is detected, the priority of the voice command is temporarily increased, and other interactive responses are paused. The intent recognition unit 34 integrates a decision tree model of multimodal input: ① When gestures (such as rotation) are combined with voice commands (such as "view the back"), the composite intent is executed first; based on knowledge graph contextual understanding: ② When a user clicks on the characters "Ouyang Xun" in the inscription, his biography and other inscription works are automatically associated.

[0032] In this embodiment, it should also be noted that the cross-platform adaptation module 4 includes a dynamic resolution adaptation unit 41, a rendering pipeline optimization unit 42, a device compatibility unit 43, and a performance monitoring unit 44, wherein: the dynamic resolution adaptation unit 41 is used to automatically adjust the rendering resolution according to the device screen parameters; the rendering pipeline optimization unit 42 is used to balance the model rendering performance and image quality of each device through GPU shader optimization and LOD level management; the device compatibility unit 43 is used to provide differentiated drivers and interface adaptations for the hardware characteristics of iOS / Android mobile terminals and AR glasses; and the performance monitoring unit 44 is used to monitor the device operating status in real time and dynamically adjust rendering parameters to ensure system smoothness.

[0033] It should be noted that the dynamic resolution adaptation unit 41 and the rendering pipeline optimization unit 42 work together to process the graphics rendering load, the device compatibility unit 43 provides underlying hardware adaptation support, and the performance monitoring unit 44 provides real-time feedback on the system status and dynamically adjusts the parameters.

[0034] Furthermore, it should be noted that the dynamic resolution adaptation unit 41 on mobile devices (GPU is Adreno 640): the rendering resolution is reduced to 720p, and then reconstructed to 1080p through the Temporal Super Resolution algorithm, reducing power consumption by 30%; AR glasses (such as HoloLens 2): adopts a native resolution of 1440×1600, and dynamically adjusts the rendering resolution to 1200×1333 to ensure a frame rate of 90fps. The rendering pipeline optimization unit 42 GPU shader optimization: uses HLSL compilation optimization technology to reduce the number of shader instructions by more than 20%; LOD level management: the tombstone model LOD level settings are: LOD0: 100,000 polygons (AR glasses); LOD1: 50,000 polygons (high-end mobile phones); LOD2: 20,000 polygons (mid-range mobile phones); automatically switching based on viewing distance and device performance. Device compatibility unit 43 includes: iOS devices: Enabled Metal rendering interface, supporting Mesh Shader acceleration; Android devices: Adapted to Vulkan 1.3 / OpenGL ES 3.2, using ASTC texture compression; AR glasses: Optimized eye-tracking interface for HoloLens 2, optimized gesture recognition algorithm for Magic Leap 2. Performance monitoring unit 44 monitors real-time metrics: GPU temperature (threshold 85℃, reducing rendering resolution when exceeding this temperature); Frame rate (target ≥30fps, triggering LOD degradation when below 25fps); Memory usage (releasing unnecessary texture resources when exceeding 80% of device memory).

[0035] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0036] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A digital interactive display system for stele culture, characterized in that, It includes a data acquisition and 3D modeling module (1), an augmented reality display module (2), a multimodal interactive control module (3), and a cross-platform adaptation module (4), among which: The data acquisition and 3D modeling module (1) is used to obtain point cloud and hyperspectral data of the stele by laser scanning and multispectral imaging fusion, generate a 3D model by Poisson surface reconstruction and GAN texture repair, and construct a knowledge graph based on LLM automatic parsing of the inscription content. The augmented reality display module (2) is based on SLAM spatial positioning technology. Through the AR engine, the three-dimensional model and multimedia information are superimposed on the real stone position in the user's field of vision in 6DoF coordinates in real time. It also integrates the virtual and real fusion rendering technology with light perception and the multi-sensor fusion positioning system to adaptively adjust the lighting effect and spatial positioning of the virtual content. The multimodal interaction control module (3) is used to integrate infrared gesture recognition, end-to-end voice noise reduction and context-aware interaction scheduling mechanism to perform robust user intent recognition and multimodal input fusion control in low light and noise environments. The cross-platform adaptation module (4) is used to balance the performance and image quality of different devices by dynamically scaling the resolution and optimizing the rendering pipeline to be compatible with iOS / Android mobile terminals and AR glasses.

2. The digital display and interactive system for stele culture according to claim 1, characterized in that, The data acquisition and 3D modeling module (1) includes a laser scanning unit (11), a multispectral imaging unit (12), a 3D model generation unit (13), and a knowledge graph construction unit (14), wherein: The laser scanning unit (11) is used to acquire point cloud data of the stele through a laser scanner, and accurately capture the geometric shape and spatial coordinates of the stele. The multispectral imaging unit (12) is used to acquire spectral information of the surface material of the stone using a multispectral camera; The three-dimensional model generation unit (13) converts point cloud data into a high-precision three-dimensional mesh model based on the Poisson surface reconstruction algorithm and repairs the missing texture through GAN technology. The knowledge graph construction unit (14) is used to automatically construct a structured cultural knowledge graph containing historical background, relationships between people, etc. by parsing the semantics of the inscription through LLM.

3. The digital display and interactive system for stele culture according to claim 2, characterized in that, The three-dimensional model generation unit (13) is based on the Poisson surface reconstruction algorithm, which converts point cloud data into a high-precision three-dimensional mesh model and repairs incomplete textures using GAN technology. The specific operation is as follows: A1: Point cloud preprocessing: The Statistical Outlier Removal algorithm is used to denoise the point cloud data obtained by laser scanning. The distance threshold is set to the mean + 2 times the standard deviation, and more than 95% of the valid points are retained. A2: Poisson Surface Reconstruction: Import the preprocessed point cloud into the Mitsuba renderer, construct the Signed Distance Field by solving the Poisson equation, set the reconstruction depth parameter to 8, and generate a triangular mesh model; A3: GAN Texture Restoration: ① Construct a training set of stone inscription textures: Collect high-definition texture images of 1000+ complete stone tablets, with a uniform size of 2048×2048 pixels; ②Design the generator network:Use the U-Net architecture with residual blocks, the input is the incomplete texture block, and the output is the repaired complete texture; ③ Training parameter settings: Use the Adam optimizer, with an adversarial loss to perceptual loss weight ratio of 1:5, and train until the FID score is ≤30.

4. The digital display and interactive system for stele culture according to claim 1, characterized in that, The augmented reality display module (2) includes a SLAM positioning unit (21), an AR rendering engine unit (22), a lighting perception unit (23), and a multi-sensor fusion unit (24), wherein: The SLAM localization unit (21) is used to register the spatial coordinates of the real scene and the virtual model through synchronous localization and mapping technology; The AR rendering engine unit (22) uses the AR engine to overlay the 3D model and multimedia information onto the user's field of vision in real time according to 6DoF coordinates; The illumination sensing unit (23) is used to analyze ambient illumination parameters using a deep learning illumination probe and dynamically adjust the shadow and reflection effects of the virtual model. The multi-sensor fusion unit (24) is used to integrate sensor data from SLAM, GPS, and inertial navigation.

5. The digital display and interactive system for stele culture according to claim 4, characterized in that, The illumination sensing unit (23) uses a deep learning illumination probe to analyze ambient illumination parameters and dynamically adjust the shadow and reflection effects of the virtual model. The specific operation is as follows: B1: Acquire environmental images of the current scene using a multispectral camera; B2: Analyze the image using a pre-trained deep learning illumination probe model to extract ambient light intensity, direction, color temperature, and shadow distribution parameters; B3: Convert the extracted illumination parameters into a spherical harmonic function representation; B4: During AR rendering, the material reflection, shadow projection, and surface specular effects of the virtual 3D model are dynamically adjusted based on the real-time lighting information. B5: Adaptive lighting fusion to achieve visual consistency between virtual content and the real environment.

6. The digital display and interactive system for stele culture according to claim 5, characterized in that, The illumination parameters are represented by second-order spherical harmonic functions, and the illumination coefficient... The specific formula is as follows: , in, For spherical harmonic basis functions, For sampling weight function, For ambient light in direction Irradiance.

7. The digital display and interactive system for stele culture according to claim 1, characterized in that, The multimodal interaction control module (3) includes an infrared gesture recognition unit (31), a voice noise reduction processing unit (32), an interaction scheduling unit (33), and an intent recognition unit (34), wherein: The infrared gesture recognition unit (31) is used to accurately capture and analyze gestures in low-light environments using infrared sensing technology. The speech noise reduction processing unit (32) is used to filter environmental noise and enhance the recognition accuracy of speech commands through an end-to-end speech noise reduction algorithm. The interactive scheduling unit (33) dynamically allocates the response priority of multimodal inputs such as gestures and voice based on a context-aware mechanism. The intent recognition unit (34) is used to fuse multimodal input data, understand user operation intents such as rotation, scaling, query explanation, and trigger corresponding interactions.

8. The digital display and interactive system for stele culture according to claim 7, characterized in that, The speech noise reduction processing unit (32) uses an end-to-end speech noise reduction algorithm to filter environmental noise and enhance the recognition accuracy of speech commands. The specific operation is as follows: C1: Data Preprocessing and Feature Extraction The input noisy speech signal is framed; Applying the Hanning window function to reduce spectral leakage and extracting log-Mel spectral features; C2: End-to-end noise reduction model construction: It adopts the Wave-U-Net architecture and includes: ① Encoder: 12 layers of one-dimensional convolution, each layer has a kernel size of 3, a stride of 2, and the number of channels increases from 16 to 512; ② Decoder: 12 symmetrical deconvolution layers, fusing multi-scale features through skip connections; ③Loss function: Combined SI-SDR loss and multi-resolution STFT loss, with a weighting ratio of 3:1; C3: Model Training and Optimization Training dataset: contains a mixture of 500 hours of museum ambient noise and 1000 hours of clean speech; Data augmentation: Randomly add noise with a signal-to-noise ratio ranging from -5dB to 20dB to simulate different environmental intensities; Optimizer: Adam, trained until SI-SDR on the validation set ≥ 15dB; C4: Real-time noise reduction inference: It adopts a streaming processing mechanism, processing 200ms audio segments at a time with a latency of ≤30ms; The speech recognition accuracy remains at ≥92% under 85dB ambient noise.

9. The digital display and interactive system for stele culture according to claim 8, characterized in that, The specific formula for the multi-resolution STFT loss function is as follows: , in, For pure speech spectrum, To predict the speech spectrum, and They represent the first Amplitude and phase spectra at a resolution of [resolution value] These are the weighting coefficients.

10. The digital display and interactive system for stele culture according to claim 1, characterized in that, The cross-platform adaptation module (4) includes a dynamic resolution adaptation unit (41), a rendering pipeline optimization unit (42), a device compatibility unit (43), and a performance monitoring unit (44), wherein: The dynamic resolution adaptation unit (41) is used to automatically adjust the rendering resolution according to the device screen parameters; The rendering pipeline optimization unit (42) is used to balance the model rendering performance and image quality of each device through GPU shader optimization and LOD level management. The device compatibility unit (43) is used to provide differentiated drivers and interface adaptations for the hardware characteristics of iOS / Android mobile terminals and AR glasses. The performance monitoring unit (44) is used to monitor the device's operating status in real time and dynamically adjust rendering parameters to ensure system smoothness.