A digital twin processing system and method for cultural and tourism virtual reality
The method addresses the lack of dynamic lighting and material deformation in virtual reality by using neural networks to generate accurate historical lighting and predict material changes, improving the realism and immersion of virtual reality experiences.
Patent Information
- Application Number
- CN202510464874.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing technology cannot accurately reproduce the dynamic light and shadow effects of historical buildings and the deformation caused by humidity changes in cultural and tourism virtual reality, resulting in insufficient realism and immersion of the user experience.
By collecting ancient building data, defining time, space and semantic nodes, calculating weights, combining Fick's diffusion law and finite element analysis, predicting deformation, and using neural networks to optimize light source parameters, performing real and historical channels rendering, mixing virtual and real light and shadow, and real to achieve accurate simulation of lighting patterns and deformation.
It realizes accurate restoration of historical lighting mode and accurate prediction of ancient building deformation, enhancing the authenticity and immersion of the user experience.
Smart Images

Figure CN119989827B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and in particular to a digital twin processing system and method for cultural and tourism virtual reality. Background Art
[0002] With the wide application of digital twin technology and virtual reality (VR) in the field of cultural tourism, how to accurately reproduce historical buildings and their environments has become a research hotspot. Existing methods can already collect three-dimensional data of ancient building components through devices such as lidar and hyperspectral cameras, and use computer graphics to achieve preliminary visual display. However, these methods mostly focus on static display, lacking the simulation of dynamic light and shadow effects and the deformation process of ancient buildings, resulting in insufficient realism and immersion in the user experience.
[0003] Although some methods attempt to combine light intensity and direction for simple rendering, they fail to fully consider the light source characteristics of different historical periods and their impact on building materials, resulting in inaccurate historical scenes presented. In addition, existing methods are lacking in dealing with the impact of humidity changes on the deformation of ancient building materials and are difficult to accurately predict and reflect the actual deformation amount. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a digital twin processing method for cultural and tourism virtual reality to solve the problems of inaccurate historical lighting patterns and deformation prediction of ancient building materials due to environmental changes.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a digital twin processing method for cultural and tourism virtual reality, which includes
[0008] collecting ancient building component data, including three-dimensional coordinates, point cloud normal vectors, reflectivity, light intensity, light direction, spectral intensity distribution, and humidity data, and preprocessing the ancient building component data;
[0009] defining time nodes, space nodes, and semantic nodes, calculating the similarity weight of time nodes, the distance weight of space nodes, the correlation weight of semantic nodes, and the comprehensive weight of historical patterns, performing mutation detection on humidity data, predicting the deformation amount of ancient buildings through Fick's diffusion law and finite element analysis, reducing the dimension of the historical light source parameters with the highest comprehensive weight of historical patterns and splicing them with real-time light source parameters, and inputting them into a generator to obtain virtual light source parameters;
[0010] Based on the virtual light source parameters, performing real-channel and historical-channel rendering, obtaining the user's line of sight direction to calculate the user's perspective, and mixing virtual and real light and shadow;
[0011] Optimize the neural network weights of the generator through incremental training.
[0012] As a preferred solution of the digital twin processing method for cultural and tourism virtual reality according to the present invention, wherein: defining time nodes, space nodes and semantic nodes, calculating the similarity weight of time nodes, the distance weight of space nodes, the correlation weight of semantic nodes and the comprehensive weight of historical patterns, the specific steps are as follows.
[0013] Identify the component type through three-dimensional coordinates and normal vectors, determine the light source type through spectral intensity, and splice the component type name and the light type name into semantic nodes.
[0014] Use the Gaussian kernel function to calculate the similarity weight of time nodes and the distance weight of space nodes.
[0015] Calculate the correlation weight of semantic nodes through the BERT model combined with spectral intensity, and calculate the comprehensive weight of historical patterns according to the similarity weight of time nodes, the distance weight of space nodes and the correlation weight of semantic nodes.
[0016] As a preferred solution of the digital twin processing method for cultural and tourism virtual reality according to the present invention, wherein: performing real channel and historical channel rendering, obtaining the user's line of sight direction to calculate the user's perspective, and mixing virtual and real light and shadow, the specific steps are as follows.
[0017] Based on the virtual light source parameters, use the lumen global illumination engine combined with the ray tracing noise reduction algorithm to perform real channel rendering.
[0018] Adopt time-based sine wave fluctuations to adjust the roughness of the ancient building materials, and map the roughness of the ancient building materials to the normal map through physically based rendering for historical channel rendering.
[0019] Adjust the real channel weight according to the user's perspective, combine the edge blur algorithm to eliminate the virtual-real boundary, capture the user's gaze area based on the eye tracker, and asynchronously calculate and delay update the light and shadow of the non-gaze area through the GPU.
[0020] As a preferred solution of the digital twin processing method for cultural and tourism virtual reality according to the present invention, wherein: reducing the dimension of the historical light source parameters with the highest comprehensive weight of historical patterns and splicing them with the real-time light source parameters, and inputting them into the generator to obtain the virtual light source parameters, the specific steps are as follows.
[0021] The real-time light source parameters refer to light intensity, light direction and spectral intensity.
[0022] The virtual light source parameters refer to virtual intensity, virtual direction and virtual spectrum.
[0023] Reduce the historical light source parameters with the highest comprehensive weight in the historical mode through principal component analysis, splice them with the real-time light source parameters, and add random noise at the same time to obtain the light source vector;
[0024] Input the light source vector into the generator, use the Sigmoid function and the Tanh function to convert it into a normalized value, obtain the virtual intensity by linearly mapping the output of the Sigmoid function, obtain the virtual direction by linearly mapping the output of the Tanh function, and restore the virtual spectrum by inverse principal component analysis of the output of the fully connected layer.
[0025] As a preferred solution of the digital twin processing method for cultural and tourism virtual reality according to the present invention, wherein: perform mutation detection on the humidity data, and predict the deformation amount of the ancient building through Fick's diffusion law and finite element analysis. The specific steps are as follows,
[0026] Perform sliding window mean filtering on the humidity data, calculate the humidity change rate per minute, and determine the environmental mutation event;
[0027] Generate a two-dimensional humidity grid using Kriging interpolation method, perform convolution on the humidity grid using the Sobel operator, and calculate the gradient vector;
[0028] Generate a humidity distribution based on Fick's diffusion law, calculate the humidity change rate according to the gradient vector and the point cloud normal vector, calculate the strain value in combination with the expansion coefficient, and calculate the overall deformation of the material through the mechanical equilibrium equation combined with finite element analysis.
[0029] As a preferred solution of the digital twin processing method for cultural and tourism virtual reality according to the present invention, wherein: the optimization of the neural network weights of the generator through incremental training refers to collecting abnormal data, setting the optimizer and parameters, inputting the abnormal data into the generator to calculate the total loss, and using the backpropagation method to update the neural network weights.
[0030] As a preferred solution of the digital twin processing method for cultural and tourism virtual reality according to the present invention, wherein: the preprocessing of the ancient building component data refers to aligning the ancient building component data on the time axis and coordinates.
[0031] In a second aspect, the present invention provides a digital twin processing system for cultural and tourism virtual reality, including,
[0032] A preprocessing module that collects ancient building component data, including three-dimensional coordinates, point cloud normal vectors, reflectivity, light intensity, light direction, spectral intensity distribution, and humidity data, and preprocesses the ancient building component data;
[0033] The generator module defines time nodes, space nodes, and semantic nodes, calculates the similarity weights of time nodes, the distance weights of space nodes, the correlation weights of semantic nodes, and the comprehensive weights of historical patterns, performs mutation detection on humidity data, predicts the deformation amount of ancient buildings through Fick's diffusion law and finite element analysis, reduces the dimension of the historical light source parameters with the highest comprehensive weight of historical patterns and concatenates them with real-time light source parameters, and inputs them into the generator to obtain virtual light source parameters;
[0034] The mixing module performs real and historical channel rendering based on the virtual light source parameters, calculates the user's viewing direction to obtain the user's perspective, and mixes virtual and real light and shadow;
[0035] The optimization module optimizes the neural network weights of the generator through incremental training.
[0036] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the digital twin processing method for cultural and tourism virtual reality as described in the first aspect of the present invention is implemented.
[0037] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the digital twin processing method for cultural and tourism virtual reality as described in the first aspect of the present invention is implemented.
[0038] The beneficial effects of the present invention are as follows: By defining time nodes, space nodes, and semantic nodes, combining the Gaussian kernel function to calculate time and space weights, and the BERT model to calculate semantic similarity, and reducing the dimension of the historical light source parameters with the highest comprehensive weight of historical patterns and inputting them into the generator to generate virtual light source parameters, the illumination parameters in historical documents (such as candlelight color temperature, spectral peak) are dynamically associated with real-time environmental data, solving the problem of the inability to accurately reproduce historical lighting patterns; By detecting humidity mutations through sliding window filtering, generating a humidity field by Kriging interpolation, calculating the gradient vector by the Sobel operator, and combining Fick's diffusion law and finite element analysis to predict the deformation amount, the problem of inaccurate prediction of the deformation of ancient buildings caused by humidity changes is solved. Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 It is a flowchart of the digital twin processing method for cultural and tourism virtual reality in Embodiment 1.
[0041] Figure 2 It is a schematic diagram of the digital twin processing system for cultural and tourism virtual reality in Embodiment 1.
[0042] Figure 3 It is the flowchart of humidity deformation prediction in Embodiment 1.
[0043] Figure 4 It is the flowchart of virtual light source generation in Embodiment 1. Specific Embodiments
[0044] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings of the specification.
[0045] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0046] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation of the present invention. The "in one embodiment" that appears in different places in this specification does not all refer to the same embodiment, nor is it a separate or selectively exclusive embodiment from other embodiments.
[0047] Embodiment 1, referring to Figure 1 , Figure 2 , Figure 3 and Figure 4 , this embodiment provides a digital twin processing method for cultural and tourism virtual reality, including the following steps:
[0048] S1: Collect ancient building component data, including three-dimensional coordinates, point cloud normal vectors, reflectivity, light intensity, light direction, spectral intensity distribution, and humidity data, and preprocess the ancient building component data.
[0049] The specific steps are as follows,
[0050] Perform phase scanning on ancient building components (such as brackets and cornices) through a phase-type lidar to collect three-dimensional point cloud data of the building components, including the three-dimensional coordinates and point cloud normal vectors of the ancient building components (such as brackets and cornices), and the point density ≥ 200 points per square meter.
[0051] It should also be noted that through high-precision point cloud acquisition of the phase-type lidar (point density ≥ 200 points / square meter), the minute structural details (such as carved patterns) of ancient building components (such as brackets and eaves corners) are ensured to be completely recorded, providing a high-fidelity geometric benchmark for subsequent spatio-temporal alignment and deformation prediction.
[0052] Use a hyperspectral camera (400 - 1000 nm band) to collect the reflectance of the material corresponding to each point cloud (wavelength interval 1 nm).
[0053] It should also be noted that the hyperspectral resolution with a 1 nm wavelength interval can accurately capture the differences in reflectance characteristics of different materials (such as wood and glazed tiles), providing spectral-level data support for the physical rendering of historical light and material interaction.
[0054] Use a high-precision digital light sensor (such as ams OSRAM TSL25911) to record the light intensity in real time, with a measurement range of 0 - 100 klux and a sampling rate of 100 Hz.
[0055] It should also be noted that the combination of 100 Hz high-frequency sampling and an azimuth angle accuracy of ±0.1° can capture the minute changes in the light direction in real time (such as the shadow offset caused by cloud movement), avoiding the light and shadow mutation distortion caused by traditional low-frequency sampling.
[0056] Collect the light direction through a six-axis IMU, including measuring the azimuth angle and elevation angle of the light, with an accuracy of ±0.1°.
[0057] Use a spectrometer to capture the spectral intensity distribution in the 380 - 780 nm band, with a resolution of 1 nm.
[0058] Use a temperature and humidity sensor to collect humidity data, with a sampling rate of once per second.
[0059] It should also be noted that the 380 - 780 nm band covers the spectral range sensitive to the human eye. Combining with the humidity sampling per second, it can dynamically correlate the environmental humidity changes with the optical properties of materials (such as the reflectance attenuation of wood after absorbing moisture).
[0060] Use OCR to parse historical documents and extract lighting parameters (such as the candlelight color temperature of 1900 K and the peak wavelength of the lantern spectrum of 560 nm).
[0061] Adopt the RoBERTa-NLP model and the knowledge graph of "Yingzao Fashi" to parse the construction rules in the documents (such as "the spacing of brackets ≤ 0.3 m"), and extract logical constraints through breadth-first search of the dependency tree to build a logical constraint library.
[0062] It should also be noted that: by combining the RoBERTa-NLP model with the Knowledge Graph of "Yingzao Fashi", implicit construction rules (such as "the spacing of dougong ≤ 0.3m") can be automatically extracted, avoiding the subjective deviation of manual annotation.
[0063] An atomic clock (such as a high-precision GPS-tamed rubidium atomic clock (±10μs)) is used to add timestamps to the 3D point cloud data, reflectivity, light intensity, light direction, and spectral intensity distribution of building components, and the time axis is aligned through the Dynamic Time Warping (DTW) algorithm to ensure that the timestamp deviation between the sensor and the point cloud ≤ 10ms.
[0064] The point cloud coordinates (UTM-50N coordinate system), high-precision digital light sensor coordinates, six-axis IMU coordinates, and spectrometer coordinates are aligned with the UTM-50N coordinate system as the reference through the ICP algorithm and encoded into the UTM coordinate format.
[0065] It should also be noted that: the combined application of a GPS-tamed rubidium atomic clock (±10μs) and the ICP algorithm ensures the strict alignment of multi-source sensor data (such as point cloud and spectrum) in the spatio-temporal dimension, eliminating artifacts caused by clock asynchronization.
[0066] The time dimension (dynasty (extracted by NLP), lunar season (converted from GPS timestamp), and hour (calculated from solar altitude angle)), space dimension (point cloud coordinates, high-precision digital light sensor coordinates, six-axis IMU coordinates, and spectrometer coordinates in UTM coordinate format), and semantic dimension (component type (dougong / eave corner) and light type (candlelight / natural light)) are fused into spatio-temporal tags.
[0067] It should also be noted that: through the high-precision collaborative acquisition of multi-modal sensors (point cloud, spectrum, light, humidity), a spatio-temporal-physical-material full-dimensional digital twin model of ancient building components is constructed, providing a millimeter-level geometric benchmark and spectral-level material feature library for subsequent historical scene reconstruction.
[0068] S2: Define time nodes, space nodes, and semantic nodes, calculate the similarity weights of time nodes, distance weights of space nodes, correlation weights of semantic nodes, and comprehensive weights of historical patterns, perform mutation detection on humidity data, predict the deformation of ancient buildings through Fick's diffusion law and finite element analysis, reduce the dimension of the historical light source parameters with the highest comprehensive weight of historical patterns and splice them with the real-time light source parameters, and input them into the generator to obtain virtual light source parameters.
[0069] Define time nodes based on the dynasty, lunar season, and hour. For example, "Ming Dynasty_Autumn_Hour of You" represents the dusk period in autumn of the Ming Dynasty.
[0070] Define spatial nodes based on the UTM-50N coordinates of ancient building components (such as brackets and eaves corners) in point cloud coordinates, the coordinates of high-precision digital light sensors, the coordinates of six-axis IMUs, and the coordinates of spectrometers. Each spatial node is labeled with the component type and the sensor number.
[0071] Define semantic nodes based on the component type (brackets / eaves corners) and the light type (candlelight / natural light): Identify the component type through three-dimensional coordinates and normal vectors, including brackets, eaves corners, beams and columns. The classification basis is the component shape database recorded in "Yingzao Fashi". Determine the light source type through the spectral intensity. Candlelight corresponds to a peak wavelength of 560-580nm, and natural light corresponds to a peak wavelength of 450-480nm. Concatenate the component type name and the light type name as the semantic node. For example, "brackets_candlelight" represents the semantic characteristics of brackets in a candlelight environment.
[0072] It should also be noted that: Integrate the dynasty, lunar season, and time of day into a time node, which can accurately associate the light characteristics recorded in historical documents (such as "candlelight at dusk in the Ming Dynasty") with the physical form of ancient building components, and avoid the dating errors of modern time sequence divisions.
[0073] Calculate the similarity weight of time nodes: If the dynasties are different (such as "Ming Dynasty" and "Qing Dynasty"), then set the similarity weight of the time node to zero; If the dynasties of two time nodes are the same (such as both being "Ming Dynasty"), then map the season to the month difference (for example, September in autumn and December in winter differ by 3 months). Use the Gaussian kernel function ( σ with σ = 12 months and the input being the month difference) to calculate the similarity weight of the time node. The smaller the month difference, the higher the similarity weight of the time node. For example, the time nodes "Ming Dynasty_autumn_You hour" (September) and "Ming Dynasty_winter_Shen hour" (December) have a month difference of 3 months, and the similarity weight of the time node is 0.97.
[0074] It should also be noted that: The non-linear attenuation characteristic of the Gaussian kernel function (σ = 12 months) can distinguish the light differences in different seasons of the same dynasty (such as strong direct sunlight in summer and scattered light in winter), and improve the spatio-temporal continuity of historical light patterns.
[0075] Calculate the distance weight of spatial nodes: Calculate the distance between spatial nodes through the Haversine algorithm (for example, the brackets are 30 meters away from the light sensor). Use the Gaussian kernel function ( σ with σ = 80 meters and the input being the distance between spatial nodes) to calculate the distance weight of spatial nodes. The distance weight of spatial nodes decreases exponentially with the increase of the distance. For example, the brackets are 30 meters away from the light sensor, and the distance weight of the spatial node is 0.82.
[0076] It should also be noted that: The Haversine algorithm combined with the Gaussian kernel function (σ = 80 meters) can quantify the impact of spatial distance on light propagation (such as the attenuation difference between the eaves corner and the ground candlelight sensor), avoiding the linear simplification of the Euclidean distance.
[0077] Calculate the semantic node correlation weight: Calculate the semantic node similarity through the BERT model (such as the similarity between "Dougong_candlelight" and "Eaves corner_candlelight" is 0.6). This includes splitting the semantic node into component types and light types (such as "Dougong" and "candlelight"), inputting them into the BERT model to generate component type vectors and light type vectors, performing weighted summation on the component type vectors and light type vectors to obtain the semantic node vector. The weight of the component type vector is 0.6, and the weight of the light type vector is 0.4. Use the cosine similarity algorithm to calculate the semantic similarity between semantic node vectors; Normalize the spectral intensity to the range of 0 - 1 (eliminating the influence of absolute brightness). Identify the peak wavelength of the spectral intensity through the local maximum algorithm (such as Savitzky-Golay filtering), convert the peak wavelength of the spectral intensity to the color temperature value (unit: K) through the CIE 1931 standard (for example, the color temperature value of a candle is 1900K, and the color temperature value of an oil lamp is 1850K), and calculate the absolute value of the color temperature difference between two light sources (for example, the color temperature difference between a candle and an oil lamp is 50K). Set the standard reference value of the color temperature difference to 100K, divide the absolute value of the color temperature difference by the standard reference value of the color temperature difference to obtain the color temperature difference ratio value. Use the exponential decay function (initial value is 1, decay constant is 0.5, input is the square of the color temperature difference ratio value to achieve non-linear decay) to calculate the decay coefficient, and multiply the decay coefficient by the semantic similarity to obtain the semantic node correlation weight. For example, the decay coefficient is 0.882, and the semantic node correlation weight is 0.53.
[0078] It should also be noted that: The double constraints of the BERT model and the exponential decay of the color temperature difference can distinguish the subtle differences between similar semantic nodes (such as the spectral peak differences between "Dougong_candlelight" and "Dougong_oil lamp"), improving the discriminative ability of the generator for historical light sources.
[0079] Multiply the time node similarity weight by the spatial node distance weight, and then multiply by the square root of the semantic node correlation weight (using the square root to avoid excessive attenuation) to obtain the comprehensive weight of the historical pattern. For example, the time node similarity weight is 0.97, the spatial node distance weight is 0.82, the semantic node correlation weight is 0.53, and the comprehensive weight of the historical pattern is 0.58.
[0080] Perform a sliding window mean filter (60 - second window) on the humidity data, calculate the humidity change rate per minute. When the change rate exceeds 5% (for example, the humidity rises from 60% to 80% within 1 minute), it is determined as an environmental mutation event. Use the Kriging interpolation method to generate a two - dimensional humidity grid with a resolution of 0.5 meters centered on the mutation area. For example, the humidity in area A increases by 2% per meter from south to north, forming a humidity field with the gradient direction pointing due north; Use the Sobel operator to convolve the humidity grid, calculate the gradient components in the east - west and north - south directions respectively, and synthesize the gradient vector. For example, a humidity field increasing from south to north will generate a gradient vector pointing north.
[0081] Predict the deformation of ancient building components through Fick's law of diffusion and finite element analysis: Based on the three - dimensional point cloud data of building components, discretize the ancient building components into tetrahedral element grids, and assign an initial humidity value (obtained by mapping the Kriging interpolation result) to each grid node. Set the expansion coefficient according to the material of the ancient building component (such as wood, stone), and calculate the humidity diffusion process through Fick's law of diffusion (the humidity diffusion coefficient is 1.2×10⁻ 6 m² / s), simulate the propagation process of humidity inside the material, consider the humidity exchange between the material surface and the environment, and predict the humidity distribution inside the material; Perform a dot - product operation on the gradient vector and the point cloud normal vector to quantify the humidity change rate along the normal direction; Multiply the humidity change rate by the expansion coefficient to obtain the strain value caused by humidity change, use the linear elastic constitutive relationship to associate the strain value with stress, establish a mechanical equilibrium equation, and set the elastic modulus and Poisson's ratio of wood and stone. Use the Newton - Raphson method to iteratively solve the displacement field to obtain the overall deformation of the material.
[0082] Use min - max normalization to normalize the peak wavelength of the candlelight intensity and spectrum to the range of 0 - 1.
[0083] Select the historical light source parameters with the highest comprehensive weight in the historical mode. Reduce the redundant information by reducing the dimensionality of the historical light source parameters (such as the spectral intensity distribution of "Ming Dynasty oil lamp") to 512 dimensions through principal component analysis (PCA). Concatenate the real-time light source parameters (256 dimensions) and the historical light source parameters (512 dimensions) into a 768-dimensional vector, add 256-dimensional random noise to generate a 1024-dimensional light source vector. Input the light source vector into the generator to obtain virtual light source parameters. The structure of the generator is a 5-layer fully connected network and a 1-layer output layer. The number of neurons in each fully connected network layer is 1024, 512, 256, 128, and 64. The activation function is LeakyReLU (α = 0.2). The output layer is divided into 4 branches: virtual intensity, virtual direction, virtual spectrum, and noise. Virtual intensity: Extract the first 32 dimensions from the 64-dimensional data output by the 5th fully connected network layer, input it into the Sigmoid function to generate a value in the range of 0-1, and map the output of the Sigmoid function (0 to 1) to the range of 800-1000 lx (for example: 0.8 → 880 lx). The virtual direction includes virtual azimuth and virtual elevation angle. Virtual azimuth: Extract the last 32 dimensions from the 64-dimensional data output by the 5th fully connected network layer, input it into the Tanh function to generate a value in the range of 0-1, and map the output of the Tanh function (-1 to 1) to 0-360° (for example: 0.25 → 45°). The virtual elevation angle maps the output of the Tanh function (-1 to 1) to 0-90° (for example: 0.33 → 30°). Virtual spectrum: Restore the 64-dimensional data output by the 5th fully connected network layer to the spectral distribution in the 400-780 nm band through inverse principal component analysis. Noise: Add Gaussian noise with a mean of 0 and a standard deviation of 0.01 to all 64-dimensional data output by the 5th fully connected network layer to simulate sensor errors (for example: ±2 lx).
[0084] It should also be noted that: Restoring the spectral distribution through inverse principal component analysis combined with Gaussian noise simulation can generate a natural spectral curve that meets the CIE standard (such as a 1900K color temperature of a candle), avoiding spectral distortion.
[0085] Calculate the point-by-point difference between the virtual spectrum and the spectrum in the 380 - 780 nm band, with a threshold ≤ 0.05. For example, if the spectrum peak is at 560 nm (intensity 0.95) and the virtual spectrum peak is at 558 nm (intensity 0.92), then the MSE (mean square error) = 0.03 (qualified). Generate a binary image comparison between the shadow mask and the real shadow. The PSNR (shadow mask quality) ≥ 40 dB. For example, if the maximum pixel value of the shadow mask is 255 and the MSE = 6.5, then the PSNR ≈ 42 dB (qualified). Embed the color temperature constraint and the direction constraint. If the color temperature of the virtual light source exceeds the documented range (e.g., for a candle, it should be in the range of 1850 - 1950 K), then regenerate the virtual intensity. If the virtual direction error exceeds the measured accuracy of the six-axis IMU (±0.1°), then regenerate the virtual direction.
[0086] It should also be noted that: the construction of the spatio-temporal - semantic three-dimensional weight system realizes the accurate spatio-temporal association between the historical lighting mode and the building components. The dimensionality reduction of historical light source parameters and the dynamic constraints of the generator ensure that the virtual light source matches the physical characteristics documented in terms of color temperature, direction, and spectral distribution, improving the restoration accuracy of historical scenes.
[0087] S3: Based on the virtual light source parameters, perform real-channel and historical-channel rendering, obtain the user's line-of-sight direction to calculate the user's perspective, and blend virtual and real light and shadow.
[0088] The specific steps are as follows.
[0089] Real-channel rendering: Input the virtual light source parameters into the Lumen Global Illumination Engine of Unreal Engine 5, bind them to the material nodes of the three-dimensional model of the ancient building, map the virtual light source parameters to the surface shader through the material editor. For example, bind the color temperature of the candlelight to the self-illumination channel of the material, and synchronize the dynamic direction light angle to the normal map offset. Use the ray tracing denoising algorithm to perform noise suppression and detail enhancement for high-density sampling of the brackets and eaves corners of the ancient building. Set 1024 ray samples per pixel, and focus on optimizing the anti-aliasing of the shadow edges and the multiple reflection calculations of the hollow structures (such as the multiple reflection calculations of light at the hollow part of the eaves corner). Reduce light noise through multiple importance sampling (MIS), and at the same time use the geometric information in the depth buffer to accelerate the ray intersection calculation; Adjust the material reflectivity according to the humidity data. For example, when the humidity exceeds 70%, the surface reflectivity of the wood decreases by 15% to simulate the enhanced diffuse reflection caused by moisture absorption; For every 5°C increase in temperature, the highlight reflection range of the metal components expands by 10% to simulate the thermal expansion effect.
[0090] It should also be noted that: 1024 light samplings per pixel and multiple importance sampling (MIS) can eliminate the multiple reflection noise of the openwork structure of the brackets and enhance the anti-aliasing effect of the shadow edges.
[0091] Historical channel rendering: Adjust the roughness of the ancient building materials according to the historical light source intensity. For example, for every 10 lux increase in the candlelight intensity, the roughness of the ancient building materials decreases by 0.03, simulating the increase in smoothness of the wood surface due to the wear of the oxide layer under low light. Adjust the roughness of the ancient building materials using a time-based sine wave, for example, at noon, the roughness of the ancient building materials decreases by 5% (thermal expansion of the wood fills the microcracks), and at midnight, it increases by 5% (cold shrinkage causes surface microcracks); map the roughness of the ancient building materials to the normal map through physically based rendering (PBR). For example, at the tenon and mortise joints of the brackets, add 0.1 roughness of the ancient building materials to simulate historical usage marks, and add periodic noise textures to the surface of the glazed tiles at the eaves corners to restore the weathering and erosion effect.
[0092] It should also be noted that: Adjusting the roughness of the ancient building materials using a time-based sine wave (±5% from noon to midnight) can simulate the changes in surface microcracks caused by the thermal expansion and contraction of the wood, enhancing the dynamic authenticity of the historical scene.
[0093] Obtain the user's line of sight direction through the built-in sensors of the head-mounted display, calculate the user's viewing angle (the angle between the user's line of sight direction and the light source direction). When the user is facing the light source directly (the angle ≤ 30°), the real channel weight is 70%; after the angle exceeds 30°, the real channel weight decays according to an exponential curve to avoid the misalignment of the virtual projection and the real scene during side viewing; when the user's line of sight is perpendicular to the normal vector of the light source (such as when viewing the side of the brackets from the side), automatically increase the historical channel weight by 15%, and blur the boundary between the virtual and real shadows through a bilateral filtering algorithm to eliminate pixel-level aliasing; use an eye tracker to capture the user's gazing area, generate a 4K high-resolution circular area with a diameter of 20° centered on the fixation point, and downsample the outer area in three levels according to the distance (2K → 1080p → 720p) to ensure rich details in the visual focus area; divide the non-gazing area into 32×32 pixel blocks, use the ETC2 compression algorithm to reduce the texture data volume, and perform parallel processing of the lighting calculation through the GPU asynchronous pipeline. For example, only retain the low-frequency shadow information in the non-gazing area, and dynamically complete the high-frequency details through motion vector interpolation.
[0094] It should also be noted that: The exponential decay rule of the user viewing angle weight (the weight decreases when the angle > 30°) can avoid the misalignment of the virtual light source projection and the physical shadow during side viewing, enhancing the immersion of the large viewing angle scene.
[0095] Combine ToF depth camera data with 3D point cloud data of building components to calculate the occlusion relationship between real objects and virtual models. For example, when the user is close to the bracket (distance <1 meter), the shadow of the real light source is only projected on the physical surface, and the projection of the virtual light source extends to the damaged part of the digital repair (such as missing carved patterns). Compare the depth values of the real shadow and the virtual projection, cover the far shadow with the near shadow, and achieve transition through Alpha blending. For example, when the shadow depth of the real bracket is 0.5 meters and the virtual projection is 1.2 meters, only the real shadow is rendered, and the occlusion boundary is Gaussian blurred with a width of 2 pixels to make the error of the virtual-real transition zone ≤0.1 mm to avoid hard boundaries visible to the naked eye. Render the picture at 4K resolution in the user's gaze area, and use anti-aliasing and detail enhancement technology to ensure that the processing time of each frame does not exceed 8 milliseconds to maintain a smooth visual effect. Intelligent rendering is used to render the picture at 1080 resolution in the non-gaze area, only half of the pixels are calculated, and the image is completed through interpolation of adjacent frame information, ensuring the continuity of the picture while reducing the GPU load.
[0096] It should also be noted that the user perspective adaptive rendering strategy combined with the virtual-real occlusion fusion algorithm solves the hard boundary problem of virtual-real light and shadow in AR / VR scenes. Through gaze point rendering and dynamic resolution allocation, it reduces the GPU load while ensuring 4K details in the visual focus area, achieving a balance between immersive interaction and hardware performance.
[0097] S4: Optimizing the generator’s neural network weights via incremental training.
[0098] The specific steps are as follows:
[0099] The time node similarity weights of the lighting mode are adjusted proportionally according to user preferences. The specific rule is: for every 10% increase in usage rate, the time node similarity weight of the lighting mode increases by 2%, but the upper limit does not exceed 150% of the time node similarity weight of the original lighting mode. For example, the user usage rate of "Tang Dynasty Dusk Candlelight" increases by 40%, and the time node similarity weight of the lighting mode increases by 8%.
[0100] It should also be noted that the dynamic growth of time weights (a 2% increase for every 10% usage rate) can strengthen historical lighting patterns with high frequency visits (such as "Tang Dynasty dusk candlelight") and avoid overfitting of unpopular patterns due to balanced weights.
[0101] The generator is compressed from 32-bit floating point (FP32) to 8-bit integer (INT8), the amount of computation is reduced by merging fully connected layers, and the generator is deployed on the NVIDIA Jetson AGX device to support synchronous calling of AR glasses and VR helmets, ensuring that the multi-terminal rendering delay is ≤15ms.
[0102] It should also be noted that: the combination of INT8 quantization compression and Jetson AGX edge computing can control the rendering latency within 15 ms and support synchronous calls of AR / VR multi-terminals.
[0103] Compare the virtual spectrum (380 - 780 nm band) with the spectrum band by band every day, and calculate the overlapping ratio of the virtual spectrum intensity distribution and the spectrum intensity distribution. For example, if the spectrum intensity of a certain band (such as 560 nm) is 0.95 and the virtual spectrum is 0.92, then the band matching degree is 97%. The average matching degree of the full band needs to be ≥ 85%; if the average matching degree is lower than 85% (such as 82%) for three consecutive days, then start the incremental training process to optimize the neural network weights of the generator, collect the abnormal data of the most recent seven days, including low-matching spectra, high-deformation error point cloud coordinates, and temperature and humidity mutation records, align them by time and store them in the training pool, lock the first three fully connected layers of the generator (1024→512→256 neurons), retain the ability to extract the spectrum intensity, direction angle, and color temperature of the historical light source, and only train the last two fully connected layers (128→64 neurons) and the output layer to adapt to the noise pattern and abnormal data in the new environment. Use the AdamW optimizer to optimize the neural network weights of the generator, set the learning rate to 0.0001, the batch size to 32, and train for at most 50 rounds. If the validation loss does not decrease for five consecutive rounds, terminate the training in advance. Input abnormal data, calculate the spectrum error (the difference between the spectrum intensity and the virtual spectrum) and the direction error (the difference between the light direction and the virtual direction), and calculate the total loss in combination with the KL divergence. Update the unfrozen neural network weights using the backpropagation method according to the total loss. Set the gradient clipping threshold to 1.0, and limit the single weight change amplitude within 5% to ensure the training stability.
[0104] It should also be noted that: when the matching degree < 85% for three consecutive days, start incremental training, and only update the last two layers of the generator network, which can quickly adapt to the new environment noise pattern (such as seasonal atmospheric scattering changes) and avoid the computational overhead of full-scale training.
[0105] Compare the three-dimensional coordinates of the ancient building components with the virtual coordinates point by point, and calculate the displacement difference of all points (such as the virtual coordinate deviates 1.2 mm from the measured coordinate). If the average error > 1 mm, then the wood expansion coefficient is increased from 0.005 mm / %RH to 0.006 mm / %RH.
[0106] It should also be noted that: the incremental training mechanism and the dynamic parameter adjustment strategy, while maintaining the core historical features, can quickly respond to long-term changes such as temperature and humidity mutations or material aging through local network updates, ensuring that the geometric deformation prediction error of the digital twin model is stably maintained at the sub-millimeter level for a long time.
[0107] This embodiment also provides a digital twin processing system for cultural and tourism virtual reality, including:
[0108] The preprocessing module collects data of ancient building components, including three-dimensional coordinates, point cloud normal vectors, reflectivity, light intensity, light direction, spectral intensity distribution, and humidity data, and preprocesses the data of ancient building components;
[0109] The generator module defines time nodes, space nodes, and semantic nodes, calculates the similarity weight of time nodes, the distance weight of space nodes, the correlation weight of semantic nodes, and the comprehensive weight of historical patterns, performs mutation detection on humidity data, predicts the deformation amount of ancient buildings through Fick's diffusion law and finite element analysis, reduces the dimension of the historical light source parameters with the highest comprehensive weight of historical patterns and stitches them with real-time light source parameters, and inputs them into the generator to obtain virtual light source parameters;
[0110] The mixing module performs real and historical channel rendering based on the virtual light source parameters, obtains the user's line of sight direction to calculate the user's perspective, and mixes virtual and real light and shadow;
[0111] The optimization module optimizes the neural network weights of the generator through incremental training.
[0112] This embodiment also provides a computer device, which is applicable to the situation of the digital twin processing method for cultural and tourism virtual reality, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the digital twin processing method for cultural and tourism virtual reality proposed in the above embodiment.
[0113] The computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0114] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the digital twin processing method for cultural and tourism virtual reality proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0115] In summary, the present invention solves the problem of the inability to accurately reproduce the historical lighting pattern by: defining time nodes, space nodes and semantic nodes, calculating time and space weights by combining the Gaussian kernel function, calculating semantic similarity by the BERT model, and reducing the dimension of the historical light source parameters with the highest comprehensive weight of the historical patterns and inputting them into the generator to generate virtual light source parameters, and dynamically associating the lighting parameters (such as candlelight color temperature, spectral peak) in historical documents with real-time environmental data; detecting humidity mutations by sliding window filtering, generating a humidity field by Kriging interpolation, calculating the gradient vector by the Sobel operator, and predicting the deformation amount by combining Fick's diffusion law and finite element analysis, thus solving the problem of inaccurate prediction of the deformation of ancient buildings caused by humidity changes.
[0116] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A digital twin processing method for cultural and tourism virtual reality, characterized in that: including Collecting data of ancient building components, including three-dimensional coordinates, point cloud normal vectors, reflectivity, light intensity, light direction, spectral intensity distribution, and humidity data, and preprocessing the data of ancient building components; Defining time nodes, space nodes, and semantic nodes, calculating the similarity weight of time nodes, the distance weight of space nodes, the correlation weight of semantic nodes, and the comprehensive weight of historical patterns, performing mutation detection on humidity data, predicting the deformation of ancient buildings through Fick's diffusion law and finite element analysis, reducing the dimension of historical light source parameters and splicing them with light source parameters, and inputting them into the generator to obtain virtual light source parameters; Performing real-channel and historical-channel rendering, obtaining the user's line-of-sight direction to calculate the user's perspective, and mixing virtual and real light and shadow; Optimizing the neural network weights of the generator through incremental training; The specific steps for performing mutation detection on humidity data and predicting the deformation of ancient buildings through Fick's diffusion law and finite element analysis are as follows Performing sliding window mean filtering on humidity data, calculating the humidity change rate per minute, and determining environmental mutation events; Generating a two-dimensional humidity grid using Kriging interpolation, performing convolution on the humidity grid using the Sobel operator, and calculating the gradient vector; Generating a humidity distribution based on Fick's diffusion law, calculating the humidity change rate according to the gradient vector and the point cloud normal vector, calculating the strain value in combination with the expansion coefficient, and calculating the overall deformation of the material through the mechanical equilibrium equation in combination with finite element analysis.
2. The digital twin processing method for cultural and tourism virtual reality according to claim 1, characterized in that: The calculation of the similarity weight of time nodes, the distance weight of space nodes, the correlation weight of semantic nodes, and the comprehensive weight of historical patterns refers to using the Gaussian kernel function to calculate the similarity weight of time nodes and the distance weight of space nodes; Calculating the correlation weight of semantic nodes through the BERT model in combination with spectral intensity, and calculating the comprehensive weight of historical patterns according to the similarity weight of time nodes, the distance weight of space nodes, and the correlation weight of semantic nodes.
3. The digital twin processing method for cultural and tourism virtual reality according to claim 1, characterized in that: The specific steps for performing real-channel and historical-channel rendering and mixing virtual and real light and shadow in combination with the user's perspective are as follows Performing real-channel rendering using the lumen global illumination engine in combination with the ray tracing denoising algorithm; Adjusting the roughness of the ancient building material using a time-based sine wave, and mapping the roughness of the ancient building material to the normal map through physically based rendering for historical-channel rendering; Adjusting the real-channel weight according to the user's perspective, eliminating the virtual-real boundary in combination with the edge blur algorithm, capturing the user's gaze area based on the eye tracker, and updating the light and shadow in the non-gaze area through GPU asynchronous calculation delay.
4. The digital twin processing method for cultural and tourism virtual reality according to claim 1, characterized in that: The specific steps for reducing the dimension of historical light source parameters, splicing them with light source parameters, and inputting them into the generator to obtain virtual light source parameters are as follows The light source parameters refer to light intensity, light direction, and spectral intensity; The virtual light source parameters refer to virtual intensity, virtual direction, and virtual spectrum; Reducing the dimension of historical light source parameters through principal component analysis, splicing them with light source parameters, and adding random noise at the same time to obtain a light source vector; Obtaining the virtual intensity by linearly mapping the output of the Sigmoid function; Obtaining the virtual direction by linearly mapping the output of the Tanh function; Restoring the output of the fully connected layer by inverse PCA to obtain the virtual spectrum.
5. The digital twin processing method for cultural and tourism virtual reality according to claim 1, characterized in that: The optimization of the neural network weights of the generator through incremental training refers to collecting abnormal data, setting an optimizer and its parameters, inputting the abnormal data into the generator to calculate the total loss, and using the backpropagation method to update the neural network weights.
6. The digital twin processing method for cultural and tourism virtual reality according to claim 1, characterized in that: The preprocessing of the ancient building component data refers to aligning the ancient building component data along the time axis and coordinates.
7. A digital twin processing system for cultural and tourism virtual reality, based on the digital twin processing method for cultural and tourism virtual reality according to any one of claims 1 to 6, characterized in that: Including, A preprocessing module that collects ancient building component data, including three-dimensional coordinates, point cloud normal vectors, reflectivity, light intensity, light direction, spectral intensity distribution, and humidity data, and preprocesses the ancient building component data; A generator module that defines time nodes, space nodes, and semantic nodes, calculates the similarity weights of time nodes, distance weights of space nodes, correlation weights of semantic nodes, and comprehensive weights of historical patterns, performs mutation detection on humidity data, predicts the deformation of ancient buildings through Fick's diffusion law and finite element analysis, reduces the dimensionality of historical light source parameters and concatenates them with light source parameters, and inputs them into the generator to obtain virtual light source parameters; A mixing module that performs real-channel and historical-channel rendering and mixes virtual and real light and shadow in combination with the user's perspective; An optimization module that optimizes the neural network weights of the generator through incremental training.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the digital twin processing method for cultural and tourism virtual reality according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the digital twin processing method for cultural and tourism virtual reality according to any one of claims 1 to 6.
Citation Information
Patent Citations
Real-time rendering method for dynamic environment illumination
CN117011446A
Smart text, travel and digital twin interaction system
CN118822792A