Artificial intelligence-driven VR cultural heritage digital display system
Through the AI-driven VR cultural heritage digital display system, combined with the analysis and optimization of multiple modules, the problem of insufficient details restoration in the reconstruction of virtual cultural heritage is solved, and a virtual display with high accuracy and authenticity is achieved, enhancing the audience's immersion and trust in the display.
Patent Information
- Application Number
- CN202510310050.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to accurately restore missing or damaged details in the reconstruction of virtual cultural heritage, resulting in insufficient authenticity and accuracy of virtual reconstruction, affecting the audience's sense of immersion and trust in the display.
The digital display system of VR cultural heritage driven by artificial intelligence is adopted, and accurate evaluation and optimization of virtual reconstruction of cultural heritage is achieved through multiple modules such as data collection, three-dimensional modeling, historical consistency parameter evaluation, physical consistency parameter evaluation, machine learning model, deviation division and dynamic optimization. Through multi-dimensional analysis and dynamic optimization, the system adjusts the details of the virtual display to ensure its consistency with historical facts and physical environment.
It improves the accuracy and historical authenticity of virtual cultural heritage display, enhances the audience's immersion and trust in the display, and enhances the value of education and entertainment.
Smart Images

Figure CN120182544A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital display, and particularly to an AI-driven VR digital display system for cultural heritage. Background Art
[0002] AI-driven VR digital display of cultural heritage refers to the process of presenting and displaying cultural heritage by combining artificial intelligence (AI) technology with virtual reality (VR) technology through digital means. This technology enables historical heritage to be reproduced in an immersive and interactive manner. Viewers can not only view digital models of cultural relics or sites but also interact with them and even experience the reproduction of historical scenarios. For example, through AI technology, the images, sounds, and structures of cultural heritage are analyzed, and virtual reality technology can integrate these analysis results into a 3D environment. Users can "visit" ancient civilizations, artworks, or sites as if they were on the spot through VR devices. For example, in virtual museum projects, the applications of reconstructing cultural heritage using 3D scanning and AI algorithms reflect the great potential of AI and VR technologies in the protection and display of cultural heritage.
[0003] The prior art has the following deficiencies:
[0004] In the reconstruction of virtual cultural heritage, although 3D scanning and AI algorithms can generate high-quality digital models, due to the lack of complete archaeological evidence or lost physical objects, it is difficult to accurately restore some details (such as the original appearance of decorations, sculptures, or architectural elements). These technologies often rely on the physical data or photos of existing sites and may not be able to accurately restore the damaged or disappeared parts, resulting in the completion based on assumptions by AI algorithms during the reconstruction process, which may not fully reflect the historical authenticity. In addition, the core attraction of VR cultural heritage display lies in its ability to provide an immersive historical experience. However, if the details of virtual reconstruction differ greatly from the real history, viewers may doubt the authenticity of the reconstruction. This lack of realistic experience will undermine the immersion of viewers, causing them to feel distrustful or even dissatisfied during the interaction process, thereby reducing the educational and entertainment value of VR cultural heritage display. Summary of the Invention
[0005] The purpose of the present invention is to provide an AI-driven VR digital display system for cultural heritage to solve the deficiencies in the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: An AI-driven VR digital display system for cultural heritage, including a data collection module, a three-dimensional modeling module, a historical consistency parameter evaluation module, a physical consistency parameter evaluation module, a machine learning model module, a deviation division module, and a dynamic optimization module;
[0007] A data collection module for obtaining 3D point cloud data and high-resolution image data of the target cultural heritage through a 3D scanner and an image acquisition device;
[0008] A 3D modeling module for generating a virtual 3D model of the cultural heritage based on the 3D point cloud data and image data, and performing texture mapping on the model;
[0009] A historical consistency parameter evaluation module for extracting historical descriptions of the cultural heritage based on historical documents, archaeological reports, and information on historical relics, comparing them with the virtual 3D model to generate historical consistency parameters, and evaluating the deviation between the virtual model and historical facts;
[0010] A physical consistency parameter evaluation module for judging the physical rationality of the virtual 3D model according to the true lighting, texture, and physical properties of the material of the target cultural heritage, generating physical consistency parameters, and evaluating the deviation between the virtual model and the actual physical environment;
[0011] A machine learning model module for receiving the historical consistency parameters and physical consistency parameters, comprehensively evaluating the deviation between the details of the virtual reconstruction and historical facts, and generating a virtual reconstruction deviation score;
[0012] A deviation classification module for classifying the virtual cultural heritage reconstruction results into different deviation levels according to the virtual reconstruction deviation score, including a completely consistent level, a slightly deviated level, and a severely deviated level;
[0013] A dynamic optimization module for adjusting the details of the virtual cultural heritage display according to the deviation classification result, including fine-tuning the texture, lighting, or geometric shape of the slightly deviated area; re-acquiring data for the severely deviated area to optimize the accuracy of the virtual display.
[0014] Preferably, in the historical consistency parameter evaluation module, after analyzing the texture consistency between the surface texture of the virtual 3D model and the actual decoration texture of the historical relics, a texture consistency analysis index is generated. The method for obtaining the texture consistency analysis index is as follows:
[0015] For an image I(x, y), setting a key point p = (x, y), then the gradient of its local area The calculation formula is: ; In the formula, represents the change rate of the image intensity in the horizontal direction x, represents the change rate of the image intensity in the vertical direction y; calculate the gradient direction distribution of each area to obtain a 128-dimensional feature vector, match the descriptors of the virtual model texture and the historical relic texture for two different images, and calculate the matching degree between the descriptors , the expression is: ; where and respectively represent the values of the two descriptors in the i-th dimension. According to the number of matching descriptors and the matching quality, the optimal descriptor matching pair is selected. According to the matching result, the texture consistency index of the two images is calculated. The number of matched key points is set as , the total number of key points in the image is , then the calculation expression of the texture consistency analysis index is: .
[0016] Preferably, in the physical consistency parameter evaluation module, the light distribution difference index is generated by comparing the light distribution differences between the virtual model and the real environment. Among them, the method for obtaining the light distribution difference index is:
[0017] First, light data is collected from the virtual three-dimensional model and the real environment. From the virtual three-dimensional model, the light intensity of each sampling point is obtained through the rendering algorithm, denoted as , where i represents the i-th sampling point; through on-site measurement or using the light data set in the real environment, the light intensity of each sampling point is obtained, denoted as ;
[0018] For each sampling point i, the error value is the absolute value of the difference between the light intensities of the virtual model and the real environment. The calculation expression is: ; where is the light error of the i-th point, and are the light intensities of the real environment and the virtual model at this point respectively. The light errors of all sampling points are summed up to obtain the overall error , the expression is: ; where N is the total number of sampling points. Calculate the light distribution difference index , the expression is: .
[0019] Preferably, in the machine learning model module, it is used to convert the texture consistency analysis index and the light distribution difference index into a comprehensive feature vector, and use the comprehensive feature vector as the input of the machine learning model. The machine learning model takes predicting the virtual reconstruction deviation score value label for each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors for all virtual reconstruction deviation score value labels as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence and then the model training is stopped. The virtual reconstruction deviation score value is determined according to the model output result. Among them, the machine learning model is a polynomial regression model.
[0020] Preferably, in the deviation division module, the obtained virtual reconstruction deviation score value is compared with the gradient deviation threshold values. The gradient deviation threshold values include a first deviation threshold value and a second deviation threshold value, and the first deviation threshold value is less than the second deviation threshold value. The virtual reconstruction deviation score value is respectively compared with the first deviation threshold value and the second deviation threshold value.
[0021] If the virtual reconstruction deviation score value is greater than the second deviation threshold value, it indicates that the degree of virtual reconstruction deviation is high, and it is classified into the severe deviation level; if the virtual reconstruction deviation score value is greater than or equal to the first deviation threshold value and less than or equal to the second deviation threshold value, it indicates that the degree of virtual reconstruction deviation is medium, and it is classified into the minor deviation level; if the virtual reconstruction deviation score value is less than the first deviation threshold value, it indicates that the degree of virtual reconstruction deviation is low, and it is classified into the minor deviation level.
[0022] Preferably, in the dynamic optimization module, it is used to adjust the details of the virtual cultural heritage display according to the deviation division result, including fine-tuning the texture, lighting or geometric shape of the minor deviation area.
[0023] When the deviation score value falls within the minor deviation level range, the texture is fine-tuned through the adjustment formula: ; is the adjusted texture value, is the texture value at the position (x, y) in the current virtual model, is the target texture value, and α is the adjustment factor;
[0024] The fine-tuning of lighting includes adjusting the light source position or intensity in the virtual scene according to the deviation of the lighting distribution difference index LDE. The expression is: ; The adjusted lighting value, is the lighting value in the current virtual model, is the target lighting value; β is the adjustment factor, controlling the amplitude of lighting fine-tuning;
[0025] The fine-tuning of geometric shape is used to correct the small errors introduced by the modeling algorithm. When the surface of the virtual model does not exactly match the geometric features of the historical relic, the adjustment expression is: ; In the formula, is the adjusted geometric shape position, is the geometric position at the position (x, y) in the current virtual model, is the target geometric shape position, and γ is the geometric shape adjustment factor, controlling the amplitude of fine-tuning.
[0026] Preferably, for the severe deviation area, new data is obtained by using a three-dimensional scanner with higher precision or a high-resolution image acquisition device, and the data obtained by re-scanning is expressed as: ; represents the geometric data obtained by rescan, is the geometric data in the original virtual model, is the newly acquired data, is a function for fusing new data;
[0027] Through new scans and data acquisition, all important areas of the virtual model are updated to ensure that the regenerated virtual model has higher accuracy. The expression is: ; represents the geometric position of the final virtual model. If is an exact value, the adjusted data is used; otherwise, the new data obtained by rescan is used .
[0028] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:
[0029] 1. By integrating multiple modules such as data collection, 3D modeling, historical and physical consistency evaluation, machine learning analysis, and dynamic optimization, the present invention can achieve precise evaluation and optimization of the virtual reconstruction of cultural heritage. Through multi-dimensional analysis, including texture consistency, light distribution, and historical and physical consistency, the system generates a virtual reconstruction deviation score, divides it into different deviation levels according to the score result, and then performs dynamic optimization processing. For slightly deviated areas, the system optimizes its accuracy through fine-tuning of texture, light, and geometric shape; while for severely deviated areas, a higher-precision data re-acquisition method is adopted to ensure the authenticity and accuracy of virtual display.
[0030] 2. For the missing or damaged parts of historical heritage, the present invention comprehensively uses modern technologies and historical evidence to optimize virtual reconstruction, providing a more accurate and realistic display of cultural heritage. With the support of AI algorithms and machine learning models, the system can intelligently evaluate and adjust the details of the virtual model, greatly enhancing the immersion and educational value of virtual display, avoiding the deficiencies of simply relying on existing data, solving the problem of lack of authenticity faced by existing technologies in virtual cultural heritage reconstruction, and thus providing a more credible and immersive historical experience for viewers. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0032] Figure 1 is the system module diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] For the embodiments, please refer to Figure 1 As shown, the AI-driven VR cultural heritage digital display system in this embodiment includes a data collection module, a 3D modeling module, a historical consistency parameter evaluation module, a physical consistency parameter evaluation module, a machine learning model module, a deviation classification module, and a dynamic optimization module;
[0035] The data collection module is used to obtain the 3D point cloud data and high-resolution image data of the target cultural heritage through a 3D scanner and an image acquisition device;
[0036] The 3D modeling module is used to generate a virtual 3D model of the cultural heritage based on the 3D point cloud data and image data, and perform texture mapping on the model;
[0037] The historical consistency parameter evaluation module is used to extract historical descriptions of the cultural heritage based on historical documents, archaeological reports, and information on historical relics, and compare them with the virtual 3D model to generate historical consistency parameters and evaluate the deviation between the virtual model and historical facts;
[0038] The physical consistency parameter evaluation module is used to judge the physical rationality of the virtual 3D model according to the true lighting, texture, and material physical properties of the target cultural heritage, generate physical consistency parameters, and evaluate the deviation between the virtual model and the actual physical environment;
[0039] The machine learning model module is used to receive the historical consistency parameters and physical consistency parameters, comprehensively evaluate the deviation between the details of the virtual reconstruction and historical facts, and generate a virtual reconstruction deviation score;
[0040] The deviation classification module is used to divide the virtual cultural heritage reconstruction results into different deviation levels according to the virtual reconstruction deviation score, including the completely consistent level, the slightly deviated level, and the severely deviated level;
[0041] The dynamic optimization module is used to adjust the details of the virtual cultural heritage display according to the deviation classification result, including fine-tuning the texture, lighting, or geometric shape of the slightly deviated area; re-acquiring data for the severely deviated area to optimize the accuracy of the virtual display.
[0042] In the data collection module, a 3D scanner is used to capture the 3D point cloud data of cultural heritage, which is an accurate digital description of the surface shape of cultural heritage objects. The specific operations of this step include:
[0043] Select appropriate 3D scanning techniques: Laser scanning: The laser scanner scans the target object with a laser beam and records the distance (the time it takes for the laser reflection to return) and position information of each laser point. Laser scanners can generate very high-precision point cloud data and are suitable for the 3D reconstruction of large cultural heritage structures, such as the exterior or interior structures of the Colosseum in ancient Rome. Structured light scanning: By using a projected light spot and an image capture device, it analyzes the deformation of the light spot shape to obtain surface information. This method performs well in the reconstruction of objects with smaller details and is suitable for scanning delicate artworks such as sculptures and reliefs.
[0044] 3D scanning process: The scanner needs to scan the cultural heritage from different angles multiple times to ensure that all surface details of the object are captured. Each scanned data point represents an accurate position on the object's surface, and multiple point cloud data are combined to form a complete 3D model. During the scanning process, the scanner records the spatial coordinates of the point cloud and may also include color information (if a color laser scanner is used), or add texture through image matching techniques at a later stage.
[0045] Data accuracy and fusion: To ensure the high precision of the 3D point cloud, the resolution of the scanner needs to be set appropriately. High resolution can capture more delicate details, although this will also increase the complexity of data processing. The data scanned from multiple angles will be fused into a complete 3D point cloud through specialized point cloud registration techniques to solve the data differences caused by different scanning angles and positions.
[0046] In addition to 3D scanning, high-resolution image acquisition devices (such as digital cameras, drones, satellite imaging) are also used to capture the detailed image data of cultural heritage, mainly to provide materials for subsequent texture mapping and historical consistency analysis. The specific steps are as follows:
[0047] Image acquisition device selection: High-resolution camera: Used to take high-resolution photos of cultural heritage, capable of capturing surface textures, carving details, ornaments, etc. These image data usually have very high clarity (e.g., 8000x6000 pixels) to ensure the capture of minute details. Drone shooting: For large cultural heritage sites (such as the Colosseum in Rome), drones equipped with high-resolution cameras can obtain aerial views from different angles, providing more comprehensive panoramic data, which is crucial especially for understanding the overall structure and scale of the building. Satellite imaging: For inaccessible sites, satellite imaging technology can provide global image data from the air, especially suitable for the collection of geographical information of large-scale sites (such as urban sites or ancient building complexes).
[0048] Take panoramic photos of the site, taking images from multiple perspectives to ensure that every detail can be captured. The images from each perspective will be processed into a seamless panoramic image later. Through multi-exposure synthesis technology, images under different lighting conditions are taken to enhance the dynamic range of the images and capture detailed information. The acquisition device needs to be used synchronously with the 3D scanner to ensure that the images and 3D data can be correctly registered to form an integrated data source.
[0049] The 3D point cloud data and image data obtained through the scanner and image acquisition device often need to go through a series of preprocessing steps to ensure that the data can be efficiently used for subsequent 3D modeling and historical analysis:
[0050] Point cloud processing: The point cloud data needs to be denoised, smoothed, and simplified, etc., to reduce data redundancy and noise while retaining sufficient detailed information. Commonly used techniques include point cloud reduction, filtering, and hole filling, etc.
[0051] Image processing: The image data needs to be color-corrected, light-adjusted, and image stitched. The stitching technology seamlessly synthesizes multiple images into a single panoramic image to ensure high-quality texture information.
[0052] Data registration and fusion: By using computer vision techniques (such as feature matching, image registration algorithms), the image and point cloud data are aligned so that the image texture can be correctly mapped onto the 3D model later.
[0053] To ensure that the quality of the obtained 3D data and images meets high-precision requirements, the data collection module needs to conduct quality assessment and repair: Data quality inspection: Evaluate the accuracy of the point cloud data and images, and detect potential scanning errors, distortions, or data loss. Supplementary scanning: For areas with missing or low-quality data, repair is carried out by re-scanning or supplementing data (such as speculating and supplementing through existing literature and historical data).
[0054] The large amount of three-dimensional point cloud data and image data collected need to be effectively stored and transmitted for subsequent modeling and processing: Store the point cloud data and image data in appropriate file formats (such as.ply,.obj,.jpg,.tif, etc.) and manage them in a cloud platform or a distributed storage system to ensure data security and easy access. For large-scale data, adopt efficient data compression and transmission technologies to reduce storage space and accelerate the data transmission process, ensuring that the data can quickly flow to the subsequent three-dimensional modeling and machine learning modules.
[0055] In this application, the data collection module uses three-dimensional scanners and image acquisition devices to ensure that the cultural heritage digital display system can obtain high-quality and accurate raw data. These data are not only the basis for virtual reconstruction but also provide a solid foundation for subsequent historical consistency analysis, physical consistency assessment, and machine learning models. Through the collection, processing, and fusion of high-precision three-dimensional point cloud and image data, the system can accurately reconstruct cultural heritage and provide a high-fidelity immersive experience for virtual reality displays.
[0056] In the three-dimensional modeling module, three-dimensional point cloud data is obtained through laser scanners or structured light scanners to describe the three-dimensional surface shape of an object. Point cloud data processing is the first step in three-dimensional modeling. During the actual scanning process, point cloud data often contains noise points and incomplete regions. These noise points may be caused by scanner errors, environmental factors, or reflectivity changes. Denoising algorithms (such as Gaussian filtering, mean filtering, etc.) need to be used to remove these unnecessary data points. After removing the noise, smoothing algorithms (such as point cloud smoothing, feature-preserving smoothing, etc.) are used to process the remaining point cloud to ensure the smoothness and realism of the point cloud surface. Point cloud data often contains a large number of redundant points, resulting in excessive computational complexity. Without losing important details, the number of points in the point cloud is reduced through point cloud simplification algorithms (such as voxel grid method, octree method, etc.) to reduce the computational complexity and speed up the processing speed. Since three-dimensional scanning is generally carried out from multiple angles, the point cloud data captured from different angles needs to be aligned through registration algorithms (such as the Iterative Closest Point (ICP) algorithm) and the point cloud data from multiple perspectives is fused into a complete three-dimensional point cloud. The preprocessed point cloud data can generate a continuous three-dimensional mesh model through surface reconstruction algorithms (such as Poisson Surface Reconstruction, Alpha Shapes, Delaunay triangulation, etc.). This mesh model is composed of polygons (usually triangles) representing the surface of the object.
[0057] The image data provides the color and texture information of the object surface. The 3D modeling module maps this image data onto a 3D mesh to enhance the realism of the model. The collected high-resolution images (such as photos, drone-captured images) are matched with the 3D point cloud data, and the correspondence between the image and the 3D point cloud is determined through image registration techniques (such as feature matching, SURF, SIFT, etc.). Texture coordinates (UV coordinates) are generated for each 3D mesh face, that is, mapping the 2D image onto the 3D mesh surface. Each triangular face corresponds to a part of the image, ensuring that the model is visually consistent with the actual site.
[0058] The processed image data is applied to each face of the 3D model to generate texture mapping. This process usually includes: Texture stitching: Stitch the textures according to the images taken from different angles to ensure seamless connection between the images.
[0059] Texture detail enhancement: By adjusting parameters such as brightness, contrast, and hue of the image, ensure the naturalness and realism of the texture mapping.
[0060] Seamless transition: Solve the color difference or discontinuity problems between different images, and eliminate the seams through smooth processing of the image edges.
[0061] Based on the texture mapping, further add lighting and material information to the model. By simulating lighting conditions (such as ambient light, point light source, spotlight, etc.) and material properties (such as reflectivity, roughness, metallicity, etc.), make the appearance of the model closer to the visual effect of the actual site.
[0062] The PBR (Physically Based Rendering) technology can be used to simulate the real reflection, refraction and other optical properties of the object, making the performance of the model in the virtual reality environment more vivid and realistic.
[0063] After generating the basic 3D model and applying the texture, further enhance the model details to ensure its consistency with the appearance of the historical site.
[0064] Detail completion: For the details missing in the point cloud or image, by combining historical documents, archaeological data and expert opinions, use AI and machine learning algorithms to automatically predict and fill in the missing parts. For example, if a certain area lacks sculptures or decorations, the AI algorithm can infer the shape and details of the missing part based on the known historical data or data of similar sites.
[0065] Detail adjustment and optimization: After the model is generated, manually adjust or optimize through automated algorithms the details with certain deviations (such as some irregular curved surfaces, missing decorations, etc.). The adjusted model ensures that each part conforms to historical facts and archaeological discoveries.
[0066] After completing the 3D model and texture mapping, the system outputs high-quality 3D models, which serve as the basis for virtual cultural heritage display. These models can be exported in various formats (such as.obj,.fbx,.ply, etc.) as needed and can be directly used for virtual reality (VR) display, augmented reality (AR) display, or other digital applications. Different 3D formats are supported to be compatible with other software tools (such as virtual reality engines, visualization platforms, game engines). Common 3D model formats include.obj,.fbx,.stl,.ply, etc. The output model files will contain information such as vertices, normals, texture coordinates, colors, etc., ensuring that the models can be correctly rendered in the virtual environment.
[0067] To ensure the accuracy of cultural heritage display and the immersive experience of users, the 3D modeling module needs to be dynamically adjusted continuously according to user feedback and actual effects. User interaction feedback: In virtual reality, users interact with the virtual site through interactions (such as clicking, rotating, zooming in and out, etc.). The system will collect the user's behavior data in real time to adjust the detailed presentation of the model. Historical consistency and physical consistency feedback: According to the evaluation results of the machine learning model on historical consistency and physical consistency, dynamically adjust the details of the virtual model to optimize its deviation from historical facts.
[0068] In this application, the 3D modeling module combines the physical form and historical texture of cultural heritage through multiple processing steps, including 3D point cloud data processing, image matching and texture mapping, detail enhancement, etc., to generate high-precision and realistic virtual 3D models. The design of this module ensures the accuracy and detailed performance of virtual reconstruction, providing a solid foundation for the display of cultural heritage in virtual reality and enabling users to experience the style of ancient sites in an immersive way.
[0069] For the historical consistency parameter evaluation module, in order to evaluate the consistency between the virtual 3D model and historical facts, it is first necessary to extract relevant information from various historical materials. These materials usually include:
[0070] Historical documents: including ancient documents, local chronicles, historical books, etc., providing information such as descriptions, functions, appearances of cultural heritage.
[0071] Archaeological reports: Archaeological excavation reports provide detailed archaeological data about the site, including the structure of buildings, the placement of relics, the form of decorations, etc.
[0072] Historical relic data: including photos, descriptions, and archaeological analysis results of the discovered relics (such as sculptures, architectural fragments, murals, etc.).
[0073] Historical documents and archaeological reports are usually natural language texts that need to be parsed, information extracted, and classified through natural language processing (NLP) techniques. Common techniques include: Named Entity Recognition (NER): identifying important entities in the documents, such as building names, historical figures, locations, etc. Keyword extraction: extracting key descriptions of the appearance, function, materials, etc. of cultural heritage. Relation extraction: analyzing the relationships in the documents regarding cultural heritage, for example, "The walls of the Colosseum were built of xx stone", and extracting information such as the building and materials. Image data processing: processing the data of historical relics to extract morphological features, dimensions, etc. of the relics through image recognition techniques (such as object detection, image classification), providing a basis for comparing the model with historical facts.
[0074] The core task of historical consistency assessment is to compare the extracted historical descriptions with the virtual 3D model and find the differences between the two. Descriptions of the building appearance, decoration, dimensions, etc. extracted from historical documents and archaeological reports are matched with the specific structures in the virtual 3D model. Using natural language processing techniques to transform historical descriptions into a comparable form, for example: if the document mentions that "the central arena of the Colosseum has four walls, two of which are arched", it is necessary to check in the virtual 3D model whether these four walls exist and whether they conform to the arched structure described in the document. Historical descriptions are usually based on two-dimensional language and image information, while virtual 3D models are three-dimensional. Therefore, the system needs to fuse the data of the two through multi-modal data fusion technology for comparison: the spatial features (such as structural location, dimensions, shape) of historical descriptions are corresponded to the three-dimensional geometric structures in the virtual 3D model. Extract descriptions of textures and decorations from archaeological reports or relics and compare them with the textures in the virtual model to ensure the reproduction of historical details.
[0075] Historical consistency parameters are quantitative indicators obtained by comparing the differences between the virtual 3D model and the descriptions in historical documents and archaeological reports, and these parameters can reflect the deviation between the virtual model and historical facts.
[0076] Parameter generation: Geometric consistency: comparing the dimensions, proportions, and shapes of the virtual 3D model with the descriptions in historical documents or archaeological reports. For example, whether the height of a certain building in the virtual model is consistent with the height mentioned in the document.
[0077] Structural consistency: evaluating whether the structural elements (such as walls, arches, columns, etc.) in the model conform to the actual structures discovered archaeologically.
[0078] Texture consistency: evaluating whether the surface textures of the virtual 3D model are consistent with the actual decorations of historical relics, such as the patterns of murals, the carvings of stone sculptures, etc.
[0079] Temporal Consistency: Whether the virtual reconstructed model conforms to the architectural styles and techniques of historical periods. For example, whether certain decorations and architectural elements conform to the styles of the ancient Roman period.
[0080] After analyzing the texture consistency between the surface texture of the virtual 3D model and the actual decoration texture of the historical relic, a texture consistency analysis index is generated. The method for obtaining the texture consistency analysis index is as follows:
[0081] First, it is necessary to preprocess the texture images of the virtual 3D model and the historical relic to ensure that they are compared under the same scale and lighting conditions. Common preprocessing steps include: Grayscale conversion: Convert the image to a grayscale image to eliminate color interference. Normalization: Adjust the brightness and contrast of the image so that lighting and color differences do not affect feature extraction. Denoising: Use methods such as Gaussian filtering to remove image noise and ensure the accuracy of feature point extraction.
[0082] For an image I(x,y), its Gaussian blur is expressed as: ; By convolving the image I(x,y) and the Gaussian kernel G(x,y,σ) at different scales, a scale space is generated. Once the key points are detected, SIFT generates a descriptor (usually a 128-dimensional vector) for each key point. The descriptor is based on the local area around the key point and describes the local texture structure of the image.
[0083] The descriptor is usually generated by calculating the gradient direction distribution in the area around the key point. Given a key point p=(x,y), the gradient of its local area The calculation formula is: ; In the formula, represents the rate of change of the image intensity in the horizontal direction x, represents the rate of change of the image intensity in the vertical direction y.
[0084] After rotating, scale normalizing, and partitioning the local area of the descriptor, the gradient direction distribution of each area is calculated to obtain a 128-dimensional feature vector. After extracting the key point descriptors, the descriptors of the two images (virtual model texture and historical relic texture) are matched. Common methods include brute-force matching (Brute-ForceMatcher) or KNN matching (K-Nearest Neighbors). Calculate the matching degree between the descriptors The expression is: ; Among them, and They respectively represent the values of two descriptors (A and B) in the i-th dimension. Based on the number of matching descriptors and the matching quality, the optimal descriptor matching pairs are selected. If the matching degree between the virtual model texture and the historical relic texture is high, it is considered that the two images have good texture consistency.
[0085] According to the matching results, calculate the texture consistency index of the two images. Let the number of matched key points be , and the total number of key points in the image be , then the calculation expression of the texture consistency analysis index is: ; The value range of the texture consistency analysis index is from 0 to 1. The closer the value is to 1, the higher the similarity and the better the consistency between the virtual 3D model texture and the historical relic texture.
[0086] The physical consistency parameter evaluation module is used to obtain the lighting information in the real scene, including the light source type (such as point light source, directional light source), light intensity, color, and physical characteristics related to the light source (such as shadows, reflections, refractions, etc.). The lighting information can be obtained through on-site shooting, environmental sensors, historical documents, or optical measurement devices. Standardize the collected lighting data to make the lighting environment of the virtual model consistent with the actual lighting environment.
[0087] Based on the lighting conditions of the virtual 3D model, simulate the lighting effects in the virtual environment, analyze the interaction between the light and the object surface, and generate a lighting model in the virtual environment. Use lighting rendering models (such as Phong model, Blinn-Phong model) to simulate the interaction between the reflection of the surface material and the light source. Use global lighting models (such as radiosity method, light transport equation) to simulate the indirect lighting effects in the environment and increase the realism of the scene. Apply techniques such as shadow mapping and light mapping to make the interaction between the light source and the surface texture more accurate.
[0088] Collect the physical characteristics (such as reflectivity, refractive index, glossiness, roughness, etc.) of the materials and texture attributes (such as stone, brick, paint, etc.) of the target cultural heritage through archaeological data or on-site investigations. Data sources: archaeological reports, on-site tests, literature records, physical measurements, etc. Standardize the collected material data to ensure its consistency with the materials in the virtual 3D model.
[0089] Accurately simulate the physical properties of materials in a virtual environment. For example, by setting parameters such as different reflectivities, refractive indices, roughnesses, etc., the virtual model presents an appearance consistent with the actual site under the action of light. Use physically based rendering (PBR) methods to simulate the reflection, refraction, scattering, etc. effects of materials based on physical properties. PBR technology is widely used in material processing in games and movies and can more realistically simulate the optical effects of materials. Reflection and refraction: Simulate the optical effects on the material surface by using environment mapping and reflectance maps. Surface roughness: Refine surface details through normal mapping and roughness maps, thereby affecting the reflection and refraction behavior of light.
[0090] Based on the lighting and material properties of the virtual 3D model, combined with the lighting and material data in the actual physical environment, calculate the physically consistent parameters. These parameters evaluate the physical differences between the virtual model and the real environment.
[0091] Lighting consistency: The consistency can be calculated by comparing the lighting distribution differences between the virtual model and the real environment.
[0092] Material consistency: By comparing the material reflection characteristics of the virtual 3D model and historical relics.
[0093] Generate a lighting distribution difference index by comparing the lighting distribution differences between the virtual model and the real environment. Among them, the method for obtaining the lighting distribution difference index is as follows:
[0094] First, collect lighting data from the virtual 3D model and the real environment. Assume that we have multiple sampling points in the environment, and obtain the lighting intensity at each point through lighting measurement.
[0095] Lighting intensity data in the virtual model: From the virtual 3D model, obtain the lighting intensity at each sampling point through the rendering algorithm, denoted as , where i represents the i-th sampling point.
[0096] Lighting intensity data in the real environment: Through on-site measurement or using the lighting dataset in the real environment, obtain the lighting intensity at each sampling point, denoted as .
[0097] For each sampling point i, the error value is the absolute value of the difference between the lighting intensities of the virtual model and the real environment, and the calculation expression is: ; where, is the lighting error at the i-th point, and are the light intensities of the real environment and the virtual model at this point respectively. The light errors of all sampling points are summed up to obtain the overall error: ; where N is the total number of sampling points, and the light distribution difference index is calculated , and the expression is: ; where the value of LDDI is in the range of [0, 1]. The closer the value is to 1, the more consistent the light distribution of the virtual model and the real environment; the closer the value is to 0, the greater the light difference between the two.
[0098] The machine learning model module is used to convert the texture consistency analysis index and the light distribution difference index into a comprehensive feature vector, take the comprehensive feature vector as the input of the machine learning model, take predicting the virtual reconstruction deviation score value label as the prediction target with each group of comprehensive feature vectors, and take minimizing the sum of the prediction errors of all virtual reconstruction deviation score value labels as the training target to train the machine learning model until the sum of the prediction errors reaches convergence and then stop the model training, and determine the virtual reconstruction deviation score value according to the model output result, where the machine learning model is a polynomial regression model.
[0099] The deviation division module is used to divide the virtual cultural heritage reconstruction result into different deviation levels according to the virtual reconstruction deviation score, including the completely consistent level, the slightly deviated level and the severely deviated level;
[0100] Compare the obtained virtual reconstruction deviation score value with the gradient deviation thresholds, where the gradient deviation thresholds include a first deviation threshold and a second deviation threshold, and the first deviation threshold is less than the second deviation threshold, and compare the virtual reconstruction deviation score value with the first deviation threshold and the second deviation threshold respectively;
[0101] If the virtual reconstruction deviation score value is greater than the second deviation threshold, it means that the virtual reconstruction deviation degree is high, and it is divided into the severely deviated level; if the virtual reconstruction deviation score value is greater than or equal to the first deviation threshold and less than or equal to the second deviation threshold, it means that the virtual reconstruction deviation degree is medium, and it is divided into the slightly deviated level; if the virtual reconstruction deviation score value is less than the first deviation threshold, it means that the virtual reconstruction deviation degree is low, and it is divided into the slightly deviated level. It can be considered that the virtual cultural heritage reconstruction is very successful, and the display effect meets the historical authenticity standard, and it can directly enter the display link.
[0102] The dynamic optimization module is used to adjust the details of the virtual cultural heritage display according to the deviation division result, including fine-tuning the texture, light or geometric shape of the slightly deviated area;
[0103] In the slightly deviated area, the goal is to make the details of the virtual model closer to the real historical environment by fine-tuning the texture, light or geometric shape.
[0104] When the deviation score value falls within the range of the minor deviation level, the texture is fine-tuned through an adjustment formula: ; is the adjusted texture value, is the texture value at the position (x, y) in the current virtual model, is the target texture value, that is, the ideal texture obtained according to historical documents or real relic information. α is the adjustment factor, which controls the amplitude of the fine-tuning. Usually, the α value in the minor deviation area is small to ensure that the adjustment is not excessive.
[0105] The adjusted texture value is obtained by and the current texture value through weighted fusion to ensure the naturalness and precision of the adjustment.
[0106] The fine-tuning of lighting usually includes adjusting the position or intensity of the light source in the virtual scene according to the deviation of the lighting distribution difference index LDE. The expression is: ; The adjusted lighting value, is the lighting value in the current virtual model, is the target lighting value, obtained according to the real lighting distribution and historical document data; β is the adjustment factor, which controls the amplitude of the lighting fine-tuning. In the minor deviation area, β is small to avoid excessive adjustment.
[0107] The fine-tuning of the geometry is usually used to correct the minor errors introduced by the modeling algorithm, especially when the surface of the virtual model does not exactly match the geometric features of the historical relic. The adjustment expression is: ; In the formula, is the adjusted geometric position, is the geometric position at the position (x, y) in the current virtual model, is the target geometric position, the ideal geometric position obtained according to historical relics or archaeological data, and γ is the geometric shape adjustment factor, which controls the amplitude of the fine-tuning. For the minor deviation area, the value of γ is small to prevent a large impact on the model structure.
[0108] For the severe deviation area, the fine-tuning may not be sufficient to achieve high-precision reconstruction. Therefore, it is necessary to optimize the accuracy of the virtual display through data re-acquisition. Data re-acquisition usually refers to re-scanning or obtaining more accurate archaeological data to improve the virtual model.
[0109] For the severe deviation area, new data can be obtained by using a three-dimensional scanner with higher precision or a high-resolution image acquisition device. The core of this process is to ensure that the details of the site are captured through more accurate scanning to reduce the deviation. The data obtained from the re-scanning is expressed as: ; represents the geometric data obtained through rescan, is the geometric data in the original virtual model, is the newly acquired data, which may include 3D point cloud data, high-resolution images, etc. from different perspectives. A function for fusing new data, which can be an algorithm based on multi-view geometric reconstruction to ensure seamless fusion of the new scan data with the existing model.
[0110] Through new scans and data acquisition, all important areas of the virtual model are updated, especially the details of texture, lighting, and geometry, to ensure that the regenerated virtual model has higher accuracy. The expression is: ; represents the geometric position of the final virtual model. If (the data after fine-tuning) is accurate enough, then the adjusted data is used; otherwise, the new data obtained through rescan is used. .
[0111] In this embodiment, first, the data collection module uses a 3D scanner and an image acquisition device to obtain the 3D point cloud data and high-resolution image data of the target cultural heritage. Then, the 3D modeling module generates a virtual 3D model of the cultural heritage based on these data and performs texture mapping on the model to enhance the detail performance. Next, the historical consistency parameter evaluation module evaluates the deviation between the virtual model and historical facts by analyzing historical documents, archaeological reports, and relic information, and generates historical consistency parameters; while the physical consistency parameter evaluation module evaluates the rationality of the virtual model in the physical environment by analyzing the physical properties of lighting, texture, and materials, and generates physical consistency parameters. The machine learning model module receives the historical and physical consistency parameters, comprehensively evaluates the details of the virtual reconstruction, and generates a virtual reconstruction deviation score. According to this score, the deviation classification module classifies the virtual reconstruction results into three levels: completely consistent, slightly deviated, and severely deviated, and differentiates and processes the areas with different deviation levels. Finally, the dynamic optimization module makes corresponding adjustments to the details of the virtual display according to the deviation classification results, performs fine-tuning of texture, lighting, or geometry for slightly deviated areas, and optimizes severely deviated areas by re-acquiring data to improve the accuracy and historical authenticity of the virtual cultural heritage display.
[0112] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. The artificial intelligence-driven VR cultural heritage digital display system is characterized by: It includes data collection module, 3D modeling module, historical consistency parameter evaluation module, physical consistency parameter evaluation module, machine learning model module, deviation partitioning module and dynamic optimization module; A data collection module, used to obtain three-dimensional point cloud data and high-resolution image data of the target cultural heritage through a three-dimensional scanner and image acquisition equipment; A three-dimensional modeling module, used to generate a virtual three-dimensional model of the cultural heritage based on the three-dimensional point cloud data and the image data, and perform texture mapping on the model; A historical consistency parameter evaluation module, for extracting historical descriptions of cultural heritage based on information from historical documents, archaeological reports and historical relics, and comparing them with the virtual three-dimensional model to generate historical consistency parameters and evaluate the deviation between the virtual model and historical facts; A physical consistency parameter evaluation module is used to determine the physical rationality of the virtual three-dimensional model according to the real lighting, texture and material physical properties of the target cultural heritage, and to generate physical consistency parameters to evaluate the deviation between the virtual model and the actual physical environment; a machine learning model module, configured to receive the historical consistency parameter and the physical consistency parameter, comprehensively evaluate the deviation between the details of the virtual reconstruction and the historical facts, and generate a virtual reconstruction deviation score; A deviation classification module, used to classify the virtual cultural heritage reconstruction results into different deviation levels according to the virtual reconstruction deviation score, including a complete consistency level, a slight deviation level and a severe deviation level; A dynamic optimization module, used to adjust the details of the virtual cultural heritage display according to the deviation division result, including fine-tuning the texture, lighting or geometry of the slightly deviation area; Data reacquisition was performed on severely deviated areas to optimize the accuracy of the virtual representation.
2. The artificial intelligence-driven VR cultural heritage digital display system according to claim 1, characterized in that: In the historical consistency parameter evaluation module, the texture consistency of the virtual 3D model and the actual decoration of the historical relics are analyzed to generate a texture consistency analysis index, wherein the texture consistency analysis index is obtained as follows: For an image I(x,y), set a key point p=(x,y), then the gradient of its local area The calculation formula is: ; In the formula, represents the rate of change of image intensity in the horizontal direction x, Represents the rate of change of image intensity in the vertical direction y; calculates the gradient direction distribution of each region to obtain a 128-dimensional feature vector, matches the descriptors of two different images of virtual model texture and historical relic texture, and calculates the matching degree between the descriptors , the expression is: ;in, and Respectively represent the values of the two descriptors in the i-th dimension. According to the number of matched descriptors and the matching quality, the optimal descriptor matching pair is selected. According to the matching results, the texture consistency index of the two images is calculated. The number of matched key point pairs is set to , the total number of key points in the image is , then the texture consistency analysis index The calculation expression is: .
3. The artificial intelligence-driven VR cultural heritage digital display system according to claim 2 is characterized by: In the physical consistency parameter evaluation module, the illumination distribution difference index is generated by comparing the illumination distribution difference between the virtual model and the real environment. The illumination distribution difference index is obtained as follows: First, the illumination data is collected from the virtual 3D model and the real environment. The illumination intensity of each sampling point is obtained from the virtual 3D model through the rendering algorithm, which is recorded as , where i represents the i-th sampling point; the light intensity of each sampling point is obtained by field measurement or using the illumination data set in the real environment, which is recorded as ; For each sampling point i, the error value is the absolute value of the difference between the virtual model and the real environment illumination intensity, and the calculation expression is: ;in, is the illumination error of the ith point, and are the illumination intensities of the real environment and the virtual model at that point respectively. The illumination errors of all sampling points are summed up to get the overall error , the expression is: ; Where N is the total number of sampling points, calculate the light distribution difference index , the expression is: .
4. The artificial intelligence-driven VR cultural heritage digital display system according to claim 3 is characterized by: In the machine learning model module, the texture consistency analysis index and the illumination distribution difference index are converted into a comprehensive feature vector, and the comprehensive feature vector is used as the input of the machine learning model. The machine learning model predicts the virtual reconstruction deviation score value label for each group of comprehensive feature vectors as the prediction target, and minimizes the sum of prediction errors for all virtual reconstruction deviation score value labels as the training target. The machine learning model is trained until the sum of prediction errors converges, and the model training is stopped. The virtual reconstruction deviation score value is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
5. The artificial intelligence-driven VR cultural heritage digital display system according to claim 4 is characterized by: In the deviation division module, the obtained virtual reconstruction deviation score value is compared with the gradient deviation threshold, the gradient deviation threshold includes a first deviation threshold and a second deviation threshold, and the first deviation threshold is less than the second deviation threshold, and the virtual reconstruction deviation score value is compared with the first deviation threshold and the second deviation threshold respectively; If the virtual reconstruction deviation score value is greater than the second deviation threshold, it means that the degree of virtual reconstruction deviation is high, and it is classified as a severe deviation level; if the virtual reconstruction deviation score value is greater than or equal to the first deviation threshold and less than or equal to the second deviation threshold, it means that the degree of virtual reconstruction deviation is moderate, and it is classified as a slight deviation level; if the virtual reconstruction deviation score value is less than the first deviation threshold, it means that the degree of virtual reconstruction deviation is low, and it is classified as a slight deviation level.
6. The artificial intelligence-driven VR cultural heritage digital display system according to claim 5 is characterized by: In the dynamic optimization module, it is used to adjust the details of the virtual cultural heritage display according to the deviation division result, including fine-tuning the texture, lighting or geometry of the slightly deviation area; When the deviation score value falls within the slight deviation level range, the texture is fine-tuned by adjusting the formula: ; is the adjusted texture value, is the texture value at position (x, y) in the current virtual model, is the target texture value, α is the adjustment factor; Fine-tuning of the lighting includes adjusting the position or intensity of the light source in the virtual scene according to the deviation of the lighting distribution difference index LDE, which is expressed as: ; Adjusted lighting values, is the lighting value in the current virtual model, is the target illumination value; β is the adjustment factor, which controls the amplitude of light fine-tuning; Fine-tuning of the geometry is used to correct the small errors introduced by the modeling algorithm. When the surface of the virtual model does not completely match the geometric features of the historical relics, the adjustment expression is: ; In the formula, is the adjusted geometric position, is the geometric position at position (x, y) in the current virtual model, is the target geometry position, and γ is the geometry adjustment factor, which controls the magnitude of fine-tuning.
7. The artificial intelligence-driven VR cultural heritage digital display system according to claim 6, characterized in that: For areas with severe deviations, new data is obtained by using a higher-precision 3D scanner or a high-resolution image acquisition device. The re-scanned data is expressed as: ; Represents the geometric data obtained by rescanning, is the geometric data in the original virtual model, For newly acquired data, Functions for integrating new data; All important areas of the virtual model are updated with new scans and data acquisition to ensure that the regenerated virtual model has higher accuracy, expressed as: ; Represents the final geometric position of the virtual model. If If the value is exact, the adjusted data is used; otherwise, the new data obtained by rescanning is used .
Citation Information
Cited By
Real-scene interaction safety teaching management method and device based on VR technology, and medium
CN120689180A
Virtual reality interaction system driven by real-time motion capture
CN121232975A