A method and system for constructing a multi-modal semantic data set oriented to cultural heritage entity goals
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]为解决现有线性文化遗产数字化保护中高质量标注样本匮乏的问题,本发明提供一种面向文化遗产实体目标的多模态语义数据集构建方法,通过构建包含多视角渲染图像、高精度几何模型及精细化语义标签的结构化数据集,为训练深度学习模型提供基础数据支撑;基于该数据集训练的模型,可进一步实现大范围线性文化遗产的自动化提取与精准识别
本发明通过融合无人机贴近摄影测量与实景三维建模技术,实现线性文化遗产高分辨率、多视角实景三维数据的高质量获取,并通过点云与网格模型的预处理与统一规范化,有效提升数据质量与一致性;通过基于自适应优化的视点选取与多模态视图渲染,生成具有高覆盖率与高信息量的深度、法线等多模态数据,增强二维样本对三维结构的表达能力;结合类型矢量数据与专家知识,实现空间约束下的三维实体要素高精度标注,并通过建立网格与点云之间的语义映射关系,实现语义信息的自动传递,降低人工标注成本;进一步基于统一视点将三维语义信息准确投影至二维视图,实现多模态数据与标签的严格对齐;最终构建“图像-点云-网格”一体化多模态语义数据集,可显著提升文化遗产场景中深度学习模型的识别精度与泛化能力,并为文化遗产数字化保护与智能解译提供可靠的数据基础与技术支撑。
Smart Images

Figure CN122550870A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine vision technology, and in particular to a method and system for constructing a multimodal semantic dataset for cultural heritage entities. Background Technology
[0002] Cultural heritage is an important carrier of human history and civilization, possessing outstanding historical, cultural, and scientific research value. Large-scale linear cultural heritage sites, such as the Great Wall and the Shu Road, typically exhibit characteristics such as wide distribution, large spatial span, complex morphological structures, and diverse environmental backgrounds. Their condition has been affected by various factors, including natural weathering, human activity, and geological disasters, leading to varying degrees of obscuring, damage, and even disappearance of some cultural heritage sites. How to systematically, meticulously, and intelligently record and model large-scale cultural heritage sites has become a crucial issue urgently needing to be addressed in the field of cultural heritage protection and utilization.
[0003] With the rapid development of technologies such as UAV photogrammetry, 3D laser scanning, and real-scene 3D reconstruction, the construction of high-precision real-scene 3D models using multi-source imagery and point cloud data provides new technical approaches for the digital preservation, virtual display, surveying and analysis, and intelligent identification of cultural heritage. Compared to traditional 2D remote sensing imagery, real-scene 3D data can more realistically express the spatial structure, geometric form, and detailed features of cultural heritage, laying the foundation for subsequent automated identification, semantic understanding, and intelligent analysis. However, the construction of high-quality sample-label datasets for cultural heritage entity elements still faces many challenges.
[0004] Existing cultural heritage datasets are mostly constructed based on 2D remote sensing imagery or single-view imagery, which generally suffers from problems such as limited perspective, severe occlusion, and insufficient representation of spatial structure, making it difficult to fully reflect the true 3D form and complex spatial relationships of cultural heritage. In recent years, multi-view rendering and multimodal data fusion methods based on real-scene 3D models have gradually attracted attention. This involves rendering 3D models from multiple viewpoints to generate sample data containing multimodal information such as original imagery, depth maps, normal maps, and white-film views. Combined with semantic annotation information in the 3D model, automatic registration of samples and labels is achieved, thereby effectively improving the efficiency of dataset construction and annotation accuracy.
[0005] However, existing methods for constructing multimodal real-world 3D datasets still have the following shortcomings: First, the association mechanism between multi-view images and multimodal data is imperfect, making it difficult to ensure high consistency of different modal data in terms of space, scale, and semantics, which easily introduces registration errors; Second, the viewpoint selection strategy in the multi-view rendering process is relatively simple, usually adopting regular distribution or manual experience, which makes it difficult to balance overall coverage and local detail expression, resulting in insufficient sample numbers for complex structural areas, slender site elements, and key components; Third, the existing sample annotation process relies heavily on manual experience, making it difficult to fully combine vector data, 3D geometric features, and expert knowledge to achieve high-precision and systematic annotation of complex cultural heritage entity elements, thereby affecting the integrity and reliability of the dataset. Summary of the Invention
[0006] To address the lack of high-quality labeled samples in the existing digital preservation of linear cultural heritage, this invention provides a method for constructing a multimodal semantic dataset for cultural heritage entities. By constructing a structured dataset containing multi-view rendered images, high-precision geometric models, and refined semantic labels, it provides basic data support for training deep learning models. The model trained based on this dataset can further achieve automated extraction and accurate identification of a large range of linear cultural heritage.
[0007] According to one aspect of the present invention, a method for constructing a multimodal semantic dataset targeting cultural heritage entities is provided, comprising: Acquire dense 3D point clouds, realistic 3D mesh models, and camera exterior orientation parameters of cultural heritage scenes; The dense 3D point cloud and the real-world 3D mesh model are preprocessed respectively; The viewpoint candidate space is initialized based on the preprocessed real-scene 3D mesh model. The viewpoint set is filtered through an adaptive decision optimization method. The real-scene 3D mesh model is then rendered from multiple perspectives using the viewpoint set to generate multimodal view data. Obtain vector data of cultural heritage types and determine the range of labeled tiles in the real-world 3D mesh model based on their spatial extension paths; Within the labeled tile area, known type entity elements in the real-world 3D mesh model are labeled based on the vector data, and unknown type entity elements are supplemented with labels based on expert knowledge to obtain a 3D semantic mesh model. Construct the spatial correspondence between the three-dimensional semantic mesh model and the dense three-dimensional point cloud, and use the spatial proximity search algorithm to transfer semantic labels from the mesh to the point cloud to generate a three-dimensional semantic point cloud. Based on the viewpoint set and camera exterior orientation parameters, the three-dimensional semantic mesh model is subjected to multi-view projection rendering to generate a two-dimensional semantic label image. Based on a unified spatial viewpoint index, multi-view original images, multimodal view data, two-dimensional semantic labeled images, three-dimensional semantic point clouds, and three-dimensional semantic mesh models are structured and organized to construct an integrated multimodal semantic dataset of "image-point cloud-mesh".
[0008] As a further technical solution, the preprocessing of the dense 3D point cloud includes denoising and downsampling; the preprocessing of the real-scene 3D mesh model includes mesh simplification, topology optimization, unified scale normalization, and spatial coordinate alignment.
[0009] As a further technical solution, the step of filtering the viewpoint set through the adaptive decision optimization method includes: Construct a comprehensive reward function Candidate viewpoints An evaluation was conducted, including For viewpoint coverage, The percentage of visible surface area. As a measure of perspective diversity For information entropy gain, The weighting coefficients are used; with the goal of maximizing the comprehensive reward function, the optimal viewpoint sequence is obtained through iterative screening.
[0010] As a further technical solution, the dense three-dimensional point cloud is subjected to noise reduction processing, including: A combination of statistical filtering and radius filtering is used, where statistical filtering is performed by calculating points. K-nearest neighbor average distance global mean with standard deviation In order to satisfy Outliers are removed periodically. Threshold coefficients; radius filtering in the search radius The number of inner neighbor points is less than the threshold Noise points are removed in real time.
[0011] As a further technical solution, the unified scale normalization and spatial coordinate alignment include: calculating the geometric center of the bounding box of the model and the maximum side length scale factor, for any point Normalization is performed, and then the model is mapped to a unified spatial reference frame through a rigid transformation.
[0012] As a further technical solution, the spatial proximity search algorithm employs K-nearest neighbor search or the KD-tree method for any point in a dense 3D point cloud. In a 3D semantic mesh model, the geometric nearest neighbor face or vertex is retrieved and the corresponding semantic label is assigned to that point.
[0013] According to one aspect of the present invention, an apparatus for constructing a multimodal semantic dataset for cultural heritage entity targets is provided, comprising: The data acquisition module is used to acquire dense 3D point clouds, real-scene 3D mesh models, and camera exterior orientation parameters of cultural heritage scenes. The preprocessing module is used to preprocess the dense 3D point cloud and the real-scene 3D mesh model respectively; The viewpoint filtering and rendering module is used to initialize the viewpoint candidate space based on the preprocessed real-scene 3D mesh model, filter the viewpoint set through an adaptive decision optimization method, and use the viewpoint set to perform multi-view rendering of the real-scene 3D mesh model to generate multimodal view data. The vector acquisition and range determination module is used to acquire vector data of cultural heritage types and determine the range of labeled tiles in the real-world 3D mesh model based on their spatial extension path. The 3D annotation module is used to annotate known type entity elements in the real-world 3D mesh model based on the vector data within the annotation tile area, and to supplement the annotation of unknown type entity elements by combining expert knowledge, so as to obtain a 3D semantic mesh model. The semantic mapping module is used to construct the spatial correspondence between the three-dimensional semantic mesh model and the dense three-dimensional point cloud, and to use a spatial proximity search algorithm to transfer semantic labels from the mesh to the point cloud to generate a three-dimensional semantic point cloud. The projection rendering module is used to perform multi-view projection rendering on the three-dimensional semantic mesh model based on the viewpoint set and camera exterior orientation parameters to generate a two-dimensional semantic label image. The dataset construction module is used to structure and organize multi-view original images, multimodal view data, two-dimensional semantic labeled images, three-dimensional semantic point clouds and three-dimensional semantic mesh models based on a unified spatial viewpoint index, and construct an integrated "image-point cloud-mesh" multimodal semantic dataset.
[0014] As a further technical solution, the preprocessing module performs denoising and downsampling on the dense 3D point cloud, and performs mesh simplification, topology optimization, unified scale normalization and spatial coordinate alignment on the real-world 3D mesh model.
[0015] As a further technical solution, the viewpoint selection and rendering module constructs a comprehensive reward function. Candidate viewpoints An evaluation is conducted, and the optimal viewpoint sequence is obtained through iterative selection with the goal of maximizing the comprehensive reward function.
[0016] According to one aspect of the present invention, a computer-readable storage medium is provided that stores an integrated multimodal semantic dataset of "image-point cloud-mesh" constructed according to the method.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention achieves high-quality acquisition of high-resolution, multi-view real-scene 3D data of linear cultural heritage by integrating UAV close-up photogrammetry and real-scene 3D modeling technology. Through preprocessing and standardization of point clouds and mesh models, data quality and consistency are effectively improved. Adaptive optimization-based viewpoint selection and multimodal view rendering generate multimodal data with high coverage and information content, including depth and normals, enhancing the expressive power of 2D samples for 3D structures. Combining type vector data and expert knowledge, high-precision annotation of 3D entity elements under spatial constraints is achieved. By establishing a semantic mapping relationship between the mesh and point cloud, automatic transfer of semantic information is realized, reducing manual annotation costs. Furthermore, based on a unified viewpoint, 3D semantic information is accurately projected onto the 2D view, achieving strict alignment between multimodal data and labels. Finally, an integrated "image-point cloud-mesh" multimodal semantic dataset is constructed, which can significantly improve the recognition accuracy and generalization ability of deep learning models in cultural heritage scenes, and provide a reliable data foundation and technical support for the digital protection and intelligent interpretation of cultural heritage. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a method for constructing a multimodal semantic dataset for cultural heritage entity targets, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the annotation of entity elements in a real-world 3D model in an embodiment of the present invention. Detailed Implementation
[0020] Existing methods for constructing multimodal real-scene 3D datasets suffer from the following shortcomings: the correlation mechanism between multi-view images and multimodal data is imperfect, making it difficult to guarantee high consistency in spatial, scale, and semantic dimensions; the viewpoint selection strategy in the multi-view rendering process is relatively simple, making it difficult to balance overall coverage with local detail representation, resulting in insufficient samples for complex structural areas; the sample annotation process relies heavily on human experience, lacking effective integration of vector data, 3D geometric features, and expert knowledge, affecting the completeness and reliability of the dataset. Therefore, there is an urgent need to propose a method for constructing a multimodal semantic dataset oriented towards cultural heritage entities. By integrating technologies such as UAV close-up photogrammetry, real-scene 3D modeling, multimodal rendering, spatial semantic indexing, and 3D semantic annotation, this method can achieve high-precision alignment and deep fusion of multi-source data, constructing a comprehensive, accurately labeled, and structurally complete multimodal real-scene 3D sample-label dataset. This will provide a reliable data foundation and technical support for the automatic identification, refined modeling, dynamic monitoring, and digital protection of cultural heritage.
[0021] The following description, in conjunction with the embodiments and accompanying drawings, will be clear and complete. Obviously, the described embodiments are only a portion, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. Furthermore, the technical features of the various embodiments or individual embodiments provided by this invention can be arbitrarily combined to form new technical solutions. Such combinations are not bound by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0022] Before formally describing the present invention, a general description of the solution of the present invention will be given first to facilitate understanding.
[0023] Please refer to Figure 1 This invention provides a method for constructing a multimodal semantic dataset targeting cultural heritage entities, comprising the following steps:
[0024] S1. Acquisition of real-scene 3D data: Using UAV close-up photography technology, the target linear cultural heritage site is photographed from multiple angles and heights to obtain close-up image data with complete coverage and high resolution. Based on the image data, close-up photogrammetry is used to reconstruct multimodal real-scene 3D basic data such as dense 3D point cloud, real-scene 3D mesh model and camera exterior orientation parameters of the cultural heritage scene.
[0025] In this embodiment, S1 specifically includes:
[0026] S11. Select a multi-UAV platform equipped with a high-resolution camera to conduct low-altitude close-range flight over the target linear cultural heritage area. Collect high-resolution close-range image data through multi-route, multi-angle, and multi-altitude circling aerial photography to ensure complete coverage of the main structure of the cultural heritage and its surrounding environment.
[0027] S12. Perform image quality screening, distortion correction and exposure equalization on the acquired multi-view image data to improve the consistency and geometric accuracy of the image data.
[0028] S13. Based on the processed multi-view images, a photogrammetric method combining aerial triangulation and dense matching of multi-view images is adopted to calculate the camera interior and exterior orientation parameters of the images, and to reconstruct and generate a dense 3D point cloud and a real-scene 3D mesh model of the cultural heritage scene.
[0029] S2. Real-world 3D data preprocessing and enhancement: The dense 3D point cloud data is subjected to denoising filtering and downsampling filtering to eliminate noise points and reduce data redundancy; the real-world 3D model is subjected to mesh simplification and topology optimization to reduce model complexity and improve geometric representation quality; on this basis, the real-world 3D model is subjected to unified scale normalization and spatial coordinate alignment to provide a standardized data foundation for subsequent multimodal rendering and semantic annotation.
[0030] In this embodiment, S2 specifically includes:
[0031] S21. A combination of statistical filtering and radius filtering is used to denoise dense 3D point clouds, remove outliers and mismatched points, and improve the quality of point cloud data.
[0032] Let the input dense 3D point cloud be:
[0033]
[0034] For any point First, calculate its K-nearest neighbor set. And calculate the average distance between this point and its K nearest neighbors:
[0035]
[0036] Further calculate the global mean of the average neighborhood distance of all points. with standard deviation :
[0037]
[0038]
[0039] Set threshold coefficient When the following conditions are met:
[0040]
[0041] At that time, the corresponding point These were identified as outliers and removed.
[0042] After statistical filtering, radius filtering is further used for additional noise reduction. The search radius is set to... For any point If its neighborhood satisfies:
[0043]
[0044] in, ,and If the minimum number of neighboring points is the threshold, then the point is identified as a noise point and removed.
[0045] By combining statistical filtering and radius filtering, isolated points, mismatched points, and scanning noise points are effectively removed, improving the structural integrity and spatial consistency of the point cloud.
[0046] S22. Use the voxel mesh downsampling algorithm to downsample the point cloud, reduce data redundancy while maintaining the integrity of the geometric structure, and improve data processing efficiency.
[0047] S23. Employ an edge-folding-based mesh simplification algorithm and a local topology reconstruction strategy for the real-world 3D model to reduce the number of model faces and improve the quality of the model's geometric representation.
[0048] S24. Perform unified scale normalization and spatial coordinate alignment on the processed 3D point cloud and the real-world 3D model to place all types of data under a unified spatial reference framework, providing a standardized data foundation for subsequent multimodal rendering and 3D semantic annotation.
[0049] Specifically, let the processed 3D point cloud or model vertex set be:
[0050]
[0051] First, calculate the spatial bounding box of the model and determine its geometric center. and the maximum side length scale factor .
[0052] For any point Perform scale normalization:
[0053]
[0054] in, The coordinates of the model center are This is the maximum side length of the model's bounding box, used to eliminate scale differences between different data and map the model to a uniform scale space.
[0055] After scale normalization, spatial coordinate alignment is further performed. A rigid transformation maps the point cloud or model to a unified spatial reference frame.
[0056]
[0057] in, It is a three-dimensional rotation matrix. It is a translation vector.
[0058] Through the above-mentioned scale normalization and spatial alignment processing, we can achieve a standardized representation of multimodal 3D data from different sources under a unified coordinate system, providing a consistent spatial reference for subsequent multimodal rendering and 3D semantic annotation.
[0059] S3. Multi-viewpoint selection strategy and model view rendering: Initialize the viewpoint candidate space based on the preprocessed real-world 3D model, and filter and sort the candidate viewpoints using an adaptive decision optimization method to select a set of viewpoints with high surface coverage and information gain; use the viewpoint set to perform multi-view rendering of the real-world 3D model to generate multi-modal view data such as depth maps, white film views, and normal maps corresponding to the original images, realizing the mapping of 3D structural information to 2D multi-modal sample data space, and providing a unified viewpoint basis for subsequent 2D semantic projection and cross-modal data alignment.
[0060] In this embodiment, S3 specifically includes
[0061] S31. Construct a 3D spherical or hemispherical viewpoint candidate space outside the bounding box of the real-world 3D model, and generate an initial set of candidate viewpoints at certain angular intervals.
[0062] S32. Construct a viewpoint selection strategy model based on adaptive decision optimization. Use viewpoint coverage, visible surface area ratio, viewpoint diversity and information entropy gain as a comprehensive reward function to perform adaptive decision optimization on candidate viewpoints and select the optimal viewpoint sequence.
[0063] Specifically, let the candidate viewpoint be... The comprehensive reward function is constructed as follows:
[0064]
[0065] in:
[0066] Indicate viewpoint Coverage of the model surface;
[0067] This represents the proportion of visible surface area corresponding to the viewpoint.
[0068] This measures the diversity of perspectives between the current viewpoint and the selected viewpoint.
[0069] This represents the information entropy gain resulting from the viewpoint.
[0070] , which is a weighting coefficient used to adjust the influence of each evaluation indicator on the overall reward.
[0071] The adaptive decision optimization strategy model uses the aforementioned comprehensive reward function As an optimization objective, the candidate viewpoint sequences are iteratively updated through decision-making, and finally the optimal viewpoint sequence with high structural coverage and information expression capabilities is obtained.
[0072] S33. Utilize the filtered viewpoint set to perform multi-view rendering of the real-world 3D model, generating multimodal view data such as depth maps, white film views, and normal maps corresponding to the original images, thereby realizing the mapping of 3D structural information to 2D multimodal sample data space.
[0073] S4. Acquisition and Labeling Range Determination of Cultural Heritage Type Vector Data: Acquire vector data of linear cultural heritage types such as the Great Wall. Based on the spatial extension path of the linear heritage, determine the corresponding labeling tile range in the real-world 3D model and construct a spatial constraint area for labeling 3D entity elements to avoid interference from irrelevant areas and improve labeling efficiency and accuracy.
[0074] In this embodiment, S4 specifically includes:
[0075] S41. Based on publicly available surveying and mapping data, historical documents, existing geographic information databases, and online data, vector data of cultural heritage types such as the Great Wall are obtained. In this embodiment, the Great Wall vector data comes from the "China Great Wall Heritage Website". The Great Wall vector data is subjected to unified coordinate transformation and accuracy correction.
[0076] S42. Project the vector data into the coordinate system of the real-world 3D model, and automatically generate corresponding labeled tile areas on the model surface according to the spatial extension path of the linear heritage, thus constructing a 3D labeled spatial constraint range.
[0077] S5. Based on vector constraints and expert knowledge, 3D entity element annotation is performed to annotate known types of entity elements in the real-world 3D model, and unknown types of cultural heritage entities not included in the vector data are manually confirmed and supplemented.
[0078] Please refer to Figure 2 , Figure 2This is a schematic diagram illustrating the annotation of entity elements in a real-world 3D model according to an embodiment of the present invention. The left image shows the overview page of the OSGB model, where the yellow boxes indicate the annotated areas. The right image is a tile view, which is the interface for annotating within the selected tile's OBJ model, where green and blue represent solid color annotation information.
[0079] In this embodiment, S5 specifically includes:
[0080] S51. The annotation tool used is the DaShiMoFang software. Based on the vector data of cultural heritage types, texture editing is performed within the annotation tile area to achieve manual three-dimensional annotation of known entity elements.
[0081] S52. Combining the knowledge of experts in the field of cultural heritage, the annotation results are manually reviewed and corrected, and supplementary annotations are made for unknown types of cultural heritage entities that are not included in the vector data, forming a complete and reliable three-dimensional semantic annotation model.
[0082] S6. Based on spatial semantic indexing, 3D point cloud semantic mapping is constructed to establish a spatial correspondence between the labeled 3D mesh model and the original dense point cloud, realizing the automatic transfer of semantic information from the mesh to the point cloud.
[0083] In this embodiment, S6 specifically includes:
[0084] S61. Based on the semantically annotated real-world 3D mesh model, construct a spatial index structure between it and the original dense point cloud.
[0085] S62. Using spatial proximity search algorithms, including but not limited to K-nearest neighbor search or KD-tree methods, for any point in the original dense point cloud... In the semantic grid model, it retrieves the nearest neighbor face or vertex;
[0086] S63. Assign semantic labels to the entity elements corresponding to the nearest neighbor face or vertex. This enables the mapping of semantic information from a continuous grid surface to a discrete point cloud space, generating three-dimensional semantic point cloud data with semantic attributes.
[0087] S7. Based on the set of viewpoints selected in S3 and the corresponding camera exterior orientation parameters, perform multi-view projection rendering on the real-world 3D model with completed semantic annotation, map the 3D semantic information onto the 2D view plane, and generate a 2D semantic label image.
[0088] In this embodiment, S7 specifically includes:
[0089] S71. Load the semantically annotated real-world 3D model into the 3D rendering engine, and construct a multi-view virtual observation scene based on the viewpoint set and camera exterior orientation parameters determined in step S3.
[0090] S72. Generate a two-dimensional image from the corresponding viewpoint through multi-view rendering, in which different categories of entity elements are expressed in different colors through preset semantic encoding;
[0091] S73. Pixel-level extraction is performed on the rendering results to generate a two-dimensional semantic label image that strictly corresponds to the multimodal views such as the depth map and normal map, thereby realizing the accurate projection of three-dimensional semantic information into the two-dimensional image space.
[0092] S8. The construction of an integrated multimodal semantic dataset of "image-point cloud-mesh" is based on a unified spatial viewpoint index to organize multi-source data in a structured manner.
[0093] In this embodiment, S8 specifically includes:
[0094] S81. Based on the unified viewpoint number, the original multi-view images, depth maps, normal maps and white film views are matched one by one with the corresponding two-dimensional semantic label images to construct two-dimensional multimodal sample data.
[0095] S82. Organize the semantically mapped 3D semantic point cloud data and 3D semantic mesh model in a unified manner, and establish a correlation with the 2D multimodal sample data;
[0096] S83. Construct a unified data organization structure to achieve collaborative storage and index management of multimodal data such as "image-point cloud-mesh", so that the dataset can directly support cross-modal deep learning training and 3D scene understanding tasks.
[0097] Based on the same inventive concept as the aforementioned method embodiments, this invention provides a device for constructing a multimodal semantic dataset for cultural heritage entities. This device addresses problems in existing multimodal real-scene 3D dataset construction, such as low alignment accuracy of multi-source data, a single viewpoint selection strategy, and reliance on manual experience for annotation. The device employs a data acquisition module, a preprocessing module, a viewpoint selection and rendering module, a vector acquisition and range determination module, a 3D annotation module, a semantic mapping module, a projection rendering module, and a dataset construction module. By integrating UAV close-up photogrammetry, real-scene 3D modeling, multimodal rendering, spatial semantic indexing, and 3D semantic annotation technologies, it achieves high-precision alignment and deep fusion of multi-source data, constructing an integrated "image-point cloud-mesh" multimodal semantic dataset. This provides a reliable data foundation and technical support for the automatic identification, refined modeling, dynamic monitoring, and digital protection of cultural heritage.
[0098] It should be noted that the device embodiments provided by the present invention, in addition to implementing the methods in the above method embodiments, are also used to implement the methods in other method embodiments provided by the present invention. The difference lies only in the setting of corresponding functional modules, and their principles are basically the same as those of the above device embodiments provided by the present invention. Based on the above device embodiments, those skilled in the art, referring to the specific technical solutions in other method embodiments, can obtain corresponding technical means and technical solutions constituted by these technical means by combining technical features. Under the premise of ensuring the practicality of the technical solutions, the modules in the above device embodiments can be improved to obtain corresponding device-type embodiments for implementing the methods in other method-type embodiments.
[0099] Based on the same inventive concept as any of the foregoing embodiments, this embodiment of the invention also provides a non-transitory computer-readable storage medium that stores computer instructions that cause the computer to execute the method for constructing a multimodal semantic dataset oriented towards cultural heritage entity targets.
[0100] Specifically, the storage medium can be any form of non-volatile memory, including but not limited to: hard disks, solid-state drives (SSDs), optical discs (CDs / DVDs), read-only memory (ROMs), programmable read-only memory (PROMs), erasable programmable read-only memory (EPROMs), electrically erasable programmable read-only memory (EEPROMs), flash memory, magnetic storage devices, etc. The computer instructions stored on the storage medium include computer program code for implementing the method.
[0101] When a computer (including servers, personal computers, mobile devices, embedded systems, etc.) reads and executes the computer instructions from this storage medium, it can perform the following functions: acquire dense 3D point clouds, realistic 3D mesh models, and camera exterior orientation parameters of cultural heritage scenes; preprocess the dense 3D point clouds and realistic 3D mesh models respectively; filter viewpoint sets and generate multimodal view data through an adaptive decision optimization method; complete 3D semantic annotation based on vector data and expert knowledge; map semantic labels from the mesh to the point cloud to generate a 3D semantic point cloud; project the 3D semantic information onto the 2D view to generate semantically labeled images; and finally construct an integrated "image-point cloud-mesh" multimodal semantic dataset. This non-transitory computer-readable storage medium enables computers to efficiently construct multimodal semantic datasets without relying on a large amount of manual annotation, significantly improving the annotation accuracy and cross-modal consistency of the dataset, and has good engineering practicality and generalization ability.
[0102] In summary, this invention discloses a method for constructing a multimodal semantic dataset for cultural heritage entities, relating to the field of machine vision technology. The method includes: acquiring original images from multiple perspectives; generating dense point clouds and a real-world 3D mesh model using photogrammetry and preprocessing them; sampling viewpoints using an adaptive decision optimization strategy to generate multimodal 2D data; semantically annotating the mesh model by combining vector data and expert knowledge; constructing a spatial semantic index; transferring the semantic labels of the mesh model to the point cloud through spatial proximity mapping to generate a point cloud with semantic attributes; and generating labeled images based on viewpoint-rendered annotation models. Finally, the method integrates multi-view 2D images, 2D semantic labels, 3D semantic meshes, and 3D semantic point clouds to construct an integrated image-point cloud-mesh multimodal semantic dataset, achieving spatial consistency and semantic alignment among multimodal data. This invention provides a streamlined reference for constructing an integrated image-point cloud-mesh multimodal semantic dataset for cultural heritage, achieving semantic alignment among multimodal data and providing key technical support for accurate identification of cultural heritage target types, 3D scene understanding, and digital protection.
[0103] The terms “comprising” and “having”, and any variations thereof, in the specification, claims, and accompanying drawings of this invention are intended to cover a non-exclusive inclusion, such as a process, method, system, product, or apparatus that includes a series of steps or units, not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a multimodal semantic dataset for cultural heritage entity targets, characterized in that, include: Acquire dense 3D point clouds, realistic 3D mesh models, and camera exterior orientation parameters of cultural heritage scenes; The dense 3D point cloud and the real-world 3D mesh model are preprocessed respectively; The viewpoint candidate space is initialized based on the preprocessed real-scene 3D mesh model. The viewpoint set is filtered through an adaptive decision optimization method. The real-scene 3D mesh model is then rendered from multiple perspectives using the viewpoint set to generate multimodal view data. Obtain vector data of cultural heritage types and determine the range of labeled tiles in the real-world 3D mesh model based on their spatial extension paths; Within the labeled tile area, known type entity elements in the real-world 3D mesh model are labeled based on the vector data, and unknown type entity elements are supplemented with labels based on expert knowledge to obtain a 3D semantic mesh model. Construct the spatial correspondence between the three-dimensional semantic mesh model and the dense three-dimensional point cloud, and use the spatial proximity search algorithm to transfer semantic labels from the mesh to the point cloud to generate a three-dimensional semantic point cloud. Based on the viewpoint set and camera exterior orientation parameters, the three-dimensional semantic mesh model is subjected to multi-view projection rendering to generate a two-dimensional semantic label image. Based on a unified spatial viewpoint index, multi-view original images, multimodal view data, two-dimensional semantic labeled images, three-dimensional semantic point clouds, and three-dimensional semantic mesh models are structured and organized to construct an integrated multimodal semantic dataset of "image-point cloud-mesh".
2. The method for constructing a multimodal semantic dataset for cultural heritage entity targets according to claim 1, characterized in that, Preprocessing of the dense 3D point cloud includes denoising and downsampling. The preprocessing of the real-world 3D mesh model includes mesh simplification, topology optimization, unified scale normalization, and spatial coordinate alignment.
3. The method for constructing a multimodal semantic dataset for cultural heritage entity targets according to claim 1, characterized in that, The process of selecting the viewpoint set using an adaptive decision optimization method includes: Construct a comprehensive reward function Candidate viewpoints An evaluation was conducted, including For viewpoint coverage, The percentage of visible surface area. As a measure of perspective diversity For information entropy gain, The weighting coefficients are used; with the goal of maximizing the comprehensive reward function, the optimal viewpoint sequence is obtained through iterative screening.
4. The method for constructing a multimodal semantic dataset for cultural heritage entity targets according to claim 2, characterized in that, The denoising process for the dense 3D point cloud includes: A combination of statistical filtering and radius filtering is used, where statistical filtering is performed by calculating points. K-nearest neighbor average distance global mean with standard deviation In order to satisfy Outliers are removed periodically. Threshold coefficients; radius filtering in the search radius The number of inner neighbor points is less than the threshold Noise points are removed in real time.
5. The method for constructing a multimodal semantic dataset for cultural heritage entity targets according to claim 2, characterized in that, The unified scale normalization and spatial coordinate alignment include: calculating the geometric center of the model bounding box and the maximum side length scale factor, for any point... Normalization is performed, and then the model is mapped to a unified spatial reference frame through a rigid transformation.
6. The method for constructing a multimodal semantic dataset for cultural heritage entity targets according to claim 1, characterized in that, The spatial proximity search algorithm employs K-nearest neighbor search or KD-tree method for any point in a dense 3D point cloud. In a 3D semantic mesh model, the geometric nearest neighbor face or vertex is retrieved and the corresponding semantic label is assigned to that point.
7. A device for constructing a multimodal semantic dataset for cultural heritage entity targets, characterized in that, include: The data acquisition module is used to acquire dense 3D point clouds, real-scene 3D mesh models, and camera exterior orientation parameters of cultural heritage scenes. The preprocessing module is used to preprocess the dense 3D point cloud and the real-scene 3D mesh model respectively; The viewpoint filtering and rendering module is used to initialize the viewpoint candidate space based on the preprocessed real-scene 3D mesh model, filter the viewpoint set through an adaptive decision optimization method, and use the viewpoint set to perform multi-view rendering of the real-scene 3D mesh model to generate multimodal view data. The vector acquisition and range determination module is used to acquire vector data of cultural heritage types and determine the range of labeled tiles in the real-world 3D mesh model based on their spatial extension path. The 3D annotation module is used to annotate known type entity elements in the real-world 3D mesh model based on the vector data within the annotation tile area, and to supplement the annotation of unknown type entity elements by combining expert knowledge, so as to obtain a 3D semantic mesh model. The semantic mapping module is used to construct the spatial correspondence between the three-dimensional semantic mesh model and the dense three-dimensional point cloud, and to use a spatial proximity search algorithm to transfer semantic labels from the mesh to the point cloud to generate a three-dimensional semantic point cloud. The projection rendering module is used to perform multi-view projection rendering on the three-dimensional semantic mesh model based on the viewpoint set and camera exterior orientation parameters to generate a two-dimensional semantic label image. The dataset construction module is used to structure and organize multi-view original images, multimodal view data, two-dimensional semantic labeled images, three-dimensional semantic point clouds and three-dimensional semantic mesh models based on a unified spatial viewpoint index, and construct an integrated "image-point cloud-mesh" multimodal semantic dataset.
8. The apparatus for constructing a multimodal semantic dataset for cultural heritage entity targets according to claim 7, characterized in that, The preprocessing module performs denoising and downsampling on the dense 3D point cloud, and mesh simplification, topology optimization, unified scale normalization, and spatial coordinate alignment on the real-world 3D mesh model.
9. The apparatus for constructing a multimodal semantic dataset for cultural heritage entity targets according to claim 7, characterized in that, The viewpoint selection and rendering module constructs a comprehensive reward function. Candidate viewpoints An evaluation is conducted, and the optimal viewpoint sequence is obtained through iterative selection with the goal of maximizing the comprehensive reward function.
10. A computer-readable storage medium storing an integrated multimodal semantic dataset of "image-point cloud-mesh" constructed according to any one of claims 1 to 6.