Intelligent surveying and mapping method and system based on AI and BIM fusion
By synchronously collecting point cloud and image data through the drone platform and combining cross-modal registration and semantic fusion, a BIM model of the ancient building with both geometric form and semantic attributes was constructed, which solved the problem of insufficient fusion of multi-source data and achieved high-precision digital management and protection of ancient buildings.
Patent Information
- Application Number
- CN202511360915.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-23
AI Technical Summary
In existing technologies, the independent collection of point clouds, images, and positioning data results in high spatial registration errors, making it impossible to form a unified semantic space. The multi-source data fusion capability is insufficient, making it difficult to meet the needs of high-precision digital protection and long-term inheritance of ancient architectural cultural relics.
Multi-angle scanning is performed through an unmanned aerial vehicle platform equipped with a lidar, RGB camera, and positioning system, and point cloud data and multi-view image data are collected simultaneously. Combined with cross-modal registration and semantic fusion, a semantically enhanced BIM model with both geometric form and semantic attributes is constructed, and integrated with the GIS base map to achieve deep fusion of multi-source data.
It breaks through the information limitations of single data, builds a more complete digital model of ancient buildings, realizes full-scale management from micro components to macro spaces, provides accurate digital carriers, and provides high-precision data support for the protection and research of ancient buildings.
Smart Images

Figure CN120852601A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of cultural relic protection and digital technology, and involves, but is not limited to, an intelligent surveying and mapping method and system based on the integration of AI and BIM. Background Art
[0002] With the continuous advancement of technology, the protection of cultural relics and ancient buildings has ushered in new opportunities and challenges. Traditional surveying techniques often suffer from inefficiency, insufficient accuracy, and difficulty in comprehensively acquiring information when dealing with complex and intricate ancient architectural relics, failing to meet the urgent needs of modern society for high-precision digital protection and long-term preservation of cultural relics and ancient buildings. Against this backdrop, the rapid development of cutting-edge technologies such as drone technology, artificial intelligence (AI), and building information modeling (BIM) has brought innovative solutions to the field of surveying ancient architectural relics.
[0003] With its advantages of high flexibility, low cost, and ability to quickly acquire high-resolution image data over large areas, drone technology has shown great potential in collecting data on the external morphology of cultural relics and ancient buildings. Equipped with advanced sensors such as LiDAR, high-definition RGB cameras, and positioning systems, drones can comprehensively scan ancient buildings and their surrounding environment from multiple angles, acquiring high-precision point cloud data, multi-view image data, and position and attitude (POS) information. This provides a rich and accurate data source for subsequent 3D modeling and analysis.
[0004] The rise of artificial intelligence (AI) technology has provided powerful tools for the efficient processing and in-depth mining of massive amounts of surveying and mapping data. Building Information Modeling (BIM), as a digital building information integration technology, can integrate and manage multi-dimensional data such as the geometric form, physical attributes, and historical information of ancient buildings and cultural relics. Based on the internationally recognized IFC standard, BIM models can not only intuitively display the three-dimensional structure of ancient buildings, but also provide standardized data support for the protection and restoration, structural analysis, and digital display of cultural relics, promoting collaboration among various professional fields and enhancing the scientific and systematic nature of cultural relic protection work.
[0005] However, point cloud, image, and localization data are collected independently, resulting in high spatial registration errors. This leads to misalignment of geometric and texture information, making it impossible to form a unified semantic space and resulting in insufficient multi-source data fusion capabilities. Summary of the Invention
[0006] In view of this, the embodiments of this application provide an intelligent surveying and mapping method based on the integration of AI and BIM, which at least solves the problem of insufficient multi-source data fusion capability.
[0007] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide an intelligent surveying method based on the fusion of AI and BIM, the method comprising: Using a drone platform equipped with lidar, RGB camera and positioning system, the ancient building and its surrounding environment are scanned from multiple angles during the first operation. Point cloud data, multi-view image data and position and attitude data are collected simultaneously. The point cloud data and multi-view image data are then associated with the position and attitude data at the corresponding time through their respective timestamps. The point cloud data and the multi-view image data are preprocessed respectively. Based on the position and pose data, cross-modal registration of the preprocessed point cloud data and multi-view image data is performed by feature matching and pose estimation to obtain a point cloud-image registration dataset. Semantic segmentation is performed on the point cloud data and image data in the registration dataset to obtain a first semantic segmentation result and a second semantic segmentation result. Based on the association relationship in the registration dataset, cross-modal semantic fusion is performed on the first semantic segmentation result and the second semantic segmentation result to obtain a first cross-modal semantic segmentation result. The first cross-modal semantic segmentation result includes multiple original point cloud components with category labels, and each original point cloud component is associated with a texture region in the corresponding image data through registration. Based on the category labels of each original point cloud component in the first cross-modal semantic segmentation result, the multiple original point cloud components are classified and aggregated to obtain segmented point cloud components, and the topological relationship between each segmented point cloud component is constructed. The assembly relationship between each segmented point cloud component is inferred through the topological relationship. Based on the assembly relationship, the template corresponding to the category of each segmented point cloud component is called from the pre-constructed BIM template library of ancient building and cultural relic components, and BIM model components are instantiated. The first cross-modal semantic segmentation result is structurally attached to the BIM model components to construct a semantically enhanced BIM model that combines geometric morphological information and semantic attribute information. The semantically enhanced BIM model and the geographic information system (GIS) base map containing geospatial benchmarks are integrated to obtain the BIM-GIS fusion model corresponding to the first operation time. The BIM-GIS fusion model is used for spatial query of components of the ancient building relics, comparison display of component status in multiple periods, and three-dimensional visualization.
[0008] Secondly, embodiments of this application provide an intelligent surveying and mapping system based on the fusion of AI and BIM, the system comprising: The data acquisition module is used to perform multi-angle flight scanning of the ancient building and its surrounding environment at the first moment of operation using a drone platform equipped with lidar, RGB camera and positioning system. Simultaneously, it collects point cloud data, multi-view image data and position and attitude data, and establishes a correlation between the point cloud data and the multi-view image data and the position and attitude data at the corresponding moment through their respective timestamps. The registration module is used to preprocess the point cloud data and the multi-view image data respectively, and based on the position and pose data, to perform cross-modal registration of the preprocessed point cloud data and multi-view image data through feature matching and pose estimation to obtain a point cloud-image registration dataset. The semantic segmentation module is used to perform semantic segmentation on the point cloud data and image data in the registration dataset to obtain a first semantic segmentation result and a second semantic segmentation result. Based on the association relationship in the registration dataset, the first semantic segmentation result and the second semantic segmentation result are subjected to cross-modal semantic fusion to obtain a first cross-modal semantic segmentation result. The first cross-modal semantic segmentation result includes multiple original point cloud components with category labels, and each original point cloud component is associated with a texture region in the corresponding image data through registration. The first generation module is used to classify and aggregate the multiple original point cloud components based on the category labels of each original point cloud component in the first cross-modal semantic segmentation result, obtain segmented point cloud components, construct the topological relationship between each segmented point cloud component, infer the assembly relationship between each segmented point cloud component through the topological relationship, and, based on the assembly relationship, call the template corresponding to the category of each segmented point cloud component from the pre-constructed building information model (BIM) template library of ancient building and cultural relic components, instantiate and generate BIM model components, and structurally attach the first cross-modal semantic segmentation result to the BIM model components to construct a semantically enhanced BIM model that combines geometric morphological information and semantic attribute information. The first integration module is used to integrate the semantically enhanced BIM model and the geographic information system (GIS) base map containing the geospatial benchmark to obtain the BIM-GIS fusion model corresponding to the first operation time. The BIM-GIS fusion model is used to perform spatial query of the components of the ancient building relics, comparative display of the component status in multiple periods, and three-dimensional visualization.
[0009] The beneficial effects of the technical solutions provided in this application include at least the following: By simultaneously collecting point cloud, image, and position and attitude data using drones, and combining cross-modal registration and semantic fusion, the information limitations of single data (point cloud / image) are overcome, while preserving geometric morphology and semantic attributes, to construct a more complete digital model of ancient buildings, achieving deep fusion of multi-source data. Through classification aggregation, topological relationship reasoning, and BIM template instantiation, the original surveying data is transformed into a structured semantically enhanced BIM model, which includes not only the geometric details of components but also attributes such as materials and categories, providing a precise digital carrier for the protection and research of ancient buildings. Through the integration of BIM and GIS base maps, spatial query, multi-period comparison, and 3D visualization of ancient building components are realized, meeting the full-scale management needs from micro-components to macro-space. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 A flowchart illustrating an intelligent surveying and mapping method based on the fusion of AI and BIM, provided for an embodiment of this application; Figure 2 This is a schematic diagram of the components of an intelligent surveying and mapping system based on the fusion of AI and BIM, provided as an embodiment of this application. DETAILED DESCRIPTION
[0011] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0012] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0013] It should be noted that the terms "first, second, and third" used in the embodiments of this application are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0014] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this application pertain. It should also be understood that terms such as those defined in general dictionaries should be understood to have a meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0015] This application provides an intelligent surveying and mapping method based on the fusion of AI and BIM, applied to electronic devices. These electronic devices include, but are not limited to, mobile phones, laptops, tablets, handheld internet devices, multimedia devices, streaming media devices, mobile internet devices, wearable devices, or other types of electronic devices. The functions implemented by this method can be achieved by a processor in the electronic device calling program code. The program code can be stored in a computer storage medium; therefore, the electronic device includes at least a processor and a storage medium. The processor can be used to process the intelligent surveying and mapping process based on the fusion of AI and BIM, and the memory can be used to store the data required and generated during the intelligent surveying and mapping process based on the fusion of AI and BIM.
[0016] Figure 1 A flowchart illustrating an intelligent surveying method based on the fusion of AI and BIM, provided for an embodiment of this application, is shown below. Figure 1 As shown, the method includes at least the following steps: Step S110: Using a drone platform equipped with a lidar, RGB camera and positioning system, the ancient building and its surrounding environment are scanned from multiple angles during the first operation. Point cloud data, multi-view image data and position and attitude data are collected simultaneously. The point cloud data and the multi-view image data are then associated with the position and attitude data at the corresponding time through their respective timestamps. Among these methods, drones equipped with LiDAR, high-definition RGB cameras, and RTK (Real-Time Kinematic) / PPK (Post-Processed Kinematic) positioning systems can conduct multi-angle surround and overhead flights over ancient architectural and cultural relic areas to collect high-density 3D point cloud data (.las / .ply format), multi-view color image data (.jpg / .png format), and position and orientation data (POS) including GPS coordinates and attitude angles calculated by IMU (Inertial Measurement Unit). The position and orientation information, which is the "image shooting state" and "point cloud acquisition state" simultaneously recorded by the drone when collecting point cloud data and multi-view image data, can provide two key "initial constraints" for matching.
[0017] To ensure data integrity, the forward overlap rate is set to be no less than 70% and the lateral overlap rate to be no less than 60% during flight. This setting aims to ensure the integrity and redundancy of the collected data, providing sufficient data support for subsequent point cloud stitching, image matching, and 3D reconstruction, thereby improving the accuracy and reliability of the mapping results. Experiments show that this overlap rate setting can effectively reduce data holes and improve the integrity of the point cloud model. In practical applications, it can significantly reduce the workload of subsequent data repair and completion; and it allows for supplementary flight strategies for areas with severe occlusion or complex structures. All collected data are timestamped to achieve accurate synchronization and registration in subsequent processes.
[0018] Step S120: Preprocess the point cloud data and the multi-view image data respectively. Based on the position and pose data, perform cross-modal registration on the preprocessed point cloud data and multi-view image data through feature matching and pose estimation to obtain a point cloud-image registration dataset. Cross-modal registration refers to the precise spatial alignment of point cloud data and image data, that is, to determine which pixel in the image data corresponds to a certain 3D point in the point cloud data.
[0019] Step S130: Semantic segmentation is performed on the point cloud data and image data in the registration dataset to obtain a first semantic segmentation result and a second semantic segmentation result. Based on the association relationship in the registration dataset, cross-modal semantic fusion is performed on the first semantic segmentation result and the second semantic segmentation result to obtain a first cross-modal semantic segmentation result. The first cross-modal semantic segmentation result includes multiple original point cloud components with category labels, and each original point cloud component is associated with the texture region in the corresponding image data through registration. The essence of semantic segmentation is to assign a category label to each element in the data. Since point cloud data and image data are different types of data, they need to be semantically segmented separately. The association relationship in the registration dataset refers to the precise correspondence between the 3D points of the point cloud and the pixels of the image. Point cloud segmentation is good at capturing the boundaries of 3D structures and components, while image segmentation is good at recognizing surface textures and material types. The purpose of cross-modal semantic fusion is to combine the segmentation results of point cloud data and image data to make up for the deficiencies of single data.
[0020] By using point-to-pixel correspondence, two semantic segmentation results can be mutually verified and supplemented.
[0021] Step S140: Based on the category labels of each original point cloud component in the first cross-modal semantic segmentation result, classify and aggregate the multiple original point cloud components to obtain segmented point cloud components, and construct the topological relationship between each segmented point cloud component. Infer the assembly relationship between each segmented point cloud component through the topological relationship. Based on the assembly relationship, call the template corresponding to the category of each segmented point cloud component from the pre-constructed BIM template library of ancient building and cultural relic components, instantiate and generate BIM model components, and structurally attach the first cross-modal semantic segmentation result to the BIM model components to construct a semantically enhanced BIM model that combines geometric morphological information and semantic attribute information. Based on the first cross-modal semantic segmentation results (i.e., cross-modal semantic point cloud and feature data), the focus is on the construction of BIM models for ancient architectural and cultural relic components. This can be achieved through three core steps: first, constructing a BIM (Building Information Modeling) model for ancient architectural and cultural relic components based on the IFC (Industry Foundation Classes) standard. The Building Information Modeling (BIM) template library, based on IFC international standards and combined with the characteristics of ancient building components, defines data structures and attribute templates for components such as brackets, beams, and column bases. For example, the "column" template presets geometric attributes such as height, radius, and verticality, as well as material attributes such as wood / stone type. Parametric design is also employed, setting variable parameters for brackets such as arch length, thickness, and number of layers to adapt to the rapid modeling needs of components of different specifications. Secondly, point cloud components are classified and their topological relationships are determined. (Topological relationship analysis can achieve automatic reasoning of assembly relationships through a four-step process: "component spatial feature quantification - proximity detection - rule base reasoning - relationship verification." First, geometric features are extracted from the segmented point cloud components (such as "column," "beam," and "bracket"), and the minimum bounding box (AABB, Axis-Aligned Bounding Box) of each component is calculated to obtain the center point coordinates.) The system first identifies the dimensions (length / width / height) and principal axis direction of components, and extracts key points on the surface of the components as feature points for relational reasoning. Then, it calculates the spatial distance between components based on the center point distance, sets a threshold (e.g., ≤0.5 meters) to filter potentially related components, and calculates the percentage of the intersecting volume of two components through point cloud collision detection. If the percentage is >5%, it is determined as "physical contact" as strong correlation evidence. Next, it constructs a rule library of topological relationships for ancient architectural components, combining component type and spatial orientation for reasoning. For example, in a support relationship, if the projections of the two endpoints of component A (beam) fall within the top bounding box of component B (column), and the angle between the principal axis of the beam and the horizontal plane is <10°, and the angle between the principal axis of the column and the vertical line is <5°, then the reasoning is "column support". In the "beam" connection relationship, if the distance between the "arch" component of the bracket set and the edge of the beam is less than 0.1 meters (unit: m), and the normal direction of the bracket set is perpendicular to the surface of the beam, and the elevation difference between their center points is less than 0.2 meters, then the reasoning is "bracket set connects beam". The rule base can cover more than 20 common assembly types of ancient buildings. Finally, the physical rationality of the preliminary reasoning relationship is verified (e.g., "support relationship" must meet the requirement that the projected area of the upper component falls within the top range of the lower component), and graph theory is used to remove redundancy. The components are used as nodes and the relationships are used as edges to construct a topological graph. Weakly related edges with a weight of less than 0.6 are deleted, and strongly related edges are retained to form the final assembly relationship network. The reasoning accuracy rate is over 95%. Using clustering algorithms such as DBSCAN and K-Means, based on the geometric features of point cloud components such as 3D coordinate distribution and curvature changes, they are aggregated into categories such as "columns," "beams," and "bracket sets," achieving a classification accuracy of over 95%. Then, a topological relationship network is constructed based on graph theory algorithms. By analyzing the spatial positions and contact boundaries between components, assembly relationships such as "beams supporting columns" and "bracket sets connecting beams" are inferred, ensuring the structural integrity of the BIM model. Finally, semantic attribute attachment and BIM model generation are completed. Using the coordinate mapping between point clouds and the BIM model, extracted material attributes such as bricks and wood, as well as structural parameters such as column height and bracket set cantilever length, are accurately attached to the classified components and written into the corresponding attribute fields of the IFC model. The model is then generated by calling the IFC standard interface. Lightweight processing such as triangular mesh simplification and texture compression, along with geometric accuracy verification with an error threshold of ±2 mm, ultimately generate an IFC format BIM model containing semantic information such as geometric shape, material, and structure. This provides standardized data support for the digital archiving, protection planning, and restoration simulation of ancient buildings.
[0022] Step S150: Integrate the semantically enhanced BIM model and the geographic information system (GIS) base map containing the geospatial benchmark to obtain the BIM-GIS fusion model corresponding to the first operation time. The BIM-GIS fusion model is used for spatial query of the components of the ancient building relics, comparison and display of the component status in multiple periods, and three-dimensional visualization.
[0023] By integrating detailed component-level BIM models with macroscopic geospatial GIS base maps, a comprehensive model that combines microscopic details with macroscopic location can be constructed. This enables full-scale management, querying, and display of ancient architectural relics, from individual components to the regional environment, providing more comprehensive digital support for the protection, research, and utilization of cultural relics.
[0024] In the above embodiments, point cloud, image, and position and attitude data are collected simultaneously by UAVs. Combined with cross-modal registration and semantic fusion, the information limitations of single data (point cloud / image) are overcome, while retaining geometric shape and semantic attributes, to construct a more complete digital model of ancient buildings and achieve deep fusion of multi-source data. Through classification aggregation, topological relationship reasoning, and BIM template instantiation, the original surveying data is transformed into a structured semantically enhanced BIM model, which includes not only the geometric details of components but also attributes such as materials and categories, providing a precise digital carrier for the protection and research of ancient buildings. Through the integration of BIM and GIS base maps, spatial query, multi-period comparison, and three-dimensional visualization of ancient building components are realized, meeting the full-scale management needs from micro-components to macro-space.
[0025] In some embodiments, the method further includes: Step S160: Repeat the steps of synchronous acquisition, cross-modal registration, and semantic segmentation at multiple second operation times according to a preset time interval to obtain the second cross-modal semantic segmentation results corresponding to the multiple second operation times. Update the semantically enhanced BIM model based on the multiple second cross-modal semantic segmentation results. Step S170: Integrate the updated semantically enhanced BIM model and GIS base map corresponding to the multiple second operation times to obtain the BIM-GIS fusion model corresponding to the multiple second operation times; Step S180: Based on the BIM-GIS fusion model corresponding to the first operation time and the plurality of second operation times, generate a digital twin animation of the building evolution process; The digital twin animation is used to display the spatiotemporal changes in the size of the components and the spatiotemporal changes in the material distribution of the ancient architectural artifacts.
[0026] This process involves periodically using drones to repeat steps S110 to S130 to acquire updated point cloud and image data of the ancient building and cultural relic area. Through point cloud registration and difference analysis, the geometric changes (such as component displacement and deformation) and material information of the BIM model are updated, generating BIM-GIS fusion models from different periods. Based on the historical model sequence, a digital twin animation of the building's evolution is generated, visually displaying the spatiotemporal changes in component size and material distribution. Details are as follows: During the periodic data collection in steps S160 to S180, data for wooden structures is collected quarterly and data for stone structures is collected annually, with void ratio controlled to <1% and consistent lighting conditions ensured to guarantee data quality. Cross-period point cloud registration (coarse registration using ISS+FPFH, rotation error ≤5°, translation error ≤5cm; fine registration using ICP with the base as the reference, RMSE ≤2mm, overlap rate >95%) and 1cm voxel mesh difference analysis are used to quantify indicators such as component displacement (early warning for column tilt >0.5°), surface deformation (reduction in relief height >1mm), and volume change (reduction rate of brackets >5%). Semantic component-level dynamic recognition and parametric updates are employed (geometric parameter accuracy ≤1mm, updating material properties such as painted fading and wood texture), with timestamp naming combined with Git LFS for version management and incremental logging. A spatiotemporal index is constructed based on keyframes of the model for each period, and digital twin animations are rendered using color mapping difference and texture blending techniques, supporting timeline control and dynamic sectioning interaction.
[0027] In the above embodiments, by periodically collecting and updating the model, the changes in the size and material of components at different periods are captured, solving the problem that static models cannot reflect the "evolution over time". Based on the multi-period BIM-GIS fusion model, evolution animations are generated to intuitively display the spatiotemporal changes of ancient buildings, providing a dynamic visualization tool for cultural relic protection and restoration (such as monitoring the development of diseases) and historical research (such as architectural change analysis).
[0028] In some embodiments, step S120, "preprocessing the point cloud data and the multi-view image data respectively", includes the following steps: Step S1201: Denoise and hole-filling processing is performed on the point cloud data; First, coarse registration of the point cloud can be achieved through FPFH (Fast Point Feature Histograms) feature matching and ICP (Iterative Closest Point) algorithm, enabling initial alignment of the point cloud from multiple perspectives. Then, statistical filters or radius filtering methods are used to remove noise and outliers, improving the overall cleanliness of the point cloud. For regions with holes or missing points, algorithms such as Poisson Surface Reconstruction are used to fill them, ensuring the integrity of the geometric structure. Finally, the surface continuity is optimized using the RANSAC (Random Sample Consensus) plane fitting algorithm. RANSAC randomly selects a subset of the point cloud to fit a plane model, and selects valid points based on the error between the model and the remaining points, iteratively obtaining the optimal plane, thereby improving the smoothness and accuracy of the point cloud surface.
[0029] Step S1202: The multi-view image data is repaired and enhanced using the posterior mean corrected stream (PMRF) algorithm.
[0030] Meanwhile, the PMRF (Posterior-Mean Rectified Flow) algorithm can be used to repair and enhance multi-view RGB images. The construction and training process of the PMRF algorithm is as follows: In its construction, the PMRF algorithm adopts a flow model architecture. It learns the posterior distribution of image restoration through an invertible transform network and solves the ordinary differential equation (ODE) based on optimal transport theory to achieve texture generation. Addressing the issues of texture defects and quality degradation in ancient architectural artifact images caused by occlusion and aging, it achieves restoration and enhancement through multi-stage processing: First, it automatically generates binary masks marking damaged areas based on computer vision algorithms. Then, it uses a UNet encoder to extract global semantic features (such as architectural style and component type) and local texture details (such as brick carving patterns, painted designs, and wood grain). A probabilistic model is used to estimate the posterior mean of the missing areas to determine the structural tendency of the restoration content. Next, it uses ODE to model the evolution path of the image from noise to clarity, generating multiple sets of restoration candidate results through iterative sampling. A perceptual loss function is then used to select the output that best matches the true texture of the ancient architectural artifact. Finally, boundary fusion and color correction eliminate restoration artifacts, restoring details such as brick carvings and painted designs while improving image contrast and color consistency, providing high-quality visual input for subsequent cross-modal registration and semantic analysis.
[0031] Texture generation is achieved by solving the ODE based on the optimal transmission theory. Simply put, this theory can find an optimal way at the mathematical level to gradually transform noisy images into clear images that conform to the real texture of ancient buildings. It is like finding the most reasonable path in a complex space, making the texture of the restored image more natural and realistic, and in line with the original appearance of the ancient buildings.
[0032] During training, in the scenario of ancient architectural artifact image restoration, the PMRF algorithm training process revolves around improving image quality and restoring realistic details. At the beginning of training, a dataset is constructed containing a large number of original clear images of ancient architectural artifacts and their corresponding degraded versions (simulating actual damage conditions such as occlusion, aging, and noise interference). These images cover a rich variety of ancient architectural artifact types, such as surface textures of ancient architectural components and partial patterns in murals. The first stage is posterior mean prediction training, aiming to minimize the mean squared error (MSE). Degraded images are input into the model, and the model, leveraging the powerful feature learning capabilities of deep neural networks, extracts and analyzes features from low-level pixel features to high-level semantic features. By continuously adjusting network parameters, it attempts to generate posterior mean prediction results that are numerically as close as possible to the original ancient architectural artifact images, thereby reducing image distortion and making the initially restored images numerically similar to the real ancient architectural artifact images. However, at this stage, the images may have problems such as blurred details and poor visual effects. The second stage involves training the correction flow model. This model learns the mapping relationship from the posterior mean prediction results to the distribution of high-quality, realistic images of ancient buildings and artifacts. This process is based on optimal transport theory and is achieved by solving the ODE. During training, the model iteratively optimizes, gradually "transporting" the posterior mean prediction images and transforming them into visually more realistic, high-quality images that conform to the true texture and characteristics of ancient buildings and artifacts. This ensures that the restored images not only have low distortion but also high perceptual quality, matching human perception of the original appearance of ancient buildings and artifacts. The entire training process uses the PyTorch deep learning framework, combined with Lightning to simplify the training process, and utilizes natten to implement a specific network architecture. After multiple rounds of training, the PMRF algorithm can accurately address the complex degradation problems of ancient building and artifact images, providing high-quality restoration and enhancement, and laying a solid data foundation for subsequent cross-modal registration and feature analysis of ancient buildings and artifacts.
[0033] To address data defects caused by severe weather, targeted algorithm optimizations are employed in the preprocessing stage: In rainy or snowy weather, point cloud denoising adds an "adaptive density filter" to the existing radius filtering, by analyzing the local point cloud density distribution (normal area density > 150 points / m²). 2The system automatically identifies low-density noise areas caused by rain and snow, and uses the K-nearest neighbor (K=10) algorithm to remove outliers, restoring the point cloud cleanliness to over 95%. Image inpainting utilizes a new "rain and snow mask automatic generation module" added to the PMRF algorithm. Based on brightness thresholds (snowflake pixel brightness > 240) and edge detection, it identifies rain and snow areas and enhances local texture coherence during inpainting (e.g., constraining the texture direction of the bracket area to be consistent with neighboring pixels), improving the signal-to-noise ratio of the inpainted image by 25%. In strong light / overexposure processing, a "dynamic range compression" step is added to overexposure images before PMRF inpainting, using the Retinex algorithm. Highlight and shadow areas are separated, with brightness reduced in highlight areas (grayscale value compressed to below 200) and contrast increased in shadow areas (texture detail enhanced by 30%) to ensure key textures are recognizable. In cloudy / foggy weather processing, point cloud completion introduces "semantic-guided completion" on top of Poisson reconstruction. Based on the geometric features of components pre-identified by PointNet++ (such as the slope angle of the eaves), the geometry of the void area is constrained, improving the completion accuracy to ±3 mm. Image enhancement is achieved by adding a "low-light sample set" (500 cloudy images) to the SegNet network training, combined with histogram equalization to expand the grayscale range and bilateral filtering to suppress noise, which improves the material classification accuracy from 75% to over 88%.
[0034] In the above embodiments, point cloud denoising and completion reduce holes and noise interference, and the PMRF algorithm repairs and enhances the image, ensuring the accuracy of subsequent registration and segmentation, and laying a high-quality data foundation for the entire process. Addressing the issues that point clouds of ancient buildings may have holes due to occlusion, and that images may suffer quality degradation due to complex environments (such as shade or backlighting), the preprocessing steps enhance the robustness of the data.
[0035] In some embodiments, step S120, "based on the position and pose data, performing cross-modal registration of the preprocessed point cloud data and multi-view image data through feature matching and pose estimation to obtain a point cloud-image registration dataset," includes the following steps: Step S1203: Using Scale Invariant Feature Transform (SIFT) or Accelerated Robust Feature Rendering (SURF) algorithm, key points are extracted from the preprocessed multi-view image data, and corresponding visual features are extracted based on the key points; geometric features are extracted from the preprocessed point cloud data; the visual features and geometric features are compared, and combined with the initial spatial mapping relationship established by the position and pose data, image-point cloud feature anchor point pairs corresponding to the same physical location are matched; based on the initial spatial mapping relationship, effective feature anchor point pairs are selected from the image-point cloud feature anchor point pairs. Step S1204: Starting from the position and pose data, based on the initial spatial mapping relationship, substitute the effective feature anchor point pairs into the perspective n-point algorithm (PnP algorithm) to determine the precise spatial pose of the RGB camera during shooting, as well as the precise spatial mapping relationship between image pixels and point cloud 3D coordinates. Step S1205: The preprocessed point cloud data, the preprocessed multi-view image data, the precise spatial pose, and the precise spatial mapping relationship are structurally integrated to obtain a point cloud-image registration dataset.
[0036] After restoration, the PMRF-enhanced image and point cloud data are registered across modally. This process extracts key point features from the image using SIFT (Scale-Invariant Feature Transform) / SURF (Speeded UpRobust Features) and combines them with POS pose information (i.e., position and pose data). The PnP (Perspective-n-Point) algorithm is then used to estimate the camera pose, thereby accurately mapping the image into the 3D point cloud space and ultimately generating an RGB point cloud (Colorized Point Cloud). This provides high-precision, high-quality input data for subsequent semantic segmentation and BIM modeling.
[0037] What is needed is that, after denoising and completion, the point cloud data is spatially registered (coordinate aligned) with the image data, but the image texture is not integrated. After the image data is repaired and enhanced, a spatial mapping relationship is established with the point cloud data, but it is not merged with the point cloud.
[0038] In the above embodiments, visual features are extracted using SIFT / SURF, and pose is optimized by combining geometric feature matching and PnP algorithm to solve the spatial alignment problem of cross-modal data (3D point cloud and 2D image), ensuring a one-to-one correspondence between "point cloud" and "image" features. An initial mapping is established based on position and pose data, and then effective anchor point pair filtering and accurate pose estimation are used to balance registration speed and accuracy, providing a reliable spatial association foundation for subsequent semantic fusion.
[0039] In some embodiments, step S130, "performing semantic segmentation on the point cloud data and image data in the registration dataset to obtain a first semantic segmentation result and a second semantic segmentation result," includes the following steps: Step S1301: Use the PointNet++ network to perform semantic segmentation on the point cloud data in the registration dataset to extract the three-dimensional boundaries and structural parameters of the ancient architectural artifact components. The PointNet++ point cloud semantic segmentation process achieves high-precision 3D analysis of ancient architectural artifact components through hierarchical feature extraction and geometric parameterization analysis. First, a multi-level network architecture is constructed, forming a point cloud pyramid through farthest point sampling (FPS) and spherical neighborhood grouping (Ball Query). FPS ensures that representative points are selected in the point cloud distribution, avoiding over-concentration of sampling points. Ball Query groups the selected representative points into neighborhoods using spherical regions, which effectively captures the local geometric features of the point cloud at different scales, providing a guarantee for the subsequent accurate extraction of the 3D boundaries and structural parameters of ancient architectural artifact components. The underlying MLP (Multi-Layer Perceptron) extracts local geometric descriptors of coordinates and normal vectors, while the higher-level network captures multi-scale features such as component corners and surface transitions. Symmetric functions such as max pooling are used to eliminate the influence of point cloud disorder. Subsequently, based on the feature space distribution, a classifier assigns semantic labels (such as "column base" and "bracket set") to each point. Combined with CRF (Conditional Random Field), the boundary continuity is optimized, and the three-dimensional contour of the component is extracted. Finally, the segmentation results are subjected to geometric parameterization analysis, automatically calculating the component dimensions (length, width, height, radius of curvature), spatial relationships (angle, offset), and surface features (roughness, depth of concavity and convexity), forming a three-dimensional boundary model containing accurate geometric contours and quantified parameters, providing basic data support for BIM modeling and disease assessment of ancient buildings and cultural relics.
[0040] Step S1302: Use the SegNet network to perform pixel-level classification on the image data in the registration dataset to identify the material type and surface detail features of the ancient architectural artifact components.
[0041] Among them, SegNet image semantic segmentation and cross-modal fusion technology achieves accurate analysis and cross-modal integration of ancient building material information through multi-stage collaboration: In the image semantic segmentation stage, the VGG16 pre-trained model can be used as the encoder, combined with the deconvolution decoder and skip connections to extract semantic features. The sample is expanded by data augmentation strategies such as ±45° rotation and ±20% brightness adjustment. The image pixels are classified into 6 material types: blue brick, wood, stone, painted, stucco, and metal by using the weighted cross-entropy and Dice loss optimization model. After morphological processing (opening / closing operation) and CRF (Conditional Random Field) optimization, a color semantic label map (PNG format) and a category intensity matrix (.npy) are generated, and quantitative features such as RGB mean and texture frequency (i.e., surface detail features) are extracted.
[0042] In the above embodiments, PointNet++ extracts the three-dimensional structural features (boundaries, parameters) of the point cloud, and SegNet identifies the material and surface details of the image, respectively analyzing the information of ancient buildings from the levels of "three-dimensional morphology" and "two-dimensional texture" to improve the comprehensiveness of semantic segmentation. To address the disordered nature of point clouds and the pixel-level distribution characteristics of images, a dedicated network model is selected to ensure that the segmentation results better match the structural and material characteristics of the ancient building components.
[0043] In some embodiments, step S130, "based on the association relationship in the registration dataset, performing cross-modal semantic fusion on the first semantic segmentation result and the second semantic segmentation result to obtain the first cross-modal semantic segmentation result," includes the following steps: Step S1303: Based on the association relationship in the registration dataset, the three-dimensional boundary and the structural parameters are mapped to the same dimension as the material type and the surface detail features through the fully connected layer of the cross-modal fusion network and then concatenated to obtain the cross-modal feature vector; Step S1304: Calculate the mutual information of the cross-modal feature vectors through the attention module of the cross-modal fusion network; perform weighted calculation on the feature pairs in the cross-modal feature vectors based on the mutual information to obtain the weighted cross-modal feature vectors; Step S1305: The joint classifier of the cross-modal fusion network is used to determine the component category and material type of the weighted cross-modal feature vector, and to correct the semantic boundary, so as to obtain the first cross-modal semantic segmentation result.
[0044] The process involves using a SegNet network to perform pixel-level classification on the enhanced image to identify material types and surface detail features. Cross-modal semantic fusion is used to achieve semantic coordination between point cloud geometric features and image texture features. The final output is a first cross-modal semantic segmentation result that integrates 3D boundaries, material labels, and category intensity information. During cross-modal fusion, the point cloud coordinates are projected onto the image plane using the registration parameters in S120. Neighboring pixels are searched to obtain material labels, and point cloud geometric features and image texture features are stitched together to form a cross-modal vector. Simultaneously, the component structure identified by PointNet++ constrains image classification, and image textures are used to supplement point cloud details. The final output includes point cloud data containing 3D coordinates, material properties, and cross-modal features, as well as a semantic label map, material statistics report, and point cloud-image association mapping table, providing geometric and material fusion information for BIM modeling.
[0045] Cross-modal semantic fusion achieves deep collaboration between point cloud geometric features and image texture features through a three-level mechanism of "spatial mapping-feature alignment-cooperative optimization". The specific process is as follows: Spatial mapping and feature association: Based on the camera extrinsic matrix (rotation matrix R, translation vector T) obtained by the PnP algorithm in S120, the spatial mapping relationship between point cloud and image is established: the three-dimensional coordinates (X,Y,Z) of point cloud are converted into the pixel coordinates (u,v) of image plane through perspective projection formula, as shown in the following formula (1): Formula (1); in, , This refers to the camera's intrinsic focal length. , The coordinates of the main point.
[0046] For each point in the point cloud, a 3×3 pixel neighborhood of its corresponding image region is found using bilinear interpolation. The material label (e.g., "pine wood" or "painted"), RGB mean (e.g., R=220, G=60, B=40), and texture frequency features (e.g., 5 lines / cm) output by SegNet are extracted to form an image texture feature vector. ( =10, including 1 tag code + 3 RGB values + 6 texture statistics).
[0047] Feature Dimension Alignment and Stitching: Extracting geometric features of the point cloud output from PointNet++: including 3D coordinates (X,Y,Z), normal vectors ( The curvature k and semantic tags (such as "dougong" and "column") form a geometric feature vector. ( =8, containing 3 coordinates + 3 normal vectors + 1 curvature + 1 label encoding), which is then processed through a fully connected layer. and Mapping to the same dimensional space (e.g., d=32), feature concatenation is used to obtain cross-modal feature vectors. (; indicates vector concatenation), achieving dimensional alignment of geometric and texture features.
[0048] Attention Mechanisms and Collaborative Optimization: Introducing a Cross-Modal Attention Module Perform weighted optimization: by calculating the mutual information between geometric features and texture features (such as MI). ()) Assign higher weights to features with high relevance (e.g., increase the weight of the geometric feature of "cantilever length" and the texture feature of "painted pattern" of the bracket by 20%), and suppress noise features (e.g. invalid textures in occluded areas).
[0049] Finally, the weighted cross-modal features are input into the joint classifier (softmax layer). The 3D contour of the component identified by PointNet++ constrains the image material classification boundary (e.g., the pixel material within the boundary of "cylinder" is only "wood" or "stone"). At the same time, the semantic labels of the point cloud are corrected by the image texture details (e.g., the label is corrected after the area of the point cloud is misclassified as "brick carving" is confirmed as "painted" by the image texture). Finally, the first cross-modal semantic segmentation result is output, which integrates the 3D boundary, material attributes and category strength (confidence ≥ 0.85). Among them, the category strength is used to evaluate the reliability of the material attribute attachment, and the category strength only corresponds to the material classification confidence.
[0050] In the above embodiments, a cross-modal network maps 3D structure and 2D texture features to the same dimension. Combined with an attention mechanism and weighted mutual information, it resolves ambiguities that may exist in single-modal segmentation (such as blurred point cloud boundaries and missing image depth) and corrects semantic boundaries. A joint classifier uniformly determines the component category and material, ensuring that the semantic labels of point clouds and images are consistent, providing a reliable basis for subsequent attribute attachment to the BIM model.
[0051] In some embodiments, the Building Information Model (BIM) template library for ancient architectural and cultural relic components is constructed based on the Industrial Foundation Class (IFC) standard, and the GIS base map contains topographic data and cadastral data in CityGML format. Step S150, "integrating the semantically enhanced BIM model and the geographic information system (GIS) base map containing a geospatial benchmark to obtain a BIM-GIS fusion model," includes the following steps: Step S1501: The semantically enhanced BIM model is converted into the CityGML format. The IFC semantic information is preserved through the LOD hierarchical modeling based on the accuracy requirements of ancient building components and the AppSchema extension mechanism of the application mode, so as to obtain the BIM model in CityGML format. Step S1502: Integrate the CityGML format BIM model and GIS base map to obtain the BIM-GIS fusion model corresponding to the first operation time.
[0052] The generated BIM model can be converted to CityGML (City Geography Markup Language) format. Through LOD (Level of Detail) hierarchical modeling and AppSchema extension mechanisms, IFC semantic information (such as material type and component ID) is preserved. The CityGML format BIM model is seamlessly integrated with the GIS (Geographic Information System) base map, enabling spatial querying of components (such as clicking on a column to display its material and dimensions), multi-source data overlay (such as comparing historical scan data), and 3D visualization (supporting WebGL interactive browsing), as detailed below: Step S150 focuses on the deep integration of the ancient building BIM model and the GIS platform, constructing a 3D geographic information system with spatial analysis and visualization capabilities through multi-dimensional technology integration. First, a standardized conversion from the BIM model to CityGML is implemented, using open-source tools such as FZK-IFC2CityGML or a custom converter. Based on ISO standards, a mapping relationship between IFC components and CityGML entities is established, mapping IfcColumns to CityGML BuildingParts. The RGB values and texture paths of IfcMaterial materials are stored through Appearance nodes, and component IDs are retained using gml:id to ensure data traceability. Simultaneously, a LOD hierarchical modeling strategy is adopted, from the simplified city-level outline of LOD1 (error ≤ 5m) to the component-level micro-details of LOD4 (error ≤ 1mm), dynamically loading detail levels through the lod1Geometry to lod4Geometry nodes. In terms of semantic information preservation, a special extended mode for ancient buildings was created based on the CityGML ApplicationSchema mechanism. AncientMaterial material attribute groups and AncientComponent component parameter groups were defined. The assembly relationships and material labels in the IFC were converted into CityGML topology links and Appearance node information through the propertySet node, achieving a two-way association between semantics and visualization. During GIS platform integration, a seven-parameter transformation method was used to convert the BIM local coordinate system to the WGS84+UTM geodetic coordinate system (error ≤10cm). Topographic data was aligned using the 1985 National Elevation Datum. Component-level query functions (clicking on a column displays material, dimensions, and maintenance records) and spatial statistical analysis (generating heatmaps by material) were developed, supporting multi-source overlay display of historical scan data, cadastral data, and environmental data. 3D visualization achieved lightweight WebGL rendering (frame rate ≥30fps) through the Draco compression algorithm (10:1 compression ratio), supporting roaming, sectioning, and measurement operations. Simultaneously, a WebGL mobile application compatible with AR mode and multi-user collaborative annotation functions was developed. This stage of technology integrates BIM geometric material information with GIS spatial analysis capabilities, and can be applied to scenarios such as digital management of cultural heritage, urban renewal planning, and emergency evacuation simulation.
[0053] Seamless integration of BIM models and GIS base maps is achieved through a three-tiered technical system: precise coordinate system transformation, structured semantic information mapping, and integration accuracy verification. For cross-coordinate system matching, the system first analyzes the translation, rotation, and scale differences between the local construction coordinate system used in the BIM model (e.g., a right-handed coordinate system with the building foundation center as the origin) and the National Geodetic Coordinate System 2000 (WGS84+UTM partition, such as UTM Zone 50N where the Yingxian Wooden Pagoda is located) used in the GIS base map. The maximum translation can reach hundreds of meters, and the rotation error is ≤3°. Then, by selecting more than three public control points around the ancient building (e.g., corners, boundary markers), and simultaneously acquiring their BIM local coordinates (total station measurement, accuracy ±2mm) and GIS geodetic coordinates (GNSS static measurement, accuracy ±5cm), the system calculates translations (ΔX, ΔY, ΔZ) and rotations (ε) based on the Bursa model. x ,ε y ,ε zThe seven parameters (including scale (m)) were optimized using the least squares method to ensure that the transformation residuals were ≤3cm (e.g., the plane error of the Yingxian Wooden Pagoda after transformation was ≤8cm, and the elevation error was ≤5cm). Finally, all geometric nodes of the CityGML model (e.g., BuildingPart, LOD4Geometry) were transformed to the WGS84+UTM coordinate system using the seven parameters. Fifty feature points (e.g., column tops, bracket ends) were randomly selected and compared with the GIS base map (1:1000 topographic map) to ensure that the matching error was ≤10cm. In terms of IFC semantic information preservation, different LOD levels were treated using "geometric nodes + attributes". The nodes have a dual-structure semantic binding. LOD1-LOD2 (simplified model) stores simplified geometry in the lod1Geometry / lod2Geometry nodes, and associates basic semantics (such as component ID and material category) through the propertySet node. LOD3-LOD4 (refined model) stores textured triangular meshes in the lod3Geometry / lod4Geometry nodes, and retains material RGB values and texture paths (mapped to IFC's IfcMaterialTexture) through the Appearance node, and uses g... ml:id corresponds one-to-one with the GlobalId of IFC (e.g., "Column-001" in IFC corresponds to "Pillar-001" in CityGML). At the same time, based on the ApplicationSchema mechanism of CityGML, the "Semantic Extension Schema of Ancient Buildings" is defined, adding the AncientMaterial group (containing "Material Type", "Moisture Content", and "Weathering Grade", corresponding to IfcMaterialProperties in IFC) and the AncientComponent group (containing "Component Age" and "Repair Record", corresponding to IfcElementQuantity in IFC). The mapping relationship is established through XSLT transformation script (e.g., "IfcColumn.Height=11.23m" in IFC is mapped to "AncientComponent / Height=11.23" in CityGML). The mapping coverage reaches 100%. After the transformation, the number of attributes is compared in batches through Python script (e.g., if IFC contains 120 attributes, CityGML needs to retain ≥118 attributes), and key attributes are manually checked 100% to ensure that the semantic loss rate is <2%.
[0054] In the above embodiments, a BIM template library is built based on the IFC standard, converted to CityGML format, and semantic information is preserved, thus overcoming the format barriers between BIM and GIS and achieving cross-platform data sharing. Through LOD grading and AppSchema extension, it is ensured that the integration process preserves both the micro-details of components (such as the texture of brackets) and meets the macro-management needs at the city level, balancing accuracy and efficiency.
[0055] In some embodiments, the method further includes: Step S101: Obtain meteorological data monitored by meteorological satellites and ground meteorological stations within a preset time period; Step S102: When the meteorological data meets the preset weather conditions suitable for UAV operation, the first operation time of the UAV platform is determined based on the meteorological data, and the adaptation scheme of the lidar, the RGB camera, the positioning system and the UAV platform is determined. The adaptation scheme includes the addition of protective equipment, switching of working modes and parameter optimization.
[0056] During data acquisition, various weather conditions such as rain, snow, strong light, overcast skies, and fog directly impact the quality of point cloud data, multi-view image data, and position and attitude data collected by drones. These weather conditions can be considered unsuitable for drone operations, while light winds, no precipitation, high visibility, and soft lighting are suitable. Rain and snow cause LiDAR laser reflection signal scattering, increasing point cloud noise (outlier ratio increases from 1%-2% to over 10%). Rain also causes image blurring and loss of texture information due to lens adhesion (e.g., the discernibility of details in dougong (bracket set) paintings decreases by more than 30%). Strong light causes image overexposure (pixel saturation in highlight areas), resulting in loss of material texture details, and LiDAR specular reflection causes local holes in the point cloud. Overcast / foggy skies reduce image contrast due to insufficient light (grayscale values are concentrated in the 50-150 range), increasing the difficulty of material classification (e.g., the accuracy of distinguishing between blue bricks and plaster decreases from 92% to 75%), while reducing laser penetration and decreasing point cloud density (e.g., from 200 points / m in distant areas). 2 Reduced to 80 points / m 2 To address this, a three-tiered approach of "weather forecasting, equipment adaptation, and parameter optimization" was adopted during the data collection phase: For 48 hours prior to data collection, meteorological data was monitored via meteorological satellites and ground weather stations to avoid extreme weather conditions, prioritizing operations during cloudy periods with humidity below 60%, and developing supplementary flight plans for mild or severe weather; in rainy or snowy weather, drones were equipped with rain covers and lenses were coated with hydrophobic coatings, and the LiDAR was put into anti-interference mode; in strong light, the camera was equipped with an adjustable ND (Neutral Density Filter) filter, and the LiDAR was put into anti-specular reflection algorithm mode; during periods of strong light, flight altitude was adjusted, exposure time was shortened, and ISO (sensitivity) and lateral overlap were increased; in cloudy / foggy weather, exposure time was extended, HDR (High Dynamic Range) mode was activated, and the LiDAR scanning frequency and lateral overlap were increased.
[0057] In the above embodiments, the timing of operations is determined based on meteorological data, and equipment adaptation plans (such as adding protection and optimizing parameters) are formulated to avoid adverse weather conditions affecting data quality or damaging equipment, ensuring that the drone data collection process is stable and controllable. By triggering operation decisions based on preset weather conditions, the cost of manual judgment is reduced, and the automation and intelligence level of ancient building surveying is improved.
[0058] This application utilizes a drone platform equipped with LiDAR, a high-definition RGB camera, and a real-time differential positioning system (RTK / PPK) to perform multi-angle aerial scanning of ancient architectural relics, collecting high-precision point cloud data, multi-view image data, and position and attitude (POS) information. The Posterior-Mean Rectified Flow (PMRF) algorithm is used to repair and enhance the images, and point cloud optimization is combined to achieve accurate cross-modal registration. PointNet++ and SegNet networks are used to perform semantic segmentation and fusion of point cloud data and image data, respectively, extracting the 3D boundaries, structural parameters, material types, and surface detail features of the ancient architectural components. A BIM template library is constructed based on the International Federation of Collaborative Classification (IFC) standard to generate semantically enhanced BIM models, which are then converted to CityGML format for seamless integration with Geographic Information System (GIS) base maps, supporting component spatial query, multi-source data overlay, and 3D visualization. This invention generates digital twin animations by regularly updating data (using a drone platform to collect new point cloud, image, and POS data; calibrating data update deviations based on the coordinate consistency of old and new POS data; and generating digital twin animations of changes in the form of ancient buildings), enabling long-term dynamic monitoring of the health status of ancient buildings. Compared with traditional surveying and existing technologies, this invention improves data acquisition efficiency and quality, and realizes automated and intelligent surveying and management of ancient architectural relics.
[0059] This application provides an intelligent surveying and mapping method based on the fusion of AI and BIM, including the following steps: S1, an unmanned aerial vehicle platform equipped with LiDAR, RGB cameras and RTK / PPK positioning system, performs multi-angle flight scanning of the ancient building and its surrounding environment, collecting high-precision point cloud data, multi-view image data and POS pose data; S2. Denoising and hole filling are performed on the original point cloud. At the same time, the PMRF algorithm is used to repair and enhance the multi-view RGB image. Accurate cross-modal registration of the image and point cloud is achieved through feature matching and pose estimation. S3. Input the point cloud and image data processed in S2. Use PointNet++ to perform semantic segmentation on the registered point cloud data to extract the three-dimensional boundaries and structural parameters of the ancient building artifact components. At the same time, use the SegNet network to perform pixel-level classification on the enhanced image to identify material type and surface detail features. And achieve semantic coordination between point cloud geometric features and image texture features through cross-modal semantic fusion. Finally, output the cross-modal semantic segmentation result that integrates three-dimensional boundary, material properties and category intensity information. S4. Construct a BIM template library for ancient building and cultural relic components based on the IFC standard. Use algorithms to classify the segmented point cloud components and infer assembly relationships through topological relationship analysis (such as "beam supporting column" and "bracket connecting beam"). Connect the material attributes and structural parameters (such as column height and bracket cantilever length) extracted in S3 to the BIM model components to generate a semantically enhanced BIM model (IFC format) containing geometric information and material attributes. S5. Convert the generated BIM model into CityGML format, retain IFC semantic information (such as material type and component ID) through LOD hierarchical modeling and AppSchema extension mechanism, and seamlessly integrate the CityGML format BIM model with GIS base map (topography and cadastral data) to realize component spatial query (such as clicking on the column to display material and size), multi-source data overlay (such as historical scan data comparison) and 3D visualization display (supports WebGL interactive browsing). S6. Regularly use drones to repeat the S1-S3 process to acquire updated point cloud and image data of ancient building and cultural relic areas. Through point cloud registration and difference analysis, update the geometric changes (such as component displacement and deformation) and material information of the BIM model, generate BIM-GIS fusion models of different periods, and generate digital twin animations of the building evolution process based on the historical model sequence. This will intuitively show the spatiotemporal changes in component size and material distribution, and support long-term dynamic monitoring of the health status of ancient buildings.
[0060] In some embodiments, taking the Yingxian Wooden Pagoda in Shanxi Province (an ancient wooden structure from the Liao Dynasty) as an example, the specific implementation process and application effects of the technical solution are described in detail.
[0061] During the data collection phase of the drone: Equipment configuration: Drones can be used, equipped with a point cloud density of ≥200 points / m². 2 It features a lidar system, a 61-megapixel high-definition camera, and an RTK / PPK positioning system with a planar accuracy of ±1cm and an elevation accuracy of ±2cm.
[0062] Flight strategy: Perform a spiral-shaped orbit around the main body of the wooden pagoda (total height 67.31m) with layered overhead views, setting a 75% forward overlap and a 65% lateral overlap. Perform two supplementary flights to complex structural areas such as the first-floor brackets and the second-floor platform. Generate .las format point clouds (approximately 1.2TB in total, point cloud spacing ≤5mm), .png format images (approximately 3000 images), and POS data including GPS / IMU (timestamp accuracy 1ms).
[0063] Supplementary measurements were taken after a rainstorm in July 2024: a rain cover was used during data collection, and the point cloud noise rate was reduced to 5% in LiDAR anti-interference mode; later, through adaptive density filtering, the point cloud cleanliness of the dougong area was restored to 96%, and the PMRF algorithm was used to repair the painted image obscured by rain, with texture clarity reaching 85% of that in normal weather.
[0064] Data collection under strong midday light in August 2024: The camera was equipped with an ND16 filter, the exposure time was 1 / 1000s, and 3-level images were synthesized in HDR mode; after processing, the proportion of overexposed areas decreased from 20% to 5%, and SegNet maintained a classification accuracy of over 90% for pine wood and painted wood.
[0065] In the point cloud and image preprocessing and cross-modal registration stages: Point cloud processing: For coarse registration, FPFH feature matching + ICP algorithm can be used to align point clouds of multiple points, with an initial rotation error ≤8° and translation error ≤10cm. For denoising and completion, outliers are removed by radius filtering (threshold 0.05m), and Poisson reconstruction algorithm is used to fill voids (volume approximately 0.3m³) in the tower body cracks. RANSAC plane fitting is used to optimize the continuity of the tower body surface (plane error ≤0.8mm).
[0066] Image Restoration: The PMRF algorithm was used to process severely occluded dougong (bracket set) painted images. Dataset Construction: 2000 high-resolution texture images of the Yingxian Wooden Pagoda were collected, along with 500 simulated aging and fading samples. Training Process: First, features such as dougong patterns and column wood grain were extracted using a UNet encoder, reducing the MSE loss to 0.012 during the posterior mean prediction stage. In the corrective flow model training, the ODE was solved based on optimal transport theory, resulting in a perceptual loss of ≤0.035 for the restored images, improving texture clarity by 40%.
[0067] Cross-modal registration: Extract SIFT keypoints (approximately 2000 per image), combine POS data with the PnP algorithm to calculate camera extrinsic parameters, map the image to point cloud space, and generate RGB point cloud (color fidelity ≥90%).
[0068] In the semantic segmentation and cross-modal fusion stage: Point cloud semantic segmentation (PointNet++): Input the registered point cloud (approximately 80 million points), construct the point cloud pyramid through FPS sampling (sampling rate 10%) and Ball Query (radius 0.2m), identify components such as "column", "bracket", and "eaves", with a 3D boundary extraction accuracy ≤2mm, and the measurement error of the bracket cantilever length ≤1.5mm.
[0069] Image semantic segmentation (SegNet): The enhanced image is classified into three categories: "pine wood", "blue brick" and "painted". The pixel-level classification accuracy is 92%. The RGB mean of the painted area (e.g., red painted R=220±5, G=60±3, B=40±2) and texture frequency (average 5 lines / cm) are extracted.
[0070] Cross-modal fusion: Project the point cloud coordinates onto the image plane, obtain the material labels of neighboring pixels, generate a cross-modal point cloud containing three-dimensional coordinates (X / Y / Z accuracy ±3mm) and material properties (such as the density of pine wood column 0.54g / cm³), and output a semantic label map and a point cloud-image association table.
[0071] During the semantically enhanced BIM model construction phase: BIM Template Library: Based on the IFC standard, create templates for wooden tower components, such as the "five-story eaves column" template with preset parameters such as height 11.2m, diameter 0.8m, and wood moisture content of 12%, and the bracket template with defined variable parameters such as arch length 45cm and overhang thickness 12cm.
[0072] Point cloud classification and topological reasoning: DBSCAN clustering (eps=0.1m, minPts=50) was used to classify point clouds into categories such as columns (89) and brackets (546 groups), with a classification accuracy of 96%. Assembly relationships were reasoned based on graph theory algorithms, such as "the east column of the second-floor main hall supports the east lintel of the second floor" and "the corner brackets connect the eaves purlins and column heads", and a topological network (edge weight ≥ 0.9) was constructed.
[0073] Attribute linking: Link the column height (11.23m±0.02m) and bracket material (pine wood) extracted from S3 to the IFC model to generate an IFC format BIM model (file size 1.8GB, geometric error ≤2mm).
[0074] In the BIM-GIS integration and visualization stage: Format conversion: Use the FZK-IFC2CityGML tool to convert the BIM model to LOD4 level CityGML, and define attribute groups such as "Wooden Pagoda Bracket" and "Liao Dynasty Painted Decoration" through AppSchema extension, while retaining component IDs (such as Gong_001) and material RGB values (painted decoration red R=218, G=56, B=32).
[0075] GIS Integration: A seven-parameter transformation converts the model coordinates to the WGS84+UTM Zone 50N coordinate system (planar error ≤8cm), overlaying 1:1000 topographic data and cadastral red lines. A WebGL interactive system was developed; clicking on the third-floor west pillar displays: material (pine), dimensions (height 9.8m, diameter 0.78m), and historical repair records (2016 anti-corrosion treatment of the pillar).
[0076] Visualization: Smooth rendering on the webpage (35fps) is achieved through Draco compression (compression ratio 12:1), and multi-source data overlay is supported (e.g., comparison of 2010 scanned point cloud with the current model, with the difference displayed using a heatmap).
[0077] In the dynamic monitoring and digital twin stage: Regular data collection: The wooden tower is scanned by drone every quarter (monitoring frequency of wooden structures). A total of 4 data collections were carried out from March to September 2024. The void rate was less than 0.8% and the lighting conditions were controlled at 500-800 lux.
[0078] Change Analysis: Point Cloud Registration: Coarse registration used ISS+FPFH (rotation error ≤3°, translation error ≤3cm), fine registration used the base as the reference (RMSE=1.5mm, overlap rate 97%). Difference Analysis: It was found that the east corner column of the second floor was tilted by 0.6° (exceeding the warning threshold of 0.5°), and the volume of the mortise and tenon joint of the bracket set was reduced by 6%, generating a deformation report (error ≤1mm).
[0079] Digital twin animation: Based on the model generated in April 2024, the animation shows the changes in the tilt of the columns with color mapping (red represents deformation >0.5°), and the texture blending technology restores the fading process of the painted decoration. It supports dragging the timeline to view the displacement trajectory of the brackets from March to September 2024 (maximum displacement 2.3mm).
[0080] Summary of implementation results: Accuracy indicators: 3D coordinate error ≤3mm, material classification accuracy 92%, BIM model geometric accuracy ±2mm, dynamic monitoring deformation recognition accuracy ≤1mm.
[0081] Application value: A digital archive containing geometry, materials, and historical evolution was established for the Yingxian Wooden Pagoda, supporting cultural relic protection units in formulating targeted restoration plans (such as supporting and reinforcing tilted columns), and realizing full-process digital management of ancient buildings from static modeling to dynamic health monitoring.
[0082] The beneficial effects of the embodiments of this application include: Data quality improvement: Using the PMRF algorithm to repair and enhance images can effectively solve problems such as occlusion, blurring or damage in images. At the same time, combined with point cloud optimization techniques, it can achieve accurate cross-modal registration between images and point clouds, thus improving the overall data quality.
[0083] Precise semantic segmentation: PointNet++ and SegNet networks are used to process the registered point cloud data and the enhanced image respectively, realizing the extraction of three-dimensional boundaries and structural parameters of ancient building and cultural relics components, as well as the identification of material type and surface detail features. Through cross-modal semantic fusion, the results are more accurate and comprehensive.
[0084] Generate a semantically enhanced BIM template library: Based on the IFC standard, construct a BIM template library for ancient building and cultural relic components (which not only includes geometric information, but also integrates material properties, topological relationships, and historical evolution information). Use algorithms to classify the segmented point cloud components, infer assembly relationships, and attach the extracted material properties and structural parameters to the BIM model components. Finally, generate a semantically enhanced BIM model containing geometric information and material properties, providing standardized data support for the digital archiving, protection planning, and restoration simulation of ancient buildings.
[0085] Deep integration of BIM and GIS: The generated BIM model is converted into CityGML format, and IFC semantic information is preserved through LOD hierarchical modeling and AppSchema extension mechanism, achieving seamless integration with GIS base map. It supports component spatial query, multi-source data overlay, and 3D visualization, improving the efficiency of information sharing and collaborative work. It solves the limitation of traditional methods that can only perform spatial location overlay.
[0086] Based on the foregoing embodiments, this application further provides an intelligent surveying and mapping system based on the fusion of AI and BIM. The system includes various modules and units included in each module, which can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0087] Figure 2 A schematic diagram illustrating the structural composition of an intelligent surveying and mapping system based on the fusion of AI and BIM, provided as an embodiment of this application, is shown below. Figure 2As shown, the system 200 includes an acquisition module 210, a registration module 220, a semantic segmentation module 230, a first generation module 240, and a first integration module 250, wherein: The acquisition module 210 is used to perform multi-angle flight scanning of the ancient building and its surrounding environment at the first moment of operation using a drone platform equipped with lidar, RGB camera and positioning system, and simultaneously acquire point cloud data, multi-view image data and position and attitude data, so that the point cloud data and the multi-view image data are associated with the position and attitude data at the corresponding moment through their respective timestamps. Registration module 220 is used to preprocess the point cloud data and the multi-view image data respectively, and based on the position and pose data, perform cross-modal registration of the preprocessed point cloud data and multi-view image data through feature matching and pose estimation to obtain a point cloud-image registration dataset. The semantic segmentation module 230 is used to perform semantic segmentation on the point cloud data and image data in the registration dataset to obtain a first semantic segmentation result and a second semantic segmentation result. Based on the association relationship in the registration dataset, the first semantic segmentation result and the second semantic segmentation result are subjected to cross-modal semantic fusion to obtain a first cross-modal semantic segmentation result. The first cross-modal semantic segmentation result includes multiple original point cloud components with category labels, and each original point cloud component is associated with a texture region in the corresponding image data through registration. The first generation module 240 is used to classify and aggregate the multiple original point cloud components based on the category labels of each original point cloud component in the first cross-modal semantic segmentation result, obtain segmented point cloud components, construct the topological relationship between each segmented point cloud component, infer the assembly relationship between each segmented point cloud component through the topological relationship, and, based on the assembly relationship, call the template corresponding to the category of each segmented point cloud component from the pre-constructed building information model (BIM) template library of ancient building and cultural relic components, instantiate and generate BIM model components, and structurally attach the first cross-modal semantic segmentation result to the BIM model components to construct a semantically enhanced BIM model that combines geometric morphological information and semantic attribute information. The first integration module 250 is used to integrate the semantically enhanced BIM model and the geographic information system (GIS) base map containing the geospatial benchmark to obtain the BIM-GIS fusion model corresponding to the first operation time. The BIM-GIS fusion model is used to perform spatial query of the components of the ancient building relics, comparison and display of the component status in multiple periods, and three-dimensional visualization.
[0088] In some embodiments, the system 200 further includes: an update module, configured to repeat the steps of synchronous acquisition, cross-modal registration, and semantic segmentation at multiple second operation times according to a preset time interval, to obtain second cross-modal semantic segmentation results corresponding to the multiple second operation times, and to update the semantically enhanced BIM model based on the multiple second cross-modal semantic segmentation results; a second integration module, configured to integrate the updated semantically enhanced BIM model and GIS base map corresponding to the multiple second operation times to obtain a BIM-GIS fusion model corresponding to the multiple second operation times; and a second generation module, configured to generate a digital twin animation of the building evolution process based on the first operation time and the BIM-GIS fusion model corresponding to the multiple second operation times; wherein the digital twin animation is used to display the spatiotemporal changes in the component size and material distribution of the ancient building artifacts.
[0089] In some embodiments, the registration module 220 includes: a first processing submodule for denoising and hole-filling processing of the point cloud data; and a second processing submodule for repairing and enhancing the multi-view image data using the posterior mean modified stream (PMRF) algorithm.
[0090] In some embodiments, the registration module 220 further includes: an extraction submodule, configured to extract key points from the preprocessed multi-view image data using Scale Invariant Feature Transform (SIFT) or Accelerated Robust Feature Rendering (SURF) algorithms, and extract corresponding visual features based on the key points; extract geometric features from the preprocessed point cloud data; compare the visual features and the geometric features, and match image-point cloud feature anchor pairs corresponding to the same physical location by combining the initial spatial mapping relationship established by the position and pose data; filter out effective feature anchor pairs from the image-point cloud feature anchor pairs based on the initial spatial mapping relationship; a determination submodule, configured to use the position and pose data as a starting point, and based on the initial spatial mapping relationship, substitute the effective feature anchor pairs into the Perspective n-Point (PnP) algorithm to determine the precise spatial pose when the RGB camera was shooting, and the precise spatial mapping relationship between image pixels and point cloud three-dimensional coordinates; and an integration submodule, configured to structurally integrate the preprocessed point cloud data, the preprocessed multi-view image data, the precise spatial pose, and the precise spatial mapping relationship to obtain a point cloud-image registration dataset.
[0091] In some embodiments, the semantic segmentation module 230 includes: a first semantic segmentation submodule, used to perform semantic segmentation on point cloud data in the registration dataset using a PointNet++ network to extract the three-dimensional boundaries and structural parameters of the ancient architectural artifact components; and a second semantic segmentation submodule, used to perform pixel-level classification on image data in the registration dataset using a SegNet network to identify the material type and surface detail features of the ancient architectural artifact components.
[0092] In some embodiments, the semantic segmentation module 230 further includes: a splicing submodule, configured to splice the three-dimensional boundary and the structural parameters, along with the material type and the surface detail features, to the same dimension based on the association relationships in the registration dataset through a fully connected layer of a cross-modal fusion network, to obtain a cross-modal feature vector; a calculation submodule, configured to calculate the mutual information of the cross-modal feature vector through the attention module of the cross-modal fusion network; and to perform weighted calculation on the feature pairs in the cross-modal feature vector based on the mutual information, to obtain a weighted cross-modal feature vector; and a judgment submodule, configured to judge the component category and material type of the weighted cross-modal feature vector and correct the semantic boundary through the joint classifier of the cross-modal fusion network, to obtain a first cross-modal semantic segmentation result.
[0093] In some embodiments, the Building Information Model (BIM) template library for ancient architectural and cultural relic components is constructed based on the Industrial Foundation Class (IFC) standard. The GIS base map contains topographic data and cadastral data in CityGML format. The first integration module 250 includes: a conversion submodule, used to convert the semantically enhanced BIM model into the CityGML format, retaining IFC semantic information through LOD hierarchical modeling based on the accuracy requirements of ancient architectural components and the AppSchema extension mechanism, to obtain a CityGML format BIM model; and an integration submodule, used to integrate the CityGML format BIM model and the GIS base map to obtain the BIM-GIS fusion model corresponding to the first operation time.
[0094] In some embodiments, the system 200 further includes: an acquisition module, configured to acquire meteorological data monitored by meteorological satellites and ground meteorological stations within a preset time period; and a determination module, configured to determine the first operating time of the UAV platform based on the meteorological data when the meteorological data meets preset weather conditions suitable for UAV operation, and to determine the adaptation scheme of the lidar, the RGB camera, the positioning system and the UAV platform, wherein the adaptation scheme includes the addition of protective equipment, switching of working modes and parameter optimization.
[0095] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0096] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0097] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0098] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected to achieve the purpose of the embodiments of this application according to actual needs. In addition, each functional unit in the embodiments of this application may be fully integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the integrated unit may be implemented in hardware or in the form of hardware plus software functional units.
[0099] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause the device automatic test line to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0100] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined to obtain new method embodiments or device embodiments without conflict.
[0101] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An intelligent surveying and mapping method based on the integration of AI and BIM, characterized in that, include: Using a drone platform equipped with lidar, RGB camera and positioning system, the ancient building and its surrounding environment are scanned from multiple angles during the first operation. Point cloud data, multi-view image data and position and attitude data are collected simultaneously. The point cloud data and multi-view image data are then associated with the position and attitude data at the corresponding time through their respective timestamps. The point cloud data and the multi-view image data are preprocessed respectively. Based on the position and pose data, cross-modal registration of the preprocessed point cloud data and multi-view image data is performed by feature matching and pose estimation to obtain a point cloud-image registration dataset. Semantic segmentation is performed on the point cloud data and image data in the registration dataset to obtain a first semantic segmentation result and a second semantic segmentation result. Based on the association relationship in the registration dataset, cross-modal semantic fusion is performed on the first semantic segmentation result and the second semantic segmentation result to obtain a first cross-modal semantic segmentation result. The first cross-modal semantic segmentation result includes multiple original point cloud components with category labels, and each original point cloud component is associated with a texture region in the corresponding image data through registration. Based on the category labels of each original point cloud component in the first cross-modal semantic segmentation result, the multiple original point cloud components are classified and aggregated to obtain segmented point cloud components, and the topological relationship between each segmented point cloud component is constructed. The assembly relationship between each segmented point cloud component is inferred through the topological relationship. Based on the assembly relationship, the template corresponding to the category of each segmented point cloud component is called from the pre-constructed BIM template library of ancient building and cultural relic components, and BIM model components are instantiated. The first cross-modal semantic segmentation result is structurally attached to the BIM model components to construct a semantically enhanced BIM model that combines geometric morphological information and semantic attribute information. The semantically enhanced BIM model and the geographic information system (GIS) base map containing geospatial benchmarks are integrated to obtain the BIM-GIS fusion model corresponding to the first operation time. The BIM-GIS fusion model is used for spatial query of components of the ancient building relics, comparison display of component status in multiple periods, and three-dimensional visualization.
2. The method according to claim 1, characterized in that, The method further includes: The steps of synchronous acquisition, cross-modal registration, and semantic segmentation in claim 1 are repeated at multiple second operation times according to a preset time interval to obtain the second cross-modal semantic segmentation results corresponding to the multiple second operation times. Based on the multiple second cross-modal semantic segmentation results, the semantically enhanced BIM model is updated. The updated semantically enhanced BIM model and GIS base map corresponding to the multiple second operation times are integrated to obtain the BIM-GIS fusion model corresponding to the multiple second operation times; Based on the BIM-GIS fusion model corresponding to the first operation time and the plurality of second operation times, a digital twin animation of the building evolution process is generated; The digital twin animation is used to display the spatiotemporal changes in the size of the components and the spatiotemporal changes in the material distribution of the ancient architectural artifacts.
3. The method according to claim 1, characterized in that, The preprocessing of the point cloud data and the multi-view image data includes: The point cloud data is subjected to denoising and hole-filling processing; The multi-view image data is repaired and enhanced using the posterior mean corrected stream (PMRF) algorithm.
4. The method according to claim 1, characterized in that, Based on the position and pose data, cross-modal registration is performed on the preprocessed point cloud data and preprocessed multi-view image data through feature matching and pose estimation to obtain a point cloud-image registration dataset, including: Key points are extracted from the preprocessed multi-view image data using Scale Invariant Feature Transform (SIFT) or Accelerated Robust Feature Rendering (SURF) algorithm, and corresponding visual features are extracted based on the key points. Geometric features are extracted from the preprocessed point cloud data. The visual features and geometric features are compared, and image-point cloud feature anchor pairs corresponding to the same physical location are matched based on the initial spatial mapping relationship established by the position and pose data. Based on the initial spatial mapping relationship, effective feature anchor pairs are selected from the image-point cloud feature anchor pairs. Starting with the position and pose data, and based on the initial spatial mapping relationship, the effective feature anchor point pairs are substituted into the perspective n-point algorithm (PnP) to determine the precise spatial pose of the RGB camera during shooting, as well as the precise spatial mapping relationship between image pixels and point cloud 3D coordinates. The preprocessed point cloud data, preprocessed multi-view image data, precise spatial pose, and precise spatial mapping relationship are structurally integrated to obtain a point cloud-image registration dataset.
5. The method according to claim 1, characterized in that, The step of performing semantic segmentation on the point cloud data and image data in the registration dataset to obtain a first semantic segmentation result and a second semantic segmentation result includes: The PointNet++ network is used to perform semantic segmentation on the point cloud data in the registration dataset in order to extract the three-dimensional boundaries and structural parameters of the ancient architectural artifact components; The image data in the registration dataset is classified at the pixel level using the SegNet network to identify the material type and surface detail features of the ancient architectural artifact components.
6. The method according to claim 5, characterized in that, The step of performing cross-modal semantic fusion on the first semantic segmentation result and the second semantic segmentation result based on the association relationship in the registration dataset to obtain the first cross-modal semantic segmentation result includes: Based on the association relationships in the registration dataset, the three-dimensional boundary and the structural parameters are mapped to the same dimension as the material type and the surface detail features through the fully connected layer of the cross-modal fusion network and then concatenated to obtain a cross-modal feature vector; The mutual information of the cross-modal feature vectors is calculated through the attention module of the cross-modal fusion network; based on the mutual information, the feature pairs in the cross-modal feature vectors are weighted to obtain the weighted cross-modal feature vectors. The joint classifier of the cross-modal fusion network is used to determine the component category and material type of the weighted cross-modal feature vector, as well as to correct the semantic boundary, to obtain the first cross-modal semantic segmentation result.
7. The method according to claim 1, characterized in that, The BIM template library for the ancient architectural and cultural relic components is built based on the Industrial Foundation Class IFC standard. The GIS base map contains topographic and cadastral data and is in CityGML format. The semantically enhanced BIM model and the GIS base map containing geospatial benchmarks are integrated to obtain a BIM-GIS fusion model, including: The semantically enhanced BIM model is converted into the CityGML format, and the IFC semantic information is preserved through the LOD hierarchical modeling based on the accuracy requirements of ancient building components and the AppSchema extension mechanism of the application mode, resulting in a BIM model in CityGML format. The CityGML format BIM model and GIS base map are integrated to obtain the BIM-GIS fusion model corresponding to the first operation time.
8. The method according to claim 1, characterized in that, The method further includes: Acquire meteorological data monitored by meteorological satellites and ground meteorological stations within a preset time period; When the meteorological data meets the preset weather conditions suitable for UAV operation, the first operating time of the UAV platform is determined based on the meteorological data, and the adaptation scheme of the lidar, the RGB camera, the positioning system and the UAV platform is determined. The adaptation scheme includes the addition of protective equipment, switching of working modes and parameter optimization.
9. An intelligent surveying and mapping system based on the integration of AI and BIM, characterized in that, include: The data acquisition module is used to perform multi-angle flight scanning of the ancient building and its surrounding environment at the first moment of operation using a drone platform equipped with lidar, RGB camera and positioning system. Simultaneously, it collects point cloud data, multi-view image data and position and attitude data, and establishes a correlation between the point cloud data and the multi-view image data and the position and attitude data at the corresponding moment through their respective timestamps. The registration module is used to preprocess the point cloud data and the multi-view image data respectively, and based on the position and pose data, to perform cross-modal registration of the preprocessed point cloud data and multi-view image data through feature matching and pose estimation to obtain a point cloud-image registration dataset. The semantic segmentation module is used to perform semantic segmentation on the point cloud data and image data in the registration dataset to obtain a first semantic segmentation result and a second semantic segmentation result. Based on the association relationship in the registration dataset, the first semantic segmentation result and the second semantic segmentation result are subjected to cross-modal semantic fusion to obtain a first cross-modal semantic segmentation result. The first cross-modal semantic segmentation result includes multiple original point cloud components with category labels, and each original point cloud component is associated with a texture region in the corresponding image data through registration. The first generation module is used to classify and aggregate the multiple original point cloud components based on the category labels of each original point cloud component in the first cross-modal semantic segmentation result, obtain segmented point cloud components, construct the topological relationship between each segmented point cloud component, infer the assembly relationship between each segmented point cloud component through the topological relationship, and, based on the assembly relationship, call the template corresponding to the category of each segmented point cloud component from the pre-constructed building information model (BIM) template library of ancient building and cultural relic components, instantiate and generate BIM model components, and structurally attach the first cross-modal semantic segmentation result to the BIM model components to construct a semantically enhanced BIM model that combines geometric morphological information and semantic attribute information. The first integration module is used to integrate the semantically enhanced BIM model and the geographic information system (GIS) base map containing the geospatial benchmark to obtain the BIM-GIS fusion model corresponding to the first operation time. The BIM-GIS fusion model is used to perform spatial query of the components of the ancient building relics, comparative display of the component status in multiple periods, and three-dimensional visualization.
10. The system according to claim 9, characterized in that, The system also includes: The update module is used to repeat the steps of synchronous acquisition, cross-modal registration, and semantic segmentation in claim 9 at multiple second operation times according to a preset time interval, to obtain the second cross-modal semantic segmentation results corresponding to the multiple second operation times, and to update the semantically enhanced BIM model based on the multiple second cross-modal semantic segmentation results; The second integration module is used to integrate the updated semantically enhanced BIM models and GIS base maps corresponding to the multiple second operation times to obtain the BIM-GIS fusion model corresponding to the multiple second operation times. The second generation module is used to generate a digital twin animation of the building evolution process based on the BIM-GIS fusion model corresponding to the first operation time and the plurality of second operation times; wherein, the digital twin animation is used to display the spatiotemporal changes in the component size and material distribution of the ancient building artifact.
Citation Information
Patent Citations
Real scene three-dimensional modeling method and system fusing laser point cloud and image
CN120147563A
Intelligent geometric reasoning and semantic understanding method based on three-dimensional large language model
CN120542438A
Multi-precision three-dimensional surveying and mapping data fusion method based on dynamic modeling
CN120563982A
Building structure intelligent design method and system based on 3D modeling, electronic equipment and storage medium
CN120579430A
Single building three-dimensional reconstruction method based on point cloud semantic segmentation and structure fitting
WO2024077812A1
Cited By
Cultural relic digital surveying and mapping method based on multi-scale point cloud fusion
CN121582320A
A digital mapping method for cultural relics based on multi-scale point cloud fusion
CN121582320B
Construction cost visualization dynamic analysis system based on multi-source heterogeneous data recognition
CN121599709A
Urban three-dimensional terrain rapid modeling method and system based on AI point cloud semantic segmentation
CN121788748A
Construction site component level difference detection method, device, equipment and medium
CN121936038A