A BIM-based intelligent surveying and mapping method and system

By simultaneously collecting point cloud and image data using drones, and combining cross-modal registration and semantic fusion, a BIM model of ancient buildings with both geometric shape and semantic attributes was constructed. This solved the problem of insufficient multi-source data fusion and enabled high-precision digital management and dynamic visualization of ancient buildings.

CN120852601BActive Publication Date: 2025-12-09XIAN UNVERSITY OF ARTS & SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511360915.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-12-09
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

In existing technologies, the spatial registration error caused by the independent acquisition of point cloud, image and positioning data is relatively high, which makes it impossible to form a unified semantic space and the multi-source data fusion capability is insufficient, making it difficult to meet the needs of high-precision digital protection and long-term inheritance of ancient architectural relics.

Method used

By using a drone platform equipped with LiDAR, RGB camera and positioning system, point cloud data and multi-view image data are collected simultaneously. Combined with cross-modal registration and semantic fusion, a semantically enhanced BIM model with both geometric shape and semantic attributes is constructed and integrated with GIS base map to achieve deep fusion of multi-source data.

Benefits of technology

It breaks through the information limitations of single data, constructs a more complete digital model of ancient buildings, realizes full-scale management from micro-components to macro-space, provides precise digital carriers and dynamic visualization tools, and meets the multi-dimensional needs of ancient building protection and research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852601B_ABST
    Figure CN120852601B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses an intelligent surveying and mapping method and system based on AI and BIM fusion, the method comprises the following steps: an unmanned aerial vehicle platform carrying a laser radar, an RGB camera and a positioning system is used to fly and scan an ancient building and its surroundings from multiple angles, point cloud data, multi-view image data and position and attitude data are synchronously collected, and the three are associated through a time stamp; after the point cloud data and the multi-view image data are preprocessed, cross-modal registration is completed through feature matching and pose estimation in combination with the position and attitude data, and a registration data set is obtained; the point cloud data and the image data in the registration data set are respectively subjected to semantic segmentation, and the semantic segmentation results are fused based on the association relationship; original point clouds are classified and aggregated according to category labels, a topological relationship is inferred, and an assembly relationship is constructed, a corresponding BIM template is called based on the assembly relationship, a model component is instantiated, the segmentation result is connected, and a semantic enhanced BIM model is generated; and the BIM model and a GIS base map are integrated to form a fusion model, so that the ancient building is surveyed and mapped.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of cultural relic protection and digitization, and relates to but is not limited to an intelligent surveying and mapping method and system based on BIM. BACKGROUND

[0002] With the continuous progress of science and technology, the protection of cultural relics and ancient buildings has ushered in new opportunities and challenges. Traditional surveying and mapping technology often has problems such as low efficiency, insufficient accuracy, and difficulty in obtaining comprehensive information when facing complex and delicate ancient buildings and cultural relics, and is difficult to meet the urgent needs of modern society for high-precision digitization protection and long-term inheritance of cultural relics and ancient buildings. Under this background, the rapid development of unmanned aerial vehicle technology, artificial intelligence (AI) and building information modeling (BIM) and other frontier technologies has brought innovative solutions to the field of ancient building and cultural relic surveying and mapping.

[0003] Unmanned aerial vehicle technology has great potential in the collection of external form data of cultural relics and ancient buildings due to its high flexibility, low cost, and ability to quickly obtain high-resolution image data of large areas. By carrying advanced sensors such as laser radar (LiDAR), high-definition RGB camera and positioning system, unmanned aerial vehicles can comprehensively scan ancient buildings and cultural relics and their surrounding environment from multiple angles, obtaining high-precision point cloud data, multi-view image data and position and attitude (POS) information, providing rich and accurate basic data sources for subsequent three-dimensional modeling and analysis.

[0004] The rise of artificial intelligence (AI) technology provides powerful tools for efficient processing and deep mining of massive surveying and mapping data. Building information modeling (BIM) as a digital building information integration technology can integrate and manage multi-dimensional data such as geometric form, physical properties and historical information of ancient buildings and cultural relics. The BIM model based on the international standard IFC can not only visually display the three-dimensional structure of ancient buildings, but also provide standardized data support for the protection and repair of cultural relics, structural analysis, digital display and other work, promote cooperation between various professional fields, and improve the scientificity and systematicness of cultural relic protection work.

[0005] However, point cloud, image and positioning data are collected independently, with high spatial registration error, resulting in misalignment of geometric and texture information, inability to form a unified semantic space, and insufficient multi-source data fusion capability. SUMMARY

[0006] Therefore, the embodiments of the present application provide an intelligent surveying and mapping method based on BIM, which at least solves the problem of insufficient multi-source data fusion capability.

[0007] The technical scheme of the embodiments of the present application is as follows:

[0008] In a first aspect, the embodiments of the present application provide a BIM-based intelligent surveying and mapping method, which comprises:

[0009] By means of a UAV platform equipped with a laser radar, an RGB camera and a positioning system, a multi-angle flight scanning is performed on an ancient building and its surrounding environment at a first operation time, and point cloud data, multi-view image data and position and attitude data are synchronously collected, so that the point cloud data and the multi-view image data are associated with the position and attitude data at the corresponding time through respective time stamps;

[0010] The point cloud data and the multi-view image data are respectively preprocessed, and the preprocessed point cloud data and multi-view image data are cross-modal registered through feature matching and pose estimation based on the position and attitude data, so as to obtain a point cloud-image registration data set;

[0011] The point cloud data and the image data in the registration data set are respectively subjected to semantic segmentation, so as to obtain a first semantic segmentation result and a second semantic segmentation result, and the first semantic segmentation result and the second semantic segmentation result are cross-modal semantic fused based on the association relationship in the registration data set, so as to obtain a first cross-modal semantic segmentation result, wherein the first cross-modal semantic segmentation result comprises a plurality of original point cloud components with class labels, and each original point cloud component corresponds to a texture region in the image data through registration association;

[0012] Based on the class labels of each original point cloud component in the first cross-modal semantic segmentation result, the plurality of original point cloud components are classified and aggregated, so as to obtain segmented point cloud components, and a topological relationship between the segmented point cloud components is constructed, an assembly relationship between the segmented point cloud components is inferred through the topological relationship, based on the assembly relationship, a template corresponding to the class of each segmented point cloud component is called from a building information model (BIM) template library of ancient building components pre-constructed, a BIM model component is instantiated, the first cross-modal semantic segmentation result is structured and connected to the BIM model component, and a semantic enhanced BIM model with geometric shape information and semantic attribute information is constructed;

[0013] The semantic enhanced BIM model and a geographic information system (GIS) base map containing a geographic space reference are integrated, so as to obtain a BIM-GIS fusion model corresponding to the first operation time, and the BIM-GIS fusion model is used for component spatial query, multi-period component state comparison display and three-dimensional visualization display of the ancient building.

[0014] In a second aspect, the embodiments of the present application provide a BIM-based intelligent surveying and mapping system, which comprises:

[0015] The collection module is configured to, by means of a UAV platform carrying a laser radar, an RGB camera and a positioning system, perform multi-angle flight scanning on an ancient building cultural relic body and a surrounding environment thereof at a first operation time, and synchronously collect point cloud data, multi-view image data and position and posture data, so that the point cloud data and the multi-view image data are associated with position and posture data at a corresponding time through respective time stamps.

[0016] The registration module is configured to respectively pre-process the point cloud data and the multi-view image data, perform cross-modal registration on the pre-processed point cloud data and multi-view image data based on the position and posture data through feature matching and pose estimation, and obtain a point cloud-image registration data set.

[0017] The semantic segmentation module is configured to respectively perform semantic segmentation on the point cloud data and the image data in the registration data set, obtain a first semantic segmentation result and a second semantic segmentation result, perform cross-modal semantic fusion on the first semantic segmentation result and the second semantic segmentation result based on an association relationship in the registration data set, and obtain a first cross-modal semantic segmentation result, wherein the first cross-modal semantic segmentation result includes a plurality of original point cloud components with class labels, and each original point cloud component corresponds to a texture region in the image data through registration association.

[0018] The first generation module is configured to classify and aggregate the plurality of original point cloud components based on class labels of each original point cloud component in the first cross-modal semantic segmentation result, obtain segmented point cloud components, construct a topological relationship between the segmented point cloud components, infer an assembly relationship between the segmented point cloud components through the topological relationship, call a template corresponding to a class of each segmented point cloud component from a building information model (BIM) template library of ancient building cultural relic components pre-constructed based on the assembly relationship, instantiate a BIM model component, structure the first cross-modal semantic segmentation result and connect the first cross-modal semantic segmentation result to the BIM model component to construct a semantic enhanced BIM model having geometric shape information and semantic attribute information.

[0019] The first integration module is configured to integrate the semantic enhanced BIM model and a geographic information system (GIS) base map containing a geographic space reference to obtain a BIM-GIS fusion model corresponding to the first operation time, and the BIM-GIS fusion model is used for component spatial query, multi-period component state comparison display and three-dimensional visualization display of the ancient building cultural relic.

[0020] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:

[0021] Through the synchronous acquisition of point cloud, image and position and attitude data by the unmanned aerial vehicle, combined with cross-modal registration and semantic fusion, the information limitation of single data (point cloud / image) is broken through, the geometric shape and semantic attribute are retained, a more complete digital model of ancient buildings is constructed, and the deep fusion of multi-source data is realized; through classification aggregation, topological relationship reasoning and BIM template instantiation, the original surveying and mapping data is converted into a structured semantic enhanced BIM model, which contains component geometric details and is associated with attributes such as material and category, providing a precise digital carrier for the protection and research of ancient buildings; through the integration of BIM and GIS base map, the spatial query, multi-period comparison and three-dimensional visualization of ancient building components are realized, meeting the full-scale management needs from micro components to macro space. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.

[0023] Figure 1 A flowchart of an intelligent surveying and mapping method based on BIM provided by the embodiments of the present application;

[0024] Figure 2 A component diagram of an intelligent surveying and mapping system based on BIM provided by the embodiments of the present application. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. The following embodiments are used to illustrate the present application, but not to limit the scope of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0026] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0027] It should be noted that the terms "first", "second", "third" involved in the embodiments of the present application are only to distinguish similar objects, and do not represent a specific order of the objects. Understandably, "first", "second", "third" can be interchanged in specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0028] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of the present application belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have meanings consistent with those in the context of the prior art, and should not be interpreted to have idealized or overly formal meanings unless specifically defined as such herein.

[0029] The embodiments of the present application provide a BIM-based intelligent mapping method, applied to an electronic device. The electronic device includes but is not limited to a mobile phone, a notebook computer, a tablet computer, a palm Internet device, a multimedia device, a streaming media device, a mobile Internet device, a wearable device or other types of electronic devices. The functions implemented by the method can be realized by calling program codes in the processor of the electronic device, and of course the program codes can be saved in the computer storage medium. Therefore, the electronic device at least includes a processor and a storage medium. The processor can be used to process the BIM-based intelligent mapping process, and the memory can be used to store the data required in the BIM-based intelligent mapping process and the generated data.

[0030] Figure 1 A flowchart of a BIM-based intelligent mapping method provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the method at least includes the following steps: Figure 1

[0031] In step S110, through the unmanned aerial vehicle platform equipped with laser radar, RGB camera and positioning system, the ancient building and its surrounding environment are scanned at multiple angles at the first operation time, and point cloud data, multi-view image data and position and attitude data are synchronously collected, so that the point cloud data, the multi-view image data and the position and attitude data of the corresponding time are associated through their respective time stamps.

[0032] ​The ancient building and cultural relic area can be surrounded and overflown from multiple angles by a UAV equipped with a laser radar (LiDAR), a high-definition RGB camera, and an RTK (Real-Time Kinematic) / PPK (Post-Processed Kinematic) positioning system, to collect high-density three-dimensional point cloud data (.las / .ply format), multi-view color image data (.jpg / .png format), and position and orientation data (Position and Orientation System, POS) including GPS coordinates and IMU (Inertial Measurement Unit) calculated attitude angles. The position and orientation information is the "image shooting state" and "point cloud collection state" recorded by the UAV synchronously when collecting point cloud data and multi-view image data, which can exactly provide two key "initial constraint conditions" for matching.

[0033] To ensure data integrity, the heading overlap rate during flight can be not less than 70%, and the lateral overlap rate can be not less than 60%. Such settings aim to ensure the integrity and redundancy of the collected data, to provide sufficient data support for subsequent point cloud stitching, image matching and three-dimensional reconstruction, and to improve the accuracy and reliability of the surveying and mapping results. Experiments show that such overlap rate settings can effectively reduce data holes and improve the integrity of the point cloud model, significantly reducing the workload of subsequent data repair and completion in practical applications; and a reflight strategy is implemented for areas with severe occlusion or complex structure; all collected data are timestamped to achieve accurate synchronization and registration in subsequent processes.

[0034] In step S120, the point cloud data and the multi-view image data are preprocessed respectively, and based on the position and orientation data, cross-modal registration is performed on the preprocessed point cloud data and multi-view image data through feature matching and pose estimation to obtain a point cloud-image registration data set.

[0035] In the cross-modal registration, the point cloud data and the image data are accurately aligned in space, i.e., it is determined which pixel point in the image data corresponds to a three-dimensional point in the point cloud data.

[0036] In step S130, the point cloud data and the image data in the registration data set are respectively subjected to semantic segmentation to obtain first and second semantic segmentation results, and cross-modal semantic fusion is performed on the first and second semantic segmentation results based on the association relationship in the registration data set to obtain a first cross-modal semantic segmentation result. The first cross-modal semantic segmentation result includes a plurality of original point cloud components with class labels, and each original point cloud component corresponds to a texture region in the image data through registration association.

[0037] The essence of semantic segmentation is to label each element in the data with a category. Since point cloud data and image data are different types of data, semantic segmentation needs to be performed separately. The association relationship in the registration dataset refers to the accurate correspondence between the three-dimensional points of the point cloud and the pixels of the image. The segmentation of point cloud is good at capturing the boundaries of three-dimensional mechanisms and components, while the segmentation of image is good at identifying surface texture and material type. The purpose of cross-modal semantic fusion is to combine the segmentation results of point cloud data and image data to make up for the defects of single data.

[0038] Through the point-pixel correspondence, the two semantic segmentation results can be verified and supplemented each other.

[0039] In step S140, based on the category labels of each original point cloud component in the first cross-modal semantic segmentation result, the plurality of original point cloud components are classified and aggregated to obtain segmented point cloud components, and a topological relationship between the segmented point cloud components is constructed. The assembly relationship between the segmented point cloud components is inferred through the topological relationship. Based on the assembly relationship, a template corresponding to the category of each segmented point cloud component is called from a building information model (BIM) template library of ancient building cultural relic components that is pre-constructed, and a BIM model component is instantiated to generate. The first cross-modal semantic segmentation result is structured and connected to the BIM model component to construct a semantic enhanced BIM model that has both geometric shape information and semantic attribute information.

[0040] Based on the first cross-modal semantic segmentation result (i.e., cross-modal semantic point cloud and feature data), the BIM model construction of ancient building cultural relic components can realize the conversion from point cloud data to semantic enhanced BIM model through three core links. First, an IFC (Industry Foundation Classes) standard-based BIM (Building Information Modeling) template library of ancient building cultural relic components is constructed. According to the IFC international standard and combined with the characteristics of ancient building components, the data structure and attribute template of components such as corbel bracket, beam, column base, etc. are defined. For example, the "column" template is pre-set with geometric attributes such as height, radius, perpendicularity, and material attributes such as wood / stone type. At the same time, parameterized design is adopted, and the arch length, arch thickness, and level number of the corbel bracket are set as variable parameters to adapt to the rapid modeling needs of different specifications of components. Second, point cloud component classification and topological relationship are performed. Topological relationship analysis can realize automatic reasoning of assembly relationship through four steps of "component spatial feature quantization-proximity detection-rule base reasoning-relation verification". First, geometric feature extraction is performed on the segmented point cloud components (such as "column", "beam", and "corbel bracket"). The minimum bounding box (AABB, Axis-Aligned Bounding Box) of each component is calculated to obtain the center point coordinates , dimensions (length / width / height) and principal axis direction, and key points on the surface of the components are extracted as feature points for relationship reasoning; then, the spatial distance between components is calculated based on the distance between the center points, a threshold (e.g. ≤0.5 meters) is set to screen potential associated components, and the intersection volume ratio of the two components is calculated through point cloud collision detection, and if the ratio is >5%, it is determined as "physical contact" as strong association evidence; then, a rule library of topological relationship of ancient building components is constructed, and reasoning is performed in combination with the component type and spatial pose, such as in the support relationship, if the two end points of component A (beam) are projected to fall within the top bounding box of component B (column), and the angle between the principal axis of the beam and the horizontal plane is <10°, and the angle between the principal axis of the column and the plumb line is <5°, then it is reasoned as "column supporting beam", and in the connection relationship, if the "arch" part of the corbel component and the edge distance of the beam are <0.1 meters (unit: m), and the normal direction of the corbel is perpendicular to the surface of the beam, and the elevation difference between the two center points is <0.2m, then it is reasoned as "corbel connecting beam", and the rule library can cover 20+ common assembly type reasoning rules of ancient buildings; finally, the physical rationality of the preliminary reasoning relationship is checked (such as "support relationship" needs to satisfy the projection area of the upper component ≥70% falling within the top range of the lower component), and graph theory is used to remove redundancy, the components are taken as nodes, the relationship is taken as edges to construct a topological graph, the weakly associated edges with a weight <0.6 are deleted, and the strongly associated edges are retained to form the final assembly relationship network, and the reasoning accuracy is above 95%. ), clustering algorithms such as DBSCAN and K-Means are used, based on the geometric features such as three-dimensional coordinate distribution and curvature change of the point cloud components, to aggregate them into categories such as "column", "beam", "corbel", etc., and the classification accuracy can be above 95%, and then a topological relationship network is constructed based on graph theory algorithms, the spatial position and contact boundary between components are analyzed, and assembly relationships such as "beam supporting column" and "corbel connecting beam" are reasoned, to ensure the structural integrity of the BIM model; finally, semantic attribute hanging and BIM model generation are completed, the coordinates of the point cloud and the BIM model are mapped, the extracted material properties such as bricks and wood, and structural parameters such as column height and corbel overhang length are accurately hung to the classified components, written into the corresponding attribute fields of the IFC model, and then the model is generated by calling the IFC standard interface, and through lightweight processing such as triangular mesh simplification and texture compression, and geometric precision verification with an error threshold of ±2mm (unit: mm), an IFC format BIM model containing semantic information such as geometric shape, material and structure is finally generated, providing standardized data support for digital archiving, protection planning and repair simulation of ancient buildings.

[0041] In step S150, the semantic enhanced BIM model and the geographic information system GIS base map containing geographic reference are integrated to obtain a BIM-GIS fusion model corresponding to the first operation time, and the BIM-GIS fusion model is used for component spatial query, multi-period component state comparison and display, and three-dimensional visualization display of the ancient building cultural relics.

[0042] By integrating the fine component-level BIM model with the macro geographical GIS base map, a comprehensive model with both micro details and macro locations can be constructed, realizing the full-scale management, query and display of ancient buildings from individual components to regional environment, and providing more comprehensive digital support for the protection, research and utilization of cultural relics.

[0043] In the above embodiments, by synchronously collecting point cloud, image and position and attitude data by the unmanned aerial vehicle, combining cross-modal registration and semantic fusion, the information limitations of single data (point cloud / image) are broken through, while the geometric shape and semantic attributes are retained, a more complete digital model of ancient buildings is constructed, and deep fusion of multi-source data is realized; by classification aggregation, topological relationship reasoning and BIM template instantiation, the original surveying and mapping data is converted into a structured semantic enhanced BIM model, which contains component geometric details and is associated with attributes such as material and category, providing a precise digital carrier for the protection and research of ancient buildings; by integrating BIM and GIS base map, spatial query, multi-period comparison and three-dimensional visualization of ancient building components are realized, meeting the full-scale management requirements from micro components to macro space.

[0044] In some embodiments, the method further comprises:

[0045] Step S160, repeating the steps of synchronous collection, cross-modal registration and semantic segmentation at multiple second operation times according to a preset time interval, obtaining second cross-modal semantic segmentation results corresponding to the multiple second operation times, and updating the semantic enhanced BIM model based on the multiple second cross-modal semantic segmentation results;

[0046] Step S170, integrating the updated semantic enhanced BIM model corresponding to the multiple second operation times and the GIS base map to obtain a BIM-GIS fusion model corresponding to the multiple second operation times;

[0047] Step S180, generating a digital twin animation of the building evolution process based on the BIM-GIS fusion models corresponding to the first operation time and the multiple second operation times;

[0048] The digital twin animation is used to show the spatial and temporal changes in the size and material distribution of the components of the ancient building.

[0049] In which, the process of steps S110 to S130 can be repeated periodically using the unmanned aerial vehicle to obtain updated point cloud and image data of the ancient building area, update the geometric changes (such as component displacement and deformation) and material information of the BIM model through point cloud registration and difference analysis, generate BIM-GIS fusion models of different periods, generate a digital twin animation of the building evolution process based on the historical model sequence, and visually display the spatial and temporal changes in the size and material distribution of the components. Specifically as follows:

[0050] In the stage of steps S160 to S180, the wooden building collects data once every quarter, and the stone building collects data once every year, and the cavity rate is controlled to be less than 1%, and the light conditions are ensured to be consistent to ensure the data quality; through cross-period point cloud registration (ISS+FPFH is used for coarse registration, rotation error is less than or equal to 5°, translation error is less than or equal to 5 cm, ICP is used for fine registration with the base as a reference, RMSE is less than or equal to 2 mm, and coincidence rate is greater than 95%), and 1 cm voxel grid difference analysis, the component displacement (column tilt is greater than 0.5°), surface deformation (relief height reduction is greater than 1 mm), volume change (corbel reduction rate is greater than 5%) and other indicators are quantified; semantic component level dynamic identification and parameterized updating (geometric parameter accuracy is less than or equal to 1 mm, updating color drawing fading, wood texture and other material properties) are adopted, and version management and incremental log recording are performed by combining Git LFS with time stamp naming; based on the key frames of the models in each period, a space-time index is constructed, a digital twin animation is rendered through color mapping difference and texture mixing technology, and time axis control and dynamic sectioning interaction are supported.

[0051] In the above embodiment, by periodically repeating collection and model updating, the size and material changes of components in different periods are captured, and the problem that a static model cannot reflect the evolution in the time dimension is solved. Based on the multi-period BIM-GIS fusion model, evolution animation is generated, and the space-time change process of the ancient building is intuitively displayed, which provides a dynamic visualization tool for cultural relic protection and repair (such as monitoring disease development) and historical research (such as building change analysis).

[0052] In some embodiments, the step S120 of “respectively pre-processing the point cloud data and the multi-view image data” includes the following steps:

[0053] Step S1201, denoising and hole completion processing are performed on the point cloud data.

[0054] Firstly, coarse registration of point clouds can be achieved by FPFH (Fast Point Feature Histograms) feature matching and ICP (Iterative Closest Point) algorithm, so that multi-view point clouds are preliminarily aligned; subsequently, statistical filter or radius filtering method is used to remove noise points and outliers, improving the overall cleanliness of the point cloud; for the region with holes or missing, Poisson surface reconstruction algorithm and other algorithms are used to fill in, to ensure the integrity of the geometric structure; finally, the surface continuity is optimized by RANSAC (Random Sample Consensus) plane fitting algorithm, RANSAC randomly selects a subset of point clouds to fit a plane model, and according to the error between the model and the remaining points, the effective points are screened, and the optimal plane is obtained by iteration, so as to improve the smoothness and accuracy of the point cloud surface.

[0055] In step S1202, the multi-view image data is repaired and enhanced by using a posterior mean rectified flow (PMRF) algorithm.

[0056] Meanwhile, the multi-view RGB image can be repaired and enhanced by using a PMRF (Posterior-Mean Rectified Flow) algorithm. The construction and training process of the PMRF algorithm is as follows:

[0057] In the construction process, the PMRF algorithm adopts a flow model architecture, learns the posterior distribution of image inpainting through a reversible transformation network, solves the ordinary differential equation (ODE) based on the optimal transport theory to realize texture generation, and realizes repair and enhancement through multi-stage processing for the problems of texture defects and quality degradation of ancient building cultural relics images caused by occlusion and aging: first, a binary mask of the marked damaged area is automatically generated based on computer vision algorithm, then UNet encoder is used to extract global semantic features (such as architectural style, component type) and local texture details (such as brick carving lines, color painting patterns, wood grain texture) of the image, the posterior mean estimation of the missing area is carried out through the probability model to determine the structure tendency of the repair content; then, the evolution path of the image from noise to clarity is modeled by ODE, a plurality of repair candidate results are generated by iterative sampling, and the output that best meets the real texture of the ancient building cultural relics is selected by combining the perceptual loss function; finally, the repair artifacts are eliminated by boundary fusion and color correction, the details such as brick carving and color painting are restored, the image contrast and color consistency are improved, and high-quality visual input is provided for subsequent cross-modal registration and semantic analysis.

[0058] The texture generation is realized by solving ODE based on optimal transport theory. In simple terms, this theory can find an optimal way to convert a noise image into a clear image that conforms to the real texture of ancient buildings at the mathematical level, just like finding the most reasonable path in a complex space. The repaired image texture is more natural, realistic, and in line with the original style of ancient buildings.

[0059] During the training process, the PMRF algorithm training process focuses on improving image quality and restoring real details in the context of ancient building artifact image restoration. At the beginning of the training, a dataset containing a large number of original clear images of ancient building artifacts and their corresponding degraded versions (simulating actual damage conditions such as occlusion, aging, and noise interference) is constructed. These images cover a wide range of ancient building artifact types, such as ancient building component surface texture and local mural patterns. The first stage is posterior mean prediction training, aiming to minimize the mean square error (MSE). The degraded image is input into the model, which uses the powerful feature learning ability of deep neural networks to extract and analyze features from low-level pixel features to high-level semantic features. By continuously adjusting network parameters, the model tries to generate a posterior mean prediction result that is as close as possible to the original ancient building artifact image in numerical value, thereby reducing image distortion and making the preliminary restored image numerically consistent with the real ancient building artifact image. However, at this stage, the image may have problems such as blurred details and poor visual effects. In the second stage of rectification flow model training, the rectification flow model learns the mapping relationship from the posterior mean prediction result to the high-quality real ancient building artifact image distribution. This process is based on the optimal transport theory and is realized by solving ODE. During the training process, the model iteratively optimizes and gradually "transports" and converts the posterior mean prediction image into a high-quality image that is visually more realistic and conforms to the real texture and feature characteristics of ancient building artifacts, ensuring that the repaired image not only has low distortion but also has high perceptual quality, consistent with human perception of the original appearance of ancient building artifacts. Throughout the training process, the PyTorch deep learning framework is used, combined with Lightning to simplify the training process, and natten is used to implement specific network architectures. After multiple rounds of training, the PMRF algorithm can accurately repair and enhance the complex degradation of ancient building artifact images, laying a solid data foundation for subsequent cross-modal registration and ancient building artifact feature analysis.

[0060] To address the data defects caused by adverse weather, targeted algorithms are used for preprocessing: In rainy and snowy weather, point cloud denoising is added to the original radius filtering based on "adaptive density filtering". By analyzing the local point cloud density distribution (normal area density > 150 points / m 2Automatic identification of low-density noise area caused by rain and snow, K-neighborhood (K=10) algorithm is used to remove outliers, so that the point cloud cleanliness is restored to more than 95%, and the image restoration is realized through the "rain and snow mask automatic generation module" added by the PMRF algorithm. Based on the brightness threshold (snowflake pixel brightness > 240) and edge detection, the rain and snow area is identified, and the local texture continuity is strengthened during restoration (such as the eave area constraint texture direction consistent with the adjacent pixel), so that the signal-to-noise ratio of the restored image is improved by 25%; In the strong light / overexposure processing, a "dynamic range compression" step is added before PMRF restoration of the overexposed image, which separates the highlight and shadow areas through the Retinex algorithm. The brightness of the highlight area is reduced (the gray value is compressed to within 200), and the contrast of the shadow area is improved (the texture details are enhanced by 30%), ensuring that the key texture can be identified; In the overcast / mist processing, the point cloud completion introduces "semantic guided completion" based on the Poisson reconstruction, which constrains the geometric form of the hollow area based on the geometric features of the components (such as the angle of the tower eave) identified by PointNet++. The completion accuracy is improved to ±3mm (unit: mm), and the image enhancement is realized by adding a "low-light sample set" (500 overcast images) in the SegNet network training, combining histogram equalization to expand the gray scale range and bilateral filtering to suppress noise, so that the material classification accuracy is restored from 75% to more than 88%.

[0061] In the above embodiment, the point cloud denoising and completion reduces the interference of holes and noise, and the PMRF algorithm restoration enhances the image, ensuring the accuracy of subsequent registration and segmentation, and laying a high-quality data foundation for the entire process. The preprocessing step enhances the robustness of the data to solve the problem that the point cloud of ancient buildings may have holes due to occlusion, and the image quality may be degraded due to complex environment (such as tree shade and backlight).

[0062] In some embodiments, the step S120 "based on the position and attitude data, the preprocessed point cloud data and multi-view image data are cross-modal registered through feature matching and pose estimation, and a point cloud-image registration data set is obtained" includes the following steps:

[0063] In step S1203, the scale-invariant feature transform (SIFT) or speeded up robust features (SURF) algorithm is used to extract key points from the preprocessed multi-view image data, and corresponding visual features are extracted based on the key points; geometric features are extracted from the preprocessed point cloud data; the visual features and the geometric features are compared, and the initial spatial mapping relationship established based on the position and attitude data is combined to match the image-point cloud feature anchor point pairs corresponding to the same physical position; based on the initial spatial mapping relationship, the effective feature anchor point pairs are selected from the image-point cloud feature anchor point pairs.

[0064] Step S1204, based on the initial spatial mapping relationship, the effective feature anchor point pair is substituted into the perspective n-point algorithm (PnP algorithm) starting from the position and attitude data, to determine the accurate spatial pose when the RGB camera is shooting, and the accurate spatial mapping relationship between the image pixel and the point cloud three-dimensional coordinate.

[0065] Step S1205, the pre-processed point cloud data, the pre-processed multi-view image data, the accurate spatial pose and the accurate spatial mapping relationship are structured and integrated to obtain a point cloud-image registration data set.

[0066] After the repair is completed, the PMRF enhanced image and the point cloud data are cross-modal registered. This process extracts key point features in the image through SIFT (Scale-Invariant Feature Transform) / SURF (Speeded Up Robust Features), and combines POS pose information (i.e., position and attitude data) to estimate the camera pose using the PnP (Perspective-n-Point) algorithm, thereby accurately mapping the image to the three-dimensional point cloud space, and finally generating an RGB point cloud (Colorized Point Cloud), which provides high-precision and high-quality input data for subsequent semantic segmentation and BIM modeling.

[0067] What is needed is that the point cloud data is spatially registered (coordinate aligned) with the image data after denoising and completion, but the image texture is not fused, and the image data is spatially mapped with the point cloud data after repair and enhancement, but is not merged with the point cloud.

[0068] In the above embodiment, the SIFT / SURF is used to extract visual features, the geometric feature matching is combined, and the PnP algorithm is used to optimize the pose, so as to solve the spatial alignment problem of cross-modal data (three-dimensional point cloud and two-dimensional image), and ensure that the "point cloud-image" features are one-to-one corresponding. Based on the position and attitude data, an initial mapping is established, and then the effective anchor point pair is screened and the accurate pose is estimated, so as to consider the registration speed and accuracy, and provide a reliable spatial correlation basis for subsequent semantic fusion.

[0069] In some embodiments, the step S130 of "performing semantic segmentation on the point cloud data and the image data in the registration data set respectively to obtain first and second semantic segmentation results" includes the following steps:

[0070] Step S1301, using a PointNet++ network to perform semantic segmentation on the point cloud data in the registration data set to extract the three-dimensional boundary and structure parameters of the ancient building cultural relic component;

[0071] The PointNet++ point cloud semantic segmentation process realizes high-precision three-dimensional analysis of ancient building cultural relics components through hierarchical feature extraction and geometric parameterization analysis: first, a multi-level network architecture is constructed, a point cloud pyramid is formed through farthest point sampling (FPS) and ball query, FPS can ensure that representative points are selected in the point cloud distribution, avoiding too concentrated sampling points; Ball Query groups the neighborhood in a spherical region around the selected representative points, which can effectively capture the local geometric features of the point cloud at different scales, providing a guarantee for the subsequent accurate extraction of the three-dimensional boundary and structural parameters of the ancient building cultural relics components. The bottom layer MLP (Multi-Layer Perceptron) extracts the local geometric descriptors of coordinates and normal vectors, and the high-level network captures multi-scale features such as component corners and curved surface transitions, and uses symmetric functions such as maximum pooling to eliminate the influence of point cloud disorder; then based on the feature space distribution, a classifier is used to assign semantic labels (such as "column base" and "couple") to each point, and CRF (Conditional Random Field) is used to optimize the boundary continuity to extract the three-dimensional contour of the component; finally, the segmentation result is analyzed by geometric parameterization, and the component size (length, width, height, curvature radius), spatial relationship (angle, offset), and surface feature (roughness, concave-convex depth) are automatically calculated, forming a three-dimensional boundary model containing accurate geometric contour and quantitative parameters, providing basic data support for ancient building cultural relics BIM modeling and disease assessment.

[0072] In step S1302, the image data in the registration data set is classified at the pixel level using the SegNet network to identify the material type and surface detail features of the ancient building cultural relics component.

[0073] In the image semantic segmentation and cross-modal fusion technology of SegNet, multi-link cooperation is used to realize accurate analysis and cross-modal integration of ancient building material information: in the image semantic segmentation stage, a VGG16 pre-trained model can be used as an encoder, combined with a deconvolution decoder and a jump connection to extract semantic features, and data augmentation strategies such as ±45° rotation and ±20% brightness adjustment can be used to expand the sample, and weighted cross-entropy and Dice loss can be used to optimize the model. The image pixels are classified into six types of materials: blue bricks, wood, stone, paintings, plaster, and metal. After morphological processing (opening / closing operation) and CRF (Conditional Random Field) optimization, a color semantic label map (PNG format) and a class intensity matrix (.npy) are generated, and quantitative features (i.e. surface detail features) such as RGB mean and texture frequency are extracted.

[0074] In the above embodiments, PointNet++ extracts the three-dimensional structural features (boundary, parameter) of the point cloud, and SegNet identifies the material and surface details of the image, respectively analyzing the ancient building information from the aspects of "three-dimensional shape" and "two-dimensional texture", thereby improving the comprehensiveness of semantic segmentation. In view of the unordered nature of the point cloud and the pixel-level distribution characteristics of the image, a special network model is selected to ensure that the segmentation result is more in line with the structural and material characteristics of the ancient building components.

[0075] In some embodiments, the step S130 of "performing cross-modal semantic fusion on the first semantic segmentation result and the second semantic segmentation result based on the association relationship in the registration data set to obtain a first cross-modal semantic segmentation result" comprises the following steps:

[0076] In step S1303, based on the association relationship in the registration data set, the three-dimensional boundary and the structural parameter are mapped to the same dimension through the full connection layer of the cross-modal fusion network, and then spliced with the material type and the surface detail feature to obtain a cross-modal feature vector;

[0077] In step S1304, the mutual information of the cross-modal feature vector is calculated through the attention module of the cross-modal fusion network; and the features in the cross-modal feature vector are weighted and calculated based on the mutual information to obtain a weighted cross-modal feature vector;

[0078] In step S1305, the weighted cross-modal feature vector is judged for component category and material type, and the semantic boundary is corrected through the joint classifier of the cross-modal fusion network, to obtain a first cross-modal semantic segmentation result.

[0079] In the above embodiments, PointNet++ extracts the three-dimensional structural features (boundary, parameter) of the point cloud, and SegNet identifies the material and surface details of the image, respectively analyzing the ancient building information from the aspects of "three-dimensional shape" and "two-dimensional texture", thereby improving the comprehensiveness of semantic segmentation. In view of the unordered nature of the point cloud and the pixel-level distribution characteristics of the image, a special network model is selected to ensure that the segmentation result is more in line with the structural and material characteristics of the ancient building components.

[0080] The cross-modal semantic fusion realizes the deep cooperation of the point cloud geometric features and the image texture features through a three-level mechanism of "space mapping-feature alignment-collaborative optimization", and the specific process is as follows:

[0081] Spatial mapping and feature association: Based on the camera extrinsic matrix (rotation matrix R, translation vector T) obtained by the PnP algorithm in S120, the spatial mapping relationship between point cloud and image is established: the three-dimensional coordinates (X,Y,Z) of point cloud are converted into the pixel coordinates (u,v) of image plane through perspective projection formula, as shown in the following formula (1):

[0082] Formula (1);

[0083] in, , This refers to the camera's intrinsic focal length. , The coordinates of the main point.

[0084] For each point in the point cloud, a 3×3 pixel neighborhood of its corresponding image region is found using bilinear interpolation. The material label (e.g., "pine wood" or "painted"), RGB mean (e.g., R=220, G=60, B=40), and texture frequency features (e.g., 5 lines / cm) output by SegNet are extracted to form an image texture feature vector. ( It includes 1 tag code + 3 RGB values ​​+ 6 texture statistics.

[0085] Feature Dimension Alignment and Stitching: Extracting geometric features of the point cloud output from PointNet++: including 3D coordinates (X,Y,Z), normal vectors ( The curvature k and semantic tags (such as "dougong" and "column") form a geometric feature vector. ( =8 (containing 3 coordinates + 3 normal vectors + 1 curvature + 1 label encoding), which is then processed through a fully connected layer. and Mapping to the same dimensional space (e.g., d=32), feature concatenation is used to obtain cross-modal feature vectors. (; indicates vector concatenation), achieving dimensional alignment of geometric and texture features.

[0086] Attention Mechanisms and Collaborative Optimization: Introducing a Cross-Modal Attention Module Perform weighted optimization: by calculating the mutual information between geometric features and texture features (e.g., Features with high relevance are given higher weights (e.g., the weights of the geometric feature of "cantilever length" and the texture feature of "painted pattern" of the bracket set are increased by 20%), and noise features (e.g. invalid textures in occluded areas) are suppressed.

[0087] Finally, the weighted cross-modal features are input into a joint classifier (softmax layer) to recognize the three-dimensional contour of the component constrained by the image material classification boundary (such as the pixel material within the "column" boundary is only "wood" or "stone"), and to correct the point cloud semantic label with image texture details (such as the area misjudged as "brick sculpture" is corrected as "color drawing" after confirming the image texture), and finally output the first cross-modal semantic segmentation result fused with three-dimensional boundary, material attribute and category intensity (confidence ≥ 0.85), wherein the category intensity is used to evaluate the reliability of the material attribute hanging, and the category intensity only corresponds to the material classification confidence.

[0088] In the above embodiment, the three-dimensional structure and two-dimensional texture features are mapped to the same dimension through the cross-modal network, the mutual information is weighted combined with the attention mechanism, the ambiguity (such as point cloud boundary blur and image depth loss) that may exist in single modal segmentation is solved, and the semantic boundary is corrected. The joint classifier is used to uniformly judge the component category and material, so as to ensure that the semantic labels of the point cloud and the image are consistent, and to provide a reliable basis for the attribute hanging of the subsequent BIM model.

[0089] In some embodiments, the building information model BIM template library of the ancient building cultural relic component is constructed based on the industry foundation class IFC standard, the GIS base map contains terrain data and cadastral data and is in the CityGML format, and the step S150 of "integrating the semantic enhanced BIM model and the geographic information system GIS base map containing geographic reference to obtain a BIM-GIS fusion model" includes the following steps:

[0090] Step S1501, converting the semantic enhanced BIM model into a city geographic markup language CityGML format, preserving IFC semantic information through a detail level of detail LOD hierarchical modeling based on the accuracy requirement of the ancient building component and an application schema AppSchema extension mechanism to obtain a BIM model in the CityGML format;

[0091] Step S1502, integrating the BIM model in the CityGML format and the GIS base map to obtain a BIM-GIS fusion model corresponding to the first operation time.

[0092] The generated BIM model can be converted into a CityGML format, the LOD (Level of Detail) hierarchical modeling and AppSchema (application schema) extension mechanism are used to retain IFC semantic information (such as material type, component ID), the CityGML format BIM model is seamlessly integrated with a GIS (Geographic Information System) base map, component spatial query (such as clicking a column to display material and size) is realized, multi-source data superposition (such as historical scanning data comparison) and three-dimensional visualization display (supporting WebGL interactive browsing) are realized, and the specific implementation is as follows:

[0093] Step S150 focuses on the deep integration of ancient building BIM model and GIS platform, and builds a three-dimensional geographic information system with spatial analysis and visualization capabilities through multi-dimensional technology integration. First, the standardized conversion of BIM model to CityGML is implemented, and open source tools such as FZK-IFC2CityGML or self-defined converters are used to establish the mapping relationship between IFC components and CityGML entities according to ISO standards. IfcColumn is mapped to CityGML BuildingPart, the RGB value and texture path of IfcMaterial are stored through the Appearance node, and the component ID is reserved through gml:id to ensure data traceability; At the same time, LOD hierarchical modeling strategy is adopted, from LOD1 city-level simplified outer contour (error ≤5m) to LOD4 component-level micro details (error ≤1mm), through lod1Geometry to lod4Geometry nodes to realize dynamic loading of detail levels. In terms of semantic information preservation, based on the ApplicationSchema mechanism of CityGML, an extension schema dedicated to ancient buildings is created, defining the AncientMaterial material attribute group and the AncientComponent component parameter group. Through the propertySet node, the assembly relationship and material label in IFC are converted to the topological link and Appearance node information of CityGML, realizing the bidirectional association of semantics and visualization. In the integration of GIS platform, seven-parameter conversion method is used to convert BIM local coordinate system to WGS84+UTM geodetic coordinate system (error ≤10cm), and through 1985 national elevation datum, the terrain data is aligned, the component-level query function (click the column to display material, size and maintenance record) and spatial statistical analysis (generate heat map according to material) are developed, supporting multi-source overlay display of historical scanning data, cadastral data and environmental data. Three-dimensional visualization realizes WebGL lightweight rendering (frame rate ≥30fps) through Draco compression algorithm (10:1 compression ratio), supports roaming, sectioning and measurement operations, and develops WebGL mobile application compatible with AR mode and multi-user collaborative labeling function. This stage of technology realizes the fusion of BIM geometric material information and GIS spatial analysis capability, and can be applied to cultural heritage digital management, urban renewal planning and emergency evacuation simulation scenarios.

[0094] The seamless integration of BIM model and GIS base map is realized by a three-level technical system of "coordinate system accurate conversion-semantic information structured mapping-integrated precision verification": in the aspect of cross-coordinate system matching, first analyze the translation, rotation and scale difference between the local construction coordinate system (such as the right-hand coordinate system with the center point of the building base as the origin) adopted by the BIM model and the national 2000 geodetic coordinate system (WGS84+UTM zone, such as the UTM Zone 50N where the Yingxian Wooden Pagoda is located) adopted by the GIS base map (the maximum translation can reach hundreds of meters, and the rotation error is ≤3°), then select more than 3 public control points (such as wall corners and boundary markers) around the ancient building, and obtain their BIM local coordinates (measured by total station, accuracy ±2mm) and GIS geodetic coordinates (GNSS static measurement, accuracy ±5cm), calculate the translation (ΔX, ΔY, ΔZ), rotation (ε x ,ε y ,ε z) and scale (m) seven parameters, the least squares optimization makes the conversion residual error ≤3 cm (such as the plane error of Yjuan Wood Tower after conversion ≤8 cm, the elevation error ≤5 cm), finally all geometric nodes (such as BuildingPart, LOD4Geometry) of CityGML model are converted to WGS84+UTM coordinate system through seven parameters, 50 feature points (such as column top, dougong endpoint) are randomly selected and compared with GIS base map (1:1000 topographic map) to ensure that the matching error is ≤10 cm; in terms of IFC semantic information retention, different LOD levels are bound by "geometric node + attribute node" double structure semantics, LOD1-LOD2 (simplified model) stores simplified geometry in lod1Geometry / lod2Geometry node, and associates basic semantics (such as component ID, material category) through propertySet node, LOD3-LOD4 (detailed model) stores triangular mesh with texture in lod3Geometry / lod4Geometry node, retains material RGB value and texture path (mapped to IFC IfcMaterialTexture) through Appearance node, and one-to-one correspondence between gml:id and IFC GlobalId (such as "Pillar-001" in CityGML corresponding to "Column-001" in IFC), at the same time, based on the ApplicationSchema mechanism of CityGML, the "Ancient Architecture Semantic Extension schema" is defined, the AncientMaterial group (including "material type", "moisture content", "weathering grade", corresponding to IFC IfcMaterialProperties) and the AncientComponent group (including "component age", "repair record", corresponding to IFC IfcElementQuantity) are added, the mapping relationship is established through XSLT conversion script (such as "IfcColumn.Height=11.23m" in IFC is mapped to "AncientComponent / Height=11.23" in CityGML), the mapping coverage reaches 100%, and after conversion, the attribute quantity is compared through Python script batch (such as IFC contains 120 attributes, CityGML needs to retain ≥118 attributes), and the key attributes are 100% manually extracted, ensuring that the semantic loss rate is <2%.

[0095] In the above embodiments, the BIM template library is constructed based on the IFC standard, converted to CityGML format and semantic information is retained, solving the format barrier between BIM and GIS and realizing cross-platform data sharing. Through LOD grading and AppSchema extension, it is ensured that both the component micro details (such as dougong texture) are retained and the urban macro management needs are met in the integration process, balancing precision and efficiency.

[0096] In some embodiments, the method further includes:

[0097] Step S101: Obtain meteorological data monitored by meteorological satellites and ground meteorological stations within a preset time period;

[0098] Step S102: When the meteorological data meets the preset weather conditions suitable for UAV operation, the first operation time of the UAV platform is determined based on the meteorological data, and the adaptation scheme of the lidar, the RGB camera, the positioning system and the UAV platform is determined. The adaptation scheme includes the addition of protective equipment, switching of working modes and parameter optimization.

[0099] During data acquisition, various weather conditions such as rain, snow, strong light, overcast skies, and fog directly impact the quality of point cloud data, multi-view image data, and position and attitude data collected by drones. These weather conditions can be considered unsuitable for drone operations, while light winds, no precipitation, high visibility, and soft lighting are suitable. Rain and snow cause LiDAR laser reflection signal scattering, increasing point cloud noise (outlier ratio increases from 1%-2% to over 10%). Rain also causes image blurring and loss of texture information due to lens adhesion (e.g., the discernibility of details in dougong (bracket set) paintings decreases by more than 30%). Strong light causes image overexposure (pixel saturation in highlight areas), resulting in loss of material texture details, and LiDAR specular reflection causes local holes in the point cloud. Overcast / foggy skies reduce image contrast due to insufficient light (grayscale values ​​are concentrated in the 50-150 range), increasing the difficulty of material classification (e.g., the accuracy of distinguishing between blue bricks and plaster decreases from 92% to 75%), while reducing laser penetration and decreasing point cloud density (e.g., from 200 points / m in distant areas). 2 Reduced to 80 points / m 2 To address this, a three-tiered approach of "weather forecasting, equipment adaptation, and parameter optimization" was adopted during the data collection phase: For 48 hours prior to data collection, meteorological data was monitored via meteorological satellites and ground weather stations to avoid extreme weather conditions, prioritizing operations during cloudy periods with humidity below 60%, and developing supplementary flight plans for mild or severe weather; in rainy or snowy weather, drones were equipped with rain covers and lenses were coated with hydrophobic coatings, and the LiDAR was put into anti-interference mode; in strong light, the camera was equipped with an adjustable ND (Neutral Density Filter) filter, and the LiDAR was put into anti-specular reflection algorithm mode; during periods of strong light, flight altitude was adjusted, exposure time was shortened, and ISO (sensitivity) and lateral overlap were increased; in cloudy / foggy weather, exposure time was extended, HDR (High Dynamic Range) mode was activated, and the LiDAR scanning frequency and lateral overlap were increased.

[0100] In the above embodiment, the operation opportunity is determined based on the meteorological data, and the equipment adaptation scheme (such as adding protection and optimizing parameters) is formulated to avoid the influence of bad weather on data quality or damage to equipment, and to ensure the stability and controllability of the unmanned aerial vehicle collection process. The operation decision is triggered by the preset weather condition, which reduces the cost of manual judgment and improves the automation and intelligent level of ancient building surveying.

[0101] The unmanned aerial vehicle platform carrying a laser radar (LiDAR), a high-definition RGB camera and a real-time differential positioning system (RTK / PPK) is used to fly and scan the ancient building relics from multiple angles, collect high-precision point cloud data, multi-view image data and position and attitude (POS) information. The Posterior-Mean Rectified Flow (PMRF) algorithm is used to repair and enhance the image, and the point cloud optimization is used to realize accurate cross-modal registration. The PointNet++ and SegNet networks are used for semantic segmentation and fusion of point cloud data and image data respectively, and the three-dimensional boundary, structure parameters, material type and surface detail features of the ancient building components are extracted. A BIM template library is constructed based on the International Cooperative Classification (IFC) standard, a semantic enhanced BIM model is generated, and it is converted into CityGML format to realize seamless integration with the geographic information system (GIS) base map, support component spatial query, multi-source data overlay and three-dimensional visualization display. Regularly update the data to generate a digital twin animation (supplement new point cloud, image and POS data by the unmanned aerial vehicle platform, calibrate the data update deviation based on the coordinate consistency of the new and old POS data, and generate a digital twin animation of the morphological changes of the ancient building), realize long-term dynamic monitoring of the health status of the ancient building. Compared with the traditional surveying and mapping and the existing technology, the application improves the data collection efficiency and quality, realizes the automatic and intelligent surveying and mapping and management of ancient building relics.

[0102] The application provides an intelligent surveying and mapping method based on AI and BIM fusion, which comprises the following steps:

[0103] S1, the unmanned aerial vehicle platform carrying the LiDAR, the RGB camera and the RTK / PPK positioning system is used to fly and scan the ancient building relics and its surrounding environment from multiple angles, and high-precision point cloud data, multi-view image data and POS position data are collected;

[0104] S2, the original point cloud is denoised and hole filled, the multi-view RGB image is repaired and enhanced by using the PMRF algorithm, and the accurate cross-modal registration of the image and the point cloud is realized through feature matching and pose estimation;

[0105] S3, input the point cloud and image data processed in S2, use PointNet++ to perform semantic segmentation on the registered point cloud data, extract the three-dimensional boundary and structural parameters of the ancient building cultural relic component; at the same time, use the SegNet network to perform pixel-level classification on the enhanced image to identify the material type and surface detail features, and realize semantic cooperation of point cloud geometric features and image texture features through cross-modal semantic fusion, and finally output the cross-modal semantic segmentation result of the fusion of three-dimensional boundary, material property and category intensity information;

[0106] S4, construct an ancient building cultural relic component BIM template library based on IFC standard, classify the segmented point cloud components by using an algorithm, infer the assembly relationship through topological relationship analysis (such as “beam supports column” “dou-gong connects beam”), and connect the material properties and structural parameters (such as column height, dou-gong overhang length) extracted in S3 to the BIM model components to generate a semantic enhanced BIM model (IFC format) containing geometric information and material properties;

[0107] S5, convert the generated BIM model into CityGML format, retain the IFC semantic information (such as material type, component ID) through LOD hierarchical modeling and AppSchema extension mechanism, seamlessly integrate the CityGML format BIM model with GIS base map (terrain, cadastral data), realize component spatial query (such as clicking on the column to display material and size), multi-source data superposition (such as historical scanning data comparison) and three-dimensional visualization display (supporting WebGL interactive browsing) ;

[0108] S6, periodically use the UAV to repeat the S1-S3 process to obtain updated point cloud and image data of the ancient building cultural relic area, update the geometric changes (such as component displacement, deformation) and material information of the BIM model through point cloud registration and difference analysis, generate BIM-GIS fusion models of different periods, generate digital twin animation of the building evolution process based on the historical model sequence, visually display the spatio-temporal changes of component size and material distribution, and support long-term dynamic monitoring of the health status of ancient buildings.

[0109] In some embodiments, taking the wooden tower in Ying County, Shanxi Province (Liaoning Dynasty full-wood structure ancient building) as an example, the specific implementation process and application effect of the technical solution are described in detail.

[0110] In the data collection stage of the UAV:

[0111] Device configuration: a UAV can be used, equipped with a laser radar with a point cloud density ≥200 points / m 2 , a 6100-megapixel high-definition camera, and an RTK / PPK positioning system with a plane accuracy of ±1 cm and a height accuracy of ±2 cm.

[0112] Flight strategy: Spiral flight around the main body of the wooden pagoda (total height 67.31 m) + layered overhead flight, set the heading overlap rate to 75%, the lateral overlap rate to 65%, implement 2 times of supplementary flight on the complex structure areas such as the first layer of dougong and the second layer of flat seat. Generate.las format point cloud (total amount about 1.2 TB, point cloud spacing ≤5 mm),.png format image (about 3000) and POS data containing GPS / IMU (timestamp accuracy 1 ms).

[0113] Supplementary measurement after rainstorm in July 2024: Rain cover was used during acquisition, the point cloud noise rate was reduced to 5% in the anti-interference mode of LiDAR; through adaptive density filtering processing, the point cloud cleanliness of the dougong area was restored to 96%, the PMRF algorithm repaired the rain-shielded painted image, and the texture clarity reached 85% of that in normal weather.

[0114] Midday strong light acquisition in August 2024: The camera was equipped with an ND16 filter, the exposure time was 1 / 1000s, and 3 images were synthesized in HDR mode; after processing, the overexposure area ratio was reduced from 20% to 5%, and the classification accuracy of SegNet for pine and painting remained above 90%.

[0115] In the point cloud and image preprocessing and cross-modal registration stage:

[0116] Point cloud processing: In the coarse registration stage, FPFH feature matching + ICP algorithm can be used to align multi-station point clouds, with an initial rotation error ≤8° and a translation error ≤10 cm. In the denoising and completion stage, radius filtering (threshold 0.05 m) is used to remove outliers, Poisson reconstruction algorithm is used to fill the cracks in the tower body (volume about 0.3 m³), and RANSAC plane fitting is used to optimize the continuity of the tower body surface (plane error ≤0.8 mm).

[0117] Image repair: PMRF algorithm is used to process the dougong painted image with severe occlusion: construct dataset: collect 2000 high-definition texture images of Yingxian Wooden Pagoda, and simulate 500 aging and fading samples. Training process: first extract the dougong pattern, column body wood pattern and other features through UNet encoder, and then reduce the MSE loss to 0.012 in the posterior mean prediction stage; in the correction flow model training, solve the ODE based on optimal transport theory, and the perception loss of the repair result is ≤0.035, and the texture clarity of the repaired painted image is improved by 40%.

[0118] Cross-modal registration: Extract SIFT key points (about 2000 per image), combine POS data to calculate camera external parameters through PnP algorithm, map images to point cloud space, and generate RGB point cloud (color restoration degree ≥90%).

[0119] In the semantic segmentation and cross-modal fusion stage:

[0120] Point cloud semantic segmentation (PointNet++): Input the registered point cloud (about 80 million points), construct the point cloud pyramid through FPS sampling (sampling rate 10%) and Ball Query (radius 0.2 m), identify components such as "column", "corbel", "eave", and the three-dimensional boundary extraction accuracy is ≤2 mm, among which the corbel overhang length measurement error is ≤1.5 mm.

[0121] Image semantic segmentation (SegNet): Classify the enhanced image, divide the tower body material into "pine", "green brick", and "color drawing" three categories, the pixel-level classification accuracy is 92%, extract the color drawing area RGB mean value (such as red color drawing R=220±5, G=60±3, B=40±2) and texture frequency (average 5 strips / cm).

[0122] Cross-modal fusion: Project the point cloud coordinates to the image plane, obtain the adjacent pixel material label, generate a cross-modal point cloud containing three-dimensional coordinates (X / Y / Z accuracy ±3 mm), material properties (such as column pine density 0.54 g / cm³), output semantic label map and point cloud-image association table.

[0123] In the semantic enhancement BIM model construction stage:

[0124] BIM template library: Based on IFC standard to create wooden tower component templates, such as "five-layer eave column" template with preset height 11.2 m, diameter 0.8 m, wood moisture content 12%, etc. The corbel template defines the arch length 45 cm, the warping thickness 12 cm, and other variable parameters.

[0125] Point cloud classification and topology reasoning: Use DBSCAN clustering (eps=0.1 m, minPts=50) to classify point clouds into categories such as columns (89) and corbels (546 groups), with a classification accuracy of 96%. Based on graph theory algorithm to infer assembly relationship, such as "two-layer east column supporting two-layer east eave architrave" and "corner corbel connecting eave rafter and column head", construct topology network (edge weight ≥0.9).

[0126] Attribute hanging: Hang the column height (11.23 m±0.02 m) and corbel material (pine) extracted by S3 to the IFC model, generate IFC format BIM model (file size 1.8 GB, geometric error ≤2 mm).

[0127] In the BIM-GIS integration and visualization stage:

[0128] Format conversion: BIM model is converted to LOD4 level CityGML using FZK-IFC2CityGML tool, and "wooden tower arch" and "Liao dynasty painting" attribute groups are defined through AppSchema extension, retaining component ID (such as Gong_001) and material RGB value (painting red R=218, G=56, B=32).

[0129] GIS integration: Seven-parameter conversion converts model coordinates to WGS84+UTM Zone 50N coordinate system (planar error ≤8cm), and superimposes 1:1000 topographic data and cadastral red line. Develop WebGL interactive system, click on the three-layer west column to display: material (pine), size (height 9.8m, diameter 0.78m), and historical repair record (2016 column body corrosion treatment).

[0130] Visualization effect: Through Draco compression (compression ratio 12:1), realize smooth rendering on web side (frame rate 35fps), support multi-source data superposition (such as 2010 scanning point cloud and current model comparison, difference is displayed by heat map).

[0131] In the dynamic monitoring and digital twin stage:

[0132] Regular collection: Every quarter, the wooden tower is scanned by unmanned aerial vehicle (wooden building monitoring frequency), 4 periods of data are collected from March to September 2024, and the void rate is less than 0.8%, and the light condition is controlled at 500-800 lux.

[0133] Change analysis: Point cloud registration: coarse registration uses ISS+FPFH (rotation error ≤3°, translation error ≤3cm), and fine registration takes the base as the reference (RMSE=1.5mm, coincidence rate 97%). Difference analysis: found that the second floor east corner column is inclined by 0.6° (exceeding the warning threshold of 0.5°), and the volume of the arch bracket mortise and tenon joint is reduced by 6%, and the deformation report (error ≤1mm) is generated.

[0134] Digital twin animation: Based on the 2024 fourth phase model, evolution animation is generated, color mapping is used to display column inclination change (red represents deformation >0.5°), texture mixing technology restores the fading process of painting, and time axis dragging is supported to view the displacement trajectory of arch bracket from March to September 2024 (maximum displacement 2.3mm).

[0135] Implementation effect summary:

[0136] Precision index: three-dimensional coordinate error ≤3mm, material classification accuracy 92%, BIM model geometric precision ±2mm, dynamic monitoring deformation recognition precision ≤1mm.

[0137] Application value: A digital archive containing geometry, material, and historical evolution of Yingxian Wooden Pagoda is established, which supports the cultural relic protection unit to develop targeted repair schemes (such as implementing support reinforcement on the inclined column), and realizes the whole-process digital management of ancient buildings from static modeling to dynamic health monitoring.

[0138] The beneficial effects of the embodiments of the present application include:

[0139] Data quality improvement: The PMRF algorithm is used to repair and enhance the image, which can effectively solve the problems of occlusion, blur or damage in the image, and combined with point cloud optimization means, precise cross-modal registration between image and point cloud is realized, and the overall quality of data is improved.

[0140] Precise semantic segmentation: PointNet++ and SegNet networks are used comprehensively to process the registered point cloud data and enhanced image, respectively, to realize the extraction of three-dimensional boundaries and structural parameters of ancient building cultural relics components, and the identification of material types and surface detail features, and through cross-modal semantic fusion, the results are more accurate and comprehensive.

[0141] Generate semantic enhanced BIM template library: Based on IFC standard, a BIM template library of ancient building cultural relics components is constructed (not only containing geometric information, but also integrating material properties, topological relationships and historical evolution information), and the segmented point cloud components are classified and reasoning assembly relationship by using algorithm, the extracted material properties and structural parameters are connected to the BIM model components, and finally the semantic enhanced BIM model containing geometric information and material properties is generated, which provides standardized data support for ancient building digital archiving, protection planning and repair simulation.

[0142] Deep integration of BIM and GIS: The generated BIM model is converted into CityGML format, the IFC semantic information is retained through LOD hierarchical modeling and AppSchema extension mechanism, the seamless integration with GIS base map is realized, the component spatial query, multi-source data overlay and three-dimensional visualization display are supported, and the efficiency of information sharing and collaborative work is improved. The limitation of traditional method that only spatial position overlay can be performed is solved.

[0143] Based on the foregoing embodiments, the embodiments of the present application further provide a BIM-based intelligent surveying and mapping system, which comprises various modules and units included in the modules, and can be realized by a processor in an electronic device. Of course, it can also be realized by a specific logic circuit. In the implementation process, the processor can be a central processing unit (CPU), a micro processing unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA).

[0144] Figure 2 A schematic diagram of the composition structure of the BIM-based intelligent surveying and mapping system provided by the embodiments of the present application is shown in Figure 2 The system 200 comprises an acquisition module 210, a registration module 220, a semantic segmentation module 230, a first generation module 240, and a first integration module 250, wherein:

[0145] The acquisition module 210 is configured to perform multi-angle flight scanning on an ancient building artifact body and its surrounding environment at a first operation time through a UAV platform loaded with a laser radar, an RGB camera, and a positioning system, and synchronously acquire point cloud data, multi-view image data, and position and attitude data, so that the point cloud data and the multi-view image data are associated with the position and attitude data at the corresponding time through respective time stamps.

[0146] The registration module 220 is configured to respectively pre-process the point cloud data and the multi-view image data, and perform cross-modal registration on the pre-processed point cloud data and multi-view image data based on the position and attitude data through feature matching and pose estimation, to obtain a point cloud-image registration data set.

[0147] The semantic segmentation module 230 is configured to respectively perform semantic segmentation on the point cloud data and the image data in the registration data set, to obtain a first semantic segmentation result and a second semantic segmentation result, and perform cross-modal semantic fusion on the first semantic segmentation result and the second semantic segmentation result based on the association relationship in the registration data set, to obtain a first cross-modal semantic segmentation result. The first cross-modal semantic segmentation result comprises a plurality of original point cloud components with class labels, and each original point cloud component corresponds to a texture region in the image data through registration association.

[0148] The first generation module 240 is configured to classify and aggregate the plurality of original point cloud components based on the category labels of the original point cloud components in the first cross-modal semantic segmentation result, to obtain segmented point cloud components, to construct a topological relationship between the segmented point cloud components, to infer an assembly relationship between the segmented point cloud components through the topological relationship, to call a template corresponding to the category of each segmented point cloud component from a building information model (BIM) template library of ancient building relic components that is pre-constructed, to instantiate a BIM model component, to structure the first cross-modal semantic segmentation result and connect the first cross-modal semantic segmentation result to the BIM model component, and to construct a semantic enhanced BIM model that has both geometric shape information and semantic attribute information.

[0149] The first integration module 250 is configured to integrate the semantic enhanced BIM model and a geographic information system (GIS) base map containing a geographic reference, to obtain a BIM-GIS fusion model corresponding to the first operation time, and to use the BIM-GIS fusion model for component spatial query, multi-period component state comparison and display, and three-dimensional visualization display of the ancient building relic.

[0150] In some embodiments, the system 200 further includes an updating module configured to repeat the steps of the synchronous acquisition, the cross-modal registration, and the semantic segmentation at a plurality of second operation times according to a preset time interval, to obtain a plurality of second cross-modal semantic segmentation results corresponding to the plurality of second operation times, and to update the semantic enhanced BIM model based on the plurality of second cross-modal semantic segmentation results; a second integration module configured to integrate the updated semantic enhanced BIM models corresponding to the plurality of second operation times and the GIS base map, to obtain BIM-GIS fusion models corresponding to the plurality of second operation times; and a second generation module configured to generate a digital twin animation of an architectural evolution process based on the BIM-GIS fusion models corresponding to the first operation time and the plurality of second operation times, and to use the digital twin animation to display the spatial and temporal changes in the size and material distribution of the components of the ancient building relic.

[0151] In some embodiments, the registration module 220 includes a first processing submodule configured to perform denoising and hole completion processing on the point cloud data, and a second processing submodule configured to perform repair and enhancement processing on the multi-view image data by using a posterior mean modification flow (PMRF) algorithm.

[0152] In some embodiments, the registration module 220 further comprises: an extraction submodule for extracting key points from the preprocessed multi-view image data using a scale-invariant feature transform (SIFT) or a speeded up robust features (SURF) algorithm, and extracting corresponding visual features based on the key points; extracting geometric features from the preprocessed point cloud data; comparing the visual features and the geometric features, and matching image-point cloud feature anchor pairs corresponding to the same physical location based on an initial spatial mapping relationship established by the position and attitude data; determining an accurate spatial pose of the RGB camera when shooting and an accurate spatial mapping relationship between image pixels and point cloud three-dimensional coordinates based on the initial spatial mapping relationship by substituting the effective feature anchor pairs into a perspective-n-point (PnP) algorithm with the position and attitude data as the starting point; and an integration submodule for structurally integrating the preprocessed point cloud data, the preprocessed multi-view image data, the accurate spatial pose, and the accurate spatial mapping relationship to obtain a point cloud-image registration dataset.

[0153] In some embodiments, the semantic segmentation module 230 comprises: a first semantic segmentation submodule for performing semantic segmentation on the point cloud data in the registration dataset using a PointNet++ network to extract three-dimensional boundaries and structural parameters of the ancient building cultural relic components; and a second semantic segmentation submodule for performing pixel-level classification on the image data in the registration dataset using a SegNet network to identify material types and surface detail features of the ancient building cultural relic components.

[0154] In some embodiments, the semantic segmentation module 230 further comprises: a splicing submodule for splicing the three-dimensional boundaries and the structural parameters to the material types and the surface detail features after mapping them to the same dimension based on the association relationship in the registration dataset through a full connection layer of a cross-modal fusion network to obtain a cross-modal feature vector; a calculation submodule for calculating mutual information of the cross-modal feature vector through an attention module of the cross-modal fusion network; performing weighted calculation on feature pairs in the cross-modal feature vector based on the mutual information to obtain a weighted cross-modal feature vector; and a judgment submodule for performing component category and material type judgment and semantic boundary correction on the weighted cross-modal feature vector through a joint classifier of the cross-modal fusion network to obtain a first cross-modal semantic segmentation result.

[0155] In some embodiments, the building information model BIM template library of the ancient building cultural relic component is constructed based on an industry foundation class IFC standard, the GIS base map contains terrain data and cadastral data and is in a CityGML format, and the first integration module 250 includes: a conversion submodule, configured to convert the semantic enhanced BIM model into a city geographic markup language CityGML format, reserve IFC semantic information through a detail level of detail LOD hierarchical modeling and application schema AppSchema extension mechanism based on accuracy requirements of the ancient building component, and obtain a BIM model in the CityGML format; and an integration submodule, configured to integrate the BIM model in the CityGML format and the GIS base map, and obtain a BIM-GIS fusion model corresponding to the first operation time.

[0156] In some embodiments, the system 200 further includes: an acquisition module, configured to acquire meteorological data monitored by a meteorological satellite and a ground meteorological station within a preset period of time; and a determination module, configured to, when the meteorological data meet preset weather conditions suitable for unmanned aerial vehicle operation, determine the first operation time of the unmanned aerial vehicle platform based on the meteorological data, and determine an adaptation scheme of the laser radar, the RGB camera, the positioning system, and the unmanned aerial vehicle platform, the adaptation scheme including adding protective equipment, switching a working mode, and optimizing parameters.

[0157] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that the size of the sequence number of each process in various embodiments of the present application does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The sequence number of the above embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments.

[0158] It should be noted that, in this document, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0159] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. The above described device embodiments are merely exemplary. After the above description of the device embodiments, it will be apparent to those skilled in the art that the unit division is only a logical function division, and there can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not implemented. In addition, the coupling or direct coupling or communication connection between the above described components can be indirect coupling or communication connection through some interfaces, devices, or units, and can be electrical, mechanical, or in other forms.

[0160] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units; they can be located in one place, or distributed on a plurality of network units; and some or all of the units can be selected as needed to achieve the purposes of the embodiments of the present application. In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware, or in the form of hardware plus software functional units.

[0161] Alternatively, the integrated units described above, if implemented in the form of software functional modules and sold or used as independent products, can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a plurality of instructions for causing an apparatus to perform all or part of the methods described in the embodiments of the present application. The storage medium mentioned above includes: mobile storage devices, ROM, magnetic disks or optical disks, and various media that can store program codes.

[0162] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments. The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.

[0163] The above description is merely an implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A BIM-based intelligent mapping method, characterized in that, The application comprises the following steps: By mounting a laser radar, an RGB camera and a positioning system on a UAV platform, multi-angle flight scanning of the ancient building and its surrounding environment is performed at a first operation time, and point cloud data, multi-view image data and position and attitude data are synchronously collected, so that the point cloud data and the multi-view image data are associated with the position and attitude data at the corresponding time through respective time stamps; The point cloud data and the multi-view image data are preprocessed respectively, and the preprocessed point cloud data and multi-view image data are cross-modal registered through feature matching and pose estimation based on the position and attitude data, to obtain point cloud-image registration data sets; The point cloud data and the image data in the registration data set are respectively subjected to semantic segmentation, to obtain first and second semantic segmentation results, and cross-modal semantic fusion is performed on the first and second semantic segmentation results based on the association relationship in the registration data set, to obtain a first cross-modal semantic segmentation result, which comprises a plurality of original point cloud components with class labels, and each original point cloud component corresponds to a texture region in the image data through registration association; Based on the class labels of each original point cloud component in the first cross-modal semantic segmentation result, the plurality of original point cloud components are classified and aggregated to obtain segmented point cloud components, and a topological relationship between the segmented point cloud components is constructed, the assembly relationship between the segmented point cloud components is inferred through the topological relationship, and based on the assembly relationship, a BIM model component is instantiated by calling a template corresponding to the class of each segmented point cloud component from a BIM template library of ancient building components pre-constructed, and the first cross-modal semantic segmentation result is structured and connected to the BIM model component, to construct a semantic enhanced BIM model with geometric shape information and semantic attribute information; The semantic enhanced BIM model and a GIS base map containing a geographic space reference are integrated, to obtain a BIM-GIS fusion model corresponding to the first operation time, which is used for component spatial query, multi-period component state comparison and display and three-dimensional visualization of the ancient building; The steps of synchronous collection, cross-modal registration and semantic segmentation are repeated at a plurality of second operation times according to a preset time interval, to obtain second cross-modal semantic segmentation results corresponding to the plurality of second operation times, and the semantic enhanced BIM model is updated based on the plurality of second cross-modal semantic segmentation results; The updated semantic enhanced BIM model corresponding to the plurality of second operation times and the GIS base map are integrated, to obtain BIM-GIS fusion models corresponding to the plurality of second operation times; Based on the BIM-GIS fusion models corresponding to the first operation time and the plurality of second operation times, a digital twin animation of the building evolution process is generated; The digital twin animation is used to show the spatial and temporal changes in the component size and the spatial and temporal changes in the material distribution of the ancient building.

2. The method of claim 1, wherein, The pre-processing of the point cloud data and the multi-view image data respectively comprises: The point cloud data is denoised and hole-completed; The multi-view image data is repaired and enhanced by using a posterior mean revision flow (PMRF) algorithm.

3. The method of claim 1, wherein, The pre-processed point cloud data and the pre-processed multi-view image data are cross-modality registered based on the position and attitude data to obtain a point cloud-image registration data set, comprising: Key points are extracted from the pre-processed multi-view image data by using a scale invariant feature transform (SIFT) or a speeded up robust features (SURF) algorithm, and corresponding visual features are extracted based on the key points; geometric features are extracted from the pre-processed point cloud data; the visual features and the geometric features are compared, and an initial spatial mapping relationship is established based on the position and attitude data to match image-point cloud feature anchor pairs corresponding to the same physical position; effective feature anchor pairs are screened from the image-point cloud feature anchor pairs based on the initial spatial mapping relationship; The effective feature anchor pairs are substituted into a perspective-n-point (PnP) algorithm based on the initial spatial mapping relationship and the position and attitude data to determine an accurate spatial attitude of the RGB camera when the image is captured and an accurate spatial mapping relationship between image pixels and point cloud three-dimensional coordinates; The pre-processed point cloud data, the pre-processed multi-view image data, the accurate spatial attitude and the accurate spatial mapping relationship are structured and integrated to obtain the point cloud-image registration data set.

4. The method of claim 1, wherein, The point cloud data and the image data in the registration data set are respectively subjected to semantic segmentation to obtain first and second semantic segmentation results, comprising: PointNet++ network is used to perform semantic segmentation on the point cloud data in the registration data set to extract three-dimensional boundaries and structure parameters of the ancient building cultural relic components; SegNet network is used to perform pixel-level classification on the image data in the registration data set to identify material types and surface detail features of the ancient building cultural relic components.

5. The method of claim 4, wherein, The first and second semantic segmentation results are cross-modality semantically fused based on the associated relationship in the registration data set to obtain a first cross-modality semantic segmentation result, comprising: The three-dimensional boundaries and the structure parameters are mapped to the same dimension as the material types and the surface detail features after being concatenated by a full connection layer of a cross-modality fusion network based on the associated relationship in the registration data set to obtain a cross-modality feature vector; Mutual information of the cross-modality feature vector is calculated by an attention module of the cross-modality fusion network; features in the cross-modality feature vector are weighted calculated based on the mutual information to obtain a weighted cross-modality feature vector; The weighted cross-modality feature vector is subjected to component category and material type judgment and semantic boundary correction by a joint classifier of the cross-modality fusion network to obtain the first cross-modality semantic segmentation result.

6. The method of claim 1, wherein, The building information model BIM template library of the ancient building cultural relic component is constructed based on an industry foundation class IFC standard, the GIS base map contains terrain data and cadastral data and is in a CityGML format, the semantic enhanced BIM model and the geographic information system GIS base map containing geographic spatial references are integrated to obtain a BIM-GIS fusion model, including: The semantic enhanced BIM model is converted into a city geographic markup language CityGML format, IFC semantic information is reserved through LOD hierarchical modeling based on accuracy requirements of ancient building components and an AppSchema extension mechanism, and a BIM model in a CityGML format is obtained; The BIM model in the CityGML format and the GIS base map are integrated to obtain a BIM-GIS fusion model corresponding to the first operation time.

7. The method of claim 1, wherein, The method further includes: acquiring meteorological data monitored by a meteorological satellite and a ground meteorological station within a preset period; when the meteorological data meet preset weather conditions suitable for unmanned aerial vehicle operation, determining the first operation time of the unmanned aerial vehicle platform based on the meteorological data, and determining an adaptation scheme of the laser radar, the RGB camera, the positioning system and the unmanned aerial vehicle platform, the adaptation scheme including adding protective equipment, switching working modes and optimizing parameters.

8. A BIM-based intelligent mapping system, characterized by, including: The acquisition module is configured to perform multi-angle flight scanning on an ancient building cultural relic body and a surrounding environment thereof at a first operation time by an unmanned aerial vehicle platform carrying a laser radar, an RGB camera and a positioning system, and synchronously acquire point cloud data, multi-view image data and position and attitude data, so that the point cloud data and the multi-view image data are associated with position and attitude data at corresponding time points through respective time stamps. The registration module is configured to respectively pre-process the point cloud data and the multi-view image data, and perform cross-modal registration on the pre-processed point cloud data and multi-view image data based on the position and attitude data through feature matching and pose estimation to obtain a point cloud-image registration data set. The semantic segmentation module is configured to respectively perform semantic segmentation on the point cloud data and the image data in the registration data set to obtain first and second semantic segmentation results, and perform cross-modal semantic fusion on the first and second semantic segmentation results based on an association relationship in the registration data set to obtain a first cross-modal semantic segmentation result, the first cross-modal semantic segmentation result including a plurality of original point cloud components with category labels, and each original point cloud component being associated with a texture region in corresponding image data through registration. The first generation module is configured to classify and aggregate the plurality of original point cloud components based on the category labels of the original point cloud components in the first cross-modal semantic segmentation result, to obtain segmented point cloud components, to construct topological relationships between the segmented point cloud components, to infer assembly relationships between the segmented point cloud components through the topological relationships, to call templates corresponding to the categories of the segmented point cloud components from a building information model (BIM) template library of historical building components pre-constructed based on the assembly relationships, to instantiate the BIM model components, to structure and link the first cross-modal semantic segmentation result to the BIM model components, and to construct a semantic-enhanced BIM model having geometric shape information and semantic attribute information; The first integration module is configured to integrate the semantic-enhanced BIM model and a geographic information system (GIS) base map containing geographic reference to obtain a BIM-GIS fusion model corresponding to the first operation time, and the BIM-GIS fusion model is used for component spatial query, multi-period component state comparison and display, and three-dimensional visualization display of the historical building; The updating module is configured to repeat the steps of the synchronous acquisition, the cross-modal registration, and the semantic segmentation at a plurality of second operation times according to a preset time interval to obtain a plurality of second cross-modal semantic segmentation results corresponding to the second operation times, and to update the semantic-enhanced BIM model based on the plurality of second cross-modal semantic segmentation results; The second integration module is configured to integrate the updated semantic-enhanced BIM models corresponding to the plurality of second operation times and the GIS base map to obtain BIM-GIS fusion models corresponding to the plurality of second operation times; The second generation module is configured to generate a digital twin animation of the building evolution process based on the BIM-GIS fusion models corresponding to the first operation time and the plurality of second operation times, and the digital twin animation is used to display the spatial and temporal changes in the component size and the spatial and temporal changes in the material distribution of the historical building.

Citation Information

Patent Citations

  • Real scene three-dimensional modeling method and system fusing laser point cloud and image

    CN120147563A

  • Intelligent geometric reasoning and semantic understanding method based on three-dimensional large language model

    CN120542438A