Robot multi-modal fusion positioning method based on BIM driving

By using a BIM-driven multimodal fusion positioning method, combined with LiDAR odometry and IMU data, a factor graph model is constructed. By utilizing BIM global constraints, the problems of low positioning accuracy and long-term drift of robot SLAM in construction environments are solved, achieving high-precision and robust positioning results.

CN122015822APending Publication Date: 2026-05-12CCCC FOURTH HIGHWAY ENG CO LTD +2
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CCCC FOURTH HIGHWAY ENG CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-12

Smart Images

  • Figure CN122015822A_ABST
    Figure CN122015822A_ABST
Patent Text Reader

Abstract

A robot multi-mode fusion positioning method based on BIM driving comprises the steps that 1, in the BIM prior map construction and initialization stage, core structure elements are extracted from a BIM model to generate a high-precision prior point cloud map, the high-precision prior point cloud map is used for rough matching with initial laser radar point cloud when a robot is started, and pose alignment under a BIM global coordinate system is completed; and 2) in a real-time data parallel processing and feature extraction stage, the system parallelly processes high-frequency data from an IMU (Inertial Measurement Unit) to perform pre-integration so as to correct laser radar point cloud distortion, and extracts environmental geometric features and BIM matching-oriented structured features from the laser radar point cloud. And 3) in a factor graph construction and optimal positioning stage, constructing a multi-constraint factor graph model, fusing motion estimation provided by an IMU, relative poses provided by scanning matching of geometric features and local sub-maps, and strong global constraints obtained by registering structured features with a global BIM prior map, so as to obtain a multi-constraint factor graph model. Therefore, high-precision and high-robustness positioning is realized, and accumulated drift is effectively inhibited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot autonomous navigation and positioning technology, specifically a BIM-driven multimodal fusion positioning method for robots. Background Technology

[0002] In recent years, with the continuous improvement of automation and intelligence, autonomous mobile robots have been increasingly widely used in fields such as building construction, post-construction maintenance, and inspection and monitoring. For example, robots that autonomously spray water to cure newly poured concrete structures can significantly improve the integrity and durability of building structures. The primary prerequisite for achieving reliable autonomous operation of robots is accurate and robust self-positioning in complex and unstructured working environments.

[0003] However, construction sites differ significantly from typical office and home environments, posing serious challenges to existing positioning technologies: Lack of infrastructure and signal interference: Construction sites (especially in the initial stages after main structure completion) often lack stable power supplies, doors, and windows. This makes positioning solutions relying on pre-deployed beacons, such as Ultra-Wideband (UWB), Wi-Fi, or Bluetooth, not only costly and difficult to deploy, but also highly susceptible to absorption, reflection, and diffraction interference from reinforced concrete structures, leading to signal instability and decreased positioning accuracy. Harsh perception environment: Construction site environments are extremely complex and variable. First, open building structures and drastically changing or even absent natural lighting pose significant challenges to visual SLAM (V-SLAM) methods that rely on lighting and texture features. Second, the large amounts of dust and smoke generated during construction severely affect the propagation of LiDAR beams, resulting in significant noise or even failure of point cloud data. Dynamic environmental changes and feature sparsity: Construction sites are filled with dynamically changing obstacles, such as moving scaffolding, stacks of building materials, and construction workers. These dynamic elements can severely interfere with traditional SLAM algorithms that rely on static environment assumptions. Furthermore, large areas of walls, floors, and other surfaces lack sufficient geometric features, which can easily lead to sensor data degradation, causing the localization algorithm to drift or even fail completely.

[0004] To address these challenges, the academic community has proposed various simultaneous localization and mapping (SLAM) techniques. Early filtering-based methods, such as the Extended Kalman Filter (EKF-SLAM), suffer from poor performance in highly dynamic and nonlinear built environments due to their linear and Gaussian noise assumptions. Furthermore, the state vector dimension expands dramatically with map size, resulting in high computational complexity and inconsistency issues. While Particle Filtering (PF-SLAM) can handle non-Gaussian distributions, it requires a massive number of particles in high-dimensional state spaces, incurring huge computational costs and making it difficult to meet real-time requirements.

[0005] Currently, optimization-based methods, especially graph-based SLAM, have become mainstream. This method treats robot pose and landmarks as nodes in a graph, and sensor observations as edges (constraints) connecting the nodes, solving for the optimal estimate of all nodes through back-end nonlinear optimization. Factor graphs, as a general representation of graph optimization, can flexibly and efficiently fuse heterogeneous information from different sensors such as LiDAR, cameras, and inertial measurement units (IMUs). By constructing tightly coupled multimodal fusion models, such as LIO-SAM and FAST-LIO, motion estimation can be provided by other sensors (such as IMUs) even when a single sensor (such as LiDAR) temporarily degrades, thus improving the robustness of the system to some extent.

[0006] However, these methods still fundamentally rely on high-quality sensor input and reliable loop closure detection to correct for long-term accumulated errors. In construction environments with similar structures (such as corridors or repeating floors), sparse features, or dynamically changing scenes, effective loop closure detection is difficult to achieve. This leads to graph optimization methods also facing serious long-term drift problems, failing to meet the stringent requirements for global position accuracy in tasks such as building maintenance and surveying.

[0007] To solve the above problems, the existing technology is as follows:

[0008] Application No.: CN201910151641.X, Patent Title: Indoor Concrete Crack Maintenance Equipment and Method Based on BIM and Computer Vision. This patent discloses an indoor concrete crack maintenance equipment and method based on BIM and computer vision. The equipment includes an intelligent mobile unit, a feature extraction module, a robot control unit, and an autonomous repair unit. It utilizes a laser scanner, ultrasonic radar, and binocular cameras for environmental scanning and modeling, crack detection, and information collection. A crack decision support module automatically determines the cause of cracks, classifies them, and matches the optimal repair process. Finally, a robotic arm performs the autonomous repair. The intelligent mobile unit (tracked vehicle) moves autonomously indoors and plans its path using SLAM intelligent navigation equipment. Its main focus is on automating the detection and repair of concrete cracks. However, the robustness and accuracy of the SLAM intelligent navigation equipment used may be challenged in complex construction environments with high dynamics, weak features, or severe dust interference, and it cannot effectively address the positioning failure problem caused by sensor data degradation.

[0009] This application proposes a novel multimodal sensor fusion localization method driven by Building Information Modeling (BIM). Its innovation lies in using building structural information from BIM as a global constraint to resist cumulative errors in complex environments. This method tightly couples LiDAR odometry and IMU data to construct a factor graph model, and obtains a robust robot state trajectory through backend optimization. Experimental verification in a real-world construction project scenario demonstrates that this method significantly improves the robot's localization robustness and accuracy in harsh environments such as varying lighting and dust interference, effectively solving the localization failure and long-term drift problems caused by sensor data degradation in traditional SLAM technology in complex construction environments.

[0010] Application No.: CN202210882638.7, Patent Title: Construction Positioning System and Method for Building Engineering Based on Posture Perception and Visual Scanning. This patent provides a construction positioning system and method for building engineering based on posture perception and visual scanning. The system integrates a visual scanning unit, a posture perception unit, a mobile platform, a control unit, and an indicator unit. It marks construction points using a BIM model module, generates 3D point cloud data using a high-precision surface structured light camera, and matches it with the BIM model to achieve automatic control and accurate positioning of construction points, thereby guiding the execution mechanism to carry out construction work. Although this patent combines posture perception and visual scanning for positioning, the performance of visual scanning is easily affected in complex environments such as drastic changes in lighting, sparse texture features, high dust concentration at the construction site, and continuously changing obstacles. This may lead to a decrease in the quality of point cloud data, thus affecting the accuracy and stability of positioning.

[0011] The BIM-driven multimodal fusion localization framework proposed in this application innovatively transforms BIM structural information into SLAM global pose constraints, combining LiDAR and IMU data for tight-coupled optimization. This not only suppresses long-term drift but also significantly improves the robot's localization robustness and accuracy in real-world construction scenarios under harsh environments such as varying lighting and dust interference, demonstrating the superior performance of the proposed framework. Summary of the Invention

[0012] The technical problem to be solved by this invention is that existing robot SLAM technology suffers from low positioning accuracy, poor robustness, and serious long-term cumulative drift in dynamic, weak feature, and harsh indoor environments such as building construction, due to sensor data degradation and lack of effective global constraints. This invention proposes a BIM-driven robot multimodal fusion positioning method.

[0013] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0014] A BIM-driven multimodal fusion localization method for robots, with the following specific steps:

[0015] Step 1, BIM prior map construction and initial pose alignment:

[0016] Core structural elements are extracted from Building Information Modeling (BIM) to generate a global prior point cloud map, and the initial pose alignment in the BIM global coordinate system is completed by registering with the initial LiDAR scanning data when the robot starts.

[0017] Step 2, Parallel processing and feature extraction of IMU and LiDAR data:

[0018] The system processes real-time data from the inertial measurement unit (IMU) and the lidar (LiDAR) in parallel, performs motion distortion correction on the LiDAR point cloud using IMU data pre-integration, and simultaneously extracts geometric features for odometry and structured features for global matching from the point cloud.

[0019] Step 3, Multi-constraint factor graph construction and localization optimization: Construct and optimize a multi-constraint factor graph model, which integrates IMU pre-integration, LiDAR odometry, and global constraints obtained by matching structured features with the global BIM prior map. Solve the problem through nonlinear optimization to output a high-precision, globally consistent robot pose.

[0020] As a further improvement to the present invention, step 1 is specifically as follows:

[0021] Step 1.1: Input the BIM model in the Industrial Basic Class IFC format, selectively extract the core structural elements related to walls, columns and floors, ignore non-structural or variable elements, and form a simplified structural frame model.

[0022] Step 1.2: Surface processing is performed on the extracted structural framework model, and a uniform and high-density three-dimensional point cloud is generated on the model surface using the Poisson disk sampling algorithm as a global BIM prior map.

[0023] Step 1.3: When the robot starts, it collects and stitches together the initial few frames of LiDAR point cloud to form an initial local map. The RANSAC global registration method is used to match it with the global BIM prior map to calculate the initial transformation matrix of the robot in the BIM global coordinate system, thus completing the initialization.

[0024] As a further improvement to the present invention, step 2 is specifically as follows:

[0025] Step 2.1: Pre-integrate all IMU readings between two consecutive lidar keyframes to obtain the relative pose increment, and use this increment to correct motion distortion within the lidar scanning cycle through interpolation to generate a distortion-free and accurate point cloud.

[0026] Step 2.2: For the distortion-free point cloud, calculate the local curvature of the points along each scan line, and classify the point cloud into edge points and planar points with related geometric features for lidar odometer matching based on the curvature threshold.

[0027] Step 2.3: Using an adaptive sliding window and line fitting method, the structured feature point set belonging to planar structures such as walls is identified and extracted from the point cloud. This process can effectively filter out interference points caused by temporary obstacles. The extracted point set is used for subsequent global matching with the BIM prior map.

[0028] As a further improvement to the present invention, step 2.3 is specifically as follows:

[0029] First, an adaptively sized sliding window moves along the point cloud scan line. Second, least-squares line fitting is performed on the point set within each window, and the fitting residual is calculated. Then, starting from the window where the residual meets a preset threshold, all continuous points whose distance to the fitted line is less than the threshold are aggregated and marked as structural points. Finally, all marked structural points are combined to form a structured feature point cloud for BIM matching. This method can automatically separate obstacle points that do not conform to the straight line model.

[0030] As a further improvement to the present invention, step 3 is specifically as follows:

[0031] Step 3.1: Construct a factor graph containing at least three types of constraint factors:

[0032] IMU pre-integration factor: Connects the pose nodes of two consecutive keyframes to provide high-frequency motion constraints;

[0033] LiDAR odometry factor: Provides high-precision relative pose constraints by matching the geometric features of the current frame with the local sub-map;

[0034] BIM Global Constraint Factor: Directly associates the pose nodes of the current frame with the global BIM coordinate system, providing global absolute pose constraints;

[0035] Step 3.2: Set a triggering mechanism. When the robot's cumulative displacement or odometry uncertainty exceeds the threshold, trigger the BIM global constraint factor. After triggering, register the structured feature point cloud extracted in the current frame with the global BIM prior map to obtain a global absolute pose as the observation of the factor.

[0036] Step 3.3: Construct all factors into a large nonlinear least squares problem and use backend optimizers such as iSAM2 to solve it in real time to obtain the optimized estimate of all robot pose nodes, thereby outputting globally consistent and high-precision positioning results.

[0037] .

[0038] The present invention, employing the above technical means, achieves the following advantages: It innovatively transforms the precise 3D structural information of the BIM model into global pose prior constraints in the SLAM factor graph, providing an "absolute benchmark" for robot trajectory. This fundamentally solves the long-term cumulative drift problem that traditional SLAM methods cannot avoid in scenarios lacking loop closures, significantly improving global positioning accuracy. Through the tight coupling and fusion of BIM priors and multimodal sensors, even when a single sensor (such as LiDAR) experiences data degradation due to dust, weak features, etc., the system can still maintain stable and reliable positioning by relying on BIM global constraints and IMU motion estimation. The proposed structured feature extraction algorithm effectively filters out dynamic obstacle interference at the construction site, further enhancing robustness. The "BIM global constraints + LiDAR odometry" mode combines global accuracy and local real-time performance. BIM factors do not need frequent triggering, only correcting global drift when necessary, while high-frequency LiDAR odometry ensures local smoothness and real-time performance of the trajectory. Compared to methods relying on dense loop closure detection, the present invention has a lower computational burden and higher efficiency. This invention directly utilizes existing BIM data from building projects, eliminating the need for additional, expensive positioning facilities, resulting in low cost and ease of implementation. The obtained high-precision global positioning results can directly serve advanced applications such as autonomous navigation, path planning, construction quality inspection, and as-built model comparison for construction robots, demonstrating extremely high engineering practical value. Attached Figure Description

[0039] Figure 1 The flowchart shows a BIM-driven multimodal fusion localization method for robots.

[0040] Figure 2 A flowchart for converting a BIM model into a global prior point cloud map;

[0041] Figure 3 A flowchart for a method of structured feature extraction and obstacle removal for lidar point clouds. Detailed Implementation

[0042] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings:

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Figure 1 The flowchart shows a BIM-driven multimodal fusion localization method for robots. Figure 2 A flowchart for converting a BIM model into a global prior point cloud map; Figure 3 A flowchart for a method of structured feature extraction and obstacle removal for lidar point clouds.

[0044] A BIM-driven multimodal fusion localization method for robots comprises three main stages: First, in the BIM prior map construction and initialization stage, core structural elements are extracted from the BIM model in the design phase to generate a high-precision prior point cloud map, which is used for coarse matching with the initial LiDAR point cloud when the robot starts, completing pose alignment in the BIM global coordinate system. Second, in the real-time data parallel processing and feature extraction stage, the system processes high-frequency data from the IMU in parallel for pre-integration to correct LiDAR point cloud distortion, and extracts environmental geometric features and structured features for BIM matching from the LiDAR point cloud. Finally, in the factor map construction and optimized localization stage, a multi-constraint factor map model is constructed, fusing motion estimation provided by the IMU, relative pose provided by geometric features and scan matching of local sub-maps, and strong global constraints obtained by registering structured features with the global BIM prior map, thereby achieving high-precision, robust localization and effectively suppressing cumulative drift.

[0045] like Figure 1 As shown, the specific steps are as follows:

[0046] (1) Construction of BIM Prior Map and Alignment of Initial Pose: First, core structural elements are extracted from the BIM model to generate a high-precision prior point cloud map. Then, by coarsely matching the initial LiDAR point cloud at robot startup with the BIM prior map, the initial pose alignment of the robot in the BIM global coordinate system is achieved, providing a global reference for subsequent positioning.

[0047] (2) Parallel processing and feature extraction of IMU and LiDAR data: The system processes real-time data from IMU and LiDAR in parallel. IMU data is pre-integrated for motion distortion correction of the LiDAR point cloud. At the same time, two types of feature extraction are performed on the point cloud: one is traditional geometric features (edge ​​points and planar points); the other is structured features for BIM matching (walls, columns, etc.), and interference points caused by temporary obstacles are filtered out through an adaptive method.

[0048] (3) Multi-constraint factor graph construction and localization optimization: A multi-constraint factor graph model is constructed to integrate all information sources to achieve high-precision and robust localization. This model combines high-frequency motion estimation provided by IMU pre-integration, relative pose estimation provided by scanning and matching environmental geometric features and local sub-maps, and most importantly, incorporates global pose estimation obtained by registering structured feature point clouds with global BIM prior maps as strong constraints, fundamentally suppressing cumulative drift.

[0049] The steps for constructing the BIM prior map and aligning the initial pose are as follows:

[0050] a. Input a BIM model in IFC (Industry Foundation Classes) or other formats. Using the BIM parsing library, we selectively extract structural elements crucial for location, such as walls, columns, floors, and beams. We ignore non-structural or volatile elements such as doors, windows, furniture, and pipes, creating a simplified BIM model containing only the permanent structural framework.

[0051] b. Since LiDAR can only perceive the surface of objects, the extracted 3D solid models need to undergo surface processing. We remove the internal geometry of all entities, retaining only their outer surfaces. Then, to ensure the uniformity and representativeness of the point cloud, we use the Poisson disk sampling algorithm to sample these surfaces, generating a high-density, uniformly distributed 3D point cloud. This point cloud is the global BIM prior map M. BIM The original BIM model is transformed into a sparse point cloud model representing the structural framework. Since the sampling density can be controlled and the subsequent matching algorithm is robust, even if there are slight deviations between the sampling points and the real physical surface, they can be regarded as part of the sensor noise.

[0052] c. When the robot starts running, acquire the first 3-5 frames of LiDAR point cloud data. Stitch these frames together to create an initial local map. Then, the RANSAC global registration method is used in... Searching for and The best match is found. Once a high-confidence match is found, the robot's position in the BIM global coordinate system can be determined. The initial transformation matrix below Complete system initialization.

[0053] The parallel processing and feature extraction steps for IMU and LiDAR data are as follows:

[0054] a. The IMU provides angular velocity ω and linear acceleration α at high frequencies (e.g., 100-200 Hz). We compare two consecutive LiDAR keyframes k i and k i+1Integrate all IMU readings between them to obtain the relative pose increment. , , This pre-integral quantity is also used to correct motion distortion during LiDAR scanning. For any point within the scan period... The time of its collection is By using IMU interpolation, the precise pose at that moment can be obtained, thereby... Transform to the coordinate system of the starting time of the frame scan to obtain a distortion-free LiDAR point cloud.

[0055] b. Calculate the local curvature of points along each scan line for the distortion-free point cloud. Based on the magnitude of curvature, classify the points into edge points (large curvature) and planar points (small curvature) for subsequent lidar odometry calculations.

[0056] c. Reduce the dimensionality of each scan line to a 2D plane. Use a sliding window method to move along the scan line. The window size w can be dynamically adjusted according to the point cloud density, and the window segmentation is adaptive to the point cloud model.

[0057] d. Perform least-squares line fitting on the point set within each window. Calculate the fitting residual R. If R is less than a preset threshold... If the point within the window is considered to belong to a potential planar structure such as a wall, then it is assumed that the point belongs to a potential planar structure such as a wall.

[0058] e. Starting with the first window that meets the conditions, add its points to a buffer queue K. Then check the subsequent points one by one. If its distance to the fitted line of the current queue K is less than If the distance is greater than 1, then add it to K. If the queue length exceeds the maximum value L, the line segment structure detection is considered complete. All points in queue K are marked as structure points, and the search for new structures restarts from subsequent points. This process effectively connects consecutive wall points and automatically separates obstacle points that do not conform to the straight-line model (such as buckets or irregular piles of materials). Finally, the set of all points marked as structure points is obtained. This will be used for subsequent BIM matching.

[0059] The steps for constructing and optimizing the multi-constraint factor graph are as follows:

[0060] a. IMU factor: Pose node connecting two consecutive keyframes and The residual is defined as the difference between the pre-integrated measurement and the predicted value between the two pose nodes. This factor provides high-frequency motion information, ensuring the smoothness of the trajectory.

[0061] b. LiDAR odometry factor: By using the current frame Environmental feature points (edge ​​points and planar points) and a local submap composed of the most recent historical keyframes. Matching is performed to obtain a relative pose transformation. This transformation serves as a constraint connection node. and This ensures the accuracy of the local trajectory.

[0062] c. BIM Global Pose Factor: A trigger mechanism is set, for example, when the robot's cumulative displacement exceeds a threshold (e.g., 5 meters) or the odometry uncertainty increases to a certain level, this factor is triggered. After triggering, the structured feature point cloud extracted from the current frame is... With global BIM prior map Perform ICP registration. After successful registration, you will obtain the current robot coordinates in the global coordinate system. Absolute position below We treat this absolute pose as a unary factor and directly connect it to the current pose node. Above. This factor acts like a powerful anchor, "pinning" the entire trajectory map to the globally correct pose, thereby correcting all historically accumulated drift.

[0063] d. Construct all the above factors into a large-scale nonlinear least squares problem. Its objective function is the weighted sum of the Mahalanobis distances of the residuals of all factors:

[0064]

[0065] in It is the first The residual function of each factor It is its covariance matrix. We use the back-end optimizer iSAM2 to solve this problem in real time, obtaining all robot pose nodes. The optimized estimation outputs globally consistent and high-precision positioning results.

[0066] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any modifications or equivalent changes made based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A BIM-driven multimodal fusion localization method for robots, characterized in that, The specific steps are as follows: Step 1, BIM prior map construction and initial pose alignment: Core structural elements are extracted from Building Information Modeling (BIM) to generate a global prior point cloud map, and the initial pose alignment in the BIM global coordinate system is completed by registering with the initial LiDAR scanning data when the robot starts. Step 2, Parallel processing and feature extraction of IMU and LiDAR data: The system processes real-time data from the inertial measurement unit (IMU) and the lidar (LiDAR) in parallel, performs motion distortion correction on the LiDAR point cloud using IMU data pre-integration, and simultaneously extracts geometric features for odometry and structured features for global matching from the point cloud. Step 3, Multi-constraint factor graph construction and localization optimization: Construct and optimize a multi-constraint factor graph model, which integrates IMU pre-integration, LiDAR odometry, and global constraints obtained by matching structured features with the global BIM prior map. Solve the problem through nonlinear optimization to output a high-precision, globally consistent robot pose.

2. The BIM-driven multimodal fusion localization method for robots according to claim 1, characterized in that: Step 1 is described in detail as follows: Step 1.1: Input the BIM model in the Industrial Basic Class IFC format, selectively extract the core structural elements related to walls, columns and floors, ignore non-structural or variable elements, and form a simplified structural frame model.

3. Step 1.2: Surface-process the extracted structural framework model and use the Poisson disk sampling algorithm to generate a uniform and high-density 3D point cloud on the model surface as a global BIM prior map. Step 1.3: When the robot starts, it collects and stitches together the initial few frames of LiDAR point cloud to form an initial local map. The RANSAC global registration method is used to match it with the global BIM prior map to calculate the initial transformation matrix of the robot in the BIM global coordinate system, thus completing the initialization.

4. The BIM-driven multimodal fusion localization method for robots according to claim 1, characterized in that: Step 2 is described in detail below: Step 2.1: Pre-integrate all IMU readings between two consecutive lidar keyframes to obtain the relative pose increment, and use this increment to correct motion distortion within the lidar scanning cycle through interpolation to generate a distortion-free and accurate point cloud. Step 2.2: For the distortion-free point cloud, calculate the local curvature of the points along each scan line, and classify the point cloud into edge points and planar points with related geometric features for lidar odometer matching based on the curvature threshold. Step 2.3: Using an adaptive sliding window and line fitting method, the structured feature point set belonging to planar structures such as walls is identified and extracted from the point cloud. This process can effectively filter out interference points caused by temporary obstacles. The extracted point set is used for subsequent global matching with the BIM prior map.

5. The BIM-driven multimodal fusion localization method for robots according to claim 3, characterized in that: Step 2.3 is as follows: First, an adaptively sized sliding window moves along the point cloud scan line. Second, least-squares line fitting is performed on the point set within each window, and the fitting residual is calculated. Then, starting from the window where the residual meets a preset threshold, all continuous points whose distance to the fitted line is less than the threshold are aggregated and marked as structural points. Finally, all marked structural points are combined to form a structured feature point cloud for BIM matching. This method can automatically separate obstacle points that do not conform to the straight line model.

6. The BIM-driven multimodal fusion localization method for robots according to claim 1, characterized in that: Step 3 is described in detail below: Step 3.1: Construct a factor graph containing at least three types of constraint factors: IMU pre-integration factor: Connects the pose nodes of two consecutive keyframes to provide high-frequency motion constraints; LiDAR odometry factor: Provides high-precision relative pose constraints by matching the geometric features of the current frame with the local sub-map; BIM Global Constraint Factor: Directly associates the pose nodes of the current frame with the global BIM coordinate system, providing global absolute pose constraints; Step 3.2: Set a triggering mechanism. When the robot's cumulative displacement or odometry uncertainty exceeds the threshold, trigger the BIM global constraint factor. After triggering, register the structured feature point cloud extracted in the current frame with the global BIM prior map to obtain a global absolute pose as the observation of the factor. Step 3.3: Construct all factors into a large nonlinear least squares problem and use backend optimizers such as iSAM2 to solve it in real time to obtain the optimized estimate of all robot pose nodes, thereby outputting globally consistent and high-precision positioning results. 。