System and method for training space intelligent large model based on BIM (Building Information Modeling) technology

By constructing a multimodal training data set and using multimodal feature extraction and representation, spatial modeling and self-supervised pre-training technologies, the spatial intelligent big model is trained and generated, which solves the shortcomings of existing BIM technology big models in spatial information processing and environmental interaction, and realizes efficient three-dimensional space cognition, reasoning and interaction, and improves operation and maintenance reliability.

CN120218128APending Publication Date: 2025-06-27CHINA CONSTR EIGHT ENG DIV CORP LTD

Patent Information

Application Number
CN202510313953.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing large models based on BIM technology have shortcomings in spatial information processing capabilities and environmental interactions, and cannot effectively realize spatial cognition, dynamic reasoning and environmental interactions, resulting in low operation and maintenance reliability.

Method used

By constructing a training data set, combining BIM data, point cloud data, RGB-D images and IoT data, multimodal feature extraction and representation, spatial modeling and self-supervised pre-training, training and generation of spatial intelligent large models to realize cognition, reasoning and interaction in three-dimensional space.

Benefits of technology

It improves the operation and maintenance reliability of three-dimensional space and enhances the spatial cognition, dynamic reasoning and environmental interaction capabilities of spatial intelligent large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218128A_ABST
    Figure CN120218128A_ABST
Patent Text Reader

Abstract

The invention discloses a system and method for training a spatial intelligent large model based on a BIM technology, and relates to the technical field of building information processing, a training data set construction module is adopted to fuse BIM data with multi-modal information such as point cloud data, RGB-D images and IoT data, so that the BIM technology is dynamically applied to obtain dynamic three-dimensional data, and the BIM technology is applied to the BIM technology; according to the spatial intelligent large model, an accurate training data set is generated, and meanwhile, spatial cognition, dynamic reasoning and environment interaction training are performed by adopting the training data set through the large model training module, so that the spatial intelligent large model can realize cognition, reasoning and interaction of a three-dimensional space, and the operation and maintenance reliability of the three-dimensional space is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of building information processing, and particularly to a large spatial model based on BIM. Background Art

[0002] Spatial intelligence refers to the understanding, reasoning, prediction, and interaction of three-dimensional space, which is widely applied in fields such as intelligent buildings, smart cities, robot navigation, and digital twins. BIM technology includes information such as building geometric structures, topological relationships, room functions, and physical material information, which can provide rich data for the realization of spatial intelligence. In the prior art, large models based on BIM technology usually focus on static three-dimensional modeling and visualization, and perform image acquisition and natural language processing on BIM.

[0003] For example, a large model based on BIM disclosed in a Chinese patent application with publication number CN 119337992 A performs static semantic processing and decomposition on BIM to obtain and extract BIM rules and specifications, resulting in poor spatial information processing capabilities of the large model, inability to effectively interact with the spatial environment, and low operation and maintenance reliability.

[0004] Therefore, how to obtain a large model with spatial intelligence functions based on BIM technology, realize spatial cognition, dynamic reasoning, and environmental interaction, and effectively improve the operation and maintenance reliability of three-dimensional spaces such as the building environment has become an urgent problem to be solved in this field. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a system and method for training a spatial intelligence large model based on BIM technology, which can realize the cognition, reasoning, and interaction of three-dimensional space through the dynamic application of BIM.

[0006] To achieve the above purpose, the system for training a spatial intelligence large model based on BIM technology provided by the present invention includes

[0007] a training dataset construction module, which includes a BIM parsing unit, a point cloud and RGB-D fusion unit, an IoT data alignment unit, an automatic annotation unit, a data enhancement unit, and a pre-training dataset construction unit. The training dataset construction module is configured to extract BIM data from the BIM model, and perform annotation, enhancement, classification, and fusion processing on the BIM data and multi-modal information to generate a training dataset;

[0008] Large model training module, which includes a multi-modal feature extraction and representation unit, a spatial modeling unit, a self-supervised pre-training unit, a reinforcement learning and simulation optimization unit, and a model evaluation and deployment unit. The large model training module is configured to perform feature extraction and spatial modeling processing on the training data set, and conduct spatial cognition, dynamic reasoning, and environment interaction training to generate a spatial intelligence large model.

[0009] Further, after the BIM parsing unit extracts BIM data and performs normalization and unified coordinate system processing, the point cloud and RGB-D fusion unit geometrically aligns the BIM data with the point cloud data, aligns the coordinate systems of the BIM data and the RGB-D images, and also jointly models the visual features of the RGB-D images and the geometric features of the point cloud data.

[0010] Further, the IoT data alignment unit uses time synchronization technology to align the data streams of different IoT sensors inside the building, and performs temporal alignment and periodic pattern recognition on the IoT data obtained by different IoT sensors through dynamic time warping and Fourier transform respectively, and then maps the IoT data into the BIM model.

[0011] Further, the automatic annotation unit combines self-supervised learning and active learning methods to compare, identify, and learn the spatial patterns of the BIM model, generates semantic labels for the BIM data, and then enhances the BIM data and multi-modal information through the data enhancement unit.

[0012] Further, the pre-training data set construction unit uses the unsupervised clustering K-Means algorithm to classify the BIM data and multi-modal information respectively, and uses a generative adversarial network and a diffusion model to fuse the BIM data and multi-modal information.

[0013] Further, the multi-modal feature extraction and representation unit extracts key information from the BIM data and multi-modal information respectively, and constructs a unified high-dimensional feature representation.

[0014] Further, the spatial modeling unit uses the GNN GraphSAGE algorithm to learn the graph structure information of the BIM data to extract the spatial topological relationship of the building, and uses Transformer for cross-modal learning, and fuses the BIM data and multi-modal information through the self-attention mechanism to establish the global spatial relationship of the building.

[0015] Further, the self-supervised pre-training unit uses the BIM data and multi-modal information to perform spatial cognition training and prediction through contrastive learning and masked auto-encoding.

[0016] Furthermore, the reinforcement learning and simulation optimization unit constructs a building simulation scenario, discretizes the building environment, and performs dynamic reasoning and environment interaction training using BIM data and multi-modal information.

[0017] To achieve the above object, the method for training a spatial intelligent large model based on BIM technology provided by the present invention, and the system for training a spatial intelligent large model based on BIM technology, the method includes:

[0018] First, the training dataset construction module extracts BIM data from the BIM model, and performs annotation, enhancement, classification, and fusion processing on the BIM data with point cloud data, RGB-D images, and IoT data respectively to generate a training dataset;

[0019] Then, the large model training module uses the training dataset for feature extraction and spatial modeling processing, and performs spatial cognition, dynamic reasoning, and environment interaction training to generate a spatial intelligent large model.

[0020] The system and method for training a spatial intelligent large model based on BIM technology provided by the present invention use the training dataset construction module to fuse multi-modal information such as BIM data, point cloud data, RGB-D images, and IoT data, so as to dynamically apply BIM technology, obtain dynamic three-dimensional data, and generate an accurate training dataset. At the same time, the large model training module uses the training dataset for spatial cognition, dynamic reasoning, and environment interaction training, so that the spatial intelligent large model can realize the cognition, reasoning, and interaction of the three-dimensional space, thereby improving the operation and maintenance reliability of the three-dimensional space. Description of the Drawings

[0021] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0022] Figure 1 It is the overall block diagram of the system for training a spatial intelligent large model based on BIM technology provided by the present invention;

[0023] Figure 2 It is the system block diagram of the BIM parsing unit in the present invention;

[0024] Figure 3 It is the system block diagram of the point cloud and RGB-D fusion unit in the present invention;

[0025] Figure 4 It is the system block diagram of the IoT data alignment unit in the present invention;

[0026] Figure 5 It is the system block diagram of the automatic annotation unit in the present invention;

[0027] Figure 6 It is the system block diagram of the pre-training dataset construction unit in the present invention;

[0028] Figure 7 It is the system block diagram of the multi-modal feature extraction and representation unit in the present invention;

[0029] Figure 8 It is the system block diagram of the spatial modeling unit in the present invention;

[0030] Figure 9 It is the system block diagram of the self-supervised pre-training unit in the present invention;

[0031] Figure 10 It is the system block diagram of the reinforcement learning and simulation optimization unit in the present invention. Detailed implementation manners

[0032] In order to make the technical means, creative features, achieved purposes and effects realized by the present invention easy to understand, the present invention will be further described below with reference to specific illustrations.

[0033] See Figure 1 , which shows an example of the system for training a spatial intelligent large model based on BIM technology provided by the present invention.

[0034] The system for training a spatial intelligent large model based on BIM technology in this example mainly includes a training dataset construction module 100 and a large model training module 200.

[0035] The training dataset construction module 100 includes a BIM parsing unit 110, a point cloud and RGB-D fusion unit 120, an IoT data alignment unit 130, an automatic annotation unit 140, a data augmentation unit 150, and a pre-training dataset construction unit 160. The training dataset construction module 100 is configured to extract BIM data from the BIM model, and perform annotation, augmentation, classification, and fusion processing on the BIM data and multi-modal information, so as to dynamically apply BIM technology, obtain dynamic three-dimensional data, and generate an accurate training dataset.

[0036] Furthermore, the large model training module 200 includes a multi-modal feature extraction and representation unit 210, a spatial modeling unit 220, a self-supervised pre-training unit 230, a reinforcement learning and simulation optimization unit 240, and a model evaluation and deployment unit 250. The large model training module 200 is configured to perform feature extraction and spatial modeling processing on the training dataset, and perform spatial cognition, dynamic reasoning, and environment interaction training, so as to generate a spatial intelligent large model, so that the spatial intelligent large model can realize the cognition, reasoning, and interaction of the three-dimensional space, and improve the operation and maintenance reliability of the three-dimensional space.

[0037] Among them, the training dataset construction module 100 constructs an accurate training dataset by using BIM data in combination with multi-modal information, and provides data input for the training of the spatial intelligent large model.

[0038] First, the training dataset construction module 100 extracts BIM data through the BIM parsing unit 110, extracts the BIM data of the building in the BIM model, including key information such as the geometric structure, topological relationship, spatial semantics, and physical attributes of the building, and converts the BIM data into a format that can be trained by the spatial intelligent large model.

[0039] Combined with Figure 2 , specifically, the BIM parsing unit 110 parses the BIM model file through an IFC parser, such as IfcOpenShell, extracts the geometric information of building components, such as walls, doors, windows, floors, rooms, etc., and classifies them hierarchically according to the building component type to obtain the geometric structure of the building.

[0040] Furthermore, the BIM parsing unit 110 constructs the topological structure of the building through a graph database, converts the topological structure of the building space into graph data, where building elements such as rooms, stairs, and corridors are modeled as nodes. At the same time, the access paths and connection relationships between various building elements are modeled as edges, and information such as communication permissions and building functions is stored in the attributes corresponding to the edges, so that building elements such as rooms, floors, access paths, doors, and windows can be modeled as input data for the graph neural network.

[0041] At the same time, the BIM parsing unit 110 extracts the physical attributes of the building, such as room area, wall thickness, floor height, building materials, etc., to enhance the spatial intelligent large model's perception ability of the building environment.

[0042] In this way, the BIM data obtained by the BIM parsing unit 110 has geometric and topological information, which can provide data support for the spatial analysis, path reasoning, and building layout optimization of the spatial intelligent large model, and can also provide high-quality input data for the fusion of multi-modal information and the training of the spatial intelligent large model.

[0043] To ensure the consistency of the BIM data and enable effective fusion with multi-modal information, the BIM parsing unit 110 performs normalization and coordinate system unification processing on the BIM data.

[0044] Specifically, the BIM parsing unit 110 first performs normalization processing, using the Min-Max normalization method, to convert the extracted building parameters, such as room area, floor height, and wall thickness, into a standardized format to avoid the impact of different data ranges on the training of the spatial intelligent large model.

[0045] Meanwhile, for non-numerical data such as materials, structural properties, and functional uses in BIM data, the BIM parsing unit 110 encodes this data using One-Hot encoding and Embedding techniques so that the spatial intelligence large model can effectively parse and learn these features.

[0046] Furthermore, the BIM parsing unit 110 performs coordinate system unification processing on BIM data. Since different BIM software may use different coordinate systems, the BIM parsing unit 110 unifies the coordinate systems of all BIM data based on the EPSG:3857 global coordinate standard to ensure that all BIM data can be calculated under the same spatial reference, improving the consistency of BIM data.

[0047] Meanwhile, to ensure the alignment of BIM data with multi-modal information, such as point cloud data, RGB-D images, and IoT data, for subsequent fusion processing, the BIM parsing unit 110 uses the ICP (Iterative Closest Point) algorithm for point cloud registration and integrates the PnP (Perspective-n-Point) method to calculate the external parameters of RGB-D images, enabling multi-modal information from different sources to be spatially aligned under the same coordinate system.

[0048] In this way, through normalization and coordinate system unification processing, the BIM parsing unit 110 can improve the consistency of BIM data and ensure seamless fusion of BIM data with multi-modal information, providing a more accurate training dataset for the spatial intelligence large model.

[0049] Due to the large scale of BIM data, when the spatial intelligence large model performs spatial cognition, reasoning, and interaction, such as in tasks like building analysis and path planning, it is necessary to efficiently query BIM data. To improve the query efficiency, the BIM parsing unit 110 also optimizes the retrieval efficiency of BIM data using spatial indexing technology.

[0050] Specifically, the BIM parsing unit 110 organizes BIM data using the R-Tree spatial index structure, dividing BIM data by spatial range. Each BIM data node corresponds to a region, enabling the spatial intelligence large model to quickly locate building components, query adjacent rooms, and optimize path rules.

[0051] As an example, the BIM parsing unit 110 uses the R-Tree to store BIM data in a hierarchical partitioning manner, stratifying BIM data by building area to reduce query time and complexity. For example, in the personnel navigation task, the spatial intelligence large model can quickly retrieve the shortest passage path, reachability analysis, and emergency evacuation plan based on the spatial index, while integrating the GNN model for optimization processing.

[0052] Furthermore, the BIM parsing unit 110 is also configured with KNN (K-Nearest Neighbor) queries to quickly calculate room adjacency, match equipment maintenance areas, and further optimize building analysis tasks by integrating the A* algorithm for path search based on topological relationships.

[0053] In this way, compared with the traditional SQL-based BIM data query method, the BIM parsing unit 110 optimizes the query of BIM data through spatial indexing technology, integrating R-Tree and KNN, which improves the retrieval efficiency of BIM data, thus effectively supporting real-time calculation, path planning, and environmental modeling of large-scale BIM data to improve the spatial dynamic reasoning efficiency of the spatial intelligence large model.

[0054] Furthermore, the BIM parsing unit 110 also performs BIM semantic analysis and automatic annotation processing on BIM data to improve the intelligence level of BIM data.

[0055] Specifically, the BIM parsing unit 110 automatically extracts the building function information of BIM data through the BIM semantic parsing technology based on BERT in a way that integrates natural language processing and machine learning.

[0056] First, the BIM parsing unit 110 uses the BERT-Graph fusion model to parse the text information in the BIM model, such as room names, uses, structural descriptions, etc., and performs context reasoning by integrating the spatial topological relationship of BIM data.

[0057] For example, when the BIM parsing unit 110 parses the BIM model and a room is named "meeting room" and is adjacent to multiple offices, it can further infer the building function of the "meeting room" as a "public office area" by integrating the spatial topological relationship of BIM data.

[0058] Furthermore, to enhance the automatic annotation ability of BIM data, the BIM parsing unit 110 integrates contrastive learning and uses the annotated BIM data for deep model training to automatically identify room types, building functions, and spatial partitions.

[0059] In addition, the BIM parsing unit 110 also uses the Transformer model for cross-modal learning, integrating the visual information of RGB-D images into BIM data to optimize the semantic annotation of BIM data, thereby identifying key areas such as elevator shafts, fire corridors, and machine rooms in the building.

[0060] In this way, compared with traditional BIM data processing methods where semantic information usually relies on manual annotation, the BIM parsing unit 110 can achieve automated semantic parsing and spatial semantic reasoning of BIM data, providing effective data support for the training of spatial intelligent large models, thereby improving the intelligent level of building information management.

[0061] The BIM parsing unit 110 thus constituted performs BIM data extraction, normalization and coordinate system unification processing, spatial index optimization, semantic analysis and automatic annotation processing respectively to obtain unified and accurate BIM data, providing data input for subsequent data fusion and large model training.

[0062] Furthermore, the training dataset construction module 100 fuses point cloud data and RGB-D images through the point cloud and RGB-D fusion unit 120 to improve the geometric accuracy and visual feature expression ability of BIM data, so as to improve the training effect of the large model.

[0063] As an example, point cloud data is mainly generated by lidar scanning of buildings, including high-precision three-dimensional coordinates, reflection intensity and surface normal vector information to accurately reflect the geometric structure of the building. For example, the actual dimensions, edge contours and curved surface forms of components such as walls, stairs and doors and windows, but lack component color and texture information.

[0064] Furthermore, the RGB-D image can be generated by collecting the building with an Intel RealSense D455 or Azure Kinect camera acquisition device. The RGB-D image includes a color image (RGB) and a depth map (Depth). Among them, the color image provides visual information such as the surface texture, material and color of the building, while the depth map provides the distance information between each pixel point and the camera acquisition device.

[0065] In this way, the point cloud and RGB-D fusion unit 120 fusing point cloud data and RGB-D images can supplement and improve the component color and texture information lacking in the point cloud data, and at the same time obtain distance information. Compared with training the spatial intelligent large model using only BIM data, fusing BIM data with point cloud data and RGB-D images can make the training dataset have spatial consistency and a sense of reality, thereby effectively improving the spatial intelligent large model's ability in spatial cognition, reasoning and interaction.

[0066] Combined Figure 3 For the fusion of BIM data and point cloud data, the point cloud and RGB-D fusion unit 120 first preprocesses the point cloud data to improve the uniformity of the point cloud data.

[0067] Specifically, the point cloud and RGB-D fusion unit 120 uses the RANSAC algorithm to denoise the point cloud, removing noise filtering and outliers to improve the quality of the point cloud data. At the same time, the point cloud and RGB-D fusion unit 120 uses the Voxel Grid Filter to downsample the point cloud to reduce the computational burden of the point cloud data, making the point cloud data sparser while still maintaining key geometric features and improving computational efficiency.

[0068] Furthermore, the point cloud and RGB-D fusion unit 120 extracts surface points of BIM data to ensure that the BIM data is comparable to the point cloud data.

[0069] Specifically, the point cloud and RGB-D fusion unit 120 extracts the corresponding geometric surface point set from the BIM model and converts the geometric surface point set into a three-dimensional point cloud format so that the BIM data can be compared with the real point cloud data.

[0070] Next, the point cloud and RGB-D fusion unit 120 performs an initial alignment of the point cloud data and the BIM data to facilitate subsequent ICP point cloud alignment processing.

[0071] As an example, the point cloud and RGB-D fusion unit 120 uses principal component analysis to calculate the main directions of the point cloud data and the BIM, and performs a preliminary rough registration to roughly align the point cloud data and the BIM data, providing better initial conditions for subsequent ICP iterations, improving the ICP convergence speed, and thus improving the alignment effect.

[0072] In cooperation with this, the point cloud and RGB-D fusion unit 120 performs ICP point cloud alignment processing. First, the point cloud and RGB-D fusion unit 120 calculates the nearest neighbor point pairs of the point cloud data and the BIM surface points through the (Iterative Closest Point) ICP algorithm, and then uses the iterative closest point matching and Euclidean distance minimization methods to gradually optimize the point cloud alignment accuracy, thereby obtaining a locally optimized matching result of the point cloud data and the BIM data.

[0073] Since there may be drift errors in the locally optimized matching result of the point cloud data and the BIM data, in order to eliminate the errors, the point cloud and RGB-D fusion unit 120 performs global optimization processing.

[0074] As an example, the point cloud and RGB-D fusion unit 120 uses GTSAM (a global optimization algorithm based on factor graphs) for optimization. First, a factor graph is constructed, with the rigid transformation matrix of the point cloud as the optimization variable and the BIM surface points as constraints. Nonlinear optimization (such as Levenberg-Marquardt optimization) is used for global error correction, thereby obtaining optimized globally consistent point cloud data, which can significantly reduce registration drift and improve matching accuracy.

[0075] In this way, the point cloud and RGB-D fusion unit 120 highly accurately fuses the point cloud data with the BIM data, enabling the point cloud data and the BIM data to be fused for tasks such as intelligent analysis, robot navigation, and automatic construction inspection, improving the geometric authenticity and spatial consistency of the BIM data, and thus optimizing the quality of the training dataset of the spatial intelligence large model.

[0076] For the fusion of BIM data and RGB-D images, the point cloud and RGB-D fusion unit 120 aligns the feature points of the RGB-D image and the BIM data through SIFT / ORB / SuperGlue feature matching algorithms, and uses the PnP (Perspective-n-Point) algorithm to calculate the external parameters of the RGB-D camera to ensure that the RGB-D image can be aligned with the coordinate system of the BIM data.

[0077] Specifically, the point cloud and RGB-D fusion unit 120 first performs feature point extraction. For the RGB-D image, the point cloud and RGB-D fusion unit 120 uses the SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF) algorithm to detect the key points of the RGB-D image and generate feature descriptors, such as feature vectors. At the same time, the point cloud and RGB-D fusion unit 120 extracts the corresponding geometric feature point sets from the BIM data, such as building facades or room floor plans, thereby obtaining the key point coordinates and feature vectors of the RGB-D image and forming a preliminary geometric feature point set with the BIM data.

[0078] Furthermore, the point cloud and RGB-D fusion unit 120 performs feature point matching on the BIM data and the RGB-D image. The point cloud and RGB-D fusion unit 120 uses the SuperGlue deep learning matching algorithm to perform key point matching between the RGB-D image and the BIM data. SuperGlue uses a self-attention mechanism to calculate cross-modal feature point associations to improve the accuracy of matching, thereby obtaining high-confidence matching point pairs between the RGB-D image and the BIM data and reducing noise point interference.

[0079] Next, the point cloud and RGB-D fusion unit 120 performs PnP extrinsic parameter calculation. As an example, based on the matching points, the point cloud and RGB-D fusion unit 120 uses the PnP (Perspective-n-Point) algorithm to calculate the extrinsic parameters of the camera acquisition device of the RGB-D image in the BIM data coordinate system, such as the rotation matrix R and the translation vector T, and adopts the RANSAC (Random Sample Consensus) method to eliminate the mismatched points to improve the stability of the calculation, so as to obtain the accurate position and orientation of the camera acquisition device in the BIM data coordinate system, which is convenient for subsequent coordinate alignment.

[0080] Correspondingly, the point cloud and RGB-D fusion unit 120 uses the extrinsic parameters (rotation matrix R and translation vector T) of the camera acquisition device obtained by PnP extrinsic parameter calculation for coordinate transformation and optimization, converts the RGB-D image and point cloud data to the global coordinate system corresponding to the BIM data, and further adjusts the alignment accuracy through the nonlinear optimization Levenberg-Marquardt to ensure that the RGB-D image is completely consistent with the BIM data.

[0081] In this way, the point cloud and RGB-D fusion unit 120 can accurately align the RGB-D image and the BIM data, so that the building texture, geometric structure and point cloud data form a unified three-dimensional space information, thereby improving the cognitive ability and training accuracy of the spatial intelligent large model for the building environment.

[0082] To improve the fusion effect of the BIM data with the point cloud data and the RGB-D image, the point cloud and RGB-D fusion unit 120 uses deep learning combined with the network DGCNN (Dynamic Graph Convolutional Neural Network) to jointly model the visual features of the RGB-D image and the geometric features of the point cloud data to improve the cross-modal learning ability of the spatial intelligent large model.

[0083] Specifically, the point cloud and RGB-D fusion unit 120 first performs feature extraction on the point cloud data and the RGB-D image. It uses PointNet++ to extract the geometric features of the point cloud data, including normal vector, curvature, point density and morphological information, to obtain the local geometric features of each point. At the same time, it uses ResNet-50 to extract the visual features of the RGB-D image, including texture, color and material information, to generate a high-dimensional visual feature vector, so as to obtain the independent feature vectors of the point cloud data and the RGB-D image, which is convenient for subsequent cross-modal feature alignment.

[0084] In coordination with this, in the cross-modal feature alignment stage, the point cloud and RGB-D fusion unit 120 adopts the Feature Alignment method. Through the Transformer attention mechanism, it learns the correlation between the RGB-D image and the point cloud data in the spatial structure, and uses the deep self-supervised method (such as contrast learning) to improve the feature matching degree between the point cloud data and the RGB-D image, enabling the point cloud data and the RGB-D image to learn in the same feature space, thereby obtaining a spatially consistent cross-modal feature representation, reducing the feature deviation between the RGB-D image and the point cloud data, and improving the fusion effect of the training dataset.

[0085] Furthermore, the point cloud and RGB-D fusion unit 120 performs DGCNN modeling on the point cloud data and the RGB-D image. As an example, the point cloud and RGB-D fusion unit 120 sends the point cloud data and the RGB-D features into the DGCNN, enabling the model to dynamically construct a K-nearest neighbor graph (KNN Graph) and adopt EdgeConv operations to extract the local and global features of the point cloud data and the RGB-D image, enabling the DGCNN to learn the mapping relationship between the local geometric relationship and the RGB-D semantic information.

[0086] At the same time, the DGCNN can dynamically update the adjacency relationship of points, understand the interaction mode between the spatial structure and visual information at different scales, so that the point cloud and RGB-D fusion unit 120 can obtain a deep feature vector combining geometric and visual information, thereby enhancing the spatial intelligent large model's cognitive understanding ability of space.

[0087] Compared with the traditional method based only on geometric matching or feature mapping, the point cloud and RGB-D fusion unit 120 uses DGCNN for joint modeling of the point cloud data and the RGB-D image, which can learn on the dynamic adjacency graph and improve the feature alignment accuracy.

[0088] To improve the fusion effect, the point cloud and RGB-D fusion unit 120 finally uses a multi-modal Transformer for cross-modal feature combination and optimization processing, further jointly learns the visual features of the RGB-D image and the geometric features of the point cloud data, and adopts a combination of the KL divergence loss function and the contrast learning loss for optimization, enabling the spatial intelligent large model to automatically learn the information complementary relationship between different modalities, and thus being able to more accurately perceive and understand the building space.

[0089] Compared with the traditional splicing multi-modal fusion, the point cloud and RGB-D fusion unit 120 using the Transformer mechanism for cross-modal combination can more effectively capture the global connection between the RGB-D image and the point cloud data, avoiding the loss of modal information.

[0090] In this way, the point cloud and RGB-D fusion unit 120 can effectively improve the fusion effect of BIM data with point cloud data and RGB-D images, and enhance the cross-modal learning ability of the spatial intelligence large model.

[0091] The point cloud and RGB-D fusion unit 120 thus constituted respectively performs the fusion of BIM data and point cloud data, the fusion of BIM data and RGB-D images, and the joint modeling of point cloud data and RGB-D images, so as to improve the fusion effect of BIM data with point cloud data and RGB-D images, and enhance the geometric accuracy and visual feature expression ability of BIM data.

[0092] In order to dynamically apply BIM technology and integrate BIM data with the spatial environment, the component module 100 of the training data set associates the IoT data inside the building, such as dynamic environment data like temperature and humidity, air quality, energy consumption, and human flow monitoring, with BIM data through the IoT data alignment unit 130, so as to enhance the dynamic environment perception ability of the spatial intelligence large model.

[0093] Combined Figure 4 , since IoT data is time series data while BIM data is usually static data, the IoT data alignment unit 130 adopts time synchronization technology and uses NTP (Network Time Protocol) and temporal interpolation to align the data streams of different IoT sensors inside the building, so that different IoT data match the same timestamp.

[0094] First of all, the IoT data alignment unit 130 synchronizes the time of the data streams of different IoT sensors inside the building, such as temperature and humidity sensors, air quality monitoring devices, energy consumption metering devices, and lighting control systems, through NTP (Network Time Protocol), to ensure that all IoT data is based on a unified time standard.

[0095] Specifically, the IoT data alignment unit 130 obtains the standard timestamp from a high-precision time server, such as GPS time synchronization or atomic clock synchronization, and assigns the standardized timestamp to all IoT data, eliminating the error caused by the asynchronous clocks of different IoT sensors.

[0096] Furthermore, since the sampling frequencies of different IoT sensors may be different, for example, the temperature and humidity sensor samples once per second while the energy consumption sensor samples once per minute, the IoT data alignment unit 130 uses temporal interpolation for IoT data filling and alignment.

[0097] Among them, temporal interpolation includes linear interpolation and spline interpolation. For equally spaced IoT data, the IoT data alignment unit 130 uses the linear interpolation method. For non-uniformly spaced IoT data, the IoT data alignment unit 130 uses the spline interpolation method to ensure that the IoT data obtained by different IoT sensors have corresponding measured values at the same time interval.

[0098] Meanwhile, for missing data points, the IoT data alignment unit 130 uses moving average or Kalman filtering for smoothing and completion to eliminate the noise in the IoT data.

[0099] In this way, the IoT data alignment unit 130 can obtain the aligned time series data of different IoT sensors, enabling all IoT data to be synchronized with the BIM data, providing an accurate time reference for subsequent fusion, effectively solving problems such as time asynchronization, sampling frequency mismatch, and data missing of multiple IoT sensors, improving the temporal consistency of the building interior environment data, and providing more accurate inputs for the spatial intelligence large model.

[0100] Furthermore, for different types of IoT data, the IoT data alignment unit 130 uses Dynamic Time Warping (DTW) for IoT data temporal alignment and combines Fourier transform to identify the periodic patterns of IoT data to optimize the fusion effect of IoT data.

[0101] Specifically, the IoT data alignment unit 130 uses Dynamic Time Warping to calculate the optimal matching path between different time series, enabling the IoT data of different IoT sensors to be aligned even if there is an offset on the time axis.

[0102] For example, the readings of a certain temperature and humidity sensor may have a lag effect, and Dynamic Time Warping can adjust the time series of this temperature and humidity sensor to make the data of this temperature and humidity sensor consistent with the data of the air quality sensor.

[0103] Furthermore, by calculating the Euclidean distance and finding the optimal matching path, Dynamic Time Warping can achieve the best match of different IoT data on the time axis, thereby obtaining the adjusted time series data and improving the alignment accuracy of different IoT data.

[0104] To identify the periodic patterns of IoT data, such as daily temperature and humidity changes, energy consumption fluctuations, etc., the IoT data alignment unit 130 uses Fourier transform (FT) to convert the time series data into frequency domain signals to analyze the main frequency components of the IoT data.

[0105] For example, the air quality sensor may exhibit specific change patterns during the morning and evening rush hours every day, while the energy consumption data may show periodic fluctuations every hour.

[0106] Furthermore, these long-term trends and periodic features can be extracted through Fourier transform and used to predict the dynamic changes in the future building environment.

[0107] In this way, the IoT data alignment unit 130 can obtain IoT data with time alignment and periodic analysis, which can be used for intelligent regulation of the building environment and improve the energy efficiency optimization and automation management capabilities of building operation and maintenance.

[0108] Compared with traditional IoT data which is usually based on simple linear interpolation, the IoT data alignment unit 130 combines dynamic time warping and Fourier transform, which can more accurately adjust the alignment of IoT data and BIM data, extract the periodic features in the time series, and effectively improve the time series prediction ability of IoT data.

[0109] Next, the IoT data alignment unit 130 performs the fusion of IoT data and BIM data. To achieve the fusion, the IoT data alignment unit 130 adopts a spatial heat map generation method based on graph convolutional network to map the IoT data into the three-dimensional space of the BIM model, so that the spatial intelligent large model can perform real-time perception and prediction in the BIM environment.

[0110] Specifically, the IoT data alignment unit 130 first uses the spatial coordinate data of the BIM data to assign a three-dimensional position (X, Y, Z) to each IoT sensor to establish the mapping relationship from the IoT sensor to the building space. For example, the temperature and humidity sensors may be located on the ceiling of the building, the air quality sensors are located at the ventilation duct openings, and the energy consumption metering devices are connected to specific building equipment, so as to obtain the spatial position information of different IoT sensors and ensure that all IoT data points can be calculated in the three-dimensional environment.

[0111] Furthermore, the IoT data alignment unit 130 adopts graph structure modeling based on GCN, regarding the rooms, corridors, ventilation systems, and equipment nodes inside the building as the nodes (Nodes) of the graph, the IoT data measured by the IoT sensors as the node attributes, and calculating the propagation mode of the IoT data in the building environment according to the spatial adjacency relationship, such as the air flow path, the temperature and humidity diffusion direction, etc.

[0112] For example, the temperature change in a certain room may affect the adjacent rooms. Therefore, the IoT data alignment unit 130 constructs an adjacency matrix based on the temperature gradient and air flow, and uses GCN for learning to predict the propagation mode of environmental changes in the building space.

[0113] Next, the IoT data alignment unit 130 uses a GCN-based spatial heatmap generation method to perform real-time prediction and visualization of the building's internal environment. For example, the PM2.5 readings of air quality sensors can be mapped into a heatmap of the building, enabling the spatial intelligent model to identify pollution sources and optimize ventilation in a 3D BIM environment.

[0114] In this way, the IoT data alignment unit 130 can obtain the IoT data heatmap in the BIM coordinate system, enabling the dynamic changes of the building environment to be presented in real time in the spatial intelligent model and improving the control level of the building.

[0115] Different from the traditional method of only relying on data table analysis of IoT data, the IoT data alignment unit 130 fuses IoT data with BIM data and uses a graph neural network (GCN) for building environment prediction to achieve spatial intelligent reasoning of IoT data.

[0116] The following takes the temperature and humidity sensors and air quality sensors inside the building as examples to illustrate the processing process of the IoT data alignment unit 130.

[0117] The IoT data alignment unit 130 first combines NTP time synchronization with time series interpolation to ensure that all temperature and humidity data and air quality data are aligned under the same timestamp. Next, the IoT data alignment unit 130 uses DTW (Dynamic Time Warping) to adjust the time series of temperature and humidity data and air quality data, making the temperature and humidity data match the air quality data in the time dimension and avoiding errors caused by data lag or asynchronization.

[0118] For example, the temperature and humidity sensor data in an office area may be sampled once a minute, while the air quality data is sampled once every five minutes. DTW can adjust their time intervals to make the measurement data of the two correspond.

[0119] Subsequently, the IoT data alignment unit 130 uses Fourier transform to analyze the periodicity of the data. For example, it is found that the air quality drops in the morning and the temperature and humidity rise in the afternoon, thus predicting possible future trends. Finally, the IoT data alignment unit 130 adopts GCN spatial heatmap generation to map the temperature and humidity data and air quality data into the BIM data environment to form a visual heatmap. The visual heatmap can analyze the changing trend of air quality in real time and perform intelligent adjustment of the HVAC (Heating, Ventilation, and Air Conditioning) system according to IoT data to optimize the comfort of the building environment.

[0120] The thus-formed IoT data alignment unit 130 respectively performs data stream alignment of different IoT sensors, IoT data time series alignment and periodic pattern recognition, and the fusion of IoT data and BIM data, fusing time series data, three-dimensional BIM geometric data, and graph neural network modeling, enabling the spatial intelligence large model to perform dynamic reasoning in the real building environment, providing more accurate intelligent prediction and control capabilities for intelligent buildings, smart cities, equipment operation and maintenance, etc.

[0121] Furthermore, the training dataset construction module 100 further includes an automatic annotation unit 140, which can improve the semantic quality of BIM data, reduce the manual annotation cost of BIM data, and improve the generalization ability of the spatial intelligence large model.

[0122] Combined with Figure 5 , specifically, the automatic annotation unit 140 combines self-supervised learning and active learning methods, enabling the spatial intelligence large model to automatically learn the spatial patterns of BIM data and infer semantic categories based on the BIM data distribution, realizing the automatic annotation of BIM data.

[0123] First, the automatic annotation unit 140 constructs a spatial component feature library to extract geometric information (area, height, shape), topological relationships (room connections, passage paths), and material properties (concrete, glass, steel structure) in BIM data to form feature vectors.

[0124] Next, the automatic annotation unit 140 uses contrastive learning technology to calculate the similarity matrix between room categories by constructing positive samples (similar room pairs, e.g., two meeting rooms) and negative samples (different room types, e.g., office and bathroom).

[0125] As an example, the automatic annotation unit 140 combines a Siamese Network with cosine similarity to measure the feature similarity of different rooms, and uses K-Means clustering or DBSCAN density clustering to infer room categories, thereby obtaining preliminary category labels for rooms such as "office", "meeting room", "bathroom", etc., automatically annotating the BIM data and reducing the dependence on manual annotation.

[0126] To optimize the annotation ability of the spatial intelligence large model, the automatic annotation unit 140 uses a contrastive learning loss function to train the spatial intelligence large model, making rooms of similar categories closer in the feature space and increasing the feature distance between rooms of different categories, thereby improving the automatic classification ability of BIM data.

[0127] In this way, the automatic annotation unit 140 can accurately identify the room types in BIM data, achieve efficient automatic annotation on a large scale of BIM data, and provide high-quality BIM data support for the building management, automatic construction review, and indoor navigation of the spatial intelligence large model.

[0128] Compared with the traditional methods that rely on manual rules or pre-trained models for room classification in BIM data, the automatic annotation unit 140 introduces the combination of contrastive learning and siamese network to learn the similarity of room categories, and accurately identifies room types by calculating the similarity matrices of geometric topological relationships, spatial distributions, and functional features, improving the classification ability of BIM data.

[0129] During the automatic annotation process, since the spatial intelligence large model may have incorrect predictions or low-confidence annotations in the early stage, the automatic annotation unit 140 combines active learning to optimize the automatic annotation, reduce error propagation, and improve the accuracy of annotation.

[0130] First, based on the prediction output of the spatial intelligence large model, the automatic annotation unit 140 calculates the confidence distribution of each room category and sets a confidence threshold. That is, when the prediction confidence of the spatial intelligence large model for a certain room is lower than a certain threshold (such as 80%), the sample is considered an uncertain sample, thus obtaining a list of low-confidence room category samples for subsequent manual review.

[0131] In coordination with this, the automatic annotation unit 140 submits these low-confidence samples to the manual reviewers. The manual reviewers only need to confirm or correct the room categories predicted by the spatial intelligence large model, rather than starting the annotation from scratch, which can effectively improve the annotation efficiency.

[0132] Then, the automatic annotation unit 140 reversely inputs the data after manual review into the training dataset of the spatial intelligence large model, and optimizes the spatial intelligence large model through Fine-tuning to continuously improve the accuracy of room category prediction.

[0133] During this process, the automatic annotation unit 140 adopts methods based on uncertainty sampling or information entropy sampling to dynamically select the data that the spatial intelligence large model most needs to optimize for manual review, so as to reduce the manual annotation cost and maximize the learning efficiency of the model.

[0134] In this way, the automatic annotation unit 140 can gradually improve the accuracy of automatic annotation, optimize the annotation quality of BIM data with the least amount of manual intervention, accelerate the automatic processing of BIM data, and provide high-precision data support for the operation and maintenance of the spatial intelligence large model and spatial intelligence computing.

[0135] Compared with the problem of traditional methods relying on a large amount of manual annotation, the automatic annotation unit 140 combines active learning with confidence screening, only conducts manual review on low-confidence data, and uses a small amount of manual feedback to reverse-optimize the spatial intelligent large model, reducing labor costs while continuously improving the annotation quality of BIM data.

[0136] Furthermore, in order to improve the accuracy of automatic annotation, the automatic annotation unit 140 also adopts BIM semantic parsing based on Transformer, combines the BERT model and BIM data for training to automatically generate semantic tags, and improves the generalization ability of the spatial intelligent large model.

[0137] First, the automatic annotation unit 140 preprocesses the text information in BIM data, such as room names, functional descriptions, component labels, etc., including stop word removal, word segmentation, word vector conversion, etc., and uses the BERT (Bidirectional Encoder Representations from Transformers) model for semantic feature extraction to generate high-dimensional text feature vectors, so as to obtain the deep semantic representation of the room text description in BIM data, enabling the spatial intelligent large model to cognitively understand the space.

[0138] Furthermore, the automatic annotation unit 140 combines the multi-head attention mechanism of the Transformer architecture for cross-modal combination learning of room geometric features, topological structures, and text descriptions.

[0139] For example, the BIM geometric information of a room may indicate that its area is large, the topological relationship shows that the room is connected to multiple offices, and the text description contains the word "meeting". The Transformer model can automatically integrate these features and infer that the function of the room is "meeting room".

[0140] In addition, the automatic annotation unit 140 adopts self-supervised learning to train BERT through masked language modeling, enabling the spatial intelligent large model to predict the missing room categories, thereby improving the generalization ability of the spatial intelligent large model.

[0141] In this way, the automatic annotation unit 140 can accurately parse the semantic information of BIM data, reduce the need for manual annotation, and thus improve the spatial intelligent large model's cognitive understanding ability of space.

[0142] Compared with the method based on geometric rule matching, the automatic annotation unit 140 is based on the Transformer-based BIM semantic parsing, combines BERT with BIM data for training, realizes the integration of text, geometry, and topology multimodality, enables the spatial intelligence large model to perform semantic reasoning based on room names, function descriptions, geometric dimensions, and spatial structures, thereby reducing human intervention and improving the automatic semantic understanding ability of BIM data.

[0143] The thus-formed automatic annotation unit 140 realizes efficient, intelligent, and automated BIM data annotation by means of contrastive learning, active learning, Transformer, and BERT geometric semantic parsing, respectively constructing a spatial component feature library, calculating a similarity matrix, optimizing the automatic annotation, and generating semantic labels. This not only reduces labor costs but also greatly improves the intelligent level of building information modeling and enhances the spatial intelligence large model's cognitive understanding ability of space.

[0144] To improve the diversity and generalization ability of the training dataset, the training dataset construction module 100 performs enhancement processing on BIM data and multimodal information through the data enhancement unit 150.

[0145] Specifically, the data enhancement unit 150 adopts various enhancement strategies according to the different data characteristics of BIM data, point cloud data, RGB-D images, and IoT data.

[0146] For BIM data, the data enhancement unit 150 adopts enhancement techniques such as data perturbation, occlusion simulation, and topological transformation. As an example, the data enhancement unit 150 first uses the data perturbation method to randomly adjust the dimensions, orientations, and storey heights of building components to simulate different architectural design styles. At the same time, it uses occlusion simulation to artificially add occlusions in the BIM data environment, for example, increasing invisible areas or deleting some components to enhance the noise resistance ability of the spatial intelligence large model. In addition, the data enhancement unit 150 allows the adjustment of the access paths, room layouts, and component connection methods of the building space through the topological transformation method to expand the adaptability of BIM data. For example, it simulates different room layout schemes of an office building to obtain various variant BIM data for training a more robust spatial intelligence large model.

[0147] For point cloud data, the data augmentation unit 150 adopts random rotation, scale transformation, noise perturbation, and combines with GAN (Generative Adversarial Network) to generate synthetic data. As an example, the data augmentation unit 150 adopts random rotation and scale transformation to enable the spatial intelligence to adapt to different perspectives and scanning scales, and combines with noise perturbation to make the point cloud data more robust through additive Gaussian noise. In addition, the Generative Adversarial Network (GAN) is used to generate synthetic point cloud data. By learning the real point cloud distribution, GAN can generate more realistic virtual point cloud building models to enhance the learning ability of the spatial intelligence model for scarce data, thereby obtaining richer point cloud data samples and enabling the spatial intelligence model to adapt to different sensors and scanning scenarios.

[0148] For RGB-D images, the data augmentation unit 150 uses color jittering, lighting simulation, and image flipping to enhance data diversity. As an example, the data augmentation unit 150 adopts color jittering, lighting simulation, and image flipping for enhancement to improve the lighting invariance and perspective adaptability of the spatial intelligence model. For example, different lighting conditions are simulated by adjusting brightness, contrast, and saturation to ensure the generalization ability of the spatial intelligence model in different scenarios, thereby obtaining diverse RGB-D images and optimizing the visual understanding ability of the spatial intelligence model.

[0149] For IoT data, the data augmentation unit 150 generates realistic IoT data samples through GAN-based time series data synthesis. By learning the patterns of historical sensing data, realistic time series data such as temperature and humidity, air quality, and energy consumption are generated to fill in missing data and enhance the time series prediction ability of the spatial intelligence model. For example, GAN can generate indoor air quality data under multiple different weather conditions or working time periods to improve the training effect of time series modeling, thereby obtaining more comprehensive IoT data samples and optimizing the environmental perception ability of the spatial intelligence model.

[0150] In this way, the data augmentation unit 150 significantly improves the diversity of BIM data, point cloud data, RGB-D images, and IoT data, and enhances the adaptability of the spatial intelligence model in different environments.

[0151] Furthermore, the training dataset construction module 100 performs classification and synthesis processing on BIM data and multimodal information through the pre-training dataset construction unit 160 to generate a training dataset and improve the generalization ability of the spatial intelligence model.

[0152] Combined Figure 6, specifically, the pre-training dataset construction unit 160 uses the unsupervised clustering K-Means algorithm to classify BIM data and multimodal information, ensuring that the BIM data and multimodal information cover different building types, such as residential, commercial, industrial, and different urban environments, such as high-density cities, suburbs, smart campuses, etc.

[0153] First, the pre-training dataset construction unit 160 extracts feature vectors based on the building geometric features (area, height, number of floors), topological features (spatial layout, access paths), and functional features (room usage, building type) of the BIM data, and uses K-Means for unsupervised classification.

[0154] For example, the pre-training dataset construction unit 160 automatically differentiates office buildings, residential communities, and industrial factories, and optimizes the distribution of the training dataset according to the characteristics of different building types.

[0155] In this way, the pre-training dataset construction unit 160 can obtain a classification training dataset covering a variety of building scenarios, ensuring the generalization ability of the spatial intelligence large model.

[0156] Compared with traditional BIM data that usually relies on manual annotation or static rule classification, resulting in uneven data distribution, the pre-training dataset construction unit 160 introduces an unsupervised clustering (K-Means) method to automatically classify based on building geometry, topological structure, and functional features, ensuring that the BIM dataset covers various building types such as residential, commercial, industrial, and different environments such as high-density cities, suburbs, and smart campuses, improving the representativeness of the training dataset.

[0157] Furthermore, in order to expand the diversity of the training dataset and the spatial cognitive understanding ability of the spatial intelligence large model, the pre-training dataset construction unit 160 uses GAN (Generative Adversarial Network) and Diffusion Model for the synthesis of BIM data and multimodal information for low-resource building data, such as rare architectural styles or special structures, to expand low-resource building data, and combines self-supervised pre-training to enable the spatial intelligence large model to automatically learn building spatial features and improve the generalization ability driven by the training dataset.

[0158] First, the pre-training dataset construction unit 160 uses GAN (Generative Adversarial Network), through the training of the generator and discriminator, to generate realistic BIM building structures, point cloud data, and RGB-D images, such as simulating different building materials, window layouts, and spatial topologies.

[0159] Meanwhile, compared with traditional GANs, the Diffusion Model has stronger detail retention and stability in building data generation. Therefore, the pre-trained dataset construction unit 160 uses the Diffusion Model to generate high-quality data in a step-by-step denoising manner for simulating complex building environments, special lighting conditions, or unseen building types, thereby obtaining richer synthetic data and effectively expanding the diversity of the training dataset.

[0160] In this way, the pre-trained dataset construction unit 160 can effectively solve the problem of insufficient low-resource building data and expand the diversity of the training dataset.

[0161] Compared with the traditional method of only relying on data augmentation to expand the dataset, the pre-trained dataset construction unit 16 combines the GAN generator and the step-by-step denoising modeling of the Diffusion Model to generate more realistic and detail-rich BIM data and multi-modal information, solving the problem of insufficient low-resource building data.

[0162] Finally, the pre-trained dataset construction unit 160 combines self-supervised pre-training, enabling the spatial intelligence large model to learn on the unlabeled training dataset and training the spatial intelligence large model to automatically fill in missing data through the masked reconstruction task, improving the understanding ability of building space features.

[0163] The pre-trained dataset construction unit 160 thus constituted conducts classification synthesis of BIM data and multi-modal information, expands the diversity of the training dataset, and performs self-supervised training, capable of generating a rich training dataset, thereby improving the generalization ability of the spatial intelligence large model.

[0164] Therefore, the training dataset construction module 100, through the cooperation of the BIM parsing unit 110, the point cloud and RGB-D fusion unit 120, and the IoT data alignment unit 130, fuses BIM data with point cloud data, RGB-D images, and IoT data, and can dynamically apply BIM to improve the accuracy of the training dataset. Meanwhile, through the automatic annotation unit 140, the data augmentation unit 150, and the pre-trained dataset construction unit 160, the BIM data and multi-modal information are annotated, enhanced, classified, and fused to improve the generalization ability of the training dataset for the spatial intelligence large model.

[0165] In order to use the training dataset for training and enable the spatial intelligence large model to achieve cognition, reasoning, and interaction in the three-dimensional space, the large model training module 200 first extracts features in the training dataset through the multi-modal feature extraction and representation unit 210 to provide a standardized input for the spatial intelligence large model.

[0166] Combined with Figure 7, specifically, the multi-modal feature extraction and representation unit 210 extracts features from BIM data, point cloud data, RGB-D images, and IoT data to ensure that the spatial intelligence large model can fully understand the spatial structure, visual information, and dynamic change patterns of the built environment.

[0167] For BIM data, the multi-modal feature extraction and representation unit 210 extracts geometric features from BIM data through the graph neural network GCN, obtains the topological structure of the building, room connection relationships, floor layouts, and building component attributes, such as wall thickness, room area, etc., and constructs a graph representation of the building space, enabling the spatial intelligence large model to understand the internal logical relationships of the building, thereby obtaining standardized BIM spatial topological features, providing a data basis for the spatial cognition, dynamic reasoning, and prediction of the spatial intelligence large model, such as intelligent navigation and path planning.

[0168] For point cloud data, the multi-modal feature extraction and representation unit 210 uses PointNet++ to extract structural information features, including the morphology of the building surface, edge features, component curvature, and three-dimensional spatial distribution, and ensures the alignment of the point cloud data with the BIM data through the ICP (Iterative Closest Point) registration algorithm to obtain high-precision three-dimensional geometric features of the building, which are used to enhance the spatial perception ability of the spatial intelligence large model.

[0169] For RGB-D images, the multi-modal feature extraction and representation unit 210 uses ResNet-50 for visual feature extraction, extracts visual textures, materials, and lighting information, and at the same time uses the depth map combined with the PnP algorithm to calculate the external parameters of the camera acquisition device to ensure that the RGB-D image can be accurately mapped into the coordinate system of the BIM data, thereby obtaining building visual features containing color, material, and depth information, which are used to optimize the understanding ability of the spatial intelligence large model for the built environment..

[0170] For IoT data, the multi-modal feature extraction and representation unit 210 extracts temporal features. Through time synchronization processing (Time Synchronization), the NTP protocol is combined with temporal interpolation to align the IoT data of different IoT sensors in time, ensuring that all IoT data is comparable on the same time scale. At the same time, the multi-modal feature extraction and representation unit 210 uses dynamic time warping (DTW) for temporal alignment to correct the time drift between different IoT sensors. For example, the time lag between temperature and humidity data and air quality data is compared. Finally, the multi-modal feature extraction and representation unit 210 inputs the IoT data into an LSTM (long short-term memory network) to learn the temporal patterns of the building environment. For example, predicting the future trends of indoor temperature changes and energy consumption trends, thereby improving the prediction ability of the spatial intelligence large model for dynamic environments. Thus, standardized and time-aligned IoT data temporal features are obtained, ensuring that the spatial intelligence large model can predict environmental changes based on historical data.

[0171] In this way, the multi-modal feature extraction and representation unit 210 can extract the features of BIM data, point cloud data, RGB-D images, and IoT data respectively, ensuring that the spatial intelligence large model can fully understand the spatial structure, visual information, and dynamic change patterns of the building environment.

[0172] Compared with traditional methods that only rely on single-modal data, such as BIM geometric structures or RGB images, the multi-modal feature extraction and representation unit 210 jointly extracts through BIM geometric features, RGB-D image visual features, point cloud data structure information, and IoT data temporal features, enabling the spatial intelligence large model to comprehensively understand the static structure and dynamic changes of the building environment.

[0173] At the same time, for IoT data, the multi-modal feature extraction and representation unit 210 introduces the combination of time synchronization and dynamic time warping (DTW) and LSTM long short-term memory network, which not only ensures the time consistency of multi-IoT data but also enhances the dynamic temporal reasoning ability of the spatial intelligence large model by learning the trends of the building environment through long-term memory, such as predicting indoor temperature or energy consumption patterns.

[0174] To enable the spatial intelligence large model to make full use of data from different modalities, the multi-modal feature extraction and representation unit 210 combines the features of BIM data, point cloud data, RGB-D images, and IoT data through a multi-modal fusion network MLP + cross-modal attention mechanism to generate a unified feature representation, providing a standardized input for subsequent modeling of the spatial intelligence large model.

[0175] First, the multi-modal feature extraction and representation unit 210 normalizes the features of each modality, that is, it uniformly converts the spatial topological relationship of BIM data, the three-dimensional geometric information of point cloud data, the texture and depth features of RGB-D images, and the temporal pattern of IoT data into high-dimensional vector representations for cross-modal fusion, so as to obtain the normalized single-modal feature vectors and ensure that different data types have the same input format.

[0176] Furthermore, the multi-modal feature extraction and representation unit 210 uses an MLP (Multi-Layer Perceptron) network to perform preliminary fusion of multi-modal features. The MLP maps the features of each modality through a fully connected layer + ReLU activation function to learn the correlation between different modalities, and assigns different importance weights to different modality features through adaptive weighting, so as to obtain the multi-modal feature vector after preliminary fusion, but there is still a semantic gap between modalities.

[0177] Therefore, the multi-modal feature extraction and representation unit 210 uses a cross-modal attention mechanism for deep fusion. The cross-modal attention mechanism calculates the information dependence relationship between different modalities based on the self-attention of Transformer, so as to fuse multi-modal features.

[0178] In this way, the multi-modal feature extraction and representation unit 210 can effectively fuse multi-modal features.

[0179] Compared with the problem of modal information loss caused by traditional feature splicing methods, the multi-modal feature extraction and representation unit 210 uses a combination of MLP (Multi-Layer Perceptron) and cross-modal attention mechanism for multi-modal fusion, enabling the BIM data and multi-modal information to be mutually correlated and improving the utilization efficiency of the training data set.

[0180] In cooperation with it, the spatial modeling unit 220 uses the GNN GraphSAGE algorithm to learn the graph structure information of BIM data to extract the topological relationship of the building space, enabling the spatial intelligent large model to understand the hierarchical relationship between rooms, the passage paths between floors, and the structural connection methods, etc., and enhancing the spatial cognitive ability of the spatial intelligent large model.

[0181] Combined Figure 8 , first, the spatial modeling unit 220 converts the BIM data into a graph data structure (Graph Representation), takes rooms, floors, passages, doors and windows as nodes (Nodes), takes the passage paths between rooms, stairs connecting floors, and the connection methods of doors and windows and passages as edges (Edges), and takes the attribute information of building components (such as room area, floor height, functional use) as node features (Node Features), so as to obtain the graph representation of BIM data and provide input for GNN to learn the building space structure.

[0182] Next, the spatial modeling unit 220 uses a GNN to perform information transmission from the local spatial neighborhood to the global building topology through a node aggregation + neighbor information update mechanism. For example, the function of a room not only depends on its own attributes but is also affected by adjacent rooms. The GNN can learn these spatial relationships to more accurately infer the global topology of the building.

[0183] Furthermore, the spatial modeling unit 220 introduces a Graph Attention Mechanism (GAT) to learn the association strength between building components through attention weights, such as the passage frequency between the main corridor and each room, thereby optimizing the representation ability of the building space structure.

[0184] In this way, the spatial modeling unit 220 can obtain the topological relationships of the building space, enabling the spatial intelligent large model to understand the hierarchical relationships between rooms, the passage paths between floors, and the structural connection methods, etc., and enhancing the spatial cognitive ability of the spatial intelligent large model.

[0185] Furthermore, the spatial modeling unit 220 uses a Transformer for cross-modal learning. By combining the self-attention mechanism with the BIM geometric features, RGB-D image visual features, point cloud data structure information, and temporal features of IoT data obtained by the multi-modal feature extraction and representation unit 210, the global spatial relationships of the building are established.

[0186] First, the spatial modeling unit 220 calculates the spatial matching degree between the geometric features of BIM data and the point cloud data using self-attention to ensure the alignment of different data sources in the coordinate system, structural information, and spatial topology, thereby obtaining the aligned BIM-point cloud spatial features and reducing data deviation.

[0187] Furthermore, the spatial modeling unit 220 uses a Transformer to calculate the correlation between the visual features of RGB-D images and the geometric features of BIM data through cross-modal feature combination. For example, by predicting the functional categories of building components through the texture, color, and depth features of RGB-D images, the spatial intelligent large model can automatically identify room types such as offices, meeting rooms, and corridors.

[0188] In addition, the spatial modeling unit 220 also uses the temporal features of IoT data and combines the time series modeling ability of the Transformer to predict the energy consumption trend, temperature and humidity changes, and air quality fluctuations in the building environment, and optimize the building operation and maintenance plan.

[0189] In this way, the spatial modeling unit 220 can perform cross-modal learning on BIM geometric features, RGB-D image visual features, point cloud data structure information, and temporal features of IoT data, establish the global spatial relationship of the building, and facilitate subsequent spatial cognitive training of the spatial intelligent large model.

[0190] To conduct spatial cognitive training, the self-supervised pre-training unit 230 uses a method that combines contrastive learning and masked autoencoders (MAE) for spatial cognitive training, enabling the spatial intelligent large model to autonomously learn the spatial structure without labeled data.

[0191] Combined Figure 9 , first of all, the self-supervised pre-training unit 230 adopts graph structure masking learning of BIM data, randomly masking some nodes or edges of building components. For example, deleting the connection information of some rooms or masking the passage paths, and requiring the spatial intelligent large model to predict the masked information, so as to learn the passage relationship between rooms and the connectivity of building topology, and thus obtain a BIM data representation with enhanced spatial topology prediction ability, enabling the spatial intelligent large model to still perform reasonable spatial reasoning in an unknown building environment.

[0192] As an example, the self-supervised pre-training unit 230 converts BIM data into a graph data structure, takes rooms, floors, doors and windows, and passage paths as nodes, takes the connection relationship between rooms, the vertical passage path of stairwells, and the opening connection method of doors and windows as edges, and takes information such as room area, functional use, and floor height as node attributes (Node Features); then, adopts a random masking method to mask the topological information of some building spaces, such as hiding the passage path of a certain room or deleting the structural connection information of some rooms, so that the spatial intelligent large model must infer the missing structure through known information during the learning process.

[0193] Next, the self-supervised pre-training unit 230 combines graph neural network (GNN) with normalized graph Laplacian transform for topology learning, and uses the message passing mechanism to infer the spatial structure of the masked part from adjacent nodes, ensuring that the spatial intelligent large model can learn the hierarchical relationship, spatial layout, and passage path between rooms.

[0194] In this way, the self-supervised pre-training unit 230 can obtain a complete building topology inference ability, enabling the spatial intelligent large model to still accurately infer the relationship between rooms in the case of missing part of the data or an unknown building environment, and improving the adaptability of the spatial intelligent large model in intelligent navigation, building information completion, and spatial planning tasks.

[0195] Furthermore, the self-supervised pre-training unit 230 uses point cloud data for contrastive learning. By constructing similar and dissimilar point cloud pairs, the spatial intelligence large model can learn the geometric features of different building components, such as differentiating walls, columns, and ceilings, improving geometric recognition accuracy, thereby obtaining a high-precision semantic understanding ability of point cloud data, and optimizing the application of the spatial intelligence large model to point cloud data in the fields of intelligent construction and 3D modeling.

[0196] As an example, the self-supervised pre-training unit 230 performs local sampling (Farthest Point Sampling, FPS) on the point cloud data, extracts the geometric features of the building surface (walls, columns, floors), and then constructs positive and negative sample pairs. For example, point clouds at different scanning angles of the same building are used as positive samples, and point cloud data of different buildings are used as negative samples; then, PointNet++ or DGCNN (Dynamic Graph Convolutional Neural Network) is used for point cloud encoding, and through a contrastive loss function (increasing the feature similarity of similar building components and reducing the feature similarity of dissimilar building components, finally obtaining the geometric feature representation of building components and improving the point cloud recognition ability.

[0197] In this way, the self-supervised pre-training unit 230 can learn and distinguish the geometric features of different building components through point cloud data.

[0198] Correspondingly, the supervised pre-training unit 230 uses RGB-D images for cross-modal contrastive learning. Utilizing the correspondence between BIM data and RGB-D images, it requires the spatial intelligence large model to predict the matching situation between BIM geometric features and RGB-D images, enabling the spatial intelligence large model to automatically understand room types, structural features, and spatial layouts, thereby obtaining the cross-modal semantic association ability of RGB-D images and enhancing the spatial cognition and visual perception abilities of the spatial intelligence large model.

[0199] As an example, the supervised pre-training unit 230 selects the visual information of a certain room in the RGB-D image, and extracts the corresponding room geometric form and room use from the BIM data as positive samples, while the BIM data of different rooms are used as negative samples. Then, the Transformer cross-modal attention mechanism is used to calculate the similarity between the RGB-D image and the BIM data, enabling the spatial intelligence large model to learn how to match visual information and BIM spatial information.

[0200] In this way, the supervised pre-training unit 230 can enhance the spatial cognition and visual perception abilities of the spatial intelligence large model through RGB-D images.

[0201] The self-supervised pre-training unit 230 thus constituted can effectively use the training dataset to conduct spatial cognition training on the spatial intelligence large model, enabling the spatial intelligence large model to understand the three-dimensional space.

[0202] In order to improve the adaptability of the spatial intelligent big model in intelligent navigation, path planning, and energy consumption optimization tasks, the reinforcement learning and simulation optimization unit 240 constructs a virtual building scene and simulates the physical environment of the building, and uses a training data set to perform dynamic reasoning and environmental interaction training on the spatial intelligent big model.

[0203] Combination Figure 10 Specifically, the reinforcement learning and simulation optimization unit 240 builds a high-precision building simulation environment based on Unity3D / Unreal Engine to simulate the real building physical environment, including indoor layout, crowd simulation, air quality changes, energy consumption and other physical parameters, so that the spatial intelligent large model can learn how to optimize paths, avoid obstacles, reduce energy consumption and other strategies through interaction with the simulation environment, thereby realizing environmental interaction.

[0204] As an example, the reinforcement learning and simulation optimization unit 240 simulates an office building scene based on Unity3D / Unreal Engine, considers the impact of population density during peak hours on the traffic path, and simulates temperature, humidity, and air circulation conditions through sensor data.

[0205] Furthermore, the reinforcement learning and simulation optimization unit 240 adopts the fusion of reinforcement learning (RL) and agent training, and continuously tries different path selection and energy consumption control schemes in the simulation environment through the spatial intelligence big model to optimize the optimal strategy.

[0206] In this way, the reinforcement learning and simulation optimization unit 240 conducts environmental interaction training on the spatial intelligence big model, and can obtain a spatial intelligence big model with stronger environmental adaptability, which can be used for tasks such as intelligent building operation and maintenance, drone cruising, and robot autonomous navigation.

[0207] Furthermore, in the path planning task, the reinforcement learning and simulation optimization unit 240 uses the reinforcement learning algorithm Deep Q-Network to optimize the optimal walking route by continuously trying different path options, thereby improving the dynamic reasoning ability of the spatial intelligent large model in terms of obstacle avoidance and path efficiency.

[0208] First, the reinforcement learning and simulation optimization unit 240 discretizes the building environment, takes rooms, passages, elevators, stairs, etc. as states, takes moving forward, turning, avoiding obstacles, and waiting as actions, and defines path efficiency, travel time, and energy consumption as reward functions.

[0209] Next, the reinforcement learning and simulation optimization unit 240 uses DQN to optimize the Q-learning strategy, continuously tries different paths, and converges to the optimal path planning solution after multiple rounds of iterations.

[0210] In this way, the reinforcement learning and simulation optimization unit 240 performs dynamic inference training on the spatial intelligent large model, enabling the spatial intelligent large model to make intelligent decisions on walking routes and optimize robot navigation and emergency evacuation path planning within a building.

[0211] The thus-formed reinforcement learning and simulation optimization unit 240 uses the training dataset for dynamic inference and environment interaction training, enabling the spatial intelligent large model to reason and interact with the three-dimensional space.

[0212] To ensure the stability of the spatial intelligent large model, the model evaluation and deployment unit 250 evaluates the spatial intelligent large model using multiple metrics such as spatial understanding accuracy, path planning success rate, and dynamic environment prediction accuracy, ensuring the feasibility of the spatial intelligent large model in actual application scenarios.

[0213] Specifically, the model evaluation and deployment unit 250 tests the spatial intelligent large model on different tasks (such as personnel flow prediction) in real building environments and simulation environments respectively to evaluate the spatial intelligent large model.

[0214] For example, the model evaluation and deployment unit 250 tests the model's spatial reasoning ability in unknown building environments, its travel efficiency in path planning tasks, and its energy consumption optimization effect in intelligent building regulation. And through multiple rounds of simulation testing + real data evaluation, it ensures that the spatial intelligent large model can efficiently execute tasks in real scenarios.

[0215] After training is completed, the spatial intelligent large model is deployed to one or more of the intelligent building management system, robot navigation system, and smart city operation platform, supporting edge computing devices to perform cognition, reasoning, and interaction with the three-dimensional space through the spatial intelligent large model, enabling the spatial intelligent large model to operate stably.

[0216] Thus, the system for training a spatial intelligent large model based on BIM technology provided by the present invention is constituted.

[0217] The present invention also provides a method for training a spatial intelligent large model based on BIM technology. Based on the system for training a spatial intelligent large model based on BIM technology constituted by the above solution, this method includes:

[0218] First, the training dataset construction module 100 extracts BIM data from the BIM model and performs annotation, enhancement, classification, and fusion processing on the BIM data with point cloud data, RGB-D images, and IoT data respectively to generate a training dataset.

[0219] Then, the large model training module 200 uses the training dataset for feature extraction and spatial modeling processing, and performs spatial cognition, dynamic inference, and environment interaction training to generate a spatial intelligent large model.

[0220] The system and method for training a spatial intelligence large model based on BIM technology provided by the present invention use the training dataset construction module 100 to fuse multi-modal information such as BIM data, point cloud data, RGB-D images, and IoT data, so as to dynamically apply BIM technology, obtain dynamic three-dimensional data, and generate an accurate training dataset. At the same time, through the large model training module 200, the training dataset is used for spatial cognition, dynamic reasoning, and environmental interaction training, enabling the spatial intelligence large model to achieve three-dimensional space cognition, reasoning, and interaction, thereby improving the operation and maintenance reliability of the three-dimensional space.

[0221] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A system for training spatial intelligent large models based on BIM technology, characterized in that: include A training data set construction module, wherein the training data set construction module includes a BIM parsing unit, a point cloud and RGB-D fusion unit, an IoT data alignment unit, an automatic annotation unit, a data enhancement unit, and a pre-training data set construction unit, wherein the training data set construction module is configured to extract BIM data from a BIM model, and annotate, enhance, classify, and fuse the BIM data with multimodal information to generate a training data set; A large model training module, the large model training module includes a multimodal feature extraction and representation unit, a spatial modeling unit, a self-supervised pre-training unit, a reinforcement learning and simulation optimization unit, and a model evaluation and deployment unit. The large model training module is configured to perform feature extraction and spatial modeling processing on the training data set, and perform spatial cognition, dynamic reasoning and environmental interaction training to generate a large spatial intelligence model.

2. The system for training spatial intelligent large models based on BIM technology according to claim 1 is characterized in that: The BIM parsing unit extracts the BIM data and performs normalization and unified coordinate system processing. Then, the point cloud and RGB-D fusion unit geometrically aligns the BIM data with the point cloud data, aligns the BIM data with the RGB-D image in coordinate system, and jointly models the visual features of the RGB-D image with the geometric features of the point cloud data.

3. The system for training spatial intelligent large models based on BIM technology according to claim 2 is characterized in that: The IoT data alignment unit uses time synchronization technology to align the data streams of different IoT sensors inside the building, and performs time alignment and periodic pattern recognition on the IoT data obtained by different IoT sensors through dynamic time warping and Fourier transform, and then maps the IoT data to the BIM model.

4. The system for training spatial intelligent large models based on BIM technology according to claim 3 is characterized in that: The automatic labeling unit combines self-supervised learning and active learning methods to compare, identify and learn the spatial patterns of the BIM model, generate semantic labels for the BIM data, and then enhance the BIM data and multimodal information through the data enhancement unit.

5. The system for training spatial intelligent large models based on BIM technology according to claim 4 is characterized in that: The pre-training data set construction unit adopts the unsupervised clustering K-Means algorithm to classify the BIM data and multimodal information respectively, and adopts the generative adversarial network and diffusion model to fuse the BIM data and multimodal information.

6. The system for training spatial intelligent large models based on BIM technology according to claim 5 is characterized in that: The multimodal feature extraction and representation unit extracts key information from BIM data and multimodal information respectively, and constructs a unified high-dimensional feature representation.

7. The system for training spatial intelligent large models based on BIM technology according to claim 6 is characterized in that: The spatial modeling unit adopts the GNN GraphSAGE algorithm to learn the graph structure information of BIM data to extract the spatial topological relationship of the building, and uses Transformer for cross-modal learning to fuse BIM data and multimodal information through the self-attention mechanism to establish the global spatial relationship of the building.

8. The system for training spatial intelligent large models based on BIM technology according to claim 7 is characterized in that: The self-supervised pre-training unit uses BIM data and multimodal information to perform spatial cognition training and prediction through contrastive learning and masked self-encoding.

9. The system for training spatial intelligent large models based on BIM technology according to claim 8 is characterized in that: The reinforcement learning and simulation optimization unit constructs a building simulation scene, discretizes the building environment, and uses BIM data and multimodal information to perform dynamic reasoning and environmental interaction training.

10. A method for training a spatial intelligent large model based on BIM technology, characterized in that: A system for training a spatial intelligent large model based on BIM technology according to any one of claims 1 to 9, the method comprising: Firstly, the BIM data is extracted from the BIM model through the training dataset construction module, and the BIM data is annotated, enhanced, classified and fused with the point cloud data, RGB-D image and IoT data to generate the training dataset. The training data set is then used through a large model training module to perform feature extraction and spatial modeling processing, and to perform spatial cognition, dynamic reasoning and environmental interaction training to generate a large spatial intelligence model.

Citation Information

Patent Citations

  • BIM rule specification extraction method based on large model

    CN119337992A

Cited By

  • Safety evaluation method, system, equipment and program for chemical storage place

    CN120579007A

  • Urban three-dimensional modeling and dynamic simulation environment generation method and device based on multi-modal large model

    CN120747380A

  • Graph structure representation and model error recognition method and device of BIM model

    CN120976673A

  • Building wind system link intelligent identification method based on graph neural network model fused with Transform

    CN121234234A

  • Remote sensing data fusion method based on locality sensitive hashing and cross-modal attention

    CN122200258A