Intelligent geometric reasoning and semantic understanding method based on three-dimensional large language model
By constructing a multimodal three-dimensional large language model and combining it with geometric perception and semantic understanding modules, the problem of insufficient functional recognition and cultural understanding of existing three-dimensional point cloud analysis methods in complex architectural scenes is solved, efficient semantic recognition and real-time response are achieved, and the intelligence level of architectural heritage protection and intelligent construction is improved.
Patent Information
- Application Number
- CN202510990464.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-18
AI Technical Summary
Existing 3D point cloud analysis methods have difficulty in achieving component function identification, construction logic reasoning, and cultural semantic understanding in complex construction scenarios. They lack the ability to fuse multimodal information, resulting in insufficient real-time semantic updates and interactive responses in dynamic construction environments.
An intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model is adopted. By constructing a multimodal three-dimensional large language model, combined with a geometric perception encoding module, a contextual semantic understanding module and a parameter efficient fine-tuning module, the fusion analysis of point cloud data and text data is realized, including denoising, standardization, and segmentation processing. Uniform downsampling and outlier screening are used, and model optimization is performed through cross-modal contrast loss and incremental fine-tuning mechanism.
It significantly improves the semantic recognition accuracy and cultural background analysis capabilities of complex building components, supports real-time updates and interactive responses in dynamic construction scenarios, reduces data redundancy and noise interference, enhances the robustness and generalization capabilities of the model, and meets the real-time requirements of construction sites.
Smart Images

Figure CN120542438B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence. More particularly, the present application relates to a method for intelligent geometric reasoning and semantic understanding based on a three-dimensional large language model. BACKGROUND
[0002] With the wide application of emerging technologies such as building information modeling (BIM), city information model (CIM) and digital twin in the field of building and city management, the expression method and processing capacity of spatial data have ushered in an unprecedented development opportunity. Among them, three-dimensional point cloud, as a highly accurate and non-contact data acquisition method, has been widely used in intelligent construction, building heritage protection, digital reconstruction and other scenes because it can record the shape, structure and spatial relationship of buildings completely. Three-dimensional point cloud not only improves the visualization and digitization of building entity space structure, but also provides key support for structure monitoring, construction assistance, virtual repair, etc. In the field of building heritage protection, the introduction of three-dimensional point cloud technology has brought fundamental changes to traditional surveying and mapping and archival recording methods, enabling the complete preservation of geometric information of historical buildings. However, only geometric information cannot meet the deep-seated needs of heritage protection for historical value, cultural connotation and semantic logic. Traditional point cloud analysis methods usually focus on geometric feature extraction and component reconstruction, relying on manual annotation and rule-driven algorithms for semantic segmentation, making it difficult to realize component function identification, structural logic reasoning and cultural context understanding, especially for large-scale and complex heritage buildings. In the field of intelligent construction, with the increasing demand for real-time perception and state recognition of construction scenes, traditional rule-based point cloud processing procedures have been difficult to meet the intelligent needs of automatic identification of construction components, state updating and semantic monitoring. Current component identification methods are mostly based on low-dimensional geometric features or shallow neural networks for classification and positioning, lacking deep modeling ability for component function semantics, resulting in low recognition accuracy and insufficient generalization ability in dynamic construction or complex environments, which seriously restricts the scalability of intelligent construction systems.
[0003] In addition, existing point cloud understanding methods are mostly limited to single-modal information processing, such as using only geometric or image data for component identification, lacking a fusion mechanism for multi-source heterogeneous information such as architectural drawings, historical texts, regulatory specifications, etc., resulting in difficulty in providing comprehensive analysis capabilities that take into account both geometric precision and semantic depth when facing building life cycle management. Especially in multi-dimensional knowledge scenarios involving architectural heritage, the model cannot effectively understand the historical background, artistic style and cultural connotation behind the building components, making the protection and reuse value based on data greatly discounted. In recent years, with the rapid development of artificial intelligence technology, especially the rise of large language models (LLM) and multi-modal learning technology, it provides a new path for semantic understanding and geometric reasoning of architectural point cloud data. Large language models have strong language understanding and generation capabilities, and can process massive language information such as historical documents, drawing instructions, and specification texts to extract semantic elements related to building components. Multi-modal learning technology can jointly model and semantically fuse multi-modal data such as point clouds, images, and texts, enabling cross-modal information collaboration and knowledge reasoning, providing a foundation for building intelligent systems with understanding, generalization, and reasoning capabilities.
[0004] Based on this, an intelligent geometric reasoning and semantic understanding method is developed for three-dimensional point cloud data, which integrates large language models and multi-modal learning capabilities. It can not only achieve efficient identification and spatial structure analysis of building components, but also deeply mine their cultural, historical and functional semantics, improve the intelligent level of architectural heritage protection, and provide general and expandable technical solutions for related fields such as intelligent construction, operation and management, digital archives, etc., with broad technical prospects and application value. SUMMARY
[0005] An object of the present application is to solve at least the above problems and to provide at least the advantages described later.
[0006] Another object of the present application is to provide an intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model, which solves the technical defects of existing three-dimensional point cloud analysis methods that rely on manual annotation and rule-driven algorithms, making it difficult to achieve functional identification, construction logic reasoning and cultural semantic understanding of components in complex architectural scenarios, and lacking multi-modal information fusion capabilities, resulting in insufficient real-time semantic updating and interactive response in dynamic construction environments. The present application is aimed at practical scenarios such as intelligent construction, digital protection of architectural heritage, intelligent analysis of the construction process, semantic identification of spatial components, and natural language interaction systems, aiming to break through the key bottlenecks of low geometric understanding precision, weak cross-modal alignment capability, and insufficient interactive intelligence in traditional three-dimensional semantic understanding, and to comprehensively improve the semantic modeling and intelligent response capabilities in complex three-dimensional scenarios.
[0007] In order to achieve the objects and other advantages according to the present application, a three-dimensional large language model-based intelligent geometric reasoning and semantic understanding method is provided, which comprises the following steps:
[0008] Step one, collecting original three-dimensional point cloud data of building components through a three-dimensional scanning device; generating text data based on the original three-dimensional point cloud data association;
[0009] Step two, preprocessing the original three-dimensional point cloud data;
[0010] Step three, constructing a multi-modal three-dimensional large language model, which comprises a geometric perception coding module, a context semantic understanding module and a parameter efficient fine-tuning module;
[0011] Among them, the geometric perception coding module calculates the geometric features by local density and global center distance, and constructs a shared semantic embedding space and a cross-modal cross-attention module;
[0012] The context semantic understanding module is pre-trained based on a bidirectional transformer encoder through a masked language model and a next sentence prediction task;
[0013] The parameter efficient fine-tuning module combines low-rank adaptation and structured prompt vectors to configure an incremental model update mechanism;
[0014] Step four, taking the preprocessed three-dimensional point cloud data obtained in step two and the text data obtained in step one as a training sample set, and training and optimizing the multi-modal three-dimensional large language model;
[0015] Step five, inputting the three-dimensional point cloud data and text data to be tested into the multi-modal three-dimensional large language model after training and optimization, and outputting the semantic recognition, attribute completion and historical background analysis results of the building components.
[0016] Preferably, the three-dimensional large language model-based intelligent geometric reasoning and semantic understanding method, step two is specifically:
[0017] The original three-dimensional point cloud data is denoised, standardized and segmented, the geometric features of each building component are extracted and a unique attribute ID is assigned;
[0018] The original three-dimensional point cloud data is down-sampled by using a uniform down-sampling algorithm, and redundant information is removed by using an isolated point detection and abnormal density point screening strategy.
[0019] Preferably, the three-dimensional large language model-based intelligent geometric reasoning and semantic understanding method, the training and optimization of the multi-modal three-dimensional large language model in step four specifically comprises:
[0020] Based on the pre-processed three-dimensional point cloud data and text data, a training sample set is formed, and through a cross-modal contrast loss function and a structure preservation loss function, a geometric perception encoding module and a context semantic understanding module are jointly pre-trained to align the geometric features and text features in a shared semantic embedding space.
[0021] During the training process, the point cloud features output by the geometric perception encoding module and the text features output by the context semantic understanding module are bidirectionally interacted through a cross-modal cross-attention module to generate a joint feature representation that fuses multi-modal semantics.
[0022] Based on the pre-trained model, a parameter efficient fine-tuning module is adaptively optimized through task instruction examples:
[0023] a. A structured prompt vector is used to dynamically encode the features of the newly added component region, and the structured prompt vector is generated based on the functional attributes and spatial region ID of the component;
[0024] b. Through an incremental model updating mechanism, the local parameters corresponding to the newly added component region are updated, and the remaining parameters are frozen.
[0025] Preferably, in the intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model, the geometric perception encoding module calculates geometric features through local density and global center distance, and constructs a shared semantic embedding space and a cross-modal cross-attention module, specifically:
[0026] 1) Based on the local density ρ i and the global center distance d i of each point in the point cloud data, a double-scale geometric modulation mechanism is constructed;
[0027] Local density , wherein, is the number of neighboring points, N ( i ) is the set of k nearest neighbors of point i , K p i , p j are the spatial coordinates of point i and its neighborhood point j , respectively;
[0028] Global center distance , wherein, c is the global center point coordinate of the point cloud;
[0029] 2) Take ρ i and d i as the geometric perception weight, and pass them through a geometric adjustment function g(di ,d j ,ρ i ,ρ j ) dynamically adjusts the self-attention score, and the specific formula is: , wherein q i is the query vector of point i; k j is the key vector of point j; is a geometric feature weight adjustment coefficient; is a geometric adjustment function responsible for calculating the spatial relationship and density relationship between two points; the index represents the vector inner product operation;
[0030] 3) Based on the attention score , further through the Softmax function to get the normalized attention weight, aggregate the value vector of the neighborhood point , and after feature transformation, output the geometric feature vector , the specific calculation formula is:
[0031]
[0032]
[0033] ,
[0034] wherein, is the attention score; v j is the value vector of the jth point, which is learned by the geometric feature and semantic feature of the point is the vector output by the self-attention mechanism, , is the network weight; , is the bias vector, is the geometric feature vector;
[0035] 4) The geometric perception feature vector is mapped to the shared semantic embedding space through linear projection, and cross-modal cross-attention calculation is performed with the text modal feature to generate a joint feature representation that fuses geometry and semantics.
[0036] Preferably, the intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model, the context semantic understanding module is based on a bidirectional transformer encoder and is pre-trained through a masked language model and a next sentence prediction task, specifically:
[0037] Pre-training based on text data, wherein:
[0038] a. In the mask language model task, 15% of the architectural professional terms in the input text are masked. The masking strategy includes: i. 80% of the masked labels are replaced with mask symbols; ii. 10% of the masked labels are replaced with random architectural terms; iii. 10% of the masked labels remain unchanged;
[0039] b. In the next sentence prediction task, the positive samples are the actual continuous sentence pairs in the architectural text, and the negative samples are the non-continuous sentence pairs randomly combined;
[0040] In the fine-tuning stage, the semantic vector output by BERT is supervised and aligned by the architectural component attribute label.
[0041] Preferably, the intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model comprises a parameter efficient fine-tuning module, a low-rank adaptive mechanism, and a structured prompt vector, and an incremental model updating mechanism.
[0042] A low-rank adaptive mechanism is introduced into the full connection layer of the multi-modal three-dimensional large language model, and specifically:
[0043] The weight matrix of the multi-modal three-dimensional large language model is decomposed into , wherein W 0 is the original frozen weight, Δ W is a trainable low-rank matrix, and Δ W= , , , and the rank ;
[0044] Only the low-rank matrix A , B is updated by gradient, and the W 0 parameter is fixed;
[0045] A structured prompt vector is generated based on the component function attribute and the spatial region ID , wherein k is the number of prompt vectors, and d is the vector dimension. The prompt vector T is embedded into the model input in the following way: the T is concatenated with the point cloud feature vector X point to form an enhanced input Xin = Concat( X point, T ); and the Xin is mapped to the model hidden space through a learnable projection matrix;
[0046] An incremental model updating mechanism is configured, and specifically: based on the spatial region ID of the newly added component, the corresponding local low-rank matrix Anew 、 B new and local prompt vector T new ; freeze original model parameters W 0 and low rank matrix of non-incremental region and prompt vector, only for A new 、 B new and T new Fine-tuning; the model output after local parameter update is normalized and fused with global model features.
[0047] Preferably, in the intelligent geometric reasoning and semantic understanding method based on the three-dimensional large language model, the original three-dimensional point cloud data collected in steps one and two and the to-be-measured data in step five are stored and managed through a spatial database, specifically including:
[0048] Construct a spatial database linked with a multi-modal three-dimensional large language model: design a component data table and a point cloud data table, the component data table is associated with the component ID field in the point cloud data table through a component identification code, forming a one-to-many data mapping relationship; the point cloud data table stores preprocessed point cloud data, including three-dimensional coordinates (X, Y, Z), color information (RGB) and associated component ID field;
[0049] Construct a spatial index on the three-dimensional coordinate field of the point cloud data table, and use R-tree or octree structure to optimize the following operations: in response to a user semantic query instruction, quickly locate the point cloud data of the target component through the spatial index, and transmit it to the multi-modal three-dimensional large language model for semantic analysis; in the incremental model updating process, based on the spatial region ID of the newly added component, the corresponding point cloud data is retrieved and loaded into the training process in a partitioned manner;
[0050] Configure a dynamic data synchronization mechanism, when the point cloud data of the newly added component is collected and preprocessed through a three-dimensional scanning device, automatically assign a spatial region ID according to the component ID field, generate a corresponding component data table record, and store it in association with the preprocessed point cloud data; then trigger the incremental updating process of the parameter efficient fine-tuning module, and update the spatial index and model parameters synchronously.
[0051] The application provides an intelligent geometric reasoning and semantic understanding device based on a three-dimensional large language model, which comprises:
[0052] A data acquisition module acquires original three-dimensional point cloud data of a building component through a three-dimensional scanning device; and acquires text data associated with the original three-dimensional point cloud data;
[0053] A data preprocessing module is configured to preprocess the original three-dimensional point cloud data;
[0054] a model construction module for constructing a multi-modal three-dimensional large language model, comprising a geometric perception encoding module, a context semantic understanding module, and a parameter-efficient fine-tuning module;
[0055] The geometric perception encoding module calculates geometric features by local density and global center distance, and constructs a shared semantic embedding space and a cross-modal cross-attention module.
[0056] The context semantic understanding module is pre-trained based on a bidirectional transformer encoder, a masked language model, and a next sentence prediction task.
[0057] The parameter-efficient fine-tuning module combines low-rank adaptation and structured prompt vectors to configure an incremental model update mechanism.
[0058] A model training optimization module for training and optimizing the multi-modal three-dimensional large language model using the preprocessed three-dimensional point cloud data obtained in step two and the text data obtained in step one as training sample sets.
[0059] An analysis module for inputting the three-dimensional point cloud data and text data to be tested into the multi-modal three-dimensional large language model after training and optimization, and outputting semantic recognition, attribute completion, and historical background analysis results of the building component.
[0060] The application also provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the program.
[0061] The application also provides a computer-readable storage medium storing a computer program, wherein the program is executed by a processor to implement the steps of the above method.
[0062] The application at least has the following advantages:
[0063] 1. The application significantly improves the semantic recognition accuracy and cultural background analysis capability of complex building components by fusing a multi-modal large language model of three-dimensional point cloud and text, supports real-time updating and interactive response in dynamic construction scenarios, and provides an efficient solution for building heritage protection and intelligent construction.
[0064] 2. The application uses uniform downsampling and outlier point screening strategies to effectively reduce data redundancy and noise interference, improve preprocessing efficiency, and enhance the robustness of the model to large-scale point cloud data, laying a high-quality data foundation for subsequent semantic segmentation and feature alignment.
[0065] 3、The application realizes efficient alignment of geometric and text features through cross-modal contrast loss and incremental fine-tuning mechanism, simultaneously supports rapid adaptation of new components, and significantly improves the generalization ability and deployment efficiency of the model in dynamic environment;
[0066] 4、The application enhances the perception ability of the model to the spatial structure of complex building components through the double-scale geometric modulation mechanism combined with local density and global distance features, improves the semantic difference capture precision of nested structures (such as dougong), and supports accurate component function reasoning;
[0067] 5、The application strengthens the understanding ability of the BERT model to professional terms and historical background through the shielding strategy and supervised alignment method in the field of architecture, improves the relevance of semantic vectors and component attributes, and provides high-discrimination text features for cross-modal fusion;
[0068] 6、The application greatly reduces the fine-tuning parameter quantity through the low-rank adaptive and structured prompt vector technology, combined with the incremental update mechanism, realizes the rapid adaptation of new components (average time <4 minutes), significantly reduces the consumption of computing resources, and meets the real-time requirements of the construction site;
[0069] 7、The application optimizes the storage and retrieval efficiency of point cloud data through the spatial index and dynamic synchronization mechanism, supports real-time semantic query and incremental learning process, ensures the stability and response speed of the system in high-concurrency scenarios (delay <500ms);
[0070] 8、The application realizes end-to-end multi-modal data processing and reasoning through modular design, reduces the system integration complexity, and improves the deployment flexibility and scalability in intelligent construction and heritage protection scenarios.
[0071] Other advantages, objects and features of the application will be partly embodied by the following description, and partly understood by those skilled in the art through research and practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0072] Figure 1 The flowchart of the intelligent geometric reasoning and semantic understanding method based on the three-dimensional large language model according to the application. DETAILED DESCRIPTION
[0073] The application will be further described in detail below with reference to the accompanying drawings and embodiments, so that those skilled in the art can implement it according to the description.
[0074] It should be understood that the terms such as "have", "contain" and "include" used herein do not exclude the presence or addition of one or more other elements or combinations thereof.
[0075] It should be noted that the experimental methods in the following embodiments are conventional methods unless otherwise specified, and the reagents and materials can be obtained from commercial sources unless otherwise specified.
[0076] In the description of the present application, the terms "transverse", "longitudinal", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.
[0077] The present application provides a kind of intelligent geometric reasoning and semantic understanding method based on three-dimensional large language model, it includes the following steps:
[0078] Step one, the original three-dimensional point cloud data of building component is collected by three-dimensional scanning equipment (Faro scanner, Leica scanner, total station or binocular depth camera);Text data is generated based on original three-dimensional point cloud data association (text data includes component name, historical background, functional attribute and dynasty information);
[0079] Step two, the original three-dimensional point cloud data is preprocessed;
[0080] Step three, a multi-modal three-dimensional large language model is constructed, which includes a geometric perception encoding module, a context semantic understanding module and a parameter efficient fine-tuning module;
[0081] The geometric perception encoding module calculates geometric features by local density and global center distance, and constructs a shared semantic embedding space and a cross-modal cross-attention module;
[0082] The context semantic understanding module (BERT) is based on a bidirectional transformer (Transformer) encoder and is pre-trained by a mask language model (MLM) and a next sentence prediction (NSP) task;
[0083] The parameter efficient fine-tuning module (PEFT) combines low-rank adaptive (LoRA) and structured prompt vectors to configure an incremental model update mechanism;
[0084] Step four, the preprocessed three-dimensional point cloud data obtained in step two and the text data obtained in step one are used as a training sample set to train and optimize the multi-modal three-dimensional large language model;
[0085] Step five, input the three-dimensional point cloud data and text data to be tested into the multi-modal three-dimensional large language model trained and optimized, and output the semantic recognition, attribute completion and historical background analysis results of the building component.
[0086] The three-dimensional scanning device of the present application can be selected as a Faro Focus S 350 laser scanner, which has a scanning accuracy of 0.1 mm and a point cloud density that can be set to 1000 to 5000 points per square meter. The device can be installed on a tripod, and the spatial consistency of the scanning data is ensured by horizontal calibration. The original point cloud data collected is transmitted to an HP Z8 G4 workstation through a USB3.0 interface, and is associated with text data, including component name (such as "corbel" "eave column"), historical era (such as "Ming Dynasty" "Qing Dynasty") and functional attribute (such as "load-bearing" "decoration"), and the text format is stored in JSON structure. The association logic is one-to-one mapping based on the point cloud file ID and the component ID in the text database, and the mapping error tolerance is set to ±2%. Similar processes can be referred to the POS (Positioning and Orientation System) and image data association process in unmanned aerial survey. Technical effect: through high-precision scanning and structured text mapping, geometric and semantic alignment input data are provided for the model, supporting subsequent cross-modal feature fusion.
[0087] The preprocessing of the point cloud data of the present application includes denoising, standardization and segmentation. The denoising adopts a statistical outlier removal algorithm, and the noise threshold σ is set to 0.05, and the isolated point density threshold is set to less than 5 points per cubic meter. The standardization scales the point cloud coordinates to the range of [-1, 1] through Z-score normalization. The segmentation adopts a region growing algorithm, and the seed point selection rule is that the curvature is less than 0.01 and the normal deviation angle is less than 15°, and the segmented components are given a unique ID (such as "COMP_001"). The downsampling adopts a voxel grid method, and the voxel size is set to 5 cm³, and the redundant point rejection rate can reach 60%. The preprocessing is run on a Dell Precision 7865 workstation, and the time consumption is about 10 minutes / GB of point cloud data. Similar processes can be found in CT data segmentation and downsampling in medical imaging. Technical effect: through noise filtering and uniform sampling, the data quality is improved, the computational load during model training is reduced, and the key geometric features are preserved.
[0088] The multimodal three-dimensional large language model of the application is composed of a geometric perception coding module, a context semantic understanding module and a parameter efficient fine-tuning module. The local density calculation of the geometric coding module uses K=32 neighboring points, and the global center distance is based on the point cloud centroid coordinates. During training, the batch size is one stage 16 and two stages 64, the learning rate is 1e-3 and 4e-5 respectively, and the training rounds are 50 rounds. The hardware uses NVIDIA A800 GPU, and the memory occupation is about 40 GB. Cross-modal alignment uses a contrast loss function, and the temperature coefficient τ is 0.07. When incrementally fine-tuning, the time consumption of updating the parameters of the newly added component region is not more than 4 minutes, and only 5% of the model parameters need to be adjusted. Similar methods can be used to compare the multimodal pre-training process of the CLIP model.
[0089] The application improves the semantic recognition accuracy of complex building components through high-precision three-dimensional scanning, text structured association, standardized preprocessing flow and multimodal large language model training, and reduces the error rate of historical background analysis. In the dynamic construction scene, the model supports 5 concurrent queries per second, and the response delay is stable within 500 ms, meeting the real-time interaction needs of building heritage protection and intelligent construction scenes.
[0090] In another technical solution, the intelligent geometric reasoning and semantic understanding method based on the three-dimensional large language model, step two is specifically:
[0091] The original three-dimensional point cloud data is denoised, standardized and segmented, the geometric features of each building component are extracted and a unique attribute ID is assigned;
[0092] The original three-dimensional point cloud data is downsampled using a uniform downsampling algorithm, and redundant information is removed through isolated point detection and abnormal density point screening strategies.
[0093] The denoising of the application uses a statistical outlier removal algorithm, and the noise threshold σ can be set to 0.05, and the isolated point density threshold is set to less than 5 points per cubic meter. Standardization is achieved by Z-score normalization to scale the point cloud coordinates to the range of [-1, 1], and the normalization parameters are calculated based on the global mean and standard deviation of the point cloud. Segmentation uses a region growing algorithm, and the seed point selection rule is a region with a curvature less than 0.01 and a normal deviation angle less than 15°, and each component is assigned a unique attribute ID (such as "COMP_001") after segmentation. The algorithm running platform can use Dell Precision 7865 workstations equipped with NVIDIA RTX A6000 graphics cards with a memory capacity of 128GB. Similar processes can refer to the point cloud preprocessing process in industrial part three-dimensional scanning. Technical effect: Through noise filtering and standardization, the interference of data outliers on model training is reduced, and the uniqueness of the component ID after segmentation provides a structured basis for subsequent semantic alignment.
[0094] The downsampling of the application adopts the voxel grid method, the voxel size can be set to 5cm³, and the redundant point elimination rate can reach 60%. The abnormal point screening is realized by density clustering (DBSCAN algorithm), the neighborhood search radius epsilon is set to 0.1 m, and the minimum neighborhood point number MinPts is set to 10. The isolated point detection is based on the KD tree space index, and the isolated points with a distance of more than 0.3 m from the nearest neighbor point are screened out. The algorithm is integrated in the CloudCompare open source software and runs in the Ubuntu 20.04 system environment, and the data storage adopts Samsung 980 PRO NVMe solid state disk. Similar methods can be found in LiDAR point cloud compression and noise filtering in the automatic driving scene. Technical effect: through uniform sampling and density clustering screening, the data size is reduced while the key geometric features are retained, the preprocessing efficiency and model robustness are improved.
[0095] The application significantly reduces the redundancy and noise interference of point cloud data through denoising, standardization, segmentation processing, uniform downsampling and abnormal point screening. The data amount after preprocessing is reduced, the model training time is shortened, and the key component geometric feature retention rate is improved, providing high-quality input for subsequent semantic segmentation and cross-modal alignment.
[0096] In another technical solution, the intelligent geometric reasoning and semantic understanding method based on the three-dimensional large language model, the training and optimization of the multi-modal three-dimensional large language model in step four specifically includes:
[0097] Based on the training sample set formed by the preprocessed three-dimensional point cloud data and text data, through the cross-modal contrast loss function and the structure preservation loss function, the geometric perception coding module and the context semantic understanding module are jointly pre-trained, and the geometric features and text features are aligned in the shared semantic embedding space;
[0098] In the training process, through the cross-modal cross-attention module, the point cloud features output by the geometric perception coding module and the text features output by the context semantic understanding module are bidirectionally interacted to generate a joint feature representation that integrates multi-modal semantics;
[0099] Based on the pre-trained model, the parameter efficient fine-tuning module is adaptively optimized through task instruction examples:
[0100] a. Adopting a structured prompt vector to dynamically encode the features of the newly added component area, the structured prompt vector is generated based on the functional attributes and spatial region ID of the component;
[0101] b. Through an incremental model updating mechanism, the local parameters corresponding to the newly added component area are updated, and the remaining parameters are frozen to maintain global semantic consistency.
[0102] In the cross-modal contrast loss function of the application, the temperature coefficient tau can be set to 0.07, and the negative sample sampling ratio is 1:3. The structure preservation loss weight lambda is set to 0.5 to constrain the integrity of the local geometric structure of the point cloud. The pre-training batch size is one stage 16, the learning rate is 1e-3, and the training epochs are 50. The hardware platform can select NVIDIA A800 GPU, the memory capacity is 80 GB, and the training data is stored in Seagate Exos X20 enterprise hard disk. In the pre-training process, the point cloud and the text feature are aligned through the shared semantic embedding space, and the embedding dimension d is set to 768. The contrast loss calculation adopts the CosineSimilarity module of the PyTorch framework, and the similarity threshold is set to 0.8. Similar methods can refer to the image-text alignment pre-training process in multi-modal retrieval. Technical effect: Through the joint optimization of geometric and text features, the understanding ability of the model to the cross-modal semantic association of building components is enhanced, laying a foundation for subsequent incremental learning.
[0103] The cross-attention module of the application adopts a bidirectional interaction mechanism, and the number of attention heads h is set to 32, and the key-value dimension is 64. The interaction layer number of the point cloud feature and the text feature is 12, and each layer includes self-attention and feedforward network. In the attention calculation, the point cloud feature is extracted through the geometric perception coding module, and the text feature is generated by the BERT module. The fusion feature dimension after interaction is kept as 768, and the residual connection weight is initialized as 0.02. The gradient clipping threshold is set to 1.0 during training to prevent gradient explosion. The module runs in the Docker container environment under the Linux system, and the container image is compiled based on CUDA 11.8. Similar processes can be found in the multi-modal feature fusion design in visual-language models. Technical effect: Through the bidirectional attention mechanism, the deep fusion of geometric and semantic information is realized, and the reasoning ability of the model to complex component functions and historical backgrounds is improved.
[0104] The parameter efficient fine-tuning module of the application adopts the low-rank adaptive (LoRA) technology, the rank r of the low-rank matrix is set to 8, the number of prompt vectors k is 10, and the dimension d is 128. During incremental updating, only the low-rank matrices A_new and B_new corresponding to the newly added component region are trained, the learning rate is set to 4e-5, and the fine-tuning epochs are 10. The structured prompt vector is dynamically generated based on the component function attribute (such as "gongche-bearing") and the spatial region ID (such as "ZONE_05"), and is fused with the point cloud feature through splicing. The fine-tuning process is completed on a single NVIDIA A800 GPU, and the time consumption of a single update is not more than 4 minutes, and the memory occupancy is less than 8 GB. Similar methods can be analogized to the local parameter update strategy in federated learning. Technical effect: Through local parameter freezing and incremental fine-tuning, the rapid adaptation of newly added components is realized, while the stability of the global model is maintained, and the deployment cycle in the dynamic construction scene is significantly shortened.
[0105] The application effectively improves the generalization ability of the model in a dynamic environment through cross-modal contrast pre-training, bidirectional attention interaction and incremental fine-tuning. The fine-tuning time of the new component is shortened to minutes, the semantic recognition consistency of the model in multiple scenes is significantly enhanced, and the resource consumption is reduced, meeting the real-time response needs of intelligent construction and heritage protection scenes.
[0106] In another technical solution, the intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model, the geometric perception coding module calculates the geometric features through local density and global center distance, and constructs a shared semantic embedding space and a cross-modal cross-attention module, specifically:
[0107] 1) Based on the local density ρ of each point in the point cloud data i and the global center distance d i , a double-scale geometric modulation mechanism is constructed;
[0108] Local density , in the formula, is the number of neighboring points, N ( i ) is the set of i nearest neighbor points of point K , p i , p j and i are the spatial coordinates of point j and its neighborhood point ;
[0109] Global center distance , in the formula, c is the global center point coordinate of the point cloud;
[0110] 2) Take ρ i and d i as the geometric perception weight, and dynamically adjust the self-attention score through the geometric adjustment function g(d i , d j , ρ i , ρ j ), the specific formula is: , in the formula, q i is the query vector of point i; is the key vector of point j; is the geometric feature weight adjustment coefficient; g is the geometric adjustment function, responsible for calculating the spatial relationship and density relationship between two points; the exponential represents the vector inner product operation;
[0111] 3) Based on the attention score , and further obtain the normalized attention weight through the Softmax function, and aggregate the value vector v of the neighborhood points j , and after feature transformation, output geometric feature vector , the specific calculation formula is:
[0112]
[0113]
[0114] ,
[0115] Where, is the attention score; v j is the value vector of the j-th point, which is learned by learning the geometric features and semantic features of the point; is the vector output by the self-attention mechanism, 、 is the network weight; 、 is the bias vector, is the geometric eigenvector;
[0116] 4) Geometry-aware feature vector It maps the image to a shared semantic embedding space through linear projection, performs cross-modal attention calculation with text modality features, and generates a joint feature representation that integrates geometry and semantics.
[0117] The local density ρ of the present invention i The calculation uses K=32 nearest neighbor points, and the neighborhood is quickly searched through the KD tree algorithm. The neighborhood radius threshold is set to 0.2 m. The global center distance d i The centroid coordinates of the point cloud are calculated using the mean of all point coordinates. In the density calculation formula, the denominator is the sum of the Euclidean distances between neighboring points, and the numerator, K, is fixed at 32. The coordinates of the global center point, c, are updated once per batch of training data. The calculation process can be run on an Intel Xeon Platinum 8360Y processor with 256GB of memory, using Samsung DDR4 ECC memory for data caching. Similar methods can be referenced in the rock porosity density calculation process used in geological exploration. Technical Effect: By leveraging the dual-scale features of local density and global distance, the model's ability to perceive the spatial distribution of building components is enhanced, providing fundamental support for the semantic differentiation of nested structures (such as brackets).
[0118] The geometric adjustment function of the present invention The expression is g = (d i + d j ) / (ρ i +ρ j +1e -5), and the weight coefficient λ is set to 0.3. In the self-attention score calculation, the query vector q i The dimension of the key vector k j is 64, and the inner product operation is realized by matrix multiplication. The attention weight normalization adopts the Softmax function, and the temperature coefficient T is set to 0.1. The output feature dimension of the geometric perception encoding module is 768, and the feature fusion is realized by a residual connection, and the residual weight is initialized to 0.01. When training the module, an NVIDIA A800 GPU can be selected, and the memory occupation is about 12GB. Similar processes can be found in the geometric feature weighted matching method in point cloud registration. Technical effect: by dynamically adjusting the attention weight, the model can more accurately capture the local details and global spatial relationship of the component, and improve the semantic difference recognition ability of complex structures such as corbel brackets.
[0119] The feature vector z i after self-attention aggregation in the application is passed through two layers of feedforward neural network (FFN), and the hidden layer dimension is 3072, and the activation function adopts ReLU. The initialization method of the feature transformation weight matrix W1, W2 is Xavier normal distribution, and the bias term b1, b2 is initialized to zero. The geometric feature vector is mapped to the shared semantic space through a linear projection layer, the projection matrix dimension is 768x768, and the parameters are updated by the Adam optimizer. The projected features and the text modal features interact in the cross-modal cross-attention module, the number of attention heads h is set to 16, and the key value dimension is 64. Similar methods can be analogized in the multi-scale feature fusion network in three-dimensional target detection. Technical effect: by joint mapping of geometric features and semantic space, accurate reasoning of the functional attributes of building components is realized, and support for structural stability analysis and historical value evaluation is provided.
[0120] The application significantly enhances the understanding ability of the model for the spatial structure of complex building components through local density and global distance calculation, geometric adjustment function optimization, and feature aggregation and projection. The double-scale geometric modulation mechanism can effectively distinguish the semantic differences of nested components (such as corbel brackets and eaves), providing reliable technical support for component function reasoning and cultural heritage protection.
[0121] In another technical solution, the intelligent geometric reasoning and semantic understanding method based on the three-dimensional large language model, the context semantic understanding module (BERT) is based on a bidirectional transformer (Transformer) encoder, and is pre-trained through a mask language model (MLM) and next sentence prediction (NSP) task, specifically:
[0122] Based on text data (including historical documents, drawing instructions and specification texts), pre-training is performed, wherein:
[0123] a. In the Masked Language Model (MLM) task, 15% of the architectural terminology tokens in the input text are masked. The masking strategies include: i. 80% of the masked tokens are replaced with mask symbols; ii. 10% of the masked tokens are replaced with random architectural terms; iii. 10% of the masked tokens remain unchanged;
[0124] b. In the next sentence prediction (NSP) task, the positive samples are actual continuous sentence pairs in architectural texts, and the negative samples are randomly combined non-continuous sentence pairs;
[0125] During the fine-tuning stage, the semantic vectors output by BERT are supervised and aligned using architectural component attribute labels (including component type, dynasty, and function) to enhance the ability to extract contextual features in the architectural field.
[0126] In the Masked Language Model (MLM) task of this invention, the input text is masked at a rate of 15%. The masking mark replacement strategy is as follows: 80% is replaced with the [MASK] symbol, 10% is replaced with random architectural terms (such as "mortise and tenon" and "eaves"), and the remaining 10% remains unchanged. The random term library is derived from the digitized text of the "Dictionary of Ancient Chinese Architecture" and the "Construction Methods," and contains over 5,000 terms. Professional terms (such as "dougong" and "liangjia") are prioritized for masking, and noun phrases are identified using part-of-speech tagging tools (such as HanLP). Training data is stored on Seagate IronWolf NAS hard drives, and the processing platform can be selected from Huawei's Atlas 800 training server equipped with the Ascend 910BAI chip. Similar methods can be referenced in the professional term masking strategy used in medical text pre-training. Technical Effect: By optimizing masking for architectural terms, the model's contextual reasoning ability for professional terms is enhanced, improving the domain discrimination of semantic vectors.
[0127] In the next sentence prediction (NSP) task of this invention, positive samples are pairs of consecutive sentences in architectural text (e.g., "Dougong is a typical component of Ming and Qing architecture" and "Its load-bearing function is achieved through layers of overhanging"). Negative samples are generated by randomly combining non-consecutive sentences (e.g., "The column base is carved with lotus patterns" and "The Song Dynasty's 'Yingzaofashi' records tile-making techniques"). The positive-to-negative sample ratio is set to 1:1, and the maximum length of a single sentence is limited to 128 characters. The text data comes from the Palace Museum archive digitization project and the electronic version of "A History of Chinese Architecture." Regular expressions are used to clean the data to filter non-Chinese symbols. During training, the batch size is 64, the learning rate is set to 2e-5, and the gradient clipping threshold is 1.0. A similar process can be seen in the task of determining the coherence of contract clauses in legal texts. Technical Effect: By learning the semantic association between the previous and next sentences, the model's ability to model long-range dependencies on the architectural historical context and construction logic is enhanced.
[0128] The fine-tuning stage of the application adopts component attribute labels (such as type, dynasty, and function) for supervised alignment. The label data is generated through an expert annotation platform (such as Label Studio), and the annotation consistency check adopts a Kappa coefficient threshold of 0.75. The alignment loss function adopts cosine similarity measurement, the similarity target value is set to 0.9, the optimizer is selected as AdamW, and the weight decay coefficient is 0.01. The fine-tuning hardware can be configured as an NVIDIA A100 GPU with a memory capacity of 40 GB, and the training data loading is accelerated through the ApacheArrow format. The semantic vector dimension of the model output is 768, and the cross-modal alignment of the point cloud features adopts a bilinear interaction layer with the interaction weight initialized as a normal distribution (μ=0, σ=0.02). Similar methods can be used for fine-tuning the process of aligning text features in cross-modal retrieval. Technical effects: Through the strong supervision signal of attribute labels, the relevance of semantic vectors and component functions, dynasties, and other attributes is improved, providing high-discrimination text feature representation for cross-modal fusion.
[0129] The application significantly improves the understanding ability of the BERT model for professional terms and historical context of ancient architecture through the shielding strategy in the field of architecture, the next sentence prediction task construction, and the supervised alignment fine-tuning. The relevance of semantic vectors and component attributes is enhanced, the cross-modal feature alignment error is reduced, and the deep fusion of three-dimensional point clouds and text data is supported, providing a reliable semantic analysis foundation for the digital protection of architectural heritage.
[0130] In another technical solution, the parameter efficient fine-tuning module (PEFT) combines low-rank adaptation (LoRA) and structured prompt vectors, and configures an incremental model updating mechanism, specifically:
[0131] The low-rank adaptation (LoRA) mechanism is introduced into the full connection layer of the multi-modal three-dimensional large language model, specifically:
[0132] The weight matrix of the multi-modal three-dimensional large language model is decomposed into , wherein W 0 is the frozen original weight, Δ W is the trainable low-rank matrix, Δ W= , , , and the rank ;
[0133] Only the low-rank matrix A , B is updated by gradient, and the W 0 parameter is fixed;
[0134] Generating structured prompt vectors based on component function attributes and spatial region IDs wherein k is the number of prompt vectors, d is the vector dimension, the prompt vector T is embedded into the model input by concatenating T with the point cloud feature vector X point to form an enhanced input Xin = Concat( X point, T ); the Xin is mapped to the model hidden space through a learnable projection matrix;
[0135] An incremental model updating mechanism is configured, specifically: based on the spatial region ID of the newly added component, the corresponding local low-rank matrix A new、 B new and the local prompt vector T new are located; the original model parameters W 0 and the low-rank matrix and prompt vector of the non-newly added region are frozen, and only A new、 B new and T new are fine-tuned; the model output after local parameter updating is normalized and fused with the global model feature.
[0136] The full connection layer weight matrix W of the multi-modal three-dimensional large language model can be decomposed into W=W0+ΔW, wherein ΔW=BA, the rank r of the low-rank matrix A and B is set to 8, and the matrix dimensions are 768*8 and 8*768 respectively. The original weight W0 is frozen during training, and only A and B are updated. The low-rank matrix is initialized by using a normal distribution with a mean of 0 and a standard deviation of 0.02, and the learning rate is set to 4e-5. The hardware can select NVIDIA A800 GPU, and the memory occupation is reduced to 15% of the original model, and the single fine-tuning memory requirement is not more than 8 GB. The matrix decomposition calculation is realized by linear layer reparameterization of the PyTorch framework, and the code is integrated in the PEFT tool library of Hugging Face. Similar methods can refer to the adapter fine-tuning technology in natural language processing. Technical effect: through low-rank decomposition, the number of trainable parameters is reduced, the efficient adaptation of newly added components is realized, and the global semantic consistency of the model is maintained.
[0137] The generation of the structured prompt vector T of the application is based on the component function attribute (such as "column-bearing") and the space region ID (such as "ZONE_03"), the vector number k is set to 10, and the dimension d is 128. The prompt vector is loaded from a predefined dictionary by a lookup table method, and the dictionary contains 50 combinations of function attributes and 100 space region IDs. After the prompt vector is spliced with the point cloud features, it is mapped to the model hidden space through a learnable projection matrix (dimension 256x768). The projection matrix is initialized using the Xavier uniform distribution, and the bias term is initialized to zero. The data storage can use Samsung 980 PRO NVMe solid state disk, with a reading speed of 7000MB / s. Similar processes can be found in user feature embedding design in recommendation systems. Technical effect: through structured prompts, the semantic coding ability of new components is enhanced, and the adaptability of the model to dynamic changes in the construction scene is improved.
[0138] In the incremental update, the application is based on the space region ID of the new component to locate the corresponding low-rank matrix A new , B new and prompt vector T new . The parameters of the non-new region (W0, Aold, Bold) are all frozen, and only Anew, Bnew and T new are fine-tuned, with 10 rounds of fine-tuning and a batch size of 64. The updated local parameters are fused with the global model through weighted averaging, and the weight coefficient a is set to 0.7. The fine-tuning process is completed on a single NVIDIA A800 GPU, with an average time consumption of no more than 4 minutes and a CPU memory occupation of less than 4 GB. The update process is automatically implemented through a Python script, which is deployed in a Docker container (the image is based on Ubuntu22.04). Similar methods can be compared with the local model update strategy in edge computing. Technical effect: through local parameter freezing and incremental fusion, the computational resource consumption is significantly reduced, supporting real-time semantic update requirements in construction sites.
[0139] Through low-rank adaptive decomposition, structured prompt embedding and incremental parameter update, the application realizes the rapid adaptation of new components. The fine-tuned parameter quantity is reduced to less than 5% of the original model, the single update time is controlled within minutes, the GPU video memory occupation is significantly reduced, and the dual requirements of real-time performance and resource efficiency in intelligent construction scenarios are met.
[0140] In another technical solution, the three-dimensional large language model-based intelligent geometric reasoning and semantic understanding method, the original three-dimensional point cloud data collected in steps one and two, and the to-be-measured data in step five are stored and managed through a spatial database, specifically including:
[0141] Constructing a spatial database linked with a multi-modal three-dimensional large language model: design a component data table and a point cloud data table, the component data table is associated with the component ID field in the point cloud data table through the component identification code (COMP_ID), forming a one-to-many data mapping relationship; the point cloud data table stores pre-processed point cloud data, including three-dimensional coordinates (X, Y, Z), color information (RGB) and associated component ID field;
[0142] Construct a spatial index on the three-dimensional coordinate field of the point cloud data table, and use R-tree or octree structure to optimize the following operations: in response to user semantic query instructions, quickly locate the point cloud data of the target component through the spatial index, and transmit it to the multi-modal three-dimensional large language model for semantic analysis; in the incremental model updating process, based on the spatial region ID of the newly added component, the corresponding point cloud data is retrieved and loaded into the training process in a partitioned manner;
[0143] Configure a dynamic data synchronization mechanism, when the point cloud data of the newly added component is collected and pre-processed through a three-dimensional scanning device, automatically assign a spatial region ID according to the component ID field, generate a corresponding component data table record, and store it in association with the pre-processed point cloud data; then trigger the incremental updating process of the parameter efficient fine-tuning module, and update the spatial index and model parameters synchronously.
[0144] The component data table and the point cloud data table of the application are associated through the component identification code (COMP_ID) and adopt a one-to-many mapping relationship. The component data table field includes component ID field, name, dynasty, function attribute, and data type is VARCHAR(64). The point cloud data table stores pre-processed point cloud data, and the field includes three-dimensional coordinates (X, Y, Z, data type FLOAT), color information (RGB, data type INT), and component ID field foreign key. The database management system can use MySQL8.0, the storage engine is InnoDB, and the transaction isolation level is set to READCOMMITTED. The data table partitioning strategy is hash sharding according to the component ID field, and the maximum capacity limit of a single table is 1 TB. The data storage hardware can be configured with Dell PowerEdge R750xa server, equipped with Seagate Exos X20 18 TB hard disk, RAID 5 redundant array. Similar methods can refer to the design of the cargo and warehouse information association database in logistics management. Technical effect: through the design of structured data table and foreign key constraint, the efficient association query of point cloud data and component attributes is ensured, and the semantic analysis requirements of multi-modal model are supported.
[0145] The three-dimensional coordinate field of the point cloud data table of the application adopts R-tree indexing, the branch factor is set to 50, and the node capacity is 100 records. The depth of the octree index is set to 5 layers, and the minimum number of points in the leaf node is 100. The index construction is realized through the spatial extension module of MySQL, and the error tolerance of the spatial query range is 0.01 m. The index update strategy is asynchronous batch update, and the index reconstruction is triggered once every 1000 newly added point cloud data. The hardware can select Intel Optane Persistent Memory PMem 200 series to accelerate the index loading speed. The index optimization algorithm is based on the greedy strategy, and the nodes with a spatial overlap rate of more than 70% are preferentially merged. Similar processes can be found in the terrain elevation data index optimization in geographic information systems (GIS). Technical effect: Through the mixed index of R-tree and octree, the spatial retrieval efficiency of massive point cloud data is improved, the response time is shortened to milliseconds, and real-time interactive query is supported.
[0146] After the new component point cloud data is collected, the spatial area ID (such as ZONE_09) is automatically allocated through the component ID field, and the corresponding component data table record is generated. The data synchronization adopts a double-writing mechanism and is written into the MySQL master library and the Redis cache at the same time. The cache expiration time is set to 24 hours. The incremental update process is realized through a trigger, and when a new record is inserted into the point cloud data table, the incremental training script of the parameter efficient fine-tuning module (PEFT) is automatically called. The script is deployed in a Kubernetes cluster, and a single node is configured with 4-core CPU and 16 GB memory. The synchronization process logs are stored in Elasticsearch 8.0, and the log retention policy is 30-day rolling deletion. Similar methods can be used in the edge-cloud synchronization architecture in the real-time processing of Internet of Things device data streams. Technical effect: Through the automatic synchronization and triggering mechanism, the end-to-end delay of the newly added data from collection to model update is ensured to be less than 10 minutes, meeting the real-time requirements of dynamic construction scenarios.
[0147] The application significantly improves the storage and retrieval efficiency of massive point cloud data through spatial database construction, mixed spatial index optimization and dynamic data synchronization. The R-tree index improves the spatial query speed, the octree optimization reduces the storage redundancy, and the dynamic synchronization mechanism ensures the continuity of the incremental learning process, providing efficient and reliable data management support for intelligent construction and heritage protection scenarios.
[0148] The application also provides an intelligent geometric reasoning and semantic understanding device based on a three-dimensional large language model, which comprises:
[0149] A data acquisition module acquires original three-dimensional point cloud data of a building component through a three-dimensional scanning device (Faro scanner, Leica scanner, total station or binocular depth camera); and acquires text data (the text data includes component name, historical background, functional attribute and dynasty information) associated with the original three-dimensional point cloud data;
[0150] a data preprocessing module for preprocessing original three-dimensional point cloud data;
[0151] a model construction module for constructing a multi-modal three-dimensional large language model, which comprises a geometry perception encoding module, a context semantic understanding module and a parameter efficient fine-tuning module;
[0152] The geometry perception encoding module calculates geometric features by local density and global center distance, and constructs a shared semantic embedding space and a cross-modal cross-attention module.
[0153] The context semantic understanding module (BERT) is based on a bidirectional transformer (Transformer) encoder and is pre-trained through a mask language model (MLM) and a next sentence prediction (NSP) task.
[0154] The parameter efficient fine-tuning module (PEFT) combines low-rank adaptive (LoRA) and structured prompt vectors to configure an incremental model update mechanism.
[0155] a model training optimization module for training and optimizing the multi-modal three-dimensional large language model by using the preprocessed three-dimensional point cloud data obtained in step two and the text data obtained in step one as training sample sets;
[0156] a parsing module for inputting the three-dimensional point cloud data and text data to be tested into the multi-modal three-dimensional large language model after training and optimization, and outputting semantic recognition, attribute completion and historical background analysis results of building components.
[0157] The data acquisition module of the application can select Faro Focus S 350 laser scanner, the scanning accuracy is 0.1mm, and the maximum scanning distance is 350m. The data preprocessing module runs on Dell Precision7865 workstation, equipped with NVIDIA RTX A6000 graphics card, and the memory capacity is 128GB. The preprocessing algorithm includes statistical outlier removal (noise threshold σ=0.05), voxel grid downsampling (voxel size 5cm³) and region growing segmentation (curvature threshold 0.01). The scanning data is transmitted through USB 3.0 interface, and the preprocessing result is stored in Samsung 980 PRO NVMe solid state disk, and the read-write speed is 7000MB / s. Similar process can refer to three-dimensional scanning and point cloud processing pipeline in industrial detection. Technical effect: through modular hardware and algorithm integration, end-to-end automation from data acquisition to preprocessing is realized, manual intervention requirement is reduced, and system deployment efficiency is improved.
[0158] The model building module of the present invention is based on the PyTorch framework. The local density calculation of the geometric perception coding module adopts K=32 nearest neighbors, and the global center distance is based on the centroid coordinates. The contextual semantic understanding module adopts the BERT-base architecture, the pre-training batch size is 16, and the learning rate is 1e-3. The parameter efficient fine-tuning module is configured with a low-rank matrix rank r=8 and a structured prompt vector dimension d=128. The training hardware can use an NVIDIA A800 GPU cluster with a single-node video memory of 80GB, and distributed training is implemented through the NCCL communication library. The model optimization uses the AdamW optimizer, the weight decay coefficient is set to 0.01, and the gradient clipping threshold is 1.0. Similar methods can be seen in the model training pipeline design of the multimodal AI platform. Technical effect: Through standardized model architecture and distributed training support, rapid iteration and optimization of complex models are achieved to adapt to the deployment requirements of scenarios of different scales.
[0159] The parsing module of this invention is deployed in a B / S architecture system. The front-end uses the Three.js library for 3D visualization, and the back-end is built with a RESTful API service based on Node.js, running on the Ubuntu 22.04 operating system. The user interface supports natural language queries (e.g., "Mark all brackets"), and the semantic parsing response latency threshold is set to 500 ms. The system hardware can be configured with the Huawei Atlas 800 inference server, equipped with the Ascend 910BAI chip, supporting ≥5 concurrent users. Data caching uses Redis 6.2 with a cache expiration time of 10 minutes, and database queries are accelerated using MySQL 8.0 spatial indexing. Similar systems can be compared to 3D visualization management platforms in smart cities. Technical Effect: Through a lightweight front-end and back-end separation design, high-concurrency real-time interaction is achieved, meeting the diverse application requirements of construction sites and heritage protection scenarios.
[0160] This invention achieves end-to-end multimodal data processing and reasoning through modular hardware and algorithm integration, a standardized model training pipeline, and a lightweight interactive system. The system automates the entire process, from data acquisition to semantic parsing, reducing integration complexity. Furthermore, through a distributed architecture and efficient caching mechanisms, it enhances deployment flexibility and scalability in intelligent construction and heritage conservation scenarios.
[0161] The application provides an electronic device, which comprises a processor and a memory, the memory stores a computer program, and the processor implements the steps of the above method when executing the program. The electronic device is a device comprising a processor (CPU / MCU / SOC), a memory (ROM / RAM), such as a desktop computer, a laptop computer, a smart phone, etc. In particular, the memory stores a computer program, and the processor implements all or part of the steps of the above method of intelligent geometric reasoning and semantic understanding based on a three-dimensional large language model when loading and executing the computer program.
[0162] The application provides a computer-readable storage medium, which stores a computer program, and the program implements the steps of the above method when executed by a processor. The storage medium includes various storage media that can store program codes, such as ROM, RAM, magnetic disk, optical disk, etc., and stores a computer program, which implements all or part of the steps of the above method of intelligent geometric reasoning and semantic understanding based on a three-dimensional large language model when loaded and executed by a processor.
[0163] As shown in Figure 1 , the application constructs a multi-modal three-dimensional large language model method framework integrating deep learning and natural language processing technology around the semantic understanding demand of three-dimensional point cloud data in complex architectural space. The overall system is divided into three core modules:
[0164] 1. Multi-modal three-dimensional large language model architecture design. The present invention first proposes a 3D-MELL (3D-Multimodal Efficient Language Learning) architecture suitable for building component point cloud semantic understanding, aiming to break the modal limitations between three-dimensional spatial geometric structure and natural language expression. Unlike existing Point-BERT, CLIP-3D and other models that only rely on shallow embedding similarity alignment, 3D-MELL innovatively introduces a spatial-semantic consistency learning mechanism, constructs a shared semantic embedding space and a cross-modal cross-attention module, and realizes the explicit mapping of building components between point cloud spatial structure and text description. The architecture is composed of a geometric perception coding module (introducing local density and global center perception mechanism), a context semantic understanding module (BERT), and a parameter efficient fine-tuning mechanism (PEFT). The model adopts a "two-stage training strategy": in the first stage, cross-modal feature alignment is performed through structure preservation and contrastive loss; in the second stage, instruction-based language interaction examples are introduced for fine-tuning to improve the model's language generation and interaction ability, thereby realizing semantic recognition, attribute completion and language description of building components and other complex tasks. To address the model failure problem caused by dynamic changes in the construction site point cloud, the present invention further proposes an incremental model updating mechanism in the PEFT architecture. This mechanism combines component ID management and regional difference analysis to quickly fine-tune only the local parameters of new components, with an average time of no more than 4 minutes, without the need for full retraining, thereby realizing real-time semantic fusion and structure updating in dynamic scenarios.
[0165] 2. Three-dimensional point cloud semantic dataset construction method. Considering the professional and high-cost problems of point cloud semantic labeling in complex building scenes, the present invention constructs a three-dimensional semantic point cloud dataset for building components, covering ancient heritage, modern construction, structural components and other types of samples. The dataset construction adopts a human-machine collaborative labeling mechanism: first, an automatic labeling tool based on rules such as component size, normal direction, spatial distribution, etc. is used for initial labeling, and then expert interactive correction is used to generate high-quality semantic labels. This mechanism supports incremental semantic expansion and continuous optimization, and can automatically access the existing labeling process when new scene or new component sampling data is introduced, effectively improving the scalability and diversity of the dataset, providing a solid semantic foundation for model training and generalization.
[0166] 3. Three-dimensional interactive system method fusing semantic question answering and spatial reasoning. Based on the semantic ability of 3D-MELL, the invention further constructs a three-dimensional semantic question answering and spatial reasoning system platform, realizes the integration of component recognition, semantic query, geometric analysis and visual response. The system adopts the browser / server (B / S) architecture, and the front end realizes the loading, interactive selection and semantic highlighting of three-dimensional scene components based on Three.js and WebGL; the back end integrates a lightweight natural language understanding module and a spatial reasoning calculation engine, supports users to ask questions in a natural language manner (such as "please mark the bracket", "which column belongs to which era", "how is this structure connected"), and the system automatically completes semantic analysis and geometric feedback, improves the interpretability and interactivity of point cloud data. At the deployment level, the system supports model pruning, attention sparsification and cache optimization strategies. In the RTXA800 GPU platform test, the average response delay of the system is less than 500ms, and the concurrent processing capacity is ≥5 people, which meets the engineering needs of high responsiveness and real-time interaction in building scenes.
[0167] Multi-modal encoding framework 3D-MELL design
[0168] 3D-MELL network is a complex architecture combining multi-source data and large language models, designed specifically for understanding and describing architectural structural elements. The overall architecture is composed of multiple modules, each with its unique function and design, ensuring that point cloud data and text information can be effectively fused. First, point cloud data applied to building scans needs to go through a series of preprocessing steps such as denoising, standardization, and segmentation, in order to extract individual architectural elements and assign them unique attribute IDs. After preprocessing, the point cloud data is represented as where is the number of points, is the feature dimension of each point, and text data is converted into vector form through natural language processing techniques and is input together with point cloud data for subsequent processing. Point cloud data is processed by a point encoder to generate a sequence of point features where is the number of point features, is the feature dimension. These point feature sequences are then further transformed into embedding vectors through point token embedding and passed to subsequent modules. Large language models can obtain high-dimensional representations of each point, providing more accurate input features for subsequent tasks. After embedding, the point tokens are input into the Geometry-aware Encoder Block. The innovation of this module is to combine self-attention mechanisms with geometric awareness. By introducing a geometric size adjustment mechanism, the module can adjust self-attention weights based on distances between points, local density, and other geometric information, resulting in better spatial structure modeling. All processed features are finally passed to the task-specific head, which generates the final interaction results based on different task requirements. The task-specific head generates precise task outputs such as building element classification and historical background through in-depth analysis of multi-modal information, enabling comprehensive description and analysis of building data.
[0169] Efficient fine-tuning PEFT network technology
[0170] The PEFT strategy proposed in the present application combines LoRA and Prompt Tuning mechanisms, and is particularly adapted for three-dimensional building component recognition and semantic generation tasks. Unlike general PEFT in the NLP field, the PEFT module in 3D-MELL introduces structured prompt vectors based on component function attributes (such as "column-eave-pendant" logic) and dynasty discrimination information to enhance semantic transfer capability. It also supports online incremental learning of new components in dynamic construction scenarios. This mechanism allows the model to quickly fine-tune and restore global semantic consistency when new point clouds are added at the construction site, improving the adaptability and deployment efficiency of the system in actual construction environments. In addition, the PEFT method only needs to adjust a small number of learnable parameters (such as Token prompts), avoiding the need for full-scale training of large pre-trained models, thereby significantly reducing the consumption of computing resources and storage overhead. When fine-tuning multiple tasks, PEFT can avoid training and storing fine-tuned models for each task, speeding up model migration and deployment. Since only newly added parameters are fine-tuned, PEFT avoids the large amount of training time required by traditional full-scale fine-tuning methods, significantly improving the efficiency of large language model training.
[0171] To address the potential for noise and redundant information in point cloud data from actual architectural scans, this paper introduces a density-controlled point screening strategy and a uniform downsampling algorithm (Formula 2-7) during the encoding phase. This strategy also improves model stability by removing isolated points and points with abnormal density. To verify the model's robustness in noisy scenarios, simulated point cloud testing experiments with varying noise levels (σ = 0.01 to 0.1) were conducted. The results show that the model maintains over 85% semantic recognition accuracy under moderate to high noise conditions, an improvement of over 9% compared to a baseline model without noise reduction, demonstrating strong anti-interference capabilities, as shown in Formula 2-1.
[0172] (2-1)
[0173] Where, Indicates the number of embedded tags; for each tag’s feature dimension, the original point cloud is converted into point feature tags that are convenient for further processing by the feature extraction module. In order to enhance the transformer model’s ability to perceive the spatial information of point cloud data, a point priority prompt module is introduced, which uses the original point cloud data to To generate a prompt token with spatial location information, the calculation formula is as follows:
[0174] (2-2)
[0175] Where, To indicate the number of tokens, after getting the point optimization prompt, compare it with the point marking feature of the previous step Combined, so that the transformer module can integrate this information for more accurate feature extraction, as shown in Formula 2-3.
[0176] (2-3)
[0177] At this time, the input feature At the same time, the local geometric features and spatial position information of the point cloud are integrated. The transformer block processes the spliced features through the self-attention mechanism and feedforward network , a total of 12 layers, each layer contains a self-attention mechanism (Self-Attention) and a feed-forward neural network (FFN). layer For the transformer module, its input characteristics are expressed as , perform self-attention calculation on the point cloud vector features, as shown in formula 2-4.
[0178] (2-4)
[0179] Further processing by the feedforward neural network, as shown in equation 2-5, each transformer block extracts features layer by layer, enhancing the model's understanding of the details and global structure in the building point cloud.
[0180] (2-5)
[0181] Output features of each layer of transformer as part of the input features of the next layer of transformer (i.e. ), so as to gradually and deeply extract the spatial structure features of the building point cloud data. During the processing of the transformer module, local detail information in the building point cloud, especially in areas with complex component decorations, will also be captured through a local feature extraction module. Let the local region be represented as a set of points , then the feature extraction formula for local features is shown in equation 2-6.
[0182] (2-6)
[0183] Further, to reduce the computational burden of large-scale point cloud processing and improve model efficiency, a uniform sampling module is used to downsample the original point cloud data The calculation formula is as follows:
[0184] (2-7)
[0185] Perform pooling operation on local features to extract global structural feature representation, as shown in equation 2-8.
[0186] (2-8)
[0187] To further explore the complex spatial dependency in the building point cloud data, a graph convolutional neural network (GCN) is introduced to analyze the topological relationship between points. Assume that the point cloud data is represented by a graph structure , where is the node set of the point set, is the edge set representing the relationship between nodes, and the graph convolution calculation formula is:
[0188] (2-9)
[0189] In the formula, is the adjacency matrix, which describes the spatial topological relationship between points; is the input point feature matrix; is a learnable parameter; is a nonlinear activation function. After GCN processing, high-order graph features The network helps to understand the structural stability and spatial form of the building more deeply. Finally, the features output by the transformer module, the local pooled features, and the graph features processed by the GCN are integrated to form the final fusion feature representation As shown in Equation 2-10, the input for the task-specific head.
[0190] (2-10)
[0191] In the formula, is a multi-dimensional feature vector, which provides a basic support for three-dimensional large language model semantic reasoning. In summary, the PEFT network architecture takes point cloud input as the source data, and gradually completes the feature extraction of building point cloud data through multiple steps such as label embedding, point priority prompt, transformer feature learning, local feature extraction, uniform sampling, pooled feature extraction, graph convolution network analysis, and task-specific head prediction, thereby realizing accurate cognition of building structure and efficient protection decision support. Compared with the application of LoRA / Adapter in general NLP, 3D-MELL introduces a structured prompt mechanism based on component function attributes and spatial region ID in the PEFT structure (Equation 2-1), and combines the characteristics of construction phase changes to propose a “local region fine-tuning + instructional Token update” strategy, which can realize incremental adaptation and rapid fusion of component semantics without retraining the entire model.
[0192] Context understanding BERT encoding structure design
[0193] The encoder transforms the input sequence into a series of contextual representation vectors and consists of multiple identical layers. Each layer consists of two sublayers: a self-attention layer and a feed-forward fully connected layer. The Transformer model represents a major innovation in natural language processing. It is based entirely on the attention mechanism, abandoning the previously widely used recurrent neural networks (RNNs) and convolutional neural networks (CNNs). The core feature of the Transformer is the self-attention layer, which processes each element of the input sequence and compares it with all other elements in the sequence to generate a contextualized representation vector. This mechanism not only improves the model's ability to handle long-range dependencies but also significantly improves efficiency through parallel computation. Building on the success of the Transformer, the BERT model further develops this foundation. Its unique feature lies in its fully bidirectional Transformer encoder structure, which allows the model to simultaneously consider contextual information at all layers. The BERT encoder utilizes a multi-head attention mechanism, which allows each attention head to focus on different aspects of the same information, thereby enriching the model's ability to capture linguistic features.
[0194] The BERT encoder uses the Transformer Encoder as its foundational language model and employs BERT to conduct in-depth analysis of architectural text, extracting information about building features, construction details, and more. The encoder employs multi-head attention, which is used to output multiple attentions and focus on different features within the text. The input query (Q), key (K), and value (V) are mapped to multiple subspaces using different linear transformation matrices. Within each subspace, the attention output of the corresponding head is independently calculated. Ultimately, the outputs of all heads are concatenated and a linear transformation is performed to produce the final attention output, as shown in Formula 2-11.
[0195]
[0196] (2-11)
[0197] Where, 、 、 is the linear transformation matrix for each head; is the output linear transformation matrix; is the number of heads, In the present application, h is set to 32. In the Pre-training task processing stage, it mainly contains two pre-training tasks of Masked Language Model (MLM) and Next Sentence Prediction (NSP). MLM selects 15% of the ancient building mark words in the input for masking, of which 80% are replaced with [MASK] marks, 10% are replaced with random words, and 10% remain unchanged, so as to understand and predict the label by using the context on both sides. NSP selects 50% of the ancient building marks to form the upper and lower sentences together as positive samples, and the other 50% of the sentences randomly select a sentence other than the next sentence to form the upper and lower sentences together as negative samples, which has the ability to abstract continuous long sequences and enhances the understanding of the Bert model to the ancient buildings. In the Fine-Tuning fine-tuning stage, by changing the learning rate, epochs and other parameter settings, overfitting is avoided.
[0198] The BERT model encoder is based on a bidirectional Transformer encoding structure, which can simultaneously focus on the context semantics of the building terminology in the text, so that the model is more comprehensive and accurate in understanding professional building terminology, structural features and historical background information. The multi-head self-attention mechanism built into BERT allows the model to capture rich and complex language features in point cloud data descriptions in parallel, and each attention head can focus on different aspects of the ancient building text, thereby achieving deep semantic modeling of point cloud data. In addition, the BERT pre-training fine-tuning strategy greatly improves the generalization ability for building field-specific terminology and expression methods, enabling the model to more effectively adapt to the association learning between point cloud scenes and text descriptions, and better understand and extract the structure and layout features of ancient buildings.
[0199] Principle of geometric perception encoding module
[0200] The geometric perception encoder is designed for building three-dimensional point clouds, aiming to accurately extract spatial geometric structure information and convert it into high-dimensional semantic representation suitable for language model processing. Unlike methods such as PointTransformer, which only construct attention weights based on the local relative position between points, 3D-MELL introduces a "double-scale geometric modulation mechanism" that considers the local density of points ( ) and the Euclidean distance of points to the global center ( ), dynamically adjusts the attention score (formula 2-13), and realizes spatial relationship modeling with more component-level understanding ability. This mechanism improves the model's ability to capture semantic differences between multi-layer nested structures such as dougong, eaves, and column bases, and is a more discriminative encoding method for building structure modeling.
[0201]
[0202] (2-12)
[0203] where, denotes the set of K nearest neighbors of point ; denotes the spatial coordinates of point ; ; denotes the spatial coordinates of the field point of point ; ; is the number of nearest neighbors; denotes the local density of point ; is the coordinate of the global center point; denotes the distance information of points to local regions or global center points. A geometric-aware weighting mechanism is introduced to combine local density and global center distance to optimize self-attention weights, and the specific formula is 2-13.
[0204] (2-13)
[0205] where, q i is the query vector of point ; k j is the key vector of point ; is the geometric feature weight adjustment coefficient; is the geometric adjustment function, which is responsible for calculating the spatial relationship and density relationship between two points; the index denotes the vector inner product operation. Through this process, the model can not only capture the semantic similarity between points, but also accurately reflect the internal relationship between spatial position and geometric structure. Based on the attention score , the normalized attention weight is further obtained through the Softmax function, and the aggregated feature of the th point cloud token is obtained, as shown in formula 2-14.
[0206]
[0207] (2-14)
[0208] where, v j is the value vector of the th point, which is learned by the geometric features and semantic features of the point together, and the vector output by the self-attention mechanism is further transformed by a two-layer feedforward neural network (with ReLU nonlinear activation function) to calculate the feature, and the calculation formula is as follows:
[0209] (2-15)
[0210] wherein, , are network weights; , is a bias vector. This process is used to strengthen the nonlinear capability of the model feature expression, enhance the perception and expression ability of the encoder to the geometric structure. The geometric feature vector processed by the feedforward neural network is further mapped to a more compact and efficient feature space through the adapter structure , which facilitates the subsequent fusion operation of the large language model. Finally, the feature vector output by the geometric perception encoder is mapped to the space of the three-dimensional large language model through the linear projector.
[0211] Examples
[0212] In a certain Ming and Qing ancient building group, the damaged dougong needs to be repaired in multiple places. The traditional method relies on manual mapping and experience judgment, which is time-consuming and easy to miss details. The technical solution is as follows:
[0213] Step 1, data acquisition and preprocessing
[0214] Equipment: Faro Focus S 350 laser scanner (accuracy 0.1mm) is used, and total station (Leica TS16) is used for spatial positioning.
[0215] Data acquisition: scan to obtain dougong point cloud (about 8 million points for single component), and synchronously collect text data (such as “dougong-Ming dynasty-south hall-cantilever story number 3”). Record the global coordinates of the component through RTK positioning (error <2 cm).
[0216] Preprocessing:
[0217] De-noising: statistical outlier removal algorithm (σ=0.05) is used to remove scaffold shielding noise points (about 15% data reduction).
[0218] Downsampling: voxel grid method (voxel size 3 cm³), data amount is compressed to 3 million points, and dougong mortise and tenon details are retained.
[0219] Segmentation: based on curvature (threshold value 0.008) and normal consistency (angle difference <10°), separate the main body and damaged part of the dougong (ID: DG_01_Defect).
[0220] Step 2. Multi-modal model training and reasoning
[0221] Model input:
[0222] Point cloud: pre-processed dougong data (ID: DG_01).
[0223] Text: Associated component attributes ("Ming Dynasty" "Load-bearing" "South Hall").
[0224] Cross-modal alignment: Calculate local density (K=32 neighbors) and global centroid distance through 3D-MELL model, generate geometric feature vector (dimension 768). BERT module parses the "Ming Dynasty dougong" historical description in the text, outputs semantic vector (dimension 768). Cross-modal contrast loss (τ=0.07) aligns the two features, and the similarity threshold is set to 0.85.
[0225] Semantic reasoning: Input query: "Highlight all Ming Dynasty dougongs with cantilever stories ≥3".
[0226] Model output: Highlight the components that meet the conditions and associate the repair solutions (such as "mortise and tenon reinforcement" "wood preservative treatment").
[0227] 3. Dynamic incremental learning and repair verification
[0228] New component adaptation: Find a Qing Dynasty eave column dougong (ID: DG_New) that has not been entered, scan the point cloud and associate the text ("Qing Dynasty - West Wing - Decorative"). Trigger PEFT incremental fine-tuning: Only update the low-rank matrix (rank r=8) and the structured prompt vector ("dougong - decoration - ZONE_12"), which takes 3 minutes and 42 seconds.
[0229] Real-time interactive verification: Construction personnel ask through B / S system: "Does this dougong have a load-bearing function?"
[0230] System feedback based on fine-tuned model: "No, Qing Dynasty west wing dougong is a decorative component, it is recommended to use light repair materials.
[0231] The technical solution provided by the present application realizes automatic semantic recognition, historical background analysis and repair scheme recommendation of dougongs. The intelligent geometric reasoning and semantic understanding method based on the three-dimensional large language model provided by the present application has been applied to a certain national-level cultural protection project, successfully repaired 23 dougongs, saved more than 60% of labor cost, and adapted to 4 types of newly discovered components through dynamic incremental learning. The consistency of the semantic labels output by the system with the historical archives reaches 98%, providing a reusable technical paradigm for digital protection of ancient buildings.
[0232] The number of devices and the scale of processing described here are used to simplify the description of the present application. Applications, modifications and variations of the present application are obvious to those skilled in the art.
[0233] While embodiments of the application have been disclosed in connection with the above specification and drawings this description is not intended to limit the scope of the application and many modifications, enhancements, alternatives, and variations will become apparent to those skilled in the art from this disclosure. Accordingly, it is intended that the application not be limited to the described embodiments, but that it include all variations falling within the scope of the claims, and their equivalents.
Claims
1. An intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model, characterized by: The following steps are involved: Step 1: Collect original 3D point cloud data of building components through 3D scanning equipment; Generate text data based on the original 3D point cloud data association; Step 2: Preprocess the original 3D point cloud data; Step 3: Build a multimodal 3D large language model, which includes a geometric perception encoding module, a contextual semantic understanding module, and a parameter efficient fine-tuning module; Among them, the geometric perception encoding module calculates geometric features through local density and global center distance, constructs a shared semantic embedding space and a cross-modal cross-attention module; The contextual semantic understanding module is based on a bidirectional transformer encoder and is pre-trained with a masked language model and the next sentence prediction task; The parameter-efficient fine-tuning module combines low-rank adaptation with structured hint vectors to configure an incremental model update mechanism; Step 4: Use the preprocessed 3D point cloud data obtained in step 2 and the text data obtained in step 1 as training sample sets to train and optimize the multimodal 3D large language model; Step 5: Input the 3D point cloud data and text data to be tested into the trained and optimized multimodal 3D large language model, and output the semantic recognition, attribute completion and historical background analysis results of the building components; Among them, the geometric perception encoding module calculates geometric features through local density and global center distance, constructs a shared semantic embedding space and a cross-modal cross-attention module, specifically: 1) Based on the local density ρ of each point in the point cloud data i Distance d from the global center i , construct a dual-scale geometric modulation mechanism; Local density , where is the number of neighboring points, N ( i ) is a point i of K The set of nearest neighbor points, p i 、 p j Points i and its neighboring points j The spatial coordinates of Global center distance , where c is the global center coordinate of the point cloud; 2) Set ρ i with d i As the geometric perception weight, the geometric adjustment function g(d i ,d j ,ρ i ,ρ j ) Dynamically adjust the self-attention score. The specific formula is: , where q i is the query vector of point i; k j is the key vector of point j; is the geometric feature weight adjustment coefficient; It is a geometric adjustment function responsible for calculating the spatial relationship and density relationship between two points; Represents a vector inner product operation; 3) Based on attention score , and further obtain the normalized attention weight through the Softmax function, and aggregate the value vector v of the neighborhood points j , and after feature transformation, output geometric feature vector , the specific calculation formula is: , Where, is the attention score; v j is the value vector of the j-th point; is the vector output by the self-attention mechanism, 、 is the network weight; 、 is the bias vector, is the geometric eigenvector; 4) Geometry-aware feature vector Through linear projection mapping to the shared semantic embedding space, cross-modal attention calculation is performed with text modality features to generate a joint feature representation that integrates geometry and semantics; The original 3D point cloud data collected in steps 1 and 2, and the data to be measured in step 5, are stored and managed in a spatial database, specifically including: Build a spatial database linked to a multimodal 3D large language model: Design component data tables and point cloud data tables. The component data table associates the component ID field in the point cloud data table with the component identification code, forming a one-to-many data mapping relationship. The point cloud data table stores pre-processed point cloud data, including 3D coordinates, color information, and associated component ID fields. A spatial index is constructed on the 3D coordinate fields of the point cloud data table, using an R-tree or octree structure to optimize the following operations: In response to user semantic query commands, the point cloud data of the target component is quickly located through the spatial index and transmitted to the multimodal 3D large language model for semantic parsing; during incremental model updates, the corresponding point cloud data is partitioned and retrieved based on the spatial region ID of the newly added component and loaded into the training process; A dynamic data synchronization mechanism is configured. When the point cloud data of a newly added component is collected and preprocessed by a 3D scanning device, a spatial region ID is automatically assigned according to the component ID field, and a corresponding component data table record is generated and stored in association with the preprocessed point cloud data. Subsequently, the incremental update process of the parameter efficient fine-tuning module is triggered to synchronously update the spatial index and model parameters.
2. The intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model according to claim 1, characterized in that: Step 2 is as follows: De-noising, standardization and segmentation of original 3D point cloud data, extracting the geometric features of each building component and assigning a unique attribute ID; The uniform downsampling algorithm is used to downsample the original 3D point cloud data, and the redundant information is removed through isolated point detection and abnormal density point screening strategies.
3. The intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model according to claim 1, characterized in that: The training optimization of the multimodal 3D large language model in step 4 specifically includes: Based on the training sample set formed by preprocessed 3D point cloud data and text data, the geometry perception encoding module and the contextual semantic understanding module are jointly pre-trained using the cross-modal contrast loss function and the structure preservation loss function to align the geometric features with the text features in a shared semantic embedding space. During the training process, the cross-modal attention module performs bidirectional interaction between the point cloud features output by the geometric perception encoding module and the text features output by the contextual semantic understanding module to generate a joint feature representation that integrates multimodal semantics. Based on the pre-trained model, the adaptability of the parameter efficient fine-tuning module is optimized through task instruction examples: a. Dynamically encode the features of the newly added component region using a structured hint vector, which is generated based on the component's functional attributes and spatial region ID; b. Update the local parameters corresponding to the newly added component area through the incremental model update mechanism, and freeze the remaining parameters.
4. The intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model according to claim 1, characterized in that: The contextual semantic understanding module is based on a bidirectional transformer encoder and is pre-trained with a masked language model and the next sentence prediction task. Specifically: Pre-training based on text data, where: a. In the masked language model task, 15% of the architectural terminology tokens in the input text are masked. The masking strategies include: i. 80% of the masked tokens are replaced with mask symbols; ii. 10% of the masked tokens are replaced with random architectural terms; iii. 10% of the masked tokens remain unchanged; b. In the next sentence prediction task, the positive samples are actual continuous sentence pairs in the architectural text, and the negative samples are randomly combined non-continuous sentence pairs; During the fine-tuning phase, the semantic vectors output by BERT are supervised and aligned using the building component attribute labels.
5. The intelligent geometric reasoning and semantic understanding method based on a three-dimensional large language model according to claim 1, characterized in that: The parameter efficient fine-tuning module combines low-rank adaptation with structured hint vectors to configure an incremental model update mechanism, specifically: A low-rank adaptive mechanism is introduced into the fully connected layer of the multimodal 3D large language model. Specifically: The multimodal three-dimensional large language model weight matrix , decomposed into ,in W 0 is the frozen original weight, Δ W is a trainable low-rank matrix, Δ W= , 、 , and order ; Only for low-rank matrices A 、 B Perform gradient updates and maintain W 0 parameter is fixed; Generate structured hint vector based on component functional attributes and spatial region ID ,in k is the number of prompt vectors, d The hint vector T is embedded into the model input in the following way: T and the point cloud feature vector X point Splicing to form enhanced input Xin =Concat( X point, T ); through the learnable projection matrix Xin Mapping to the model latent space; Configure the incremental model update mechanism, specifically: locate the corresponding local low-rank matrix based on the spatial region ID of the newly added component A new 、 B new and local hint vector T new , only for A new 、 B new and T new Fine-tune the model output after local parameter update and normalize and fuse it with the global model features.
6. Intelligent geometric reasoning and semantic understanding device based on a three-dimensional large language model, characterized by: include: A data acquisition module, which collects original three-dimensional point cloud data of building components through a three-dimensional scanning device; Obtain text data associated with original 3D point cloud data; A data preprocessing module, which is used to preprocess the original three-dimensional point cloud data; A model building module, which is used to build a multimodal three-dimensional large language model, including a geometric perception encoding module, a contextual semantic understanding module, and a parameter efficient fine-tuning module; Among them, the geometric perception encoding module calculates geometric features through local density and global center distance, constructs a shared semantic embedding space and a cross-modal cross-attention module; The contextual semantic understanding module is based on a bidirectional transformer encoder and is pre-trained with a masked language model and the next sentence prediction task; The parameter-efficient fine-tuning module combines low-rank adaptation with structured hint vectors to configure an incremental model update mechanism; A model training and optimization module is used to train and optimize the multimodal three-dimensional large language model using the pre-processed three-dimensional point cloud data obtained in step 2 and the text data obtained in step 1 as training sample sets; The parsing module is used to input the 3D point cloud data and text data to be tested into the trained and optimized multimodal 3D large language model, and output the semantic recognition, attribute completion and historical background analysis results of building components; Among them, the geometric perception encoding module calculates geometric features through local density and global center distance, constructs a shared semantic embedding space and a cross-modal cross-attention module, specifically: 1) Based on the local density ρ of each point in the point cloud data i Distance d from the global center i , construct a dual-scale geometric modulation mechanism; Local density , where is the number of neighbor points, N ( i ) is a point i of K The set of nearest neighbor points, p i 、 p j Points i and its neighboring points j The spatial coordinates of Global center distance , where c is the global center coordinate of the point cloud; 2) Set ρ i with d i As the geometric perception weight, the geometric adjustment function g(d i ,d j ,ρ i ,ρ j ) Dynamically adjust the self-attention score. The specific formula is: , where q i is the query vector of point i; k j is the key vector of point j; is the geometric feature weight adjustment coefficient; It is a geometric adjustment function responsible for calculating the spatial relationship and density relationship between two points; Represents a vector inner product operation; 3) Based on attention score , and further obtain the normalized attention weight through the Softmax function, and aggregate the value vector v of the neighborhood points j , and after feature transformation, output geometric feature vector , the specific calculation formula is: , Where, is the attention score; v j is the value vector of the j-th point; is the vector output by the self-attention mechanism, 、 is the network weight; 、 is the bias vector, is the geometric eigenvector; 4) Geometry-aware feature vector Through linear projection mapping to the shared semantic embedding space, cross-modal attention calculation is performed with text modality features to generate a joint feature representation that integrates geometry and semantics; The original 3D point cloud data collected in steps 1 and 2, and the data to be measured in step 5, are stored and managed in a spatial database, specifically including: Build a spatial database linked to a multimodal 3D large language model: Design component data tables and point cloud data tables. The component data table associates the component ID field in the point cloud data table with the component identification code, forming a one-to-many data mapping relationship. The point cloud data table stores pre-processed point cloud data, including 3D coordinates, color information, and associated component ID fields. A spatial index is constructed on the 3D coordinate fields of the point cloud data table, using an R-tree or octree structure to optimize the following operations: In response to user semantic query commands, the point cloud data of the target component is quickly located through the spatial index and transmitted to the multimodal 3D large language model for semantic parsing; during incremental model updates, the corresponding point cloud data is partitioned and retrieved based on the spatial region ID of the newly added component and loaded into the training process; A dynamic data synchronization mechanism is configured. When the point cloud data of a newly added component is collected and preprocessed by a 3D scanning device, a spatial region ID is automatically assigned according to the component ID field, and a corresponding component data table record is generated and stored in association with the preprocessed point cloud data. Subsequently, the incremental update process of the parameter efficient fine-tuning module is triggered to synchronously update the spatial index and model parameters.
7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the processor executes the program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that A computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
3D visual question and answer method based on three-mode knowledge distillation
CN117216225A
Intelligent media asset multi-modal knowledge graph management method based on large model
CN119294506A
3D point cloud large model dialogue interaction presentation method and device and electronic equipment
CN120318825A
Cited By
Pose temporal evolution and anomaly detection method considering spatial semantic prior
CN122695689A