Big language model-based parameterized three-dimensional model reverse modeling method and medium

By using coaxial calibration of LiDAR and camera and multimodal large language model to identify targets, light cone extraction point cloud clusters are generated, solving the problems of low efficiency, poor accuracy and insufficient generalization in existing technologies, and realizing efficient and accurate parametric BIM model generation.

CN120930233APending Publication Date: 2025-11-11CHINA CONSTR FOURTH ENG DIV CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511071417.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

The existing Scan2BIM technology is inefficient when dealing with complex structures, irregular shapes, or large-scale point clouds. Its modeling accuracy is affected by noise, and it has poor generalization and adaptability. It cannot edit the generated models, has low computational efficiency, and requires a lot of manual intervention.

Method used

The system employs coaxial calibration of LiDAR and camera, synchronous acquisition of images and point clouds, target identification and light cone generation using a multimodal large language model, extraction of point cloud clusters, and completion of attribute information through internet retrieval to generate parametric BIM components.

Benefits of technology

It improves modeling accuracy and efficiency, enables generalized reverse modeling of complex scenarios, allows editing of native BIM models, reduces computational load, and minimizes manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930233A_ABST
    Figure CN120930233A_ABST
Patent Text Reader

Abstract

The invention discloses a parameterized three-dimensional model reverse modeling method based on a large language model and a medium. The parameterized three-dimensional model reverse modeling method comprises the following steps: acquiring an internal reference matrix of a camera and an external reference matrix of a laser radar and the camera; when the time deviation of the data output by the laser radar and the camera exceeds a set threshold value, the laser radar and the camera are triggered to reset synchronously; inputting image data to the multi-modal large model, identifying building components in the image data, and outputting a segmentation mask with a label; carrying out distortion correction on pixel coordinates of the segmented masks, and converting the pixel coordinates to a laser radar coordinate system by using an internal reference matrix and an external reference matrix; extracting a point cloud cluster from the point cloud data, and executing Euclidean clustering segmentation on the point cloud cluster; semantic tags are added to the point cloud clusters, geometric parameters of components are extracted by inputting the semantic tags to the multi-modal large model, and attribute information is complemented; and generating the parameterized BIM component according to the geometric parameters and the attribute information of the component and the API specification of the BIM software. According to the invention, the modeling precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of building information technology, and in particular to a method and medium for reverse modeling of parametric 3D models based on a large language model. Background Technology

[0002] Currently, the Scan2BIM technology, a reverse engineering method for Building Information Modeling (BIM), is a process that converts point cloud data acquired through 3D laser scanning into a BIM model. It is widely used in architecture, engineering, construction (AEC), and infrastructure management to improve modeling efficiency, reduce manual intervention, and enhance the accuracy of digital models. Existing Scan2BIM technologies are mainly divided into two categories: methods based on traditional algorithms rely on feature matching and geometric shape analysis. While suitable for simple scenarios, these methods are inefficient and require significant manual intervention when dealing with complex structures, irregular shapes, or large-scale point clouds, leading to excessive time and cost. In contrast, methods based on deep learning (such as convolutional neural networks) automatically analyze diverse geometric features through learning from massive amounts of data, significantly improving the level of automation and accuracy in modeling, and have become the mainstream solution for handling complex scenarios.

[0003] The existing technology has the following specific problems:

[0004] 1. Data quality and acquisition issues: Laser scanning may introduce noise due to occlusion, reflection, or moving objects (such as pedestrians and vehicles), affecting modeling accuracy.

[0005] 2. The generated model is based on point cloud reconstruction and cannot be edited like the original model;

[0006] 3. Poor generalization and adaptability; if new features appear in a new scenario, targeted retraining is required.

[0007] 4. The computational efficiency is low due to the large amount of data being processed. Each time, global data needs to be registered, segmented, and modeled, which involves unnecessary redundant and inefficient calculations. Summary of the Invention

[0008] In view of this, the purpose of this invention is to propose a parametric 3D modeling reverse modeling method based on a large language model.

[0009] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows:

[0010] This invention provides a method for reverse modeling of parametric 3D models based on a large language model, comprising the following steps:

[0011] Step 1: Align the LiDAR and the camera's optical axis to obtain the camera's intrinsic parameter matrix and the LiDAR and camera's extrinsic parameter matrices.

[0012] Step 2: Add a GPS timestamp to each frame of point cloud data output by the LiDAR and each frame of image data output by the camera. When the time deviation exceeds the set threshold, trigger the LiDAR and camera to reset synchronously.

[0013] Step 3: Input image data to a multimodal large model, identify the architectural components in the image data and output a segmentation mask with labels;

[0014] Step 4: After distortion correction, the pixel coordinates of the segmented mask are converted into rays in the camera coordinate system using the intrinsic parameter matrix to generate a light cone, and the light cone is converted into the lidar coordinate system using the extrinsic parameter matrix.

[0015] Step 5: Extract subsets located within the light cone from the point cloud data to form point cloud clusters, perform Euclidean clustering on the point cloud clusters and remove outliers;

[0016] Step 6: Add semantic tags to the point cloud clusters, input the geometric parameters of the components from the multimodal large model, and complete the attribute information by searching the Internet;

[0017] Step 7: Generate parametric BIM components based on the geometric parameters, attribute information, and API specifications of the BIM software.

[0018] Furthermore, step 1 specifically includes:

[0019] Step 11: Fix the lidar and camera on the same bracket, and make the lidar and camera optical axis parallel;

[0020] Step 12: Use a microcontroller to trigger the lidar and camera, and set a time deviation threshold as the triggering condition;

[0021] Step 13: Using a checkerboard calibration board, obtain the camera's intrinsic parameter matrix K through Zhang Zhengyou's calibration method:

[0022]

[0023] Among them, f x f represents the horizontal focal length. y s represents the vertical focal length, s represents the distortion coefficient, and c represents the vertical focal length. x c represents the horizontal principal point. y Indicates the vertical principal point;

[0024] s = [k1, k2, p1, p2, k3], where k1, k2, and k3 represent radial distortion coefficients, and p1 and p2 represent tangential distortion coefficients;

[0025] Step 14: Using a checkerboard calibration board, calculate the extrinsic parameter matrix T of the LiDAR and camera using the ICP algorithm. L2C ,T L2C It is a 4×4 homogeneous matrix, and the formula is:

[0026]

[0027] Among them, R L2C It is a 3x3 rotation matrix, representing the rotation relationship from the lidar coordinate system to the camera coordinate system; t L2C It is a 3x1 translation vector, representing the coordinates of the origin of the lidar coordinate system in the camera coordinate system; 1 and 0 T Used to uniformly represent transformations of rotation and translation, 0 T =[0,0,0] is the canonical form of homogeneous coordinates.

[0028] Furthermore, step 2 specifically includes:

[0029] Step 21: LiDAR outputs point cloud data Among them, each point cloud p i =(x i y i , z i Camera output image data I m1×n1 Where m1 represents the horizontal resolution and n1 represents the vertical resolution;

[0030] Step 22: Set the frequency and time point of the output data of the LiDAR and camera to be the same, and add a GPS timestamp to each frame of point cloud data and image data. When the time deviation between the GPS timestamps of two adjacent point cloud data and the time deviation between the GPS timestamps of two adjacent image data are both greater than the threshold, the LiDAR and camera are triggered to reset synchronously through the microcontroller.

[0031] Step 23: Perform voxel grid downsampling and statistical outlier removal filtering on the point cloud data and image data.

[0032] Furthermore, step 3 specifically includes:

[0033] Step 31, transfer image data I m1×n1 Input into a large multimodal model;

[0034] Step 32: Call the multimodal large model to execute the prompt: identify all building components in the image data;

[0035] Step 33: After recognition, output the segmentation mask with labels. The output format is as follows:

[0036]

[0037] Among them, the label list: [l1, l2, ..., lx] is the semantic label of the identified building components;

[0038] Mask list: In, each It is a binary mask matrix:

[0039] The dimensions are the same as the input image, with a horizontal resolution of m1 and a vertical resolution of n1.

[0040] Tag elements are either 0 or 1, where 1 marks the tag l. j In the corresponding area, 0 is marked as the background;

[0041] Correspondence: Labels and masks are matched one-to-one using indices; the j-th label l j Corresponding to the j-th mask Where j represents the index, and the range of j is 1≤j≤k and j is a positive integer.

[0042] Furthermore, step 4 specifically includes:

[0043] Step 41: Divide the pixel coordinates (u) of the segmented mask. c ,v c After distortion correction, the coordinates are converted to the corrected coordinates (u′, v′):

[0044]

[0045] Where r represents the radial distance from the distorted pixel to the principal point; u c v represents the horizontal coordinate of the distorted pixel. c U represents the vertical coordinate of the distorted pixel. d v represents the normalized offset in the horizontal direction after distortion correction. d This represents the normalized offset in the vertical direction after distortion correction, where u′ represents the horizontal coordinate of the corrected pixel and v′ represents the vertical coordinate of the corrected pixel.

[0046] Step 42: Use the camera's intrinsic parameter matrix to convert the corrected coordinates (u′, v′) into 3D rays in the camera coordinate system.

[0047]

[0048] Where d is the depth value corresponding to the point cloud data;

[0049] Step 43: Generate a light cone F with the camera's optical center as the vertex and the mask contour as the base. k ;

[0050] Step 44: Utilize the extrinsic parameter matrix T of the lidar and camera L2C , light cone F k Transform from the camera coordinate system to the lidar coordinate system to determine the spatial region in the point cloud data space that corresponds to the target component in the image data.

[0051] Furthermore, step 5 specifically includes:

[0052] Step 51: Extract the point cloud data located at the light cone F from the point cloud data P. k Subsets within: Forming point cloud clusters;

[0053] Step 52, for Perform Euclidean clustering segmentation with a clustering threshold of 0.05m to remove outliers.

[0054] Furthermore, step 6 specifically includes:

[0055] Step 61: For point cloud clusters Add attribute: semantic tag = l k ;

[0056] Step 62: Cluster the point clouds The corresponding image patches are input into the multimodal large model, and multiple prompts are set, including geometric parameters, output format, and attribute information, as follows:

[0057] (1) Extract geometric parameters from point cloud data and image data;

[0058] (2) Output the result in JSON format, including type, length, width, height, angle and radius;

[0059] (3) Use internet search capabilities to compare with standard models to obtain material information and correct the final parameter dimensions of the component;

[0060] Step 63: The multimodal large model extracts the geometric parameters of the corresponding components by analyzing the spatial distribution of point cloud data and the visual features of image data. At the same time, based on the recognition results, the multimodal large model uses its Internet retrieval capabilities to retrieve the type, brand, and model attribute information of the corresponding components from the standard library to complete the attribute description of the point cloud cluster.

[0061] Furthermore, step 7 specifically includes:

[0062] Step 71: Organize the geometric parameters, attribute information and API specifications of the component into structured prompts and input them into the multimodal large model;

[0063] Step 72: Based on structured prompts, the multimodal large model generates a BIM software script that can be executed directly, realizing the automated creation of parametric BIM components.

[0064] Furthermore, the multimodal large model is GPT-4V or Qwen2.5-VL.

[0065] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for reverse modeling of a parameterized 3D model based on a large language model.

[0066] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art:

[0067] This invention utilizes coaxial images and point cloud data to identify targets using a multimodal large model, calculates light cones, and obtains target point cloud clusters. The target point cloud clusters are then reported to the multimodal large model to extract set features and retrieve parameters from the internet. Finally, using cue word engineering, based on the previously obtained target information and geometric features, a parametric geometric model is generated by calling a BIM modeling software interface (such as Rvit) through the multimodal large model.

[0068] 1. This invention utilizes the reasoning capabilities of multimodal large models to analyze data, and searches the Internet to match it with standard component parameters, thereby improving modeling accuracy;

[0069] 2. This invention uses multimodal large model parametric reverse modeling to obtain the native BIM model, which can be adjusted using BIM software;

[0070] 3. The target recognition and point cloud geometric feature recognition adopt a multimodal large model, which has a strong generalization ability and can realize reverse modeling for a wide range of scenes;

[0071] 4. Using light cones to divide the global point cloud into local segments reduces the amount of computation. By obtaining point cloud clusters, the number of point clouds to be computed is reduced, thus improving processing efficiency. Attached Figure Description

[0072] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0073] Figure 1 This is an execution flowchart of a parametric 3D model reverse modeling method based on a large language model provided in an embodiment of the present invention.

[0074] Figure 2 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of the present invention. Detailed Implementation

[0075] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0076] Please see Figure 1 This invention discloses a parametric 3D model reverse modeling method based on a large language model. This method is achieved through sensor calibration, simultaneous image-point cloud acquisition, large model target recognition and mask generation, light cone constraint space calculation, local point cloud cluster extraction, semantic assignment and feature parsing, and automated generation of BIM components. Specifically, it includes the following steps:

[0077] Step 1, Sensor Calibration: Align the LiDAR with the camera's optical axis and obtain the camera's intrinsic parameter matrix and the extrinsic parameter matrices of the LiDAR and camera. The purpose of this step is to ensure that the data from the LiDAR and the RGB camera are synchronized spatially and temporally. Calibration obtains the camera's intrinsic parameters (such as focal length and distortion coefficients) and the extrinsic parameter transformation matrix between the LiDAR and camera, providing a precise geometric alignment basis for subsequent image and point cloud fusion.

[0078] In this embodiment, step 1 specifically includes:

[0079] Step 11: Fix the LiDAR and the camera (RGB camera) on the same bracket, and make the optical axis of the LiDAR parallel to that of the camera; the industrial camera model is Robot MV-CU120-10UC, and the LiDAR model is Livox.

[0080] Step 12: Use a microcontroller (STM32) to trigger the LiDAR and camera, and set a time deviation threshold as the trigger condition. In this embodiment, the threshold is set to 1ms, and the value can be set according to the user's needs; for example, the threshold is set to 1ms, and the specific value can be set according to the user's needs.

[0081] Step 13: Using a checkerboard calibration board (a commonly used tool in computer vision and photogrammetry for camera calibration, aiming to determine the camera's intrinsic parameters (focal length, principal point, distortion coefficients, etc.) and extrinsic parameters (position, pose)), obtain the camera's intrinsic parameter matrix K using Zhang's Method (Zhang's Method is a camera calibration method based on a planar checkerboard, proposed by Zhengyou Zhang in 1998. This method uses multiple checkerboard images, employing homography and nonlinear optimization to solve for the camera's intrinsic and extrinsic parameters; it is one of the most classic calibration methods in the field of computer vision).

[0082]

[0083] Among them, f x f represents the horizontal focal length. y s represents the vertical focal length, s represents the distortion coefficient, and c represents the vertical focal length. x This represents the horizontal principal point (the principal point is the origin of the image coordinate system, usually the center of the image. It defines the offset between the camera coordinate system and the image coordinate system), c y This represents the vertical principal point; the specific values ​​are as follows:

[0084] Horizontal focal length f x = 4528.6px, vertical focal length f y = 4527.25653px, horizontal principal point c x =1983.24876px, vertical principal point c y =1526.74923px, distortion coefficient s=0.

[0085] s = [k1, k2, p1, p2, k3], where k1, k2, and k3 represent radial distortion coefficients, and p1 and p2 represent tangential distortion coefficients; the specific values ​​are as follows:

[0086] Radial distortion coefficient k1 = -4.61298122 × 10 -2 k2 = 1.44714462, k3 = -1.83630160 × 10, tangential distortion coefficient p1 = 2.31570104 × 10 -4 p2 = -5.53243723 × 10 -5 .

[0087] Step 14: Using a checkerboard calibration board (grid size 90mm × 90mm), calculate the extrinsic parameter matrix T of the LiDAR and camera using the ICP (Iterative Closest Point) algorithm. L2C The registration mean square error (RMSE) is ≤0.03m, T L2CIt is a 4×4 homogeneous matrix, and the formula is:

[0088]

[0089] Among them, R L2C It is a 3x3 rotation matrix, representing the rotation relationship from the lidar coordinate system to the camera coordinate system; t L2C It is a 3x1 translation vector, representing the coordinates of the origin of the lidar coordinate system in the camera coordinate system; 1 and 0 T Used to uniformly represent transformations of rotation and translation, 0 T =[0,0,0] is the canonical form of homogeneous coordinates.

[0090] Step 2, Image-Point Cloud Synchronous Acquisition: Add a GPS timestamp to each frame of point cloud data output by the LiDAR and each frame of image data output by the camera. When the time deviation exceeds the set threshold, the LiDAR and camera are triggered to synchronize and reset. Ideally, the time deviation between each frame of point cloud data output by the LiDAR and each frame of image data output by the camera is ≤0.5ms. When the time deviation exceeds the threshold, a reset will be triggered.

[0091] In this embodiment, step 2 specifically includes:

[0092] Step 21: LiDAR outputs point cloud data Among them, each point cloud p i =(x i y i , z i Camera output image data I m1×n1 Where m1 represents the horizontal resolution and n1 represents the vertical resolution;

[0093] Step 22: Set the frequency and time point of the output data of the LiDAR and camera to be the same, and add a GPS timestamp to each frame of point cloud data and image data. When the time deviation between the GPS timestamps of two adjacent point cloud data and the time deviation between the GPS timestamps of two adjacent image data are both greater than the threshold (e.g., >1ms), the LiDAR and camera are triggered to reset synchronously through the microcontroller.

[0094] Step 23: Perform voxel grid downsampling and statistical outlier removal filtering on the point cloud and image data. Voxel grid downsampling (leaf_size = 0.01m) reduces density, and statistical outlier removal filtering (kNN = 50, std_dev = 1.0) removes noisy points. The purpose of this step is to optimize the original point cloud and image data, including point cloud downsampling (reducing data volume), noise removal (improving data quality), and timestamp synchronization (ensuring data consistency), providing clean and efficient data input for subsequent analysis.

[0095] Step 3, Large-Scale Target Recognition and Mask Generation: Input image data into a multimodal large-scale model, identify architectural components in the image data, and output a list of labels and corresponding mask lists. The purpose of this step is to use a multimodal large-scale model (such as GPT-4V) to identify architectural components in RGB images and generate segmentation masks with semantic labels. This step transforms visual information into structured target regions, providing semantic and spatial constraints for subsequent point cloud cluster extraction.

[0096] In this embodiment, step 3 specifically includes:

[0097] Step 31, transfer image data I m1×n1 In the input multimodal large model, the multimodal large model is GPT-4V or Qwen2.5-VL;

[0098] Step 32: Call the multimodal large model to execute the prompt: identify all building components in the image data;

[0099] Step 33: After recognition, output the segmentation mask with labels. The output format is as follows:

[0100]

[0101] The label list is: [l1, l2, ..., l k [] represents the semantic tags of the identified building components (such as "column", "steel beam", "window");

[0102] Mask list: In, each It is a binary mask matrix:

[0103] The dimensions are the same as the input image, with a horizontal resolution of m1 and a vertical resolution of n1.

[0104] Tag elements are either 0 or 1, where 1 marks the tag l. j In the corresponding area, 0 is marked as the background;

[0105] Correspondence: Labels and masks are matched one-to-one using indices; the j-th label l j Corresponding to the j-th mask Where j represents the index, and the range of j is 1≤j≤k and j is a positive integer.

[0106] Step 4: Calculation of the light cone constraint space: After distortion correction, the pixel coordinates of the segmented mask are converted into rays in the camera coordinate system using an intrinsic parameter matrix to generate a light cone. Then, the light cone is transformed to the LiDAR coordinate system using an extrinsic parameter matrix. The purpose of this step is to generate a view frustum space based on the target mask and camera parameters in the image, mapping the target region in the image to the LiDAR coordinate system. This step ensures that the point cloud data and the image target are aligned in 3D space, providing a precise spatial range for local point cloud cluster extraction.

[0107] In this embodiment, step 4 specifically includes:

[0108] Step 41, Distortion Removal: The pixel coordinates (u) of the segmented mask are... c ,v c After distortion correction, the coordinates are converted to the corrected coordinates (u′, v′):

[0109]

[0110] Where r represents the radial distance from the distorted pixel to the principal point; u c v represents the horizontal coordinate of the distorted pixel. c U represents the vertical coordinate of the distorted pixel. d v represents the normalized offset in the horizontal direction after distortion correction. d This represents the normalized offset in the vertical direction after distortion correction, where u′ represents the horizontal coordinate of the corrected pixel and v′ represents the vertical coordinate of the corrected pixel.

[0111] Step 42: Use the camera's intrinsic parameter matrix to convert the corrected coordinates (u′, v′) into 3D rays in the camera coordinate system.

[0112]

[0113] Where d is the depth value corresponding to the point cloud data (obtained through nearest neighbor interpolation with an interpolation radius of 0.03m);

[0114] Step 43: Generate a light cone F with the camera's optical center as the vertex and the mask contour as the base. k ;

[0115] Step 44: Utilize the extrinsic parameter matrix T of the lidar and camera L2C , light cone F k The coordinate system is transformed from the camera coordinate system to the lidar coordinate system, and the spatial region corresponding to the target component in the image data is determined in the point cloud data space, which provides spatial constraints for the subsequent extraction of local point cloud clusters.

[0116] Step 5, Local Point Cloud Cluster Extraction: Extract subsets located within the light cone from the point cloud data to form point cloud clusters. Perform Euclidean clustering on these clusters and remove outliers. The purpose of this step is to extract local point cloud clusters corresponding to the image target from the full point cloud and remove outliers through clustering. This step significantly reduces computation, focusing only on point cloud data relevant to the target, thus improving the efficiency and accuracy of subsequent processing.

[0117] In this embodiment, step 5 specifically includes:

[0118] Step 51: Extract the point cloud data located at the light cone F from the point cloud data P. k Subsets within: Forming point cloud clusters;

[0119] Step 52, for Perform Euclidean clustering segmentation with a clustering threshold of 0.05m to remove outliers.

[0120] Step 6, Semantic Assignment and Feature Parsing: Add semantic tags to the point cloud clusters, extract the geometric parameters of the components from the multimodal large model, and complete the attribute information through internet retrieval. The purpose of this step is to add semantic tags to the point cloud clusters and extract geometric parameters (such as dimensions, angles, etc.) from the point cloud and images using the large model. At the same time, supplement the standard attributes of the components (such as model and material) through internet retrieval to generate complete component description information, providing a parametric basis for BIM modeling.

[0121] In this embodiment, step 6 specifically includes:

[0122] Step 61: For point cloud clusters Add attribute: semantic tag = l k ;

[0123] Step 62: Cluster the point clouds The corresponding image patches are input into the multimodal large model, and multiple prompts are set, including geometric parameters, output format, and attribute information, as follows:

[0124] (1) Extract geometric parameters from point cloud data and image data;

[0125] (2) Output the result in JSON format, including type, length, width, height, angle and radius;

[0126] (3) Use internet search capabilities to compare with standard models to obtain material information and correct the final parameter dimensions of the component;

[0127] Step 63: The multimodal large model extracts the geometric parameters of the corresponding components by analyzing the spatial distribution of point cloud data and the visual features of image data. At the same time, based on the recognition results, the multimodal large model uses its Internet retrieval capabilities to retrieve the type, brand, and model attribute information of the corresponding components from the standard library, completes the attribute description of the point cloud cluster, and provides detailed parameter basis for generating BIM components.

[0128] Step 7: Automated Generation of BIM Components: Based on the geometric parameters, attribute information, and API specifications of the BIM software, generate parametric BIM components. The purpose of this step is to convert the extracted geometric parameters and semantic information into API call scripts for BIM software (such as Revit), automatically generating parametric components. This step achieves direct conversion from data to model, reducing manual intervention and improving modeling efficiency and accuracy.

[0129] In this embodiment, step 7 specifically includes:

[0130] Step 71: Organize the geometric parameters, attribute information and API specifications of the target BIM software (such as Revit API) of the components into structured prompts and input them into the multimodal large model (such as GPT-4 or DeepSeek R1).

[0131] Step 72: Based on structured prompts, the multimodal large model generates a BIM software script that can be executed directly, realizing the automated creation of parametric BIM components.

[0132] Example prompt template: "Please develop an automation script based on Revit 2025 API to generate a parametric H-section steel column conforming to GB / T11263-2017 standard. This component is required to use Baosteel HW300×300×10×15 steel (material Q345B), with geometric parameters including total length 5.2 meters, section width 0.3 meters, section height 0.4 meters, flange thickness 0.016 meters, web thickness 0.01 meters, and precise positioning at coordinates (X=10.5mm, Y=8.3mm, Z=0.0mm); during implementation, priority should be given to matching the "HW300×300" standard family type in the project family library (allowing 5% geometric tolerance), dynamically binding dimensions, and marking as a load-bearing column."

[0133] Based on the aforementioned prompts, the large model generates a directly executable BIM software script. The geometric parameters output by the large model are parsed and converted into data structures required by the Revit API. Taking steel column modeling as an example: A transaction is started. All family symbols in the document are obtained using the FilteredElementCollector, and then filtered by name to select family symbols of type "H-Shape Steel Column". The base point and axis of the steel column are set. The base point is obtained from the centroid of the point cloud cluster (x, y, z), and the axis is a straight line extending the height from the base point along the Z-axis. A steel column family instance is created at the base point location using the doc.Create.NewFamilyInstance method, specifying the family type and structural type as Column.

[0134] like Figure 2 As shown, this embodiment of the invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for reverse modeling of a parameterized 3D model based on a large language model.

[0135] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0136] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0137] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for reverse modeling of parametric 3D models based on a large language model, characterized in that, Includes the following steps: Step 1: Align the LiDAR and the camera's optical axis to obtain the camera's intrinsic parameter matrix and the LiDAR and camera's extrinsic parameter matrices. Step 2: Add a GPS timestamp to each frame of point cloud data output by the LiDAR and each frame of image data output by the camera. When the time deviation exceeds the set threshold, trigger the LiDAR and camera to reset synchronously. Step 3: Input image data to a multimodal large model, identify the architectural components in the image data and output a segmentation mask with labels; Step 4: After distortion correction, the pixel coordinates of the segmented mask are converted into rays in the camera coordinate system using the intrinsic parameter matrix to generate a light cone, and the light cone is converted into the lidar coordinate system using the extrinsic parameter matrix. Step 5: Extract subsets located within the light cone from the point cloud data to form point cloud clusters, perform Euclidean clustering on the point cloud clusters and remove outliers; Step 6: Add semantic tags to the point cloud clusters, input the geometric parameters of the components from the multimodal large model, and complete the attribute information by searching the Internet; Step 7: Generate parametric BIM components based on the geometric parameters, attribute information, and API specifications of the BIM software.

2. The method for reverse modeling of a parametric 3D model based on a large language model as described in claim 1, characterized in that, Step 1 specifically includes: Step 11: Fix the lidar and camera on the same bracket, and make the lidar and camera optical axis parallel; Step 12: Use a microcontroller to trigger the lidar and camera, and set a time deviation threshold as the triggering condition; Step 13: Using a checkerboard calibration board, obtain the camera's intrinsic parameter matrix K through Zhang Zhengyou's calibration method: Among them, f x f represents the horizontal focal length. y s represents the vertical focal length, s represents the distortion coefficient, and c represents the vertical focal length. x c represents the horizontal principal point. y Indicates the vertical principal point; s = [k1, k2, p1, p2, k3], where k1, k2, and k3 represent radial distortion coefficients, and p1 and p2 represent tangential distortion coefficients; Step 14: Using a checkerboard calibration board, calculate the extrinsic parameter matrix T of the LiDAR and camera using the ICP algorithm. L2C ,T L2C It is a 4×4 homogeneous matrix, and the formula is: Among them, R L2C It is a 3x3 rotation matrix, representing the rotation relationship from the lidar coordinate system to the camera coordinate system; t L2C It is a 3x1 translation vector, representing the coordinates of the origin of the lidar coordinate system in the camera coordinate system; 1 and 0 T Used to uniformly represent transformations of rotation and translation, 0 T =[0,0,0] is the canonical form of homogeneous coordinates.

3. The method for reverse modeling of a parametric 3D model based on a large language model as described in claim 1, characterized in that, Step 2 specifically includes: Step 21: LiDAR outputs point cloud data Among them, each point cloud p i =(x i y i , z i Camera output image data I m1×n1 Where m1 represents the horizontal resolution and n1 represents the vertical resolution; Step 22: Set the frequency and time point of the output data of the LiDAR and camera to be the same, and add a GPS timestamp to each frame of point cloud data and image data. When the time deviation between the GPS timestamps of two adjacent point cloud data and the time deviation between the GPS timestamps of two adjacent image data are both greater than the threshold, the LiDAR and camera are triggered to reset synchronously through the microcontroller. Step 23: Perform voxel grid downsampling and statistical outlier removal filtering on the point cloud data and image data.

4. The method for reverse modeling of a parametric 3D model based on a large language model as described in claim 1, characterized in that, Step 3 specifically includes: Step 31, transfer image data I m1×n1 Input into a large multimodal model; Step 32: Call the multimodal large model to execute the prompt: identify all building components in the image data; Step 33: After recognition, output the segmentation mask with labels. The output format is as follows: The label list is: [l1, l2, ..., l k ] is the semantic tag of the identified building component; Mask list: In, each It is a binary mask matrix: The dimensions are the same as the input image, with a horizontal resolution of m1 and a vertical resolution of n1. Tag elements are either 0 or 1, where 1 marks the tag l. j In the corresponding area, 0 is marked as the background; Correspondence: Labels and masks are matched one-to-one using indices; the j-th label l j Corresponding to the j-th mask Where j represents the index, and the range of j is 1≤j≤k and j is a positive integer.

5. The method for reverse modeling of a parametric 3D model based on a large language model as described in claim 1, characterized in that, Step 4 specifically includes: Step 41: Divide the pixel coordinates (u) of the segmented mask. c ,v c After distortion correction, the coordinates are converted to the corrected coordinates (u′, v′): Where r represents the radial distance from the distorted pixel to the principal point; u c v represents the horizontal coordinate of the distorted pixel. c U represents the vertical coordinate of the distorted pixel. d v represents the normalized offset in the horizontal direction after distortion correction. d This represents the normalized offset in the vertical direction after distortion correction, where u′ represents the horizontal coordinate of the corrected pixel and v′ represents the vertical coordinate of the corrected pixel. Step 42: Use the camera's intrinsic parameter matrix to convert the corrected coordinates (u′, v′) into 3D rays in the camera coordinate system. Where d is the depth value corresponding to the point cloud data; Step 43: Generate a light cone F with the camera's optical center as the vertex and the mask contour as the base. k ; Step 44: Utilize the extrinsic parameter matrix T of the lidar and camera L2C , light cone F k Transform from the camera coordinate system to the lidar coordinate system to determine the spatial region in the point cloud data space that corresponds to the target component in the image data.

6. The method for reverse modeling of a parametric 3D model based on a large language model as described in claim 1, characterized in that, Step 5 specifically includes: Step 51: Extract the point cloud data located at the light cone F from the point cloud data P. k Subsets within: Forming point cloud clusters; Step 52, for Perform Euclidean clustering segmentation with a clustering threshold of 0.05m to remove outliers.

7. The method for reverse modeling of a parametric 3D model based on a large language model as described in claim 1, characterized in that, Step 6 specifically includes: Step 61: For point cloud clusters Add attribute: semantic tag = l k ; Step 62: Cluster the point clouds The corresponding image patches are input into the multimodal large model, and multiple prompts are set, including geometric parameters, output format, and attribute information, as follows: (1) Extract geometric parameters from point cloud data and image data; (2) Output the result in JSON format, including type, length, width, height, angle and radius; (3) Use internet search capabilities to compare with standard models to obtain material information and correct the final parameter dimensions of the component; Step 63: The multimodal large model extracts the geometric parameters of the corresponding components by analyzing the spatial distribution of point cloud data and the visual features of image data. At the same time, based on the recognition results, the multimodal large model uses its Internet retrieval capabilities to retrieve the type, brand, and model attribute information of the corresponding components from the standard library to complete the attribute description of the point cloud cluster.

8. The method for reverse modeling of a parametric 3D model based on a large language model as described in claim 1, characterized in that, Step 7 specifically includes: Step 71: Organize the geometric parameters, attribute information and API specifications of the component into structured prompts and input them into the multimodal large model; Step 72: Based on structured prompts, the multimodal large model generates a BIM software script that can be executed directly, realizing the automated creation of parametric BIM components.

9. A method for reverse modeling of a parametric 3D model based on a large language model as described in claim 4 or 7, characterized in that, The multimodal large model is either GPT-4V or Qwen2.5-VL.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements a parametric 3D model reverse modeling method based on a large language model as described in any one of claims 1 to 9.