A progress monitoring method and BIM information management system based on digital twin
Through digital twin technology and semantic segmentation of deep learning, the real-time and accuracy issues in traditional progress monitoring methods are solved, and real-time progress monitoring and visual management of construction activities are achieved.
Patent Information
- Application Number
- CN202510137024.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-02-07
AI Technical Summary
Traditional progress monitoring methods have difficulty in reporting the progress status of construction activities in real time and accurately, especially when construction activities are not completed, resulting in untimely project management and improper resource allocation.
A digital twin-based progress monitoring method is adopted to detect and report the completion percentage of construction activities through 3D information elevation view and deep learning semantic segmentation technology, and use BIM color coding to visualize the progress.
It enables real-time progress monitoring during the construction process, early detection of potential delays or problems, and improves construction efficiency and the accuracy of resource allocation.
Smart Images

Figure CN120047873B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of progress monitoring and BIM information management, and specifically to a progress monitoring method and BIM information management system based on digital twins. Background Art
[0002] The maturation of building information management (BIM) technology and advances in real-world data collection, storage, and processing techniques have created a unique opportunity for construction researchers and practitioners to automate various project management tasks, such as construction progress monitoring (CPM), quality monitoring, productivity assessment, and safety management. Effective progress monitoring on the construction site is essential to ensuring successful project delivery. By timely monitoring the progress of construction projects, construction professionals can identify and resolve any issues that may arise during construction, such as delays or quality issues, which helps improve the overall efficiency of the construction process. Providing all stakeholders with up-to-date information on project status also increases the opportunity for proactive project control and reduces the risk of delays and cost overruns. Traditionally, progress monitoring is accomplished through manual on-site measurements of constructed work and paper-based daily progress reports (DPRs). The project schedule is then updated based on the actual progress percentages captured in the DPRs. However, this traditional approach, when creating 4D BIM for construction, often results in multiple schedule activities being associated with a single BIM element. Consequently, inferring the progress status of a single activity becomes difficult, and progress monitoring of that element can only be performed after all planned activities have been completed. Second, existing visual-based progress monitoring methods mostly report progress in binary form, such as built / not built and completed / not completed. Therefore, the progress status of "in progress" activities is often not reported. In this case, progress reporting using existing methods is less useful because the construction activities are partially completed at the progress monitoring date. The progress status of an activity can only be reported after a construction activity is completed. However, recording the progress of partially completed work and reporting it as a percentage of completion is necessary for effective project control, and the same percentage of completion is required when updating the project schedule. Summary of the Invention
[0003] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a digital twin-based progress monitoring method and BIM information management system, which provides a method for automating the progress monitoring process and providing more accurate and objective information on the progress of construction activities. By using three-dimensional information elevation views and semantic segmentation based on deep learning, the completion percentage of construction activities can be detected and reported. It can help construction managers track the progress of main structures such as beams, columns, and floor slabs, and detect potential delays or problems as early as possible during the construction process. Through BIM color coding and point cloud segmentation, the progress of each activity can be visualized, helping project managers make more informed decisions and improve resource allocation.
[0004] This application first discloses a progress monitoring method based on digital twins, which includes:
[0005] Step 1: Take images or videos of the construction site captured by the device as input. After the images or videos are input, detect key frames within each second interval and extract them as an image sequence. Convert the 360-degree spherical images into equirectangular images through internal image processing. Map each equirectangular image into a 3D cube, extract perspective images of the top, bottom, left, right, front, and back sides, and pre-process the images to remove distortion. The devices include smartphone cameras, commercial digital cameras, drones, and cameras equipped with spherical 360-degree lenses or fisheye lenses.
[0006] Step 2: Based on overlapping images of the same construction site from different perspectives, feature points in the corresponding overlapping images are detected and matched. A sparse 3D scene is reconstructed using image registration, triangulation, and bundle adjustment. This results in a sparse built-in point cloud model of the construction site and the intrinsic (focal length, optical center, image sensor format) and extrinsic (rotation and translation) camera parameters of all input images. The depth value and normal direction of each pixel are calculated from the built-in point cloud model and camera parameters to obtain a depth map and normal map. Multiple images are then fused to generate a dense point cloud model of the 3D scene.
[0007] Step 3: Detect whether the element under construction has been constructed; obtain the element's construction status and the number of points in the element's bounding box through the BIM model and point cloud model; if the number of points exceeds the threshold per square meter of surface area, the element is considered to be under construction and marked as "under construction"; otherwise, it is marked as "not built";
[0008] Step 4: Determine whether the face of the element under construction is dynamically occluded. If not, select the camera with the largest coverage area for projection transformation, and synthesize an orthogonal view through projection transformation. If occluded, use the neural radiation field to synthesize an orthogonal view of the element surface.
[0009] Step 5: Use a convolutional neural network based on the mask region to detect the progress status of the orthogonal view;
[0010] Step 6: Calculation of completion percentage;
[0011]
[0012] The completion percentage is obtained by taking the ratio of the total pixel area of the activity mask to the total pixel area of the face, which is displayed in real time and updated as the activity progresses.
[0013] Step 7: Elements of the planning model are color-coded to show progress status using a traffic signal metaphor; elements with recent ongoing activity similar to the planned progress are marked in green, elements with previous activity are marked in yellow, and elements with a “no progress” status are marked in red.
[0014] Preferably, a QR code is set up at the construction site, and the corresponding construction information of the site can be obtained in real time by scanning the QR code, and the content of the QR code can be added and modified; wherein the construction information includes: site progress, production of production factors, transportation and installation information.
[0015] Preferably, in the coarse registration step, three or more corresponding points are manually selected between the planned BIM and the as-built point cloud; the point cloud is roughly scaled and converted to the BIM coordinate system by solving the seven-degree-of-freedom similarity transformation; in the fine registration step, the BIM and point cloud are tightly registered using an iterative nearest neighbor algorithm; before fine registration, the BIM is converted to an equivalent point cloud by sampling points around the BIM surface; the point cloud is cleaned by applying statistical outlier removal; during the statistical outlier removal process, the P in the 3D point cloud is calculated. i The average distance d between a point and its k nearest neighbors i ; If point P i If the average distance d to its k neighboring points exceeds the threshold T, the point is removed; the threshold is T = μ d +α*σ d Calculated, where α is the standard deviation multiplier, μ d and σ d is the distance d i At the end of the registration process, a 4x4 refined transformation matrix is obtained, which is a combination of rotation, scaling and translation.
[0016] Preferably, the virtual camera is placed at the center of the surface f, at a distance d from the normal direction of the surface f. f At the face of face f, the distance d between the camera and face f is f It is calculated based on the diagonal length of the face f and the vertical sensor height of the camera; the vertex V of the face f f Back-projection to camera θ f and θ pixel coordinates, and calculate the isomorphism matrix H by mapping the corresponding vertex coordinates of the pixel coordinate system f ; Finally, apply the isomorphism matrix H f θ f The actual image captured is converted into a front view, and the area within the region of interest is cropped to obtain an orthographic view.
[0017] Preferably, light emitted by the camera passes through the pixel positions and the three-dimensional scene to generate a set of three-dimensional point samples; the positions and viewing directions of these points are input into a multi-layer perceptron to generate colors and densities corresponding to these points; the colors and densities are accumulated into a two-dimensional image through volume rendering technology; the scene representation is optimized by minimizing the loss between the synthesized image and the original image; and the target facial area is cropped out from these views using a binary ROI mask generated from the planned BIM.
[0018] Preferably, the convolutional neural network model of the mask area consists of five parts: a backbone network, a region proposal network (RPN), a region of interest alignment layer (ROI Align), a class prediction and bounding box regression module, and a mask generation module; when the orthogonal view is input, the image passes through different layers of the backbone network and becomes a feature map; the RPN proposes areas in the feature map where class objects may exist; the ROI Align layer extracts feature vectors from the feature map areas proposed by the RPN and converts them into a tensor of fixed size; the class prediction and bounding box regression module is a fully connected layer that predicts class identification and bounding box based on the ROI aligned feature vector; the mask generation module consists of a series of transposed convolutional layers and convolutional layers, which generates a predicted binary mask for each class detected by the class prediction module
[0019] According to another aspect of the present application, a BIM information management system is provided, which adopts the above-mentioned digital twin-based progress monitoring method, and the system includes: an element input module, which inputs lightweight data and structured data; a three-dimensional reconstruction module, which uses the SFM three-dimensional reconstruction method to create a dense point cloud and estimates the position of the camera based on the input image; a detection module, which is used to detect whether the elements under construction have been constructed; an orthogonal image synthesis module, which synthesizes orthogonal images through projection transformation and NeRF; a semantic segmentation module, which uses the mask RCNN model for semantic segmentation; a calculation module, which calculates the construction percentage; and a data display module, which is used to display various data.
[0020] Preferably, lightweight data includes dynamic data and static data; structured data includes process, procedure, and working time information.
[0021] This solution takes 4D BIM and site imagery as input and reports the percentage completion of activities. The system uses structure-from-motion and multi-view stereo algorithms to create an as-built reality capture model (point cloud) from the input imagery and aligns this model with the planned 4D BIM in the same coordinate system. This system then creates a digital twin of the construction project by integrating BIM, reality capture, and the project schedule to continuously monitor project progress. Simultaneously, a deep learning-based semantic segmentation model is trained to perform activity-level progress inference. The synthetic image is then passed through the trained segmentation model to create an activity segmentation mask. The progress percentage of the schedule activity is calculated based on the ratio of the masked area to the apparent area of the entire face of the element in the orthographic image. Finally, the planned BIM is color-coded, and the as-built point cloud is semantically segmented by projecting the 2D segmentation mask onto a 3D point cloud for 3D progress visualization. Given the reality model, it is first positioned in a world collaborative system using the 3D BIM or predefined ground control points. Position information is then inferred from the BIM or ground control points. If a 3D BIM is available, occupancy checks can be used to detect objects, namely building components (beams, columns, walls, ceilings, and floors). Otherwise, a deep learning-based point cloud segmentation algorithm is used to detect them. Appearance-based segmentation of point clouds or images is used to identify tasks and materials, such as pouring concrete and fixing formwork, and estimate the progress percentage for each activity. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 To visualize BIM progress through color coding;
[0023] Figure 2 Schematic diagram of image projection transformation;
[0024] Figure 3 This is the network architecture of the NeRF model. DETAILED DESCRIPTION
[0025] The present invention will be further described below with reference to the embodiments.
[0026] BIM creation is typically a one-time effort, completed during the initial project phase. Project schedules compiled with commercial project scheduling software are exported in file formats supported by BIM coordination tools. Coordination tools link schedule activities to 3D model elements by querying tasks and elements, respectively. Links can be categorized into three types, depending on the model element's Location of Dimension (LOD) and the project schedule's Work-Based Structure (WBS) hierarchy. In 4D BIM, activities associated with each element can be represented using a Gantt chart. By filtering model elements by start or end date, a planned model for a specific date can be generated. Construction site images or videos captured by any device can be used as input, including smartphone cameras, commercial digital cameras, drones, and cameras equipped with spherical 360-degree or fisheye lenses. Given the video input, the system detects keyframes within one-second intervals and extracts them as image sequences. An internal image processing pipeline first converts the 360-degree spherical images into equirectangular images. Each equirectangular image is then mapped into a 3D cube, and perspective images are extracted for the top, bottom, right, left, front, and back sides. Similarly, fisheye camera images undergo preprocessing to remove distortion before being used in the next module. 360-degree cameras and fisheye cameras have a wider field of view and are typically used for indoor construction monitoring; drones are used to capture outdoor construction data.
[0027] Based on overlapping images of the same construction site from different perspectives, feature points in the corresponding overlapping images are detected and matched, and a sparse three-dimensional scene is reconstructed using image registration, triangulation, and bundle adjustment. A sparse built-in point cloud model of the construction site and the intrinsic (focal length, optical center, image sensor format) and extrinsic (rotation and translation) camera parameters of all input images are obtained. The depth value and normal direction of each pixel are calculated from the built-in point cloud model and camera parameters to obtain a depth map and normal map, and a dense point cloud model of the three-dimensional scene is generated by fusing multiple images. Preferably, in the coarse registration step, three or more corresponding points are manually selected between the planned BIM and the as-built point cloud; the point cloud is roughly scaled and converted to the BIM coordinate system by solving the seven-degree-of-freedom similarity transformation; in the fine registration step, the BIM and point cloud are tightly registered using an iterative nearest neighbor algorithm; before fine registration, the BIM is converted to an equivalent point cloud by sampling points around the BIM surface; the point cloud is cleaned by applying statistical outlier removal; during the statistical outlier removal process, the P in the 3D point cloud is calculated. i The average distance d between a point and its k nearest neighbors i ; If point P i If the average distance d to its k neighboring points exceeds the threshold T, the point is removed; the threshold is T = μ d +α*σ d Calculated, where α is the standard deviation multiplier, μ d and σ d is the distance di At the end of the registration process, a 4x4 refined transformation matrix is obtained, which is a combination of rotation, scaling and translation.
[0028] Detect whether an element under construction has been constructed. The construction status of the element and the number of points within the element's bounding box are obtained from the BIM model and point cloud model. If the number of points exceeds a threshold per square meter of surface area, the element is considered under construction and marked as "under construction"; otherwise, it is marked as "not built." Specifically, the system traverses each element e in the BIM model, calculates the coordinates co of its bounding box, calculates the center point cog of the bounding box, and calculates the bounding box's size extent. To accommodate registration errors and the presence of temporary structures (such as templates) around the element, the element's bounding box size is increased by a threshold. A new bounding box is obtained by increasing the bounding box size by a threshold. Points within the new bounding box are extracted from the point cloud model and the number of points within the new bounding box is calculated. If the number of points exceeds a threshold per square meter of surface area n, the element is considered under construction and marked as "under construction"; otherwise, it is marked as "not built." Points within the element's bounding box are retained and counted. Finally, a filtered point cloud within the BIM boundary is obtained, as shown in the example in the experimental section. This filtering method removes unnecessary points outside the region of interest, improving system efficiency. Elements marked as "under construction" are only passed to subsequent modules for checking progress status at the activity level.
[0029] In practice, the progress status of the activity level is determined by observing the appearance of the element surface. To accurately measure the progress percentage, the element surface must be observed from the front view position. Oblique views often lead to measurement errors. Determine whether the face of the element under construction is dynamically occluded. If not, select the camera with the largest coverage area for projection transformation and synthesize orthogonal views through projection transformation. If occluded, use the neural radiation field to synthesize orthogonal views of the element surface.
[0030] like Figure 2 As shown, the virtual camera is placed at the center of the surface f, at a distance d from the normal direction of the surface f. f At the face of face f, the distance d between the camera and face f is f It is calculated based on the diagonal length of the face f and the vertical sensor height of the camera; the vertex V of the face f f Back-projection to camera θ f and θ pixel coordinates, and calculate the isomorphism matrix H by mapping the corresponding vertex coordinates of the pixel coordinate system f ; Finally, apply the isomorphism matrix H f θ f The actual image captured is converted into a front view, and the area within the region of interest is cropped to obtain an orthographic view.
[0031] Among them, it is necessary to select an optimal view from multiple camera views, which can capture the surface of the target element completely and without obstruction, and the coverage area of the surface in the view is the largest; the specific steps are: 1. Initialization; traverse each element e in the BIM model, and for each face f of each element e, initialize a list cf to store the cameras that can capture the face f; 2. Detect visibility; move from camera c to the center O of the target face f. f Cast a ray, if the ray hits O f , then it is considered that the camera c captures the face f, and the vertex V of the face f f Back-project from the world coordinate system to the pixel coordinate system of camera c; 3. Check whether the vertex is within the field of view; Check whether each vertex is within the field of view of camera c. If all vertices Vf are within the field of view of camera c, add camera c to the list cf; 4. Check for dynamic occlusion; Compare the depth map of the real model and the depth map of the planned BIM to check whether the target surface f is dynamically occluded. If there is dynamic occlusion, remove the camera c from the list cf; 5. Select the best camera; For each face f, calculate the area of face f in the view of each candidate camera c, and select the camera with the largest area for projection transformation.
[0032] like Figure 3 As shown, the light emitted by the camera passes through the pixel position and the three-dimensional scene to generate a set of three-dimensional point samples; the position and viewing direction of these points are input into a multi-layer perceptron to generate the color and density corresponding to these points; the color and density are accumulated into a two-dimensional image through volume rendering technology; the scene representation is optimized by minimizing the loss between the synthesized image and the original image; the target facial area is cropped out from these views using a binary ROI mask generated from the planned BIM.
[0033] A convolutional neural network based on mask regions is used to detect the progress status of orthogonal views. The convolutional neural network for mask regions consists of five parts: a backbone network, a region proposal network (RPN), a region of interest alignment layer (ROIAlign), a class prediction and bounding box regression module, and a mask generation module. When an orthogonal view is input, the image passes through different layers of the backbone network and becomes a feature map. The RPN proposes regions in the feature map where class objects may exist. The ROI Align layer extracts feature vectors from the feature map regions proposed by the RPN and converts them into a tensor of fixed size. The class prediction and bounding box regression module is a fully connected layer that predicts class identifiers and bounding boxes based on the ROI-aligned feature vectors. The mask generation module consists of a series of transposed convolutional layers and convolutional layers, which generates a predicted binary mask for each category detected by the class prediction module.
[0034] Calculation of percentage of completion;
[0035]
[0036] The completion percentage is calculated by taking the ratio of the total pixel area of the activity mask to the total pixel area of the face, which is displayed in real time and updated as the activity progresses. Percentage completion estimation rules are set for different progress monitoring scenarios. For outdoor monitoring, the progress percentage of the latest activity associated with the component is estimated based on the appearance of the component's outer surface. In indoor construction monitoring, components can be divided into two categories: internal elements and external elements. Activity progress is estimated on the inner largest face of the external element on the two largest faces of the internal element. If both faces of the internal element have the same latest activity, the progress percentage is calculated as the average progress of the two faces. If the activity progress on the two faces is different, the activity progress is recorded separately.
[0037] A QR code can also be set up at the construction site. By scanning the QR code, the corresponding construction information of the site can be obtained in real time, and the content of the QR code can be added and modified. Among them, the construction information includes: on-site progress, production, transportation and installation information of production factors. Specifically, the completed and qualified production factors are transported to the construction unit's warehouse for storage, and the QR code on the production factors is scanned before and after transportation to the warehouse to add the transportation information and storage information of the production factors; the situation that occurs to each production factor during transportation, such as whether it is exposed to rain or collision, etc., so as to facilitate subsequent review by production factor inspectors; for installed production factors, their QR codes are scanned and the installation information is added to the QR code content; the situations that occur during the installation process, such as collisions, drops, rain, etc., and the remedial measures taken for these situations are filled in the QR code content for later review; progress information such as completion percentage is added to the QR code content so that inspectors can connect to the project progress in real time.
[0038] Figure 1 As shown, elements of the planning model are color-coded to indicate progress status, similar to a traffic signal metaphor; elements with the latest ongoing activity similar to the planned progress are indicated in green, elements with previous activity are indicated in yellow, and "no progress" status is indicated in red. The constructed point cloud is also categorized based on activity-level progress information. All captured images are passed through an image segmentation network, generating masks color-coded by activity. Precomputed depth maps and camera poses from the 3D reconstruction module are used to project 2D masks onto the 3D point cloud, resulting in a semantically segmented built-in point cloud.
[0039] According to another aspect of the present application, a BIM information management system is provided, which adopts the above-mentioned digital twin-based progress monitoring method, and the system includes: an element input module, which inputs lightweight data and structured data; a three-dimensional reconstruction module, which uses the SFM three-dimensional reconstruction method to create a dense point cloud and estimates the position of the camera based on the input image; a detection module, which is used to detect whether the elements under construction have been constructed; an orthogonal image synthesis module, which synthesizes orthogonal images through projection transformation and NeRF; a semantic segmentation module, which uses the mask RCNN model for semantic segmentation; a calculation module, which calculates the construction percentage; and a data display module, which is used to display a variety of data.
[0040] Specifically, lightweight data includes dynamic data and static data; structured data includes process, procedure, and working time information.
[0041] This solution takes 4D BIM and site imagery as input and reports the percentage completion of activities. The system uses structure-from-motion and multi-view stereo algorithms to create an as-built reality capture model (point cloud) from the input imagery and aligns this model with the planned 4D BIM in the same coordinate system. This system then creates a digital twin of the construction project by integrating BIM, reality capture, and the project schedule to continuously monitor project progress. Simultaneously, a deep learning-based semantic segmentation model is trained to perform activity-level progress inference. The synthetic image is then passed through the trained segmentation model to create an activity segmentation mask. The progress percentage of the schedule activity is calculated based on the ratio of the masked area to the apparent area of the entire face of the element in the orthographic image. Finally, the planned BIM is color-coded, and the as-built point cloud is semantically segmented by projecting the 2D segmentation mask onto a 3D point cloud for 3D progress visualization. Given the reality model, it is first positioned in a world collaborative system using the 3D BIM or predefined ground control points. Position information is then inferred from the BIM or ground control points. If a 3D BIM is available, occupancy checks can be used to detect objects, namely building components (beams, columns, walls, ceilings, and floors). Otherwise, a deep learning-based point cloud segmentation algorithm is used to detect them. Appearance-based segmentation of point clouds or images is used to identify tasks and materials, such as pouring concrete and fixing formwork, and estimate the progress percentage for each activity.
[0042] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this document.
[0043] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0044] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices, or units, or can be an electrical, mechanical, or other form of connection.
[0045] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments herein.
[0046] In addition, the functional units in the various embodiments herein may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0047] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this article is essentially or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this article. The aforementioned storage medium includes: various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0048] The above is a detailed description of an embodiment of the present invention. However, the content described is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A progress monitoring method based on digital twins, the method comprising: Step 1: Take images or videos of the construction site captured by the device as input. After the images or videos are input, key frames are detected within each second interval and extracted as an image sequence. The 360-degree spherical images are converted into equirectangular images through internal image processing. Each equirectangular image is mapped into a 3D cube, and perspective images of the top, bottom, left, right, front, and back sides are extracted. The images are then pre-processed to remove distortion. The devices include smartphone cameras, commercial digital cameras, drones, and cameras equipped with spherical 360-degree lenses or fisheye lenses. Step 2: Based on overlapping images of the same construction site from different perspectives, feature points in the corresponding overlapping images are detected and matched. A sparse 3D scene is reconstructed using image registration, triangulation, and bundle adjustment. This results in a sparse built-in point cloud model of the construction site and the intrinsic and extrinsic camera parameters of all input images. The depth value and normal direction of each pixel are calculated from the built-in point cloud model and camera parameters to obtain a depth map and normal map. Multiple images are then fused to generate a dense point cloud model of the 3D scene. Step 3: Detect whether the element under construction has been constructed. The construction status of the element and the number of points in the element's bounding box are obtained through the BIM model and point cloud model. If the number of points exceeds the threshold per square meter of surface area, the element is considered to be under construction and marked as "under construction"; otherwise, it is marked as "not constructed"; Step 4: Determine whether the face of the element under construction is dynamically occluded. If not, select the camera with the largest coverage area for projection transformation, and synthesize an orthogonal view through projection transformation. If occluded, use the neural radiation field to synthesize an orthogonal view of the element surface. Step 5: Use a convolutional neural network based on the mask region to detect the progress status of the orthogonal view; The specific method of step 5 is as follows: the convolutional neural network model of the mask region consists of five parts: the backbone network, the region proposal network, the region of interest alignment layer, the category prediction and bounding box regression module, and the mask generation module; when the orthogonal view is input, the image will pass through the different layers of the backbone network and become a feature map; the RPN proposes the region in the feature map where the class object exists; the ROI Align layer extracts the feature vector from the feature map region proposed by the RPN and converts it into a fixed-size tensor; the category prediction and bounding box regression module is a fully connected layer that predicts the category identity and bounding box based on the ROI aligned feature vector; the mask generation module consists of a series of transposed convolutional layers and convolutional layers, which generates a predicted binary mask for each category detected by the category prediction module; Step 6: Calculation of completion percentage; ; The completion percentage is obtained by taking the ratio of the total pixel area of the activity mask to the total pixel area of the face, which is displayed in real time and updated as the activity progresses; Step 7: Elements of the planning model are color-coded to show progress status using a traffic signal metaphor; elements with the latest ongoing activity similar to the planned progress are marked in green, elements with previous activity are marked in yellow, and "no progress" status is marked in red.
2. A digital twin-based progress monitoring method according to claim 1, further comprising: A QR code is set up at the construction site. By scanning the QR code, the corresponding construction information of the site can be obtained in real time, and the content of the QR code can be added and modified; among them, the construction information includes: site progress, production of production factors, transportation and installation information.
3. A progress monitoring method based on digital twins according to claim 1, wherein: Step 2 also includes: converting the initial camera position to BIM coordinates, which is specifically done as follows: in the coarse registration step, three or more corresponding points are manually selected between the planned BIM and the as-built point cloud; the point cloud is roughly scaled and converted to the BIM coordinate system by solving the seven-degree-of-freedom similarity transformation; in the fine registration step, the BIM and point cloud are tightly registered using an iterative nearest neighbor algorithm; before fine registration, the BIM is converted to an equivalent point cloud by sampling points around the BIM surface; the point cloud is cleaned by applying statistical outlier removal; during the statistical outlier removal process, the number of points in the 3D point cloud is calculated. Point and The average distance between neighboring points ; If the point with it The average distance between neighboring points Exceeding the threshold , then the point is removed; the threshold is used Calculate, where is the standard deviation multiplier, and It's distance At the end of the registration process, a 4x4 refined transformation matrix is obtained, which is a combination of rotation, scaling and translation.
4. A progress monitoring method based on digital twins according to claim 1, wherein: In step 4, the camera with the largest coverage area is selected for projection transformation. The specific steps of synthesizing orthogonal views through projection transformation are as follows: The virtual camera is placed on the surface The center position of the distance surface The distance in the normal direction At the right side , camera and face distance It is based on the surface The diagonal length of the camera and the vertical sensor height are calculated; vertex Back-projection to camera and The pixel coordinates are mapped to the corresponding vertex coordinates of the pixel coordinate system and the isomorphism matrix is calculated. ; Finally, apply the isomorphism matrix Will The actual image captured is converted into a front view, and the area within the region of interest is cropped to obtain an orthographic view.
5. A progress monitoring method based on digital twins according to claim 1, wherein: In step 4, when occluded, the neural radiation field is used to synthesize the orthogonal view of the element surface. The specific steps are as follows: the light emitted by the camera passes through the pixel position and the three-dimensional scene to generate a set of three-dimensional point samples; the position and viewing direction of these points are input into the multi-layer perceptron to generate the color and density corresponding to these points; the color and density are accumulated into a two-dimensional image through volume rendering technology; the scene representation is optimized by minimizing the loss between the synthesized image and the original image; the target facial area is cropped out from these views by generating a binary ROI mask from the planned BIM.
6. A BIM information management system, the system adopting the digital twin-based progress monitoring method according to any one of claims 1 to 5, the system comprising: Element input module, input lightweight data and structured data; The 3D reconstruction module uses the SFM 3D reconstruction method to create a dense point cloud and estimates the camera position based on the input image. The detection module is used to detect whether the elements under construction have been constructed. The orthogonal image synthesis module synthesizes orthogonal images through projection transformation and NeRF. The semantic segmentation module uses the mask RCNN model for semantic segmentation; the calculation module calculates the construction percentage; and the data display module is used to display various data.
7. A BIM information management system according to claim 6, wherein: Lightweight data includes dynamic data and static data; structured data includes process, procedure, and working time information.
Citation Information
Patent Citations
Intelligent construction site monitoring method and device based on digital twinning and electronic equipment
CN118886846A
Digital twin modeling method and system for assembling a robotic teleoperation environment
US20230230312A1