Progress monitoring method based on digital twinning and BIM information management system

Through the progress monitoring method based on digital twins, using three-dimensional information front view and deep learning technology, the problem that traditional progress monitoring methods are difficult to accurately detect construction activity progress is solved, and more accurate and timely progress monitoring is achieved, and construction efficiency is improved.

CN120047873AActive Publication Date: 2025-05-27YUNTU DATA TECH (ZHENGZHOU) CO LTD

Patent Information

Application Number
CN202510137024.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-27
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

Traditional progress monitoring methods are difficult to accurately detect the progress status of construction activities, especially when multiple progress planning activities are associated with a single BIM element, resulting in insufficient timely and accurate progress monitoring.

Method used

Using a progress monitoring method based on digital twins, the image or video of the construction site is input into the system, and the three-dimensional information front view and deep learning technology are used to perform semantic segmentation to detect and report the percentage of construction activities.

Benefits of technology

It realizes automated monitoring of construction activity progress, provides more accurate and objective progress information, helps construction managers to promptly detect potential delays or problems, and improve construction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047873A_ABST
    Figure CN120047873A_ABST
Patent Text Reader

Abstract

The invention relates to a progress monitoring method based on digital twinning and a BIM (Building Information Modeling) information management system, which can realize automation of a progress monitoring process and can provide more accurate and more objective building activity progress information. By using a three-dimensional information front view and a deep learning-based semantic segmentation technique, the completion percentage of the construction activity can be detected and reported. According to the invention, construction management personnel can be helped to track the progress of main structures such as beams, columns, floor plates and the like, and potential delay or problems can be found as early as possible in the construction process. Through BIM color coding and a point cloud segmentation technology, the progress of each activity can be visualized, so that a project manager is helped to make a more intelligent decision, and resource allocation is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of progress monitoring and BIM information management, and particularly to a progress monitoring method and a BIM information management system based on digital twin. Background Art

[0002] The maturity of building information management technology and the development of real-world data collection, storage, and processing technologies have created a unique opportunity for building researchers and practitioners to automate various project management tasks, such as construction progress monitoring (CPM), quality monitoring, productivity assessment, and safety management. To ensure the successful delivery of a project, effective progress monitoring must be carried out at the construction site. By monitoring the progress of a construction project in a timely manner, construction professionals can identify and resolve any problems that may arise during the construction process, such as delays or quality issues, which helps to improve the overall efficiency of the construction process. Providing up-to-date information on the project status to all stakeholders also increases the opportunity to actively control the project and reduces the risk of project delays and cost overruns. Traditional progress monitoring is accomplished by manual on-site measurement of the constructed works and paper-based daily progress reports (DPRs). Subsequently, the project schedule is updated based on the actual progress percentage obtained from the DPRs. However, in creating a 4D BIM for construction, multiple schedule activities are often associated with a single BIM element using this traditional method. Therefore, it becomes very difficult to infer the progress status of an individual activity, and progress monitoring of the element can only be carried out after all the planned activities are completed. Secondly, most of the existing vision-based progress monitoring methods report progress in a binary form, such as built / unbuilt and completed / uncompleted. Therefore, the progress status of "in-progress" activities is usually not reported. In this case, the role of progress reporting by existing methods is limited because the construction activities are partially completed on the progress monitoring date. The progress status of an activity can only be reported after a particular construction is completed. However, recording the partially completed progress and reporting it as a percentage of completion is necessary for effective project control, and the same progress percentage is required when updating the project schedule. Summary of the Invention

[0003] The object of the present invention is to overcome the deficiencies existing in the prior art and provide a progress monitoring method and a BIM information management system based on digital twin, which provide a way to automate the progress monitoring process and provide more accurate and objective building activity progress information. By using a three-dimensional information front view and deep learning-based semantic segmentation, the percentage of completion of construction activities can be detected and reported. It can help construction managers track the progress of main structures such as beams, columns, and floor slabs and identify potential delays or problems at an early stage during the construction process. Through BIM color coding and point cloud segmentation, the progress of each activity can be visualized, which helps project managers make more informed decisions and improve resource allocation.

[0004] This application first discloses a progress monitoring method based on digital twin. The method includes:

[0005] Step 1: Take the images or videos of the construction site captured by the device as input. After the input of the images or videos, detect the key frames within each second interval and extract them as an image sequence; convert the 360-degree spherical image into an equiangular image through internal image processing, map each equiangular image to a three-dimensional cube, and extract the perspective images of the six faces of up, down, left, right, front, and back, and preprocess the images to eliminate distortion; where the device includes a smartphone camera, a commercial digital camera, a drone, a camera equipped with a spherical 360-degree lens or a fish-eye lens;

[0006] Step 2: According to the overlapping images from different perspectives of the same construction site, detect and match the feature points in the corresponding overlapping images, and use image registration, triangulation, and bundle adjustment to reconstruct a sparse three-dimensional scene, obtain the sparse built-in point cloud model of the construction site and the intrinsic (focal length, optical center, image sensor format) and extrinsic (rotation and translation) camera parameters of all input images; calculate the depth value and normal direction of each pixel from the built-in point cloud model and camera parameters to obtain the depth map and normal map, and generate a dense point cloud model of the three-dimensional scene by fusing multiple images;

[0007] Step 3: Detect whether the elements under construction have been constructed; obtain the construction status of the elements and the number of points in the element bounding box through the BIM model and the point cloud model; if the number of points exceeds the threshold per square meter of surface area, it is considered that the element is under construction and marked as "under construction"; otherwise, it is marked as "not built";

[0008] Step 4: Determine whether the face of the element under construction is dynamically occluded; when there is no dynamic occlusion, select the camera with the largest coverage area for projective transformation and synthesize an orthographic view through projective transformation; when it is occluded, use the neural radiance field to synthesize the orthographic view of the element surface;

[0009] Step 5: Use a convolutional neural network based on masked regions to detect the progress status of the orthographic view;

[0010] Step 6: Calculate the percentage of completion;

[0011]

[0012] Obtain the percentage of completion by taking the ratio of the total pixel area of the active mask to the total pixel area of this face, display it in real time and update it as the activity progresses.

[0013] Step 7: The elements of the planning model are color-coded with traffic signal metaphors to display the progress status; the latest ongoing activities and elements similar to the planned progress are marked in green, elements with previous activities are marked in yellow, and the "no progress" status is marked in red.

[0014] Preferably, a QR code is set at the construction site, and the corresponding construction information on-site can be obtained in real time by scanning the QR code, and the content of the QR code can be added and modified; among them, the construction information includes: on-site progress, production, transportation, and installation information of production factors.

[0015] Preferably, in the rough registration step, three or more corresponding points are manually selected between the planned BIM and the built point cloud; by solving the seven-degree-of-freedom similarity transformation, the point cloud is roughly scaled and transformed into the BIM coordinate system; in the fine registration step, the iterative closest point algorithm is used to closely register the BIM and the point cloud; before the fine registration, the BIM is first converted into an equivalent point cloud by sampling the points around the BIM surface; the point cloud is cleaned by applying statistical outlier removal; during the statistical outlier removal process, the average distance d of point P i in the three-dimensional point cloud from its k nearest neighbor points is calculated i ; if the average distance d of point P i from its k neighboring points exceeds the threshold T, then the point is removed; the threshold is calculated using T = μ d + α * σ d , where α is the standard deviation multiplier, μ d and σ d are the average value and standard deviation of the distance d i ; at the end of the registration process, a 4x4 fine transformation matrix is obtained, which is a combination of rotation, scaling, and translation.

[0016] Preferably, the virtual camera is placed at the center position of the face f, at a distance d f from the normal direction of the face f, looking directly at the face f, and the distance d f between the camera and the face f is calculated based on the diagonal length of the face f and the height of the camera's vertical sensor; the vertex V f of the face f is back-projected onto the pixel coordinates of the camera θ f and θ, and the homography matrix H f is calculated by mapping the corresponding vertex coordinates of the pixel coordinate system; finally, the homography matrix H f is applied to convert the actual image captured by θ f into a front view, and the area within the region of interest is cropped to obtain an orthographic view.

[0017] Preferably, the light emitted by the camera passes through the positions of the pixels and the three-dimensional scene to generate a set of three-dimensional point samplings; the positions and viewing directions of these points are input into a multi-layer perceptron to generate colors and densities corresponding to these points; through volume rendering technology, the colors and densities are accumulated into a two-dimensional image; the scene representation is optimized by minimizing the loss between the synthesized image and the original image; the target face region is cropped from these views by a binary ROI mask generated from the planned BIM.

[0018] Preferably, the convolutional neural network model of the mask region consists of five parts: backbone network, region proposal network (RPN), region of interest alignment layer (ROI Align), class prediction and bounding box regression module, and mask generation module; when an orthogonal view is input, the image passes through different layers of the backbone network and becomes a feature map; the RPN proposes regions in the feature map where class objects may exist; the ROI Align layer extracts feature vectors from the feature map regions proposed by the RPN and converts them into tensors of a fixed size; the class prediction and bounding box regression module is a fully connected layer that predicts class labels and bounding boxes based on the feature vectors aligned by the ROI; the mask generation module consists of a series of transposed convolutional layers and convolutional layers, and this module generates predicted binary masks for each class detected by the class prediction module.

[0019] According to another aspect of the present application, a BIM information management system is provided. This system adopts the above-mentioned progress monitoring method based on digital twin. The system includes: an element input module for inputting lightweight data and structured data; a three-dimensional reconstruction module that uses the SFM three-dimensional reconstruction method to create a dense point cloud and estimates the position of the camera based on the input images; a detection module for detecting whether the elements under construction have been constructed; an orthogonal image synthesis module that synthesizes orthogonal views through projection transformation and NeRF; a semantic segmentation module that uses the mask RCNN model for semantic segmentation; a calculation module for calculating the construction percentage; and a data display module for displaying various data.

[0020] Preferably, the lightweight data includes dynamic data and static data; the structured data includes process, operation, and man-hour information.

[0021] This solution takes 4D BIM and on-site images as inputs and reports the percentage of activity completion. The system uses Structure from Motion and Multi-View Stereo algorithms to create a built reality capture model (point cloud) from the input images and aligns this model with the planned 4D BIM in the same coordinate system. Then, by integrating BIM, reality capture, and the project schedule, a digital twin of the construction project is created to continuously monitor the project progress. At the same time, a deep learning-based semantic segmentation model is trained for activity-level progress inference. The synthetic images are then passed through the trained segmentation model, and an activity segmentation mask is created. The progress percentage of the scheduled activities is calculated based on the ratio of the masked area to the appearance area of the entire face of the elements in the front-facing image. Finally, the planned BIM is color-coded, and the built point cloud is semantically segmented by projecting the 2D segmentation mask onto the 3D point cloud to achieve 3D progress visualization; after a given reality model, it is first necessary to be positioned in the world coordination system with the help of 3D BIM or predefined ground control points. Then, the position information is inferred based on the BIM or ground control points. If there is 3D BIM, objects, i.e., building components (beams, columns, walls, ceilings, and floors), can be detected through occupancy checks. Otherwise, a deep learning-based point cloud segmentation algorithm will be used to detect them. Tasks and materials, such as pouring concrete and fixing formwork, are identified through appearance-based segmentation of the point cloud or images, and the progress percentage of each activity is estimated. Description of the Drawings

[0022] Figure 1 For visualizing the BIM progress through color coding;

[0023] Figure 2 For the schematic diagram of the projective transformation of the image;

[0024] Figure 3 For the network architecture of the NeRF model. Detailed Implementation Manner

[0025] The present invention will be further described below in conjunction with embodiments.

[0026] BIM creation is usually a one-time task, completed in the initial stage of the project. The project schedule prepared by commercial project scheduling software is exported in a file format supported by the BIM coordination tool, which links the schedule activities and 3D model elements by querying tasks and elements respectively. According to the LoD of the model elements and the WBS hierarchy of the project schedule, the links can be divided into three types. In 4D BIM, the activities associated with each element can be represented by a Gantt chart. By filtering the model elements by start or end date, a planned model for a specific date can be obtained. Images or videos of the construction site taken by any device can be used as input, including smartphone cameras, commercial digital cameras, drones, cameras equipped with spherical 360-degree lenses or fisheye lenses. Given video input, the system detects key frames at one-second intervals and extracts them as an image sequence. First, the 360-degree spherical image is converted into an equirectangular image through an internal image processing pipeline. Then, each equirectangular image is mapped to a 3D cube, and perspective images of the six faces, namely top, bottom, right, left, front, and back, are extracted. Similarly, fisheye camera images also need to be preprocessed to eliminate distortion before being used in the next module. 360-degree cameras and fisheye cameras have a high field of view and are usually used for indoor construction monitoring; drones are used to capture outdoor construction data.

[0027] Based on overlapping images from different perspectives of the same construction site, feature points in the corresponding overlapping images are detected and matched, and image registration, triangulation, and bundle adjustment are used to reconstruct a sparse 3D scene, obtaining a sparse in-built point cloud model of the construction site and the intrinsic (focal length, optical center, image sensor format) and extrinsic (rotation and translation) camera parameters of all input images; the depth value and normal direction of each pixel are calculated from the in-built point cloud model and camera parameters to obtain a depth map and a normal map, and a dense point cloud model of the 3D scene is generated by fusing multiple images. Preferably, in the coarse registration step, three or more corresponding points are manually selected between the planned BIM and the built point cloud; the point cloud is roughly scaled and transformed into the BIM coordinate system by solving a seven-degree-of-freedom similarity transformation; in the fine registration step, the iterative closest point algorithm is used to tightly register the BIM and the point cloud; before fine registration, the BIM is converted into an equivalent point cloud by sampling points around the BIM surface; the point cloud is cleaned by applying statistical outlier removal; during the statistical outlier removal process, the average distance d between a 3D point P i and its k nearest neighbor points is calculated i ; if the average distance d between point P i and its k neighboring points exceeds the threshold T, then the point is removed; the threshold is calculated using T = μ d + α * σ d where α is the standard deviation multiplier, μ d and σ d are the distances di The mean value and standard deviation; At the end of the registration process, a 4x4 fine transformation matrix is obtained, which is a combination of rotation, scaling, and translation.

[0028] Detect whether the elements in the construction have been constructed; Obtain the construction status of the elements and the number of points in the element bounding box through the BIM model and the point cloud model; If the number of points exceeds the threshold per square meter of surface area, the element is considered to be under construction and marked as "under construction"; Otherwise, it is marked as "not built"; Specifically: Traverse each element e in the BIM model, calculate the coordinates co of the bounding box of the element e, calculate the center of gravity cog of the bounding box, calculate the size extent of the bounding box, To accommodate registration errors and the presence of temporary structures (such as formwork) around the element, a threshold is added to the element border size, the size of the bounding box is increased by a threshold to obtain a new bounding box, extract the points located within the new bounding box from the point cloud model and calculate the number of points within the new bounding box; If the number of points exceeds the threshold n per square meter of surface area, the element is considered to be under construction and marked as "under construction", otherwise it is marked as "not built". The points within the element border are retained and counted. Finally, a filtered point cloud within the BIM boundary is obtained, as shown in (the example in the experimental section). This filtering method removes unnecessary points outside the region of interest and improves the efficiency of the system. Elements marked as "under construction" will only be passed to the subsequent module for checking the progress status at the activity level.

[0029] In practice, the progress status at the activity level is determined by observing the appearance of the element surface. To accurately measure the progress percentage, the element surface must be observed from the front view position. Oblique views usually result in measurement errors. Determine whether the face of the element under construction is dynamically occluded; When there is no dynamic occlusion, select the camera with the largest coverage area for projective transformation, and synthesize the orthographic view through projective transformation; When it is occluded, use the neural radiance field to synthesize the orthographic view of the element surface;

[0030] As Figure 2 shown, the virtual camera is placed at the center position of the face f, at a distance d from the normal direction of the face f f to face the face f directly, and the distance d between the camera and the face f f is calculated based on the diagonal length of the face f and the height of the camera's vertical sensor; Back-project the vertex V f of the face f to the pixel coordinates of the camera θ f and θ, and calculate the homography matrix H f through the corresponding vertex coordinates of the mapped pixel coordinate system; Finally, apply the homography matrix H f to convert the actual image captured by θ f into a front view and crop the region within the region of interest to obtain an orthographic view.

[0031] Among them, it is necessary to select an optimal view from multiple camera views, which can capture the surface of the target element completely and without occlusion, and the coverage area of this surface in the view is the largest. The specific steps are as follows: 1. Initialization: Traverse each element e in the BIM model. For each face f of each element e, initialize a list cf to store the cameras that can capture face f; 2. Detect visibility: Project a ray from the camera c to the center O of the target face f. f If the ray hits O f , it is considered that the camera c captures face f, and project the vertex V of face f f from the world coordinate system back to the pixel coordinate system of camera c; 3. Check if the vertex is within the field of view: Check if each vertex is within the field of view of camera c. If all vertices Vf are within the field of view of camera c, add camera c to the list cf; 4. Check for dynamic occlusion: Compare the depth maps of the real model and the planned BIM, and check if the target face f is dynamically occluded. If there is dynamic occlusion, remove the camera c from the list cf; 5. Select the best camera: For each face f, calculate the area of face f in the view of each candidate camera c, and select the camera with the largest area for projection transformation.

[0032] As Figure 3 shown, the positions where the light rays emitted by the camera pass through the pixels and the three-dimensional scene generate a set of three-dimensional point samples; input the positions and viewing directions of these points into a multi-layer perceptron to generate the colors and densities corresponding to these points; through volume rendering technology, accumulate the colors and densities into a two-dimensional image; optimize the scene representation by minimizing the loss between the synthesized image and the original image; crop the target face area from these views through the binary ROI mask generated from the planned BIM.

[0033] Use a convolutional neural network based on masked regions to detect the progress status of the orthographic view. Among them, the convolutional neural network of the masked region consists of 5 parts: the backbone network, the region proposal network (RPN), the region of interest alignment layer (ROIAlign), the class prediction and bounding box regression module, and the mask generation module; when the orthographic view is input, the image passes through different layers of the backbone network and becomes a feature map; the RPN proposes regions in the feature map where class objects may exist; the ROI Align layer extracts feature vectors from the feature map regions proposed by the RPN and converts them into tensors of a fixed size; the class prediction and bounding box regression module is a fully connected layer that predicts the class label and bounding box according to the feature vectors aligned by the ROI; the mask generation module consists of a series of transposed convolutional layers and convolutional layers, and this module generates a predicted binary mask for each class detected by the class prediction module.

[0034] Calculation of the percentage of completion;

[0035]

[0036] The percentage of completion is obtained by taking the ratio of the total pixel area of the activity mask to the total pixel area of the face, and is displayed in real time and updated as the activity progresses. For different progress monitoring scenarios, percentage of completion estimation rules are set. For outdoor monitoring, only the progress percentage of the latest activity related to the component needs to be estimated based on the appearance of the outer surface of the component. In indoor construction monitoring, components can be divided into two categories: internal elements and external elements. The estimation of the activity progress is carried out on the inner maximum surface of the external elements on the two largest surfaces of the internal elements. If the two surfaces of the internal element have the same latest activity, the progress percentage is calculated based on the average progress of the two surfaces. If the activity progress on the two surfaces is different, the activity progress is recorded separately.

[0037] QR codes can also be set up at the construction site. By scanning the QR codes, the corresponding construction information on site can be obtained in real time, and the content of the QR codes can be added or modified; among them, the construction information includes: on-site progress, production, transportation, and installation information of production factors. Specifically, the production factors that are manufactured and qualified are transported to the warehouse of the construction unit for storage, and the QR codes on the production factors are scanned before transportation and after being transported to the warehouse to add the transportation information and storage information of the production factors; the situations that occur during the transportation of each production factor are recorded, such as whether it gets rained on, whether there is a collision, etc.; thus facilitating subsequent viewing by the production factor inspectors; for the installed production factors, scan their QR codes and add the installation information to the QR code content; fill in the situations that occur during the installation process, such as collisions, drops, getting rained on, etc., and the remedial measures taken for such situations into the QR code content for later personnel to view; add progress information such as the percentage of completion to the QR code content for the inspectors to connect to the project progress in real time.

[0038] Figure 1 As shown, the elements of the planning model are color-coded with traffic signal metaphors to display the progress status; elements with the latest activity in progress and similar to the planned progress are marked in green, elements with previous activities are marked in yellow, and the "no progress" status is marked in red. The constructed point cloud is also classified according to the progress information at the activity level. All images captured in reality will pass through an image segmentation network and generate a mask color-coded by activity. The pre-computed depth map and the camera pose from the 3D reconstruction module are used to project the 2D mask onto the 3D point cloud, thus generating a semantically segmented in-built point cloud.

[0039] According to another aspect of the present application, a BIM information management system is provided. This system adopts the above-mentioned progress monitoring method based on digital twin. The system includes: an element input module for inputting lightweight data and structured data; a 3D reconstruction module for creating a dense point cloud using the SFM 3D reconstruction method and estimating the position of the camera based on the input images; a detection module for detecting whether the elements under construction have been constructed; an orthographic image synthesis module for synthesizing orthographic images through projective transformation and NeRF; a semantic segmentation module for performing semantic segmentation using the Mask R-CNN model; a calculation module for calculating the construction percentage; and a data display module for displaying various data.

[0040] Specifically, the lightweight data includes dynamic data and static data; the structured data includes process, procedure, and man-hour information.

[0041] This solution takes 4D BIM and on-site images as inputs and reports the percentage of activities completed. The system uses structure from motion and multi-view stereo algorithms to create a built reality capture model (point cloud) from the input images and aligns this model with the planned 4D BIM in the same coordinate system. Then, by integrating BIM, reality capture, and the project schedule, a digital twin of the construction project is created to continuously monitor the project progress. At the same time, a deep learning-based semantic segmentation model is trained for activity-level progress inference. The synthesized images are then passed through the trained segmentation model, and then an activity segmentation mask is created. The progress percentage of the schedule activities is calculated based on the ratio of the masked area to the appearance area of the entire face of the element in the front view image. Finally, the planned BIM is color-coded, and the built point cloud is semantically segmented by projecting the 2D segmentation mask onto the 3D point cloud to achieve 3D progress visualization; after a given reality model is available, it is first necessary to locate it in the world coordination system with the help of 3D BIM or predefined ground control points. Then, the position information is inferred based on the BIM or ground control points. If there is 3D BIM, objects, i.e., building components (beams, columns, walls, ceilings, and floors), can be detected through occupancy checks. Otherwise, a deep learning-based point cloud segmentation algorithm will be used to detect them. Tasks and materials, such as pouring concrete and fixing formwork, are identified through appearance-based segmentation of the point cloud or images, and the progress percentage of each activity is estimated.

[0042] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.

[0043] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0044] In the several embodiments provided in this article, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.

[0045] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments in this article.

[0046] In addition, the functional units in the various embodiments of this article can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0047] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution herein, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments herein. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0048] The above has described in detail one embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made in accordance with the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.

Claims

1. A progress monitoring method based on digital twins, the method comprising: Step 1: Take the construction site image or video taken by the equipment as input, detect the key frames in the second interval after the image or video is input, and extract them as an image sequence; convert the 360-degree spherical image into an equirectangular image through internal image processing, map each equirectangular image into a three-dimensional cube, and extract the perspective images of the six faces of top, bottom, left, right, front and back, and pre-process the image to eliminate distortion; the equipment includes a smartphone camera, a commercial digital camera, a drone, and a camera equipped with a spherical 360-degree lens or a fisheye lens; Step 2: Based on overlapping images of the same construction site from different perspectives, detect and match feature points in the corresponding overlapping images, and use image registration, triangulation, and bundle adjustment to reconstruct a sparse 3D scene, obtain a sparse built-in point cloud model of the construction site, and the intrinsic and extrinsic camera parameters of all input images; calculate the depth value and normal direction of each pixel from the built-in point cloud model and camera parameters to obtain a depth map and a normal map, and generate a dense point cloud model of the 3D scene by fusing multiple images; Step 3: Detect whether the element under construction has been constructed; obtain the element's construction status and the number of points in the element's bounding box through the BIM model and point cloud model; if the number of points exceeds the threshold per square meter of surface area, the element is considered to be under construction and marked as "under construction"; otherwise, it is marked as "not built"; Step 4: Determine whether the face of the element under construction is dynamically blocked; if it is not dynamically blocked, select the camera with the largest coverage area for projection transformation, and synthesize the orthogonal view through projection transformation; if it is blocked, use the neural radiation field to synthesize the orthogonal view of the element surface; Step 5: Use a convolutional neural network based on the mask region to detect the progress status of the orthogonal view; Step 6: Calculation of completion percentage; By taking the ratio of the total pixel area of ​​the activity mask to the total pixel area of ​​the face, the completion percentage is obtained, which is displayed in real time and updated as the activity progresses; Step 7: Elements of the planning model are color-coded using a traffic signal metaphor to show progress status; elements with the latest ongoing activity similar to the planned progress are marked in green, elements with previous activity are marked in yellow, and "no progress" status is marked in red.

2. According to a digital twin-based progress monitoring method according to claim 1, the method further comprises: A QR code is set up at the construction site, and the corresponding construction information of the site can be obtained in real time by scanning the QR code, and the content of the QR code can be added and modified; among them, the construction information includes: site progress, production of production factors, transportation and installation information.

3. A progress monitoring method based on digital twins according to claim 1, wherein: Step 2 also includes: converting the initial camera position to BIM coordinates, which is specifically done as follows: in the coarse registration step, manually selecting three or more corresponding points between the planned BIM and the built point cloud; roughly scaling the point cloud and converting it to the BIM coordinate system by solving the seven-degree-of-freedom similarity transformation; in the fine registration step, tightly registering the BIM and the point cloud using an iterative nearest neighbor algorithm; before fine registration, converting the BIM to an equivalent point cloud by sampling points around the BIM surface; cleaning the point cloud by applying statistical outlier removal; in the statistical outlier removal process, calculating the P in the 3D point cloud. i The average distance d between a point and its k nearest neighbors i ; If point P i If the average distance d to its k neighboring points exceeds the threshold T, the point is removed; the threshold is T = μ d +α*σ d Calculate, where α is the standard deviation multiplier, μ d and σ d is the distance d i At the end of the registration process, we get a 4x4 refined transformation matrix, which is a combination of rotation, scaling, and translation.

4. A progress monitoring method based on digital twins according to claim 1, wherein: In step 4, the camera with the largest coverage area is selected for projection transformation. The specific steps of synthesizing the orthogonal view through projection transformation are as follows: the virtual camera is placed at the center of the surface f, at a distance d from the normal direction of the surface f. f At the face of face f, the distance d between the camera and face f f It is calculated based on the diagonal length of face f and the vertical sensor height of the camera; the vertex V f Back-projection to camera θ f and θ pixel coordinates, and calculate the isomorphism matrix H by mapping the corresponding vertex coordinates of the pixel coordinate system f ; Finally, apply the isomorphism matrix H f θ f The actual image captured is converted into a front view, and the area within the region of interest is cropped to obtain an orthographic view.

5. A progress monitoring method based on digital twins according to claim 1, wherein: In step 4, when occluded, the orthogonal views of the element surface are synthesized using the neural radiation field. The specific steps are as follows: the light emitted by the camera passes through the pixel position and the three-dimensional scene to generate a set of three-dimensional point samples; the position and viewing direction of these points are input into the multi-layer perceptron to generate the color and density corresponding to these points; the color and density are accumulated into a two-dimensional image through volume rendering technology; the scene representation is optimized by minimizing the loss between the synthesized image and the original image; the target facial area is cropped out from these views through the binary ROI mask generated from the planned BIM.

6. A progress monitoring method based on digital twins according to claim 1, wherein: The specific method of step 5 is as follows: the convolutional neural network model of the mask area consists of five parts: the backbone network, the region proposal network (RPN), the region of interest alignment layer (ROI Align), the category prediction and bounding box regression module, and the mask generation module; when the orthogonal view is input, the image will pass through different layers of the backbone network and become a feature map; RPN proposes the area where the class object exists in the feature map; the ROIAlign layer extracts the feature vector from the feature map area proposed by RPN and converts it into a tensor of fixed size; the category prediction and bounding box regression module is a fully connected layer that predicts the category identity and bounding box based on the ROI aligned feature vector; the mask generation module consists of a series of transposed convolutional layers and convolutional layers, which generates a predicted binary mask for each category detected by the category prediction module.

7. A BIM information management system, the system adopts a progress monitoring method based on digital twins according to any one of claims 1 to 6, the system comprising: Element input module, input lightweight data and structured data; The 3D reconstruction module uses the SFM 3D reconstruction method to create a dense point cloud and estimates the camera position based on the input image. The detection module is used to detect whether the elements under construction have been constructed. The orthogonal image synthesis module synthesizes orthogonal images through projection transformation and NeRF. The semantic segmentation module uses the mask RCNN model for semantic segmentation; the calculation module calculates the construction percentage; the data display module is used to display various data.

8. A BIM information management system according to claim 7, wherein: Lightweight data includes dynamic data and static data; structured data includes process, procedure and working time information.

Citation Information

Patent Citations

  • Construction progress visualization method based on BIM model and monitoring image comparison

    CN113360583A

  • Offshore wind power plant construction supervision method and system using image recognition

    CN118570726A

  • Intelligent construction site monitoring method and device based on digital twinning and electronic equipment

    CN118886846A

  • Digital twin modeling method and system for assembling a robotic teleoperation environment

    US20230230312A1

Cited By

  • Building engineering construction quality progress monitoring method and system based on artificial intelligence

    CN120782216A

  • AI intelligent detection system and method based on panorama and BIM

    CN121921642A

  • Water conservancy project three-dimensional perception method and system based on cooperation of unmanned aerial vehicle and video monitoring

    CN122265294A

  • Unmanned aerial vehicle and video monitoring cooperative water conservancy engineering stereoscopic perception method and system

    CN122265294B