An image-based construction project progress analysis method and device

By using multimodal data fusion and temporal modeling techniques, the occlusion areas at the construction site were identified and reconstructed, thus solving the problem of the impact of dynamic occlusions on construction progress analysis and enabling real-time and accurate assessment of construction progress.

CN122388906APending Publication Date: 2026-07-14HANGZHOU WEST RAILWAY STATION HUB DEV CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU WEST RAILWAY STATION HUB DEV CO LTD
Filing Date
2026-04-10
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Dynamic obstructions at construction sites make it difficult for traditional progress analysis methods based on single-view RGB images to accurately reconstruct the structural state and construction completion status of the obstructed areas, resulting in problems such as poor real-time performance, data redundancy, and low accuracy in obstruction recovery.

Method used

By collecting multimodal data from the construction site, spatiotemporal alignment and data-level fusion are performed. Target detection networks and temporal modeling methods are used to identify dynamic occlusions, generate spatiotemporal distribution features, and reconstruct the image content of the occluded area through generative adversarial networks. Combined with 3D point cloud data, 3D spatial structure reconstruction is performed to generate a compensated complete scene model.

Benefits of technology

It enables real-time and accurate assessment of construction progress, improves the robustness and accuracy of construction progress analysis, and effectively addresses the impact of dynamic obstructions on progress analysis in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122388906A_ABST
    Figure CN122388906A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of engineering management, and discloses a construction project progress analysis method and device based on images. The method comprises the following steps: collecting image data, depth data and three-dimensional point cloud data of a construction site, performing space-time alignment and data-level fusion, and constructing a real-time fusion scene representation; based on a target detection network and a time sequence modeling method, a dynamic occlusion object is identified and tracked to generate space-time distribution features; according to the features, the spatial position of an occluded area is extracted, and based on the image and depth data of the area that is not occluded, the image content of the occluded area is generated through a generative adversarial network, and three-dimensional space structure reconstruction is performed in combination with three-dimensional point cloud data to obtain a complete scene model after compensation; and the scene model and a construction progress plan are input into a progress reasoning network to output a construction progress evaluation result. The application can effectively overcome the interference of dynamic occlusion objects on progress monitoring in a complex construction environment, and improve the accuracy and robustness of construction progress evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of engineering management technology, specifically to a method and equipment for analyzing construction project progress based on images. Background Technology

[0002] In image-based construction progress monitoring, occlusion is a core bottleneck affecting analysis accuracy. Occlusion at construction sites comes not only from static objects such as building structures and material piles, but also from dynamic objects such as cranes, construction workers, and transport vehicles. Dynamic occlusion is characterized by randomness, complex movement trajectories, and time-varying occlusion ranges, leading to frequent omissions of key construction areas in images. Traditional progress analysis methods based on single-view RGB images struggle to accurately reconstruct the structural state and construction completion status of occluded areas. Existing methods typically rely on manual re-capture or multi-viewpoint deployment, resulting in poor real-time performance, data redundancy, and low occlusion recovery accuracy.

[0003] Therefore, there is an urgent need for a construction progress assessment method that can detect dynamic obstructions in real time and intelligently compensate for the obstructed areas, so as to improve the robustness and accuracy of progress monitoring in complex construction environments. Summary of the Invention

[0004] To address the technical problems mentioned in the background section, the purpose of this application is to provide a method, device, electronic device, storage medium, and computer program product for image-based construction project progress analysis.

[0005] According to the first aspect of this application, a method for analyzing the progress of construction projects based on images is provided, comprising the following steps: S1. Collect multimodal data from the construction site, including image data, depth data, and 3D point cloud data. Perform spatiotemporal alignment and data-level fusion on the multimodal data to construct a real-time fused scene representation of the construction site. S2, based on the target detection network, performs real-time identification of dynamic occlusions in the fused scene representation, and uses the temporal modeling method to continuously track the motion trajectory and occlusion range of dynamic occlusions, generating the spatiotemporal distribution characteristics of dynamic occlusions; S3. Based on the spatiotemporal distribution characteristics of the dynamic occlusion, the spatial location of the occluded area is extracted from the fused scene representation. Using the unoccluded area image and depth data as conditions, the image content of the occluded area is generated through a generative adversarial network. The image content and the 3D point cloud data are combined to reconstruct the 3D spatial structure of the occluded area to obtain the compensated complete scene model. S4. Input the compensated complete scenario model and construction schedule plan into the schedule inference network, and output the construction schedule evaluation result.

[0006] According to a second aspect of this application, an image-based construction project progress analysis device is provided, the device comprising: The multimodal data acquisition module is used to collect multimodal data from the construction site, including image data, depth data, and 3D point cloud data. It performs spatiotemporal alignment and data-level fusion on the multimodal data to construct a real-time fused scene representation of the construction site. The data fusion module uses a target detection network to identify dynamic occlusions in the fused scene representation in real time, and uses a temporal modeling method to continuously track the motion trajectory and occlusion range of dynamic occlusions, generating the spatiotemporal distribution characteristics of dynamic occlusions. The dynamic occlusion detection and tracking module is used to extract the spatial location of the occluded area from the fused scene representation based on the spatiotemporal distribution characteristics of the dynamic occlusion. Using the unoccluded area image and depth data as conditions, it generates the image content of the occluded area through a generative adversarial network. Combining the image content and 3D point cloud data, it reconstructs the 3D spatial structure of the occluded area to obtain a compensated complete scene model. The occlusion compensation and reconstruction module is used to input the compensated complete scene model and the construction schedule plan into the schedule inference network and output the construction schedule evaluation results.

[0007] According to a third aspect of this application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements the method as described in any of the preceding claims.

[0008] According to a fourth aspect of this application, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the preceding claims.

[0009] According to a fifth aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method as described in any of the preceding claims.

[0010] This invention acquires real-time image data, depth data, and 3D point cloud data from the construction site. Utilizing spatiotemporal alignment and data-level fusion technologies, it constructs a real-time representation of the construction scene, enabling the system to accurately capture the impact of dynamic occlusions on construction progress. Furthermore, through object detection networks and temporal modeling methods, it identifies dynamic occlusions in real time and tracks their trajectories, generating spatiotemporal distribution characteristics to accurately predict the impact of occlusions on construction progress.

[0011] Compared with traditional methods, this invention can effectively solve the impact of dynamic obstructions on progress analysis in complex construction site environments, improve the accuracy and robustness of construction progress assessment, and has significant technical advantages. It can be widely applied to real-time progress monitoring and management in fields such as construction and engineering. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0013] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A flowchart illustrating a construction project progress analysis method based on imagery, provided as an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the progress inference network provided in an embodiment of this application; Figure 3 A schematic diagram of the structure of an image-based construction progress analysis device provided in an embodiment of this application; Figure 4 Another structural schematic diagram of an image-based construction progress analysis device provided in this application embodiment. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0015] In this embodiment, the construction progress monitoring system can be deployed at the construction site and includes a multimodal sensor array, a data processing server, and a visualization terminal. The multimodal sensor array may include RGB cameras, depth sensors (such as binocular cameras or structured light cameras), and LiDAR, used to simultaneously acquire image data, depth data, and 3D point cloud data from the construction site. The data processing server is used to execute image processing and progress analysis algorithms and may include an embedded processor, an industrial computer, or a computing platform equipped with a GPU. The visualization terminal is used to display progress assessment results and early warning information to construction management personnel.

[0016] like Figure 1 As shown in the figure, this application discloses a construction project progress analysis method based on images, including the following steps: S1. Collect multimodal data from the construction site, including image data, depth data, and 3D point cloud data. Perform spatiotemporal alignment and data-level fusion on the multimodal data to construct a real-time fused scene representation of the construction site. Image data, depth data, and 3D point cloud data are simultaneously acquired using a multimodal sensor array deployed at the construction site. Image data, acquired by RGB cameras, provides color texture information of the construction site; depth data, acquired by depth sensors, reflects the distance information from object surfaces to the camera; and 3D point cloud data, acquired by LiDAR, provides high-precision 3D spatial geometry information. Due to differences in sampling frequency, coordinate system, and data format among the various sensors, direct fusion would lead to spatial misalignment and temporal asynchrony. Therefore, spatiotemporal alignment and data-level fusion of the multimodal data are necessary.

[0017] In some embodiments, the multimodal data is spatiotemporally aligned and fused at the data level to construct a real-time fused scene representation of the construction site, including: S11 performs time stamp synchronization and spatial coordinate system registration on the acquired image data, depth data and 3D point cloud data to achieve spatiotemporal alignment of multimodal data; Specifically, timestamp synchronization can be achieved through hardware triggering, where the synchronization controller sends a unified clock signal to each sensor to ensure that each sensor collects data at the same time; or through software interpolation, using the sensor with the highest frequency as a reference to align the timestamps of the data from other sensors.

[0018] Spatial coordinate system registration requires pre-calibrating the intrinsic and extrinsic parameters of each sensor to unify depth data and 3D point cloud data into the camera coordinate system of the RGB camera or into the world coordinate system of the construction site. For example, Zhang Zhengyou's calibration method can be used to perform dual-target calibration of the RGB camera and depth sensor to obtain the rotation matrix and translation vector between them. For LiDAR, the transformation relationship between its coordinate system and the RGB camera coordinate system can be calculated through point cloud registration algorithms (such as the iterative nearest point algorithm), thereby achieving spatial alignment of multimodal data.

[0019] S12 employs pixel-level or feature-level fusion methods to integrate spatiotemporally aligned multimodal data into a unified fusion scene representation, and uses Kalman filtering to eliminate sensor measurement errors and environmental interference in order to generate a real-time updated fusion scene representation.

[0020] Specifically, pixel-level fusion aligns the RGB values ​​of image data with depth data, directly combining them at the pixel level into RGB-D data, which is then spatially mapped with 3D point cloud data to form a dense fused scene representation. Feature-level fusion extracts key features (such as edges, corners, and planes) from image data, depth data, and 3D point cloud data respectively, and associates and fuses the extracted features to form a compact scene feature representation.

[0021] Considering the measurement noise present in the sensor during the acquisition process, and the environmental interference such as changes in lighting and vibration at the construction site, this embodiment uses the Kalman filter algorithm to perform recursive filtering on the fused scene representation. Specifically, based on the scene state estimate of the previous moment and the observation value of the current moment, the optimal state estimate of the current moment is iteratively calculated through two steps of prediction and update, thereby eliminating random noise and generating a smooth, real-time updated fused scene representation.

[0022] S2, based on the target detection network, performs real-time identification of dynamic occlusions in the fused scene representation, and uses the temporal modeling method to continuously track the motion trajectory and occlusion range of dynamic occlusions, generating the spatiotemporal distribution characteristics of dynamic occlusions; Dynamic occlusions include objects such as construction workers, transport vehicles, and lifting machinery whose positions change over time at the construction site. The aforementioned fused scene representation is input into an object detection network, which can employ a convolutional neural network-based object detection model (such as the YOLO series, Faster R-CNN, etc.) to identify dynamic occlusions in the fused scene representation in real time and output the occlusion category, bounding box location, and confidence score.

[0023] Because dynamic occlusions exhibit continuous motion and predictable trajectories, single-frame recognition results are easily affected by factors such as occlusion and changes in illumination. Therefore, a temporal modeling method is used to continuously track the motion trajectory and occlusion range of the occluded objects to generate stable spatiotemporal distribution features. Specifically: In some embodiments, a temporal modeling method is used to continuously track the motion trajectory and occlusion range of dynamic occlusion objects to generate spatiotemporal distribution characteristics of the dynamic occlusion objects, including: S21, extract the bounding box position and category information of dynamic occluders in each frame of the image, construct the detection sequence of dynamic occluders, and input the detection sequence into the temporal modeling network to perform temporal modeling and future position prediction of dynamic occluders to obtain the predicted motion trajectory; wherein, the temporal modeling network adopts LSTM or Transformer architecture. Specifically, in the fused scene representation of consecutive frames, the object detection network outputs the detection results of dynamic occlusions frame by frame, and each detection record includes the coordinates of the top-left corner of the bounding box. lower right corner coordinates Category confidence and category label. For each dynamically occluded object being tracked, its historical detection results are organized into a detection sequence in chronological order. ,in For a moment The detected feature vector is composed of the bounding box center coordinates, width, height, and class embedding vector. This sequence serves as input to a temporal modeling network to learn the motion patterns of occluded objects over time and predict their positions in the future.

[0024] (1) When using LSTM as the temporal modeling network, the network consists of two layers of bidirectional LSTM units stacked together, with each layer having a hidden state dimension of 128. Detection sequence The inputs are sequentially fed into the LSTM at time steps. Each LSTM unit at each time step is configured based on the current input. Compared to the previous hidden state Calculate the current hidden state and through a fully connected output layer Mapped to the next time step (i.e. Bounding box offset of a frame The predicted bounding box position for the future time step is obtained by superimposing it with the bounding box coordinates of the current frame. During the training phase, the network uses the next frame bounding box of the real trajectory as the supervision signal and adopts a smooth L1 loss function to optimize the network parameters, so that the prediction result gradually approximates the real motion trajectory.

[0025] During the inference phase, for occluded objects that have been continuously tracked for multiple frames, LSTM outputs the bounding box positions for several future frames (e.g., frames 1-3) based on the historical detection sequence, forming a predicted motion trajectory. When the occluded object is temporarily occluded or the detector temporarily fails, this prediction result can be used to maintain the continuity of the trajectory and prevent trajectory interruption.

[0026] (2) When using the Transformer architecture, the detection sequence is converted into a fixed-dimensional feature vector through a linear embedding layer, and then added to the positional encoding before being input into the Transformer encoder. The encoder consists of multiple multi-head self-attention layers and a feedforward network stacked together. The self-attention mechanism can capture complex motion patterns of occluded objects over a long time range by calculating the attention weights between any two time steps in the detection sequence, such as periodic back-and-forth motion, variable speed motion, or trajectory intersection. The contextual features output by the encoder are mapped to the bounding box prediction values ​​of future frames through the decoder or linear projection layer. Compared with LSTM, Transformer can more accurately model motion trends in long historical windows through parallel computing and a global receptive field, and is especially suitable for nonlinear trajectory scenarios such as tower crane rotation and vehicle detours.

[0027] The aforementioned temporal modeling network generates a predicted motion trajectory for each dynamic occluder, containing its position in multiple future frames. This trajectory is represented as a sequence of bounding boxes, and the prediction result for each frame is accompanied by a confidence estimate.

[0028] S22 combines the predicted motion trajectory with the real-time recognition results of the target detection network, performs data association and trajectory matching through Kalman filtering or Hungarian algorithm, dynamically updates the occlusion range of dynamic occlusion objects, and generates spatiotemporal distribution features containing motion trajectory, occlusion range and time series information.

[0029] The detection results output by the object detection network in each frame are uncertain, and may include false positives, false negatives, and jitter in the detection box position. While the aforementioned generated predicted motion trajectory is based on historical sequences and has a certain degree of smoothness and foresight, it may accumulate errors over time. Therefore, it is necessary to perform data association and trajectory matching between the real-time detection results of the current frame and the predicted results of the existing trajectory, combining the advantages of both to form a stable, continuous, and accurate occlusion tracking trajectory, and dynamically update its occlusion range.

[0030] This step offers two optional implementation methods: a data association method based on Kalman filtering or a data association method based on the Hungarian algorithm. The choice can be made based on the computational resources, real-time requirements, and motion characteristics of the obstructing objects in the actual application scenario. Details are as follows: (1) Data association and trajectory matching based on Kalman filtering For each established trajectory, a Kalman filter is maintained to model and estimate the motion state of the occluded object. The state vector of the Kalman filter is defined as follows: ,in Represents the coordinates of the center point of the bounding box. This indicates the width and height of the bounding box. These represent the rates of change of the center point coordinates, width, and height, respectively. A uniform motion model is used as the state transition equation, which assumes that the occluder maintains a constant velocity change between adjacent frames.

[0031] In each frame, the Kalman filter first performs a prediction step: based on the optimal state estimate from the previous frame... and its covariance matrix Through the state transition matrix Predict the state of the current frame Covariance Predicted state Observation matrix After mapping, the predicted bounding box position of the current frame is obtained.

[0032] Subsequently, all detection boxes output by the object detection network in the current frame are associated and matched with the predicted boxes of each trajectory. Specifically, for each trajectory, the Mahalanobis distance between its predicted box and all detection boxes in the current frame is calculated:

[0033] in, For the first The observation vector of each detection box. This is the measurement residual covariance matrix of the Kalman filter. Mahalanobis distance comprehensively considers the uncertainty of prediction and can more accurately measure the correlation between the detection box and the trajectory. For each trajectory, the detection box with the smallest Mahalanobis distance that is less than a preset threshold (e.g., 3.0) is selected as the matching object.

[0034] For successfully matched trajectory-detection pairs, perform the Kalman filter update step: using the detection boxes as observations. Calculate Kalman gain The optimal state estimate is obtained by fusing predicted values ​​and observed values. The covariance matrix is ​​updated. Simultaneously, the boundary position of the detection box is recorded as the occlusion range of the occluder in the current frame and appended to the trajectory's history.

[0035] For trajectories that do not match a detection box, the trajectory is not immediately terminated. Instead, its Kalman filter prediction state is retained as an estimate for the current frame, and the predicted box is recorded as an occlusion area. Simultaneously, the consecutive lost frame counter for the trajectory is incremented. If the number of consecutive lost frames does not exceed a preset threshold (e.g., 5 frames), the trajectory continues to exist; if it exceeds the threshold, it is determined that the occlusion has left the monitoring field of view or that tracking has failed, and the trajectory is terminated.

[0036] For detection boxes that do not match a trajectory, they are considered as new candidate targets. If a detection box appears stably in multiple consecutive frames (e.g., 3 frames) and does not match an existing trajectory, a new trajectory is initialized for it, and a corresponding Kalman filter is created.

[0037] (2) Data association and trajectory matching based on the Hungarian algorithm For each established trajectory, a uniform motion model or linear interpolation method can be used to predict the bounding box position of the current frame based on the bounding box positions of historical frames. Specifically, a displacement vector can be calculated based on the actual positions of the two most recent frames, and the position of the previous frame can be superimposed on this displacement vector to obtain the predicted position of the current frame; or a simpler constant velocity assumption can be used, extrapolating the average velocity of the most recent few frames.

[0038] In each frame, construct a Cost matrix ,in This represents the number of detection boxes in the current frame. This represents the number of active trajectories. Each element in the cost matrix... Represents the detection box With trajectory The matching cost between predicted bounding boxes can be expressed as the intersection-over-union (IoU) cost, defined as follows: ,in This is the intersection-over-union (IoU) ratio between the detected bounding box and the predicted bounding box. A higher IoU value results in lower costs and is suitable for matching pairs that are close in location and have similar scales.

[0039] After constructing the cost matrix, the Hungarian algorithm is used to solve for the optimal matching. The Hungarian algorithm finds the allocation scheme that minimizes the total matching cost in polynomial time through row and column reduction and augmenting path search, and outputs a set of matching pairs that satisfy the constraints that each detection box matches at most one trajectory and each trajectory matches at most one detection box.

[0040] For a successfully matched trajectory-detection pair, the boundary position of the detection box is recorded as the occlusion range of the occluder in the current frame and appended to the history of the trajectory. The prediction model of the trajectory is then updated using the detection box (e.g., the average velocity is updated for prediction in the next frame).

[0041] For trajectories that do not match a detection box, the predicted bounding box of the trajectory is recorded as the occlusion range of the current frame, and the consecutive lost frame counter for that trajectory is incremented. If the number of consecutive lost frames does not exceed a preset threshold, the trajectory remains active and continues to extrapolate predictions based on historical positions; if the threshold is exceeded, the trajectory is terminated.

[0042] For detection boxes that do not match a trajectory, they are considered as candidate new targets. If a detection box appears stably in multiple consecutive frames and does not match an existing trajectory, a new trajectory is initialized for it.

[0043] After completing data association and trajectory matching using one of the above two methods, a complete spatiotemporal distribution feature is generated for each dynamic occlusion, including: Motion trajectory, i.e., the temporal sequence of the occluder's position information in each frame, is defined by the coordinates of the bounding box center point. The record, in the form of a document, reflects the movement path and speed changes of the obstruction in the construction site space.

[0044] Occlusion range, i.e., the set of areas covered by the bounding box of the occluding object in each frame, expressed in pixel coordinates. The record is in the form of a document used to precisely define the spatial occlusion relationship of the obstruction on the construction area. The occlusion range corresponds one-to-one with the movement trajectory, together forming a complete description of the obstruction in the spatiotemporal domain.

[0045] The time series information, including the start time, end time, duration, and timestamp of the trajectory of the occluded object, as well as metadata such as the number of consecutive lost frames and trajectory confidence, provides an accurate temporal reference for the spatial positioning and temporal alignment of the occluded area in subsequent steps.

[0046] S3. Based on the spatiotemporal distribution characteristics of the dynamic occlusion, the spatial location of the occluded area is extracted from the fused scene representation. Using the unoccluded area image and depth data as conditions, the image content of the occluded area is generated through a generative adversarial network. The image content and the 3D point cloud data are combined to reconstruct the 3D spatial structure of the occluded area to obtain the compensated complete scene model. In this step, the spatiotemporal distribution features include the occlusion range of each dynamic occluder in each frame (i.e., the area covered by the bounding box). The pixel region of this occlusion range in the image coordinate system is mapped to the three-dimensional spatial coordinate system. Combined with depth data and three-dimensional point cloud data, the spatial location and geometric range of the occluded area are determined.

[0047] Image information in the occluded area is missing due to occlusion, but the image and depth data of the unoccluded area remain intact. These visible areas can be used as conditions to reasonably infer and complete the image content of the occluded area through generative adversarial networks, as follows: In some embodiments, image content of the occluded region is generated by a generative adversarial network, conditioned on the unoccluded region image and depth data, including: S31, take the unoccluded region image and depth data as joint condition input, extract the texture feature map of the region image and the geometric structure feature map of the depth data respectively, and construct a joint condition feature space based on the texture feature map and the geometric structure feature map; Specifically, based on the spatial location of the occluded region, the unoccluded region image and its corresponding depth data are segmented from the fused scene representation. A pre-trained feature extraction network (such as the shallow part of VGG-16 or ResNet-50) is used to perform convolution operations on the unoccluded region image to extract multi-scale texture feature maps. It should be understood that this feature extraction network is pre-trained on the ImageNet dataset and has good texture and edge feature extraction capabilities; its shallow convolutional kernels can capture low-level visual information such as color, edges, and texture.

[0048] Meanwhile, a similar coding network is used to process the depth data, which is represented in grayscale. The pixel values ​​reflect the distance from the object surface to the camera. The coding network extracts geometric feature maps from these maps, which contain geometric information such as the depth gradient, plane orientation, and concavity / convexity of the object surface.

[0049] Texture feature maps and geometric feature maps are concatenated along the channel dimension to form a joint conditional feature space. This joint conditional feature space contains both the apparent texture information of the visible region and the three-dimensional geometric constraints of the visible region, providing rich conditional inputs for the generator. This ensures that the generated image content of the occluded region is consistent with the surrounding region in texture style, while conforming to the depth distribution of the real scene in geometric structure.

[0050] However, in practical engineering applications, relying solely on the texture and depth conditions of the visible area for image generation still faces the following technical problems: Structural components at construction sites have clear geometric shapes and design specifications. However, without prior structural constraints, generative adversarial networks (GANs) are prone to generating structural forms that do not conform to the design drawings. For example, they may generate frame columns with incorrect dimensions, floor slab edges as irregular curves, or confuse features of different structural types, resulting in generated images that fail to meet the geometric accuracy requirements for construction progress assessment. Furthermore, the boundaries of occluded areas are often blurred due to the irregular shapes of dynamically occluding objects, making it difficult for the generative network to accurately define the boundary positions of the occluded structures, easily leading to boundary drift or structural misalignment. To address these issues, this application further establishes a structure-prior guidance method based on the construction schedule plan. Standard geometric templates are used to apply geometric constraints to the generation process, ensuring that the generated results maintain consistency with design specifications in terms of structural form.

[0051] In some embodiments, a joint conditional feature space is constructed based on the texture feature map and the geometric structure feature map, including: S311, Identify the structural type and expected completion status of the occluded area in the construction schedule, retrieve a matching standard geometric template from the preset structural prior library, spatially register the standard geometric template with the texture feature map, and determine the precise boundary position and geometric shape of the occluded area. The construction schedule includes information on structural components to be completed in the current construction phase, such as shear walls, floor slabs, frame columns, and beams. Each component has clearly defined design dimensions, spatial coordinates, and expected completion status (e.g., poured, installed, awaiting construction). Based on the spatial coordinates of the obscured area, the corresponding structural type and expected completion status are retrieved from the construction schedule, and a standard geometric template matching that structural type is retrieved from a pre-built structural prior knowledge library. It should be understood that the structural prior knowledge library can be built based on a BIM model, storing standard outlines, standard dimensions, and typical geometric shapes of various structural components.

[0052] The retrieved standard geometric template is spatially registered with the texture feature map to determine the precise boundary position and geometric shape of the occluded region. The registration process can employ feature point matching or iterative nearest-point algorithms. Specifically: edge features of the visible region, such as visible edges and corners of structural components, are extracted from the texture feature map; the standard geometric template is projected onto the image plane and aligned with the extracted edge features; by minimizing the distance between the projected contour and the visible edges, the optimal spatial transformation parameters (translation, rotation, scaling) of the template are solved, thereby achieving spatial registration between the template and the visible region. It should be understood that after registration, the standard geometric template defines the precise boundary position and geometric shape of the occluded region, serving as a strong geometric constraint for subsequent feature fusion.

[0053] S312, the standard geometric template, geometric structure feature map and texture feature map are fused according to the precise boundary position and geometric shape; wherein, the weight ratio of geometric structure features and texture features in the fusion process is dynamically adjusted according to the criticality of the occluded area in the construction progress; In this step, the standard geometric template, geometric structure feature map, and texture feature map obtained after spatial registration are fused to form a fused joint feature representation. The standard geometric template provides prior structural information about the occluded region, including precise geometric boundaries, structural contours, and dimensional information. The geometric structure feature map provides geometric cues from the depth data, including depth gradients, plane orientation, and concavity / convexity of the object surface. The texture feature map provides appearance information for the visible region, including low-level visual features such as color, edges, and texture.

[0054] Meanwhile, the importance of different feature types to the generation task is not constant, but varies depending on the importance of the structural components corresponding to the occluded area in the construction progress. If a uniform fusion weight is applied to all structural types, it may result in insufficient geometric accuracy of critical load-bearing structures, or stiff and unnatural textures of non-critical decorative structures. Therefore, this application sets a dynamic adjustment of the weight ratio between geometric features and texture features in the fusion process based on their criticality.

[0055] Specifically, the criticality score of the obstructed area in the construction schedule is determined. The score is calculated based on a comprehensive quantitative assessment of the following factors: Structure type coefficient The determination is based on the mechanical action of the components in the building structure. For example, load-bearing components (such as frame columns, shear walls, core tubes, and main beams) are taken as... Secondary structural members (such as secondary beams and structural columns) are taken as follows: Non-load-bearing components (such as infill walls, decorative layers, doors and windows) should be taken as follows: .

[0056] Critical path coefficient The critical path is determined based on whether the component is located on the critical path. For example, if the construction progress of this component directly affects the overall project duration, then... If it is on a non-critical path, then If it is on a line with a large time difference, then .

[0057] Schedule Deviation Sensitivity Coefficient The determination is based on the degree of deviation between the current progress and the planned progress of the component. For example, if the deviation exceeds 10% of the planned duration, then... If the deviation is between 5% and 10%, then If the deviation is less than 5%, then .

[0058] Overall criticality score ,in The preset weighting coefficients are used, and they satisfy the following conditions: It can be calibrated according to the actual project conditions, for example, by taking... .

[0059] Calculated criticality score A higher value indicates that the structural components in the occluded area are more important in the construction progress assessment, requiring higher geometric accuracy when generating the image, while relatively lowering the requirement for texture naturalness; conversely, a lower value indicates a lower requirement for texture naturalness. The lower the value, the more relaxed the geometric constraints can be, and the greater emphasis can be placed on the visual continuity of the texture.

[0060] Based on keyness scoring Dynamically adjust the weight ratio of geometric structure features and texture features during the fusion process: weight of geometric structure features Texture feature weights .

[0061] when When it is large (e.g.) Geometric structure feature weights Texture feature weights The fusion result is primarily based on geometric structure, ensuring that the generated structure's dimensions and shape strictly conform to the standard geometric template; when When smaller (e.g.) Geometric structure feature weights Texture feature weights The fusion result is dominated by texture appearance, which makes the generated content visually transition naturally with the surrounding area.

[0062] The fusion process is implemented as follows: Assume that a template feature map is generated after spatial registration of a standard geometric template. ,in The template feature channel number is typically described using a binary boundary mask or multi-level contour; the geometric structure feature map is... Texture feature map is First, convolutional layers are used to map each feature map to the same number of channels. ,get , , Subsequently, a weighted fusion process is performed according to the following formula to generate a fused feature map. :

[0063] in These are the internal weights of the geometric structure features, used to adjust the ratio between the standard geometric template and the deep geometric structure features. For example, taking... This allows the template morphology prior to dominate, while also taking into account the actual geometric measurements in the depth data. The fused feature map It carries precise geometric boundary information, depth geometric cues, and texture information of the visible area, and the weight ratio is dynamically and adaptively adjusted according to the criticality of the occluded area.

[0064] Through the aforementioned dynamic weight fusion mechanism, differentiated processing for different structural types can be achieved: for critical load-bearing components, the generated results strictly follow the geometric dimensions and shapes of the design drawings to ensure the structural accuracy of the progress assessment; for non-critical decorative components, the generated results pay more attention to visual continuity, making the completed scene model visually natural and coordinated, and avoiding the problem of unnatural textures caused by excessive pursuit of geometric precision.

[0065] S313, using the structural form defined by the standard geometric template as the geometric constraint benchmark, the texture feature map is spatially aligned and feature enhancement guided by the structural form is performed through the cross attention mechanism to generate a joint conditional feature space containing precise geometric boundaries, continuous texture information and progress semantic labels.

[0066] After completing the weighted fusion of multi-source features, the fused feature map Although the image already contains geometric and texture information, spatial misalignment may still exist between the precise boundaries defined by the standard geometric template and the texture feature map. This is mainly due to: pixel-level deviations in the bounding boxes of occluded objects output by the object detection network; residuals in the spatial registration process of the standard geometric template; and resolution differences between depth data and image data. Directly using unaligned fused features for image generation can lead to defects such as blurred boundaries, structural shifts, and mismatches between texture and geometric contours in the generated occluded areas.

[0067] To address the aforementioned issues, this application further employs a cross-attention mechanism, using the structural morphology defined by the standard geometric template as the geometric constraint benchmark, to perform spatial boundary alignment and feature enhancement guided by the structural morphology on the texture feature map. Specifically, the precise boundary information provided by the standard geometric template is used as a query condition to retrieve texture features corresponding to their spatial locations in the texture feature map. Through attention weight redistribution, texture features are enhanced near the geometric boundaries, and the spatial distribution of texture features is aligned with the standard geometric morphology.

[0068] The specific implementation steps are as follows: (1) Project the standard geometric template onto the image plane to generate a binary boundary mask. In this model, boundary pixels are set to 1, and non-boundary pixels are set to 0. Position encoding is performed on the boundary mask to generate a position encoding matrix. Position encoding employs sine / cosine coding or learnable position embedding to ensure that each pixel location receives a unique position representation. The position encoding matrix... Query terms as cross-attention .

[0069] Texture feature map As the source of the value terms, the texture feature map is linearly projected and used as the key terms. Specifically, two different convolutional layers are used to... Perform the transformation to obtain the key terms. Sum of values ,in This represents the embedding dimension of the attention mechanism.

[0070] (2) Query items Key items Value Item The data is flattened into a sequence, where each spatial location corresponds to a feature vector. For each location in the query term... (For a boundary pixel on the boundary mask), calculate its relationship with all positions. The similarity of the key vectors can be achieved using scaled dot product attention:

[0071] in Indicates position Position The attention weights reflect the boundary locations. With each position in the texture feature map The correlation between them is determined. Through this attention mechanism, each boundary pixel retrieves the texture feature most relevant to its spatial location from the texture feature map. For regions located inside the boundary, geometry-based guided attention is also computed: the internal region of the standard geometric template is meshed to generate an internal sampling point position code, which is also used as a query term to retrieve the texture features of the internal region from the texture feature map, ensuring that the spatial distribution of the internal texture is consistent with the topological structure of the standard geometric shape.

[0072] (3) Based on attention weights , value term Weighted aggregation is performed to obtain the enhanced feature map. :

[0073] For boundary pixels, this aggregation process enhances the features. By fusing texture information most relevant to geometric boundaries from the texture feature map, spatial alignment of texture features with geometric boundaries is achieved. For internal pixels, this aggregation process guides the aggregation of texture features according to the morphological structure defined by the standard geometric template, ensuring that the distribution of internal textures is consistent with the geometric shape (such as rectangles, circles, L-shapes, etc.). For example, for a rectangular frame column, the aggregation direction of internal textures is guided along the column axis, ensuring that the vertical direction of the texture is perpendicular to the edge of the column.

[0074] (4) Enhance the texture feature map Geometric structure feature map Standard geometric template features The features are concatenated to generate a joint conditional feature space. The joint conditional feature space has the following characteristics: Precise geometric boundaries: Through the cross-attention mechanism, the spatial distribution of the texture feature map is aligned with the precise boundaries defined by the standard geometric template at the pixel level. The texture response near the boundary is enhanced, and the boundary position is clear and without offset.

[0075] Continuous texture information: Enhanced texture features Based on boundary alignment, the continuity and richness of the original texture feature map are preserved. Through the attention aggregation mechanism, the texture inside the boundary achieves a smooth transition without breakage or repetition artifacts.

[0076] Progress semantic labels: The joint conditional feature space additionally concatenates a progress semantic encoding vector in the channel dimension. This encoding vector contains the structural type (e.g., "frame column", "shear wall", "floor slab"), construction status (e.g., "poured", "installed", "awaiting construction"), and criticality score of the structural components corresponding to the occluded area. Information such as these are injected into the feature space in the form of learnable embeddings, enabling the generator to perceive the semantics of the construction progress in the region when generating image content, thereby generating a reasonable structural form that corresponds to the construction stage.

[0077] S32, the joint conditional feature space is input into the generator of the generative adversarial network. The generator adopts an encoder-decoder structure. The encoder extracts high-level semantic features step by step through multiple convolutions. The decoder restores the spatial resolution step by step through deconvolution. Skip connections are used to pass the texture features and geometric structure features of different scales in the encoder to the corresponding levels of the decoder to achieve the fusion and guidance of multi-scale features. In this step, the joint conditional feature space will be... The generator of the input generative adversarial network employs an encoder-decoder structure to learn the mapping relationship from the visible region to the occluded region from the joint conditional feature space, generating image content of the occluded region that is cooperative with the surrounding region in terms of texture continuity and geometric consistency.

[0078] The generator adopts a U-Net architecture, consisting of an encoder, a decoder, and skip connections. The encoder expands the receptive field through progressive downsampling, extracting high-level semantic features to enable the generated image to understand the global structure and contextual relationships of the scene. The decoder restores spatial resolution through progressive upsampling, giving the generated image fine texture details. Skip connections directly transmit high-resolution features from the encoder's shallow layers to the corresponding levels of the decoder, compensating for spatial detail lost during downsampling and ensuring the quality of the generated image in terms of edge sharpness and texture clarity.

[0079] Encoder Structure: The encoder consists of multiple stacked downsampling blocks. Each downsampling block contains two convolutional layers, a batch normalization layer, and a Leaky ReLU activation function. The convolutional kernel size is 3×3 with a stride of 1. A max-pooling layer or convolutional layer with a stride of 2 is added at the end of each downsampling block for spatial downsampling. The encoder contains a total of 5 downsampling levels, and the feature map size is determined by the input image... Gradually shrink to The number of channels increases progressively from 64 to 512. Specifically, the first-level output has 64 channels, and the feature map size is... The second-level output channel has 128 channels, and the feature map size is [size missing]. The third-level output channel has 256 channels, and the feature map size is [size missing]. The fourth level has 512 output channels and a feature map size of [size missing]. The fifth level has 512 output channels and a feature map size of [size missing]. In the output feature maps of each level of the encoder, shallow feature maps have high spatial resolution and contain rich texture details, edge information and geometric boundary responses; deep feature maps have a large receptive field and contain high-level semantic features, global structural information and contextual relationships.

[0080] Decoder Structure: The decoder consists of multiple stacked upsampling blocks. Each upsampling block contains a deconvolution layer, a batch normalization layer, and a ReLU activation function. The deconvolution layer (also called a transposed convolution layer) has a 4×4 kernel size and a stride of 2, used to progressively restore the spatial resolution of the feature map. The decoder also contains 5 upsampling layers, symmetrically distributed with the encoder layers. The feature map size is determined by... Gradually recover to The number of channels gradually decreases from 512 to the number of output channels. The last stage of the decoder outputs 3 channels, corresponding to an RGB three-channel image, which is then normalized to the range [-1, 1] after passing through the Tanh activation function.

[0081] Skip connections: Skip connections concatenate the feature map output from each downsampled level in the encoder with the feature map of the corresponding level in the decoder along the channel dimension. Specifically, the output of the first-level encoder is concatenated with the input of the fifth-level decoder (the feature map after 5 upsampling steps) via skip connections; the output of the second-level encoder is concatenated with the input of the fourth-level decoder; and so on.

[0082] It should be understood that the core role of skip connections is that the encoder's shallow feature map retains high-frequency information such as fine textures, edge directions, and geometric boundaries around the occluded area, which is gradually lost during downsampling. Through skip connections, the decoder can simultaneously utilize shallow texture details and deep semantic information when restoring spatial resolution, achieving multi-scale feature fusion and guidance. For example, when generating floor slab edges, the edge responses in the shallow feature map can guide the decoder to generate sharp boundaries; when generating concrete textures, the texture patterns in the shallow feature map can guide the decoder to generate a surface texture consistent with the surrounding area.

[0083] Through the encoder-decoder and skip connection structure described above, the generator can learn the complex mapping relationship from the visible region to the occluded region from the joint conditional feature space, so that the generated result maintains both global structural consistency and local texture realism.

[0084] S33, in the output stage of the decoder, firstly, the generated image content is edge-enhanced and structurally constrained through a spatial attention mechanism, and then the texture of the generated region is filled and smoothed based on the statistical distribution of the surrounding texture, so as to generate occluded region image content that is coordinated with the surrounding region in terms of texture continuity and geometric structure consistency.

[0085] While the initial generated image output by the decoder already possesses a reasonable structure and texture, it still has shortcomings in the following aspects: First, the structural edges in the generated image may become blurred or jagged due to the deconvolution operation, resulting in insufficiently sharp boundaries; second, there may be subtle texture differences or seam marks between the generated region and the surrounding visible region, affecting visual continuity. To address these issues, this step further utilizes a spatial attention mechanism for edge enhancement and structural constraints, as well as texture filling and smoothing based on the statistical distribution of surrounding textures, during the decoder output stage.

[0086] Spatial attention mechanisms are used to enhance the sharpness of structural edges in the generated image, aligning them precisely with the boundaries defined by a standard geometric template. The specific implementation steps are as follows: Boundary information of the standard geometric template is extracted from the joint conditional feature space to generate a binary boundary mask. The boundary pixels are set to 1, and the non-boundary pixels are set to 0. It should be understood that this boundary mask reflects the precise geometric contour of the structural components corresponding to the occluded area.

[0087] The initial generated image of the decoder output Perform edge detection and extract edge response maps from the generated image. Edge detection can employ the Sobel operator, Canny edge detection, or differentiable edge detection networks, outputting the probability that each pixel location belongs to an edge. A boundary mask is then applied. With edge response map Perform pixel-by-pixel comparisons and calculate edge difference loss:

[0088] in Boundary mask The set of pixel locations with a value of 1. This loss function constrains the edge response of the generated image to approach 1 at the boundary locations defined by the standard geometric template, thereby guiding the generated image to form clear and sharp edges at these locations.

[0089] Then, the spatial attention weight map It is used to generate images. The spatial attention weight map assigns weights based on the importance of pixel positions: boundary pixels are given higher weights, and non-boundary pixels are given lower weights. Specifically, the weight map can be generated using a boundary mask. Gaussian blur expansion is applied to give pixels near the boundary a certain weight, forming a transition zone. The spatial attention weight map is multiplied pixel-by-pixel with the generated image, and then jointly optimized with the boundary constraints of the standard geometric template. Gradient descent is used to fine-tune the pixel values ​​in the boundary region, causing the edges of the generated image to gradually converge towards the precise boundary defined by the standard geometric template. The edge enhancement process can be iterated several times until edge difference loss occurs. Converging to below the preset threshold.

[0090] After edge enhancement, the texture of the generated region is filled and smoothed based on the statistical distribution of the surrounding texture to eliminate texture differences between the generated region and the surrounding visible region, ensuring natural texture transition and visual continuity. The specific implementation steps are as follows: Extract texture statistical features from the visible area surrounding the occluded region. Specifically, starting from the boundary of the occluded region, construct a ring-shaped texture sampling region by extending outwards by a certain width (e.g., 50 pixels). Within this sampling region, extract the following texture features: Gray-level co-occurrence matrix (GLCM): Calculates four texture parameters: contrast, correlation, energy, and homogeneity, describing the roughness, directionality, and uniformity of the texture.

[0091] Local Binary Pattern (LBP): Extracts local texture pattern histograms to describe the microstructure and local contrast of the texture.

[0092] Histogram of Oriented Gradients (HOG): Extracts the gradient direction distribution, describing the edge direction and structure orientation of the texture.

[0093] It should be understood that the above texture features constitute a texture statistical distribution model of the visible area, reflecting the typical appearance characteristics of the structural component in a real construction scenario.

[0094] Next, a sample-based texture synthesis algorithm is used to fill the generated region with texture. This algorithm divides the generated region into multiple overlapping texture blocks, each with a size that can be set to 16×16 or 32×32 pixels. For each texture block to be filled, the best-matching texture block is searched in the texture sampling space of the visible region, using the Euclidean distance or cosine similarity between texture feature vectors as the matching criterion. The matched texture block is copied and pasted to the corresponding position in the generated region, and adjacent texture blocks are seamlessly stitched together using linear interpolation or Poisson fusion to eliminate inter-block boundaries.

[0095] Finally, a Poisson image editing method is employed to perform gradient domain fusion at the boundary between the generated and visible regions. Specifically, while preserving the internal texture structure of the generated region, the pixel values ​​at the boundary of the generated region are smoothly transitioned to those at the boundary of the visible region by solving the Poisson equation. Specifically, using the pixel values ​​at the boundary of the visible region as Dirichlet boundary conditions and guided by the gradient field of the generated region, the pixel values ​​within the generated region are reconstructed, ensuring continuous gradient changes at the boundary and eliminating seam artifacts. After Poisson fusion, the generated region and the surrounding visible region achieve a natural transition in texture, making it visually difficult to distinguish between the generated region and the original visible region.

[0096] After generating the image content of the occluded area, the image content is fused with the original 3D point cloud data to reconstruct the 3D spatial structure of the occluded area, so as to restore the missing 3D structure of the occluded area and form a compensated scene model with complete geometric shape and texture information.

[0097] The specific implementation is as follows: The generated image content of the occluded region is used as the texture prior input. Since the generated image content contains precise geometric boundaries and structural morphology information under the constraints of a standard geometric template, a depth map can be extracted from the generated image through a monocular depth estimation network, and then back-projected to generate a sparse 3D point cloud as geometric cues. Simultaneously, the original 3D point cloud surrounding the occluded region, i.e., the 3D point cloud of the unoccluded region, is extracted as the geometric context. These two types of point clouds are then concatenated to form the input point cloud set for the completion network.

[0098] The point cloud completion network adopts an improved architecture of PCN or Point Completion Network, consisting of an encoder and a decoder. The encoder uses PointNet++ or DGCNN to perform multi-level feature extraction on the input point cloud, learning the local geometric features and global shape features of each point. The decoder uses a folded network or a Transformer-based point cloud generator to generate a complete point cloud step by step, conditioned on the global features output by the encoder, with the number of points matching the expected number of points that the occluded region should contain.

[0099] During the training phase, the network uses complete point clouds derived from CAD or BIM models of complete structural components as supervision signals, and employs chamfer distance or bulldozer distance as loss functions to optimize network parameters so that the generated complete point clouds are geometrically consistent with design specifications.

[0100] Subsequently, the point cloud of the occluded region generated by the completed network is fused with the original scene point cloud to form a complete geometric model. During the fusion process, redundant points in the overlapping areas need to be removed and smoothed. Voxel mesh filtering or Poisson surface reconstruction can be used to ensure the continuity and smoothness of the fusion boundary.

[0101] Finally, the generated image content of the occluded area is used as a texture map. Through UV unwrapping and texture mapping techniques, the texture is mapped onto the completed point cloud surface. Based on the spatial location of the point cloud and the geometric relationship with camera imaging, the corresponding pixel coordinates of each 3D point in the image are calculated, and the RGB values ​​are extracted as the texture attributes of that point, forming a 3D model with complete texture information. Furthermore, for regions in the completed point cloud that cannot be directly mapped to the generated image, nearest-neighbor interpolation or texture synthesis methods can be used for filling.

[0102] After the above processing, the finally compensated complete scene model is obtained, which has both complete geometric structure (restored through point cloud completion) and continuous texture information (through generated image mapping).

[0103] S4. Input the compensated complete scenario model and construction schedule plan into the schedule inference network, and output the construction schedule evaluation result.

[0104] In this step, this application constructs a progress inference network, which employs multimodal feature fusion and temporal modeling techniques to achieve end-to-end mapping from 3D scene models to progress indicators.

[0105] The progress inference network employs an architecture combining graph neural networks and attention mechanisms, such as... Figure 2 As shown, it includes a feature extraction module, a multimodal fusion module, a time series analysis module, and a progress regression module, as detailed below: The feature extraction module extracts actual construction state features from the compensated complete scene model. The complete scene model is represented as a 3D point cloud or a textured mesh, containing both geometric and textural information. The feature extraction module encodes the following two types of information respectively: For geometric information, point cloud processing networks such as PointNet++ or DGCNN are used to extract features point-by-point from the scene point cloud. These point cloud processing networks extract local geometric features (such as curvature, normal vector, and local density) and global structural features (such as the overall shape of components and their spatial distribution) for each point through multi-level abstraction and feature aggregation, outputting a geometric feature map. ,in This refers to the number of point clouds or the number of sampling points. For geometric feature dimensions.

[0106] For texture information, a convolutional neural network (such as ResNet-50) is used to extract features from the texture mapping image of the scene to obtain a texture feature map. ,in For feature map size, This refers to the texture feature dimension. Texture features reflect the material properties, construction quality, and finishing status of the structural surface.

[0107] Simultaneously, the feature extraction module also extracts spatial location information from the complete scene model, mapping the three-dimensional coordinates of each structural component to spatial partitions in the construction schedule, and generating location-coded features. This is used to spatially align geometric elements in the scene with their corresponding components in the construction schedule.

[0108] The multimodal fusion module integrates geometric features, texture features, and location-encoded features to form a unified component-level feature representation. Considering the presence of various structural components on a construction site, and the different contributions of each component to progress assessment, this module employs a graph neural network to model the spatial relationships and dependencies between components, as detailed below: Construct a scene graph based on the spatial relationships in the complete scene model. , where nodes Representing each structural component, edge This indicates the spatial adjacency relationship between components. Each node... The initial feature is the concatenation of the geometric features, texture features, and location-encoded features of its corresponding region: .

[0109] A graph convolutional network is used for message passing and feature updating. In the... In layer graph convolution, nodes The feature update formula is:

[0110] in For nodes The set of adjacent nodes, and The weight matrix is ​​a learnable matrix. This is a non-linear activation function. Through multi-layer graph convolution, the features of each node are fused with its own state and the state information of adjacent components, enabling the network to understand the construction logic relationships between components.

[0111] Time Series Analysis Module: Construction site progress assessment relies not only on the current scenario state but also on historical construction progress data to determine progress trends and identify delay risks. Therefore, the time series analysis module introduces Long Short-Term Memory (LSTM) networks or Transformers to model the temporal evolution of construction progress.

[0112] The complete scene model of multiple time series frames is sequentially input into the feature extraction module and the multimodal fusion module to obtain the component-level feature sequence at each time step. ,in For a moment The feature matrix of all components. This sequence is input into the temporal modeling network, which outputs the temporal state features of each component. .

[0113] When using LSTM, the feature sequence of each component is input into the LSTM unit sequentially according to time steps. LSTM captures the temporal dependencies of the component's progress state through forget gates, input gates, and output gates, and outputs the temporal features at the current moment. When using Transformer, the component feature sequence is added to the position encoding and then input into the Transformer encoder. The dependencies between time steps are calculated in parallel through a self-attention mechanism.

[0114] The schedule regression module compares the time-series status characteristics output by the time-series analysis module with the expected status in the construction schedule plan, and outputs the final construction schedule evaluation result.

[0115] The construction schedule is input in a structured format, containing information such as the planned start time, planned completion time, resource requirements, and construction sequence for each structural component. The expected states in the construction schedule are encoded into feature vectors. The expected state characteristics of each component include its planned completion percentage, current construction stage, and construction logic constraints.

[0116] The progress regression module employs an attention mechanism to fuse and compare actual state features with expected state features. Specifically, it integrates and compares time-series state features. As a query term, the expected state characteristics As key and value terms, calculate cross-attention:

[0117]

[0118] Through cross-attention, the network automatically learns the degree of matching between the actual state and the expected state of each component, and outputs comparative features. .

[0119] Then, the comparative features will be... The input is processed by a multilayer perceptron for schedule regression, and the output is a schedule evaluation index for each component, including: Percentage of completion: The ratio of the actual amount of work completed for this component to the total planned amount of work; Schedule deviation: The difference between the actual progress completed and the planned progress completed; Estimated completion time: The estimated completion date based on the current construction speed and remaining workload.

[0120] The progress assessment indicators for each component are weighted and aggregated (the weights are determined based on the component's criticality coefficient) to obtain the overall construction progress assessment results, including: Overall progress completion percentage: a weighted average of the completion percentages of all components; Critical path schedule deviations: Summary of component schedule deviations on the critical path; Schedule Delay Warning: When the schedule deviation of any component exceeds a preset threshold, a warning message is output, identifying the delayed component and its location; Construction status visualization: The progress status of each component is rendered in the form of color marks on the complete scene model.

[0121] The construction progress assessment results output by the progress inference network are displayed to construction management personnel through a visualization terminal. The visualization terminal provides functions such as 3D scene overlay display, progress report generation, and early warning push, assisting management personnel in decision-making and scheduling.

[0122] Reference Figure 3 As shown in the illustration, this application also provides an image-based construction project progress analysis device 200, the device comprising: The multimodal data acquisition module 10 is used to acquire multimodal data from the construction site, including image data, depth data and 3D point cloud data, and to perform spatiotemporal alignment and data-level fusion on the multimodal data to construct a real-time fused scene representation of the construction site. The data fusion module 20 uses a target detection network to identify dynamic occlusions in the fused scene representation in real time, and uses a temporal modeling method to continuously track the motion trajectory and occlusion range of the dynamic occlusions to generate the spatiotemporal distribution characteristics of the dynamic occlusions. The dynamic occlusion detection and tracking module 30 is used to extract the spatial location of the occluded area from the fused scene representation based on the spatiotemporal distribution characteristics of the dynamic occlusion. Using the unoccluded area image and depth data as conditions, it generates the image content of the occluded area through a generative adversarial network. It then combines the image content and the 3D point cloud data to reconstruct the 3D spatial structure of the occluded area to obtain a compensated complete scene model. The occlusion compensation and reconstruction module 40 is used to input the compensated complete scene model and the construction schedule plan into the schedule inference network and output the construction schedule evaluation result.

[0123] In some embodiments, refer to Figure 4 As shown, the data fusion module 20 includes: The spatiotemporal alignment unit 201 is used to synchronize the timestamps and register the spatial coordinates of the acquired image data, depth data and 3D point cloud data, so as to realize the spatiotemporal alignment of multimodal data. The data-level fusion unit 202 is used to integrate spatiotemporally aligned multimodal data into a unified fusion scene representation by employing pixel-level fusion or feature-level fusion methods, and to eliminate sensor measurement errors and environmental interference through Kalman filtering to generate a real-time updated fusion scene representation.

[0124] In some embodiments, refer to Figure 4 As shown, the dynamic occlusion detection and tracking module 30 includes: The detection sequence construction unit 301 is used to extract the bounding box position and category information of dynamic occluders in each frame of the image and construct the detection sequence of dynamic occluders. The temporal prediction unit 302 is used to input the detection sequence into the temporal modeling network to perform temporal modeling and future position prediction on the dynamic occlusion to obtain the predicted motion trajectory; wherein, the temporal modeling network adopts an LSTM or Transformer architecture. The association matching unit 303 is used to combine the predicted motion trajectory with the real-time recognition results of the target detection network, perform data association and trajectory matching through Kalman filtering or Hungarian algorithm, dynamically update the occlusion range of dynamic occlusion objects, and generate spatiotemporal distribution features containing motion trajectory, occlusion range and time series information.

[0125] In some embodiments, refer to Figure 4 As shown, the occlusion compensation and reconstruction module 40 includes: The conditional feature construction subunit 401 is used to take the unoccluded region image and depth data as joint conditional inputs, extract the texture feature map of the region image and the geometric structure feature map of the depth data respectively, and construct a joint conditional feature space based on the texture feature map and the geometric structure feature map; The generator subunit 402 is used to input the joint conditional feature space into the generator of the generative adversarial network. The generator adopts an encoder-decoder structure. The encoder extracts high-level semantic features step by step through multiple convolutions, and the decoder restores the spatial resolution step by step through deconvolution. It also uses skip connections to pass the texture features and geometric structure features of different scales in the encoder to the corresponding levels of the decoder, thereby realizing the fusion and guidance of multi-scale features. The post-processing subunit 403 is used to first perform edge enhancement and structural constraints on the generated image content through a spatial attention mechanism during the output stage of the decoder, and then fill and smooth the texture of the generated region based on the statistical distribution of the surrounding texture, so as to generate occluded region image content that maintains coordination with the surrounding region in terms of texture continuity and geometric structure consistency.

[0126] In some embodiments, the condition feature construction subunit 401 is specifically used for: Identify the structural type and expected completion status of the occluded area in the construction schedule, retrieve matching standard geometric templates from the preset structural prior library, spatially register the standard geometric templates with the texture feature map, and determine the precise boundary position and geometric shape of the occluded area. The standard geometric template, geometric structure feature map, and texture feature map are fused based on the precise boundary position and geometric shape; wherein, the weight ratio of geometric structure features and texture features in the fusion process is dynamically adjusted according to the criticality of the occluded area in the construction progress; Using the structural form defined by the standard geometric template as the geometric constraint benchmark, the texture feature map is spatially aligned and feature enhancement guided by the structural form is performed through the cross attention mechanism to generate a joint conditional feature space containing precise geometric boundaries, continuous texture information and progress semantic labels.

[0127] This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the method as described in any of the preceding claims.

[0128] This application also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the preceding claims.

[0129] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any of the preceding claims.

[0130] The above description represents the preferred embodiments of the present invention. It should be noted that, for those skilled in the art, various improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for analyzing construction project progress based on images, characterized in that, Includes the following steps: S1. Collect multimodal data from the construction site, including image data, depth data, and 3D point cloud data. Perform spatiotemporal alignment and data-level fusion on the multimodal data to construct a real-time fused scene representation of the construction site. S2, based on the target detection network, performs real-time identification of dynamic occlusions in the fused scene representation, and uses the temporal modeling method to continuously track the motion trajectory and occlusion range of dynamic occlusions, generating the spatiotemporal distribution characteristics of dynamic occlusions; S3. Based on the spatiotemporal distribution characteristics of the dynamic occlusion, the spatial location of the occluded area is extracted from the fused scene representation. Using the unoccluded area image and depth data as conditions, the image content of the occluded area is generated through a generative adversarial network. The image content and the 3D point cloud data are combined to reconstruct the 3D spatial structure of the occluded area to obtain the compensated complete scene model. S4. Input the compensated complete scenario model and construction schedule plan into the schedule inference network, and output the construction schedule evaluation result.

2. The image-based construction project progress analysis method according to claim 1, characterized in that, The multimodal data is spatiotemporally aligned and fused at the data level to construct a real-time fused scene representation of the construction site, including: S11 performs time stamp synchronization and spatial coordinate system registration on the acquired image data, depth data and 3D point cloud data to achieve spatiotemporal alignment of multimodal data; S12 employs pixel-level or feature-level fusion methods to integrate spatiotemporally aligned multimodal data into a unified fusion scene representation, and uses Kalman filtering to eliminate sensor measurement errors and environmental interference in order to generate a real-time updated fusion scene representation.

3. The image-based construction project progress analysis method according to claim 1, characterized in that, Temporal modeling methods are used to continuously track the motion trajectory and occlusion range of dynamic occlusion objects, generating spatiotemporal distribution characteristics of these objects, including: S21, extract the bounding box position and category information of dynamic occluders in each frame of the image, construct the detection sequence of dynamic occluders, and input the detection sequence into the temporal modeling network to perform temporal modeling and future position prediction of dynamic occluders to obtain the predicted motion trajectory; wherein, the temporal modeling network adopts LSTM or Transformer architecture. S22 combines the predicted motion trajectory with the real-time recognition results of the target detection network, performs data association and trajectory matching through Kalman filtering or Hungarian algorithm, dynamically updates the occlusion range of dynamic occlusion objects, and generates spatiotemporal distribution features containing motion trajectory, occlusion range and time series information.

4. The image-based construction project progress analysis method according to claim 1, characterized in that, Using unoccluded region images and depth data as conditions, a generative adversarial network is used to generate image content for occluded regions, including: S31, take the unoccluded region image and depth data as joint condition input, extract the texture feature map of the region image and the geometric structure feature map of the depth data respectively, and construct a joint condition feature space based on the texture feature map and the geometric structure feature map; S32, the joint conditional feature space is input into the generator of the generative adversarial network. The generator adopts an encoder-decoder structure. The encoder extracts high-level semantic features step by step through multiple convolutions. The decoder restores the spatial resolution step by step through deconvolution. Skip connections are used to pass the texture features and geometric structure features of different scales in the encoder to the corresponding levels of the decoder to achieve the fusion and guidance of multi-scale features. S33, in the output stage of the decoder, firstly, the generated image content is edge-enhanced and structurally constrained through a spatial attention mechanism, and then the texture of the generated region is filled and smoothed based on the statistical distribution of the surrounding texture, so as to generate occluded region image content that is coordinated with the surrounding region in terms of texture continuity and geometric structure consistency.

5. The image-based construction project progress analysis method according to claim 4, characterized in that, Constructing a joint conditional feature space based on the texture feature map and the geometric structure feature map includes: S311, Identify the structural type and expected completion status of the occluded area in the construction schedule, retrieve a matching standard geometric template from the preset structural prior library, spatially register the standard geometric template with the texture feature map, and determine the precise boundary position and geometric shape of the occluded area. S312, the standard geometric template, geometric structure feature map and texture feature map are fused according to the precise boundary position and geometric shape; wherein, the weight ratio of geometric structure features and texture features in the fusion process is dynamically adjusted according to the criticality of the occluded area in the construction progress; S313, using the structural form defined by the standard geometric template as the geometric constraint benchmark, the texture feature map is spatially aligned and feature enhancement guided by the structural form is performed through the cross attention mechanism to generate a joint conditional feature space containing precise geometric boundaries, continuous texture information and progress semantic labels.

6. A construction project progress analysis device based on image processing, characterized in that, The device includes: The multimodal data acquisition module is used to collect multimodal data from the construction site, including image data, depth data, and 3D point cloud data. It performs spatiotemporal alignment and data-level fusion on the multimodal data to construct a real-time fused scene representation of the construction site. The data fusion module uses a target detection network to identify dynamic occlusions in the fused scene representation in real time, and uses a temporal modeling method to continuously track the motion trajectory and occlusion range of dynamic occlusions, generating the spatiotemporal distribution characteristics of dynamic occlusions. The dynamic occlusion detection and tracking module is used to extract the spatial location of the occluded area from the fused scene representation based on the spatiotemporal distribution characteristics of the dynamic occlusion. Using the unoccluded area image and depth data as conditions, it generates the image content of the occluded area through a generative adversarial network. Combining the image content and 3D point cloud data, it reconstructs the 3D spatial structure of the occluded area to obtain a compensated complete scene model. The occlusion compensation and reconstruction module is used to input the compensated complete scene model and the construction schedule plan into the schedule inference network and output the construction schedule evaluation results.

7. The image-based construction project progress analysis device according to claim 6, characterized in that, The data fusion module includes: The spatiotemporal alignment unit is used to synchronize the timestamps and register the spatial coordinates of the acquired image data, depth data and 3D point cloud data, so as to realize the spatiotemporal alignment of multimodal data. The data-level fusion unit is used to integrate spatiotemporally aligned multimodal data into a unified fusion scene representation using pixel-level fusion or feature-level fusion methods, and to eliminate sensor measurement errors and environmental interference through Kalman filtering to generate a real-time updated fusion scene representation.

8. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the method as described in any one of claims 1-5.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.