Method for automatically detecting construction progress of engineering main body structure based on unmanned aerial vehicle vision
By using an improved YOLOv12 model, combined with the A2C2f_Mcva module, the C3k2-SLBlock module, and the GDSA feature fusion module, the problems of automated identification and accuracy in construction progress monitoring are solved, realizing automated and accurate detection of construction progress, which is suitable for complex construction scenarios.
Patent Information
- Application Number
- CN202510928481.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-11
AI Technical Summary
Existing image processing-based construction progress monitoring methods struggle to accurately identify the construction progress of the main structure, especially in high-altitude operations and complex construction scenarios. They cannot automatically determine construction activities and progress percentages, and traditional methods suffer from large errors and long processing times.
An improved YOLOv12 model is adopted. By replacing modules in the backbone and neck networks and combining the A2C2f_Mcva module, C3k2-SLBlock module and GDSA feature fusion module, the visual perception capability and feature fusion are enhanced. Combined with construction sequence coding and CAD software to calculate the progress correction coefficient, the automated detection of construction progress is realized.
It improves the accuracy and automation of construction progress monitoring, enabling accurate identification of construction activities and calculation of specific progress in complex construction scenarios, reducing manual intervention, and adapting to different construction projects.
Smart Images

Figure CN120932088A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing technology and architectural engineering, specifically to an automatic detection method for the construction progress of the main structure of an engineering project based on UAV vision. Background Technology
[0002] Construction projects are characterized by long construction periods, high degree of parallelism in tasks, and complex resource scheduling. Real-time and accurate monitoring of construction progress is a crucial aspect of project management. Currently, construction progress monitoring mainly relies on project management personnel summarizing data through on-site inspections, photographic records, and subjective judgment. This method is inefficient and prone to errors, especially during the main structure phase, where high-altitude work and repetitive components make progress statistics even more difficult, easily leading to a disconnect between plans and actual progress.
[0003] To achieve automated construction progress monitoring, image analysis methods based on computer vision have developed rapidly in recent years. Point cloud-based and image-based methods are two commonly used techniques in construction progress monitoring. Point cloud-based methods typically acquire point cloud data using photogrammetry or laser scanning. Through point cloud denoising and registration, the point cloud model is aligned with the building model, enabling component location identification and timestamp marking to complete construction progress monitoring. Although high-precision point cloud models can accurately represent construction progress, the matching process between the point cloud and the building model is difficult to automate fully, requiring significant manual intervention and is time-consuming. Furthermore, the construction of the main building structure often involves high-altitude operations, and construction materials need to be transported to the site using tower cranes. Construction workers and materials can easily obscure some components, preventing laser scanners from capturing complete data. This makes obtaining complete, continuous, and high-quality point cloud models for main structure progress monitoring extremely difficult.
[0004] In contrast, image-based methods are increasingly being used in construction progress monitoring. These methods acquire construction images using cameras, drones, or monitoring equipment, and then apply computer vision technology to process and analyze the images, directly or indirectly extracting construction progress information. For example, feature extraction and object classification algorithms can be used to identify specific building materials in images, thereby inferring construction activities. In terms of data acquisition, compared to the collection and modeling of point cloud data, acquiring construction images is faster and more economical. Within acceptable progress error ranges, image processing-based construction progress monitoring becomes a more advantageous choice.
[0005] However, existing image processing-based construction progress monitoring methods still have shortcomings. Traditional methods indirectly determine construction activities by identifying building materials in images, but this approach is not suitable for identifying the progress of the main structure construction because the same building material may be used in the construction of multiple construction elements, making it difficult to accurately determine the construction progress. Secondly, monitoring methods using recognition algorithms often report progress in binary form, failing to obtain the percentage of completion for each construction element, which is not conducive to judging the amount of construction materials used and subsequent resource allocation. Although segmentation algorithms can calculate the percentage of construction progress by calculating the ratio of the mask area of the construction element to the image area, in actual construction, the shape of the construction surface is irregular, and the captured images often contain a large number of non-construction areas. Without image segmentation, it is impossible to directly use image segmentation methods to accurately identify the construction progress of the main structure. Moreover, the main structure includes different construction activities such as formwork erection, rebar tying, and concrete pouring in columns, beams, and slabs. Furthermore, there is a sequence between different construction activities; for example, concrete pouring should follow rebar tying. Segmentation algorithms only output all construction procedures, therefore, further algorithm design is needed to automatically determine which construction activity is in progress and calculate the specific progress value. Summary of the Invention
[0006] The purpose of this invention is to provide an automatic detection method for the construction progress of the main structure of an engineering project based on UAV vision in order to solve the above problems. This method achieves automated and accurate detection of construction progress and is applicable to complex construction scenarios.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] An automatic detection method for the construction progress of the main structure of an engineering project based on UAV vision includes the following steps:
[0009] The first step is to use drones to collect high-resolution images of the main structure of the building construction site and build a dataset.
[0010] The second step is to label and preprocess the dataset, dividing it into training, validation and test sets, and converting it into the YOLOv12 network model recognition format.
[0011] The third step involves training the dataset using an improved YOLOv12 model, the improvements of which include:
[0012] In the backbone network and neck network, the A2C2f module is replaced with the A2C2f_Mcva module, which is composed of A2C2f and the multi-cognitive visual adapter Mcva module;
[0013] Replace the C3k2 module with the C3k2-SLBlock module. The C3k2-SLBlock module models neighborhood relationships by introducing the SLConv module and combining large kernel static convolution (LKP) and small kernel dynamic convolution (SKA).
[0014] In the neck network, the Concat module is replaced with a GDSA feature fusion module, which uses Ctm as the hierarchical feature modulator;
[0015] The fourth step is to train the improved YOLOv12 model using the training and validation sets to obtain the optimal engineering main structure detection model based on the improved YOLOv12.
[0016] The fifth step is to use drones to capture real-time footage of the construction activities;
[0017] The sixth step is to connect the drone to the computer, retrieve real-time images of the main construction structure of the construction area, and extract the construction area.
[0018] Step 7: Use the optimal main structure detection model obtained in step 4 to segment the main structure image of the construction project. In the segmented categories, determine the current construction activity based on the construction sequence code and calculate the area of the current construction activity area.
[0019] Step 8: Calculate the preliminary construction progress based on the current construction activity area. Based on the assessment of current construction activities and the calculation of the progress correction factor δ using CAD software, the preliminary construction progress is determined. And the schedule correction factor δ is used to calculate the final construction schedule. .
[0020] Furthermore, the improved YOLOv12 model described in the third step is an improvement upon the YOLOv12 network, and the improvements include:
[0021] In the backbone network and neck network, the A2C2f module is replaced with the A2C2f_Mcva module, which is composed of A2C2f and the multi-cognitive visual adapter Mcva module;
[0022] Replace the C3k2 module with the C3k2-SLBlock module. The C3k2-SLBlock module models neighborhood relationships by introducing the SLConv module and combining large kernel static convolution (LKP) and small kernel dynamic convolution (SKA).
[0023] In the neck network, the Concat module is replaced with a GDSA feature fusion module, which uses Ctm as the hierarchical feature modulator.
[0024] Furthermore, the A2C2f_Mcva module combines A2C2f with a multimodal visual converter and the Mcva module. Its implementation process is as follows: The input tensor X is first fed into an initial convolutional layer constructed based on the parent class initialization mechanism. This initial convolutional layer compresses the channel dimension of the input feature map through low-rank mapping. Using a sliding window convolution operation in the feature map space, the weights are multiplied and accumulated point-by-point with the corresponding elements, completing the dimensionality transformation from the input channel to the hidden channel. This extracts feature information with basic representational capabilities. The feature map after convolution is temporarily stored in a specific list data structure. Subsequently, the feature processing flow enters a loop processing stage composed of a sequence of Mcva modules. The last element of the list is used as input and passed sequentially to each module for processing. The feature vector X2 generated after processing by the Mcva module is fused with the original input feature vector X through a residual mapping mechanism, ultimately outputting the fused feature Y. The mathematical expression of its forward propagation process is:
[0025] .
[0026] Furthermore, the multi-cognitive visual adapter Mcva module is a lightweight nonlinear channel transformation module, consisting of normalization, channel compression and restoration, lightweight convolution, adaptive fusion, and residual connections. The multi-cognitive visual adapter Mcva module first receives the input feature X, performs normalization on it, and introduces two learnable parameters. and This process achieves adaptive fusion between normalized features and original features. The fused features are then reduced to 64 channels via 1×1 convolution, and then input into the lightweight convolution module McvaOp for cross-channel information modeling. Subsequently, overfitting is suppressed by activation functions and regularization techniques, and the channel dimensions are restored by 1×1 convolution. Finally, a residual connection is made with the original input X to obtain the output result. The forward propagation mathematical expression is as follows:
[0027] Furthermore, the lightweight convolutional module McvaOp is a lightweight convolutional submodule within the Mcva module used to enhance cross-channel perception capabilities. Its core consists of multi-scale depthwise separable convolutions. Input features are first passed through three depthwise convolutional branches with different kernel sizes (3×3, 5×5, and 7×7) to capture multi-scale contextual information. Subsequently, the results from each branch are fused using mean fusion, and the initial input features are introduced into skip connections to retain key information. The fused result is then passed through a 1×1 point convolution to achieve cross-channel information interaction, and finally, residual fusion is performed again with the input features to improve the expressive power and stability of the features. Its forward propagation mathematical expression is:
[0028]
[0029] Where DW represents depthwise convolution, PW represents pointwise convolution, and X represents the input feature. This represents the intermediate features after multi-scale fusion, and the final output is the result after two residual connections.
[0030] Furthermore, the C3k2_SLBlock module, by introducing the SLConv module, uses Large Kernel Perception (LKP) to model the neighborhood relationship of the expanded receptive field through large kernel static convolution. The implementation process is as follows: the feature vector is first processed by the convolution module to generate a feature map; then the input is sequentially processed by the SLBlock module to extract and transform features; then the outputs of all sub-modules are accumulated through residual connections to generate the final output feature map.
[0031] Furthermore, the SLBlock module is used for deep modeling and nonlinear enhancement of input features. Its structure adaptively selects local or global modeling methods according to different network layers and stages. The implementation process is as follows: When the depth is even, the SLBlock module uses the RepVGGDW module for local modeling; then, a channel attention mechanism is introduced to enhance the response of salient feature channels; when the depth is odd, the module distinguishes according to the stage: if it is currently in a high-level stage, the Attention module is used for long-distance dependency modeling; otherwise, a lightweight local spatial convolutional structure SLConv is used to extract spatial mixture features; finally, feature reconstruction and transformation are performed through a feedforward neural network, and residual connections are used to improve training stability and information retention. The mathematical expression of the above process is as follows:
[0032]
[0033] in, Represents the input image features; depth represents the depth index of the current network layer, used to control module branch switching; stage represents the stage number of the network; Attention(X) represents the multi-head attention mechanism, which enhances global modeling capabilities; used for channel attention enhancement; FFN(•) represents the feedforward network module, which includes nonlinear mapping and residual connections; This represents the intermediate features output by the multi-branch module; This represents the features introduced by channel attention; This indicates the final output of the module, which includes the fusion result of two residual connections.
[0034] Furthermore, the RepVGGDW module is based on the RepVGG multi-branch structure reparameterization concept, integrating 3×3 depthwise convolution, 1×1 depthwise convolution, and identity mapping to improve feature representation while maintaining efficient structural representation. During the training phase, the RepVGGDW module employs a three-branch structure: a 3×3 depthwise separable convolution, a 1×1 depthwise separable convolution, and an identity mapping branch to process the input feature X, thereby enhancing the model's local perception and residual learning capabilities. Then, the outputs of the three branches are element-wise summed during forward propagation to output the feature vector Y. Secondly, to accelerate the inference process, a structure reparameterization mechanism is introduced, converting the multi-branch structure of the training phase into a single 3×3 depthwise convolution, thus simplifying and accelerating the computation graph during deployment. The output of the RepVGGDW module during the forward propagation phase is:
[0035] .
[0036] Furthermore, the SLConv module uses Large Kernel Perception (LKP) to model the neighborhood relationship of the expanded receptive field through large kernel static convolution, and Small Kernel (SKA) to aggregate surrounding features adaptively through small kernel dynamic convolution. The implementation process is as follows: First, the input feature map is processed by the LKP module to generate spatially adaptive dynamic convolution kernel weights to guide the subsequent sparse convolution process; then, the SKA module is used to perform sparse attention convolution operations based on input features and dynamic kernel weights to enhance salient features in local regions; finally, batch normalization is introduced to stabilize the training process and improve the convergence speed.
[0037] Furthermore, the LKP module adopts a large-core bottleneck block design framework, and its implementation process is as follows: Given a visual feature map First, pointwise convolution (PW) is used to project the token to a lower channel dimension, which is the default value. This reduces computational costs and makes the model more lightweight; for Then, a kernel size of [missing information] was used. Large kernel depthwise separable convolution (DW) to effectively capture The large receptive field context information; large kernel depthwise separable convolution can expand the receptive field and enhance context awareness at minimal cost; then, pointwise convolution (PW) is used to model the spatial relationship between tokens to generate context-adaptive weights for the aggregation step; the whole process is represented as:
[0038]
[0039] in, Represented as the generated weights; Represented as adaptive weights; Represented as Surrounding size is The neighborhood of.
[0040] Furthermore, the SKA module adopts a grouped dynamic convolution design framework, and its implementation process is as follows: For visual feature maps Its channels are divided into Groups, each group contains Each channel shares aggregate weights within the same group to reduce memory overhead and computational cost, thereby achieving a lightweight model; for each The corresponding weights Reconstruction Then, using Aggregate its highly correlated Neighborhood, through and The convolution operation between them yields their aggregated features. This method effectively represents adaptive, fine-grained features, allowing the model to remain sensitive to dynamic and complex changes in various contexts. The entire process can be represented as:
[0041]
[0042] in, Represented as aggregate features; Represented as the generated weights; Represented as small core size; It is expressed as being based on Centered on, size is ; Represented as a channel group; Indicated as The One channel.
[0043] Furthermore, in the neck network, the Concat module is replaced with a GDSA feature fusion module, which uses Ctm as the hierarchical feature modulator. The implementation process is as follows: First, the input feature map is processed, which involves two sets of... Convolution generates queries respectively s and keys The original spatial dimensions are preserved for subsequent association calculations; K compresses the spatial dimension through adaptive average pooling to efficiently capture contextual information; then, the channels of Q and K are grouped, uniformly divided into G groups, each corresponding to an independent sub-channel; after flattening Q and K into two-dimensional matrices for each sub-channel group, the affinity matrix is calculated through matrix multiplication, which represents the association strength between each spatial location (token) and the compressed context region; next, dynamic context fusion is achieved based on the affinity matrix. After normalizing the affinity matrix, it is weighted and aggregated with K to obtain the contextual features of each token; the contextual features are reshaped into convolutional kernels and applied to Q through a sliding window to achieve dynamic contextual information injection, thereby modulating the long-term dependencies of each token; subsequently, channel attention is enhanced through an SE layer; the features output by Ctm are compressed through global average pooling, and then channel weights are generated through two fully connected layers and Sigmoid activation to recalibrate the features to enhance the response of key channels; simultaneously, gated branch signals are generated; the input features are processed by... Convolutional channel adjustment and SiLU activation function processing generate a gated signal to filter out invalid contextual noise; finally, feature fusion is performed; the SE layer output is then element-wise multiplied with the gated signal to suppress noise, and then... The number of channels is adjusted by convolution to obtain the final fused features, thus completing the global dynamic self-attention modeling.
[0044] Furthermore, the Ctm represents the relationship between a token and its context using the set of association values between this single token and all tokens in a set of region centers in the feature map; then, these association values are aggregated to define a token-level dynamic convolutional kernel in a learnable manner, thereby injecting contextual knowledge into each weight of the convolutional kernel; once such a dynamic kernel is applied to the feature map via a sliding window, each token in the feature map is modulated by approximately global information collected through the region centers; a token-based global contextual representation, given an input feature map. First, it is transformed into two parts, namely:
[0045] and
[0046] in, and express Convolutional layer This represents the reshaping operation, where K represents the aggregation of X to... through adaptive average pooling. Regional center
[0047] Next, the channels of Q and K are evenly divided into Group, get and , making and Because each pair and All of them have been flattened into two-dimensional matrices, and simple matrix multiplication between them can be used to calculate... Affinity matrix ,in Affinity matrix The Okay, that is ,save The Middle Each token and The affinity value between all tokens in the set.
[0048] Furthermore, the eighth step specifically includes: First, coding the different construction activities according to the actual construction sequence; when multiple construction activities are identified, prioritizing them to determine the ongoing activities; second, using CAD software to outline the external scaffolding of the building at the filming location, individual construction activities, and the area to be completed, and using the CAD software to calculate the total area of the outlined outlines and obtain the progress correction factor. :
[0049] ;
[0050] in This represents the area of the external scaffolding. The relevant area is the area that should be completed during construction activities;
[0051] The construction activity area was initially used in the above segmentation algorithm. Divide by the area of the external scaffolding Calculate construction progress :
[0052] ;
[0053] Based on preliminary construction progress And the schedule correction factor δ is used to calculate the final construction schedule. :
[0054] .
[0055] Compared with the prior art, the present invention has the following beneficial effects:
[0056] I. Enhance the model's visual perception capability
[0057] By designing the A2C2f_Mcva module (which integrates the A2C2f module with the multimodal vision converter Mcva module), a visually friendly filter is used to replace the traditional language-friendly linear filter, and a scaling and normalization layer is added to the adapter to adjust the input feature distribution. Only a few parameters need to be fine-tuned to improve model performance, avoiding the overfitting problem of full-scale fine-tuning, while also enhancing the ability to model the boundaries of the main engineering structure.
[0058] II. Optimizing Large-Scale Structural Modeling Capabilities
[0059] To address the limitations of traditional convolutional algorithms with their fixed receptive field and static weights, the C3k2-SLBlock module was designed. It expands the receptive field through large-kernel static convolution (LKP) to model neighborhood relationships, and adaptively integrates surrounding features using small-kernel dynamic convolution (SKA). This enables more accurate differentiation of repetitive structures such as wall templates, improves robustness to local occlusion and lighting changes, and effectively models the spatial topological relationships between structures, aiding in judgment during the construction phase.
[0060] III. Enhancing the semantic selectivity of feature fusion
[0061] To address the shortcomings of the YOLOv12 standard feature fusion module Concat (which fails to explicitly distinguish semantic importance, leading to background texture interference with the modeling of the main structure), a GDSA feature fusion module (using Ctm as a hierarchical feature modulator) is introduced into the neck network. By injecting contextual knowledge through dynamic convolutional kernels, global information modulation of feature map labels is achieved, enhancing long-term dependency modeling capabilities, optimizing semantically selective fusion and feature guidance, and improving the detection accuracy and recall of main structures (especially those with blurred outlines and small object defects).
[0062] IV. Achieving Adaptive Calculation of Construction Progress
[0063] By using the sequential coding of construction activities after network segmentation, the priority of construction activities and currently ongoing activities can be automatically determined. Combined with the construction activity area, the area of external scaffolding, and the progress correction factor calculated from CAD, specific progress values are automatically output. Custom sequential coding and correction factor updates are supported, significantly improving the method's adaptability to different construction projects. Attached Figure Description
[0064] Figure 1 This is a flowchart of an automatic detection method for the construction progress of an engineering main structure based on UAV vision, according to an embodiment of the present invention.
[0065] Figure 2 This is a diagram of an engineering main structure detection model based on YOLOv12, according to an embodiment of the present invention.
[0066] Figure 3 This is a schematic diagram of the A2C2f_Mcva module structure;
[0067] Figure 4 This is a schematic diagram of the Mcva module structure;
[0068] Figure 5 This is a schematic diagram of the McvaOp module structure;
[0069] Figure 6 This is a schematic diagram of the C3k2-SLBlock module structure;
[0070] Figure 7 This is a schematic diagram of the SLBlock module structure;
[0071] Figure 8 This is a schematic diagram of the RepVGGDW module structure;
[0072] Figure 9 This is a schematic diagram of the SLConv module structure;
[0073] Figure 10 This is a schematic diagram of the GDSA module structure;
[0074] Figure 11 This is a schematic diagram of the Ctm module structure. Detailed Implementation
[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0076] like Figure 1 As shown in the figure, this embodiment of the invention provides an automatic detection method for the construction progress of the main structure of an engineering project based on UAV vision. The specific steps are as follows:
[0077] The first step is to use drones to collect high-resolution images of the main structure of the building construction site and build a dataset.
[0078] This invention constructs a dataset by using aerial photography taken by a DJI Mavic 4 Pro drone of a residential building under construction at a construction site in Wuhan.
[0079] The second step is to label and preprocess the dataset, dividing it into training, validation, and test sets, and converting it into the YOLOv12 network model recognition format.
[0080] To obtain construction progress data, it is necessary to identify the Space and construction areas. This invention replicates images to obtain two identical image datasets, named Data-Space and Data-Construction, respectively. Data-Space contains only one category: the overall main structure area of the project, labeled as Space, used to train the network to recognize scaffolding on the building's exterior. Data-Construction contains seven classes representing construction progress: Column-RB (for column-reinforced concrete), Column-FC (for column-formwork erection), Scaffold (for scaffolding erection), Beam&Slab-FC (for beam-slab-formwork erection), Beam-RB (for beam-reinforced concrete), Slab-RB (for slab-reinforced concrete), and Column&Beam&Slab-CP (for column-beam-slab-concrete pouring). These are used to train the network to recognize construction stages and their extent. The two datasets differ in their class settings, but the other processing methods are identical. This invention employs supervised machine learning, requiring image category labeling using Roboflow and a polygon annotation method. The Data-Space dataset labels one class, while the Data-Construction dataset labels seven recognition classes. After image labeling, data augmentation is used to increase the dataset size. The training images are randomly rotated 90 degrees clockwise or counter-clockwise; then, 0.1% noise is added to the images, meaning 0.1% of pixels may display abnormal colors or brightness. This data augmentation triples the size of the training set. The output Space dataset includes 246 training images, 18 validation images, and 12 test images. The Construction dataset includes 489 training images, 32 validation images, and 24 test images. Finally, the data is exported in YOLOv12 format.
[0081] The third step is to train the dataset using the improved YOLOv12 model.
[0082] The improved YOLOv12 model is an improvement on the YOLOv12 network to build an engineering main structure identification and detection model.
[0083] The improved YOLOv12 design and uses an A2C2f_Mcva module that combines A2C2f with the multimodal vision converter Mcva module at the backbone and neck network to replace the original model's A2C2f module.
[0084] In the design and neck network, the C3k2-SLBlock module replaced the original C3k2 module. Small-Large Convolution was introduced, using Large Kernel Perception (LKP) to model the neighborhood relationships of the expanded receptive field through static convolution of a large kernel. Small Kernel Aggregation (SKA) adaptively integrates surrounding features through dynamic convolution of a small kernel. Finally, in the neck network, a GDSA feature fusion module using Ctm as the hierarchical feature modulator replaced the original Concat module. Figure 2 As shown.
[0085] The A2C2f_Mcva module enhances its ability to process visual signals by introducing multiple visually friendly filters into the adapter, abandoning the previous reliance on language-friendly linear filters. Secondly, a scaling normalization layer is added to the adapter to adjust the distribution of input features for the visual filters. Fine-tuning only a few parameters allows the model to achieve better results, avoiding overfitting caused by full-scale fine-tuning; the visual perception filter enhances the modeling ability of the engineering structure boundaries. The input tensor is first fed into an initial convolutional layer constructed according to the parent class initialization mechanism. This layer compresses the channel dimension of the input feature map through low-rank mapping. Using a sliding window convolution operation in the feature map space, the weights are multiplied and accumulated point-by-point with the corresponding elements, completing the dimensionality transformation from input channels to hidden channels, thereby extracting feature information with basic representation capabilities. The feature map after this convolution is temporarily stored in a specific list data structure. Subsequently, the feature processing flow enters a loop processing stage consisting of a sequence of Mcva modules. In this stage, the last element of the list is used as input and sequentially passed to each module for processing. The new feature vector generated after processing by the Mcva module will be fused with the original input feature vector through a residual mapping mechanism, ultimately outputting the fused feature. For example... Figure 3 As shown.
[0086] The Mcva module of the multi-cognitive visual adapter is a lightweight non-linear channel transformation module, mainly composed of normalization, channel compression and restoration, lightweight convolution, adaptive fusion, and residual connections. This module first receives input features, normalizes them, and introduces two learnable parameters to achieve adaptive fusion between the normalized features and the original features. The fused features are then reduced to 64 channels through a 1×1 convolution, and then input into the lightweight convolution module McvaOp for cross-channel information modeling. Subsequently, overfitting is suppressed by activation functions and regularization techniques, and the channel dimensions are restored through a 1×1 convolution. Finally, a residual connection is made with the original input feature vector to obtain the output result. Figure 4 As shown.
[0087] McvaOp is a lightweight convolutional submodule within the Mcva module of the multi-cognitive visual adapter, designed to enhance cross-channel perception capabilities. Its core consists of multi-scale depthwise convolutions. Specifically, input features are first passed through three depthwise convolutional branches with different kernel sizes (3×3, 5×5, and 7×7) to capture multi-scale contextual information. Subsequently, the results from each branch are fused using mean fusion, and the initial input features are introduced into a skip connection to preserve key information. The fused result is then passed through a 1×1 point convolution to achieve cross-channel information interaction, and finally, residual fusion is performed again with the input features to improve the expressive power and stability of the features. Figure 5 As shown.
[0088] Furthermore, the C3k2_SLBlock module, by introducing the SLConv module, uses Large Kernel Perception (LKP) to model the neighborhood relationship of the expanded receptive field through large kernel static convolution. The implementation process is as follows: the feature vector first passes through a convolution module to generate a feature map; then, the input is sequentially processed and transformed by the SLBlock module; finally, the outputs of all sub-modules are accumulated through residual connections to generate the final output feature map. For example... Figure 6 As shown.
[0089] The SLBlock module is used for deep modeling and nonlinear enhancement of input features. Its structure adaptively selects between local and global modeling methods based on the network layer and stage. When the depth is even, the module uses the RepVGGDW module for local modeling; subsequently, a channel attention mechanism is introduced to enhance the response of salient feature channels. When the depth is odd, the module differentiates based on the current stage: if it is a high-level stage, an attention mechanism module is used for long-distance dependency modeling; otherwise, a lightweight local spatial convolutional structure, SLConv, is used to extract spatial mixture features. Finally, a feedforward neural network is used for feature reconstruction and transformation, and residual connections are employed to improve training stability and information retention. Figure 7 As shown.
[0090] The RepVGGDW module is based on the RepVGG multi-branch structure reparameterization concept, integrating 3×3 depthwise convolution, 1×1 depthwise convolution, and identity mapping to improve feature representation while maintaining efficient structural representation. During training, the RepVGGDW module employs a three-branch structure: a 3×3 depthwise separable convolution, a 1×1 depthwise separable convolution, and an identity mapping branch to process input features, enhancing the model's local perception and residual learning capabilities. Then, the outputs of the three branches are element-wise summed during forward propagation to output a feature vector. Furthermore, to accelerate the inference process, a structural reparameterization mechanism is introduced, converting the multi-branch structure of the training phase into a single 3×3 depthwise convolution, thereby simplifying and accelerating the computation graph during deployment. Figure 8 As shown.
[0091] The SLConv module integrates a Locally Learnable Kernel Generator (LKP) and a Sparse Kernel Attention (SKA) mechanism. It primarily uses Large Kernel Perception (LKP) to model the neighborhood relationships of the expanded receptive field through large kernel static convolutions. Small Kernel (SKA) aggregates surrounding features adaptively through small kernel dynamic convolutions. The implementation process is as follows: First, the input feature map is processed by the LKP module to generate spatially adaptive dynamic convolution kernel weights to guide the subsequent sparse convolution process; then, the SKA module performs sparse attention convolution operations based on the input features and dynamic kernel weights to enhance salient features in local regions; finally, batch normalization is introduced to stabilize the training process and improve convergence speed. Figure 9 As shown.
[0092] The large kernel processing module employs the Large Convolutional Kernel Bottleneck Block (LKP) design framework. Its implementation process is as follows: Given a visual feature map, firstly, pointwise convolution is used to project the feature vector to a lower channel dimension, thereby reducing computational cost and making the model more lightweight. Then, for this feature vector, a large convolutional kernel depthwise separable convolution is used to effectively capture its large receptive field contextual information. Large convolutional kernel depthwise separable convolution can expand the receptive field and enhance context awareness at minimal cost. Next, pointwise convolution is used to model the spatial relationships between feature vectors, generating context-adaptive weights for the aggregation step.
[0093] The SKA module employs a grouped dynamic convolution design framework. For visual feature maps, their channels are divided into G groups. Each group contains the number of channels divided by G channels, and channels within the same group share aggregate weights, thereby reducing memory overhead and computational cost, thus achieving a lightweight model. For each feature vector, its corresponding weights are reshaped. Then, its aggregated features are obtained by performing a convolution operation on the neighborhood and dimension. In this way, adaptive fine-grained features can be effectively represented, allowing the model to remain sensitive to dynamic and complex changes in various contexts.
[0094] The GDSA feature fusion module, using CTM as a hierarchical feature modulator, represents the relationship between a token and its context by using the set of association values between this single token and all tokens in a set of region centers in the feature map. These affinity values are then aggregated to define a token-level dynamic convolutional kernel in a learnable manner, thus injecting contextual knowledge into each weight of the kernel. Once such a dynamic kernel is applied to the feature map via a sliding window, each token in the feature map is modulated by approximately global information collected through the region centers. Therefore, long-term dependencies can be effectively modeled. This is based on a global contextual representation of informational features, enabling semantically selective fusion and feature-guided enhancement. First, the input feature map is processed. The input feature map is processed through two sets of... Convolution generates query s and keys The original spatial dimensions are preserved for subsequent association calculations; K compresses the spatial dimension through adaptive average pooling to efficiently capture contextual information. Then, the channels of Q and K are grouped. The channels of both are evenly divided into G groups, each corresponding to an independent sub-channel. After flattening Q and K into two-dimensional matrices for each sub-channel group, an affinity matrix is calculated through matrix multiplication. This matrix represents the association strength between each spatial location token and the compressed context region. Next, dynamic context fusion is implemented based on the affinity matrix. After normalizing the affinity matrix, it is weighted and aggregated with K to obtain the contextual features of each token. This feature is reshaped into a convolutional kernel and applied to Q through a sliding window to inject dynamic contextual information, thereby modulating the long-term dependencies of each token. Subsequently, channel attention is enhanced through an SE layer. The features output by Ctm are compressed through global average pooling, and then channel weights are generated through two fully connected layers and a Sigmoid activation layer. The features are recalibrated to enhance the response of key channels. Simultaneously, gated branch signals are generated. The input features are processed... Convolutional channel adjustment and SiLU activation function processing generate a gated signal to filter out invalid contextual noise. Finally, feature fusion is performed to output the feature. The SE layer output is then element-wise multiplied with the gated signal to suppress noise, and then... Convolution adjusts the number of channels to obtain the final fused features, completing global dynamic self-attention modeling. For example... Figure 10 As shown.
[0095] Furthermore, for the CTM module, the main idea is to represent the relationship between a token and its context by utilizing the set of association values between a single token and all tokens in a set of region centers in the feature map. Subsequently, these affinity values can be aggregated, and token-level dynamic convolutional kernels can be defined in a learnable manner, thereby injecting contextual knowledge into each weight of the convolutional kernel. When such dynamic convolutional kernels are applied to the feature map through a sliding window, each token in the feature map is modulated by approximately global information collected through the region centers. Therefore, long-term dependencies can be effectively modeled. Given an input feature map, it is first transformed into two parts: the first part is transformed by a one-to-one convolutional layer and then reshaped to obtain a query tensor. The second part is aggregated from the input feature map by adaptive average pooling, then transformed by a one-to-one convolutional layer and reshaped into a key tensor. Next, the channels of the query tensor and key tensor are evenly divided into G groups, resulting in G groups of query sub-tensors and G groups of key sub-tensors. For each group of query sub-tensors and key tensors, the corresponding affinity matrix is calculated through matrix multiplication. Specifically, after flattening each group into a two-dimensional matrix, the product of the transposed key tensor and the query tensor is calculated, ultimately yielding G affinity matrices. For example... Figure 11 As shown.
[0096] The fourth step is to train the improved YOLOv12 model using the training and validation sets to obtain the optimal engineering main structure detection model based on the improved YOLOv12.
[0097] The training and validation sets are input into the improved YOLOv12 model (i.e., the engineering main structure recognition and detection model), and the number of training iterations is set. As the number of training iterations increases, the loss function curve of the model gradually converges. When the loss function curve converges and stabilizes, the recognition and detection model is trained to the optimal state, and its optimal model weight file is saved. The images to be detected in the test set are input into the trained optimal engineering main structure detection model based on the improved YOLOv12, and the segmented images of the engineering main structure can be output. The segmented images include label information and detection information.
[0098] The fifth step is to use drones to capture real-time footage of the construction activities;
[0099] The sixth step is to connect the drone to the computer, retrieve real-time images of the main construction structure of the construction area, and extract the construction area.
[0100] This invention uses a DJI Mavic 4 Pro drone, which is connected to a computer's data transmission interface via the HDMI port of the RCPro remote controller to retrieve real-time footage of the construction site captured by the drone.
[0101] The sixth step is to use the optimal main structure detection model obtained in the fourth step to segment the construction main structure image, determine the current construction activity based on the construction sequence code in the segmented categories, and calculate the area of the current construction activity area.
[0102] First, the YOLO algorithm is used to segment the image, identify and delineate the mask contours of each target region; then, OpenCV is used to extract the contours and calculate the pixel area of each region; combined with parameters such as image resolution and shooting distance, the proportion is converted, and finally the area results of each region are output.
[0103] Step 7: Calculate the preliminary construction schedule based on the current construction activity area. Based on the assessment of current construction activities and the calculation of the progress correction factor δ using CAD software, the preliminary construction progress is determined. And the schedule correction factor δ is used to calculate the final construction schedule. .
[0104] Specifically, different construction activities are coded sequentially according to the actual construction sequence. When multiple construction activities are identified, the ongoing activity is determined by priority ranking. Next, the outlines of the building's external scaffolding, individual construction activities, and the area to be completed are drawn in CAD software. The CAD software then calculates the total area of the drawn outlines and determines the progress correction factor. The calculation formula is as follows: ,in To determine the area value of the external scaffolding through a specific method (such as CAD hand-drawing), This refers to the area that should be completed for the relevant construction activities.
[0105] The construction activity area was initially used in the above segmentation algorithm. Divide by the area of the external scaffolding Calculate construction progress . The calculation formula is Using the aforementioned progress correction parameters and combining them with the model recognition results, the final construction progress is calculated. .
[0106] To verify the effectiveness of the invention, comparative experiments were conducted, and the datasets used in the experiments were all self-built engineering main structure datasets.
[0107] Experimental example:
[0108] Experimental environment for this invention:
[0109] The hardware and software environment required for this system to run is shown in Table 1.
[0110] Table 1 Hardware and Software Environment
[0111]
[0112] Comparative experiments of the present invention:
[0113] Table 2 Comparison of experimental results
[0114]
[0115] Wherein, Net represents the network model name; Style refers to the deep learning framework on which the model is based; mAP / IoU=0.5(space) represents the average precision calculated for space-related tasks under the standard of Intersection over Union (IoU) of 0.5, reflecting the model's detection accuracy; the higher the value, the more accurate the detection; mAP(space) represents the average precision for space tasks without specifying an IoU threshold, comprehensively reflecting the model's performance in this task; mAP / IoU=0.5(construction) represents the average precision for construction-related tasks with an IoU of 0.5; mAP(construction) represents the average precision for construction tasks without specifying an IoU threshold.
[0116] Here, "space" refers to the external scaffolding, and "Construction" refers to a single construction activity.
[0117] As shown in Table 2 and the comparative experimental results, the improved YOLOv12 algorithm of this invention achieves a mAP / IoU=0.5 of 99.0% and an mAP of 99.5% in the space task, significantly higher than YOLOv8, Mask R-CNN, and DeepLabv3. This indicates that the improved YOLOv12 is more accurate in target localization and recognition, and can more accurately capture the position and category information of target objects, reducing false detections and false negatives. In the construction task, the improved YOLOv12 algorithm achieves an mAP / IoU=0.5 of 91.5% and an mAP of 70.1%, also outperforming some of the compared algorithms. This demonstrates its good adaptability and detection performance even in complex building-related target detection scenarios. In contrast, algorithms such as Mask R-CNN have relatively low accuracy in this task, further highlighting the effectiveness of the improved algorithm of this invention.
[0118] This invention has the following characteristics:
[0119] (1) Enhance visual perception: By introducing a visually friendly filter and normalization layer through the A2C2f_Mcva module, the model can be improved to perceive the boundaries, materials and geometric contours of engineering structures by only fine-tuning a small number of parameters, thus avoiding the overfitting problem of full fine-tuning.
[0120] (2) Optimize spatial relationship modeling: The C3k2-SLBlock module combines large kernel static convolution (LKP) and small kernel dynamic convolution (SKA) to expand the receptive field and adaptively integrate surrounding features, improve the discrimination of repetitive structures such as wall templates, enhance the robustness to local occlusion and lighting changes, and more accurately capture the spatial topological relationship of the structure.
[0121] (3) Improve the quality of feature fusion: The GDSA feature fusion module uses Ctm as the hierarchical feature modulator. Through dynamic context information injection and channel attention enhancement, it achieves semantic selective fusion, reduces background texture interference, and solves the problems of blurred main structure boundaries and easy omission of small targets in the early construction stage.
[0122] (4) Adaptive construction progress calculation: The current activity is determined by the construction sequence code, and the progress correction coefficient is calculated by combining CAD software. The coding rules and correction coefficients can be customized to adapt to the needs of different construction projects and realize the automatic output of specific progress values. Experimental verification shows that the detection accuracy is significantly better than that of traditional algorithms.
[0123] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for automatically detecting the construction progress of the main structure of an engineering project based on UAV vision, characterized in that, Includes the following steps: The first step is to use drones to collect high-resolution images of the main structure of the building construction site and build a dataset. The second step is to label and preprocess the dataset, dividing it into training, validation and test sets, and converting it into the YOLOv12 network model recognition format. The third step is to train the dataset using the improved YOLOv12 model; The fourth step is to train the improved YOLOv12 model using the training and validation sets to obtain the optimal engineering main structure detection model based on the improved YOLOv12. The fifth step is to use drones to capture real-time footage of the construction activities; The sixth step is to connect the drone to the computer, retrieve real-time images of the main construction structure of the construction area, and extract the construction area. Step 7: Use the optimal main structure detection model obtained in step 4 to segment the main structure image of the construction project. In the segmented categories, determine the current construction activity based on the construction sequence code and calculate the area of the current construction activity area. Step 8: Calculate the preliminary construction progress based on the current construction activity area. Based on the assessment of current construction activities and the calculation of the progress correction factor δ using CAD software, the preliminary construction progress is determined. And the schedule correction factor δ is used to calculate the final construction schedule. .
2. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 1, characterized in that: The improved YOLOv12 model described in step three is an improvement upon the YOLOv12 network, and the improvements include: In the backbone network and neck network, the A2C2f module is replaced with the A2C2f_Mcva module, which is composed of A2C2f and the multi-cognitive visual adapter Mcva module; Replace the C3k2 module with the C3k2-SLBlock module. The C3k2-SLBlock module models neighborhood relationships by introducing the SLConv module and combining large kernel static convolution (LKP) and small kernel dynamic convolution (SKA). In the neck network, the Concat module is replaced with a GDSA feature fusion module, which uses Ctm as the hierarchical feature modulator.
3. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 2, characterized in that: The A2C2f_Mcva module combines A2C2f with a multimodal visual converter and the Mcva module. Its implementation process is as follows: The input tensor X is first fed into an initial convolutional layer constructed based on the parent class initialization mechanism. This initial convolutional layer compresses the channel dimension of the input feature map through low-rank mapping. Using a sliding window convolution operation in the feature map space, the weights are multiplied and added point-by-point with the corresponding elements, completing the dimensionality transformation from the input channel to the hidden channel. This extracts feature information with basic representational capabilities. The convolutionally processed feature map is temporarily stored in a specific list data structure. Subsequently, the feature processing flow enters a loop processing stage composed of a sequence of Mcva modules. The last element of the list is used as input and passed sequentially to each module for processing. The feature vector X2 generated after processing by the Mcva module is fused with the original input feature vector X through a residual mapping mechanism, ultimately outputting the fused feature Y. The mathematical expression of its forward propagation process is: 。 4. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 2, characterized in that: The multi-cognitive visual adapter Mcva module is a lightweight nonlinear channel transformation module, consisting of normalization, channel compression and restoration, lightweight convolution, adaptive fusion, and residual connections. The multi-cognitive visual adapter Mcva module first receives the input feature X, normalizes it, and introduces two learnable parameters. and This process achieves adaptive fusion between normalized features and original features. The fused features are then reduced to 64 channels via 1×1 convolution, and then input into the lightweight convolution module McvaOp for cross-channel information modeling. Subsequently, overfitting is suppressed by activation functions and regularization techniques, and the channel dimensions are restored by 1×1 convolution. Finally, a residual connection is made with the original input X to obtain the output result. The forward propagation mathematical expression is as follows: 。 5. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 4, characterized in that: The lightweight convolutional module McvaOp is a lightweight convolutional submodule in the Mcva module used to enhance cross-channel perception capabilities. Its core consists of multi-scale depthwise separable convolutions. Input features are first passed through three depthwise convolutional branches with different kernel sizes (3×3, 5×5, 7×7) to capture multi-scale contextual information. Subsequently, the results from each branch are fused by mean fusion, and the initial input features are introduced into a skip connection to preserve key information; The fused result is then subjected to a 1×1 point convolution to achieve cross-channel information interaction, and finally residual fusion is performed again with the input features to improve the expressive power and stability of the features. The forward propagation mathematical expression is: ; Where DW represents depthwise convolution, PW represents pointwise convolution, and X represents the input feature. This represents the intermediate features after multi-scale fusion, and the final output is the result after two residual connections.
6. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 2, characterized in that: The C3k2_SLBlock module, by introducing the SLConv module, uses Large Kernel Perception (LKP) to model the neighborhood relationship of the expanded receptive field through large kernel static convolution. The implementation process is as follows: the feature vector is first processed by the convolution module to generate a feature map; then the input is sequentially processed by the SLBlock module to extract and transform features; then the outputs of all sub-modules are accumulated through residual connections to generate the final output feature map.
7. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 2, characterized in that: The SLBlock module is used for deep modeling and nonlinear enhancement of input features. Its structure adaptively selects local or global modeling methods according to different network layers and stages. The implementation process is as follows: When the depth is even, the SLBlock module uses the RepVGGDW module for local modeling; then, a channel attention mechanism is introduced to enhance the response of salient feature channels; when the depth is odd, the module distinguishes according to the stage: if it is currently in a high-level stage, the Attention module is used for long-distance dependency modeling; otherwise, a lightweight local spatial convolutional structure SLConv is used to extract spatial mixture features; finally, feature reconstruction and transformation are performed through a feedforward neural network, and residual connections are used to improve training stability and information retention. The mathematical expression of the above process is as follows: ; in, Represents the input image features; depth represents the depth index of the current network layer, used to control module branch switching; stage represents the stage number of the network; Attention(X) represents the multi-head attention mechanism, which enhances global modeling capabilities; used for channel attention enhancement; FFN(•) represents the feedforward network module, which includes nonlinear mapping and residual connections; This represents the intermediate features output by the multi-branch module; This represents the features introduced by channel attention; This indicates the final output of the module, which includes the fusion result of two residual connections.
8. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 7, characterized in that: The RepVGGDW module is based on the multi-branch structure reparameterization idea of RepVGG, integrating 3×3 depthwise convolution, 1×1 depthwise convolution and identity mapping to improve feature representation while maintaining efficient structural representation. During the training phase, the RepVGGDW module adopts a three-branch structure, namely 3×3 depthwise separable convolution, 1×1 depthwise separable convolution and identity mapping branches to process the input feature X, so as to improve the model's local perception ability and residual learning ability. Then, the outputs of the three branches are element-wise summed during forward propagation to output the feature vector Y; secondly, to accelerate the inference process, a structural reparameterization mechanism is introduced to convert the multi-branch structure in the training phase into a single 3×3 Depthwise convolution, thereby simplifying and accelerating the computation graph in the deployment phase; the output of the RepVGGDW module in the forward phase is: 。 9. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 2, characterized in that: The SLConv module consists of a Large Kernel Perception (LKP) module that models the neighborhood relationships of the expanded receptive field through large kernel static convolution, and a Small Kernel (SKA) module that aggregates surrounding features adaptively through small kernel dynamic convolution. The implementation process is as follows: First, the input feature map is processed by the LKP module to generate spatially adaptive dynamic convolution kernel weights, which are used to guide the subsequent sparse convolution process. Then, the SKA module is used to perform sparse attention convolution operations based on the input features and dynamic kernel weights to enhance the salient features of local regions. Finally, batch normalization is introduced to stabilize the training process and improve convergence speed.
10. The automatic detection method for construction progress of engineering main structure based on UAV vision as described in claim 1, characterized in that: The eighth step specifically includes: First, according to the actual construction sequence, the different construction activities are coded sequentially; when multiple construction activities are identified, the ongoing activities are determined by priority ranking; second, the outlines of the building's external scaffolding, individual construction activities, and the area to be completed are drawn in CAD software, and the CAD software calculates the total area of the drawn outlines to obtain the progress correction factor. : ; in This represents the area of the external scaffolding. The relevant area is the area that should be completed during construction activities; The construction activity area was initially used in the above segmentation algorithm. Divide by the area of the external scaffolding Calculate construction progress : ; Based on preliminary construction progress And the schedule correction factor δ is used to calculate the final construction schedule. : 。
Citation Information
Cited By
Machine vision-based automatic identification method for assembly quality of assembly type station component
CN121366386A
Photovoltaic panel number detection method and device based on unmanned aerial vehicle vision and medium
CN121767351A