Method and system for construction fence structural change detection and multi-modal assisted disposition
Patent Information
- Application Number
- CN202610230240.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-02-26
AI Technical Summary
[0003]然而,上述现有技术在实际应用中仍存在以下问题和不足:首先,人工成本高且效率低下,无论是现场人工巡查还是视频人工盯防,均需占用大量人力,且受人员疲劳、注意力分散影响,难以实现持续、高精度的监控
[0053]本发明提供的施工围挡结构性变化检测及多模态辅助处置方法,通过结合视觉变换器的深度特征比对与单目深度估计,显著提升了变化检测的精度与语义理解能力,有效区分真实结构变化与光影、遮挡等干扰,大幅降低了误报率。同时,该方法创新性地将图像变化区域转换为近似物理尺寸,为安全风险评估和维修资源调度提供了客观、可量化的决策依据,实现了从定性观察到定量分析的跨越。通过引入多模态大语言模型,系统能够基于结构化检测结果自动生成专业、易懂的语义解释与维修建议,并将检测、分析、派单、复检流程整合为自动化闭环,极大地提升了施工安全管理的智能化水平和响应效率,减少了对专业人工巡检的依赖,具有良好的实用价值与经济效益。
Smart Images

Figure CN122024170B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of construction engineering safety management technology, and in particular relates to a method and system for detecting structural changes in construction site enclosures and for multimodal auxiliary handling. Background Technology
[0002] Construction site fencing is an important safety protection and isolation facility at urban construction sites. Its main function is to separate the construction area from the external environment, preventing unauthorized personnel from accidentally entering the construction site and causing safety accidents. It also plays a role in environmental beautification and noise and dust reduction. The stability and integrity of the fencing are directly related to the safety of pedestrians and vehicles in the surrounding area. Currently, the supervision of construction site fencing mainly relies on the following methods: 1. Regular manual inspections: Safety officers or project management personnel conduct daily or weekly foot inspections of the fencing at the construction site, visually checking for problems such as damage, tilting, collapse, graffiti, or missing components; 2. Manual analysis of video surveillance: Fixed cameras are installed at the construction site, transmitting on-site video to a monitoring center, where on-duty personnel review the footage to determine if the fencing is abnormal; 3. Image detection algorithms based on computer vision: Some studies use frame difference, background modeling, optical flow field estimation, or traditional machine learning algorithms and convolutional neural networks to perform target detection or semantic segmentation on single-frame images, attempting to automatically identify damaged areas of the fencing; 4. Sensor-based monitoring methods: Vibration sensors, tilt sensors, and switch sensors are installed on the fencing structure to detect whether the fencing is tilted, whether there has been a strong impact, or whether the fencing gate has been opened.
[0003] However, the aforementioned existing technologies still have the following problems and shortcomings in practical applications: First, they are labor-intensive and inefficient. Both on-site manual inspections and video monitoring require a large amount of manpower and are susceptible to fatigue and distraction, making it difficult to achieve continuous and high-precision monitoring. Second, they lack semantic understanding capabilities. Existing image analysis methods based on frame difference, optical flow, or CNNs mostly focus on pixel-level changes or texture anomalies, lacking semantic judgment on whether the changes constitute structural damage to the enclosure itself. They are easily affected by factors such as changes in lighting, temporary obstructions, and interference from pedestrians and vehicles, resulting in a high false alarm rate. Third, sensor solutions have a high false alarm rate and are complex to deploy and maintain. Vibration sensors and other sensors are easily affected by environmental noise, wind, and mechanical vibration. Hardware deployment and long-term maintenance costs are high, making large-scale application difficult. Finally, they lack closed-loop management support. Most systems only remain at the "detect anomalies and alarm" stage, failing to form an automated closed loop with maintenance, re-inspection, and other stages, thus failing to provide complete decision support and process optimization for safety management. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes a method and system for detecting structural changes in construction site enclosures and for multimodal assisted handling, thereby resolving the issues present in the prior art.
[0005] To achieve the above objectives, the present invention provides a method for detecting structural changes in construction site enclosures and for multimodal assisted handling, comprising:
[0006] Images of the construction enclosure area are acquired to obtain the current image. The current image and the reference image are then processed using an image registration method to obtain the registered current image.
[0007] Based on the reference image and the registered current image, a feature extraction model is used to process the image to obtain the reference image feature map and the current image feature map.
[0008] Based on the current image feature map, a depth prediction model is used to process it to obtain the current image depth map;
[0009] A change detection model is used to process the feature map of the reference image, the feature map of the current image, and the depth map of the current image to obtain the bounding box coordinates and change type labels of the changed area of the enclosure;
[0010] The physical dimensions of the changing region are obtained by calculating and processing the current image depth map, camera intrinsic parameters, and bounding box coordinates.
[0011] Based on the bounding box coordinates, the change type label, the physical dimensions, and the current image, a structured detection result is constructed;
[0012] Based on the structured detection results, a multimodal large language model is used for processing to output semantic explanations and maintenance suggestions;
[0013] Based on the structured detection results and the semantic interpretation and maintenance suggestions, a maintenance work order is generated.
[0014] Optionally, the process of processing the current image and the reference image using an image registration method includes:
[0015] Based on the reference image and the current image, the SIFT feature point detection algorithm is used to process and extract corresponding feature point pairs.
[0016] Based on the corresponding feature point pairs, the RANSAC robust fitting method is used to estimate the two-dimensional homography matrix.
[0017] Based on the two-dimensional homography matrix and the current image, a geometric transformation is performed to obtain the registered current image.
[0018] Optionally, the process of processing the reference image feature map, the current image feature map, and the current image depth map using a change detection model includes:
[0019] In the change region proposal stage, based on the feature map of the reference image and the feature map of the current image, the feature difference is calculated, and then processed in combination with high frequency enhancement and the enclosure interest region mask to generate candidate bounding boxes for change regions.
[0020] In the change category determination stage, feature and depth fusion is performed based on the current image feature map, the current image depth map, and the candidate bounding boxes of the change region. The change type label is obtained by retrieving a pre-built feature library.
[0021] Optionally, the process of calculating feature differences includes:
[0022] The feature vectors in the reference image feature map and the corresponding feature vectors in the current image feature map are subjected to L2 normalization to obtain normalized feature vectors.
[0023] Based on the normalized feature vector, calculate the cosine distance and L2 difference norm;
[0024] After normalizing the cosine distance plot and the L2 difference norm plot respectively, they are linearly fused to obtain the significance plot of basic changes.
[0025] Optionally, the feature library construction process includes:
[0026] Based on images of historical fence change samples, feature extraction, depth estimation, and change region proposal steps are performed to obtain candidate regions.
[0027] Based on the current image feature map and the current image depth map of the candidate region, feature and depth fusion processing is performed to obtain fused features;
[0028] The fusion features are flattened and subjected to dimensionality reduction based on principal component analysis to obtain dimensionality-reduced features;
[0029] The reduced-dimensional features are subjected to supervised feature transformation based on linear discriminant transformation to obtain the embedding vector;
[0030] The embedded vectors and their corresponding variation type labels are stored to form a feature library.
[0031] Optionally, the process of calculating and processing the current image depth map, camera intrinsic parameters, and bounding box coordinates to obtain the physical dimensions of the changed region includes:
[0032] Calculate the pixel width and pixel height based on the bounding box coordinates;
[0033] Based on the current image depth map and the bounding box coordinates, calculate the average depth value of the pixels within the bounding box region;
[0034] The physical width, physical height, and physical area are calculated based on the pixel width, pixel height, average depth value, and focal length parameter in the camera intrinsics.
[0035] Optionally, the process of processing the structured detection results using a multimodal large language model includes:
[0036] Based on the bounding box coordinates, the change type label, and the physical dimensions, construct the text prompt information;
[0037] The text prompt information and the current image with the bounding box drawn are input into the multimodal large language model;
[0038] Based on the processing of the multimodal large language model, natural language text containing risk level assessment and maintenance recommendations is output.
[0039] Optionally, after generating a maintenance work order, the process further includes performing a re-inspection process based on the maintenance work order. The re-inspection process includes: after the maintenance work order status changes to pending review, acquiring a re-inspection image based on a fixed camera; performing image registration, feature extraction, depth estimation, change detection, and physical size calculation steps sequentially based on the reference image and the re-inspection image; if no structural changes are detected near the bounding box coordinates, the maintenance is deemed valid; after the maintenance is deemed valid, the re-inspection image is updated to a new reference image.
[0040] This embodiment also provides a systematic method for detecting structural changes in construction site enclosures and for multimodal assisted handling, used to implement the method described above. The system includes:
[0041] The image acquisition module is used to acquire images of the construction site enclosure area based on deployed fixed cameras, and obtain the current image;
[0042] The image registration module is used to process the current image based on a pre-stored reference image when the fence is intact, using an image registration method to obtain the registered current image.
[0043] The feature extraction module is used to process the reference image and the registered current image using a feature extraction model to obtain the reference image feature map and the current image feature map.
[0044] The depth prediction module is used to process the current image feature map using a depth prediction model to obtain the current image depth map;
[0045] The change detection module is used to process the reference image feature map, the current image feature map, and the current image depth map using a change detection model to obtain the bounding box coordinates and change type label of the changed area of the fence.
[0046] The physical size calculation module is used to perform calculations based on the current image depth map, camera intrinsic parameters, and the bounding box coordinates to obtain the physical size of the changing region.
[0047] The result construction module is used to construct structured detection results based on the bounding box coordinates, the change type label, the physical size, and the current image;
[0048] The multimodal large language model module is used to process the structured detection results and output semantic explanations and maintenance suggestions.
[0049] The work order generation module is used to generate maintenance work orders based on the structured detection results and the semantic interpretation and maintenance suggestions;
[0050] The closed-loop management module is used to perform a re-inspection process based on the reference image and the new re-inspection image after maintenance, and to update the reference image according to the re-inspection results.
[0051] The present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.
[0052] Compared with the prior art, the present invention has the following advantages and technical effects:
[0053] The structural change detection and multimodal assisted handling method for construction site enclosures provided by this invention significantly improves the accuracy and semantic understanding capabilities of change detection by combining depth feature comparison with visual transformers and monocular depth estimation. It effectively distinguishes between real structural changes and interference from light and shadow, occlusion, etc., greatly reducing the false alarm rate. Simultaneously, this method innovatively converts the image change area into an approximate physical size, providing objective and quantifiable decision-making basis for safety risk assessment and maintenance resource scheduling, achieving a leap from qualitative observation to quantitative analysis. By introducing a multimodal large language model, the system can automatically generate professional and easy-to-understand semantic explanations and maintenance suggestions based on structured detection results, and integrate the detection, analysis, dispatch, and re-inspection processes into an automated closed loop. This greatly improves the intelligence level and response efficiency of construction safety management, reduces reliance on professional manual inspections, and has good practical value and economic benefits. Attached Figure Description
[0054] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0055] Figure 1 This is a schematic diagram of the system operation according to an embodiment of the present invention;
[0056] Figure 2 This is a flowchart illustrating the implementation of an embodiment of the present invention. Detailed Implementation
[0057] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0058] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0059] The technical problems solved by this invention are as follows: 1. How to automatically and accurately detect structural changes in construction site fencing and their approximate range under fixed high-definition network camera conditions, and preliminarily determine the type of changes in the construction site fencing (such as missing, damaged, tilted, or obstructed), reducing the pressure of manual inspection and lowering false alarms about fencing changes. 2. Existing multimodal large-scale models directly rely on images for construction safety judgments, which are easily affected by occlusion, lighting, and changes in viewing angle, resulting in output results lacking stability and engineering verifiability. This invention provides "engineering fact input" to the multimodal large-scale model through image change detection and physical quantity constraints, allowing the model to perform anomaly semantic interpretation and maintenance suggestion generation based on high-confidence engineering facts. This improves the generalization ability and event interpretation ability of cross-modal models in construction site fencing monitoring scenarios, supporting the formation of a closed-loop management process of "detection—dispatch—maintenance—re-inspection".
[0060] like Figure 1 As shown, this embodiment provides a system for detecting structural changes in construction site enclosures and for multimodal assisted handling, including:
[0061] The image acquisition module is used to acquire images of the construction site enclosure area based on deployed fixed cameras, and obtain the current image;
[0062] The image registration module is used to process the current image based on a pre-stored reference image when the fence is intact, using an image registration method to obtain the registered current image.
[0063] The feature extraction module is used to process the reference image and the registered current image using a feature extraction model to obtain the feature map of the reference image and the feature map of the current image.
[0064] The depth prediction module is used to process the current image feature map using a depth prediction model to obtain the current image depth map.
[0065] The change detection module is used to process the feature map of the reference image, the feature map of the current image, and the depth map of the current image using a change detection model to obtain the bounding box coordinates and change type labels of the changed area of the fence.
[0066] The physical size calculation module is used to perform calculations based on the current image depth map, camera intrinsic parameters, and bounding box coordinates to obtain the physical size of the changing region.
[0067] The results construction module is used to construct structured detection results based on bounding box coordinates, change type labels, physical dimensions, and the current image.
[0068] The multimodal large language model module is used to process structured detection results and output semantic interpretations and maintenance suggestions.
[0069] The work order generation module is used to generate maintenance work orders based on structured detection results, semantic interpretation, and maintenance suggestions.
[0070] The closed-loop management module is used to perform a re-inspection process based on the baseline reference image and the new re-inspection image after maintenance, and to update the baseline reference image according to the re-inspection results.
[0071] like Figure 2 As shown, this embodiment also provides a method for detecting structural changes in construction site enclosures and for multimodal assisted handling, including:
[0072] 1. System hardware and deployment architecture.
[0073] (1) Front-end acquisition equipment: high-definition network camera. One or more fixed cameras are deployed in each construction enclosure area. The resolution of the video and images acquired by a single camera is not less than 1920×1080. The cameras are fixed in the installation position and the acquired video and images can cover the target enclosure area.
[0074] (2) Backend computing resources: servers / computing units. Among them, deep neural network computing units are included.
[0075] Used to run the image change detection model model_1 and the visual language model model_2;
[0076] Model_1 consists of three modules: feature extraction (model_1.1), image depth estimation (model_1.2), and change region detection (model_1.3).
[0077] model_1.1 is a neural network based on the Visual Transformers (Vit) architecture. It is the backbone network of the entire model_1 and is used for feature extraction from the input image.
[0078] model_1.2 is a convolutional neural network and is the depth prediction and detection module of the entire model_1. It predicts the image depth based on the feature maps extracted from model_1.1.
[0079] model_1.3 is a convolutional neural network that is the change region detection module of the entire model_1. Based on the feature map of model_1.1 and the depth map of model_1.2, it outputs the coordinates of the change region (bounding box) and the fence change category.
[0080] model_2 is a multimodal large language neural network model (VLM) based on the Transformers architecture, which is mainly responsible for the auxiliary interpretation of changes in the enclosure structure and the generation of maintenance suggestions.
[0081] The parameters of the computing unit, such as CPU, video memory, RAM, and storage, need to be configured according to the number of cameras and the detection frequency on site.
[0082] The software system configured in the computing unit includes an image acquisition module, a feature extraction module, a change detection module, a depth estimation module, a VLM module, a work order management module, and a database.
[0083] (3) Network and management platform. The camera and computing node are connected via wired or wireless network; the computing node is connected to the construction unit's management platform, work order system, database, and mobile terminal notification system via intranet or secure extranet.
[0084] 2. Initialization phase: ROI calibration and reference image acquisition.
[0085] (1) ROI calibration of the construction site enclosure. After the system is deployed and the cameras are installed, the current camera image is displayed through the system's management interface. The administrator uses polygons to draw the "Region of Interest (ROI) of the construction site enclosure" on the system management interface. After the drawing is submitted, the system stores the ROI in the form of a binary mask as a spatial constraint for subsequent detection and filtering.
[0086] (2) Reference Image Acquisition. When the fence is in good condition, the system acquires one or more frames of images from each camera as the reference reference image I_ref for the area monitored by that camera, and updates the image periodically. Periodic image updates mean that if the system determines that the fence is intact and unchanged, the system will periodically acquire and update the reference reference image; if the system determines that the fence has changed, the system will stop periodic acquisition, and after the administrator confirms that the fence has been repaired, the system will resume periodic acquisition and update the reference reference image.
[0087] 3. Overview of the overall online monitoring process.
[0088] During system operation, the following steps are performed for each camera at fixed time intervals: S1: Periodically acquire the latest image I_cur of the monitored area and read the reference image I_ref; S2: Perform image registration between I_cur and I_ref to eliminate camera micro-shakes, and obtain I_cur' after I_cur registration; S3: Use the backbone network model_1.1, input I_ref to obtain the reference image feature map F_ref, and input I_cur' to obtain the current image feature map F_cur; S4: Use the image depth estimation model model_1.2, input F_cur to obtain the current image depth map D_cur; S5: Use the fence change detection model model_1.3, input F_ref, F_cur, and D_cur, first extract the change area from F_ref and F_cur. The domain image coordinates are bounding_box, and then the feature maps and depth maps of the changing regions F_cur and D_cur are extracted. After feature fusion, feature matching is performed in a vector library built based on historical data to obtain the type labels of the changing regions of the fence; S6: The physical width W_cur, height H_cur, and approximate area S_cur of the changing region are estimated using the depth map D_cur and camera intrinsic parameters; S7: Based on bounding_box, labels, W_cur, H_cur, and Area_cur, the structured detection results are constructed and output along with I_ref and I_cur' to the visual language model model_2; S8: model_2 outputs semantic interpretation and maintenance suggestions; S9: The results are used to generate work orders and dispatch maintenance, and a closed loop is formed through automatic re-inspection and manual confirmation; S10: After the re-inspection is passed, the reference image and status are updated, the work order status is set to completed, and the data is archived.
[0089] Furthermore, the image alignment process in S2 includes: to reduce interference from slight camera shake, mechanical offset, etc., on image change detection, I_ref is registered with the currently acquired I_cur: (1) the SIFT feature point detection algorithm is used to extract corresponding point pairs in the two images; (2) the RANSAC robust fitting method is used to estimate the 2D affine or homography matrix H: ;in These are the pixel coordinates in reference image I_ref. It is the pixel coordinates corresponding to the current image I_cur; (3) Use this matrix to perform geometric transformation on I_cur to obtain the latest image I_cur' of the aligned monitoring area, thereby realizing that the reference base image I_ref and the latest acquired image I_cur' are in the same image coordinate system, which is convenient for subsequent analysis.
[0090] Furthermore, existing change detection methods based on convolutional features mainly rely on local texture differences. Under conditions of lighting changes, shadows, temporary occlusion, or slight camera shake, their feature distribution will drift significantly, making it difficult to maintain the consistency of the overall structure of the fence in the feature space. A model trained based on a self-supervised Transformer architecture, such as DINOv3, has a global feature space that is "semantically consistent, structurally stable, and reusable across scenes." This ensures that when the fence structure does not change, its representation in the feature space remains stable. The feature extraction module in S3 includes: using model_1.1 to extract features from the image. (1) Input I_ref and I_cur' into model_1.1 respectively; (2) The model_1.1 encoder contains several layers of Transformer Blocks. The image features of lower layers focus on image texture, the image features of medium layers focus on image region-level structure, and the later layers focus more on the overall semantics. This invention selects one layer from each of the low, medium, and high-level Blocks, retains the tokens corresponding to the image patches, and rearranges them into two-dimensional feature maps F_cur and F_ref; assuming that the low, medium, and high levels selected from model_1.1 are L1, L2, and L3 layers respectively, the following results are obtained: ; ; ; and the corresponding ,in For the number of channels, The feature map space size. (3) Select high-level features ( , ) as the main input features for change detection. Select all features of the current image ( , , () as input features for depth prediction.
[0091] Furthermore, the depth estimation process in S4 uses model_1.2, which is input from S3. , , Output the depth map D_cur of the current image I_cur'. Used for subsequent estimation of the physical extent of the changed region and semantic input for change detection. The depth is expressed in meters after appropriate scale calibration. Used to exclude the discriminative constraint between foreground occlusion and changes in the enclosure body. If the depth constraint is not introduced, the change detection cannot distinguish between the missing enclosure and the foreground occlusion, resulting in the failure of the change category determination, and thus the multimodal model output does not have engineering credibility. Specifically, it includes the following steps: (1) Multi-layer feature channel alignment: where, With 64 output channels The convolution kernel is used to process the data to obtain a 64-channel feature map. . and Through 128 and 256 output channels respectively The convolution kernel is used to process the data to obtain the feature representation for the corresponding number of channels. , (2) U-Net cascaded decoding: for , , Progressive feature fusion is performed to obtain = , , .
[0092] The first step is to integrate the characteristics of high-level and mid-level management. Concatenate along the channel dimension, then use Convolutional blocks fuse channels and reduce dimensions to output fused features in the middle layer. : The second step is to incorporate low-level features, and... and Concatenate along the channel dimension, then use Convolutional blocks fuse channels and reduce dimensions to output fully fused features. : The third step, depth prediction, first involves two feature refinement convolutions. use Convolution dimensionality reduction outputs deep features : Finally, in deep features Use 1×1 convolution to output single-channel depth prediction The image is upsampled to the original image size and then subjected to ReLU mapping to obtain a non-negative depth map D_cur.
[0093] ;
[0094] ;
[0095] in , is the upsampling factor, using bilinear interpolation.
[0096] The training process includes: an image depth prediction model based on a convolutional neural network structure. During training, the model_1.1 feature extraction model is frozen, and only model_1.2 is trained. The model uses a dataset with ground truth depth values (including outdoor scenes and structures like fences, without the need for a specific dataset), L1 depth error + structural similarity (SSIM) as the loss function, and the objective function is to minimize the scale error of depth prediction for supervised training.
[0097] Furthermore, the change detection process in S5 includes: a method for determining structural changes in the feature space of Visual Transformers based on a reference frame. The change detection method proposed in this invention is divided into two stages: change region proposal and change category determination. Specifically, it includes the following steps:
[0098] (1) Proposal for change area.
[0099] Phase 1: Use model_1.1 to extract feature maps of the reference frame and the current frame. , By combining multi-index feature difference with high-frequency / edge enhancement and ROI constraint + connected component analysis, candidate boxes for regions with significant changes are generated.
[0100] 1) Multi-index feature difference: at each feature location Extract the feature vectors corresponding to the two frames:
[0101] ;
[0102] First, L2 normalization is performed to obtain the unit vector:
[0103] ;
[0104] Based on this, several change measures are constructed: (i) cosine distance (direction change): (ii) L2 difference norm (amplitude variation): Normalize the two difference plots above to the [0,1] interval using min–max normalization, and we get: Then, a significance plot of the basic changes is constructed using linear fusion: .in Furthermore, these values can be set to constants that sum to 1 (such as 0.6, 0.4) to pre-determine weights for engineering applications. This multi-index feature difference method considers both changes in feature direction and overall magnitude, which is beneficial for capturing various structural change patterns.
[0105] 2) High-frequency / edge enhancement: To further emphasize high-frequency variation details such as fine cracks and edge damage, the basic saliency map is enhanced. High-frequency / edge enhancement processing is performed. A LoG (Laplacian of Gaussian) filter is used for... Perform convolution: ;in This indicates a convolution operation using a 3×3 LoG kernel, without trainable parameters. This indicates taking the absolute value. Then, the high-frequency enhancement image E is normalized to [0,1]: Then, a weighted fusion is performed with the basic saliency map to obtain the saliency map of changes after detail enhancement: .in For preset weights (e.g.) This step is used to balance overall changes with local high-frequency changes. Through this step, the response to detailed changes such as cracks and gaps on the fence surface can be effectively enhanced, making the candidate area more closely resemble the actual damage contour in terms of spatial structure.
[0106] 3) ROI constraint + connected component morphology analysis to generate candidate boxes: This results in a saliency map with enhanced details. Then, candidate bounding boxes for changes are generated by combining the ROI mask of the construction site and the morphological analysis of connected components. (i) ROI constraint: The ROI mask of the construction site is downsampled to the feature map size to obtain... ,Then: This ensures that subsequent searches are conducted only within the fenced area. (ii) Upsampling and threshold segmentation: Upsampled to the original image resolution, resulting in Choose a threshold T and construct a binary transformation mask: (iii) Connectivity analysis and morphological filtering: Connectivity analysis is performed on the binary mask B to obtain several connected regions. For each connected component: calculate the minimum bounding rectangle of the set of pixels to obtain the candidate bounding box. Filtering is based on area thresholds and shape constraints: discard areas that are too small (less than a given pixel threshold); discard noise areas with extreme aspect ratios (too narrow or too flat) or that are elongated and thin.
[0107] Preserved This is the set of change candidate boxes output from stage one. These candidate boxes are concentrated in regions where there are significant changes in the image feature space and where high-frequency detail enhancement and ROI definition have been applied, providing highly informative input objects for the stage two classification submodule.
[0108] (2) Change category determination. The goal of Phase 2 is to determine the change category of each candidate box based on the feature map and depth estimation results, given the change candidate boxes.
[0109] 1) Input: Feature map of the current frame Current frame depth map The set of change candidate boxes output in stage one A database of fence change features built based on historical data.
[0110] 2) Feature cropping and fusion per bounding box: For each candidate bounding box Perform the following steps: (i) ROI feature alignment (feature map dimension): Map the bounding_box coordinates from the original image coordinates to the feature map coordinate system (scaled according to the patch downsampling factor P); Using RoIAlign to crop the corresponding region into a fixed-size feature block: ;in (ii) Depth submap cropping and scaling: Crop the corresponding depth block from the depth map D_cur according to the bounding_box region: .Will Scaled to the same spatial size as the ROI feature using bilinear interpolation. ,get: (iii) Feature and depth fusion: At the channel dimension... and To splice: .
[0111] This fusion feature simultaneously includes: high-level semantic features of the image (texture, shape, edges, etc.); and depth distribution information of the corresponding region, which can distinguish between foreground occlusion, surface changes of enclosures, and background exposure due to missing elements. These features serve as the comprehensive input features for the candidate region and are used for subsequent classification.
[0112] 3) Flattening and PCA Pre-Dimensionality Reduction: To facilitate supervised feature transformation and nearest neighbor retrieval, the fused features are first flattened and PCA pre-dimensionality reduced: (i) Flattening and Centering: Use the mean vector obtained during the offline training phase. Center the features: (ii) PCA pre-dimensionality reduction: In the offline phase, based on historical samples... Perform principal component analysis (PCA) to obtain the projection matrix: Where d is the target dimension (typically 1024 dimensions), used for dimensionality reduction while preserving as much variation information as possible. In the online phase, Projected onto PCA space: .
[0113] 4) Supervised Feature Transformation: Based on PCA pre-dimensionality reduction, a linear discriminant transformation trained using historical label information is introduced to make different transformation types more separable in the embedding space. (i) Form of Linear Discriminant Transformation. Define a linear transformation matrix: ;in This refers to the transformed feature dimension. (For) Perform a linear transformation to obtain the supervised embedding vector: (ii) Offline training method: In the historical data, for each sample j, there is a corresponding feature vector. and change type tags ;Supervisory feature transformation matrix This can be learned through linear metric learning methods (linear mappings based on Center Loss and Triplet Loss): the optimization objective is to shorten the distance between samples of the same class and widen the distance between samples of different classes in the transformation space. Formally, this can be represented as learning a matrix. , making .for Smaller, for The value is relatively large. After training, a fixed supervised feature transformation matrix is obtained. This is used to map PCA features to a more discriminative k-dimensional embedding space.
[0114] 5) Feature Library Construction (Offline Phase): During system deployment or training, a feature library of change types is constructed based on historical data. (i) Sample Preparation: Select several change area samples from historical monitoring data and label each sample with its change type (e.g., missing, damaged, tilted, occluded, slight change, etc.). (ii) Feature Generation: For each historical sample, perform the same process as in the online phase: feature / deep fusion to obtain... Flattening + PCA projection yields The supervised feature transformation yields the embedding vector. (iii) Data Entry: The embedding vectors of all historical samples and their labels are combined to form a feature library. Cluster or average the embedding vectors of each class of samples to generate a small number of class prototype vectors to reduce storage and retrieval overhead.
[0115] 6) Nearest neighbor retrieval and change type determination during online monitoring: During online monitoring, for each candidate change box... The resulting embedding vector Utilizing feature libraries Perform nearest neighbor retrieval to determine the type of change. (i) Similarity calculation: For each sample j in the database, calculate the similarity with the current embedding. The similarity is calculated using cosine similarity: (ii) Nearest Neighbor or Top-K Voting: Nearest Neighbor Strategy: .Will as candidate boxes Predicting the type of change; Top-K voting strategy: selecting the K most similar historical sample indices to form a set. Weighted voting is applied to the tags within the set (weighted by similarity): . This is the category label for the area where the fence changes (labels), and the Top-K average similarity is used as the prediction confidence score (labels_conf).
[0116] Furthermore, S6 uses depth + intrinsic parameters to approximate the physical area estimation of changing bounding boxes. Area estimation is performed for regions identified as having missing, damaged, tilted, or deformed fencing. Given: camera intrinsic parameter matrix. : Current frame depth map The depth unit is meters; a filtered and retained bounding box with certain variations: Assuming the entire enclosure lies on an approximately vertical plane, this invention uses the average depth of the bounding_box to approximate the distance to this local plane.
[0117] A. Mean depth estimation: ;in This is the set of pixel coordinates within the bounding box.
[0118] B. Pixel width and height: .
[0119] C. Physical width and height approximation: The physical length corresponding to one pixel horizontally at depth Z is approximately... ,therefore: One vertical pixel corresponds to a physical length of... ,therefore: .
[0120] D. Approximate estimation of the area of the changed fence zone: The area S_cur is an approximate physical area (square meters) in the camera coordinate system, used to quantify the "approximate scale" of the changing area, supporting risk level assessment and maintenance priority ranking, and is not used as a precise engineering measurement value.
[0121] Furthermore, in S7 and S8, multimodal large model (VLM) auxiliary interpretation and maintenance suggestion generation are performed. The specific process includes: (1) Structured input construction: construct structured data for VLM containing the following: 1) Camera ID, timestamp; 2) Each change region: pixel coordinates bounding_box; change type labels (missing / damaged / tilted / occluded, etc.); change confidence labels_conf; estimate approximate size W_cur, H_cur, S_cur; 3) Current image, and draw the change bounding_box and short labels on it. (2) Prompt design: convert structured information into natural language prompts, input the prompt text and image into VLM, and let VLM execute the output. The natural language prompts given to VLM are as follows: "You are a construction site fence safety inspection assistant. Below are monitoring images of construction site fences, along with fence anomaly information detected by the model: Area 1: Right center, marked with a red rectangle in the image, type 'Fence Missing', estimated missing area approximately 1.2 square meters, missing area height approximately 2 meters, width approximately 0.6 meters, model confidence level 0.92. Based on the image content and the above detection results, please explain the potential safety risks posed by this anomaly and generate a human-readable risk assessment and maintenance recommendations for safety management."
[0122] (3) VLM output: 1) Natural language description of the anomaly type (e.g., "There is a clear gap in the right fence, and the construction area behind can be seen"); 2) Assess the safety risk level (e.g., "High risk: pedestrians may accidentally enter the construction area"); 3) Provide maintenance suggestions (e.g., "It is recommended to set up temporary fences immediately and replace the boards within 2 hours").
[0123] Furthermore, the closed-loop process of work order dispatch and automatic review in S9 and S10 includes:
[0124] (1) Work order generation and dispatch: The system integrates VLM output and structured detection results to generate maintenance work orders: 1) Site information; 2) Camera location and description; 3) Abnormal area location and type; 4) Variation range and area estimation; 5) Risk assessment and suggested processing time limit; 6) The work order is sent to the construction unit management platform and the mobile terminal of the responsible personnel through the interface to realize automatic dispatch.
[0125] (2) Repair reporting: After the repair personnel arrive at the site to handle the problem, they mark the work order system as "repaired" and can upload "before and after comparison photos" or videos of the site.
[0126] (3) Automatic re-inspection and review: After the work order status changes to "pending review", the system: 1) captures one or more new images from the corresponding camera in a short time; 2) runs the entire change detection pipeline (S3-S7) again; 3) if no structural changes are detected near the original abnormal area and the change intensity has obviously fallen back to the normal range, the repair is automatically determined to be effective; 4) otherwise it is marked as "re-inspection failed" and pushed to the management personnel for manual review; 5) the images before and after the repair and the detection results are sent back to VLM, prompting it to judge "whether the fence has been restored to the normal safe state" as an auxiliary reference.
[0127] (4) Reference image update and data archiving: 1) When the fence has been continuously tested without any abnormalities for a period of time and the maintenance work order has been approved, a certain latest frame can be set as the new reference image I_ref to update the subsequent comparison benchmark; 2) Store all detection results, work order processing records and automatic review results in the database for statistical analysis and model optimization.
[0128] This embodiment also provides a storage medium on which a computer program is stored, which, when executed by a processor, implements the method described thereon.
[0129] This embodiment has the following technical effects:
[0130] 1. High detection accuracy and semantic understanding depth: By improving change detection based on the Visual Transformers backbone network, the deep features of the current frame and historical reference frames are directly compared, which can accurately identify real structural changes (such as missing or damaged parts) at the semantic level. This effectively overcomes interference from changes in lighting and temporary occlusion, and significantly reduces the false alarm rate.
[0131] 2. Provides quantifiable engineering decision-making basis: It innovatively combines monocular depth estimation with change detection to output the approximate physical size of the changed area, providing objective and quantifiable data support for risk assessment and maintenance resource scheduling, surpassing the traditional method that only provides image location.
[0132] 3. High level of intelligence and closed-loop management: The structured inspection results are professionally interpreted through the multimodal large model (VLM) to generate easy-to-understand maintenance suggestions. It also realizes the full-process automated closed-loop management from automatic detection, intelligent analysis, work order dispatch to maintenance and inspection, which greatly improves the efficiency and intelligence level of construction safety management.
[0133] 4. Excellent overall cost-effectiveness: It makes full use of existing surveillance cameras and general computing platforms, avoiding the deployment and maintenance costs of a large number of dedicated sensors. The system is flexible in deployment, highly scalable, and has significant economic and social benefits.
[0134] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for detecting structural changes in construction site enclosures and for multimodal assisted handling, characterized in that, Includes the following steps: Images of the construction enclosure area are acquired to obtain the current image. The current image and the reference image are then processed using an image registration method to obtain the registered current image. Based on the reference image and the registered current image, a feature extraction model is used to process the image to obtain the reference image feature map and the current image feature map. Based on the current image feature map, a depth prediction model is used to process it to obtain the current image depth map; A change detection model is used to process the feature map of the reference image, the feature map of the current image, and the depth map of the current image to obtain the bounding box coordinates and change type labels of the changed area of the enclosure; The physical dimensions of the changing region are obtained by calculating and processing the current image depth map, camera intrinsic parameters, and bounding box coordinates. Based on the bounding box coordinates, the change type label, the physical dimensions, and the current image, a structured detection result is constructed; Based on the structured detection results, a multimodal large language model is used for processing to output semantic explanations and maintenance suggestions; Based on the structured detection results and the semantic interpretation and maintenance suggestions, a maintenance work order is generated; After generating a repair work order, the process also includes a re-inspection process based on the repair work order. The re-inspection process includes: after the repair work order status changes to pending review, a re-inspection image is acquired using a fixed camera; based on the reference image and the re-inspection image, the steps of image registration, feature extraction, depth estimation, change detection, and physical size calculation are executed sequentially; if no structural changes are detected near the bounding box coordinates, the repair is deemed valid; after the repair is deemed valid, the re-inspection image is updated to a new reference image. The process of processing the current image and the reference image using an image registration method includes: Based on the reference image and the current image, the SIFT feature point detection algorithm is used to process and extract corresponding feature point pairs. Based on the corresponding feature point pairs, the RANSAC robust fitting method is used to estimate the two-dimensional homography matrix. Based on the two-dimensional homography matrix and the current image, a geometric transformation is performed to obtain the registered current image; The process of processing the reference image feature map, the current image feature map, and the current image depth map using a change detection model includes: In the change region proposal stage, based on the feature map of the reference image and the feature map of the current image, the feature difference is calculated, and the high-frequency enhancement and the enclosure interest region mask are combined for processing to generate candidate bounding boxes for change regions. In the change category determination stage, feature and depth fusion is performed based on the current image feature map, the current image depth map, and the candidate bounding boxes of the change region, and the change type label is obtained by retrieving a pre-built feature library. The process of calculating and processing the current image depth map, camera intrinsic parameters, and bounding box coordinates to obtain the physical dimensions of the changed region includes: Calculate the pixel width and pixel height based on the bounding box coordinates; Based on the current image depth map and the bounding box coordinates, calculate the average depth value of the pixels within the bounding box region; Based on the pixel width, the pixel height, the average depth value, and the focal length parameter in the camera intrinsics, calculate the physical width, physical height, and physical area; Based on the structured detection results, the process of processing using a multimodal large language model includes: Based on the bounding box coordinates, the change type label, and the physical dimensions, construct the text prompt information; The text prompt information and the current image with the bounding box drawn are input into the multimodal large language model; Based on the processing of the multimodal large language model, natural language text containing risk level assessment and maintenance recommendations is output.
2. The method for detecting structural changes in construction site enclosures and providing multimodal assisted treatment according to claim 1, characterized in that, The process of calculating feature differences includes: The feature vectors in the reference image feature map and the corresponding feature vectors in the current image feature map are subjected to L2 normalization to obtain normalized feature vectors. Based on the normalized feature vector, calculate the cosine distance and L2 difference norm; After normalizing the cosine distance plot and the L2 difference norm plot respectively, they are linearly fused to obtain the significance plot of basic changes.
3. The method for detecting structural changes in construction site enclosures and providing multimodal assisted treatment according to claim 1, characterized in that, The process of building a feature library includes: Based on images of historical fence change samples, feature extraction, depth estimation, and change region proposal steps are performed to obtain candidate regions. Based on the current image feature map and the current image depth map of the candidate region, feature and depth fusion processing is performed to obtain fused features; The fusion features are flattened and subjected to dimensionality reduction based on principal component analysis to obtain dimensionality-reduced features; The reduced-dimensional features are subjected to supervised feature transformation based on linear discriminant transformation to obtain the embedding vector; The embedded vectors and their corresponding variation type labels are stored to form a feature library.
4. A systematic method for detecting structural changes in construction site enclosures and for multimodal assisted treatment, characterized in that, The system for implementing the method of claim 1, wherein the system comprises: The image acquisition module is used to acquire images of the construction site enclosure area based on deployed fixed cameras, and obtain the current image; The image registration module is used to process the current image based on a pre-stored reference image when the fence is intact, using an image registration method to obtain the registered current image. The feature extraction module is used to process the reference image and the registered current image using a feature extraction model to obtain the reference image feature map and the current image feature map. The depth prediction module is used to process the current image feature map using a depth prediction model to obtain the current image depth map; The change detection module is used to process the reference image feature map, the current image feature map, and the current image depth map using a change detection model to obtain the bounding box coordinates and change type label of the changed area of the fence. The physical size calculation module is used to perform calculations based on the current image depth map, camera intrinsic parameters, and the bounding box coordinates to obtain the physical size of the changing region. The result construction module is used to construct structured detection results based on the bounding box coordinates, the change type label, the physical size, and the current image; The multimodal large language model module is used to process the structured detection results and output semantic explanations and maintenance suggestions. The work order generation module is used to generate maintenance work orders based on the structured detection results and the semantic interpretation and maintenance suggestions; The closed-loop management module is used to perform a re-inspection process based on the reference image and the new re-inspection image after maintenance, and to update the reference image according to the re-inspection results.
5. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in claim 1.
Citation Information
Patent Citations
Image-based change detection method
CN112365462A
Unmanned aerial vehicle remote sensing image change detection method and system
CN121438131A