Construction hidden danger automatic detection method and system based on point cloud and image multi-mode

By fusing point cloud and RGB image multimodal data, and utilizing an improved YOLOv5 and 3D-CNN model, safety hazard detection at construction sites is achieved. This solves the problems of missed detection and misjudgment in traditional monitoring methods, and realizes high-precision, real-time safety hazard analysis and three-dimensional visualization.

CN120823552APending Publication Date: 2025-10-21CHINA CONSTR THIRD ENG BUREAU GRP SOUTH CHINA CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510710081.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Traditional construction site safety monitoring methods cannot achieve all-weather and full-area coverage, and there are risks of missed detections and misjudgments. A single sensor cannot meet the needs of high-precision safety monitoring in complex environments.

Method used

By adopting multimodal data fusion of point cloud and RGB image, the improved YOLOv5 model is used to detect safety equipment, combined with the 3D-CNN model to judge abnormal personnel behavior, combined with the semantic segmentation model to calculate the height of the guardrail and coordinate conversion, and the three-dimensional real scene model is generated by using point cloud feature matching and Poisson surface reconstruction.

Benefits of technology

It realizes the rapid qualitative and quantitative analysis of safety hazards at construction sites, improves the accuracy and real-time performance of detection, provides an intuitive three-dimensional real-scene model, and enhances the intelligence and precision of construction safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823552A_ABST
    Figure CN120823552A_ABST
Patent Text Reader

Abstract

The invention discloses a construction hidden danger automatic detection method and system based on point cloud and image multi-mode. The method comprises the following steps: step 1, acquiring point cloud data and RGB image data and respectively preprocessing the point cloud data and the RGB image data; 2, safety equipment detection is carried out by establishing an improved YOLOv5 model; establishing a 3D-CNN model to qualitatively judge the abnormal behavior of the personnel; 3, combining a semantic segmentation model and point cloud data processing to carry out quantitative judgment of height calculation and coordinate conversion of the fence; and 4, realizing three-dimensional live-action model reconstruction based on point cloud feature matching, Poisson surface reconstruction and GPU rendering, and detecting potential safety hazards of a construction site. The accuracy of qualitative detection is further improved, the precision error of quantitative detection is reduced, the recognition efficiency and precision of potential safety hazards are greatly improved, intelligent and all-around guarantee is provided for safety management of a construction site, remote decision is assisted, the management cost is reduced, and the method is suitable for various construction scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent construction and building safety monitoring, and specifically relates to a method and system for automatically detecting construction hazards based on point cloud and multimodal image analysis. Background Art

[0002] The construction industry is a high-risk industry, with numerous safety hazards present on construction sites, including falls from heights, impacts from objects, and mechanical injuries. Furthermore, the complex and ever-changing working environment (such as deep foundation pits, confined spaces, and cross-construction) makes it difficult for traditional manual inspections to achieve all-weather, all-area coverage, and they are prone to missed inspections due to subjective negligence or environmental interference. As projects scale and process complexity increase, monitoring systems that rely solely on manual experience or a single sensor (such as cameras or lidar) are no longer sufficient. Visual technology is significantly affected by light and dust, and radar technology lacks semantic understanding capabilities, both of which can lead to misjudgments or omissions of key risks (such as failure to wear a seatbelt or deformation of the surrounding structure). Furthermore, intelligent resolution capabilities are limited.

[0003] Existing construction safety monitoring methods, such as a construction site safety monitoring system based on BIM technology (Application Number: CN202410086861X), include a data acquisition module, a data processing module, and a BIM data processing module. The data acquisition module is used to collect RGB images of the construction site; the data processing module is used to optimize the RGB images of the construction site; and the BIM information processing module is used to convert the optimized RGB images of the construction site into a three-dimensional monitoring model to determine abnormal conditions at the construction site. However, this technology does not integrate multimodal data for comprehensive judgment, the degree of intelligent analysis dataization is not high, there is no precise qualitative analysis, the amount of useless data processed is large, the system response is slow, and it is prone to missed and misjudgment. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of the existing technology and propose an automatic detection method and system for construction hazards based on point cloud and image multimodality. With the synergistic advantages of point cloud and RGB image, it can quickly identify safety hazards qualitatively and accurately analyze them quantitatively, while providing an intuitive three-dimensional real-scene model for on-site management, thereby comprehensively ensuring construction safety and management efficiency.

[0005] The present invention provides a method for automatically detecting construction hazards based on point cloud and image multimodality, comprising:

[0006] Step 1: Obtain point cloud data and RGB image data and preprocess them separately;

[0007] Step 2: Build an improved YOLOv5 model to detect safety equipment and build a 3D-CNN model to qualitatively judge abnormal human behavior.

[0008] Step 3: Combine the semantic segmentation model and point cloud data processing to perform quantitative judgment on guardrail height calculation and coordinate transformation;

[0009] Step 4: Reconstruct the 3D real-scene model based on point cloud feature matching, Poisson surface reconstruction, and GPU rendering to detect safety hazards at the construction site.

[0010] Furthermore, in step 1: statistical filtering, dynamic downsampling adaptive adjustment and ground segmentation are sequentially performed on the point cloud data;

[0011] Gaussian filtering is performed on the RGB image data to remove noise, and then adaptive histogram equalization is performed to enhance the image contrast.

[0012] Furthermore, in step 2: the improved YOLOv5 model adopts the CBAM attention mechanism module to be embedded in the Backbone network, introduces the cross-modal feature pyramid (CMFPN) in the Neck layer, and jointly screens key features through channel attention and spatial attention to enhance adaptability to the dynamic environment of the construction site; the image feature input is the feature map extracted by CSPDarknet53; the point cloud feature input is the concatenation vector of the normal vector histogram and the spatial coordinate feature;

[0013] The weighted fusion weight α of the cross-modal feature pyramid (CMFPN) is calculated by the normal vector direction cosine:

[0014]

[0015] Where θ is the angle between the point cloud normal vector and the horizontal plane, α∈[0.6,0.8].

[0016] The 3D-CNN model adopts the spatiotemporal pyramid pooling (SPP) module, takes RGB images as input, and detects abnormal behaviors.

[0017] Furthermore, in step 3: the edge enclosing area is located by using the semantic segmentation model, and the plane equation Ax+By+Cz+D=0 is fitted after extracting the mask;

[0018] Calculate the maximum vertical distance difference h=|p in the point cloud max -p min |, compared with the safety standard threshold h s ,when The coordinate transformation is optimized using the ICP algorithm.

[0019] Furthermore, in step 4: the point cloud feature matching adopts the FPFH algorithm;

[0020] The Poisson surface reconstruction algorithm adaptively adjusts the octree depth according to the point cloud density and performs a smoothing process after generating a mesh.

[0021] Furthermore, the improved YOLOv5 model adds an attention mechanism module (CBAM), and its attention weight calculation function is

[0022]

[0023] Where Q, K, and V are query vectors, key vectors, and value vectors respectively; d k For the key vector dimension, use transfer learning to load the ImageNet pre-trained weights and train on the self-built dataset with the learning rate.

[0024] The present invention also provides an automatic detection system for construction hazards based on point cloud and image multimodality, which includes the following modules:

[0025] Data acquisition module, used to collect point cloud data and RGB image data of the construction site;

[0026] A data processing module, comprising a noise filtering unit, a dynamic downsampling unit, a ground segmentation unit, and an image fusion unit connected in sequence; and used to process the point cloud data and RGB image data collected at the construction site;

[0027] The qualitative detection module includes a safety equipment detection unit based on an improved YOLOv5 model and an abnormal behavior detection unit based on a 3D-CNN model. It is used to detect safety equipment by establishing an improved YOLOv5 model and to qualitatively judge abnormal human behavior by establishing a 3D-CNN model.

[0028] Quantitative analysis module, including edge protection height detection unit and position accuracy assurance unit;

[0029] The visualization display and early warning platform module includes a point cloud registration unit, a mesh reconstruction unit, a real-time rendering unit, and an early warning unit.

[0030] Furthermore, in the qualitative detection module, the improved YOLOv5 model adopts the CBAM attention mechanism module to be embedded in the Backbone network, introduces the cross-modal feature pyramid (CMFPN) in the Neck layer, and jointly screens key features through channel attention and spatial attention to enhance adaptability to the dynamic environment of the construction site;

[0031] The 3D-CNN model adopts the spatiotemporal pyramid pooling (SPP) module, takes RGB images as input, and detects abnormal behaviors.

[0032] Furthermore, the system uses the processed point cloud data and RGB image data to reconstruct a three-dimensional real-scene model of the construction site, uses a point cloud-based three-dimensional reconstruction algorithm to generate a three-dimensional mesh model, and maps the texture information of the RGB image to the three-dimensional model, making the model more realistic and intuitive.

[0033] Furthermore, the identified safety hazards are marked on the three-dimensional real-scene model. The marked information includes the type, location, and severity of the safety hazard, which makes it easier for management personnel to intuitively understand the safety status of the construction site.

[0034] Compared with the prior art, the automatic detection method and system of the present invention have the following technical advantages:

[0035] 1. Use deep fusion of multimodal data to break through the limitations of a single sensor. Through collaborative perception of LiDAR point clouds (construction site geometric information) and RGB images, that is, through computer vision technologies such as semantic segmentation, image pixels or regions are given clear "semantic labels", overcoming the technical deficiencies of traditional detection solutions, such as visual technology being affected by light / dust interference and radar technology lacking semantic understanding. This invention focuses on utilizing multimodal data fusion of point clouds and RGB images to create an intelligent system that can automatically and accurately identify safety hazards on construction sites and achieve multi-faceted judgment.

[0036] 2. Leveraging the technical advantages of the improved YOLOv5 model and the attention mechanism. First, the CBAM attention mechanism module is embedded in the Backbone network of the YOLOv5 model. Through the joint screening of key features by channel attention and spatial attention, the adaptability to the dynamic environment of the construction site is enhanced. Secondly, Ghost convolution is introduced in the Neck layer of YOLOv5 to reduce the amount of computation by cheap feature reuse. At the same time, the number of feature channels is increased, and spatiotemporal pyramid pooling (SPP) is combined to capture contextual information at different scales, optimizing and improving the detection capability of small targets. Finally, a hybrid scaling strategy is adopted to enhance the feature extraction capability while maintaining the lightweight model. The input resolution is increased from 640×640 to 896×896, and the voxel side length is adjusted by dynamic downsampling. The lightweight design of the weighted bidirectional feature pyramid network (BiFPN) is used to integrate multi-scale features to effectively balance the real-time performance and accuracy of detection. The system can accurately identify the irregular placement of equipment on construction sites, the positions of fences, safety nets, scaffolding, etc., and after frame extraction and denoising of RGB images, use deep learning to identify the wearing of workers' safety equipment, the construction movements of personnel and machines, etc. The solution of the present invention can detect in real time, intelligently judge risks, and give safety reminders. Its rapid response and accurate judgment capabilities are much higher than those of existing technical solutions.

[0037] The solution of the present invention uses artificial intelligence algorithms to make qualitative and quantitative judgments. The qualitative judgment includes safety equipment monitoring and abnormal behavior detection; the quantitative judgment includes edge protection detection and position accuracy calibration, and comprehensively and intelligently checks for safety hazards.

[0038] 3. Utilizing the technical advantages of a 3D-CNN model and multi-scale spatiotemporal feature extraction, the system employs a sparse spatiotemporal sampling strategy. During construction site videography, invalid frames are automatically skipped during occasional occlusions, retaining only valid action segments. This system also incorporates a hybrid LSTM + Transformer temporal model, leveraging a self-attention mechanism to complete the action logic in occluded areas. A cross-modal feature interaction module is added before the final fully connected layer of the 3D-CNN model. This dynamically adjusts feature contributions through attention weights, improving the accuracy of identifying and analyzing safety hazards.

[0039] 4. Deep collaborative optimization of point cloud registration and semantic segmentation enables the system to achieve millimeter-level accuracy. First, the FPFH (Fast Point Feature Histogram) algorithm is used to extract local features of the point cloud, and the ICP (Iterative Closest Point) algorithm is combined to perform global optimization of the multi-view point cloud. Through iterative convergence of feature matching and rigid body transformation, the registration error is controlled to 0.001m. At the same time, multi-scale feature fusion of RGB images is performed based on the DeepLabv3+ semantic segmentation model to accurately locate key areas such as guardrails and scaffolding and extract their geometric contours. Finally, the semantic segmentation results are spatially aligned with the point cloud data, and the Kalman filter is used to dynamically compensate for the sensor timing error. Combined with bilateral filtering to suppress local noise, the measurement accuracy of guardrail height and other measurements is much higher than that of existing technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a flow chart of the method for automatically detecting construction hazards based on point cloud and image multimodality of the present invention;

[0041] Figure 2 This is an architecture diagram of the automatic detection system for construction hazards based on point cloud and image multimodality of the present invention. DETAILED DESCRIPTION

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0043] Embodiment 1 of the present invention provides a method for automatically detecting construction hazards based on point cloud and image multimodality, including:

[0044] Step 1: Obtain point cloud data and RGB image data and preprocess them separately.

[0045] Step 2: Build an improved YOLOv5 model to detect safety equipment and build a 3D-CNN model to make qualitative judgments on abnormal human behavior.

[0046] The improved YOLOv5 model uses the CBAM attention mechanism module embedded in the Backbone network and introduces a cross-modal feature pyramid (CMFPN) in the Neck layer. It uses channel attention and spatial attention to jointly filter key features to enhance its adaptability to the dynamic environment of the construction site.

[0047] The 3D-CNN model adopts the spatiotemporal pyramid pooling (SPP) module, takes RGB images as input, and detects abnormal behaviors.

[0048] Step 3: Combine the semantic segmentation model and point cloud data processing to perform quantitative judgment of guardrail height calculation and coordinate transformation.

[0049] The semantic segmentation model is used to locate the edge enclosure area, extract the mask and fit the plane equation Ax+By+Cz+D=0.

[0050] Where A, B, and C are the X, Y, and Z (vertical) components of the plane normal vector, respectively, and D is the offset from the plane to the origin.

[0051] Calculate the maximum vertical distance difference h=|p in the point cloud max -p min |, compared with the safety standard threshold h s ,when When the alarm is triggered;

[0052] in

[0053] The coordinate transformation is optimized using the ICP algorithm, and the convergence threshold is set to 0.001m.

[0054] Step 4: Reconstruct the 3D real-scene model based on point cloud feature matching, Poisson surface reconstruction, and GPU rendering to detect safety hazards at the construction site.

[0055] The point cloud feature matching adopts the FPFH algorithm; the Poisson surface reconstruction algorithm adaptively adjusts the octree depth according to the point cloud density, and performs a Laplacian smoothing process after generating a mesh.

[0056] Specifically, such as Figure 1 As shown, the second embodiment of the present invention provides an automatic detection method for construction hazards based on point cloud and image multimodality, which includes the following process.

[0057] Step S1: Perform multimodal data collection including LiDAR group deployment, point cloud data collection, and RGB image data collection.

[0058] The deployment of the laser radar group and the collection of point cloud data are carried out according to the construction site area S and the complexity C. The complexity C can be quantitatively set by factors such as the type of construction technology and the number of building floors. For example, when S>5000m2 and C>0.8, ≥8 laser radars are arranged along the outer contour of the construction area and key internal points (such as tower cranes and unloading platforms), distributed at heights such as tower cranes and the four corners of buildings to ensure that the point cloud covers the entire scene. The radar installation height is ≥5m, the acquisition frequency is set to 10-20Hz, and the point cloud data is transmitted to the data processing center in real time via wired Ethernet or wireless network.

[0059] The RGB image data collection is divided into monitoring grids according to the construction area, and every 100-200m 2 Set up a high-definition camera and encrypt the key areas to 100m 2 / unit; the camera is installed at a height of 2.5-3.5m, with an acquisition frame rate of 25-30 frames per second, and uses H.265 encoding to transmit video streams to edge computing nodes or data processing centers.

[0060] Step S2: Data preprocessing includes point cloud data preprocessing and RGB image preprocessing.

[0061] Furthermore, point cloud data preprocessing is implemented through noise filtering, dynamic downsampling, ground segmentation and image fusion.

[0062] To filter the noise of point cloud data, a statistical filtering algorithm is used. The number of neighborhood points k (usually 16-32) and the distance threshold d (such as 0.1-0.3m) are set, and the average distance from a point to its neighborhood points is calculated. If the mean distance of a point is greater than d, it is judged as a noise point and removed, thereby removing about 75% of the noise interference.

[0063] Dynamic downsampling uses voxel grid downsampling and adjusts the voxel side length according to the scene point cloud density using the following formula:

[0064]

[0065] where l min 、l max where ρ represents the minimum and maximum voxel edge lengths, ρ represents the point cloud density, ρ0 represents the density threshold, and k is a scaling factor that determines how steeply voxel edge lengths change with density. This approach dynamically adjusts voxel edge lengths based on point cloud density ρ, compressing data in dense areas (where ρ is high) while preserving detail in sparse areas (where ρ is low), achieving adaptive adjustment.

[0066] For ground segmentation, the Random Sample Consensus Algorithm (RANSAC) is used to fit the ground plane model, and the maximum number of iterations N (e.g., 500-1000 times) and the distance threshold (0.05-0.1m) are set to separate the ground and non-ground point clouds for subsequent target extraction.

[0067] Image fusion, extract feature points of point cloud and RGB image, use SURF feature descriptor matching, and achieve preliminary fusion by calculating the transformation matrix H to enhance the richness of scene information.

[0068] Furthermore, for RGB image preprocessing, key frames of the video stream are captured at fixed time intervals (e.g., 5 seconds / frame), and Gaussian filtering (σ value is 0.5-1.5) is performed on the RGB image data to remove noise. Then, adaptive histogram equalization is used to enhance the contrast. The image is scaled to 416×416 pixels and divided into training set / validation set / test set (ratio 8:1:1).

[0069] Optimally, in preparation for deep learning recognition, the preprocessed images are scaled to [W, H], proportionally divided into training, validation, and test sets, annotated with worker safety equipment information, and fed into the improved YOLOv5 model. In the model's backbone network, some convolutional layers are replaced with Ghost convolutions to reduce computational effort and improve training and recognition efficiency.

[0070] Step S3: Build an improved YOLOv5 model to detect safety equipment and build a 3D-CNN model to qualitatively identify abnormal personnel behavior. Qualitative safety hazard analysis includes both safety equipment detection and abnormal personnel behavior detection.

[0071] Preferably, the safety equipment detection inputs the preprocessed RGB image into the improved YOLOv5 model, which includes an attention mechanism module, and the real-time image is preprocessed and input into the model, and the model outputs the target category probability p i With the bounding box coordinates (x i ,y i ,w i ,h i ), after the non-maximum suppression (NMS) algorithm, the intersection-over-union ratio threshold IoU is set th Screening: If it detects behaviors such as not wearing a helmet / safety belt, it will trigger an audible and visual alarm and record the time and space information in the database.

[0072] As an implementation method, safety equipment detection includes model building and training, and real-time detection.

[0073] The model construction and training are based on the improved YOLOv5 model, adding the attention mechanism module (CBAM), and its attention weight calculation function is

[0074]

[0075] Where Q, K, and V are query vectors, key vectors, and value vectors respectively; d k The key vector dimension is usually set to 64. Use transfer learning to load the ImageNet pre-trained weights and train on a self-built dataset with a learning rate (initial 0.001, adjusted according to the cosine annealing strategy).

[0076] In the real-time detection, the real-time image is pre-processed and input into the model, and the model outputs the target category probability p i With the bounding box coordinates, the non-maximum suppression (NMS) algorithm is used to set the intersection-over-union threshold IoU th (e.g. 0.5) screening, identifying behaviors such as not wearing a safety helmet / safety belt, triggering an alarm, and recording abnormal information to the database, including timestamp t, location coordinates (x, y), etc.

[0077] Preferably, the abnormal behavior detection of personnel includes behavior feature learning and real-time behavior detection.

[0078] The behavioral feature learning method collects n sets of abnormal behavior video data, divides the video into segments of duration T, annotates key action frames, and trains a 3D-CNN model to learn the spatiotemporal characteristics of the actions. The 3D-CNN model uses a spatiotemporal pyramid pooling (SPP) module and inputs continuous RGB images to detect abnormal behaviors.

[0079] The real-time behavior detection uses continuous camera frames as input into the model. When the model outputs an abnormal behavior confidence level conf>0.8, an alarm is triggered, and surrounding cameras are linked for tracking. The behavior category and the location of the person involved are output, the trajectory of the person involved is locked, and pushed to the management terminal.

[0080] Step S4, quantitative analysis of safety hazards includes edge guard height detection and position accuracy calibration.

[0081] The edge guardrail height detection uses the DeepLabv3+ model to segment the guardrail area in the RGB image, extract the corresponding point cloud data, fit the guardrail plane equation Ax+By+Cz+D=0, and calculate the highest point P max With the lowest point P min The vertical distance h=|p max -p min |, compared with the safety standard threshold h s ,when The alarm is triggered.

[0082] As a specific implementation method, in the calculation of the guardrail height, the semantic segmentation model adopts the DeepLabv3+ architecture, and the expansion rate of the ASPP module is set to 6, 12, and 18; the plane fitting error tolerance is set to 0.05m, and a secondary check is triggered when the maximum distance from the point cloud to the fitting plane is greater than 0.05m.

[0083] The position accuracy calibration uses the camera calibration parameters to convert the image pixel coordinates (u, v) into the camera coordinate system coordinates (X c ,Y c ,Z c ), and then projected to the point cloud coordinate system through the extrinsic parameter matrices R and T. The ICP algorithm is used to iteratively optimize the point cloud alignment accuracy (convergence threshold ε = 0.001m), combined with the Kalman filter to suppress dynamic errors.

[0084] Step S5: The 3D real scene reconstruction module includes point cloud registration, mesh reconstruction and rendering.

[0085] Point cloud registration is based on the FPFH algorithm to extract point cloud features. The search radius is set to 0.1-0.2m. The KD-Tree is used to search for matching feature point pairs. After calculating the initial transformation matrix T0, the ICP algorithm is used for iterative optimization. The maximum number of iterations n icp =100, error threshold δ = 0.0001m.

[0086] Furthermore, the mesh reconstruction and rendering adopts a Poisson surface reconstruction algorithm, adaptively adjusts the octree depth according to the point cloud density, and utilizes the OpenGL graphics API to dynamically switch the model accuracy through the level of detail (LOD) technology to ensure the rendering frame rate ≥30fps.

[0087] Step S6: The system operation and maintenance module includes data storage and update, and system self-check.

[0088] Data storage and update: The detection results (including timestamp, coordinates, and hidden danger types) are stored in a distributed database, and the safety equipment detection model is updated regularly, such as once a month.

[0089] The system self-check performs a sensor status self-check when it is started every day. If the laser radar point cloud missing rate is greater than 10% or the camera frame rate is less than 20fps, an alarm is triggered and the backup device is switched.

[0090] The proposed method is highly intelligent, combining multiple detection methods to detect potential safety hazards at construction sites. It utilizes deep multimodal data fusion to overcome the limitations of single sensors. By collaboratively sensing the LiDAR point cloud (construction site geometry) with RGB images, it overcomes the technical limitations of traditional approaches, such as visual technology being susceptible to light and dust interference and radar technology lacking semantic understanding.

[0091] After years of research and development, the team has particularly leveraged the technical advantages of the improved YOLOv5 model and its attention mechanism. First, the CBAM attention mechanism module is embedded in the YOLOv5 model's Backbone network. Through the combined use of channel attention and spatial attention to screen key features, this method enhances adaptability to the dynamic environment of construction sites. Secondly, Ghost convolution is introduced in the Neck layer of YOLOv5 to reduce computational complexity through cheap feature reuse. At the same time, the number of feature channels is increased, combined with spatiotemporal pyramid pooling (SPP) to capture contextual information at different scales and optimize small target detection capabilities. Finally, a hybrid scaling strategy is adopted to enhance feature extraction capabilities while maintaining the model's lightweight, effectively balancing real-time detection and accuracy.

[0092] like Figure 2 As shown, the third embodiment of the present invention provides an automatic detection system for construction hazards based on point cloud and image multimodality, using the above method, the system includes:

[0093] Data acquisition module: In this embodiment, this module is responsible for collecting point cloud data and RGB image data at the construction site. According to the above data acquisition scheme, the lidar group is deployed at key points to collect point cloud data, and the high-definition cameras are arranged by area to collect RGB image data to achieve full-scene data acquisition.

[0094] Data processing module: includes a noise filtering unit, a dynamic downsampling unit, a ground segmentation unit and an image fusion unit connected in sequence; the noise filtering unit uses a statistical filtering algorithm to set the number of neighborhood points k and the distance threshold d as described above to eliminate noise points; the dynamic downsampling unit is based on the voxel grid method and adjusts the voxel side length according to the point cloud density; the ground segmentation unit uses the RANSAC algorithm to fit the ground plane model to separate the ground and non-ground point clouds; the image fusion unit calculates the transformation matrix H through SURF feature matching to complete the preliminary fusion of the point cloud and the RGB image.

[0095] Qualitative detection module: This module includes a safety equipment detection unit based on an improved YOLOv5 model and an abnormal behavior detection unit based on a 3D-CNN model. It is used to detect safety equipment by establishing an improved YOLOv5 model and to make qualitative judgments on abnormal personnel behavior by establishing a 3D-CNN model.

[0096] The safety equipment detection unit uses a modified YOLOv5 model, embedded with a CBAM attention mechanism module, and introduces a cross-modal feature pyramid (CMFPN) at the Neck layer. It takes a preprocessed RGB image as input, processes it through the model, and filters it using the NMS algorithm to detect when safety equipment is not being worn and generates an alarm. The 3D-CNN abnormal behavior detection unit collects abnormal behavior video data to train a model, inputs continuous camera frames, and triggers an alarm and tracks the person involved when the output confidence exceeds a threshold.

[0097] The quantitative analysis module includes a perimeter guard height detection unit and a position accuracy assurance unit. The perimeter guard height detection unit uses the DeepLabv3+ semantic segmentation model to locate the guard area, extracts the point cloud, fits the plane equation, calculates the height, and compares it to safety regulations and thresholds to generate an alarm. The position accuracy assurance unit uses camera calibration parameters and an extrinsic matrix to transform coordinates, employs the ICP algorithm to optimize alignment accuracy, and combines Kalman filtering and bilateral filtering to suppress errors.

[0098] Visualization display and early warning platform module: includes point cloud registration unit, mesh reconstruction unit, real-time rendering unit and early warning unit.

[0099] The point cloud registration unit uses the FPFH algorithm to extract point cloud features and optimizes the alignment through KD-Tree search and ICP algorithm; the mesh reconstruction unit uses the Poisson surface reconstruction algorithm to generate meshes and smooth them; the real-time rendering unit uses GPU parallel computing, OpenGL graphics API and LOD technology to achieve real-time rendering, providing an intuitive three-dimensional real-scene visualization model, and pushes warning information to the platform and management personnel terminals through the warning unit.

[0100] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A method for automatic detection of construction hazards based on point cloud and image multimodality, characterized in that: include: Step 1: Obtain point cloud data and RGB image data and preprocess them separately; Step 2: Build an improved YOLOv5 model to detect safety equipment and build a 3D-CNN model to qualitatively judge abnormal human behavior. Step 3: Combine the semantic segmentation model and point cloud data processing to perform quantitative judgment on fence height calculation and coordinate transformation; Step 4: Reconstruct the 3D real-scene model based on point cloud feature matching, Poisson surface reconstruction, and GPU rendering to detect safety hazards at the construction site.

2. The method for automatic detection of construction hazards based on point cloud and image multimodality according to claim 1 is characterized in that: In step 1: Statistical filtering, dynamic downsampling adaptive adjustment and ground segmentation are performed on the point cloud data in sequence; Gaussian filtering is performed on the RGB image data to remove noise, and then adaptive histogram equalization is performed to enhance the image contrast.

3. The method for automatic detection of construction hazards based on point cloud and image multimodality according to claim 1 is characterized in that: In step 2: The improved YOLOv5 model uses the CBAM attention mechanism module embedded in the Backbone network and introduces a cross-modal feature pyramid (CMFPN) in the Neck layer. It uses channel attention and spatial attention to jointly filter key features to enhance its adaptability to the dynamic environment of the construction site. The 3D-CNN model adopts the spatiotemporal pyramid pooling (SPP) module, takes RGB images as input, and detects abnormal behaviors.

4. The method for automatic detection of construction hazards based on point cloud and image multimodality according to claim 1 is characterized in that: In step 3: The semantic segmentation model is used to locate the edge enclosure area, extract the mask and fit the plane equation Ax+By+Cz+D=0; Calculate the maximum vertical distance difference h=|p in the point cloud max -p min |, compared with the safety standard threshold h s ,when When the alarm is triggered; The coordinate transformation is optimized using the ICP algorithm.

5. The method for automatic detection of construction hazards based on point cloud and image multimodality according to claim 1 is characterized in that: In step 4: The point cloud feature matching adopts the FPFH algorithm; the Poisson surface reconstruction algorithm adaptively adjusts the octree depth according to the point cloud density, and performs a smoothing process after generating a mesh.

6. According to the method for automatic detection of construction hazards based on point cloud and image multimodality in claim 3, the improved YOLOv5 model adds an attention mechanism module (CBAM), and its attention weight calculation function is Where Q, K, and V are query vectors, key vectors, and value vectors respectively; d k For the key vector dimension, use transfer learning to load the ImageNet pre-trained weights and train on the self-built dataset with the learning rate.

7. An automatic detection system for construction hazards based on point cloud and image multimodality, characterized by: Includes the following modules: Data acquisition module, used to collect point cloud data and RGB image data of the construction site; A data processing module, comprising a noise filtering unit, a dynamic downsampling unit, a ground segmentation unit, and an image fusion unit connected in sequence; and used to process the point cloud data and RGB image data collected at the construction site; The qualitative detection module includes a safety equipment detection unit based on an improved YOLOv5 model and an abnormal behavior detection unit based on a 3D-CNN model. It is used to detect safety equipment by establishing an improved YOLOv5 model and to qualitatively judge abnormal human behavior by establishing a 3D-CNN model. Quantitative analysis module, including edge protection height detection unit and position accuracy assurance unit; The visualization display and early warning platform module includes a point cloud registration unit, a mesh reconstruction unit, a real-time rendering unit, and an early warning unit.

8. The automatic detection system for construction hazards based on point cloud and image multimodality according to claim 7 is characterized in that: The qualitative detection module, wherein the improved YOLOv5 model adopts the CBAM attention mechanism module to embed the Backbone network, introduces the cross-modal feature pyramid (CMFPN) in the Neck layer, and jointly screens key features through channel attention and spatial attention to enhance adaptability to the dynamic environment of the construction site; The 3D-CNN model adopts the spatiotemporal pyramid pooling (SPP) module, takes RGB images as input, and detects abnormal behaviors.

9. The automatic detection system for construction hazards based on point cloud and image multimodality according to claim 7 is characterized in that: The system uses the processed point cloud data and RGB image data to reconstruct a 3D real-scene model of the construction site, uses a point cloud-based 3D reconstruction algorithm to generate a 3D mesh model, and maps the texture information of the RGB image onto the 3D model, making the model more realistic and intuitive.

10. The automatic detection system for construction hazards based on point cloud and image multimodality according to claim 7 is characterized in that: The identified safety hazards are marked on the 3D real-life model. The marked information includes the type, location, and severity of the safety hazard, which helps managers to intuitively understand the safety status of the construction site.

Citation Information

Cited By

  • High-place operation risk real-time judgment method based on image and point cloud bimodal data

    CN121459296A

  • High-altitude operation risk real-time determination method based on image and point cloud dual-mode data

    CN121459296B