Construction behavior identification method based on 3D point cloud and visible light image data fusion

By fusing 3D point cloud and visible light image data, optimizing the RANSAC algorithm and DBSCAN clustering, and combining it with crane boom straight line fitting, high-precision positioning of small target workers and identification of dangerous behaviors in oil construction scenarios were achieved. This solved the problems of insufficient accuracy and environmental adaptability in existing technologies and improved construction safety.

CN121661708APending Publication Date: 2026-03-13CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In oilfield construction scenarios, existing technologies struggle to balance accuracy and environmental adaptability with data from a single sensor, especially in power distribution network safety monitoring scenarios where the accuracy of locating and judging the behavior of personnel working on small targets is insufficient, leading to high safety risks.

Method used

A method of fusing 3D point cloud and visible light image data is adopted. The RANSAC algorithm is optimized by sector division and small-range point selection strategy. Combined with multi-frame point cloud accumulation and projection quantization error deduplication, the spatial location of personnel is accurately located. The point cloud-image is combined to perform straight line fitting on the crane arm to achieve two-dimensional hazard judgment.

Benefits of technology

It significantly improves the accuracy and safety of construction behavior recognition, reduces the impact of non-ground point interference and dense background points, adapts to complex construction environments, and achieves high-precision dangerous behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661708A_ABST
    Figure CN121661708A_ABST
Patent Text Reader

Abstract

The invention discloses a construction behavior identification method based on 3D point cloud and visible light image data fusion. The method comprises the following steps: synchronously acquiring a multi-modal data set of point cloud and visible light images; optimizing a to-be-fitted point set by adopting a fan-shaped regional and small-range point selection strategy; introducing a multi-frame point cloud accumulation and projection quantization error de-duplication strategy to optimize a point cloud background point dense problem; and combining the point cloud-image to carry out straight line fitting on the suspension arm, and judging whether an operator is under the suspension arm or in a rotating radius area in a two-dimensional manner. According to the method, the point clouds are divided in a fan-shaped mode, low-high points are selected in small areas to fit the ground, dense non-ground point interference is remarkably reduced, projection quantization error de-duplication processing is conducted on DBSCAN clustering results, accurate space positioning information is provided for personnel, the point clouds and image data are fused to fit a suspension arm straight line, and the positioning accuracy of the suspension arm is improved. The rotation radius is calculated through projection, and whether the personnel are in a dangerous area is judged, so that the accuracy and safety of construction behavior recognition in a distribution network safety supervision scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a construction behavior recognition method based on the fusion of 3D point cloud and visible light image data, belonging to the field of construction safety early warning technology in the petroleum system. Background Technology

[0002] In petroleum system construction and inspection scenarios, the working environments for power distribution network maintenance and line hoisting are complex and variable, involving the coordinated operation of high-voltage equipment, large cranes, and personnel, resulting in extremely high safety risks. Real-time and accurate identification of construction activities is a core requirement for ensuring operational safety. However, existing technologies have significant shortcomings in dealing with small target detection and interference from complex environments. Visible light images are easily affected by changes in lighting and occlusion, leading to misjudgments or omissions by personnel or equipment.

[0003] Early oilfield construction activity recognition relied on manual inspections, which were inefficient and lacked real-time performance, making it difficult to handle large-scale operation scenarios. With the development of computer vision technology, visible light image recognition has been gradually applied, but it is limited by two-dimensional information and cannot accurately obtain spatial distance and positional relationships. The recognition accuracy drops sharply when there is insufficient lighting or complex backgrounds.

[0004] A search revealed that Chinese patent CN118840397A discloses a point cloud-based method for identifying the construction status of components. Targeting the information-based supervision of building construction projects, it uses a single 3D point cloud sensor to acquire point cloud data from the construction site. Through a process of "coarse registration to generate a first target component point cloud → cropping to obtain a second target component point cloud → fine registration of the second target component point cloud → identification of the construction status based on the finely registered point cloud," it solves the problems of high cost and inaccurate construction status identification in traditional image processing solutions, enabling precise monitoring of component construction progress. Chinese patent CN119533422B discloses an intelligent detection and identification method and system for building construction. It uses a laser scanner to collect indoor laser point clouds, calculates the spatial coordinate information of the point clouds, and identifies the interface to be detected by comparing the point cloud with the display reference interface. This achieves detection and identification of each stage of building construction, reducing the time spent on repetitive detection and improving detection efficiency.

[0005] With the introduction of 3D point cloud technology, although the positioning capability has been greatly improved by spatial coordinates, the non-repeating scanning of LiDAR leads to uneven distribution of point clouds. Traditional random sampling consistency algorithms are easily affected by dense non-ground points when fitting the ground, resulting in insufficient fitting accuracy. At the same time, relying solely on point cloud or image data makes it difficult to solve problems such as the number of background points exceeding the number of target points, leading to inaccurate judgment of the spatial relationship between personnel and equipment, which cannot meet the high-precision safety monitoring needs of oil construction.

[0006] In summary, while 3D point clouds can provide spatial location information, data from a single sensor is difficult to balance accuracy and environmental adaptability. Especially in power distribution network safety monitoring scenarios, the accuracy of locating and judging the behavior of personnel working on small targets urgently needs to be improved, which directly threatens the safety of oil construction and the stable operation of the system.

[0007] A search revealed a Chinese patent with application number 202411684389.6 that discloses a real-time detection method for power line safety distance. The method involves installing monitoring equipment on the boom of a power line inspection crane. The equipment includes a binocular camera and a lidar. It acquires power line information and calculates the safe distance using point cloud data, registers this data with image data, draws a bounding box around the power line on the target image, and displays the distance to determine the distance between the power line and the crane's boom. This patent addresses situations where the crane is operating in an environment with high-voltage power lines. It relies on traditional RANSAC (Random Sample Consensus) global sampling for ground fitting, which suffers from severe non-ground interference. Furthermore, it uses single-frame point cloud clustering and distance calculation based on a preset safety threshold for hazard assessment. While it can measure power line distance in real time, it is limited by single-modal data and has a high false detection rate in complex construction scenarios (such as boom obstruction). Summary of the Invention

[0008] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a construction behavior recognition method based on the fusion of point cloud and visible light image data, so as to improve the accuracy of small target image detection results in power distribution network safety monitoring scenarios.

[0009] This invention provides a construction behavior recognition method based on the fusion of 3D point cloud and visible light image data, comprising the following steps: Step 1: Synchronously collect multimodal datasets of 3D point cloud and visible light images related to the crane boom and workers at the construction site; Step 2: Optimize the RANSAC set of points to be fitted by adopting a sector-shaped regional division and small-range point selection strategy, reduce interference from non-ground points, and improve fitting accuracy. Step 3: Introduce multi-frame point cloud accumulation and projection quantization error deduplication strategy to optimize the dense point cloud background and improve the accurate positioning of personnel in space. Step 4: Combine point cloud images to perform straight line fitting on the boom, and determine in two dimensions whether the operator is under the boom or within the rotation radius area.

[0010] This invention filters out non-target interference layer by layer from ground fitting and zoning point selection to personnel positioning projection deduplication, adapting to complex environments with many equipment and cluttered backgrounds at construction sites, and has strong anti-interference capabilities. At the same time, it adopts point cloud accumulation to adapt to real-time monitoring of dynamic construction scenarios, and sets up dual-dimensional judgment of dangerous distance and risk scenarios based on the safety requirements of power grid hoisting operations, quantifies the dangerous range of the boom, and accurately identifies high-risk behaviors under the boom or within the rotation radius.

[0011] The following is a further optimized technical solution of the present invention: In step 1, the 3D point cloud is continuously acquired by LiDAR, and the visible light image is continuously acquired by a visible light camera, so as to simultaneously obtain the three-dimensional coordinates of the boom and the action information of the operator, forming a multimodal dataset containing point cloud frames and image frames. In the multimodal dataset, point cloud frames and image frames are stored in pairs, and each pair of data contains the three-dimensional coordinates of the point cloud, image pixel information and a unified timestamp to ensure data synchronization.

[0012] In step 2, the specific steps to reduce non-ground point interference and improve fitting accuracy are as follows: Step 2.1: Based on the preset threshold of ground height at the construction site, perform height filtering on the original point cloud to initially screen out points with lower heights, so as to reduce the amount of calculation for non-ground points. Step 2.2: Divide the selected point cloud into several fan-shaped regions according to a certain step size, determine the fan-shaped region to which each point belongs, divide each fan-shaped region into small cells according to a certain radial distance, determine the small cell to which each point belongs, and form a gridded spatial unit. Step 2.3: Traverse each small cell and extract the point with the lowest height within the cell. The set of the lowest points in all small cells together constitutes the set of points to be fitted in RANSAC. Step 2.4: Execute the RANSAC algorithm on the constructed set of points to be fitted. In each iteration, three points are randomly selected to fit the plane. The plane equation is:

[0013] Where a, b, and c are the three components of the plane's normal vector, which together describe the plane's spatial orientation; d is the plane's intercept parameter, which, together with the normal vector, determines the plane's position in space; x, y, and z are the three-dimensional coordinates of any point in space. The distance from each point in the point cloud to the plane is calculated using the following formula:

[0014] In the formula, Let be the perpendicular distance from the i-th point in the point cloud to the fitting plane. x i , y i ,z i For the first in space i The three-dimensional coordinates of each point; Distance Points smaller than a set threshold are identified as interior points. The plane with the most interior points in multiple iterations is used as the fitted ground model; By merging the ground models of multiple sector regions, a complete global ground model is obtained.

[0015] This invention addresses the problems of non-ground point interference and uneven point cloud distribution that exist when directly using RANSAC for ground fitting, leading to poor extraction results (see...). Figure 2 This paper proposes a two-layer optimization strategy of first partitioning and then selecting points to screen high-quality point sets for fitting from the source, thereby improving the accuracy of ground point cloud extraction. Point clouds acquired by lidar contain a large number of low-altitude, dense non-ground points (such as stacked building materials and low-lying equipment), which may be misclassified as interior points, interfering with the planar fitting results. Simultaneously, the non-repetitive scanning characteristic of lidar results in uneven point cloud distribution in space, with some areas having sparse points and others dense points. Direct fitting will lead to local biases and fail to cover the entire ground. This invention employs sector partitioning to address the uneven distribution problem. Based on the horizontal field of view of the lidar, the point cloud is divided into several sector regions with a preset step size. The viewing range of each region is fixed, avoiding the influence of cross-region point cloud density differences. At the same time, small-cell point selection is used to solve non-ground point interference. Within each sector sub-region, small cells are divided radially (i.e., the radar transmission direction) with a preset step size to further reduce the point cloud range. Within each small cell, only the point with the lowest height direction (Z-axis) is selected as the fitting point, eliminating non-ground high points within the small cell from the source, ensuring that the set of points to be fitted is dominated by ground points.

[0016] In step 3, the specific operations to improve the accuracy of personnel spatial positioning are as follows: Step 3.1: Accumulate point clouds within a certain time period according to a preset frequency to obtain a set of point clouds superimposed with multiple frames, and at the same time acquire the visible light image within the corresponding time window; Step 3.2: Process the visible light image using a target tracking algorithm and output the image tracking box for each person; Step 3.3: Calculate the projection matrix from the point cloud to the image using the extrinsic and intrinsic parameters of the lidar and visible light camera, retain the point cloud that falls within the human tracking box after projection, and obtain the candidate human point cloud set; Step 3.4: Perform DBSCAN clustering on the candidate human point cloud set to obtain multiple cluster categories. Project the point cloud of each cluster onto the image pixel coordinate system. After deduplicating the projected coordinates, count the number of deduplicated projected pixels in each cluster. Select the cluster with the most deduplicated projected pixels as the final human point cloud cluster. Step 3.5: For the selected human body cluster point cloud, calculate the spatial position using the centroid formula and output the three-dimensional coordinates of the person.

[0017] in, , , These are the average coordinates of the human body cluster point cloud in the x, y, and z axes, respectively; m is the total number of point clouds contained in the human body cluster. , , These are the x-axis, y-axis, and z-axis coordinates of the j-th point in the human body cluster, respectively.

[0018] To address the issue of density-based background interference in traditional DBSCAN clustering, this invention employs multi-frame accumulation and projection quantization error deduplication to achieve precise spatial positioning of personnel, making personnel positioning more reliable in dynamic construction scenarios. Traditional DBSCAN clustering and point cloud quantity determination methods have significant shortcomings in construction scenarios, directly leading to positioning errors. This is due to two reasons: first, uneven point cloud density. LiDAR scanning exhibits a "dense in the center, sparse at the edges" phenomenon. If a human target bounding box falls on the edge area while background points are concentrated in the center, the number of background points will exceed the number of human points. Second, background point accumulation. When multiple frames are superimposed, the point cloud of static backgrounds (such as ground equipment or fixed supports) continuously accumulates, further amplifying the numerical advantage of background points, causing DBSCAN to mistakenly cluster the background as a human. This invention optimizes the method from two aspects: data input and clustering selection. It first uses multi-frame point cloud accumulation to address the problem of uneven point cloud density, and then uses projection quantization error deduplication to address background interference.

[0019] To balance point cloud density and real-time performance, personnel spatial localization employs an accumulation strategy that divides the point cloud into several parts within a certain time frame. Visible light images at corresponding time points are acquired using time-aligned methods, and human tracking boxes are obtained from the images using a target tracking algorithm. Based on the intrinsic and extrinsic parameters of radar and camera, and the projection matrix, the point cloud set mapped to the tracking box is extracted. To address the interference from certain dense background points in the space, a point cloud-image projection quantization error mechanism is introduced into the personnel spatial localization algorithm to further process the DBSCAN clustering results. Specifically, the point cloud of each cluster category is projected onto the image pixel space. However, many projected point coordinates are duplicated, which is caused by uneven point cloud spatial distribution density and the existence of quantization errors.

[0020] In step 4, the point cloud image is used to fit a straight line to the boom, and a two-dimensional determination is made as to whether the operator is under the boom or within the rotation radius area. The specific operation is as follows: Step 4.1: Obtain the crane image detection box through image target detection. Combine the radar and camera extrinsic and intrinsic parameters to filter out the crane point cloud projected into the detection box in order to eliminate background point interference. Step 4.2: Execute the RANSAC algorithm on the point cloud of the boom. In each iteration, select two points to generate the equation of a straight line, calculate the distance from other points to the straight line, count the number of interior points, and the straight line with the most interior points after multiple iterations is the straight line in the boom space. Step 4.3: Traverse the point cloud along the straight line of the boom and filter out the points with the highest and lowest heights, which are the boom apex and base points. Step 4.4: Project the top of the boom and the base point onto the ground (Z=0 is the plane), and calculate the Euclidean distance between the two points, which is the boom rotation radius; Step 4.5: Retrieve the three-dimensional coordinates of the personnel and project them onto the ground to obtain the personnel's ground projection points; Step 4.6, Two-dimensional hazard assessment—Calculate the vertical distance from the personnel's ground projection point to the straight line of the crane's ground projection point. If the vertical distance is less than the preset value, it is determined to be a hazard under the crane. Calculate the Euclidean distance between the personnel's ground projection point and the base point's ground projection. If the Euclidean distance is less than the rotation radius, it is determined to be within the rotation radius and therefore dangerous. An alarm will be triggered if any of the conditions are met.

[0021] This invention transforms abstract spatial relationships into quantifiable distance judgments, making the identification of hazardous behaviors more operational. Based on safety regulations for oil lifting operations, personnel directly beneath the boom face the highest impact risk should the boom structurally break or heavy objects fall. Personnel within the boom's rotation radius are also at risk of horizontal collisions during rotation, as they are easily grazed by the boom or auxiliary equipment. Therefore, this invention achieves accurate identification of hazardous behaviors through two steps: fitting a straight line to the boom to define the risk range and projecting personnel's distance to determine the risk. The first step, fitting a straight line to the boom to define the risk range, abstracts the boom as a straight line and quickly calculates its ground projection and rotation radius using the line's parameters, avoiding errors in range judgment caused by complex boom structures (such as multi-section booms or hooks). Specifically, it first obtains an image detection box of the boom through image target detection, then extracts the boom point cloud projected onto the detection box using radar and camera intrinsic and extrinsic parameters, and finally uses the RANSAC algorithm to fit a straight line (selecting two points per round to define the line, and filtering for the optimal line based on the number of interior points) to ensure fitting accuracy. The dual-dimensional distance-based hazardous behavior assessment transforms 3D spatial relationships into 2D ground projection distance calculations, simplifying the judgment logic and aligning with the actual risks in construction scenarios, where hazards primarily occur at the ground level. Specifically, hazard assessment under the boom is based on the ground projection of the boom's straight line. The vertical distance from the personnel's ground projection point to the boom's projected straight line is calculated, and a preset threshold is applied; if the threshold is less than the threshold, a hazard is deemed present. Hazard assessment within the rotation radius involves first calculating the distance between the boom's apex (maximum height) and the base point (minimum height) on the ground projection; this distance is the boom's rotation radius. Then, the Euclidean distance from the personnel's ground projection to the base point's projection is calculated; if the distance is less than the radius, a hazard is deemed present.

[0022] Compared with the prior art, the present invention has the following beneficial effects: (1) Optimize the random sampling consistency algorithm. By dividing the point cloud into fan shapes according to the horizontal field of view of the lidar and selecting low and high points in a small area to fit the ground, the interference of dense non-ground points is significantly reduced, the ground fitting accuracy is improved, and the foundation is laid for subsequent target processing. (2) Improve the spatial positioning strategy for operators by accumulating point cloud frames and combining them with target tracking, and perform projection quantization error deduplication on DBSCAN clustering results. This effectively uses multiple sensors to solve the problem of sudden changes in ranging caused by dense background points, and provides accurate spatial positioning information for personnel in safety monitoring algorithms. (3) By integrating point cloud and image data to fit the straight line of the boom, the rotation radius is calculated by projection and it is determined whether personnel are in the danger zone. This accurately identifies dangerous behaviors under the boom and within the rotation radius, effectively improving the accuracy and safety of construction behavior identification in power distribution network safety supervision scenarios. Attached Figure Description

[0023] Figure 1This is a flowchart of the present invention; Figure 2 This is a RANSAC ground fitting result diagram; Figure 2 The blue dots represent the fitted ground points; Figure 3 A schematic diagram of point cloud sector division; Figure 4 The improved ground fitting result diagram; Figure 4 The blue dots represent the fitted ground points; Figure 5 Before and after removing ground points from the original point cloud; in the figure, a is the original point cloud image, and b is the point cloud image after removing ground points; Figure 6 This diagram illustrates the abrupt change in distance during spatial positioning of personnel in data fusion; in the diagram, a is the first frame, b is the second frame, and c is the third frame. Figure 7 A diagram illustrating the pixel proportions of the target bounding box occupied by the human body; Figure 7 Yellow represents the human body, and red represents the background; Figure 8 The image shows the optimized personnel spatial positioning algorithm; in the image, a is the first frame, b is the second frame, and c is the third frame. Figure 9 This is a schematic diagram of the fitted boom line and rotation radius in the point cloud; Figure 9 The red dot in the middle represents the boom point; a is the boom rotation radius, and b is the projection of the boom rotation radius. Figure 10 A schematic diagram showing people standing below the boom and within the rotation radius; Figure 11 This is a schematic diagram of a personnel alarm within the rotation radius. Detailed Implementation

[0024] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings: This embodiment is implemented under the premise of the technical solution of the present invention, and provides detailed implementation methods and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments. Example 1

[0025] A construction behavior recognition method based on the fusion of 3D point cloud and visible light image data, such as Figure 1 As shown, it includes the following steps: Step 1: Synchronously collect multimodal datasets of 3D point cloud and visible light images related to crane booms and workers at the construction site.

[0026] The 3D point cloud data was continuously acquired using an AVIA LiDAR system, utilizing its 70° horizontal field of view and non-repeating scanning characteristics to cover the main spatial range of the boom's rotation. Visible light images were continuously acquired using a high-resolution visible light camera, focusing on capturing the visual characteristics of workers and the boom, such as clothing and boom color markings. Both the AVIA LiDAR and the visible light camera were horizontally fixed at high points on the construction site, such as tower cranes and supports, ensuring unobstructed views. The AVIA LiDAR and visible light camera needed to cover the boom's working radius (≥50 meters) and the workers' activity area to avoid missing key targets due to limited viewing angles. Considering the dynamic operation characteristics of the boom, the AVIA LiDAR and visible light camera used continuous acquisition mode, recording the boom's spatial posture and workers' movement trajectories in real time, while simultaneously acquiring the boom's 3D coordinates and workers' motion information, forming a multimodal dataset containing point cloud frames and image frames. In the multimodal dataset, the collected data is stored in pairs as point cloud frames and image frames. Each pair of data contains the three-dimensional coordinates of the point cloud, image pixel information, and a unified timestamp, providing complete and synchronized raw data support and ensuring data synchronization.

[0027] Step 2: Optimize the RANSAC set of points to be fitted by adopting a sector-shaped regional division and small-range point selection strategy, reduce interference from non-ground points, and improve fitting accuracy.

[0028] The specific steps are as follows: Step 2.1 Initial point cloud screening—Based on the preset threshold of ground height at the construction site (the preset threshold is ≤1.5m), the original point cloud is filtered by height to initially screen out points with lower height in the height direction (i.e., Z-axis) in order to reduce the amount of calculation for non-ground points. Step 2.2, Spatial Mesh Division—Divide the selected point cloud into 14 sector sub-regions with a step size of 5°, and determine the sector region to which each point belongs; Divide each sector region into small cells with a radial distance of 1m, and determine the small cell to which each point belongs; Step 2.3: Construction of the set of points to be fitted - Traverse each small cell and extract the point with the lowest Z-axis within the cell. The set of the lowest points of all small cells together constitutes the set of points to be fitted in RANSAC. Step 2.4, Optimize RANSAC Fitting—Execute the RANSAC algorithm on the constructed set of points to be fitted. In each iteration, three points are randomly selected to fit the plane. The plane equation is:

[0029] Where a, b, and c are the three components of the plane's normal vector, which together describe the plane's spatial orientation; d is the plane's intercept parameter, which, together with the normal vector, determines the plane's position in space; x, y, and z are the three-dimensional coordinates of any point in space. The distance from each point in the point cloud to the plane is calculated using the following formula:

[0030] In the formula, Let be the perpendicular distance from the i-th point in the point cloud to the fitting plane. x i , y i , z i For the first in space i The three-dimensional coordinates of each point; Distance Points smaller than a set threshold are identified as interior points. The plane with the most interior points in multiple iterations is used as the fitted ground model; By merging the ground models of 14 sector regions, a complete global ground model is obtained.

[0031] Accurate extraction of ground point clouds is a crucial prerequisite for lidar point cloud preprocessing. Removing ground point interference allows for more targeted segmentation and analysis of point clouds containing targets such as workers and cranes. In practical applications, considering that most equipment is horizontally placed, ground points are located at a relatively low position in the acquired point cloud. Therefore, to improve algorithm efficiency, points at lower elevations in the point cloud are input into the algorithm for ground extraction. The classic Random Sample Consensus (RANSAC) algorithm is used to fit a plane to the lower elevations (z-axis) of the point cloud. Therefore, in each iteration of the RANSAC algorithm, three points are randomly selected to fit a plane, and the distance from each point in the point cloud to the plane is calculated. If the value is less than a set threshold, the point is considered an interior point in the current iteration. The plane with the most interior points across multiple iterations is the fitted ground model. The point cloud ground extraction result is as follows. Figure 2 As shown, the blue dots in the figure (ground points obtained by fitting points with lower values ​​in the overall height direction) do not cover most of the ground. Considering the non-repeating scanning characteristics of lidar, which leads to uneven distribution of point clouds in space, some dense non-ground points in space severely interfere with interior point statistics. Therefore, the strategy for selecting the set of points to be fitted needs further improvement.

[0032] Therefore, this embodiment addresses the problems of non-ground point interference and uneven point cloud distribution that exist when directly using RANSAC for ground fitting, which lead to poor extraction results (see...). Figure 2This paper proposes a two-layer optimization strategy of first partitioning and then selecting points to screen high-quality point sets for fitting from the source, thereby improving the accuracy of ground point cloud extraction. Point clouds acquired by lidar contain a large number of low-altitude, dense non-ground points (such as stacked building materials and low-lying equipment), which are misclassified as interior points, interfering with the planar fitting results. Simultaneously, the non-repetitive scanning characteristic of lidar results in uneven point cloud distribution in space, with some areas having sparse points and others dense points. Direct fitting will lead to local biases and fail to cover the entire ground. This invention addresses the uneven distribution problem by employing fan-shaped partitioning. Based on the lidar's 70° horizontal field of view, the original point cloud is divided into 14 fan-shaped sub-regions with a step size of 5°. (See Figure 3 Each region has a fixed field of view to avoid the impact of cross-regional point cloud density differences. On the other hand, small-cell point selection is used to address non-ground point interference. Within each fan-shaped sub-region, small cells are divided radially (i.e., in the radar transmission direction) at 1-meter intervals. This forms a gridded spatial unit, further reducing the point cloud range. Within each cell, only the point with the lowest height (Z-axis) is selected as the fitting point, eliminating non-ground high points within the cell from the source, ensuring that the set of points to be fitted is mainly composed of ground points.

[0033] For each spatial cell, points with lower elevation in the point cloud direction within the region are selected to form the set of points to be fitted. The RANSAC algorithm is then executed to fit the ground. Finally, the fitting results from each region are merged to obtain a complete ground model. Figure 4 As shown, the blue points (the fitted ground points) cover most of the ground points. After improving the strategy for the set of points to be fitted, the interference from dense non-ground points is significantly reduced. The point cloud after removing ground points (see...) Figure 5 It can accurately preserve the three-dimensional features of targets such as crane booms and workers, providing a high-quality data foundation for subsequent human positioning and dangerous behavior assessment.

[0034] Step 3: Introduce multi-frame point cloud accumulation and projection quantization error deduplication strategy to optimize the dense point cloud background and improve the accurate positioning of personnel in space.

[0035] The specific steps are as follows: Step 3.1, Multi-frame point cloud accumulation—Accumulate point clouds within 1 second at a frequency of 10Hz (1 frame every 0.1 seconds) to obtain a set of 10 superimposed point clouds, and at the same time acquire the visible light image within the corresponding time window; Step 3.2, Image Human Tracking—Use a target tracking algorithm (such as DeepSORT) to process the visible light image and output the image tracking box (i.e., the range of pixel coordinates) for each person; Step 3.3, Point Cloud-Image Mapping Filtering—Calculate the projection matrix from the point cloud to the image using the extrinsic parameters (such as rotation matrix and translation vector) and intrinsic parameters (such as camera focal length and principal point coordinates) of the LiDAR and visible light camera. Retain the point clouds that fall within the human tracking box after projection to obtain the candidate human point cloud set. Step 3.4, DBSCAN Clustering + Projection Deduplication Filtering: Perform DBSCAN clustering on the candidate human point cloud set to obtain multiple cluster categories, including human cluster and background cluster. Project the point cloud of each cluster onto the image pixel coordinate system. After deduplicating the projected coordinates, count the number of deduplicated projected pixels in each cluster. Select the cluster with the most deduplicated projected pixels as the final human point cloud cluster. Step 3.5, Personnel 3D Coordinate Calculation—For the selected human cluster point cloud, the spatial position is calculated using the centroid formula, and the 3D coordinates of the personnel are output.

[0036] in, , , These are the average coordinates of the human body cluster point cloud in the x, y, and z axes, respectively; m is the total number of point clouds contained in the human body cluster. , , These are the x-axis, y-axis, and z-axis coordinates of the j-th point in the human body cluster, respectively.

[0037] This embodiment divides the point cloud within one second into 10 parts at a frequency of 10Hz by setting a time window, and accumulates the point cloud in a time-aligned manner. This ensures both an increase in the density of the personnel point cloud and avoids clustering failure caused by sparse point clouds in a single frame, while also controlling the time window to ensure real-time positioning and adapting to the needs of dynamic personnel movement in construction scenarios. In addition, by pre-screening the point cloud and obtaining the visible light image at the corresponding time through point cloud-image time synchronization in step 1, the human tracking box in the image is obtained using a target tracking algorithm (such as SORT, DeepSORT). Only the point cloud mapped to the tracking box is extracted and accumulated, reducing the initial input of irrelevant background points.

[0038] Traditional methods directly perform DBSCAN clustering on the point cloud within the human bounding box, performing density-based clustering and selecting the cluster with the most points as the human body category. However, actual testing results show that the obtained human body category is not the true human body point cloud, such as... Figure 6The image shows the human ranging results for three consecutive frames. The ranging results in the first and third frames are correct, while the second frame shows a sudden change. This is because the point cloud density obtained from the AVIA LiDAR spatial scan is not uniform; the point cloud is denser closer to the center. In the second frame, the background within the human target frame is located in an area with high LiDAR scanning density, and there is also background accumulation within the point cloud. Therefore, the number of background points is greater than the number of human point clouds. Relying solely on the number of points in each cluster to determine the human target category will severely interfere with positioning accuracy.

[0039] This embodiment leverages the fact that the human body occupies a much larger proportion of the image tracking box than the background. This means that background points are spatially dispersed within the tracking box, resulting in low pixel repetition when projected onto the image, while the human body point cloud has a high pixel repetition rate. First, each cluster of point clouds output by DSBCAN is mapped to the image pixel coordinate system using radar and camera intrinsic and extrinsic parameters and a projection matrix. Then, the projected pixel coordinates are deduplicated, retaining only one instance of the same pixel coordinate. Finally, the number of deduplicated projected pixels is counted, and the cluster with the most duplicates is selected as the human target, rather than relying on the traditional point cloud count method.

[0040] Through the Figure 7 Analysis revealed that background points, due to their dispersed spatial distribution, had a low pixel repetition rate in their projections, while the human body occupied the vast majority of pixels within the target bounding box. Therefore, the pixel coordinates projected from the point cloud onto the image were deduplicated. Then, the number of projection points for each cluster was counted. Leveraging the prior knowledge that "the human body occupies a much larger proportion of the image than the background," the cluster with the most projected pixels was selected as the human body category. Finally, the spatial position of the selected human body cluster point cloud was calculated using the centroid formula, ultimately outputting the three-dimensional coordinates of the person.

[0041] The optimized ranging algorithm performs as follows: Figure 8 As shown, the ranging result in the second frame is normal, and there is no sudden change in distance.

[0042] Step 4: Combine point cloud images to perform straight line fitting on the boom, and determine in two dimensions whether the operator is under the boom or within the rotation radius area.

[0043] In oil lifting operations, identifying hazardous behaviors of construction workers requires considering not only whether they are wearing protective equipment, but also the potential dangers posed by the lifting equipment. Safety regulations state that there is a danger if personnel are located under the boom or within its rotation radius during lifting operations. Therefore, it is necessary to determine the spatial relationship between the personnel and the boom.

[0044] The specific steps are as follows: Step 4.1, Crane Boom Point Cloud Extraction—Obtain the crane boom image detection box through image target detection, and combine radar and camera extrinsic and intrinsic parameters to filter out the crane boom point cloud projected into the detection box in order to eliminate background point interference; Step 4.2, RANSAC Line Fitting—Execute the RANSAC algorithm on the boom point cloud. In each iteration, select two points to generate a line equation. The line equation is:

[0045] in, Let be the three-dimensional coordinates of any point on a straight line in space; , These are the three-dimensional coordinates of any two points on the line; Calculate the distance from other points to the line, count the number of interior points, and after multiple iterations, the line with the most interior points is the crane space line; Step 4.3, Determining the key points of the boom—Traverse the point cloud along the straight line of the boom and filter out the points with the highest and lowest height in the direction of height, which are the boom apex and base points; Step 4.4, Rotation radius calculation—Project the boom apex and base point onto the ground (Z=0 plane), calculate the Euclidean distance between the two points, which is the boom rotation radius; Step 4.5, Personnel Location Acquisition—Retrieve the three-dimensional coordinates of the personnel output in Step 3 and project them onto the ground to obtain the personnel's ground projection points; Step 4.6, Two-dimensional hazard assessment—Calculate the vertical distance from the personnel's ground projection point to the straight line of the crane's ground projection point. If the vertical distance is less than the preset value, it is determined to be a hazard under the crane. Calculate the Euclidean distance between the personnel's ground projection point and the base point's ground projection. If the Euclidean distance is less than the rotation radius, it is determined to be within the rotation radius and therefore dangerous. An alarm will be triggered if any of the conditions are met.

[0046] This embodiment abstracts the boom as a straight line in space, similar to the spatial positioning of a person. Combining the boom rotation detection frame and point cloud, the point cloud within the boom frame is extracted through intrinsic and extrinsic parameter mapping of radar and cameras. Two points can determine a straight line. The RANSAC algorithm is used to fit the line. In each iteration, two points are randomly selected, and the line equation is determined based on the two-point equation. The distance from each point in the point cloud to the line is calculated. If the distance is less than a set threshold, then this point is an interior point in this iteration. The line with the most interior points in multiple iterations is the boom line. The fitting result is as follows: Figure 9 As shown, by fusing point cloud and image data, a straight line of the boom can be accurately fitted in space, and it can be projected onto the image using intrinsic and extrinsic parameters. After fitting the straight line of the boom, the highest point in the height direction can be determined as the boom vertex, and the lowest point in the height direction as the boom base point, based on height features. These two points are then projected onto the ground (…). The radius of rotation of the boom can be calculated by looking at the plane.

[0047] The spatial location of personnel is obtained based on step 3, combining the human detection bounding box and point cloud data, through intrinsic and extrinsic parameter mapping of LiDAR and camera. For example... Figure 10 As shown, based on the spatial positioning results of personnel, two types of hazard assessments are performed through ground projection: (1) Whether there are people standing under the boom: Project the position of the person in space obtained from the distance measurement onto the ground, calculate the vertical distance from the projection point of the person to the projection line of the boom. If the distance is less than 1 meter, it is determined that the person is directly under the boom, which is a dangerous behavior. (2) Whether there are people standing within the rotation radius: Project the spatial position of the people obtained from the personnel distance measurement onto the ground, calculate the Euclidean distance between the personnel projection point and the lowest point of the boom's projection point on the ground. If the distance is less than the boom's rotation radius, it is determined that the personnel are within the boom's rotation radius, which is a dangerous behavior. Figure 11 The diagram shown illustrates the alarm for people stationed within the rotation radius.

[0048] The rotation radius is determined by fitting the straight line of the boom, and the position of personnel in the construction scene is determined by spatial positioning of personnel. Then, it is determined whether the personnel are under the boom or within the rotation radius, and an alarm is issued in time.

[0049] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any transformations or substitutions that can be conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A construction behavior recognition method based on the fusion of 3D point cloud and visible light image data, characterized in that, Includes the following steps: Step 1: Synchronously collect multimodal datasets of 3D point cloud and visible light images related to the crane boom and workers at the construction site; Step 2: Optimize the RANSAC set of points to be fitted by adopting a sector-shaped regional division and small-range point selection strategy, reduce interference from non-ground points, and improve fitting accuracy. Step 3: Introduce multi-frame point cloud accumulation and projection quantization error deduplication strategy to optimize the dense point cloud background and improve the accurate positioning of personnel in space. Step 4: Combine point cloud images to perform straight line fitting on the boom, and determine in two dimensions whether the operator is under the boom or within the rotation radius area.

2. The construction behavior recognition method based on the fusion of 3D point cloud and visible light image data according to claim 1, characterized in that: In step 1, the 3D point cloud is continuously acquired by LiDAR, and the visible light image is continuously acquired by a visible light camera, so as to simultaneously obtain the three-dimensional coordinates of the boom and the action information of the operator, forming a multimodal dataset containing point cloud frames and image frames. In the multimodal dataset, point cloud frames and image frames are stored in pairs, and each pair of data contains the three-dimensional coordinates of the point cloud, image pixel information and a unified timestamp to ensure data synchronization.

3. The construction behavior recognition method based on the fusion of 3D point cloud and visible light image data according to claim 1, characterized in that: In step 2, the specific steps to reduce non-ground point interference and improve fitting accuracy are as follows: Step 2.1: Based on the preset threshold of ground height at the construction site, perform height filtering on the original point cloud to initially select points with lower heights; Step 2.2: Divide the selected point cloud into several fan-shaped regions according to a certain step size, determine the fan-shaped region to which each point belongs, divide each fan-shaped region into small cells according to a certain radial distance, determine the small cell to which each point belongs, and form a gridded spatial unit. Step 2.3: Traverse each small cell and extract the point with the lowest height within the cell. The set of the lowest points in all small cells together constitutes the set of points to be fitted in RANSAC. Step 2.4: Execute the RANSAC algorithm on the constructed set of points to be fitted. In each iteration, three points are randomly selected to fit the plane. The plane equation is: , in, , , These are the three components of the plane normal vector; The intercept parameter of the plane; , , Let be the three-dimensional coordinates of any point in space; The distance from each point in the point cloud to the plane is calculated using the following formula: , In the formula, Let be the perpendicular distance from the i-th point in the point cloud to the fitting plane. x i , y i , z i For the first in space i The three-dimensional coordinates of each point; Distance Points smaller than a set threshold are identified as interior points. The plane with the most interior points in multiple iterations is used as the fitted ground model; By merging the ground models of multiple sector regions, a complete global ground model is obtained.

4. The construction behavior recognition method based on the fusion of 3D point cloud and visualized image data according to claim 1, characterized in that: In step 3, the specific operations to improve the accuracy of personnel spatial positioning are as follows: Step 3.1: Accumulate point clouds within a certain time period according to a preset frequency to obtain a set of point clouds superimposed with multiple frames, and at the same time acquire the visible light image within the corresponding time window; Step 3.2: Process the visible light image using a target tracking algorithm and output the image tracking box for each person; Step 3.3: Calculate the projection matrix from the point cloud to the image using the extrinsic and intrinsic parameters of the lidar and visible light camera, retain the point cloud that falls within the human tracking box after projection, and obtain the candidate human point cloud set; Step 3.4: Perform DBSCAN clustering on the candidate human point cloud set to obtain multiple cluster categories. Project the point cloud of each cluster onto the image pixel coordinate system. After deduplicating the projected coordinates, count the number of deduplicated projected pixels in each cluster. Select the cluster with the most deduplicated projected pixels as the final human point cloud cluster. Step 3.5: For the selected human body cluster point cloud, calculate the spatial position using the centroid formula and output the three-dimensional coordinates of the person. ,, in, , , These are the average coordinates of the human body cluster point cloud in the x, y, and z axes, respectively; m is the total number of point clouds contained in the human body cluster. , , These are the x-axis, y-axis, and z-axis coordinates of the j-th point in the human body cluster, respectively.

5. The construction behavior recognition method based on the fusion of 3D point cloud and visible light image data according to claim 1, characterized in that: In step 4, the point cloud image is used to fit a straight line to the boom, and a two-dimensional determination is made as to whether the operator is under the boom or within the rotation radius area. The specific operation is as follows: Step 4.1: Obtain the crane image detection box through image target detection, and filter out the crane point cloud projected into the detection box by combining radar, camera extrinsic and intrinsic parameters; Step 4.2: Execute the RANSAC algorithm on the point cloud of the boom. In each iteration, select two points to generate the equation of a straight line, calculate the distance from other points to the straight line, count the number of interior points, and the straight line with the most interior points after multiple iterations is the straight line in the boom space. Step 4.3: Traverse the point cloud along the straight line of the boom and filter out the points with the highest and lowest heights, which are the boom apex and base points. Step 4.4: Project the top of the boom and the base point onto the ground, and calculate the Euclidean distance between the two points, which is the boom rotation radius; Step 4.5: Retrieve the three-dimensional coordinates of the personnel and project them onto the ground to obtain the personnel's ground projection points; Step 4.6, Two-dimensional hazard assessment—Calculate the vertical distance from the personnel's ground projection point to the straight line of the crane's ground projection point. If the vertical distance is less than the preset value, it is determined to be a hazard under the crane. Calculate the Euclidean distance between the personnel's ground projection point and the base point's ground projection. If the Euclidean distance is less than the rotation radius, it is determined to be within the rotation radius and therefore dangerous.

Citation Information

Patent Citations

  • Component construction state identification method based on point cloud

    CN118840397A

  • An intelligent detection and identification method and system for building construction

    CN119533422B

  • Real-time detection method for safe distance of power line

    CN119648789A

  • Object detection and tracking method of fusing laser point clouds and images

    CN108509918A

  • Fusion calibration method of three-dimensional laser radar and binocular visible light sensor

    CN110349221A