Tower crane hook tracking and monitoring system and method based on point cloud target detection algorithm

Through the tower crane hook tracking and monitoring system based on point cloud target detection algorithm, safety hazards caused by manual negligence during tower crane construction are solved, real-time monitoring and collision warning are achieved, and construction and safety levels are improved.

CN120088457APending Publication Date: 2025-06-03CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510174353.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

During the construction process, safety hazards and accidents caused by manual negligence occur frequently, and the prior art cannot conduct real-time and accurate collision warning and control.

Method used

The tower crane hook tracking and monitoring system based on point cloud object detection algorithm is adopted to obtain 3D point cloud data at the tower crane site frame by frame, generate pseudo-images, and use feature extraction network and dual-path feature fusion network for object detection and bounding box regression, combine the key-yali algorithm and Kalman filter for trajectory matching and prediction, and monitor in real time and perform collision warning.

Benefits of technology

Real-time monitoring of the cargo transportation process, accurately perform collision warning and control, avoid dangerous situations of manual monitoring and safety problems caused by negligence, reduce labor costs for safe operation and maintenance, and improve lifting construction and safety levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088457A_ABST
    Figure CN120088457A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of tower crane hook monitoring and tracking, in particular to a tower crane hook tracking and monitoring method based on a point cloud target detection algorithm, which comprises the following steps: S1, generating a pseudo image, and carrying out feature extraction on the pseudo image by utilizing a trained feature extraction network to obtain a multi-level feature map with more than three levels and gradually reduced scales; using the trained dual-path feature fusion network to perform dual-path feature fusion on the last three-level feature maps to obtain a fused feature map, and obtaining a target bounding box detection result including a hook based on the fused feature map; s2, matching a track id by using a Hungary algorithm based on the size of an overlapping degree IOU between detection results of one or more hook target bounding boxes in the first m frames; s3, when the detection result of the hook target bounding box of the (m + 1) th frame is obtained, based on the IOU size between the # imgabs0 # and the # imgabs1 #, a Hungary algorithm is used for matching; in the (m + 1) th frame, on the basis of the hook target bounding box detection result, hook target bounding box prediction of the next frame is carried out on the basis of Kalman filtering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tower crane hook monitoring and tracking, and particularly to a tower crane hook tracking and monitoring system and method based on a point cloud target detection algorithm. Background Art

[0002] In construction sites, freight terminals, etc., large tower cranes are needed. At present, most of the handling of large heavy objects is completed by manually operating tower cranes, and safety hazards and accidents may occur due to human negligence during the work process.

[0003] The tower crane (referred to as tower crane for short), as an important tool for vertical and horizontal load transportation in cargo transportation, has the working characteristics of large lifting capacity, high height, and long amplitude. During the operation of the tower crane, in addition to possible collisions and scratches with known obstacles (such as tower cranes, buildings, etc.), it may also collide and scratch with unknown and suddenly appearing obstacles (such as people, temporarily stacked materials, etc.), posing a great safety hazard. The existing technologies mainly rely on manual monitoring and simple sensors, and cannot perform collision warning and control in real time and accurately.

[0004] At the same time, with the development of global trade, the cargo throughput of port terminals is increasing continuously, and the traditional manual handling method has been difficult to meet the growing operation requirements. The automated terminal cargo handling system has gradually attracted attention due to its high efficiency and safety.

[0005] Due to its high-altitude operation advantages, tower cranes are also widely used in the construction of high-rise buildings for vertical transportation of materials and equipment. With the increasing global demand for construction and infrastructure, it is necessary to replace manual labor with automated construction operations to solve the technical problems of existing safety hazards. Summary of the Invention

[0006] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a tower crane hook tracking and monitoring system and method based on a point cloud target detection algorithm, which is used to solve the technical problem of low safety caused by human negligence during the construction process of the tower crane hook.

[0007] To achieve the above purpose, the present invention provides a tower crane hook tracking and monitoring method based on a point cloud target detection algorithm, including:

[0008] S1: Obtain the 3D point cloud data of the construction site where the tower crane is located frame by frame and generate a pseudo-image. Use the trained feature extraction network to extract features from the pseudo-image to obtain multi-level feature maps with gradually decreasing scales above the third level. Use the trained dual-path feature fusion network to perform dual-path feature fusion on the last three-level feature maps to obtain a fused feature map. Finally, perform object classification detection and object 3D bounding box regression based on the fused feature map to obtain the detection results of the object bounding boxes including the hook.

[0009] S2: Based on the detection results of one or more hook object bounding boxes in the previous m frames Between the overlap degree IOU sizes, use the Hungarian algorithm to match the trajectory id for the detection results of the object bounding boxes in the previous m frames of the object; at the same time, according to the number of the detection results of the object bounding boxes in the previous m frames, determine the total number N of the hook objects to be tracked and the corresponding trajectories.

[0010] S3: When obtaining the detection results of the hook object bounding boxes in the (m + 1)-th frame Based on And Between the IOU sizes, use the Hungarian algorithm to match Id; in the (m + 1)-th frame, based on the detection results of the hook object bounding boxes, predict the hook object bounding boxes in the next frame based on the Kalman filter to obtain the prediction results of the hook object bounding boxes.

[0011] In each frame after the (m + 1)-th frame, based on each detection result of the hook object bounding box and the corresponding confidence of the hook object box, use the Kalman filter to predict the prediction results of the hook object bounding boxes in the next frame To obtain the prediction results of the hook object bounding boxes After that, based on And Between the IOU sizes, use the Hungarian algorithm to match out Id;

[0012] S4: Based on the detection results of the hook object bounding boxes in the previous frame And the prediction results of the hook object bounding boxes Of the IOU sizes, use the Hungarian algorithm again for id matching, and then use the prediction results of the hook object bounding boxes with the same id in the current frame and several previous frames to draw the prediction trajectory band;

[0013] S5: If the predicted trajectory band of the hook object passes through other items, an alarm is issued.

[0014] The present invention also provides a tower crane hook tracking and monitoring system based on a point cloud object detection algorithm, including:

[0015] A point cloud data acquisition module, which is used to acquire real-time point cloud data during the operation of the tower crane;

[0016] A point cloud image processing module. The point cloud image processing module obtains point cloud data from the point cloud data acquisition unit, detects the target position, predicts the position of the target in the next frame, and draws a trajectory band. At the same time, it uses a three-dimensional conversion toolbox to convert the three-dimensional point cloud data result into a two-dimensional video for real-time display and real-time monitoring of the detection and tracking results;

[0017] An alarm module. When other objects appear in the predicted trajectory band, the alarm module gives an alarm and automatically brakes or stops the tower crane.

[0018] The beneficial effects of the present invention are as follows: The method of the present invention can achieve real-time monitoring during the cargo transportation process, accurately conduct collision warnings and controls, control all working conditions within a safe range, avoid safety problems caused by the inability to react in time or negligence in dangerous situations that may occur during manual monitoring, and at the same time reduce the labor cost in terms of safety operation and maintenance, and improve the hoisting construction level and safety level. Description of the Drawings

[0019] Figure 1 It is a schematic diagram of the network structure of the 3D object detection model in the embodiment;

[0020] Figure 2 It is a schematic diagram of the structure of the pseudo-image generation network in the embodiment;

[0021] Figure 3 It is a schematic diagram of the structure of the pseudo-image feature extraction network and the dual-path feature fusion network in the embodiment;

[0022] Figure 4 It is a schematic diagram of the target trajectory motion prediction process in the embodiment;

[0023] Figure 5 It is a schematic diagram of the structure of the second embodiment. Specific Embodiments

[0024] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0025] It should be noted that the illustrations provided in this embodiment only schematically illustrate the basic concept of the present invention. Therefore, only the units related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the units in actual implementation. The types, quantities, and proportions of the units in actual implementation can be arbitrarily changed, and the unit layout type may also be more complex.

[0026] Embodiment 1

[0027] In view of the complexity of the existing tower crane operation and the variability of the goods transportation site environment, there is a risk of collision during the operation of the tower crane, especially the collision between the hook and the surrounding structures or equipment. The existing technologies mainly rely on manual monitoring and simple sensors, and cannot perform collision warning and control in real time and accurately. The present invention provides a method for tracking and monitoring the tower crane hook based on a point cloud target detection algorithm, including:

[0028] S1: Obtain the 3D point cloud data of the tower crane site frame by frame and generate a pseudo-image. Use the trained feature extraction network to extract features from the pseudo-image to obtain a multi-level feature map with the scales of three levels or more gradually decreasing; use the trained dual-path feature fusion network to perform dual-path feature fusion on the last three-level feature maps to obtain a fusion feature map, and finally perform target classification detection and target 3D bounding box regression based on the fusion feature map to obtain the target bounding box detection results including the hook.

[0029] S2: Based on the overlap degree IOU size between one or more hook target bounding box detection results in the first 3 (in this example, m = 3) frames Use the Hungarian algorithm to match the target bounding box detection results of the first 3 frames of the target with the trajectory id; at the same time, according to the number of target bounding box detection results in the first 3 frames, determine the total number N of hook targets to be tracked and the corresponding trajectories. Specifically, the number of the frame with the largest number of hook target bounding box detection results is used as the total number N of hook targets, and then the predicted trajectory bands are drawn for the hook target bounding box detection results with the same id in the first three frames.

[0030] S3: When obtaining the hook target bounding box detection result of the 4th frame Based on And Use the Hungarian algorithm to match And In the 4th frame, based on the hook target bounding box detection result, perform prediction of the hook target bounding box for the next frame based on the Kalman filter to obtain the hook target bounding box prediction result.

[0031] In each frame after the 4th frame, input each hook target bounding box detection result and its corresponding hook target box confidence into the Kalman filter (if the confidence corresponding to the hook target bounding box detection result in the previous frame is low, use the hook target bounding box prediction result obtained in the previous frame), and use the Kalman filter to predict the hook target bounding box in the next frame Obtain the hook target bounding box prediction result After that, based on Or And The IOU size between, use the Hungarian algorithm to match the id;

[0032] S4: Based on the hook target bounding box detection result in the previous frame And the hook target bounding box prediction result Of the IOU size, use the Hungarian algorithm to match the id again, and then use the hook target bounding box prediction results with the same id in the current frame and several previous frames to draw the prediction trajectory band. It can be understood that if the frames required to draw the trajectory band do not contain prediction results, use the hook target bounding box detection results in this frame;

[0033] S5: Use a 3D conversion tool to convert the 3D point cloud data result into a 2D video, and realize the real-time display and real-time monitoring of the detection (current frame hook target border) and the tracking result (prediction trajectory band); if the hook target prediction trajectory band passes by other items, give an alarm, and automatically brake or stop the tower crane.

[0034] The method of the present invention can achieve real-time monitoring during the goods transportation process, accurately conduct collision warning and control, control all working conditions within a safe range, avoid safety problems caused by the inability to react in time or negligence of dangerous situations that may occur in manual monitoring, and at the same time reduce the labor cost in safety operation and maintenance, and improve the hoisting construction level and safety level.

[0035] In this example, a 3D object detection model is used to detect the hook and output the hook target bounding box detection result. In addition, a 3D object detection model is also used to detect other items, but it is not limited to this. In some other embodiments, existing object detection models known in the prior art can also be used for implementation.

[0036] Such as Figure 1 Shown, the 3D object detection model adopted is divided into a pseudo-image generation module, a pseudo-image feature extraction network, a dual-path feature fusion network, and a dense detection head.

[0037] Such as Figure 2As shown, the pseudo-image generation module is used to preprocess the point cloud data, divide the point cloud data into Pillar structures, and perform feature encoding on the points within each pillar, including the coordinates, intensity, and relative position to the pillar center of the points. Then, through a series of multi-layer perceptrons and pooling operations, the features of the points within the pillar are aggregated into the features of the pillar, thereby converting the original point cloud data into a format suitable for processing by a 2D convolutional network.

[0038] The pseudo-image feature extraction network extracts features from the 2D image data output by the pseudo-image generation module, such as Figure 3 As shown, in this example, it uses an FPN (Feature Pyramid Network) to obtain four levels of feature maps with decreasing scales, namely the first-level feature map, the second-level feature map, the third-level feature map, and the fourth-level feature map. The second-level feature map is a two-fold downsampled feature map of the first-level feature map, the third-level feature map is a four-fold downsampled feature map of the first-level feature map, and the fourth-level feature map is an eight-fold downsampled feature map of the first-level feature map.

[0039] The dual-path feature fusion network (PAN) performs dual-path feature fusion on the second-level feature map, the third-level feature map, and the fourth-level feature map to obtain a fused feature map; in this example, the dual-path feature fusion network first performs a top-down path fusion of low-resolution feature maps on the last three levels of feature maps, then performs a bottom-up path fusion of high-resolution images, and finally fuses the high-resolution images fused by the bottom-up path. The dual-path feature fusion network fuses the feature information in multiple ways to further enrich the feature information at different scales.

[0040] Among them, the top-down path fusion of low-resolution feature maps is specifically as follows: The dual-path feature fusion network obtains the fourth-level feature map of the pseudo-image feature extraction network, performs a convolution operation with a convolution kernel of size 3×3 and a stride of 1 to obtain feature map A; then uses nearest neighbor interpolation to perform an upsampling operation on feature map A to double the size of feature map A, and then splices and fuses it with the third-level feature map, and then performs a convolution operation with a convolution kernel of size 3×3 and a stride of 1 to obtain feature map B; then uses nearest neighbor interpolation to perform an upsampling by a factor of 1 and splice and fuse it with the second-level feature map to obtain feature map C.

[0041] The bottom-up path fusion of high-resolution images is specifically as follows: Feature map C is subjected to a convolution operation with a convolution kernel of size 3×3 and a stride of 2 while being downsampled by a factor of 1 to obtain feature map D. Feature map D is spliced and fused with feature map B to obtain feature map E. Feature map E is subjected to a convolution operation with a convolution kernel of size 3×3 and a stride of 2 for downsampling, and then spliced and fused with feature map A to obtain feature map F.

[0042] Finally, for the bottom-up path fusion of high-resolution images, specifically, perform deformable convolution operations (Deformable Conv) on feature maps D, E, and F, adjust them to the same eight-fold downsampled size, and then perform splicing fusion to obtain feature map G. Feature map G is fed into the dense detection head, and feature map G is used for detection.

[0043] Finally, the dense detection head performs detection classification and regression based on the fused feature map to obtain the detection results of the hook target bounding box. The dense detection head mainly includes a target classification detection branch and a target 3D bounding box regression branch. Among them, the target classification detection branch is used to predict the category of each anchor, and the target 3D bounding box regression branch is used to output the target 3D bounding box prediction results, including position and size.

[0044] In this example, the pseudo-image generation module, the pseudo-image feature extraction network, and the dense detection head adopt the original modules of the pointpillar point cloud object detection model, and a dual-path feature fusion network is added between the pseudo-image feature extraction network and the dense detection head, which further improves the effect of point cloud object detection.

[0045] Adopt the combination of FPN (Feature Pyramid Network) and PAN (Path Aggregation Network). FPN is responsible for top-down feature fusion, while PAN introduces a bottom-up path, strengthening the transmission from the bottom layer features to the top layer, enabling the model to better handle objects of different scales. By repeatedly extracting features, the feature fusion is strengthened. Fusing features of different scales extracted from the backbone network enhances the model's detection ability for objects of different sizes. At the same time, by fusing feature maps from different stages, the model's robustness to tower crane hook occlusion and trajectory mutation can be improved.

[0046] As Figure 4 shown, in this example, taking the tower crane hook as one of the detection targets, when the 3D object detection model outputs the detection results of the first 3 frames of hook target bounding boxes, based on the overlap degree (IOU) of the first 3 frames of hook target bounding box detection results output by the 3D object detection model between them, use the Hungarian algorithm to match the id of the first 3 frames of hook target bounding box detection results That is, the maximum overlap degree of the first 3 frames of hook target bounding box detection results of the same target is matched together. After matching, the id of the first 3 frames of hook target bounding box detection results of the same target is the same. The number of tower crane hooks is equal to the number of ids. According to the total number of target detection box results in the first 3 frames, the maximum value is the total number N of targets to be tracked. Each hook target has its own exclusive id number to achieve real-time monitoring during the cargo transportation process and accurately perform collision warning and control.

[0047] When the 3D object detection model outputs the detection results of the hook object bounding boxes in the 4th frame first, match the of multiple objects with That is, based on the IOU size between and use the Hungarian algorithm to match and Then use the detection results of the hook object bounding boxes in the detected 4th frame and the speed calculated using the 3rd and 4th frames input into the Kalman filter to obtain the predicted for the next frame. And so on. In the prediction stage of each subsequent frame, for each detection result of the hook object bounding box, predict the next frame to obtain the prediction result of each hook object bounding box. Among them, the speed is calculated based on the detection results of the hook object bounding boxes in the current frame and the previous frame, and the predicted speed is calculated based on the detection result of the hook object bounding box in the current frame and the predicted result of the hook object bounding box in the next frame. The specific calculation process

[0048] In this example, the specific prediction is implemented using the ByteTrackv2 algorithm, but not limited to this; input each detection result of the hook object bounding box and the corresponding object box confidence into the ByteTrackv2 algorithm module. ByteTrackv2 is based on the detection result of the hook object bounding box or the predicted result of the hook object bounding box in the current frame, and uses the Kalman filter to predict the object bounding box in the next frame. Among them, a confidence threshold is introduced to control the input information of the Kalman filter, and then correct the predicted result of the hook object bounding box by the Kalman filter. This is to prevent abnormal detection boxes with too low confidence from interfering with the model's prediction of the tower crane hook trajectory. Specifically, when the confidence C t,id of the detection result of the hook object bounding box is greater than the set threshold (such as 0.2), the detection result of the hook object bounding box in the current frame can be input into the Kalman filter for predicting the object bounding box in the next frame to obtain the predicted result of the hook object bounding box in the next frame. Otherwise, the predicted result of the hook object bounding box obtained by the previous round of Kalman filtering and the predicted speed are input into the Kalman filter for predicting the object bounding box in the next frame.

[0049] After obtaining the predicted result of the hook object bounding box, based on the IOU size between and use the Hungarian algorithm for matching. Specifically, the current frame of multiple objects and the next frame The ones with the largest IOU are matched together. If, during the matching process, multiple predicted target bounding box results match the same detected target bounding box of the hook, the predicted target bounding box result with a higher confidence is selected to match the detected target bounding box of the hook. After matching, the IDs of the predicted target bounding box of the hook and the detected target bounding box of the hook are kept consistent.

[0050] To ensure the accuracy of target trajectory prediction, a trajectory confirmation mechanism is added. Specifically, a queue is set up to temporarily store the target detection results of the current frame and the target detection results of the previous frame After multiple predicted target bounding box results of the hook are matched with the detected target bounding box results of the hook, based on the detected target bounding box results of the hook in the previous frame and the predicted target bounding box results of the hook in terms of the IOU size, the Hungarian algorithm is used again to match the IDs, and then it is judged whether the total number of matched IDs is equal to the total number of targets N. If they are equal, a trajectory band is drawn according to the prediction If not, taking the predicted target bounding box result of the hook in the previous frame as the starting point, multiplying the predicted speed in the previous frame by the time of two frames, moving a certain distance at a uniform speed, a bounding box A is drawn. The size of bounding box A is the same as that of the target bounding box in the previous frame, and a trajectory band is drawn with this bounding box A. Here, the case where the total number of matched IDs is not equal to the total number of targets is that the total number of matched IDs is less than the total number of targets. In this example, to eliminate the influence of self-movement, a standard Kalman filter frame with a constant speed motion and a linear observation model is directly adopted. Through the added trajectory confirmation mechanism of the present invention, the trajectory can be confirmed accurately, the total number of trajectories is accurate, and the missing of predicted frames caused by the missed detection in the current frame is avoided.

[0051] It should be noted that in actual work, different confidence thresholds can be selected according to the three processes of lifting, transporting, and lowering the load, which can be manually selected according to the on-site situation; or preset and switched according to the current process.

[0052] Before step S1 of the present invention, there is step S0: training of the 3D target detection model:

[0053] S01: Collect the point cloud data for training. In this example, the pure radar scan data of the tower crane device at different processes and different time periods is used. The data is framed and the target boxes of the tower crane hook are carefully labeled to obtain the labeled data set. In this example, the panoramic data of the tower crane hook at different processes and different time periods is collected by lidar, the radar data packets are framed and saved, and each frame of data is labeled to mark the exact position of the hook on the tower crane boom. The specific labeling form is as follows: Column 1: target category (type), currently "hook"; Column 2: truncation degree (truncated), set to 0; Column 3: occlusion degree (occluded), set to 0; Column 4: observation angle (alpha), set to 0 because the camera has been registered; Columns 5-8: 2D detection box (bbox), since we only use 3D point cloud data and do not mark the 2D detection box, all are set to 0; Columns 9-11: dimensions of the 3D object, height, width, length, in meters; Columns 12-14: center coordinates (location), (x, y, z); Column 15: rotation angle (rotation y), with a value range of (-π, π). The present invention provides a target detection and tracking data set of pure radar data for common working scenarios of tower cranes for training a 3D target detection model. The data set includes three process scenarios of tower crane lifting, transporting, and lowering. After data augmentation operations, it includes a total of 10,000 radar data. The data set is labeled by following the large-scale point cloud Kitti target detection data set.

[0054] S02: Use the labeled data set to train a 3D target detection model so that the 3D target detection model can accurately obtain the exact position information of the hook on the boom. Specifically, each frame of data in the labeled data set is generalized using rotation, translation, scaling, and adding Gaussian noise to increase the diversity and generalization of the data set, and the generalized data set is put into the network for training.

[0055] Since the tower crane hook is very small relative to the background, for the training of the target classification detection branch in the 3D target detection model, the Focal Loss is used to calculate the target classification detection loss l cls for the supervision and optimization of the classification results. The target classification detection loss l cls is calculated by the formula:

[0056] P cls = softmax(Z cls )

[0057]

[0058] where Z cls is the target classification output vector of the model, and y clsis a true label vector with one-hot encoding. K represents the total number of target categories, and y cls,i indicates whether the i-th target classification category is the true label; P cls is the target classification probability distribution vector predicted by the model, where p cls,i represents the probability that the model predicts the i-th target category. is the weight for balancing positive and negative samples, and γ is the adjustment factor.

[0059] It can make the model pay more attention to small and difficult-to-classify targets by reducing the weights of easy-to-classify samples and increasing the weights of difficult-to-classify targets.

[0060] For the training of the target 3D bounding box regression branch in the 3D object detection model, the regression loss includes the localization loss and the orientation loss. The localization loss uses Smooth L1Loss, which is a smoothing of the traditional L1 loss and is used to handle the difference between the predicted box and the true box. This helps to maintain the stability of the loss function in different situations. The localization loss L loc is calculated as follows:

[0061]

[0062] In the formula, s is the traditional L1 loss between the predicted value and the true value of the detection box; both the predicted value and the true value of the detection box include the length, width, height, center point position, and orientation angle of the detection box.

[0063] The orientation loss is used to estimate the orientation of the target, that is, the angle between the forward direction of the object and the X-axis of the camera coordinate system. Cross Entropy Loss is used, and one degree in the angle is set as one category, and the decimal part is rounded to an integer. A total of 360 categories are set for 360 degrees. For multi-classification problems, first use the softmax function to convert the original output of the model into a probability distribution. Then use the Cross Entropy Loss formula to calculate the loss value. The orientation loss L dir is calculated as follows:

[0064] P dir = softmax(Z dir )

[0065]

[0066] In the formula, Z dir is the original output vector of the orientation angle of the model, and y dir is a true label vector with one-hot encoding, and y dir,i indicates whether the i-th angle category is the true label; P diris the angle classification probability distribution vector predicted by the model, where p dir,i represents the probability that the model predicts the i-th angle category.

[0067] The expression of the total loss function is:

[0068] L = β cls L cls + β loc L loc + β dir L dir

[0069] In the formula, β cls is the weight factor of the target classification detection loss, β loc is the weight factor of the localization loss, β dir is the weight factor of the orientation loss. The weight factors are the weights optimized for the training stage of the 3D object detection model. In this example, β cls = 0.3, β loc = 0.55, β dir = 0.15.

[0070] Embodiment 2

[0071] As Figure 5 shown, this embodiment provides a tower crane hook tracking and monitoring system based on a point cloud object detection algorithm, including:

[0072] A point cloud data acquisition module for acquiring real-time point cloud data during the operation of the tower crane;

[0073] A point cloud image processing module for drawing a predicted trajectory band of the hook according to the method described in Embodiment 1 to obtain a three-dimensional point cloud tracking result;

[0074] An alarm module for detecting whether there are other objects in the predicted trajectory band. If so, an alarm is issued and an instruction to automatically brake or stop the tower crane is sent.

[0075] This system can implement monitoring of the tower crane operation, avoid a series of safety problems caused by the appearance of other objects in front of the predicted operation trajectory of the tower crane hook, and improve the tower crane construction level and safety level.

[0076] In this example, the point cloud image acquisition module includes a laser radar camera, which is configured to obtain real-time point cloud data during the operation of the tower crane. The laser radar camera is located around the working environment. The operation process of the tower crane includes lifting, transporting and dropping. The point cloud image processing module includes a point cloud target detection unit 21, a point cloud target tracking post-processing unit 22 and a tracking result visualization unit 23. The point cloud target detection unit 21 is used to detect the real-time position information of the tower crane hook in real time according to the point cloud data, specifically to process the contour of the real-time point cloud data of the tower crane operation, amplify the signal, dot matrix and obtain the dot matrix target position information. In this example, the cloud target detection unit is a 3D target detection model; the point cloud target tracking post-processing unit 22 is used to predict the real-time operation trajectory of the hook according to the real-time position information of the tower crane hook; the tracking result visualization unit 23 is used to convert the three-dimensional point cloud tracking results that are not easy to observe into two-dimensional video data that is easy for the human eye to observe, and to display and monitor the detection and tracking results in real time. The alarm module 3 includes an alarm 31, which sends an alarm signal.

[0077] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A tower crane hook tracking and monitoring method based on point cloud target detection algorithm, characterized in that: include: S1: Obtain 3D point cloud data of the site where the tower crane is located frame by frame and generate a pseudo image. Use the trained feature extraction network to extract features from the pseudo image to obtain a multi-level feature map with a scale of more than three levels decreasing step by step. Use the trained dual-path feature fusion network to perform dual-path feature fusion on the last three levels of feature maps to obtain a fused feature map. Finally, perform target classification detection and target 3D bounding box regression based on the fused feature map to obtain the target bounding box detection result including the hook. S2: Based on one or more hook target bounding box detection results in the previous m frames The overlap IOU size between t=1,2,…m, uses the Hungarian algorithm to match the trajectory id with the target bounding box detection results of the target m frames before; at the same time, according to the number of target bounding box detection results of the first m frames, the total number of hook targets and corresponding trajectories N to be tracked is determined; S3: Get the bounding box detection result of the hook target in the m+1th frame When, based on and The IOU size between them is matched using the Hungarian algorithm id; in the m+1th frame, based on the hook target bounding box detection result, the hook target bounding box of the next frame is predicted based on the Kalman filter to obtain the hook target bounding box prediction result In each frame after the m+1th frame, based on each hook target bounding box detection result and the corresponding hook target box confidence, the Kalman filter is used to predict the hook target bounding box prediction result of the next frame. Get the hook target bounding box prediction result Afterwards, based on and The IOU size between them is matched using the Hungarian algorithm id; S4: Hooking the target bounding box detection results based on the previous frame Hooking target bounding box prediction results The IOU size is 100%, and the Hungarian algorithm is used again for id matching. Then, the prediction trajectory is drawn using the bounding box prediction results of the hook target with the same id in the current frame and the previous frames. S5: If the predicted trajectory of the hook target passes through other objects, an alarm is issued.

2. The method according to claim 1, characterized in that The dual-path feature fusion network first fuses the low-resolution feature maps of the last three levels from the top-down path, then fuses the high-resolution images from the bottom-up path, and finally fuses the high-resolution images from the bottom-up path.

3. The method according to claim 2, characterized in that The top-down path fusion of low-resolution feature maps is as follows: the dual-path feature fusion network obtains the fourth-level feature map of the pseudo image feature extraction network, passes it through a convolution kernel of size 3×3, and performs a convolution operation with a step size of 1 to obtain feature map A; then the feature map A is upsampled using the nearest neighbor interpolation to double the size of feature map A, and then concatenated and fused with the third-level feature map, and then a convolution operation with a step size of 1 is performed using a convolution kernel of size 3×3 to obtain feature map B; then the nearest neighbor interpolation is used to upsample it by one time and then concatenated and fused with the second-level feature map to obtain feature map C; The bottom-up path fusion of high-resolution images is as follows: a convolution kernel of size 3×3 is used on feature map C, and a convolution operation with a step size of 2 is performed while downsampling by one time to obtain feature map D. Feature map D is concatenated and fused with feature map B to obtain feature map E. A convolution kernel of size 3×3 is used on feature map E, and a convolution operation with a step size of 2 is performed on feature map E to downsample it, and then it is concatenated and fused with feature map A to obtain feature map F. Finally, the bottom-up path fusion high-resolution image is fused by performing deformable convolution operations on feature maps D, E, and F, adjusting them to a uniform eight-fold downsampling size, and then concatenating and fusion to obtain feature map G, which is sent to the dense detection head and used for detection.

4. The method according to claim 1, characterized in that: Step 4 also includes, after matching the ids, determining whether the total number of matched ids is equal to the target total number N. If they are equal, Draw the track belt. If they are not equal, take the prediction result of the hook target bounding box of the previous frame as the starting point, multiply the predicted speed of the previous frame by the time of two frames, move a certain distance at a uniform speed, draw a bounding box A, the size of which is the same as the target bounding box of the previous frame, and use this bounding box A to draw the track belt.

5. The method according to claim 1, characterized in that: In step 3, it is also included to set a confidence threshold, when the confidence C of the target bounding box detection result is hooked t,id When it is less than the set threshold, the hook target bounding box prediction result obtained by the previous frame Kalman filter is used instead. and prediction speed Make target bounding box prediction for the next frame.

6. The method according to claim 1, characterized in that For the training of the target classification detection branch in the 3D target detection model, FocalLoss is used to calculate the target classification detection loss L cls Supervise and optimize the classification results, target classification detection loss L cls The calculation formula is: P cls =softmax(Z cls ) In the formula, Z cls is the target classification output vector of the model, y cls is a one-hot encoding of the true label vector, K represents the total number of target categories, and y cls,i Indicates whether the i-th target classification category is the true label; P cls is the target classification probability distribution vector predicted by the model, where p cls,i Represents the probability that the model predicts the i-th target category is the weight to balance positive and negative samples, and γ is the adjustment factor.

7. The method according to claim 6, characterized in that For the training of target 3D bounding box regression in 3D target detection model, the regression loss includes positioning loss and orientation loss. The positioning loss L loc The calculation formula is: Where s is the traditional L1 loss between the detection box prediction value and the detection box true value; the detection box prediction value and the true value both include the length, width, height, center point position, and direction angle of the detection box.

8. The method according to claim 7, characterized in that In the training of directional angles, one degree in the angle is set as one category, and the decimal part is rounded to an integer. A total of 360 categories are set for 360 degrees. First, the softmax function is used to convert the original output of the model into a probability distribution, and then the Cross Entropy Loss formula is used to calculate the loss value. The directional loss L dir The calculation formula is: P dir =softmax(Z dir ) In the formula, Z dir is the original output vector of the model’s orientation angle, y dir is a one-hot encoding of the true label vector, y dir,i Indicates whether the i-th angle category is a true label; P dir is the angle classification probability distribution vector predicted by the model, where p dir,i Represents the probability that the model predicts the i-th angle category; The total loss function expression is: L=β cls L cls +b loc L loc +b dir L dir In the formula, β cls is the weight factor of target classification detection loss, β loc is the weight factor of the positioning loss, β dir is the weight factor of the directional loss.

9. A tower crane hook tracking and monitoring system based on point cloud target detection algorithm, characterized in that: Point cloud data acquisition module, used to obtain real-time point cloud data during the operation of the tower crane; A point cloud image processing module, used for drawing a predicted trajectory of the hook according to any method described in claims 1-8 to obtain a three-dimensional point cloud tracking result; The alarm module is used to detect whether other objects appear in the predicted trajectory. If so, an alarm will be issued and a command will be issued for the tower crane to automatically brake or stop.

10. The system according to claim 9, characterized in that The point cloud image processing module also includes a tracking result visualization unit, which is used to convert the three-dimensional point cloud tracking results into two-dimensional video data, and to display and monitor the detection and tracking results in real time.