A method and device for detecting and locating obstacles in tower crane construction areas

Through the multi-model fusion target detection algorithm and camera calibration technology, the problem of obstacle detection accuracy in tower crane construction scenarios was solved, real-time positioning and alarm of obstacles were achieved, and the construction safety and intelligence level were improved.

CN115375756BActive Publication Date: 2025-10-10JIANGSU XCMG CONSTRUCTION MACHINERY RESEARCH INSTITUTE LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210627699.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-10-10
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

Existing monitoring solutions for tower crane construction scenes cannot accurately detect the three-dimensional coordinates of pedestrians or vehicles, making it impossible to determine whether they are in dangerous areas. Target detection also suffers from missed and misidentified targets, affecting construction safety.

Method used

A multi-model fusion target detection algorithm is adopted, combined with YOLOv5, DETR and Faster-RCNN models. The ground plane image of the tower crane hook is collected through the image acquisition unit to determine the boundary of the dangerous area. The three-dimensional coordinates are converted into a two-dimensional image. Camera calibration and EPNP algorithm are used to perform target positioning and status judgment.

Benefits of technology

It improves the recall rate and position accuracy of target detection, can detect and locate obstacles in real time and at low cost, improves the safety and intelligence level of tower crane construction, and provides alarm signals to avoid danger.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375756B_ABST
    Figure CN115375756B_ABST
Patent Text Reader

Abstract

The application discloses a tower crane technical field, and relates to a tower crane construction area obstacle detection and positioning method and device, which comprises the following steps: collecting a ground plane image projected by a hook of a tower crane through an image acquisition unit installed at a root of a large arm of the tower crane, and determining a dangerous area circle boundary of tower crane construction; detecting pedestrian and vehicle targets in the ground plane image, and determining pixel point positions of the detected targets in the ground plane image; converting three-dimensional coordinates of edge points of the dangerous area on the ground plane projected by the hook in a three-dimensional coordinate system of the tower crane into two-dimensional coordinates in a pixel coordinate system of the image; and judging whether the targets are in the dangerous area circle of the tower crane construction according to the pixel point positions of the targets in the ground plane image. The application can detect whether there are obstacles such as personnel or vehicles in a dangerous area affected by the hook during tower crane construction, and ensures the safety of the tower crane construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method and a device for detecting and locating obstacles in a tower crane construction area, and belongs to the technical field of tower cranes. Background Art

[0002] With the development of society, the number of high-rise building projects in scenarios such as power construction, civil engineering, and large-scale construction sites continues to increase. Tower cranes offer numerous advantages, including high construction efficiency, a small footprint, and a wide operating range. Therefore, tower cranes have become an essential piece of construction machinery for high-rise building sites. As high-risk environments, construction sites are subject to increasing demands for site safety. However, construction sites are not only complex, but also require extensive personnel management and monitoring. With the increasing adoption of smart construction sites, video analysis and monitoring methods are gaining increasing attention. Using video to analyze construction sites can effectively improve the comprehensiveness and accuracy of site information perception for operators of tower cranes and other construction machinery.

[0003] During tower crane operation, pedestrians and vehicles are strictly prohibited from passing under heavy objects to prevent falling objects and potential injuries. At tower crane construction sites, pedestrians and vehicles frequently pass through the crane's construction environment. Completely avoiding or prohibiting this practice would significantly reduce crane efficiency, and thus prohibiting access to the crane's construction area is also undesirable. Accurately detecting when people or vehicles approach the crane's hook and heavy object area, determining their distance from the danger zone, and providing reminders or alarms as needed are crucial to ensuring safe tower crane operation.

[0004] Existing monitoring solutions for tower crane construction scenarios only detect the presence of people or vehicles in the monitoring area, but are unable to determine the specific three-dimensional coordinates of pedestrians or vehicles on the construction site where the tower crane is operating. Consequently, it is impossible to determine the pedestrian's location or the distance from the projection of the load being hoisted by the tower crane hook, making it impossible to determine whether a person is in a dangerous area. Furthermore, existing monitoring solutions often rely on a single model for target detection, resulting in numerous missed and misidentified targets in complex scenarios like construction sites. Detected targets also suffer from significant deviations between the target frame and the specific position in the image. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method and device for detecting and locating obstacles in a tower crane construction area, which can detect the distance between obstacles such as people or vehicles and the projected position of the tower crane hook on the ground, thereby determining whether there are obstacles such as people or vehicles in the dangerous area affected by the hook during tower crane construction, thereby improving the intelligence level of the tower crane monitoring system and ensuring the safety of tower crane construction.

[0006] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0007] In a first aspect, the present invention provides a method for detecting and locating obstacles in a tower crane construction area, comprising:

[0008] The image acquisition unit installed at the base of the tower crane's boom captures the ground plane image projected by the tower crane's hook and determines the boundary of the dangerous area circle where the tower crane is being constructed;

[0009] Detecting pedestrian and vehicle targets in the ground plane image, and determining pixel locations of the detected targets in the ground plane image;

[0010] Convert the three-dimensional coordinates of the edge point of the dangerous area on the ground plane projected by the hook in the three-dimensional coordinate system of the tower crane into two-dimensional coordinates in the pixel coordinate system of the image;

[0011] According to the pixel position of the target in the ground plane image, it is determined whether it is within the dangerous area circle of the tower crane construction.

[0012] Furthermore, the three-dimensional coordinate system of the tower crane takes the center point of the tower crane base located on the horizontal ground as the origin, the axis along the direction of the tower crane boom is the Zw axis, the axis perpendicular to the tower crane base and the Zw axis is the Xw axis, and the straight line parallel to the tower crane body is the Yw axis; the image acquisition unit is a monitoring camera, a dome camera or a gun camera, which is installed at the junction of the tower crane boom and the tower crane base, and rotates with the rotation of the tower crane boom, and its optical axis is always located in the plane formed by the Yw axis and the Zw axis, and the angle between its optical axis and the Yw axis of the three-dimensional coordinate system of the tower crane is α, the angle between the optical axis and the Zw axis is 90-α degrees, and the angle between the optical axis and the Xw axis is 90 degrees, and the value range of the angle α is 20 degrees to 80 degrees.

[0013] Furthermore, the dangerous area for tower crane construction is the area within the boundary of a circle formed by rotating 360 degrees with the projection point Ow2 of the tower crane hook on the ground as the center and the safe distance R between the target and the projection point Ow2 as the radius. The dangerous area for tower crane construction starts at an angle of 0 degrees in the positive direction of the Zw axis, and a marking point for the dangerous area is set at every fixed angle on the boundary of the circle, and the fixed angle is not greater than 45 degrees; there are no less than 8 marking points for the dangerous area.

[0014] Furthermore, the image acquisition unit is an industrial camera, and the tower crane three-dimensional coordinate system imaging system establishes a pinhole camera imaging model of the tower crane construction plane, and performs mapping between the three-dimensional coordinates (xw, yw, zw) of the marking points and target position points in the dangerous area on the tower crane construction plane to the two-dimensional coordinates (ui, vi) in the image coordinate system generated by the image acquisition unit. The mapping relationship is shown in the following formula:

[0015]

[0016] Where Z is the scale factor; f is the focal length of the camera; dX and dY represent the physical lengths of a pixel on the photosensitive plate in the X-axis and Y-axis directions, respectively; (u0, v0) represent the coordinates of the center of the camera photosensitive plate in the pixel coordinate system; R is the rotation matrix from the tower crane's 3D coordinate system to the camera's 3D coordinate system; t is the translation vector from the tower crane's 3D coordinate system to the camera's 3D coordinate system; K is the camera's intrinsic parameter matrix; and T is the transformation matrix.

[0017] Furthermore, for the same image sample of the tower crane construction plane area collected, three models, namely the YOLOv5 model, the DETR model and the Faster-RCNN model, were used for detection to generate three sets of different target detection results. The weights of each set of results were β1, β2, and β3, respectively. The sum of these three weights was 1, that is, β1+β2+β3=1. Each set of target detection results included the target box size, two-dimensional coordinates in the image, and target type information, where β1=0.5, β2=0.3, and β3=0.2.

[0018] Furthermore, in the target frame results detected by the YOLOv5 model, the DETR model, and the Faster-RCNN model, the center points of the target frames are (iu1, iv1), (iu2, iv2), and (iu3, iv3), respectively. The length and width of the target frames are (h1, w1), (h2, w2), and (h3, w3), respectively. The difference in the coordinates of the u-axis and v-axis directions in these three sets of coordinates is less than 5% of the image width and height, and the difference in the height and width of the target frame is less than 20% of the height and width of the target frame. It is considered that the detected target is the same obstacle. At this time, the following formula is used to calculate the position of the corresponding target in the image:

[0019]

[0020]

[0021] In the above formula, C ou and C ov are the pixel coordinates of the detected obstacle on the u-axis and v-axis in the image respectively;

[0022] If any model misses an object, resulting in only two sets of detection results for the same object, the weights of each set are β1 and β2, and the sum of these two weights is 1, that is, β1+β2=1. In this case, the position of the corresponding object in the image is calculated using the following formula, where β1=0.6, β2=0.4, the weight of the model ranked first is β1, and the weight of the model ranked second is β2;

[0023]

[0024]

[0025] In the above formula, C ou and C ov are the pixel coordinates of the detected obstacles on the u-axis and v-axis in the image, respectively.

[0026] Furthermore, the following steps are used to convert the three-dimensional coordinates of the edge point of the dangerous area on the ground plane projected by the hook in the three-dimensional coordinate system of the tower crane into two-dimensional coordinates in the pixel coordinate system of the image:

[0027] Calculate the transformation matrix T from the tower crane's three-dimensional coordinate system to the camera's coordinate system: Use the Rodriguez formula to calculate the rotation matrix R between the two coordinate systems based on the rotation vector rvec. Combined with the three-dimensional translation vector tvec between the two coordinate systems, the transformation matrix T between the two coordinate systems is obtained using the following formula:

[0028]

[0029] Calculate the coordinates Pc of the target in the camera coordinate system: Use the following formula to calculate the coordinates of the target in the camera coordinate system:

[0030]

[0031] Calculate the coordinates Pi of the target in the pixel coordinate system:

[0032]

[0033] Furthermore, the following steps are used to obtain the camera intrinsic parameter matrix K and the rotation vectors rvec and tvec from the tower crane 3D coordinate system to the camera coordinate system:

[0034] Monitoring camera debugging: Install the monitoring camera at the point where it connects to the tower crane body under the tower crane boom, and debug the monitoring camera with different focal lengths, different camera image resolutions, and camera angles to obtain the camera parameters such as focal length, image resolution, and angle that meet monitoring requirements;

[0035] Offline calibration of surveillance cameras: Using Zhang Zhengyou's camera calibration algorithm, 20 checkerboard images are used to calibrate the surveillance cameras. The intrinsic parameter matrix K and the distortion parameter vector D of the surveillance camera are obtained. The distortion parameter vector D contains three radial distortion parameters k1, k2, and k3 and two tangential distortion parameters p1 and p2.

[0036] Installing a camera in a tower crane: The surveillance camera is installed at the junction of the tower crane boom and the tower crane base, and rotates with the rotation of the tower crane boom;

[0037] Adjusting camera parameters: adjust the angle between the camera and the Y axis, the focal length of the camera and other parameters, so that the range that the camera can shoot covers the construction range of the tower crane that needs to be monitored;

[0038] Designated target detection: four people or other designated targets stand at four designated positions on the tower crane construction ground plane that needs to be monitored, and the coordinates of the four people in the tower crane three-dimensional coordinate system (xw1, 0, zw1), (xw2, 0, zw2), (xw3, 0, zw3), (xw4, 0, zw4) are manually measured and obtained;

[0039] Image target position detection: a target position detection algorithm based on multi-model fusion is used to detect and obtain the corresponding pixel coordinates of the four people in the image (Cou1, Cov1), (Cou2, Cov2), (Cou3, Cov3), (Cou4, Cov4);

[0040] Calculate the rotation vector and translation vector from the tower crane three-dimensional coordinate system to the camera coordinate system: according to the obtained coordinates of the four people in the tower crane three-dimensional coordinate system (xw1, 0, zw1), (xw2, 0, zw2), (xw3, 0, zw3), (xw4, 0, zw4) and the corresponding pixel coordinates in the image (Cou1, Cov1), (Cou2, Cov2), (Cou3, Cov3), (Cou4, Cov4), as well as the intrinsic matrix K and the distortion parameter vector D of the camera, the EPNP algorithm is used to calculate and obtain the rotation vector rvec and the translation vector tvec from the tower crane three-dimensional coordinate system to the camera coordinate system.

[0041] Further, the target positioning and state determination in the image are carried out according to the following steps:

[0042] Obtain the two points p1 and p2 closest to the target position point among the circle mark points of the dangerous area: calculate the distances between the 90 mark points constituting the dangerous area circle and the target position point and sort them to obtain the closest point p1 and the second closest point p2; obtain the distance d1 between the target position point pt and the p1 point, and the distance d2 between the target position point pt and the p2 point;

[0043] Calculate the equivalent dangerous distance: calculate the distance R1 between the hook ground projection point and the closest point p1, and the distance R2 between the hook ground projection point and the second closest point p2, and calculate the safety distance safe_dist of the hook ground projection point corresponding to the target position point pt according to the following formula:

[0044]

[0045] Determine the state of the target detected in the image: calculate the distance Rt between the projection point of the hook on the ground and the target position point pt; if Rt is greater than safe_dist, it means that the target is outside the dangerous area of ​​the hook projection and the target is in a safe state; if Rt is less than safe_dist, it means that the target is inside the dangerous area of ​​the hook projection and the target is in a dangerous state.

[0046] The second invention provides a tower crane construction area obstacle detection and positioning device, including a processor and a memory, wherein the processor is used to implement the tower crane construction area obstacle detection and positioning method as described in any one of the above items when executing the computer program stored in the memory.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] The present invention uses a target detection algorithm based on multi-model fusion of deep learning to integrate the detection capabilities of multiple models to improve the recall rate of target detection and the accuracy of target position. Discrete marker points are used to mark the range of the circle boundary of the dangerous construction area with the projection point of the tower crane hook on the ground as the center and the safe distance as the radius. In addition, the dangerous area boundary marker points and the detected target position on the tower crane construction ground plane are mapped to the image taken by the camera through camera calibration and EPNP and other algorithms to determine whether the detected target is within the dangerous area of ​​the tower crane construction. The camera is used to capture images and process the images to conveniently detect targets in the tower crane construction ground plane in real time and at low cost, and effectively determine whether targets such as pedestrians and vehicles in the construction site are within the dangerous area of ​​the tower crane construction. It effectively promotes the intelligent level of information processing of the camera-based tower crane monitoring system and improves the efficiency of tower crane construction. The present invention can detect targets such as pedestrians and vehicles and effectively locate them by only using a camera. When these targets are within the dangerous area of ​​tower crane construction, an alarm signal is provided to the operator or the system. This can significantly improve the intelligence level of the monitoring system in the construction site where the tower crane is constructed, and can also effectively improve the safety of the tower crane construction, with good effect and low cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a side view of the three-dimensional coordinate system of the tower crane provided in the first embodiment of the present invention;

[0050] Figure 2 This is a top view of the three-dimensional coordinate system of the tower crane provided in the first embodiment of the present invention;

[0051] Figure 3 This is a schematic diagram of a pinhole camera imaging model of a tower crane construction site provided by the first embodiment of the present invention;

[0052] Figure 4This is a schematic diagram of corresponding points on the image plane of the boundary points of the dangerous area on the tower crane construction ground plane provided by the first embodiment of the present invention;

[0053] Figure 5 This is a flowchart of ground level target detection and dangerous area determination for tower crane construction provided by the first embodiment of the present invention.

[0054] In the figure: 1. Image acquisition unit; 2. Tower crane boom; 3. Hook trolley; 4. Hook; 5. Hoisted object; 6. Obstacle; 7. Safety boundary; 8. Top of tower crane. DETAILED DESCRIPTION

[0055] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0056] Example 1:

[0057] A method for detecting and locating obstacles in a tower crane construction area, see Figure 1-2 First, a three-dimensional coordinate system for the tower crane is established, with the center point of the tower crane base on horizontal ground as the origin. The axis along the direction of the tower crane's boom is the Zw axis, and the axis perpendicular to the tower crane base and the Zw axis is the Xw axis. The line parallel to the tower crane body is the Yw axis. Detected objects such as vehicles or people on the ground are converted to three-dimensional coordinates in the tower crane's three-dimensional coordinate system based on their position in the image. The linear distance between the detected object and the tower crane hook, as well as the distance along the three coordinate axes, is then determined. This allows for determination of whether objects such as vehicles or people are within the dangerous construction zone of the tower crane.

[0058] The monitoring camera is installed below the tower crane cab. The coordinates of the optical center of the monitoring camera are [0, h, 0], where h is the vertical height between the bottom of the tower crane cab and the base. The side view of the tower crane's three-dimensional coordinate system is as follows: Figure 1 The top view of the tower crane's three-dimensional coordinate system is shown in Figure 2 When the surveillance camera's visible area covers the scene from the hook's starting position to the end position, the projection point of the hook on the ground is the center and the specified safety distance is the radius of all circles.

[0059] See also Figure 1 , O w1The image acquisition unit 1 is the origin of the tower crane's three-dimensional coordinate system. The image acquisition unit 1 transmits the image to an industrial control computer for processing via a fiber optic switch or other device. In this embodiment, the image acquisition unit 1 is used to capture images of the ground within the tower crane's construction area. The image acquisition unit 1 is mounted at the junction of the tower crane's boom 2 and the crane's base, and rotates with the boom 2. A hook trolley 3 is provided at the bottom of the boom 2. The hook trolley 3 can move along the boom 2 while carrying the weight of the hook and any objects loaded. A hook 4 is provided at the bottom of the hook trolley 3, and a hoisted object 5 is hung from the hook 4. Obstacles 6 exist on the ground within the tower crane's construction area, primarily including pedestrians and vehicles, two of the most common obstacles in tower crane construction environments. The optical axis of the image acquisition unit 1 is always located within the plane formed by the Yw and Zw axes. The optical axis forms an angle α with the Yw axis of the tower crane's three-dimensional coordinate system, an angle of 90-α degrees with the Zw axis, and a 90-degree angle with the Xw axis. The value range of the angle α is 20 degrees to 80 degrees. The image acquisition unit 1 can be an industrial camera, or an infrared camera such as a dome camera or a box camera.

[0060] See also Figure 2 The Xw axis is perpendicular to the main arm of the tower crane, with the origin at the center of the tower crane base. The Zw axis is the axis in the same direction as the main arm of the tower crane. The top of the tower crane 8 is the main body of the tower crane viewed from the top of the tower crane. The safety boundary 7 is a circular area with a radius of a specified length R and the projection of the hook on the ground as the center. This circular area is the dangerous area for tower crane construction. If there are pedestrians or vehicles in the dangerous area, an alarm should be given to the operator. If the tower crane hook has lifted a heavy object and is in the process of moving, the operator can be prompted to take measures such as stopping the machine to avoid danger to people or vehicles. Set the projection point of the hook on the ground to O w2 , the object is far from the projection point O w2 The safety distance is R, with O w2 The area within the circle with O as the center and R as the radius is the dangerous area for tower crane construction. If a person or vehicle appears inside the dangerous area, an alarm signal will be triggered. w2 The center of the circle is R, and the radius is rotated 360 degrees. Starting from the positive direction of the Zw axis at 0 degrees, a danger zone marker is set every 4 degrees, resulting in a total of 90 danger zone markers S. They are labeled S0, S1, S2, and S89, corresponding to the markers at 0 degrees, 4 degrees, 8 degrees, ..., 352 degrees, and 356 degrees, respectively.

[0061] In order to determine the position of the person or vehicle detected on the tower crane construction plane in the three-dimensional coordinate system of the tower crane through the image collected by the image acquisition unit, and thus determine whether the detected person or vehicle is in a dangerous area, it is necessary to use a camera to collect images of the tower crane construction plane. In order to accurately establish the corresponding relationship between the coordinates corresponding to the tower crane construction plane and each pixel in the image, it is necessary to establish a pinhole camera imaging model of the tower crane construction plane, such as Figure 3 shown.

[0062] See also Figure 3 , Pw point is a point on the ground plane of the tower crane construction, this point can be the corresponding point of the vehicle or personnel on the ground plane. Pi is the point corresponding to Pw point in the image plane. Oc is the origin of the camera's three-dimensional coordinate system, and the Xc axis, Yc axis, and Zc axis are the three mutually perpendicular coordinate axes of the camera coordinate system; Oi is the origin of the image's two-dimensional coordinate system, and the u axis and v axis are the two mutually perpendicular coordinate axes of the image coordinate system; O w1 The origin of the tower crane's three-dimensional coordinate system is represented by the Xw, Yw, and Zw axes, which are perpendicular to each other. The line connecting points Pw, Pi, and Oc is the projection of point Pw on the crane's construction plane onto the image plane. The projection line is obtained by connecting points Pw and Oc, and the intersection of the projection line and the image plane is the image point of point Pw on the image.

[0063] The following steps are used to obtain the camera intrinsic parameter matrix K and the rotation vectors rvec and tvec from the tower crane 3D coordinate system to the camera coordinate system:

[0064] (a) Industrial camera debugging: Install the industrial camera at the point where it connects to the tower crane's main body below the boom. Debug the industrial camera with different focal lengths, image resolutions, and angles to obtain the camera's focal length, image resolution, and angle parameters that meet monitoring requirements.

[0065] (b) Offline calibration of industrial cameras: Using the Zhang Zhengyou camera calibration algorithm, 20 checkerboard images are used to calibrate the industrial camera. The intrinsic parameter matrix K and the distortion parameter vector D are obtained. The distortion parameter vector D contains three radial distortion parameters k1, k2, and k3, and two tangential distortion parameters p1 and p2.

[0066] (c) Installing the camera in the tower crane: Install the industrial camera in the required position and angle of the tower crane.

[0067] (d) Adjusting camera parameters: Adjust the angle between the camera and the Y axis and the focal length of the camera so that the camera can capture the entire construction range of the tower crane to be monitored.

[0068] (e) Designated target detection: Four persons or other designated targets are placed at four designated positions on the construction site of the tower crane to be monitored. The coordinates of these four persons in the three-dimensional coordinate system of the tower crane are manually measured and obtained: (xw1, 0, zw1), (xw2, 0, zw2), (xw3, 0, zw3), and (xw4, 0, zw4).

[0069] (f) Target position detection in the image: A target position detection algorithm based on multi-model fusion is used to detect and obtain the corresponding pixel coordinates (Cou1, Cov1), (Cou2, Cov2), (Cou3, Cov3), and (Cou4, Cov4) of the four people in the image.

[0070] Calculate the rotation vector and translation vector from the three-dimensional coordinate system of the tower crane to the camera coordinate system: According to the coordinates of the four personnel in the three-dimensional coordinate system of the tower crane obtained in step (e) and step (f) (xw1, 0, zw1), (xw2, 0, zw2), (xw3, 0, zw3), (xw4, 0, zw4) and the corresponding pixel coordinates in the image (Cou1, Cov1), (Cou2, Cov2), (Cou3, Cov3), (Cou4, Cov4), as well as the intrinsic parameter matrix K and distortion parameter vector D of the camera, use the EPNP algorithm to calculate the rotation vector rvec and translation vector tvec from the three-dimensional coordinate system of the tower crane to the camera coordinate system.

[0071] Specifically, the following steps are used to convert the three-dimensional coordinates of the point on the ground plane projected by the hook in the three-dimensional coordinate system of the tower crane into the two-dimensional coordinates in the pixel coordinate system of the image:

[0072] (a) Calculate the transformation matrix T from the tower crane's 3D coordinate system to the camera's coordinate system: Use the Rodriguez formula to calculate the rotation matrix R between the two coordinate systems based on the rotation vector rvec. Combined with the 3D translation vector tvec between the two coordinate systems, the transformation matrix T between the two coordinate systems is obtained using the following formula:

[0073]

[0074] (b) Calculate the coordinates Pc of the target in the camera coordinate system: Use the following formula to calculate the coordinates of the target in the camera coordinate system:

[0075]

[0076] (c) Calculate the coordinates Pi of the target in the pixel coordinate system:

[0077]

[0078] The conversion relationship between the three-dimensional coordinates (xw, yw, zw) of point Pw in the three-dimensional coordinate system of the tower crane and the two-dimensional coordinates (ui, vi) in the image coordinate system is shown in the following formula:

[0079]

[0080] Among them, Z is called the scale factor; dX and dY represent the physical lengths of a pixel on the photosensitive plate in the X-axis and Y-axis directions respectively; (u0, v0) represent the coordinates of the center of the camera photosensitive plate in the pixel coordinate system; f is the focal length of the camera; R is the rotation matrix from the tower crane's 3D coordinate system to the camera's 3D coordinate system; t is the translation vector from the tower crane's 3D coordinate system to the camera's 3D coordinate system; K is the camera's intrinsic parameter matrix; and T is the transformation matrix.

[0081] In addition, Z in the above formula is the Z coordinate of point Pw in the camera coordinate system. Multiplying this 1 / Z and the three-dimensional coordinates of point Pw will give the normalized coordinates of P in the camera coordinate system, which is located on the plane with Z = 1 in front of the camera.

[0082] The transformation matrix composed of R and t represents the rigid body transformation, while the transformation from the camera coordinate system to the image coordinate system is the perspective transformation, and the transformation from the image coordinate system to the pixel coordinate system is the affine transformation.

[0083] For the same image sample of a tower crane construction area, we used the YOLOv5 model, the DETR model, and the Faster-RCNN model to detect the target. Three different sets of target detection results were generated, each with weights β1, β2, and β3. The sum of these three weights is 1, i.e., β1 + β2 + β3 = 1. Each set of target detection results includes the target box size, 2D coordinates in the image, and target type information. β1 = 0.5, β2 = 0.3, and β3 = 0.2.

[0084] In the target box detection results of the three models, the YOLOv5 model, the DETR model, and the Faster-RCNN model, the center points of the target boxes are (iu1, iv1), (iu2, iv2), and (iu3, iv3), respectively, and the length and width of the target boxes are (h1, w1), (h2, w2), and (h3, w3), respectively. In addition, the difference between the u-axis and v-axis coordinates in these three sets of coordinates is less than 5% of the image width and height, and the difference in the height and width of the target box is less than 20% of the height and width of the target box. It is considered that the detected target is the same obstacle. At this time, the following formula is used to calculate the position of the corresponding target in the image:

[0085]

[0086]

[0087] In the above formula, C ou and C ov are the pixel coordinates of the detected obstacles on the u-axis and v-axis in the image, respectively.

[0088] If any model misses an object, resulting in only two sets of detected positions, lengths, and widths for the same object, the weights for each set are β1 and β2, and the sum of these two weights is 1, i.e., β1 + β2 = 1. In this case, the position of the corresponding object in the image is calculated using the following formula.

[0089]

[0090]

[0091] In the above formula, C ou and C ov are the pixel coordinates of the detected obstacle on the u-axis and v-axis in the image, respectively. Without loss of generality, in this embodiment, β1 = 0.6 and β2 = 0.4. The weight of the model ranked first is β1, and the weight of the model ranked second is β2.

[0092] Please refer to Figure 4 The danger zone of a tower crane construction project is a circle on the ground plane of the tower crane construction area. However, the angle between the camera located under the tower crane boom and the ground plane is not perpendicular, so the corresponding image of the danger zone circle on the ground plane is projected. The danger zone circle captured by the camera is not a circle, but an irregular ellipse. The following scheme is used to determine whether the detected target is in the hook projection danger zone in the image with projected transformation:

[0093] (a) Obtain the two points p1 and p2 closest to the target location in the danger zone circle: Calculate the distances between the 90 points that make up the danger zone circle and the target location and sort them to obtain the closest point p1 and the second closest point p2 to the target location; obtain the distance d1 between the target location point pt and point p1, and obtain the distance d2 between the target location point pt and point p2;

[0094] (b) Calculate the equivalent dangerous distance: Calculate the distance R1 between the projection point of the hook on the ground and the closest point p1, and the distance R2 between the projection point of the hook on the ground and the next closest point p2. Calculate the safe distance safe_dist of the projection point of the hook on the ground corresponding to the target position point pt according to the following formula:

[0095]

[0096] (c) Determine the state of the target detected in the image: calculate the distance Rt between the projection point of the hook on the ground and the target position point pt; if Rt is greater than safe_dist, it means that the target is outside the dangerous area of ​​the hook projection and the target is in a safe state; if Rt is less than safe_dist, it means that the target is inside the dangerous area of ​​the hook projection and the target is in a dangerous state.

[0097] See also Figure 5 ,Based on the above functional modules, the process of obstacle target detection and positioning method in the tower crane construction area is as follows:

[0098] (a) Input the hook projection point coordinates and safety distance: Based on the real-time length of the tower crane hook trolley from the tower crane body (arm_length) and other actual working conditions, generate the real-time tower crane hook projection coordinates on the ground and the safety distance R between the projection point and the obstacle target;

[0099] (b) Generate the boundary of the danger zone and 90 landmarks: With the projection coordinates of the hook on the ground as the center and the safe distance R between the object and the projection point of the tower crane hook on the ground as the radius, generate the three-dimensional coordinates of the danger zone circle of the tower crane construction and the 90 corresponding danger zone landmarks; if the length arm_length of the hook trolley from the tower crane body changes, then with the projection of the hook on the ground as the center point and the safe distance R as the radius, generate a new danger zone circle; and sample the new danger zone circle to generate 90 new landmarks;

[0100] (c) Capturing an image of the tower crane construction ground plane with a camera: Setting appropriate parameters such as focal length and resolution for the camera to capture an image of the tower crane construction ground plane; then simultaneously performing steps (d) and (e);

[0101] (d) Obtain the hook projection coordinates and the corresponding coordinates of the 90 markers on the image: Based on the camera's intrinsic parameter matrix K and transformation matrix T, the target 3D coordinate to image coordinate system coordinate conversion process is used to convert the 3D coordinates of the hook projection coordinates and the 90 dangerous area markers in the tower crane 3D coordinate system into the corresponding 2D coordinates in the camera image;

[0102] (e) Detecting the target in the image and outputting its position coordinates in the image: Using a multi-model fusion target position detection algorithm, the target position in the image is detected, its position in the image is determined, and the position coordinates of the target in the image are output;

[0103] (f) determining whether the position of the target in the image is within the dangerous area: based on the position coordinates of the detected target in the image outputted in step (d), the hook projection coordinates outputted in step (e), and the corresponding coordinates of the 90 landmarks on the image, a method for determining the position of the target in the image and determining its state is used to determine whether the target is within the dangerous area corresponding to the projection of the hook on the ground in the image;

[0104] (g) If the position of the target in the image is within the dangerous area, then proceed to step (h); if the position of the target in the image is outside the dangerous area, then proceed to step (a);

[0105] (h) Providing an alarm signal to the operator: If the target in the image is within the danger zone, the object dropped from the hook could pose a serious threat to the obstacle. Therefore, an alarm signal is provided to the operator, allowing the operator to take appropriate measures to avoid the danger. After this is completed, the process proceeds to step (a).

[0106] The tower crane's 3D coordinate system implemented in this invention can also incorporate 3D coordinate data of other types of obstacles. For 3D coordinate data of obstacles such as buildings within the tower crane's construction environment, collected by sensors such as lidar, a conversion matrix can be calculated from the lidar coordinate system to the camera coordinate system. This data is then transformed so that the 3D coordinates of the buildings and obstacles are in the same coordinate system as those of pedestrians and vehicles located at the construction site. This effectively integrates the data and provides a reliable basis for automatic trajectory planning during tower crane construction.

[0107] Example 2:

[0108] The embodiment of the present invention further provides a tower crane construction area obstacle detection and positioning device, which can implement the tower crane construction area obstacle detection and positioning method described in the first embodiment, including a processor and a storage medium;

[0109] The storage medium is used to store instructions;

[0110] The processor is configured to operate according to the instructions to execute the steps of the following method:

[0111] The image acquisition unit installed at the base of the tower crane's boom captures the ground plane image projected by the tower crane's hook and determines the boundary of the dangerous area circle where the tower crane is being constructed;

[0112] Detecting pedestrian and vehicle targets in the ground plane image, and determining pixel locations of the detected targets in the ground plane image;

[0113] Convert the three-dimensional coordinates of the edge point of the dangerous area on the ground plane projected by the hook in the three-dimensional coordinate system of the tower crane into two-dimensional coordinates in the pixel coordinate system of the image;

[0114] According to the pixel position of the target in the ground plane image, it is determined whether it is within the dangerous area circle of the tower crane construction.

[0115] In this embodiment, the three-dimensional coordinate system of the tower crane takes the center point of the tower crane base located on the horizontal ground as the origin, the axis along the direction of the tower crane boom is the Zw axis, the axis perpendicular to the tower crane base and the Zw axis is the Xw axis, and the straight line parallel to the tower crane body is the Yw axis; the image acquisition unit is a monitoring camera, a dome camera or a gun camera, which is installed at the junction of the tower crane boom and the tower crane base, and rotates with the rotation of the tower crane boom, and its optical axis is always located in the plane formed by the Yw axis and the Zw axis, and the angle between its optical axis and the Yw axis of the three-dimensional coordinate system of the tower crane is α, the angle between its optical axis and the Zw axis is 90-α degrees, and the angle between its optical axis and the Xw axis is 90 degrees, and the value range of the angle α is 20 degrees to 80 degrees.

[0116] In this embodiment, the dangerous area for tower crane construction is the area within the boundary of a circle formed by rotating 360 degrees with the projection point Ow2 of the tower crane hook on the ground as the center and the safe distance R between the target and the projection point Ow2 as the radius. The dangerous area for tower crane construction takes the positive direction of the Zw axis as the starting angle of 0 degrees, and a marking point of the dangerous area is set at every fixed angle on the boundary of the circle, and the fixed angle is not greater than 45 degrees; there are no less than 8 marking points in the dangerous area.

[0117] In this embodiment, the image acquisition unit is an industrial camera, and the tower crane three-dimensional coordinate system imaging system establishes a pinhole camera imaging model of the tower crane construction plane, and performs mapping between the three-dimensional coordinates (xw, yw, zw) of the marker points and target position points in the dangerous area on the tower crane construction plane to the two-dimensional coordinates (ui, vi) in the image coordinate system generated by the image acquisition unit. The mapping relationship is shown in the following formula:

[0118]

[0119] Where Z is the scale factor; f is the focal length of the camera; dX and dY represent the physical lengths of a pixel on the photosensitive plate in the X-axis and Y-axis directions, respectively; (u0, v0) represent the coordinates of the center of the camera photosensitive plate in the pixel coordinate system; R is the rotation matrix from the tower crane's 3D coordinate system to the camera's 3D coordinate system; t is the translation vector from the tower crane's 3D coordinate system to the camera's 3D coordinate system; K is the camera's intrinsic parameter matrix; and T is the transformation matrix.

[0120] In this embodiment, for the same collected image sample of the tower crane construction plane area, three models, namely the YOLOv5 model, the DETR model and the Faster-RCNN model, are used for detection to generate three sets of different target detection results. The weights of each set of results are β1, β2, and β3, respectively. The sum of these three weights is 1, that is, β1+β2+β3=1. Each set of target detection results includes the target box size and the two-dimensional coordinates in the image, and the target type information, where β1=0.5, β2=0.3, and β3=0.2.

[0121] In this embodiment, the YOLOv5 model, the DETR model, and the Faster-RCNN model detect the target frame results. The center points of the target frames are (iu1, iv1), (iu2, iv2), and (iu3, iv3), respectively. The length and width of the target frame are (h1, w1), (h2, w2), and (h3, w3), respectively. The difference in the coordinates of the u-axis and v-axis directions in these three sets of coordinates is less than 5% of the image width and height, and the difference in the height and width of the target frame is less than 20% of the height and width of the target frame. It is considered that the detected target is the same obstacle. At this time, the following formula is used to calculate the position of the corresponding target in the image:

[0122]

[0123]

[0124] In the above formula, C ou and C ov are the pixel coordinates of the detected obstacle on the u-axis and v-axis in the image respectively;

[0125] If any model misses an object, resulting in only two sets of detection results for the same object, the weights of each set are β1 and β2, and the sum of these two weights is 1, that is, β1+β2=1. In this case, the position of the corresponding object in the image is calculated using the following formula, where β1=0.6, β2=0.4, the weight of the model ranked first is β1, and the weight of the model ranked second is β2;

[0126]

[0127]

[0128] In the above formula, C ou and C ov are the pixel coordinates of the detected obstacles on the u-axis and v-axis in the image, respectively.

[0129] In this embodiment, the following steps are used to convert the three-dimensional coordinates of the edge point of the dangerous area on the ground plane projected by the hook in the three-dimensional coordinate system of the tower crane into two-dimensional coordinates in the pixel coordinate system of the image:

[0130] Calculate the transformation matrix T from the tower crane's three-dimensional coordinate system to the camera's coordinate system: Use the Rodriguez formula to calculate the rotation matrix R between the two coordinate systems based on the rotation vector rvec. Combined with the three-dimensional translation vector tvec between the two coordinate systems, the transformation matrix T between the two coordinate systems is obtained using the following formula:

[0131]

[0132] Calculate the coordinates Pc of the target in the camera coordinate system: Use the following formula to calculate the coordinates of the target in the camera coordinate system:

[0133]

[0134] Calculate the coordinates Pi of the target in the pixel coordinate system:

[0135]

[0136] In this embodiment, the following steps are used to obtain the camera intrinsic parameter matrix K and the rotation vectors rvec and tvec from the tower crane 3D coordinate system to the camera coordinate system:

[0137] Monitoring camera debugging: Install the monitoring camera at the point where it connects to the tower crane body under the tower crane boom, and debug the monitoring camera with different focal lengths, different camera image resolutions, and camera angles to obtain the camera parameters such as focal length, image resolution, and angle that meet monitoring requirements;

[0138] Offline calibration of surveillance cameras: Using Zhang Zhengyou's camera calibration algorithm, 20 checkerboard images are used to calibrate the surveillance cameras. The intrinsic parameter matrix K and the distortion parameter vector D of the surveillance camera are obtained. The distortion parameter vector D contains three radial distortion parameters k1, k2, and k3 and two tangential distortion parameters p1 and p2.

[0139] Installing a camera in a tower crane: The surveillance camera is installed at the junction of the tower crane boom and the tower crane base, and rotates with the rotation of the tower crane boom;

[0140] Adjust camera parameters: adjust the angle between the camera and the Y axis, the focal length and other parameters of the camera so that the camera can capture the construction range of the tower crane to be monitored;

[0141] Designated target detection: Four people or other designated targets are placed at four designated locations on the construction site of the tower crane to be monitored. The coordinates of these four people in the three-dimensional coordinate system of the tower crane are manually measured and obtained: (xw1, 0, zw1), (xw2, 0, zw2), (xw3, 0, zw3), and (xw4, 0, zw4).

[0142] Target position detection in the image: Using a target position detection algorithm based on multi-model fusion, the corresponding pixel coordinates of the four people in the image are detected and obtained: (Cou1, Cov1), (Cou2, Cov2), (Cou3, Cov3), and (Cou4, Cov4);

[0143] Calculate the rotation vector and translation vector from the three-dimensional coordinate system of the tower crane to the camera coordinate system: According to the obtained coordinates of the four personnel in the three-dimensional coordinate system of the tower crane (xw1, 0, zw1), (xw2, 0, zw2), (xw3, 0, zw3), (xw4, 0, zw4) and the corresponding pixel coordinates in the image (Cou1, Cov1), (Cou2, Cov2), (Cou3, Cov3), (Cou4, Cov4), as well as the intrinsic parameter matrix K and distortion parameter vector D of the camera, the EPNP algorithm is used to calculate the rotation vector rvec and translation vector tvec from the three-dimensional coordinate system of the tower crane to the camera coordinate system.

[0144] In this embodiment, the following steps are performed to locate the target in the image and determine its status:

[0145] Obtain the two points p1 and p2 closest to the target location in the danger zone circle markers: Calculate the distances between the 90 markers that make up the danger zone circle and the target location and sort them to obtain the closest point p1 and the second closest point p2 to the target location; obtain the distance d1 between the target location point pt and point p1, and obtain the distance d2 between the target location point pt and point p2;

[0146] Calculate the equivalent dangerous distance: Calculate the distance R1 between the projection point of the hook on the ground and the closest point p1, and the distance R2 between the projection point of the hook on the ground and the second closest point p2. Calculate the safe distance safe_dist of the projection point of the hook on the ground corresponding to the target position point pt according to the following formula:

[0147]

[0148] Determine the state of the target detected in the image: calculate the distance Rt between the projection point of the hook on the ground and the target position point pt; if Rt is greater than safe_dist, it means that the target is outside the dangerous area of ​​the hook projection and the target is in a safe state; if Rt is less than safe_dist, it means that the target is inside the dangerous area of ​​the hook projection and the target is in a dangerous state.

[0149] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0150] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0151] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0152] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0153] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for detecting and locating obstacles in a tower crane construction area, characterized in that: include: The image acquisition unit installed at the base of the tower crane's boom captures the ground plane image projected by the tower crane's hook and determines the boundary of the dangerous area circle where the tower crane is being constructed; Detecting pedestrian and vehicle targets in the ground plane image, and determining pixel locations of the detected targets in the ground plane image; Convert the three-dimensional coordinates of the edge point of the dangerous area on the ground plane projected by the hook in the three-dimensional coordinate system of the tower crane into two-dimensional coordinates in the pixel coordinate system of the image; According to the pixel position of the target in the ground plane image, it is determined whether it is within the dangerous area circle of the tower crane construction; Follow these steps to locate the target in the image and determine its status: Obtain the two points p1 and p2 closest to the target location in the danger zone circle markers: Calculate the distances between the 90 markers that make up the danger zone circle and the target location and sort them to obtain the closest point p1 and the second closest point p2 to the target location; obtain the distance d1 between the target location point pt and point p1, and obtain the distance d2 between the target location point pt and point p2; Calculate the equivalent dangerous distance: Calculate the distance R1 between the projection point of the hook on the ground and the closest point p1, and the distance R2 between the projection point of the hook on the ground and the second closest point p2. Calculate the safe distance safe_dist of the projection point of the hook on the ground corresponding to the target position point pt according to the following formula: Determine the state of the target detected in the image: calculate the distance Rt between the projection point of the hook on the ground and the target position point pt; if Rt is greater than safe_dist, it means that the target is outside the dangerous area of ​​the hook projection and the target is in a safe state; if Rt is less than safe_dist, it means that the target is inside the dangerous area of ​​the hook projection and the target is in a dangerous state.

2. The tower crane construction area obstacle detection and positioning method according to claim 1 is characterized in that: The three-dimensional coordinate system of the tower crane takes the center point of the tower crane base located on the horizontal ground as the origin, the axis along the direction of the tower crane boom is the Zw axis, the axis perpendicular to the tower crane base and the Zw axis is the Xw axis, and the straight line parallel to the tower crane body is the Yw axis; the image acquisition unit is a monitoring camera, a dome camera or a gun camera, which is installed at the junction of the tower crane boom and the tower crane base, and rotates with the rotation of the tower crane boom, and its optical axis is always located in the plane formed by the Yw axis and the Zw axis, and the angle between the optical axis and the Yw axis of the three-dimensional coordinate system of the tower crane is α, the angle between the optical axis and the Zw axis is 90-α degrees, and the angle between the optical axis and the Xw axis is 90 degrees, and the value range of the angle α is 20 degrees to 80 degrees.

3. The tower crane construction area obstacle detection and positioning method according to claim 2 is characterized in that: The dangerous area for tower crane construction is the area within the boundary of a circle formed by rotating 360 degrees with the projection point Ow2 of the tower crane hook on the ground as the center and the safe distance R between the target and the projection point Ow2 as the radius. The dangerous area for tower crane construction takes the positive direction of the Zw axis as the starting angle of 0 degrees, and a marking point of the dangerous area is set at every fixed angle on the boundary of the circle, and the fixed angle is not greater than 45 degrees; there are no less than 8 marking points for the dangerous area.

4. The method for detecting and locating obstacles in a tower crane construction area according to claim 3, wherein: The image acquisition unit is an industrial camera, and the tower crane three-dimensional coordinate system imaging system is used to establish a pinhole camera imaging model of the tower crane construction plane. The three-dimensional coordinates (xw, yw, zw) of the marker points and target position points in the dangerous area on the tower crane construction plane are mapped to the two-dimensional coordinates (ui, vi) in the image coordinate system generated by the image acquisition unit. The mapping relationship is shown in the following formula: Where Z is the scale factor; f is the focal length of the camera; dX and dY represent the physical lengths of a pixel on the photosensitive plate in the X-axis and Y-axis directions, respectively; (u0, v0) represent the coordinates of the center of the camera photosensitive plate in the pixel coordinate system; R is the rotation matrix from the tower crane's 3D coordinate system to the camera's 3D coordinate system; t is the translation vector from the tower crane's 3D coordinate system to the camera's 3D coordinate system; K is the camera's intrinsic parameter matrix; and T is the transformation matrix.

5. The method for detecting and locating obstacles in a tower crane construction area according to claim 4, wherein: For the same image sample of the tower crane construction plane area collected, three models, namely the YOLOv5 model, the DETR model and the Faster-RCNN model, are used for detection to generate three different sets of target detection results. The weights of each set of results are β1, β2, and β3, respectively. The sum of these three weights is 1, that is, β1+β2+β3=1. Each set of target detection results includes the target box size, two-dimensional coordinates in the image, and target type information, where β1=0.5, β2=0.3, and β3=0.

2.

6. The method for detecting and locating obstacles in a tower crane construction area according to claim 5, wherein: In the target frame results detected by the YOLOv5 model, DETR model, and Faster-RCNN model, the center points of the target frames are (iu1, iv1), (iu2, iv2), and (iu3, iv3), respectively. The length and width of the target frames are (h1, w1), (h2, w2), and (h3, w3), respectively. The difference between the u-axis and v-axis coordinates in these three sets of coordinates is less than 5% of the image width and height, and the difference between the height and width of the target frame is less than 20% of the height and width of the target frame. It is considered that the detected target is the same obstacle. At this time, the following formula is used to calculate the position of the corresponding target in the image: In the above formula, C ou and C ov are the pixel coordinates of the detected obstacle on the u-axis and v-axis in the image respectively; If any model misses an object, resulting in only two sets of detection results for the same object, the weights of each set are β1 and β2, and the sum of these two weights is 1, that is, β1+β2=1. In this case, the position of the corresponding object in the image is calculated using the following formula, where β1=0.6, β2=0.4, the weight of the model ranked first is β1, and the weight of the model ranked second is β2; In the above formula, C ou and C ov are the pixel coordinates of the detected obstacles on the u-axis and v-axis in the image, respectively.

7. The method for detecting and locating obstacles in a tower crane construction area according to claim 5, wherein: Use the following steps to convert the three-dimensional coordinates of the edge point of the dangerous area on the ground plane projected by the hook in the three-dimensional coordinate system of the tower crane into the two-dimensional coordinates in the pixel coordinate system of the image: Calculate the transformation matrix T from the tower crane's three-dimensional coordinate system to the camera's coordinate system: Use the Rodriguez formula to calculate the rotation matrix R between the two coordinate systems based on the rotation vector rvec. Combined with the three-dimensional translation vector tvec between the two coordinate systems, the transformation matrix T between the two coordinate systems is obtained using the following formula: Calculate the coordinates Pc of the target in the camera coordinate system: Use the following formula to calculate the coordinates of the target in the camera coordinate system: Calculate the coordinates Pi of the target in the pixel coordinate system:

8. The tower crane construction area obstacle detection and positioning method according to claim 7 is characterized in that: The following steps are used to obtain the camera intrinsic parameter matrix K and the rotation vectors rvec and tvec from the tower crane 3D coordinate system to the camera coordinate system: Monitoring camera debugging: Install the monitoring camera at the point where it connects to the tower crane body under the tower crane boom, and debug the monitoring camera with different focal lengths, different camera image resolutions, and camera angles to obtain the camera parameters such as focal length, image resolution, and angle that meet monitoring requirements; Offline calibration of surveillance cameras: Using Zhang Zhengyou's camera calibration algorithm, 20 checkerboard images are used to calibrate the surveillance cameras. The intrinsic parameter matrix K and the distortion parameter vector D of the surveillance camera are obtained. The distortion parameter vector D contains three radial distortion parameters k1, k2, and k3 and two tangential distortion parameters p1 and p2. Installing a camera in a tower crane: The surveillance camera is installed at the junction of the tower crane boom and the tower crane base, and rotates with the rotation of the tower crane boom; Adjust camera parameters: adjust the angle between the camera and the Y axis, the focal length and other parameters of the camera so that the camera can capture the construction range of the tower crane to be monitored; Designated target detection: Four people or other designated targets are placed at four designated locations on the construction site of the tower crane to be monitored. The coordinates of these four people in the three-dimensional coordinate system of the tower crane are manually measured and obtained: (xw1, 0, zw1), (xw2, 0, zw2), (xw3, 0, zw3), and (xw4, 0, zw4). Target position detection in the image: Using a target position detection algorithm based on multi-model fusion, the corresponding pixel coordinates of the four people in the image are detected and obtained: (Cou1, Cov1), (Cou2, Cov2), (Cou3, Cov3), and (Cou4, Cov4); Calculate the rotation vector and translation vector from the three-dimensional coordinate system of the tower crane to the camera coordinate system: According to the obtained coordinates of the four personnel in the three-dimensional coordinate system of the tower crane (xw1, 0, zw1), (xw2, 0, zw2), (xw3, 0, zw3), (xw4, 0, zw4) and the corresponding pixel coordinates in the image (Cou1, Cov1), (Cou2, Cov2), (Cou3, Cov3), (Cou4, Cov4), as well as the intrinsic parameter matrix K and distortion parameter vector D of the camera, the EPNP algorithm is used to calculate the rotation vector rvec and translation vector tvec from the three-dimensional coordinate system of the tower crane to the camera coordinate system.

9. A device for detecting and locating obstacles in a tower crane construction area, characterized in that: The method comprises a processor and a memory, wherein the processor is used to implement the method for detecting and locating obstacles in a tower crane construction area as described in any one of claims 1 to 8 when executing a computer program stored in the memory.

Citation Information

Patent Citations

  • Hoisting safe distance detection method based on deep learning

    CN109019335A