Intelligent tower crane human-computer interaction control system and method for complex scene
By combining sensor and vision camera technologies, precise control of tower cranes in complex construction sites has been achieved, solving the problem of insufficient operating precision of traditional tower cranes and improving the system's adaptability and safety.
Patent Information
- Application Number
- CN202511446235.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Traditional tower crane operation relies on the operator's visual judgment, making it difficult to accurately place the hoisted object. Existing automated systems lack positioning accuracy in complex construction site environments and cannot effectively cope with multi-tower crane operations and complex scenarios.
By combining sensor detection from accelerometers and gyroscopes with image processing technology from vision cameras, and through gesture recognition and wireless channel matching, the camera focal length is automatically adjusted to recognize the operator's posture and gestures, and to generate and transmit control commands.
It improves the accuracy of gesture recognition and the robustness of the system, adapts to complex construction sites, simplifies operation procedures, improves work efficiency, and ensures construction safety.
Smart Images

Figure CN120922764B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a tower crane control system, more particularly to an intelligent tower crane human-machine interaction control system and method for complex scenarios. BACKGROUND
[0002] In the field of construction, tower cranes serve as important vertical transportation equipment, responsible for lifting and placing materials and components. Traditionally, the operation of a tower crane relies on the driver manually operating the controller in the cab to complete actions such as lifting, amplitude changing, and rotating. However, in actual operation, especially for the requirement of precise placement of the hoisted object, relying solely on the visual judgment of the tower crane driver often fails to achieve the desired accuracy. Typically, ground command personnel are required to use a walkie-talkie or hand signals to guide the tower crane driver for fine-tuning, ensuring that the hoisted object accurately reaches the designated location.
[0003] In recent years, with the advancement of automation technology and intelligent control systems, automatic operation of tower cranes has emerged. These systems can achieve automatic lifting to a certain extent, but still face challenges in the final precise positioning phase, mainly due to the inertia of the tower crane itself and positioning accuracy issues. In addition, a single control system cannot cope with complex construction site environments, such as multi-tower crane operations and complex scenarios, which add additional difficulties to automatic lifting.
[0004] Therefore, it is necessary to design a new method that improves the accuracy of gesture recognition, enhances the robustness of the system, effectively adapts to the complex and variable conditions of the construction site, simplifies the operation process, improves work efficiency, and ensures construction safety. SUMMARY
[0005] The present application aims to overcome the shortcomings of the prior art and provide an intelligent tower crane human-machine interaction control system and method for complex scenarios.
[0006] To achieve the above-mentioned purpose, the present application adopts the following technical solution: an intelligent tower crane human-machine interaction control method for complex scenarios, comprising:
[0007] Through a specific wireless channel matching process, the gesture detection device is bound to the corresponding tower crane; wherein the gesture detection device includes an accelerometer and a gyroscope;
[0008] The gesture detection device installed on the hand is used to capture the gesture changes of the operator, determine the control instruction, and obtain the sensor detection result;
[0009] A vision camera is used to capture images, and the operator facing the corresponding tower crane is determined by analyzing the direction the operator is facing and the size of the operator in the vision image;
[0010] The image processing technology is used to position and identify the gestures of the specific dress operator by automatically adjusting the camera focus, monitor and identify the predetermined mode gesture and its direction, and determine the control instruction to obtain the camera detection result;
[0011] The final control instruction is selected from the sensor detection result and the camera detection result to obtain the selected instruction; wherein the sensor detection result is transmitted through a wireless network mode; and the camera detection result is transmitted through a Modbus protocol;
[0012] The control signal is generated according to the selected instruction, transmitted to the tower crane control system wirelessly, and the corresponding action is performed.
[0013] Further technical solutions thereof are as follows: the final control instruction is selected from the sensor detection result and the camera detection result to obtain the selected instruction, including:
[0014] When the sensor detection result and the camera detection result both exist, the sensor detection result is selected as the final control instruction to obtain the selected instruction; when only one of the sensor detection result and the camera detection result exists, the existing result is selected as the final control instruction to obtain the selected instruction.
[0015] Further technical solutions thereof are as follows: the gesture detection device is bound to the corresponding tower crane through a specific wireless channel matching process, including:
[0016] The gesture detection device button is pressed to make the gesture detection device enter the pairing mode, and the gesture detection device starts to send matching data packets one by one from the first channel until the response of the tower crane is received, so as to complete the binding of the gesture detection device and the corresponding tower crane.
[0017] Further technical solutions thereof are as follows: the gesture detection device installed on the hand is used to capture the gesture change of the operator, determine the control instruction, and obtain the sensor detection result, including:
[0018] An initial coordinate system is set;
[0019] The rotation angle is calculated by using the gyroscope and the initial coordinate system, and the hand orientation is determined and verified to be effective;
[0020] When the hand orientation is effective, the accelerometer is used to judge whether there is an effective swing action according to the set condition;
[0021] When there is an effective swing action, the gesture change condition is determined based on the hand orientation and the swing action, and the corresponding tower crane control instruction is determined according to the pre-defined mapping relationship to obtain the sensor detection result.
[0022] Further technical solutions are as follows: The initial coordinate system is set, including:
[0023] The initial coordinate system is set according to the current posture of the operator, wherein the positive direction of the z-axis corresponds to the direction of the palm facing upwards, the positive direction of the y-axis corresponds to the direction of the four fingers pointing to the tower crane, and the positive direction of the x-axis corresponds to the direction of the thumb.
[0024] Further technical solutions are as follows: The rotation angle is calculated by using the gyroscope and the initial coordinate system, and it is determined and verified whether the palm direction is valid, including:
[0025] The rotation angle is calculated by using the gyroscope to measure the rotation angular velocity around the x, y and z axes and by integrating, so as to determine the palm direction, and when the palm direction meets a specific threshold condition, the palm direction is valid.
[0026] Further technical solutions are as follows: The set conditions include that the acceleration change amplitude exceeds a threshold, the swing switching time meets the rapidity requirement, and a swing mode appears at least three times.
[0027] Further technical solutions are as follows: The image processing technology is used to automatically adjust the camera focal length to locate and identify the posture of the specific dressed operator, monitor and identify the gesture action and its direction of the predetermined mode, and determine the control instruction, so as to obtain the camera detection result, including:
[0028] The camera focal length is automatically adjusted, and the operator is located step by step until the operator is located, so as to obtain a positioning picture;
[0029] The specific dressed operator in the picture is identified and located by using a pre-trained human body detection algorithm, so as to obtain a target picture;
[0030] Based on the angle of the visual camera and the body symmetry, the posture and standing direction of the operator in the target picture are analyzed and determined;
[0031] The action of the arm and hand in the target picture is monitored based on the posture and standing direction of the operator, so as to determine whether there is a gesture action meeting the predetermined mode;
[0032] If there is a gesture action meeting the predetermined mode, the position of the left hand of the operator in the target picture is identified, so as to determine the direction of the gesture;
[0033] The gesture action and the direction are mapped into specific tower crane control instructions, so as to obtain the camera detection result.
[0034] Further technical solutions are as follows: Based on the angle of the visual camera and the body symmetry, the posture and standing direction of the operator in the target picture are analyzed and determined, including:
[0035] Based on the visual camera angle and human body symmetry, the posture and side angle of the operator in the camera visual angle in the target picture are calculated by using a convolution neural network to obtain the posture and standing direction of the operator.
[0036] The application also provides an intelligent tower crane human-computer interaction control system for a complex scene, which comprises:
[0037] A sensor binding unit is configured to bind the gesture detection device and the corresponding tower crane through a specific wireless channel matching process, wherein the gesture detection device comprises an accelerometer and a gyroscope.
[0038] A sensor detection unit is configured to capture the gesture change of the operator by using the gesture detection device installed on the hand, determine the control instruction, and obtain the sensor detection result.
[0039] A personnel positioning unit is configured to capture the image by using the visual camera, determine the operator of the corresponding tower crane by analyzing the direction faced by the operator and the size of the operator in the visual image.
[0040] A visual detection unit is configured to position and identify the posture of the specific dressed operator by automatically adjusting the camera focal length by using the image processing technology, monitor and identify the gesture action and direction of the predetermined mode, determine the control instruction, and obtain the camera detection result.
[0041] A selection unit is configured to select the final control instruction from the sensor detection result and the camera detection result to obtain the selected instruction, wherein the sensor detection result is transmitted through a wireless network mode, and the camera detection result is transmitted through a Modbus protocol.
[0042] A control unit is configured to generate a control signal according to the selected instruction, transmit the control signal to the tower crane control system through wireless transmission, and execute the corresponding action.
[0043] Compared with the prior art, the present application has the beneficial effects that: the present application realizes the double recognition and verification of gesture action by combining the sensor detection of accelerometer and gyroscope with the image processing technology of visual camera, thereby significantly improving the accuracy of gesture recognition and the robustness of the system. Specifically, the system first binds the gesture detection device with the corresponding tower crane through specific wireless channel matching, and uses these sensors to capture the gesture changes of the operator to generate control instructions; at the same time, the visual camera automatically adjusts the focal length to locate and identify the operator with specific dress and his / her predetermined mode gesture action, and further confirms the control instructions. Finally, the most reliable control instruction is selected from the two sources of sensors and cameras, and transmitted to the tower crane control system through appropriate communication protocol to execute corresponding action. This method not only effectively adapts to the complex and variable conditions of the construction site, simplifies the operation process, improves the work efficiency, but also ensures the construction safety, so that the operation is more accurate and reliable.
[0044] The present application will be further described below in conjunction with the drawings and specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0046] Figure 1 The flowchart of the intelligent tower crane human-computer interaction control method for complex scenes provided by the embodiment of the present application;
[0047] Figure 2 The sub-flowchart of the intelligent tower crane human-computer interaction control method for complex scenes provided by the embodiment of the present application;
[0048] Figure 3 The sub-flowchart of the intelligent tower crane human-computer interaction control method for complex scenes provided by the embodiment of the present application;
[0049] Figure 4 The motion curve diagram of the swing detected by the accelerometer provided by the embodiment of the present application;
[0050] Figure 5 The angle diagram of the plane and the human body actually shot by the visual camera provided by the embodiment of the present application;
[0051] Figure 6 The diagram of the swing direction and swing action of the operator provided by the embodiment of the present application Figure 1 ;
[0052] Figure 7The schematic diagram of the swing direction and swing action of the operator provided by the embodiment of the present application Figure 2 ;
[0053] Figure 8 The schematic block diagram of the intelligent tower crane human-computer interaction control system for a complex scene provided by the embodiment of the present application
[0054] Figure 9 The schematic block diagram of a computer device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0056] It should be understood that, when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0057] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing the specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, unless otherwise clearly indicated by the context, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0058] It should be further understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0059] Please refer to Figure 1 , Figure 1The schematic flowchart of the intelligent tower crane human-computer interaction control method for complex scenes provided by the embodiment of the present application. The intelligent tower crane human-computer interaction control method for complex scenes is applied to a tower crane controller, which interacts with a gesture detection device and a visual camera. By combining the gesture detection device (including an accelerometer and a gyroscope) and the visual camera technology, the wireless channel matching, sensor data capture and image processing technology are used to identify the gesture action and direction of the operator, and the corresponding tower crane control instruction is generated according to the information. This method not only accurately identifies the human posture and gesture in the complex and variable construction site environment, but also automatically adjusts the camera focal length to adapt to the operator at different distances, ensuring high accuracy and robustness. When the sensor and camera detection results coexist, the sensor detection result is used as the final control instruction, further enhancing the reliability of the system. The whole process simplifies the operation process, improves the work efficiency, and at the same time guarantees the construction safety.
[0060] Figure 1 The flowchart of the intelligent tower crane human-computer interaction control method for complex scenes provided by the embodiment of the present application. As shown in Figure 1 , the method comprises the following steps S110 to S160.
[0061] S110, through a specific wireless channel matching process, the gesture detection device is bound with the corresponding tower crane; wherein the gesture detection device includes an accelerometer and a gyroscope.
[0062] In this embodiment, the gesture detection device button is pressed to make the gesture detection device enter the pairing mode, and the matching data packet is sent from the first channel set in advance until the response of the tower crane is received, so as to complete the binding of the gesture detection device and the corresponding tower crane.
[0063] In the tower crane intelligent control system, in order to ensure that the gesture detection device of the operator interacts with the correct tower crane, a specific wireless channel matching process is adopted to bind the gesture detection device (including an accelerometer and a gyroscope) with the corresponding tower crane. This process not only enhances the safety of the system, but also improves the accuracy of the operation, which is particularly important in the complex construction environment where multiple tower cranes work at the same time.
[0064] For the case where multiple tower cranes exist in the scene. The operator's orientation and distance are used as the main basis for the tower crane and the operator in visual detection. Specifically, when the gesture detection device needs to be bound with the tower crane, the operator first needs to press the button on the gesture detection device to make the device enter the pairing mode. This step indicates that the user intends to start the connection process between the device and the tower crane.
[0065] Once in pairing mode, the gesture detection device starts sending matching data packets from a pre-set first channel. These data packets contain information identifying and authenticating the gesture detection device, aiming to find and match the corresponding tower crane.
[0066] Since there can be multiple tower cranes in the same construction site, each tower crane can use a different wireless communication channel to avoid interference. Therefore, the gesture detection device will try each channel in sequence, sending matching data packets, until it finds a matching tower crane.
[0067] When a certain tower crane receives the matching data packet and confirms that it is the signal from the expected gesture detection device, it sends a response signal to the gesture detection device through the same channel. This response signal marks a successful match between the two.
[0068] Once the gesture detection device receives the response signal from the target tower crane, it indicates that the gesture detection device and the tower crane have established an effective connection, completing the binding process. After that, the gesture detection device can be used to control the corresponding tower crane.
[0069] Using the above method of binding the tower crane and the operator, an operator can flexibly operate different tower cranes at different times, and the relationship between the tower crane and the operator is clearly corresponding, and there is no confusion in control.
[0070] This binding method based on wireless channel matching effectively solves the problem of correctly pairing the gesture detection device and the tower crane in a complex construction scene, ensuring the safety and accuracy of the operation. In addition, by using accelerometers and gyroscopes to capture the gesture changes of the operator, the operation of the tower crane is more intuitive and convenient, and the robustness and adaptability of the entire system are also improved.
[0071] S120, capture the gesture changes of the operator using the gesture detection device installed on the hand, determine the control instruction to obtain the sensor detection result.
[0072] In this embodiment, the sensor detection result refers to the gesture changes captured by the gesture detection device (including accelerometers and gyroscopes) installed on the operator's hand, and the specific tower crane control instructions generated based on these gesture changes. Specifically, this result contains the following information:
[0073] Validity verification of palm orientation: based on the data provided by the gyroscope, the rotation angle of the hand relative to the initial coordinate system is calculated to determine the current orientation of the palm and verify whether it meets the pre-set validity conditions.
[0074] Effective swing motion recognition: using accelerometer data to determine whether there is a swing motion that meets certain conditions, such as acceleration change amplitude, swing switching time, and repetition pattern, etc.
[0075] Final gesture change and control instruction: combining the palm orientation and effective swing motion, according to the pre-defined mapping relationship to determine the specific tower crane control instruction, such as lifting, luffing or slewing operation, etc.
[0076] Therefore, the sensor detection result is essentially the result of converting human gestures into machine-understandable and executable operation commands.
[0077] In an embodiment, referring to Figure 2 The above step S120 can include steps S121-S124.
[0078] S121, set the initial coordinate system.
[0079] In this embodiment, the initial coordinate system refers to a reference coordinate system set for accurate analysis of gesture actions. This coordinate system is established based on the specific posture of the operator wearing the gesture detection device to ensure consistency and accuracy of subsequent gesture analysis. The specific definition is as follows:
[0080] z-axis positive direction: corresponding to the direction of the palm up. This is usually the state when the operator is ready to start gesture interaction with the palm completely upwards.
[0081] y-axis positive direction: corresponding to the direction of the four fingers pointing to the tower crane. This means that when the operator stands facing the tower crane, the direction of the fingers is the positive direction of the y-axis.
[0082] x-axis positive direction: corresponding to the direction of the thumb. In most cases, the direction of the thumb is perpendicular to the direction of the four fingers, forming a right-handed orthogonal coordinate system.
[0083] The setting of this coordinate system helps to accurately describe the rotation and translation in the gesture action, which is crucial for correctly understanding and executing the operator's intention. For example, during gesture recognition, by comparing the current position of the hand relative to the initial coordinate system, various predetermined gesture actions can be accurately recognized, and then the corresponding tower crane control instructions are triggered.
[0084] Specifically, the initial coordinate system is set according to the current posture of the operator, where the z-axis positive direction corresponds to the direction of the palm up, the y-axis positive direction corresponds to the direction of the four fingers pointing to the tower crane, and the x-axis positive direction corresponds to the direction of the thumb.
[0085] Before starting gesture detection, an initial coordinate system needs to be set according to the current posture of the operator. The establishment of this coordinate system is crucial for subsequent gesture recognition.
[0086] Specifically, the positive direction of the z-axis corresponds to the direction of the palm facing upwards, the positive direction of the y-axis corresponds to the direction of the four fingers pointing towards the tower crane, and the positive direction of the x-axis corresponds to the direction of the thumb. The setting of such a coordinate system is based on the principle of ergonomics, ensuring that different gestures can be accurately recognized.
[0087] S122, calculate the rotation angle using the gyroscope and the initial coordinate system, determine and verify whether the palm orientation is valid.
[0088] In this embodiment, the gyroscope is used to measure the rotation angular velocity around the x, y, and z axes, and integration is performed to calculate the rotation angle, determine the palm orientation, and when the palm orientation meets a certain threshold condition, the palm orientation is valid.
[0089] The gyroscope is used to measure the rotation angular velocity around the x, y, and z axes, and integration is performed to calculate the rotation angle of the palm relative to the initial coordinate system.
[0090] According to these angle data, the specific orientation of the palm is determined. For example, when is 90 degrees, and are both 0 degrees, it indicates that the palm is facing backward; is -90 degrees, and are both 0 degrees, it indicates that the palm is facing forward, etc.
[0091] A certain threshold condition (such as ±10 degrees) is set to verify whether the palm orientation is valid. If the actual orientation of the palm falls within the pre-set threshold range, it is considered that the palm orientation is valid, otherwise it needs to be adjusted or continue to be detected.
[0092] S123, when the palm orientation is valid, use the accelerometer to determine whether there is an effective swing motion according to the set conditions.
[0093] Specifically, after confirming that the palm orientation is valid, the next step is to use the accelerometer to detect whether there is an effective swing motion.
[0094] The set conditions include but are not limited to:
[0095] The change amplitude of acceleration must exceed a certain threshold to exclude slight motion interference;
[0096] The time of swing switching should meet the requirement of rapidity, i.e. the swing should be a relatively rapid process;
[0097] At least three repeated swing patterns appear to ensure that it is an intentional operation rather than an accidental motion.
[0098] Only when all the above conditions are met, it is considered that there is an effective swing action.
[0099] S124, when there is an effective swing action, determining the gesture change case based on the palm orientation and the swing action, and determining the corresponding tower crane control instruction according to the pre-defined mapping relationship to obtain the sensor detection result.
[0100] In this embodiment, based on the determined palm orientation and effective swing action, the specific case of gesture change is further analyzed.
[0101] According to the pre-defined mapping relationship (for example, hand up and down swing represents the hook up, hand down and up swing represents the hook down, etc.), the gesture change is converted into specific tower crane control instruction.
[0102] The final sensor detection result contains the tower crane action information intended to be executed by the operator, which will be sent to the intelligent control system of the tower crane to drive the corresponding motor to complete the specified action.
[0103] In summary, the step S120 realizes the conversion from the operator's gesture to the tower crane control instruction through accurate gesture capture and analysis, greatly improving the convenience and accuracy of tower crane operation. This method is not only suitable for single tower crane operation, but also has important application value in multi-tower crane collaborative operation scene.
[0104] In this embodiment, the gesture detection device in the above step S120 is composed of an accelerometer and a gyroscope, which is specially used to control the action adjustment of the tower crane according to the gestures of the construction personnel in the intelligent human-computer interaction scene. Considering that the operation of the tower crane is limited to three actions of lifting, amplitude changing and rotating, and only needs to run at the lowest speed when fine-tuning at the target position of the hoisted object, the gestures required to be performed by the construction personnel are relatively simple, including up, down, trolley forward, trolley backward, clockwise rotation, counterclockwise rotation and emergency stop.
[0105] The gesture detection device is attached to the palm of the operator by a fixed belt. The device is equipped with a button. When the button is pressed, the system considers that the intelligent human-computer interaction process starts; and when the button is released, it is considered to end the interaction with the tower crane intelligent control system and trigger an emergency stop signal to make the tower crane control system stop all current actions.
[0106] Based on the six actions of the tower crane (including emergency stop) and the habitual posture of the human body, it is agreed to use the right palm as the detection area, and the operator faces the tower crane that needs to be interacted. The specific gesture definition is as follows:
[0107] Hand horizontal swing: hand left swing represents clockwise rotation of the tower crane; hand right swing represents counterclockwise rotation of the tower crane.
[0108] Hand up-down swing: Swing hand up controls the hook to go up; swing hand down controls the hook to go down.
[0109] Hand front-back swing: Swing hand forward represents the trolley going forward; swing hand backward represents the trolley going backward.
[0110] In order to determine the initial orientation of the hand, at the moment when the start interaction button is pressed, the operator is required to keep the hand up and the four fingers pointing towards the tower crane. At this moment, the gesture detection device will take this state as the initial state, and all subsequent gesture recognitions are based on this.
[0111] In the initial state, the gesture detection device sets:
[0112] The positive direction of the z-axis corresponds to the direction of the hand up;
[0113] The positive direction of the y-axis corresponds to the direction of the four fingers pointing towards the tower crane;
[0114] The positive direction of the x-axis corresponds to the direction of the thumb.
[0115] On this basis, the orientation of the hand is determined by the rotation angle around the three axes. In the initial state, the rotation angle of all axes is 0 degrees. During the interaction, the angular velocity data provided by the gyroscope is used to calculate the angle change of the hand relative to the initial state by time integration, that is , and Specifically, taking rotation around the x-axis as an example, the rotation angle of the hand around the x-axis is determined by Similarly, the y and z axes are and When is 90 degrees, and are both 0 degrees, the hand is facing backward, that is, facing the face; when is -90 degrees.
[0116] Because the operator will not accurately place the orientation at 90 degrees, a threshold is set for the orientation. The threshold is set to 10 degrees, so when is between 80 and 100 degrees, and are both between -10 and 10 degrees, it is considered that the hand is facing backward.
[0117] In order to analyze the specific swing action, the system relies on the data of the accelerometer. As shown in Figure 4 When the hand swings back and forth, the acceleration will fluctuate between the maximum and minimum values. However, due to the existence of non-target actions such as natural shaking in actual operation, the system sets strict constraints on the change of acceleration:
[0118] The maximum and minimum values must exceed a certain threshold to filter out slight movements.
[0119] The switching time between maximum and minimum values should be less than a certain threshold, ensuring that the speed of the movement is as expected.
[0120] Three similar swing changes need to be detected consecutively to be considered as one valid tower crane interaction movement.
[0121] Once the palm orientation and corresponding swing movement are confirmed, the gesture detection device will package this information according to the preset protocol and send it to the intelligent control system of the tower crane through the wireless data transmission module, thereby triggering the corresponding action instruction. When the swing stops or the button is released, the tower crane will stop the current operation accordingly.
[0122] S130, capturing images with a vision camera, determining the operator corresponding to the tower crane by analyzing the direction the operator faces and the size of the operator in the vision image.
[0123] In this embodiment, in order to ensure the accuracy and safety of tower crane control, it is crucial to correctly identify which operator is interacting with a specific tower crane in the complex and variable construction site environment. The vision camera system is used to capture site images, and through analysis of these images, it is determined which operator is corresponding to which tower crane.
[0124] The system for detecting human gestures through a vision camera mainly consists of a vision camera and a processor. The vision camera is installed outside the tower crane cab and, with the assistance of a pan-tilt head, its field of view can cover most of the scene below the tower crane jib, thus effectively monitoring the behavior of ground personnel. Since the behavior of ground personnel needs to be detected, a vision camera with high zoom function is selected so that the details of the ground personnel's actions can still be clearly captured even at a high tower crane. In addition, the vision camera has already implemented a hook tracking function, which can track the position near the target position when the load is lifted to the vicinity of the target position, thereby facilitating the detection of people near the target position.
[0125] Specifically, first, the vision camera is installed outside the tower crane cab, and the angle of view can be adjusted using a pan-tilt head to cover most of the area below the tower crane jib. Considering the need to monitor the behavior of ground construction personnel, the camera is equipped with a high zoom function to clearly capture human movements at different distances. In addition, the camera has already implemented a hook tracking function, which can automatically adjust the angle of view to the vicinity of the target position, further improving detection accuracy and efficiency.
[0126] Before starting detection, due to the height difference of tower cranes and the change of the position of the hoisted object, the size of the human body in the camera's field of view may vary. Therefore, appropriate zoom processing is needed first. The system gradually adjusts from the far focus to the near focus, and after each adjustment, it tries to detect whether the human body exists. If not found, continue to reduce the focus until a suitable human body model is found or all focus options are traversed.
[0127] Using a graphics processing framework such as YOLOv, the system can recognize and track the person in the picture. To distinguish the operators who are truly involved in the control of the tower crane, these personnel are required to wear clothes of a specific color.
[0128] As shown in Figure 5 , the posture of the human body in the picture is determined, especially the direction the operator is facing. This involves calculating the inclination angle of the human body in the vertical direction , and the angle that breaks the overall symmetry when the human body is sideways . A convolutional neural network (CNN) trained model is used to estimate these angle values, so as to more accurately understand the operator's posture and orientation.
[0129] According to the orientation of the operator's body (i.e., the direction he / she is facing), it can be inferred whether they are preparing to interact with a certain tower crane. For example, if a person is facing a certain tower crane, it is likely to indicate that he / she intends to issue instructions to that tower crane.
[0130] If there are multiple tower cranes operating simultaneously within the same scene, in addition to considering the orientation, the distance of the operator relative to each tower crane also needs to be evaluated. Typically, the operator will try to get as close as possible to the tower crane they want to control, which means that in the visual image, that operator will appear larger than the personnel near other tower cranes. This size difference can help the system more accurately associate the operator with the corresponding tower crane.
[0131] In practical applications, there may be multiple construction personnel appearing in the same picture at the same time. At this time, combined with the above orientation judgment and size comparison method, irrelevant personnel can be effectively filtered out to lock the real operator. In addition, when multiple tower cranes are working together, not only rely on visual information, but also combine sensor data (such as wireless communication channel binding) to enhance the accuracy of matching, to ensure that each operator can only control the tower crane paired with them.
[0132] In summary, by capturing images through visual cameras and deeply analyzing factors such as the orientation and relative size of the operator, accurate identification of tower crane operators can be effectively achieved, improving the safety and reliability of human-machine interaction. This process not only enhances the fault tolerance of the system, but also provides strong support for efficient management in complex construction environments.
[0133] S140, using image processing technology, locating and identifying the posture of the specific dressed operator by automatically adjusting the camera focal length, monitoring and identifying the predetermined mode of gesture action and its direction, and determining the control instruction to obtain the camera detection result.
[0134] In the embodiment, the camera detection result refers to the analysis result of the operator gesture position, shape and dynamic information obtained after processing the image data captured by the visual camera.
[0135] In an embodiment, referring to Figure 3 The above step S140 can include steps S141-S146.
[0136] S141, automatically adjusting the camera focal length, gradually searching until the operator is located to obtain the positioning picture.
[0137] In the embodiment, the positioning picture refers to the picture captured by the camera after automatic zooming, which contains one or more possible operators.
[0138] This process first starts from a far focal length, gradually adjusts to a near focal length, and tries to detect whether a human body exists after each adjustment. If not found, continue to reduce the focal length until a suitable human body model is found or all focal length options are traversed. Once a required human model (i.e. operator) is found, the frame picture is considered as the positioning picture.
[0139] Specifically, the first step of detecting gestures by a visual camera is to automatically adjust the camera focal length, gradually search until the operator is located to obtain the positioning picture. Because the height of the tower crane and the position of the hoisted object are different, the size of the human body in the visual camera will also change, so it is necessary to first adjust the zoom to make the human model within the range that can be detected by the algorithm. In the case of not knowing the distance of the hoisted object, the method of traversing the zoom is used to determine the appropriate zoom range. The specific logic is to start from a far focal length, gradually reduce the focal length, and use the algorithm to detect the human body after each reduction of the focal length. If no human body is detected, continue to reduce the focal length until a human model is found. If no human model is detected during the entire zooming process, it can be determined that there is no operator in the current scene. On the contrary, if a human body is successfully found, the focal length is used as the subsequent focal length. In the subsequent detection process, if a human body is not found again in a certain detection, the zooming search is performed again to ensure that the human body can always be accurately detected.
[0140] S142, using a pre-trained human body detection algorithm to identify and locate the specific dressed operator in the picture to obtain the target picture.
[0141] In the embodiment, the target picture refers to a picture captured by the visual camera and adjusted to a proper focal length, which focuses on the operator and his / her gesture action, so as to accurately recognize and analyze the operation instruction.
[0142] On the basis of locating the picture, the picture of the operator wearing clothes of a specific color is further screened out. This step uses a graphic processing framework such as YOLOv combined with a pedestrian algorithm model for detection. In this way, the operator who really needs to interact with the tower crane can be effectively distinguished from many construction personnel.
[0143] Specifically, the pedestrian algorithm model is mainly used to detect the workers near the hook. Since there can be multiple workers near the hook, in order to accurately identify the operator who interacts with the tower crane, it is agreed that the operator wears clothes of a specific color. After detecting the workers, the processor will make a judgment according to the clothing color of the workers, so as to determine the operator from many workers to obtain the target picture.
[0144] Many construction personnel will increase the difficulty of visual detection. In order to solve this problem, the operator is distinguished by wearing clothes of a specific color. When there are multiple tower cranes in the scene, the operator corresponding to a tower crane is determined by the orientation of the operator and the size of the operator in the vision. The operator must face the tower crane when operating the corresponding tower crane, and the corresponding operator should be the closest to the corresponding tower crane, so in the visual image, the size of the operator is relatively larger. Only when the operator wears clothes of a specific color and faces the tower crane and is the closest, the operator is considered to be the operator corresponding to the tower crane.
[0145] S143, based on the angle of the visual camera and the body symmetry, the posture and standing direction of the operator in the target picture are analyzed and determined.
[0146] In the embodiment, based on the angle of the visual camera and the body symmetry, the posture and standing direction of the operator in the target picture are analyzed and determined.
[0147] The posture and standing direction refer to the body posture and facing direction of the operator determined according to the angle of the visual camera relative to the ground and the body side angle of the body itself
[0148] Specifically, the convolutional neural network (CNN) is used to calculate these angle values, so as to more accurately understand the posture and standing direction of the operator. For example, when the operator faces a tower crane, it indicates that he / she is preparing to issue an instruction to the tower crane.
[0149] AsFigure 5 As shown, based on the angle of the visual camera and the body symmetry, the posture and standing direction of the operator in the target picture are analyzed and determined. In actual cases, there is an angle between the plane shot by the visual camera and the human body , which will cause the human body to appear tilted in the vertical direction in the camera vision, so that the human body appears to be shorter in the camera vision. Although the human body as a whole is symmetrical, if the human body has a sideways action in the camera vision, the symmetry of the human body will be destroyed. By using a convolutional neural network to train the sideways human body, and using the trained model to obtain the angle of the sideways human body in the camera vision .
[0150] S144, based on the posture and standing direction of the operator, the motion of the arm and hand in the target picture is monitored to determine whether there is a gesture action conforming to a predetermined mode.
[0151] In this embodiment, this step is mainly to detect whether the arm swing of the operator conforms to the pre-defined gesture mode. By tracking the motion trajectory of the arm and hand, and analyzing the swing amplitude and direction, the system can judge the effectiveness of the gesture action. For example, if the motion trajectory of the hand is to move back and forth on a straight line and the swing amplitude is within the normal threshold, it is considered as an effective arm swing action.
[0152] Specifically, based on the posture and standing direction of the operator, the motion of the arm and hand in the target picture is monitored to determine whether there is a gesture action conforming to a predetermined mode. After the human body posture is determined, the arms and hands on the human body are further detected. The convolutional neural network under the Yolov framework is also used for training and detection, and the motion trajectory of the arm and hand in the visual camera is tracked and observed. If the motion trajectory of the hand is to move back and forth on a straight line and the swing amplitude is within the normal threshold, this process is considered as the swing of the arm. However, due to the tilt of the human body in the camera vision, the actual swing amplitude needs to be corrected, and the specific method is to divide the detected swing amplitude by or .
[0153] S145, if there is a gesture action conforming to a predetermined mode, the position of the left hand of the operator in the target picture is recognized to determine the direction of the gesture.
[0154] In this embodiment, since the visual detection cannot directly distinguish the specific direction of the arm swing (such as left and right or up and down), it is agreed to use the position of the left hand to indicate the direction. For example, when the operator wants to rotate the big arm clockwise, the left hand should be horizontally spread out and point to the left; and to indicate the lifting of the hook, the left hand needs to be stretched upwards.
[0155] Specifically, if there is a gesture action that meets the predetermined pattern, the position of the operator's left hand in the target picture is identified to determine the direction of the gesture. Although the algorithm can detect the swinging motion of the arm, it is difficult to directly detect the direction of the swing. For example, when the operator wants to rotate the large arm, the processor can detect the horizontal swing through vision, but it is difficult to determine whether it is swinging to the left or to the right. To solve this problem, it is agreed to use the left hand to distinguish the direction of the swing. The specific rules are as follows:
[0156] When the operator needs to rotate the large arm clockwise, the left hand should be horizontally spread out to the left, and the right hand should swing. At this time, the processor will not only detect the swinging action of the right hand, but also detect the posture of the left hand.
[0157] Similarly, for the indication of the lifting direction, the left hand is used to point up or down; for the indication of the amplitude, the left hand is used to point away from the chest or close to the chest.
[0158] S146, according to the gesture action and the direction, mapping to a specific tower crane control instruction to obtain a camera detection result.
[0159] In this embodiment, the corresponding control instruction is generated according to the recognized gesture action and its direction, and is sent to the intelligent control system through RS485 or other communication protocols, so that the tower crane performs the corresponding action according to the instruction. For example, the operator makes a gesture of moving the trolley inward (the left hand stretches forward, and the right hand swings forward and backward), and the system will analyze this series of gesture actions and convert them into commands for the movement of the tower crane.
[0160] Specifically, according to the gesture action and the direction, mapping to a specific tower crane control instruction to obtain a camera detection result. After determining the swinging direction and swinging action of the operator, the processor sends the corresponding instruction to the intelligent control system through RS485, and the intelligent control system makes the tower crane move according to the received instruction.
[0161] Through the above steps, the whole process from automatically adjusting the camera focal length to finally determining the control instruction is realized, which ensures the safety and accuracy of the tower crane control in the complex construction site environment. This multi-level detection mechanism not only improves the fault tolerance of the system, but also enhances the overall stability and reliability.
[0162] For the visual detection process of step S140 described above, in the tower crane construction operation, in order to ensure the safety and accuracy of the operation, a gesture recognition method based on a visual camera is used to assist the operation of the tower crane, which combines specific color dressing rules, human body posture detection algorithm and gesture direction recognition mechanism, so as to realize the precise control of the tower crane in the complex construction site environment.
[0163] Specifically, the visual camera is installed outside the tower crane cab and uses a pan-tilt head to cover most of the scene below the tower crane jib. Considering the need to detect the behavior of ground personnel, a camera with high zoom function is selected. In addition, the camera has realized the hook tracking function, which can automatically adjust the angle of view to the target nearby, so as to better detect the gesture action of the operator.
[0164] For the determination of gestures, as Figure 6 to Figure 7 shown, Figure 6 The left operator stretches the left hand forward and swings the right hand back and forth, indicating that the trolley amplitude moves inward towards the cab. If the left hand is retracted in front of the chest, it indicates that the trolley amplitude moves outward.
[0165] Figure 6 The middle operator stretches the left hand outward, and because the operator is facing the tower crane, the left hand indicates the clockwise direction, and the right hand swings left and right, indicating that the jib rotates clockwise. If the left hand is directed to the right, it indicates that the jib rotates counterclockwise.
[0166] Figure 6 The right operator stretches the left hand upward and swings the right hand up and down, indicating the action of lifting the hook. If the left hand is stretched downward, it indicates that the hook is lowered.
[0167] The direction detection of the sensor can be analogous to the direction detection of the vision, but there are differences in the rules between the two. In visual detection, the left hand is used to indicate the direction; while in sensor detection, as Figure 7 shown, the direction is indicated by the orientation of the right hand palm. In the six figures, the right hand palm of the sensor detection should be oriented forward, left, upward, backward, right, and downward, respectively.
[0168] In practical applications, in order to prevent false detection, it is agreed that after completing a gesture action, the tower crane action must be completely stopped before the next gesture action can be performed. The processor of the visual detection process will record the gesture action in the current period of time. If the detected gesture changes without stopping, it is considered as a detection error or an operation error. At this time, the processor of the visual detection process will immediately stop the action of the tower crane to ensure the safety and accuracy of the operation.
[0169] S150, selecting a final control instruction from the sensor detection result and the camera detection result to obtain a selected instruction; wherein the sensor detection result is transmitted through a wireless network; and the camera detection result is transmitted through a Modbus protocol.
[0170] In this embodiment, when both the sensor detection result and the camera detection result exist, the sensor detection result is selected as the final control instruction to obtain the selected instruction; when only one of the sensor detection result and the camera detection result exists, the existing result is selected as the final control instruction to obtain the selected instruction.
[0171] Specifically, in this embodiment, in order to ensure the safety and accuracy of the tower crane operation, two independent but complementary gesture detection methods are adopted in the gesture recognition process: sensor detection based on accelerometer and gyroscope, and image analysis based on visual camera. These two methods transmit their detection results to the intelligent control system through wireless network and Modbus protocol respectively, so as to further process and generate corresponding control instructions.
[0172] When only one of the sensor or the camera successfully detects the gesture and determines the corresponding control instruction, the system will directly adopt the detection result as the final control instruction. This means that if the sensor detects a valid gesture while the camera fails to detect, or vice versa, the system will perform the corresponding action according to the detection result that exists.
[0173] In an ideal case, both the sensor and the camera can successfully detect the same gesture and generate corresponding control instructions respectively. However, in order to ensure the stability and reliability of the system, when both provide valid detection results, the system preferentially adopts the sensor detection result as the final control instruction. This is because the sensor detection can more accurately reflect the gesture intention of the operator, especially in complex and variable construction environments.
[0174] The sensor (including accelerometer and gyroscope) is installed on the operator's hand to monitor the hand movement and orientation change in real time. These data are transmitted to the control system through wireless network, and the control system converts them into specific control instructions after analysis. The visual camera is responsible for capturing the overall image of the scene, and uses algorithms such as convolutional neural network to analyze the image content and identify the specific operator and his gesture. Then, the recognition result is sent to the control system through Modbus protocol.
[0175] This dual gesture detection strategy not only improves the fault tolerance of the system, but also enhances the adaptability in complex environments. By combining the advantages of sensors and visual cameras, it can effectively solve the problems that may be encountered by single detection method, such as occlusion, false detection, etc., thereby ensuring the safety and accuracy of tower crane operation.
[0176] In summary, the core of S150 step is to comprehensively evaluate the sensor detection result and the camera detection result, and flexibly select the most reliable control instruction source according to the actual situation, in order to realize efficient and accurate tower crane control.
[0177] S160, generate control signals according to the selected instructions, transmit wirelessly to the tower crane control system, and execute corresponding actions.
[0178] In the tower crane intelligent control system, S160 step will convert the final gesture control instruction after screening and verification into specific control signals, and send them to the driving module of the tower crane through wireless communication technology to execute corresponding actions.
[0179] First, in step S150, the final control instruction has been selected from the sensor detection results and camera detection results. This stage ensures the accuracy and reliability of the selected instruction, thereby laying the foundation for the next operation.
[0180] According to the selected instruction, there is a preset mapping table in the control system, which maps the gesture corresponding action (such as lifting, lowering, amplitude change, etc.) to specific control parameters. For example, "hand up and down" is mapped to the specific speed and acceleration value of the hook lifting.
[0181] In order to adapt to different working environments and needs, the control system also needs to adjust these parameters according to the current working state. For example, if the tower crane is handling heavy loads, it may need to slow down to ensure safety.
[0182] The generated control signal needs to be packaged according to a specific communication protocol (such as MODBUS or a custom wireless communication protocol) for transmission through a wireless network. This step includes adding necessary check codes (such as CRC) to ensure the integrity and accuracy of data transmission.
[0183] The packaged control signal is sent out through the wireless data transmission module. Considering the possibility of multiple tower cranes operating simultaneously on site, each tower crane has its own independent wireless channel to ensure that the instructions can be accurately and correctly transmitted to the target tower crane.
[0184] After receiving the wireless signal at the tower crane end, the received data is first decoded and verified. If the data is complete and correct, the control parameters contained therein are further parsed.
[0185] Finally, according to the parsed control parameters, the driving system of the tower crane starts to operate, driving the mechanical structure to complete actions such as lifting, amplitude change, rotation, etc. In this process, the control system also continuously monitors various operating parameters to ensure the safety of the operation.
[0186] In summary, the S160 step realizes the closed-loop control from gesture recognition to actual operation of the tower crane, not only improving the flexibility and accuracy of the tower crane operation, but also greatly enhancing the safety during construction. Through this way, the operator can more intuitively and conveniently remotely control the tower crane without relying on traditional remote controllers or other complex interfaces.
[0187] In addition, the tower crane needs to be fine-tuned in the last stage of automatic operation to accurately determine the landing site of the hoisted object. The operation precision requirement in this stage is high, and quick response is required. Traditional tower crane operation methods usually rely on manual control by the operator in the cab or remote operation through a remote controller. However, these methods may have problems such as inconvenient operation and delayed response in complex construction environments. In order to improve the convenience and efficiency of operation, the method of the present embodiment proposes a remote control method based on the gesture actions of the operator, so that the on-site construction personnel can directly remotely control the tower crane through gesture actions on site without using a remote controller or complex interfacing with other operators.
[0188] In order to realize the automatic operation and remote operation of the tower crane, the method involves an intelligent control system. The system is the basis for the automatic operation and remote operation of the tower crane, and can receive gesture instructions from the operator and convert them into action instructions for the tower crane, thereby realizing the automatic control of the tower crane. The system needs to have high-precision sensors, fast signal processing capability and reliable communication modules to ensure that the operation instructions can be accurately and correctly transmitted to the drive system of the tower crane and execute the corresponding actions.
[0189] Specifically, the intelligent control system includes an instruction module, a control module and a drive module, realizing the intelligent operation and automatic control of the tower crane.
[0190] The instruction module is the core part of the interaction between the tower crane intelligent control system and the outside world, mainly responsible for receiving external instructions and parsing these instructions into a data format that the control system can understand and process, and then distributing the data to the control module. The role of the instruction module is to ensure that external instructions can accurately and correctly enter the control system and prepare for subsequent processing.
[0191] The instruction module receives instructions through external interfaces, which mainly include wireless data transmission modules and wired RS485 interfaces.
[0192] The wireless data transmission module supports 2.4G and 5G transmission frequencies. This design enables the tower crane intelligent control system to select the appropriate transmission frequency according to different scenarios and needs. For example, in an environment with more interference, a more stable frequency can be selected; when high data transmission rate is required, 5G frequency can be selected. This flexibility greatly improves the adaptability and reliability of the system.
[0193] The wired RS485 interface is a commonly used serial communication interface, with the advantages of long transmission distance and strong anti-interference ability. In the operating environment of the tower crane, due to the possible existence of strong electromagnetic interference, the RS485 interface can ensure the stable transmission of the command and avoid the command error or loss caused by interference.
[0194] The command is encapsulated through the MODBUS protocol. MODBUS is a widely used communication protocol in the field of industrial automation, with the characteristics of simplicity, reliability and easy implementation. After entering the command module, the command will first be verified by CRC (Cyclic Redundancy Check) to ensure the integrity and accuracy of the command. CRC verification is a commonly used error detection method, which can effectively detect errors that may occur during transmission, such as data loss or damage.
[0195] If the CRC verification is passed, the command module will further analyze the command content according to the protocol agreed by the intelligent control system. The basic operation commands of the tower crane include lifting, amplitude changing, rotating, and corresponding setting speed and direction, etc. These commands are the basic control signals for the operation of the tower crane, and the command module needs to accurately analyze these commands and convert them into a data format that the control system can understand.
[0196] The control module is the core of the intelligent control system of the tower crane, responsible for receiving data from the command module and further determining the action mode of the tower crane according to these data. The main task of the control module is to realize the intelligent operation and automatic control of the tower crane, ensuring that the tower crane can safely and efficiently complete various operation tasks.
[0197] The control module verifies the received command parameters to ensure that these parameters are within a reasonable range and meet the safety operation requirements of the tower crane. For example, the speed setting cannot exceed the maximum allowable speed of the tower crane, and the direction setting must be legal. If the command parameters are found to have problems, the control module will refuse to execute the command and may issue an alarm or error information. The control module needs to implement a series of systematic functions to ensure the safe and efficient operation of the tower crane. These functions include obstacle avoidance, emergency stop and motion mode optimization, etc.
[0198] The tower crane may encounter various obstacles during operation, such as buildings, other tower cranes or other equipment. The control module monitors the environment around the tower crane in real time through sensors or other detection means, and when an obstacle is detected, it will automatically adjust the running track or speed of the tower crane to avoid collision.
[0199] In emergency situations, such as when the operator finds abnormal conditions or the system detects potential dangers, the control module can quickly start the emergency stop mechanism to make the tower crane stop running immediately, ensuring the safety of personnel and equipment.
[0200] The control module can automatically optimize the motion mode according to the operating state and task requirements of the tower crane, improving the operating efficiency. For example, when lifting heavy objects, the control module can automatically adjust the speed and acceleration to ensure the stability and safety of the lifting process.
[0201] The control module converts the parsed instruction content into corresponding drive signals according to different driving requirements, and distributes these signals to different drivers. This process requires precise control and coordination to ensure that each driver can accurately execute the instructions.
[0202] The drive module is the power part of the tower crane intelligent control system, responsible for providing power for the lifting, amplitude changing, and rotating of the tower crane. The drive module controls the operation of the motor or other driving devices according to the drive signals from the control module, thereby realizing the mechanical movement of the tower crane.
[0203] The drive module is the power source for the operation of the tower crane, and its performance directly affects the operating efficiency and safety of the tower crane. By precisely controlling the drive signals, the drive module can achieve smooth operation, fast response, and precise operation of the tower crane.
[0204] The intelligent control system based on the method of this embodiment is different from traditional tower crane control systems, and the main purpose is to provide support for the intelligent operation and automatic operation of the tower crane. In traditional control systems, the operation of the tower crane usually requires the operator to manually control each action, which is complex and prone to errors. The intelligent control system receives simple external instructions through the instruction module, and then the control module and the drive module cooperate to complete complex operation control tasks. This design greatly simplifies the operation process, improves the operation efficiency, and also reduces the labor intensity and operation difficulty of the operator.
[0205] Under the intelligent control system, the motion control of the tower crane can be completed with simple external instructions. For example, the operator can issue simple instructions such as "lifting", "amplitude changing", and "rotating" through a wireless remote control or a wired control console, and the intelligent control system will automatically analyze these instructions and complete the corresponding actions according to the pre-set control logic and safety rules. This intelligent control method not only improves the operating efficiency of the tower crane, but also enhances the safety and reliability of the tower crane.
[0206] In order to deal with complex scenes and improve fault tolerance, this method proposes two gesture detection methods: sensor detection and visual detection. These two methods complement each other to ensure that at least one gesture detection method can work normally in complex scenes, and improve the reliability and accuracy of detection through mutual verification.
[0207] The sensor detection method detects the operator's gesture actions through accelerometers and gyroscopes. These sensors can capture the hand's motion acceleration and rotation angle in real time, thus inferring the type and direction of the gesture. The specific implementation details are as follows:
[0208] A series of basic gesture actions are defined, which correspond to different operation instructions of the tower crane, such as lifting, luffing, rotating, etc. The operator controls the operation of the tower crane by performing these pre-defined gesture actions.
[0209] Accelerometers are used to measure the linear acceleration of the hand, and gyroscopes are used to measure the angular velocity of the hand. Through these data, the motion trajectory and posture of the hand can be calculated.
[0210] Based on sensor data, specific algorithms are used to recognize gesture actions. These algorithms can analyze the change pattern of acceleration and angular velocity to determine which gesture action the operator is performing.
[0211] The recognized gesture actions are converted into control instructions for the tower crane and sent to the automatic control system of the tower crane through the communication module, thus realizing remote control of the tower crane.
[0212] The visual detection method captures the operator's gesture actions through a visual camera and uses convolutional neural networks (CNN) and object trajectory analysis techniques for gesture recognition. The specific implementation details are as follows:
[0213] A series of basic gesture actions are also defined, which correspond to the actions in the sensor detection, ensuring that the two detection methods can be used in combination. For example, the operator can indicate the lifting, luffing or rotating of the tower crane through specific gesture actions.
[0214] A high-resolution visual camera is used to capture the operator's hand movements. The camera is installed on the tower crane or other suitable locations to ensure that the operator's gestures can be clearly captured.
[0215] A pre-trained convolutional neural network is used to analyze the visual data and identify the type of gesture action. CNN can automatically learn the features of gesture actions, thus achieving high-precision gesture recognition.
[0216] In addition to recognizing the type of gesture action, the trajectory of the gesture action also needs to be analyzed. By tracking the motion trajectory of the hand, the direction and amplitude of the gesture can be further determined, thus more accurately controlling the action of the tower crane.
[0217] The recognized gesture actions and their trajectories are converted into control instructions for the tower crane and sent to the automatic control system of the tower crane through the communication module, thus realizing remote control of the tower crane.
[0218] To improve the reliability and fault tolerance of gesture detection, the method of the embodiment combines sensor detection and visual detection. The two detection methods can complement each other to ensure that at least one method can work properly in a complex scene. At the same time, through mutual verification, the accuracy of detection can be further improved. For example, when the sensor detects a certain gesture action, visual detection can perform secondary verification to ensure the authenticity of the gesture action. Conversely, the same is true. This fusion and verification mechanism can effectively reduce misjudgment and improve the overall performance of the system.
[0219] The method of the embodiment is suitable for the automatic operation and remote operation of the tower crane in a complex construction environment. By controlling the tower crane through the gesture action of the operator, the convenience and efficiency of operation can be greatly improved. The operator does not need to leave the construction site or use a remote control or other complex equipment, but only needs to perform a simple gesture action to achieve precise control of the tower crane. This method is particularly suitable for use when fine-tuning the landing position of the hoisted object in the last stage, which can significantly improve the precision and response speed of the operation.
[0220] The above-mentioned intelligent tower crane human-computer interaction control method for complex scenes combines sensor detection of accelerometers and gyroscopes with image processing technology of visual cameras to achieve double recognition and verification of gesture actions, thereby significantly improving the accuracy of gesture recognition and the robustness of the system. Specifically, the system first binds the gesture detection device to the corresponding tower crane through specific wireless channel matching, and uses these sensors to capture the gesture changes of the operator to generate control instructions; at the same time, the visual camera automatically adjusts the focal length to locate and identify the operator in a specific dress and his or her predetermined mode gesture action, further confirming the control instructions. Finally, the most reliable control instructions are selected from the two sources of sensors and cameras, and transmitted to the tower crane control system through appropriate communication protocols to perform corresponding actions. This method not only effectively adapts to the complex and variable conditions of the construction site, simplifies the operation process, improves work efficiency, but also ensures construction safety, making the operation more precise and reliable.
[0221] Figure 8 is a schematic block diagram of an intelligent tower crane human-computer interaction control system 300 provided by an embodiment of the present application. As Figure 8 shown, corresponding to the above-mentioned intelligent tower crane human-computer interaction control method for complex scenes, the present application also provides an intelligent tower crane human-computer interaction control system 300 for complex scenes. The intelligent tower crane human-computer interaction control system 300 for complex scenes includes units for executing the above-mentioned intelligent tower crane human-computer interaction control method for complex scenes, and the system can be configured in a server. Specifically, please refer to Figure 8The intelligent tower crane human-computer interaction control system 300 for complex scenes includes a sensor binding unit 301, a sensor detection unit 302, a personnel positioning unit 303, a visual detection unit 304, a selection unit 305, and a control unit 306.
[0222] The sensor binding unit 301 is configured to bind a gesture detection device to a corresponding tower crane through a specific wireless channel matching process; the gesture detection device includes an accelerometer and a gyroscope; the sensor detection unit 302 is configured to capture gesture changes of an operator by using the gesture detection device installed on the hand, determine a control instruction, and obtain a sensor detection result; the personnel positioning unit 303 is configured to capture an image by using a visual camera, determine an operator of a corresponding tower crane by analyzing a direction faced by the operator and a size of the operator in the visual image; the visual detection unit 304 is configured to locate and identify a posture of a specific dressed operator by automatically adjusting a camera focal length through an image processing technology, monitor and identify a gesture action and a direction of the gesture action of a predetermined mode, and determine a control instruction to obtain a camera detection result; the selection unit 305 is configured to select a final control instruction from the sensor detection result and the camera detection result to obtain a selected instruction; the sensor detection result is transmitted through a wireless network mode; the camera detection result is transmitted through a Modbus protocol; and the control unit 306 is configured to generate a control signal according to the selected instruction, transmit the control signal to a tower crane control system through wireless transmission, and perform a corresponding action.
[0223] In an embodiment, the selection unit 305 is configured to select the sensor detection result as the final control instruction to obtain the selected instruction when both the sensor detection result and the camera detection result exist; and select an existing result as the final control instruction to obtain the selected instruction when only one of the sensor detection result and the camera detection result exists.
[0224] In an embodiment, the sensor binding unit 301 is configured to press a gesture detection device button to make the gesture detection device enter a pairing mode, send matching data packets one by one from a first channel set in advance until a response of the tower crane is received, and complete the binding of the gesture detection device to the corresponding tower crane.
[0225] In an embodiment, the sensor detection unit 302 includes:
[0226] The setting subunit is configured to set an initial coordinate system; the angle calculation subunit is configured to calculate a rotation angle by using the gyroscope and the initial coordinate system, determine and verify whether the palm orientation is valid; the valid action judgment subunit is configured to, when the palm orientation is valid, determine whether there is a valid swing action according to a set condition by using the accelerometer; the gesture change detection subunit is configured to, when there is a valid swing action, determine a gesture change condition based on the palm orientation and the swing action, and determine a corresponding tower crane control instruction according to a predefined mapping relationship, to obtain a sensor detection result.
[0227] In an embodiment, the setting subunit is configured to set the initial coordinate system according to a current posture of an operator, wherein a positive direction of a z-axis corresponds to a direction in which a palm is upward, a positive direction of a y-axis corresponds to a direction in which four fingers point to a tower crane, and a positive direction of an x-axis corresponds to a direction of a thumb.
[0228] In an embodiment, the angle calculation subunit is configured to measure rotation angular velocities around the x-axis, the y-axis and the z-axis by using the gyroscope, and perform integration to calculate the rotation angle, so as to determine the palm orientation, and when the palm orientation satisfies a specific threshold condition, the palm orientation is valid.
[0229] In an embodiment, the visual detection unit 304 includes:
[0230] The positioning subunit is configured to automatically adjust a camera focal length, and gradually search until an operator is positioned, to obtain a positioning picture; the recognition subunit is configured to recognize and position a specific dressed operator in the picture by using a pre-trained human body detection algorithm, to obtain a target picture; the posture analysis subunit is configured to analyze and determine a posture and a standing direction of the operator in the target picture based on an angle of a visual camera and human body symmetry; the monitoring subunit is configured to monitor actions of arms and hands in the target picture based on the posture and the standing direction of the operator, to determine whether there is a gesture action conforming to a predetermined mode; the gesture direction determination subunit is configured to, if there is a gesture action conforming to the predetermined mode, recognize a position of a left hand of the operator in the target picture, to determine a direction of the gesture; and the mapping subunit is configured to map the gesture action and the direction into a specific tower crane control instruction, to obtain a camera detection result.
[0231] In an embodiment, the posture analysis subunit is configured to calculate a posture and a side angle of the operator in a camera view in the target picture by using a convolutional neural network based on the angle of the visual camera and the human body symmetry, to obtain the posture and the standing direction of the operator.
[0232] It should be noted that the specific implementation process of the intelligent tower crane human-computer interaction control system 300 and each unit facing the complex scene can be clearly understood by those skilled in the art, which can be referred to the corresponding description in the foregoing method embodiments, and for the convenience and brevity of description, it will not be repeated here.
[0233] The intelligent tower crane human-computer interaction control system 300 facing the complex scene can be implemented in the form of a computer program, which can run on a computer device as shown in the figure. Figure 9
[0234] Please refer to Figure 9 , Figure 9 is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a server, wherein the server can be a stand-alone server or a server cluster composed of multiple servers.
[0235] Refer to Figure 9 , the computer device 500 includes a processor 502, a memory and a network interface 505 connected through a system bus 501, wherein the memory can include a non-volatile storage medium 503 and an internal memory 504.
[0236] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which when executed, can cause the processor 502 to execute an intelligent tower crane human-computer interaction control method facing a complex scene.
[0237] The processor 502 is configured to provide computing and control capabilities to support the operation of the entire computer device 500.
[0238] The internal memory 504 provides an environment for the running of the computer program 5032 in the non-volatile storage medium 503, which when executed by the processor 502, can cause the processor 502 to execute an intelligent tower crane human-computer interaction control method facing a complex scene.
[0239] The network interface 505 is configured to perform network communication with other devices. Those skilled in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 500 to which the scheme of the present application is applied. The specific computer device 500 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0240] The processor 502 is configured to run the computer program 5032 stored in the memory to implement all the steps of the intelligent tower crane human-computer interaction control method for complex scenarios.
[0241] It should be understood that, in the embodiments of the present application, the processor 502 can be a central processing unit (CPU), and the processor 502 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0242] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the above-mentioned embodiments.
[0243] Therefore, the present application also provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program is executed by a processor to make the processor execute all the steps of the intelligent tower crane human-computer interaction control method for complex scenarios.
[0244] The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, and various computer-readable storage media that can store program codes.
[0245] It can be realized by electronic hardware, computer software or a combination of both that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0246] In several embodiments provided by the present application, it should be understood that the disclosed system and method can be implemented in other manners. For example, the division of the system embodiments is merely illustrative. For example, the division of the units can be divided in another manner, for example, a plurality of units or componental can be combined or can be integrated into another system, or some features can be ignored or not executed.
[0247] The steps in the method embodiments of the present application can be adjusted, combined and deleted in sequence according to actual needs. The units in the system embodiments of the present application can be combined, divided and deleted according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0248] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part of the prior art that makes a contribution, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a terminal or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0249] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for intelligent tower crane human-machine interaction control for complex scenes, characterized in that, The application relates to a gesture control system for a tower crane, comprising: binding a gesture detection device to a corresponding tower crane through a specific wireless channel matching process; wherein the gesture detection device comprises an accelerometer and a gyroscope; capturing gesture changes of an operator by using the gesture detection device installed on the hand, determining a control instruction, and obtaining a sensor detection result; capturing an image by using a visual camera, determining the operator of the corresponding tower crane by analyzing the direction faced by the operator and the size of the operator in the visual image; positioning and recognizing the posture of a specific dressed operator by automatically adjusting the focal length of the camera through image processing technology, monitoring and recognizing gesture actions in a predetermined mode and the direction thereof, and determining a control instruction, and obtaining a camera detection result; selecting a final control instruction from the sensor detection result and the camera detection result, and obtaining a selected instruction; wherein the sensor detection result is transmitted through a wireless network mode; and the camera detection result is transmitted through a Modbus protocol; generating a control signal according to the selected instruction, and transmitting the control signal to a tower crane control system through wireless transmission to execute corresponding actions; the gesture detection device installed on the hand is used to capture gesture changes of an operator, determine a control instruction, and obtain a sensor detection result, comprising: setting an initial coordinate system; calculating a rotation angle by using the gyroscope and the initial coordinate system, determining and verifying whether a palm orientation is valid; when the palm orientation is valid, judging whether there is a valid swing action according to a set condition by using the accelerometer; when there is a valid swing action, determining a gesture change condition based on the palm orientation and the swing action, and determining a corresponding tower crane control instruction according to a pre-defined mapping relationship, and obtaining a sensor detection result; image processing technology is used to position and recognize the posture of a specific dressed operator by automatically adjusting the focal length of the camera, monitor and recognize gesture actions in a predetermined mode and the direction thereof, and determine a control instruction, and obtain a camera detection result, comprising: automatically adjusting the focal length of the camera, gradually searching until the operator is positioned, and obtaining a positioning picture; recognizing and positioning a specific dressed operator in the picture by using a pre-trained human body detection algorithm, and obtaining a target picture; analyzing and determining the posture and standing direction of the operator in the target picture based on the angle of the visual camera and the human body symmetry; monitoring the actions of arms and hands in the target picture based on the posture and standing direction of the operator, and judging whether there is a gesture action in a predetermined mode; if there is a gesture action in a predetermined mode, recognizing the position of the left hand of the operator in the target picture, and determining the direction of the gesture; mapping the gesture action and the direction into a specific tower crane control instruction, and obtaining a camera detection result.
2. The intelligent tower crane human-machine interaction control method for complex scenes according to claim 1, characterized in that, the final control instruction is selected from the sensor detection result and the camera detection result, and a selected instruction is obtained, comprising: When both the sensor detection result and the camera detection result exist, the sensor detection result is selected as the final control instruction to obtain a selected instruction; when only one of the sensor detection result and the camera detection result exists, the existing result is selected as the final control instruction to obtain the selected instruction.
3. The intelligent tower crane human-machine interaction control method for complex scenes according to claim 1, characterized in that, The gesture detection device is bound to the corresponding tower crane through a specific wireless channel matching process, which comprises: The gesture detection device is pressed to enter a pairing mode, and matching data packets are sent from a first preset channel one by one until a response of the tower crane is received, so as to complete the binding of the gesture detection device and the corresponding tower crane.
4. The intelligent tower crane human-machine interaction control method for complex scenes according to claim 3, characterized in that, The initial coordinate system is set, which comprises: The initial coordinate system is set according to the current posture of the operator, wherein the positive direction of the z-axis corresponds to the direction of the palm up, the positive direction of the y-axis corresponds to the direction of the four fingers pointing to the tower crane, and the positive direction of the x-axis corresponds to the direction of the thumb.
5. The intelligent tower crane human-machine interaction control method for complex scenes according to claim 4, characterized in that, The rotation angle is calculated by using the gyroscope and the initial coordinate system, and the hand heart direction is determined and verified, which comprises: The rotation angle is calculated by using the gyroscope to measure the rotation angular velocity around the x, y and z axes and performing integration, the hand heart direction is determined, and the hand heart direction is valid when the hand heart direction meets a specific threshold condition.
6. The intelligent tower crane human-machine interaction control method for complex scenes according to claim 4, characterized in that, The set conditions comprise that the acceleration change amplitude exceeds a threshold value, the swing switching time meets the rapidity requirement, and at least three repeated swing modes appear.
7. The intelligent tower crane human-machine interaction control method for complex scenes according to claim 6, characterized in that, The posture and standing direction of the operator in the target picture are analyzed and determined based on the angle of the visual camera and the human body symmetry, which comprises: The posture and standing direction of the operator in the target picture are calculated by using a convolutional neural network based on the angle of the visual camera and the human body symmetry to obtain the posture and standing direction of the operator.
8. The intelligent tower crane human-machine interaction control system for complex scenes, based on the method of any one of claims 1-7, characterized in that, It comprises: A sensor binding unit is configured to bind the gesture detection device to the corresponding tower crane through a specific wireless channel matching process; wherein the gesture detection device comprises an accelerometer and a gyroscope; A sensor detection unit is configured to capture the gesture change of the operator by using the gesture detection device installed on the hand, determine a control instruction, and obtain a sensor detection result; A personnel positioning unit is configured to capture an image by using a visual camera, determine the operator of the corresponding tower crane by analyzing the direction faced by the operator and the size of the operator in the visual image; A visual detection unit is configured to position and identify the posture of a specific dressed operator by automatically adjusting the focal length of the camera by using image processing technology, monitor and identify the gesture action and direction of the predetermined mode, and determine a control instruction to obtain a camera detection result; A selection unit is configured to select a final control instruction from the sensor detection result and the camera detection result to obtain a selected instruction; wherein the sensor detection result is transmitted through a wireless network; and the camera detection result is transmitted through a Modbus protocol. A control unit is configured to generate a control signal according to the selected instruction, transmit the control signal to the tower crane control system through wireless transmission, and perform corresponding actions.
Citation Information
Patent Citations
System and methods for on-body gestural interfaces and projection displays
US20170123487A1
Hand gesture recognition system for vehicular interactive control
US20190302895A1