Tower crane steel structure hanging object posture prediction method and system based on multiple cameras
Through the multi-camera system, the key nodes of tower crane objects are identified and tracked, and the problem of inaccurate monitoring of the posture and motion status of the hanging object in traditional methods is solved, accurate prediction and safety control of the movement trajectory of the hanging object is achieved, and the safety of tower crane operations is improved.
Patent Information
- Application Number
- CN202510366750.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional methods are difficult to accurately monitor the posture and movement status of tower crane objects, resulting in low prediction accuracy of key points of steel structure crane objects.
A multi-camera system is used to collect videos of hanging objects, identify key nodes of steel structure hanging objects, determine their two-dimensional and three-dimensional spatial coordinates, predict the motion trajectory in combination with the external environment, and perform safe control based on risk results.
Accurate prediction and safety control of the movement trajectory of steel structure lifting objects is achieved, and the safety and reliability of tower crane operations are improved.
Smart Images

Figure CN120298947A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tower cranes, and in particular to a method and system for predicting the attitude of a steel structure lifting object of a tower crane based on multiple cameras. Background Art
[0002] With the development of technology, tower cranes are gradually applied to people's lives and are used to hang steel structure lifting objects. In terms of the safety monitoring of tower crane operations, the traditional methods relying on manual observation and simple sensors have limitations and are difficult to meet the comprehensive and accurate monitoring of the attitude and motion state of the lifting object, resulting in a low prediction accuracy of the motion trajectory of the key points of the steel structure lifting object. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art, and the present invention provides a method and system for predicting the attitude of a steel structure lifting object of a tower crane based on multiple cameras.
[0004] An embodiment of the present invention provides a method for predicting the attitude of a steel structure lifting object of a tower crane based on multiple cameras, including: when multiple cameras are installed at the tower crane operation site, the multiple cameras collect lifting object videos at different angles; determining multiple key nodes of the steel structure lifting object based on the lifting object videos and determining their two-dimensional spatial coordinates in different videos; determining the coordinates of the steel structure lifting object in three-dimensional space according to the two-dimensional spatial coordinates of the key nodes of the steel structure lifting object from at least two different perspectives; predicting the motion trajectory of the key points of the steel structure lifting object based on the spatial positions, attitudes of the lifting object, and the external environment; predicting the corresponding risk results according to the motion trajectory of the key points of the steel structure lifting object and the attitude of the steel structure lifting object; determining the dynamic measures of the steel structure lifting object according to the risk results, the steel structure lifting object, and the tower crane, and performing safety control on the steel structure lifting object.
[0005] An embodiment of the present invention provides a system for predicting the attitude of a steel structure lifting object of a tower crane based on multiple cameras. The system for predicting the attitude of a steel structure lifting object of a tower crane based on multiple cameras is applied to the above-mentioned method for predicting the attitude of a steel structure lifting object of a tower crane based on multiple cameras. The system for predicting the attitude of a steel structure lifting object of a tower crane based on multiple cameras includes:
[0006] A collection module, configured to collect lifting object videos at different angles by multiple cameras when multiple cameras are installed at the tower crane operation site;
[0007] A key node module, configured to determine multiple key nodes of the steel structure lifting object based on the lifting object videos and determine their two-dimensional spatial coordinates in different videos;
[0008] A spatial position module, configured to determine the coordinates of the steel structure lifting object in three-dimensional space according to the two-dimensional spatial coordinates of the key nodes of the steel structure lifting object from at least two different perspectives;
[0009] A trajectory module, configured to predict the movement trajectory of key points of a steel structure lifting object based on the spatial coordinates of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment;
[0010] A risk result module, configured to predict corresponding risk results according to the movement trajectory of the key points of the steel structure lifting object and the attitude of the steel structure lifting object;
[0011] A control module, configured to determine dynamic measures for the steel structure lifting object according to the risk results, the steel structure lifting object, and the tower crane, and perform safety control on the steel structure lifting object.
[0012] Compared with the prior art, the beneficial effects of the present invention are:
[0013] In the embodiment of the present invention, by means of the method in the embodiment of the present invention, when multiple cameras are installed at the tower crane operation site, the multiple cameras collect lifting object videos at different angles; determine multiple key nodes of the steel structure lifting object based on the lifting object videos and determine their two-dimensional spatial coordinates in different videos; determine the coordinates of the steel structure lifting object in the three-dimensional space according to the two-dimensional spatial coordinates of the key nodes of the steel structure lifting object from at least two different perspectives; predict the movement trajectory of the key points of the steel structure lifting object based on the spatial position of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment, which comprehensively considers the spatial position of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment, and realizes accurate prediction of the movement trajectory of the key points of the steel structure lifting object.
[0014] Therefore, predict corresponding risk results according to the movement trajectory of the key points of the steel structure lifting object and the attitude of the steel structure lifting object; determine dynamic measures for the steel structure lifting object according to the risk results, the steel structure lifting object, and the tower crane, and perform safety control on the steel structure lifting object, so as to realize the safety prevention and control of the steel structure lifting object under dynamic suspension. Description of the Drawings
[0015] Figure 1 It is a schematic diagram of an application scenario of a method for predicting the attitude of a tower crane steel structure lifting object based on multiple cameras in an embodiment;
[0016] Figure 2 It is a schematic flowchart of a method for predicting the attitude of a tower crane steel structure lifting object based on multiple cameras in an embodiment of the present invention;
[0017] Figure 3 It is a schematic diagram of the structural composition of a system for predicting the attitude of a tower crane steel structure lifting object based on multiple cameras in an embodiment of the present invention;
[0018] Figure 4 It is a schematic diagram of a tower crane in an embodiment of the present invention. Detailed Embodiments
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0020] Embodiment 1
[0021] The method for predicting the attitude of the steel structure lifting object of the tower crane based on multiple cameras provided by this application is applied to the application environment as Figure 1 shown. Among them, the computer 102 communicates with the server 104 through the network. Among them, the terminal 102 is not limited to various personal computers, servers, and tower cranes, and the server 104 is implemented by an independent server or a server cluster composed of servers.
[0022] Embodiment 2
[0023] Please refer to Figures 1 to 4 , a method for predicting the attitude of the steel structure lifting object of the tower crane based on multiple cameras, which is applied to the control scenario of the tower crane for the steel structure lifting object; the method for predicting the attitude of the steel structure lifting object of the tower crane based on multiple cameras includes:
[0024] Step S11: When multiple cameras are installed at the tower crane operation site, the multiple cameras collect the video of the lifting object at different angles;
[0025] Step S12: Determine multiple key nodes of the steel structure lifting object based on the lifting object video and determine their two-dimensional spatial coordinates in different videos;
[0026] Step S13: Determine its coordinates in the three-dimensional space according to the two-dimensional spatial coordinates of the key nodes of the steel structure lifting object from at least two different perspectives;
[0027] Step S14: Predict the movement trajectory of the key points of the steel structure lifting object based on the spatial position, lifting object attitude and external environment of the key points of the steel structure lifting object;
[0028] Step S15: Predict the corresponding risk results according to the movement trajectory of the key points of the steel structure lifting object and the attitude of the steel structure lifting object;
[0029] Step S16: Determine the dynamic measures of the steel structure lifting object according to the risk results, the steel structure lifting object and the tower crane, and perform safety control on the steel structure lifting object.
[0030] In step S11, when multiple cameras are installed at the tower crane operation site, the multiple cameras collect the video of the lifting object at different angles;
[0031] In the specific implementation process of the present invention, the specific steps are as follows:
[0032] S111: Determine the positions of multiple cameras according to the hanging position of the tower crane and the shape of the tower crane;
[0033] S112: When multiple cameras are installed at the tower crane operation site, all the multiple cameras face the hanging position of the tower crane;
[0034] S113: The multiple cameras photograph the hanging position of the tower crane, and the multiple cameras collect video of the suspended load at different angles; the multiple videos of the suspended load contain steel structure suspended loads.
[0035] In an embodiment of the present application, a three-dimensional model of the tower crane is collected; the hanging position of the tower crane is marked based on the three-dimensional model of the tower crane; the positions of the multiple cameras are determined according to the hanging position of the tower crane and the shape of the tower crane, the positions of the multiple cameras are controlled, and the overall consideration of the hanging position of the tower crane and the shape of the tower crane is compatible.
[0036] At this time, it is necessary to obtain a three-dimensional model of the tower crane. This can usually be obtained through three-dimensional scanning, CAD design, or an existing three-dimensional model library; on the three-dimensional model of the tower crane, the hanging position is marked according to the actual hanging point or the expected hanging area, and based on the hanging position and the shape of the tower crane, the positions of the cameras that can comprehensively cover the hanging position and capture the necessary information are calculated. At the same time, the determination of the camera positions is to ensure that clear and comprehensive videos of the suspended load can be collected, taking into account factors such as the viewing angle, focal length, and installation conditions of the cameras.
[0037] Therefore, when multiple cameras are installed at the tower crane operation site, all the multiple cameras face the hanging position of the tower crane; the multiple cameras photograph the hanging position of the tower crane, and the multiple cameras collect video of the suspended load at different angles; the multiple videos of the suspended load contain steel structure suspended loads, and multiple videos of the suspended load are introduced for further control.
[0038] At this time, according to the positions determined in the previous step, select a suitable camera model and install it at the corresponding position on the tower crane. When installing the cameras, ensure that each camera faces the hanging position to capture the dynamic changes of the hanging point. Start the cameras to start shooting videos of the hanging position. These videos will contain the dynamic changes of the steel structure suspended load. At the same time, ensure that the orientation of the cameras is to maximize the accuracy and effectiveness of video collection, avoiding blind spots or overlapping areas. Collecting videos of the suspended load provides data support for subsequent steps such as video analysis, target recognition, and motion trajectory prediction.
[0039] In addition, system initialization and camera deployment. Install multiple cameras at key positions of the tower crane, adjust their positions and parameters to ensure coverage of the working area and clear video streams, and at the same time configure the system operating environment to lay a foundation for subsequent data collection and processing.
[0040] By rationally installing multiple cameras on and around the tower crane and accurately calibrating their positional relationships and parameters, comprehensive coverage of the steel structure hoisting working area can be achieved. The distance between the camera and the hoisted object can be calculated in real time, providing basic data support for subsequent precise analysis.
[0041] In step S12, multiple key nodes of the steel structure hanging object are determined based on the hanging object video and their two-dimensional spatial coordinates in different videos are determined;
[0042] In the specific implementation process of the present invention, the specific steps are:
[0043] S121: Acquire multiple videos of hanging objects;
[0044] S122: Based on the recognition of the shape of the steel structure hoist, multiple key nodes of the steel structure hoist are determined, and the two-dimensional spatial coordinates of the steel structure hoist in different videos are determined, and the multiple key nodes are two-dimensional key points.
[0045] In an embodiment of the present application, multiple videos of suspended objects are acquired; and the outer contour of the steel structure suspended object is determined based on the recognition of the multiple videos of suspended objects, thereby achieving the recognition of the multiple videos of suspended objects and ensuring the accuracy of the outer contour of the steel structure suspended object.
[0046] At this point, the installation and configuration of the camera in the early stage are completed. Multiple cameras are installed on the tower crane, which are responsible for shooting the hanging position of the tower crane, thereby capturing the dynamic video of the steel structure hoisting objects. These videos are transmitted to the central processing unit or remote monitoring center in real time through the data transmission system for subsequent analysis and processing. When obtaining the video of the hoisting object, the following points should be noted: Ensure that the clarity and viewing angle of the camera can meet the recognition requirements. The camera should be installed in a position that can fully cover the hanging position to avoid blind spots. Video data should be transmitted in real time to ensure the timeliness and accuracy of the analysis.
[0047] After acquiring multiple videos of suspended objects, image processing technology is used to process and analyze these videos. First, each frame of the image is extracted from the video; then, edge detection, contour extraction and other algorithms are used to identify the outer contour of the steel structure suspended object in each frame of the image.
[0048] Specifically, assume that at a construction site, there is a tower crane in operation, responsible for hoisting steel structure components. Four cameras are installed on the tower crane, facing the front, back, left, and right directions of the hanging position respectively. These cameras captured the entire process of the steel structure components being lifted from the ground, rising, moving, and descending. Through the videos captured by these four cameras, the dynamic changes of the steel structure components can be observed comprehensively. Assume that a large number of image frames are extracted from the videos captured by the four cameras. Then, image processing algorithms are used to process these image frames to identify the outer contours of the steel structure components in each frame of the image. These outer contours are represented in the form of two-dimensional coordinate points and are connected into continuous curves or broken lines to form the outer contour map of the steel structure components.
[0049] By comparing the outer contour maps at different time points, the morphological changes of the steel structure components during the hoisting process can be found, such as elongation during rising and swinging during movement, etc. This information provides important inputs for subsequent steps such as determining the spatial position, predicting the movement trajectory, and controlling risks.
[0050] Therefore, the morphology of the steel structure lifting object is determined based on its outer contour; multiple key nodes of the steel structure lifting object are determined based on the recognition of the morphology of the steel structure lifting object. The multiple key nodes are two-dimensional key points, realizing the recognition of the morphology of the steel structure lifting object and ensuring the accuracy of the multiple key nodes of the steel structure lifting object.
[0051] At this time, a detailed analysis of the outer contour is carried out, including features such as the length, width, height, and curvature of the contour. The features obtained from the analysis are matched with a preset steel structure morphology library. The morphology library may contain various common steel structure morphologies, such as rectangles, circles, triangles, irregular shapes, etc. According to the matching results, the most likely morphology of the steel structure lifting object is determined. If the matching degree is very high, the preset morphology is directly adopted; if the matching degree is low, it may be necessary to combine manual judgment or further analysis to determine the morphology. In practical applications, the determination of the morphology may also need to consider factors such as the material, weight, and hoisting method of the steel structure lifting object to ensure that the determined morphology is consistent with the actual situation.
[0052] Based on this morphology, multiple key nodes of the steel structure lifting object are identified. These key nodes are usually important connection points, turning points, or feature points on the steel structure lifting object, and they are represented in the form of coordinate points on the two-dimensional image. At the same time, the steel structure lifting object is segmented into different parts or components according to its morphology. Feature points are extracted on each part or component. These feature points may be connection points, turning points, edge points, etc. The two-dimensional spatial coordinates in different videos are determined, and multiple key nodes of the steel structure lifting object are determined according to the positions and importance of the feature points. These nodes will be used for subsequent steps such as determining the spatial position, predicting the movement trajectory, and controlling risks.
[0053] In step S13, the coordinates of the key nodes of the steel structure lifting object in three-dimensional space are determined based on the two-dimensional space coordinates of the key nodes from at least two different perspectives of the steel structure lifting object;
[0054] In the specific implementation process of the present invention, the specific steps are as follows:
[0055] S131: Obtain multiple key nodes of the steel structure lifting object;
[0056] S132: Mark the corresponding coordinates for the multiple key nodes;
[0057] S133: Use the coordinates of the multiple key nodes from at least two different perspectives to perform triangulation to determine the positions of the key points in three-dimensional space;
[0058] S134: Determine the coordinates of the steel structure lifting object in three-dimensional space based on the continuous iteration of the positions of the key points in three-dimensional space, and the spatial position of the steel structure lifting object matches the corresponding three-dimensional coordinates.
[0059] In the embodiment of the present application, obtaining multiple key nodes of the steel structure lifting object; marking the corresponding coordinates for the multiple key nodes; using the coordinates of the multiple key nodes from at least two different perspectives to perform triangulation to determine the positions of the key points in three-dimensional space ensures the accuracy of the positions of the key points in three-dimensional space.
[0060] At this time, obtain multiple key nodes of the steel structure lifting object in the two-dimensional image. These key nodes are usually the feature points on the steel structure lifting object, such as connection points, turning points or edge points, and they have clear and recognizable positions in the image. The method of obtaining key nodes may include image processing techniques, such as edge detection, corner detection, feature matching, etc. First, preprocess the image, such as denoising, enhancing contrast, etc., to improve the recognition accuracy of key nodes. Then, use image processing algorithms to automatically detect or manually select key nodes and record their position information.
[0061] Mark the coordinates of the key nodes obtained in step S131. This means assigning a unique coordinate value to each key node, and this coordinate value represents the position of the key node in the image coordinate system. Coordinate marking is usually automatically completed by image processing software or algorithms, but manual intervention may also be required to ensure accuracy. The process of coordinate marking may involve converting the image into a digital matrix, where each pixel has a corresponding coordinate value. Then, use image processing algorithms to determine the position of the key node in the matrix and use this position as the coordinate of the key node.
[0062] Triangulation is performed using the coordinates of key nodes in images taken from at least two different perspectives. Triangulation is a geometric method that uses the image coordinates of the same object in two or more perspectives to determine the position of the object in three-dimensional space. In actual operation, we need to first ensure that there are enough common key nodes in the images of the two perspectives, and these nodes have clear coordinates in both images. Then, use the triangulation algorithm to calculate the positions of these key nodes in three-dimensional space.
[0063] Specifically, assume that in an image of a steel structure lifting object, we can see the outline of the lifting object and some obvious feature points, such as the connection point between the hook and the lifting object, a certain protruding part on the lifting object, etc. Through image processing algorithms, we can automatically detect these feature points and use them as key nodes. Each key node has a corresponding coordinate in the image, indicating its position. Assume that we have detected several key nodes on the steel structure lifting object through image processing algorithms. Now, we need to label coordinates for these key nodes. In the image coordinate system, each pixel has a unique coordinate (such as x, y), indicating its position in the image. We can use image processing software or algorithms to automatically determine the coordinates of each key node in the image coordinate system and record these coordinates.
[0064] Assume that we have two images of a steel structure lifting object taken from different perspectives, each image contains some key nodes, and these nodes are visible in both images. Now, we want to determine the positions of these key nodes in three-dimensional space. For this purpose, we can use the triangulation algorithm. First, we need to determine the relative position relationship between the two perspectives (such as the rotation matrix and translation vector), which can be obtained through methods such as camera calibration or feature matching. Then, we use the triangulation algorithm to calculate the coordinates of each key node in three-dimensional space. These coordinates represent the positions of the key nodes in the real world and can be used for subsequent spatial position analysis, motion trajectory prediction, and other applications.
[0065] Therefore, the coordinates of the steel structure lifting object in three-dimensional space are determined based on the continuous iteration of the positions of the key points in three-dimensional space. The spatial position of the steel structure lifting object matches the corresponding three-dimensional coordinates, realizing the continuous iteration of the positions of the key points in three-dimensional space and ensuring the accuracy of the spatial position of the steel structure lifting object.
[0066] At this time, based on the key point positions in the three-dimensional space that have been determined through triangulation, continuously iterate to determine the spatial position of the entire steel structure lifting object. Based on the key point positions, first estimate the initial spatial position of the steel structure lifting object. Use iterative algorithms (such as the Iterative Closest Point algorithm ICP, non-linear least squares method, etc.) to continuously optimize the spatial position estimation of the steel structure lifting object. These algorithms compare the difference between the predicted key point positions and the actually observed key point positions, and accordingly adjust the spatial position of the steel structure lifting object until the convergence condition is reached (such as the position change is less than the preset threshold). During the iteration process, it is necessary to consider the possible deformation of the steel structure lifting object and measurement errors. This can be handled by introducing a deformation model or an error correction method. Determine when the iterative algorithm converges, that is, when to stop the iteration. This is usually judged based on the amount of position change, the size of the residual, or other convergence indicators. When the iterative algorithm converges, the final spatial position estimation of the steel structure lifting object is obtained. This position estimation matches the corresponding three-dimensional coordinates and can be used for subsequent analysis, control, or visualization tasks.
[0067] Specifically, assume we have a complex steel structure lifting object composed of multiple components, and these components may deform during the lifting process. To determine the spatial position of this steel structure lifting object, use cameras from two different perspectives to photograph the steel structure lifting object, and extract the two-dimensional coordinates of the key points through image processing technology. Using the triangulation algorithm, calculate the positions of these key points in the three-dimensional space based on their two-dimensional coordinates. Based on the calculated three-dimensional positions of the key points, we estimate the initial spatial position of the steel structure lifting object.
[0068] We use the Iterative Closest Point algorithm (ICP) to optimize the spatial position estimation of the steel structure lifting object. In each iteration, we calculate the difference between the predicted key point positions and the actually observed key point positions, and adjust the spatial position of the steel structure lifting object based on this difference. To handle the deformation problem, we introduce a deformation model that allows the steel structure lifting object to deform to a certain extent during the iteration process. During the iteration process, we monitor the amount of position change and the size of the residual. When the amount of position change is less than the preset threshold and the size of the residual is stable, we consider that the iteration has converged. When the iterative algorithm converges, we obtain the final spatial position estimation of the steel structure lifting object. This position estimation matches the corresponding three-dimensional coordinates and takes into account the effects of deformation and measurement errors. Through this process, we can accurately determine the spatial position of the steel structure lifting object, providing strong support for subsequent analysis, control, or visualization tasks.
[0069] In step S14, predict the motion trajectory of the key points of the steel structure lifting object based on the spatial positions of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment;
[0070] In the specific implementation process of the present invention, the specific steps are as follows:
[0071] S141: Obtain the spatial positions of the key points of the steel structure lifting object;
[0072] S142: Determine the attitude of the lifting object based on the spatial positions of the key points;
[0073] S143: Collect multiple environmental parameters through environmental detection based on the spatial position of the steel structure lifting object;
[0074] S144: Determine the corresponding external environment according to multiple environmental parameters;
[0075] S145: Associate the spatial positions of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment;
[0076] S146: Predict the movement trajectory of the key points of the steel structure lifting object based on the spatial positions of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment.
[0077] In the embodiments of the present application, the spatial positions of the key points of the steel structure lifting object are obtained; the attitude of the lifting object is determined based on the spatial positions of the key points, and the attitude of the lifting object is introduced for further control. At the same time, multiple environmental parameters are collected through environmental detection based on the spatial position of the steel structure lifting object; the corresponding external environment is determined according to multiple environmental parameters, and the external environment and the attitude are introduced.
[0078] Furthermore, the spatial positions of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment are associated; the movement trajectory of the key points of the steel structure lifting object is predicted based on the spatial positions of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment, which accommodates the overall consideration of the spatial positions of the key points of the steel structure lifting object, the attitude of the lifting object, and the external environment, and realizes the accurate prediction of the movement trajectory of the key points of the steel structure lifting object.
[0079] In the embodiments of the present application, multiple cameras are installed at key positions of the tower crane, and their positions and parameters are adjusted to ensure that the working area is covered and the video stream is clear, and at the same time, the system operating environment is configured. As Figure 2 shown.
[0080] 1) Calculation of the number of cameras
[0081] For a camera, its field of view (Field of View, FOV) can be expressed in angles, generally divided into the horizontal field of view (FOV h ) and the vertical field of view (FOV v ). If the focal length f of the camera and the sensor size (the horizontal size is Ws and the vertical size is Hs) are known, the field of view can be calculated through trigonometric functions:
[0082] Accuracy rate:
[0083] Recall rate:
[0084] For a working area covering a length of L and a width of W, based on the installation height h and the field of view angle of the camera, the number of cameras required can be estimated. For example, in the horizontal direction, the number of cameras n h can be approximated as:
[0085]
[0086] where represents rounding up.
[0087] Similarly, in the vertical direction, the number of cameras required can be calculated as:
[0088]
[0089] ① Considering the field of view overlap
[0090] To ensure a non-blind spot monitoring, there usually needs to be a certain overlap in the field of view of the cameras. Assume the horizontal field of view overlap rate is p h (0 < p h < 1), and the vertical field of view overlap rate is p v (0 < p v < 1), and the specific value is determined according to the on-site installation situation.
[0091] The actual horizontal coverage width of a single camera becomes:
[0092]
[0093] The actual vertical coverage height of a single camera becomes:
[0094]
[0095] ② Considering the influence of installation position and angle
[0096] Assume that due to the changes in installation position and angle, the effective field of view angle in the horizontal direction becomes FOV heff = k h × FOV h , and the effective field of view angle in the vertical direction becomes FOV veff = k v × FOV v , where k h and k v are correction factors (0 < k h < 1, 0 < k v < 1), and the specific values need to be evaluated and determined according to the actual on-site installation situation.
[0097] ③Comprehensively consider the influence of field of view overlap, installation position and angle
[0098] Number of cameras in the horizontal direction:
[0099]
[0100] Number of cameras in the vertical direction:
[0101]
[0102] Arrange them in a matrix layout.
[0103] At this time, the total number of cameras N is the product of the number of cameras in the horizontal direction and the number of cameras in the vertical direction, that is: N = n h ×n v
[0104] 2) Calculation of camera resolution requirements
[0105] Image resolution is usually expressed in pixels, such as M×N (the number of horizontal pixels is M, and the number of vertical pixels is N). According to the Nyquist sampling theorem, in order to clearly distinguish the details in the working area, it is necessary to ensure that the resolution of the camera meets certain conditions.
[0106] Assume that the minimum size of the object to be distinguished in the working area is d. According to the working distance h and the required minimum resolution r (the number of pixels per unit length), the required minimum number of pixels in the horizontal and vertical directions are respectively:
[0107]
[0108] Ideally, it can be considered that each distinguishable minimum object corresponds to at least one pixel, that is, the relationship between r and d can be expressed as: (One pixel corresponds to the minimum distinguishable size):
[0109] The formula can be written as:
[0110]
[0111] II. Pose Estimation Module
[0112] Select the HRNet model and adjust the architecture according to the characteristics of steel structures. Use the PEFT method to effectively fine-tune the parameters. Combine the professional data collection and annotation process, covering the video streams of steel structure objects under various working conditions and the detailed annotation of key parts, so that the model has high adaptability and accuracy in the pose estimation of the lifted objects of tower crane steel structures.
[0113] 1) Data collection and annotation
[0114] Data collection: Use the camera installed on the tower crane in the actual working environment to capture the video streams of different types of steel structure objects under various working conditions. Ensure that all possible angles, lighting conditions, and weather conditions are covered.
[0115] Normalization: Normalize the image pixel values to the range of [0, 1] or [-1, 1].
[0116] Formula:
[0117] For RGB images, I min = 0, I max = 255.
[0118] Where:
[0119] I norm : The normalized image pixel value, that is, the pixel value within the specified range (such as [0, 1]) after processing;
[0120] I: The original image pixel value, that is, the pixel value without normalization processing;
[0121] I min : The minimum value of the image pixel value. For RGB images, I min = 0, indicating the minimum possible value of the pixel value;
[0122] I max : The maximum value of the image pixel value. For RGB images, I max = 255, indicating the maximum possible value of the pixel value.
[0123] Size adjustment: Ensure that all input images have the same resolution.
[0124] Scaling ratio: Where W and W t are the original width and the target width respectively. For the case of maintaining the aspect ratio, adjust the height H to H t .
[0125] Data annotation: Use professional image annotation tools (such as LabelImg) to annotate the key parts of each steel structure object. Define a set of pose key point sets suitable for steel structures and develop detailed annotation guidelines for each type. The annotation results should be saved in a standard format file (such as XML or JSON) for subsequent processing.
[0126] 2) Model selection and adjustment
[0127] Pre-trained model selection: Select the HRNet human pose estimation model.
[0128] Model Architecture Adjustment: Modify the model architecture according to the characteristics of steel structures:
[0129] Refined Plan for HRNet Model Architecture Adjustment
[0130] ① Design of New Convolution Module Structure: Aiming at the unique edge and texture features of steel structures and the objects lifted by tower cranes in tower crane operation images, design an attention module based on dilated convolution (DCA-Module). This module is stacked by multiple dilated convolution layers, combined with the channel attention mechanism, to enhance the ability to extract specific features. Dilated Convolution Layers: Dilated convolution can expand the receptive field of the convolution kernel without increasing the number of parameters and computational complexity. In the DCA-Module, set 3 dilated convolution layers. The dilation rate of the first dilated convolution layer is 2, the convolution kernel size is 3×3, and the stride is 1; the dilation rate of the second layer is 4, and the convolution kernel and stride remain unchanged; the dilation rate of the third layer is 8, and the convolution kernel and stride parameters are also kept. This can capture edge and texture information at different scales in sequence.
[0131] Channel Attention Mechanism: After passing through the dilated convolution layers, introduce the channel attention mechanism. First, perform global average pooling on the convolved feature map to compress the feature map into a 1×1×C vector (C is the number of channels). Then, through two fully connected layers (the number of neurons in the first fully connected layer is C / r, r is the compression rate, set to 16; the number of neurons in the second fully connected layer is restored to C) for feature fusion and weight adjustment. Finally, multiply the generated channel attention weights by the original feature map to enhance the attention to key channel features.
[0132] ② Parameter Settings of New Module: Taking the common resolution of tower crane operation images (such as 1080×1920) as an example, assume the number of input feature map channels is 64. In the dilated convolution layers, the output channel number of each convolution layer is set to 128 to ensure the richness and diversity of feature information. For the fully connected layers in the channel attention mechanism, according to the above settings, the weight matrix size of the first fully connected layer is 64×4, and the bias vector size is 4; the weight matrix size of the second fully connected layer is 4×64, and the bias vector size is 64.
[0133] ③ Location of the new module in HRNet: The DCA-Module is inserted into the backbone network of HRNet, specifically after the last bottleneck layer in stage2. This location is chosen because stage2 has already performed preliminary feature extraction on the image. Introducing the DCA-Module at this point can further enhance the extraction of the edge and texture features of the steel structure, while not adding excessive computational burden due to premature introduction. After inserting the DCA-Module, its output features will be used as the input for subsequent stage3 and other stages, participating in the entire feature fusion and pose estimation process.
[0134] ④ Principle of accuracy improvement: In tower crane operation images, the edge and texture features of the steel structure are important bases for judging its pose. The dilated convolutional layer of the DCA-Module can capture edges and textures at different scales through different dilation rates, effectively perceiving from fine structures to overall contours. The channel attention mechanism weights the channels according to the importance of the features, highlighting the key features related to the steel structure and suppressing irrelevant background information. In this way, the adjusted HRNet model can more accurately extract the features of the steel structure when processing tower crane operation images, providing a more reliable feature representation for subsequent pose estimation, thereby improving the accuracy of steel structure lifting object pose estimation.
[0135] PEFT fine-tuning strategy:
[0136] The PEFT method is used to fine-tune only a part of the model parameters. For example, using LoRA (such as a low-rank value r = 8, and the rank decomposition dimensions of the fine-tuning weight matrix are m×8 and 8×n) technology: LoRA realizes this by adding a low-rank decomposition perturbation matrix to the original weight matrix W: Let W = W0 + ΔW, where ΔW = AB T is a low-rank matrix, and A and B are two smaller matrices respectively.
[0137] The increased number of parameters is: r×(d in +d out ), where r is the low-rank value, and d in and d out are the input and output dimensions respectively.
[0138] 3) Training process
[0139] Data preprocessing: Preprocess the collected data, including operations such as cropping, scaling, and normalization, to ensure that the input data meets the model requirements.
[0140] Dataset division: Divide the dataset into a training set, a validation set, and a test set, with typical ratios of 70% for the training set, 15% for the validation set, and 15% for the test set.
[0141] Hyperparameter Settings: Set hyperparameters such as the learning rate η, batch size B, and number of iterations N. The initial values can be set according to experience or recommended values in the literature and adjusted dynamically during the training process based on the performance on the validation set.
[0142] Training Execution: Start the training process using the selected optimization algorithm (such as Adam or SGD). Regularly save the model weights and monitor the trend of the loss function to evaluate the training effect.
[0143] Loss Function: For the keypoint detection task, the commonly used loss function is the mean squared error (MSE):
[0144]
[0145] where:
[0146] L: The loss value, representing the error degree between the model prediction result and the true result;
[0147] N: The number of samples, i.e., the total number of keypoints participating in the loss calculation;
[0148] y i : The true value, referring to the actual annotation value of the keypoint coordinates;
[0149] Predicted value, referring to the keypoint coordinate value predicted by the model.
[0150] Optimization Algorithm: Update using the Adam optimizer:
[0151] m t =β1m t-1 +(1 - β1)g t
[0152]
[0153] where:
[0154] m t : The first - moment estimate value at the t - th moment, used to store the first - moment (mean) information of the gradient;
[0155] β1: The decay coefficient of the first - moment estimate, usually set to a value close to 1 (such as 0.9), controlling the retention ratio of historical first - moment information;
[0156] g t : The gradient value at the current moment, obtained by taking the derivative of the loss function with respect to the parameters;
[0157] v t : The second - moment estimate value at the t - th moment, used to store the second - moment (variance) information of the gradient;
[0158] β2: The attenuation coefficient of the second-order moment estimate, usually set to a value close to 1 (such as 0.999);
[0159] θ t : Model parameters updated at time t;
[0160] θ t-1 : Model parameters at time t-1;
[0161] η: learning rate, which controls the step size of parameter update;
[0162] ∈: A very small positive number (smoothing term) that prevents the denominator from being zero and ensures calculation stability.
[0163] Model evaluation: After training, evaluate the model performance on an independent test set. Calculate accuracy, recall, mAP and other indicators to measure the effectiveness of the model.
[0164] Accuracy:
[0165] Recall:
[0166] in:
[0167] TP (True Positive): True positive, that is, the number of samples that are actually positive and correctly predicted as positive;
[0168] TN (True Negative): True negative examples, that is, the number of samples that are actually negative and correctly predicted to be negative;
[0169] FP (False Positive): False positive examples, that is, the number of samples that are actually negative but are mistakenly predicted as positive;
[0170] FN (False Negative): False negative examples, that is, the number of samples that are actually positive but are mistakenly predicted as negative.
[0171] Mean Average Precision (mAP): Among them, AP i is the average precision of the i-th category.
[0172] 4) Integration and deployment
[0173] The fine-tuned model is integrated into the existing tower crane control system to achieve real-time attitude estimation capabilities.
[0174] In step S15, the corresponding risk results are predicted according to the motion trajectory of the key points of the steel structure hanging object and the posture of the steel structure hanging object;
[0175] In the specific implementation process of the present invention, the specific steps are:
[0176] S151: Obtain the motion trajectory of the key points of the steel structure lifting object;
[0177] S152: Determine the posture of the steel structure lifting object according to the relative positions of the key points of the steel structure lifting object;
[0178] S153: Associate the motion trajectory of the key points of the steel structure lifting object and the posture of the steel structure lifting object;
[0179] S154: Predict the corresponding risk result based on the motion trajectory of the key points of the steel structure lifting object and the posture of the steel structure lifting object;
[0180] At this time, risk assessment based on distance:
[0181] Assume that the position coordinates of the steel structure lifting object are (x lift , y lift , z lift ), and the position coordinates of the obstacle are (x obs , y obs , z obs ). The three-dimensional space Euclidean distance formula can be used to calculate the distance between the two:
[0182]
[0183] Hazard assessment based on speed:
[0184] If the velocity vector of the lifting object at a certain moment is its speed magnitude can be calculated:
[0185]
[0186] Set the speed threshold v thresh , when v > d thresh , it indicates that the speed is abnormal, which is one of the conditions for triggering a hazard warning;
[0187] Hazard assessment based on posture:
[0188] Calculation of the inclination angle: Assume that the posture of the steel structure lifting object can be measured by the angle between it and the direction of gravity. Taking the pitch angle θ as an example, if a unit vector of a certain axis of the lifting object is and the unit vector of the direction of gravity is the cosine value of the pitch angle can be calculated according to the vector dot product formula:
[0189]
[0190] Furthermore, obtain
[0191]
[0192] When θ exceeds a certain threshold θthresh An early warning can be triggered.
[0193] In an embodiment of the present application, the movement trajectory of the key points of the steel structure lifting object is obtained; multiple movement nodes are determined according to the division of the movement trajectory of the key points of the steel structure lifting object, realizing the division of the movement trajectory of the key points of the steel structure lifting object, introducing multiple movement nodes, and realizing the control of multiple movement nodes.
[0194] At this time, the movement trajectory of the key points of the steel structure lifting object is obtained, and multiple movement nodes are divided according to the movement trajectory of the key points of the steel structure lifting object. These nodes usually represent the position or state of the lifting object at specific time points, such as the starting point, key points (such as turning, accelerating, decelerating, etc.) and the ending point. By dividing the movement nodes, we can more clearly understand the movement state of the lifting object during the entire lifting process and provide precise control points for subsequent lifting operations.
[0195] First, we need to conduct a detailed analysis of the obtained movement trajectory to understand the overall movement trend and key change points of the lifting object. According to the analysis results, multiple key nodes are determined on the movement trajectory of the lifting object. These nodes should be able to comprehensively reflect the movement state of the lifting object and facilitate subsequent control and adjustment. For each determined node, we need to record its position information (such as coordinates), time stamp, and other relevant parameters (such as speed, acceleration, etc.). These information will be used for subsequent lifting operation control and risk assessment.
[0196] Specifically, after obtaining the movement trajectory of the steel structure prefabricated component, the project team began to divide the movement nodes according to the trajectory. They first analyzed the overall movement trend of the lifting object and found that the lifting object needs to go through multiple stages such as lifting, translation, turning, and landing during the lifting process. Then, they determined multiple movement nodes at the key positions of each stage, such as the lifting point, translation midpoint, turning point, and landing point, etc. For each determined node, the project team recorded its position information, time stamp, and corresponding parameters such as speed and acceleration. These information will be used for subsequent lifting operation control to ensure that the lifting object can move along the predetermined trajectory and speed. At the same time, these information will also be used for risk assessment to identify potential safety hazards and take corresponding preventive measures.
[0197] Furthermore, among each movement node, the aerial image of the steel structure lifting object is collected; the posture of the steel structure lifting object is determined according to the relative position of the key points of the steel structure lifting object, realizing the recognition of the aerial image of the steel structure lifting object and ensuring the accuracy of the posture of the steel structure lifting object.
[0198] At this time, aerial images of the steel structure lifting object are collected at each motion node. This step is usually completed using various cameras. The purpose of image collection is to obtain the visual information of the lifting object at a specific time point for subsequent analysis and determination of its posture. At the same time, cameras or drones with high resolution and good field of view are selected to ensure that clear images of the lifting object can be captured. Cameras or drones are arranged at the lifting site to ensure that they can cover the field of view of all motion nodes. At the same time, considering the safety and feasibility of the lifting operation, the equipment should be avoided being arranged in overly dangerous positions or positions that affect the lifting operation. When each motion node arrives, the image collection device is started in a timely manner to ensure that real-time images of the lifting object can be captured. At the same time, considering the time delay in image collection and processing, the device may need to be started in advance to obtain sufficient image data. During the collection process, attention should be paid to the impact of factors such as light conditions and background interference on the image quality. If necessary, image enhancement techniques can be used to improve the clarity and contrast of the images.
[0199] Image processing or computer vision techniques are used to identify and analyze the posture of the steel structure lifting object in the aerial images. The purpose of this step is to determine the posture information of the lifting object at a specific time point, such as the tilt angle, bending degree, etc. At the same time, according to the characteristics of the shape, size, and texture of the lifting object, appropriate image processing algorithms are selected to extract the contour and features of the lifting object. Commonly used algorithms include edge detection, shape matching, feature extraction, etc. After the contour and features of the lifting object are identified, corresponding algorithms are needed to extract the posture information of the lifting object. This usually involves locating and analyzing key points in the image to determine parameters such as the tilt angle and bending degree of the lifting object. To ensure the accuracy of the extracted posture information, the results need to be verified and calibrated. Other sensor data (such as accelerometers, gyroscopes, etc.) can be used to assist in verification, or multiple images can be comprehensively analyzed to improve accuracy. The extracted posture information is processed and analyzed to generate a visual report or for subsequent lifting operation control.
[0200] Therefore, the motion trajectories of the key points of the steel structure lifting object are associated with the posture of the steel structure lifting object; based on the motion trajectories of the key points of the steel structure lifting object and the posture of the steel structure lifting object, the corresponding risk results are predicted, which incorporates the overall consideration of the motion trajectories of the key points of the steel structure lifting object and the posture of the steel structure lifting object, ensuring the accurate prediction of the corresponding risk results.
[0201] At this time, associate the motion trajectory of the key points of the steel structure lifting object with its posture. The purpose of this step is to establish a complete motion model of the lifting object in space, so as to more comprehensively understand its motion state and potential risks. Ensure that the collected motion trajectory data and posture data are synchronized in time. This may require using timestamps or other synchronization mechanisms to align different data sources. Convert the motion trajectory and posture data into the same coordinate system for spatial association and analysis. Use data fusion technology to combine the motion trajectory and posture data to form a comprehensive motion model. This may involve data processing techniques such as interpolation, filtering, and smoothing. To visually display the association results, three-dimensional visualization tools can be used to present the motion trajectory and posture data graphically.
[0202] It is necessary to use a risk prediction model or algorithm to predict potential risk results based on the motion trajectory and posture of the key points of the steel structure lifting object. The purpose of this step is to identify possible safety hazards in advance and take corresponding preventive measures to reduce risks. First, it is necessary to identify potential risk factors related to the lifting operation, such as equipment failures, operation errors, environmental interferences, etc. Use mathematical models or algorithms to describe the relationship between risk factors and risk results. This may involve techniques such as statistical analysis, machine learning, and deep learning. Input the collected motion trajectory and posture data into the risk prediction model as part of the model input. Use the risk prediction model to calculate potential risk results, such as the probability of an accident occurring and the possible losses. Analyze the calculated risk results and formulate corresponding countermeasures based on the analysis results to reduce risks.
[0203] 1) Two-dimensional key point detection
[0204] First, use the optimized HRNet to process the images captured by each camera to obtain the two-dimensional key point coordinates of the lifting object from various perspectives. Suppose there are n cameras, and each camera can provide a set of two-dimensional key point coordinates where m is the number of key points.
[0205] 2) Camera parameter calibration
[0206] Ensure that the internal parameter matrix K and external parameter matrices (rotation matrix K and translation vector T) of all cameras have been accurately calibrated. These parameters are used to convert two-dimensional image coordinates into three-dimensional world coordinates.
[0207] For the jth camera, its camera parameters include:
[0208] Internal parameter matrix:
[0209]
[0210] where:
[0211] f x,j : Focal length of the j-th camera in the x direction of the image, reflecting the scaling ability of the lens in the x direction.
[0212] f y,j : Focal length of the j-th camera in the y direction of the image, reflecting the scaling ability of the lens in the y direction.
[0213] c x,j : Abscissa of the principal point of the j-th camera image, i.e., the pixel coordinate of the intersection of the optical axis and the imaging plane in the x direction of the image.
[0214] c y,j : Ordinate of the principal point of the j-th camera image, i.e., the pixel coordinate of the intersection of the optical axis and the imaging plane in the y direction of the image.
[0215] Extrinsic parameter matrix: Rotation matrix R j and translation vector T j
[0216] 3) Triangulation to calculate 3D coordinates
[0217] Use the 2D key point coordinates from at least two different viewpoints for triangulation to determine the key point positions in 3D space. Here, the method of minimizing the reprojection error is used to estimate the optimal 3D coordinates.
[0218] Assume that the 3D coordinates of the key point observed from the j-th camera in the world coordinate system are (X, Y, Z), then the projected coordinates (u j , v j ) on the image plane can be calculated by the following formula:
[0219]
[0220] where s j is the scaling factor, K j is the intrinsic parameter matrix of the j-th camera, R J is the rotation matrix, and T J is the translation vector. Expanding:
[0221] s j u j = f x,j (r 11,j X + r 12,j Y + r 13,j Z + t 1,j ) + c x,j s j
[0222] s j v j = f y,j (r21,j X + r 22,j Y + r 23,j Z + t 2,j ) + c y,j s j
[0223] Wherein:
[0224] s j : The scaling factor of the j-th camera, used for the conversion from homogeneous coordinates to pixel coordinates;
[0225] u j 、v j : The pixel coordinates (horizontal and vertical coordinates) of the three-dimensional space point projected onto the image plane of the j-th camera;
[0226] f x,j 、f y,j : The focal lengths of the j-th camera in the x and y directions (parameters of the internal parameter matrix);
[0227] r 11,j 、r 12,j 、r 13,j 、r 21,j 、r 22,j 、r 23,j ; The elements of the rotation matrix of the j-th camera,
[0228] describing the rotation relationship of the three-dimensional space point from the world coordinate system to the camera coordinate system;
[0229] t 1,j 、t 2,j : The components of the translation vector Tj, representing the translation amount of the camera coordinate system relative to the world coordinate system;
[0230] X, Y, Z: The coordinates of the three-dimensional space point in the world coordinate system;
[0231] To solve for (X, Y, Z), a non-linear optimization problem needs to be solved, with the goal of minimizing the sum of the reprojection errors of all views. This can be achieved by non-linear optimization methods such as the Levenberg-Marquardt algorithm.
[0232] The mathematical expression is as follows:
[0233]
[0234] Wherein, π represents the perspective projection operation, i.e.:
[0235]
[0236] The Levenberg - Marquardt (LM) algorithm is an iterative method that combines the advantages of gradient descent and the Gauss - Newton method and is applicable to solving nonlinear least - squares problems. In each iteration, the LM algorithm attempts to find a new parameter estimate that reduces the cost function E(X, Y, Z).
[0237] The basic update rule can be expressed as:
[0238] Δx=-(J T J+λI) -1 J T r
[0239] where,
[0240] J is the Jacobian matrix, which contains the partial derivatives with respect to each observation:
[0241] Let Taking the partial derivative of the reprojection error function with respect to x gives the Jacobian matrix J. Taking the u - coordinate
[0242] as an example (similarly for v):
[0243]
[0244] The calculation for the v - coordinate is similar. Finally, the Jacobian matrix J is obtained as:
[0245]
[0246] λ is a parameter that controls the behavior of the algorithm. It is large in the early iterations to ensure stability and gradually decreases as it approaches the optimal solution.
[0247] r is the residual vector, which is the difference between the current predicted value and the actual observed value.
[0248] The updated parameter is:
[0249] x new =x old +Δx
[0250] By continuously iterating the above process until the convergence condition is reached or the preset maximum number of iterations is reached, the values of (X, Y, Z) that minimize E(X, Y, Z) are finally found, and the coordinates of the key points of the lifted object in three - dimensional space can be obtained.
[0251] According to experience and experimental tests, the maximum number of iterations is set to N max =200. In most cases, after 200 iterations, the algorithm can converge to a relatively reasonable result. If the number of iterations is too small, accurate three - dimensional coordinates may not be obtained; while if the number of iterations is too large, the calculation time will increase, affecting real - time performance.
[0252] Set the reprojection error change threshold ε = 10 -6 . In each iteration process, calculate the difference between the reprojection error of this iteration and the previous iteration. When this difference is less than, it is considered that the algorithm has converged, and the three-dimensional coordinates obtained at this time are the final result.
[0253] IV. Motion Prediction Model Module
[0254] Construct a Seq2Seq model based on the Bahdanau attention mechanism, fuse the pose estimation results with external factors such as wind speed and operator input to construct a dataset, use LSTM or GRU units to process sequence data, capture long-term dependencies with the help of the attention mechanism, train and optimize with the mean squared error as the loss function, and accurately predict the motion trajectory of the suspended object.
[0255] 1) Data Preparation
[0256] Use the previously obtained pose estimation results as part of the input data, and combine other external factors (such as wind speed, operator input, etc.) to construct a dataset containing time series features.
[0257] Suppose there is a time series dataset It contains external factors such as pose estimation results, wind speed, and operator input. The data at each time t can be represented as a vector
[0258] x t = [p t , v t , w t , o t
[0259] Wherein:
[0260] p t is the pose estimation result (such as the position of key points),
[0261] v t is the speed information,
[0262] w t is the wind speed,
[0263] o t is the operator's input.
[0264] 2) Model Architecture Design
[0265] Construct a Seq2Seq (Sequence-to-Sequence) model based on the Bahdanau attention mechanism. This model structure allows capturing the long-term dependencies of the suspended object's motion, thereby more accurately predicting its future position and speed.
[0266] In the encoder part, LSTM or GRU cells can be used to process the input sequence; the decoder part is responsible for generating the predicted output.
[0267] Bahdanau Attention Mechanism: When applying the Bahdanau attention mechanism in the Seq2Seq model, calculate the context vector C t as the weighted average at the current time step t, where the weights are determined by the similarity between the query vector Q t and all encoder hidden states H i
[0268] Attention Weight Calculation:
[0269]
[0270] where: e t,i = v T tanh(W h H i + W q Q t )
[0271] Context Vector: C t = ∑ i α t,i H i
[0272] Loss Function: The mean squared error (MSE) is used as the loss function to measure the difference between the predicted result and the true value:
[0273]
[0274] where y i represents the true value, represents the predicted value.
[0275] 3) Training Process
[0276] Input the prepared dataset into the Seq2Seq model for training.
[0277] Initialize Model Parameters: Initialize the weight matrices and other parameters of the LSTM or GRU cells.
[0278] Forward Propagation: For a given time series input {x1, x2,..., x T}, generate a sequence of hidden states {H1, H2,..., H T} through the encoder
[0279] Attention Mechanism Application: For each decoder time step t, calculate the context vector C t and generate the output in combination with the current decoder hidden state.
[0280] Loss calculation: Calculate the difference between the predicted output and the true label using the MSE loss function.
[0281] Backpropagation and parameter update: Adjust the model parameters according to the loss gradient.
[0282] Use an appropriate loss function (such as Mean Squared Error, MSE) to measure the difference between the predicted value and the actual value, and update the model parameters through backpropagation.
[0283] Dynamically adjust the model structure and hyperparameters through the iterative training process until the prediction accuracy meets the preset threshold requirements.
[0284] 4) Integration and application
[0285] Deploy the trained Seq2Seq model to the tower crane control system, use it to predict the future movement trajectory of the lifted object in real time, provide decision-making support for the operator, and help prevent potential safety hazards.
[0286] Hazard discrimination and warning module:
[0287] Based on preset safety rules covering multiple aspects such as the position, speed, and attitude of the lifted object, use machine learning algorithms such as decision trees, support vector machines, or random forests to accurately judge dangerous situations, set up a multi-level warning mechanism based on attitude and trajectory information, and automatically trigger alarms or take emergency measures according to the danger level to effectively prevent tower crane operation accidents.
[0288] 1) Danger situation judgment and warning mechanism
[0289] Rule definition: According to the preset safety rules (such as the lifted object exceeding the predetermined area, approaching an obstacle, abnormal speed, etc.), the system will automatically trigger an alarm to notify the operator or take emergency measures to stop the tower crane operation. Machine learning algorithms such as decision trees, support vector machines (SVM), or random forests can be used to define and classify different dangerous situations.
[0290] Combined with attitude and trajectory information: When judging dangerous situations, not only consider the current position and speed of the lifted object, but also combine its attitude information (such as rotation angle, inclination degree, etc.). For example, if the lifted object is approaching an obstacle and in an unstable attitude (such as excessive inclination), the system should immediately trigger an emergency stop.
[0291] Multi-level warning mechanism: According to the attitude and trajectory information of the lifted object, set multiple warning levels (such as warning, serious warning, emergency stop), and take corresponding countermeasures according to different levels.
[0292] 2) Safety rule definition
[0293] According to preset safety rules (such as the lifted object exceeding the predetermined area, approaching obstacles, abnormal speed, etc.), the system will automatically trigger an alarm to notify the operator or take emergency measures to stop the tower crane operation.
[0294] ① Distance-based risk assessment
[0295] Assume the position coordinates of the lifted object are (x lift , y lift , z lift ), and the position coordinates of the obstacle are (x obs , y obs , z obs ). The three-dimensional Euclidean distance formula can be used to calculate the distance between the two:
[0296]
[0297] In the safety rules, a threshold d thresh can be set. When d < d thresh , the corresponding early warning is triggered.
[0298] ② Speed-based risk assessment
[0299] If the velocity vector of the lifted object at a certain moment is , its speed magnitude can be calculated:
[0300]
[0301] Set a speed threshold v thresh . When v > d thresh , it indicates abnormal speed, which is one of the conditions for triggering a risk early warning.
[0302] ③ Attitude-related risk assessment
[0303] Tilt angle calculation: Assume that the attitude of the lifted object can be measured by the angle between it and the direction of gravity. Taking the pitch angle θ as an example, if a unit vector of a certain axis of the lifted object is and the unit vector of the gravity direction is , the cosine value of the pitch angle can be calculated according to the vector dot product formula:
[0304]
[0305] Then
[0306]
[0307] When θ exceeds a certain threshold θ thresh , the early warning can be triggered.
[0308] Comprehensive Attitude Hazard Assessment: Construct a comprehensive attitude hazard assessment index A, considering multiple attitude parameters such as pitch angle θ, roll angle φ, and yaw angle γ, and set weights w1, w2, w3. Then A = w1θ + w2φ + w3γ. When the set threshold is exceeded, a corresponding level of warning is triggered.
[0309] Safety Rule Assessment:
[0310]
[0311] Where c is the hazard level,
[0312] state is the current state of the system (including attitude and trajectory information),
[0313] P(c|state) is the probability distribution based on the machine learning model.
[0314] 3) Multi-level Warning Mechanism
[0315] Based on the assessment of the attitude and trajectory information of the lifted object and various hazard parameters, multiple warning levels are set comprehensively, and corresponding countermeasures are taken according to different levels.
[0316] Warning Level Judgment:
[0317]
[0318] Where:
[0319] "Normal" means normal, and the facility is working normally;
[0320] "Warning" means warning, and the staff should eliminate the warning;
[0321] "Severe Warning" means severe warning, and the staff should immediately check the situation and eliminate the warning;
[0322] "Emergency Stop" means emergency situation, and the equipment stops urgently or enters the emergency state.
[0323] In step S16, determine the dynamic measures for the steel structure lifted object according to the risk result, the steel structure lifted object, and the tower crane, and perform safety control on the steel structure lifted object;
[0324] In the specific implementation process of the present invention, the specific steps are:
[0325] S161: Obtain the risk result;
[0326] S162: Associate the risk result, the steel structure lifted object, and the tower crane;
[0327] S163: Determine the first measure parameter based on the risk result and the steel structure lifted object;
[0328] S164: Determine the second measure parameter according to the risk result and the tower crane.
[0329] S165: Based on the first measure parameter and the second measure parameter, determine the dynamic measure for the steel structure lifted object, and conduct safety control over the steel structure lifted object.
[0330] In the embodiment of the present application, obtain the risk result; associate the risk result, the steel structure lifted object, and the tower crane, introduce the risk result, the steel structure lifted object, and the tower crane, and overall control the risk result, the steel structure lifted object, and the tower crane.
[0331] Furthermore, determine the first measure parameter based on the risk result and the steel structure lifted object; determine the second measure parameter according to the risk result and the tower crane; based on the first measure parameter and the second measure parameter, determine the dynamic measure for the steel structure lifted object, and conduct safety control over the steel structure lifted object, which is compatible with the overall consideration of the first measure parameter and the second measure parameter, ensures the accuracy of the dynamic measure for the steel structure lifted object, and facilitates the safety control of the steel structure lifted object.
[0332] At this time, determine the first set of measure parameters according to the risk assessment result and the specific characteristics of the steel structure lifted object. These parameters may involve aspects such as the selection of the lifting plan, the setting of the lifting points, the selection of the slings and rigging, and the dynamic adjustment during the lifting process. The purpose is to ensure the safety and stability of the steel structure lifted object during the lifting process.
[0333] It is necessary to determine the second set of measure parameters according to the risk assessment result and the performance characteristics of the tower crane. These parameters may involve aspects such as the selection of the tower crane, the working radius, the lifting capacity, the stability, and the qualifications of the operators. The purpose is to ensure the safety and reliability of the tower crane during the lifting process.
[0334] Integrate the first measure parameter and the second measure parameter to formulate a dynamic measure plan for the steel structure lifted object. These dynamic measures may involve aspects such as real-time monitoring during the lifting process, early warning mechanism, and emergency handling. The purpose is to conduct all-round safety control over the steel structure lifted object during the lifting process, ensure the safety and smooth progress of the lifting operation. Use sensors and monitoring systems to real-time monitor information such as the shape, position, and height of the lifted object, as well as the working state and stability of the tower crane. Set the early warning threshold, and trigger an early warning signal when the monitored data exceeds the threshold to remind the operator to take measures for adjustment in a timely manner. Formulate an emergency handling plan, clarify the response measures and responsible persons in case of emergencies, and ensure that various emergencies can be quickly and effectively responded to.
[0335] Embodiment Three
[0336] Please refer to Figure 3 , Figure 3It is a schematic structural diagram of a lifting object attitude prediction system for tower crane steel structures based on multiple cameras in an embodiment of the present invention.
[0337] As Figure 3 shown, a lifting object attitude prediction system for tower crane steel structures based on multiple cameras, the lifting object attitude prediction system for tower crane steel structures based on multiple cameras includes:
[0338] An acquisition module 21, configured to collect hoisting object videos at different angles when multiple cameras are installed at the tower crane operation site;
[0339] A key node module 22, configured to determine multiple key nodes of the steel structure hoisting object based on the hoisting object video and determine their two-dimensional spatial coordinates in different videos;
[0340] A spatial position module 23, configured to determine the coordinates of the steel structure hoisting object in three-dimensional space according to the two-dimensional spatial coordinates of the key nodes of the steel structure hoisting object from at least two different perspectives;
[0341] A trajectory module 24, configured to predict the movement trajectory of the key points of the steel structure hoisting object based on the spatial coordinates of the key points of the steel structure hoisting object, the hoisting object attitude, and the external environment;
[0342] A risk result module 25, configured to predict the corresponding risk result according to the movement trajectory of the key points of the steel structure hoisting object and the attitude of the steel structure hoisting object;
[0343] A control module 26, configured to determine the dynamic measures of the steel structure hoisting object according to the risk result, the steel structure hoisting object, and the tower crane, and perform safety control on the steel structure hoisting object.
[0344] Arbitrary combinations of the technical features of the above embodiments are made. For the sake of brevity of description, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
Claims
1. A method for predicting the attitude of the lifted object of the tower crane steel structure based on multiple cameras, characterized in that, Including: When multiple cameras are installed at the tower crane operation site, the multiple cameras collect hoisted object videos at different angles; Based on the hoisted object videos, determine multiple key nodes of the steel structure hoisted object and determine their two-dimensional spatial coordinates in different videos; According to the two-dimensional spatial coordinates of the key nodes of the steel structure hoisted object from at least two different perspectives, determine its coordinates in the three-dimensional space; Based on the spatial coordinates of the key points of the steel structure hoisted object, the hoisted object attitude, and the external environment, predict the movement trajectory of the key points of the steel structure hoisted object; According to the movement trajectory of the key points of the steel structure hoisted object and the attitude of the steel structure hoisted object, predict the corresponding risk results; According to the risk results, the steel structure hoisted object, and the tower crane, determine the dynamic measures for the steel structure hoisted object and perform safety control on the steel structure hoisted object.
2. The method for predicting the attitude of the lifted object of the tower crane steel structure based on multiple cameras according to claim 1, wherein, When multiple cameras are installed at the tower crane operation site, the multiple cameras collect hoisted object videos at different angles, including: Determine the positions of multiple cameras according to the hanging position of the tower crane and the shape of the tower crane; When multiple cameras are installed at the tower crane operation site, the multiple cameras are all oriented towards the hanging position of the tower crane; The multiple cameras photograph the hanging position of the tower crane, and the multiple cameras collect hoisted object videos at different angles; the multiple hoisted object videos contain the steel structure hoisted object.
3. The method for predicting the attitude of the lifted object of the tower crane steel structure based on multiple cameras according to claim 2, wherein Based on the hoisted object videos, determine multiple key nodes of the steel structure hoisted object and determine their two-dimensional spatial coordinates in different videos, including: Obtain multiple hoisted object videos; Based on the recognition of the shape of the steel structure hoisted object, determine multiple key nodes of the steel structure hoisted object and determine their two-dimensional spatial coordinates in different videos, and the multiple key nodes are two-dimensional key points.
4. The method for predicting the attitude of the lifted object of the tower crane steel structure based on multiple cameras according to claim 3, wherein According to the two-dimensional spatial coordinates of the key nodes of the steel structure hoisted object from at least two different perspectives, determine its coordinates in the three-dimensional space, including: Obtain multiple key nodes of the steel structure hoisted object; Mark the corresponding coordinates for the multiple key nodes; Use the coordinates of the multiple key nodes from at least two different perspectives for triangulation to determine the position of the key point in the three-dimensional space; According to the continuous iteration of the position of the key point in the three-dimensional space, determine the coordinates of the steel structure hoisted object in the three-dimensional space, and the spatial position of the steel structure hoisted object matches the corresponding three-dimensional coordinates.
5. The method for predicting the attitude of the hoisted object of the tower crane steel structure based on multiple cameras according to any one of claims 1 to 4, characterized in that, Based on the spatial position of the key points of the steel structure hoisted object, the hoisted object attitude, and the external environment, predict the movement trajectory of the key points of the steel structure hoisted object, including: Obtain the spatial position of the key points of the steel structure hoisted object; Based on the determined spatial position of the key points, determine the hoisted object attitude; Collect multiple environmental parameters based on the environmental detection of the spatial position of the steel structure hoisted object; Determine the corresponding external environment according to the multiple environmental parameters; Associate the spatial position of the key points of the steel structure hoisted object, the hoisted object attitude, and the external environment; According to the spatial position of the key points of the steel structure hoisted object, the hoisted object attitude, and the external environment, predict the movement trajectory of the key points of the steel structure hoisted object.
6. The method for predicting the attitude of the lifted object of the tower crane steel structure based on multiple cameras according to claim 5, wherein According to the movement trajectory of the key points of the steel structure hoisted object and the attitude of the steel structure hoisted object, predict the corresponding risk results, including: Obtain the movement trajectory of the key points of the steel structure hoisted object; Determine the attitude of the steel structure hoisted object according to the relative position of the key points of the steel structure hoisted object.
7. The method for predicting the attitude of the hoisted object of the tower crane steel structure based on multiple cameras according to claim 6, wherein, According to the movement trajectory of the key points of the steel structure hoisted object and the attitude of the steel structure hoisted object, predicting the corresponding risk results also includes: Associate the motion trajectory of the key points of the steel structure lifting object and the posture of the steel structure lifting object; Predict the corresponding risk results based on the motion trajectory of the key points of the steel structure lifting object and the posture of the steel structure lifting object; At this time, risk assessment based on distance: Assume that the position coordinates of the steel structure hoisted object are (x lift , y lift , z lift ), and the position coordinates of the obstacle are (x obs , y obs , z obs ). The three-dimensional space Euclidean distance formula can be used to calculate the distance between the two: Hazard assessment based on speed: If the velocity vector of the steel structure lifting object at a certain moment is Its velocity magnitude can be calculated as follows: Set the speed threshold v thresh When v > d thresh it indicates that the speed is abnormal, which is one of the conditions for triggering a danger warning; Hazard assessment based on posture: Calculation of the inclination angle: Assume that the attitude of the steel structure lifting object can be measured by the angle between it and the direction of gravity. Taking the pitch angle θ as an example, if a unit vector in a certain axial direction of the lifting object is and the unit vector in the direction of gravity is the cosine value of the pitch angle can be calculated according to the vector dot product formula: Furthermore, obtain When θ exceeds a certain threshold value θ thresh a warning can be triggered.
8. The method for predicting the attitude of the lifted object of the tower crane steel structure based on multiple cameras according to claim 7, characterized in that, Determine the dynamic measures of the steel structure lifting object according to the risk results, the steel structure lifting object and the tower crane, and perform safety control on the steel structure lifting object, including: Obtain the risk results; Associate the risk results, the steel structure lifting object and the tower crane.
9. The method for predicting the attitude of the lifted object of the tower crane steel structure based on multiple cameras according to claim 8, wherein, Determine the dynamic measures of the steel structure lifting object according to the risk results, the steel structure lifting object and the tower crane, and perform safety control on the steel structure lifting object, and further include: Determine the first measure parameter based on the risk results and the steel structure lifting object; Determine the second measure parameter according to the risk results and the tower crane; Determine the dynamic measures of the steel structure lifting object based on the first measure parameter and the second measure parameter, and perform safety control on the steel structure lifting object.
10. A tower crane steel structure lifting object attitude prediction system based on multiple cameras, characterized in that, The multi-camera-based tower crane steel structure lifting object posture prediction system is applied to the multi-camera-based tower crane steel structure lifting object posture prediction method as described in any one of claims 1-9. The multi-camera-based tower crane steel structure lifting object posture prediction system includes: An acquisition module, configured to collect the video of the lifting object at different angles by multiple cameras when the multiple cameras are installed at the tower crane operation site; A key node module, configured to determine multiple key nodes of the steel structure lifting object based on the lifting object video and determine their two-dimensional spatial coordinates in different videos; A spatial position module, configured to determine the coordinates of the steel structure lifting object in the three-dimensional space according to the two-dimensional spatial coordinates of the key nodes of at least two different perspectives of the steel structure lifting object; A trajectory module, configured to predict the motion trajectory of the key points of the steel structure lifting object based on the spatial coordinates of the key points of the steel structure lifting object, the posture of the lifting object, and the external environment; A risk result module, configured to predict the corresponding risk results according to the motion trajectory of the key points of the steel structure lifting object and the posture of the steel structure lifting object; A control module, configured to determine the dynamic measures of the steel structure lifting object according to the risk results, the steel structure lifting object and the tower crane, and perform safety control on the steel structure lifting object.
Citation Information
Cited By
A tower crane operation posture anomaly identification processing method, system, device and medium
CN122530937A