A hoisting operation dynamic risk identification method based on space-time trajectory prediction and 3D reconstruction

CN122551271APending Publication Date: 2026-08-11JIANGSU ANSHENG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]尽管上述已有技术在一定程度上实现了吊装作业的安全监控,但在实际工程应用中,这两类技术一方面均试图在2D图像平面解决3D空间的碰撞防护问题,缺乏有效的深度感知能力,受透视投影原理限制,无法区分目标前后遮挡关系,易出现危险区划定逻辑悖论,引发频繁的误报与漏报;另一方面风险评估仅基于当前帧静态位置信息,缺乏对吊物动态惯性风险的预判能力,风险量化指标无明确物理因果关系,不仅评估精度有限,还存在严重的响应滞后问题,难以适配复杂吊装场景的实时安全管控需求

Benefits of technology

1、通过像素级深度图,将2D坐标转化为3D坐标,同时通过深度信息区分遮挡与重叠,误报率降低约85%,漏报率降低约90%;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551271A_ABST
    Figure CN122551271A_ABST
Patent Text Reader

Abstract

This application relates to the technical field of intelligent monitoring of construction and logistics equipment, and particularly to a dynamic risk identification method for hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction. The method includes real-time acquisition of video stream data of the hoisting operation scene, extracting structured data of the 3D spatial state of the hoisted object and personnel targets frame by frame from the video stream; calculating the spatiotemporal proximity parameters of the current spatial distance and future collision trend between the hoisted object and personnel targets based on the structured data; comparing the dynamic trigger threshold of the current operation risk level with the spatiotemporal proximity parameters to determine whether the current frame meets the triggering conditions for full trajectory prediction; performing preset lightweight state monitoring on frames that do not meet the prediction triggering conditions and non-critical steady-state frames, and determining whether to trigger full trajectory prediction based on the monitoring results. This application has the advantage of improving the accuracy and timeliness of risk identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent monitoring of building construction and logistics equipment, and in particular to a method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction. Background Technology

[0002] Currently, safety monitoring during hoisting operations generally uses 2D cameras to capture images of the operation. The mainstream technical solutions are divided into two categories: one is visual perimeter protection technology based on 2D geometric rules, which uses Canny edge detection and background subtraction to extract the outline of the hoisted object, and generates a closed-loop "virtual control area" with a fixed pixel width on the image plane based on the 2D outline through morphological dilation. The algorithm logic relies on hard-coded rules, such as "the distance between the hook and the hoisted object remains unchanged for N consecutive frames" to determine hoisting, and only determines whether the 2D pixel coordinates of the personnel fall into the virtual control area to trigger an alarm.

[0003] Another type is a risk identification technique that combines general object detection with statistical correction. It uses an improved YOLO algorithm for object detection, and to address the issue of missed detections, it employs cascaded co-occurrence matrix statistical feature matching and LightGlue feature point matching algorithms, utilizing inter-frame displacement to infer target positions. For risk assessment, a "region rejection space" is defined, but in actual calculations, Harris corner detection and DBSCAN clustering algorithms are used to statistically analyze the corner density within specific areas of workers and cranes. When the density exceeds a threshold, it is considered a risk.

[0004] While the aforementioned technologies have achieved safety monitoring of hoisting operations to some extent, in practical engineering applications, both types of technologies, on the one hand, attempt to solve the collision protection problem in 3D space on a 2D image plane, lacking effective depth perception capabilities. Limited by the principle of perspective projection, they cannot distinguish the occlusion relationship between the target and the front and back, easily leading to logical paradoxes in the delineation of dangerous zones, resulting in frequent false alarms and missed alarms. On the other hand, risk assessment is based only on the static position information of the current frame, lacking the ability to predict the dynamic inertial risks of the hoisted object. The risk quantification indicators do not have a clear physical causal relationship, resulting in limited assessment accuracy and serious response lag, making it difficult to adapt to the real-time safety management needs of complex hoisting scenarios. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a method for dynamic risk identification in hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction.

[0006] Firstly, this application provides a method for dynamic risk identification in hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction, employing the following technical solution: A dynamic risk identification method for hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction includes the following steps: Real-time acquisition of video stream data of hoisting operation scene, and extraction of structured data of 3D spatial state of hoisted target and personnel target in video stream frame by frame; Based on structured data, the spatiotemporal proximity parameters of the current spatial distance and future collision trend between the suspended object target and the personnel target are calculated. The dynamic trigger threshold of the current operation risk level is obtained and compared with the spatiotemporal proximity parameter to determine whether the current frame meets the trigger conditions for full trajectory prediction. If the current frame meets the prediction triggering conditions, the current frame will be used as the first key frame, and the structured data corresponding to the key frame will be used for the preset full trajectory prediction. A risk score will be obtained based on the prediction results. If the risk score is higher than or equal to the preset warning threshold, a control signal is output based on the risk score; if the risk score is lower than the preset warning threshold, the state change is identified for subsequent consecutive frames in the triggered state, the frames with state changes greater than the preset change threshold are locked as new key frames, and the full trajectory prediction is retried; and the frames with state changes less than the preset change threshold are defined as non-key steady-state frames. For frames that do not meet the prediction trigger conditions and non-critical steady-state frames, perform preset lightweight state monitoring, and determine whether to trigger full trajectory prediction based on the monitoring results.

[0007] In one embodiment, the full trajectory prediction step specifically includes: Historical trajectories are generated based on the historical structured data corresponding to the suspended target, and the final predicted trajectory is obtained by combining physical flow deduction and neural network residual correction. A two-dimensional Gaussian distribution is defined for each point on the final predicted trajectory. Its covariance matrix deforms in real time with the velocity vector of the suspended object, and the potential field strength is proportional to the potential energy of the suspended object. Kalman filtering or linear extrapolation is used to predict the short-term simple trajectory of personnel, and the spatiotemporal overlap integral between the personnel trajectory and the dynamic potential field is calculated to obtain a quantitative risk score.

[0008] In one embodiment: the step of obtaining the final predicted trajectory specifically includes: The suspended object is abstracted as a simple pendulum system. Based on the current structured data and historical trajectory of the suspended object, the pure physical theoretical trajectory of the pendulum system at future moments is deduced. Based on a pre-trained neural network model, the position residual for future moments is output using the current structured data and historical trajectory of the suspended target. The final predicted trajectory is obtained by fusing the physical theoretical trajectory with the position residual predicted by the neural network.

[0009] In one embodiment: the step of acquiring the structured data includes: Analyze the image frames of the video stream, output the 2D bounding box and category of the target, and generate a pixel-level relative depth map corresponding to the image frame. The categories include suspended objects and personnel targets. Based on the camera's intrinsic and extrinsic parameters, the 2D pixel coordinates and depth values ​​of each target's center point are converted into 3D positions in the world coordinate system with the tower crane base as the origin, and structured data is output. The structured data includes the target's ID, category, 3D coordinates, and velocity.

[0010] In one embodiment: a dynamic trigger threshold adapted to the current operation risk level is generated based on the real-time operating parameters of the hoisting operation and the inherent rated parameters of the hoisting equipment.

[0011] In one embodiment, the lightweight condition monitoring step specifically includes: Based on the target's 3D position and velocity in the previous frame, the target position in the current frame is linearly extrapolated and state smoothed by extended Kalman filtering to predict structured data for a preset number of frames. Based on the structured data obtained from the prediction, the trigger conditions and state change identification are performed frame by frame for full trajectory prediction.

[0012] In one embodiment: if the lightweight state monitoring does not trigger full trajectory prediction, and the next frame is a frame that does not meet the prediction triggering conditions or a non-critical steady-state frame, then the consistency between the predicted structured data and the actual structured data is verified, and a determination is made on whether to trigger full trajectory prediction based on the verification results.

[0013] In one embodiment: when the lightweight state monitoring triggers full trajectory prediction, the structured data of the current frame is used for full trajectory prediction; If the risk score after the current frame prediction is lower than the preset warning threshold, perform full trajectory prediction on the prediction frame that triggers full trajectory prediction. If the risk score predicted in the current frame is higher than or equal to the preset warning threshold, the consistency between the subsequent actual frames and the predicted frames is verified. If the consistency requirements are met, a control signal is output based on the risk score.

[0014] In one embodiment: if the lightweight state monitoring triggers full trajectory prediction, and the risk scores of the current frame and the predicted frame are both lower than the preset warning threshold, the consistency of the actual frame and the predicted frame is verified. If the consistency requirement is not met, full trajectory prediction is triggered.

[0015] In one embodiment, the training method for the pre-trained neural network model includes: Structured data and real-time wind speed of the suspended target are collected. The actual position of the suspended target and the position deviation of the trajectory in pure physical theory are used as labels to construct a dataset and divide it into training set and test set. A neural network model was built using a long short-term memory network. The input of the neural network model was defined as the structured data of the suspended target and the real-time wind speed, and the output was the position deviation of the suspended target at a future time. Construct a composite loss function with hoisting physical constraints, calculate the loss value with the fit between the position deviation of the model output and the position deviation of the label as the core, and perform phased iterative training on the model based on the training set until the loss value converges to the preset threshold. The accuracy of the trained model is verified using a test set. Once the verification is successful, the pre-trained neural network model is obtained.

[0016] In summary, this application has the following beneficial effects: 1. By using pixel-level depth maps, 2D coordinates are transformed into 3D coordinates. At the same time, depth information is used to distinguish between occlusion and overlap, reducing the false alarm rate by about 85% and the false negative rate by about 90%. 2. By combining lightweight state monitoring with full trajectory prediction, full trajectory prediction is performed only on key frames, which significantly reduces the computing load and shortens the time edge to reduce the occurrence of prediction lag. 3. By predicting the motion trajectory of the suspended target through full trajectory prediction, the inertial risks caused by the swing of the suspended object can be effectively addressed, and the problem of lag in risk identification can be solved. Attached Figure Description

[0017] Figure 1 This is a flowchart of a dynamic risk identification method for hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction in this embodiment; Figure 2 This is a flowchart of panoramic spatiotemporal perception in a dynamic risk identification method for hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction in this embodiment; Figure 3 This is a flowchart illustrating the prediction process of the final predicted trajectory in a dynamic risk identification method for hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction, as described in this embodiment. Detailed Implementation

[0018] The present application will be further described in detail below with reference to the accompanying drawings.

[0019] like Figure 1 As shown, this application discloses a method for dynamic risk identification in hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction, including the following steps: S100: Real-time acquisition of video stream data of hoisting operation scene, and extraction of structured data of the 3D spatial state of hoisted target and personnel target in video stream frame by frame.

[0020] This step forms the perceptual foundation of the entire method. Its core is to reconstruct the 3D spatial information of the hoisting operation scene from the 2D video stream, and to detect and output the precise 3D world coordinates and motion status data of the hoisted objects and personnel in real time, achieving panoramic spatiotemporal perception. Figure 2 As shown, the specific implementation process is as follows: S101. Real-time acquisition of video stream data of hoisting operation scene, and image frames obtained after preprocessing the video stream.

[0021] A monocular high-definition camera deployed at the front end of the tower crane boom or the top of the cab collects video streams of the hoisting operation area in real time. The video stream parameters are required to be resolution ≥1080P and frame rate ≥30FPS. The video stream is transmitted to the edge computing node through a gigabit network port or fiber optic cable. In this embodiment, NVIDIA Jetson Orin series modules are used.

[0022] The edge computing node decodes the received video stream frame by frame and preprocesses the decoded RGB image frames, including scaling the image to the model input size of 640×640 and normalizing the pixel values ​​to eliminate the interference of uneven lighting and image noise on subsequent detection.

[0023] S102: Analyze the image frames of the video stream, output the 2D bounding box and category of the target, and generate a pixel-level relative depth map corresponding to the image frame.

[0024] For the preprocessed image frames, an improved Swin-YOLO-v9 deep learning detection network is used to perform object detection, outputting the 2D bounding boxes and categories of the objects in the image. The object categories include at least suspended objects and personnel, and can also be extended to related objects such as hooks and booms. Simultaneously, a lightweight depth estimation module is used to generate a pixel-level relative depth map corresponding to the input image size.

[0025] In this embodiment, the specific implementation architecture of the Swin-YOLO-v9 network is as follows: the network is divided into four modules: Input, Backbone, Neck, and Head. The Input module is responsible for the above-mentioned image preprocessing operations. While retaining the basic CBS module, the Backbone module introduces SwinTransformerBlock to replace the original CSP convolution module in the last two stages (C4, C5). SwinTransformer divides the image into several non-overlapping windows, calculates self-attention within the windows, and establishes cross-window connections by cyclically shifting the window positions. Its self-attention calculation formula is: In the formula, B represents the relative position bias, and d represents the feature dimension. The Neck module employs a weighted bidirectional feature pyramid network (BiFPN), introducing learnable weights wi to achieve bidirectional fusion of multi-scale features. The fusion formula is as follows: This structure enhances the ability to simultaneously detect small targets (such as hooks in the air) and large targets (such as storage yards on the ground). By using a top-down path and a bottom-up path, the high-level strong semantic information output by the Backbone is fully and bidirectionally fused with the low-level precise localization information, generating three information-rich fused feature maps, laying a solid foundation for the detection of targets of different sizes.

[0026] The Head module employs a dual-branch decoupled prediction structure. It uses a parallel CBS standard convolutional branch and a DSC depthwise separable convolutional lightweight branch to handle classification and regression tasks respectively, ultimately outputting the target's bounding box, confidence score, and category information. Using the DSC branch helps reduce the number of parameters and computational complexity of the head network while maintaining accuracy, achieving more efficient detection.

[0027] S103: Based on camera intrinsic and extrinsic parameters, convert the 2D pixel coordinates and depth values ​​of each target center point into 3D positions in the world coordinate system with the tower crane base as the origin, and output structured data.

[0028] By combining the pre-calibrated intrinsic and extrinsic parameters of the camera, the 2D pixel coordinates of each target center point, along with the corresponding depth value, are converted into a 3D position in the world coordinate system with the tower crane base as the origin. At the same time, the instantaneous three-dimensional resultant velocity of the target is obtained by calculating the 3D coordinate difference between consecutive frames, and finally, structured data is output.

[0029] In this embodiment, the structured data is standardized data in JSON format, which includes at least the ID, category, 3D coordinates (x, y, z), and velocity of each target. The 3D coordinates are in the world coordinate system with the tower crane base as the origin, and the velocity is calculated by the difference between the 3D coordinates of consecutive frames. The unit of 3D coordinates is meters (m), and the unit of velocity is m / s.

[0030] For the target object being lifted, the structured data is also synchronously linked to the actual lifting weight data of the object, which is synchronized in real time by the tower crane PLC system, providing data support for subsequent risk level adaptation.

[0031] S200. Based on structured data, calculate the spatiotemporal proximity parameters of the current spatial distance and future collision trend between the suspended object target and the personnel target.

[0032] This step is the core of the entire method's judgment benchmark, providing a unified quantitative basis for subsequent trigger condition determination. The spatiotemporal proximity between the suspended object and the person is quantified using a spatiotemporal proximity parameter, the specific formula of which is: In the formula, S is the spatiotemporal proximity parameter. The smaller the value of S, the closer the spatiotemporal distance between the suspended object and the person, and the higher the potential risk of a collision in the future.

[0033] D hor The horizontal projected straight-line distance (unit: m) between the suspended object and the personnel reflects the current actual spatial distance between them. The calculation formula is: The 3D coordinates of the suspended target are (x... o ,y o ,z o ), where z o The height of the suspended object above the ground is given by the 3D coordinates of the personnel target (x, y). p ,y p ,z p ).

[0034] D safe The basic safety distance for hoisting operations (unit: m) is set according to the "Safety Regulations for Lifting Machinery Part 1: General Rules" (GB6067.1), with a fixed benchmark value of 6m, which is the statutory minimum safety radius for hoisting operations.

[0035] K is the safety redundancy coefficient, with a value ranging from 1.2 to 2.0. It is preset in advance according to the risk level of the hoisting operation. The coefficient is 2.0 for heavy hoisting operations and high-altitude operations, and 1.2 for light-load conventional operations, in order to adapt to the safety redundancy requirements of different operation scenarios.

[0036] V rel v is the relative approach velocity (in m / s) between the suspended object and the person. When they are moving towards each other, v rel =v o +v p When the two move in opposite directions, v rel =|v o -v p |;If the relative motion between the two is a continuous moving away, v rel Set the value to 0 directly to avoid false predictions triggered when the target is far away from the target.

[0037] T pre The preset trajectory prediction lead time (in seconds) for the system is consistent with the prediction duration of the full trajectory prediction.

[0038] It should be noted that, for situations where there are multiple personnel targets in the same scenario, the spatiotemporal proximity S value between the suspended object and each personnel target is calculated separately, and the minimum value is taken as the final output of this step to ensure coverage of the personnel target with the highest risk.

[0039] S300: Obtain the dynamic trigger threshold of the current operation risk level and compare it with the spatiotemporal proximity parameter to determine whether the current frame meets the trigger conditions for full trajectory prediction.

[0040] In this step, based on the real-time operating parameters of the hoisting operation and the inherent rated parameters of the lifting equipment, a dynamic trigger threshold adapted to the current operation risk level is generated. The calculation formula is as follows: In the formula, S th For dynamic trigger threshold; z o This represents the current height of the suspended load above the ground, in meters (m). max This refers to the maximum rated lifting height of the tower crane, in meters (m). This is an inherent parameter of the tower crane equipment and should be entered into the system in advance. o The actual weight of the load being lifted, in tons, is real-time synchronized data from the tower crane's PLC system; m max The maximum rated lifting capacity of the tower crane is expressed in tons (t). These are inherent parameters of the tower crane equipment and should be entered into the system in advance.

[0041] The higher the height of the load off the ground and the greater the weight lifted, the higher the risk level of the lifting operation, and the corresponding S th The higher the value, the higher the threshold for triggering prediction, and the earlier the system will start full trajectory prediction, reserving a longer warning and risk avoidance time for high-risk operations.

[0042] For example, when the target object is at its maximum lifting height and maximum rated lifting capacity, Sth = 1 + 1 + 1 = 3. At this time, prediction will be triggered as long as S ≤ 3, which means that prediction will start in advance when the horizontal distance between the object and the personnel reaches 3 times the safe distance. When the object is on the ground and unloaded, Sth = 1 + 0 + 0 = 1. Prediction will only be triggered when S ≤ 1, that is, when the distance between the object and the personnel enters the basic safe distance, thus avoiding the waste of computing power in low-risk scenarios.

[0043] Specifically, if the current frame meets the prediction triggering conditions, it is taken as the first keyframe. The structured data corresponding to the keyframe is then used for a preset full trajectory prediction, and a risk score is obtained based on the prediction results. A preset warning threshold is compared with the risk score. If the risk score is higher than or equal to the preset warning threshold, a control signal is output based on the risk score. This control signal triggers warning lights and on-site voice prompts, reminding personnel to evacuate to a safe area. Furthermore, for precise control, an emergency threshold can be set. When the risk score is greater than or equal to the emergency threshold, warning lights and on-site voice prompts are triggered, and deceleration or emergency stop commands are sent to the tower crane PLC via dry contacts or Modbus protocol, while simultaneously recording a detailed risk event log.

[0044] If the risk score is lower than the preset warning threshold, the subsequent consecutive frames in the triggered state will be identified for state changes. Frames with state changes greater than the preset change threshold will be locked as new key frames and the full trajectory prediction will be retried. Frames with state changes less than the preset change threshold will be defined as non-key steady-state frames.

[0045] In this embodiment, state change recognition specifically refers to: recording the structured data corresponding to the first keyframe as the comparison benchmark for subsequent frames; for subsequent consecutive frames in the triggered state, calculating the target state change rate, which is obtained by weighted calculation of three core dimensions, namely the change rate of the distance between the suspended object and the personnel, the change rate of the speed of the suspended object, and the change rate of the movement direction of the suspended object. Preferably, the weight of the change rate of the distance between the suspended object and the personnel is greater than the weight of the change rate of the speed of the suspended object, which is greater than the weight of the change rate of the movement direction of the suspended object, and the sum of the weights of the three is 1.

[0046] With this setup, when the relative motion state between the suspended object and the personnel does not change significantly, the predicted trajectory output by the previous keyframe remains valid, eliminating the need to repeat the high-computing-power full prediction. Only when the motion state undergoes a sudden change will the predicted trajectory be updated through a new keyframe, thereby significantly reducing the computing power load in the triggered state while ensuring prediction accuracy.

[0047] In this embodiment, the steps for full trajectory prediction specifically include: S310. Generate historical trajectories based on the historical structured data corresponding to the suspended target, and obtain the final predicted trajectory by combining physical flow deduction and neural network residual correction.

[0048] like Figure 3 As shown, the specific steps for obtaining the final predicted trajectory include: S311. Abstract the suspended object into a simple pendulum system. Based on the current structured data and historical trajectory of the suspended object, deduce the pure physical theoretical trajectory of the simple pendulum system at future moments.

[0049] The historical motion trajectory is constructed by retrieving the historical sequence of the 3D positions of the suspended target corresponding to the past 30 consecutive image frames. The motion law of the pendulum system is described by differential equations. In the formula, L is the rope length, θ is the swing angle, g is the gravitational acceleration, and c is the damping coefficient. At each time step t, the system solves the differential equation using the fourth-order Runge-Kutta (RK4) numerical integration method, iteratively calculating the theoretical position sequence for future times, and obtaining the purely physical predicted position P. phy (t+1). This physical model ensures that the prediction results conform to the basic physical laws, providing a stable benchmark framework for the overall prediction.

[0050] S312. Based on a pre-trained neural network model, output the position residual for future moments using the current structured data and historical trajectory of the suspended target.

[0051] In this step, the training methods for the pre-trained neural network model include: Structured data and real-time wind speed of the suspended target are collected. The actual position of the suspended target and the position deviation of the trajectory in pure physical theory are used as labels to construct a dataset and divide it into training set and test set. The real-time wind speed data can be obtained by devices such as ultrasonic anemometers and wind vanes, cup anemometers, etc. A neural network model is built using a long short-term memory network. The input of the neural network model is defined as the structured data of the suspended target and the real-time wind speed, and the output is the position deviation of the suspended target at a future time. Construct a composite loss function with hoisting physical constraints, calculate the loss value with the fit between the position deviation of the model output and the position deviation of the label as the core, and perform phased iterative training on the model based on the training set until the loss value converges to the preset threshold. The accuracy of the trained model is verified using a test set. Once the verification is successful, the pre-trained neural network model is obtained.

[0052] During training, a composite loss function is used to optimize the network. The first term represents the fitting error between the predicted and actual values, ensuring prediction accuracy. The second term is a regularization term, used to constrain the neural network to learn only the "residual" part, preventing the network from overfitting and deviating from physical laws. α and β are weighting coefficients.

[0053] Through this physical and data collaborative architecture, PI-LSTM can ensure that the prediction results conform to physical laws and flexibly adapt to complex and ever-changing field environments, significantly improving the accuracy and robustness of trajectory prediction.

[0054] S313. The physical theoretical trajectory is added to and fused with the position residual predicted by the neural network to obtain the final predicted trajectory.

[0055] S320. Define a two-dimensional Gaussian distribution for each point on the final predicted trajectory and construct a dynamic Gaussian danger potential field.

[0056] Among them, the covariance matrix of the Gaussian distribution deforms in real time with the velocity vector of the suspended object, and is stretched along the direction of the object's movement. The potential field strength is proportional to the potential energy of the suspended object. The higher the height and the greater the weight of the suspended object, the greater the coverage and strength of the potential field.

[0057] S330. Use Kalman filtering or linear extrapolation to predict the short-term simple trajectory of personnel, calculate the spatiotemporal overlap integral of the personnel trajectory and the dynamic potential field, and obtain a quantitative risk score.

[0058] S400 performs preset lightweight state monitoring on frames that do not meet the prediction trigger conditions and non-critical steady-state frames, and determines whether to trigger full trajectory prediction based on the monitoring results.

[0059] In one embodiment, the lightweight state monitoring steps specifically include: Based on the target's 3D position and velocity in the previous frame, the target position in the current frame is linearly extrapolated and state smoothed by extended Kalman filtering to predict structured data for a preset number of frames. Based on the structured data obtained from the prediction, the trigger conditions and state changes for full trajectory prediction are identified frame by frame. When either identification is valid, full trajectory prediction is triggered, i.e., when the spatiotemporal proximity is greater than the dynamic trigger threshold or the state change is greater than the preset change threshold.

[0060] If lightweight state monitoring does not trigger full trajectory prediction, and the next frame is a frame that does not meet the prediction trigger conditions or a non-critical steady-state frame, then the consistency between the predicted structured data and the actual structured data is verified, and a determination is made on whether to trigger full trajectory prediction based on the verification results.

[0061] The logic of consistency verification is as follows: calculate the positional deviation between the predicted structured data and the actual structured data. When the positional deviation exceeds the preset deviation threshold, it is determined that the consistency does not meet the requirements, and the current frame is immediately locked as the key frame to trigger full trajectory prediction. When the positional deviation is less than or equal to the deviation threshold, it is determined that the consistency meets the requirements, and the lightweight state monitoring process continues.

[0062] Specifically, when lightweight status monitoring triggers full trajectory prediction, the structured data of the current frame is used for full trajectory prediction. If the risk score after the current frame prediction is lower than a preset warning threshold, full trajectory prediction is performed on the predicted frame that triggered the full trajectory prediction, achieving two-frame cross-validation. If the risk score after the current frame prediction is higher than or equal to the preset warning threshold, the consistency of subsequent actual frames and predicted frames is verified. If the consistency requirements are met, a control signal is output based on the risk score.

[0063] In addition, if lightweight status monitoring triggers full trajectory prediction, and the risk scores of both the current frame and the predicted frame are lower than the preset warning threshold, the consistency between the actual frame and the predicted frame is verified. If the consistency requirements are not met, full trajectory prediction is triggered.

[0064] Example 1: At a construction site of a commercial center, a tower crane is lifting a 6-meter-long, 2-ton bundle of steel bars. The tower crane has a maximum rated lifting height of 80 meters, a maximum rated lifting capacity of 10 tons, and a safety redundancy factor k of 2.0. At this time, a worker is bending down to organize tools 5 meters ahead of the load's path. The execution flow of this method is as follows: Cameras deployed at the front end of the tower crane boom collect real-time video streams of the operation scene. The "steel bar" target and the "worker" personnel target are identified through the Swin-YOLO-v9 network. Monocular depth estimation calculates that the height of the steel bar above the ground is 8m, and the worker is located on the ground. The 3D coordinates and velocity structured data of both are output. The calculated horizontal projected distance between the suspended object and the person is Dhor=5m, the relative approach speed is vrel=1.5m / s, and the spatiotemporal proximity is S=5 / (6+2.0×1.5×2)=0.357; Based on the suspended object height of 8m and weight of 2t, the dynamic trigger threshold Sth is calculated as 1 + 8 / 80 + 2 / 10 = 1.3. It is determined that S≤Sth, which satisfies the trigger condition for full trajectory prediction, and the current frame is locked as the first key frame. Based on the historical trajectory of the suspended object in 30 frames, the physical trajectory is deduced through a pendulum dynamics model. Combined with the wind residual correction output by the LSTM network, the final predicted trajectory for the next 2 seconds is obtained. It is predicted that the steel bar will reach the area directly above the worker's head due to inertial swing after 2 seconds. A dynamic Gaussian hazard potential field extending forward with the speed of the hoisted object is constructed. The spatiotemporal overlap integral of the personnel trajectory and the potential field is calculated, resulting in a risk score of 85. The preset warning threshold is 30 and the emergency threshold is 70, so the emergency threshold is exceeded. A high-level output from the edge node's GPIO triggers a flashing red light and a voice alarm. Simultaneously, a slewing braking command is sent to the tower crane's PLC via Modbus, causing the tower crane's boom to decelerate and the steel bar's swing to be controlled. Upon hearing the alarm, workers quickly run away to avoid collisions.

[0065] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A method for dynamic risk identification in hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction, characterized in that, include: Real-time acquisition of video stream data of hoisting operation scene, and extraction of structured data of 3D spatial state of hoisted target and personnel target in video stream frame by frame; Based on structured data, the spatiotemporal proximity parameters of the current spatial distance and future collision trend between the suspended object target and the personnel target are calculated. The dynamic trigger threshold of the current operation risk level is obtained and compared with the spatiotemporal proximity parameter to determine whether the current frame meets the trigger conditions for full trajectory prediction. If the current frame meets the prediction triggering conditions, the current frame will be used as the first key frame, and the structured data corresponding to the key frame will be used for the preset full trajectory prediction. A risk score will be obtained based on the prediction results. If the risk score is higher than or equal to the preset warning threshold, a control signal is output based on the risk score; if the risk score is lower than the preset warning threshold, the state change is identified for subsequent consecutive frames in the triggered state, and the frames with state changes greater than the preset change threshold are locked as new key frames, and the full trajectory prediction is retried. Frames whose state changes are less than a preset change threshold are defined as non-critical steady-state frames. For frames that do not meet the prediction trigger conditions and non-critical steady-state frames, perform preset lightweight state monitoring, and determine whether to trigger full trajectory prediction based on the monitoring results.

2. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 1, characterized in that, The steps for full trajectory prediction specifically include: Historical trajectories are generated based on the historical structured data corresponding to the suspended target, and the final predicted trajectory is obtained by combining physical flow deduction and neural network residual correction. A two-dimensional Gaussian distribution is defined for each point on the final predicted trajectory. Its covariance matrix deforms in real time with the velocity vector of the suspended object, and the potential field strength is proportional to the potential energy of the suspended object. Kalman filtering or linear extrapolation is used to predict the short-term simple trajectory of personnel, and the spatiotemporal overlap integral between the personnel trajectory and the dynamic potential field is calculated to obtain a quantitative risk score.

3. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 2, characterized in that, The steps for obtaining the final predicted trajectory specifically include: The suspended object is abstracted as a simple pendulum system. Based on the current structured data and historical trajectory of the suspended object, the pure physical theoretical trajectory of the pendulum system at future moments is deduced. Based on a pre-trained neural network model, the position residual for future moments is output using the current structured data and historical trajectory of the suspended target. The final predicted trajectory is obtained by fusing the physical theoretical trajectory with the position residual predicted by the neural network.

4. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 1, characterized in that, The steps for obtaining the structured data include: Analyze the image frames of the video stream, output the 2D bounding box and category of the target, and generate a pixel-level relative depth map corresponding to the image frame. The categories include suspended objects and personnel targets. Based on the camera's intrinsic and extrinsic parameters, the 2D pixel coordinates and depth values ​​of each target's center point are converted into 3D positions in the world coordinate system with the tower crane base as the origin, and structured data is output. The structured data includes the target's ID, category, 3D coordinates, and velocity.

5. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 1, characterized in that: Based on the real-time operating parameters of the hoisting operation and the inherent rated parameters of the hoisting equipment, a dynamic trigger threshold adapted to the current operation risk level is generated.

6. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 1, characterized in that, The lightweight condition monitoring steps specifically include: Based on the target's 3D position and velocity in the previous frame, the target position in the current frame is linearly extrapolated and state smoothed by extended Kalman filtering to predict structured data for a preset number of frames. Based on the structured data obtained from the prediction, the trigger conditions and state change identification are performed frame by frame for full trajectory prediction.

7. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 6, characterized in that: If the lightweight state monitoring does not trigger full trajectory prediction, and the next frame is a frame that does not meet the prediction triggering conditions or a non-critical steady-state frame, then the consistency between the predicted structured data and the actual structured data is verified, and a determination is made on whether to trigger full trajectory prediction based on the verification results.

8. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 7, characterized in that: When the lightweight state monitoring triggers full trajectory prediction, the structured data of the current frame is used to perform full trajectory prediction. If the risk score after the current frame prediction is lower than the preset warning threshold, perform full trajectory prediction on the prediction frame that triggers full trajectory prediction. If the risk score predicted in the current frame is higher than or equal to the preset warning threshold, the consistency between the subsequent actual frames and the predicted frames is verified. If the consistency requirements are met, a control signal is output based on the risk score.

9. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 8, characterized in that: If the lightweight status monitoring triggers full trajectory prediction, and the risk scores of both the current frame and the predicted frame are lower than the preset warning threshold, the consistency between the actual frame and the predicted frame is verified. If the consistency requirement is not met, full trajectory prediction is triggered.

10. The method for dynamic risk identification of hoisting operations based on spatiotemporal trajectory prediction and 3D reconstruction according to claim 1, characterized in that, The training method for the pre-trained neural network model includes: Structured data and real-time wind speed of the suspended target are collected. The actual position of the suspended target and the position deviation of the trajectory in pure physical theory are used as labels to construct a dataset and divide it into training set and test set. A neural network model was built using a long short-term memory network. The input of the neural network model was defined as the structured data of the suspended target and the real-time wind speed, and the output was the position deviation of the suspended target at a future time. Construct a composite loss function with hoisting physical constraints, calculate the loss value with the fit between the position deviation of the model output and the position deviation of the label as the core, and perform phased iterative training on the model based on the training set until the loss value converges to the preset threshold. The accuracy of the trained model is verified using a test set. Once the verification is successful, the pre-trained neural network model is obtained.