Obstacle positioning-based crane obstacle avoidance early warning method and system

CN122809331APending Publication Date: 2026-09-25HUAIBEI SPECIAL EQUIP SUPERVISION & INSPECTION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610990211.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

针对现有技术存在的在动态变形与自遮挡工况下,难以兼顾自身结构,精确滤除与盲区障碍有效感知的问题,本申请通过基于障碍定位的起重机避障预警方法及系统,实现对起重机作业环境中显性及隐性障碍物的全方位精准感知与预警

Benefits of technology

[0014]本申请提供基于障碍定位的起重机避障预警方法及系统,能够通过实时采集基座倾斜角与吊载重量等参数,对数字几何模型进行倾斜补偿与挠曲变形修正,使得自身点云滤除基准能够随车体姿态与臂架形变动态更新,消除因坐标系偏差与刚体假设导致的滤除误差,降低了误报警率与漏检率。同时,利用运动学状态参数作为先验条件引导点云补全网络,使网络能够依据起重机特有的运动约束推理出自遮挡盲区内的障碍物分布,有效恢复了被自身结构遮挡的环境信息,提升感知系统的空间完整性。此外,通过实时计算尾部摆动、副臂外摆、最大工作半径及净空高度等隐性障碍动态包络,将传感器无法直接探测的物理尺寸限制转化为虚拟障碍物纳入决策,填补了传统感知方案的固有盲区,预防了隐性碰撞风险。最终,通过融合候选环境点云、补全环境点云与隐性障碍动态包络,并结合运动预测生成全局障碍物感知结果,实现了从单一传感器感知向多源知识融合感知的质变,在保证毫秒级动态响应的同时,提升起重机在倾斜地面、重载变形、自遮挡等恶劣工况下的避障预警精度与可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122809331A_ABST
    Figure CN122809331A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent perception and safety control of engineering machinery, and provides a crane obstacle avoidance early warning method and system based on obstacle positioning, comprising: acquiring base attitude parameters, body stress deformation parameters and environmental three-dimensional point cloud data of the working machine; correcting a digital geometric model based on the base attitude parameters and the body stress deformation parameters; filtering out self-structure point clouds from the environmental three-dimensional point cloud data based on the corrected digital geometric model to obtain candidate environmental point clouds; identifying self-occlusion hollow regions and inputting kinematic state parameters as prior conditions into a point cloud completion network to obtain completed environmental point clouds; generating a dynamic envelope of hidden obstacles; and fusing each point cloud and the envelope to generate a global obstacle perception result. The present application solves the problem of perception failure under dynamic deformation and self-occlusion working conditions, and improves the integrity and reliability of obstacle avoidance early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent perception and safety control of construction machinery, and in particular to a method and system for crane obstacle avoidance and early warning based on obstacle location. Background Technology

[0002] Currently, in scenarios such as building construction, port loading and unloading, and heavy manufacturing, the intelligent obstacle avoidance and perception technology of mobile cranes mainly relies on lidar point cloud processing or self-point cloud filtering schemes based on known geometric models. The former identifies obstacles by collecting 3D point clouds of the environment and performing clustering and segmentation, while the latter uses offline CAD models combined with real-time attitude sensors to calculate the theoretical position of its own structure and remove it from the lidar point cloud to retain external environmental information.

[0003] However, during actual operation, cranes frequently experience complex conditions such as body tilting, boom elastic deformation, and large-amplitude rotation. When the crane body tilts due to uneven ground, the fixed rigid body geometric model cannot compensate for coordinate system deviations, causing the point cloud filtering window to shift. This can lead to false alarms due to residual point clouds or incorrect filtering of real obstacles. When the boom rotates at a large elevation angle or is used for heavy lifting, the boom's structure itself can severely obstruct the radar field of view. Existing technologies only perform removal processing and cannot fill in the real obstacle information in the blind spot, resulting in a very high risk of missed detection. At the same time, the implicit collision areas formed by the boom tail swing and the outward swing of the jib, constrained by physical dimensions, become perception blind spots because they are beyond the direct detection range of the sensors. In summary, existing technologies cannot simultaneously take into account the crane's own structure, accurately filter obstacles, and effectively perceive obstacles in blind spots under dynamic and complex conditions. This results in a high false alarm rate and safety blind spots in the obstacle avoidance system, failing to meet the requirements for high-reliability operation. Summary of the Invention To address the problem that existing technologies struggle to accurately filter out and effectively perceive blind-spot obstacles while maintaining their own structure under dynamic deformation and self-shading conditions, this application proposes a crane obstacle avoidance and early warning method and system based on obstacle localization, which enables comprehensive and accurate perception and early warning of both visible and hidden obstacles in the crane's operating environment.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a crane obstacle avoidance and early warning method based on obstacle localization, comprising: acquiring the base attitude parameters, body stress deformation parameters, and environmental three-dimensional point cloud data of the operating machinery; correcting the digital geometric model based on the base attitude parameters and body stress deformation parameters; filtering out the self-structure point cloud from the environmental three-dimensional point cloud data based on the corrected digital geometric model to obtain candidate environmental point clouds; identifying self-occluding void regions in the candidate environmental point clouds, using the kinematic state parameters of the operating machinery as prior conditions, inputting them into a point cloud completion network to complete the point cloud, and obtaining a completed environmental point cloud; generating a hidden obstacle dynamic envelope based on the real-time attitude parameters and physical size constraints of the operating machinery; and fusing the candidate environmental point cloud, the completed environmental point cloud, and the hidden obstacle dynamic envelope to generate a global obstacle perception result.

[0005] In conjunction with the first aspect mentioned above, in one possible implementation, the process of correcting the digital geometric model based on the base attitude parameters and the body stress deformation parameters includes: calling the base attitude parameters and the body stress deformation parameters, wherein the base attitude parameters include the base tilt angle and the body stress deformation parameters include the load weight; based on the base tilt angle, performing real-time correction on the coordinate transformation matrix between the sensor coordinate system and the base coordinate system of the acquired 3D point cloud data to obtain a tilt-compensated coordinate transformation relationship; based on the load weight, estimating the body deflection deformation of the working machinery under the load state, and offsetting and correcting the geometric position of the corresponding deformable component in the digital geometric model based on the body deflection deformation; and generating the corrected digital geometric model based on the tilt-compensated coordinate transformation relationship and the offset-corrected geometric position.

[0006] In conjunction with the first aspect mentioned above, in one possible implementation, the process of estimating the amount of body deflection deformation of the operating machinery under the load state based on the load weight, and offsetting and correcting the geometric position of the corresponding deformable component in the digital geometric model based on the amount of body deflection deformation includes: obtaining the telescopic length and luffing angle of the operating machinery; calculating the spatial offset of each section of the boom of the operating machinery based on beam deformation theory or a finite element reduced-order model, combined with the load weight, telescopic length, and luffing angle; and superimposing the spatial offset onto the corresponding boom section node in the digital geometric model to update the three-dimensional spatial pose of the boom.

[0007] In conjunction with the first aspect mentioned above, in one possible implementation, the kinematic state parameters of the operating machinery are used as prior conditions and input into a point cloud completion network for point cloud completion. This process includes: acquiring the kinematic state parameters of the operating machinery, which include body posture parameters and working condition feature parameters. The body posture parameters include at least rotation angle, amplitude angle, and telescopic length, while the working condition feature parameters include at least base tilt angle and load weight; identifying self-occluding void regions in the candidate environment point cloud and extracting sparse point cloud features from preset regions adjacent to the self-occluding void regions; inputting the body posture parameters and working condition feature parameters into the kinematic coding branch of the point cloud completion network to generate a kinematic prior feature vector; inputting the sparse point cloud features and the kinematic prior feature vector into a cross-modal attention fusion module, embedding the kinematic prior feature vector into the point cloud features through a cross-attention mechanism to obtain fused features; and inputting the fused features into the decoder of the point cloud completion network to generate a completed point cloud for the self-occluding void regions, which is then merged with the candidate environment point cloud to obtain a completed environment point cloud.

[0008] In conjunction with the first aspect mentioned above, in one possible implementation, the input / output format and confidence filtering process of the point cloud completion network include: converting the spatial location of the self-occluded hole region into a polar coordinate grid index, dividing the polar coordinate grid index into N grid cells according to horizontal angular intervals and radial distance intervals, with each grid cell uniquely corresponding to a spatial sector, and N being a preset value; aligning the sparse point cloud features and kinematic prior feature vectors according to the polar coordinate grid index and inputting them into the cross-modal attention fusion module; the decoder of the point cloud completion network outputs the predicted point cloud block corresponding to each grid cell and generates a confidence value for each predicted point, where the confidence value represents the probability that the predicted point belongs to a real obstacle; discarding predicted points with confidence values ​​lower than a preset threshold, retaining predicted points with confidence values ​​not lower than the preset threshold, and merging them with the candidate environment point cloud to obtain the completed environment point cloud.

[0009] In conjunction with the first aspect mentioned above, in one possible implementation, the process of generating a dynamic envelope of hidden obstacles based on the real-time attitude parameters and physical size constraints of the operating machinery includes: acquiring the real-time attitude parameters and physical size constraints of the operating machinery, wherein the real-time attitude parameters include at least the slewing angle, luffing angle, and telescopic length, and the physical size constraints include at least the geometry of the boom tail, the geometry of the jib, and the load-amplitude characteristic curve; calculating the tail swing envelope based on the slewing angle and the geometry of the boom tail, wherein the tail swing envelope is the boundary of the annular area swept by the boom tail during slewing; calculating the jib outswing envelope based on the luffing angle, telescopic length, and the geometry of the jib, wherein the jib outswing envelope is the boundary of the horizontal projection trajectory of the outer end point of the jib in the deployed state; querying the load-amplitude characteristic curve based on the telescopic length and the load weight of the operating machinery to obtain the safe working radius, and generating a maximum working radius limiting envelope with the slewing center as the center and the safe working radius as the radius; and outputting at least one of the tail swing envelope, the jib outswing envelope, and the maximum working radius limiting envelope as the dynamic envelope of hidden obstacles.

[0010] In conjunction with the first aspect mentioned above, in one possible implementation, the dynamic envelope of the hidden obstacle also includes a clearance height envelope, wherein: the bending deformation of the working machinery under the load state is obtained, and the bending deformation of the working machinery characterizes the degree of elastic deflection of the boom under the load; the luffing angle and extension length of the working machinery are obtained, and the spatial coordinates of the lowest point of the boom after elastic deformation are calculated in combination with the bending deformation of the working machinery; based on the spatial coordinates of the lowest point, the clearance height boundary below the boom is generated as the clearance height envelope.

[0011] In conjunction with the first aspect mentioned above, in one possible implementation, the process of fusing candidate environmental point clouds, complete environmental point clouds, and dynamic envelopes of hidden obstacles to generate a global obstacle perception result includes: acquiring the current motion speed parameters and motion acceleration parameters of the operating machinery, wherein the motion speed parameters include at least rotational angular velocity, amplitude angular velocity, and extension speed, and the motion acceleration parameters include at least rotational angular acceleration, amplitude angular acceleration, and extension acceleration; aligning the candidate environmental point clouds, complete environmental point clouds, and dynamic envelopes of hidden obstacles in spatial coordinates and merging them into an occupied grid map under the same coordinate system to generate a global obstacle map; predicting the motion trajectory of the operating machinery within a preset time interval in the future based on the motion speed parameters and motion acceleration parameters, and in conjunction with the global obstacle map, and calculating the minimum collision time with each obstacle in the global obstacle map along the trajectory; determining the current warning level based on the comparison results between the minimum collision time and multiple preset warning thresholds, and outputting the corresponding warning signal or control command.

[0012] In conjunction with the first aspect mentioned above, in one possible implementation, the process of filtering out the self-structure point cloud from the environmental 3D point cloud data based on the modified digital geometric model includes: acquiring the real-time swing amplitude parameter of the working machine boom, which is used to characterize the instantaneous offset of the actual spatial position of the boom relative to the theoretical motion trajectory caused by inertia, wind load, or hydraulic impact during the movement; generating the theoretical self-point cloud set at the current moment based on the modified digital geometric model; setting an initial matching distance threshold, performing nearest neighbor matching between each point in the environmental 3D point cloud data and the theoretical self-point cloud set, and calculating the minimum Euclidean distance; when the real-time swing amplitude parameter exceeds the preset swing threshold, widening the matching distance threshold from the initial value to a widened value, which is greater than the initial value; determining points with a minimum Euclidean distance less than the current matching distance threshold as self-structure point clouds and removing them from the environmental 3D point cloud data, with the remaining point clouds being output as candidate environmental point clouds.

[0013] Secondly, an obstacle avoidance and early warning system for cranes based on obstacle localization is provided, comprising: a sensor module, including an inertial measurement unit, an attitude sensor, and a load sensor, for acquiring the base attitude parameters, body stress deformation parameters, and kinematic state parameters of the multi-degree-of-freedom operating machinery; a point cloud acquisition module for acquiring environmental 3D point cloud data; a memory for storing program instructions; and a processor configured to execute the program instructions stored in the memory to correct the digital geometric model based on the base attitude parameters and body stress deformation parameters; based on the corrected digital geometric model, filtering out the self-structure point cloud from the environmental 3D point cloud data to obtain candidate environmental point clouds; identifying self-occluding void regions in the candidate environmental point clouds, using the kinematic state parameters of the operating machinery as prior conditions, inputting them into a point cloud completion network for point cloud completion to obtain a completed environmental point cloud; generating a dynamic envelope of hidden obstacles based on the real-time attitude parameters and physical size constraints of the operating machinery; and fusing the candidate environmental point clouds, the completed environmental point clouds, and the dynamic envelope of hidden obstacles to generate a global obstacle perception result.

[0014] This application provides a crane obstacle avoidance and early warning method and system based on obstacle localization. It can compensate for tilt and correct flexural deformation of the digital geometric model by real-time acquisition of parameters such as base tilt angle and load weight. This allows the point cloud filtering benchmark to be dynamically updated with the vehicle body attitude and boom deformation, eliminating filtering errors caused by coordinate system deviation and rigid body assumptions, thus reducing false alarm and missed detection rates. Simultaneously, kinematic state parameters are used as prior conditions to guide the point cloud completion network, enabling the network to infer the obstacle distribution within the self-occluding blind zone based on the crane's unique motion constraints. This effectively recovers environmental information obscured by the crane's own structure, improving the spatial integrity of the perception system. Furthermore, by real-time calculation of the dynamic envelope of hidden obstacles such as tail swing, boom outward swing, maximum working radius, and clearance height, physical size limitations that sensors cannot directly detect are transformed into virtual obstacles and incorporated into the decision-making process. This fills the inherent blind spots of traditional perception schemes and prevents hidden collision risks. Ultimately, by fusing candidate environmental point clouds, completing environmental point clouds, and dynamic envelopes of hidden obstacles, and combining motion prediction to generate global obstacle perception results, a qualitative leap from single-sensor perception to multi-source knowledge fusion perception was achieved. While ensuring millisecond-level dynamic response, the accuracy and reliability of obstacle avoidance and early warning for cranes under harsh working conditions such as inclined ground, heavy load deformation, and self-occlusion were improved. Attached Figure Description

[0015] Figure 1 A flowchart illustrating the obstacle avoidance and early warning method for cranes based on obstacle localization provided in this application embodiment; Figure 2 Detailed flowchart of digital geometric model correction in the obstacle avoidance and early warning method for cranes based on obstacle localization provided in the embodiments of this application; Figure 3 The flowchart of the self-structure point cloud adaptive filtering process in the obstacle avoidance and early warning method for cranes based on obstacle localization provided in the embodiments of this application; Figure 4 The flowchart of the kinematic prior-guided point cloud completion network processing in the obstacle avoidance and early warning method for cranes based on obstacle localization provided in the embodiments of this application is as follows: Figure 5 The flowchart for generating the dynamic envelope of hidden obstacles in the crane obstacle avoidance and early warning method based on obstacle localization provided in the embodiments of this application is as follows: Figure 6 The flowchart of global obstacle perception and hierarchical early warning in the crane obstacle avoidance and early warning method based on obstacle localization provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of a crane obstacle avoidance and early warning system based on obstacle location provided in an embodiment of this application. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0018] Example 1: like Figure 1 As shown, this embodiment provides a crane obstacle avoidance and early warning method based on obstacle localization. This method constructs a closed-loop processing flow from multi-dimensional state perception to global risk decision-making, aiming to solve the technical challenge of coexisting structural interference and perception blind spots in complex working conditions. The method of this embodiment mainly includes the following steps: Step 100: Obtain the base attitude parameters, body stress deformation parameters, and environmental 3D point cloud data of the operating machinery.

[0019] Among them, the base attitude parameter refers to the physical quantity that characterizes the spatial tilt state of the crane chassis or turntable relative to the horizontal reference plane of the earth.

[0020] The deformation parameters of the crane body under stress refer to the mechanical state quantities that reflect the degree of elastic deformation of the crane's metal structure under external loads.

[0021] Environmental 3D point cloud data is a set of discrete points collected by active sensors such as LiDAR, which contains spatial geometric information of the work scene.

[0022] In some implementations, the roll and pitch angles of the base are collected in real time by an inertial measurement unit as base attitude parameters. The load weight is obtained by a pin force sensor or hydraulic pressure sensor as a core indicator of the body's stress deformation parameters. At the same time, a lidar installed on the turntable or boom root is used to scan the surrounding environment at a preset frequency to generate a three-dimensional point cloud.

[0023] To ensure the spatiotemporal consistency of multi-source data, hardware-level or software-level timestamp alignment and coordinate system preprocessing are performed on the data acquisition stage to ensure that subsequent processing is based on a complete state snapshot at the same time and in the same coordinate system.

[0024] Step 200: Based on the base attitude parameters and the body stress deformation parameters, the digital geometric model is corrected.

[0025] Among them, digital geometric model correction is the process of dynamically adjusting the static CAD model of the crane to a virtual mirror that is highly consistent with the real physical entity according to the real working conditions. Its essence is to compensate for the geometric position deviation caused by the tilt of the crane body and the elastic deformation of the structure through mathematical transformation, so that the virtual model can accurately reproduce the actual spatial position of the crane at the current moment.

[0026] In some implementations, the processor first reads the base attitude parameters and calculates the real-time rotation and translation matrix between the sensor coordinate system and the mechanical base coordinate system to eliminate the installation reference offset caused by uneven ground. Then, based on the body's stress deformation parameters (such as lifting weight, boom length, and luffing angle), it uses a preset mechanical model to estimate the deflection deformation of each boom segment and superimposes this deformation onto the corresponding nodes of the digital geometric model, thereby generating a double-corrected digital geometric model that reflects the true pose.

[0027] The correction process is carried out in real time. As the crane's posture and load change, the digital geometric model is continuously and dynamically updated, always keeping in sync with the physical entity. For example, when the crane lifts a heavy object on a slope, the model will not only rotate with the tilt of the crane body, but will also bend due to the pull of the heavy object. The corrected model edges will fit closely to the surface of the actual boom.

[0028] It should be noted that through real-time correction, the system can accurately distinguish which point clouds belong to its own structure and which belong to the external environment. Even under extreme conditions such as vehicle tilt or heavy load deformation, it can ensure the accuracy of its own point cloud filtering, effectively preventing false alarms or missed detections caused by model mismatch and improving safety decision-making.

[0029] Step 300: Based on the corrected digital geometric model, filter out the self-structure point cloud from the environmental 3D point cloud data to obtain candidate environmental point clouds.

[0030] Among them, self-structure point cloud filtering refers to the process of using a modified digital geometric model as a reference mask in three-dimensional space to identify and remove points belonging to the crane body from the original environmental point cloud, thereby separating a subset of point clouds that purely represent external environmental obstacles.

[0031] Candidate environment point clouds are the input data that are retained after their own structure has been cleaned and are used for subsequent obstacle analysis and completion.

[0032] In some implementations, the modified digital geometric model is projected onto the same field of view and resolution as the lidar to generate a theoretical self-point cloud set at the current moment. Then, each point in the original environmental 3D point cloud data is traversed, and its Euclidean distance to the nearest neighbor in the theoretical self-point cloud set is calculated. If the distance is less than a preset matching threshold, the point is determined to be a self-structured point cloud and is discarded; otherwise, it is retained as a candidate environmental point cloud.

[0033] It should be understood that the matching threshold can be adaptively adjusted based on the sensor noise level and the model accuracy to balance the thoroughness of filtering with the protection of environmental points.

[0034] It should be noted that by introducing a dynamically modified model in real time as the filtering benchmark, the accuracy and robustness of self-point cloud removal are improved. Compared with traditional methods that use fixed models or simple geometric rules, this solution can effectively deal with the drift of self-point cloud positions caused by dynamic factors such as vehicle tilt and boom deformation, avoiding the situation where residual self-point clouds are misjudged as obstacles. It also reduces the risk of mistakenly deleting nearby real obstacles due to excessive model deviation, providing clean and reliable background data for subsequent completion and fusion.

[0035] Step 400: Identify self-occluding void regions in the candidate environmental point cloud, use the kinematic state parameters of the operating machinery as prior conditions, input them into the point cloud completion network to complete the point cloud, and obtain the completed environmental point cloud.

[0036] Point cloud completion is a process that addresses the perception blind spots caused by crane self-occlusion. It utilizes deep learning technology combined with the unique kinematics of cranes to infer and recover the environmental geometric information of the area occluded by the crane's own structure.

[0037] Kinematic state parameters as prior conditions refer to encoding parameters that characterize the current configuration of the crane, such as slewing angle, luffing angle, and telescopic length, into feature vectors to guide the neural network to generate complete results that conform to physical laws and environmental continuity.

[0038] In some implementations, the spatial distribution of the candidate environment point cloud is first analyzed to identify the missing point cloud regions caused by occlusion from structures such as booms and towers, i.e., self-occluding void regions. Then, sparse point cloud features of the surrounding adjacent regions can be extracted. At the same time, the current kinematic state parameters are input into a dedicated encoding branch of the point cloud completion network to generate a kinematic prior feature vector.

[0039] By employing a cross-modal attention fusion mechanism within the network, kinematic priors are deeply coupled with local geometric features, driving the decoder to generate dense point clouds for filling in empty regions. This allows the high-confidence filled point clouds to be merged with candidate environment point clouds to form a complete environment point cloud. For example, when the boom is positioned at a large elevation angle and obstructing scaffolding behind it, the network can reasonably infer the scaffolding structure of the obstructed portion based on the boom's current posture and the orientation of the surrounding point clouds, rather than fabricating it out of thin air or simply filling in a plane.

[0040] Step 500: Based on the real-time attitude parameters and physical size constraints of the operating machinery, generate a dynamic envelope of hidden obstacles; fuse candidate environmental point clouds, complete environmental point clouds and dynamic envelope of hidden obstacles to generate global obstacle perception results.

[0041] Among them, the hidden obstacle dynamic envelope refers to the virtual safety boundary that exists objectively but cannot be directly detected by sensors, calculated based on the physical size limitations and real-time motion status of the crane itself. Examples include the tail swing sweep area, the outward swing trajectory of the jib, the maximum working radius limit, and the clearance height boundary.

[0042] The global obstacle perception result is a comprehensive map that fully represents the risks in the work space by fusing explicit environmental point clouds (including candidate and complete parts) with implicit dynamic envelopes under a unified spatiotemporal framework.

[0043] In some implementations, the dynamic envelope of various hidden obstacles is calculated in real time based on real-time attitude parameters and physical size constraints, and then transformed into an occupancy grid or geometric boundary representation.

[0044] Simultaneously, the candidate environmental point clouds, the completed environmental point clouds, and these latent envelopes are spatially aligned and merged into the occupied grid map under the same coordinate system. This allows for the prediction of the crane's trajectory over a future period by combining the crane's current speed and acceleration, and the calculation of the minimum collision time with each obstacle in the global obstacle map along this trajectory. Based on this, multi-level warning signals or control commands are generated.

[0045] It should be understood that the fusion process is not a simple data overlay, but a logical integration based on risk assessment. Implicit envelopes and explicit point clouds have equal or even higher priority at the decision-making level to ensure zero tolerance for potential collision risks.

[0046] It should be noted that this step constructs a comprehensive safety protection system covering both explicit and implicit risks. By transforming physical size limitations into calculable virtual obstacles, it effectively compensates for the inherent limitations of sensor field of view and prevents accidents such as rear-end collisions and overturning due to exceeding limits, which are difficult to detect with traditional sensing solutions. Simultaneously, the fused global perception results provide a complete and reliable environmental model for downstream path planning and motion control, ensuring the safety and efficiency of crane operations under complex dynamic conditions.

[0047] Based on the above technical solution, this embodiment achieves a higher-level conceptual verification of the crane obstacle avoidance and early warning method by constructing a closed-loop perception framework that includes dynamic model correction, self-occlusion completion, and implicit envelope generation. This framework does not rely on specific sensor models or algorithm implementations, but rather defines the data flow and logical relationships between functional modules, providing a robust foundation and ensuring that the dynamic accuracy of the perception benchmark, the recoverability of blind spot information, and the comprehensiveness of risk coverage are maintained under different working conditions and crane models, thereby improving the robustness and practicality of the intelligent obstacle avoidance system for cranes.

[0048] Example 2: like Figure 2 As shown, based on Example 1, this example further refines the specific implementation process of correcting the digital geometric model based on the base attitude parameters and the body stress deformation parameters. Addressing the two major interference sources commonly found in crane operations—body tilt and boom elastic deformation—this example establishes a dual correction mechanism that couples coordinates and deformation to ensure that the digital geometric model can accurately represent the crane's actual spatial occupancy under complex working conditions.

[0049] Step 201: Based on the base tilt angle, the coordinate transformation matrix between the sensor coordinate system and the base coordinate system of the operating machinery for collecting 3D point cloud data of the environment is corrected in real time to obtain the tilt-compensated coordinate transformation relationship.

[0050] The tilt-compensated coordinate transformation relationship refers to the mathematical transformation operator recalculated to accurately map the raw point cloud data collected by the lidar from the sensor's local coordinate system to the crane's motion coordinate system, taking into account the tilt of the crane base relative to the horizontal ground plane. This transformation relationship is no longer a fixed factory calibration constant, but a dynamic matrix that changes in real time with the base's tilt angle to eliminate the systematic coordinate deviation introduced by the non-parallelism between the installation reference plane and the gravity plumb line.

[0051] In some implementations, an inertial measurement unit mounted on a turntable or chassis acquires the roll and pitch angles of the base in real time, constructing a rotation matrix characterizing the current tilt state of the base. Simultaneously, pre-calibrated sensor mounting extrinsic parameters are retrieved, namely, the fixed rotation and translation matrix of the sensor coordinate system relative to the ideal horizontal base coordinate system.

[0052] When the base tilts, the sensor's actual spatial orientation becomes a composite transformation of the rotation matrix and the fixed rotation-translation matrix. Therefore, within each frame of data processing, the system uses the current rotation matrix to perform left or right multiplication on the original extrinsic parameters (depending on the coordinate system definition order) to generate a tilt-compensated coordinate transformation matrix specific to the current moment.

[0053] For example, when the crane is parked on a 3-degree slope, the system will automatically calculate the rotational offset of the radar point cloud in the base coordinate system and apply this offset to the coordinate transformation of all the original point clouds, so that the transformed point cloud distribution is kept in the same reference system as the crane's own kinematic model.

[0054] It should be noted that in actual construction site environments, ground flatness is often difficult to guarantee. Without this correction, even if the boom itself is not deformed, the radar field of view deflection caused by ground tilt alone can result in misalignment of tens of millimeters or even centimeters in its point cloud filtering window, leading to numerous false alarms or missed detections. By dynamically updating the coordinate transformation relationship, it is ensured that regardless of the crane's tilt state, the sensor data can be converted into a reference frame consistent with the digital geometric model.

[0055] Step 202: Based on the load weight, estimate the amount of body deflection deformation of the working machinery under the load state, and offset and correct the geometric position of the corresponding deformable parts in the digital geometric model according to the amount of body deflection deformation.

[0056] Specifically, the steps include: obtaining the telescopic length and luffing angle of the operating machinery; calculating the spatial offset of each section of the boom of the operating machinery based on beam deformation theory or a reduced-order finite element model, combined with the load weight, telescopic length, and luffing angle; and superimposing the spatial offset onto the corresponding boom section nodes in the digital geometric model to update the three-dimensional spatial pose of the boom.

[0057] Among them, the body flexural deformation refers to the elastic displacement vector of a slender and flexible component such as a crane boom, which deviates from its ideal rigid geometric position under the combined action of its own weight and external load.

[0058] Spatial offset refers to the component of the deformation in three-dimensional space along a specific direction (usually vertically downward or along the normal to the boom axis).

[0059] In some implementations, instead of relying on time-consuming full-scale finite element real-time simulation, a proven, engineered reduced-order model is used to balance accuracy and real-time performance. Specifically, the processor first reads three key state variables: the current load weight, the telescopic length, and the luffing angle.

[0060] Input these variables into the preset beam deformation calculation module or the finite element reduced-order model: If beam deformation theory is adopted, the multi-section telescopic arm is equivalent to a variable cross-section cantilever beam, and the deflection of each section caused by bending moment is calculated using the formulas of mechanics of materials. If a finite element reduced-order model is used, the modal basis or surrogate model constructed in the offline stage is used to reconstruct the full-field deformation through a fast solution of a small number of generalized coordinates.

[0061] The calculated spatial offsets are a set of discrete numerical sequences, each corresponding to a set of pre-defined key cross-section nodes on the boom. Subsequently, the boom mesh vertices in the digital geometric model are traversed, and based on their respective cross-section intervals, the corresponding spatial offsets are superimposed onto the original coordinates of each vertex using linear interpolation or spline interpolation, thereby driving the entire boom model to undergo bending deformation in accordance with physical laws.

[0062] For example, when lifting a 30-ton load and with the boom fully extended, the deflection at the boom tip may reach more than 100 millimeters. The system will smoothly distribute this deformation to all model nodes from the root to the tip, so that the virtual boom exhibits a sag curve consistent with the real boom.

[0063] It should be noted that under heavy loads or long-arm conditions, the boom will undergo elastic deformation. If the digital geometric model maintains an ideal, straight shape, there will be a significant positional difference between the model's edges and the actual boom surface. This will cause the point cloud filtering algorithm to either retain a large number of boom point clouds, triggering false alarms, or mistakenly delete nearby real obstacles due to the excessively large model envelope. By using node-level offset correction, the digital twin model gains the ability to deform dynamically, ensuring that the model can still closely conform to the real structure under high-load operating scenarios, thus improving the accuracy and robustness of its point cloud filtering.

[0064] Step 203: Based on the tilt-compensated coordinate transformation relationship and the offset-corrected geometric position, generate the corrected digital geometric model.

[0065] The corrected digital geometric model refers to a high-fidelity virtual image that fully represents the current real spatial posture of the crane, integrating both base tilt compensation and body flexural deformation correction effects.

[0066] In some implementations, after executing steps 201 and 202, the processor applies the tilt-compensated coordinate transformation relationship to the digital geometric model that has already undergone flexural deformation correction, or conversely, it first transforms the model to the tilted base coordinate system and then applies deformation offset within that coordinate system.

[0067] Regardless of the execution order, the final corrected digital geometric model output simultaneously reflects the coupling effect of the two states: "the crane is tilted" and "the boom is bent." This is because in the physical world, a tilted ground changes the direction of the gravity vector relative to the boom axis, thus affecting the distribution of flexural deformation; and the boom deformation, in turn, affects the center of gravity and overturning moment of the entire vehicle. Although these two are often decoupled in engineering simplification, it is essential to ensure that they are strictly aligned in the spatiotemporal dimensions when generating the final model. For example, when a crane laterally lifts a heavy object on a slope, the boom not only deflects vertically but may also experience slight lateral bending due to the lateral component force; the corrected model should be able to reflect this complex deformation characteristic.

[0068] It should be noted that if tilt compensation is based solely on inertial measurement units while ignoring load-induced deflection, severe model mismatch will still occur under heavy-load conditions. Conversely, if deformation correction is based solely on load sensors while ignoring ground tilt, coordinate deviations will also occur when operating on non-level ground. Only by organically combining the two can the spatial pose of the boom under any complex working conditions be accurately and completely reflected.

[0069] Based on the above technical solution, this embodiment achieves high-fidelity reproduction of the crane's real physical state by constructing a dual dynamic correction mechanism combining tilt compensation and deflection correction. This mechanism not only eliminates the systematic deviation of the coordinate system caused by uneven ground, but also accurately compensates for the structural elastic deformation caused by loads, enabling the virtual model to adapt to the crane's non-rigid motion characteristics. This paradigm shift from static rigid body assumptions to dynamic soft body perception improves the accuracy and robustness of its point cloud filtering, providing a geometric basis for achieving accurate obstacle avoidance and early warning with zero false alarms and zero missed detections under complex working conditions.

[0070] Example 3: like Figure 3 As shown, based on Example 1, this example further refines the specific implementation process of filtering out the self-structure point cloud from the environmental 3D point cloud data based on the modified digital geometric model. Addressing the uncertainty in the self-point cloud position caused by dynamic interference during crane operation, this example proposes a filtering mechanism based on real-time swing amplitude adaptive adjustment of the matching threshold to solve the problem of residual false alarms or accidental deletion of nearby obstacles that are easily generated under dynamic conditions when using fixed thresholds.

[0071] Step 301: Obtain the real-time swing amplitude parameters of the working robot boom.

[0072] Among them, the real-time swing amplitude parameter refers to the physical quantity that characterizes the instantaneous dynamic offset of the actual spatial position of the boom relative to the theoretical motion trajectory caused by non-steady-state factors such as inertial force, wind load, or hydraulic system impact during the movement. Unlike the quasi-static flexural deformation caused by gravity and vehicle tilt described in Example 2, the real-time swing amplitude parameter has the characteristics of high frequency, randomness, and transient change. It is usually manifested as a slight vibration or sway in three-dimensional space at the end or middle of the boom, and its amplitude and frequency fluctuate in real time with the intensity of the operation and environmental conditions.

[0073] In some implementations, the angular velocity or linear acceleration integral value of the boom relative to the theoretical attitude is calculated in real time by using a high-frequency sampling inertial measurement unit or an accelerometer installed at a key node of the boom, thereby obtaining the current swing amplitude.

[0074] The real-time sway deviation can also be estimated by tracking and analyzing the feature points on the boom surface in a continuous multi-frame lidar point cloud. For example, when the operator quickly turns around and stops, the boom will generate significant residual vibration due to inertia, and the sway amplitude parameter detected by the sensor will increase instantaneously; while in a stable lifting and stationary state, this parameter remains at an extremely low level of ambient noise.

[0075] It should be noted that traditional solutions often assume the boom position is fixed after model correction, neglecting the positional ambiguity caused by high-frequency vibrations. This leads to either insufficient threshold coverage of the vibration range during dynamic operations, resulting in residual point clouds, or excessively large thresholds that remain loose even in static conditions, leading to the risk of accidental deletion. By introducing real-time swing amplitude parameters, the system can perceive the current dynamic confidence level.

[0076] Step 302: Based on the corrected digital geometric model, generate the theoretical self-point cloud set at the current moment; set the initial matching distance threshold, perform nearest neighbor matching between each point in the environmental 3D point cloud data and the theoretical self-point cloud set, and calculate the minimum Euclidean distance.

[0077] Among them, the theoretical point cloud refers to the set of spatial points that the crane should occupy in its current pose, generated by virtual rendering or geometric projection based on the ideal rigid body kinematic model and the static or quasi-static deformation correction of Example 2.

[0078] The initial matching distance threshold refers to the reference Euclidean distance tolerance used to determine whether a real point cloud belongs to its own structure under stable conditions or low dynamic disturbance conditions. Its value is usually determined by a combination of the sensor's ranging accuracy, calibration error, and the model's manufacturing tolerance.

[0079] In some implementations, the processor uses the graphics rendering pipeline or analytical geometry algorithms to discretize the modified digital geometric model into a theoretical point set with a density comparable to the real radar point cloud. Then, for each point in the environmental 3D point cloud data, the minimum Euclidean distance between it and its nearest neighbor in the theoretical point cloud set is calculated. This allows the system to use an initial matching distance threshold for discrimination by default when starting up or detecting low oscillation amplitude.

[0080] For example, when the sensor accuracy is high and the model calibration is good, the initial matching distance threshold can be set to 5 mm. This ensures that only when the deviation between the real point cloud and the theoretical model surface is within 5 mm will it be recognized as its own structure. This makes the processing logic focus on high precision, aiming to protect real obstacles adjacent to the boom from being mistakenly deleted under static or steady-state conditions.

[0081] It should be noted that setting the initial matching distance threshold reflects the pursuit of static accuracy, which is a prerequisite for ensuring that the system has high-resolution sensing capabilities under normal operating conditions. However, relying solely on this fixed benchmark cannot cope with all operating conditions; it must be combined with subsequent dynamic adjustment mechanisms to form a complete robust filtering scheme.

[0082] Step 303: When the real-time swing amplitude parameter exceeds the preset swing threshold, the matching distance threshold is relaxed from the initial value to the relaxed value.

[0083] The preset swing threshold is the critical criterion for distinguishing between a stable boom state and a state of significant dynamic disturbance. Its value is pre-calibrated based on the structural characteristics and safety margin of the crane.

[0084] The relaxation value is an extended tolerance, greater than the initial matching distance threshold, set to accommodate the positional uncertainty of the boom during violent swings. Its size is usually positively correlated with or piecewise mapped to the real-time swing amplitude parameter.

[0085] In some implementations, the swing amplitude parameter obtained in step 301 is monitored in real time and compared with a preset swing threshold. Once the swing amplitude is detected to exceed the threshold, it indicates that there is a large instantaneous uncertainty in the current boom position, and the original initial matching distance threshold is no longer sufficient to cover the actual boom positioning range. At this time, the system automatically triggers the threshold switching logic to dynamically adjust the matching distance threshold to a relaxed value.

[0086] Specifically, the adjustment can be step-wise, such as directly switching the threshold from 5 mm to 15 mm; it can also be continuous and linear, that is, increasing the threshold proportionally according to the increase in swing amplitude; or it can be a non-linear mapping based on a lookup table to adapt to the actual envelope characteristics of different models under different amplitudes.

[0087] For example, on a certain type of truck crane, when the swing speed at the boom end is detected to exceed 0.2 m / s or the peak swing displacement exceeds 10 mm, the matching threshold is smoothly transitioned to 15 mm to ensure that all self-point clouds during the vibration process can be effectively captured.

[0088] It should be understood that the upper limit of the relaxation value should be subject to safety constraints and cannot be expanded indefinitely to avoid swallowing up information about adjacent real obstacles under extreme swings. Generally, this upper limit does not exceed three times the nominal accuracy of the sensor or a specific proportion of the boom cross-sectional size.

[0089] It should be noted that by dynamically adjusting the tolerance, the system actively adapts to the dynamic position ambiguity of the boom, effectively solving the problem of residual point cloud caused by vibration.

[0090] Step 304: Points whose minimum Euclidean distance is less than the current matching distance threshold are identified as their own structure point clouds and removed from the environmental 3D point cloud data. The remaining point clouds are output as candidate environmental point clouds.

[0091] The current matching distance threshold refers to the real-time threshold that is actually applied to the current frame data processing after being dynamically adjusted in step 303. It may be an initial value, a relaxed value, or a transitional value between the two.

[0092] Candidate environment point clouds refer to the set of point cloud data that, after the above adaptive filtering process, are considered not to belong to the crane's own structure and are retained for subsequent obstacle analysis and completion.

[0093] In some implementations, the current matching distance threshold determined in step 303 is used to make a final classification decision on the environmental 3D point cloud data. Points whose minimum Euclidean distance to their theoretical point cloud set is less than the current threshold are marked as part of their own structure and removed from the original data stream; conversely, points whose distance is greater than or equal to the current threshold are retained.

[0094] This process is executed in real time within each frame of data processing. As the boom's swing state changes, the tightness of the filtering is also adjusted synchronously within milliseconds. For example, when the boom recovers from violent swinging to stability, the swing amplitude parameter falls back below the preset threshold, and the matching distance threshold will automatically tighten back to its initial value, thereby immediately restoring the high-precision perception capability of the adjacent area and avoiding unnecessary missed detection risks caused by maintaining a loose threshold for a long time.

[0095] It should be noted that through this dynamic and static filtering mechanism that combines leniency and strictness, when the swing is violent, priority is given to ensuring the complete removal of its own point cloud to prevent false alarms from interfering with normal operations; when the swing is gentle, priority is given to ensuring the precision of environmental perception to prevent real obstacles from being mistakenly deleted.

[0096] Based on the above technical solution, this embodiment introduces a real-time swing amplitude parameter to drive an adaptive adjustment mechanism for the matching threshold, thus constructing a self-point cloud filtering method that can adapt to the dynamic characteristics of the boom. This method overcomes the performance bottleneck of traditional fixed-threshold filtering schemes under dynamic conditions, avoiding false alarms due to self-point cloud residue caused by vibration and preventing the accidental deletion of real obstacles due to overly relaxed thresholds. Therefore, it can ensure thorough filtering while preserving the integrity of environmental perception to the greatest extent. This not only improves the robustness of the algorithm but also provides a cleaner and more reliable data foundation for subsequent point cloud completion and fusion decisions.

[0097] Example 4: like Figure 4 As shown in Example 1, this example further refines the identification of self-occluding void regions in the candidate environment point cloud and uses the kinematic state parameters of the operating machinery as prior conditions, inputting them into the point cloud completion network for point cloud completion. Addressing the problem of geometric illusions easily generated by general algorithms in crane self-occlusion blind zone completion, this example constructs a kinematic prior deep embedding completion mechanism. Through polar coordinate gridding representation and confidence quality control, it ensures that the completion results both conform to environmental continuity and strictly adhere to mechanical motion constraints.

[0098] Step 401: Obtain the kinematic state parameters of the operating machinery, and input the body posture parameters and working condition feature parameters into the kinematic coding branch in the point cloud completion network to generate kinematic prior feature vectors.

[0099] Among them, kinematic state parameters refer to a set of physical quantities that can uniquely determine the current spatial configuration and stress state of the crane, including body attitude parameters and working condition characteristic parameters.

[0100] The body attitude parameters include at least rotation angle, amplitude angle, and extension length, which are used to describe the geometric pose of the mechanical structure.

[0101] The operating condition characteristic parameters include at least the base tilt angle and the load weight, which are used to characterize the influence of the external environment on the mechanical form.

[0102] The kinematic prior feature vector transforms the aforementioned scalar parameters with clear physical meaning into a dense representation in a high-dimensional latent space through nonlinear mapping. This vector semantically encodes the spatial occupancy pattern of the crane's own structure, the range of line-of-sight obstruction, and the possible elastic deformation trend of the boom at the current moment.

[0103] In some implementations, the raw data from the rotary encoder, amplitude angle sensor, telescopic length sensor, inclinometer, and load sensor are first read in real time through the sensor interface and then normalized to eliminate dimensional differences.

[0104] The processed parameters are input into the kinematic coding branch of the point cloud completion network. This branch is usually composed of multiple fully connected layers or one-dimensional convolutional layers. Its internal weights are trained with a large amount of simulation data and have learned to associate the combination of physical parameters with the corresponding spatial geometric constraints.

[0105] For example, when the input parameter combination is "large elevation angle + long boom length + heavy load", the encoding branch will output a specific feature vector, which strongly activates the semantic patterns of "large area occlusion under boom" and "significant boom tip deflection" in the latent space.

[0106] It should be understood that the design of the kinematic coding branch is not a simple numerical concatenation, but rather a bridge connecting the physical world and the feature space of the neural network through end-to-end training. In other implementations, lookup table interpolation or symbolic regression models can also be used to generate prior features, as long as these features can effectively convey the geometric constraint information of the crane.

[0107] It should be noted that in crane operation scenarios, the shape and position of the self-occluding area are entirely determined by the kinematic state of the machine itself, rather than by the environmental texture. Introducing kinematic prior feature vectors enables the network to infer reasonable geometric structures based on mechanical principles even when faced with sparse or even completely missing point cloud data, reducing the probability of generating false obstacles or distorting real obstacles.

[0108] Step 402: Identify self-occluding hole regions in the candidate environment point cloud, extract sparse point cloud features from adjacent preset regions, convert the spatial location of the self-occluding hole regions into polar coordinate grid indexes, and align the sparse point cloud features with the kinematic prior feature vectors according to the indexes.

[0109] Among them, the self-occluding void region refers to a continuous spatial sector in the candidate environmental point cloud that lacks effective measurement points due to the crane's own structure blocking the line of sight of the lidar.

[0110] Sparse point cloud features refer to the local descriptors of point clouds that remain within a certain range outside the boundary of the void region and can provide environmental geometric clues.

[0111] The polar coordinate grid index is a two-dimensional discretized spatial representation system with the crane's rotation center as the origin and divided according to horizontal angular intervals and radial distance intervals. The N grid cells in the index uniquely correspond to the spatial sector, and N is a preset value.

[0112] Feature alignment refers to spatially registering heterogeneous kinematic prior feature vectors with sparse point cloud features within a unified polar coordinate grid framework, so that each grid cell is simultaneously associated with the corresponding geometric observation information and physical prior information.

[0113] In some implementations, the theoretical self-occlusion boundary at the current moment is calculated based on the modified digital geometric model and radar field of view model, and the start and end angles and radial range of the hole region are accurately marked by combining the actual distribution of the candidate environmental point cloud.

[0114] Point clouds within a preset angle (e.g., ±5°) on both sides of the hole are extracted, and sparse point cloud features are encoded using feature extractors such as PointNet++ or DGCNN. Simultaneously, the spatial coordinates of the hole region and its neighborhood are converted into polar coordinate grid indices. For example, the horizontal 360° is divided into 720 0.5° sectors, and the radial 0-50m is divided into 250 0.2m rings, forming a 720×250 grid matrix.

[0115] For each grid cell that falls into a hole or neighborhood, the corresponding prior sub-feature is retrieved or interpolated from the kinematic prior feature vector based on its angle index, and then concatenated or added to the sparse point cloud features of the grid cell in the channel dimension to complete the alignment operation.

[0116] Step 403: Input the sparse point cloud features and the kinematic prior feature vector into the cross-modal attention fusion module. Through the cross-attention mechanism, the kinematic prior feature vector is embedded into the point cloud features to obtain the fused features. Then, input the fused features into the decoder of the point cloud completion network to generate the completed point cloud for the self-occluded hole region.

[0117] Among them, the cross-modal attention fusion module is the core reasoning unit of the point cloud completion network. Its function is to realize the dynamic interaction and complementarity between geometric observation and physical knowledge at the feature level.

[0118] Cross-attention mechanism is a computational process that allows features (queries) of one modality to be adaptively reweighted and aggregated based on features (keys / values) of another modality. In this embodiment, it is specifically manifested as using kinematic prior features to "query" and "filter" information in sparse point cloud features that is compatible with mechanical constraints, while using point cloud features to "correct" and "concretize" the abstract constraints in prior features.

[0119] The fusion feature is an enhanced feature representation that combines environmental geometric details and the rationality of mechanical motion, generated after the above two-way interaction.

[0120] The decoder is a generative network that progressively upsamples the fused features and maps them back to the 3D point cloud space. Its output, the completed point cloud, geometrically fills in the gaps and semantically maintains coherence with the surrounding environment and consistency with its own kinematics.

[0121] In some implementations, the cross-modal attention fusion module contains multiple cascaded cross-attention layers. In each layer, sparse point cloud features serve as the query, and kinematic prior feature vectors serve as the key and value (or vice versa, depending on the specific attention flow design). An attention map is obtained by scaling dot product attention, which reflects the degree of attention each local geometric point receives to different kinematic constraints.

[0122] For example, at the edge of a hole, point cloud features may focus more on the component of the prior features related to the "boom cross-sectional shape," while deeper in the hole, they may focus more on the component related to the "boom extension direction." After multi-head attention aggregation and residual connection, the resulting fused features are fed into the decoder.

[0123] Decoders typically employ transposed convolution or upsampling MLP structures to progressively restore spatial resolution, ultimately outputting the predicted point cloud block corresponding to each polar coordinate grid cell.

[0124] It should be noted that in crane self-occlusion scenarios, the interior of the openings often lacks any direct observation information. Purely geometric methods can only guess at smooth surfaces or repetitive textures, easily missing critical obstacles such as slender rods or scaffolding. This solution, however, injects kinematic priors into every stage of feature generation through cross-attention, ensuring that the completion process is always constrained and guided by the laws of mechanical motion. This "knowledge-driven perception" paradigm ensures that even in extreme blind spots with no observation, the completed point cloud roughly matches the possible distribution of the boom or obstacles, improving the robustness and completeness of the perception system under complex conditions.

[0125] Step 404: The decoder of the point cloud completion network outputs the predicted point cloud block corresponding to each grid cell and generates a confidence value for each predicted point. Predicted points with confidence values ​​lower than a preset threshold are discarded, and predicted points with confidence values ​​not lower than the preset threshold are retained and merged with the candidate environment point cloud to obtain the completed environment point cloud.

[0126] The confidence score is a self-evaluation metric by which the point cloud completion network assesses the probability that each generated predicted point belongs to a real obstacle. Its value is usually between 0 and 1, reflecting the network's confidence in the geometric location and existence of the point.

[0127] The preset threshold is a confidence level filtering threshold pre-defined according to the security policy, used to distinguish between reliable completion points and potential noise points.

[0128] The completed environmental point cloud is a highly reliable set of completed points after confidence screening. Its union with the original candidate environmental point cloud forms the complete data foundation for subsequent global obstacle perception and obstacle avoidance decisions.

[0129] In some implementations, the decoder, while outputting the 3D coordinates, predicts a confidence value for each point using an additional sigmoid output header. This confidence value is obtained during training through joint optimization with the chamfer distance or occupancy loss of the real point cloud, and can reflect the distribution of prediction errors well.

[0130] During the inference phase, all predicted points are iterated through, and only those with a confidence level greater than or equal to a preset threshold (e.g., 0.7) are retained. Points with a confidence level below the threshold are discarded, regardless of how reasonable their geometric location may seem. They are considered unreliable predictions.

[0131] For example, in cases where the boom tail swings violently or abnormal operating parameters cause prior failure, the network may generate some ambiguous or contradictory completion points. The confidence of these points is usually reduced, and they are automatically filtered out.

[0132] It should be understood that the preset threshold needs to strike a balance between "recall rate" and "false alarm rate": an excessively high threshold will filter out a large number of real obstacles, reducing the integrity of perception; an excessively low threshold will introduce too many false alarms, increasing the risk of false alarms. In actual deployment, this threshold can be dynamically calibrated based on historical false alarm statistics or operator feedback.

[0133] Based on the above technical solutions, this embodiment achieves a highly reliable point cloud completion method specifically for crane self-occlusion scenarios by constructing a complete technical closed loop of kinematic prior coding, polar coordinate grid alignment, cross-modal attention fusion, and confidence-based safe filtering. This method deeply integrates the unique kinematic knowledge of cranes into the neural network architecture, overcoming the incompatibility of general completion algorithms in strongly constrained scenarios. Polar coordinate gridding representation efficiently adapts to the rotational perception characteristics of cranes. Simultaneously, the cross-attention mechanism achieves an organic unity between physical laws and geometric observations. Finally, confidence filtering ensures the safe availability of the completion results, enabling the system to maintain a high recall rate and low false alarm rate for real obstacles in blind spots even under extreme self-occlusion conditions such as large boom tilt angles.

[0134] Example 5: like Figure 5 As shown, based on Example 1, this example further refines the specific implementation process of generating a dynamic envelope of hidden obstacles based on the real-time attitude parameters and physical size constraints of the operating machinery. Addressing the problem of blind spots in the perception of the crane's own motion limit area due to the limited field of view of sensors during crane operation, this example constructs a mechanism that transforms physical size constraints into virtual safety boundaries in real time. By calculating four types of dynamic envelopes—tail swing, boom outward swing, maximum working radius, and clearance height—it achieves explicit early warning of hidden collision risks.

[0135] Step 501: Obtain the real-time attitude parameters and physical dimensional constraints of the operating machinery.

[0136] Among them, real-time attitude parameters refer to the set of dynamic kinematic variables that characterize the current spatial configuration of the working machinery, including at least rotation angle, amplitude angle and extension length, which together determine the instantaneous position and orientation of each component of the machine body in three-dimensional space.

[0137] Physical dimensional constraints refer to the inherent properties or dynamic thresholds that limit the range of motion of machinery, as determined by mechanical design specifications and safe operating procedures. These constraints include at least the geometry of the boom tail, the geometry of the jib, and the load and amplitude characteristic curves, and are used to define the safety boundary conditions of the machinery under different operating conditions.

[0138] In some implementations, a high-precision absolute encoder is used to collect the rotation angle and amplitude angle in real time, and the extension length is obtained by a draw rope displacement sensor or a laser rangefinder. The sampling frequency is usually no less than 50Hz to ensure the real-time update of the envelope.

[0139] For physical dimensional constraints, the geometry of the boom tail (such as the tail turning radius and counterweight overhang distance) and the geometry of the jib (such as the jib length and folding angle range) are stored as fixed constants in the system configuration file; while the load and amplitude characteristic curve is preset in the memory in the form of a two-dimensional lookup table or a piecewise fitting function. This curve reflects the maximum allowable working amplitude under different lifting weights and is the key safety red line to prevent overturning accidents.

[0140] For example, when the operator performs a compound action, the current slewing angle, amplitude angle, telescopic length, and lifting weight value are read synchronously in each control cycle, and the corresponding safety radius limit is immediately retrieved from the characteristic curve.

[0141] It should be noted that this step, by establishing a real-time correlation between dynamic states and static constraints, enables the system to perceive motion exclusion zones that objectively exist even though they are not within the sensor's field of view. This solves the problem of perception blind spots caused by limitations in sensor installation location (such as radar being located in the middle of the boom and unable to detect the tail) or physical limitations (such as limited amplitude under heavy load).

[0142] Step 502: Based on real-time attitude parameters and physical size constraints, calculate the tail swing envelope, the arm outward swing envelope, and the maximum working radius limit envelope.

[0143] Among them, the tail swing envelope refers to the boundary of the annular area swept by the tail of the boom during rotation, which is used to warn personnel or equipment behind to avoid entering the danger sector.

[0144] The outer swing envelope of the boom refers to the horizontal projection trajectory boundary of the outer end point of the boom in the deployed state, which is used to define the space occupied by the boom when it is extended laterally.

[0145] The maximum working radius limiting envelope is a circular boundary generated with the rotation center as the center and the safe working radius corresponding to the current load as the radius, used to prevent the risk of overturning caused by overload torque.

[0146] In some implementations, for the tail swing envelope, based on the current slewing angle and the pre-stored boom tail geometry (including tail length and width), the radial distance and tangential range of the tail profile relative to the slewing center at the current moment are calculated using polar coordinate transformation, thereby generating a fan-shaped or ring-shaped warning zone that rotates in real time with the slewing motion.

[0147] For the outer swing envelope of the auxiliary arm, when the auxiliary arm is detected to be in the extended state, the projection coordinate sequence of the outer endpoint of the auxiliary arm on the horizontal plane is calculated by combining the amplitude angle, extension length and the folding angle of the auxiliary arm itself through positive kinematics. These coordinate points are then connected to form a closed polygonal boundary, which accurately describes the lateral encroachment space of the auxiliary arm in complex postures.

[0148] For the maximum working radius limit envelope, the system uses the current load weight and telescopic length as indexes to query the load and amplitude characteristic curves to obtain the current allowable safe working radius, and then draws a circular boundary with radius in the occupied grid map with the rotation center as the origin.

[0149] For example, in a certain lifting operation, when the lifting weight increases from 10 tons to 25 tons, the maximum working radius limit envelope is reduced from 30 meters to 18 meters. If the boom continues to move outward and approach this boundary at this time, it is determined that there is a hidden collision (overturning) risk.

[0150] Step 503: Based on the body's flexural deformation, amplitude angle, and telescopic length, calculate the spatial coordinates of the lowest point of the boom after elastic deformation, and generate the net height envelope.

[0151] The clearance height envelope refers to the vertical safety boundary of the actual lowest point of the boom relative to the ground or obstacles below, after taking into account the elastic deflection of the boom under load.

[0152] The body deflection deformation characterizes the elastic deflection of the boom under load. This parameter can directly reuse the calculation result of step 202 in Example 2 without repeated calculation. This envelope is specifically designed to address the vertical collision risk caused by the actual boom height being lower than the theoretical geometric height under heavy load conditions.

[0153] In some implementations, the bending deformation of the main body under the current working condition is obtained from the dynamic geometric model correction module of Embodiment 2. This deformation has taken into account the nonlinear effects of the load weight, telescopic length, and luffing angle. Then, combined with the current luffing angle and telescopic length, the theoretical height of the boom end or key section in the undeformed state is calculated using trigonometric functions. The corresponding bending deformation is then subtracted to obtain the actual spatial coordinates of the lowest point after elastic deformation.

[0154] Using this lowest point as a reference, a preset safety margin (e.g., 0.5 meters) is extended downwards to generate a planar boundary parallel to the ground or a curved boundary distributed along the boom axis as the clearance height envelope. For example, when a 45-meter-long boom is operating under a 30-ton load, the boom tip deflection can reach 112 millimeters. If this deformation is not considered, it might be assumed that there is still 2 meters of clearance below the boom, but in reality, only 1.89 meters remain, making it extremely easy to touch the scaffold crossbar below. After introducing the clearance height envelope, this 112-millimeter deflection will be included in the calculation, and the vertical warning line will be updated in real time.

[0155] It should be understood that, in addition to calculating the minimum height at a single point, multiple cross-sectional nodes can be selected along the boom length to calculate the height after deformation, generating a continuous clearance height curve to more precisely describe the overall sag shape of the boom.

[0156] It should be noted that existing technologies typically assume the boom is a rigid body and calculate the theoretical height solely based on the angle feedback from the encoder, neglecting elastic deformation under heavy loads. This can lead to serious misjudgments of clearance in large-tonnage lifting scenarios. This embodiment extends the deformation correction results from Embodiment 2 to the envelope generation stage, ensuring that the clearance height boundary always reflects the true physical state. This effectively prevents the middle or end of the boom from accidentally touching the material piles, pipelines, or temporary facilities below during luffing or slewing, thus improving vertical safety protection capabilities.

[0157] Step 504: Output at least one of the tail swing envelope, the auxiliary arm outward swing envelope, the maximum working radius limit envelope, and the clearance height envelope as the hidden obstacle dynamic envelope.

[0158] The hidden obstacle dynamic envelope is a collective term or combination of the aforementioned envelopes. Its output form can be a geometric boundary equation, a sequence of polygon vertices, an occupied grid mask, or a symbolic distance field (SDF), depending on the data interface requirements of the downstream fusion decision module. This output signifies that the physical limitations of the operating machinery itself have been successfully transformed into a digital risk entity that can be treated in the same way as environmental obstacles.

[0159] In some implementations, the type of envelope to be activated is intelligently selected based on the current operating mode and security policy. For example: In rotary operation mode, the tail swing envelope and the maximum working radius limit envelope are output first; When the boom is deployed, the outward swing envelope of the boom is automatically superimposed. In heavy-load hoisting mode, the headroom envelope is forcibly activated.

[0160] All activated envelopes are transformed to the same global coordinate system as the candidate environment point cloud and the completed environment point cloud, and are labeled with their own constraints so that they can be given higher priority or different collision penalty weights in the subsequent fusion stage.

[0161] It should be noted that by outputting a standardized dynamic envelope of hidden obstacles, not only are sensor blind spots filled, but a knowledge-driven safety protection paradigm is also established. This allows the system to move beyond simple real-time sensor detection, instead encoding human experts' understanding of the crane's mechanical characteristics and motion laws into algorithmic logic, giving the system predictive capabilities. Even in extreme environments such as dust obstruction, strong light interference, or complete visual obstruction, the virtual envelope generated based on physical laws can still provide a reliable safety baseline, ensuring the robustness and integrity of the obstacle avoidance and early warning system under all weather conditions and operating conditions.

[0162] Based on the above technical solution, this embodiment constructs a four-dimensional implicit obstacle dynamic envelope system, including tail swing, boom outward swing, maximum working radius, and clearance height, thereby realizing the real-time digitization and explicitness of the physical limitations of the operating machinery itself. This, combined with the dynamic model correction in Embodiment 2 and the point cloud completion in Embodiment 4, forms a comprehensive perception network covering both the explicit environment and implicit constraints. By transforming motion limit areas that sensors cannot directly detect into calculable virtual obstacles, the inherent blind spots of pure perception solutions are effectively compensated for, preventing typical accidents such as tail scraping, boom interference, overload overturning, and insufficient clearance, thus improving the intrinsic safety level of the crane under complex dynamic working conditions.

[0163] Example 6: like Figure 6 As shown, based on Example 1, this example further refines the specific implementation process of fusing candidate environmental point clouds, completing environmental point clouds, and dynamic envelopes of latent obstacles to generate global obstacle perception results. Addressing the difficulty in uniformly evaluating multi-source heterogeneous perception data in crane operations, and the inability of static distance warnings to adapt to dynamic operational risks, this example constructs a fusion decision-making mechanism based on spatiotemporal joint analysis. By uniformly mapping explicit environment and implicit constraints to an occupied grid map and combining kinematic prediction to calculate the minimum collision time, it achieves a leap from passive perception to proactive risk management.

[0164] Step 601: Obtain the current motion speed parameters and motion acceleration parameters of the working machinery.

[0165] Among them, motion speed parameters refer to the set of physical quantities that characterize the instantaneous rate of change of each degree of freedom of the working machinery, including at least rotational angular velocity, amplitude angular velocity and extension speed, which are used to describe the real-time motion trend of the machine body in space.

[0166] Motion acceleration parameters refer to the rate of change of the aforementioned velocity parameters, including at least angular acceleration, amplitude acceleration, and extension acceleration, which are used to reflect the inertial state of mechanical motion and the intensity of the control intention.

[0167] In some implementations, the real-time velocity and acceleration values ​​are obtained by sampling the original position signals of the rotary encoder, amplitude angle sensor, and telescopic length sensor at high frequency, and then smoothing them using a sliding window differential algorithm or Kalman filter.

[0168] Alternatively, the angular velocity and linear acceleration data output by the inertial measurement unit can be directly read and mapped to each joint degree of freedom through coordinate transformation. To ensure synchronization with the sensing data, the update frequency of these motion parameters is usually no less than 50Hz, and they are strictly timestamped.

[0169] For example, when performing a rapid slewing braking operation, it can not only detect that the current slewing angular velocity is decreasing, but also predict that the boom will continue to slide forward a certain distance due to inertia through negative angular acceleration, thereby correcting the risk prediction model in advance.

[0170] It should be noted that traditional obstacle avoidance schemes often only compare the current distance with a fixed safety threshold, ignoring the machine's own kinetic energy and braking distance. This leads to late warnings during high-speed operations or frequent false alarms during low-speed operations. This embodiment, however, uses speed and acceleration as a fusion decision, enabling it to distinguish between different risk levels of stationary approach and high-speed collision.

[0171] Step 602: Align the candidate environment point cloud, the completed environment point cloud, and the dynamic envelope of hidden obstacles in spatial coordinates, and merge them into the occupied grid map under the same coordinate system to generate a global obstacle map.

[0172] Among them, the occupied grid map is a probabilistic environmental representation model that discretizes a continuous three-dimensional space into regular grid cells. Each grid cell stores a value representing the probability that the space is occupied by an obstacle.

[0173] A global obstacle map is a unified digital copy of the environment used for downstream path planning and collision detection, which integrates all known explicit obstacles (from sensor measurements and inferences) and implicit constraints (from physical model inferences). This map is not only a container for geometric information, but also a crystallization of multi-source knowledge fusion.

[0174] In some implementations, a global coordinate system is first established with the crane's slewing center as the origin, and a three-dimensional grid space is divided according to a preset resolution (e.g., 0.1m × 0.1m horizontally, 0.2m vertically). The candidate environmental point cloud is then directly projected onto the corresponding grid, and its occupancy probability is set to a high confidence value (e.g., 0.9). For completing the environmental point cloud, a weighted mapping is performed based on its associated confidence value. The higher the confidence value, the closer the probability of the point cloud is to the actual point cloud. Conversely, a lower uncertainty value (such as 0.4-0.6) is set to retain uncertainty information for decision-making reference. For the dynamic envelope of latent obstacles, since it represents an absolutely insurmountable physical boundary, all rasters within its coverage area are forcibly marked as fully occupied (probability 1.0) and have the highest priority. Even if there are low-confidence complete point clouds or holes in the area, the latent envelope shall prevail.

[0175] For example, when the network completion infers that there may be scaffolding behind the boom but the confidence level is only 0.5, if the location happens to fall within the tail swing envelope, the grid will be immediately set to 1.0, because the area itself is a no-movement zone regardless of whether the scaffolding exists.

[0176] It should be noted that candidate point clouds are discrete, precise, but sparse; completed point clouds are dense, continuous, but probabilistic; and latent envelopes are abstract, absolute, but unobservable. Mapping them uniformly onto an occupied grid map not only achieves precise spatial alignment but also completes logical integration at the semantic level. In particular, assigning the highest priority to latent envelopes ensures that, under any circumstances, the machine's own safety baseline will not be obscured by uncertain environmental perception.

[0177] Step 603: Based on the motion speed parameters and motion acceleration parameters, and combined with the global obstacle map, predict the motion trajectory of the working machinery within a preset time interval in the future, and calculate the minimum collision time with each obstacle in the global obstacle map along the trajectory.

[0178] The motion trajectory refers to the continuous sequence of positions of key points of the mechanical body (such as the arm end, tail, and auxiliary arm end) in three-dimensional space, which is deduced based on the current motion state and dynamic constraints within the time window.

[0179] Minimum collision time is the time required for any point on the trajectory of a machine to first touch any occupied grid cell in the global obstacle map, assuming the machine maintains its current or expected motion pattern. It is a core indicator for quantifying the urgency of collision risk.

[0180] In some implementations, the dynamic window method or forward simulation algorithm is used for trajectory prediction: the processor takes the current attitude, velocity and acceleration as initial conditions, and combines the kinematic equations and dynamic constraints of the crane (such as maximum acceleration and deceleration, hydraulic response characteristics) to iteratively extrapolate the pose sequence of the next N times in a fixed time step (such as 20ms) to generate one or more possible motion trajectories.

[0181] For each sampling point on the trajectory, query the occupancy probability of the corresponding location in the global obstacle map. If the occupancy probability of a point exceeds a safety threshold (e.g., 0.8), record the timestamp corresponding to that point. After traversing the entire trajectory, take the minimum value of all collision times as the current minimum collision time.

[0182] If no collision is detected within the entire prediction window, the minimum collision time is set to infinity or the upper limit of the window. For example, if a crane is rotating at a speed of 0.5 rad / s with an angular acceleration of 0, and the prediction is that the boom tail will sweep across an area with an occupancy probability of 1.0 within the next 2 seconds, and the calculation shows that the contact occurs after 1.2 seconds, then the output TTC = 1.2s.

[0183] It should be noted that traditional closest distance indicators cannot distinguish the essential difference between slow approach and high-speed impact. In contrast, minimum collision time directly couples distance and speed, and based on the calculation method of predicted trajectory, it takes into account the inertia and delay of machinery, avoiding the embarrassing situation of "although an alarm has been sounded, a collision still occurs" due to untimely braking, thus improving the effectiveness of safety protection.

[0184] Step 604: Based on the comparison results between the minimum collision time and multiple preset warning thresholds, determine the current warning level and output the corresponding warning signal or control command.

[0185] Among them, the preset multiple warning thresholds are a set of time nodes pre-calibrated based on the crane's dynamic performance, operator reaction time and safety margin, used to divide the continuous minimum collision time into several discrete risk intervals.

[0186] The warning level is a security protection level that corresponds one-to-one with each risk zone, reflecting the system's assessment of the urgency of the current situation.

[0187] Warning signals or control commands are specific intervention measures taken for different levels, including but not limited to sound and light prompts, voice broadcasts, automatic deceleration, restriction of movement direction, or emergency braking, aiming to achieve maximum safety with minimal interference.

[0188] In some implementations, a three- or four-level early warning strategy is set. For example: The first threshold is 3 seconds. When the TTC is greater than 3 seconds, it is considered a safe state and no intervention is taken. The second threshold is 2 seconds. When 2 seconds is less than or equal to TTC and less than or equal to 3 seconds, a yellow warning is triggered. The operator is alerted by intermittent buzzer sounding and flashing yellow light on the instrument panel, but no control intervention is performed. The third threshold is 1 second. When 1 second is less than or equal to TTC and less than 2 seconds, an orange warning is triggered, and the current action speed is automatically reduced proportionally (e.g., reduced to 30% of the original speed) while issuing a continuous rapid alarm. When TTC is less than 1 second, a red emergency warning is triggered. The system immediately cuts off the power supply to the hydraulic valve group of the relevant action or sends an emergency stop command to force the machine to stop moving until the risk is eliminated.

[0189] The above thresholds are not fixed and can be dynamically adjusted according to the current load weight, boom length or working mode. For example, the threshold can be appropriately increased to increase the safety margin under heavy load and long boom working conditions.

[0190] Based on the above technical solution, this embodiment achieves a comprehensive upgrade of crane obstacle avoidance early warning from static geometric protection to dynamic spatiotemporal safety protection by constructing a complete closed loop of motion state perception, multi-source data fusion, spatiotemporal risk prediction, and hierarchical response decision-making. This mechanism not only unifies the expression paradigm of explicit environment and implicit constraints, but also deeply integrates the motion laws of the physical world into the decision-making logic through the minimum collision time index. This enables the system to provide intelligent protection that conforms to human cognitive habits and engineering safety specifications while ensuring millisecond-level real-time response, thus solving the technical problems of delayed early warning, frequent false alarms, and coexistence of safety blind spots in traditional solutions under complex dynamic working conditions.

[0191] Example 7: like Figure 7 As shown, this embodiment provides a crane obstacle avoidance and early warning system based on obstacle localization. This system serves as the physical carrier of the aforementioned method embodiments, and through the coordinated configuration of hardware sensing and computing units, it provides physical support for the all-around obstacle perception and early warning functions of the operating machinery. Specifically, the crane obstacle avoidance and early warning system includes a sensor module, a point cloud acquisition module, a memory, and a processor.

[0192] The sensor module is the physical basis for the system to perceive its own state. It includes not only inertial measurement units, attitude sensors and load sensors, but also various integrated sensing units that correspond to the key state variable inputs required in the previous method embodiments, and are used to obtain the base attitude parameters, body force deformation parameters and kinematic state parameters of the multi-degree-of-freedom working machine.

[0193] Specifically, the inertial measurement unit typically uses a six-axis or nine-axis industrial-grade MEMS device, which is installed at the center of the crane turntable or chassis. It outputs raw data of three-axis acceleration and three-axis angular velocity at a frequency of not less than 2kHz. After internal calculation, it provides high-precision roll angle and pitch angle as base attitude parameters to support the real-time correction of the tilt compensation matrix in step 201 of embodiment 2.

[0194] The attitude sensors include an absolute encoder installed on the slewing mechanism, an inclinometer installed on the luffing cylinder, and a rope displacement sensor inside the telescopic arm. They are used to accurately measure the slewing angle, luffing angle, and telescopic length. They are the source of the kinematic prior feature vector in Example 4 and the geometric reference for calculating the dynamic envelope of the hidden obstacle in Example 5.

[0195] The preferred load cell is a pin-type load cell or a main oil circuit pressure sensor, which provides real-time feedback of the load weight signal, providing a mechanical basis for estimating the body deflection in step 202 of Example 2 and querying the maximum working radius limit envelope in Example 5.

[0196] The sensing units are connected to the central processing unit via CAN FD bus or RS422 differential communication link, and microsecond-level time synchronization is achieved by hardware triggering or PTP protocol to ensure that all status parameters point to the same physical snapshot at the same moment.

[0197] It should be noted that this module provides high-confidence state input to the upper-layer algorithm through a high-frequency, low-latency multi-source heterogeneous data acquisition architecture, enabling core methods and steps such as dynamic geometric model correction, adaptive threshold adjustment, and kinematic prior guidance to operate stably in the real physical system, fundamentally ensuring the real-time accuracy of the sensing benchmark.

[0198] The point cloud acquisition module is the only direct window for the system to perceive the external environment. Its performance determines the original quality and spatial coverage of the candidate environment point cloud, and it is used to acquire 3D point cloud data of the environment.

[0199] Specifically, the module typically consists of one or more lidar units. 16-line or 32-line mechanical rotating lidars can be selected to obtain higher vertical resolution, while solid-state flash lidars can be used to improve vibration resistance and lifespan. The installation location is generally at the top of the turntable or at the base of the boom to ensure effective coverage of the area in front of and to the side of the work.

[0200] The point cloud acquisition module connects to the processor via gigabit Ethernet or a dedicated data interface, and outputs a 3D point cloud stream containing XYZ coordinates and reflection intensity information at a frame rate of 10Hz to 20Hz.

[0201] During the system initialization phase, precise extrinsic parameter calibration of this module is required to determine its rotational and translational relationship relative to the mounting reference of the inertial measurement unit in the sensor module. This calibration result will serve as the initial value for the coordinate transformation matrix correction in Example 2. Furthermore, the point cloud acquisition module also receives synchronization pulse signals from the sensor module to ensure that each frame of point cloud data carries a precise global timestamp, thereby ensuring strict alignment with parameters such as the base attitude and body deformation acquired simultaneously.

[0202] For example, when performing the self-point cloud filtering in Example 3, if there is a millisecond-level asynchrony between the point cloud and the attitude data, a positional deviation of several centimeters may occur under the condition of rapid boom rotation, resulting in filtering failure.

[0203] The memory is a persistent carrier of the system software logic and knowledge model. It carries the executable code and pre-set data tables that implement all the aforementioned methods and steps, and is used to store program instructions.

[0204] Specifically, memory typically consists of two parts: non-volatile flash memory and high-speed dynamic random access memory. The non-volatile flash memory contains core program instructions such as the operating system kernel, sensor driver, point cloud preprocessing algorithm, dynamic geometric model correction module, point cloud completion network weight file, implicit envelope calculation function library, and multi-level early warning decision logic. It also stores static knowledge bases such as the crane's factory calibration parameters, load-amplitude characteristic curve table, and beam deformation reduced-order model coefficients. High-speed dynamic random access memory serves as a data exchange buffer during runtime, temporarily storing real-time sensor data streams, intermediate calculation results (such as corrected digital geometric models, candidate environment point clouds, occupied grid maps), and activation values ​​during neural network inference.

[0205] During system startup, the processor loads key programs and models from non-volatile flash memory into memory, and dynamically calls corresponding data table entries according to changes in operating conditions during operation. For example, when generating the maximum working radius limit envelope in Embodiment 5, the processor quickly retrieves the corresponding safe radius value from the load-amplitude characteristic curve table in memory based on the current lifting weight and boom length index.

[0206] It should be noted that this module transforms abstract methods and processes into a sequence of binary instructions that can be executed by machines, and encodes domain expert knowledge into structured data, thereby solidifying the technical solution from theory to engineering and ensuring the consistency and reproducibility of the system's behavior in different work cycles.

[0207] The processor is the computational hub and decision engine of the entire obstacle avoidance and early warning system. Its hardware architecture and computing power configuration directly determine whether the aforementioned complex methods and steps can be completed under strict real-time constraints. It is configured to execute program instructions stored in memory: to correct the digital geometric model based on the base attitude parameters and the body's stress deformation parameters; based on the corrected digital geometric model, to filter out the self-structure point cloud from the environmental 3D point cloud data to obtain candidate environmental point clouds; to identify self-occluding void regions in the candidate environmental point clouds, and to input the kinematic state parameters of the operating machinery as prior conditions into the point cloud completion network to complete the point cloud, thereby obtaining the completed environmental point cloud; to generate a dynamic envelope of hidden obstacles based on the real-time attitude parameters and physical size constraints of the operating machinery; and to fuse the candidate environmental point cloud, the completed environmental point cloud, and the dynamic envelope of hidden obstacles to generate a global obstacle perception result.

[0208] Specifically, processors typically employ heterogeneous computing platforms. For example, a quad-core ARM Cortex-A series CPU is used to handle serial tasks such as business logic scheduling, sensor data fusion, and early warning decision-making. At the same time, an FPGA or GPU acceleration unit is integrated to specifically handle high-parallel computing loads such as point cloud completion network inference, large-scale point cloud filtering, and occupied grid map updates.

[0209] During execution, the processor first reads synchronized data from the sensor module and the point cloud acquisition module, and calls the correction algorithm in memory to generate a dynamic geometric model. Then, it uses this model to perform adaptive filtering on the original point cloud, and sends the residual point cloud and kinematic priors to the hardware acceleration unit to run the network completion. At the same time, the CPU calculates the dynamic envelopes of various hidden obstacles in parallel. Finally, the three results are fused in a unified coordinate system to generate a global obstacle map, and combined with motion prediction, it outputs a warning level or control command.

[0210] End-to-end latency across the entire processing chain is controlled to meet the safety response requirements of dynamic crane operations. For example, in the cross-modal attention fusion computation of Example 4, the FPGA acceleration unit can compress the time consumed in a single inference, so that even when the boom rotates rapidly and causes frequent changes in the self-occlusion area, the completed point cloud can still be updated in real time without lag.

[0211] It should be understood that the specific processor model and architecture can be flexibly selected according to the performance requirements of the machine. For lightweight models, a pure SoC solution can be used, while for large crawler cranes, a multi-card parallel computing cluster can be expanded. In addition, the system adopts a modular quick-installation design, and the installation interfaces of the sensor module and point cloud acquisition module are compatible with the ISO 13000 standard size, which improves the adaptability to cranes of different brands and tonnages.

[0212] Based on the above technical solutions, this embodiment provides a solid and reliable physical foundation for the aforementioned obstacle avoidance and early warning method for cranes based on obstacle localization by constructing a hardware system architecture with high-precision sensing, real-time acquisition, heterogeneous computing, and modular adaptation. The spatiotemporal synchronization design of the sensor module and the point cloud acquisition module ensures data consistency required for dynamic model correction and adaptive filtering; the hardware-software co-configuration of the memory and processor supports millisecond-level real-time response for point cloud completion network and implicit envelope calculation; and the standardized modular interface breaks the limitation of a single model, enhancing the universality and engineering value of the technology. This system is not merely a simple mapping of the methodological process, but also a specialized engineering optimization for the harsh working conditions and stringent safety requirements of cranes, ensuring seamless integration from theoretical innovation to field application and comprehensively improving the inherent safety level and intelligence of crane operations.

[0213] Example 8: To further verify the effectiveness and robustness of the present invention in complex real-world operating environments, this embodiment provides a composite working condition application verification based on a real crane operation scenario. This verification aims to demonstrate the synergistic effect of the technical branches in the aforementioned embodiments, such as dynamic geometric model correction, adaptive point cloud filtering, kinematic prior-guided completion, and dynamic envelope of hidden obstacles, through system performance under extreme conditions, and to provide factual evidence for the inventiveness of the invention.

[0214] In this application verification, a set of extremely challenging composite working conditions were set: the operating machinery was a certain type of mobile truck crane with a lifting capacity of 30 tons, the main boom fully extended to 45 meters, parked on an inclined ground with a slope of 3°, and performing a rapid slewing maneuver at a large elevation angle (75°). This working condition simultaneously superimposed multiple interference factors such as coordinate system deviation caused by vehicle tilt, significant elastic deformation caused by heavy-load long boom, dynamic swaying of the boom caused by rapid slewing, and severe self-blocking caused by the large elevation angle, which is an extreme test of the comprehensive performance of the obstacle avoidance and warning system.

[0215] Under this combined working condition, the system first addresses the reference mismatch problem through the dynamic geometric model correction mechanism of Example 2: due to the 3° tilt of the ground, if the traditional horizontal assumption is used, a lateral position deviation of about 2.3 meters will occur between the radar point cloud and the boom model (at the 45-meter boom end), which is enough to cause its own point cloud filtering to completely fail.

[0216] This system reads the roll and pitch angles output by the inertial measurement unit in real time and performs tilt compensation on the coordinate transformation matrix to eliminate this systemic bias. Simultaneously, for the approximately 112 mm boom tip deflection resulting from the combination of a 30-ton load and a 45-meter boom length, a reduced-order beam deformation model is used to calculate the offset of each section in real time and correct the digital geometric model, ensuring that the virtual boom curve closely matches the actual physical boom height.

[0217] Following this, during the high-angle rapid slewing operation, the adaptive threshold filtering mechanism described in Example 3 addresses dynamic swaying interference. When the operator performs an emergency stop during slewing, the boom sways momentarily due to inertia, and the real-time sway amplitude parameter instantly exceeds the preset threshold. The system automatically and smoothly widens its own point cloud matching distance threshold from 5 mm in static mode to 15 mm, effectively accommodating the positional uncertainty of the boom during vibration. During this stage, point clouds of the boom's own structure that deviate from their theoretical positions due to swaying are successfully eliminated, and no false alarms triggered by residual point clouds occur. In contrast, if a fixed 5 mm threshold is used, an average of 43 residual point clouds per frame would be misjudged as obstacles under the same operating conditions, leading to continuous false alarms. The introduction of the adaptive mechanism allows the system to maintain the purity of its perception even during periods of drastic dynamic change.

[0218] Meanwhile, to address the severe self-occlusion area behind the boom caused by large-angle slewing, the kinematic prior-guided point cloud completion network of Example 4 was activated. In this scenario, the boom's own structure occludes the scaffolding steel pipes within a 60° sector behind it, an area where traditional methods have no perception capability. This system encodes the current slewing angle, luffing angle, telescopic length, and load weight into kinematic prior feature vectors and injects them into the completion process through a cross-modal attention fusion module. Based on the prior constraints of "large angle + long boom," combined with sparse point cloud cues from the occlusion edges, the network successfully infers and generates the scaffolding completion point cloud for the occluded area.

[0219] Furthermore, during the slewing process, the hidden obstacle dynamic envelope and fusion decision-making mechanism of Embodiments 5 and 6 provides safety protection beyond the sensor's field of view. Although the lidar cannot directly detect the sweeping area at the boom's tail, the system dynamically generates a tail swing envelope based on the real-time slewing angle and tail geometry, and injects it as a high-priority virtual obstacle into the global grid map. When the boom's tail approaches the material stacking area behind the operator's cab, the system calculates that the minimum collision time is less than the warning threshold before a collision occurs, automatically triggering an orange warning and reducing the slewing speed, successfully avoiding a potential tail scraping accident. This warning does not rely on real-time sensor detection but is based on proactive reasoning of the mechanical motion laws, filling the inherent blind spots of pure perception solutions.

[0220] It should be noted that this application example is for illustrative purposes only and is not intended to limit the scope of protection of this invention. In actual engineering applications, the above performance indicators may fluctuate due to differences in sensor selection, computing platform power, specific crane model, and environmental conditions. However, as long as the core technical concept disclosed in this invention is adopted, namely the closed-loop processing architecture driven by multi-source state perception for dynamic model correction, adaptive filtering, prior completion, and implicit envelope generation, it should be considered to fall within the protection scope of this invention.

[0221] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of variations or substitutions within the technical scope disclosed in this invention. For example, equivalent sensing substitutions can be made for the acquisition methods of base attitude parameters and body stress deformation parameters; different mechanical reduction models or neural network fittings can be used for the correction algorithm of the digital geometric model; structural modifications can be made to the cross-modal attention fusion mechanism in the point cloud completion network; the types of hidden obstacle dynamic envelopes can be added, subtracted, or combined; and adaptive adjustments can be made to the warning threshold and grading strategy. As long as these do not depart from the core technical concept defined by the claims of this invention, i.e., the overall scheme for achieving all-round accurate obstacle avoidance warning for cranes through multi-source state perception-driven dynamic model correction, adaptive filtering, prior-guided completion, and hidden envelope generation, should be covered within the scope of protection of this invention. Therefore, the scope of protection of this invention should be determined by the scope of the claims.

Claims

1. A crane obstacle avoidance and early warning method based on obstacle localization, characterized in that, include: Acquire the base attitude parameters, body stress deformation parameters, and environmental 3D point cloud data of the operating machinery; The digital geometric model is corrected based on the base attitude parameters and the body stress deformation parameters; Based on the corrected digital geometric model, the self-structure point cloud is filtered out from the environmental 3D point cloud data to obtain the candidate environmental point cloud. Identify self-occluding void regions in the candidate environmental point cloud, use the kinematic state parameters of the operating machinery as prior conditions, input them into the point cloud completion network to complete the point cloud, and obtain the completed environmental point cloud. Based on the real-time attitude parameters and physical size constraints of the operating machinery, a dynamic envelope of hidden obstacles is generated; By fusing the candidate environmental point cloud, the completed environmental point cloud, and the dynamic envelope of the hidden obstacle, a global obstacle perception result is generated.

2. The crane obstacle avoidance and early warning method based on obstacle localization according to claim 1, characterized in that, The process of correcting the digital geometric model based on the base attitude parameters and the body stress deformation parameters includes: The base attitude parameters and the body stress deformation parameters are invoked, wherein the base attitude parameters include the base tilt angle and the body stress deformation parameters include the load weight; Based on the base tilt angle, the coordinate transformation matrix between the sensor coordinate system and the work machinery base coordinate system that collects the environmental three-dimensional point cloud data is corrected in real time to obtain the tilt-compensated coordinate transformation relationship. Based on the load weight, estimate the amount of body deflection deformation of the working machinery under the load state, and offset and correct the geometric position of the corresponding deformable component in the digital geometric model based on the amount of body deflection deformation. Based on the tilt-compensated coordinate transformation relationship and the offset-corrected geometric position, a corrected digital geometric model is generated.

3. The crane obstacle avoidance and early warning method based on obstacle localization according to claim 2, characterized in that, The process of estimating the amount of body deflection deformation of the working machinery under lifting conditions based on the lifting weight, and then offsetting and correcting the geometric position of the corresponding deformable component in the digital geometric model based on the amount of body deflection deformation, includes: Obtain the telescopic length and amplitude angle of the operating machinery; Based on beam deformation theory or finite element reduced-order model, and combined with the lifting weight, the telescopic length and the luffing angle, the spatial offset of each section of the boom of the operating machinery is calculated; The spatial offset is superimposed on the corresponding boom section node in the digital geometric model to update the three-dimensional spatial pose of the boom.

4. The crane obstacle avoidance and early warning method based on obstacle localization according to claim 1, characterized in that, The process of using the kinematic state parameters of the operating machinery as prior conditions and inputting them into the point cloud completion network for point cloud completion includes: Obtain the kinematic state parameters of the operating machinery. The kinematic state parameters include body posture parameters and working condition characteristic parameters. The body posture parameters include at least rotation angle, luffing angle and telescopic length. The working condition characteristic parameters include at least base tilt angle and load weight. Identify self-occluding hole regions in the candidate environment point cloud and extract sparse point cloud features of the preset regions adjacent to the self-occluding hole regions; The body posture parameters and the working condition feature parameters are input into the kinematic coding branch of the point cloud completion network to generate a kinematic prior feature vector. The sparse point cloud features and the kinematic prior feature vector are input into the cross-modal attention fusion module. Through the cross-attention mechanism, the kinematic prior feature vector is embedded into the point cloud features to obtain the fused features. The fused features are input into the decoder of the point cloud completion network to generate a completed point cloud for the self-occluding hole region, and then merged with the candidate environment point cloud to obtain a completed environment point cloud.

5. The crane obstacle avoidance and early warning method based on obstacle localization according to claim 4, characterized in that, The input / output format and confidence filtering process of the cloud completion network includes: The spatial location of the self-occluding void region is converted into a polar coordinate grid index. The polar coordinate grid index is divided into N grid cells according to the horizontal angle interval and the radial distance interval. The grid cell corresponds uniquely to the spatial sector, and N is a preset value. The sparse point cloud features and the kinematic prior feature vectors are aligned according to the polar coordinate grid index and then input into the cross-modal attention fusion module. The decoder of the point cloud completion network outputs the predicted point cloud block corresponding to each grid cell and generates a confidence value for each predicted point, wherein the confidence value represents the probability that the predicted point belongs to a real obstacle. Predicted points with confidence values ​​below a preset threshold are discarded, while predicted points with confidence values ​​not below the preset threshold are retained and merged with the candidate environmental point cloud to obtain a complete environmental point cloud.

6. The crane obstacle avoidance and early warning method based on obstacle localization according to claim 1, characterized in that, The process of generating a dynamic envelope of hidden obstacles based on the real-time attitude parameters and physical size constraints of the operating machinery includes: The real-time attitude parameters and physical dimensional constraints of the operating machinery are obtained. The real-time attitude parameters include at least the slewing angle, luffing angle, and telescopic length. The physical dimensional constraints include at least the geometry of the boom tail, the geometry of the jib, and the load-amplitude characteristic curve. Based on the slewing angle and the geometry of the boom tail, the tail swing envelope is calculated, which is the boundary of the annular area swept by the boom tail during slewing. Based on the amplitude angle, the telescopic length and the geometric parameters of the auxiliary arm, the outer swing envelope of the auxiliary arm is calculated. The outer swing envelope of the auxiliary arm is the horizontal projection trajectory boundary of the outer end point of the auxiliary arm in the deployed state. Based on the telescopic length and the lifting weight of the operating machinery, the load and amplitude characteristic curve is queried to obtain the safe working radius, and a maximum working radius limiting envelope with the rotation center as the center and the safe working radius as the radius is generated; At least one of the tail swing envelope, the auxiliary arm outward swing envelope, and the maximum working radius limiting envelope is output as the hidden obstacle dynamic envelope.

7. The crane obstacle avoidance and early warning method based on obstacle localization according to claim 6, characterized in that, The dynamic envelope of the hidden obstacle also includes a headroom envelope, wherein: The bending deformation of the working machinery under load is obtained, and the bending deformation of the machinery body represents the degree of elastic deflection of the boom under load. Obtain the luffing angle and telescopic length of the operating machinery, and calculate the spatial coordinates of the lowest point of the boom after elastic deformation by combining the deflection deformation of the main body. Based on the spatial coordinates of the lowest point, a clearance height boundary is generated below the boom, which serves as the clearance height envelope.

8. The crane obstacle avoidance and early warning method based on obstacle localization according to claim 1, characterized in that, The process of fusing the candidate environment point cloud, the completed environment point cloud, and the dynamic envelope of the latent obstacle to generate a global obstacle perception result includes: The current motion speed parameters and motion acceleration parameters of the working machinery are obtained. The motion speed parameters include at least rotational angular velocity, amplitude angular velocity, and extension speed. The motion acceleration parameters include at least rotational angular acceleration, amplitude angular acceleration, and extension acceleration. The candidate environment point cloud, the completed environment point cloud, and the hidden obstacle dynamic envelope are aligned in spatial coordinates and merged into an occupied grid map in the same coordinate system to generate a global obstacle map. Based on the motion speed parameters and the motion acceleration parameters, combined with the global obstacle map, the motion trajectory of the operating machinery within a future preset time interval is predicted, and the minimum collision time with each obstacle in the global obstacle map along the trajectory is calculated. Based on the comparison results between the minimum collision time and multiple preset warning thresholds, the current warning level is determined, and the corresponding warning signal or control command is output.

9. The crane obstacle avoidance and early warning method based on obstacle localization according to claim 1, characterized in that, The process of filtering out its own structural point cloud from the environmental 3D point cloud data based on the modified digital geometric model includes: The real-time swing amplitude parameter of the working machine boom is obtained. The real-time swing amplitude parameter is used to characterize the actual spatial position of the boom during the movement due to inertia, wind load or hydraulic impact, and the instantaneous offset relative to the theoretical motion trajectory. Based on the revised digital geometric model, generate the theoretical self-point cloud at the current moment; Set an initial matching distance threshold, perform nearest neighbor matching between each point in the environmental 3D point cloud data and the theoretical point cloud set, and calculate the minimum Euclidean distance. When the real-time swing amplitude parameter exceeds the preset swing threshold, the matching distance threshold is relaxed from the initial value to a relaxed value, where the relaxed value is greater than the initial value. Points whose minimum Euclidean distance is less than the current matching distance threshold are identified as their own structural point clouds and removed from the environmental 3D point cloud data. The remaining point clouds are output as candidate environmental point clouds.

10. A crane obstacle avoidance and early warning system based on obstacle localization, characterized in that, The crane obstacle avoidance warning system, applied to the obstacle location-based method for cranes according to any one of claims 1-9, specifically includes: The sensor module, including an inertial measurement unit, an attitude sensor, and a load sensor, is used to acquire the base attitude parameters, body stress deformation parameters, and kinematic state parameters of a multi-degree-of-freedom operating machine. Point cloud acquisition module, used to acquire 3D point cloud data of the environment; Memory, used to store program instructions; The processor is configured to execute program instructions stored in the memory to correct the digital geometric model based on the base attitude parameters and the body force deformation parameters; based on the corrected digital geometric model, filter out the self-structure point cloud from the environmental 3D point cloud data to obtain candidate environmental point clouds; identify self-occluding void regions in the candidate environmental point clouds, use the kinematic state parameters of the operating machinery as prior conditions, input them into the point cloud completion network to complete the point cloud, and obtain the completed environmental point cloud; generate a hidden obstacle dynamic envelope based on the real-time attitude parameters and physical size constraints of the operating machinery; and fuse the candidate environmental point cloud, the completed environmental point cloud, and the hidden obstacle dynamic envelope to generate a global obstacle perception result.