An object detection and positioning system based on sling vision
By integrating RTK positioning equipment and surround-type binocular cameras, combined with data processing units and dynamic obstacle avoidance optimization modules, the positioning accuracy and environmental perception problems of front-mounted cranes in complex environments have been solved, enabling precise positioning of the cranes and dangerous targets and safe operation.
Patent Information
- Application Number
- CN202511484532.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Traditional reach stacker positioning technology lacks accuracy in complex environments and has limited environmental perception capabilities, making it difficult to achieve high-precision operations and safely identify potentially dangerous targets.
It integrates RTK positioning equipment and surround-type binocular cameras, and realizes the attitude calculation of the lifting device and the precise spatial positioning of dangerous targets through the data processing unit. Combined with the dynamic obstacle avoidance optimization module, it optimizes the motion trajectory in real time.
It achieves synchronous and precise positioning of the spreader and the dangerous target, improving the safety and automation level of the front-end crane in complex environments.
Smart Images

Figure CN120943139B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of machine vision and automated control technology, and in particular to an object detection and positioning system based on lifting device vision. Background Technology
[0002] With the increasing demand for automation in large cargo handling scenarios such as ports and logistics parks, reach stackers, as a highly efficient container handling equipment, have become key to improving operational efficiency and safety through their precise positioning and safe operation capabilities.
[0003] Traditional reach stacker positioning mainly relies on technologies such as inertial navigation or laser ranging. However, these methods often suffer from positioning errors in complex environments (such as those with multiple obstacles or high dynamic scenarios) due to insufficient accuracy or environmental interference, making it difficult to meet the requirements of high-precision operations. In addition, the real-time perception capability of the surrounding environment of the spreader is insufficient. Existing systems mostly use a single sensor or a low-resolution camera, which makes it difficult to accurately identify and locate potential dangerous targets (such as personnel, obstacles, etc.), increasing safety hazards during operations.
[0004] Currently, RTK (Real-Time Kinematic) positioning technology, with its centimeter-level accuracy, is gradually being applied in the field of engineering machinery, while binocular cameras, with their depth perception capabilities, demonstrate significant advantages in 3D reconstruction and target detection. However, how to effectively integrate high-precision RTK positioning with multi-binocular camera systems, and achieve precise spatial positioning of lifting equipment attitude calculation, camera coordinate transformation, and hazardous targets through data processing units, remains a current technical challenge.
[0005] To address this issue, there is an urgent need for an integrated system that combines RTK positioning equipment, surround-view binocular cameras, and a data processing unit. This system can acquire precise spatial information about the spreader and surrounding hazardous targets in real time through high-precision positioning and multi-view image processing, thereby improving the intelligence level and operational safety of reach stackers. Summary of the Invention
[0006] The purpose of this invention is to provide an object detection and positioning system based on the vision of a lifting device to solve the problems mentioned in the background art.
[0007] This invention provides an object detection and positioning system based on lifting device vision, applied to front-facing lifting devices, comprising:
[0008] The RTK positioning device is installed above the front crane head to obtain the latitude and longitude of the front crane and establish a three-dimensional spatial coordinate system with its center as the origin.
[0009] Multiple binocular cameras are mounted around the hoist to capture real-time images of the surrounding environment.
[0010] The data processing unit is connected to the RTK positioning device, the multiple binocular cameras, and the front-end crane via a PLC.
[0011] The data processing unit is configured as follows:
[0012] Based on the data obtained from the PLC, the coordinates and rotational attitude of the bottom center of the lifting device in the three-dimensional spatial coordinate system are calculated.
[0013] Based on the pre-defined relative positional relationship between the camera center point of each binocular camera and the center point of the rigging, the coordinates of the camera center point of each binocular camera in the three-dimensional spatial coordinate system are calculated.
[0014] The system receives images captured by the binocular camera, identifies dangerous targets in the images using a loaded detection algorithm, and calculates the first coordinates P1 of the feature points on the dangerous targets in the camera coordinate system.
[0015] Based on the transformation relationship between the camera coordinate system and the three-dimensional spatial coordinate system, the first coordinate P1 is converted into the second coordinate P2 of the feature point in the three-dimensional spatial coordinate system;
[0016] The spatial location information of the dangerous target is output based on the second coordinate P2.
[0017] Optionally, before the operation begins, the data processing unit is further configured to perform a camera calibration operation to determine the intrinsic and extrinsic parameters of the left and right cameras of each of the binocular cameras.
[0018] Optionally, before the operation begins, the data processing unit is further configured to receive and load configuration information for personnel, vehicle, and foreign object detection algorithms.
[0019] Optionally, the relative positional relationship is obtained by physically measuring the relative distance between the center point of the camera and the center point of the rigging.
[0020] Optionally, calculating the first coordinate P1 of the feature point on the dangerous target in the camera coordinate system includes:
[0021] When a dangerous target is identified, one or more pixels near the center point of the dangerous target are selected as feature points.
[0022] The disparity of feature points is calculated using the principle of binocular vision, and camera parameters are obtained.
[0023] Based on the parallax and camera parameters, the coordinate value P1 of the feature point in the camera coordinate system is directly calculated.
[0024] Optionally, the transformation relationship between the camera coordinate system and the three-dimensional spatial coordinate system is determined by the coordinates and rotation attitude of the camera center point of each binocular camera in the three-dimensional spatial coordinate system.
[0025] Optionally, it may also include an alarm unit and / or a visualization unit for alarm processing and / or visualization display based on the output spatial location information of the hazardous target.
[0026] Optionally, the number of binocular cameras is four, which are respectively installed at the midpoint of the four sides of the rigging, so that the shooting range covers the perimeter of the rigging to form a detection area without blind spots.
[0027] Optionally, the object detection and positioning system based on spreader vision also includes a dynamic obstacle avoidance optimization module, used to optimize the spreader's movement trajectory in real time during reach stacking operations to address potential collision risks caused by the dynamic movement of hazardous targets. The dynamic obstacle avoidance optimization module is configured to perform the following operations:
[0028] The data processing unit obtains the second coordinate P2 of the hazardous target in the three-dimensional spatial coordinate system, as well as the real-time coordinates and rotation attitude of the bottom center of the lifting device in the three-dimensional spatial coordinate system.
[0029] Based on the second coordinate P2 sequence of historically collected hazardous targets, the velocity vector V and acceleration vector A of the hazardous targets are calculated, where V represents the displacement rate of the hazardous target in the three-dimensional spatial coordinate system, in meters per second, reflecting the direction and magnitude of the hazardous target's movement; A represents the acceleration of the hazardous target, in meters per square second, reflecting the change in the hazardous target's motion trend.
[0030] Based on the real-time coordinates and rotational attitude of the bottom center of the spreader, the expected motion trajectory T of the spreader is calculated. The expected motion trajectory T consists of a set of continuous three-dimensional spatial coordinate points, representing the predicted position of the spreader in the future time period.
[0031] Using the velocity vector V and acceleration vector A of the hazardous target, and the expected trajectory T of the lifting device, the dynamic obstacle avoidance optimization index D is calculated using the following formula:
[0032] D=[k1×||V||² / V0+k2×||A|| / A0+k3×min(d(T, P2)) / d0] / [1+exp(-Q)];
[0033] Wherein, D is the dynamic obstacle avoidance optimization index, ranging from 0 to 1. A higher value indicates a higher collision risk, requiring priority optimization of the spreader's trajectory; ||V||² is the square norm of the velocity vector of the hazardous target, measured in square meters per second squared, reflecting the kinetic energy of the hazardous target; ||A|| is the norm of the acceleration vector of the hazardous target, measured in meters per second squared, reflecting the drastic change in the hazardous target's motion; min(d(T, P2)) represents the minimum Euclidean distance between the expected trajectory T of the spreader and the second coordinate P2 of the hazardous target, measured in meters, reflecting the closest spatial proximity between the spreader and the hazardous target; k1, k2, and k3 are weighting coefficients, corresponding to the weighted influence of velocity, acceleration, and minimum distance, respectively. The range is 0 to 1, and k1+k2+k3=1. It is determined by training a machine learning model based on historical operation data; Q is the environmental complexity factor, representing the density and distribution complexity of hazardous targets in the operation environment. It ranges from 0 to 10, with higher values indicating more complex environments. It is calculated by the data processing unit based on the number and distribution uniformity of hazardous targets identified in the binocular camera images; exp(-Q) is an exponential function used to adjust the nonlinear influence of the environmental complexity factor on the dynamic obstacle avoidance optimization index, ensuring that the optimization index is more sensitive in complex environments; V0 is the square norm reference value, in square meters per square second; A0 is the norm reference value, in meters per square second; d0 is the Euclidean distance reference value, in meters.
[0034] Based on the calculated dynamic obstacle avoidance optimization index D, and by referring to the pre-set mapping relationship between the optimization index and the trajectory adjustment strategy, the optimized motion trajectory T' of the spreader is determined. The optimized motion trajectory T' is achieved by adjusting the spreader's moving speed, direction, or height to avoid the movement path of dangerous targets.
[0035] The optimized motion trajectory T' is transmitted to the PLC control system of the front-end crane to adjust the motion state of the spreader in real time, thereby effectively avoiding collisions with dynamically moving dangerous targets.
[0036] Optionally, determining the optimized motion trajectory T' of the spreader based on the calculated dynamic obstacle avoidance optimization index D, and by referring to the pre-set mapping relationship between the optimization index and the trajectory adjustment strategy, includes:
[0037] During the front-end crane operation, the data processing unit collects dynamic environmental change data in real time, and combines it with a quantum random number generator to generate a high-entropy random seed, which is used to initialize the trajectory planning state machine of the dynamic obstacle avoidance optimization module to solve the local convergence problem of trajectory optimization caused by dynamic environmental changes.
[0038] Based on the dynamic obstacle avoidance optimization index D, a multi-dimensional decision space is constructed, which includes the real-time coordinates of the lifting device, its rotational attitude, the velocity vector V and acceleration vector A of the dangerous target, and the environmental complexity factor Q.
[0039] By performing high-dimensional interpolation on the multidimensional decision space, a set of candidate optimized motion trajectories T'i are generated. Each candidate trajectory T'i consists of a set of adjusted three-dimensional spatial coordinate points, velocity vectors, and attitude angles, satisfying that the minimum Euclidean distance min(d(T'i, P2)) is greater than the safe distance threshold Dsafe. Dsafe is dynamically adjusted according to the dynamic hazard density of the work scenario, ranging from 0.5 meters to 2 meters.
[0040] Using a high-entropy random seed generated by a quantum random number generator, Monte Carlo simulations are performed on the candidate optimized motion trajectories T'i to calculate the collision probability Pcoll of each candidate trajectory T'i in a future time period. The collision probability Pcoll is calculated using the following formula:
[0041] Pcoll=Σ[wi×(1-exp(-α×||T'i-P2|| / σ))];
[0042] Where Pcoll represents the collision probability of candidate trajectory T'i, ranging from 0 to 1; wi is the weight of the i-th candidate trajectory, determined by a combination of trajectory smoothness and energy consumption, ranging from 0 to 1; ||T'i-P2|| represents the Euclidean distance between candidate trajectory T'i and the second coordinate P2 of the hazardous target; α is the collision sensitivity coefficient, dynamically adjusted according to the magnitude of the hazardous target's velocity vector V, ranging from 0.1 to 1; σ is the environmental uncertainty factor, calculated by regression analysis of the environmental complexity factor Q and historical collision data, in meters.
[0043] Candidate optimized motion trajectories T'i with collision probabilities Pcoll lower than a preset threshold Pth are selected, and the trajectory with the lowest dynamic obstacle avoidance optimization index D is chosen as the final optimized motion trajectory T'.
[0044] The present invention has achieved the following beneficial effects:
[0045] This invention integrates a high-precision RTK positioning device, a surround-type binocular camera, and a data processing unit to construct a three-dimensional perception and positioning solution for reach stackers. The system's workflow is as follows: First, an RTK device at the front of the vehicle establishes a centimeter-level precision three-dimensional spatial coordinate system as a global positioning reference. Then, four binocular cameras installed at the midpoints of the four sides of the spreader simultaneously capture images of the surrounding environment from different perspectives, forming a comprehensive detection coverage. The data processing unit, as the core, acquires RTK and vehicle attitude data via a PLC to calculate the precise coordinates and attitude of the spreader's bottom center. Simultaneously, it runs a detection algorithm to process the binocular images, identify hazardous targets, and calculate their depth information using binocular vision principles. Finally, through a series of coordinate transformations, the identified hazardous targets are positioned in a unified global three-dimensional coordinate system. This method fundamentally solves the problems of insufficient positioning accuracy and limited environmental perception in traditional solutions in dynamic and complex environments, achieving synchronous and accurate positioning of the spreader and hazardous targets. It provides key technical support for the automated and safe operation of reach stackers in scenarios such as ports and logistics parks.
[0046] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0047] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0048] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0049] Figure 1 This is a schematic diagram of the structure of an object detection and positioning system based on lifting device vision in an embodiment of the present invention;
[0050] Figure 2 This is a flowchart illustrating the steps involved in configuring a data processing unit to perform operations in an embodiment of the present invention. Detailed Implementation
[0051] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0052] The research and development approach of this application stems from solving the problem of precise positioning and safety perception of reach stackers in highly dynamic and complex environments. Addressing the insufficient accuracy and limited environmental perception capabilities of traditional positioning technologies, the research team proposed using a high-precision RTK positioning device as the core to construct a three-dimensional spatial coordinate system and achieving multi-view environmental perception through surround-type binocular cameras. During the development process, firstly, the latitude and longitude data of the reach stacker are acquired using the RTK device, establishing a coordinate system with the device's center as the origin. Then, multiple binocular cameras capture images of the surrounding environment of the spreader, and combined with the pre-calibrated relative positional relationship between the cameras and the spreader, the camera's position in the coordinate system is calculated. Finally, a detection algorithm is loaded through a data processing unit to identify hazardous targets in the images and calculate their spatial coordinates, completing the coordinate system transformation and position output. This approach, through multi-source data fusion and real-time processing, achieves precise positioning of the spreader's attitude and hazardous targets, providing a new technical path for intelligent operation of reach stackers.
[0053] Figure 1 This application provides a schematic diagram of an object detection and positioning system based on a lifting device's vision, as shown in the embodiment of the present application. Figure 1 As shown, this system is applied to a front-end crane and includes:
[0054] The RTK positioning device is installed above the front crane head to obtain the latitude and longitude of the front crane and establish a three-dimensional spatial coordinate system with its center as the origin.
[0055] Multiple binocular cameras are installed around the lifting device to capture real-time images of the surrounding environment. There are four binocular cameras, each installed at the midpoint of one of the four sides of the lifting device, so that the shooting range covers the entire perimeter of the lifting device to form a detection area without blind spots.
[0056] The data processing unit communicates with the RTK positioning device, the multiple binocular cameras, and the front-end crane via a PLC.
[0057] Among them, such as Figure 2 As shown, the data processing unit is configured as follows:
[0058] Step A: Based on the data obtained from the PLC, calculate the coordinates and rotational attitude of the bottom center of the lifting device in the three-dimensional spatial coordinate system.
[0059] This step aims to use the high-precision RTK positioning data acquired by the PLC to calculate the position coordinates and rotational attitude of the spreader's bottom center in a three-dimensional spatial coordinate system with the RTK device's center as the origin. The RTK positioning device refers to a positioning module employing real-time dynamic differential technology, capable of providing centimeter-level latitude and longitude positioning accuracy. Its output data includes latitude, longitude, and altitude, in the WGS84 coordinate system (longitude in degrees), latitude in degrees, and altitude in meters. The three-dimensional spatial coordinate system has the RTK device's center as the origin (0, 0, 0), with the X-axis pointing geographically east, the Y-axis pointing geographically north, and the Z-axis pointing vertically upward. The PLC (Programmable Logic Controller) acts as a data relay device, acquiring latitude and longitude data from the RTK device and performing coordinate transformation by combining it with the mechanical structural parameters of the front-end crane (such as the installation offset of the spreader relative to the RTK device). To ensure consistent time references for all sensor data, the system adopts a hardware time synchronization scheme based on PTP (Precision Time Protocol). The RTK device, binocular camera, attitude sensor, and PLC are all connected to a network switch supporting the PTP protocol. The PLC acts as the master clock, periodically issuing synchronization signals to control the clock synchronization error of each slave device to within 1 millisecond, thereby ensuring the time alignment of position, image, and attitude data and avoiding calculation errors caused by time asynchrony. The coordinates of the center of the spreader bottom are defined as (Xh, Yh, Zh), in meters, representing the position of the center of the spreader bottom in the three-dimensional spatial coordinate system; the rotation attitude is defined as Euler angles (α, β, γ), in degrees, representing the rotation angles around the X, Y, and Z axes, respectively. The rotation sequence is specified as ZYX (i.e., first rotate γ around the Z axis, then rotate β around the new Y axis, and finally rotate α around the new X axis). This sequence conforms to the conventions in vehicle dynamics, ensuring the uniqueness of the transformation result from the spreader coordinate system to the three-dimensional spatial coordinate system. The coordinate calculation process first uses latitude and longitude data acquired by the RTK device to convert them into coordinates in a local Cartesian coordinate system using a Gaussian projection algorithm. Then, it combines this with the physical offset between the RTK device and the center of the spreader bottom (obtained from mechanical design drawings, unit: meters) for translation transformation to obtain (Xh, Yh, Zh). Rotational attitude is calculated by integrating attitude sensor data from the front-end crane (such as angular velocity output from a gyroscope, unit: degrees / second) acquired by the PLC, combined with the initial calibration angle (obtained through level calibration, unit: degrees) to obtain Euler angles (α, β, γ). To ensure centimeter-level positioning and stable attitude perception in the dynamic operating environment of the front-end crane, the core calculation uses an extended Kalman filter (EKF) algorithm to fuse multi-source data from the RTK, gyroscope, and accelerometer. The state vector of the EKF filter is explicitly defined as x=[X,Y,Z,Vx,Vy,Vz,α,β,γ,ωα,ωβ,ωγ]^T, which has a total of 12 dimensions, and fully describes the translational and rotational motion of the lifting device in three-dimensional space.Wherein, [X, Y, Z] and [Vx, Vy, Vz] respectively represent the position (unit: meters) and velocity (unit: meters / second) of the bottom center of the lifting device in the three-dimensional spatial coordinate system established in this invention; [α, β, γ] represent the Euler rotation attitude angles (unit: degrees) of the lifting device following the ZYX sequence; and [ωα, ωβ, ωγ] are the corresponding three-axis angular velocities (unit: degrees / second), which serve as the direct observation state of the gyroscope.
[0060] The values of the main diagonal elements of the process noise covariance matrix Q are directly determined by the inherent white noise characteristics of the core sensors. Specifically, the angular random walk parameter of the gyroscope is set to 0.01 deg / s / √Hz, which mainly couples into the angular velocity state (ω) and propagates to the attitude angles (α, β, γ) through integration; the velocity random walk parameter of the accelerometer is set to 200 μg / √Hz, which mainly affects the velocity state (V) and further propagates to the position state (X, Y, Z). This design ensures that the filter can accurately assess the uncertainty of the system dynamics model in the prediction step.
[0061] The observation noise covariance matrix R is empirically assigned based on the measured accuracy of each sensor, with its diagonal elements representing the uncertainty variance of different observations. The position observations (X, Y, Z) provided by the RTK module are assigned a noise standard deviation of 0.02 meters; the velocity observations (Vx, Vy, Vz) obtained through RTK carrier phase difference decomposition are assigned a noise standard deviation of 0.05 meters per second; and during periods when the vehicle is stationary or moving at a constant speed, the pitch angle (β) and roll angle (α) observations calculated using the accelerometer gravity vector components are assigned a noise standard deviation of 0.1 degrees. This configuration allows the filter to intelligently balance the reliability of predicted values with observations from different sources.
[0062] Step B: Based on the pre-calibrated relative positional relationship between the camera center point of each binocular camera and the center point of the lifting device, calculate the coordinates of the camera center point of each binocular camera in the three-dimensional spatial coordinate system; the relative positional relationship is obtained by physically measuring the relative distance between the camera center point and the center point of the lifting device.
[0063] The relative attitude relationship is defined by the direction angles (Δα, Δβ, Δγ) of the camera optical axis relative to the rigging coordinate system, all in degrees. The calibration process uses a high-precision total station (such as a Leica TS60). First, stable measurement control points are established on the rigging, and the coordinates and attitude reference of the rigging bottom center are accurately measured. Then, the total station is aligned with the pre-set targets on the housings of each binocular camera mounted on the rigging, and the coordinates (ΔXc, ΔYc, ΔZc) of each camera center point (defined as the optical center points of the left and right cameras) are accurately measured. Finally, by measuring two points along the camera optical axis, the camera's mounting attitude angles (Δα, Δβ, Δγ) are calculated. Each camera requires at least 10 independent measurements, and the average value is taken to eliminate random errors. This step, using the pre-calibrated relative positional relationship between the binocular camera center points and the rigging bottom center point, combined with the rigging bottom center coordinates and attitude obtained in step A, calculates the coordinates of the four binocular camera center points in the three-dimensional coordinate system. The relative positional relationship refers to the spatial offset of the binocular camera's center point relative to the center point of the spreader's bottom, defined as (ΔXc, ΔYc, ΔZc), in meters, obtained through physical measurement. The binocular camera's center point coordinates are defined as (Xc, Yc, Zc), in meters, representing the camera's position in a three-dimensional coordinate system with the RTK device's center as the origin. The calibration process uses a high-precision laser rangefinder to measure the distance from the camera's center point to the center point of the spreader's bottom (by averaging multiple measurements (at least 10 times), achieving millimeter-level accuracy), and records the relative direction (measured using an angle gauge, in degrees). The calculation process first obtains the coordinates (Xh, Yh, Zh) of the bottom center of the spreader and its rotational attitude (α, β, γ) from step A. A rotation matrix R (a 3×3 matrix calculated based on Euler angles) is then constructed. The relative positional relationships (ΔXc, ΔYc, ΔZc) are transformed into a three-dimensional coordinate system using the rotation matrix to obtain the local coordinates of the camera center point. These coordinates are then added to the coordinates of the bottom center of the spreader (Xh, Yh, Zh) to calculate (Xc, Yc, Zc). To ensure accuracy, the calibration data is averaged through multiple measurements during installation (at least 10 times), and floating-point operations (with 6 decimal places retained) are used in the calculation to reduce error accumulation.
[0064] Considering that large steel structures (such as front-end crane booms) will experience thermal deformation under diurnal temperature variations or uneven sunlight, directly affecting the stability of calibration parameters (ΔXc, ΔYc, ΔZc), a dynamic calibration compensation strategy is adopted to achieve robust positioning. Temperature sensor arrays (accuracy ±0.5°C) are installed at key locations on the lifting device (such as the midpoints of the four sides) to monitor the temperature field of the steel structure in real time. Through preliminary experiments, a temperature-deformation compensation model is established (for example, for steel, the coefficient of linear expansion is approximately 1.2 × 10⁻⁶). -5 / °C). This model maps the temperature change ΔT to correction values (δXc, δYc, δZc) for calibration parameters. The data processing unit receives temperature data in real time and performs online compensation on the original (ΔXc, ΔYc, ΔZc) based on the model to control the positioning error caused by thermal deformation, thereby ensuring system accuracy while conforming to engineering practice.
[0065] Step C: Receive the image captured by the binocular camera, identify dangerous targets in the image using the applied detection algorithm, and calculate the first coordinate P1 of the feature points on the dangerous targets in the camera coordinate system. Step C specifically includes the following sub-steps:
[0066] Step C1: When a dangerous target is identified, select one or more pixels near the center point of the dangerous target as feature points.
[0067] This step uses a loaded detection algorithm to identify hazardous targets (such as obstacles, people, or containers) from images captured by a stereo camera, and selects pixels near the center point of the hazardous target as feature points for subsequent localization calculations. A hazardous target is defined as an object that may affect the safe operation of the spreader. A feature point is defined as one or more pixels near the center point of the hazardous target, with coordinates (u, v) on the image plane, in pixels. The detection algorithm uses the YOLOv5 (YouOnlyLookOnce) model based on deep learning, which extracts image features through a convolutional neural network (CNN) and outputs the bounding box (containing the top-left and bottom-right pixel coordinates) and category label (e.g., "personnel") of the hazardous target. The YOLOv5 model is pre-trained on a port scene dataset (containing images of containers, people, vehicles, etc.). The training process uses a publicly available dataset labeled with hazardous targets (such as the COCO dataset) combined with private port data (approximately 5000 images), optimizing the loss function (cross-entropy loss and bounding box regression loss), and training for approximately 100 epochs, achieving an mAP@0.5 accuracy of 0.85. The feature point selection strategy is as follows: 1 to 3 pixels are selected from the center point of the bounding box (calculated as the average of the coordinates of the top-left and bottom-right corners). (If the bounding box is large, multiple points are selected to improve robustness.) The specific number is determined by the area of the bounding box (unit: pixels², threshold set to 10000 pixels²). If the area is greater than the threshold, 3 points are selected (the center point and points offset 5 pixels vertically); otherwise, 1 point is selected. The coordinates (u, v) of the selected pixels are recorded as feature points for subsequent distance calculations.
[0068] Step C2: Calculate the disparity of feature points using the principle of binocular vision and obtain camera parameters.
[0069] This step utilizes the principle of binocular vision to calculate the disparity of feature points in the left and right camera images using a feature point matching algorithm, and obtains pre-calibrated camera parameters to provide data for subsequent coordinate calculations. Disparity (disp) is defined as the difference between the abscissa uL of a feature point in the left camera image and the abscissa uR of the corresponding matching point in the right camera image, i.e., disp = uL - uR (unit: pixels). The disparity is calculated using the SGBM (Semi-Global Block Matching) algorithm, implemented through the OpenCV library. A disparity search range is set (e.g., 64 pixels) to ensure matching accuracy (error less than 0.5 pixels). Camera parameters include: the baseline length of the binocular cameras (B, unit: meters, obtained through calibration measurement), the focal length of the left camera (f, unit: pixels, obtained through camera intrinsic parameter calibration), and the principal point coordinates (cx, cy, unit: pixels). These parameters are all pre-calibrated using a high-precision checkerboard calibration method and stored in a configuration file for later use.
[0070] Step C3: Based on the parallax and camera parameters, directly calculate the coordinate value P1 of the feature point in the camera coordinate system.
[0071] This step, based on the disparity (disp) calculated in step C2 and the acquired camera parameters, uses the direct linear transformation method to calculate the 3D coordinates P1 (x1, y1, z1) of the feature point in the camera coordinate system, in meters. The camera coordinate system has its origin at the optical center of the left camera, with the Z-axis pointing forward along the optical axis, the X-axis pointing horizontally to the right, and the Y-axis pointing vertically downward. Coordinate P1 is calculated directly using the following formula:
[0072] z1=(fB) / disp;
[0073] x1 = z1(u-cx) / f;
[0074] y1=z1(v-cy) / f;
[0075] Where (u, v) are the pixel coordinates of the feature point on the left camera image (in pixels), f is the focal length of the left camera (in pixels), B is the baseline length (in meters), (cx, cy) are the principal point coordinates (in pixels), and disp is the disparity (in pixels). Calculations use floating-point operations (with 6 decimal places) to reduce errors. To improve robustness, if multiple feature points are selected in step C1, the P1 coordinates of each feature point are calculated, and the average is taken as the final P1 (if there are 3 feature points, the average is x1, y1, z1). Camera intrinsic parameters (focal length f, principal point coordinates (cx, cy)) are used to verify the consistency of the calculations, ensuring that the P1 coordinates are consistent with the image plane projection (through backprojection error checks, with an error threshold of 0.1 pixels). The calculation process uses the open-source mathematical library NumPy to ensure efficiency and accuracy.
[0076] Step D: Based on the transformation relationship between the camera coordinate system and the three-dimensional spatial coordinate system, the first coordinate P1 is converted into the second coordinate P2 of the feature point in the three-dimensional spatial coordinate system; the transformation relationship between the camera coordinate system and the three-dimensional spatial coordinate system is determined by the coordinates and rotation attitude of the camera center point of each binocular camera in the three-dimensional spatial coordinate system.
[0077] This step transforms the feature point coordinates P1 (x1, y1, z1) calculated in step C3 into coordinates P2 (X2, Y2, Z2) in the 3D spatial coordinate system using the transformation relationship between the camera coordinate system and the 3D spatial coordinate system, in meters. The transformation relationship is determined by the camera center point coordinates (Xc, Yc, Zc) obtained in step B and the camera rotation attitude (based on the hoist attitude (α, β, γ) and the camera mounting angle). The camera rotation attitude is calculated by combining the hoist attitude (α, β, γ) with the fixed mounting angle of the camera relative to the hoist (measured using an angle gauge, in degrees), constructing a rotation matrix Rc (3×3 matrix). The transformation process first transforms P1 to local coordinates in the 3D spatial coordinate system using the rotation matrix Rc, with the formula P1' = Rc × P1, where P1' is the local coordinate (x1', y1', z1'). Then, P2 is calculated through translation transformation: X2 = Xc + x1', Y2 = Yc + y1', Z2 = Zc + z1'. The camera mounting angles are obtained through calibration (accuracy 0.1°), for example, the pitch angle ψ and yaw angle ω of camera 1 relative to the rigging (unit: degrees). The rotation matrix Rc is calculated based on Euler angles (α+ψ, β+ω, γ) and implemented using the Eigen library. To ensure accuracy, the conversion process uses double-precision floating-point operations, with errors controlled within 0.001 meters.
[0078] Step E: Output the spatial location information of the dangerous target based on the second coordinate P2.
[0079] This step outputs the spatial location information of the hazardous target based on the feature point coordinates P2 (X2, Y2, Z2) calculated in step D, which is used for obstacle avoidance or operational decisions of the frontal crane. The spatial location information is defined as the position of the hazardous target in the three-dimensional coordinate system (X2, Y2, Z2, unit: meters), its relative distance (D, unit: meters) to the center of the spreader bottom, and its azimuth angle (θh, φh, unit: degrees). The relative distance D is calculated using the Euclidean distance formula: D = sqrt((X2-Xh)² + (Y2-Yh)² + (Z2-Zh)²), where (Xh, Yh, Zh) are the coordinates of the spreader bottom center from step A. The azimuth angle θh = atan2(Y2-Yh, X2-Xh), φh = atan((Z2-Zh) / sqrt((X2-Xh)² + (Y2-Yh)²)), implemented using the NumPy library. The output information is transmitted to the control system of the reach stacker via PLC. The format is structured data containing coordinates, distance, and orientation (encoded as a binary stream, with units uniformly in meters and degrees). To ensure real-time performance, the calculation time is controlled within 10ms, and the output frequency is 10Hz. If multiple feature points exist, the average value of each P2 is calculated as the hazard marker location to reduce noise impact. The output also includes a confidence score (based on YOLOv5 detection confidence, ranging from 0 to 1), with scores below 0.8 indicating low reliability.
[0080] This technical solution integrates high-precision RTK positioning equipment, surround-type binocular cameras, and a data processing unit to construct a three-dimensional perception and positioning solution for reach stackers. The system's workflow is as follows: First, an RTK device at the front of the vehicle establishes a centimeter-level precision three-dimensional spatial coordinate system as a global positioning reference. Then, four binocular cameras installed at the midpoints of the four sides of the spreader simultaneously capture images of the surrounding environment from different perspectives, forming a comprehensive detection coverage. The data processing unit, as the core, acquires RTK and vehicle attitude data via a PLC to calculate the precise coordinates and attitude of the spreader's bottom center; simultaneously, it runs a detection algorithm to process the binocular images, identify hazardous targets, and calculate their depth information using binocular vision principles. Finally, through a series of coordinate transformations, the identified hazardous targets are located in a unified global three-dimensional coordinate system. This method fundamentally solves the problems of insufficient positioning accuracy and limited environmental perception in traditional solutions in dynamic and complex environments, achieving synchronous and accurate positioning of the spreader and hazardous targets, and providing key technical support for the automated and safe operation of reach stackers in scenarios such as ports and logistics parks.
[0081] In some embodiments, prior to the start of the operation, the data processing unit is further configured to perform a camera calibration operation to determine the intrinsic and extrinsic parameters of the left and right cameras of each of the stereo cameras.
[0082] Before the operation begins, the data processing unit performs camera calibration to determine the intrinsic parameters (such as focal length, principal point position, and distortion coefficient) and extrinsic parameters (such as the relative position and attitude between the cameras) of the left and right cameras of each binocular camera.
[0083] In some embodiments, before the operation begins, the data processing unit is further configured to receive and load configuration information for personnel, vehicle, and foreign object detection algorithms.
[0084] Before the operation begins, the data processing unit receives and loads the detection algorithm configuration information for personnel, vehicles and foreign objects to optimize the identification performance of dangerous targets.
[0085] In some embodiments, the object detection and positioning system based on the vision of the lifting device further includes an alarm unit and / or a visualization unit, used for alarm processing and / or visualization display based on the output spatial location information of the hazardous target.
[0086] The object detection and positioning system based on spreader vision uses alarm units and / or visualization units to process alarms and / or visualize the spatial location information of hazardous targets. The alarm unit triggers audible and visual alarms or sends alarm signals to the reach stack control system based on the distance between the hazardous target and the spreader, and the threat level, alerting the operator or automatically adjusting the operation status. The visualization unit presents the spatial location information of the hazardous target in the form of 3D graphics or 2D projection on the display interface, intuitively showing the position and trajectory of the hazardous target relative to the spreader, assisting the operator in making safety decisions.
[0087] In some embodiments, the object detection and positioning system based on spreader vision further includes a dynamic obstacle avoidance optimization module, used to optimize the spreader's movement trajectory in real time during reach stacking operations to address potential collision risks caused by the dynamic movement of hazardous targets. The dynamic obstacle avoidance optimization module is configured to perform the following operations:
[0088] The data processing unit obtains the second coordinate P2 of the hazardous target in the three-dimensional spatial coordinate system, as well as the real-time coordinates and rotation attitude of the bottom center of the lifting device in the three-dimensional spatial coordinate system.
[0089] The dynamic obstacle avoidance optimization module is based on real-time data calculated by the data processing unit, including: the coordinates P2 (X2, Y2, Z2) of the hazardous target in the three-dimensional spatial coordinate system, and the coordinates (Xh, Yh, Zh) and rotational attitude (α, β, γ) of the center of the spreader bottom. These data constitute the input for trajectory planning and collision risk prediction. The second coordinate P2 of the hazardous target represents the position of the hazardous target (such as obstacles, personnel, or containers) in the three-dimensional spatial coordinate system with the center of the RTK equipment as the origin, and is calculated by the data processing unit through binocular camera image processing and coordinate transformation (refer to step D). The real-time coordinates (Xh, Yh, Zh) of the center of the spreader bottom are calculated by converting the high-precision latitude and longitude data (accuracy to the centimeter level, formatted as longitude, latitude, and altitude in the WGS84 coordinate system) provided by the RTK positioning equipment into local Cartesian coordinates through a Gaussian projection algorithm, and then combined with the physical offset between the RTK equipment and the spreader (measured through mechanical design drawings, accuracy 0.1 mm). The rotational attitude (α, β, γ) is acquired by the PLC from the attitude sensor (such as a gyroscope, outputting angular velocity in degrees / second) suspended at the front. This angular velocity data is combined with the initial calibration angle (measured with a level, accuracy 0.1 degrees) and calculated using time integration. A Kalman filter algorithm is then used to fuse the RTK data and attitude data, filtering out environmental noise (such as vibration or electromagnetic interference) to ensure the accuracy of coordinates and attitude (error controlled within 0.01 meters and 0.1 degrees, respectively). The data processing unit acquires this data in real-time at a frequency of 10Hz via the PLC (Programmable Logic Controller), stores it as a structured data stream (containing timestamps, coordinates, and attitude), and transmits it to the dynamic obstacle avoidance optimization module through a memory cache (at least 1MB) to ensure data real-time performance and consistency.
[0090] Based on the second coordinate P2 sequence of historically collected hazardous targets, the velocity vector V and acceleration vector A of the hazardous targets are calculated, where V represents the displacement rate of the hazardous target in the three-dimensional spatial coordinate system, in meters per second, reflecting the direction and magnitude of the hazardous target's movement; A represents the acceleration of the hazardous target, in meters per square second, reflecting the change in the hazardous target's motion trend.
[0091] This step utilizes the historical second coordinate P2 sequence of hazard targets stored in the data processing unit (with a time interval of 0.1 seconds, storing at least the most recent 10 seconds of data, approximately 100 sets of coordinates) to calculate the velocity vector V (unit: m / s) and acceleration vector A (unit: m / s²) of the hazard target, to characterize the dynamic behavior of the hazard target. The velocity vector V is defined as the displacement rate of the hazard target in the three-dimensional coordinate system, calculated as V = (ΔX² / Δt, ΔY² / Δt, ΔZ² / Δt), where ΔX², ΔY², and ΔZ² are the coordinate differences between adjacent time steps, and Δt is the time interval (0.1 seconds). The acceleration vector A is defined as the rate of change of the velocity vector, calculated as A = (ΔVx / Δt, ΔVy / Δt, ΔVz / Δt), where ΔVx, ΔVy, and ΔVz are the velocity differences between adjacent time steps. The historical P2 sequence is obtained from the memory cache of the data processing unit (stored in timestamp and coordinate pairs) and the velocity and acceleration are calculated using the finite difference method. To improve accuracy, a five-point difference method was used (based on two sets of data, smoothing out noise effects, with errors controlled within 0.01 m / s and 0.01 m / s²), and a low-pass filter (cutoff frequency 10 Hz, implemented using the SciPy library) was applied to remove high-frequency noise (such as camera shake). The calculation results were stored in vector form (V=(Vx, Vy, Vz), A=(Ax, Ay, Az)) with an added confidence level (based on the variance of the coordinate data, ranging from 0 to 1) for use in subsequent steps.
[0092] Based on the real-time coordinates and rotational attitude of the bottom center of the spreader, the expected motion trajectory T of the spreader is calculated. The expected motion trajectory T consists of a set of continuous three-dimensional spatial coordinate points, representing the predicted position of the spreader in the future time period.
[0093] This step calculates the expected motion trajectory T based on the real-time coordinates (Xh, Yh, Zh) and rotational attitude (α, β, γ) of the spreader's bottom center, combined with the motion control parameters of the reach stack. This trajectory represents the predicted position sequence of the spreader within a future time period (e.g., 5 seconds). The expected motion trajectory T is defined as a set of continuous three-dimensional spatial coordinate points (Xt, Yt, Zt, unit: meters), with a time step of 0.1 seconds, determined by the reach stack's current speed (Vs, unit: meters per second), acceleration (As, unit: meters per second²), and control commands (acquired via PLC). The speed Vs and acceleration As are read from the PLC from the reach stack's drive system data (e.g., motor speed and acceleration, with an accuracy of 0.01 meters per second and 0.01 meters per second²), and calculated using a kinematic model: Xt = Xh + Vsx × t + 0.5 × Asx × t² (similar to the Y and Z directions). Rotational attitude (α, β, γ) is used to correct the trajectory direction. A rotation matrix R (3×3 matrix, based on Euler angles) is constructed to transform the velocity and acceleration vectors to a three-dimensional coordinate system. Motion control parameters include maximum velocity (Vmax, e.g., 5 m / s) and maximum acceleration (Amax, e.g., 1 m / s²), obtained from a mechanical design manual. To ensure accuracy, the trajectory calculation uses the Runge-Kutta 4th order method (implemented via SciPy.integrate) to solve the kinematic equations, with the error controlled within 0.01 meters. The output trajectory point sequence is stored as an array (approximately 50 points).
[0094] Using the velocity vector V and acceleration vector A of the hazardous target, and the expected trajectory T of the lifting device, the dynamic obstacle avoidance optimization index D is calculated using the following formula:
[0095] D=[k1×||V||² / V0+k2×||A|| / A0+k3×min(d(T, P2)) / d0] / [1+exp(-Q)];
[0096] Wherein, D is the dynamic obstacle avoidance optimization index, ranging from 0 to 1. A higher value indicates a higher collision risk, requiring priority optimization of the spreader's trajectory; ||V||² is the square norm of the velocity vector of the hazardous target, measured in square meters per second squared, reflecting the kinetic energy of the hazardous target; ||A|| is the norm of the acceleration vector of the hazardous target, measured in meters per second squared, reflecting the drastic change in the hazardous target's motion; min(d(T, P2)) represents the minimum Euclidean distance between the expected trajectory T of the spreader and the second coordinate P2 of the hazardous target, measured in meters, reflecting the closest spatial proximity between the spreader and the hazardous target; k1, k2, and k3 are weighting coefficients, corresponding to the weighted influence of velocity, acceleration, and minimum distance, respectively. The range is 0 to 1, and k1+k2+k3=1. It is determined by training a machine learning model based on historical operation data. Q is the environmental complexity factor, which represents the density and distribution complexity of dangerous targets in the operation environment. It ranges from 0 to 10. The higher the value, the more complex the environment. It is calculated by the data processing unit based on the number and distribution uniformity of dangerous targets identified in the binocular camera images. exp(-Q) is an exponential function used to adjust the nonlinear effect of the environmental complexity factor on the dynamic obstacle avoidance optimization index, ensuring that the optimization index is more sensitive in complex environments. V0 is the square norm reference value, in square meters per square second. A0 is the norm reference value, in meters per square second. d0 is the Euclidean distance reference value, in meters.
[0097] This step uses the velocity vector V and acceleration vector A of the hazardous target and the expected trajectory T of the spreader to calculate the dynamic obstacle avoidance optimization index D (range 0-1, higher values indicate higher collision risk), which is used to assess the potential collision risk between the spreader and the hazardous target. The formula for D is D=[k1×||V||² / V0+k2×||A|| / A0+k3×min(d(T,P2)) / d0] / [1+exp(-Q)], where ||V||² is the square norm of the velocity vector (unit: square meters / second², reflecting kinetic energy, calculated by NumPy.linalg.norm), ||A|| is the norm of the acceleration vector (unit: meters / second², reflecting the degree of motion change), and min(d(T,P2)) is the minimum Euclidean distance between trajectory T and P2 (unit: meters, calculated by point-by-point comparison). The weighting coefficients k1, k2, and k3 (range 0-1, satisfying k1+k2+k3=1) are determined by a machine learning model (random forest, trained based on historical operation data, containing at least 10,000 collision scenarios, features being velocity, acceleration, and distance, mAP accuracy 0.9). The environmental complexity factor Q (range 0-10) is obtained by the data processing unit based on the number of hazard markers in the binocular camera image (detected via YOLOv5) and the uniformity of their distribution (calculated via image entropy, range 0-10), with the formula Q=0.6×N+0.4×E, where N is the number of hazard markers and E is the image entropy (calculated using OpenCV). exp(-Q) is an exponential function (implemented using NumPy.exp) that adjusts the nonlinear effect of Q. D is calculated using floating-point operations (6 decimal places precision) and stored as a single value. The square norm reference value V0, the norm reference value A0, and the Euclidean distance reference value d0 can be easily preset by the technical team according to actual needs, historical operation data, or domain knowledge.
[0098] As an exemplary implementation, the reference values V0, A0, and d0 can be obtained by analyzing historical safe operation data: V0 can be taken as the 85th percentile of the square norm of the velocity of the dangerous target in the historical data, A0 can be taken as the 90th percentile of the norm of the acceleration of the dangerous target in the historical data, and d0 can be taken as 1.2 times the minimum safe distance between the lifting equipment and the dangerous target in the historical data.
[0099] Based on the calculated dynamic obstacle avoidance optimization index D, and by referring to the pre-set mapping relationship between the optimization index and the trajectory adjustment strategy, the optimized motion trajectory T' of the spreader is determined. The optimized motion trajectory T' is the movement path of the spreader to avoid dangerous targets by adjusting its moving speed, direction, or height.
[0100] This step determines the optimized motion trajectory T' based on the dynamic obstacle avoidance optimization index D (range 0-1), referring to the pre-set mapping relationship between the optimization index and trajectory adjustment strategy (stored as a lookup table containing the D value range and adjustment strategy, such as speed reduction, direction shift, or height increase), to avoid the movement path of dangerous targets. The optimized motion trajectory T' is defined as a set of adjusted three-dimensional spatial coordinate points (Xt', Yt', Zt', unit: meters), velocity vector (Vs', unit: meters / second), and attitude angles (α', β', γ', unit: degrees), generated by adjusting the spreading equipment's moving speed (range 0-5 meters / second), direction (yaw angle change ±30 degrees), or height (lift 0-2 meters). The mapping relationship is constructed by training a support vector machine (SVM) model (based on scikit-learn, kernel function RBF, C=1.0, training set mAP=0.85) using historical operation data (10,000 records). The input is the D value, and the output is the adjustment strategy (e.g., 50% speed reduction when D>0.7). T' was recalculated using a kinematic model (Runge-Kutta 4th order method, implemented in SciPy) to ensure that the minimum Euclidean distance min(d(T', P2)) is greater than the safe distance threshold Dsafe (0.5-2 meters). The calculation process used double-precision floating-point operations, with the error controlled within 0.01 meters.
[0101] The optimized motion trajectory T' is transmitted to the PLC control system of the front-end crane to adjust the motion state of the spreader in real time, thereby effectively avoiding collisions with dynamically moving dangerous targets.
[0102] This step optimizes the motion trajectory T' (including coordinates, velocity vectors, and attitude angles) as a structured data stream (binary format, including timestamps, coordinates, velocity, and attitude) and transmits it to the reach stack's control system via PLC. This allows for real-time adjustment of the spreader's motion state to prevent collisions with hazard markers. The PLC (supporting Modbus protocol, 100Mbps transmission rate) parses T' into control commands (e.g., motor speed, direction angle) to drive the servo motors (accuracy 0.01 m / s) and hydraulic system (accuracy 0.1 m) to execute the trajectory. Transmission delay is controlled within 5ms, and the output frequency is 10Hz. The control system adjusts the motor output using a PID controller (parameters Kp=0.5, Ki=0.1, Kd=0.05, based on the PLC's built-in algorithm) to ensure the trajectory execution error is less than 0.01 meters. To prevent command conflicts, the PLC uses a priority queue (based on timestamps) to manage the T' data, prioritizing the execution of the most recent trajectory.
[0103] To address the dynamic movement of hazardous targets in the operational environment, the system integrates a dynamic obstacle avoidance optimization module. This module continuously receives real-time position and attitude data of the spreader and hazardous targets, and estimates the target's velocity V and acceleration A based on its historical coordinate sequence. Simultaneously, the module predicts the spreader's expected trajectory T based on its current state. By comprehensively calculating a dynamic obstacle avoidance optimization index D that includes kinetic energy, acceleration, and minimum distance, the module can assess potential collision risks. Subsequently, based on a preset optimization strategy, the module adjusts the spreader's speed, direction, or height to generate an optimized trajectory T', which is then transmitted to the control system via PLC for execution, thereby achieving active obstacle avoidance.
[0104] In some embodiments, in the object detection and positioning system based on the vision of the lifting device, step S5, determining the optimized motion trajectory T' of the lifting device according to the calculated dynamic obstacle avoidance optimization index D and the mapping relationship between the pre-set optimization index and the trajectory adjustment strategy, includes:
[0105] During the front-end crane operation, the data processing unit collects dynamic environmental change data in real time, and combines it with a quantum random number generator to generate a high-entropy random seed, which is used to initialize the trajectory planning state machine of the dynamic obstacle avoidance optimization module to solve the local convergence problem of trajectory optimization caused by dynamic environmental changes.
[0106] This step involves collecting dynamic environmental change data (including the number and location of hazard markers and image entropy, updated at a frequency of 10Hz) through a data processing unit. A high-entropy random seed (256 bits, entropy > 7.8 bits / byte) is then generated using a quantum random number generator to initialize the trajectory planning state machine of the dynamic obstacle avoidance optimization module, preventing local convergence issues. The dynamic environmental change data includes the number of hazard markers N (detected via YOLOv5 with an accuracy of 0.85) and image entropy E (calculated via OpenCV, range 0-10), acquired from binocular camera images. A quantum random number generator (based on photon polarization measurement, such as the IDQuantiqueQuantis model, with an output rate of 1Mbps) generates a random seed and initializes the state machine (a finite state machine, with states including initial, search, and optimization, and transition probabilities based on the random seed). The state machine sets the initial search point using the random seed, increasing the globality of trajectory optimization and preventing getting trapped in local optima (e.g., repeatedly selecting invalid paths). The seed generation and state machine initialization time are controlled within 1ms, and the data is stored in a memory cache (Redis, 1MB capacity).
[0107] Based on the dynamic obstacle avoidance optimization index D, a multi-dimensional decision space is constructed, which includes the real-time coordinates of the lifting device, its rotational attitude, the velocity vector V and acceleration vector A of the dangerous target, and the environmental complexity factor Q.
[0108] This step constructs a multi-dimensional decision space based on the dynamic obstacle avoidance optimization index D (range 0-1), including the real-time coordinates of the lifting device (Xh, Yh, Zh), rotational attitude (α, β, γ), the velocity vector V of the hazard marker, the acceleration vector A, and the environmental complexity factor Q, for subsequent trajectory optimization. The multi-dimensional decision space is defined as a 6-dimensional vector (Xh, Yh, Zh, α, β, γ, Vx, Vy, Vz, Ax, Ay, Az, Q), processed by normalization (Min-Max normalization, range 0-1), and stored as an array. Data is obtained from steps S1, S2, and S4, and the normalization parameters (e.g., coordinate range 0-500 meters, attitude range -180 to 180 degrees) are preset according to the work scenario. The construction process uses NumPy array operations, with the time controlled within 0.5ms and the error controlled within 0.001. By performing high-dimensional interpolation on the multidimensional decision space, a set of candidate optimized motion trajectories T'i are generated. Each candidate trajectory T'i consists of a set of adjusted three-dimensional spatial coordinate points, velocity vectors, and attitude angles, satisfying that the minimum Euclidean distance min(d(T'i, P2)) is greater than the safe distance threshold Dsafe. Dsafe is dynamically adjusted according to the dynamic hazard density of the work scenario, ranging from 0.5 meters to 2 meters.
[0109] This step generates a set of candidate optimized motion trajectories T'i by performing high-dimensional interpolation on the multi-dimensional decision space (using trilinear interpolation, implemented using SciPy.interpolate). Each T'i contains adjusted coordinate points (Xt'i, Yt'i, Zt'i), a velocity vector, and an attitude angle, satisfying min(d(T'i, P2)) > Dsafe (0.5-2 meters, dynamically adjusted based on the hazard density). The interpolation is based on a 6-dimensional decision space, generating 100 candidate trajectories (50 points each), adjusted by speed (0-5 m / s), direction (±30 degrees), or height (0-2 meters). Dsafe is linearly mapped by Q (Dsafe=2 meters when Q=0, Dsafe=0.5 meters when Q=10). The computation time is controlled within 10 ms, with an error of 0.01 meters.
[0110] Using a high-entropy random seed generated by a quantum random number generator, Monte Carlo simulations are performed on the candidate optimized motion trajectories T'i to calculate the collision probability Pcoll of each candidate trajectory T'i in a future time period. The collision probability Pcoll is calculated using the following formula:
[0111] Pcoll=Σ[wi×(1-exp(-α×||T'i-P2|| / σ))];
[0112] Where Pcoll represents the collision probability of candidate trajectory T'i, ranging from 0 to 1; wi is the weight of the i-th candidate trajectory, determined by a combination of trajectory smoothness and energy consumption, ranging from 0 to 1; ||T'i-P2|| represents the Euclidean distance between candidate trajectory T'i and the second coordinate P2 of the dangerous target; α is the collision sensitivity coefficient, dynamically adjusted according to the magnitude of the dangerous target's velocity vector V, ranging from 0.1 to 1; σ is the environmental uncertainty factor, calculated by regression analysis of the environmental complexity factor Q and historical collision data, in meters.
[0113] This step utilizes a high-entropy random seed (256 bits, generated by IDQuantiqueQuantis) generated by a quantum random number generator to perform Monte Carlo simulations on candidate trajectories T'i, calculating the collision probability Pcoll (range 0-1). The formula is Pcoll=Σ[wi×(1-exp(-α×||T'i-P2|| / σ))], where wi is the trajectory weight (based on smoothness and energy consumption, range 0-1, calculated by an SVM model), α is the collision sensitivity coefficient (0.1-1, based on a ||V|| linear mapping), and σ is the environmental uncertainty factor (unit: meters, calculated by regression from Q and historical collision data, implemented using scikit-learn). The Monte Carlo simulation is run 1000 times, generating perturbations based on the random seed, calculating Pcoll for each T'i within 20ms.
[0114] Candidate optimized motion trajectories T'i with collision probabilities Pcoll lower than a preset threshold Pth are selected, and the trajectory with the lowest dynamic obstacle avoidance optimization index D is chosen as the final optimized motion trajectory T'.
[0115] This step filters candidate trajectories T'i whose Pcoll is lower than the threshold Pth (set to 0.3), and selects the trajectory with the lowest dynamic obstacle avoidance optimization index D as T'. Pth is determined statistically from historical operation data (10,000 records) (average collision probability + 1 standard deviation). The filtering is performed using NumPy array operations, comparing Pcoll and D, with the time controlled within 1ms and the error 0.001.
[0116] The dynamic obstacle avoidance optimization module addresses the local convergence problem caused by dynamic environmental changes by optimizing trajectory planning through a multi-dimensional decision space and a quantum random number generator. Specifically, the system initializes the trajectory planning state machine using real-time environmental data and a high-entropy random seed, constructing a multi-dimensional decision space that includes the spreader's coordinates, attitude, and hazard marker motion information. Candidate optimized trajectories T'i are generated through high-dimensional interpolation, and Monte Carlo simulation is used to calculate the collision probability Pcoll for each trajectory. Finally, the system selects the trajectory with a collision probability below a threshold and the lowest dynamic obstacle avoidance optimization index D as the final optimized trajectory T', ensuring that the spreader completes safe operations in complex environments with minimal collision risk and optimal energy efficiency, providing reliable support for the intelligent upgrade of scenarios such as ports and logistics parks.
[0117] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. An object detection and positioning system based on lifting device vision, applied to front-facing lifting devices, characterized in that, include: The RTK positioning device is installed above the front crane head to obtain the latitude and longitude of the front crane and establish a three-dimensional spatial coordinate system with its center as the origin. Multiple binocular cameras are mounted around the hoist to capture real-time images of the surrounding environment. The data processing unit is connected to the RTK positioning device, the multiple binocular cameras, and the front-end crane via a PLC. The data processing unit is configured as follows: Based on the data obtained from the PLC, the coordinates and rotational attitude of the bottom center of the lifting device in the three-dimensional spatial coordinate system are calculated. Based on the pre-defined relative positional relationship between the camera center point of each binocular camera and the center point of the rigging, the coordinates of the camera center point of each binocular camera in the three-dimensional spatial coordinate system are calculated. The system receives images captured by the binocular camera, identifies dangerous targets in the images using a loaded detection algorithm, and calculates the first coordinates P1 of the feature points on the dangerous targets in the camera coordinate system. Based on the transformation relationship between the camera coordinate system and the three-dimensional spatial coordinate system, the first coordinate P1 is converted into the second coordinate P2 of the feature point in the three-dimensional spatial coordinate system; Based on the second coordinate P2, output the spatial location information of the dangerous target; The dynamic obstacle avoidance optimization module is used to optimize the movement trajectory of the spreader in real time during reach stacking operations. The dynamic obstacle avoidance optimization module is configured to perform the following operations: The data processing unit obtains the second coordinate P2 of the hazardous target in the three-dimensional spatial coordinate system, as well as the real-time coordinates and rotation attitude of the bottom center of the lifting device in the three-dimensional spatial coordinate system. Based on the historically collected second coordinate P2 sequence of dangerous targets, calculate the velocity vector V and acceleration vector A of the dangerous targets; Based on the real-time coordinates and rotational attitude of the bottom center of the spreader, the expected motion trajectory T of the spreader is calculated; Using the velocity vector V and acceleration vector A of the hazardous target, and the expected trajectory T of the lifting device, the dynamic obstacle avoidance optimization index D is calculated using the following formula: D=[k1×||V||² / V0+k2×||A|| / A0+k3×min(d(T, P2)) / d0] / [1+exp(-Q)]; Where D is the dynamic obstacle avoidance optimization index; ||V||² is the square norm of the velocity vector of the dangerous target; ||A|| is the norm of the acceleration vector of the dangerous target; min(d(T, P2)) represents the minimum Euclidean distance between the expected trajectory T of the spreader and the second coordinate P2 of the dangerous target; k1, k2, and k3 are weighting coefficients; Q is the environmental complexity factor; exp(-Q) is the exponential function; V0 is the square norm reference value, A0 is the norm reference value, and d0 is the Euclidean distance reference value. Based on the calculated dynamic obstacle avoidance optimization index D, and by comparing the pre-set mapping relationship between the optimization index and the trajectory adjustment strategy, the optimized motion trajectory T' of the spreader is determined. The optimized motion trajectory T' is transmitted to the PLC control system of the front-end crane to adjust the motion state of the spreader in real time, thereby effectively avoiding collisions with dynamically moving dangerous targets.
2. The object detection and positioning system based on lifting device vision as described in claim 1, characterized in that, Before the operation begins, the data processing unit is further configured to perform camera calibration operations to determine the intrinsic and extrinsic parameters of the left and right cameras of each of the binocular cameras.
3. The object detection and positioning system based on lifting device vision as described in claim 1, characterized in that, Before the operation begins, the data processing unit is further configured to receive and load configuration information for personnel, vehicle and foreign object detection algorithms.
4. The object detection and positioning system based on lifting device vision as described in claim 1, characterized in that, The relative positional relationship is obtained by physically measuring the relative distance between the center point of the camera and the center point of the hoist.
5. The object detection and positioning system based on lifting device vision as described in claim 1, characterized in that, Calculating the first coordinate P1 of the feature point on the dangerous target in the camera coordinate system includes: When a dangerous target is identified, one or more pixels near the center point of the dangerous target are selected as feature points. The disparity of feature points is calculated using the principle of binocular vision, and camera parameters are obtained. Based on the parallax and camera parameters, the coordinate value P1 of the feature point in the camera coordinate system is directly calculated.
6. The object detection and positioning system based on lifting device vision as described in claim 1, characterized in that, The transformation relationship between the camera coordinate system and the three-dimensional spatial coordinate system is determined by the coordinates and rotation attitude of the camera center point of each binocular camera in the three-dimensional spatial coordinate system.
7. The object detection and positioning system based on lifting device vision as described in claim 1, characterized in that, It also includes an alarm unit and / or a visualization unit, used for alarm processing and / or visualization display based on the output spatial location information of the dangerous target.
8. The object detection and positioning system based on lifting device vision as described in claim 1, characterized in that, The binocular cameras consist of four units, each installed at the midpoint of one of the four sides of the rigging, so that the shooting range covers the perimeter of the rigging to form a detection area without blind spots.
9. The object detection and positioning system based on lifting device vision as described in claim 1, characterized in that, The step of determining the optimized motion trajectory T' of the spreader based on the calculated dynamic obstacle avoidance optimization index D and the pre-set mapping relationship between the optimization index and the trajectory adjustment strategy includes: During the front-end crane operation, the data processing unit collects dynamic environmental change data in real time and combines it with a quantum random number generator to generate a high-entropy random seed; Based on the dynamic obstacle avoidance optimization index D, a multi-dimensional decision space is constructed, which includes the real-time coordinates of the lifting device, its rotational attitude, the velocity vector V and acceleration vector A of the dangerous target, and the environmental complexity factor Q. By performing high-dimensional interpolation on the multidimensional decision space, a set of candidate optimized motion trajectories T'i are generated, wherein each candidate optimized motion trajectory T'i consists of a set of adjusted three-dimensional spatial coordinate points, velocity vectors and attitude angles, satisfying that the minimum Euclidean distance min(d(T'i, P2)) is greater than the safe distance threshold Dsafe; Using a high-entropy random seed generated by a quantum random number generator, Monte Carlo simulations are performed on the candidate optimized motion trajectories T'i to calculate the collision probability Pcoll of each candidate optimized motion trajectory T'i in a future time period. The collision probability Pcoll is calculated using the following formula: Pcoll=Σ[wi×(1-exp(-α×||T'i-P2|| / σ))]; Where Pcoll represents the collision probability of the candidate optimized trajectory T'i; wi is the weight of the i-th candidate optimized trajectory; ||T'i-P2|| represents the Euclidean distance between the candidate optimized trajectory T'i and the second coordinate P2 of the dangerous target; α is the collision sensitivity coefficient; and σ is the environmental uncertainty factor. Candidate optimized motion trajectories T'i with collision probabilities Pcoll lower than a preset threshold Pth are selected, and the trajectory with the lowest dynamic obstacle avoidance optimization index D is chosen as the final optimized motion trajectory T'.
Citation Information
Patent Citations
Crane obstacle-avoidance system based on stereoscopic vision
CN102175222A
Bridge crane hoisting safety anti-collision system and method based on dynamic binocular vision
CN112418103A