Portal crane grab bucket control system and method based on multi-sensor fusion
By laying multiple types of sensors and image acquisition equipment in the door machine operation area, combined with the hollow convolution and feature pyramid alignment mechanism in the YOLOv3-tiny structure, the problem of feature alignment error in grab detection is solved, accurate detection of grabbing and bounding box regression is realized, recognition accuracy and system stability are improved, and it is suitable for complex operation scenarios, significantly improving the intelligence level and efficiency of door machine operations.
Patent Information
- Application Number
- CN202510496382.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
In the existing grab detection based on improved YOLOv3-tiny network, the introduction of hollow convolution causes feature alignment errors, affecting the positioning accuracy of the grab edge, especially when grab recognition with small sizes or special postures, the recognition rate may decrease, and may cause boundary overlap and positioning offset in multi-target scenarios, affecting detection accuracy and system stability.
Image acquisition equipment and multi-type sensors are arranged in the door machine operation area to obtain the video images, attitude information, position coordinates and rotation angle data of the grab in real time, perform enhancement and normalization processing, and synchronously fuse with the sensor data under a unified time stamp to build a multi-dimensional input data set. A hollow convolution layer is introduced in the original YOLOv3-tiny structure, and spatial alignment between features of different scales is achieved based on the feature pyramid alignment mechanism to generate grab position and bounding box information. Use the attention mechanism to continuously track and abnormal judgment of the grab state, and automatically generate accurate grab paths and posture adjustment instructions to control grab actions in real time. According to the capture results and the pose deviation of the sensor back-pass, the control model parameters are dynamically adjusted and the detection accuracy and control strategy are continuously optimized.
Accurate detection and bounding box regression of grabs of different scales are achieved, and the accuracy of grab recognition and the stability of system operation are improved. It is especially suitable for complex or dynamically changing operation scenarios, such as port loading and unloading or bulk material handling, which significantly improves the intelligence level and operation efficiency of door machine operations.
Smart Images

Figure CN120004149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gantry crane grab bucket control, and in particular to a gantry crane grab bucket control system and method based on multi-sensor fusion. Background Art
[0002] Multi-sensor fusion-based gantry crane grab control refers to the comprehensive analysis and processing of the grab's operating status, material position and environmental information by fusing data collected by various types of sensors (such as position sensors, force sensors, image sensors, acceleration sensors, etc.) in the gantry crane (gantry crane) grab control system, thereby achieving precise control and optimization of the grab's action. This method can improve the grabbing accuracy, operating efficiency and the intelligence level of the system, and is particularly suitable for complex or dynamically changing operating scenarios, such as port loading and unloading or bulk material handling.
[0003] The prior art has the following deficiencies: In the grab detection based on the improved YOLOv3-tiny network, the introduction of dilated convolution can enhance high-level semantic information, but it is easy to cause feature alignment errors. Since the feature map of YOLOv3-tiny itself is relatively shallow, while dilated convolution expands the receptive field, it causes spatial coordinate deviations between high-level features and low-level details in the multi-scale fusion process, resulting in inaccurate positioning of the grab edge. In particular, the recognition rate drops significantly when dealing with grabs of small size or special postures (such as the open state), and it may cause boundary overlap and positioning offset in multi-target scenes, affecting detection accuracy and system stability. Summary of the invention
[0004] The purpose of the present invention is to provide a gantry crane grab bucket control system and method based on multi-sensor fusion to solve the shortcomings of the background technology.
[0005] In order to achieve the above object, the present invention provides the following technical solution: a gantry crane grab control method based on multi-sensor fusion, comprising: Image acquisition equipment and multiple types of sensors are deployed in the gantry crane operation area to obtain the grab bucket's video images, posture information, position coordinates, and rotation angle data in real time; Enhance and normalize the acquired video images, and synchronize them with the sensor data at a unified timestamp to construct a multi-dimensional input data set; A hole convolution layer is introduced into the original YOLOv3-tiny structure. Before the high-level and low-level features are fused, the spatial alignment between features of different scales is achieved based on the feature pyramid alignment mechanism. The fused feature map is convolved to generate the grab position and bounding box information. Based on the image detection results and multi-source sensor position data, the attention mechanism is used to achieve continuous tracking and abnormal judgment of the grab state; The detection and tracking results are input into the grab control system to automatically generate accurate grabbing paths and posture adjustment instructions, and control the grab action in real time; According to the grasping results and the posture deviation sent back by the sensor, the control model parameters are dynamically adjusted to continuously optimize the detection accuracy and control strategy.
[0006] Preferably, the multi-type sensors include a posture sensor, a position coordinate sensor, a rotation angle sensor and a laser ranging and obstacle avoidance sensor.
[0007] Preferably, the step of enhancing and normalizing the acquired video image and synchronously fusing it with the sensor data at a unified timestamp includes: Perform histogram equalization, gamma correction, filtering noise reduction and sharpening enhancement on the image; Normalize the image pixel values to the range of [0,1] and unify the image size to meet the YOLOv3-tiny network input requirements; Normalize and unify the dimensions of all sensor data and encapsulate them into a standardized tensor format; Align image frames and sensor data timestamps based on a unified clock system; Construct multimodal datasets of image-sensor joint input by splicing or network fusion.
[0008] Preferably, a hole convolution layer is introduced into the original YOLOv3-tiny structure and spatial alignment is performed based on a feature pyramid alignment mechanism, including: Introducing dilated convolution after deep feature extraction; Use upsampling operations to upsample high-level feature maps to the same spatial resolution as shallow feature maps; Perform 1×1 convolution on the shallow feature map to match the number of channels; Introduce deformable convolution or spatial transformation network modules before feature fusion to spatially align feature maps; The fused feature map is convolved using a 3×3 standard convolution to output a feature map of uniform scale for bounding box detection.
[0009] Preferably, the use of the attention mechanism to achieve continuous tracking and abnormal judgment of the grab state includes: constructing a fusion feature vector, inputting a channel attention mechanism, introducing a spatial attention mechanism, using LSTM combined with a Self-Attention mechanism to construct a state time series model, and judging the dynamic stability trend of the grab; and identifying abnormal states according to preset abnormal judgment rules.
[0010] Preferably, the automatic generation of accurate grasping paths and posture adjustment instructions includes: Based on the current position of the grab bucket, the target position and the distribution of obstacles in the working area, a path planning algorithm is used to generate the optimal falling trajectory; Set critical control points along the route, including grab points, closure points, and transfer points; Combine the IMU attitude and visual detection results to calculate the required attitude adjustment angle; Output motor control voltage or hydraulic valve opening command to achieve real-time adjustment of grab bucket posture; The control period is set to no more than 100 milliseconds to meet high-frequency real-time control requirements.
[0011] Preferably, the lateral drift rate of the grab bucket's motion trajectory is generated by analyzing the trajectory deviation amplitude and trend in the horizontal direction when the grab bucket moves from the starting point to the target position. The acquisition method is: The point sequences of the ideal motion trajectory and the actual motion trajectory are respectively represented as two-dimensional coordinate point sets sampled in time order, where the ideal trajectory point sequence is , N is the total number of ideal trajectory points, and the actual trajectory point sequence is , M is the total number of actual trajectory points, where each point contains the coordinate values of the X-axis and the Y-axis; for each pair of trajectory points, the absolute distance difference in the Y-axis direction is calculated, that is, a two-dimensional distance matrix is constructed to represent the lateral offset between any pair of ideal points and actual points, and each element represents the distance error between two points in the Y direction; The minimum cumulative cost path algorithm in the DTW algorithm is used. Starting from the starting point, the minimum cumulative cost value from the starting point to any point pair is recursively calculated through dynamic programming. The cumulative cost is based on the previous state, and the minimum cost path of the adjacent points is selected to superimpose the current Y-direction deviation until all point pairs are traversed to form a complete cost matrix. After the cost matrix calculation is completed, the optimal alignment path is extracted by backtracking from the end point of the cost matrix. That is, the actual trajectory points are paired with the ideal trajectory points at the optimal time. The offset of each point pair in the Y direction during the entire trajectory is extracted through the path. The Y deviations of all matching point pairs are averaged to obtain the lateral drift rate of the grab movement.
[0012] Preferably, the terminal posture stabilization time is generated by analyzing the time taken for the posture of the grab bucket to converge from the swaying state to the tolerable range before it descends and closes. The acquisition method is: Collect continuous attitude data of the grab bucket from the IMU or attitude solution system: Pitch angle sequence: , roll angle sequence: ; Set a sliding time window length , corresponding to several frames of data, at any time t, calculate the average absolute error of the posture in the past time window ; The average absolute error of the posture in the current window should satisfy: < and < , that is, judging that the grab bucket posture enters a stable state, is the angle error range, and records the time ts from the start of descent to the first time the condition is met, then: ; Among them: t0 is the time point when the grab bucket starts to descend or start to swing; Tstable is the terminal posture stabilization time.
[0013] Preferably, the lateral drift rate and the attitude stabilization time are used as the input of the fuzzy logic system; the control error level is used as the output of the fuzzy logic system, and the graded parameters are adjusted according to the error level: the current control parameters are maintained at low errors; the proportional control gain Kp is increased and the waiting stabilization time is increased at the medium error stage; the adaptive compensation mechanism is triggered at the high error stage to improve the response rate of the PID controller or activate the trajectory correction strategy.
[0014] The present invention also provides a gantry crane grab control system based on multi-sensor fusion, including an image acquisition module, an image preprocessing module, a grab visual detection module, a grab state tracking module, a path planning and motion control module and an adaptive control optimization module; Image acquisition module: Image acquisition equipment and multiple types of sensors are deployed in the gantry crane operation area to obtain the grab bucket’s video image, posture information, position coordinates, and rotation angle data in real time; Image preprocessing module: enhances and normalizes the acquired video images, and synchronizes them with the sensor data at a unified timestamp to construct a multi-dimensional input data set; Grab bucket visual detection module: A hole convolution layer is introduced into the original YOLOv3-tiny structure. Before the high-level and low-level features are fused, the spatial alignment between features of different scales is achieved based on the feature pyramid alignment mechanism. The fused feature map is convoluted to generate the grab bucket position and bounding box information. Grab bucket state tracking module: Based on image detection results and multi-source sensor position data, the attention mechanism is used to achieve continuous tracking and abnormal judgment of the grab bucket state; Path planning and motion control module: inputs the detection and tracking results into the grab control system, automatically generates accurate grabbing paths and posture adjustment instructions, and controls the grab action in real time; Adaptive control optimization module: dynamically adjusts control model parameters based on the grasping results and the posture deviation sent back by the sensor, and continuously optimizes detection accuracy and control strategies.
[0015] In the above technical solution, the technical effects and advantages provided by the present invention are: 1. This invention breaks through the limitations of traditional grab bucket recognition and control based on single visual information, integrates video images and multiple sensor data, builds a unified multi-dimensional input model in image preprocessing, time synchronization, data normalization, etc., and introduces hole convolution and feature alignment mechanism by improving the YOLOv3-tiny structure, realizing accurate detection and bounding box regression of grab buckets of different scales. By integrating the attention mechanism with the timing network, the system can continuously track the state of the grab bucket and make intelligent judgments on abnormal situations during the operation process, such as sway, jamming, collision risk, etc., effectively improving the accuracy of grab bucket recognition and the stability of system operation.
[0016] 2. The present invention realizes real-time evaluation of control error level by performing fuzzy logic analysis on key control feedback quantities such as lateral drift rate and terminal attitude stabilization time, and dynamically adjusts control model parameters based on error level, including adjusting PID gain, adaptive compensation strategy and grab waiting time, thereby realizing closed-loop optimization and adaptive enhancement of the control system. The system is particularly suitable for large-scale material handling scenarios such as ports and mining areas. It can still maintain high grabbing accuracy and system robustness under small-sized grab buckets, multi-target interference or complex dynamic operating conditions, significantly improving the intelligence level and operating efficiency of gantry crane operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0018] Figure 1 The figure is a mind map of the method of the present invention.
[0019] Figure 2 This is a system mind map of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Example 1, please refer to Figure 1 As shown, the gantry crane grab control method based on multi-sensor fusion described in this embodiment includes: Image acquisition equipment and multiple types of sensors are deployed in the gantry crane operation area to obtain the grab bucket's video images, posture information, position coordinates, and rotation angle data in real time; Enhance and normalize the acquired video images, and synchronize them with the sensor data at a unified timestamp to construct a multi-dimensional input data set; A hole convolution layer is introduced into the original YOLOv3-tiny structure. Before the high-level and low-level features are fused, the spatial alignment between features of different scales is achieved based on the feature pyramid alignment mechanism. The fused feature map is convolved to generate the grab position and bounding box information. Based on the image detection results and multi-source sensor position data, the attention mechanism is used to achieve continuous tracking and abnormal judgment of the grab state; The detection and tracking results are input into the grab control system to automatically generate accurate grabbing paths and posture adjustment instructions, and control the grab action in real time; According to the grasping results and the posture deviation sent back by the sensor, the control model parameters are dynamically adjusted to continuously optimize the detection accuracy and control strategy.
[0022] In the intelligent control system of the gantry crane grab, in order to achieve accurate perception and dynamic control of the grab's operating status, it is necessary to deploy image acquisition equipment and multiple types of sensors in the gantry crane's operating area to perform multi-dimensional real-time monitoring of the grab. The following is a detailed description of the deployment process and the data obtained: Camera type selection: Use industrial-grade high-definition cameras (such as CMOS cameras with more than 2 million pixels) with wide dynamic range and low latency characteristics; optional infrared night vision cameras or thermal imaging cameras can ensure image acquisition capabilities at night or in smoky environments; if depth information is required, an RGB-D camera (such as Intel RealSense) or a stereo camera system can be integrated.
[0023] Installation location layout principles: The camera is installed on the gantry crane trolley beam, grab bracket or hoist, with a bird's-eye view angle of about 45°, covering the up and down movement area of the grab; if the working environment is large (such as a bulk cargo warehouse in a port), multiple fixed cameras are deployed in key areas for regional monitoring, and PTZ (pan-tilt rotation) cameras are used to achieve dynamic tracking when necessary; the camera is connected to the industrial Ethernet or edge computing unit to support real-time transmission of video streams to the control host or AI detection module.
[0024] Collected image content: Real-time capture of grab shape (closed, open), posture changes, and cargo status in the operating area; cooperate with the AI detection module to identify the grab outline, edge, shadow, cargo type, etc.
[0025] Attitude sensor (IMU – Inertial Measurement Unit): Installed on the grab frame or the connection part of the spreader, it collects the acceleration, angular velocity, inclination and yaw angle of the grab in real time; it senses the movement attitude of the grab in three-dimensional space (including swing amplitude and direction); and the data is uploaded to the central processing module via the CAN or RS485 interface.
[0026] Position coordinate sensor: GNSS-RTK high-precision positioning module (centimeter-level accuracy) is used and installed on the gantry body and the top of the grab bucket. In indoor or blocked scenes, it is supplemented by UWB positioning base station and tag system to perform high-precision indoor positioning. The absolute or relative position coordinates of the grab bucket in the XYZ three-dimensional space are obtained in real time.
[0027] Rotation angle sensor (encoder): installed on the grab bucket rotating shaft, motor or spreader connection; obtains the grab bucket's rotation angle information in real time on the vertical and horizontal planes (such as spreader rotation angle, grab bucket flip angle); can choose absolute encoder or incremental encoder, and cooperate with the controller for feedback control.
[0028] Laser ranging and obstacle avoidance sensors: installed at the four corners of the grab or on the arm, used to measure the relative distance between the grab and the ground, cargo or obstacles; LiDAR or ToF sensors are used to build a contour model of the working area to assist in path planning and collision avoidance.
[0029] Synchronous acquisition and fusion mechanism: All sensors and image acquisition devices synchronize data through timestamps and unify sampling frequencies; Use multi-sensor fusion algorithms (such as Kalman filter, extended Kalman filter EKF or deep fusion network) to jointly process image, pose and position data.
[0030] Edge computing and real-time transmission: Deploy some detection models and data processing logic on the edge computing device at the door machine end (such as the NVIDIA Jetson series); high-frequency data is transmitted to the central control platform via industrial bus (CAN, EtherCAT) or wireless communication (5G, Wi-Fi 6).
[0031] The acquired video images are enhanced and normalized, and synchronously fused with the sensor data at a unified timestamp to construct a multidimensional input data set. The detailed process is as follows: Image enhancement processing: In order to improve the recognizability of images in complex environments (such as backlight, low illumination, rain, fog, dust, etc.), the collected video frames are enhanced, including: Histogram equalization: improve image contrast and highlight the edge and contour of the grab; Gamma correction: adjust the brightness range to avoid overexposure or underexposure; Gaussian filtering or bilateral filtering: reduce noise while retaining edge information; Image sharpening processing: enhance the grab structure lines and improve the detection model's sensitivity to edges; Gamma enhancement + color channel reconstruction: used to improve feature visibility at night or in complex color temperature environments.
[0032] Image normalization: Normalize the pixel values of the image to the range of [0, 1] or [-1, 1]. According to the input requirements of the YOLOv3-tiny model, unify the image size (such as 416×416 or 320×320). Normalize the channel dimension (RGB) of the image so that the network input has good distribution characteristics.
[0033] Original sensor data acquisition: including the grab's three-axis acceleration, angular velocity (IMU), GPS / UWB coordinates, attitude angle (Pitch, Roll, Yaw), grab angle (encoder), etc.; data is transmitted to the main control system via high-speed serial port or CAN bus.
[0034] Data normalization and format unification: Unify the dimensions of physical quantities (angle, position, acceleration) and standardize them (such as unit conversion and normalization to the interval [-1, 1]); all data are encapsulated as structured arrays or tensors to facilitate subsequent neural network processing.
[0035] Timestamp alignment: All image frames and sensor data acquisition modules are connected to a unified system clock (NTP server synchronization or PPS pulse synchronization can be used); the image acquisition frame rate (such as 30Hz) and the sensor data sampling frequency (such as 100Hz) are time-aligned through interpolation, downsampling or Kalman filtering; ensure that each frame of the image is accompanied by a set of multi-sensor data consistent with its timestamp.
[0036] Data fusion: Use synchronized timestamps to concatenate image features and sensor features into a multi-dimensional input tensor. The input structure can be multi-channel input (image + numerical channel), or fused after parallel processing of branch neural networks.
[0037] Build multimodal input data that can be used for training or real-time reasoning; support the improvement of the YOLOv3-tiny model to perform deep feature and sensor feature fusion (such as intermediate fusion or post-fusion); improve the robustness of the grab detection system to complex environments and the response accuracy of the control system.
[0038] A dilated convolution layer is introduced into the original YOLOv3-tiny structure. Before the high-level and low-level features are fused, the spatial alignment between features of different scales is achieved based on the feature pyramid alignment mechanism. The fused feature map is convolved to generate the grab position and bounding box information. Specifically, the input image size is 416×416×3; it is preprocessed (enhanced and normalized); and it is input into the system together with the synchronized sensor data (see the aforementioned multimodal input construction).
[0039] Use the YOLOv3-tiny backbone network (simplified version of Darknet-19) to extract shallow and deep features.
[0040] Shallow feature map (C3): from the lower convolutional layer, size is about 52×52×256; Deep feature map (C5): comes from the end of the network, with a size of about 13×13×512.
[0041] Insert a dilated convolution layer after the deep feature extraction module, such as dilated conv3×3, with dilation rate = 2 or 4; Function: Expand the receptive field; No need to increase parameters or significantly reduce resolution; Enhance the model's recognition ability for large-scale objects (such as an unfolded grab or a covered cargo pile); Output is an extended perceptual feature map, and the size is maintained at 13×13×512.
[0042] In order to alleviate the spatial alignment error caused by dilated convolution, an FPN-style alignment mechanism is introduced before multi-scale feature fusion: Use upsampling to upsample the deep feature map from 13×13 to 26×26 or 52×52; At the same time, 1×1 convolution is performed on the shallow feature map to compress the number of channels and maintain channel consistency.
[0043] Use one of the following techniques to perform spatial alignment: Deformable convolution: Automatically learn displacement and enhance spatial adaptation between different scales; Attention mechanism (such as SE or CBAM): guides the shallow layer to focus on high-level semantic areas; Spatial Transformer Module (STN): transforms the spatial coordinates of shallow feature maps before merging; Purpose: Ensure that feature maps of different scales are aligned in geometric coordinates to solve the deviation problem caused by hole convolution.
[0044] The processed shallow and deep feature maps are fused, and the methods include: Element-wise addition or concatenation; Then it is followed by 3×3 standard convolution for fusion denoising; A fused feature map of uniform scale and rich information is obtained, with a size of, for example, 52×52×512.
[0045] Use the YOLO detection head to further convolve the fused feature map: The number of convolution kernel output channels is B × (5 + C), where B is the number of prediction boxes for each grid (such as 3), and C is the number of grab categories (if only the grab is identified, C=1); Output the position (x, y, w, h), confidence and category probability of each candidate box. x is the relative coordinate of the center point of the candidate box in the image width direction (in units of image width). It indicates the horizontal position of the center of the box. y is the relative coordinate of the center point of the candidate box in the image height direction (in units of image height). It indicates the vertical position of the center of the box. w is the width of the candidate box, which indicates the ratio of the width of the grab or target box to the width of the entire image. It is usually the normalized value of the width of the candidate box. h is the height of the candidate box, which indicates the ratio of the height of the grab or target box to the height of the entire image. It is usually the normalized value of the height of the candidate box.
[0046] Application of non-maximum suppression (NMS) to the prediction results to remove overlapping boxes; confidence screening threshold (such as 0.5); three-dimensional position mapping can be performed in combination with sensor spatial information (such as GPS, IMU); and finally output of the grab’s two-dimensional image bounding box coordinates and spatial position estimation for use by the control system.
[0047] The 2D bounding box coordinates (x, y, w, h) and detection confidence of the grab in each frame are obtained from the improved YOLOv3-tiny model. If the posture estimation module is integrated, the orientation and opening and closing state of the grab can also be extracted.
[0048] Multi-source sensor data acquisition: IMU sensor: obtain acceleration, angular velocity, tilt angle and other data; GPS / UWB positioning: obtain grab bucket spatial coordinates (X, Y, Z); encoder: grab bucket rotation angle, spreader swing angle; laser ranging: the distance between the grab bucket and the cargo / ground. Feature vector fusion: synchronize all the above data according to the unified timestamp; construct a fusion feature vector.
[0049] Input attention network module: Use the channel attention mechanism (SE module) to highlight key physical quantities, such as grab height, angular velocity changes, etc.; use the spatial attention mechanism to strengthen the attention to the changed areas in the grab image of the current frame (such as tilt, stretch or offset).
[0050] State trend modeling: Input multi-frame feature vector sequences into a time series network with attention (such as LSTM + Self-Attention); learn the dynamic change trend of the grab state and identify irregular movement behaviors.
[0051] In the grab state tracking system based on image recognition and multi-source sensor fusion, a set of systematic abnormal judgment rules must be set to monitor the grab operation status in real time and trigger control intervention in time when an abnormality is detected. Specifically, the following typical abnormal states and their judgment conditions are included: Abnormal excessive sway of grab bucket: If the grab bucket swings significantly left and right or forward and backward during lifting, it may cause problems such as grab bucket collision, misalignment or structural fatigue. Such abnormalities can be judged by the tilt angle obtained by the attitude sensor (such as IMU): When the roll angle (Roll) or pitch angle (Pitch) of the grab exceeds ±15° and lasts for more than 1 second, it is considered as abnormal grab swing; the system will trigger an early warning and slow down the running speed or perform anti-swing actions.
[0052] Grab bucket is not closed or stuck abnormally: When the grab bucket performs the closing action, if it is not completely closed or stuck due to foreign matter, mechanical failure or no load, it will cause the grab to fail. The judgment method includes: the recognition confidence of the grab bucket closing state in the image detection is significantly reduced; at the same time, the change range of the grab bucket rotation angle (obtained by the encoder) is lower than the set threshold, indicating that the expected mechanical action has not occurred; this combined state is judged as a stuck abnormality, which can prompt manual inspection or execution of a re-closing command.
[0053] Abnormal positioning drift: In dynamic operations, if there is a significant difference between the grab position detected by the image and the coordinates fed back by the positioning system such as GNSS or UWB, it may indicate a sensor failure or interference with the positioning signal. The judgment logic is: calculate the spatial distance between the predicted coordinates of the center point of the image and the GNSS positioning coordinates; when the error value exceeds 30 cm (or other set thresholds) continuously, the "abnormal positioning drift" is triggered; the system can temporarily disable some positioning signals and use the backup data source instead.
[0054] Abnormal posture fluctuation: If abnormal vibration or rotation is detected during the period when the grab is stationary (such as hovering, waiting for instructions), it may be due to loose structure, control error or external interference. It can be monitored by IMU: if there is a sudden change in acceleration or angular velocity jump in a short period of time in the stationary state (such as exceeding the set threshold Δa or Δω), it is judged as "abnormal posture fluctuation"; such abnormalities will trigger the dynamic compensation mechanism or suspend the current operation instruction.
[0055] Grab failure exception: If the grab bucket performs a complete grab action but fails to grab the material, it needs to be identified and corrected in time. The judgment method is: after the grab action is completed, it is judged by the weighing sensor or the change in the hoisting load; if the load mass is approximately 0 after the grab bucket is closed, it means that the material has not been grabbed, which triggers the "grab failure" exception; the system can re-identify the target position and initiate a re-grab command.
[0056] Abnormal collision risk: To prevent the grab from colliding with the ground, equipment or materials during the falling process, the system needs to evaluate the distance and movement status between the grab and the obstacle in real time: if the distance measurement data at the bottom of the grab is lower than the set minimum safety distance (such as <30cm), and the current descent speed exceeds the threshold (such as >0.5m / s), it is regarded as a "collision risk"; at this time, the system should immediately perform emergency braking, alarm or path correction operations.
[0057] The detection and tracking results are input into the grab control system to automatically generate accurate grabbing paths and posture adjustment instructions, and realize the real-time control of the grab action process, specifically: The input includes: the grab bucket bounding box coordinates (x, y, w, h) obtained from the YOLOv3-tiny improved network; the dynamic trajectory of the grab bucket in the image sequence (which can be tracked through Kalman filtering or attention mechanism); the spatial position information provided by multi-source sensors (such as GNSS / UWB coordinates, IMU attitude angle, encoder angle); the position and state information of the target material (such as the outline of the loading area, the target center point); Target state analysis: determine whether the current position of the grab is aligned with the center of the target material; analyze whether the grab's posture is suitable for grabbing (whether the angle is horizontal, whether it is swinging); if multiple targets are identified, select the optimal target area as the current grab object.
[0058] Target path planning input: Starting point: current grab 3D coordinates + posture; Target point: material surface or center position (combined with image detection and laser ranging modeling); Constraints: obstacle avoidance area, maximum speed / acceleration limit, anti-sway control boundary. Path generation algorithm: Use common path planning methods (such as A*, RRT*, DWA) combined with grab working conditions to build motion trajectories; Insert key control points (Key Poses) in the path, including descending sections, closing points, and ascending sections; Fit the path with a speed curve to achieve smooth start and precise braking.
[0059] According to the target direction, combined with the IMU attitude and visual correction results, the required attitude correction angle (Pitch, Roll) is calculated; the corresponding motor / cylinder action instructions are generated to achieve automatic horizontal or vertical adjustment of the grab.
[0060] The control system is based on a closed-loop structure. After receiving the path and posture instructions, it generates control quantities (such as motor voltage and valve opening); Use PID control, adaptive control or model predictive control (MPC) algorithm to adjust the execution unit to achieve precise control of position, speed, angle, etc.
[0061] All control commands are transmitted to the door machine execution unit through the CAN bus, EtherCAT or PLC command system; the control cycle is generally 50ms~100ms to meet the high-frequency dynamic response requirements; the system has built-in buffering and emergency stop mechanisms to ensure immediate cessation of operation under abnormal circumstances.
[0062] The execution actions include: the grab moves to the target area; automatically leveling and falling vertically; the grab closes; and after completion, it automatically rises and transfers to the target unloading position.
[0063] Feedback signal collection: real-time monitoring of grab position, posture, speed and load; if the error deviation is found to be greater than the set tolerance, the trajectory correction mechanism is triggered; after the operation is completed, the status record is updated and fed back to the dispatching system or upper-level management platform.
[0064] According to the grasping results and the posture deviation sent back by the sensor, the control model parameters are dynamically adjusted to continuously optimize the detection accuracy and control strategy. Specifically: The lateral drift rate of the grab bucket's motion trajectory is generated by analyzing the trajectory deviation amplitude and trend in the horizontal direction (especially the Y axis) when the grab bucket moves from the starting point to the target position. The acquisition method is as follows: The point sequences of the ideal motion trajectory and the actual motion trajectory are respectively represented as two-dimensional coordinate point sets sampled in time order. The ideal trajectory point sequence is , N is the total number of ideal trajectory points, and the actual trajectory point sequence is , M is the total number of actual trajectory points, where each point contains the coordinate values of the X-axis and the Y-axis.
[0065] For each pair of trajectory points, the absolute distance difference in the Y-axis direction is calculated, that is, a two-dimensional distance matrix is constructed to represent the lateral offset between any pair of ideal points and actual points. Each element represents the distance error between two points in the Y direction.
[0066] The minimum cumulative cost path algorithm in the DTW algorithm is used. Starting from the starting point (the first set of trajectory points), the minimum cumulative cost value from the starting point to any point pair is recursively calculated through dynamic programming. The cumulative cost is based on the previous state, and the minimum cost path of the adjacent points is selected to superimpose the current Y-direction deviation until all point pairs are traversed to form a complete cost matrix.
[0067] After the cost matrix is calculated, trace back from the end point of the cost matrix (i.e., the last point pair) to extract the optimal alignment path, that is, to pair the actual trajectory points with the ideal trajectory points at the optimal time. Through this path, the offset of each point pair in the Y direction during the entire trajectory can be extracted. The Y deviations of all matching point pairs are averaged to obtain the lateral drift rate of the grab movement.
[0068] The terminal posture stabilization time is generated by analyzing the time taken for the posture of the grab bucket to converge from the sway state to the tolerable range (such as ±2°) before it descends and closes. The acquisition method is: Collect continuous attitude data of the grab bucket from the IMU or attitude solution system: Pitch angle sequence (Pitch): , rolling angle sequence (Roll): ; Define the angle error range allowed for the grab to stabilize, for example: , which means that when the attitude error is continuously within ±2°, it is considered stable. Set a sliding time window length , in seconds (e.g. = 1s), corresponding to several frames of data (such as 50 frames). At any time t, calculate the average absolute error of the posture in the past time window: ; In the formula, The average absolute error of the posture in the current window satisfies: < and < , that is, to judge that the grab bucket posture enters a stable state, and record the time ts from the beginning of the descent to the first time the above conditions are met, then: ; Among them: t0 is the time point when the grab bucket starts to descend or start to swing; Tstable is the terminal posture stabilization time.
[0069] The lateral drift rate of the grab bucket motion trajectory and the terminal posture stabilization time are used as the input items of the fuzzy logic, and the error level of the gantry crane grab bucket control is used as the output item of the fuzzy logic. Input variable 1: lateral drift rate, which indicates the average deviation of the grab bucket from the ideal trajectory in the Y-axis direction during movement. Its fuzzy language value is set as "small", "medium" or "large".
[0070] Input variable 2: End posture stabilization time refers to the time required for the grab bucket to converge from the sway state to the set tolerance range (such as ±2°) after it is lowered. Its fuzzy language value is set to "fast", "normal", and "slow".
[0071] Output variable: Control error level, used to quantify the performance of the control system in the current operation process. The output value range is 0 to 1, and the fuzzy language values are "low error", "medium error" and "high error".
[0072] According to the control experience and experimental data, a fuzzy rule base is constructed. The rule form is "if...then...", for example: If the drift rate is "small" and the settling time is "fast", the control error level is "low error"; If the drift rate is "medium" and the settling time is "slow", the control error level is "high error"; If the drift rate is "large" and the settling time is "normal", the control error level is "high error".
[0073] By combining different input situations, a two-dimensional fuzzy reasoning table containing 9 core rules is established.
[0074] Fuzzy processing: The actually measured drift rate and stabilization time are input and mapped into the membership degree of the corresponding language value through the membership function.
[0075] Rule activation and weight calculation: Calculate the activation strength of all rules, usually using the "minimum value" or "multiplication method" for fuzzy intersection processing.
[0076] Fuzzy output aggregation: superimpose the output parts of all activated rules to synthesize the final fuzzy output set.
[0077] Defuzzification: The fuzzy output set is defuzzified using the Center of Gravity method to obtain a clear value representing the control error level.
[0078] According to the error level of the fuzzy system output, a hierarchical adjustment strategy is implemented to dynamically optimize the control model parameters: Low error (output value ∈ [0, 0.3]): The control system is in good condition and maintains the current parameter configuration without adjustment.
[0079] Mean error (output value ∈ (0.3, 0.7]): appropriately increase the controller response gain, such as increasing the proportional term Kp in the path tracking PID; at the same time, extend the stable waiting time before the grab closing action to improve the posture stability.
[0080] High error (output value ∈ (0.7, 1.0]): indicates that the system deviation is significant and a stronger compensation mechanism needs to be activated. For example, activate the trajectory offset adaptive compensation module, increase the differential gain of the angle closed-loop controller, limit the maximum moving speed of the grab, issue a system abnormality prompt or perform recalibration when necessary.
[0081] The above fuzzy reasoning and parameter adjustment process can be embedded in the real-time scheduling process of the grab control system to form a closed-loop adaptive control mechanism of perception, evaluation, adjustment and feedback to achieve dynamic control performance optimization.
[0082] Example 2, please refer to Figure 2As shown, a gantry crane grab control system based on multi-sensor fusion described in this embodiment includes an image acquisition module, an image preprocessing module, a grab visual detection module, a grab state tracking module, a path planning and motion control module, and an adaptive control optimization module; Image acquisition module: Image acquisition equipment and multiple types of sensors are deployed in the gantry crane operation area to obtain the grab bucket’s video image, posture information, position coordinates, and rotation angle data in real time; Image preprocessing module: enhances and normalizes the acquired video images, and synchronizes them with the sensor data at a unified timestamp to construct a multi-dimensional input data set; Grab bucket visual detection module: A hole convolution layer is introduced into the original YOLOv3-tiny structure. Before the high-level and low-level features are fused, the spatial alignment between features of different scales is achieved based on the feature pyramid alignment mechanism. The fused feature map is convoluted to generate the grab bucket position and bounding box information. Grab bucket state tracking module: Based on image detection results and multi-source sensor position data, the attention mechanism is used to achieve continuous tracking and abnormal judgment of the grab bucket state; Path planning and motion control module: inputs the detection and tracking results into the grab control system, automatically generates accurate grabbing paths and posture adjustment instructions, and controls the grab action in real time; Adaptive control optimization module: dynamically adjusts control model parameters based on the grasping results and the posture deviation sent back by the sensor, and continuously optimizes detection accuracy and control strategies.
[0083] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0084] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0085] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0086] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A gantry crane grab bucket control method based on multi-sensor fusion, characterized in that: include: Image acquisition equipment and multiple types of sensors are deployed in the gantry crane operation area to obtain the grab bucket's video images, posture information, position coordinates, and rotation angle data in real time; Enhance and normalize the acquired video images, and synchronize them with the sensor data at a unified timestamp to construct a multi-dimensional input data set; A hole convolution layer is introduced into the original YOLOv3-tiny structure. Before the high-level and low-level features are fused, the spatial alignment between features of different scales is achieved based on the feature pyramid alignment mechanism. The fused feature map is convolved to generate the grab position and bounding box information. Based on the image detection results and multi-source sensor position data, the attention mechanism is used to achieve continuous tracking and abnormal judgment of the grab state; The detection and tracking results are input into the grab control system to automatically generate accurate grabbing paths and posture adjustment instructions, and control the grab action in real time; According to the grasping results and the posture deviation sent back by the sensor, the control model parameters are dynamically adjusted to continuously optimize the detection accuracy and control strategy.
2. The method for controlling a gantry crane grab bucket based on multi-sensor fusion according to claim 1, characterized in that: The multi-type sensors include a posture sensor, a position coordinate sensor, a rotation angle sensor, and a laser ranging and obstacle avoidance sensor.
3. The method for controlling a gantry crane grab bucket based on multi-sensor fusion according to claim 1, characterized in that: The step of enhancing and normalizing the acquired video image and synchronously fusing it with the sensor data at a unified timestamp includes: Perform histogram equalization, gamma correction, filtering noise reduction and sharpening enhancement on the image; Normalize the image pixel values to the range of [0,1] and unify the image size to meet the YOLOv3-tiny network input requirements; Normalize and unify the dimensions of all sensor data and encapsulate them into a standardized tensor format; Align image frames and sensor data timestamps based on a unified clock system; Construct multimodal datasets of image-sensor joint input by splicing or network fusion.
4. The gantry crane grab bucket control method based on multi-sensor fusion according to claim 1 is characterized in that: Introducing the hole convolution layer into the original YOLOv3-tiny structure and performing spatial alignment based on the feature pyramid alignment mechanism, including: Introducing dilated convolution after deep feature extraction; Use upsampling operations to upsample high-level feature maps to the same spatial resolution as shallow feature maps; Perform 1×1 convolution on the shallow feature map to match the number of channels; Introduce deformable convolution or spatial transformation network modules before feature fusion to spatially align feature maps; The fused feature map is convolved using a 3×3 standard convolution to output a feature map of uniform scale for bounding box detection.
5. The method for controlling a gantry crane grab bucket based on multi-sensor fusion according to claim 4 is characterized in that: The use of the attention mechanism to achieve continuous tracking and abnormal judgment of the grab state includes: constructing a fusion feature vector, inputting a channel attention mechanism, introducing a spatial attention mechanism, using LSTM combined with a Self-Attention mechanism to construct a state time series model, and judging the dynamic stability trend of the grab; and identifying abnormal states according to preset abnormal judgment rules.
6. The method for controlling a gantry crane grab bucket based on multi-sensor fusion according to claim 5, characterized in that: The automatic generation of accurate grasping paths and posture adjustment instructions includes: Based on the current position of the grab bucket, the target position and the distribution of obstacles in the working area, a path planning algorithm is used to generate the optimal falling trajectory; Set critical control points along the route, including grab points, closure points, and transfer points; Combine the IMU attitude and visual detection results to calculate the required attitude adjustment angle; Output motor control voltage or hydraulic valve opening command to achieve real-time adjustment of grab bucket posture; The control period is set to no more than 100 milliseconds to meet high-frequency real-time control requirements.
7. The method for controlling a gantry crane grab bucket based on multi-sensor fusion according to claim 6, characterized in that: The lateral drift rate of the grab bucket's motion trajectory is generated by analyzing the trajectory deviation amplitude and trend in the horizontal direction when the grab bucket moves from the starting point to the target position. The acquisition method is as follows: The point sequences of the ideal motion trajectory and the actual motion trajectory are respectively represented as two-dimensional coordinate point sets sampled in time order, where the ideal trajectory point sequence is , N is the total number of ideal trajectory points, and the actual trajectory point sequence is , M is the total number of actual trajectory points, where each point contains the coordinate values of the X-axis and the Y-axis; for each pair of trajectory points, the absolute distance difference in the Y-axis direction is calculated, that is, a two-dimensional distance matrix is constructed to represent the lateral offset between any pair of ideal points and actual points, and each element represents the distance error between two points in the Y direction; The minimum cumulative cost path algorithm in the DTW algorithm is used. Starting from the starting point, the minimum cumulative cost value from the starting point to any point pair is recursively calculated through dynamic programming. The cumulative cost is based on the previous state, and the minimum cost path of the adjacent points is selected to superimpose the current Y-direction deviation until all point pairs are traversed to form a complete cost matrix. After the cost matrix calculation is completed, the optimal alignment path is extracted by backtracking from the end point of the cost matrix. That is, the actual trajectory points are paired with the ideal trajectory points at the optimal time. The offset of each point pair in the Y direction during the entire trajectory is extracted through the path. The Y deviations of all matching point pairs are averaged to obtain the lateral drift rate of the grab movement.
8. The method for controlling a gantry crane grab bucket based on multi-sensor fusion according to claim 7, characterized in that: The terminal posture stabilization time is generated by analyzing the time taken for the grab bucket to converge from the swaying state to the tolerable range before it descends and closes. The acquisition method is as follows: Collect continuous attitude data of the grab bucket from the IMU or attitude solution system: Pitch angle sequence: , roll angle sequence: ; Set a sliding time window length , corresponding to several frames of data, at any time t, calculate the average absolute error of the posture in the past time window ; The average absolute error of the posture in the current window should satisfy: < and < , that is, judging that the grab bucket posture enters a stable state, is the angle error range, and records the time ts from the start of descent to the first time the condition is met, then: ; Where: t0 is the time point when the grab bucket starts to descend or start to swing; Tstable is the terminal attitude stabilization time.
9. The method for controlling a gantry crane grab bucket based on multi-sensor fusion according to claim 8, characterized in that: The lateral drift rate and attitude stabilization time are used as the input of the fuzzy logic system; the control error level is used as the output of the fuzzy logic system, and the graded parameters are adjusted according to the error level: the current control parameters are maintained at low errors; the proportional control gain Kp is increased and the waiting stabilization time is increased in the medium error stage; the adaptive compensation mechanism is triggered in the high error stage to improve the response rate of the PID controller or activate the trajectory correction strategy.
10. A gantry crane grab bucket control system based on multi-sensor fusion, used to implement a gantry crane grab bucket control method based on multi-sensor fusion as claimed in any one of claims 1 to 9, characterized in that: It includes image acquisition module, image preprocessing module, grab bucket visual detection module, grab bucket state tracking module, path planning and motion control module and adaptive control optimization module; Image acquisition module: Image acquisition equipment and multiple types of sensors are deployed in the gantry crane operation area to obtain the grab bucket’s video image, posture information, position coordinates, and rotation angle data in real time; Image preprocessing module: enhances and normalizes the acquired video images, and synchronizes them with the sensor data at a unified timestamp to construct a multi-dimensional input data set; Grab bucket visual detection module: A hole convolution layer is introduced into the original YOLOv3-tiny structure. Before the high-level and low-level features are fused, the spatial alignment between features of different scales is achieved based on the feature pyramid alignment mechanism. The fused feature map is convoluted to generate the grab bucket position and bounding box information. Grab bucket state tracking module: Based on image detection results and multi-source sensor position data, the attention mechanism is used to achieve continuous tracking and abnormal judgment of the grab bucket state; Path planning and motion control module: inputs the detection and tracking results into the grab control system, automatically generates accurate grabbing paths and posture adjustment instructions, and controls the grab action in real time; Adaptive control optimization module: dynamically adjusts control model parameters based on the grasping results and the posture deviation sent back by the sensor, and continuously optimizes detection accuracy and control strategies.
Citation Information
Patent Citations
Port grab bucket detection method based on improved YOLOv3-tiny algorithm
CN110826520A
Portal crane grab bucket control system and method based on multi-sensor fusion
CN117105098A
Method for realizing multi-dimensional image processing by using robot
CN117671626A
Robot control system and method, storage medium, controller and robot
CN118927246A
Intelligent grab bucket control method and system based on multi-sensor fusion
CN119822238A
Cited By
Sampling positioning precision calibration method and system based on sample library
CN120245003A
Calibration method and system for sampling positioning accuracy based on sample library
CN120245003B
Robot system for engineering monitoring
CN120245082A
A robotic system for engineering monitoring
CN120245082B
Material grabbing offset real-time compensation method and device based on multi-sensor data fusion
CN120287313A