Tray operation method, device and equipment of unmanned forklift and medium

By integrating a multi-sensor system and IMU into the unmanned forklift, and combining reverse compensation and neural network models, the problem of pallet recognition deviation in unmanned forklifts has been solved, enabling precise pallet operation control and improving the accuracy and safety of operations.

CN121415366APending Publication Date: 2026-01-27ANHUI JIUYAO INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511323414.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing unmanned forklifts are prone to image acquisition deviations due to vehicle movement during pallet recognition, resulting in misalignment between the fork teeth and the pallet, which in turn leads to risks such as operation failure, cargo damage, and equipment collision.

Method used

By integrating RGB color cameras, binocular depth cameras, and millimeter-wave radar sensors onto an unmanned forklift, combining real-time speed measurement information from an IMU for reverse compensation, and utilizing a pre-trained neural network model for pose prediction, precise control commands are generated.

Benefits of technology

It significantly improves the accuracy and reliability of pallet recognition and forklift operation of unmanned forklifts in dynamic working environments, and reduces the risk of operation failure and safety accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415366A_ABST
    Figure CN121415366A_ABST
Patent Text Reader

Abstract

The invention discloses a tray operation method, device and equipment of an unmanned forklift and a medium, and the method comprises the steps that when the unmanned forklift operates a target tray in a logistics park, identification information related to the target tray is acquired; measuring the speed information of the unmanned forklift in real time through an IMU (Inertial Measurement Unit); performing reverse compensation on the identification information based on the speed information to obtain attitude compensation identification information; inputting the attitude compensation identification information into a pre-trained neural network model, and predicting to obtain pose data of the target tray; and converting the pose data into a coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift generates a corresponding control instruction based on the converted pose data, and completes the operation of the target tray through the control instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, equipment and medium for pallet operation of an unmanned forklift. Background Technology

[0002] In modern logistics warehousing and industrial park settings, unmanned forklifts have gradually replaced manual labor in tasks such as pallet handling, stacking, and loading / unloading, becoming key equipment for improving logistics efficiency and reducing labor costs. In actual operation, unmanned forklifts need to use visual sensors (such as cameras, lidar, etc.) to identify the position and posture (i.e., "pose") of the target pallet and generate motion control commands accordingly to achieve precise picking or placing.

[0003] However, existing vision-based pallet recognition methods have certain limitations, which can easily lead to misalignment between the forklift forks and the pallet, resulting in risks such as operation failure, cargo damage, or even equipment collision. Summary of the Invention

[0004] This specification provides one or more embodiments of a pallet handling method, apparatus, equipment, and medium for unmanned forklifts, which are used to solve the technical problems mentioned in the background art.

[0005] One or more embodiments of this specification employ the following technical solutions: This specification provides one or more embodiments of a pallet handling method using an unmanned forklift, the method comprising: When an unmanned forklift operates on a target pallet in a logistics park, it acquires identification information related to the target pallet. The speed information of the unmanned forklift is measured in real time using an IMU; Based on the velocity information, the recognition information is reverse-compensated to obtain attitude compensation recognition information; The posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray; The pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

[0006] It should be noted that this invention introduces an IMU to measure the speed information of the unmanned forklift in real time and uses this information to perform reverse compensation on the recognition information, effectively eliminating the image acquisition deviation caused by the movement of the vehicle. Subsequently, the compensated information is input into a neural network model for pose prediction and transformed into the forklift coordinate system, so that the final generated control commands can accurately correspond to the actual spatial posture of the pallet. This significantly improves the accuracy and reliability of pallet recognition and forklift operation of the unmanned forklift in dynamic working environments and reduces the risk of operation failure and safety accidents.

[0007] Furthermore, the identification information includes the feature information of the target pallet, the scene depth information in the logistics park, and the distance measurement information of the target pallet; The acquisition of the identification information related to the target tray includes: Two-dimensional color images are acquired using an RGB color camera to extract the feature information, which includes texture, color, and visual key point information. Binocular parallax images are acquired using a binocular depth camera to extract the scene depth information; The point cloud image of the target tray is acquired using millimeter-wave radar to extract the ranging information.

[0008] It should be noted that this invention integrates the advantages of three sensors—RGB color camera, binocular depth camera, and millimeter-wave radar—in texture perception, depth calculation, and accurate ranging to construct a multi-dimensional recognition information system. This allows the characteristics of the target pallet, environmental depth, and its own distance information to complement and verify each other, thus providing a more comprehensive and reliable data foundation for subsequent motion compensation and neural network pose prediction. This greatly enhances the system's perception robustness and accuracy in complex scenarios such as changes in lighting, missing textures, or occlusion, providing a solid perception guarantee for the precise control of unmanned forklifts.

[0009] Furthermore, after obtaining the identification information related to the target tray, the method further includes: The extrinsic parameter matrix between the feature information, the scene depth information, and the ranging information is determined by a joint calibration algorithm, so as to transform the feature information, the scene depth information, or the ranging information to the same coordinate system. The extrinsic parameter matrix includes the relative position and attitude relationship between the feature information, the scene depth information, and the ranging information.

[0010] It should be noted that this invention, by introducing a joint calibration algorithm, accurately obtains the extrinsic parameter matrix between different sensor information, thereby unifying the originally heterogeneous and dispersed feature, depth, and ranging information in their respective coordinate systems to the same spatial reference. This effectively eliminates data fusion errors caused by differences in installation position and viewing angle of multi-source sensors, providing a highly consistent and spatially aligned reliable data foundation for subsequent motion compensation and pose prediction. This improves the accuracy and coordination of the entire perception system from the source, ultimately ensuring the accuracy and reliability of unmanned forklift operation decisions.

[0011] Furthermore, after obtaining the identification information related to the target tray, the method further includes: The synchronization controller sends a unified hardware trigger pulse signal to the RGB color camera, the binocular depth camera, and the millimeter-wave radar to acquire data from the RGB color camera, the binocular depth camera, and the millimeter-wave radar at the same time.

[0012] It should be noted that this invention sends a unified hardware trigger pulse signal to the RGB color camera, binocular depth camera, and millimeter-wave radar through a synchronization controller, forcing all sensors to acquire data at the exact same moment. This fundamentally eliminates the data time difference and spatiotemporal misalignment caused by the asynchronous acquisition time of each sensor, ensuring that the acquired multi-source perception information (features, depth, and ranging) has a high degree of temporal consistency and spatial synchronization. This provides a precise time-aligned data foundation for subsequent multi-source data fusion, motion compensation, and pose calculation, thereby significantly improving the temporal consistency and pose estimation accuracy of the entire perception system and enhancing the operational reliability of the unmanned forklift in dynamic scenarios.

[0013] Furthermore, the step of performing reverse compensation on the recognition information based on the velocity information to obtain attitude compensation recognition information includes: Based on the speed information, the motion of the unmanned forklift between two adjacent synchronous triggers is integrated to estimate the attitude change; Based on the posture change, the feature information, scene depth information and ranging information at the current moment are inversely compensated to obtain posture compensation recognition information.

[0014] It should be noted that this invention utilizes the speed information measured in real time by the IMU to integrate the vehicle's own motion within the interval between two adjacent synchronous triggering events, accurately estimating the attitude change of the unmanned forklift from the data acquisition moment to the current processing moment. Based on this change, it performs reverse compensation on the synchronously acquired multi-source identification information, effectively eliminating the dynamic distortion and offset of feature, depth, and ranging information caused by the continuous movement of the vehicle. This makes the input information received by the subsequent neural network model closer to the data state under static or ideal acquisition conditions, thereby significantly improving the accuracy and stability of target pallet pose prediction and enhancing the accuracy and reliability of the unmanned forklift in real-time operation during movement.

[0015] Furthermore, the neural network model includes an RGB branch corresponding to the feature information, a depth branch corresponding to the scene depth information, and a radar branch corresponding to the ranging information. The neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch through a fusion encoder to form a depth feature tensor. The RGB branch uses a first convolutional neural network to extract visual feature maps related to the tray from the two-dimensional color image; The deep branch uses a second convolutional neural network to extract tray-related geometric feature maps from the binocular parallax image; The radar branch uses a third convolutional neural network to extract a distance feature map related to the tray from the point cloud image.

[0016] It should be noted that this invention constructs a dedicated neural network model that includes RGB, depth, and radar branches, and uses a fusion encoder to deeply fuse the extracted visual features, geometric features, and distance features. This enables the model to comprehensively utilize multi-dimensional information such as texture, spatial structure, and precise distance for joint reasoning, thereby significantly enhancing the model's ability to perceive pallet pose and its robustness in complex environments. It effectively avoids errors caused by missing or interference from a single sensor, and ultimately greatly improves the accuracy of pallet pose estimation and the success rate of unmanned forklift operations.

[0017] Furthermore, before inputting the pose compensation recognition information into the pre-trained neural network model, the method further includes: Based on the current characteristics of the posture compensation recognition information, the signal-to-noise ratio of each posture compensation recognition information is evaluated in real time; Based on the signal-to-noise ratio, real-time weights are assigned to the feature information, the scene depth information, and the ranging information, respectively. The neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch using a fusion encoder to form a depth feature tensor, including: The neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch using a fusion encoder and real-time weights to form a depth feature tensor.

[0018] It should be noted that this invention evaluates the signal-to-noise ratio of each posture compensation recognition information in real time and dynamically allocates fusion weights accordingly. This enables the neural network model to adaptively adjust its dependence on data from different sources (vision, depth, radar). When environmental interference causes a decline in the quality of data from a certain type of sensor, its contribution is automatically reduced, while the weight of more reliable data sources is increased. This ensures that multi-source feature fusion is always based on the most reliable information, significantly enhancing the adaptability and anti-interference ability of the pose estimation system in dynamically changing environments. Ultimately, this improves the stability and success rate of unmanned forklifts operating under various complex conditions.

[0019] This specification provides one or more embodiments of a pallet handling device for an unmanned forklift, comprising: The acquisition unit acquires identification information related to the target pallet when an unmanned forklift operates on the target pallet in the logistics park. The measurement unit measures the speed information of the unmanned forklift in real time via an IMU; The compensation unit performs reverse compensation on the identification information based on the velocity information to obtain attitude compensation identification information; The prediction unit inputs the posture compensation recognition information into a pre-trained neural network model to predict the pose data of the target tray. The work unit converts the pose data into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

[0020] This specification provides one or more embodiments of a pallet handling device for an unmanned forklift, comprising: At least one processor and bus; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: When an unmanned forklift operates on a target pallet in a logistics park, it acquires identification information related to the target pallet. The speed information of the unmanned forklift is measured in real time using an IMU; Based on the velocity information, the recognition information is reverse-compensated to obtain attitude compensation recognition information; The posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray; The pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

[0021] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, which, when executed by a computer, can perform the following: When an unmanned forklift operates on a target pallet in a logistics park, it acquires identification information related to the target pallet. The speed information of the unmanned forklift is measured in real time using an IMU; Based on the velocity information, the recognition information is reverse-compensated to obtain attitude compensation recognition information; The posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray; The pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

[0022] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: This invention introduces an IMU to measure the speed information of the unmanned forklift in real time and uses this information to perform reverse compensation on the recognition information, effectively eliminating the image acquisition deviation caused by the movement of the vehicle. Subsequently, the compensated information is input into a neural network model for pose prediction and transformed into the forklift coordinate system, so that the final generated control commands can accurately correspond to the actual spatial posture of the pallet. This significantly improves the accuracy and reliability of pallet recognition and forklift operation of the unmanned forklift in dynamic working environments and reduces the risk of operation failure and safety accidents. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A flowchart illustrating a pallet handling method using an unmanned forklift, provided for one or more embodiments of this specification; Figure 2 A schematic diagram of the structure of a pallet handling device for an unmanned forklift provided for one or more embodiments of this specification; Figure 3 This is a structural schematic diagram of a pallet handling device for an unmanned forklift provided for one or more embodiments of this specification. Detailed Implementation

[0024] This specification provides an embodiment of a pallet handling method, apparatus, equipment, and medium for an unmanned forklift.

[0025] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0026] Figure 1 This diagram illustrates a process flow for one or more embodiments of a pallet handling method using an unmanned forklift, which can be executed by a pallet handling system. Certain input parameters or intermediate results in the process can be manually adjusted to help improve accuracy.

[0027] The method flow steps of the embodiments in this specification are as follows: S101, when an unmanned forklift is operating on a target pallet in a logistics park, the identification information related to the target pallet is obtained.

[0028] In the embodiments described in this specification, a multi-sensor system can be deployed on an unmanned forklift. This system includes an RGB color camera, a binocular depth camera, and a millimeter-wave radar. A synchronization controller sends a unified hardware trigger pulse signal to these three sensors to ensure that they acquire data simultaneously. The RGB color camera acquires two-dimensional color images of the target pallet and its environment; the binocular depth camera acquires binocular parallax images to calculate scene depth information; and the millimeter-wave radar acquires point cloud data of the target pallet to obtain ranging information. Subsequently, by using a multi-sensor extrinsic parameter matrix determined in advance using a joint calibration algorithm, the feature information, scene depth information, and ranging information from different sensors are uniformly transformed into the same coordinate system, forming a set of time-synchronized and spatially aligned multi-source recognition information.

[0029] S102, the speed information of the unmanned forklift is measured in real time via IMU.

[0030] In the embodiments described in this specification, an inertial measurement unit (IMU) fixedly mounted on the body of the unmanned forklift can be used to continuously measure and output raw data of the linear acceleration and angular velocity of the unmanned forklift in three-dimensional space at a fixed high frequency. This data is sent to the host computer controller of the unmanned forklift in real time, where the built-in algorithm of the controller filters and preprocesses the raw data to calculate the real-time speed, angular velocity, and other key motion state information of the unmanned forklift.

[0031] S103, based on the speed information, reverse compensation is performed on the recognition information to obtain attitude compensation recognition information.

[0032] In the embodiments described in this specification, the host computer controller receives velocity information from the IMU and multi-source recognition information from S101. Based on the time interval between two sensor synchronization hardware triggers, the controller integrates the velocity information continuously measured by the IMU during this period to estimate the attitude changes (including displacement and rotation) of the unmanned forklift during this short time. Subsequently, based on this estimated attitude change, the controller performs a mathematical inverse transformation and compensation on the feature information (such as key points in the image), scene depth information, and ranging information acquired in S101 at the current moment. This compensates for errors such as image blurring and point cloud shift caused by vehicle movement, generating a set of attitude compensation recognition information corrected for motion distortion.

[0033] S104, the posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray.

[0034] In the embodiments of this specification, the pose compensation recognition information obtained in S103 (i.e., motion-compensated color image, depth information, and radar point cloud) is input into a pre-trained multi-branch neural network model. The RGB branch of this model processes the color image to extract texture features, the depth branch processes the depth information to extract geometric features, and the radar branch processes the point cloud to extract distance features. The feature maps extracted by each branch are weighted and fused by a fusion encoder, and finally, the model output layer predicts the six-degree-of-freedom pose data (including three-dimensional position and three-dimensional attitude) of the target pallet in a coordinate system with a fixed reference point as the origin.

[0035] S105, the pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

[0036] In the embodiments of this specification, the target pallet pose data predicted by the neural network in S104 is acquired. This data is based on a preset fixed coordinate system. Utilizing the real-time pose provided by the unmanned forklift's own positioning system (such as a SLAM system), the target pallet pose data is transformed into the current coordinate system of the unmanned forklift through a fixed coordinate transformation matrix, resulting in transformed pose data. The unmanned forklift's path planning and motion controller receives this transformed pose data, uses it as input, and generates precise control commands (including travel direction, speed, forklift lifting and tilting angle, etc.) to drive the unmanned forklift to perform the final pallet handling tasks such as picking up, transporting, or placing.

[0037] It should be noted that this invention introduces an IMU to measure the speed information of the unmanned forklift in real time and uses this information to perform reverse compensation on the recognition information, effectively eliminating the image acquisition deviation caused by the movement of the vehicle. Subsequently, the compensated information is input into a neural network model for pose prediction and transformed into the forklift coordinate system, so that the final generated control commands can accurately correspond to the actual spatial posture of the pallet. This significantly improves the accuracy and reliability of pallet recognition and forklift operation of the unmanned forklift in dynamic working environments and reduces the risk of operation failure and safety accidents.

[0038] Furthermore, the identification information includes the feature information of the target pallet, the scene depth information in the logistics park, and the ranging information of the target pallet. When acquiring the identification information related to the target pallet, a two-dimensional color image can be acquired using an RGB color camera to extract the feature information, which includes texture, color, and visual key point information; a binocular parallax image can be acquired using a binocular depth camera to extract the scene depth information; and a point cloud image of the target pallet can be acquired using millimeter-wave radar to extract the ranging information.

[0039] It should be noted that an RGB color camera, a binocular depth camera, and a millimeter-wave radar are installed on the unmanned forklift, and their physical installation is ensured to be secure. A central synchronization controller sends a unified hardware trigger pulse signal to these three sensors, forcing them to acquire data at the same precise moment. This step ensures that the raw data acquired from sensors with different physical characteristics are completely synchronized in time, laying the foundation for temporal consistency for subsequent data fusion. Based on the synchronization trigger signal, the sensors operate in parallel.

[0040] For the RGB color camera, a two-dimensional color image of the target tray and its surrounding environment is acquired. This image is mainly used for subsequent extraction of feature information such as texture, color, and visual key points of the target.

[0041] For a stereo depth camera, a pair of images is acquired simultaneously by its two cameras to form a stereo parallax image. This image pair is mainly used to calculate the depth of objects in the scene, thereby extracting scene depth information.

[0042] For millimeter-wave radar, electromagnetic waves are emitted and echoes are received to generate a point cloud image of the target tray. This image consists of a series of three-dimensional spatial points, each containing distance information, primarily used to extract high-precision ranging information.

[0043] The raw data (2D color images, binocular parallax images, and radar point clouds) acquired from different sensors and existing in their own independent coordinate systems are transformed using a pre-determined extrinsic parameter matrix determined through a joint calibration algorithm. This matrix defines the relative positions and attitude relationships between the sensors. Through this transformation, the data from all sensors are unified into a single coordinate system, ultimately forming a set of fused recognition information that is aligned in both time and space and includes feature information, scene depth information, and ranging information, for use in subsequent processing steps.

[0044] It should be noted that this invention integrates the advantages of three sensors—RGB color camera, binocular depth camera, and millimeter-wave radar—in texture perception, depth calculation, and accurate ranging to construct a multi-dimensional recognition information system. This allows the characteristics of the target pallet, environmental depth, and its own distance information to complement and verify each other, thus providing a more comprehensive and reliable data foundation for subsequent motion compensation and neural network pose prediction. This greatly enhances the system's perception robustness and accuracy in complex scenarios such as changes in lighting, missing textures, or occlusion, providing a solid perception guarantee for the precise control of unmanned forklifts.

[0045] Furthermore, after obtaining the identification information related to the target tray, a joint calibration algorithm can be used to determine the extrinsic matrix between the feature information, the scene depth information, and the ranging information, so as to transform the feature information, the scene depth information, or the ranging information to the same coordinate system. The extrinsic matrix includes the relative position and attitude relationship between the feature information, the scene depth information, and the ranging information.

[0046] It should be noted that before the unmanned forklift was put into use, data from an RGB color camera, a binocular depth camera, and a millimeter-wave radar were simultaneously collected in a known calibration site with specific calibration objects (such as a checkerboard calibration board or specific markers). A specialized joint calibration algorithm was used to process these multiple sets of data collected from different perspectives, each containing common calibration objects, to accurately calculate the relative position and attitude parameters between each pair of sensors. These parameters were ultimately integrated and expressed as a unified set of extrinsic parameter matrices. This set of matrices defines the spatial relationships between feature information (from the RGB camera), scene depth information (from the binocular camera), and ranging information (from the millimeter-wave radar).

[0047] During the actual operation of the unmanned forklift, whenever new raw recognition information in the coordinate system of each sensor is acquired through the aforementioned steps, the system calls the pre-calculated and stored extrinsic parameter matrices from those steps. Using these matrices, the feature information acquired by the RGB camera, the scene depth information acquired by the binocular depth camera, and the ranging information acquired by the millimeter-wave radar are all transformed into the same predefined unified coordinate system (e.g., a coordinate system established with the optical center of the binocular depth camera as the origin). This step ensures that the multi-source information used in subsequent processing has a consistent spatial reference, laying the foundation for effective data fusion and processing.

[0048] It should be noted that this invention, by introducing a joint calibration algorithm, accurately obtains the extrinsic parameter matrix between different sensor information, thereby unifying the originally heterogeneous and dispersed feature, depth, and ranging information in their respective coordinate systems to the same spatial reference. This effectively eliminates data fusion errors caused by differences in installation position and viewing angle of multi-source sensors, providing a highly consistent and spatially aligned reliable data foundation for subsequent motion compensation and pose prediction. This improves the accuracy and coordination of the entire perception system from the source, ultimately ensuring the accuracy and reliability of unmanned forklift operation decisions.

[0049] Furthermore, after acquiring the identification information related to the target tray, a unified hardware trigger pulse signal can be sent to the RGB color camera, the binocular depth camera, and the millimeter-wave radar through a synchronization controller, so that the RGB color camera, the binocular depth camera, and the millimeter-wave radar can acquire data at the same time.

[0050] It should be noted that a dedicated synchronization controller (or synchronization trigger unit) is integrated into the control system of the unmanned forklift. This controller is connected to the external trigger interfaces of the RGB color camera, binocular depth camera, and millimeter-wave radar via physical cables (such as GPIO lines or trigger cables) or a high-speed bus, forming a master-slave hardware trigger architecture. The firmware or drivers of the RGB color camera, binocular depth camera, and millimeter-wave radar are configured to set their data acquisition mode to "external trigger" or "slave mode." In this mode, each sensor no longer operates autonomously according to its internal clock but enters a waiting state, ready to receive instructions from the synchronization controller. When the unmanned forklift system needs to acquire sensing data, the master control system sends an instruction to the synchronization controller. Upon receiving the instruction, the synchronization controller immediately generates a unified hardware trigger pulse signal and sends it simultaneously (or within a very small time difference) to the RGB color camera, binocular depth camera, and millimeter-wave radar through its connected physical lines. The RGB color camera, binocular depth camera, and millimeter-wave radar immediately initiate a data acquisition process upon receiving the same hardware trigger pulse signal. An RGB color camera captures a frame of two-dimensional color image, a binocular depth camera captures a pair of photos used to generate a parallax image, and a millimeter-wave radar acquires a frame of point cloud image. This ensures that the acquired feature information, scene depth information, and ranging information have identical timestamps, achieving strictly simultaneous data acquisition.

[0051] It should be noted that this invention sends a unified hardware trigger pulse signal to the RGB color camera, binocular depth camera, and millimeter-wave radar through a synchronization controller, forcing all sensors to acquire data at the exact same moment. This fundamentally eliminates the data time difference and spatiotemporal misalignment caused by the asynchronous acquisition time of each sensor, ensuring that the acquired multi-source perception information (features, depth, and ranging) has a high degree of temporal consistency and spatial synchronization. This provides a precise time-aligned data foundation for subsequent multi-source data fusion, motion compensation, and pose calculation, thereby significantly improving the temporal consistency and pose estimation accuracy of the entire perception system and enhancing the operational reliability of the unmanned forklift in dynamic scenarios.

[0052] Furthermore, when performing reverse compensation on the identification information based on the speed information to obtain attitude compensation identification information, the attitude change can be estimated by integrating the motion of the unmanned forklift between two adjacent synchronous triggers based on the speed information; and the attitude change is used to perform reverse compensation on the feature information, scene depth information and ranging information at the current moment to obtain attitude compensation identification information.

[0053] It should be noted that the host computer controller receives real-time velocity information from the inertial measurement unit (IMU). The controller records the timestamps of two consecutive hardware trigger pulse signals issued by the synchronization controller and calculates the precise time interval between them. Within this time interval, the linear velocity and angular velocity data continuously measured by the IMU are integrated to estimate the displacement and rotation of the unmanned forklift itself from the last data acquisition moment to the current processing moment, i.e., the attitude change. The host computer controller acquires multi-source recognition information (i.e., feature information, scene depth information, and ranging information) acquired at the latest synchronization trigger moment and unified by coordinates. Subsequently, using the attitude change (displacement and rotation) calculated in the aforementioned steps, a mathematical inverse transformation is performed on this batch of recognition information. For the feature information (two-dimensional data derived from RGB color images), the compensation process is equivalent to performing an inverse affine transformation on the image according to the vehicle's movement to compensate for pixel shifts or blurring caused by the vehicle's movement. For scene depth information (derived from binocular depth cameras) and ranging information (derived from point cloud data from millimeter-wave radar), the compensation process is equivalent to performing a reverse rotation and translation transformation on each point in three-dimensional space to correct the coordinate deviation caused by the vehicle's movement.

[0054] Through the above processing, a set of attitude compensation recognition information that has eliminated the errors introduced by the vehicle's own movement since the acquisition time is finally output, which is closer to the ideal static acquisition state.

[0055] It should be noted that this invention utilizes the speed information measured in real time by the IMU to integrate the vehicle's own motion within the interval between two adjacent synchronous triggering events, accurately estimating the attitude change of the unmanned forklift from the data acquisition moment to the current processing moment. Based on this change, it performs reverse compensation on the synchronously acquired multi-source identification information, effectively eliminating the dynamic distortion and offset of feature, depth, and ranging information caused by the continuous movement of the vehicle. This makes the input information received by the subsequent neural network model closer to the data state under static or ideal acquisition conditions, thereby significantly improving the accuracy and stability of target pallet pose prediction and enhancing the accuracy and reliability of the unmanned forklift in real-time operation during movement.

[0056] Furthermore, the neural network model includes an RGB branch corresponding to the feature information, a depth branch corresponding to the scene depth information, and a radar branch corresponding to the ranging information. The neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch using a fusion encoder to form a depth feature tensor. The RGB branch uses a first convolutional neural network to extract visual feature maps related to the tray from the two-dimensional color image. The depth branch uses a second convolutional neural network to extract geometric shape feature maps related to the tray from the binocular parallax image. The radar branch uses a third convolutional neural network to extract distance feature maps related to the tray from the point cloud image.

[0057] It should be noted that the pose compensation recognition information obtained from S103 (i.e., the motion-compensated 2D color image, binocular parallax image, and radar point cloud image) can be input into three independent branches of the pre-trained neural network model for parallel processing: For the RGB branch processing, this branch contains a first convolutional neural network. Its input is a compensated two-dimensional color image, and through layers of convolution and pooling operations, it extracts high-dimensional visual feature maps such as texture, edges, and semantics related to the tray.

[0058] For the depth branch processing, this branch contains a second convolutional neural network. Its input is a compensated binocular parallax image or a depth map computed by it. This network is specifically designed to extract geometric feature maps representing the spatial relationship between the pallet and fork.

[0059] For the radar branch processing, this branch contains a third convolutional neural network. Its input is a compensated and coordinate-unified radar point cloud image. This network is used to learn the sparse structure of the point cloud and extract distance feature maps that accurately reflect the tray distance and contour.

[0060] The feature maps output from the three branches (visual feature map, geometric shape feature map, and distance feature map) are simultaneously input into a fusion encoder. This fusion encoder uses a specific network structure (such as convolution, attention mechanisms, etc.) to fuse and enhance feature maps from different perceptual modalities, ultimately generating a unified and information-rich deep feature tensor. This tensor integrates the target's appearance, geometry, and precise location information.

[0061] The fused deep feature tensor is input into the prediction layer (usually a fully connected layer or a convolutional layer) at the end of the neural network model. The prediction layer directly regresses and calculates the six-degree-of-freedom pose data of the target tray in the current unified coordinate system.

[0062] It should be noted that this invention constructs a dedicated neural network model that includes RGB, depth, and radar branches, and uses a fusion encoder to deeply fuse the extracted visual features, geometric features, and distance features. This enables the model to comprehensively utilize multi-dimensional information such as texture, spatial structure, and precise distance for joint reasoning, thereby significantly enhancing the model's ability to perceive pallet pose and its robustness in complex environments. It effectively avoids errors caused by missing or interference from a single sensor, and ultimately greatly improves the accuracy of pallet pose estimation and the success rate of unmanned forklift operations.

[0063] Furthermore, before inputting the pose compensation recognition information into the pre-trained neural network model, the signal-to-noise ratio of each pose compensation recognition information can be evaluated in real time based on the current characteristics of the pose compensation recognition information; real-time weights are assigned to the feature information, the depth information, and the ranging information based on the signal-to-noise ratio; when the neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch through the fusion encoder to form a depth feature tensor, the neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch through the fusion encoder in combination with the real-time weights to form a depth feature tensor.

[0064] It should be noted that before inputting the pose compensation and recognition information into the neural network model, the system first performs a real-time quality assessment of the three types of data it contains (feature information, scene depth information, and ranging information). This assessment is based on the inherent characteristics of each data point in the current environment (such as image sharpness and contrast, point cloud density and completeness), and calculates the real-time signal-to-noise ratio (SNR) for each type of information using a built-in evaluation algorithm. Subsequently, a dynamic real-time weight is assigned to the feature information, scene depth information, and ranging information according to their SNR levels. Information with a high SNR receives a higher weight, and information with a low SNR receives a lower weight, thus achieving adaptive importance allocation.

[0065] When the fusion encoder of a neural network model performs multi-source feature fusion, its fusion strategy incorporates the real-time weights calculated in the preceding steps. When integrating the visual feature map output from the RGB branch, the geometric feature map output from the depth branch, and the distance feature map output from the radar branch, the fusion encoder performs weighted fusion processing on the different feature maps according to the real-time weights corresponding to their respective information sources. This means that high-weight feature maps will dominate in the fused depth feature tensor, while the influence of low-weight feature maps will be correspondingly weakened, thus ensuring that the fusion result relies more on the most reliable data source in the current environment.

[0066] It should be noted that this invention evaluates the signal-to-noise ratio of each posture compensation recognition information in real time and dynamically allocates fusion weights accordingly. This enables the neural network model to adaptively adjust its dependence on data from different sources (vision, depth, radar). When environmental interference causes a decline in the quality of data from a certain type of sensor, its contribution is automatically reduced, while the weight of more reliable data sources is increased. This ensures that multi-source feature fusion is always based on the most reliable information, significantly enhancing the adaptability and anti-interference ability of the pose estimation system in dynamically changing environments. Ultimately, this improves the stability and success rate of unmanned forklifts operating under various complex conditions.

[0067] Figure 2 This specification provides a schematic diagram of the structure of a pallet operating device for an unmanned forklift, which includes one or more embodiments: an acquisition unit 201, a measurement unit 202, a compensation unit 203, a prediction unit 204, and an operating unit 205.

[0068] The acquisition unit 201 acquires identification information related to the target pallet when the unmanned forklift is operating on the target pallet in the logistics park. Measurement unit 202 measures the speed information of the unmanned forklift in real time via IMU; The compensation unit 203 performs reverse compensation on the identification information based on the speed information to obtain attitude compensation identification information; The prediction unit 204 inputs the posture compensation recognition information into a pre-trained neural network model to predict the pose data of the target tray. The work unit 205 converts the pose data into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

[0069] Figure 3 A schematic diagram of the structure of a pallet handling device for an unmanned forklift, provided for one or more embodiments of this specification, includes: At least one processor and bus; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: When an unmanned forklift operates on a target pallet in a logistics park, it acquires identification information related to the target pallet. The speed information of the unmanned forklift is measured in real time using an IMU; Based on the velocity information, the recognition information is reverse-compensated to obtain attitude compensation recognition information; The posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray; The pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

[0070] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, which, when executed by a computer, can perform the following: When an unmanned forklift operates on a target pallet in a logistics park, it acquires identification information related to the target pallet. The speed information of the unmanned forklift is measured in real time using an IMU; Based on the velocity information, the recognition information is reverse-compensated to obtain attitude compensation recognition information; The posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray; The pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

[0071] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0072] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0073] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0074] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0075] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0076] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The aforementioned units can be implemented in hardware or software.

[0077] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0078] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for pallet handling using an unmanned forklift, characterized in that, The method includes: When an unmanned forklift operates on a target pallet in a logistics park, it acquires identification information related to the target pallet. The speed information of the unmanned forklift is measured in real time using an IMU; Based on the velocity information, the recognition information is reverse-compensated to obtain attitude compensation recognition information; The posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray; The pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

2. The method according to claim 1, characterized in that, The identification information includes the feature information of the target pallet, the scene depth information in the logistics park, and the distance measurement information of the target pallet; The acquisition of the identification information related to the target tray includes: Two-dimensional color images are acquired using an RGB color camera to extract the feature information, which includes texture, color, and visual key point information. Binocular parallax images are acquired using a binocular depth camera to extract the scene depth information; The point cloud image of the target tray is acquired using millimeter-wave radar to extract the ranging information.

3. The method according to claim 2, characterized in that, After obtaining the identification information related to the target tray, the method further includes: The extrinsic parameter matrix between the feature information, the scene depth information, and the ranging information is determined by a joint calibration algorithm, so as to transform the feature information, the scene depth information, or the ranging information to the same coordinate system. The extrinsic parameter matrix includes the relative position and attitude relationship between the feature information, the scene depth information, and the ranging information.

4. The method according to claim 2, characterized in that, After obtaining the identification information related to the target tray, the method further includes: The synchronization controller sends a unified hardware trigger pulse signal to the RGB color camera, the binocular depth camera, and the millimeter-wave radar to acquire data from the RGB color camera, the binocular depth camera, and the millimeter-wave radar at the same time.

5. The method according to claim 4, characterized in that, The step of performing reverse compensation on the recognition information based on the velocity information to obtain attitude compensation recognition information includes: Based on the speed information, the motion of the unmanned forklift between two adjacent synchronous triggers is integrated to estimate the attitude change; Based on the posture change, the feature information, scene depth information and ranging information at the current moment are inversely compensated to obtain posture compensation recognition information.

6. The method according to claim 2, characterized in that, The neural network model includes an RGB branch corresponding to the feature information, a depth branch corresponding to the scene depth information, and a radar branch corresponding to the ranging information. The neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch through a fusion encoder to form a depth feature tensor. The RGB branch uses a first convolutional neural network to extract visual feature maps related to the tray from the two-dimensional color image; The deep branch uses a second convolutional neural network to extract tray-related geometric feature maps from the binocular parallax image; The radar branch uses a third convolutional neural network to extract a distance feature map related to the tray from the point cloud image.

7. The method according to claim 6, characterized in that, Before inputting the pose compensation recognition information into the pre-trained neural network model, the method further includes: Based on the current characteristics of the posture compensation recognition information, the signal-to-noise ratio of each posture compensation recognition information is evaluated in real time; Based on the signal-to-noise ratio, real-time weights are assigned to the feature information, the scene depth information, and the ranging information, respectively. The neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch using a fusion encoder to form a depth feature tensor, including: The neural network model fuses the feature maps of the RGB branch, the depth branch, and the radar branch using a fusion encoder and real-time weights to form a depth feature tensor.

8. A pallet handling device for an unmanned forklift, characterized in that, include: The acquisition unit acquires identification information related to the target pallet when an unmanned forklift operates on the target pallet in the logistics park. The measurement unit measures the speed information of the unmanned forklift in real time via an IMU; The compensation unit performs reverse compensation on the identification information based on the velocity information to obtain attitude compensation identification information; The prediction unit inputs the posture compensation recognition information into a pre-trained neural network model to predict the pose data of the target tray. The work unit converts the pose data into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

9. A pallet handling device for an unmanned forklift, characterized in that, include: At least one processor and bus; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: When an unmanned forklift operates on a target pallet in a logistics park, it acquires identification information related to the target pallet. The speed information of the unmanned forklift is measured in real time using an IMU; Based on the velocity information, the recognition information is reverse-compensated to obtain attitude compensation recognition information; The posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray; The pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.

10. A non-volatile computer storage medium, characterized in that, It stores computer-executable instructions, which, when executed by a computer, can achieve the following: When an unmanned forklift operates on a target pallet in a logistics park, it acquires identification information related to the target pallet. The speed information of the unmanned forklift is measured in real time using an IMU; Based on the velocity information, the recognition information is reverse-compensated to obtain attitude compensation recognition information; The posture compensation recognition information is input into a pre-trained neural network model to predict the pose data of the target tray; The pose data is converted into the coordinate system of the unmanned forklift to obtain converted pose data, so that the unmanned forklift can generate corresponding control commands based on the converted pose data and complete the operation of the target pallet through the control commands.