Machine vision-based automatic unloading method and system
By fusing multi-view images and point cloud data with dynamic modeling, the accuracy problems of pallet recognition and pose estimation in unmanned unloading were solved, enabling stable grasping and safe path planning in complex environments, thus improving the efficiency and safety of unloading operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LONGHE INTELLIGENT EQUIP MFG CO LTD
- Filing Date
- 2026-03-18
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies in the field of unmanned unloading face problems such as low pallet recognition accuracy, pose estimation deviation, and inaccurate path obstacle avoidance in complex unloading environments, leading to operation interruptions, grasping failures, and equipment collisions. In particular, they lack environmental adaptability and dynamic compensation capabilities in unstructured scenarios.
A multi-view industrial camera array is used to simultaneously acquire visible light images and depth point cloud data. Through multi-stage image preprocessing and a dual-stream feature extraction network, combined with a prior knowledge base of pallet structure and dynamic operation plane modeling, highly robust identification and 3D pose calculation of pallet targets are achieved. A hierarchical path planning and force control prediction mechanism is constructed to ensure the reliability and safety of grasping.
Achieving stable identification and accurate grasping of pallet targets in complex unloading environments improves the success rate and safety of unloading operations, shortens the single cycle time to within 45 seconds, adapts to any vehicle type, stacking pattern and membrane coverage, and significantly improves the throughput efficiency and safety of the logistics system.
Smart Images

Figure CN121871998B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence, and specifically relates to a method and system for unmanned unloading automation based on machine vision. Background Technology
[0002] With the popularization and upgrading of automated warehousing and logistics systems in industries such as petrochemicals, chemicals, grain and oil, and beverages, the outbound efficiency of palletized materials and the level of automation in loading and unloading have become key bottlenecks restricting the response speed of the supply chain. Traditional manual loading and unloading operations rely on high-intensity human input and have inherent defects such as low efficiency, high safety risks, and poor operational consistency, making them difficult to meet the core requirements of modern intelligent warehousing for high throughput, high flexibility, and high reliability. Although automated loading technology has been applied on a large scale and is becoming mature in some scenarios, automated solutions for vehicle unloading are still lacking in the industry. In particular, guiding handling robots to complete high-precision pallet unloading in complex unloading environments through visual algorithms has become a core technical challenge restricting the implementation of closed-loop logistics automation.
[0003] Among them, the machine vision-based unmanned unloading automated operation system aims to achieve fully unmanned operation of autonomously identifying, locating, and grabbing pallets from transport vehicles through a closed loop of perception-decision-execution. The core challenge of this technology lies in how to stably extract the spatial coordinates of the pallet fork holes in highly unstructured unloading scenarios to guide the mechanical actuators to complete collision-free and interference-free precise docking. Its basic principle relies on multimodal sensor data fusion and geometric feature extraction to provide real-time and reliable operational benchmarks for loading and unloading equipment by constructing a three-dimensional spatial mapping of the pallet hole positions.
[0004] Existing technologies exhibit multiple systemic defects when dealing with actual unloading conditions: First, there is no uniform standard for vehicle type and pallet placement, and the truck bed often presents non-planar shapes such as front-high-rear-low, local concavity, or tilt, causing traditional visual positioning methods based on fixed coordinate systems to fail; second, the pallet surface is generally covered by plastic wrapping film, with blurred hole edges and light reflection interference, resulting in high false detection rates and poor robustness of edge detection algorithms based on single RGB images; third, adjacent pallets form asymmetrical gaps due to inertial displacement during transportation, and without dynamic offset compensation, it is very easy to cause interference between the forks and pallet structure, resulting in equipment jamming or cargo damage.
[0005] Finally, existing solutions lack real-time assessment of the confidence level of hole location recognition and an anomaly alarm mechanism. They fail to trigger safety redundancy strategies in scenarios where key features are missing or occluded, seriously threatening operational continuity and system security. These problems are amplified dramatically in high-density, high-frequency unloading operations across multiple industries, necessitating a visual guidance solution for unmanned unloading that possesses strong environmental adaptability, high recognition accuracy, and dynamic compensation capabilities. Summary of the Invention
[0006] This invention provides a machine vision-based unmanned unloading automatic operation method and system. By constructing a multimodal visual perception architecture and an adaptive spatial modeling mechanism, it achieves highly robust recognition of pallet targets, accurate 3D pose calculation, and dynamic path autonomous planning in unstructured unloading scenarios. Thus, even under complex working conditions such as irregular fluctuations in vehicle plane height, pallet surfaces covered with wrapping film or plastic film, and pallet holes being obscured, it stably guides the handling robot to complete pallet grasping and transfer operations. This completely solves the problems of operation interruption, grasping failure, and equipment collision caused by environmental perception failure, pose estimation deviation, and inaccurate path obstacle avoidance in the field of unmanned unloading in the existing technology.
[0007] This invention provides a machine vision-based unmanned automated unloading operation method, comprising:
[0008] A multi-view industrial camera array deployed above the unloading area is used to simultaneously collect visible light image sequences and depth point cloud data inside the cargo compartment of the vehicle to be unloaded. The industrial camera array contains at least three imaging units, which cover the entire cargo compartment from top-down, front-oblique, and side-oblique angles, respectively. The imaging units achieve microsecond-level time synchronization through hardware trigger signals.
[0009] Multi-stage image preprocessing is performed on the acquired visible light image sequence, including using an adaptive histogram equalization algorithm to enhance image contrast, applying the Gaussian-Laplacian operator for edge sharpening, filling edge breakage areas through morphological closing operations, and finally outputting image data with enhanced texture features.
[0010] Spatial filtering and coordinate normalization are performed on the collected depth point cloud data. First, a statistical outlier removal algorithm is used to remove noise points. Then, data density is reduced by voxel grid downsampling. Finally, the point cloud coordinate system is transformed to a unified world coordinate system with the lower left corner of the inner side of the rear tailgate of the vehicle cargo box as the origin.
[0011] A dual-stream feature extraction network is constructed, wherein the first branch receives the preprocessed visible light image and uses an improved residual convolutional neural network to extract the surface texture features and contour edge features of the tray; the second branch receives the normalized depth point cloud data and uses a point cloud transformation network to extract the three-dimensional geometric structure features and spatial distribution features of the tray; the dual-stream feature extraction network performs cross-modal attention fusion at the feature layer to generate a tray candidate region feature map that integrates visual and geometric semantics.
[0012] Based on the fused feature map, a joint task of tray instance segmentation and pose regression is performed. A region proposal network is used to generate tray candidate boxes. Then, a fully convolutional segmentation head is used to output a pixel-level tray mask. At the same time, a rotation-invariant pose regression head is used to predict the three-dimensional coordinates, length, width and height parameters and yaw angle around the vertical axis of the tray center point. Finally, the six-degree-of-freedom pose parameters of the tray in the world coordinate system are output.
[0013] To address the issue of pallet holes being partially obscured by stretch film, making gripping points invisible, a prior knowledge base for pallet structure is constructed. This knowledge base stores a template for the distribution of holes in a standard pallet and a topological map of the load-bearing area. When the vision system cannot directly detect holes, the theoretical location of the holes is calculated by matching the current pallet outer contour dimensions with the knowledge base template. The feasibility of gripping at this location is then verified by combining the local curvature analysis of the depth point cloud, thereby generating virtual gripping point coordinates.
[0014] A dynamic operation plane modeling module is established. A three-dimensional surface equation is generated by fitting the point cloud data of the cargo box floor. The surface equation is expressed by a cubic B-spline surface. Its control vertex is determined by interpolation of the measured elevation values of the four corners and the central area of the cargo box, thereby accurately describing the deformation state of the cargo box floor that is higher in the front and lower in the back or has local concavity.
[0015] Based on the pallet pose parameters and the working plane surface equation, the actual contact area between the bottom of the pallet and the cargo box floor is calculated, and the initial approach height and gripping tilt angle of the end effector of the handling robot are adjusted accordingly to ensure that the gripping action has a 5mm safety margin in the vertical direction and is aligned with the center line of the pallet load-bearing beam in the horizontal direction.
[0016] A hierarchical path planning architecture is constructed. The upper layer uses an improved A* algorithm to search for a global collision-free path from the robot base to the pallet gripping point in a 3D grid map. The grid map is generated by voxelizing the cargo box point cloud data, and obstacle areas are marked as impassable. The lower layer uses a dynamic window method to adjust the robot joint trajectory in real time to avoid temporary obstacles caused by pallet stacking shaking or slight vehicle displacement during path execution.
[0017] Before the handling robot performs the grasping action, the force control prediction module is activated. The module calculates the minimum clamping force required for grasping based on the estimated value of the pallet weight and the empirical value of the friction coefficient of the cargo box floor. The six-dimensional force sensor at the end of the robot monitors the change of clamping force in real time. When the clamping force fluctuation is detected to exceed the preset threshold, the grasping posture fine adjustment command is immediately triggered to compensate for the loss of grasping force caused by the slippage of the stretch film or the deformation of the pallet.
[0018] After the pallet is picked up, a dynamic center of gravity compensation algorithm is activated. The algorithm adjusts the torque output of the wrist joint of the handling robot in real time according to the center of gravity offset caused by uneven distribution of materials in the pallet, so as to ensure that the pallet remains in a horizontal position during transportation and avoid the materials from tipping over.
[0019] Remove the coordinates of the unloaded pallets from the work map and trigger the vision system to rescan and update the pose of the remaining pallet stacks, then enter the next unloading cycle until all pallets in the cargo compartment are emptied.
[0020] This invention provides an automated unmanned unloading system based on machine vision, comprising:
[0021] The multi-view synchronous imaging module is used to simultaneously acquire visible light image sequences and depth point cloud data inside the cargo compartment of the vehicle to be unloaded by using a multi-view industrial camera array deployed above the unloading operation area. The industrial camera array contains at least three imaging units, which cover the entire cargo compartment from top-down, front-oblique, and side-oblique angles, respectively. The imaging units achieve microsecond-level time synchronization through hardware trigger signals.
[0022] The image and point cloud preprocessing module performs multi-stage image preprocessing on the acquired visible light image sequence. This includes enhancing image contrast using an adaptive histogram equalization algorithm, sharpening edges using the Gaussian-Laplacian operator, filling edge breakage areas through morphological closing operations, and finally outputting image data with enhanced texture features. The module also performs spatial filtering and coordinate normalization on the acquired depth point cloud data. First, it removes noise points using a statistical outlier removal algorithm, then reduces data density through voxel grid downsampling, and finally transforms the point cloud coordinate system to a unified world coordinate system with the lower left corner of the inner side of the rear tailgate of the vehicle cargo box as the origin.
[0023] The dual-stream feature fusion and recognition module is used to construct a dual-stream feature extraction network. The first branch receives the preprocessed visible light image and uses an improved residual convolutional neural network to extract the surface texture features and contour edge features of the tray. The second branch receives the normalized depth point cloud data and uses a point cloud transformation network to extract the three-dimensional geometric structure features and spatial distribution features of the tray. The dual-stream feature extraction network performs cross-modal attention fusion at the feature layer to generate a tray candidate region feature map that fuses visual and geometric semantics.
[0024] The pose calculation and virtual grasping point generation module is used to perform joint tasks of pallet instance segmentation and pose regression based on fused feature maps. It uses a region proposal network to generate pallet candidate boxes, and then outputs a pixel-level pallet mask through a fully convolutional segmentation head. At the same time, it uses a rotation-invariant pose regression head to predict the 3D coordinates, length, width, and height parameters of the pallet center point and the yaw angle around the vertical axis. Finally, it outputs the six-degree-of-freedom pose parameters of the pallet in the world coordinate system. In order to address the situation where the grasping point is not visible due to the partial occlusion of the pallet holes by the wrapping film, a prior knowledge base of pallet structure is constructed. The knowledge base stores the hole distribution template of standard pallets and the topology map of the load-bearing area. When the vision system cannot directly detect the hole, the theoretical position of the hole is estimated by matching the current pallet outer contour size with the knowledge base template. The feasibility of grasping at the position is verified by combining the local curvature analysis of the depth point cloud, thereby generating the coordinates of the virtual grasping point.
[0025] The dynamic operation plane modeling module is used to establish a dynamic operation plane modeling module. It generates a three-dimensional surface equation by fitting the point cloud data of the cargo box floor. The surface equation is expressed by a cubic B-spline surface. Its control vertices are determined by interpolation of the measured elevation values of the four corners and the central area of the cargo box, thereby accurately describing the deformation state of the cargo box floor that is higher in the front and lower in the back or has local concavity.
[0026] The adaptive adjustment module for gripping parameters is used to calculate the actual contact area between the bottom of the pallet and the floor of the cargo box based on the pallet pose parameters and the working plane surface equation, and adjust the initial approach height and gripping tilt angle of the end effector of the handling robot accordingly, so as to ensure that the gripping action has a safety margin of five millimeters in the vertical direction and is aligned with the center line of the pallet load-bearing beam in the horizontal direction.
[0027] The hierarchical path planning module is used to construct a hierarchical path planning architecture. The upper layer uses an improved A* algorithm to search for a globally collision-free path from the robot base to the pallet gripping point in a 3D grid map. The grid map is generated by voxelizing the cargo box point cloud data, and obstacle areas are marked as impassable. The lower layer uses a dynamic window method to adjust the robot joint trajectory in real time to avoid temporary obstacles caused by pallet stacking swaying or slight vehicle displacement during path execution.
[0028] The force control prediction and center of gravity compensation module is used to activate the force control prediction module before the handling robot performs the grasping action. The module calculates the minimum clamping force required for grasping based on the estimated pallet weight and the empirical value of the friction coefficient of the cargo box floor. It also monitors the changes in clamping force in real time through a six-dimensional force sensor at the end of the robot. When the clamping force fluctuation exceeds a preset threshold, it immediately triggers a fine-tuning command for the grasping posture to compensate for the loss of grasping force caused by the slippage of the stretch film or the deformation of the pallet. After the pallet is grasped, the dynamic center of gravity compensation algorithm is activated. The algorithm adjusts the torque output of the wrist joint of the handling robot in real time according to the center of gravity offset caused by the uneven distribution of materials in the pallet, ensuring that the pallet maintains a horizontal posture during transportation and avoiding the tipping of materials.
[0029] The operation cycle control module is used to remove the coordinates of unloaded pallets from the operation map and trigger the vision system to rescan and update the pose of the remaining pallet stacking status, and enter the next unloading cycle until all pallets in the cargo compartment are emptied.
[0030] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0031] By constructing a multi-view synchronous imaging and dual-stream feature fusion architecture, the defect of feature loss of a single sensor under strong reflective film interference was overcome, and stable identification of pallet targets under the condition of wrapping film cover was achieved.
[0032] By introducing a prior knowledge base of tray structure and a virtual gripping point generation mechanism, the problem of gripping point positioning failure caused by invisible holes was solved.
[0033] By establishing a dynamic working plane model based on a cubic B-spline surface, the gripping height error caused by unevenness of the vehicle floor was accurately compensated.
[0034] Through a hierarchical path planning and force control prediction coordination mechanism, the robot's motion safety and grasping reliability in dynamic obstacle environments are ensured.
[0035] The real-time center of gravity compensation algorithm avoids the risk of pallets tipping over during handling due to non-uniform loads.
[0036] The entire system can adapt to unloading scenarios with any vehicle type, any stacking configuration, and any membrane material coverage without human intervention, improving the success rate of unloading operations and reducing the single unloading cycle time to less than 45 seconds, significantly improving the throughput efficiency and operational safety of the logistics automation system. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the overall technical solution architecture of the unmanned unloading automatic operation method and system based on machine vision proposed in this invention;
[0038] Figure 2 This is a schematic diagram illustrating the core principle framework of the dual-stream feature fusion and cross-modal attention mechanism in this invention;
[0039] Figure 3 This is a flowchart illustrating the logical flow of pallet pose calculation and virtual gripping point generation in this invention.
[0040] Figure 4 This is a schematic diagram of the multi-level interaction relationship and data flow of dynamic operation plane modeling and hierarchical path planning in this invention;
[0041] Figure 5 This is a logical flowchart of the coordinated control of force control prediction and center of gravity compensation in this invention; Detailed Implementation
[0042] Example 1: Please refer to Figure 1-5 This invention provides an automated unmanned unloading operation method based on machine vision. Its core lies in using a multi-view industrial camera array deployed above the unloading operation area to simultaneously collect visible light image sequences and depth point cloud data of the cargo compartment of the vehicle to be unloaded. The industrial camera array contains at least three imaging units, which cover the entire cargo compartment from top-down, front-oblique, and side-oblique angles, respectively. The imaging units achieve microsecond-level time synchronization through hardware trigger signals.
[0043] In actual deployment, the physical installation positions of the three imaging units were precisely calibrated to ensure that their fields of view had no blind spots or overlap within the cargo compartment. The angles between the optical axes of each camera and the cargo compartment floor plane were 90°, 45°, and 45° respectively, maximizing the acquisition of information on the pallet's top, sides, and edges. Hardware trigger signals were uniformly issued by a central timing controller with a trigger interval of 33mm, corresponding to a sampling frequency of 30 frames per second. All cameras completed exposure and data latching within 5μs after receiving the trigger pulse, ensuring strict temporal alignment of multi-view data. The acquired visible light images had a resolution of 1920×1080 pixels, and the spatial resolution of the depth point cloud data was 0.5mm, with an effective measurement range of 0.5m to 5m, meeting the full coverage requirements of the cargo compartment's depth and stacking height.
[0044] To address the degradation in imaging quality caused by strong light reflection and film material shading, each imaging unit is equipped with a polarizing filter and an adjustable aperture lens. The polarization direction is preset according to the distribution of ambient light sources, and the aperture opening is adjusted in real time by the ambient light sensor to ensure that clear contrast and well-suppressed noise original images and point cloud data can be obtained under different lighting conditions.
[0045] Multi-stage image preprocessing is performed on the acquired visible light image sequence, including using an adaptive histogram equalization algorithm to enhance image contrast, applying the Gaussian-Laplacian operator for edge sharpening, filling edge breakage areas through morphological closing operations, and finally outputting image data with enhanced texture features. The adaptive histogram equalization algorithm divides the image into 16×16 local regions, calculates the cumulative distribution function independently for each region, and smooths the gray-level transition between adjacent regions through bilinear interpolation to avoid blocky artifacts.
[0046] The Gaussian-Laplacian operator uses a Gaussian kernel with a standard deviation of 1.5 pixels and a Laplacian kernel for convolution, highlighting high-frequency details of the tray edges and membrane wrinkles while suppressing low-frequency background interference. The morphological closing operation uses a circular structuring element with a radius of 3 pixels, first performing a dilation operation to connect broken edges, and then performing an erosion operation to restore the original contour width, ensuring the continuous closure of the tray boundary.
[0047] The preprocessed image data is cropped to the effective working area, irrelevant background from the cargo box panels and vehicle frame structure is removed, and the data is normalized to a floating-point value range of 0 to 1 as the input tensor for the subsequent neural network. This preprocessing process is executed in a pipelined manner in the embedded image processing unit, with a single-frame processing latency of less than 20ms, meeting real-time requirements.
[0048] Spatial filtering and coordinate normalization were performed on the collected depth point cloud data. First, a statistical outlier removal algorithm was used to remove noise points. Then, voxel grid downsampling was used to reduce the data density. Finally, the point cloud coordinate system was transformed to a unified world coordinate system with the lower left corner of the inner side of the rear tailgate of the vehicle cargo box as the origin. The statistical outlier removal algorithm calculates the average distance of the 30 nearest neighbors of each point. If the distance deviates from the global mean by more than two standard deviations, it is identified as an outlier and removed. Voxel grid downsampling divides the 3D space into a cubic grid with a side length of 20mm. Each grid retains the point closest to the geometric center, reducing the point cloud density from 500,000 points per cubic meter to 12,500 points per cubic meter, significantly reducing the subsequent computational load.
[0049] Coordinate normalization relies on a pre-calibrated origin marker for the cargo compartment coordinate system. This marker is a high-reflectivity circular target fixed to the lower left corner of the inner side of the rear baffle. Before each operation, the vision system locates the center of this target through template matching and constructs a right-handed Cartesian coordinate system based on it. The X-axis points forward along the length of the cargo compartment, the Y-axis points to the right along the width, and the Z-axis points vertically upward. All point cloud data is mapped to this coordinate system through a rigid body transformation matrix. The transformation matrix parameters are pre-determined and stored in the hand-eye calibration process to ensure the consistency and traceability of the spatial data.
[0050] A dual-stream feature extraction network is constructed, wherein the first branch receives the preprocessed visible light image and uses an improved residual convolutional neural network to extract the surface texture features and contour edge features of the tray; the second branch receives the normalized depth point cloud data and uses a point cloud transformation network to extract the three-dimensional geometric structure features and spatial distribution features of the tray. The dual-stream feature extraction network performs cross-modal attention fusion at the feature layer to generate a tray candidate region feature map that integrates visual and geometric semantics.
[0051] The improved residual convolutional neural network is based on the ResNet50 architecture. It removes the last global average pooling layer, retaining the multi-scale feature maps output by the five levels of residual blocks. The spatial resolutions of these residual blocks are 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input size, respectively. The point cloud transformation network uses a PointNet architecture, comprising three multilayer perceptron layers and max pooling operations. It outputs a 1024-dimensional global feature vector and restores it to the same spatial resolution as the image feature map through deconvolution upsampling.
[0052] The cross-modal attention fusion mechanism is implemented at the fourth-level feature map level. First, the point cloud feature map is compressed to 256 dimensions through a 1×1 convolution, concatenated with the image feature map in the channel dimension, and then input into the attention module. The attention module contains two parallel branches: one calculates the channel attention weights, and the other calculates the spatial attention weights. Finally, the weighted fusion generates a 512-dimensional fused feature map.
[0053] Channel attention weights are generated through global average pooling and two fully connected layers, while spatial attention weights are generated through convolutional layers with a kernel size of 7×7. The fused feature map serves as input to the region proposal network for subsequent tray detection and segmentation tasks. This two-stream feature extraction network is optimized end-to-end using a dataset containing 100,000 labeled samples during offline training. The loss function is a weighted sum of detection loss, segmentation loss, and pose regression loss, with weight coefficients of 1, 0.5, and 0.8, respectively. The training cycle is 200 epochs, with the learning rate exponentially decaying from 0.001 to 0.00001.
[0054] The system performs a joint task of tray instance segmentation and pose regression based on the fused feature map. A region proposal network (RPN) generates candidate tray bounding boxes, which are then output as pixel-level tray masks via a fully convolutional segmentation head. Simultaneously, a rotation-invariant pose regression head predicts the 3D coordinates, length, width, and height dimensions of the tray's center point, as well as its yaw angle around the vertical axis. Finally, the system outputs the six-DOF pose parameters of the tray in the world coordinate system. The RPN slides 3×3 anchor boxes on the fused feature map, with preset anchor box sizes of 800mm×1200mm, 1000mm×1400mm, and 1200mm×1600mm to cover common tray sizes.
[0055] Each anchor box outputs four bounding box offsets and an object confidence score. Candidate boxes with scores higher than 0.7 proceed to subsequent processing. The fully convolutional segmentation head adopts a U-Net architecture, containing four downsampling layers and four upsampling layers. Skip connections add the encoded and decoded features of the corresponding layers, ultimately outputting a binary mask map with the same resolution as the input image. Pixels with a mask value of one belong to the tray region. The rotation-invariant pose regression head performs global average pooling on the feature regions corresponding to the candidate boxes, inputting it into a 3-layer fully connected network and outputting a seven-dimensional vector. The first three dimensions are the X, Y, and Z coordinates of the tray center point in the world coordinate system, the middle three dimensions are the length, width, and height, and the last dimension is the yaw angle around the Z-axis.
[0056] To eliminate the impact of rotation on regression accuracy, the network applies random rotation enhancement to the input image and point cloud data during training, with rotation angles ranging from -180° to +180°, forcing the network to learn rotation-invariant features. After the pose parameters are output, the system automatically verifies the reasonableness of the dimensions. If the length and width values deviate from the standard pallet dimensions by more than 10%, a re-detection mechanism is triggered to re-acquire data and perform inference. This joint task takes approximately 150ms during the inference phase, meeting the requirements of real-time unloading cycle time.
[0057] To address the issue of pallet holes being partially obscured by stretch film, making gripping points invisible, a prior knowledge base for pallet structure is constructed. This knowledge base stores a template for the distribution of holes in a standard pallet and a topological map of the load-bearing area. When the vision system cannot directly detect a hole, the theoretical location of the hole is calculated by matching the current pallet outer contour dimensions with the knowledge base template. This calculation is then combined with local curvature analysis of the depth point cloud to verify whether the location is feasible for gripping, thereby generating virtual gripping point coordinates.
[0058] The knowledge base contains five mainstream pallet sizes: 1200mm×1000mm European pallet, 1100mm×1100mm Japanese pallet, 1067mm×1067mm American pallet, 1400mm×1400mm chemical pallet, and 800mm×1200mm light-duty pallet. For each pallet, the system stores the offset of the hole center coordinates relative to the pallet's geometric center, the hole diameter, and the orientation and width of the load-bearing beam. When the pallet dimensions output by the pose regression module have an error of less than 5% compared to the dimensions of a template in the knowledge base, the system automatically retrieves the hole distribution parameters from that template, combines them with the current pallet center point coordinates and yaw angle, and calculates the theoretical position of the hole in the world coordinate system through rigid body transformation.
[0059] To verify the physical feasibility of the theoretical location, the system extracts a local point cloud from the depth point cloud, with the center of the theoretical hole as the sphere and a radius of 50mm, and calculates its average curvature. If the average curvature is less than 0.05, the area is determined to be a plane or gentle slope, meeting the grasping conditions; if the average curvature is greater than 0.05, it is determined to be an obstacle or depression, the hole is abandoned, and an attempt is made to match adjacent holes. The final output of the virtual grasping point coordinates is a three-dimensional spatial point, whose Z-coordinate is determined by the local elevation value provided by the dynamic operation plane modeling module, ensuring that the grasping point is located on the pallet's physical structure. This mechanism can still guarantee a grasping point positioning success rate of over 95% even under extreme working conditions where the wrapping film coverage exceeds 70%.
[0060] A dynamic operational planar modeling module is established. A three-dimensional surface equation is generated by fitting the point cloud data of the cargo box floor. This surface equation is expressed using a cubic B-spline surface, and its control vertices are determined by interpolation of the measured elevation values of the four corners and the central area of the cargo box, thus accurately describing the deformation state of the cargo box floor, which is higher at the front and lower at the back, or has local concavity. Before each unloading operation, the system first collects the exposed point cloud data of the cargo box floor, extracts the floor point cloud using a planar segmentation algorithm, and removes interference from pallet and cargo point clouds. The floor point cloud is divided into a 5x5 grid, and the mean Z-coordinate of the point cloud within each grid is calculated as the elevation value of that grid node. The cubic B-spline surface is defined as:
[0061] ;
[0062] in To control the vertex, its X and Y coordinates are determined by the grid node positions, and its Z coordinate is assigned by the measured elevation value; and The system uses a cubic B-spline basis function with uniformly distributed node vectors. After surface fitting, the system samples the bottom region of each pallet and calculates its minimum distance from the surface, which is taken as the actual height of the pallet off the ground. This height value is used to correct the initial approach height of the robot's end effector, ensuring a 5mm safety margin in the vertical direction for the grasping action.
[0063] Meanwhile, the normal vector of the curved surface in the projection area of the pallet's load-bearing beam is used to calculate the gripping tilt angle, aligning the end effector's posture with the local slope of the base plate and preventing lateral slippage caused by tilting. The dynamic working plane model is automatically updated after slight vehicle displacement or load changes, with an update cycle of each pallet removal to ensure the model always reflects the current real-world state.
[0064] Based on the pallet pose parameters and the equation of the working plane surface, the actual contact area between the bottom of the pallet and the cargo compartment floor is calculated. Accordingly, the initial approach height and gripping angle of the end effector of the handling robot are adjusted to ensure a 5mm safety margin in the vertical direction and alignment with the centerline of the pallet's load-bearing beam in the horizontal direction. The actual contact area is determined by projecting the pallet bottom surface onto a B-spline surface and calculating the set of intersection points; the convex hull of the intersection point set is the contact area. The system extracts the geometric center of the contact area and calculates its normal vector on the surface, which serves as the target pose of the end effector.
[0065] The initial approach height is set to the Z-coordinate of the highest point of the contact area plus 5 mm, and is resolved into a joint space trajectory by the robot motion controller. The gripping tilt angle is determined by the angle between the normal vector and the Z-axis of the world coordinate system, and is converted into the wrist joint angle increment by the inverse kinematics solver. Horizontal alignment is achieved by projecting the centerline of the pallet support beam onto the XY plane and aligning it with the X-axis of the robot base coordinate system. The alignment error is compensated by visual servo closed-loop control, with a maximum allowable deviation of ±2 mm. The adjusted gripping parameters are encapsulated into motion commands and sent to the robot controller for execution. This adjustment mechanism ensures the stability and repeatability of the gripping action even under extreme conditions where the cargo box floor slope reaches 15%.
[0066] A hierarchical path planning architecture is constructed. The upper layer uses an improved A* algorithm to search for a globally collision-free path from the robot base to the pallet gripping point in a 3D grid map. The grid map is generated from cargo point cloud data through voxelization, and obstacle areas are marked as impassable. The lower layer uses a dynamic window method to adjust the robot's joint trajectory in real time to avoid temporary obstacles caused by pallet stacking swaying or slight vehicle displacement during path execution. The 3D grid map has a resolution of 50mm and is generated from cargo point cloud data through voxelization and dilation operations. The dilation radius is the robot's maximum outer radius plus a 50mm safety distance. The improved A* algorithm uses Euclidean distance plus a heuristic function. The search space is a discretized grid of the robot's reachable workspace, and each grid node stores its parent node pointer and cumulative cost. Path search starts from the robot's current position and expands to the pallet gripping point. If a path exists, the node sequence is output; otherwise, a replanning mechanism is triggered. The lower-level dynamic window method is executed in the robot's joint space, with a sampling period of 10ms and a prediction window length of 0.5s. The evaluation function includes path tracking error, obstacle distance, joint velocity, and acceleration constraints. Temporary obstacles are identified by differential detection between real-time point cloud data and historical maps, with a differential threshold set to 50ms. Regions of difference lasting longer than 0.2s are marked as new obstacles. The dynamic window method generates multiple candidate trajectories within each control cycle and selects the trajectory with the highest evaluation score for execution. This hierarchical architecture enables the robot to safely reach the target point at an average speed of 0.8m / s even in crowded environments with a pallet stacking density of up to 80%, with a path replanning frequency of less than once per minute.
[0067] Before the handling robot performs its grasping action, a force control prediction module is activated. This module calculates the minimum clamping force required for grasping based on the estimated pallet weight and the empirical value of the friction coefficient of the cargo box floor. It also monitors changes in the clamping force in real time using a six-dimensional force sensor at the robot's end effector. When a fluctuation in the clamping force exceeds a preset threshold, a fine-tuning command for the grasping posture is immediately triggered to compensate for the loss of grasping force caused by slippage of the stretch film or deformation of the pallet. The estimated pallet weight is determined jointly by a material type database and the pallet dimensions. The database includes the unit volume weight of four types of materials: bagged, drummed, bottled, and canned, each with a value of 0.8 t / m³. 3 1.2t / m 3 0.6t / m 3 0.9t / m 3 Multiplying this by the pallet volume gives an estimated total weight. The coefficient of friction for the cargo box floor is empirically set based on the floor material: 0.3 for steel, 0.4 for wood, and 0.25 for plastic. The formula for calculating the minimum clamping force is:
[0068] ;
[0069] in The coefficient of friction, This is an estimated value for the pallet's weight. It is the acceleration due to gravity. For safety, a value of 1.5 is used. The six-dimensional force sensor has a sampling frequency of 1000Hz and outputs force components and torque components in three directions in real time. The clamping force fluctuation detection is based on the sliding window standard deviation of the Z-direction force component, with a window length of 0.1s. If the standard deviation exceeds 5N, it is judged as an abnormal fluctuation, triggering a fine-tuning command.
[0070] The fine-tuning commands include translational compensation along the X and Y directions and rotational compensation around the Z-axis. The compensation amount is calculated by integrating the torque components, with a maximum compensation range of ±5mm and ±5°. The compensation action is completed within 0.2s, ensuring gripping stability. Even when the gripping force drops instantaneously by 30% due to the slippage of the wrapping film, this force control mechanism can still maintain reliable gripping through dynamic compensation, preventing the pallet from falling off.
[0071] After pallet grasping is complete, a dynamic center of gravity compensation algorithm is activated. This algorithm adjusts the torque output of the robot's wrist joints in real time based on the center of gravity offset caused by uneven material distribution within the pallet, ensuring the pallet remains horizontal during transport and preventing material tipping. The center of gravity offset is calculated by back-calculating the torque components measured by a six-dimensional force sensor. Assuming the pallet mass is uniformly distributed, the torque should be zero; the difference between the measured torque and the theoretical torque, divided by the pallet mass, is the center of gravity offset vector. The dynamic center of gravity compensation algorithm maps the center of gravity offset vector to the three rotary joints of the robot's wrist and calculates the required joint torque increment using a Jacobian matrix. The torque increment application period is 5ms, synchronized with the robot's underlying controller. To avoid over-compensation leading to oscillation, a low-pass filter with a cutoff frequency of 5Hz is introduced to smooth the torque command.
[0072] Simultaneously, the system monitors the pallet's attitude angle. If the pitch or roll angle exceeds 3°, a deceleration command is triggered, reducing the handling speed to 50% of the original speed until the attitude returns to a horizontal position. This compensation mechanism ensures that even in extreme cases where the material is unbalanced by up to 40%, the pallet's attitude angle error remains less than 1°, guaranteeing transportation safety.
[0073] Once the unloaded pallets are removed from the work map, the vision system is triggered to rescan and update the pose of the remaining pallets, initiating the next unloading cycle until all pallets in the cargo compartment are emptied. Pallet removal is performed after the robot confirms the pallets have been placed in the target location, which is the specified coordinates of the warehouse inlet, pre-assigned by the warehouse management system. The vision system rescans in a fast mode, using only the top-view and front-angle cameras, increasing the acquisition frequency to 60 frames per second, and limiting the scanning range to the local area where the remaining pallets are located, reducing data processing load.
[0074] The pose update process is the same as the first round of detection, but skips the dynamic operation plane modeling step and reuses the existing model. The system maintains a pallet status list, recording the unique identifier, pose parameters, and grasping status of each pallet. Grabbed pallets are marked as completed, and ungrabbed pallets are marked as pending. When all pallet statuses in the list are completed, the unloading operation ends, and the system outputs an operation report, including the total number of pallets, average cycle time, and abnormal event records. This loop mechanism supports continuous operation, with a single cycle time controlled within 45 seconds, of which vision processing accounts for 50 seconds, path planning for 10 seconds, robot motion for 15 seconds, and force control and compensation for 5 seconds, meeting the requirements for efficient unloading.
[0075] Example 2: Please refer to Figure 1-5 This invention provides an automated unmanned unloading system based on machine vision, the core of which is:
[0076] The multi-view synchronous imaging module is used to simultaneously acquire visible light image sequences and depth point cloud data inside the cargo compartment of the vehicle to be unloaded by using a multi-view industrial camera array deployed above the unloading operation area. The industrial camera array contains at least three imaging units, which cover the entire cargo compartment from top-down, front-oblique, and side-oblique angles, respectively. The imaging units achieve microsecond-level time synchronization through hardware trigger signals.
[0077] The image and point cloud preprocessing module performs multi-stage image preprocessing on the acquired visible light image sequence. This includes enhancing image contrast using an adaptive histogram equalization algorithm, sharpening edges using the Gaussian-Laplacian operator, filling edge breakage areas through morphological closing operations, and finally outputting image data with enhanced texture features. The module also performs spatial filtering and coordinate normalization on the acquired depth point cloud data. First, it removes noise points using a statistical outlier removal algorithm, then reduces data density through voxel grid downsampling, and finally transforms the point cloud coordinate system to a unified world coordinate system with the lower left corner of the inner side of the rear tailgate of the vehicle cargo box as the origin.
[0078] The dual-stream feature fusion and recognition module is used to construct a dual-stream feature extraction network. The first branch receives the preprocessed visible light image and uses an improved residual convolutional neural network to extract the surface texture features and contour edge features of the tray. The second branch receives the normalized depth point cloud data and uses a point cloud transformation network to extract the three-dimensional geometric structure features and spatial distribution features of the tray. The dual-stream feature extraction network performs cross-modal attention fusion at the feature layer to generate a tray candidate region feature map that fuses visual and geometric semantics.
[0079] The pose calculation and virtual grasping point generation module is used to perform joint tasks of pallet instance segmentation and pose regression based on fused feature maps. It uses a region proposal network to generate pallet candidate boxes, and then outputs a pixel-level pallet mask through a fully convolutional segmentation head. At the same time, it uses a rotation-invariant pose regression head to predict the 3D coordinates, length, width, and height parameters of the pallet center point and the yaw angle around the vertical axis. Finally, it outputs the six-degree-of-freedom pose parameters of the pallet in the world coordinate system. In order to address the situation where the grasping point is not visible due to the partial occlusion of the pallet holes by the wrapping film, a prior knowledge base of pallet structure is constructed. The knowledge base stores the hole distribution template of standard pallets and the topology map of the load-bearing area. When the vision system cannot directly detect the hole, the theoretical position of the hole is estimated by matching the current pallet outer contour size with the knowledge base template. The feasibility of grasping at the position is verified by combining the local curvature analysis of the depth point cloud, thereby generating the coordinates of the virtual grasping point.
[0080] The dynamic operation plane modeling module is used to establish a dynamic operation plane modeling module. It generates a three-dimensional surface equation by fitting the point cloud data of the cargo box floor. The surface equation is expressed by a cubic B-spline surface. Its control vertices are determined by interpolation of the measured elevation values of the four corners and the central area of the cargo box, thereby accurately describing the deformation state of the cargo box floor that is higher in the front and lower in the back or has local concavity.
[0081] The adaptive adjustment module for gripping parameters is used to calculate the actual contact area between the bottom of the pallet and the floor of the cargo box based on the pallet pose parameters and the working plane surface equation, and adjust the initial approach height and gripping tilt angle of the end effector of the handling robot accordingly, so as to ensure that the gripping action has a safety margin of five millimeters in the vertical direction and is aligned with the center line of the pallet load-bearing beam in the horizontal direction.
[0082] The hierarchical path planning module is used to construct a hierarchical path planning architecture. The upper layer uses an improved A* algorithm to search for a globally collision-free path from the robot base to the pallet gripping point in a 3D grid map. The grid map is generated by voxelizing the cargo box point cloud data, and obstacle areas are marked as impassable. The lower layer uses a dynamic window method to adjust the robot joint trajectory in real time to avoid temporary obstacles caused by pallet stacking swaying or slight vehicle displacement during path execution.
[0083] The force control prediction and center of gravity compensation module is used to activate the force control prediction module before the handling robot performs the grasping action. The module calculates the minimum clamping force required for grasping based on the estimated pallet weight and the empirical value of the friction coefficient of the cargo box floor. It also monitors the changes in clamping force in real time through a six-dimensional force sensor at the end of the robot. When the clamping force fluctuation exceeds a preset threshold, it immediately triggers a fine-tuning command for the grasping posture to compensate for the loss of grasping force caused by the slippage of the stretch film or the deformation of the pallet. After the pallet is grasped, the dynamic center of gravity compensation algorithm is activated. The algorithm adjusts the torque output of the wrist joint of the handling robot in real time according to the center of gravity offset caused by the uneven distribution of materials in the pallet, ensuring that the pallet maintains a horizontal posture during transportation and avoiding the tipping of materials.
[0084] The operation cycle control module is used to remove the coordinates of unloaded pallets from the operation map and trigger the vision system to rescan and update the pose of the remaining pallet stacking status, and enter the next unloading cycle until all pallets in the cargo compartment are emptied.
[0085] This system, through the coordinated execution of the aforementioned methods, adapts to unloading scenarios with any vehicle type, stacking configuration, and membrane material coverage, without requiring manual intervention. This improves the success rate of unloading operations, shortens the single unloading cycle time, and enhances the throughput efficiency and operational safety of the logistics automation system. The system hardware platform utilizes an industrial-grade embedded computer and a real-time operating system, while the software modules are deployed in a microservice architecture, supporting hot-swapping and online upgrades.
[0086] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0087] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A machine vision-based unmanned unloading automated operation method, characterized in that, include: Visible light image sequences and depth point cloud data of the interior of the cargo compartment of the vehicle to be unloaded are simultaneously acquired by a multi-view industrial camera array deployed above the unloading area. Perform multi-stage image preprocessing on the acquired visible light image sequence; Spatial filtering and coordinate normalization are performed on the collected depth point cloud data; A dual-stream feature extraction network is constructed, wherein the first branch receives the preprocessed visible light image and uses an improved residual convolutional neural network to extract the surface texture features and contour edge features of the tray; the second branch receives the normalized depth point cloud data and uses a point cloud transformation network to extract the three-dimensional geometric structure features and spatial distribution features of the tray; the dual-stream feature extraction network performs cross-modal attention fusion at the feature layer to generate a tray candidate region feature map that integrates visual and geometric semantics. Perform a joint task of tray instance segmentation and pose regression based on fused feature maps; To address the issue of pallet holes being partially obscured by stretch film, making gripping points invisible, a prior knowledge base for pallet structure is constructed. This prior knowledge base stores a template for the distribution of holes in a standard pallet and a topological map of the load-bearing area. When the vision system cannot directly detect holes, the theoretical location of the holes is calculated by matching the current pallet outer contour dimensions with the knowledge base template. The feasibility of gripping at this location is then verified by combining the local curvature analysis of the depth point cloud, thereby generating virtual gripping point coordinates. Establish a dynamic operation planar modeling module; Construct a hierarchical path planning architecture; Remove the coordinates of the unloaded pallets from the work map and trigger the vision system to rescan and update the pose of the remaining pallet stacks, then enter the next unloading cycle until all pallets in the cargo compartment are emptied.
2. The automated unmanned unloading method based on machine vision according to claim 1, characterized in that, The acquired visible light image sequence is subjected to multi-stage image preprocessing, including using an adaptive histogram equalization algorithm to enhance image contrast, applying the Gaussian-Laplacian operator for edge sharpening, and then filling the edge breakage area through morphological closing operation, finally outputting image data with enhanced texture features.
3. The automated unmanned unloading method based on machine vision according to claim 2, characterized in that, Spatial filtering and coordinate normalization are performed on the collected depth point cloud data. First, a statistical outlier removal algorithm is used to remove noise points. Then, data density is reduced by voxel grid downsampling. Finally, the point cloud coordinate system is transformed to a unified world coordinate system with the lower left corner of the inner side of the rear tailgate of the vehicle cargo box as the origin.
4. The automated unmanned unloading method based on machine vision according to claim 1, characterized in that, The system performs a joint task of tray instance segmentation and pose regression based on fused feature maps. It uses a region proposal network to generate tray candidate boxes, and then outputs a pixel-level tray mask through a fully convolutional segmentation head. At the same time, it uses a rotation-invariant pose regression head to predict the 3D coordinates, length, width and height parameters of the tray center point and the yaw angle around the vertical axis. Finally, it outputs the six-degree-of-freedom pose parameters of the tray in the world coordinate system.
5. The automated unmanned unloading method based on machine vision according to claim 4, characterized in that, Construct a prior knowledge base for the tray structure and generate virtual grab point coordinates, including: The knowledge base stores the offset of the hole center coordinates relative to the geometric center, hole diameter, and the orientation and width of the load-bearing beam for five mainstream pallet specifications. When the error between the pose regression output size and the knowledge base template is less than 5%, the corresponding template parameters are called to calculate the theoretical position of the hole by combining the current center point coordinates and yaw angle. Extract a local point cloud with a radius of 50mm centered at the theoretical hole center, calculate the average curvature, and if it is less than 0.05, it is determined that the grasping conditions are met; The Z-coordinate of the virtual grab point is determined by the local elevation value provided by the dynamic operation plane modeling module.
6. The automated unmanned unloading method based on machine vision according to claim 5, characterized in that, A dynamic operation plane modeling module is established. A three-dimensional surface equation is generated by fitting the point cloud data of the cargo box floor. The surface equation is expressed by a cubic B-spline surface. Its control vertex is determined by interpolation of the measured elevation values of the four corners and the central area of the cargo box, thereby accurately describing the deformation state of the cargo box floor that is higher in the front and lower in the back or has local concavity. Based on the pallet pose parameters and the working plane surface equation, the actual contact area between the bottom of the pallet and the cargo box floor is calculated, and the initial approach height and gripping tilt angle of the end effector of the handling robot are adjusted accordingly to ensure that the gripping action has a safety margin of 5 mm in the vertical direction and is aligned with the center line of the pallet load-bearing beam in the horizontal direction.
7. The automated unmanned unloading method based on machine vision according to claim 6, characterized in that, A hierarchical path planning architecture is constructed. The upper layer uses an improved A* algorithm to search for a global collision-free path from the robot base to the pallet gripping point in a 3D grid map. The grid map is generated by voxelizing the cargo box point cloud data, and obstacle areas are marked as impassable. The lower layer uses a dynamic window method to adjust the robot joint trajectory in real time to avoid temporary obstacles caused by pallet stacking shaking or slight vehicle displacement during path execution.
8. The automated unmanned unloading method based on machine vision according to claim 1, characterized in that, Also includes: Before the handling robot performs the grasping action, the force control prediction module is activated. Based on the estimated value of the pallet weight and the empirical value of the friction coefficient of the cargo box floor, the minimum clamping force required for grasping is calculated. The six-dimensional force sensor at the end of the robot monitors the change of clamping force in real time. When the clamping force fluctuation is detected to exceed the preset threshold, the grasping posture fine adjustment command is immediately triggered to compensate for the loss of grasping force caused by the slippage of the stretch film or the deformation of the pallet. After the pallet is picked up, a dynamic center of gravity compensation algorithm is activated. The algorithm adjusts the torque output of the wrist joint of the handling robot in real time according to the center of gravity offset caused by uneven distribution of materials in the pallet, so as to ensure that the pallet remains in a horizontal position during transportation and avoid the material from tipping over.
9. A machine vision-based unmanned unloading automated operation system, comprising: The multi-view synchronous imaging module is used to simultaneously acquire visible light image sequences and depth point cloud data inside the cargo compartment of the vehicle to be unloaded by using a multi-view industrial camera array deployed above the unloading operation area. The industrial camera array contains at least three imaging units, which cover the entire cargo compartment from top-down, front-oblique, and side-oblique angles, respectively. The imaging units achieve microsecond-level time synchronization through hardware trigger signals. The image and point cloud preprocessing module performs multi-stage image preprocessing on the acquired visible light image sequence. This includes enhancing image contrast using an adaptive histogram equalization algorithm, sharpening edges using the Gaussian-Laplacian operator, filling edge breakage areas through morphological closing operations, and finally outputting image data with enhanced texture features. The module also performs spatial filtering and coordinate normalization on the acquired depth point cloud data. First, it removes noise points using a statistical outlier removal algorithm, then reduces data density through voxel grid downsampling, and finally transforms the point cloud coordinate system to a unified world coordinate system with the lower left corner of the inner side of the rear tailgate of the vehicle cargo box as the origin. The dual-stream feature fusion and recognition module is used to construct a dual-stream feature extraction network. The first branch receives the preprocessed visible light image and uses an improved residual convolutional neural network to extract the surface texture features and contour edge features of the tray. The second branch receives the normalized depth point cloud data and uses a point cloud transformation network to extract the three-dimensional geometric structure features and spatial distribution features of the tray. The dual-stream feature extraction network performs cross-modal attention fusion at the feature layer to generate a tray candidate region feature map that fuses visual and geometric semantics. The pose calculation and virtual grasping point generation module is used to perform joint tasks of pallet instance segmentation and pose regression based on fused feature maps. It uses a region proposal network to generate pallet candidate boxes, and then outputs a pixel-level pallet mask through a fully convolutional segmentation head. At the same time, it uses a rotation-invariant pose regression head to predict the 3D coordinates, length, width, and height parameters of the pallet center point and the yaw angle around the vertical axis. Finally, it outputs the six-degree-of-freedom pose parameters of the pallet in the world coordinate system. In order to address the situation where the grasping point is not visible due to the partial occlusion of the pallet holes by the wrapping film, a prior knowledge base of pallet structure is constructed. The prior knowledge base stores the hole distribution template and the topology map of the load-bearing area of the standard pallet. When the vision system cannot directly detect the hole, the theoretical position of the hole is estimated by matching the current pallet outer contour size with the knowledge base template. The feasibility of grasping at the position is verified by combining the local curvature analysis of the depth point cloud, thereby generating the coordinates of the virtual grasping point. The dynamic operation plane modeling module is used to establish a dynamic operation plane modeling module. It generates a three-dimensional surface equation by fitting the point cloud data of the cargo box floor. The surface equation is expressed by a cubic B-spline surface. Its control vertices are determined by interpolation of the measured elevation values of the four corners and the central area of the cargo box, thereby accurately describing the deformation state of the cargo box floor that is higher in the front and lower in the back or has local concavity. The adaptive adjustment module for gripping parameters is used to calculate the actual contact area between the bottom of the pallet and the floor of the cargo box based on the pallet pose parameters and the working plane surface equation, and adjust the initial approach height and gripping tilt angle of the end effector of the handling robot accordingly, so as to ensure that the gripping action has a safety margin of five millimeters in the vertical direction and is aligned with the center line of the pallet load-bearing beam in the horizontal direction. The hierarchical path planning module is used to construct a hierarchical path planning architecture. The upper layer uses an improved A* algorithm to search for a globally collision-free path from the robot base to the pallet gripping point in a 3D grid map. The grid map is generated by voxelizing the cargo box point cloud data, and obstacle areas are marked as impassable. The lower layer uses a dynamic window method to adjust the robot joint trajectory in real time to avoid temporary obstacles caused by pallet stacking swaying or slight vehicle displacement during path execution. The force control prediction and center of gravity compensation module is used to activate the force control prediction module before the handling robot performs the grasping action. The module calculates the minimum clamping force required for grasping based on the estimated value of the pallet weight and the empirical value of the friction coefficient of the cargo box floor. It also monitors the change of clamping force in real time through the six-dimensional force sensor at the end of the robot. When the clamping force fluctuation is detected to exceed the preset threshold, the grasping posture fine adjustment command is immediately triggered to compensate for the loss of grasping force caused by the slippage of the stretch film or the deformation of the pallet. After completing the pallet grabbing, the dynamic center of gravity compensation algorithm is activated. The algorithm adjusts the torque output of the wrist joint of the handling robot in real time according to the center of gravity offset caused by the uneven distribution of materials in the pallet, so as to ensure that the pallet maintains a horizontal posture during transportation and avoids the material tipping over. The operation cycle control module is used to remove the coordinates of unloaded pallets from the operation map and trigger the vision system to rescan and update the pose of the remaining pallet stacking status, and enter the next unloading cycle until all pallets in the cargo compartment are emptied.