A method and system for layered positioning and insertion of mother and child pallets by an unmanned forklift
Patent Information
- Application Number
- CN202611142769.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]为了解决现有成熟3D视觉技术仅适配单层标准托盘作业,在子母托盘堆叠场景存在层级分不清、间隙测不准、叉孔找不到、工况适配不了的根本性缺陷,本申请提供一种无人叉车子母托盘分层定位插取方法及系统
1、通过同步控制2D和3D相机曝光与采集,完成联合标定与坐标转换,经图像识别平面特征、点云分割分层点云簇;设计双向数据互相补偿融合算法、双视觉交叉校验容错逻辑,进行坐标映射与校准、数据校验与互补拟合;配套子母托专属分层判定与自适应插取控制流程,实现更为精准子母托盘分层定位插取;
Smart Images

Figure CN122667321A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of logistics and warehousing automation technology, specifically to a method and system for layered positioning and insertion of mother and child pallets by an unmanned forklift. Background Technology
[0002] With the increasing demand for high-density and flexible operations in smart warehousing, mother-daughter pallets have become the mainstream load-bearing equipment in e-commerce, manufacturing, and automated warehouses due to their advantages of high space utilization, flexible disassembly and transfer, and low warehousing costs. Mother-daughter pallet operations include three typical scenarios: independent mother pallet placement, independent daughter pallet placement, and stacked placement with daughter pallets on top of the mother pallet. Currently, the demand for visual positioning of unmanned forklifts is growing, and mainstream technologies, after long-term iteration, have achieved good results in some common scenarios, providing strong support for the automation development of warehousing and logistics.
[0003] Currently, visual positioning solutions for unmanned forklifts are widely adopted in the industry, with the mainstream technology being a single 3D structured light / laser visual recognition solution. This solution uses only a single 3D structured light camera in hardware, and relies on mature point cloud segmentation and template matching algorithms in software. Its workflow is as follows: First, the single 3D camera acquires all 3D point cloud data of the pallet in the work area; then, filtering, denoising, and segmentation algorithms are used to extract the overall point cloud contour of the pallet, fitting the pallet's 3D dimensions, ground clearance, and placement posture; next, a standard pallet 3D model is pre-loaded, and the real-time point cloud is compared with the model to determine the pallet type and output positioning coordinates; finally, based on the single 3D positioning data, fixed parameters are used to control the forklift to complete the pallet insertion and removal operation.
[0004] However, existing mature 3D vision technologies are only suitable for single-layer standard pallet operations, and have fundamental defects in the core scenario of mother-daughter pallet stacking. First, they cannot accurately distinguish between mother-daughter pallets of the same shape, resulting in layer recognition failure and easy misjudging of stacked mother-daughter pallets as single-layer thick pallets, leading to forklifts bumping, hitting, or damaging goods during insertion. Second, 3D point cloud overlap interference in stacking scenarios makes it impossible to extract layer gaps and effective insertion positions, significantly reducing the success rate of operations. Third, the lack of fine feature recognition capabilities in pure 3D vision makes it unsuitable for non-standard stacking conditions, frequently resulting in misinsertion, empty insertion, and collision accidents. Fourth, single 3D vision lacks a fault-tolerant compensation mechanism, resulting in extremely poor adaptability to various working conditions. It is prone to recognition failure and interruption of automated operations when faced with uneven warehouse lighting, pallet surface reflections, or slight dust obstruction. Fifth, traditional 3D solutions lack layered operation logic, lack control strategies adapted to mother-daughter pallet operations, cannot distinguish different working conditions, and cannot adaptively adjust lifting height and forklift insertion depth, resulting in a lack of safe operation control capabilities in stacking scenarios. Summary of the Invention
[0005] To address the fundamental shortcomings of existing mature 3D vision technologies, which are only compatible with single-layer standard pallet operations and suffer from unclear layer distinctions, inaccurate gap measurements, inability to locate forkholes, and incompatibility with working conditions in scenarios involving stacked mother and daughter pallets, this application provides a method and system for layered positioning and insertion of mother and daughter pallets using an unmanned forklift.
[0006] In a first aspect, this application provides a method for layered positioning and insertion of mother and child pallets using an unmanned forklift, comprising: By controlling the synchronous exposure of a 2D industrial camera and a structured light 3D camera through synchronous hardware trigger pulse control, the pallet plane image and 3D point cloud are acquired at the same time and from the same perspective, and the joint calibration is completed and a unified coordinate transformation matrix from the 2D pixel coordinate system to the 3D world coordinate system is established. Based on the pallet planar image, the planar features of the pallet are identified, and the pixel coordinates of the fork hole center and the pixel coordinates of the pallet edge are fitted; based on the three-dimensional point cloud, the point cloud clusters corresponding to each layer of the pallet are segmented, and the three-dimensional pose and geometric parameters of each layer of the pallet are obtained. Based on the unified coordinate transformation matrix obtained by joint calibration, the identified tray edge pixel coordinates are mapped to a three-dimensional coordinate system to generate a region mask for filtering and completing the three-dimensional point cloud, and the three-dimensional point cloud filtering and completion are performed; the fork hole center pixel coordinates are mapped to initial three-dimensional coordinates, and spatial coordinate calibration is performed based on the point cloud clusters of the corresponding layer to correct the two-dimensional plane positioning drift error; Calculate the absolute difference between the pallet size identified and fitted based on the pallet planar image and the 3D point cloud, respectively. When the absolute difference exceeds the set difference threshold, initiate 2D or 3D single-path recognition failure complementary fitting. The current working condition is determined based on the number of layered point cloud clusters. When the working condition is a parent-child stacking condition, the vertical safety gap between layers is detected and a fork arm avoidance adjustment command is generated based on the detection result. The corresponding forklift motion parameter set is called according to the working condition to execute the insertion action.
[0007] By adopting the above scheme, synchronous hardware trigger pulses are used to enable 2D and 3D cameras to acquire data synchronously and complete joint calibration, establish a unified coordinate transformation matrix, and achieve accurate fusion of 2D and 3D data; planar images and 3D point clouds are processed separately to accurately obtain pallet features and pose parameters; positioning accuracy can be improved through region masking and coordinate calibration, and single-path failure complementary fitting can ensure uninterrupted recognition; the working condition is determined, interlayer gaps are detected, avoidance instructions are generated, and matching parameters are called to improve the accuracy and safety of mother and child pallet insertion and removal, and adapt to complex working conditions.
[0008] Preferably, the joint calibration is completed by multiple sets of synchronously acquired chessboard calibration board images and point cloud data; the calibration parameters include 2D camera intrinsic parameters, distortion coefficients, 3D camera point cloud world coordinate system parameters, and rotation and translation matrices and homogeneous transformation matrices from the two-dimensional pixel coordinate system to the three-dimensional world coordinate system, and the calibration parameters are permanently stored in the forklift on-board motion controller.
[0009] By adopting the above scheme, the joint calibration completes the unbiased mapping between the two-dimensional pixel coordinate system and the three-dimensional world coordinate system, enabling the image and point cloud data to be synchronized with high precision, ensuring complete data matching, avoiding data fragmentation, and at the same time, the calibration parameters are solidified and stored so that no secondary manual calibration is required on site, which facilitates mass production of forklifts and rapid deployment on site.
[0010] Preferably, the step of identifying the planar features of the tray based on the tray planar image and fitting the pixel coordinates of the center of the fork hole and the pixel coordinates of the tray edge includes: preprocessing the tray planar image and extracting the length and width of the fork hole, the thickness of the border and the texture features of the board surface; classifying the tray type according to the pre-stored two-dimensional templates of the mother tray and the child tray; and fitting the pixel coordinates of the center of the fork hole and the pixel coordinates of the tray edge. The steps of segmenting the point cloud clusters corresponding to each layer of trays based on the three-dimensional point cloud and obtaining the three-dimensional pose and geometric parameters of each layer of trays include: filtering and denoising the three-dimensional point cloud and performing multi-layer clustering segmentation; separating two independent point cloud clusters of the upper sub-tray and the lower mother tray based on the height threshold and spatial connectivity analysis; and calculating the ground clearance of each layer, the vertical gap between layers, and the spatial attitude angle of each layer.
[0011] By adopting the above scheme, the classification problem of mother and child pallets is solved by using the 2D camera's ability to recognize fine planar features, and the problem of stacking and gap measurement is solved by using the 3D camera's depth detection capability. The advantages of the two perception dimensions complement each other, thereby improving the recognition accuracy of mixed mother and child pallets and stacked conditions. Moreover, the 2D and 3D recognition boundaries are completely defined, and there is no duplicate recognition content, which greatly reduces the computing load of the vehicle controller and meets the real-time motion control cycle of the unmanned forklift.
[0012] Preferably, the step of performing 3D point cloud screening and completion includes: based on the generated region mask, determining all point clouds outside the region mask as ground, cargo expansion, and debris noise and removing them, retaining the original 3D point cloud inside the region mask, and completing the point cloud holes and missing edge areas according to the identified pallet edge pixel coordinates; The steps of mapping the center pixel coordinates of the fork hole to initial three-dimensional coordinates and performing spatial coordinate calibration based on the point cloud cluster of the corresponding layer to correct the two-dimensional plane positioning drift error include: mapping the center pixel coordinates of the two-dimensional fork hole to initial three-dimensional coordinates; searching for a point set in the sphere domain with the initial three-dimensional coordinates as the center; performing least squares plane fitting on the searched point set; taking the center point of the fitted plane as the final three-dimensional coordinates of the fork hole center to complete the spatial coordinate calibration; mapping the identified edge pixel coordinates of the tray to three-dimensional spatial lines; intersecting the plane obtained by fitting the corresponding point cloud cluster to obtain the corrected three-dimensional spatial lines of the tray edge.
[0013] By adopting the above solution, the problems of 3D point cloud coupling, edge blurring, and local voids in the stacking scenario are solved. The true boundary of each pallet layer is accurately locked, false point clouds are eliminated, and the calculation accuracy of layer gap and layer height is improved. At the same time, the positioning error caused by perspective distortion and viewpoint offset of a single 2D image is eliminated, and the high-precision two-dimensional fork hole position is given real depth information to obtain accurate three-dimensional insertion point position, thereby improving the accuracy and safety of forklift insertion.
[0014] Preferably, the step of initiating 2D or 3D single-path identification failure complementary fitting includes: When the 3D point cloud is distorted, the contour and fork hole positions of the pallet plane image are used as a reference, and the 3D spatial coordinates of the layered pallet are fitted in combination with historical layer height parameters; when the pallet plane image feature recognition is abnormal, the 3D point cloud cluster is used as a reference, and the theoretical plane position of the fork hole of each layer of pallet is fitted in reverse in combination with the pre-stored standard pallet 2D dimensions; an alarm command is output only when both visual recognition channels fail.
[0015] By adopting the above scheme, when the recognition of the 3D point cloud or pallet plane image is abnormal, the other normal data can be used for complementary fitting to maintain the positioning of the layered pallet. An alarm will only be triggered when both vision channels fail, ensuring the continuity and reliability of the unmanned forklift's layered positioning and insertion operation of the mother and child pallets under complex working conditions.
[0016] Preferably, the method further includes: For the fitted pixel coordinates of the fork hole center and the pixel coordinates of the tray edge, a pre-built detection network is used to output the covariance matrix or probability distribution representing the detection uncertainty, respectively; for each point in the 3D point cloud, a measurement noise model is established based on its depth value and surface characteristics, and depth uncertainty is assigned to each point. The step of mapping the identified tray edge pixel coordinates to a three-dimensional coordinate system based on the unified coordinate transformation matrix obtained by joint calibration, generating a region mask for filtering and completing the three-dimensional point cloud, and performing three-dimensional point cloud filtering and completion further includes: generating a probabilistic region mask based on the uncertainty of the tray edge pixel coordinates, and using the probabilistic region mask to perform weighted filtering and completion of the three-dimensional point cloud; the step of mapping the fork center pixel coordinates to initial three-dimensional coordinates, and performing spatial coordinate calibration based on the point cloud clusters of the corresponding layer to correct the two-dimensional plane positioning drift error further includes: fusing the detection uncertainty of the fork center pixel coordinates with the depth uncertainty of the corresponding point cloud clusters in the vicinity of the fork center pixel coordinates through an uncertainty propagation algorithm to obtain a three-dimensional fork position estimate; The obtained three-dimensional fork hole position estimate is used as an observation and input into a pre-constructed extended Kalman filter to perform multi-frame temporal fusion estimation of the tray pose, smooth single-frame noise and remove abnormal observations; After multi-frame temporal fusion estimation, physical consistency verification is performed. The physical consistency verification includes: calculating the spatial distance between the centers of the fork holes obtained in the current frame and comparing it with the known standard size of the pallet. If the deviation exceeds a preset threshold, the fusion result of the current frame is discarded and relocation is triggered.
[0017] By adopting the above scheme, the output detection uncertainty and the assigned point cloud depth uncertainty more accurately reflect the reliability of the measurement; the probabilistic region mask is used to perform weighted filtering and completion of the 3D point cloud, which more accurately corrects the defects of the 3D point cloud; the detection uncertainty is fused to obtain the 3D fork hole position estimate, which improves the accuracy of the fork hole position estimate; multi-frame temporal fusion estimation is performed by extending Kalman filter to smooth single-frame noise and remove abnormal observations, which improves the stability of pallet pose estimation; physical consistency verification is performed to ensure the accuracy of the fusion result, and repositioning is triggered when the deviation exceeds the preset threshold to ensure the reliability of the operation.
[0018] Preferably, the step of performing inter-layer vertical safety clearance detection and generating a forklift avoidance adjustment command based on the detection result includes: A set of parameters for the forklift tooth structure and a set of parameters for the pallet type are pre-established. The set of parameters for the forklift tooth structure includes tooth thickness, tooth width, and maximum pitch adjustment angle. The set of parameters for the pallet type includes fork hole entry height, fork hole width, bottom structure type, and standard dimensions. Based on the reconstruction of the highest envelope surface of the top of the mother bracket and the lower edge surface of the bottom structure of the child bracket using 3D point cloud, and combined with the thickness of the fork tooth and the insertion position of the current operation plan, a fork tooth passage sweep body is generated. Calculate the minimum vertical distance between the upper edge of the fork-tooth sweep body and the lower edge of the bottom structure of the sub-support, and the minimum vertical distance between the lower edge of the fork-tooth sweep body and the highest envelope surface of the top of the mother support, and take the smaller of the two values as the dynamic effective safety clearance. The dynamic effective safety gap is compared with the dynamic safety threshold. The dynamic safety threshold is dynamically calculated based on the fork tooth thickness, the current insertion speed, and the preset safety margin. If the dynamic effective safety gap is greater than or equal to the dynamic safety threshold, it is determined that the upper sub-support can be directly inserted. If the dynamic effective safety gap is less than the dynamic safety threshold, a fork tooth adjustment command or a stop insertion command is generated. The step of calling the corresponding set of forklift motion parameters to perform the insertion action according to the working conditions includes: dynamically adjusting the insertion speed according to the ratio of the dynamic effective safety clearance to the dynamic safety threshold.
[0019] By adopting the above scheme, when inserting the pallet, the dynamic effective safety clearance and dynamic safety threshold are accurately calculated by comprehensively considering various parameters of the forklift teeth and the pallet, and compared to determine whether direct insertion is possible. An avoidance adjustment command is generated, and the insertion speed is dynamically adjusted according to the ratio of clearance to threshold, thereby further improving the safety and accuracy of the insertion operation.
[0020] Preferably, the method further includes: After each insertion and removal action is completed, the actual spatial offset of the fork teeth relative to the pallet fork hole is collected by a micro-sensor set at the end of the fork arm, and the recognition error is calculated. By utilizing the accumulated recognition error, the joint calibration extrinsic parameters between the pallet planar image and the 3D point cloud are incrementally corrected online, and the joint calibration parameters are adjusted. The identification error is correlated with the number of layered point cloud clusters and the height difference features before insertion, and the classifier feature weights used to determine the current working condition are updated online. The revised joint calibration parameters and classifier feature weights are applied to the perception and condition determination of the next round of insertion operations.
[0021] By adopting the above scheme, the recognition error can be calculated based on the actual spatial offset collected after the insertion action is completed. This error can be used to incrementally correct the joint calibration extrinsic parameters and adjust the joint calibration parameters online. At the same time, the recognition error and related features are correlated to update the classifier feature weights online, thereby improving the perception accuracy and working condition judgment of the next round of insertion operation.
[0022] Secondly, this application provides an unmanned forklift pallet layering positioning and insertion system, comprising: The dual-camera joint calibration and spatiotemporal synchronization module is used to control the synchronous exposure of a 2D industrial camera and a structured light 3D camera through synchronous hardware trigger pulses, to acquire tray plane images and 3D point clouds at the same time and from the same perspective, to complete joint calibration and establish a unified coordinate transformation matrix from the 2D pixel coordinate system to the 3D world coordinate system. The differentiated division of labor independent recognition module is used to identify the planar features of the tray based on the tray planar image, and fit the pixel coordinates of the center of the fork hole and the pixel coordinates of the tray edge; based on the three-dimensional point cloud, it segments the point cloud clusters corresponding to each layer of the tray, and obtains the three-dimensional pose and geometric parameters of each layer of the tray. The bidirectional fusion calibration module is used to map the identified tray edge pixel coordinates to a three-dimensional coordinate system based on the unified coordinate transformation matrix obtained by joint calibration, generate a region mask for filtering and completing the three-dimensional point cloud, and perform three-dimensional point cloud filtering and completion; map the center pixel coordinates of the fork hole to the initial three-dimensional coordinates, and perform spatial coordinate calibration based on the point cloud clusters of the corresponding layer to correct the two-dimensional plane positioning drift error; The dual-vision cross-tolerance verification module is used to calculate the absolute difference between the pallet size identified and fitted by the pallet plane image and the 3D point cloud, respectively. When the absolute difference exceeds the set difference threshold, 2D or 3D single-path recognition failure complementary fitting is initiated. The parent-child pallet layered adaptive insertion control module is used to determine the current working condition based on the number of layered point cloud clusters. When the working condition is parent-child stacking, it performs vertical safety clearance detection between layers and generates fork arm avoidance adjustment instructions based on the detection results. It calls the corresponding forklift motion parameter set to perform the insertion action according to the working condition.
[0023] By adopting the above solution, 2D and 3D vision are integrated, which solves the problems of the existing single 3D vision in the scenario of not being able to distinguish the hierarchy, measure the gap accurately, find the fork holes, and adapt to the working conditions. This improves the positioning and insertion accuracy of unmanned forklifts in the operation of pallets and the stability of automated operation, reduces the frequency of manual intervention, and improves the efficiency of warehouse transfer.
[0024] In summary, this application has the following beneficial effects: 1. By synchronously controlling the exposure and acquisition of 2D and 3D cameras, joint calibration and coordinate transformation are completed. After image recognition of planar features and point cloud segmentation into layered point cloud clusters, a bidirectional data mutual compensation and fusion algorithm and a dual-vision cross-verification fault-tolerant logic are designed to perform coordinate mapping and calibration, data verification and complementary fitting. A dedicated layer judgment and adaptive insertion control process for mother and child trays is provided to achieve more accurate layered positioning and insertion of mother and child trays. 2. Joint calibration ensures accurate coordinate transformation; image preprocessing accurately identifies planar features, point cloud clustering divides the point cloud into hierarchical point cloud clusters, coordinate mapping can filter and complete the point cloud and correct two-dimensional drift, improving the positioning accuracy in stacking conditions; the design of single-path failure complementary fitting can ensure continuous operation and significantly reduce the probability of forklift downtime. 3. Consider uncertainty to more accurately correct and calibrate coordinates; complete multi-frame fusion estimation to smooth noise and remove outliers; ensure insertion safety based on safety gap detection and speed adjustment, and correct calibration parameters after insertion to improve the accuracy of subsequent operations. Attached Figure Description
[0025] Figure 1 This is a flowchart of the unmanned forklift pallet layer positioning and insertion method described in a specific embodiment; Figure 2 This is a schematic diagram of the unmanned forklift pallet layer positioning and insertion system described in a specific embodiment. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0027] To address the fundamental shortcomings of existing mature 3D vision technology, which is only compatible with single-layer standard pallet operations and suffers from unclear layer distinctions, inaccurate gap measurements, inability to locate fork holes, and inability to adapt to working conditions in scenarios involving stacked mother and daughter pallets.
[0028] This application abandons the conventional approach of single 3D perception and simple parallel acquisition by dual sensors, and constructs a differentiated collaborative fusion architecture of 2D planar fine perception and 3D spatial depth perception. It introduces a low-cost 2D industrial camera specifically to compensate for the deficiencies in planar detail recognition by 3D vision. The two types of sensors perform their respective functions, with complete task decoupling. Spatiotemporal alignment is achieved through hardware synchronous triggering and unified coordinate system calibration. A bidirectional data mutual compensation fusion algorithm and dual-vision cross-verification fault-tolerant logic are designed. A dedicated layer judgment and adaptive insertion control process for mother-child trays is also included. This approach retains the mature, reliable, and accurate depth measurement advantages of traditional 3D cameras for single-layer tray recognition, while utilizing 2D cameras to supplement the capabilities of mother-child tray classification, fine fork hole positioning, and stacked contour completion. It simultaneously solves the problem of mother-child stack recognition, which pure 3D vision cannot overcome, from three layers: hardware perception, data processing, and logic control. The following is a further detailed description of this application.
[0029] Example 1
[0030] like Figure 1As shown in the figure, this application discloses a method for layered positioning and insertion of mother and child pallets in an unmanned forklift, including steps such as: joint calibration and spatiotemporal synchronization of dual cameras, differentiated division of labor and independent recognition, bidirectional fusion calibration, dual-vision cross-tolerance verification, and layered adaptive insertion control of mother and child pallets. It employs 2D and 3D vision fusion to achieve layered positioning and insertion of mother and child pallets, overcoming the limitations of pure 3D vision and adapting to the technical effects of mixed mother and child pallet scenarios. Each step will be described in detail below.
[0031] S1. Perform joint calibration and spatiotemporal synchronization of dual cameras to collect operational data for the target area.
[0032] Specifically, the dual-camera joint calibration and spatiotemporal synchronization process involves a 2D industrial camera and a structured light 3D camera. It employs a synchronously triggered 2D high-definition industrial camera, a structured light 3D camera, and an onboard motion controller, all mounted on the front end of the unmanned forklift mast. To differentiate itself from existing devices that only perform rough software timestamp alignment, lack a unified joint calibration process, and whose two visual coordinate systems are independent, resulting in simple stitching of recognition results and fixed system errors, this application designs the following detailed execution process, including: During the factory calibration phase, a standard checkerboard calibration board is placed in the forklift operation recognition area, and 2D and 3D cameras are simultaneously triggered to acquire multiple sets of paired images and point cloud data. The intrinsic parameters of the 2D camera, distortion coefficients, and the world coordinate system parameters of the 3D camera point cloud are solved respectively. Then, a unified coordinate transformation matrix and rotation and translation matrix between the two-dimensional pixel coordinate system and the three-dimensional world coordinate system are established through multiple sets of paired feature points, and the transformation parameters are solidified and stored in the vehicle controller.
[0033] During the on-site data acquisition phase, the controller outputs the same synchronous hardware trigger pulse signal, which is simultaneously sent to the 2D camera and the 3D camera. This forces the two types of sensor devices to expose and acquire target area data synchronously, obtaining pallet planar images (two-dimensional image data) and three-dimensional point cloud data. Each set of 2D image frames and the corresponding 3D point cloud data packets are marked with a unified timestamp to ensure that the pallet target is acquired from the same perspective at the same time, thus preventing target misalignment caused by acquisition time difference.
[0034] Based on the above steps, precise calibration is first performed to ensure the uniformity of the coordinate system and the accurate correspondence of the data. The calibration matrix achieves a one-to-one mapping between pixel coordinates and three-dimensional spatial coordinates, eliminating physical installation deviations of the two sensors. Then, data is collected synchronously during operation, and hardware synchronous triggering ensures that the two sets of data are completely matched, eliminating time dimension deviations. This allows the 2D image and 3D point cloud to accurately reflect the pallet information at the same moment, providing an accurate data foundation for subsequent processing. Therefore, compared with existing technologies, this application adopts image-point cloud paired joint calibration, rather than separate 2D and 3D calibration, generating a globally unified transformation matrix to achieve unbiased mapping of two-dimensional and three-dimensional coordinates. Hardware synchronous pulse triggering replaces asynchronous software acquisition, achieving frame-level synchronization at the hardware level, rather than software post-matching, improving synchronization accuracy by an order of magnitude. The calibration parameters are solidified into the vehicle controller, eliminating the need for secondary manual calibration on-site, and adapting to mass production of forklifts and rapid on-site deployment.
[0035] S2, based on the differentiated division of labor between 2D and 3D cameras for independent recognition.
[0036] Specifically, the differentiated division of labor and independent identification process includes two-dimensional identification and three-dimensional identification, which correspond to completely independent and non-overlapping identification tasks, achieving complementary advantages and decoupling of computing power.
[0037] For 2D recognition: After synchronous triggering, the 2D camera outputs a high-resolution visible light image of the tray's planar surface. This image is then processed through a series of image processing algorithms. The execution flow includes: First, image preprocessing. This step removes reflections, filters out dust and noise, and enhances edges, such as using Gaussian filtering to remove noise and histogram equalization to enhance edges. Second, fine feature extraction. This step extracts planar features of the tray, such as the length and width of the fork openings, the thickness of the border, the printed texture on the board surface, and the hollow structure. These features can be extracted using edge detection, morphological operations, and other methods. Third, mother-child tray classification and matching. Considering multiple scenarios, this step compares the extracted planar features with pre-stored standard 2D templates for mother and child trays, outputting a tray type identifier, including single-layer mother tray / single-layer child tray / stacked upper-layer child tray. Fourth, 2D positioning calculation. The 2D positioning calculation fits the pixel coordinates of the center of the fork openings and the pixel positions of the left and right boundaries of the tray, outputting high-precision 2D planar positioning data.
[0038] Throughout the entire 2D recognition process, the 2D camera is limited to the single task of planar classification and fine 2D localization, without participating in depth or layer height calculations, thus avoiding wasted computing power. A dedicated 2D matching model is designed for the differentiated planar features of the mother and child trays, relying solely on 2D details to stably distinguish between them without depending on their 3D shape, and unaffected by the approximate 3D contour of the tray. The main purpose is to leverage its high resolution and ability to recognize planar textures and fine structures to solve the root cause problem of pure 3D vision being unable to distinguish between mother and child trays or find precise fork openings.
[0039] For 3D recognition: 3D cameras and 2D images are simultaneously acquired to obtain a complete 3D point cloud of the tray area. A series of processing steps are performed, including: First, point cloud preprocessing. This step uses statistical filtering to remove ground debris and noise from distant irrelevant objects, segmenting the target point cloud clusters of the tray. Second, multi-layer point cloud clustering and segmentation. For stacked trays, based on point cloud height thresholds and spatial connectivity algorithms, two independent point clouds are automatically split: the upper tray and the lower tray. Specifically, this includes: setting a z-axis height interval based on the standard tray thickness, segmenting the point cloud according to the z-coordinate; performing eight-neighbor spatial connectivity judgment on the point cloud within each height interval; when the projection distance between adjacent points in the xy plane is less than a first preset distance threshold and the z-axis difference is less than a second preset distance threshold, they are determined to be the same tray point cloud cluster; finally, two unconnected tray point cloud clusters are separated, corresponding to the upper tray and the lower tray, respectively. Third, 3D solution parameter calculation. The 3D parameter calculation will calculate the ground clearance of each pallet, the pallet thickness, the vertical gap between the upper and lower layers, and the pallet tilt angle (front / back / left / right). It will output the 3D pose and geometric parameters, including: the independent 3D contour point set, depth, layer height, and gap data of each pallet.
[0040] Throughout the entire 3D recognition process, the 3D camera is limited to handling 3D spatial layering and depth parameter analysis only, without participating in tray model classification and recognition, which significantly reduces the difficulty of training point cloud AI models and the consumption of computing power. A dedicated double-layer point cloud clustering and segmentation logic for parent-child stacking is designed, which is specifically optimized for the coupling and adhesion problem of stacked point clouds. It can separate tightly fitted upper and lower tray point cloud clusters, achieving a layering effect that a single 3D algorithm cannot stably achieve.
[0041] Based on the above steps, the boundaries between 2D and 3D recognition are completely defined, and there is no duplicate recognition content, which greatly reduces the computing load on the vehicle controller and meets the real-time motion control cycle of the unmanned forklift. The 2D two-dimensional details are used to solve the problem of classifying mother and child pallets, and the 3D three-dimensional depth is used to solve the problems of stacking and layering and gap measurement. The two perception dimensions are complementary in principle, rather than simply hardware superposition. The entire recognition division logic is specifically designed for the mixed storage and stacking of mother and child pallets, which is different from the existing single 3D solution and dual vision solution without division of labor.
[0042] S3. Perform bidirectional fusion calibration of 2D image and 3D point cloud.
[0043] Specifically, the two-way fusion calibration process includes positive correction of 3D point cloud defects using 2D planar features and negative compensation of 2D planar positioning deviations using 3D depth data.
[0044] To address the shortcomings of 3D point cloud correction using 2D planar features, this primarily resolves issues such as 3D point cloud coupling, edge blurring, and localized voids in stacked scenarios. Considering the overlap of point clouds between upper and lower pallets during stacking, pure 3D algorithms cannot distinguish between valid pallet point clouds and adhering noise. By leveraging clear and complete 2D contours, the true boundaries of each pallet layer are accurately located, eliminating false point clouds and improving the accuracy of layer gap and layer height calculations. The specific execution process includes: Based on the unified coordinate transformation matrix obtained by joint calibration, the complete outline of the pallet and the boundary of the fork opening obtained by 2D recognition are mapped to the 3D world coordinate system to form a standard effective recognition area mask. Based on this area mask, the 3D point cloud is filtered: based on the generated area mask, all point clouds outside the area mask are identified as ground, cargo extension, and debris noise and are removed, while the original 3D point cloud inside the area mask is retained, and the point cloud holes and missing edge areas are filled in according to the recognized pallet edge pixel coordinates.
[0045] 3D depth data is used to compensate for 2D planar positioning deviations. The main purpose is to eliminate positioning errors caused by perspective distortion and viewpoint shift in a single 2D image, assigning true depth information to the high-precision 2D forkhole positions to obtain 3D interpolation points that can be directly used for forklift motion control. The specific execution process includes: Using precise depth values output by a 3D camera and a unified coordinate transformation matrix, the pixel-level fork coordinates of the 2D image are converted into true 3D spatial coordinates, which serve as the initial 3D coordinates. To address the 2D image planar shift caused by overexposure and local occlusion, the center coordinates fitted from the same layer of the tray point cloud are used to correct the 2D positioning drift error. This includes: searching for a point set within a sphere in the corresponding layer of the point cloud cluster using the initial 3D coordinates as the center; performing least-squares plane fitting on the searched point set; and taking the center point of the fitted plane as the final 3D coordinates of the fork center to complete the spatial coordinate calibration. Additionally, the identified tray edge pixel coordinates are mapped to 3D spatial lines (including curves or straight lines), and their intersection with the plane obtained from the fitting of the corresponding layer of the point cloud cluster is calculated to obtain the corrected 3D spatial lines of the tray edge.
[0046] Based on the above steps, unlike the industry's one-way data overlay mode, it achieves 2D repair of 3D point cloud defects and 3D correction of 2D plane drift, with bidirectional complementarity, which greatly improves the positioning accuracy in stacking conditions; the layered filtering algorithm based on the two-dimensional contour to generate point cloud masks is specifically designed to solve the industry pain point of adhesion between parent and child stacked point clouds, which cannot be accurately screened by relying on a single 3D device; after fusion, it outputs a unified single-group layered three-dimensional coordinate, rather than two sets of independent recognition results, simplifying the backend motion control logic.
[0047] S4. Perform dual-vision cross-tolerance verification.
[0048] Specifically, considering that traditional single-channel 3D vision only has one source of perception data, once reflections, dust, or cargo occlusion cause point cloud distortion, the target is directly judged as lost, the forklift stops and awaits manual processing, resulting in poor automation continuity; existing simple dual-camera solutions lack cross-validation logic, the two recognition channels operate independently, and there is no complementary means after failure. Based on the above fusion results, a reliability assessment and failure handling are conducted, and a dual-vision cross-fault tolerance verification step is designed, including a consistency verification step and single-channel vision failure complementary fitting; the specific execution flow is as follows: For the consistency verification step: extract the tray plane dimensions fitted by 2D recognition (i.e., the tray plane dimensions identified and fitted based on the tray plane image, such as the length and width dimensions of the tray plane if the tray is rectangular) and the tray plane dimensions fitted by 3D point cloud (i.e., the tray dimensions identified and fitted based on 3D point cloud); calculate the absolute difference between the two sets of dimensions. If the absolute difference is less than the preset difference threshold (any relative dimension, such as the length or width of the tray), it is determined that the recognition results of the two vision sensors are consistent and the data is reliable, and the fused coordinates are directly output; if the absolute difference exceeds the preset difference threshold, it is determined that one of the vision sensors is obstructed or has recognition abnormalities due to reflection interference, and 2D or 3D single-channel recognition failure complementary fitting is initiated, entering the single-channel complementary fitting process.
[0049] The complementary fitting mechanism for single-path vision failure includes the following scenarios: First, when 3D point cloud recognition is abnormal (one vision sensor fails), such as distortion or severe voids (due to reflection / dust interference), the 2D complete planar contour and fork hole positions (i.e., the contour and fork hole positions based on pallet planar image recognition) are used as a benchmark, combined with historical frame pallet average layer height, gap parameters, and other layer height parameters to fit the 3D spatial coordinates of the layered pallet. Second, when 2D recognition is abnormal (one vision sensor fails), such as loss of fork hole features due to strong light overexposure or cargo occlusion in the 2D image (loss of pallet planar image features), the complete point cloud cluster after 3D layered segmentation is used as a benchmark, combined with pre-stored standard 2D structural dimensions of the parent and child pallets, to backfit the theoretical planar position of the fork holes of each layer of pallet, completing the positioning output. Third, the controller only outputs a shutdown alarm command when both 2D and 3D vision recognition fail simultaneously.
[0050] Based on the above steps, a cross-consistency judgment rule for two-dimensional and three-dimensional dimensions is adopted to quickly identify anomalies in single-channel sensor recognition. A bidirectional complementary fitting logic is designed, with the two vision channels serving as backup sensing sources for each other. There is no primary or secondary dependency relationship, which significantly reduces the probability of forklift downtime and solves the inherent defects of poor fault tolerance and weak adaptability of pure 3D vision. This greatly improves the continuous operation time of automation, reduces the frequency of manual intervention, and improves the efficiency of warehouse transfer.
[0051] S5, Execute parent and child tray layered adaptive insertion control.
[0052] Specifically, a layered adaptive insertion and removal control mechanism for mother and child pallets is set up to provide dedicated safety control logic for mother-child stacking scenarios, avoiding risks of top-loading, collisions, and cargo tipping from a control perspective; it is compatible with single-layer ordinary pallet operations, balancing versatility and safety in stacking scenarios. The layered adaptive insertion and removal control mechanism for mother and child pallets includes automatic working condition determination, layered safety gap detection, adaptive motion parameter matching, and output control commands. The specific execution process includes: Automatic determination of working conditions: Based on the number of layered point cloud clusters obtained from 3D segmentation, three types of working conditions are automatically distinguished. If only one set of point cloud clusters is determined to be a single-layer pallet, it is further distinguished as a single-layer mother pallet / single-layer child pallet by combining the 2D classification results. If two independent sets of point cloud clusters are determined to be a mother-child stacking working condition.
[0053] For layered safety gap detection: When the working condition is a mother-child stacking condition, the vertical safety gap between the bottom of the upper child support and the top of the lower mother support is calculated in real time under this stacking condition to determine the feasibility of prioritizing the insertion of the upper child support; if the gap is less than the safe passage threshold of the fork arm, an avoidance adjustment command is output.
[0054] For adaptive motion parameter matching: The system calls the corresponding forklift motion parameter set based on the operating condition to execute the insertion action. Specifically, this includes: if the current operating condition is a single-layer mother pallet, reusing mature and universal insertion parameters ensures efficient single-layer pallet transfer; if the current operating condition is a single-layer child pallet, automatically lowering the forklift's base lifting height to match the lower size of the child pallet; if the current operating condition is a mother-child stacking operation, adaptively adjusting the forklift's forward offset, forklift lifting avoidance height, and fork insertion depth based on the fused target layer's 3D coordinates and pallet tilt angle; the forklift is first raised to a safe clearance height above the target pallet before horizontally inserting the forklift forward to avoid impacting the lower / upper pallet.
[0055] Regarding output control commands: Continuous motion control commands are output to the forklift drive system to complete the precise and safe insertion and removal of layered pallets.
[0056] Based on the above steps, an automatic three-condition judgment mechanism for mother and child trays based on 2D+3D fusion layered data is established to achieve accurate condition differentiation; a safety prediction of layered gaps in stacked scenarios and a lifting avoidance control strategy are adopted to solve the problem that existing technologies lack gap interference detection logic for mother and child trays; and a dynamic matching of adaptive parameters for multiple working conditions is set to break the control mode of fixed motion parameters for traditional tray insertion and removal, and adapt to mother and child trays with different sizes and specifications.
[0057] By employing steps S1-S5 above, this approach breaks through the industry's inherent reliance on a single 3D vision method. It adopts a dual-sensor architecture that integrates 2D and 3D perception with differentiated division of labor, addressing the fundamental shortcomings of pure 3D vision in distinguishing between mother and child trays and facilitating layered identification. A dual-camera joint calibration, synchronous hardware acquisition, and bidirectional 2D / 3D calibration fusion algorithm are designed to resolve issues of stacked point cloud coupling and positioning offset. A redundant perception mechanism with dual-vision cross-verification and single-path failure complementarity is established to address the poor adaptability and frequent downtime issues associated with single 3D vision. Furthermore, a dedicated layered identification, gap detection, and adaptive avoidance insertion control process for mother and child trays is implemented, filling the technological gap in the industry for automated and safe operation of stacked mother and child trays.
[0058] Example 2 The difference from Embodiment 1 above is that a three-stage probabilistic spatiotemporal fusion and verification engine are integrated into bidirectional fusion calibration to improve the accuracy and stability of pallet pose estimation and reduce the impact of outliers on the system. The method includes: Specifically, considering that the original bidirectional fusion calibration method (2D forward correction of 3D defects, 3D reverse compensation of 2D deviations) is a hard cross-correction, if there are uncertainties in the 2D planar features or 3D, it may lead to erroneous data correction and compensation. To avoid this, a three-stage engine is embedded into the original bidirectional fusion calibration logic to optimize the bidirectional fusion calibration. This three-stage engine includes: a single-frame probabilistic fusion stage, a multi-frame temporal verification and filtering stage, and a physical rule arbitration stage. The optimization steps of bidirectional fusion calibration are described in detail below: First, complete the modeling of the source uncertainty.
[0059] Specifically, for 2D feature uncertainty modeling, a pre-built detection network model (e.g., using a neural network) is constructed. The tray image captured by the 2D camera is input into the detection network model. The model outputs the pixel coordinates of the center of the fork opening and the pixel coordinates of the tray edge, as well as the covariance matrix or probability distribution representing the detection uncertainty (belonging to the tray category). For 3D point cloud uncertainty modeling, a measurement noise model is calculated for each 3D point cloud based on the depth camera principle, such as the depth measurement variance of each point. It is a quadratic function of its depth value z, and the formula is: The greater the distance, the greater the noise, and k is set to 0.005; therefore, a measurement noise model is established based on its depth value and surface characteristics, and depth uncertainty is assigned to each point.
[0060] Secondly, based on uncertainty propagation and single-frame probability fusion, 2D correction to 3D and 3D compensation to 2D are completed simultaneously.
[0061] Among them, the 2D contour correction for 3D point cloud layering: the main goal is to guide the layering of point cloud by using the clear boundaries of the 2D image when the mother and child trays are stuck together in the 3D point cloud. Therefore, the following is set: based on the uncertainty of the pixel coordinates of the tray edge, a probability region mask is generated, and the probability region mask is used to perform weighted filtering and completion of the 3D point cloud, so as to positively correct the defects of the 3D point cloud through the 2D planar features.
[0062] Specifically, the probability map obtained based on the detection network model, which is the same size as the input tray image, is used... The value of each pixel represents the probability that the pixel belongs to the tray category, and is back-projected into 3D space through camera intrinsics. Since monocular vision lacks depth, it forms a cone-shaped probability volume emanating from the optical center. This probability volume is intersected with the 3D point cloud region to be corrected. For each point in the point cloud, a weight is assigned based on its position within the probability volume, indicating its association with a boundary. In regions with high probabilities (greater than a preset probability threshold), strong constraints are applied to establish the segmentation boundary. Under these constraints, the point cloud is re-segmented or fitted to a plane. The algorithm forces the splitting of different planes where the 2D probability indicates a strong boundary, thus separating adhered child trays from the parent tray. The entire process introduces 2D probabilities as weights, running a weighted point cloud segmentation algorithm.
[0063] Specifically, for 3D depth reverse compensation 2D fork hole localization: the main goal is to use 3D depth information to compensate for the deviation of 2D fork hole coordinates due to illumination or viewing angle, and accurately locate it to the correct spatial position. Therefore, the following is set up: through the uncertainty propagation algorithm, the detection uncertainty of the fork hole center pixel coordinates is fused with the depth uncertainty of the corresponding point cloud cluster to obtain a 3D fork hole position estimate with covariance matrix.
[0064] Specifically, the detection uncertainty for obtaining the center pixel coordinates of the fork hole: 2D fork hole mean. With covariance Sampling 3D depth with uncertainty: Within the 3D point cloud neighborhood corresponding to the center of the 2D fork hole, sample local depth values z, and calculate the variance of this depth as the uncertainty according to the above model. An uncertainty propagation algorithm is employed, such as using the classic camera projection model. ; This indicates the specific location of the center of the fork in the camera coordinate system; it is a three-dimensional point coordinate in space. For the camera intrinsic parameter matrix, input: coordinates of the fork hole center. and depth The nonlinear transformation propagation law of Gaussian distribution (such as the unscented transformation UT algorithm) is selected to select a set of Sigma points (containing these mean and covariance information), project them one by one into 3D space, and statistically analyze the distribution of these 3D points to obtain the three-dimensional cross hole position estimate with covariance matrix.
[0065] Then, multi-frame temporal fusion is performed.
[0066] Specifically, the pallet is modeled as a rigid body running system, and optimal state estimation is performed using EKF. The state vector is defined as the six-DOF pose (position + attitude) of the pallet in the global coordinate system. And maintain a state covariance matrix P. Perform motion model prediction, using the forklift's own odometer as input to establish a motion model. , This is process noise; To control the input vector, measurements from the forklift chassis odometer, IMU, etc., are used; the observation model is updated by using a 3D forkhole position estimate with a covariance matrix as the observation. Establish observation equations , To observe the noise, This function transforms the fork hole positions in the currently estimated tray pose to the sensor coordinate system. It inputs the 3D fork hole positions with covariance matrix into a pre-built extended Kalman filter, performs EKF update, updates the state X and covariance P, realizes multi-frame temporal fusion estimation of tray pose, smooths single-frame noise (continuous uncertain observations will be automatically weighted and averaged by EKF) and removes abnormal observations (transient erroneous observations caused by vibration or occlusion, whose covariance is necessarily huge, will be automatically suppressed or even removed by Kalman gain).
[0067] Finally, after multi-frame temporal fusion estimation, physical consistency verification is performed.
[0068] Specifically, the physical consistency verification includes: rigid body invariance check: calculate the spatial distance between the centers of the left and right fork holes in the current frame and compare it with the known standard size of the tray. If the deviation exceeds the preset threshold, discard the fusion result of the current frame and trigger repositioning; or geometric feasibility check: check whether the z-coordinate of the bottom surface of the upper sub-tray is strictly greater than the z-coordinate of the top surface of the lower mother tray to prevent model interpenetration.
[0069] Example 3 The difference from embodiments 1 and 2 above lies in that, based on the forklift tooth structure and pallet type parameters, combined with the surface of the mother pallet and child pallet reconstructed from 3D point cloud, the dynamic effective safety clearance is accurately calculated and compared with a dynamic safety threshold. This enables safety judgment and forklift tooth elevation angle adjustment for inserting the upper child pallet. Simultaneously, the insertion speed is dynamically adjusted based on the ratio of the clearance to the threshold, improving the safety and flexibility of unmanned forklift insertion operations under mother-child pallet stacking conditions. The method includes: the step of performing inter-layer vertical safety clearance detection and generating forklift avoidance adjustment commands based on the detection results includes: First, establish the parameter set for the forklift tooth structure and the parameter set for the pallet type.
[0070] Specifically, the forklift tooth structure parameter set includes: tooth thickness, tooth width, tooth length, maximum pitch adjustment angle, and maximum pitch adjustment angle of the tooth; the pallet type parameter set includes: total pallet height, fork hole entry height, fork hole width, center distance between two fork holes, bottom structure type (pallet bottom ground clearance, specifically whether the pallet bottom has a bottom plate or crossbeam) and standard dimensions (standard dimensions of various pallets).
[0071] Secondly, considering geometric constraints, we completed the real-time calculation and feasibility assessment of the layer gaps.
[0072] Specifically, in stacking operations, the core of determining whether to prioritize the insertion of the upper-layer sub-carrier is to dynamically verify the three-dimensional physical interference on the fork tooth entry path. Therefore, the design is as follows: First, a three-dimensional safety clearance model is reconstructed, no longer calculating only a single vertical distance, but constructing a complete interference inspection space. The input includes: the point set of the bottom surface of the sub-support and the point set of the top surface of the mother support obtained by 3D point cloud segmentation, which are used to complete the modeling of the top surface of the mother support, the bottom surface of the sub-support, and the fork-tooth passageway, respectively.
[0073] The modeling process includes: Modeling the top surface of the main pallet: Fitting the highest plane of the main pallet top, considering the possibility of uneven cargo, and taking the highest envelope surface within a local area (such as the fork projection range) as the safety benchmark. Modeling the bottom of the sub-pallet: For the sub-pallet, identifying its bottom structural features. If the sub-pallet has no bottom slab, the bottom is an open cavity, and the danger boundary is the lower edge of the bottom crossbeam of the sub-pallet; if it has a complete bottom slab, the danger boundary is the entire bottom plane. Based on the pallet model parameters, matching the known bottom crossbeam position and height, the protruding structure is accurately segmented in the point cloud. Modeling the fork passageway: Combining the fork insertion position, spacing, and fork width of the current operation plan, a bounding box representing the fork sweep body, i.e., the fork passage sweep body, is generated in three-dimensional space.
[0074] Based on the reconstructed 3D safety clearance model, the dynamic minimum clearance is calculated. This includes calculating the minimum vertical distance between the upper edge of the fork-tooth sweep body and the lower edge of the sub-support bottom structure, and the minimum vertical distance between the lower edge of the fork-tooth sweep body and the highest envelope surface of the top of the main support. The smaller of the two values is taken as the dynamic effective safety clearance.
[0075] Second, an adaptive feasibility assessment logic is implemented. The safe passage threshold is no longer set to a fixed value, but rather a dynamically calculated value. It is dynamically calculated based on the fork tooth thickness, current insertion speed, and preset safety margin. The formula is: In the formula, The thickness of the fork teeth ensures that even at the point of minimum clearance, The preset fork tooth thickness is set to 10mm; This represents the current fork insertion speed. For the dynamic safety factor, vibration and overshoot settings during motion are taken into account; A constant safety margin set by humans.
[0076] A three-level judgment is performed: the dynamic effective safety clearance is compared with the dynamic safety threshold. If the dynamic effective safety clearance is greater than or equal to the dynamic safety threshold, it indicates that there is sufficient space at the bottom of the upper sub-staple, and it is determined that the upper sub-staple can be directly inserted. If the dynamic effective safety clearance is less than the dynamic safety threshold, the insertion command needs to be adjusted or an avoidance command needs to be output directly. The insertion command after adjustment refers to creating space by adjusting the fork tooth posture. For example, if the sub-staple is stacked incorrectly, resulting in a front-lower-rear height difference, an elevation angle is calculated so that the fork tooth tip is slightly lower than the root, thus safely entering the narrow front gap. Within the maximum pitch adjustment angle range, the new sweeping body after adjusting the fork tooth elevation angle is calculated, the clearance is rechecked, and an avoidance adjustment command containing the elevation angle adjustment is generated when the new sweeping body passes. If the adjusted clearance is less than the dynamic safety threshold, an avoidance command is output, such as suggesting that the mother tray stacking be adjusted first, or issuing an alarm prompting manual reorganization.
[0077] Then, optimization of adaptive motion parameter matching is performed.
[0078] Specifically, different forklift tooth structures and pallet types determine the optimal motion control parameters. The insertion speed is dynamically adjusted based on the ratio of the dynamic effective safety clearance to the dynamic safety threshold. When this ratio approaches 1, the insertion speed is automatically reduced to a preset jogging speed. This optimizes the execution of the insertion action by calling the corresponding forklift motion parameter set based on the working conditions.
[0079] Dynamically adjusting the insertion speed includes: setting reference parameters. This means setting the forklift feed speed based on the pallet type and implementing dynamic speed limiting related to the clearance. The formula is: In the formula, Indicates the dynamic effective safety clearance. Indicates the dynamic safe passage threshold; This represents the buffer spacing, a calibration parameter used to adjust the smoothness of speed changes.
[0080] Example 4 The difference from embodiments 1 and 2 above lies in that the recognition error is calculated using the actual spatial offset collected by the macro sensor, and the joint calibration extrinsic parameters are incrementally corrected online and the classifier feature weights are updated. This improves the perception accuracy and working condition judgment accuracy of subsequent insertion operations, making the unmanned forklift more adaptable and its operation more precise and reliable during the layered positioning and insertion of mother and child pallets. The method includes: First, obtain online feedback values and calculate the recognition error.
[0081] Specifically, after each insertion / removal action, a micro-sensor located at the end of the fork arm, such as a high-precision laser displacement sensor or a contact micro-switch installed at the tip or side of the fork tooth, collects the actual spatial offset of the fork tooth relative to the pallet fork hole when the fork tooth smoothly enters the fork hole and completes insertion. This includes: Actual 3D offset and attitude angle deviation. At the same time point of successful interpolation, the original 2D and 3D data of that frame and the fusion calculation results are extracted to obtain the output 3D position of the fork hole and the pallet plane normal vector, which are used as theoretical output values. Based on the actual pose of the fork teeth and the geometric relationship of the pallet, the true 3D position of the fork hole is deduced as the true value. The recognition error is calculated, which includes position error, depth error, and attitude error.
[0082] Secondly, by utilizing the accumulated recognition error, the joint calibration extrinsic parameters between the pallet planar image and the 3D point cloud are incrementally corrected online, and the joint calibration parameters are adjusted.
[0083] Specifically, assuming the error mainly originates from minute rotations and translations of the extrinsic parameters, the online learning process includes: Load the offline calibration parameters and the learned correction values. The offline calibration parameters include: extrinsic parameters. and Initial depth scale factor The learned corrections include: parameterization of the defined variables to be optimized and extrinsic parameters; and corrections for translational components. (translation component) And depth scale factor s, and a tiny rotation correction. This constitutes an effective extrinsic parameter. The forklift uses the currently effective extrinsic parameters to perform 2D-3D fusion and complete the insertion. After successful insertion, the macro sensor captures the actual offset, calculates the 3D ground truth value by combining it with the forklift kinematics, and then calculates the cumulative recognition error. If the cumulative recognition error is within a preset reasonable range, the learning step begins; otherwise, it is discarded. The stochastic gradient descent (SGD) formula is used to update the parameters with minimal step size. ,s, The update amount is truncated to ensure that the parameters are within the physically feasible domain.
[0084] Then, the identification error is correlated with the number of layered point cloud clusters and the height difference features before insertion, and the classifier feature weights used to determine the current working condition are updated online.
[0085] Specifically, the feature vectors at the instant before insertion are recorded, such as the average height of the upper layer point cloud cluster, the average height of the lower layer, the gap, the point density, etc., as well as the real working condition labels after insertion: single-layer mother support, single-layer child support, stacked child support, etc.; online learning of weights is performed. For example, if the current working condition classifier uses a logistic regression classifier, given the feature vectors and real labels, the real working condition is predicted. If the predicted output of the real working condition matches the real standard, the weight vector is updated according to the learning rate.
[0086] The revised joint calibration parameters and classifier feature weights are applied to the perception and condition determination of the next round of insertion operations.
[0087] Example 5 The method differs from embodiments 1 and 2 above in that it establishes an arbitration state machine to manage different recognition result states, more accurately handling the consistency problem of dual-vision recognition results; performing false consistency verification can avoid erroneous decisions caused by superficial consistency but actual deviations; and multi-evidence accumulation arbitration can more reasonably select trust sources for complementary fitting output, achieving reliability and robustness in complex situations. The method includes: the step of initiating complementary fitting for 2D or 3D single-path recognition failure further includes: First, define the arbitration state machine.
[0088] Specifically, four distinct states are defined, with transitions between each state triggered by both "conflict level" and "evidence score." State S0: High-confidence fusion state, primarily describing consistent matching of the fitting data from the two visual sensors (2D / 3D), directly outputting fused 2D / 3D matching data; State S1: Single-path degraded fitting state, indicating that one path's matching data is deemed unreliable, trusting the other, and outputting a compensated fitting result based on the reliable path; State S2: Conflict pending observation state, primarily describing severe conflict between the two matching data paths, or suspicion of false consistency, entering a questionable state, temporarily not outputting new data, instead maintaining the previous cycle's safe state or relying on short-term predictions from the motion model, continuously collecting evidence within a preset time window including the quality, temporal stability, and conformity with physical priors of each matching result; State S3: Safe parking alarm state, where the conflict persists and cannot be arbitrated, or false consistency is confirmed, triggering an emergency stop and requesting manual intervention.
[0089] Thus, the state transition methods include: S0-S1: A clear single-path anomaly is detected; S0-S2: A serious collision or false consistency is detected; S1-S0: Data on the abnormal path is restored, and consistency is re-established with the trusted path in the next frame; S2-S0: Accumulated evidence indicates that the collision is transient noise, and the data returns to normal; S2-S3: Within the time window, the collision cannot be resolved or the false consistency is absolutely verified.
[0090] Secondly, quantify the conflict level and the cumulative confidence level of multiple pieces of evidence.
[0091] Specifically, conflicts include not only inconsistent fitted output dimensions, but also inconsistent poses. In this embodiment, conflicts are divided into three levels: Level 0: no conflict, the size difference is less than the preset size difference value, and the pose similarity is greater than the preset similarity value; Level 1: slight conflict, the size difference is less than the preset size difference value, and the pose similarity is not greater than the preset similarity value; Level 2: the size difference is not less than the preset size difference value, and the pose similarity is not greater than the preset similarity value.
[0092] Specifically, a multi-evidence cumulative confidence score is used. Dynamic confidence scores are constructed for both 2D and 3D evidence. , The score is a weighted sum of multiple pieces of evidence, including: instantaneous quality score (single frame): 2D: confidence weighted calculation of image feature extraction; 3D: product of point rate in the point cloud of the fork region and normalized value of plane fitting residual; temporal stability score (multiple frames): calculation of the variance of each output (size, pose) within the sliding window, the smaller the variance, the higher the score; physical prior compliance score: the degree of matching between the measured size and the standard library size, and whether the rigid body constraint (constant fork spacing) is met.
[0093] Then, after collecting the current frame's 2D / 3D data and performing fusion calibration and consistency verification steps, the arbitration and false consistency detection process driven by the state machine is executed.
[0094] First, when the current state is S0, calculate the conflict level. When the conflict level is 0, perform a false consistency check. The false consistency check includes comparing the pallet size identified, fitted, and fused based on the pallet plane image and the 3D point cloud with the pre-stored standard pallet size. If the size deviation exceeds the preset standard deviation threshold, it is judged as a suspected false consistency.
[0095] If the false consistency check passes, maintain state S0 and output the fused coordinates. If the false consistency check fails and the state is suspected of being falsely consistent, drive the arbitration state machine to transition from the high-confidence fusion state S0 to the conflict-pending observation state S2.
[0096] Second, when the state enters state S2, the output of new coordinates is paused, and short-term prediction is performed using the effective pose and motion model of the previous S0 cycle. In the next N frames, the dynamic confidence score is continuously observed, that is, the observation evidence is accumulated within a preset time window, and the multi-evidence cumulative confidence score is calculated.
[0097] If, after accumulating N frames, the conflict level drops to 0 and the suspicion of false consistency is eliminated, the system transitions back to S0. If, after accumulating N frames, the cumulative confidence score of one multi-evidence path remains below the preset score threshold, and the cumulative confidence score of another multi-evidence path remains above the preset score threshold (maintaining health and trust), the trusted path is selected, and it is used to fit the other data for compensation output, and the system enters state S1. If, after accumulating N frames, the conflict level rises to 2, or the suspicion of false consistency is still not eliminated, the arbitration state machine is driven to transition to the safe parking alarm state S3.
[0098] Third, when the state enters S1, if the conflict level drops to 0 for N consecutive frames and the cumulative confidence score of the failed path is greater than the preset score threshold, the system transitions back to S0. If the cumulative confidence score of the trusted path remains below the preset score threshold or the conflict between the two data paths escalates further, the system transitions back to S3.
[0099] like Figure 2 As shown, this application provides an unmanned forklift pallet layering positioning and insertion system, comprising: The dual-camera joint calibration and spatiotemporal synchronization module 100 is used to control the synchronous exposure of a 2D industrial camera and a structured light 3D camera through synchronous hardware trigger pulses, to acquire tray plane images and 3D point clouds at the same time and from the same perspective, to complete joint calibration and establish a unified coordinate transformation matrix from the two-dimensional pixel coordinate system to the three-dimensional world coordinate system. The differentiated division of labor independent recognition module 200 is used to identify the planar features of the tray based on the tray planar image, and fit the pixel coordinates of the center of the fork hole and the pixel coordinates of the tray edge; based on the three-dimensional point cloud, it segments the point cloud clusters corresponding to each layer of the tray, and obtains the three-dimensional pose and geometric parameters of each layer of the tray. The bidirectional fusion calibration module 300 is used to map the identified tray edge pixel coordinates to a three-dimensional coordinate system based on the unified coordinate transformation matrix obtained by joint calibration, generate a region mask for filtering and completing the three-dimensional point cloud, and perform three-dimensional point cloud filtering and completion; map the center pixel coordinates of the fork hole to the initial three-dimensional coordinates, and perform spatial coordinate calibration based on the point cloud cluster of the corresponding layer to correct the two-dimensional plane positioning drift error; The dual-vision cross-tolerance verification module 400 is used to calculate the absolute difference between the pallet size identified and fitted by the pallet plane image and the 3D point cloud, respectively. When the absolute difference exceeds the set difference threshold, 2D or 3D single-path recognition failure complementary fitting is initiated. The parent-child pallet layered adaptive insertion control module 500 is used to determine the current working condition based on the number of layered point cloud clusters. When the working condition is parent-child stacking, it performs vertical safety clearance detection between layers and generates fork arm avoidance adjustment instructions based on the detection results. It calls the corresponding forklift motion parameter set to perform the insertion action according to the working condition.
[0100] In one specific embodiment, the system further includes: a mother-child pallet layered insertion control feedback module 600, used to collect the actual spatial offset of the fork teeth relative to the pallet fork holes by a micro-sensor set at the end of the fork arm after each insertion action, and calculate the recognition error; the recognition error includes position error and depth error; using the accumulated recognition error, the joint calibration extrinsic parameters between the pallet planar image and the three-dimensional point cloud are incrementally corrected online, and the joint calibration parameters are adjusted; the recognition error is associated with the number of layered point cloud clusters and height difference features corresponding to the insertion before insertion, and the classifier feature weights used to determine the current working condition are updated online; wherein, the corrected joint calibration parameters and classifier feature weights are applied to the perception and working condition determination of the next round of insertion operation.
[0101] This application also discloses a computer-readable storage medium.
[0102] Specifically, the computer-readable storage medium stores a computer program that can be loaded by a processor and executed, such as the unmanned forklift pallet layering positioning and insertion method described above. The computer-readable storage medium includes, for example, various media that can store program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0103] This application also discloses a computer device.
[0104] Specifically, the computer device includes a memory and a processor, and the memory stores a computer program that can be loaded by the processor and executed to perform the above-described unmanned forklift pallet layering positioning and insertion method.
[0105] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A method for layered positioning and insertion of mother and child pallets using an unmanned forklift, characterized in that, include: By controlling the synchronous exposure of a 2D industrial camera and a structured light 3D camera through synchronous hardware trigger pulse control, the pallet plane image and 3D point cloud are acquired at the same time and from the same perspective, and the joint calibration is completed and a unified coordinate transformation matrix from the 2D pixel coordinate system to the 3D world coordinate system is established. Based on the pallet planar image, the planar features of the pallet are identified, and the pixel coordinates of the fork hole center and the pixel coordinates of the pallet edge are fitted; based on the three-dimensional point cloud, the point cloud clusters corresponding to each layer of the pallet are segmented, and the three-dimensional pose and geometric parameters of each layer of the pallet are obtained. Based on the unified coordinate transformation matrix obtained by joint calibration, the identified tray edge pixel coordinates are mapped to a three-dimensional coordinate system to generate a region mask for filtering and completing the three-dimensional point cloud, and the three-dimensional point cloud filtering and completion are performed; the fork hole center pixel coordinates are mapped to initial three-dimensional coordinates, and spatial coordinate calibration is performed based on the point cloud clusters of the corresponding layer to correct the two-dimensional plane positioning drift error; Calculate the absolute difference between the pallet size identified and fitted based on the pallet planar image and the 3D point cloud, respectively. When the absolute difference exceeds the set difference threshold, initiate 2D or 3D single-path recognition failure complementary fitting. The current working condition is determined based on the number of layered point cloud clusters. When the working condition is a parent-child stacking condition, the vertical safety gap between layers is detected and a fork arm avoidance adjustment command is generated based on the detection result. The corresponding forklift motion parameter set is called according to the working condition to execute the insertion action.
2. The unmanned forklift pallet layer positioning and insertion method according to claim 1, characterized in that, The joint calibration is completed by acquiring multiple sets of chessboard calibration board images and point cloud data simultaneously. The calibration parameters include 2D camera intrinsic parameters, distortion coefficients, 3D camera point cloud world coordinate system parameters, and rotation and translation matrices and homogeneous transformation matrices from the two-dimensional pixel coordinate system to the three-dimensional world coordinate system. The calibration parameters are stored in the forklift's onboard motion controller.
3. The unmanned forklift pallet layer positioning and insertion method according to claim 1, characterized in that, The steps of identifying the planar features of the tray based on the tray planar image and fitting the pixel coordinates of the center of the fork hole and the pixel coordinates of the tray edge include: preprocessing the tray planar image and extracting the length and width of the fork hole, the thickness of the border and the texture features of the board surface; classifying the tray type according to the pre-stored two-dimensional templates of the mother tray and the child tray; and fitting the pixel coordinates of the center of the fork hole and the pixel coordinates of the tray edge. The steps of segmenting the point cloud clusters corresponding to each layer of trays based on the three-dimensional point cloud and obtaining the three-dimensional pose and geometric parameters of each layer of trays include: filtering and denoising the three-dimensional point cloud and performing multi-layer clustering segmentation; separating two independent point cloud clusters of the upper sub-tray and the lower mother tray based on the height threshold and spatial connectivity analysis; and calculating the ground clearance of each layer, the vertical gap between layers, and the spatial attitude angle of each layer.
4. The unmanned forklift pallet layer positioning and insertion method according to claim 1, characterized in that, The steps of performing 3D point cloud screening and completion include: based on the generated region mask, all point clouds outside the region mask are identified as ground, cargo expansion, and debris noise and are removed, the original 3D point cloud inside the region mask is retained, and the point cloud holes and missing edge areas are completed according to the identified pallet edge pixel coordinates. The steps of mapping the center pixel coordinates of the fork hole to initial three-dimensional coordinates and performing spatial coordinate calibration based on the point cloud clusters of the corresponding layer to correct the two-dimensional planar positioning drift error include: The two-dimensional fork hole center pixel coordinates are mapped to initial three-dimensional coordinates. Using these initial three-dimensional coordinates as the center, a point set within the sphere domain is searched in the corresponding layer point cloud cluster. The searched point set is fitted with a least-squares plane, and the center point of the fitted plane is taken as the final three-dimensional coordinates of the fork hole center to complete the spatial coordinate calibration. The identified tray edge pixel coordinates are mapped to three-dimensional spatial lines, and the intersection with the plane obtained by fitting the corresponding layer point cloud cluster is calculated to obtain the corrected tray edge three-dimensional spatial lines.
5. The unmanned forklift pallet layer positioning and insertion method according to claim 1, characterized in that, The steps for initiating 2D or 3D single-path identification failure complementary fitting include: When the 3D point cloud is distorted, the contour and fork hole positions of the pallet plane image are used as a reference, and the 3D spatial coordinates of the layered pallet are fitted in combination with historical layer height parameters; when the pallet plane image feature recognition is abnormal, the 3D point cloud cluster is used as a reference, and the theoretical plane position of the fork hole of each layer of pallet is fitted in reverse in combination with the pre-stored standard pallet 2D dimensions; an alarm command is output only when both visual recognition channels fail.
6. The unmanned forklift pallet layer positioning and insertion method according to claim 1, characterized in that, The method further includes: For the fitted pixel coordinates of the fork hole center and the pixel coordinates of the tray edge, a pre-built detection network is used to output the covariance matrix or probability distribution representing the detection uncertainty, respectively; for each point in the 3D point cloud, a measurement noise model is established based on its depth value and surface characteristics, and depth uncertainty is assigned to each point. The step of mapping the identified tray edge pixel coordinates to a three-dimensional coordinate system based on the unified coordinate transformation matrix obtained by joint calibration, generating a region mask for filtering and completing the three-dimensional point cloud, and performing three-dimensional point cloud filtering and completion further includes: generating a probabilistic region mask based on the uncertainty of the tray edge pixel coordinates, and using the probabilistic region mask to perform weighted filtering and completion of the three-dimensional point cloud; the step of mapping the fork center pixel coordinates to initial three-dimensional coordinates, and performing spatial coordinate calibration based on the point cloud clusters of the corresponding layer to correct the two-dimensional plane positioning drift error further includes: fusing the detection uncertainty of the fork center pixel coordinates with the depth uncertainty of the corresponding point cloud clusters in the vicinity of the fork center pixel coordinates through an uncertainty propagation algorithm to obtain a three-dimensional fork position estimate; The obtained three-dimensional fork hole position estimate is used as an observation and input into a pre-constructed extended Kalman filter to perform multi-frame temporal fusion estimation of the tray pose, smooth single-frame noise and remove abnormal observations; After multi-frame temporal fusion estimation, physical consistency verification is performed. The physical consistency verification includes: calculating the spatial distance between the centers of the fork holes obtained in the current frame and comparing it with the known standard size of the pallet. If the deviation exceeds a preset threshold, the fusion result of the current frame is discarded and relocation is triggered.
7. The unmanned forklift pallet layer positioning and insertion method according to claim 1, characterized in that, The step of performing inter-layer vertical safety clearance detection and generating a forklift avoidance adjustment command based on the detection result includes: A set of parameters for the forklift tooth structure and a set of parameters for the pallet type are pre-established. The set of parameters for the forklift tooth structure includes tooth thickness, tooth width, and maximum pitch adjustment angle. The set of parameters for the pallet type includes fork hole entry height, fork hole width, bottom structure type, and standard dimensions. Based on the reconstruction of the highest envelope surface of the top of the mother bracket and the lower edge surface of the bottom structure of the child bracket using 3D point cloud, and combined with the thickness of the fork tooth and the insertion position of the current operation plan, a fork tooth passage sweep body is generated. Calculate the minimum vertical distance between the upper edge of the fork-tooth sweep body and the lower edge of the bottom structure of the sub-support, and the minimum vertical distance between the lower edge of the fork-tooth sweep body and the highest envelope surface of the top of the mother support, and take the smaller of the two values as the dynamic effective safety clearance. The dynamic effective safety gap is compared with the dynamic safety threshold. The dynamic safety threshold is dynamically calculated based on the fork tooth thickness, the current insertion speed, and the preset safety margin. If the dynamic effective safety gap is greater than or equal to the dynamic safety threshold, it is determined that the upper sub-support can be directly inserted. If the dynamic effective safety gap is less than the dynamic safety threshold, a fork tooth adjustment command or a stop insertion command is generated. The step of calling the corresponding set of forklift motion parameters to perform the insertion action according to the working conditions includes: dynamically adjusting the insertion speed according to the ratio of the dynamic effective safety clearance to the dynamic safety threshold.
8. The unmanned forklift pallet layer positioning and insertion method according to claim 1, characterized in that, The method further includes: After each insertion and removal action is completed, the actual spatial offset of the fork teeth relative to the pallet fork hole is collected by a micro-sensor set at the end of the fork arm, and the recognition error is calculated. By utilizing the accumulated recognition error, the joint calibration extrinsic parameters between the pallet planar image and the 3D point cloud are incrementally corrected online, and the joint calibration parameters are adjusted. The identification error is correlated with the number of layered point cloud clusters and the height difference features before insertion, and the classifier feature weights used to determine the current working condition are updated online. The revised joint calibration parameters and classifier feature weights are applied to the perception and condition determination of the next round of insertion operations.
9. A layered positioning and insertion system for unmanned forklift mother and child pallets, characterized in that, include: The dual-camera joint calibration and spatiotemporal synchronization module is used to control the synchronous exposure of a 2D industrial camera and a structured light 3D camera through synchronous hardware trigger pulses, to acquire tray plane images and 3D point clouds at the same time and from the same perspective, to complete joint calibration and establish a unified coordinate transformation matrix from the 2D pixel coordinate system to the 3D world coordinate system. The differentiated division of labor independent recognition module is used to identify the planar features of the tray based on the tray planar image, and fit the pixel coordinates of the center of the fork hole and the pixel coordinates of the tray edge; based on the three-dimensional point cloud, it segments the point cloud clusters corresponding to each layer of the tray, and obtains the three-dimensional pose and geometric parameters of each layer of the tray. The bidirectional fusion calibration module is used to map the identified tray edge pixel coordinates to a three-dimensional coordinate system based on the unified coordinate transformation matrix obtained by joint calibration, generate a region mask for filtering and completing the three-dimensional point cloud, and perform three-dimensional point cloud filtering and completion; map the center pixel coordinates of the fork hole to the initial three-dimensional coordinates, and perform spatial coordinate calibration based on the point cloud clusters of the corresponding layer to correct the two-dimensional plane positioning drift error; The dual-vision cross-tolerance verification module is used to calculate the absolute difference between the pallet size identified and fitted by the pallet plane image and the 3D point cloud, respectively. When the absolute difference exceeds the set difference threshold, 2D or 3D single-path recognition failure complementary fitting is initiated. The parent-child pallet layered adaptive insertion control module is used to determine the current working condition based on the number of layered point cloud clusters. When the working condition is parent-child stacking, it performs vertical safety clearance detection between layers and generates fork arm avoidance adjustment instructions based on the detection results. It calls the corresponding forklift motion parameter set to perform the insertion action according to the working condition.