Multi-camera visual guidance adaptive stacking pose planning method for automatic loading and unloading scene

By using multi-camera vision guidance and local environment reconstruction technology, the problems of low global search efficiency, missing data, and insufficient adaptability in traditional automated loading and unloading scenarios have been solved, achieving efficient and accurate stacking posture planning that can adapt to complex environmental changes.

CN120735001BActive Publication Date: 2026-04-24BEIJING ADVANCED DIGITAL TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ADVANCED DIGITAL TECH
Filing Date
2025-05-15
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional automated loading and unloading scenarios suffer from problems such as low global search efficiency, inaccurate planning due to missing data, insufficient adaptability, and lack of dynamic verification mechanisms, making it difficult to meet the stacking requirements in real-time and complex environments.

Method used

Employing multi-camera visual guidance and local environment reconstruction techniques, this method optimizes stacking posture planning by combining multi-camera data acquisition, semantic segmentation, point cloud fusion, and a real-time feedback closed-loop mechanism, through initial pose generation, visual verification, precise positioning, target verification, and replanning.

Benefits of technology

It significantly improves the efficiency and accuracy of stacking planning, shortens planning time to the millisecond level, controls pose error within ±2mm, adapts to complex stacking scenarios, and responds to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120735001B_ABST
    Figure CN120735001B_ABST
Patent Text Reader

Abstract

The application discloses a kind of automatic loading and unloading scene multi-camera visual guidance adaptive stacking posture planning method, it is successively including initial pose generation, including analysis pose parameter, generates preliminary placement pose;Visual verification, including multi-camera data acquisition, target area idle state verification, pose rationality evaluation;Precise positioning, including virtual box construction, space matching optimization, generates final coordinate;Target verification and re-planning, based on real-time feedback closed-loop mechanism, whether target area is idle state is verified again, and when it is verified that no, then the posture planning of re is carried out.The application is shortened to millisecond level from the planning time of second level of traditional method by setting local search and optimization algorithm;Through point cloud completion and space matching optimization, pose error is controlled within ±2mm, and through visual verification, environmental occlusion and stacking change can be coped with;The method can support complex stacking scene, and is widely applicable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot automation technology, specifically to a multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios. Background Technology

[0002] In automated loading and unloading scenarios, robots need to accurately stack goods (such as boxes) in designated locations. Traditional stacking planning methods have the following problems:

[0003] Global search is inefficient: Traditional methods rely on global space search, which is computationally intensive and time-consuming, making it difficult to meet real-time requirements.

[0004] Data gaps lead to errors: Due to camera view limitations or environmental occlusion, 3D cameras and vision sensors may not be able to fully capture the point cloud of the target area, resulting in inaccurate planning results.

[0005] Insufficient adaptability: Existing methods are poorly adapted to complex stacking environments (such as irregular stacking and occlusion), and are prone to placement failure due to missing local data.

[0006] Lack of dynamic verification mechanism: The planning process lacks real-time visual verification of the target area, making it impossible to dynamically adjust the pose to cope with environmental changes.

[0007] This invention solves the above problems by introducing multi-camera visual guidance and local environment reconstruction technology, and significantly improves the efficiency and accuracy of stacking planning. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides a multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios, thereby solving the problems mentioned in the background section.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] In a first aspect, embodiments of the present invention provide a multi-camera visual-guided adaptive stacking posture planning method for automatic loading and unloading scenarios, which includes four specific steps: initial pose generation, visual verification, precise positioning, target verification and replanning.

[0011] Initial pose generation includes receiving pose commands from the robot control system, parsing pose parameters, and generating an initial placement pose.

[0012] Visual verification includes multi-camera data acquisition, target area idle state verification, and pose rationality assessment.

[0013] Precise positioning includes virtual box construction, spatial matching optimization, and generation of final coordinates;

[0014] Target verification and replanning involves the robot's robotic arm moving to the target pose and, based on a real-time feedback closed-loop mechanism, re-verifying whether the target area is in an idle state. If the verification fails, replanning the pose is performed.

[0015] To further optimize this technical solution, in the initial pose generation step:

[0016] The pose parameters in the pose command include the target point coordinates and the box size parameters;

[0017] Preprocess the pose parameters, including coordinate transformation and Z-axis calculation.

[0018] The generated initial placement pose includes the placement position and orientation of the box.

[0019] To further optimize this technical solution, in the visual verification step:

[0020] Multi-camera data acquisition: Simultaneously acquires images and point cloud data of the target area using multiple 3D cameras and RGB cameras;

[0021] Target area idle status verification: Use semantic segmentation algorithm to identify target areas with existing boxes in the image, combine point cloud data to calculate the actual position and size of existing boxes, and verify whether the target area is idle;

[0022] The pose rationality assessment involves reconstructing a 3D spatial model of the target area based on local point clouds and detecting the minimum safe distance around the target point.

[0023] To further optimize this technical solution, the 3D spatial model is reconstructed using local environment reconstruction technology, which includes image segmentation and point cloud fusion, and local point cloud completion.

[0024] To further optimize this technical solution, in the local environment reconstruction technology:

[0025] Image segmentation and point cloud fusion: The point cloud data processing is guided by the RGB image segmentation results to fill in missing areas;

[0026] Local point cloud completion: Predicting the point cloud distribution of occluded areas in the target region based on generative adversarial networks (GANs).

[0027] To further optimize this technical solution, in the step of precise positioning:

[0028] Virtual box construction: Based on the dimensions of the box to be placed, a virtual box is generated to simulate its placement state;

[0029] Spatial matching optimization uses the iterative nearest point algorithm to match the virtual box with the 3D spatial model, optimizes the placement position, and adjusts the box pose through gradient descent to maximize space utilization and meet safety constraints.

[0030] Generate final coordinates: Output the optimized target point coordinates and yaw angle, and send them to the robot controller for execution.

[0031] To further optimize this technical solution, the real-time feedback closed-loop mechanism in the target verification and replanning steps includes:

[0032] The robot's robotic arm moves to the target pose and scans the environment again;

[0033] If the environment remains unchanged, i.e. the target area is verified to be in an idle state, then the placement action is executed;

[0034] If the environment changes, such as the box placed in the target area being moved or a new obstacle being detected, the two steps of visual verification and precise positioning are retried.

[0035] To further optimize this technical solution, the 3D camera used in this method is installed on the top and sides of the robot, with a field of view covering the loading and unloading area to which the target area belongs. The resolution of the 3D camera is 0.5mm and the frame rate is 30Hz. The RGB camera is coaxially installed with the 3D camera and has a resolution of 1920×1080.

[0036] To further optimize this technical solution, the robot used in this method is a six-axis robotic arm, and the end effector of the robot is equipped with a suction cup or gripper.

[0037] To further optimize this technical solution, in the data preprocessing:

[0038] Coordinate transformation converts the target point coordinates from the world coordinate system to the robot's base coordinate system;

[0039] Z-axis calculation processing: The height of the Z-axis is calculated as the current stacking layer height plus the height of the box.

[0040] In a second aspect, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, they implement the steps of a multi-camera vision-guided adaptive stacking posture planning method for an automatic loading and unloading scenario as described in the first aspect of the present invention.

[0041] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, they implement the steps of a multi-camera vision-guided adaptive stacking posture planning method for an automatic loading and unloading scenario as described in the first aspect of the present invention.

[0042] Compared with existing technologies, this invention provides a multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios, which has the following beneficial effects:

[0043] This multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios reduces the planning time from seconds to milliseconds by setting local search and optimization algorithms. Through point cloud completion and spatial matching optimization, the pose error is controlled within ±2mm. Visual verification shows that it can cope with environmental occlusion and stacking changes. This method can support complex stacking scenarios (such as tilted stacking and multi-layer gaps) and has wide applicability. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating a multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios proposed in this invention.

[0046] Figure 2 This is a schematic diagram of the initial pose generation process in a multi-camera vision-guided adaptive stacking posture planning method for automatic loading and unloading scenarios proposed in this invention.

[0047] Figure 3 This is a flowchart illustrating the visual verification process in a multi-camera visual-guided adaptive stacking posture planning method for automated loading and unloading scenarios proposed in this invention.

[0048] Figure 4 This is a flowchart illustrating the precise positioning process in a multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios proposed in this invention.

[0049] Figure 5 This is a schematic diagram of the target verification and replanning process in the multi-camera vision-guided adaptive stacking posture planning method for automatic loading and unloading scenarios proposed in this invention. Detailed Implementation

[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0051] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0052] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.

[0053] Example 1:

[0054] Reference Figures 1-5 This is the first embodiment of the present invention, which provides a multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios. The method uses 3D cameras mounted on the top and sides of the robot, with a field of view covering the loading and unloading area of ​​the target region. The 3D camera has a resolution of 0.5mm and a frame rate of 30Hz. An RGB camera is coaxially mounted with the 3D camera, with a resolution of 1920×1080, used for image segmentation and texture analysis. The robot used is a six-axis robotic arm with a repeatability accuracy of ±0.1mm. The robot's end effector is equipped with a suction cup or gripper and integrates a real-time controller (such as ROS2+MoveIt). The method includes four specific steps: initial pose generation, visual verification, precise positioning, target verification, and replanning.

[0055] S1. Initial pose generation, including receiving pose instructions from the robot control system, parsing pose parameters, and generating a preliminary placement pose.

[0056] In the initial pose generation step:

[0057] The pose parameters in the pose command include the target point coordinates and the box size parameters.

[0058] The pose parameters are preprocessed, including coordinate transformation and Z-axis calculation.

[0059] The generated initial placement pose includes the placement position and orientation of the box.

[0060] S2, Visual verification, including multi-camera data acquisition, target area idle state verification, and pose rationality assessment.

[0061] In the steps of the visual verification:

[0062] Multi-camera data acquisition: Simultaneously acquires images and point cloud data of the target area using multiple 3D cameras and RGB cameras.

[0063] To verify the idle status of the target region, a semantic segmentation algorithm (such as Mask R-CNN) is used to identify the target region in the image where there are already boxes. Combined with point cloud data, the actual position and size of the existing boxes are calculated to verify whether the target region is idle.

[0064] The pose rationality assessment involves reconstructing a 3D spatial model of the target area based on the local point cloud, and detecting the minimum safe distance around the target point (e.g., box gap ≥ 50mm) to ensure there is no risk of collision.

[0065] The visual verification process can handle environmental occlusion and stacking changes, increasing the planning success rate to 99%.

[0066] S3. Precise positioning, including virtual box construction, spatial matching optimization, and generation of final coordinates.

[0067] In the precise positioning step:

[0068] Virtual container construction: Based on the dimensions of the container to be placed, a virtual container is generated to simulate its placement state.

[0069] Spatial matching optimization uses the Iterative Closest Point (ICP) algorithm to match the virtual box with the 3D spatial model, optimize the placement position, and adjust the box pose through gradient descent to maximize space utilization and meet safety constraints.

[0070] Generate final coordinates: Output the optimized target point coordinates and yaw angle, and send them to the robot controller for execution.

[0071] In this embodiment, the 3D spatial model is reconstructed using local environment reconstruction technology, which includes image segmentation and point cloud fusion, and local point cloud completion.

[0072] Among them, the local environment reconstruction technology includes:

[0073] Image segmentation and point cloud fusion: The point cloud data processing is guided by the RGB image segmentation results (such as box edges) to fill in missing areas.

[0074] Local point cloud completion: Based on generative adversarial networks (GANs), predict the point cloud distribution of occluded areas in the target region to improve data integrity.

[0075] In this embodiment, the code for image segmentation and point cloud fusion is as follows:

[0076] # Image segmentation using Mask R-CNN

[0077] from detectron2 import model_zoo

[0078] from detectron2.engine import DefaultPredictor

[0079] cfg = model_zoo.get_config("COCO-InstanceSegmentation / mask_rcnn_R_50_FPN_3x.yaml")

[0080] predictor = DefaultPredictor(cfg)

[0081] outputs = predictor(image)

[0082] masks = outputs["instances"].pred_masks.cpu().numpy()

[0083] # Point cloud and segmentation mask fusion

[0084] for mask in masks:

[0085] point_cloud_segment = extract_point_cloud_by_mask(depth_map,mask)

[0086] reconstructed_cloud += point_cloud_segment

[0087] In this embodiment, the pseudocode for the Iterative Closest Point (ICP) algorithm is as follows:

[0088] def optimize_placement(virtual_box, local_point_cloud):

[0089] # ICP coarse matching

[0090] icp_transform = ICP(local_point_cloud, virtual_box)

[0091] coarse_pose = apply_transform(virtual_box, icp_transform)

[0092] # Gradient Descent Fine-Tuning

[0093] for _ in range(max_iterations):

[0094] error = calculate_placement_error(coarse_pose, local_point_cloud)

[0095] gradient = compute_gradient(error)

[0096] coarse_pose -= learning_rate * gradient

[0097] if error < threshold:

[0098] break

[0099] return coarse_pose

[0100] This method uses point cloud completion and ICP matching to control the pose error within ±2mm, which is 60% better than the traditional method.

[0101] S4. Target verification and replanning: Before the robot's robotic arm moves to the target pose, it re-verifies whether the target area is in an idle state based on a real-time feedback closed-loop mechanism. If the verification fails, the pose is replanned.

[0102] The real-time feedback closed-loop mechanism in the target verification and replanning steps includes:

[0103] The robot's robotic arm moves to the target pose and scans the environment again.

[0104] If the environment remains unchanged, i.e., the target area is verified to be in an idle state, then the placement action is executed.

[0105] If the environment changes, such as the box placed in the target area being moved or a new obstacle being detected, the two steps of visual verification and precise positioning are retried.

[0106] Example 2:

[0107] Based on the method described in Embodiment 1, the detailed steps of the implementation process are as follows:

[0108] Step 1: Initial pose generation.

[0109] Receive control command: {"target_x": 1000, "target_y": 500, "box_size": [600,400, 300]} (unit: mm).

[0110] Coordinate transformation: Transform the target point from the world coordinate system to the robot's base coordinate system.

[0111] Preprocessing: Generate initial pose based on box dimensions (Z-axis height = current stacking layer height + box height).

[0112] Step 2: Visual verification.

[0113] Data acquisition: A 3D camera scans the target area to obtain point clouds; an RGB camera captures images.

[0114] Idle detection:

[0115] Image segmentation identifies existing boxes and generates a mask.

[0116] Point cloud cropping: Only retain data within a 500mm radius around the target point.

[0117] If an obstacle is detected occupying the target point, a replanning process is triggered.

[0118] Step 3: Precise positioning.

[0119] Virtual box construction: Generate virtual box point cloud based on box_size.

[0120] ICP matching: Align the virtual box with the local point cloud and calculate the initial transformation matrix.

[0121] Gradient descent optimization: fine-tune the pose to ensure that the distance between the boxes is ≥50mm and the center of gravity is stable.

[0122] Output: The final pose parameters are sent to the robotic arm controller.

[0123] Step 4: Dynamic verification and replanning.

[0124] Before the robotic arm moves to the target pose, it scans the environment again:

[0125] If the environment remains unchanged, perform the placement action.

[0126] If a new obstacle is detected, steps 2-3 will be triggered again.

[0127] Example 3:

[0128] This embodiment also provides a computer device applicable to a multi-camera visual guidance adaptive stacking posture planning method for an automated loading and unloading scenario, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the multi-camera visual guidance adaptive stacking posture planning method for an automated loading and unloading scenario as proposed in the above embodiment.

[0129] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements a multi-camera vision-guided adaptive stacking posture planning method for automatic loading and unloading scenarios as proposed in the above embodiments.

[0130] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0131] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0132] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0133] More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0134] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0135] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios, characterized in that, The process includes four steps: initial pose generation, visual verification, precise localization, target verification, and replanning. Initial pose generation includes receiving pose commands from the robot control system, parsing pose parameters, and generating an initial placement pose. Visual verification includes multi-camera data acquisition, target area idle state verification, and pose rationality assessment. Precise positioning includes virtual box construction, spatial matching optimization, and generation of final coordinates; Target verification and replanning involves the robot's robotic arm moving to the target pose and, based on a real-time feedback closed-loop mechanism, re-verifying whether the target area is in an idle state. If the verification fails, replanning the pose is performed.

2. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 1, characterized in that, In the initial pose generation step: The pose parameters in the pose command include the target point coordinates and the box size parameters; Preprocess the pose parameters, including coordinate transformation and Z-axis calculation. The generated initial placement pose includes the placement position and orientation of the box.

3. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 1, characterized in that, In the steps of the visual verification: Multi-camera data acquisition: Simultaneously acquires images and point cloud data of the target area using multiple 3D cameras and RGB cameras; Target area idle status verification: Use semantic segmentation algorithm to identify target areas with existing boxes in the image, combine point cloud data to calculate the actual position and size of the existing boxes, and verify whether the target area is idle. The pose rationality assessment involves reconstructing a 3D spatial model of the target area based on local point clouds and detecting the minimum safe distance around the target point.

4. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 3, characterized in that, The 3D spatial model is reconstructed using local environment reconstruction technology, which includes image segmentation and point cloud fusion, and local point cloud completion.

5. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 4, characterized in that, In the local environment reconstruction technology: Image segmentation and point cloud fusion: The point cloud data processing is guided by the RGB image segmentation results to fill in missing areas; Local point cloud completion: Predicting the point cloud distribution of occluded areas in the target region based on generative adversarial networks (GANs).

6. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 1, characterized in that, In the precise positioning step: Virtual box construction: Based on the dimensions of the box to be placed, a virtual box is generated to simulate its placement state; Spatial matching optimization uses the iterative nearest point algorithm to match the virtual box with the 3D spatial model, optimizes the placement position, and adjusts the box pose through gradient descent to maximize space utilization and meet safety constraints. Generate final coordinates: Output the optimized target point coordinates and yaw angle, and send them to the robot controller for execution.

7. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 1, characterized in that, The real-time feedback closed-loop mechanism in the target verification and replanning steps includes: The robot's robotic arm moves to the target pose and scans the environment again; If the environment remains unchanged, i.e. the target area is verified to be in an idle state, then the placement action is executed; If the environment changes, such as the box placed in the target area being moved or a new obstacle being detected, the two steps of visual verification and precise positioning are retried.

8. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 1, characterized in that, The method uses a 3D camera mounted on the top and sides of the robot, with a field of view covering the loading and unloading area of ​​the target area. The 3D camera has a resolution of 0.5mm and a frame rate of 30Hz. An RGB camera is mounted coaxially with the 3D camera and has a resolution of 1920×1080.

9. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 1, characterized in that, The robot used in this method is a six-axis robotic arm, and the end effector of the robot is equipped with a suction cup or gripper.

10. The multi-camera vision-guided adaptive stacking posture planning method for automated loading and unloading scenarios according to claim 2, characterized in that, In the data preprocessing: Coordinate transformation converts the target point coordinates from the world coordinate system to the robot's base coordinate system; Z-axis calculation processing: The height of the Z-axis is calculated as the current stacking layer height plus the height of the box.

Citation Information

Patent Citations

  • Depth 6D pose estimation network model and workpiece pose estimation method

    CN114299150A

  • Online boxing method, and terminal and storage medium

    WO2022011979A1