A target navigation and grasping method and system for mobile robot piece flow sorting

CN122500707APending Publication Date: 2026-08-04SHANBEN (SHANGHAI) AUTOMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANBEN (SHANGHAI) AUTOMATION TECH CO LTD
Filing Date
2026-05-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

现有的视觉检测算法在边缘计算设备上往往面临推理速度慢、实时性不足的问题,而传统的跟踪算法在目标被遮挡或快速移动时容易丢失目标,缺乏自主恢复机制

Benefits of technology

[0017]1. Fully autonomous operation: Adopting a composite structure of "wheeled mobile platform + robotic arm", it integrates SLAM navigation, target tracking and automatic grasping functions, eliminating the dependence on fixed routes and manual loading and unloading, and significantly reducing the labor cost and operational risks of logistics sorting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122500707A_ABST
    Figure CN122500707A_ABST
Patent Text Reader

Abstract

The application discloses a target navigation and grabbing method and system for mobile robot logistics sorting, and relates to the technical fields of mobile robots and intelligent logistics. The method comprises the following steps: collecting logistics storage environment data through a multi-modal sensor, constructing a high-precision grid map by using a laser SLAM algorithm, and realizing autonomous path planning and dynamic obstacle avoidance of a mobile robot in combination with a NAV2 navigation system; adopting a YOLOv12-KCF dynamic fusion algorithm to realize real-time detection and tracking of a logistics sorting target, using TensorRT technology to realize lightweight deployment of a deep learning model, and realizing high-frame-rate inference on an edge computing device; after the robot navigates to a target area, recognizing a target pose by using RGB-to-HSV color space processing and a contour extraction algorithm, and combining inverse kinematics calculation to control a mechanical arm to realize precise grabbing; and realizing collaborative control of chassis movement and mechanical arm operation by using a double-ROS master control system and a cascade PID control algorithm. The application solves the problems of traditional logistics sorting AGVs, such as dependence on fixed routes, inability to autonomously grab and poor adaptability to dynamic environments, and significantly improves the automation degree, operation efficiency and system robustness of logistics sorting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mobile robot and intelligent logistics technology, specifically to a target navigation and grasping method and system for mobile robot logistics sorting. Background Technology

[0002] With the rapid development of e-commerce and intelligent manufacturing, the sorting volume in the logistics and warehousing industry is growing exponentially. Traditional logistics sorting mainly relies on manual operation or semi-automated equipment, which suffers from high labor costs, low operational efficiency, high error rates, and high labor intensity. Although Automated Guided Vehicles (AGVs) have been applied in some warehousing scenarios, existing AGV systems typically rely on fixed navigation modes such as magnetic tracks, QR codes, or laser reflectors, resulting in poor path planning flexibility and difficulty in adapting to the frequent adjustments required for warehouse layouts. In addition, traditional AGVs mostly only have handling functions, and manual assistance or fixed robotic arm workstations are still required in the loading and unloading process, making it impossible to independently complete the fully automated operation of "navigation-identification-grabbing-transfer".

[0003] In real-world logistics sorting scenarios, the environment is highly dynamic and unstructured. Warehouses contain numerous moving obstacles (such as workers, forklifts, and other robots), and sorting targets (parcels, boxes, etc.) are often randomly placed, varying in shape, color, and size. This poses significant challenges to robots' environmental perception, dynamic obstacle avoidance, target tracking, and precise grasping capabilities. Existing visual detection algorithms often suffer from slow inference speeds and insufficient real-time performance on edge computing devices, while traditional tracking algorithms are prone to losing targets when they are occluded or moving rapidly, lacking an autonomous recovery mechanism. Furthermore, high-quality labeled datasets relevant to logistics scenarios are relatively scarce, limiting the generalization ability and practical application of intelligent vision algorithms in real-world warehousing environments.

[0004] Therefore, there is an urgent need to develop a mobile robot logistics sorting method and system with high autonomy, strong robustness and real-time processing capabilities to solve the problems of efficient navigation, stable tracking and accurate grasping in complex dynamic environments, and promote the intelligent upgrading of logistics sorting operations. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a target navigation and grasping method and system for mobile robot logistics sorting. Through a composite robot architecture, multimodal perception fusion, lightweight deep learning model, and efficient collaborative control strategy, the entire logistics sorting process is automated and intelligent.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A target navigation and grasping method for logistics sorting using mobile robots includes the following steps:

[0008] Multimodal data collection is performed on the logistics and warehousing environment, an environmental map is constructed using the laser SLAM algorithm, and the optimal navigation path is generated based on the NAV2 navigation system to control the mobile robot chassis to navigate autonomously and dynamically avoid obstacles.

[0009] During navigation, the environment is perceived by depth camera and lidar, and the YOLOv12-KCF dynamic fusion algorithm is used to detect and track logistics sorting targets in real time, and control the mobile robot to adjust its posture to keep the target in the center of the field of vision and approach the target.

[0010] Once the mobile robot reaches the preset grasping area, it switches to the robotic arm operation mode, uses the robotic arm vision recognition algorithm to obtain the precise pose information of the target, and generates control commands for each joint of the robotic arm through inverse kinematics calculation.

[0011] Control the robotic arm to perform the grasping action, and after the grasping is completed, send a status signal back to the mobile chassis, and resume navigation mode to transport the sorted target to the designated destination.

[0012] Preferably, the YOLOv12-KCF dynamic fusion algorithm introduces an attention mechanism to improve detection accuracy and uses TensorRT technology for model lightweighting; a detection-triggered tracking recovery mechanism is designed to automatically call YOLOv12 for relocalization when KCF tracking is lost, thus solving the tracking failure problem caused by occlusion.

[0013] Preferably, the robotic arm visual recognition algorithm uses RGB to HSV color space processing, combined with morphological filtering and contour extraction technology, to stably recognize targets under complex lighting conditions; through inverse kinematics calculation and hand-eye calibration, it achieves accurate mapping from image coordinates to robotic arm joint angles.

[0014] Preferably, the method uses a cascaded PID control algorithm to realize the robot's motion control. The outer loop positional PID calculates the target speed based on visual deviation, and the inner loop incremental PID outputs the motor PWM signal based on the speed error. In the approach phase, seamless switching between depth camera and lidar ranging is achieved.

[0015] This invention also provides a target navigation and grasping system for mobile robot logistics sorting, including a mobile chassis system, a robotic arm subsystem, and a perception and control unit. The system adopts a dual ROS main control architecture and achieves efficient collaboration between the chassis and the robotic arm through a serial communication protocol, forming a closed loop of "perception-decision-execution".

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] 1. Fully autonomous operation: Adopting a composite structure of "wheeled mobile platform + robotic arm", it integrates SLAM navigation, target tracking and automatic grasping functions, eliminating the dependence on fixed routes and manual loading and unloading, and significantly reducing the labor cost and operational risks of logistics sorting.

[0018] 2. Highly Robust Target Tracking: The YOLOv12-KCF dynamic fusion algorithm is proposed, which combines the advantages of high-precision detection of deep learning with high-frame-rate tracking of correlation filtering. Through the detection-triggered recovery mechanism, the target loss problem caused by package stacking and personnel occlusion in logistics scenarios is effectively solved, and the stability of the system in dynamic environments is improved.

[0019] 3. Real-time inference at the edge: TensorRT is used to deploy the YOLOv12 model in a lightweight manner. Through layer fusion, operator optimization and accuracy calibration, the inference speed is doubled on resource-constrained edge computing devices, which meets the stringent real-time requirements of logistics sorting.

[0020] 4. Precise Collaborative Control: The design of dual ROS master control communication protocol and cascade PID control strategy realizes precise collaboration between the mobile chassis and the robotic arm; combined with RGB-HSV visual recognition and inverse kinematics calculation, it ensures stable grasping of sorting targets in different poses.

[0021] 5. Strong scenario adaptability: A dedicated dataset for logistics sorting scenarios has been constructed, covering multiple categories of logistics items, which enhances the generalization ability and adaptability of the algorithm model in real warehousing environments. Attached Figure Description

[0022] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0023] Figure 1 This is a general design block diagram of the system of the present invention;

[0024] Figure 2 This is a general hardware block diagram of the present invention;

[0025] Figure 3 This is the overall flowchart of the software of this invention;

[0026] Figure 4 This is a block diagram of the YOLOv12-KCF fusion tracking algorithm of the present invention;

[0027] Figure 5 This is a schematic diagram of the target tracking state machine of the present invention;

[0028] Figure 6 This is a block diagram of the cascaded PID control algorithm of the present invention;

[0029] Figure 7 This is a comparison chart of the lightweight inference speed of the model of this invention. Detailed Implementation

[0030] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way.

[0031] Example 1: System Hardware Architecture and Dual ROS Communication Protocol

[0032] like Figure 1 and Figure 2 As shown, the system adopts a hierarchical heterogeneous control architecture, consisting of a mobile chassis system, a robotic arm subsystem, and a sensing and computing unit.

[0033] Computing platform configuration:

[0034] The first main control module (mobile chassis) adopts the NVIDIA Jetson Orin Nano edge computing module, is equipped with 4GB LPDDR5 memory, has an AI computing power of 20 TOPS, runs the ROS2 (Humble) operating system, and is responsible for SLAM mapping, NAV2 navigation, YOLOv12-KCF fusion tracking and chassis motion decision-making.

[0035] The second main control module (robotic arm): adopts an NVIDIA Jetson Nano module, is equipped with 4GB of memory, has a computing power of 0.5 TOPS, runs the ROS1 (Noetic) operating system, and is responsible for the robotic arm's visual recognition, inverse kinematics calculation and grasping control.

[0036] The underlying controllers are STM32F407VET6 (Cortex-M4, 168MHz) and STM32F103C8T6 (Cortex-M3, 72MHz) microcontrollers, respectively. They drive the actuators through PWM signals and connect to the MPU6050 inertial measurement unit via the IIC interface to provide real-time IMU data feedback.

[0037] Dual ROS cross-platform serial communication protocol:

[0038] To achieve efficient collaboration between the ROS2 and ROS1 master controllers, a full-duplex communication link based on a three-wire serial port (TX / RX / GND) was designed, with a baud rate configured at 115200 bps, 8 data bits, 1 stop bit, and no parity check. A general data frame format was defined as shown in Table 1 to ensure the reliability and anti-interference capability of command transmission.

[0039] Table 1 Serial Communication Data Frame Format Table

[0040] Fields Frame header aisle Data length Data content Validation Additional verification length 2 Byte 1 Byte 1 Byte n Byte 1 Byte 1 Byte illustrate 0xAA, 0xFF Equipment identification Data length Commands / Status Cumulative verification Cumulative verification

[0041] Verification algorithm implementation:

[0042] Checksum: Starting from frame header 0xAA to the end of the Data area, each byte is accumulated, and the lower 8 bits are taken as the checksum value.

[0043] Additional checksum: During the calculation and checksum process, each byte addition operation is synchronously added to the AC register, and the lower 8 bits are used as the additional checksum value. This dual checksum mechanism effectively reduces the communication error rate and ensures the accurate transmission of capture instructions and status feedback.

[0044] Example 2: YOLOv12 Object Detection and Area Attention Mechanism

[0045] The system uses the YOLOv12 algorithm as the detector for logistics sorting targets. Its network structure consists of three parts: Backbone, Neck, and Head. To improve the model's feature extraction capability in complex warehousing environments, YOLOv12 introduces the AreaAttention mechanism.

[0046] Area Attention principle:

[0047] Traditional Self-Attention calculates global relevance, with a computational complexity of O(n log n). Area Attention divides the feature map into n regions along the horizontal or vertical direction (n=4 in this example), reducing computational overhead while maintaining a large receptive field. The attention weights are calculated as follows:

[0048]

[0049] in, These are query, key, and value matrices, respectively. The key vector dimension is denoted as . Area Attention generates an "excitation-inhibition" weight matrix by calculating the correlation between elements within the region, enabling the model to focus on key target features such as packages and express boxes, and suppressing background noise interference.

[0050] Model training parameter configuration:

[0051] Based on the constructed logistics sorting scenario dataset (including categories such as parcels, express boxes, pallets, and irregularly shaped items), the model was trained in an Ubuntu 22.04 environment using dual NVIDIA RTX 3080 (10GB) GPUs for acceleration. Specific training parameters are shown in Table 2.

[0052] Table 2 YOLOv12 Model Training Parameter Configuration Table

[0053] Parameters Setting value illustrate operating system Ubuntu 22.04 Deep learning development environment CUDA version 12.4 GPU acceleration library version Epochs 1000 Training rounds Batch Size 64 Batch size Image Size 640×640 Input image resolution Data Augmentation Mosaic, Mixup, Flip Improve model generalization ability

[0054] Data augmentation strategies include scaling, random cropping, vertical and horizontal flipping, and mosaic enhancement, which effectively improve the model's robustness in recognizing logistics items of different sizes, angles, and under occlusion conditions.

[0055] Example 3: Lightweighting and Deployment of TensorRT Models

[0056] To meet the real-time requirements of edge devices, a lightweight deployment of the YOLOv12 model was implemented using the NVIDIA TensorRT inference engine. Specific optimization steps included:

[0057] Model conversion: Convert the .pt weight files trained by PyTorch into TensorRT engine files (.engine) to adapt to the GPU architecture of Jetson Orin Nano.

[0058] Layer Fusion: Enables layer fusion technology, which combines convolutional layers (Conv), batch normalization layers (BN), and activation layers (ReLU / SiLU) into a single computing node, significantly reducing memory access frequency and kernel startup overhead.

[0059] Precision Calibration: Select FP32 precision mode for inference. While maintaining almost lossless detection accuracy (mAP@0.5 > 90%), maximize inference speed through operator optimization and automatic kernel tuning.

[0060] Dynamic Tensor Memory Management: Enables dynamic memory allocation strategies, reuses intermediate tensor memory, reduces video memory usage, and improves GPU utilization.

[0061] Test results show that the lightweight TensorRT model improves the inference frame rate on Jetson Orin Nano from 9 FPS of the PyTorch model to nearly 20 FPS, with an inference speed improvement of about 100%, meeting the 30ms real-time response requirements of logistics sorting scenarios.

[0062] Example 4: YOLOv12-KCF Dynamic Fusion Tracking Algorithm

[0063] To address the limitations of single algorithms in dynamic environments, a YOLOv12-KCF dynamic fusion tracking algorithm was designed, the process of which is as follows: Figure 4 As shown.

[0064] KCF tracing principle:

[0065] The KCF algorithm utilizes correlation filters for fast target localization in the frequency domain. (Filter) The learning objective is to minimize the regression error:

[0066]

[0067] in, For cyclically shifted samples of the target image, For the expected response, is the regularization coefficient. By introducing a Gaussian kernel function to map features to a high-dimensional space and utilizing FFT for acceleration, KCF can achieve high frame rate tracking.

[0068] Fusion strategy and state machine:

[0069] Initialization phase: The YOLOv12 detector performs a global search on the video stream, and outputs bounding boxes after identifying the logistics targets. And initialize the KCF tracker.

[0070] Tracking phase: The system switches to KCF mode and outputs the target position based on the correlation filter. and confidence score .

[0071] Re-detection mechanism: Real-time monitoring ,when When the threshold is set to 0.6 or the target is occluded, causing tracking drift, the state machine automatically switches back to the "no target detected" state, reactivates YOLOv12 for global search and target relocation, and updates the KCF state. This mechanism achieves millisecond-level recapture after tracking interruption, effectively dealing with occlusion interference such as personnel movement and goods stacking in warehouse scenarios.

[0072] Example 5: Cascade PID Control Algorithm and Parameter Tuning

[0073] The robot motion control employs a cascaded PID algorithm, with the outer loop being a positional PID and the inner loop an incremental PID. The control block diagram is shown below. Figure 6 As shown.

[0074] Outer loop positional PID:

[0075] The input is the pixel coordinate deviation of the target center point. and distance deviation The output is the target linear velocity. With angular velocity The control law is as follows:

[0076]

[0077] in, For the current deviation, This represents the deviation from the previous time step. The outer loop PID parameters, after tuning, are shown in Table 3.

[0078] Table 3 Outer Loop Position PID Parameter Table

[0079] parameter Linear speed control Angular velocity control illustrate 3.0 0.5 Proportional coefficient, response speed 0.0 0.0 Integral coefficient, not used 1.0 2.0 Differential coefficients, suppressing overshoot

[0080] Inner-loop incremental PID:

[0081] The input is the deviation between the target speed and the encoder feedback speed, and the output is the motor PWM increment. Since the system only requires PI control, the incremental PID controller simplifies to:

[0082]

[0083] The inner loop PID parameters are shown in Table 4.

[0084] Table 4 Inner Loop Incremental PID Parameter Table

[0085] parameter Setting value illustrate 300 proportionality coefficient 300 Integral coefficient 0 Differential coefficients, not used

[0086] Ranging switching strategy:

[0087] During the approach to the target, distance is estimated based on the depth camera. .when When the range is less than the effective range (e.g., 0.6m) or the depth data is unavailable, the system automatically switches to lidar ranging data. Continue calculating the forward linear velocity to ensure the robot's smooth approach and precise stopping during the close-range phase.

[0088] Example 6: Robotic Arm Visual Recognition and Inverse Kinematics Grasping

[0089] RGB-HSV color space processing:

[0090] The robotic arm camera acquires RGB images of the target area and converts them to the HSV color space. Taking advantage of the H channel's insensitivity to lighting changes, a binary mask is generated based on preset color thresholds for logistics packages (e.g., yellow express bags, brown cardboard boxes). Noise is eliminated through morphological opening and closing operations, the contours of the largest connected components are extracted, and the centroid pixel coordinates are calculated. And the smallest bounding rectangle.

[0091] Inverse kinematics solution:

[0092] Combining depth information Solve the three-dimensional coordinates of the target in the camera coordinate system Hand-eye calibration matrix Transform the coordinates to the robot arm base coordinate system:

[0093]

[0094] Calculate the target angles of each joint of the robotic arm using inverse kinematics algorithms. Plan a smooth motion trajectory, avoid singularities and self-collisions, and control the end effector gripper of the robotic arm to perform precise grasping.

[0095] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A target navigation and grasping method for logistics sorting using mobile robots, characterized in that, The method includes the following steps: Multimodal data collection is performed on the logistics and warehousing environment, an environmental map is constructed using the laser SLAM algorithm, and the optimal navigation path is generated based on the NAV2 navigation system to control the mobile robot chassis to navigate autonomously and dynamically avoid obstacles. During navigation, the environment is perceived by depth camera and lidar, and the YOLOv12-KCF dynamic fusion algorithm is used to detect and track logistics sorting targets in real time, and control the mobile robot to adjust its posture to keep the target in the center of the field of vision and approach the target. Once the mobile robot reaches the preset grasping area, it switches to the robotic arm operation mode, uses the robotic arm vision recognition algorithm to obtain the precise pose information of the target, and generates control commands for each joint of the robotic arm through inverse kinematics calculation. Control the robotic arm to perform the grasping action, and after the grasping is completed, send a status signal back to the mobile chassis, and resume navigation mode to transport the sorted target to the designated destination; The mobile robot and the robotic arm communicate and synchronize their data through a dual master control communication protocol.

2. The target navigation and grasping method for mobile robot logistics sorting according to claim 1, characterized in that, The YOLOv12-KCF dynamic fusion algorithm specifically includes: The YOLOv12 object detection algorithm is used to perform a global search on the input image to identify the category and bounding box of the logistics sorting target, and the KCF tracker is initialized. During the tracking process, the KCF tracker is used to continuously track the target at a high frame rate based on the correlation filtering principle, and the tracking confidence is calculated in real time. Design a detection-triggered tracking recovery mechanism: When the confidence of the KCF tracker is lower than the preset threshold or the target is occluded and tracking is lost, the YOLOv12 detector is automatically activated to re-perform a global search and target relocation, update the initial state of the KCF tracker, and achieve millisecond-level re-acquisition after tracking interruption. The YOLOv12 model is lightweighted using the TensorRT inference acceleration engine, including operator optimization, layer fusion, and accuracy calibration, and then deployed on edge computing devices.

3. The target navigation and grasping method for mobile robot logistics sorting according to claim 1, characterized in that, The robotic arm visual recognition algorithm specifically includes: The target area is captured as an RGB image, converted to the HSV color space, and a binary mask is generated using preset hue, saturation, and brightness threshold parameters to segment the target object area. Morphological opening and closing operations are performed on the binary mask to eliminate noise interference and extract the contour of the largest connected component. Calculate the centroid pixel coordinates and minimum bounding rectangle of the target contour, and combine the depth information from the depth camera to solve the three-dimensional spatial coordinates of the target in the camera coordinate system. The three-dimensional coordinates in the camera coordinate system are transformed to the robot arm base coordinate system using the hand-eye calibration matrix, which serves as the input for inverse kinematics calculation.

4. The target navigation and grasping method for mobile robot logistics sorting according to claim 1, characterized in that, The control of the mobile robot to adjust its posture to keep the target in the center of its field of vision and to approach the target specifically employs a cascaded PID control algorithm: The outer loop uses a positional PID controller. The inputs are the deviation between the pixel coordinates of the target center point and the image center, as well as the target distance deviation. The outputs are the target linear velocity and angular velocity of the robot. The inner loop uses an incremental PID controller, with the input being the real-time speed deviation between the target speed and the encoder feedback, and the output being the PWM control signal for motor drive; During the approach process, the forward linear velocity is calculated based on the depth estimation information. When the distance is less than the effective range of the depth camera, the system automatically switches to the lidar ranging information to continue calculating the forward linear velocity until the preset grasping distance is reached.

5. The target navigation and grasping method for mobile robot logistics sorting according to claim 1, characterized in that, The dual-master communication protocol adopts serial communication and defines a general data frame format, which includes a frame header, channel identifier, valid data length, data content, checksum, and additional checksums. The data content includes grasping instructions and status feedback signals: when the mobile robot reaches the target position, it sends grasping instruction bytes to the robotic arm master controller; after the robotic arm completes grasping or fails to recognize, it sends status feedback bytes to the mobile robot master controller to realize the state machine jump of the task flow.

6. The target navigation and grasping method for mobile robot logistics sorting according to claim 1, characterized in that, The method also includes the step of constructing a dataset for logistics sorting scenarios: Collect image data of common sorting objects in logistics and warehousing scenarios, including packages of different sizes, express boxes, pallets and irregularly shaped parts; Image data is cleaned, labeled, and structured to construct a dedicated dataset containing multiple categories of logistics items for training and validation of the YOLOv12 model, thereby improving the model's generalization ability and recognition accuracy in complex logistics environments.

7. A target navigation and grasping system for mobile robot logistics sorting, characterized in that, The system includes a mobile chassis system, a robotic arm subsystem, and a sensing and control unit. The mobile chassis system includes a wheeled mobile platform, a low-level motion controller, and a drive motor, used to perform navigation and movement tasks; The robotic arm subsystem includes a multi-degree-of-freedom robotic arm, a servo controller, and an end effector gripper, used to perform target grasping tasks. The perception and control unit includes a lidar, a depth camera, an inertial measurement unit, and a dual ROS main control computing platform; the dual ROS main control computing platform runs on the mobile chassis and the robotic arm respectively, and is connected through a serial communication module. The system is configured to perform the target navigation and grasping method for mobile robot logistics sorting as described in any one of claims 1 to 6.

8. The target navigation and grasping system for mobile robot logistics sorting according to claim 7, characterized in that, The dual ROS main control computing platform includes a first main control module and a second main control module; The first main control module is mounted on the mobile chassis and runs the SLAM mapping algorithm, NAV2 navigation algorithm and YOLOv12-KCF fusion tracking algorithm, and is responsible for environmental perception, path planning and chassis motion decision-making. The second main control module is mounted on the robotic arm and runs the robotic arm visual recognition algorithm and inverse kinematics solution algorithm, which are responsible for the precise positioning of the target and the motion planning of the robotic arm; The first main control module and the second main control module work together through a cross-platform communication interface to form a closed-loop control of "perception-decision-execution".