Quadruped robot autonomous navigation and mechanical arm collaborative operation method

By employing deep collaborative control between a quadruped robot and a robotic arm, combined with YOLOv8 target detection and polarization imaging technology, the problems of mobility and reflective target recognition for mobile robots in unstructured environments have been solved, enabling efficient and safe operation in scenarios such as chemical plants and medical sorting facilities.

CN121821384APending Publication Date: 2026-04-10CHINA MACHINERY DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MACHINERY DIGITAL TECHNOLOGY CO LTD
Filing Date
2026-02-03
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing mobile robots suffer from poor maneuverability in unstructured environments, low accuracy in recognizing reflective or dynamic targets, and insufficient safety redundancy in hazardous scenarios, resulting in limited operating range, low success rate, and high system survival risk.

Method used

By employing deep collaborative control of a quadruped robot and a robotic arm, combined with the YOLOv8 target detection model, polarization imaging technology, and multi-sensor fusion, the system achieves accurate identification and pose estimation of reflective targets. Furthermore, the stability and safety of the system are enhanced through leg-arm collaborative stability control and the Intelligent Safety Decision Module (ISS 2.0).

Benefits of technology

It has achieved a leap in all-round operation capabilities in complex and dangerous environments, improved the operational accuracy of highly reflective and dynamic targets, enhanced the system's autonomous risk avoidance capability, expanded the stable working domain, and reduced the risk of secondary accidents in dangerous scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121821384A_ABST
    Figure CN121821384A_ABST
Patent Text Reader

Abstract

The invention relates to a quadruped robot autonomous navigation and mechanical arm collaborative operation method, which comprises the following steps that 1, a quadruped robot control system is designed, and after a quadruped robot is powered, real-time pictures acquired by a color camera at the front end of the robot are transmitted to a video picture display area of the quadruped robot control system through a UDP protocol of socket communication; 2, after the quadruped robot enters the working area, detecting the current environment through a YOLOv8 target detection model, and if no working target exists, controlling the quadruped robot to continuously turn left until the working target is detected; 3, the target detection picture is 640 * 480 pixels, and the mechanical arm is controlled to complete grabbing operation through a UDP protocol of socket communication. According to the target recognition operation method for cooperation of the quadruped robot and the mechanical arm, through deep fusion and collaborative innovation of the four dimensions of perception, decision making, execution and safety, the target recognition efficiency is improved, and the target recognition efficiency is improved. And the systematic breakthrough of the operation capability in a complex and dangerous environment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for autonomous navigation of a quadruped robot and collaborative work of a mechanical arm, and belongs to the technical field of robots. BACKGROUND

[0002] Currently, in the field of dangerous environment operation (such as toxic gas leakage in chemical plant, material sorting in medical pollution area), it mainly relies on wheeled mobile robot carrying mechanical arm or remote manual control device to execute task, the closest prior art is the following two schemes:

[0003] 1. Description of prior art scheme

[0004] Scheme A (wheeled mobile operation robot): AGV chassis is adopted to carry 6-DOF mechanical arm, indoor navigation is realized through laser SLAM, and target object is recognized in combination with OpenCV vision library. A typical representative is KUKA KMR iiwa system, and its working process is: mapping and positioning → mechanical arm preset path grabbing → wheeled platform moving to next station.

[0005] Scheme B (remote control humanoid robot): such as Boston Dynamics Atlas biped robot, the operator remotely controls the arm of the robot through VR device to execute valve rotation action, and relies on high-definition camera to return real-time picture.

[0006] 2. Defects of prior art

[0007] (1) Poor terrain adaptability leads to limited operation range

[0008] The wheeled platform (scheme A) cannot cross stairs, pipeline ruins or oil-polluted ground, and more than 40% of the valves in the chemical plant leakage accident are located in non-flat areas, so that the robot cannot reach the target position;

[0009] The biped robot (scheme B) has terrain adaptability, but the dynamic balance algorithm has huge calculation amount (more than 200W GPU is needed to solve in real time), and cannot recover automatically after falling down, so the failure rate is > 25%;

[0010] (2) Insufficient vision-action coordination accuracy

[0011] The recognition failure rate of traditional vision recognition (such as OpenCV + SIFT feature matching) on reflective surface (medicine bottle glass) or low light environment (toxic gas leakage area) is as high as 35%-40%, which leads to a deviation of > 3cm in the calculation of the mechanical arm grabbing position;

[0012] The movement of the mechanical arm depends on the preset trajectory, and cannot respond to the dynamic target displacement (such as medicine bottle on the conveyor belt) in real time, and the grabbing delay is > 500ms;

[0013] (3) Lack of safety redundancy mechanism

[0014] The existing system lacks an emergency decision-making module for multi-sensor fusion: the platform stability is insufficient when the mechanical arm is operating, and external force disturbance (such as airflow impact in a chemical plant) can easily cause rollover; there is no linkage response to dangerous substance (toxic gas / radiation) concentration, and manual intervention delay increases the risk of secondary accidents by 60%. The remote control scheme (Scheme B) has a communication delay problem (average > 300ms), which may miss the best operation window in an explosion critical scene.

[0015] 3. Core problems caused by defects,

[0016] The above defects result in three key bottlenecks in the existing technology:

[0017] ① Limited operation scene: only works in structured environment (flat ground, fixed lighting), cannot cover real dangerous scenes such as chemical plant and post-disaster ruins;

[0018] ② Low operation success rate: the failure rate of reflective target grabbing exceeds the industry safety threshold (medical sorting requires > 95% success rate), and insufficient valve control precision causes leakage aggravation;

[0019] ③ High system survival risk: no self-protection ability in emergency scenes, causing equipment damage and secondary disasters.

[0020] The present application aims to solve the three core problems of "poor passability in unstructured environment", "low operation precision for high-reflective / dynamic targets" and "lack of safety redundancy in dangerous scenes" existing in the existing mobile operation robot, and improve the reliability in medical sorting, chemical plant rescue and other scenes through deep collaborative control of quadruped robot and mechanical arm. SUMMARY

[0021] The present application is exactly aimed at the technical problems existing in the prior art, and provides a method for autonomous navigation of a quadruped robot and collaborative operation of a mechanical arm. The technical scheme provides a target recognition and operation system and control method for collaborative operation of a quadruped robot and a mechanical arm, to solve the technical problems of poor passability of the existing mobile operation robot in unstructured terrain, low operation precision for reflective or dynamic targets, and insufficient safety redundancy in dangerous environments, so as to realize reliable and autonomous execution of tasks in complex dangerous scenes (such as medical material sorting and chemical plant emergency disposal).

[0022] In order to achieve the above purpose, the technical scheme of the present application is as follows: a method for autonomous navigation of a quadruped robot and collaborative operation of a mechanical arm, comprising the following steps:

[0023] Step 1, design a four-legged robot control system, after the four-legged robot is powered on, the real-time picture obtained by the front-end color camera of the robot is transmitted to the video picture display area of the four-legged robot control system through the UDP protocol of socket communication;

[0024] Step 2, when the four-legged robot enters the working area, the current environment is detected by the YOLOv8 target detection model, and if there is no work target, the four-legged robot is continuously turned left until the work target is detected;

[0025] Step 3, the target detection picture is 640*480 pixels, in the horizontal X-axis direction, when the left upper corner horizontal coordinate of the work target detection rectangle is not in the interval of 220 to 420 pixels, the four-legged robot is continuously adjusted to the effective work interval through ROS2 communication control, and the four-legged robot is moved to the front of the work target point, and the mechanical arm is controlled to complete the grabbing work through the UDP protocol of socket communication.

[0026] In step 2, the YOLOv8 target detection model detection process is as follows,

[0027] When using a depth camera for target detection, the processing flow of YOLOv8 is as follows: the depth camera first captures a three-channel color image in RGB format (resolution of 640x480 pixels), which is used as the original input of the detection system. In the preprocessing stage, the image will be normalized in size, scaled by maintaining the aspect ratio, and padded on the edges to adjust to the standard input size of YOLOv8. At the same time, color value normalization (mapping pixel values from 0-255 to 0-1 range) and channel order conversion (from HWC format to CHW format) are completed;

[0028] Subsequently, the preprocessed image is sent to the YOLOv8 network for inference: the backbone network first extracts primary features through the basic convolution module, then realizes cross-stage feature fusion to enhance small target recognition capability through the C2f module, and finally captures multi-scale context information through the SPPF spatial pyramid pooling module; the neck network fuses the feature maps from the three scales of the backbone network, enlarges the high-level semantic features through upsampling operation, and then concatenates and fuses with the bottom layer position features to construct a feature pyramid containing rich spatial and semantic information; the detection head adopts decoupling design, the classification branch outputs the class probability distribution of each anchor point, and the regression branch predicts the accurate coordinates of the bounding box (including the center point position x, y and width w and height h), and optimizes the positioning accuracy of the bounding box through the DFL distribution focal loss module,

[0029] After model inference, the post-processing stage is entered: first, the confidence threshold (default 0.25) is applied to filter out low-confidence prediction boxes, and then non-maximum suppression is used to eliminate overlapping detection boxes.

[0030] The final output detection results contain structured information: each target corresponds to a detection box, and the bounding box position is given in the form of pixel coordinates (usually represented as [x_min, y_min, x_max, y_max] or [x_center, y_center, width, height]), while the target category (such as "bottle", "cutton" and other labels) and detection confidence (a probability value of 0-1) are labeled.

[0031] In step 3, after powering on, the robotic arm is first calibrated at its zero point to ensure the subsequent processes can proceed normally. After zero-point calibration, it enters the initial controllable state and receives the start command from the quadruped robot. The robotic arm moves along a predetermined path. There are six target poses for this movement. Upon reaching each pose, the robotic arm pauses briefly, at which point the target recognition task is initiated. After receiving depth camera information, the target is identified and matched using a recognition model. If a matching model exists in the camera information, the target recognition result is first transformed to the robotic arm's coordinate system to obtain the target pose quadruple for subsequent tasks. Then, it is determined which type of object it belongs to. This project trained two types of objects for recognition, corresponding to two different types of robotic arm motion tasks: The first type of object is a bottle. After recognizing this object, the robotic arm performs a grasping task, first moving to the target using the airbot.move_to_cart_pose() function. The first type of object to be identified is a button. When this type of object is identified, the robot first moves to the target pose using the `airbot.move_eef_pos()` function, then grips the object using the `airbot.move_eef_pos()` function. The robot then uses the `airbot.move_to_cart_pose()` function to control the end joint of the robotic arm to rotate the object by a predetermined angle using the `airbot.move_to_joint_pos()` function. Finally, the robot releases the object. After this task is completed, the robot continues to process other objects. If all identified objects have been processed, the robot moves to the next predetermined pose until the last pose is reached and the identification, gripping, or rotation task is completed. If all poses have been traversed, the robot returns to its origin and stops moving, indicating the task is complete. The power can then be turned off and the process terminated.

[0032] The above scheme constructs a precise identification and pose estimation method for reflective targets based on the fusion of polarization imaging and a deep learning visual model (YOLOv8 / v9). Existing technologies typically use ordinary RGB cameras combined with traditional image processing algorithms (such as SIFT, SURF) or unoptimized deep learning models for target recognition. In environments such as chemical plants and laboratories, the surfaces of target objects (such as glass medicine bottles and metal valves) often produce strong reflections, leading to failure in feature point extraction or false detections by the model, resulting in extremely low recognition rates.

[0033] The above-mentioned solution proposes and implements a dynamic stability control mechanism for active "leg-arm" coordination. Currently, the mobile platform (whether wheeled or legged) and robotic arm of mobile manipulators are typically controlled as two independent systems, or only have simple static coordination. When the robotic arm performs large-scale, high-speed, or heavy-load operations, the resulting reaction forces and torques cause the center of gravity of the mobile platform to shift. This can affect operational accuracy or even cause the entire system to become unstable and overturn, especially in systems with fewer support points like quadruped platforms.

[0034] I. Composition of the Control System

[0035] The control system comprises hardware and software components. The hardware provides the physical carrier and signal interaction basis for system operation, while the software provides logic control and algorithm support. Together, they realize the full-process operation function.

[0036] The hardware components encompass core computing hardware, a perception hardware cluster, mobile execution hardware, task execution hardware, communication hardware, safety monitoring hardware, and display hardware. The core computing hardware utilizes an NVIDIA development board, serving as the core of the system's data processing and logical decision-making, handling YOLOv8 model inference, navigation algorithm execution, collaborative control logic operations, and multi-sensor data fusion processing. The perception hardware cluster includes a robot front-end color camera, a depth camera (RGB-D camera), and radar (LiDAR). The color camera captures real-time images of the work scene, the depth camera acquires RGB images and depth information to support target pose calculation, and the radar is used for environmental scanning and mapping, as well as static and dynamic obstacle detection. The mobile execution hardware consists of a quadruped robot body and a platform drive unit, receiving navigation commands to complete turning, movement, and posture adjustment actions, adapting to unstructured terrain operation requirements. The task execution hardware comprises a robotic arm and a matching gripper, used to perform target grasping, rotation, and other tasks, and is equipped with a zero-point calibration mechanism to ensure operational accuracy. The communication hardware includes a Socket / UDP communication module and a ROS2 communication module, respectively enabling data transmission and command issuance in different scenarios. The safety monitoring hardware includes toxic gas / Radiation sensors, platform attitude sensors, and robotic arm force sensors are used to collect environmental hazard parameters and equipment operating status parameters in real time; the display hardware is the video display area of ​​the control system, used to receive and visualize the real-time operation images transmitted by the color camera.

[0037] The software comprises a communication protocol stack module, a vision processing module, a navigation control module, a collaborative control module, a robotic arm control module, a safety decision module (ISS 2.0), and a display control module. The communication protocol stack module integrates UDP and ROS2 protocol processing functions. UDP is used for real-time image transmission and robotic arm control command issuance, while ROS2 is used for quadruped robot orientation adjustment and movement control. The vision processing module includes a YOLOv8 target detection model, a polarization imaging fusion algorithm, and a coordinate transformation module. The YOLOv8 model is used for target detection, the polarization imaging fusion algorithm is used to suppress glare interference to improve target recognition rate, and the coordinate transformation module converts target pixel coordinates to the robotic arm coordinate system. The navigation control module integrates a Google Cartographer mapping algorithm, an end-to-end navigation algorithm, and a command safety correction module. The mapping algorithm constructs a 2D raster map and a 3D point cloud semantic map of the environment, the end-to-end navigation algorithm plans the optimal operation path, and the command safety correction module avoids static and dynamic obstacles. The collaborative control module includes a leg-arm collaborative stability management unit and a robotic arm trajectory planning unit. The robotic arm coordination stability management unit generates joint torque compensation commands for the quadruped platform by real-time calculation of the robotic arm's kinematics and dynamics model to counteract the interference of the robotic arm's operation on the platform's center of gravity. The robotic arm trajectory planning unit controls the robotic arm to move to six target poses along a preset path. The robotic arm control module includes a zero-point calibration program and motion control functions (airbot.move_to_cart_pose(), airbot.move_eef_pos(), airbot.move_to_joint_pos(), etc.). The zero-point calibration program initializes the robotic arm's working state, and the motion control functions execute grasping and rotating actions for two types of targets: bottles and buttons. The safety decision module (ISS 2.0) monitors environmental parameters and equipment status in real time through multi-sensor fusion evaluation algorithms and emergency response algorithms. When parameters exceed limits, an emergency response is triggered. The display control module includes a real-time image receiving and rendering program, which receives camera images transmitted via UDP protocol and outputs them to the display area.

[0038] II. Control Relationships of Each Functional Module

[0039] The control system's functional modules follow a closed-loop control logic of "sensing-transmission-processing-decision-execution-feedback," forming a hierarchical control relationship according to the work process, and clearly defining control priorities and protocol division of labor, as detailed below:

[0040] Initialization and Data Acquisition Phase: After the quadruped robot is powered on, the display control module triggers the image acquisition command, which is sent to the color camera via the UDP communication module. After the color camera acquires the real-time image, it is transmitted to the display control module via the Socket / UDP protocol. The display control module completes the image rendering and outputs it to the video display area. At the same time, the robotic arm control module automatically starts the zero-point calibration program. After the calibration is completed, the robotic arm enters the initial controllable state. The safety decision module (ISS 2.0) simultaneously starts the data acquisition work of the toxic gas / radiation sensor, platform attitude sensor, and robotic arm force sensor, and monitors the environment and equipment status throughout the process.

[0041] Target detection and navigation phase: After the quadruped robot enters the work area, the navigation control module collects environmental data through radar, calls the Google Cartographer mapping algorithm to build an environmental map, and then plans the optimal path to the target area through an end-to-end navigation algorithm. Subsequently, it sends movement commands to the platform drive unit to drive the quadruped robot to move along the planned path. When the quadruped robot approaches the work area, the vision processing module loads the YOLOv8 target detection model, receives RGB images and depth information transmitted from the depth camera, and performs target detection on the work environment. If the horizontal coordinate of the upper left corner of the detected target rectangle is not in the range of 220 to 420 pixels, the vision processing module feeds back the target coordinate information to the navigation control module. The navigation control module sends direction adjustment commands to the platform drive unit through the ROS2 communication module, controlling the quadruped robot to continuously adjust its direction until the target enters the effective work area. Then, the navigation control module continues to send movement commands to drive the quadruped robot to move to the target point.

[0042] Collaborative Operation and Execution Phase: Before the quadruped robot reaches the target point, the collaborative control module receives the target pose quadruple set (after coordinate transformation) output by the vision processing module and the platform posture data transmitted by the safety decision module. The leg-arm collaborative stability management unit calculates the torque of the robotic arm's operation on the quadruped platform's center of gravity, generates a torque compensation command, and sends it to the platform drive unit to achieve quadruped platform posture compensation. Simultaneously, the collaborative control module sends a start command to the robotic arm control module. The robotic arm control module calls the trajectory planning unit to control the robotic arm to move along a preset path to six target poses. During a brief pause at each pose, a target recognition task is initiated. If the vision processing module recognizes a bottle-like target, the robotic arm control module sequentially calls the `airbot.move_to_cart_pose()` function to control the robotic arm to move to the target pose and the `airbot.move_eef_pos()` function to control the gripper's opening and closing, grasping the target and placing it in a predetermined position. If a button-like target is recognized, the robotic arm control module first uses `airbot.move_to_cart_pose()` to... The `airbot.move_eef_pos()` function moves the robot to the target pose, the `airbot.move_eef_pos()` function clamps the target, and the `airbot.move_to_joint_pos()` function controls the end joint of the robot arm to rotate the target by a predetermined angle. Finally, the target is released. After the target processing of a single pose is completed, the robot arm control module sends a completion signal to the collaborative control module. The collaborative control module controls the robot arm to move to the next predetermined pose until all poses have been traversed. Then, the robot arm returns to the origin and stops moving.

[0043] Safety monitoring and emergency control phase: Throughout the entire operation process, the safety decision module (ISS 2.0) continuously integrates and analyzes environmental parameters and equipment status parameters. When safety risks such as excessive toxic gas concentration, excessive platform tilt angle, or abnormal joint torque of the robotic arm are detected, the safety decision module (ISS 2.0) immediately interrupts the current instructions of the navigation control module and the robotic arm control module, prioritizes calling the navigation control module to plan the evacuation path, and simultaneously issues an emergency evacuation command to the platform drive unit to control the quadruped robot to quickly evacuate to a safe area. The response time from risk detection to the start of evacuation does not exceed 100ms. In the control relationship of the control system, the UDP protocol is specifically used to meet the low latency requirements of real-time image transmission and robotic arm motion control, while the ROS2 protocol is specifically used to meet the high-precision coordination requirements of quadruped robot orientation adjustment and movement control. The control priority follows the rule of Safety Decision Module (ISS 2.0) > Collaborative Control Module > Navigation Control Module, Vision Processing Module, and Robotic Arm Control Module, ensuring priority for emergency response in hazardous scenarios. The outputs of all modules serve as inputs for relevant subsequent modules, forming a closed-loop control without data gaps. Dynamic coordination between the robotic arm and the quadruped platform is achieved through the leg-arm collaborative stability management unit, ensuring operational stability and accuracy in unstructured environments. The above scheme designs an intelligent safety system (ISS 2.0) driven by multi-sensor fusion and possessing autonomous decision-making capabilities. Existing safety strategies mostly rely on remote operator monitoring and emergency remote control, or simple single-threshold trigger emergency stop. This approach suffers from communication delays, personnel reaction time delays, and biased decision-making, and cannot guarantee the safety of the system and environment in hazardous scenarios where every second counts.

[0044] Compared with existing technologies, this invention has the following advantages: the target recognition operating system and method for quadruped robots and robotic arms provided by this invention achieves a systematic breakthrough in operational capabilities in complex and dangerous environments through deep integration and collaborative innovation in four dimensions: perception, decision-making, execution, and safety. Its beneficial effects are mainly reflected in the following three aspects:

[0045] First, it achieves a comprehensive leap in operational capabilities from "structured" to "unstructured" environments. This invention is not simply a combination of a mobile platform and a robotic arm; rather, it fundamentally solves the terrain adaptation problem for mobile robots by organically combining the high mobility of a quadruped robot with active leg-arm cooperative stability control. Specifically, the quadruped platform endows the system with the ability to climb stairs, traverse obstacles (passage rate >95%), and withstand external impacts (such as airflow in chemical plants), enabling it to reach treacherous locations inaccessible to wheeled or tracked platforms. More importantly, when the robotic arm performs operations, the cooperative stability management unit can calculate dynamic changes in real time and drive the quadruped joints for torque compensation, controlling the platform's posture deviation during operation to within <2°. This effectively avoids the predicament of "being able to reach the location but unable to perform the task," expanding the robot's stable working range by approximately 70%, laying a solid foundation for performing tasks in real unstructured environments.

[0046] Secondly, it achieves precise and robust operation on highly reflective and dynamic targets, solving the final centimeter challenge of visual servoing. The core advantage of this invention lies in integrating cutting-edge visual perception algorithms with a high-speed real-time control closed loop. Addressing the pain point of low recognition rates (<65%) for reflective surfaces (such as medicine bottles and metal valves) in existing technologies, the system innovatively combines polarization imaging technology with the YOLOv8 key point detection model. By analyzing the polarization characteristics of light, specular reflection interference is effectively suppressed, significantly improving the recognition rate of reflective targets to over 95%. Based on this, the visual servo control unit can calculate the 6D pose of targets with accuracy better than ±1cm and ±1.5°. Simultaneously, for dynamic targets, the trajectory prediction algorithm controls the total delay from recognition to the start of grasping to within 120ms, more than 75% faster than traditional solutions. This achieves a high success rate (>98%) for grasping dynamic targets such as moving medicine bottles on conveyor belts, meeting the actual accuracy and cycle time requirements of industrial sorting.

[0047] Third, it constructs an integrated proactive safety paradigm of "perception-decision-execution," greatly enhancing the robot's survivability in hazardous scenarios. This invention transcends the passive safety model that relies solely on remote emergency stops. Through a multi-sensor fusion-based Intelligent Safety Decision Unit (ISS 2.0), it empowers the system with the ability to autonomously assess risks and respond rapidly. The system continuously monitors environmental parameters (such as toxic gas concentration), and if limits are exceeded, it can interrupt the mission within 100ms and autonomously plan an evacuation route, controlling the quadruped platform to quickly retreat to a safe zone. Compared to purely manual remote control operation (response delay >300ms), this system reduces the time personnel and equipment are exposed to danger by more than 60%, achieves a 100% mission success rate for evacuation, effectively avoids secondary accidents, and provides crucial safety assurance for unmanned operations in extremely hazardous environments such as nuclear, biological, and chemical (NBC) environments.

[0048] In summary, the beneficial effects of this invention are not merely improvements to a single module, but rather, through systematic integration and innovation, it achieves a qualitative leap in three core dimensions: operational scope, operational accuracy, and safety redundancy. Its overall performance far surpasses existing technologies, providing a comprehensive and reliable technical solution for replacing manual labor in hazardous environments. These effects have been verified through joint simulation and physical platform testing. Attached Figure Description

[0049] Figure 1 This is a roadmap for collaborative work technology between quadruped robots and robotic arms.

[0050] Figure 2 Here is a flowchart of the YOLOv8 object detection method.

[0051] Figure 3 This is a roadmap for end-to-end navigation technology for quadruped robots.

[0052] Figure 4 This is a flowchart of the robotic arm's workflow.

[0053] Figure 5 System workflow diagram

[0054] Figure 6 System architecture diagram. Detailed Implementation

[0055] To enhance understanding of the present invention, the embodiments will be described in detail below with reference to the accompanying drawings.

[0056] Example 1: See Figure 1 , reference Figure 1 This invention provides a technology for collaborative operation between a quadruped robot and a robotic arm. A quadruped robot control system is designed. After the quadruped robot is powered on, the real-time image captured by the color camera at the robot's front end is transmitted to the video display area of ​​the quadruped robot control system via UDP socket communication. When the quadruped robot enters the working area, it detects the current environment using a YOLOv8 target detection model. If no target is found, the quadruped robot is controlled to continuously turn left until a target is detected. The target detection image is 640*480 pixels. In the horizontal X-axis direction, if the horizontal coordinate of the upper left corner of the target detection rectangle is not within the range of 220 to 420 pixels, the quadruped robot is controlled via ROS2 communication to continuously adjust its direction to within the effective working area and move to the target point. The robotic arm is controlled via UDP socket communication to complete the grasping operation.

[0057] Reference Figure 2The YOLOv8 object detection model is used to detect the current environment. Specifically, when using a depth camera for object detection, the YOLOv8 processing flow is as follows: The depth camera first captures a three-channel color image in RGB format (resolution 640×480 pixels), which serves as the raw input to the detection system. In the preprocessing stage, the image undergoes size normalization, adjusting it to the standard input size of YOLOv8 by maintaining aspect ratio through proportional scaling and edge padding. Simultaneously, color value normalization (mapping pixel values ​​from 0-255 to the 0-1 range) and channel order conversion (from HWC format to CHW format) are performed. Subsequently, the preprocessed image is fed into the YOLOv8 network for inference: the backbone network first extracts primary features through basic convolutional modules, then uses the C2f module to achieve cross-stage feature fusion to enhance the recognition capability of small targets, and finally captures multi-scale contextual information through the SPPF spatial pyramid pooling module; the neck network fuses feature maps from the three scales of the backbone network, amplifies high-level semantic features through upsampling operations, and then concatenates and fuses them with the low-level positional features to construct a feature pyramid containing rich spatial and semantic information; the detection head adopts a decoupled design, with the classification branch outputting the class probability distribution of each anchor point, and the regression branch predicting the precise coordinates of the bounding box (including the center point position x, y, as well as the width w and height h), and optimizing the bounding box localization accuracy through the DFL distributed focus loss module.

[0058] After model inference is complete, the post-processing stage begins: First, a confidence threshold (default 0.25) is applied to filter out low-confidence predicted boxes. Then, non-maximum suppression is used to eliminate overlapping detection boxes. The final output detection results contain structured information: each target corresponds to a detection box, with the bounding box position given in pixel coordinates (usually represented as [x_min, y_min, x_max, y_max] or [x_center, y_center, width, height]), along with the target category (such as "bottle", "cutton", etc.) and detection confidence (a probability value between 0 and 1).

[0059] Reference Figure 3 This invention provides an end-to-end navigation technology for quadruped robots. The input is data obtained from the robot's sensors, and the output is a safety-corrected command (vx, vy). The robot then reaches a target point via ROS2 communication. Environmental information is divided into static obstacles and dynamic obstacles, which require dilation processing and geometric approximation, respectively. Command safety correction aims to prevent collisions caused by the robot executing inferred commands.

[0060] Reference Figure 4After powering on, the robotic arm is first calibrated at its zero point to ensure the smooth operation of subsequent processes. After zero-point calibration, it enters an initial controllable state. At this point, it receives the start command from the quadruped robot and moves along a predetermined path. There are six target poses for this movement. Upon reaching each pose, the robotic arm pauses briefly, at which point the target recognition task begins. After receiving depth camera information, the target is identified and matched using a different model. If a matching model exists in the camera information, the target recognition result is first transformed to the robotic arm's coordinate system to obtain the target pose quadruple for subsequent tasks. Then, it is determined which type of object it belongs to. This project trained two types of objects for recognition, corresponding to two different types of robotic arm motion tasks: The first type of object is a bottle. Upon recognition of this object, the robotic arm performs a grasping task. First, it moves to the target pose using the `airbot.move_to_cart_pose()` function, then controls the opening and closing of the gripper using the `airbot.move_eef_pos()` function to grasp the object and place it in a predetermined position. The second type of object is a button. When this type of object is recognized, it similarly moves to the target pose using the `airbot.move_to_cart_pose()` function, grasps the object using the `airbot.move_eef_pos()` function, then controls the end joint of the robotic arm to rotate the object by a predetermined angle using the `airbot.move_to_joint_pos()` function, and finally releases the object. After completing this task, it continues to process other objects. If all recognized items have been processed, the robotic arm continues to move to the next predetermined pose until it reaches the last pose and completes the recognition, grasping, or rotation task. Once all poses have been traversed, control the robotic arm to return to the origin and stop moving. This completes the task, and the power can be turned off to end the process.

[0061] This invention integrates a polarization imaging lens at the end effector of a robotic arm to actively acquire image sequences of a target at different polarization angles. Subsequently, the multi-polarization angle images and RGB images are input together into a YOLOv8 model trained on a special dataset (containing a large number of reflective target samples) for fusion processing. Polarization information effectively suppresses specular highlights, highlighting the material and texture details of the object's surface. This fusion scheme fundamentally solves the industry problem of high reflectivity interference, significantly improving the recognition rate of reflective targets from approximately 65% ​​in existing technologies to over 95%, and enabling subsequent 6D pose estimation accuracy to achieve translation <1cm and rotation <1.5°. This provides unprecedentedly reliable visual guidance for the precise operation of the robotic arm, an unexpected effect that cannot be achieved with a single visual modality.

[0062] This invention incorporates a collaborative stability management unit within the central collaborative controller. This unit, acting as a core middleware, calculates the kinematics and dynamics model of the robotic arm in real time, accurately predicts the disturbance torque its motion will generate on the center of gravity of the quadruped platform, and generates feedforward compensation commands. These commands, fused with the platform's own balance control commands, drive the hip and shoulder joint motors of the quadruped platform in real time to collaboratively output torque, actively "counteracting" the effects of the robotic arm's motion. This proactive compensation, with its "predictive" nature, ensures that the platform's attitude deviation angle remains stable within <2° during high-speed robotic arm operation, expanding the stable working domain by approximately 70%. This achieves true "operation in motion," resolving the long-standing core contradiction in the field of mobile robotics where "mobility" and "operation" mutually constrain each other, representing a significant technological advancement.

[0063] This invention elevates safety from a mere "function" to a "system." ISS 2.0 deeply integrates multi-source information, including toxic gas / radiation sensors, platform status sensors, and robotic arm force sensors. It no longer passively awaits instructions but actively and continuously assesses the environmental hazard level and its own status. When any parameter (such as toxic gas concentration, platform tilt, or robotic arm joint torque) exceeds a safety threshold, the system can autonomously make optimal decisions within 100ms (such as immediately terminating the task, releasing the load, or planning an evacuation route) and control the quadruped platform to execute an emergency evacuation. This improves hazard response time from the "seconds" dependent on manual intervention to the "milliseconds," and increases the successful evacuation rate from less than 80% to nearly 100%, significantly enhancing the survival rate of robots operating in extremely dangerous environments and preventing major property losses and secondary disasters. This shift from "passive protection" to "active survival" brings an unexpected leap in safety performance.

[0064] Example 2: Refer to Figure 5 The system workflow diagram shown is as follows: Figure 6 The system structure diagram shown in this embodiment details the specific implementation of the system in performing the task of sorting hazardous chemical reagents.

[0065] S101: System Initialization and Environment Mapping

[0066] (1) Activate the navigation and perception unit 05 (LiDAR) located on the quadruped robot 01 to scan the high-risk environment.

[0067] (2) The navigation planning unit of the central collaborative controller adopts the Google Cartographer algorithm to process the lidar point cloud data in real time, construct a two-dimensional grid map and a three-dimensional point cloud semantic map, and mark key areas such as shelves, aisles, and charging piles.

[0068] S102: Task assignment and target identification

[0069] (1) The operator issues sorting tasks through the host computer software.

[0070] (2) The navigation planning unit plans the optimal path to the target shelf. The platform drive unit of the quadruped robot 01 drives the platform body 01 to move along the path in a trot gait.

[0071] (3) When approaching the target shelf, the arm-mounted sensing unit 04 (RGB-D camera) of the robotic arm operation module 03 begins to acquire images. The visual recognition unit loads a pre-trained YOLOv8 model (input resolution 640x640) to perform real-time inference on the images, identify the target, and output its pixel coordinates and confidence score.

[0072] S103: Precise Positioning and Platform Stance

[0073] (1) The visual servo control unit identifies and matches the target based on the pixel coordinates of the target and the depth information provided by the RGB-D camera using a target recognition model. If a matching model exists in the camera information, the target recognition result is first transformed to the coordinate system of the robotic arm to obtain the target pose quadruple for subsequent tasks, and then it is determined which category the target belongs to. This project trained two categories of objects, corresponding to two different types of robotic arm motion tasks: grasping bottles and pressing buttons.

[0074] (2) The navigation planning unit controls the quadruped robot 01 to fine-tune its final position so that the target medicine bottle is located in the optimal working space of the robotic arm body 03, and the platform body maintains a safe distance of about 50cm from the shelf.

[0075] S104: Robotic arm grasping execution

[0076] (1) The robotic arm trajectory planning unit, based on the target pose and the current environmental point cloud, completes different robotic arm motion tasks corresponding to two different types of objects. The first type of object is a bottle. The robotic arm performs a grasping task. First, it moves to the target pose using the airbot.move_to_cart_pose() function, and then controls the opening and closing of the gripper using the airbot.move_eef_pos() function. After grasping the object, it places it in a predetermined position. The second type of object is a button. When this type of object is detected, it also moves to the target pose using the airbot.move_to_cart_pose() function, and then grips the object using the airbot.move_eef_pos() function. Then, it controls the end joint of the robotic arm to rotate the object by a predetermined angle using the airbot.move_to_joint_pos() function, and finally releases the object.

[0077] (2) After the task is completed, control the robotic arm to return to the recognition position and continue to recognize and process other objects. If there are no more unprocessed objects in the field of view of the depth camera, control the robotic arm to continue to move to the next predetermined pose until the last pose is reached and the recognition, grasping or rotation task is completed.

[0078] S105: Placement and Quest Completion

[0079] (1) After successful grabbing, the robotic arm will transfer the high-risk reagent to the designated collection basket.

[0080] (2) The system will prompt the task to be completed via voice or status light and wait for the next instruction.

[0081] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention. Equivalent transformations or substitutions made based on the above technical solutions all fall within the scope of protection of the claims of the present invention.

Claims

1. A method for autonomous navigation of a quadruped robot and collaborative operation with a robotic arm, characterized in that, The method includes the following steps: Step 1: Design a quadruped robot control system. After the quadruped robot is powered on, the real-time image captured by the color camera at the front of the robot is transmitted to the video display area of ​​the quadruped robot control system via UDP protocol of socket communication. Step 2: After the quadruped robot enters the work area, the YOLOv8 target detection model is used to detect the current environment. If there is no target, the quadruped robot is controlled to turn left continuously until the target is detected. Step 3: The target detection screen is 640*480 pixels. In the horizontal X-axis direction, when the horizontal coordinate of the upper left corner of the target detection rectangle is not in the range of 220 to 420 pixels, the quadruped robot is continuously adjusted to the effective working range through ROS2 communication and controlled to move to the front of the target point. The robotic arm is then controlled to complete the grasping operation through the UDP protocol of socket communication.

2. The method for autonomous navigation of a quadruped robot and collaborative operation of a robotic arm according to claim 1, characterized in that, The detection process of the YOLOv8 object detection model in step 2 is as follows: When using a depth camera for object detection, the YOLOv8 processing flow is as follows: The depth camera first captures a three-channel color image in RGB format (resolution of 640×480 pixels). This image serves as the raw input to the detection system. In the preprocessing stage, the image undergoes size normalization. By maintaining the aspect ratio and scaling proportionally and filling the edges, it is adjusted to the standard input size of YOLOv8. At the same time, color value normalization and channel order conversion are completed. Subsequently, the preprocessed image is fed into the YOLOv8 network for inference: the backbone network first extracts primary features through basic convolutional modules, then uses the C2f module to achieve cross-stage feature fusion to enhance small target recognition capabilities, and finally captures multi-scale contextual information through the SPPF spatial pyramid pooling module; the neck network fuses feature maps from the three scales of the backbone network, amplifies high-level semantic features through upsampling, and then concatenates and fuses them with low-level positional features to construct a feature pyramid containing rich spatial and semantic information; the detection head adopts a decoupled design, with the classification branch outputting the class probability distribution of each anchor point, and the regression branch predicting the precise coordinates of the bounding box (including the center point position x, y, width w, and height h), and optimizing the bounding box localization accuracy through the DFL distributed focus loss module. After the model inference is completed, the post-processing stage begins: first, a confidence threshold is applied to filter out low-confidence prediction boxes, and then non-maximum suppression is used to eliminate overlapping detection boxes; The final output detection results contain structured information: each target corresponds to a detection box, the bounding box position is given in pixel coordinates, and the target category is labeled.

3. The method for autonomous navigation of a quadruped robot and collaborative operation of a robotic arm according to claim 2, characterized in that, In step 3, after powering on, the robotic arm is first calibrated at its zero point to ensure the smooth operation of subsequent processes. After zero-point calibration, it enters the initial controllable state. At this time, it receives the start command from the quadruped robot, and the robotic arm moves along a predetermined path. There are a total of 6 target poses for this movement. When reaching each pose, the robotic arm will pause briefly, at which point the target recognition task is initiated. After receiving depth camera information, the target is identified and matched using a recognition model. If a matching model exists in the camera information, the target recognition result is first transformed to the coordinate system of the robotic arm to obtain the target pose quadruple for subsequent tasks. Then, it is determined which type of object it belongs to. This project trained two types of recognition objects, corresponding to two different types of robotic arm motion tasks: The first type of recognition object is a bottle. After recognizing this object, the robotic arm performs a grasping task, first moving to the target position using the airbot.move_to_cart_pose() function. The first type of object to be identified is a button. When this type of object is identified, the robot first moves to the target pose using the `airbot.move_to_cart_pose()` function, then grips the object using the `airbot.move_eef_pos()` function, and then rotates the object by a predetermined angle using the `airbot.move_to_joint_pos()` function. Finally, the object is released. After the task is completed, other objects are processed. If all identified objects have been processed, the robot moves to the next predetermined pose until the last pose is reached and the identification, gripping, or rotation task is completed. If all poses have been traversed, the robot returns to the origin and stops moving, thus completing the task. The power is then turned off and the process ends.