A minimalist configuration mobile manipulation robot based on pure visual perception

By eliminating depth sensors and LiDAR, integrating an RGB camera and IMU, and employing a 2-DOF robotic arm and a high-performance GPU computing platform, the structural redundancy and insufficient computing power of existing mobile operation robot platforms are resolved, achieving a simplification of robot structure and an improvement in intelligence.

CN224310622UActive Publication Date: 2026-06-02TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
Filing Date
2026-04-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing mobile robot platforms rely on depth sensors and LiDAR, resulting in structural redundancy, dispersed sensor placement, excessive freedom of the robotic arm, and insufficient computing power of the computing platform, making it difficult to support next-generation visual intelligence algorithms.

Method used

It adopts an overall configuration of a wheeled mobile chassis, column, multimodal perception module and high computing power GPU platform, eliminating depth sensor and LiDAR, integrating RGB camera and IMU, the robotic arm has 2 degrees of freedom, and the chassis has built-in GPU computing unit to realize 3D environment reconstruction and intelligent decision-making.

Benefits of technology

It achieves simplified robot structure, reduced weight, and improved mobility stability, supports real-time operation of next-generation vision algorithms, and is suitable for widespread indoor service scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN224310622U_ABST
    Figure CN224310622U_ABST
Patent Text Reader

Abstract

The utility model discloses a kind of minimalist configuration mobile operation robots based on pure vision perception, comprising: wheeled chassis, is equipped with driving wheel and at least two auxiliary universal casters, the driving wheel is driven by motor and is equipped with encoder;Stand, fixedly installed on wheeled chassis, is equipped with lifting platform that can be along its vertical motion;Mechanical arm, installed on lifting platform, at least including horizontal telescopic degree of freedom and end gripper opening and closing degree of freedom;Multi-modal perception module, installed at the top of the stand, located above lifting platform, including a shell and at least one RGB camera and at least one inertial measurement unit integrated in shell, RGB camera and inertial measurement unit are fixed in shell by rigid mounting seat to keep fixed relative pose relationship;Local computing unit, installed inside wheeled chassis, local computing unit has graphics processor GPU computing capacity, and is communicatively connected to encoder, RGB camera and inertial measurement unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This utility model relates to the field of service robot structural design technology, and in particular to a minimalist mobile operation robot based on pure visual perception. Background Technology

[0002] Mobile operating robots are one of the core forms of service robots, requiring them to possess both autonomous movement and navigation capabilities as well as object manipulation abilities. In recent years, the demand for mobile operating robots in application scenarios such as indoor home services has been growing, leading to the emergence of a number of representative commercial and research platforms.

[0003] In terms of hardware architecture, existing products and platforms mainly include:

[0004] (1) Hello Robot Stretch RE1 / RE2 (See: Kemp, CC, et al. "Design of Stretch: A Compact Mobile Manipulation Robot for Indoor Environments." IEEE International Conference on Robotics and Automation (ICRA), 2022.): This platform uses a square differential wheel chassis, a synchronous belt lifting column, and a retractable slender robotic arm. It is equipped with an Intel RealSense D435i RGB-D depth camera and an Intel RealSense T265 tracking camera on top, as well as a LiDAR (RPLiDAR A1) on the chassis for navigation. The computing platform is an Intel NUC mini-PC.

[0005] (2) Fetch Robotics Fetch Mobile Manipulator (see: Wise, M., et al. "Fetch and Freight: Standard Platforms for Service Robot Applications." Workshop on Autonomous Mobile Service Robots, IJCAI, 2016.): This platform adopts a circular differential wheel chassis, a fixed-height torso structure, and is equipped with a 7-DOF industrial-grade robotic arm, a PrimeSense RGB-D depth camera, and a chassis LiDAR. The computing platform is an embedded x86 industrial computer.

[0006] (3) PAL Robotics TIAGo (see: Pages, J., et al. "TIAGo: the modular robot that adapts to different research needs." International Workshop on Robot Modularity, IROS, 2016.): This platform uses a differential chassis, a height-adjustable torso and a 7-DOF robotic arm, equipped with an RGB-D depth camera (Orbbec Astra or Intel RealSense) and chassis LiDAR, and the computing platform is an Intel i7 industrial computer.

[0007] While the aforementioned existing platforms have achieved some success in the field of mobile operation, they still have the following shortcomings in terms of hardware architecture:

[0008] (1) The perception system relies on depth sensors, resulting in structural redundancy and inherent limitations: The aforementioned platforms generally use RGB-D depth cameras as the core perception method. RGB-D cameras are large in size and have high power consumption (usually 3-5W), which increases the top load and center of gravity of the robot. More importantly, consumer-grade RGB-D cameras have inherent depth range limitations (generally an effective distance of 0.3m-4m). When facing scenarios such as specular reflection, transparent objects, and strong light interference, depth data is frequently missing or severely distorted. The hardware defects directly limit the robustness of the upper-layer perception algorithm.

[0009] (2) Low integration of sensing modules and scattered sensor layout: Existing platforms typically install multiple independent sensors scattered throughout the robot (for example, the Stretch RE1 has an RGB-D camera installed in the head, another RGB-D camera installed in the wrist, and a LiDAR installed in the chassis), and there is a lack of unified structured integration design among the sensors. This scattered layout increases the complexity of the overall wiring and the difficulty of calibrating the external parameters between sensors, and is not conducive to maintenance and upgrades.

[0010] (3) The robotic arm has too many degrees of freedom, complex structure and high cost: Although the 7-DOF industrial-grade robotic arms on platforms such as Fetch and TIAGo are highly flexible, they are complex in structure, heavy (usually more than 5kg per arm) and expensive. They are over-designed for common operations such as "picking-carrying-placing" in indoor home environments. In addition, the excessive weight of the robotic arm also raises the center of gravity of the whole machine, affecting the stability of movement.

[0011] (4) Insufficient GPU computing power on the computing platform makes it difficult to support next-generation visual intelligence algorithms: Existing platforms mostly use Intel NUC or general-purpose x86 industrial control computers as computing platforms, lacking high-performance GPUs. However, current advanced environmental perception algorithms (such as real-time 3D reconstruction based on 3D Gaussian sputtering, 3D object detection deep learning networks, etc.) and decision reasoning algorithms (such as large visual language models) all have rigid requirements for GPU parallel computing capabilities. The computing power bottleneck of existing platforms makes it difficult to run these next-generation algorithms locally in real time, limiting the level of robot intelligence.

[0012] The above background information is provided only to aid in understanding the concept and technical solution of this utility model. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above information was disclosed on the filing date of this patent application, the above background information should not be used to evaluate the novelty and inventiveness of this application. Utility Model Content

[0013] The technical problem this invention aims to solve is: how to build a mobile operation robot platform with a minimalist structure, high integration, and powerful local computing power to support advanced pure vision algorithms without relying on depth sensors and LiDAR.

[0014] A minimalist mobile manipulation robot based on pure visual perception, comprising:

[0015] A wheeled mobile chassis is provided with drive wheels and at least two auxiliary swivel casters, wherein the drive wheels are driven by a motor and equipped with an encoder;

[0016] The column is fixedly installed on the wheeled mobile chassis and is equipped with a lifting platform that can move vertically along it;

[0017] A robotic arm is mounted on the lifting platform, and the robotic arm includes at least a horizontal extension degree of freedom and an end effector gripper opening and closing degree of freedom.

[0018] A multimodal sensing module is installed at the top of the column, above the lifting platform. The multimodal sensing module includes a housing and at least one RGB camera and at least one inertial measurement unit integrated in the housing. The RGB camera and the inertial measurement unit are fixed in the housing by a rigid mounting bracket to maintain a fixed relative pose relationship.

[0019] A local computing unit is installed inside the wheeled mobile chassis. The local computing unit has a graphics processing unit (GPU) computing capability and is communicatively connected to the encoder, the RGB camera, and the inertial measurement unit.

[0020] Preferably, the housing of the multimodal sensing module is detachably mounted on the top of the column.

[0021] Preferably, the robotic arm is a 2-degree-of-freedom structure, including an upper arm with the horizontal extension degree of freedom and a gripper with the end gripper opening and closing degree of freedom mounted at the end of the upper arm.

[0022] Preferably, the column is provided with a synchronous belt lifting transmission mechanism, which is driven by a stepper motor; the outside of the column is provided with a linear guide rail, and the lifting platform is slidably mounted on the linear guide rail via a linear slider, and the synchronous belt lifting transmission mechanism is connected to the linear slider.

[0023] Preferably, the optical axis of the RGB camera is set to tilt forward and downward at 15 to 20 degrees.

[0024] Preferably, the wheeled mobile chassis has a battery pack, a local computing unit, and a power management module arranged in layers from bottom to top inside; the local computing unit is fixed to a metal mounting plate inside the wheeled mobile chassis, and the wheeled mobile chassis is also provided with heat dissipation ventilation holes and a cooling fan at the position corresponding to the local computing unit.

[0025] Preferably, the local computing unit is an NVIDIA Jetson AGX Orin computing platform.

[0026] Preferably, the inertial measurement unit is a six-axis inertial measurement unit, comprising a three-axis accelerometer and a three-axis gyroscope.

[0027] Preferably, the mobile operating robot does not include depth sensors and lidar.

[0028] Preferably, the battery pack is a lithium-ion battery pack, and the power management module includes a multi-channel DC-DC converter for providing power outputs of different voltages to the wheeled mobile chassis, the lifting platform, the robotic arm, the local computing unit, and the multimodal sensing module, respectively.

[0029] Compared with existing mobile robot platforms in the background art, the present invention has the following technical effects:

[0030] This invention adopts a column-based integrated configuration consisting of a wheeled mobile chassis, a column, an integrated multimodal sensing module, and a high-performance GPU computing platform built into the chassis. In this configuration, the column simultaneously performs the dual functions of vertical adjustment of the robotic arm and multi-height, multi-view acquisition by the multimodal sensing module. The robotic arm has two degrees of freedom—horizontal extension and a gripper—achieving extreme structural simplification. The multimodal sensing module highly integrates an RGB camera and an IMU in a detachable housing, completely eliminating the need for depth sensors and LiDAR. The integrated GPU computing platform within the chassis provides local computing power support for pure vision-based 3D perception algorithms. Each component has a clear division of labor yet works closely together, achieving a complete functional loop of movement, perception, operation, and intelligent decision-making with the simplest structure. This "subtraction" approach to overall design contrasts sharply with the existing approach of "stacking more sensors and degrees of freedom," and offers the following advantages:

[0031] (1) Eliminating the depth sensor to achieve lightweight perception system: This utility model completely eliminates the RGB-D depth camera and LiDAR. It can achieve 3D environment reconstruction at the algorithm level by simply integrating an RGB camera with a six-axis IMU and cooperating with the encoder on the chassis. This structure makes the total weight of the multimodal perception module on the top less than 150g. Compared with the combination of RGB-D camera (usually about 72g-120g, and requires additional mounting brackets and heat dissipation structures) and LiDAR (usually about 170g-400g) on ​​the existing platform, the total weight and volume of the multimodal perception module are greatly reduced, the center of gravity of the whole machine is lowered, the movement stability is improved, and the depth sensor is fundamentally avoided in the case of mirror reflection, transparent objects and distant scenes.

[0032] (2) The multimodal sensing module is highly integrated and modular: the RGB camera and IMU are integrated in the same compact housing and mounted on the top of the column. This design eliminates the wiring complexity and calibration difficulty caused by the dispersed installation of multiple sensors. Preferably, the multimodal sensing module can be disassembled and upgraded independently, and the maintenance convenience is significantly better than the distributed sensor layout of the existing platform.

[0033] (3) Multi-height perspective acquisition achieved through linkage between the column and the multimodal sensing module: Since the multimodal sensing module is fixed at the top of the column, and the robotic arm moves with the lifting platform, when the lifting platform rises and falls to adjust the operating height of the robotic arm, the height difference between the fixed-position RGB camera and the end of the robotic arm and its operating target changes accordingly. This design allows the robot to obtain multi-view observations of the same operating scene in the vertical dimension (e.g., switching from a high overhead view to a low eye-level view) without the need for an additional camera gimbal. This is particularly beneficial for visual algorithms to more accurately estimate the three-dimensional pose of the target object and the relative relationship between the gripper and the object during grasping operations. Compared with the existing platform, which requires the additional installation of a gimbal, this solution is simpler and more reliable.

[0034] (4) Minimalist robotic arm with lifting column achieves large workspace coverage: The robotic arm of this utility model includes 2 degrees of freedom (horizontal extension + gripper opening and closing), and the overall weight is about 1.5kg, which is much lighter than the existing 6 to 7 degree of freedom industrial-grade robotic arms (usually exceeding 5kg). The robotic arm does not have a wrist lifting joint, and the vertical height adjustment is entirely handled by the lifting platform (about 100cm travel), achieving a large operating space coverage from near the ground to the high cabinet shelves with minimal degrees of freedom. This design, which concentrates the vertical movement function entirely in the column, eliminates the redundant structure of the robotic arm and achieves a good balance between structural simplicity, weight control and workspace.

[0035] (5) High-performance local computing platform supports new generation of intelligent algorithms: This utility model integrates a local computing unit with GPU computing capabilities (such as NVIDIA Jetson AGX Orin (AI computing power 275 TOPS)) in the chassis, which can run new generation algorithms such as 3D reconstruction based on 3D Gaussian sputtering, deep learning 3D object detection, semantic scene graph construction and visual language large model inference in real time locally, without relying on cloud computing, thus taking into account both real-time performance and data privacy.

[0036] In summary, the synergistic effect of the aforementioned technologies enables the robot of this invention to achieve a significant simplification of the overall structure, a substantial reduction in weight, and an improvement in reliability and maintainability without relying on expensive and easily malfunctioning depth sensors. At the same time, the local high-performance computing power ensures the real-time operation of advanced visual AI algorithms, ultimately achieving a high-performance, cost-effective, and highly intelligent mobile operation robot overall solution that is particularly suitable for popular indoor service scenarios. Attached Figure Description

[0037] Figure 1 This is a front view of the robot of this utility model;

[0038] Figure 2 This is a side view of the robot of this utility model;

[0039] Figure 3 This is a structural schematic diagram of the column and robotic arm of the robot of this utility model. Detailed Implementation

[0040] To make the technical problems, technical solutions, and beneficial effects of the embodiments of this utility model clearer, the present utility model will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present utility model and are not intended to limit the present utility model.

[0041] This utility model provides a minimalist mobile manipulation robot (hereinafter referred to as "robot") based on pure visual perception, comprising:

[0042] A wheeled mobile chassis is provided with drive wheels and at least two auxiliary swivel casters, wherein the drive wheels are driven by a motor and equipped with an encoder;

[0043] The column is fixedly installed on the wheeled mobile chassis and is equipped with a lifting platform that can move vertically along it;

[0044] A robotic arm is mounted on the lifting platform, and the robotic arm includes at least a horizontal extension degree of freedom and an end effector gripper opening and closing degree of freedom.

[0045] A multimodal sensing module is installed at the top of the column, above the lifting platform. The multimodal sensing module includes a housing and at least one RGB camera and at least one inertial measurement unit integrated in the housing. The RGB camera and the inertial measurement unit are fixed in the housing by a rigid mounting bracket to maintain a fixed relative pose relationship.

[0046] A local computing unit is installed inside the wheeled mobile chassis. The local computing unit has a graphics processing unit (GPU) computing capability and is communicatively connected to the encoder, the RGB camera, and the inertial measurement unit.

[0047] In some implementations, there are two drive wheels, symmetrically arranged on the left and right sides of the wheeled mobile chassis, each driven by an independent DC servo motor. Forward, backward, and stationary rotation movements are achieved by controlling the speed difference between the two drive wheels. An optical encoder is installed on the motor shaft of each drive wheel to measure the speed and cumulative travel of each drive wheel in real time, providing basic data for the robot's odometer positioning. There are two auxiliary omnidirectional casters, which are unpowered and are installed on the front and rear sides of the wheeled mobile chassis respectively, providing stable support.

[0048] In some implementations, the wheeled mobile chassis has a circular structure, which makes it less likely for the robot to scrape against obstacles when rotating in place in a narrow space.

[0049] In some embodiments, the housing of the multimodal sensing module is detachably mounted on the top of the column. For example, the housing of the multimodal sensing module is mounted to the flange connection surface of the column via a standardized bolt interface, and the electrical connection uses an aviation plug, thereby configuring the multimodal sensing module as a detachable independent structural unit.

[0050] In some embodiments, the robotic arm is a two-degree-of-freedom structure, comprising an upper arm with the horizontal extension-retraction degree of freedom and a gripper mounted at the end of the upper arm with an end gripper opening and closing degree of freedom. For example, the upper arm can be a slender rod-like structure, driven by a stepper motor or DC motor, using a synchronous belt or lead screw linear transmission mechanism to achieve horizontal forward and backward extension movement. For example, the gripper can be a two-finger parallel gripper, driven by a single micro linear motor or servo motor, performing symmetrical opening and closing movements to grasp objects. More preferably, the gripper's finger surfaces are also covered with silicone anti-slip pads to increase friction. The robotic arm does not have a wrist lifting joint; the vertical height adjustment of the gripper at the end of the upper arm is provided by the lifting platform of the column.

[0051] In some embodiments, the column is equipped with a synchronous belt lifting transmission mechanism driven by a stepper motor; a linear guide rail is provided on the outer side of the column, and the lifting platform is slidably mounted on the linear guide rail via a linear slider, with the synchronous belt lifting transmission mechanism connected to the linear slider. Preferably, the lifting platform is configured to be driven at a height (from the ground) of 50-150cm on the column, and the stepper motor has a built-in encoder to accurately measure the current height position of the lifting platform. The lifting platform is used to mount a robotic arm, which is mounted on the lifting platform via a robotic arm base. The vertical working range of the robotic arm can be expanded by changing the height of the robotic arm base.

[0052] In some implementations, the optical axis of the RGB camera is set to tilt forward and downward at 15 to 20 degrees.

[0053] In some implementations, the RGB camera employs an industrial-grade compact USB 3.0 RGB camera module with a resolution of at least 1280x720 pixels and a frame rate of at least 30fps. The camera module is rigidly mounted inside the housing of the multimodal sensing module, and its optical axis is tilted slightly forward and downward at 15 to 20 degrees to allow the field of view to simultaneously cover the mid-to-long-distance scene in front and the operating area in front of the robotic arm. The camera is mounted at the very top of the column (approximately 165cm from the ground), at the highest point of the robot, providing a broad, overlooking view and reducing near-field obstruction.

[0054] In some embodiments, the wheeled mobile chassis has a battery pack, a local computing unit, and a power management module arranged in layers from bottom to top inside. The local computing unit is fixed to a metal mounting plate inside the wheeled mobile chassis, and the wheeled mobile chassis also has ventilation holes and a cooling fan at the positions corresponding to the local computing unit. The bottom layer houses the heavier battery pack, which lowers the robot's center of gravity and improves movement stability. The metal mounting plate can be made of aluminum alloy to facilitate heat dissipation for the local computing unit. In some embodiments, the local computing unit is an NVIDIA Jetson AGX Orin computing platform.

[0055] In some embodiments, the inertial measurement unit is a six-axis inertial measurement unit, comprising a three-axis accelerometer and a three-axis gyroscope.

[0056] In some implementations, the inertial measurement unit can be welded onto the same rigid mounting bracket inside the housing of the multimodal sensing module that houses the RGB camera, so as to maintain a fixed relative position and attitude relationship with the RGB camera, ensuring that the calibration accuracy of the vision-inertial fusion remains constant during use.

[0057] In some implementations, the mobile manipulator does not include depth sensors and lidar.

[0058] In some embodiments, the battery pack is a lithium-ion battery pack, and the power management module includes multiple DC-DC converters for providing power outputs of different voltages to the wheeled mobile chassis, the lifting platform, the robotic arm, the local computing unit, and the multimodal sensing module, respectively.

[0059] The preferred embodiment of the mobile operating robot of this utility model will be further described below with reference to specific parameters.

[0060] like Figure 1-3 As shown, the mobile operation robot in this embodiment includes the following parts:

[0061] 1. Wheeled Mobile Chassis 1 (hereinafter referred to as "Chassis"): A circular structure with a diameter of 340mm and a height of 80mm, made of stamped aluminum alloy sheet, covered with an ABS engineering plastic shell, and with three layers of internal storage space. It is equipped with two DC servo motors (rated power 30W, rated speed 300rpm), each driving a 100mm diameter rubber drive wheel 11, which are symmetrically mounted on both sides of the chassis diameter. The chassis has one 50mm diameter auxiliary omnidirectional caster at the front and rear for stable support. The DC servo motors drive and control the speed difference between the two drive wheels to achieve forward, backward, and stationary rotation. Each drive motor shaft is equipped with an incremental photoelectric encoder with a resolution of 2048 lines / revolution to measure the speed and cumulative travel of each drive wheel in real time, providing basic data for the robot's odometry positioning.

[0062] The chassis's internal space is divided into three layers from bottom to top, separated by an aluminum alloy plate: the bottom layer houses a rechargeable lithium-ion battery pack (24V / 15Ah, 7S configuration) to lower the overall center of gravity and improve stability; the middle layer houses the local computing unit (NVIDIA Jetson AGX Orin computing platform), which is bolted to the aluminum alloy plate, which also serves as a passive cooling base; the chassis shell has an array of ventilation holes and a small, low-noise axial fan for active cooling; the top layer houses the power management module (PMU) and motor drive board, with space reserved for cable routing.

[0063] 2. Column 2: Made of 6063-T5 aluminum alloy rectangular profile, with a cross-section of 40mm x 40mm and a total length of 1550mm. The column can be fixed to the reinforcing plate on the top surface of the chassis via the bottom flange using four M6 bolts, and is located slightly rear of the center of the top surface of the chassis. An MGN12H type linear guide rail 21 (1200mm long) is installed on one side of the column, internally equipped with a GT2-6mm synchronous belt lifting transmission mechanism 22, driven by a NEMA 17 stepper motor 23 (step angle 1.8 degrees, holding torque 0.45N*m, built-in 1000-line encoder). The encoder built into the stepper motor can accurately measure the current height position of the lifting platform 24. The lifting platform 24 is mounted on the linear guide rail via the MGN12H linear slider 25, and can move vertically along the column. The synchronous belt lifting transmission mechanism 22 is connected to the linear slider 25. The lifting stroke of the lifting platform 24 is 1000mm (the lowest position is 500mm from the ground and the highest position is 1500mm from the ground), and the positioning accuracy is better than 1mm.

[0064] 3. Two-DOF robotic arm 3: Robotic arm 3 is fixed to the lifting platform 24 via a robotic arm base and rises and falls together with the lifting platform 24. The vertical working range of the robotic arm can be expanded by changing the height of the robotic arm base. Robotic arm 3 includes a large arm 31 and a gripper 32 mounted at the end of the large arm. The large arm 31 is a slender rod-like structure, and the gripper 32 is a two-finger parallel gripper. Wherein:

[0065] The first degree of freedom (DOF-1) is the horizontal extension and retraction of the boom, which uses a T8 lead screw linear transmission mechanism (8mm lead, 500mm effective stroke), driven by a NEMA 14 stepper motor. The main body of the boom is a 20mm x 20mm aluminum alloy square tube. The effective extension and retraction stroke of the boom is 50cm. When the boom is retracted, the end does not exceed the outer contour of the chassis. When fully extended, the end of the boom can reach about 50cm in front of the chassis.

[0066] The second degree of freedom (DOF-2) is the opening and closing of the end effector gripper. The two parallel grippers are driven by an MG996R digital servo motor, with a maximum opening width of 80mm and a gripping force of approximately 10N. The finger surfaces are covered with silicone anti-slip pads (Shore hardness 40A). The two parallel grippers can grasp common small household objects (such as cups, remote controls, toys, etc.).

[0067] The robotic arm has no wrist lifting or rotation joints, weighs approximately 1.5 kg, and boasts an extremely simple and reliable structure. The vertical height adjustment of the end effector gripper is entirely handled by the column. The column provides approximately 1000 mm of lifting stroke, which, combined with the horizontal extension and retraction of the robotic arm, achieves a wide operating space coverage from near the ground to tall cabinet shelves with minimal degrees of freedom.

[0068] 4. Multimodal Sensing Module 4: Installed at the very top of column 2, above lifting platform 24, its height relative to the ground remains constant. The housing is made of ABS engineering plastic, with dimensions of 100mm x 80mm x 60mm and a total weight of approximately 120g. The housing of Multimodal Sensing Module 4 is mounted on the aluminum alloy flange at the top of the column using four M3x8 stainless steel hex bolts. Electrical connections use a GX12-4 pin aviation connector (including two pairs of USB 3.0 differential signal cables + one pair of 5V / 2A power supply cables), which can be plugged in and out to complete all electrical connections. The multimodal sensing module integrates multiple sensors into a compact, independent housing, forming a detachable, standardized sensing unit. Specifically, the housing integrates the following two sensors:

[0069] RGB camera: FLIR Blackfly S BFS-U3-04S2C-CS, 1 / 3-inch sensor, 720x540 pixel resolution, maximum frame rate 120fps (30fps in this system), CS mount with M12 fixed-focus lens (3.6mm focal length, approximately 82 degrees horizontal field of view), USB 3.0 interface. The camera's optical axis is mounted at approximately a 15-degree downward and forward tilt.

[0070] Six-axis IMU (Inertial Measurement Unit): Bosch BMI088 (or equivalent ICM-42688-P) is selected. The three-axis accelerometer has a range of ±16g, and the three-axis gyroscope has a range of ±2000dps. It is connected to the STM32F103 microcontroller inside the casing via the SPI interface. The microcontroller reads data at a frequency of 200Hz and uploads it to the host via the USB CDC virtual serial port.

[0071] The RGB camera and the six-axis IMU are soldered onto the same rigid aluminum alloy mounting plate (60mm x 50mm x 3mm) inside the housing, and their relative pose is fixed after factory calibration.

[0072] The two types of sensors mentioned above, together with the photoelectric encoder on the chassis, constitute the robot's three-modal perception system: the visual modality (provided by an RGB camera) provides visual image data (high-resolution color image sequence information of the environment); the inertial measurement modality (provided by a six-axis IMU) provides angular velocity and linear acceleration information of the robot body; and the odometry modality (provided by an encoder on the chassis) provides displacement increment information in the chassis coordinate system. The data from the three modalities are synchronized and fused through a local computing unit, and the algorithm estimates the precise camera pose corresponding to each frame of image, collaboratively achieving precise robot localization and 3D reconstruction of the environment under conditions without depth sensors.

[0073] This robot can operate without any depth sensors (RGB-D cameras, structured light sensors, ToF sensors, or stereo cameras) or LiDAR. The 3D depth information of the environment is fully recovered through algorithms from a multi-view color image sequence captured by a single RGB camera during robot movement, combined with pose information provided by the IMU and encoder.

[0074] Within the housing of the multimodal sensing module, the sensors maintain a fixed geometric calibration relationship through rigid mounting brackets. A rigid transformation relationship is established between the multimodal sensing module and the chassis coordinate system via the fixed geometric structure of the column. When the lifting platform rises and falls on the column, the displacement of the lifting platform is precisely measured by a stepper motor encoder, which can correct the vertical offset between the multimodal sensing module and the chassis in real time. Therefore, the coordinate transformation relationship between the three modal sensors remains fixed after factory calibration (only one-dimensional known change occurs during platform lifting), eliminating the need for dynamic calibration during operation.

[0075] This multimodal sensing module can operate without a gimbal and lacks independent rotation or pitch adjustment capabilities. Multi-view data acquisition relies on the overall movement of the robot: chassis rotation in place enables 360-degree horizontal scanning; chassis movement allows for multi-view observation from different positions; and the lifting platform's elevation changes the robot arm's height relative to the ground, thereby altering the RGB camera's vertical observation height of the object grasped by the robot arm.

[0076] 5. Local Computing Unit: Utilizes an NVIDIA Jetson AGX Orin 64GB development kit (GPU: 2048-core NVIDIA Ampere architecture CUDA cores, CPU: 12-core Arm Cortex-A78AE, AI computing power up to 275 TOPS). It is secured to a 3mm thick aluminum alloy mounting plate inside the chassis via four M3 bolts. The mounting plate also serves as a passive cooling base. Corresponding ventilation holes are located on the chassis shell, housing a 40mm x 40mm x 10mm low-noise axial fan (4500rpm, noise level below 25dB). This robot can function solely as a local computing platform using only one NVIDIA Jetson AGX Orin, eliminating the need for cloud computing; all perception, decision-making, and control algorithms are completed locally.

[0077] The communication connections between the modules are as follows:

[0078] The RGB camera connects to the local computing unit via a USB 3.0 interface to transmit high frame rate, high resolution image data;

[0079] The IMU reads data through the SPI or I2C interface on the internal PCB of the multimodal sensing module and then connects it to the local computing unit via USB.

[0080] The photoelectric encoder on the chassis is connected to the chassis motor drive board via the ABZ pulse interface. The chassis motor drive board reports the photoelectric encoder data to the local computing unit via CAN bus or UART.

[0081] The local computing unit sends control commands to the chassis motor drive board via CAN bus or UART to the rubber drive wheel and the lifting stepper motor (NEMA 17 stepper motor);

[0082] The local computing unit sends control commands for the extension and retraction of the boom and the opening and closing of the grippers to the robotic arm via UART or RS-485.

[0083] The software communication framework is based on ROS 2 (Robot Operating System 2), with each sensor and actuator encapsulated as an independent ROS node. The local computing unit has a built-in WiFi 6 module, supporting remote debugging and visual monitoring via WiFi, but it does not rely on wireless communication to complete core tasks.

[0084] 6. Power Supply System: Includes a 7-series 2-parallel (7S2P) 18650 lithium-ion battery pack, nominal voltage 25.2V (29.4V when fully charged), total capacity 10Ah (approximately 252Wh), installed at the bottom of the chassis. The Power Management Module (PMU) centrally manages power distribution, integrating 4 DC-DC converters to provide different voltage outputs to the chassis, lifting platform, robotic arm, local computing unit, and multimodal sensing module. Specifically: 24V / 5A (for the motors on the chassis and the lifting stepper motor), 19V / 4A (for the local computing unit), 12V / 2A (for the motors of the robotic arm), and 5V / 3A (for low-power devices such as the RCB camera, IMU, and communication module). The PMU features overcurrent protection (independent fuse for each circuit), overvoltage protection (30V cutoff), undervoltage protection (21V automatic shutdown), and short-circuit protection. Under typical indoor mobile operation tasks, the overall runtime is expected to be approximately 2 to 3 hours.

[0085] The main dimensions of the robot in this preferred embodiment are as follows: height approximately 1700mm (the uprights are of fixed height, so the overall height is constant), chassis diameter 340mm, overall width (when the arm is retracted) approximately 340mm (same width as the chassis), overall depth (when the arm is fully extended) approximately 840mm, and overall weight not exceeding 20kg (approximately 18kg including the battery). The robot can easily pass through standard indoor door frames (800mm wide, 2000mm high). The main frame is made of 6063 aluminum alloy profiles (uprights and linear guides) and aluminum alloy plates (base plate of the chassis), and the outer shell is made of ABS engineering plastic injection molded parts.

[0086] The working process of this utility model is as follows:

[0087] The robot moves within the indoor environment driven by a chassis. During movement, encoders on the chassis continuously measure the rotational speed of each drive wheel, providing odometer data. A six-axis IMU inside the multimodal sensing module housing at the top of the column continuously measures angular velocity and linear acceleration. An RGB camera continuously captures environmental color images at a frame rate of no less than 30fps. The data from these three modalities are aggregated through their respective communication interfaces onto the local computing unit inside the chassis.

[0088] The local computing unit fuses encoder odometry data and IMU data to estimate the precise spatial pose of the camera corresponding to each frame of RGB image. Then, using a sequence of multiple RGB image frames with precise pose markers, it reconstructs a high-fidelity 3D scene representation of the environment online using a 3D Gaussian sputtering-based algorithm, and performs object detection and semantic scene graph construction within this 3D scene. Upon receiving an operation task command, the system locates the target object in the scene graph, generates candidate operation poses around the target, evaluates them using a virtual viewpoint image rendered in the 3D scene, and selects the optimal pose. Subsequently, the local computing unit sends navigation commands to the chassis motor drive board to move the chassis to the target position, sends lifting commands to the lifting stepper motor to adjust the height of the robotic arm, and sends extension and gripper control commands to the robotic arm drive board to complete the object grasping operation.

[0089] While adjusting the height of the robotic arm to match the height of the target object, the column also enables the multimodal sensing module at the top to obtain the observation perspective at that height, realizing the linkage between perception and operation in the vertical dimension.

[0090] For example, the complete workflow of using the mobile robot of this embodiment to perform a typical indoor object handling task (such as "moving the water glass on the table to the coffee table") is as follows:

[0091] Step 1 (Power-on Initialization): The user presses the power switch on the chassis. The PMU starts supplying power, and the local computing unit starts and loads the ROS 2 system and various functional nodes. The column automatically returns to zero (the lifting platform returns to zero after triggering the limit switch when it descends to the lowest position), the robotic arm retracts to its initial position, and the grippers open. The RGB camera and IMU begin data acquisition. The initialization process takes approximately 30 seconds.

[0092] Step Two (Environmental Exploration and Mapping): The user sends exploration commands to the robot via a WiFi remote terminal (laptop or mobile app), or the robot autonomously executes a preset exploration strategy. The chassis's drive wheels activate, and the robot moves within the indoor environment while the chassis rotates in place to achieve a 360-degree field of view scan. During this process, the RGB camera continuously acquires images at 30fps, the IMU continuously acquires inertial data at 200Hz, and the encoders on the drive wheels continuously provide odometer data feedback. Algorithms running on the local computing unit utilize this multimodal data to construct a 3D Gaussian scene representation and semantic Gaussian scene map of the environment in real time. The pillars can adjust their height during exploration, allowing the camera to acquire images from different vertical heights, enriching the field of view coverage of the 3D reconstruction.

[0093] Step 3 (Receiving Task Instructions): The user inputs task instructions via text or voice through a remote terminal (e.g., "Move the water glass on the table to the coffee table"). The Visual Language Model (VLM) on the local computing unit parses the instructions, searches for and locates the target object (water glass) and its placement position (coffee table) in the constructed scene graph.

[0094] Step 4 (Navigation to the vicinity of the target object): The algorithm generates candidate operation poses around the target object and uses a 3D Gaussian scene for counterfactual rendering to generate virtual view images at each candidate pose. The VLM scores and selects the best operation pose. The local computing unit sends navigation commands to the chassis motor drive board (via the CAN bus), and the drive wheels on the chassis rotate differentially, moving the robot to the position and orientation corresponding to the best operation pose.

[0095] Step 5 (Height Adjustment and Arm Extension): Upon reaching the target position, the local computing unit calculates the required lifting platform height and arm extension distance based on the 3D coordinates of the target object in the scene diagram. Control commands are then sent to the lifting stepper motor, causing the lifting platform to rise or fall along the linear guide rail on the column to the target height (e.g., if the target table height is 75cm, the lifting platform rises to approximately 70cm above the ground). Then, an extension command is sent to the arm's stepper motor, and the lead screw drives the arm to extend horizontally forward above the target object.

[0096] Step Six (Object Grabbing): After the upper arm is in position, the grippers, in their open state, are aligned with the target object. The local computing unit sends a gripper closing command to the servo motor, and the two fingers symmetrically close to clamp the object. The silicone pads on the gripper fingers provide sufficient friction to prevent slippage.

[0097] Step 7 (Transfer to Target Location): After gripping the object, the upper arm retracts, and the differential drive of the chassis navigates the robot to the target placement location (next to the coffee table). The lifting platform is adjusted to the placement height, and the upper arm extends again above the coffee table.

[0098] Step 8 (Object Placement): The local computing unit issues a gripper opening command, and the object is released to the target position. The boom retracts, and the lifting platform returns to its default height, completing the task. The system reports the task execution result to the user via a remote terminal.

[0099] The comparison results between this embodiment and existing products are as follows:

[0100] (1) Comparison of this embodiment with Hello Robot Stretch RE1:

[0101] (a) Sensing configuration: In this embodiment, there is one RGB camera + one IMU (integrated housing, approximately 120g). The HelloRobot Stretch RE1 has two RGB-D depth cameras + one LiDAR (distributed installation, total sensor weight approximately 350g or more). This embodiment has fewer sensors, lighter total weight, and higher integration.

[0102] (b) Robotic arm: In this embodiment, it has 2 degrees of freedom (horizontal extension + gripper) and weighs approximately 1.5 kg; the Hello RobotStretch RE1 has a telescopic rod + 3 degrees of freedom DexWrist + gripper and weighs approximately 2.5 kg. This embodiment has a simpler and lighter structure.

[0103] (c) Computing platform: In this embodiment, the computing power is Jetson AGX Orin (275 TOPS GPU computing power), and the Hello RobotStretch RE1 is Intel NUC (without a dedicated GPU). The GPU computing power of this embodiment is significantly superior.

[0104] (d) Depth sensor dependence: This embodiment does not rely on a depth sensor at all, thus avoiding the problem of missing depth; Hello Robot Stretch RE1 relies on RealSense D435i, which has problems with depth range and reflective surface failure.

[0105] (2) Comparison of this embodiment with Fetch Mobile Manipulator (abbreviated as Fetch):

[0106] (a) Robotic arm: This embodiment has 2 degrees of freedom and weighs about 1.5 kg, while Fetch has 7 degrees of freedom and weighs about 7 kg. This embodiment is significantly lighter and simpler.

[0107] (b) Perception configuration: This embodiment does not have a depth sensor and LiDAR. Fetch relies on the PrimeSense depth camera and LiDAR.

[0108] (c) Overall weight: Approximately 18kg in this embodiment, and approximately 113kg in Fetch. This embodiment is suitable for home environments.

[0109] (3) Comparison of this embodiment with PAL Robotics TIAGo (abbreviated as TIAGo):

[0110] (a) Robotic arm: In this embodiment, it has 2 degrees of freedom and weighs about 1.5 kg, while TIAGo has 7 degrees of freedom and weighs about 5 kg or more.

[0111] (b) Sensing configuration: This embodiment is an integrated pure RGB solution, while TIAGo is a distributed RGB-D + LiDAR solution.

[0112] (c) Computing platform: In this embodiment, the computing power of the Jetson Orin GPU far exceeds that of TIAGo's Intel i7 industrial control computer.

[0113] Therefore, the advantages of this preferred embodiment are: (1) the whole machine is lighter and more compact, and the number of sensors and total weight are reduced by eliminating the depth camera and lidar; (2) the perception module has higher integration and better maintainability, and can be quickly disassembled and replaced through modular design; (3) the robotic arm has only 2 degrees of freedom and weighs about 1.5 kg, which is far lower than the weight and complexity of the traditional 6 or 7 degrees of freedom scheme; (4) the GPU computing power far exceeds that of the existing platform, and can run advanced 3D reconstruction and large model inference algorithms locally in real time.

[0114] The basic configuration of this invention, relying on an advanced Vision-Language-Motion (VLA) model and 3D perception algorithm, can already cover the vast majority of basic indoor operational tasks. However, when facing scenarios with special spatial geometric constraints, as a modular functional extension of the minimalist basic configuration, the following variant schemes can also be adopted. It should be noted that such variant schemes only make minimal degree-of-freedom compensation in hardware, and their overall architecture still maintains the lightweight and minimalist characteristics that distinguish it from traditional multi-degree-of-freedom heavy-duty platforms: in other embodiments, the upper arm of the robotic arm can be replaced with a multi-segment telescopic sleeve structure to increase the extension stroke, and a wrist rotation degree of freedom can be added at the end to upgrade it to 3 degrees of freedom.

[0115] In other implementations, the differential drive chassis can be replaced with a Mecanum wheel omnidirectional chassis, enabling the robot to perform lateral translation.

[0116] In other embodiments, two 2-DOF robotic arms can be symmetrically installed on the lifting platform to achieve collaborative operation of the two arms.

[0117] In other implementations, a single-axis rotating gimbal can be added to the top of the column to enable the multimodal sensing module to have independent horizontal rotation capability, achieving a wider field of view scanning without moving the chassis.

[0118] In other embodiments, a single RGB camera in the multimodal sensing module can be replaced with a ring array of 2 to 4 small RGB cameras, fixed in the same housing, to achieve multi-directional synchronous observation at the same time.

[0119] In other embodiments, the structural relationship between the column and the robotic arm can be changed from "the column is in the center and the arm extends forward from the lifting platform" to "the column is offset to one side of the chassis and the arm extends from the side", thereby changing the workspace distribution.

[0120] In summary, this utility model is superior to existing representative platforms in terms of the lightweight nature of the multimodal perception module, the simplicity of the robotic arm structure, the local GPU computing power, and the overall compactness. It is especially suitable for indoor service scenarios with high computing power requirements but simple structural requirements, such as: (1) Elderly care: assisting the elderly or those with mobility difficulties to retrieve daily items such as medicines, water cups, and remote controls at home. (2) Hospitals and laboratories: delivering, picking up, and organizing small items in wards or laboratories. (3) Retail and convenience stores: organizing shelves, picking up and placing goods, and performing simple inventory management in small retail stores. (4) Education and scientific research: serving as a physical experimental platform for mobile operation algorithms to verify research results in areas such as visual perception, 3D reconstruction, semantic understanding, and operation planning. (5) Hotel and catering services: delivering simple items and collecting tableware in hotel rooms or restaurants. (6) Exhibition halls and museums: serving as an interactive guide assistant, moving autonomously in exhibition halls and assisting in simple display operations of exhibits. (7) Special environment inspection: Conduct inspections and simple equipment operations (such as pressing buttons, plugging and unplugging cables) in controlled environments such as cleanrooms and data centers.

[0121] The above description, in conjunction with specific / preferred embodiments, provides a further detailed explanation of the present invention and should not be construed as limiting the specific implementation of the present invention to these descriptions. For those skilled in the art, various substitutions or modifications can be made to these described embodiments without departing from the concept of the present invention, and all such substitutions or modifications should be considered within the protection scope of the present invention. In the description of this specification, the reference to terms such as "an embodiment," "some embodiments," "preferred embodiment," "example," "specific example," or "some examples," etc., indicates that the specific features, structures, materials, or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples. Although embodiments of the present invention and their advantages have been described in detail, it should be understood that various changes, substitutions and alterations may be made herein without departing from the scope defined by the appended claims.

Claims

1. A minimalist mobile manipulation robot based on pure visual perception, characterized in that, include: A wheeled mobile chassis is provided with drive wheels and at least two auxiliary swivel casters, wherein the drive wheels are driven by a motor and equipped with an encoder; The column is fixedly installed on the wheeled mobile chassis and is equipped with a lifting platform that can move vertically along it; A robotic arm is mounted on the lifting platform, and the robotic arm includes at least a horizontal extension degree of freedom and an end effector gripper opening and closing degree of freedom. A multimodal sensing module is installed at the top of the column, above the lifting platform. The multimodal sensing module includes a housing and at least one RGB camera and at least one inertial measurement unit integrated in the housing. The RGB camera and the inertial measurement unit are fixed in the housing by a rigid mounting bracket to maintain a fixed relative pose relationship. A local computing unit is installed inside the wheeled mobile chassis. The local computing unit has a graphics processing unit (GPU) computing capability and is communicatively connected to the encoder, the RGB camera, and the inertial measurement unit.

2. The minimalist mobile operation robot according to claim 1, characterized in that, The housing of the multimodal sensing module is detachably mounted on the top of the column.

3. The minimalist mobile operation robot according to claim 1, characterized in that, The robotic arm is a 2-degree-of-freedom structure, including an upper arm with the horizontal extension degree of freedom and a gripper with the end gripper opening and closing degree of freedom mounted at the end of the upper arm.

4. The minimalist mobile operation robot according to claim 1, characterized in that, The column is equipped with a synchronous belt lifting transmission mechanism, which is driven by a stepper motor; the outside of the column is equipped with a linear guide rail, and the lifting platform is slidably mounted on the linear guide rail via a linear slider, with the synchronous belt lifting transmission mechanism connected to the linear slider.

5. The minimalist mobile operation robot according to claim 1, characterized in that, The optical axis of the RGB camera is set to tilt forward and downward at 15 to 20 degrees.

6. The minimalist mobile operation robot according to claim 1, characterized in that, The wheeled mobile chassis has a battery pack, a local computing unit, and a power management module arranged in layers from bottom to top inside. The local computing unit is fixed to a metal mounting plate inside the wheeled mobile chassis. The wheeled mobile chassis also has heat dissipation ventilation holes and a cooling fan at the position corresponding to the local computing unit.

7. The minimalist mobile manipulation robot according to claim 1, characterized in that, The local computing unit is the NVIDIA Jetson AGX Orin computing platform.

8. The minimalist mobile operation robot according to claim 1, characterized in that, The inertial measurement unit is a six-axis inertial measurement unit, which includes a three-axis accelerometer and a three-axis gyroscope.

9. The minimalist mobile operation robot according to claim 1, characterized in that, The mobile robot does not include depth sensors or lidar.

10. The minimalist mobile manipulation robot according to claim 6, characterized in that, The battery pack is a lithium-ion battery pack, and the power management module includes multiple DC-DC converters for providing power outputs of different voltages to the wheeled mobile chassis, the lifting platform, the robotic arm, the local computing unit, and the multimodal sensing module, respectively.