Wheel type desktop accompanying robot, control system and control method thereof
By employing multimodal perception and a dual-control hierarchical architecture, the system addresses the perception and tracking challenges of desktop companion robots in complex environments, enabling precise and coordinated motion control and safe response, thereby enhancing the system's scalability and real-time performance.
Patent Information
- Application Number
- CN202610081347.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-21
- Publication Date
- 2026-03-20
AI Technical Summary
Existing desktop companion robots suffer from insufficient perception accuracy in complex environments, have a single tracking mechanism, poor action coordination, limited safety protection, high hardware-software coupling, poor scalability, and difficulty in achieving multi-dimensional data fusion and real-time interaction.
A multimodal perception fusion architecture is adopted, which combines visual and auditory data to build a dual-control hierarchical collaborative architecture. Accurate tracking is achieved through the ROS2 communication framework and YOLOv8 target detection algorithm. Combined with multi-axis synchronous servo control and safety protection mechanisms, data transmission latency is reduced.
It improves perception accuracy and tracking stability in complex environments, enables natural coordination of actions and real-time response, reduces collision risk, and enhances system scalability and data transmission efficiency.
Smart Images

Figure CN121696973A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of desktop robot technology, specifically a wheeled desktop companion robot and its control system. Background Technology
[0002] With the development of smart hardware and artificial intelligence technologies, desktop companion robots are increasingly being applied in close-range scenarios such as home entertainment and office assistance. The core requirements focus on the integrated realization of environmental perception, target tracking, and human-like interaction. Existing desktop companion robots mostly employ single-modal perception solutions. Pure visual tracking is susceptible to light and occlusion, while pure auditory positioning suffers from large angle errors and weak anti-interference capabilities, making it difficult to adapt to complex desktop environments. The tracking mechanism only achieves single-dimensional chassis movement, lacking coordination between head and torso movements, resulting in target loss and insufficient accuracy. Regarding human-like interaction, existing products mostly use single-axis independent drive servos, resulting in limited preset movements, stiff motion, and a lack of synchronization between screen animations and mechanical actions, leading to poor immersion. Safety protection relies on a single sensor, lacking multiple verifications, resulting in delayed responses and a high risk of collisions. The hardware and software architecture is mostly a direct-drive mode for the main control board, with high coupling between upper-layer algorithms and lower-layer drivers, leading to large data transmission latency, poor scalability, and difficulty in balancing the operation of intelligent algorithms with the real-time performance of hardware drivers.
[0003] The perception mode is singular and lacks multi-dimensional data fusion, resulting in a significant decrease in tracking accuracy when affected by environmental factors. During tracking, the movements of the chassis, head, and arm are independent, and no motion-attitude-interaction coordination logic is established, leading to a high target loss rate. The servo control lacks a multi-axis synchronization strategy, the motion parameters are not optimized in conjunction with human movement habits, and there is a lack of linkage between the screen and mechanical movements, resulting in insufficient interactive immersion. Safety protection relies solely on a single sensor and lacks a detection-verification-stop closed-loop design, leading to a high risk of collision. The hardware and software adopt a single control unit direct-drive architecture, with high coupling between upper-layer logic and lower-layer drivers, resulting in large data latency, poor scalability, and hindering future technology upgrades. Summary of the Invention
[0004] The purpose of this invention is to provide a wheeled desktop companion robot to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A wheeled desktop companion robot includes a robot frame consisting of a power torso, a mobile part, limbs, and a head. A camera module is embedded in the nose protrusion of the head, and a lidar is installed on the forehead area of the head. The head and the power supply torso are movably connected by tilt servo motor and rotation servo motor, and the limbs are provided with elbow servo motor and shoulder servo motor. A sound positioning motherboard is installed at the rear of the power supply body. The sound positioning motherboard is equipped with an array of sound acquisition microphones. The moving part includes two moving wheels with collinear axes arranged in a triangle and a microphone wheel. The microphone wheel is located below the sound positioning motherboard.
[0006] As a further aspect of the present invention: the eyes of the head are provided with an LED display screen.
[0007] A control system for a wheeled desktop companion robot includes a main control board located at the back of the head, a microcontroller located within the power supply torso, and a voice wake-up module, wherein the microcontroller is located between the sound positioning main board and the microphone wheel.
[0008] As a further aspect of the present invention: the main control board is a Jetson Nano, which is responsible for receiving the data collected by the LiDAR, the camera module, and the sound positioning motherboard, and running intelligent algorithms and issuing control commands.
[0009] As a further aspect of the present invention: the microcontroller is an STM32F407, which serves as an auxiliary control unit to precisely control the motion drive of the tilt servo motor, the rotation servo motor, the elbow servo motor, the shoulder servo motor, and the moving part.
[0010] A control method for a wheeled desktop companion robot's control system, wherein the human tracking function relies on the ROS2 communication framework and the YOLOv8 target detection algorithm, the camera module acquires environmental images in real time, and the main control board uses an algorithm to filter out human targets with a confidence level in the center region of the image, prioritizing targets closest to the image center. Precise control commands are generated by analyzing the horizontal position and area parameters of the target bounding box. When the target is on the right side of the image, the chassis wheel motors are driven to rotate in the opposite direction at a differential speed to achieve a left turn; if the target is to the left, the wheels are controlled to rotate in the opposite direction at a differential speed to achieve a right turn; if the target is in a horizontal dead zone, for example... Within 0.1% of the image width, the robot immediately stops turning to avoid frequent jitter. For distance control, if the target box area is smaller than the preset target ratio of 0.6 and exceeds the dead zone of 0.2% of the target area, the wheels are controlled to rotate in the same direction to drive the robot forward. If the target box area is too large and exceeds the dead zone, the robot is controlled to move backward. When the area is within the dead zone, the robot remains stationary. At the same time, the head dual-axis servo and the neck servo work together to adjust the head pitch and rotation angle in real time to ensure that the human target is always in the center of the camera's field of view. This forms a dual tracking mechanism of chassis movement + dynamic head posture adjustment, which improves tracking stability and accuracy.
[0011] As a further aspect of this invention: the sound source tracking function uses a microphone array and a sound positioning motherboard to collaboratively collect ambient sound source signals, accurately calculates the sound source angle information within the 0~360° range, and then transmits the data to the motion_control node. The node quickly calculates the required turning angle and duration based on the sound source angle, and controls the chassis to smoothly turn towards the sound source direction at a turning speed of 0.5 rad / s. During the turning process, the lidar continuously scans the surrounding environment. If an obstacle is detected within 0.75 meters three times consecutively, a safety protection mechanism is immediately triggered, the turning action is stopped, and a prompt is issued to the user through the voice module. After the obstacle is removed or safety is confirmed, the turning command can be executed again. After the turning is completed, the head synchronously turns to the sound source direction and enters an interactive waiting state, ready to respond to the user's subsequent voice commands or action requests at any time.
[0012] As a further aspect of this invention: the anthropomorphic interaction function is activated by the sherpa-onnx-kws voice wake-up module after receiving the user's wake-up word. The intelligent agent triggers corresponding action groups according to different interaction scenarios: in the greeting scenario, the shoulder servo motor drives the right arm to rise 45°, while the elbow servo motor controls the right elbow to bend 90°, simulating a natural waving motion, and simultaneously controls the tilt servo motor to drive the head to slowly swing left and right, while the eye screen plays a greeting animation; when expressing happiness, the tilt servo motor and the rotation servo motor drive the head to nod up and down or rotate, while the shoulder servo drives the arms to swing back and forth in a small amplitude, and the eye screen simultaneously displays the corresponding happy expression animation; when entering standby sleep mode, all servos are precisely reset to their initial positions, the wheel motors stop running, and the eye screen switches to a low-power state, reducing energy consumption while preparing for the next wake-up.
[0013] Compared with the prior art, the beneficial effects of the present invention are: This technical solution constructs a multimodal perception fusion architecture, integrating visual, auditory, and environmental perception data. Through a communication framework, it achieves real-time interaction and complementary verification, solving the problem that single-modal perception is susceptible to interference from light, noise, and occlusion. It builds a dual-control layered collaborative architecture, with the upper main control board focusing on intelligent algorithm operation and decision logic output, and the lower auxiliary control board responsible for hardware driving and status feedback. The software modules adopt a modular design and are decoupled by topic to reduce data transmission latency, while precisely matching the control requirements of the mechanical structure to improve the synchronization of command execution.
[0014] Based on the above advantages, an innovative mechanical collaborative structure design simplifies the number of data cables at the joints. Vision and radar are integrated into the head and connected to the main control board. This design reduces obstacles to image acquisition without affecting the appearance or enhancing aesthetics. The wiring is encapsulated in the head, while the microcontroller is located in the power supply section. The servo motor control and power supply lines are also encapsulated in the power supply section. The main control board, sound positioning motherboard, and microcontroller are distributed at the rear. The data interconnection between the main control board, sound positioning motherboard, and microcontroller does not affect the aesthetics of the robot's front, ensuring that the application of this wiring distribution method is not limited by appearance. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A 3D schematic diagram of a wheeled desktop companion robot; Figure 2 A three-dimensional schematic diagram of the head of a wheeled desktop companion robot; Figure 3 A three-dimensional schematic diagram of the torso of a wheeled desktop companion robot; Figure 4 A three-dimensional schematic diagram of the torso of a wheeled desktop companion robot from another perspective; Figure 5 This is a core motion logic diagram for a wheeled desktop companion robot. In the diagram: 1. Power supply torso; 2. Motion unit; 21. Wheel; 3. Limbs; 31. Elbow servo motor; 32. Shoulder servo motor; 4. Head; 41. Tilt servo motor; 42. Rotation servo motor; 5. LiDAR; 6. Camera module; 7. Main control board; 8. Sound positioning mainboard; 9. Microcontroller. Detailed Implementation
[0017] All standard parts used in this application can be purchased from the market, and irregular parts can be customized according to the description and drawings. The specific connection methods of each part adopt conventional methods such as bolts, rivets, welding, and bonding that are mature in the prior art. The components used for circuit connection are all conventional models in the prior art.
[0018] Meanwhile, in order to clearly express the connection relationship and working principle between the components and highlight the key points, the accompanying drawings in the instruction manual are organized and drawn in the form of simplified diagrams. One simplified diagram can correspond to multiple materials and actual external structural shapes.
[0019] Please see Figures 1-5 Example 1: In this embodiment, a robot frame is composed of a power supply torso 1, a moving part 2, limbs 3, and a head 4. A camera module 6 is embedded in the nose protrusion of the head 4, and a lidar 5 is installed on the forehead area of the head 4. The head 4 and the power supply torso 1 are connected by an angle servo motor 41 and a rotation servo motor 42. An elbow servo motor 31 and a shoulder servo motor 32 are provided inside the limbs 3. A sound positioning motherboard 8 is installed at the rear of the power supply torso 1. An array of sound acquisition microphones is provided on the sound positioning motherboard 8. The moving part 2 includes two moving wheels with collinear axes arranged in a triangular distribution and a microphone wheel 21. The microphone wheel 21 is located below the sound positioning motherboard 8. An LED display screen is provided in the eyes of the head 4.
[0020] The head 4 is connected to the neck base of the power supply torso 1 via tilt servo motor 41 and rotation servo motor 42. The neck base 21 is located above the power supply torso 1. The power supply torso 1 contains a battery that provides power for the entire device. Dual-axis servo motor 1 is integrated into the tilt servo motor 41 and is responsible for driving the head to achieve pitch movement. The servo motor output shaft is rigidly connected to the internal support of the head. By adjusting the servo motor angle (range 0~4095°, corresponding to 0~360°, the actual pitch angle is set to -30°~30° through a limit switch), servo motor 2 is installed between the rotation servo motor 42 and the neck base to drive the head to achieve left and right rotation movement, with a rotation angle range of 0~180°, ensuring that the lidar 5 and camera module 6 can cover a wider field of view. The outer shell of the head 4 encloses the internal support 13 and various sensing components. The ears, as decorative interactive components, move synchronously with the head.
[0021] Secondly, by placing the camera module 6 on the nose protrusion of the head 4, the camera's image acquisition will not be interfered with, and the lens color and lens gloss will blend in with the appearance.
[0022] Furthermore, the sound positioning motherboard 8 is equipped with an array of distributed sound acquisition microphones. The target location is calculated based on the time difference between the acquisitions from multiple microphones, using this array of microphones combined with a time-delay positioning algorithm. The main reason for placing the sound positioning motherboard 8 at the rear is to avoid the airflow generated when a person speaks directly affecting the microphones, thus preventing damage. Being at the rear avoids this direct exposure to airflow.
[0023] The limb part 3 is connected to the power supply torso 1 via the elbow servo motor movement part 31 and the shoulder servo motor movement part 32. The left and right servo motor output ends are connected to the arm via the shaft connector, responsible for driving the arm to complete the shoulder lifting, lowering, and forward and backward swinging movements. The shoulder servo motor 3 is installed inside the shoulder servo motor movement part 32. The shoulder servo motor 3 uses position mode control, with an angle range of 0~4095. By adjusting the target position of the servo motor, the arm lifting angle can be 0~90° and the forward and backward swinging angle can be -45°~45°. The shaft servo motor 4 is installed in the elbow servo motor movement part 31, one for each arm, responsible for driving the arm to achieve elbow bending and extension, with a bending angle range of 0~120°, and returning to the initial 0° position when extended. The hand is fixed to the end of the limb part 3 and moves synchronously with the arm. It can realize anthropomorphic movements such as waving through preset action groups. Its movement is controlled by the shoulder servo motor 3 and the shaft servo motor 4 in coordination. All servos are Feite STS models, which are controlled by STM32F407 and ft_driver function package. They support synchronous writing of instructions to achieve coordinated movement of multiple servos. By adjusting the servo speed (0~1000) and acceleration (0~100) parameters, smooth and natural movement is ensured.
[0024] The mobile unit 2 consists of a chassis shell, wheel motors, wheels, a wheel bracket, a wheel 21, and motor mounting hardware. The wheel 21 is mounted on the wheel bracket on the chassis and is fixedly connected to the wheel motors via the motor mounting hardware. The wheel motors receive speed control commands from the main control board, driving the wheels to achieve forward, backward, left, right, and diagonal movement of the robot. Forward / backward movement is achieved by dual motors rotating in the same direction, while turning is achieved by dual motors rotating in opposite directions at different speeds. The linear velocity is fixed at 0.25 m / s, and the angular velocity is fixed at 0.3 rad / s. The lidar 5 monitors the surrounding environment in real time. When an obstacle is detected within 0.75 meters, a stop command is sent from below to control the wheel motors to stop moving, ensuring safe operation. Example 2
[0025] In this embodiment, a main control board 7 is installed at the back of the head 4, a microcontroller 9 is installed in the power supply torso 1, and a voice wake-up module is included. The microcontroller 9 is located between the sound positioning motherboard 8 and the microphone 21. The main control board 7 is a Jetson Nano. The main control board 7 is responsible for receiving the data collected by the lidar 5, the camera module 6, and the sound positioning motherboard 8, and running intelligent algorithms and issuing control commands. The microcontroller 9 is an STM32F407. The microcontroller 9 serves as an auxiliary control unit, precisely controlling the motion drive of the tilt servo motor 41, the rotation servo motor 42, the elbow servo motor 31, the shoulder servo motor 32, and the moving part 2.
[0026] The specific usage method is as follows: The human tracking function relies on the ROS2 communication framework and the YOLOv8 target detection algorithm. Camera module 6 acquires environmental images in real time. The main control board 7 uses algorithms to filter human targets with a confidence level in the image's central region, prioritizing targets closest to the image center. Precise control commands are generated by analyzing the target's horizontal position and area parameters. When the target is on the right side of the image, the chassis wheel motors are driven to rotate in the opposite direction at a differential speed to achieve a left turn; if the target is slightly to the left, the wheels are controlled to rotate in the opposite direction at a differential speed to achieve a right turn. If the target is in a horizontal dead zone, for example, within 0.1% of the image width... The robot immediately stops turning to avoid frequent shaking. In terms of distance control, if the target frame area is smaller than 0.6 times the preset target ratio and exceeds the dead zone (0.2 times the target area), the wheels are controlled to rotate in the same direction to drive the robot forward. If the target frame area is too large and exceeds the dead zone, the robot is controlled to move backward. When the area is within the dead zone, the robot remains stationary. At the same time, the head dual-axis servo motor and the neck servo motor work together to adjust the head pitch and rotation angles in real time to ensure that the human target is always in the center of the camera's field of view. This forms a dual tracking mechanism of chassis movement + dynamic head posture adjustment, which improves tracking stability and accuracy.
[0027] The sound source tracking function uses a microphone array and a sound positioning motherboard to collect ambient sound source signals, accurately calculates the sound source angle information within a 0~360° range, and then transmits the data to the motion_control node. The node quickly calculates the required turning angle and duration based on the sound source angle, and controls the chassis to smoothly turn towards the sound source at a turning speed of 0.5 rad / s. During the turning process, the LiDAR continuously scans the surrounding environment. If an obstacle is detected within 0.75 meters three times consecutively, the safety protection mechanism is immediately triggered, the turning action is stopped, and a prompt is issued to the user via the voice module. After the obstacle is removed or safety is confirmed, the turning command can be executed again. After the turning is completed, the head synchronously turns to the sound source direction and enters an interactive waiting state, ready to respond to the user's subsequent voice commands or action requests at any time.
[0028] The anthropomorphic interaction function is activated by the sherpa-onnx-kws voice wake-up module after receiving the user's wake-up word. The intelligent agent triggers corresponding action groups according to different interaction scenarios: In the greeting scenario, the shoulder servo motor 32 drives the right arm to rise 45°, while the elbow servo motor 31 controls the right elbow to bend 90°, simulating a natural waving motion. Simultaneously, the tilt servo motor 41 drives the head to slowly swing left and right, accompanied by a greeting animation playing on the eye screen. When expressing happiness, the tilt servo motor 41 and the rotation servo motor 42 drive the head to nod up and down or rotate, while the shoulder servo drives the arms to swing back and forth in a small amplitude. The eye screen simultaneously displays the corresponding happy expression animation. When entering standby sleep mode, all servos are precisely reset to their initial positions, the wheel motors stop running, and the eye screen switches to a low-power state, reducing energy consumption while preparing for the next wake-up.
Claims
1. A wheeled desktop companion robot, comprising a robot frame consisting of a power torso (1), a mobile part (2), limbs (3), and a head (4), characterized in that: A camera module (6) is embedded in the nose protrusion of the head (4), and a lidar (5) is installed on the forehead area of the head (4). The head (4) and the power supply torso (1) are connected by tilt servo motor movement part (41) and rotation servo motor movement part (42). The limb part (3) is provided with elbow servo motor movement part (31) and shoulder servo motor movement part (32). The power supply body (1) is equipped with a sound positioning motherboard (8) at the rear. The sound positioning motherboard (8) is provided with an array of sound acquisition microphones. The moving part (2) includes two moving wheels with collinear axes arranged in a triangular distribution and a microphone wheel (21). The microphone wheel (21) is located below the sound positioning motherboard (8).
2. The wheeled desktop companion robot according to claim 1, characterized in that: The head (4) is equipped with an LED display screen in the eye area.
3. A control system for a wheeled desktop companion robot, characterized in that: Used to control a wheeled desktop companion robot as proposed in claim 2.
4. The control system for a wheeled desktop companion robot according to claim 3, characterized in that: It includes a main control board (7) located at the back of the head (4), a microcontroller (9) located in the power supply torso (1), and a voice wake-up module. The microcontroller (9) is located between the sound positioning main board (8) and the microphone (21).
5. The control system for a wheeled desktop companion robot according to claim 4, characterized in that: The main control board (7) is a Jetson Nano. The main control board (7) is responsible for receiving the data collected by the lidar (5), the camera module (6), and the sound positioning motherboard (8), and running intelligent algorithms and issuing control commands.
6. The control system for a wheeled desktop companion robot according to claim 5, characterized in that: The microcontroller (9) is an STM32F407. The microcontroller (9) serves as an auxiliary control unit and precisely controls the motion drive of the tilt servo motor (41), the rotation servo motor (42), the elbow servo motor (31), the shoulder servo motor (32), and the moving part (2).
7. The control method for the control system of a wheeled desktop companion robot according to claim 4, characterized in that: The human tracking function relies on the ROS2 communication framework and the YOLOv8 target detection algorithm. The camera module (6) collects environmental images in real time. The main control board (7) filters out human targets with a confidence level in the center area of the image through an algorithm, prioritizing targets closest to the center of the image. By analyzing the horizontal position and area parameters of the target bounding box, precise control commands are generated. When the target is on the right side of the image, the chassis wheel motor is driven to rotate in the opposite direction to achieve a left turn; when the target is to the left, the wheels are controlled to rotate in the opposite direction to achieve a right turn; if the target is in a horizontal dead zone, for example, occupying 0.1 of the image width... Within the specified ratio, the robot immediately stops turning to avoid frequent shaking. In terms of distance control, if the target frame area is smaller than 0.6 times the preset target ratio and exceeds the dead zone (0.2 times the target area), the wheels are controlled to rotate in the same direction to drive the robot forward. If the target frame area is too large and exceeds the dead zone, the robot is controlled to move backward. When the area is within the dead zone, the robot remains stationary. At the same time, the head dual-axis servo motor and the neck servo motor work together to adjust the head pitch and rotation angles in real time, ensuring that the human target is always in the center of the camera's field of view. This forms a dual tracking mechanism of chassis movement + dynamic head posture adjustment, improving tracking stability and accuracy.
8. The control method for the control system of a wheeled desktop companion robot according to claim 3, characterized in that: The sound source tracking function uses a microphone array and a sound positioning motherboard to collect ambient sound source signals, accurately calculates the sound source angle information within a 0~360° range, and then transmits the data to the motion_control node. The node quickly calculates the required turning angle and duration based on the sound source angle, and controls the chassis to smoothly turn towards the sound source at a turning speed of 0.5 rad / s. During the turning process, the LiDAR continuously scans the surrounding environment. If an obstacle is detected within 0.75 meters three times consecutively, the safety protection mechanism is immediately triggered, the turning action is stopped, and a prompt is issued to the user via the voice module. After the obstacle is removed or safety is confirmed, the turning command can be executed again. After the turning is completed, the head synchronously turns to the sound source direction and enters an interactive waiting state, ready to respond to the user's subsequent voice commands or action requests at any time.
9. The control method for the control system of a wheeled desktop companion robot according to claim 3, characterized in that: The anthropomorphic interaction function is activated by the sherpa-onnx-kws voice wake-up module after receiving the user's wake-up word. The intelligent agent triggers the corresponding action group according to different interaction scenarios: In the greeting scenario, the shoulder servo motor (32) drives the right arm to rise 45°, while the elbow servo motor (31) controls the right elbow to bend 90° to simulate a natural waving action. Simultaneously, the tilt servo motor (41) drives the head to swing slowly left and right, and the eye screen plays a greeting animation. When expressing a happy emotion, the tilt servo motor (41) and the rotation servo motor (42) drive the head to nod up and down or rotate, and the shoulder servo drives the arms to swing back and forth in a small amplitude. The eye screen displays the corresponding happy expression animation simultaneously. When entering the standby sleep mode, all servos are precisely reset to the initial position, the wheel motor stops running, and the eye screen switches to a low power consumption state to reduce energy consumption and prepare for the next wake-up.