Large model visual impairment auxiliary robot system based on edge cloud collaboration
By using a large-scale model-based visually impaired assistive robot system with edge-cloud collaboration, combining depth cameras, LiDAR, and cloud-based large language models, the system addresses the problem of insufficient information transmission in complex environments by existing guide devices and guide dogs. It achieves efficient autonomous navigation and accurate perception, reduces hardware costs, and enhances the user experience.
Patent Information
- Application Number
- CN202511096189.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
Existing guide devices and guide dogs are limited in their information delivery dimension, unable to provide semantic descriptions of complex environments, lack the support of large language models, cannot perform natural voice interaction, and their algorithms rely on local embedded computing resources, resulting in high costs and a high barrier to entry for users. They also fail to meet the requirements of content completeness and experience consistency in professional scenarios.
The system employs a large-scale model-based visually impaired assistive robot system with edge-cloud collaboration. It integrates walking components, skeleton support, and electrical control system. It combines depth cameras, LiDAR, and microphones for environmental perception and utilizes a large cloud-based language model for image recognition and voice interaction to achieve autonomous navigation and information broadcasting.
It achieves accurate perception and information acquisition in complex environments, autonomous navigation capabilities, reduced hardware costs, improved user experience, and is suitable for the elderly and visually impaired. Voice response latency is controlled within 1.2 seconds, obstacle avoidance success rate is high, and resource computing efficiency is optimized.
Smart Images

Figure CN120983249A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent robot technology, and more specifically to a large-scale visually impaired assistive robot system based on edge-cloud collaboration. Background Technology
[0002] Existing methods of guiding the visually impaired primarily involve guide canes and guide dogs. Guide canes are typically lightweight, straight-barreled structures with a ground-sensing device at the end for detecting obstacles. Guide dogs, after training, possess abilities such as path memory and obstacle avoidance, and visually impaired users interact with them via a leash. However, both methods suffer from the following problems: information transmission is limited, only reflecting physical obstacle information and failing to provide semantic descriptions of complex environments; real-time perception and interpretation of the environment are poor, failing to meet the assistance needs of high-information-density scenarios such as cultural spaces (e.g., museums); guide dog training is lengthy and costly (approximately 200,000 yuan per dog), resulting in low efficiency in the allocation of social resources; and they lack interactive feedback mechanisms, preventing real-time adjustments based on user needs.
[0003] To address the aforementioned issues, existing technologies include electronic assistive devices for the visually impaired, such as ultrasonic canes and hexapod guide robots. These devices often integrate infrared or ultrasonic sensors, coupled with vibration motors or voice feedback systems, to provide users with obstacle warnings. For example, the hexapod guide robot developed by Shanghai Jiao Tong University possesses environmental perception and path-following capabilities, enabling it to move using its mechanical legs. However, the perception methods of these devices are limited to geometric obstacle recognition, failing to achieve image-level content recognition or visual understanding; they lack the support of large language models, hindering natural voice interaction or intelligent content explanation; their algorithms rely on local embedded computing resources, resulting in high robot costs and difficult upgrades; their complex construction presents a high barrier to entry for users, making them unsuitable for the elderly and visually impaired. AI-interactive smart speakers or mobile apps, by connecting to remote volunteers or customer service personnel, provide users with voice assistance and text explanations based on camera images. However, these systems rely on remote connections or human assistance, making fully autonomous operation impossible; they lack physical guidance, requiring users to hold the device, hindering autonomous navigation; image recognition capabilities and language description quality vary, failing to meet the requirements of content completeness and user experience consistency in professional scenarios; and they cannot be integrated with physical devices, leading to issues such as latency, security, and privacy breaches. Therefore, providing a large-scale model-based visually impaired assistive robot system based on edge-cloud collaboration is a pressing problem for those skilled in the art. Summary of the Invention
[0004] In view of this, the present invention provides a large-scale visually impaired assistive robot system based on edge-cloud collaboration, which has functions such as image recognition, voice interaction, autonomous navigation, and information broadcasting, and is suitable for visually impaired users to travel independently and obtain information in complex public spaces.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A large-scale visually impaired assistive robot system based on edge-cloud collaboration includes a walking component, a skeleton support, a sitting / lying structure, and an electrical control system. The walking component is located below the skeleton support and is used to enable the robot to move and turn. The sitting / lying structure is located in the middle of the skeleton support and is used to provide a seating function. The electrical control system is used to provide sensing and control functions.
[0007] Optionally, the electrical control system includes a main control module, a drive module, a steering module, a sensing module, an audio module, and a cloud computing module; wherein, the drive module and the steering module are used to control the movement and steering of the walking components, respectively, the audio module is used to play voice prompts, the sensing module is used to collect environmental data, and the cloud computing module is used to identify environmental information and generate voice prompts.
[0008] Optionally, the main control module is located inside the main control box, which is located below the seated / lying structure. The main control module includes a main controller, a lithium battery, a power management unit, a wireless communication unit, and an edge computing unit. The edge computing unit integrates a local path planning algorithm.
[0009] Optionally, the drive module is a DC motor connected to the walking component, and the main control module controls the speed and direction of the DC motor through PWM signals; the steering module is a servo motor connected to the walking component, used to control the left and right steering of the walking component.
[0010] Optionally, the perception module includes a depth camera, a LiDAR, and a microphone mounted on the skeleton support. The depth camera acquires 3D environmental information, the LiDAR scans the surrounding environment in real time, and the microphone acquires voice information.
[0011] Optionally, the cloud computing module is equipped with a large language model, an image recognition module, and a TTS speech generation model. The image recognition module recognizes image information, and the recognition result is converted into auditory language text by the large language model. The auditory language text is then converted into prompt speech by the TTS speech synthesis model.
[0012] Optionally, the frame support includes a front support rod, a rear support rod, and a telescopic rod, which together form a triangular mechanical support.
[0013] Optionally, a handle connector and handle are provided at the top of the rear support rod to form a two-hand grip structure.
[0014] Optionally, the walking assembly includes drive wheels, a steering mechanism, and front wheels. The drive wheels are steered by the steering mechanism, which is connected to a telescopic rod. The front wheels are connected to a front support rod.
[0015] Optionally, the sitting / reclining structure includes a support plate, a seat cushion, a backrest, and a footrest. The footrest is located at the lower end of the front support rod, the support plate is located in the middle of the front and rear support rods, the seat cushion is located on the support plate, and the backrest is located at the upper end of the rear support rod.
[0016] As can be seen from the above technical solution, compared with the prior art, the present invention provides a large-scale visually impaired assistive robot system based on edge-cloud collaboration, which has the following beneficial effects:
[0017] 1. Significantly Enhanced Precision Perception and Information Acquisition Capabilities: This invention introduces a multimodal AI model to achieve automatic image content recognition and auditory language output. Based on a visual recognition API, the image-to-speech module, combined with a trained and optimized auditory language style, enables real-time recognition and broadcasting of museum paintings, exhibits, scenes, and other environments. The accuracy of the speech description reaches 92.6%; the speech response latency is controlled within 1.2 seconds; and the trained language content more closely resembles the actual auditory cognition of visually impaired users (e.g., perceptual descriptions such as "what time is it?" or "click sound").
[0018] 2. Achieving high-precision autonomous navigation and obstacle avoidance control: This invention integrates Karto SLAM and TEB local path planning algorithms, enabling mapping and real-time obstacle avoidance in dynamic environments. In simulation and field tests, the average mapping error is less than 5cm; the obstacle avoidance success rate reaches 96.3% in a conventional indoor environment; and it can independently complete autonomous navigation tasks within a range of 200m.
[0019] 3. Edge-cloud collaborative optimization of system resources and computing efficiency: This invention adopts a combination of lightweight local hardware and powerful cloud services, avoiding the high cost and high power consumption problems of traditional robots that require high-performance local GPUs. The power consumption of a single device is less than 20W; the average CPU utilization rate of the system is <40%; the stability of cloud API calls is >99.5% and the average response latency is <1.1 seconds. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of the electrical control system of the visually impaired assistive robot of the present invention;
[0022] Figure 2 This is a schematic diagram of the structure of the visually impaired assistive robot of the present invention;
[0023] In the diagram: 1-Drive wheel, 2-Steering mechanism, 3-Telescopic rod, 4-Main control box, 5-Rear support rod, 6-Backrest, 7-Handle connector, 8-Handle, 9-Audio module, 10-Depth camera, 11-Handbrake, 12-Seat cushion, 13-Support plate, 14-Front support rod, 15-Foot pedals, 16-LiDAR, 17-Front wheel. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] This invention discloses a large-scale visually impaired assistive robot system based on edge-cloud collaboration, including a walking component, a skeleton support, a sitting / lying structure, and an electrical control system; the walking component is located below the skeleton support and is used to realize the movement and steering of the assistive robot; the sitting / lying structure is located in the middle of the skeleton support and is used to realize the seating function; the electrical control system is used to realize the perception and control functions.
[0026] Furthermore, the electrical control system includes a main control module, a drive module, a steering module, a sensing module, an audio module 9, and a cloud computing module; among which, the drive module and the steering module are used to control the movement and steering of the walking components, the audio module 9 is used to play voice prompts, the sensing module is used to collect environmental data, and the cloud computing module is used to identify environmental information and generate voice prompts.
[0027] Furthermore, the main control module is located inside the main control box 4, which is situated below the seated / reclining structure, as shown below. Figure 1 As shown, the main control module includes a main controller, a lithium battery, a power management unit, a wireless communication unit, and an edge computing unit. The edge computing unit integrates a local path planning algorithm.
[0028] In this embodiment of the invention, the main control module uses a low-power edge MCU (STM32F407, 168MHz, based on ARM Cortex-M4 core) to control the underlying motion, motors, and sensors, while delegating complex tasks such as image recognition, speech synthesis, and language understanding to the cloud computing module. Edge-cloud data transmission is achieved through a wireless communication unit (WiFi / 4G), reducing local hardware costs, breaking through the traditional limitations of robot computing power, and realizing a lightweight terminal, cloud-intensive intelligent deployment approach.
[0029] In this embodiment of the invention, the edge computing unit combines LiDAR 16 and depth camera 10 to achieve simultaneous localization and mapping (SLAM), and generates a global path through Dijkstra's algorithm and performs local dynamic obstacle avoidance through TEB algorithm. The path is updated in real time, and the autonomous walking trajectory planning ensures that the robot can operate stably in dynamic environments, replacing the manual companionship and guidance method for the blind.
[0030] 4G / WiFi can be replaced with NB-IoT or LoRa communication to adapt to different signal coverage scenarios.
[0031] Furthermore, the drive module is a DC motor connected to the walking component, and the main control module controls the speed and direction of the DC motor through PWM signals; the steering module is a servo motor connected to the walking component, used to control the left and right steering of the walking component.
[0032] Furthermore, the perception module includes a depth camera 10, a lidar 16, and a microphone mounted on the skeleton support. The depth camera 10 acquires three-dimensional environmental information, the lidar 16 scans the surrounding environment in real time, and the microphone acquires voice information.
[0033] Furthermore, the cloud computing module is equipped with a large language model, an image recognition module, and a TTS speech generation model. The image recognition module recognizes image information, and the recognition results are converted into auditory language text by the large language model. The auditory language text is then converted into prompt speech by the TTS speech synthesis model.
[0034] In this embodiment of the invention, the cloud-based large language model uses a visual large model (such as BLIP-2 or Baidu Visual Recognition SDK) to perform image content recognition on the user's environment. Then, it calls an auditory language large model to generate auditory language text optimized for visually impaired users from the recognition results. Finally, it is converted into speech audio by PaddleSpeech-TTS and played back locally.
[0035] The visual model can also use VisualGPT, MiniGPT, or Baidu Wenxin visual model. All models with image content understanding capabilities are suitable. The TTS speech generation model can also use iFlytek Open Platform or EdgeTTS model, which have API call capabilities.
[0036] Furthermore, such as Figure 2 As shown, the frame support includes a front support rod 14, a rear support rod 5, and a telescopic rod 3, which together form a triangular mechanical support.
[0037] Furthermore, such as Figure 2 As shown, the top of the rear support rod 5 is provided with a handle connector 7 and a handle 8, forming a two-hand grip structure.
[0038] In this embodiment of the invention, a handbrake 11 is also included. The handbrake 11 is disposed on the handle 8 for manual braking control by the user, thereby enhancing safety.
[0039] Furthermore, such as Figure 2 As shown, the walking assembly includes a drive wheel 1, a steering mechanism 2, and a front wheel 17. The drive wheel 1 is steered by the steering mechanism 2, which is connected to the telescopic rod 3. The front wheel 17 is connected to the front support rod 14.
[0040] Furthermore, such as Figure 2 As shown, the sitting / reclining structure includes a support plate 13, a seat cushion 12, a backrest 6, and a footrest 15. The footrest 15 is located at the lower end of the front support rod 14, the support plate 13 is located in the middle of the front support rod 14 and the rear support rod 5, the seat cushion 12 is located on the support plate 13, and the backrest 6 is located at the upper end of the rear support rod 5.
[0041] In this embodiment of the invention, a robot-integrated dual-armrest chair-type support structure is provided. This structure includes two symmetrically arranged upper armrest frames, an embedded soft backrest, and a flip-up central support plate 13. The overall design adopts an ergonomic hand-grip angle. When the user is standing or walking, this structure serves as a psychological safety armrest, enhancing orientation and stability. When the user is fatigued or needs to pause briefly, the central support plate 13 can be flipped down to form a temporary seat, enabling switching between sitting and lying down. The entire support is constructed of lightweight, high-strength composite materials and uses a detachable linkage mechanism with the robot chassis, exhibiting high modularity and replaceability.
[0042] In this embodiment of the invention, the chassis drive structure can be replaced with a tracked or omnidirectional wheel chassis to adapt to different terrain requirements.
[0043] In one embodiment of the present invention, after the robot is started, the following operation process is performed:
[0044] 1. Map building and environmental perception: The LiDAR 16 and IMU (Inertial Measurement Unit) acquire pose data in real time; the Karto-SLAM algorithm is used to build an environmental map as the basis for navigation; the depth camera 10 acquires images of the front and sends them to the edge computing unit and cloud computing module for content recognition.
[0045] 2. Image Recognition and Voice Broadcasting: Image information is preprocessed by the edge computing unit and then uploaded to the cloud computing module, calling the AppBuilder visual recognition API; the recognition result is converted into auditory language text through a large language model (such as Baidu Qianfan); the text is converted into speech by the TTS speech synthesis model and output by the speaker module 9; the voice prompts include spatial descriptions (such as "the display board is three meters ahead, and the vase is on the left") and safety guidance.
[0046] 3. Navigation and obstacle avoidance control: The user sets the target point via voice; the oROS system calls the Dijkstra algorithm for global path planning and the TEB algorithm for local obstacle avoidance; the main control module controls the drive motor and servo motor to work together according to the navigation instructions to complete forward movement and obstacle avoidance.
[0047] 4. User operation and interaction: The user controls the forward direction or starts and stops through the handle 8; emergency braking can be completed through the handbrake 11; when a rest is needed, the user can sit on the seat cushion 12, with the backrest 6 providing support.
[0048] It can be seen that the auxiliary robot disclosed in this embodiment has the following advantages:
[0049] Edge-cloud collaborative structure: By combining local lightweight control with large model processing in the cloud, the local hardware computing power requirements are reduced and the recognition capability is improved;
[0050] Dual-purpose structure for sitting and walking: The unique foldable support structure can be used for standing assistance and temporary rest, making it especially suitable for visually impaired elderly people;
[0051] Multimodal perception and interaction: Combining image recognition, voice broadcasting, and LiDAR16, to create a humanized, safe, and efficient travel experience;
[0052] Highly modular and maintainable: The structure and control system are highly separated, which facilitates later maintenance, upgrades or adaptation to other scenarios (such as hospitals and airports).
[0053] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0054] Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A large-scale model-based visually impaired assistive robot system based on edge-cloud collaboration, characterized in that, It includes a walking component, a skeleton support, a sitting / reclining structure, and an electrical control system; the walking component is located below the skeleton support and is used to enable the robot to move and turn; the sitting / reclining structure is located in the middle of the skeleton support and is used to provide a seating function; the electrical control system is used to provide sensing and control functions.
2. The large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 1, characterized in that, The electrical control system includes a main control module, a drive module, a steering module, a sensing module, an audio module (9), and a cloud computing module; among which, the drive module and the steering module are used to control the movement and steering of the walking components, the audio module (9) is used to play voice prompts, the sensing module is used to collect environmental data, and the cloud computing module is used to identify environmental information and generate voice prompts.
3. The large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 2, characterized in that, The main control module is located inside the main control box (4), which is located below the sitting and lying structure. The main control module includes a main controller, a lithium battery, a power management unit, a wireless communication unit, and an edge computing unit. The edge computing unit integrates a local path planning algorithm.
4. The large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 2, characterized in that, The drive module is a DC motor connected to the walking component, and the main control module controls the speed and direction of the DC motor through PWM signals; the steering module is a servo motor connected to the walking component, used to control the left and right steering of the walking component.
5. A large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 2, characterized in that, The perception module includes a depth camera (10), a lidar (16), and a microphone mounted on the skeleton support. The depth camera (10) collects three-dimensional information of the environment, the lidar (16) scans the surrounding environment in real time, and the microphone collects voice information.
6. A large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 2, characterized in that, The cloud computing module is equipped with a large language model, an image recognition module, and a TTS speech generation model. The image recognition module recognizes image information, and the recognition results are converted into auditory text by the large language model. The auditory text is then converted into prompt speech by the TTS speech synthesis model.
7. A large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 1, characterized in that, The frame support includes a front support rod (14), a rear support rod (5), and a telescopic rod (3), which together form a triangular mechanical support.
8. A large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 7, characterized in that, The top of the rear support rod (5) is provided with a handle connector (7) and a handle (8) to form a two-hand grip structure.
9. A large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 7, characterized in that, The walking assembly includes a drive wheel (1), a steering mechanism (2) and a front wheel (17). The drive wheel (1) is controlled by the steering mechanism (2) and the steering mechanism (2) is connected to the telescopic rod (3). The front wheel (17) is connected to the front support rod (14).
10. A large-scale visually impaired assistive robot system based on edge-cloud collaboration according to claim 7, characterized in that, The sitting / lying structure includes a support plate (13), a seat cushion (12), a backrest (6), and a footrest (15). The footrest (15) is located at the lower end of the front support rod (14), the support plate (13) is located in the middle of the front support rod (14) and the rear support rod (5), the seat cushion (12) is located on the support plate (13), and the backrest (6) is located at the upper end of the rear support rod (5).
Citation Information
Cited By
Human-machine collaborative navigation system and method based on visual tactile sense-language-action model
CN122108167A