Blind guiding robot system based on end-to-end large model
The guide robot system, which combines an end-to-end large model with low-cost hardware, solves the problems of high cost and poor adaptability to dynamic environments of traditional guide equipment, realizes efficient and low-cost autonomous navigation and natural interaction, and improves the popularity of guide products and user experience.
Patent Information
- Application Number
- CN202510981897.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-05
AI Technical Summary
Traditional guide devices are costly, rely on complex hardware, have poor adaptability to dynamic environments, have low efficiency in multi-module collaboration, and provide a poor user experience.
A guide robot system based on an end-to-end large model, combined with low-cost hardware, integrates multimodal input through several low-cost RGB cameras and an end-to-end large model (GRNM), achieves autonomous navigation, integrates touch vibration feedback and voice interaction, simplifies the architecture and improves response speed.
Significantly reduce hardware costs, achieve autonomous obstacle avoidance and path optimization in dynamic environments, improve user experience and navigation success rate, enhance the natural perception and interaction of blind users, and promote the popularization of smart guide products.
Smart Images

Figure CN120588232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent robots, and in particular to a blind-guiding robot system based on an end-to-end large model. Background Art
[0002] Traditional guidance devices, such as guide dogs and LiDAR-guided robots, suffer from high costs, reliance on pre-built maps, complex sensors, and poor adaptability to dynamic environments. Existing navigation systems based on depth cameras or LiDAR are characterized by high hardware costs, and traditional algorithms require the coordination of multiple modules, such as positioning, planning, and control, resulting in response delays and high system complexity. Furthermore, wearable devices suffer from short battery life, and mechanical guidance devices lack flexibility.
[0003] In addition, existing equipment is expensive, relies on complex hardware, has poor map pre-construction and dynamic adaptability, and has low efficiency in multi-module collaboration and a poor user experience. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention provides a guide robot system based on an end-to-end large model. The present invention integrates multimodal input through the end-to-end large model and combines it with low-cost hardware to achieve efficient and low-cost autonomous navigation.
[0005] The present invention achieves the above technical objectives through the following technical means.
[0006] The guide robot system based on an end-to-end large model includes a hand grip, a telescopic rod and a robot main structure. A touch vibration sensor is placed on the inside of the hand grip, which vibrates when an obstacle is detected and accelerates the vibration frequency as the obstacle approaches. A high-definition camera is set below the hand grip to indicate important objects in the current environment and recognize the faces of people communicating at close range. One end of the telescopic rod is connected to the hand grip and the other end is connected to the robot main structure. The telescopic rod can be extended and retracted to accommodate different users. The robot main structure adopts a two-wheeled balanced dismounted drive structure and is equipped with a high-torque motor and a gyroscope. Several cameras are set on the outside of the robot main structure to cover a 360-degree field of view for real-time environmental image acquisition. The main control chip board is embedded in the robot main structure to support real-time multimodal data processing.
[0007] In the above solution, there are four cameras, including a left navigation camera, a right navigation camera, a front left navigation camera and a front right navigation camera, which are respectively arranged on the left and right sides and the middle of the front side of the robot body structure.
[0008] In the above solution, a microphone array is provided on the hand grip, and the high-definition camera is an AI camera.
[0009] The above solution also includes a power supply system with a built-in high-density lithium battery pack.
[0010] In the above solution, the main control chip board has a built-in end-to-end navigation model (GRNM), adaptive training and deployment, and human-computer interaction system.
[0011] In the above scheme, the end-to-end navigation model (GRNM) includes an input layer, a core processing module and an output layer; the input layer receives four-camera images, voice commands and GPS data, and generates a multimodal embedding vector through the feature extraction layer; the core processing module includes a visual language module, a bird's-eye view generation module and a diffusion strategy module; the visual language module is used to fuse images and voice commands and parse navigation targets, and the bird's-eye view generation module is used to convert multi-view images into a top-down semantic map, marking obstacles, feasible areas and target locations; the diffusion strategy module generates navigation action sequences based on the bird's-eye view, optimizes the path by occupying the network, and avoids dynamic obstacles; the output layer generates robot control instructions and environment description voice prompts, and provides real-time feedback to the user.
[0012] In the above solution, the adaptive training and deployment includes support for zero-shot cross-platform deployment, automatic data collection and Transformer architecture.
[0013] In the above solution, the human-computer interaction system includes voice interaction support for natural language recognition, a vibration feedback module and cloud data management.
[0014] In the above solution, the vibration feedback module provides gradient tactile prompts according to the distance of the obstacle.
[0015] In the above solution, the cloud data management is used to anonymize and store navigation anomaly cases to achieve model iterative optimization.
[0016] Beneficial effects:
[0017] 1. This invention significantly reduces hardware costs by more than 80% by combining several low-cost RGB cameras with an end-to-end large model (GRNM).
[0018] 2. The present invention adopts map-free navigation technology and realizes autonomous obstacle avoidance and path optimization in dynamic environments through real-time semantic understanding and diffusion strategy modules.
[0019] 3. The present invention directly maps multimodal inputs such as images, voice, and GPS to navigation instructions through an end-to-end model, simplifying the architecture and improving response speed.
[0020] 4. Among existing guide products, wearable devices are not comfortable, headphone-type guide devices interfere with the blind person's hearing senses, and most products have limited interactive experience and a single voice feedback method. The present invention integrates traction-type guide, voice command recognition, vibration gradient prompts at the handle, and voice output of environmental descriptions to enhance the naturalness of the blind user's physical perception and language interaction.
[0021] 5. Through an end-to-end large model, the present invention enables users to conveniently control forward and backward movement and flexibly adjust the direction, thus realizing a revolutionary human-computer interaction method; at the same time, it realizes touch feedback interaction. Specifically, more than 200 sound effects are built-in and encoded, and each sound effect corresponds to a situation. For example, there is one sound effect when there is a pedestrian in front, another sound effect when there is a bicycle in front, and there is also a dedicated vibration sound effect when there is a red light in front. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a schematic diagram of a guide robot;
[0023] Figure 2 for Figure 1 Schematic diagram of the robot's main structure;
[0024] Figure 3 For end-to-end large models;
[0025] Figure 4 This is the navigation module architecture diagram;
[0026] Figure 5 This is a flowchart for multimodal data processing in the navigation system;
[0027] Figure 6 This is a schematic diagram of a guide robot with a camera installed in the robot's main structure;
[0028] Figure 7 This is a schematic diagram of a guide robot with two cameras installed on the robot's main structure;
[0029] Figure 8 This is a schematic diagram of a guide robot with three cameras installed on its main body.
[0030] Figure 9 This is a schematic diagram of a guide robot with six cameras installed in its main body structure.
[0031] Reference numerals:
[0032] 1-Microphone; 2-AI camera; 3-Touch and vibration sensor; 4-Left navigation camera; 5-Ultrasonic sensor; 6-Right navigation camera; 7-Front left navigation camera; 8-Front right navigation camera; 9-Left motor; 10-Right motor; 11-Battery pack; 12-Main control chip board. DETAILED DESCRIPTION
[0033] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0034] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "up", "down", "axial", "radial", "vertical", "horizontal", "inside", "outside" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In the present invention, unless otherwise clearly specified and limited, the terms "install", "connect", "connect", "fix" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0035] Combined with attachment Figure 1 and Figure 2 As shown, the robot consists of a handle, a telescopic rod, and a main body. The handle has a built-in touch vibration sensor 3 that vibrates when an obstacle is detected, and the vibration frequency increases as the obstacle approaches. A camera, or AI camera 2, is located at the bottom of the handle to indicate important objects in the current environment and recognize the faces of people in close communication. The telescopic rod is adjustable in length.
[0036] The robot's main structure adopts a two-wheeled balancing cart-style drive structure, equipped with a high-torque motor and gyroscope, supports dynamic balance control, and has a battery life of ≥6 hours. The body is integrated with four wide-angle RGB cameras: two in the front, one on the left, and one on the right, covering a 360-degree field of view for real-time environmental image acquisition. Equipped with a Jetson AGX Orin 64GB embedded computing module with a computing power of 275TOPS, it supports real-time multimodal data processing. It is equipped with a high-sensitivity microphone array, GPS module, and vibration feedback device for voice interaction and user tactile prompts. The power system supports fast charging technology, has a built-in high-density lithium battery pack, and supports hot-swappable replacement. Safety protection design includes an anti-collision buffer layer, an emergency brake button, and IP65-level waterproof and dustproof.
[0037] Combined with attachment Figure 6-9 As shown, the number of RGB cameras is set according to the actual situation, generally not less than 1. For example, if there is only one camera, it can be placed in the front. Two cameras can be placed in the front and back or left and right at the same time. The placement of three cameras can be one in the front or back and one in the left and right. If there are six cameras, there are two in the front or back, two on the left and two on the right. However, any setting can collect information about the external environment. In short, it is not limited to the above description.
[0038] The calculation module may also be other types of algorithm chips, which may be adjusted or replaced according to actual conditions, and the present invention is not limited to the above description.
[0039] Combined with attachment Figure 3-5 As shown, the software part of the present invention is set in the main control chip board 12.
[0040] Software
[0041] 1) End-to-end large model (GRNM)
[0042] Input layer: Receives four-camera images, voice commands, and GPS data, and generates multimodal embedding vectors through the feature extraction layer.
[0043] Core processing module:
[0044] Vision Language Module (VLM): fuses images and voice commands to parse navigation goals, such as "go to the bus stop."
[0045] Bird's-eye view generation (BEV): Convert multi-view images into a bird's-eye view semantic map, annotating obstacles, feasible areas, and target locations.
[0046] Diffusion strategy module: Generates navigation action sequences based on BEV, such as steering, acceleration, braking, etc., optimizes the path by occupying the network, and avoids dynamic obstacles.
[0047] Output layer: Generates robot control instructions, such as steering angle, speed, and voice prompts describing the environment, and provides real-time feedback to the user.
[0048] 2) Adaptive training and deployment
[0049] It supports zero-shot cross-platform deployment, and the model can be adapted to various robot architectures such as wheeled and legged robots.
[0050] The automatic data collection system generates a 2,000-hour navigation dataset, combining supervised learning with reinforcement learning to optimize the model's generalization capabilities.
[0051] It uses a Transformer architecture with 4 layers and 4 attention heads to achieve multimodal feature fusion. The loss function integrates motion prediction and spatiotemporal distance error:
[0052] L = MSE(ε k ,ε θ )+λMSE(d(ot,og),fd(ct))
[0053] Where λ = 10 -4 , optimizer is AdamW, learning rate is 10 -4 .
[0054] Combined with attachment Figure 5 As shown, 3) Human-computer interaction system
[0055] Voice interaction supports natural language recognition and dynamically adjusts navigation modes, such as exploratory and directional.
[0056] The vibration feedback module provides gradient tactile prompts based on the distance of obstacles, such as high-frequency vibrations indicating close obstacles.
[0057] Cloud data management: Anonymize and store navigation anomaly cases for model iteration and optimization.
[0058] Technical effects:
[0059] 1) High navigation success rate and high obstacle avoidance capabilities, improving the quality of travel for the blind
[0060] In real-world tests indoors and outdoors, the robot achieved navigation success rates of 80.2% indoors and 84.4% outdoors, respectively. Its long-distance navigation success rate was 82.8%, significantly outperforming traditional guide dogs or LiDAR robots. The robot's obstacle avoidance success rates were 78.2% indoors and 87.6% outdoors, respectively, effectively identifying dynamic obstacles such as pedestrians and cyclists, reducing collision risks.
[0061] 2) Reduce hardware costs and increase the social penetration rate of smart guide products
[0062] Hardware costs are reduced by replacing LiDAR with four RGB cameras, reducing total hardware costs by 80%, lowering costs for both users and the industry (the average annual cost of a guide dog is approximately $15,000). Annual sales are expected to reach 200,000 units within five years, reaching 700,000 visually impaired users worldwide and increasing the penetration rate to 3%. Currently, the penetration rate for guide dogs is only 0.3%. This will expand user coverage and enable widespread adoption and adoption of intelligent navigation solutions.
[0063] 3) Promote the construction of a barrier-free society
[0064] Through its friendly appearance, natural language interaction, vibration feedback and traction navigation functions, it enhances users' perception and trust in the environment (user satisfaction test reaches 90%), reduces dependence on others, reduces the strange looks from people around them caused by the tapping of the guide stick when traveling, and promotes the social integration of blind users.
[0065] 4) Collaborative development of the industrial chain
[0066] It will drive the demand for low-cost sensors, high-computing-power embedded chips such as Jetson AGX Orin, and end-to-end model optimization technologies, thereby driving the upgrade of the AI robot industry chain.
[0067] Analysis of the feasibility and beneficial effects of the present invention
[0068] Feasibility analysis:
[0069] First, hardware design feasibility:
[0070] Low-cost sensor replacement: Four wide-angle RGB cameras replace traditional LiDAR, reducing hardware costs by 80%. Field tests show that the RGB cameras, using the GRNM model, can accurately interpret environmental semantics, such as obstacles and feasible areas, with a 92% accuracy rate, meeting navigation requirements.
[0071] Two-wheeled chassis optimization: The self-balancing vehicle drive structure is combined with a high-density lithium battery, with a battery life of ≥6 hours. In complex road conditions such as slopes and gravel tests, the stability reaches 95%, which is significantly better than four-wheeled or legged robots. Among them, the competing guide robot dog has a battery life of only 2 hours.
[0072] Secondly, the feasibility of software algorithms:
[0073] Simplified end-to-end large-scale model architecture: The GRNM model directly maps multimodal inputs, such as images, speech, and GPS, to navigation commands, eliminating the need for pre-built maps or multi-module coordination. Experiments show that model inference time is as low as 50ms, compared to the 2-second or greater delay of traditional algorithms with multiple modules connected in series, resulting in a 40x improvement in efficiency.
[0074] Adaptability to dynamic environments: The diffusion strategy module combines the occupancy network with real-time path optimization. In tests with dynamic obstacles such as pedestrians and vehicles, the obstacle avoidance success rate was 87.6% and the collision rate was only 1.0, while the collision rate of traditional algorithms such as VIB was 4.0.
[0075] Cross-platform deployment capability: Through zero-shot adaptation technology, the GRNM model can be deployed on wheeled platforms, such as automatic wheelchairs, and foot-based platforms, such as household robots. The cross-platform navigation success rate is >85%, verifying the model's generalization capability.
[0076] Technical solution effect experiment
[0077] Experiment 1: Navigation success rate and path accuracy test
[0078] Experimental design:
[0079] Environment: Indoor and outdoor, such as office and shopping mall scenes, including static obstacles such as tables, chairs, and trees
[0080] Task: Set a target point, such as a bathroom, and require the robot to navigate to the target point with a position error of ≤0.5m and an angle deviation of ≤30°.
[0081] result:
[0082] Success rate: 80.2% indoors, of which 80 / 100 were valid tests; 84.4% outdoors, of which 84 / 10 were valid tests; 82.8% long-distance orientation tasks, of which 82 / 100 were valid tests.
[0083] Path deviation: Indoors, the average deviation is 7.3%, specifically 285.3m vs. 265m ideal path. Outdoors, the deviation is 23.5%, specifically 508.7m vs. 412m ideal path. This is mainly due to differences in environmental complexity.
[0084] Experiment 2: Dynamic obstacle avoidance performance test
[0085] Experimental design:
[0086] Scenario: Simulate a pedestrian suddenly crossing the road, for example, 10 times, a bicycle approaching, for example, 10 times, and an obstacle moving, for example, 10 times.
[0087] Metrics: obstacle avoidance success rate, such as no collision, and response time from obstacle detection to braking time.
[0088] result:
[0089] Obstacle avoidance success rate: 78.2% indoors, of which 7 / 9 were effective tests; 87.6% outdoors, of which 9 / 10 were effective tests.
[0090] Response time: average 0.3 seconds, fastest 0.2 seconds, significantly better than traditional algorithms. Traditional VIB averages 1.2 seconds.
[0091] Experiment 3: Low-light environment performance test
[0092] Experimental design:
[0093] Scene: Nighttime streets, light level <10 lux, underground parking lot.
[0094] Task: Navigate to the target point, for example, 10 times, and calculate the obstacle avoidance success rate and navigation deviation.
[0095] result:
[0096] Obstacle avoidance success rate: 60%-70%, a 20% decrease compared to daytime, mainly due to the camera's low-light performance limitations.
[0097] Improvement plan: After adding active infrared lighting, the success rate increased to 85% in preliminary tests.
[0098] Data Analysis and Conclusion
[0099] Navigation and obstacle avoidance capabilities: Navigation success rate in complex environments is >80%, and obstacle avoidance success rate is >78%, significantly better than traditional systems, such as Masked ViNT^m^, which has a success rate of 30%.
[0100] Economic feasibility verification: Hardware costs are reduced by 80%, and the average annual cost for users is only 1 / 75 of that of guide dogs, which has the potential for large-scale promotion.
[0101] Technical advantages: The end-to-end model has a simplified architecture, real-time response, latency ≤ 0.5 seconds, cross-platform adaptability, and a success rate > 85%, verifying the comprehensive feasibility of the technical solution.
[0102] Through systematic experiments and data analysis, the present invention has achieved the expected results in terms of navigation accuracy, dynamic adaptability, cost control and user experience. The experimental results are highly consistent with the theoretical design and have clear commercial and social application value.
[0103] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0104] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and purpose of the present invention.
Claims
1. The blind guide robot system based on the end-to-end large model is characterized by: It includes a hand grip, a telescopic rod and a robot main structure; among them, a touch vibration sensor is placed on the inside of the hand grip, which issues a vibration prompt when an obstacle is detected, and the vibration frequency increases as the obstacle approaches. A high-definition camera is set under the hand grip to prompt important objects in the current environment and recognize the faces of people communicating at close range; one end of the telescopic rod is connected to the hand grip, and the other end is connected to the robot main structure. The telescopic rod can be extended and retracted to adapt to different users; the robot main structure adopts a two-wheeled balanced dismounted drive structure and is equipped with a high-torque motor and a gyroscope. Several cameras are set on the outside of the robot main structure to cover a 360° field of view for real-time environmental image acquisition. The main control chip board is embedded in the robot main structure to support real-time multimodal data processing.
2. The blind-guiding robot system based on an end-to-end large model according to claim 1, characterized in that: There are four cameras, including a left navigation camera, a right navigation camera, a front left navigation camera and a front right navigation camera, which are respectively arranged on the left and right sides and the middle of the front side of the robot body structure.
3. The blind-guiding robot system based on an end-to-end large model according to claim 1, characterized in that: The handle is provided with a microphone array, and the high-definition camera is an AI camera.
4. The blind-guiding robot system based on an end-to-end large model according to claim 1, characterized in that: It also includes a power supply system with a built-in high-density lithium battery pack.
5. The blind-guiding robot system based on an end-to-end large model according to claim 1, characterized in that: The main control chip board has a built-in end-to-end navigation model (GRNM), adaptive training and deployment, and human-computer interaction system.
6. The blind-guiding robot system based on an end-to-end large model according to claim 5, characterized in that: The end-to-end navigation model (GRNM) includes an input layer, a core processing module, and an output layer; the input layer receives four-camera images, voice commands, and GPS data, and generates a multimodal embedding vector through a feature extraction layer; the core processing module includes a visual language module, a bird's-eye view generation module, and a diffusion strategy module; the visual language module is used to fuse images and voice commands and parse navigation targets, and the bird's-eye view generation module is used to convert multi-view images into a top-down semantic map, annotating obstacles, feasible areas, and target locations; the diffusion strategy module generates navigation action sequences based on the bird's-eye view, optimizes paths by occupying the network, and avoids dynamic obstacles; the output layer generates robot control instructions and environment description voice prompts, which are fed back to the user in real time.
7. The blind-guiding robot system based on an end-to-end large model according to claim 5, characterized in that: The adaptive training and deployment includes support for zero-shot cross-platform deployment, automatic data collection and Transformer architecture.
8. The blind-guiding robot system based on an end-to-end large model according to claim 5, characterized in that: The human-computer interaction system includes voice interaction support, natural language recognition, vibration feedback module and cloud data management.
9. The blind-guiding robot system based on an end-to-end large model according to claim 8, characterized in that: The vibration feedback module provides a gradient tactile prompt according to the distance of the obstacle.
10. The blind-guiding robot system based on an end-to-end large model according to claim 8, characterized in that: The cloud data management is used to anonymize and store navigation anomaly cases to achieve iterative model optimization.