Multi-mode sensing indoor blind person navigation system based on mixed Agent
By integrating LiDAR, camera, and IMU sensors into a multimodal perception system based on hybrid agents, and combining it with SLAM algorithms for environmental perception and decision-making, the system solves the problems of incomplete perception, unrobust decision-making, and unfriendly user interaction in indoor navigation for the blind. It achieves high-precision navigation and safe obstacle avoidance, and improves the ability of the blind to travel independently in complex environments.
Patent Information
- Application Number
- CN202511146243.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing indoor navigation systems for the blind suffer from incomplete perception, unreliable decision-making, inaccurate navigation, and unfriendly user interaction in complex environments, making it difficult to effectively cope with dynamic obstacles and the lack of stable external positioning signals.
A multimodal perception system based on hybrid agents is adopted, integrating LiDAR, camera and IMU sensors, combining SLAM algorithm for environmental perception and localization, using hybrid agent architecture for intelligent decision making and path planning, and providing navigation instructions and environmental feedback through multimodal interaction methods such as voice and bone conduction.
It significantly improves navigation safety and accuracy, enhances the ability of blind people to travel independently in indoor environments, reduces reliance on external infrastructure, and optimizes user experience and navigation efficiency.
Smart Images

Figure CN120970653A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a navigation system, more particularly to a multi-modal perception indoor blind navigation system based on hybrid Agent. BACKGROUND
[0002] With the acceleration of urbanization and the increasing complexity of building structures, blind people face severe challenges in independent and safe travel in indoor environments. Traditional guide tools (such as walking sticks, guide dogs) have limited perception ability and navigation accuracy in complex indoor environments, making it difficult to effectively deal with dynamic obstacles, complex layouts, and lack of stable external positioning signals (such as GNSS). Existing indoor navigation systems often rely on a single technology or lack deep adaptation to the needs of blind users, such as insufficient positioning accuracy, limited obstacle avoidance ability, unnatural user interaction, or information overload. SUMMARY
[0003] In view of the deficiencies in the prior art, the purpose of the present application is to provide a multi-modal perception indoor blind navigation system based on hybrid Agent, aiming to solve the technical problems of existing blind indoor navigation methods in complex indoor environments, such as incomplete perception, non-robust decision-making, inaccurate navigation, and user-unfriendly interaction.
[0004] To achieve the above-mentioned purpose, the present application provides the following technical solution: a multi-modal perception indoor blind navigation system based on hybrid Agent, comprising: a sensor module integrated with multiple sensors for collecting environmental data and motion data; a perception module in communication with the sensor module for processing sensor raw data and performing environmental perception, including positioning, mapping, obstacle detection and identification, and landmark identification; a decision module in communication with the perception module for intelligent decision-making based on information provided by the perception module, combining navigation targets and environmental uncertainties, including selecting navigation strategies, assessing risks, and handling unexpected situations; a planning module in communication with the decision module for path planning based on instructions from the decision module and environmental information, including global path planning, local obstacle avoidance planning, and landmark-based path adjustment; a user interaction module for information interaction with the user, receiving user instructions, and providing navigation instructions and environmental feedback to the user through voice and bone conduction; a memory module storing environmental maps, landmark information, user preferences, and historical navigation data; a coordination module in communication with each of the above modules for managing information flow and collaboration between modules to ensure efficient and coherent operation of the system.
[0005] As a further improvement of the present invention, the sensor module includes: LiDAR is used to provide accurate depth information and environmental point clouds for building high-precision indoor maps and detecting obstacles. Cameras are used to provide rich texture and semantic information for recognizing landmarks, detecting obstacles, and identifying traffic lights; The IMU is used to provide motion information of the device, perform dead reckoning, provide continuous positioning information in a short period of time, and assist the SLAM algorithm to reduce drift accumulation.
[0006] As a further improvement of the present invention, the specific method by which the perception module processes the raw sensor data is as follows: Real-time localization and mapping are performed using the fused sensor data; a graph-optimized SLAM algorithm is adopted to fuse LiDAR point clouds, visual features, and IMU data to construct an accurate indoor map and estimate device pose; obstacles in the environment, including static and dynamic obstacles, are detected and identified in real time using LiDAR point clouds and camera images, combined with deep learning algorithms; camera images are processed using CNN- or Transformer-based object detection models, and 3D obstacle localization and size estimation are performed using LiDAR depth information; predefined or automatically selected environmental landmarks are identified using camera images, combined with feature matching or deep learning methods; while constructing a geometric map, a semantic map containing landmarks, room types, and accessibility facilities is constructed using the semantic information of the camera images, providing the Agent with a higher level of environmental understanding.
[0007] As a further improvement of the present invention, the decision module includes: Reactive components are used to handle real-time, low-level tasks; The thoughtful component is responsible for high-level planning and decision-making, specifically map-based global path planning, task decomposition, and strategy selection. This component uses the environment map and landmark information in the memory module to plan the path from the current location to the target. Learning components continuously optimize the agent's behavioral strategies by interacting with the environment and receiving user feedback; The coordination module integrates the above-mentioned different types of components, so that the deliberate component plans the global path, the reactive component is responsible for local obstacle avoidance, and the learning component adjusts the strategy according to the obstacle avoidance results.
[0008] As a further improvement of the present invention, the deliberate component performs global path planning based on a map in the following way: using a high-precision indoor map constructed by the perception module, the optimal path from the current location to the target location is calculated using path planning algorithms such as A* and D*. While performing path planning, the reactive component uses obstacle information detected in real time by the perception module to perform local obstacle avoidance. The deliberate component also uses landmark information identified by the perception module to assist in positioning and path adjustment.
[0009] As a further improvement of the present invention, the user interaction module includes: The voice interaction module uses speech recognition technology to receive the user's navigation target and query instructions, and also uses speech synthesis technology to convert navigation instructions, environmental information and warnings into voice output. Bone conduction feedback module, which transmits voice information directly to the user through bone conduction headphones; The multimodal feedback module combines feedback from multiple modalities, including speech and bone conduction, to provide richer and more intuitive information. The adaptive and personalized module adjusts the volume, speech rate, and level of detail of voice prompts, as well as the intensity and mode of bone conduction feedback, based on the user's hearing level, walking speed, and usage habits. The environmental information feedback module is used to provide environmental information to users.
[0010] As a further improvement of the present invention, the decision-making module makes decisions in uncertain states as follows: Risk quantification and dissemination involves quantifying uncertainties in perceived information, positioning results, and environmental predictions, and tracking the propagation of uncertainty in the decision-making and planning process. Decisions should be made based on uncertainty. When the positioning uncertainty is high, the speed should be reduced, the environmental scanning frequency should be increased, or known landmarks should be prioritized for positioning correction. Maintain a safety margin by considering uncertainties and increasing the safety margin in path planning and obstacle avoidance; Emergency handling: Design response mechanisms for emergency situations.
[0011] The beneficial effects of this invention are: 1. Significantly improve navigation safety: By integrating multimodal sensor data such as LiDAR, cameras and IMU, combined with robust obstacle detection and recognition algorithms and risk assessment-based decision models, the system can perceive the environment more comprehensively and accurately, effectively identify static and dynamic obstacles, and make safe obstacle avoidance decisions under uncertainty, greatly reducing the collision risk for blind people in indoor environments.
[0012] 2. Improve indoor positioning and mapping accuracy: By integrating data from multiple sensors and employing advanced SLAM algorithms, the limitations of a single sensor in indoor environments are overcome, achieving high-precision indoor positioning and environmental map construction, laying the foundation for accurate navigation.
[0013] 3. Enhanced navigation robustness: Multi-sensor fusion and hybrid agent architecture enable the system to cope with complex indoor environmental changes, sensor data quality degradation and unexpected situations, thus improving the robustness and reliability of navigation.
[0014] 4. Optimized User Experience: Through multimodal and adaptive user interaction methods such as voice and bone conduction, the system can provide users with navigation instructions and environmental information in a natural and intuitive way, reducing the user's cognitive load and improving navigation efficiency and comfort. Personalized settings and learning capabilities further enhance user satisfaction.
[0015] 5. Facilitating Independent Travel for the Blind: The safe, accurate, robust, and user-friendly navigation service provided by this invention can significantly enhance the confidence and ability of blind people to travel independently in indoor environments, expand their activity range, and promote their better integration into society.
[0016] Reduced reliance on external infrastructure: Although infrastructure such as Bluetooth beacons can be used to assist in positioning, the system mainly relies on the device's built-in sensors for environmental perception and positioning, which reduces reliance on external deployments and improves the system's versatility. Attached Figure Description
[0017] Figure 1 This is a system architecture diagram of the multimodal perception indoor navigation system for the blind based on hybrid agents according to the present invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the embodiments shown in the accompanying drawings.
[0019] Reference Figure 1 As shown in this embodiment, a multimodal perception indoor navigation system for the blind based on a hybrid agent adopts a modular design. The core is a hybrid agent, whose architecture can be viewed as a complex intelligent agent containing multiple specialized sub-modules or sub-agents that collaboratively complete navigation tasks. Specifically, it includes the following modules: Sensor module: Integrates sensors such as LiDAR, camera and IMU, responsible for collecting environmental data and motion data.
[0020] Perception Module (PerceptionAgent): Processes raw sensor data to perform environmental perception, including localization, mapping, obstacle detection and recognition, landmark recognition, etc.
[0021] DecisionAgent: Based on the information provided by the perception module, and combined with the navigation target and environmental uncertainties, it makes intelligent decisions, such as selecting navigation strategies, assessing risks, and handling emergencies.
[0022] Planning Agent: Based on the instructions from the decision-making module and environmental information, it performs path planning, including global path planning, local obstacle avoidance planning, and landmark-based path adjustment.
[0023] The user interaction module (InteractionAgent) is responsible for interacting with users, receiving user commands, and providing navigation instructions and environmental feedback to users through voice and bone conduction.
[0024] Memory module: Stores environment maps, landmark information, user preferences, and historical navigation data.
[0025] Orchestration module: Manages the information flow and collaboration between modules to ensure efficient and consistent system operation.
[0026] Furthermore, in the aforementioned perception module, the system integrates data from LiDAR, camera, and IMU to overcome the limitations of a single sensor and improve the accuracy and robustness of environmental perception. Therefore, the perception module in this embodiment further includes: LiDAR: Provides accurate depth information and environmental point clouds for building high-precision indoor maps (such as multi-resolution raster maps) and obstacle detection. LiDAR is unaffected by lighting conditions, but may perform poorly in low-feature environments (such as long corridors).
[0027] Cameras provide rich texture and semantic information for landmark recognition, obstacle detection (especially dynamic and high-altitude obstacles), and traffic light identification. Cameras are significantly affected by lighting conditions.
[0028] IMU (In-memory Unit): Provides motion information (acceleration, angular velocity) for dead reckoning, offering continuous positioning information over a short period and assisting SLAM algorithms to reduce drift accumulation. Fusion of IMU data with visual or LiDAR data (such as VIO-SLAM) can improve positioning robustness, especially when sensor data quality deteriorates.
[0029] The aforementioned sensor fusion is achieved using a variety of technologies, including the following: SLAM (Simultaneous Localization and Mapping): Real-time localization and mapping using fused sensor data. For example, graph-optimized SLAM algorithms can be used to fuse LiDAR point clouds, visual features, and IMU data to build accurate indoor maps and estimate device pose.
[0030] Obstacle detection and recognition: Utilizing LiDAR point clouds and camera images, combined with deep learning algorithms, this system detects and recognizes obstacles in the environment in real time, including static obstacles (walls, furniture) and dynamic obstacles (pedestrians, moving vehicles). For example, CNN- or Transformer-based object detection models can be used to process camera images, combined with LiDAR depth information for 3D obstacle localization and size estimation.
[0031] Landmark recognition and localization: Utilizing camera images and combining feature matching or deep learning methods, this feature identifies predefined or automatically selected environmental landmarks (such as doorways, elevators, and specific signs). Landmark information is used to assist in localization and navigation.
[0032] Semantic map construction: While building a geometric map, semantic information from camera images is used to construct a semantic map containing information such as landmarks, room types, and accessibility facilities, providing the agent with a higher level of environmental understanding.
[0033] Furthermore, the Agent in the aforementioned decision-making module adopts a hybrid architecture, combining the characteristics of reactive, deliberate, and learning agents to address the complexity and dynamism of indoor environments, as detailed below: Reactive components: These handle real-time, low-level tasks, such as rapid obstacle avoidance based on sensor data. When sensors detect nearby obstacles, reactive components can quickly trigger obstacle avoidance behavior to ensure safety.
[0034] The deliberate component is responsible for high-level planning and decision-making, such as map-based global path planning, task decomposition, and strategy selection. This component utilizes the environment map and landmark information in its memory module to plan the path from the current location to the target.
[0035] Learning-based components: By interacting with the environment and receiving user feedback, the agent continuously optimizes its behavioral strategies, such as improving obstacle avoidance strategies, adapting to user preferences, and enhancing the robustness of decision-making. Reinforcement learning and other methods can be used to train the agent to make optimal decisions in complex scenarios.
[0036] Hybridization and Coordination: The coordination module is responsible for integrating these different types of components. For example, the deliberate component plans the global path, the reactive component handles local obstacle avoidance, and the learning component adjusts the strategy based on the obstacle avoidance results. The decision-making module selects and switches between different strategies, such as using map-based navigation in open areas and switching to real-time obstacle avoidance and landmark navigation in congested areas.
[0037] In the above decision-making process, the agent considers environmental uncertainty and security, as follows: Uncertainty handling: Sensor fusion and filtering techniques (such as Kalman filtering and particle filtering) are used to reduce uncertainty in localization and perception. The decision-making module can employ probabilistic models (such as Bayesian networks) to infer environmental states and predict the behavior of dynamic obstacles, thereby making more robust decisions under uncertainty.
[0038] Risk assessment: The decision-making module assesses risks during navigation in real time, such as collision risk and the risk of getting lost. Risk assessment can be based on sensor data, map information, and predictive models.
[0039] Safety First: The decision-making process prioritizes safety. For example, it prioritizes accessible routes during path planning, ensures a safe distance from obstacles during obstacle avoidance, and reduces navigation speed or provides more frequent warnings when uncertainty is high. Safety constraints can be introduced into path planning and decision-making algorithms.
[0040] Furthermore, the decision-making module integrates map-based path planning, real-time obstacle avoidance, and landmark-based navigation strategies, as detailed below: Map-based path planning: Utilizing a high-precision indoor map constructed using a perception module, path planning algorithms such as A* and D* are employed to calculate the optimal path from the current location to the target location. The map can contain geometric and semantic information, such as marking accessible facilities and hazardous areas.
[0041] Real-time obstacle avoidance: While performing path planning, obstacle information detected in real time by the perception module is used for local obstacle avoidance. Dynamic window method (DWA), vector field histogram (VFH) or deep learning-based obstacle avoidance algorithms can be used to quickly adjust the direction of movement and avoid static and dynamic obstacles.
[0042] Landmark-based navigation utilizes landmark information identified by the perception module to assist in positioning and path adjustment. Landmarks serve as reference points during navigation, helping users build a cognitive map of the environment and correcting for positioning drift. The agent can adjust the local path based on landmark information, guiding the user towards the next landmark. In this embodiment, the above strategies do not operate independently but are integrated and coordinated. For example, global path planning provides the general direction, real-time obstacle avoidance handles local obstacles, and landmark navigation provides assistance and verification at key nodes. The decision module dynamically adjusts the weights and priorities of different strategies based on the current environmental state and task requirements.
[0043] Furthermore, this embodiment provides the following user interaction module, which is responsible for receiving the user's voice commands and providing navigation instructions and environmental feedback to the user through voice and bone conduction. This module mainly includes the following modules: Voice interaction: This feature utilizes speech recognition technology to receive user navigation goals and query commands. Speech synthesis technology converts these commands, environmental information, and warnings into voice output. Natural language interaction can be achieved using intelligent voice modules from companies like iFlytek.
[0044] Bone conduction feedback: Bone conduction headphones transmit voice information directly to the user without occupying the auditory channel, allowing the user to perceive ambient sounds simultaneously, thus improving safety. Bone conduction can also be used to provide non-voice feedback, such as using different vibration patterns or frequencies to indicate direction, distance, or danger level.
[0045] Multimodal feedback: Combining feedback from multiple modalities such as voice and bone conduction provides richer and more intuitive information. For example, voice prompts indicate direction, while bone conduction vibrations indicate the distance to obstacles or the level of danger.
[0046] Adaptive and Personalized: The system can adjust the volume, speed, and level of detail of voice prompts, as well as the intensity and pattern of bone conduction feedback, based on factors such as the user's hearing level, walking speed, and usage habits. The system can learn user preferences to provide more personalized navigation services.
[0047] Environmental information feedback: In addition to navigation instructions, the system can also provide users with environmental information, such as the type and distance of obstacles ahead, nearby landmarks, the room or area they are currently in, etc., to help users better understand their surroundings.
[0048] In indoor environments, sensor data, positioning, and environmental conditions all present uncertainties. The agent's decision-making module needs to pay particular attention to safety under uncertainty; specifically, the decision-making module should employ the following decision-making methods under uncertain conditions: Risk Quantification and Propagation: Quantifying uncertainties in perceived information, positioning results, and environmental predictions, and tracking the propagation of uncertainty in the decision-making and planning process.
[0049] Uncertainty-based decision-making: The decision-making module can adjust its strategy based on the level of uncertainty. For example, when the location uncertainty is high, the agent may reduce its speed, increase the frequency of environmental scanning, or prioritize finding known landmarks for location correction.
[0050] Safety margin: In path planning and obstacle avoidance, uncertainties are considered, and a safety margin is increased. For example, maintaining a greater distance from obstacles and choosing a wider passage.
[0051] Emergency handling: Design response mechanisms for emergencies (such as sensor failure, sudden danger), such as triggering an emergency stop, issuing high-priority warnings, or guiding users to a safe area.
[0052] User Trust and Feedback: Clear and timely feedback allows users to understand the system's status and uncertainty level, such as prompts like "Location signal is weak, please be aware of your surroundings." Simultaneously, collecting user feedback allows for continuous optimization of uncertainty handling and security strategies, leading to the following formula: P(S_t|O_{1:t},A_{1:t-1})\proptoP(O_t|S_t)\sum_{S_{t-1}}P(S_t|S_{t-1},A_{t-1})P(S_{t-1}|O_{1:t-1},A_{1:t-2}) The above formula represents the agent's belief update process for the current state S_t in a partially observable Markov decision process (POMDP), where O_{1:t} is the observation sequence up to time t, A_{1:t-1} is the action sequence up to time t-1, P(O_t|S_t) is the observation model, and P(S_t|S_{t-1},A_{t-1}) is the state transition model. Although this embodiment may not directly adopt the complete POMDP framework, its decision-making process draws on the idea of state estimation and decision-making under uncertainty.
[0053] Ultimately, the system in this embodiment can be deployed on glasses for the blind or wearable devices that work with smartphones. Considering the resource limitations of the devices (computing power, memory, power consumption), the algorithm needs to be optimized. For example, lightweight model techniques (knowledge distillation, quantization) can be used to compress deep learning models. The development of this system can be based on frameworks such as the Robot Operating System (ROS), utilizing its provided sensor drivers, SLAM libraries, navigation stacks, and other tools. During system development, a large amount of indoor environmental data (LiDAR point clouds, images, IMU data) needs to be collected for training map building, obstacle recognition, and landmark recognition models. Large-scale training and testing can be conducted in simulated environments to reduce the risks of real-world testing. Continuous user research is also necessary, inviting blind users to participate in usability testing and evaluation, collecting feedback, and continuously optimizing user experience and safety. Evaluation metrics should include navigation accuracy, obstacle avoidance success rate, number of collisions, navigation time, user satisfaction, and cognitive load.
[0054] Based on the system provided in the above embodiments, this embodiment provides the following three specific examples: Example 1: Navigating to a specific store within a large shopping mall 1. Scenario: A blind user enters an unfamiliar multi-story shopping mall and wants to be navigated to a specific clothing store on the third floor.
[0055] 2. Process: ○ The user inputs the target into the system via voice command: "Navigate to the [clothing store name] on the third floor".
[0056] The system uses a pre-built or real-time built indoor map of the shopping mall, combined with the current location (obtained through sensor fusion SLAM), to plan an optimal path, including taking an elevator or escalator.
[0057] During navigation, the system continuously integrates LiDAR, camera, and IMU data to detect dynamic obstacles such as pedestrians, shopping carts, and temporarily stacked goods in real time.
[0058] When an obstacle is detected, the decision-making module assesses the risk, and the planning module calculates the obstacle avoidance path. The user interaction module guides the user to avoid the obstacle through bone conduction vibration and voice prompts (e.g., "There is a pedestrian ahead, please adjust slightly to the left").
[0059] The system identifies landmarks within the shopping mall (such as elevator entrances, floor signs, and specific shop signs) to help users confirm their current location and direction, and informs users through voice prompts (e.g., "You have arrived at the elevator entrance").
[0060] ○ When riding the elevator, the system recognizes the floor buttons and prompts the user to select the correct floor via voice.
[0061] ○ After reaching the third floor, the system continues to guide the user to the target store, identifies the landmark at the store entrance, and finally prompts the user that they have arrived at their destination.
[0062] 3. Results: Users can safely and independently find their target stores in complex shopping mall environments without relying on others. Real-time obstacle avoidance effectively prevents collisions with dynamic obstacles. Landmark navigation helps users develop an understanding of the mall's layout.
[0063] Example 2: Navigating from your office to a meeting room within an office building, avoiding temporarily placed cleaning equipment. 1. Scenario: A blind user is in a familiar office building, heading from their office to a pre-booked meeting room. A cleaning machine has been temporarily placed in the corridor by cleaning staff.
[0064] 2. Process: ○ Users set their destination via voice command: "Go to [meeting room name]".
[0065] The system uses the office building's indoor map to plan routes.
[0066] ○ As users walk along the corridor, the sensing module detects temporarily placed cleaning equipment using LiDAR and cameras.
[0067] The decision-making module determines that the cleaning equipment is an obstacle, and the planning module calculates the detour path.
[0068] The user interaction module guides users safely through obstacles by providing voice and bone conduction prompts (e.g., "There is an obstacle ahead, please go around to the right").
[0069] The system can identify landmarks such as office doorplates and meeting room doorplates to help users confirm their location.
[0070] 3. Results: The system can identify temporary obstacles not present on the map and guide users to avoid them safely, demonstrating its ability to avoid obstacles in real time and adapt to dynamic environmental changes. Users were able to reach the conference room smoothly without being affected by temporary obstacles.
[0071] Example 3: Navigation in areas with poor indoor positioning signals 1. Scenario: A blind user is walking in the underground area of a large building. The GPS signal is completely lost and the indoor Bluetooth beacon coverage is sparse, resulting in a decrease in positioning accuracy.
[0072] 2. Process: The system primarily relies on LiDAR, camera, and IMU data for SLAM localization.
[0073] ○ The decision module detected an increase in positioning uncertainty (e.g., an increase in the covariance matrix of the SLAM algorithm).
[0074] The decision-making module adjusts its strategy, reducing navigation speed, increasing environmental scanning frequency, and prioritizing the search for known landmarks for positioning correction.
[0075] ○ The user interaction module prompts the user that the current location signal may be poor and suggests walking slowly and cautiously (e.g., "Location signal is weak, please watch your step and the surrounding environment").
[0076] When the perception module identifies a known landmark (such as a specific pillar, wall feature, or sign), it uses the landmark information to correct the positioning and reduce cumulative errors.
[0077] The planning module uses location information with uncertainties to plan more conservative routes and avoid potentially dangerous areas.
[0078] 3. Results: Even in challenging environments with poor positioning signals, the system can maintain a certain level of navigation capability and safety through multi-sensor fusion, uncertainty handling, and landmark assistance, preventing users from becoming completely lost or encountering danger.
[0079] These embodiments demonstrate how the present invention, in different indoor scenarios, utilizes hybrid agents, multimodal perception, fusion navigation strategies, and user interaction to provide safe, independent, and effective navigation services for the blind. In summary, the core of the hybrid agent-based multimodal perception indoor navigation system for the blind in this embodiment lies in constructing a hybrid agent capable of sensing the environment, making intelligent decisions, planning safe routes, and interacting naturally with the user. This agent integrates data from multiple sensors and combines map-based, real-time obstacle avoidance, and landmark-based navigation strategies.
[0080] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A multimodal perception indoor navigation system for the blind based on a hybrid agent, characterized in that: include: The sensor module integrates multiple sensors and is responsible for collecting environmental and motion data. The perception module communicates with the sensor module and is used to process raw sensor data to perform environmental perception, including localization, mapping, obstacle detection and recognition, and landmark recognition.
2. Decision module, which communicates with the perception module, makes intelligent decisions based on the information provided by the perception module, combined with the navigation target and environmental uncertainties. The decision includes selecting a navigation strategy, assessing risks, and handling emergencies. The planning module communicates with the decision-making module and is used to perform path planning based on the instructions from the decision-making module and environmental information, including global path planning, local obstacle avoidance planning, and landmark-based path adjustment.
3. User interaction module, which is responsible for interacting with users, receiving user commands, and providing navigation instructions and environmental feedback to users through voice and bone conduction; The memory module stores environment maps, landmark information, user preferences, and historical navigation data.
4. Coordination module, which communicates with the above modules and is used to manage the information flow and collaboration between the modules to ensure the efficient and coherent operation of the system.
5. The multimodal perception indoor navigation system for the blind based on hybrid agents according to claim 1, characterized in that: The sensor module includes: LiDAR is used to provide accurate depth information and environmental point clouds for building high-precision indoor maps and detecting obstacles. Cameras are used to provide rich texture and semantic information for recognizing landmarks, detecting obstacles, and identifying traffic lights; The IMU is used to provide motion information of the device, perform dead reckoning, provide continuous positioning information in a short period of time, and assist the SLAM algorithm to reduce drift accumulation.
6. The multimodal perception indoor navigation system for the blind based on hybrid agents according to claim 2, characterized in that: The specific method by which the perception module processes the raw sensor data is as follows: Real-time localization and mapping are performed using the fused sensor data. A graph-optimized SLAM algorithm is employed, fusing LiDAR point clouds, visual features, and IMU data to construct an accurate indoor map and estimate device pose. Using LiDAR point clouds and camera images, combined with deep learning algorithms, obstacles in the environment are detected and identified in real time, including static and dynamic obstacles. Camera images are processed using CNN- or Transformer-based object detection models, and LiDAR depth information is combined to perform 3D obstacle localization and size estimation. Camera images, combined with feature matching or deep learning methods, are used to identify predefined or automatically selected environmental landmarks. While constructing a geometric map, semantic information from camera images is used to construct a semantic map containing landmarks, room types, and accessibility facilities, providing the agent with a higher level of environmental understanding.
7. The multimodal perception indoor navigation system for the blind based on a hybrid agent according to claim 3, characterized in that: The decision-making module includes: Reactive components are used to handle real-time, low-level tasks; The thoughtful component is responsible for high-level planning and decision-making, specifically map-based global path planning, task decomposition, and strategy selection. This component uses the environment map and landmark information in the memory module to plan the path from the current location to the target. Learning components continuously optimize the agent's behavioral strategies by interacting with the environment and receiving user feedback; The coordination module integrates the above-mentioned different types of components, so that the deliberate component plans the global path, the reactive component is responsible for local obstacle avoidance, and the learning component adjusts the strategy according to the obstacle avoidance results.
8. The multimodal perception indoor navigation system for the blind based on hybrid agents according to claim 4, characterized in that: The deliberate component's map-based global path planning is implemented as follows: using a high-precision indoor map constructed by the perception module, path planning algorithms such as A* and D* are used to calculate the optimal path from the current location to the target location. While executing path planning, the reactive component uses obstacle information detected in real time by the perception module to perform local obstacle avoidance. The deliberate component also uses landmark information identified by the perception module to assist in positioning and path adjustment.
9. The multimodal perception indoor navigation system for the blind based on a hybrid agent according to any one of claims 1 to 5, characterized in that: The user interaction module includes: The voice interaction module uses speech recognition technology to receive the user's navigation target and query instructions, and also uses speech synthesis technology to convert navigation instructions, environmental information and warnings into voice output. Bone conduction feedback module, which transmits voice information directly to the user through bone conduction headphones; The multimodal feedback module combines feedback from multiple modalities, including speech and bone conduction, to provide richer and more intuitive information. The adaptive and personalized module adjusts the volume, speech rate, and level of detail of voice prompts, as well as the intensity and mode of bone conduction feedback, based on the user's hearing level, walking speed, and usage habits. The environmental information feedback module is used to provide environmental information to users.
10. The multimodal perception indoor navigation system for the blind based on a hybrid agent according to any one of claims 1 to 5, characterized in that: The decision-making module makes decisions under uncertainty as follows: Risk quantification and dissemination involves quantifying uncertainties in perceived information, positioning results, and environmental predictions, and tracking the propagation of uncertainty in the decision-making and planning process. Decisions should be made based on uncertainty. When the positioning uncertainty is high, the speed should be reduced, the environmental scanning frequency should be increased, or known landmarks should be prioritized for positioning correction. Maintain a safety margin by considering uncertainties and increasing the safety margin in path planning and obstacle avoidance; Emergency handling: Design response mechanisms for emergency situations.