Multi-mode sensing blind guiding system considering self-positioning and target guiding
Through a multimodal perception guide system, combined with voice interaction and deep learning models, the problems of self-positioning and target guidance of existing guide devices in unfamiliar environments are solved, enabling effective navigation and safe movement of blind people in unknown environments.
Patent Information
- Application Number
- CN202511049969.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-14
AI Technical Summary
Existing guide devices that do not rely on prior maps cannot achieve both self-positioning and goal guidance in unfamiliar environments, causing blind people to get lost in unknown environments.
A multimodal perception guidance system that takes into account both self-positioning and goal guidance is designed. It includes a multimodal interaction module, a spatial mapping module, a planning guidance module, and a trajectory tracking module. It obtains environmental information through voice interaction, uses a large language model and a deep learning model for environmental description and path planning, and combines IMU data and camera data for navigation.
It provides natural environmental descriptions and navigation feedback for the blind in unknown environments, improves the adaptability of guide equipment, reduces pathfinding time and collision frequency, and enhances the reliability and safety of navigation.
Smart Images

Figure CN120771022A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of perception and guidance technology, and in particular relates to a multimodal perception and guidance system for the blind that takes both self-positioning and target guidance into consideration. Background Art
[0002] Research on assistive technologies for environmental perception for the blind focuses on two key areas. The first is the acquisition of environmental information. The rapid development of machine vision technology has provided diverse solutions for this purpose, with the emergence of advanced algorithms ranging from target detection to scene analysis, enabling accurate perception of complex environments. The second is the effective transmission of acquired information to the user. Currently, a variety of modal strategies exist, including voice, audio, vibration, and electrical stimulation. The voice modality, due to its natural reliance on text, allows for greater density and accuracy in the expression of information. Furthermore, large-scale visual language models have made significant progress in recent years. Although customizing personalized models requires a lot of training (Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. (2018) Speaker-Follower Models for Vision-and-LanguageNavigation. ComputerVision and Pattern Recognition.), offline operation requires strong computing power support (Shah, D., Sridhar, AK, Dashora, N., Stachowicz, K., Black, K., Hirose, N., & Levine, S (2023) ViNT: A Foundation Model for Visual Navigation. 7th Conference on Robot Learning.). Fortunately, some models have opened online API calls, making them easy to use in indoor wifi environments. It is analogous to a robot reading a map, because blind people cannot read the map like mobile robots. Figure 1 It is still difficult for blind people to understand environmental information naturally because the map information is decoded mechanically according to universal standards.
[0003] Current artificial vision-based guide devices can be divided into two categories based on whether they require a prior map. In known environments, guide devices that require a prior map (M. Zhao et al. (2021) A General Framework for Lifelong Localization and Mapping in Changing Environment. IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS). 3305-3312. and G. Li, J. Xu, Z. Li, C. Chen and Z. Kan. (2023) Sensing and Navigation of Wearable Assistance Cognitive Systems for the Visually Impaired. IEEE Transactions on Cognitive and Developmental Systems. 15(1): 122-133.) can achieve relatively accurate positioning and path planning by matching prior maps with sensor data, providing accurate navigation information for the blind. However, its limitation is that in unknown environments such as unfamiliar rooms, due to the lack of a map, it is difficult to effectively carry out the guide task. Guide devices that don't rely on prior maps are more universal and can be divided into two categories. One type uses positioning and obstacle avoidance as its core technology (Zhang and C.Ye. (2020) A Visual Positioning System for Indoor Blind Navigation. IEEE International Conference on Robotics and Automation (ICRA) 9079-9085. and Sung-Jae Kang, Young Ho and In Hyuk Moon. (2001) Development of an intelligent guide-stick for the blind. IEEE International Conference on Robotics and Automation 4: 3208-3213.). These devices often rely on the user's memory to find the target. If the user cannot accurately recall the target location or environmental information, the device's navigation function will be limited.A method based on visual recognition and segmentation (J. Zhang, K. Yang, A. Constantinescu, K. Peng, K. Müller and R. Stiefelhagen (2021) Trans4Trans: Efficient Transformer for Transparent Object Segmentation to Help Visually Impaired People Navigate in the Real World. IEEE / CVF International Conference on Computer Vision Workshops (ICCVW) 1760-1770. and Zhang, K., Wang, Y., Shi, S. et al (2024) Improved yolov5 algorithm combined with depth camera and embedded system for blind indoor visual assistance. Sci Rep 14: 23000.) uses deep learning algorithms to identify and segment objects and roads in the scene, thereby guiding the blind to the target direction and having clear target guidance capabilities. However, when the target is not visible, the device cannot provide clear directional guidance to the user, causing the blind to lose their way.
[0004] In summary, the adaptability of existing guide devices that do not rely on prior maps is still poor. When faced with unknown environments such as unfamiliar rooms or when the target is invisible, existing guide devices that do not rely on prior maps still cannot take into account both self-positioning and target guidance functions. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem of poor adaptability of existing blind-guiding devices that do not rely on prior maps, and to propose a multimodal perception blind-guiding system that takes into account both self-positioning and target guidance.
[0006] The technical solution adopted by the present invention to solve the above technical problems is: a multimodal perception guidance system that takes into account both self-positioning and target guidance, the system includes a multimodal interaction module, a spatial mapping module, a planning and guidance module and a trajectory tracking module, wherein:
[0007] In the awakening stage of the multimodal perception blind guide system:
[0008] The user wakes up the system by inputting voice into the system's multimodal interaction module in the current environment, and then inputs questions about the entire environment information into the multimodal interaction module through voice.
[0009] The multimodal interaction module is used to obtain the environment image and convert the user's input voice into text, and then send the text and image to the spatial mapping module;
[0010] The spatial mapping module processes the input text and images and outputs a macro-description of the environment layout. The macro-description is sent to the multimodal interaction module, which converts the macro-description into speech and broadcasts it to the user.
[0011] In the guidance stage of the multimodal perception guidance system:
[0012] The user inputs voice into the system's multimodal interaction module to inquire about the location information of the target object;
[0013] The multimodal interaction module is used to obtain the environment image and convert the user's input voice into text, and then send the text and image to the spatial mapping module;
[0014] The spatial mapping module processes the input text and images, outputs a sentence containing the specific location of the target object, and sends the output sentence to the multimodal interaction module, which converts the output sentence into speech and broadcasts it to the user;
[0015] After the user inputs a voice message to start navigation into the system's multimodal interaction module, if the image acquired by the multimodal interaction module contains the target object, the planning and guidance module processes the image acquired by the multimodal interaction module to obtain the location of the target object, then performs path planning and sends movement instructions to the multimodal interaction module based on the path planning results. The multimodal interaction module then converts the movement instructions into voice and broadcasts them to the user.
[0016] If the image acquired by the multimodal interaction module does not contain the target object, the trajectory tracking module obtains the user's coordinates and orientation based on the image and IMU data, and outputs movement instructions or a reminder of the end of navigation based on the user's coordinates, orientation, and the most recently obtained target coordinates.
[0017] The beneficial effects of the present invention are:
[0018] The system of the present invention avoids the risk of failure of guide devices that do not rely on prior maps when facing special working conditions, and can provide natural environmental descriptions and navigation feedback for the blind, achieving the goal of allowing the blind to become familiar with an unfamiliar room and easily complete indoor navigation tasks. The strategy of the present invention is based on the actual needs of users, helping users to quickly become familiar with unfamiliar environments and helping blind people complete indoor navigation tasks; based on engineering technology, it solves the problem that guide devices without prior maps cannot take into account both self-positioning and target guidance during use, thereby improving the adaptability of guide devices. The feasibility of the guidance system of the present invention was verified by the improvement of pathfinding efficiency and the reduction of collision frequency in Experiment 1; the higher sketch similarity after voice interaction in Experiment 2 verified the effectiveness of the spatial perception of the present invention in helping blind people adapt to indoor environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is an architecture diagram of a multimodal perception blind guidance system that takes into account both self-positioning and target orientation of the present invention;
[0020] Figure 2 It is a schematic diagram of the system hardware composition and wear;
[0021] Figure 3 It is a schematic diagram of neighborhood expansion strategy;
[0022] Each grid in the figure represents the unit pixel size of the compressed image, black represents obstacles, white represents passable areas, red points represent the current position, and orange represents candidate neighboring points;
[0023] Figure 4a This is a schematic diagram of the path before post-processing;
[0024] The horizontal and vertical coordinates are pixel coordinates, black represents randomly generated obstacles, white represents the passable area, the green dot below represents the starting point, the red dot above represents the end point specified in the test phase, and the blue line segment is the resulting path;
[0025] Figure 4b This is a schematic diagram after path post-processing;
[0026] Figure 5 This is the workflow diagram of the blind guide system;
[0027] In the bird's-eye view, the green arrow indicates the user's actual direction during movement, the red arrow indicates the direction based on the system's guidance, and the black area indicates the inaccessible area where obstacles are located;
[0028] Figure 6 It is a diagram of information flow;
[0029] Figure 7 It is a guide blind mobile experimental scene;
[0030] Figure 8There are 9 experimental scenarios with similar spatial complexity;
[0031] In the figure, the central obstacle is a table, and the remaining 16 obstacles are boxes or chairs that occupy an area of 50 × 50 cm;
[0032] Figure 9 Schematic diagram of the subject's environmental cognition;
[0033] Figure 10 The data for Experiment 1;
[0034] a is the actual path length, b is the ratio of the actual path to the optimal path length, c is the collision frequency, and d is the success score; the line where the triangle is located is the experimental group, the line where the dot is located is the control group, and the corresponding dotted line is the average value;
[0035] Figure 11 The data for Experiment 2. DETAILED DESCRIPTION
[0036] Specific implementation method 1: Combination Figure 1 and Figure 6 This embodiment describes a multimodal perception guidance system for the blind that takes into account both self-positioning and target guidance. The system includes a multimodal interaction module, a spatial mapping module, a planning and guidance module, and a trajectory tracking module, wherein:
[0037] In the awakening stage of the multimodal perception blind guide system:
[0038] The user wakes up the system by inputting voice into the system's multimodal interaction module in the current environment, and then inputs questions about the entire environment information into the multimodal interaction module through voice.
[0039] The multimodal interaction module is used to obtain the environment image and convert the user's input voice into text, and then send the text and image to the spatial mapping module;
[0040] The spatial mapping module processes the input text and images and outputs a macro-description of the environment layout. The macro-description is sent to the multimodal interaction module, which converts the macro-description into speech and broadcasts it to the user.
[0041] In the guidance stage of the multimodal perception guidance system:
[0042] The user inputs voice into the system's multimodal interaction module to inquire about the location information of the target object;
[0043] The multimodal interaction module is used to obtain the environment image and convert the user's input voice into text, and then send the text and image to the spatial mapping module;
[0044] The spatial mapping module processes the input text and images, outputs a sentence containing the specific location of the target object, and sends the output sentence to the multimodal interaction module, which converts the output sentence into speech and broadcasts it to the user;
[0045] After the user inputs a voice message to start navigation into the system's multimodal interaction module, if the image acquired by the multimodal interaction module contains the target object, the planning and guidance module processes the image acquired by the multimodal interaction module to obtain the location of the target object, then performs path planning and sends movement instructions to the multimodal interaction module based on the path planning results. The multimodal interaction module then converts the movement instructions into voice and broadcasts them to the user.
[0046] If the image acquired by the multimodal interaction module does not contain the target object, the trajectory tracking module obtains the user's coordinates and orientation based on the image and IMU data, and outputs a movement instruction or a reminder of the end of navigation based on the user's coordinates, orientation and the most recently obtained target coordinates (the target's world coordinates are updated once new perception information is received, and if no valid perception information is received, the target's world coordinates remain unchanged).
[0047] like Figure 5 As shown, the blind guiding task of the present invention includes three stages: spatial perception, long-distance movement and short-distance approach.
[0048] 1. Spatial perception stage:
[0049] The user stands still and wakes up the entire system by calling a customized system nickname. The user then expresses his or her cognitive needs in the language logic he or she is accustomed to. The system uses an online API to process the user's natural language, and after recognizing the spatial representation intention, it begins to make a macro description of the layout of objects in the room according to the construction elements of the cognitive map. The system fully stimulates the user's spatial imagination and subjective initiative, helping users to become familiar with unfamiliar environments from scratch. This method enables users to build a more concrete cognitive map in their minds, and at the same time trust the system more, which helps to improve the efficiency of subsequent navigation. After having a certain cognitive foundation, users can further inquire about the location information of the target object and obtain a targeted introduction output by the system's spatial mapping module (for example, the person is 2.2 meters in front of the 11 o'clock direction, and the table is 1.5 meters directly in front).
[0050] 2. Long distance movement stage:
[0051] The user moves and turns according to the system's voice guidance. Voice guidance is generated in two ways: when the target object is within the camera's field of view, the system plans a path based on the current target position and accessible area, generating movement instructions. When the target is out of view, the system generates movement instructions based on the target's historical location and the user's current posture. The system not only provides users with movement instructions containing direction and location information, but also includes a description of the current accessible area, improving the user's cognitive map and self-positioning, answering the user's questions of "how should I walk," "why do I walk this way," and "where have I walked?"
[0052] 3. Short distance approach stage:
[0053] Due to the camera's field of view (FOV), path planning is difficult to implement when the user is too close to the target, preventing the system from recognizing the target through a partial image or missing the target. In these cases, the system provides feedback on the target's relative position based on the target's historical position and the user's current posture. Based on this feedback, the user approaches the target object by taking small steps and turning at small angles until the distance is within 50 cm. This concludes navigation, allowing the user to interact with the target simply by swinging their arms.
[0054] Specific embodiment 2: This embodiment differs from specific embodiment 1 in that the hardware part of the multimodal interaction module includes headphones with an integrated microphone and a camera (the cameras that can be used include but are not limited to the Intel RealSense D435i camera).
[0055] Other steps and parameters are the same as those in the first embodiment.
[0056] Hardware composition and wearing method Figure 2 As shown in the figure, the multimodal interaction module builds an interactive bridge between people, devices, and the environment. Its functions include wake-up, voice recognition, TTS (text-to-speech), image capture, and IMU data recording, and it is responsible for deciding whether to activate the entire system. The multimodal interaction module receives the user's voice and converts it into text, which is then sent along with the image to the spatial mapping module for processing. It is also responsible for receiving the results (spatial descriptions and navigation instructions) from the system's internal spatial mapping module, planning and guidance module, and trajectory tracking module, and converting these results into speech to broadcast to the user.
[0057] Specific embodiment three: This embodiment differs from specific embodiments one or two in that the software portion of the spatial mapping module includes a large language model and a Yolact deep learning model;
[0058] The large language model is fine-tuned by prompt, and the fine-tuned large language model is used to process text and images;
[0059] The Yolact deep learning model is used to identify objects in images;
[0060] The spatial mapping module outputs a macro-description text of the environment layout and a sentence containing the specific location of the target object based on the processing results of the large language model and the recognition results of the Yolact deep learning model.
[0061] Other steps and parameters are the same as those in the first or second embodiment.
[0062] The present invention adjusts the large language model by prompting so that it can output cognitive Figure 5 The scene description of elements (roads, signs, nodes, areas, boundaries) completes the conversion of spatial scenes into sentence groups. An abstract concept formed by people's perception and reasoning of the environment is called a cognitive map, which specifically refers to the psychological representation of the spatial environment formed in people's minds (CHEN Xiaomeng, LIU Chunling, QIAO Fuqiang, QI Kemin (2016) The process of construction of spatial representation in the unfamiliar environment in the blind: The role of strategies and its effect. Acta Psychologica Sinica 48 (6): 637-647.), and cognitive maps are the basis of navigation behavior (M. Allritz, et al (2022) Chimpanzees (Pantroglodytes) navigate to find hidden fruit in a virtual environment. Sci. Adv 8: 4754.). The description output by the present invention can help users build a cognitive map, which is a psychological spatial representation that can help people find directions in complex environments and effectively find routes (Burkut, EB & B(2024)Cognitive mapping and space syntax analysis of universaldesign principles:The case of the Barrier-Free Life Center. Journal of Architectural Sciences and Applications 9(1):422-443). Cognitive maps can be initially and vaguely established in the human mind through macroscopic natural language descriptions, and can also be further supplemented and refined by obtaining the orientation of specific objects. Therefore, the present invention allows users to ask questions about the orientation of objects in the room, identify objects through Yolact, and learn their orientation using depth information. Studies have found that blind people prefer to encode environmental information in an egocentric reference frame (Ruggiero G, Ruotolo F, Iachini T (2022) How ageing and blindness affect egocentric and allocentric spatial memory. QJ Exp Psychol (Hove) 75(9):1628-1642.), so the present invention compiles object information into a descriptive statement that uses a clock to represent the orientation (for example, a person is 2.3 meters at 11 o'clock).
[0063] Specific embodiment 4: This embodiment differs from any one of specific embodiments 1 to 3 in that the working process of the planning guidance module is as follows:
[0064] The Yolact deep learning model is used to segment each object in the image. Then, based on the user's current location, the target location, and the locations of various objects in the current environment, a bidirectional A* path planning algorithm is used to plan a path from the user's current location to the target location.
[0065] And according to the path planning results, movement instructions are sent to the multimodal interaction module, and the multimodal interaction module converts the movement instructions into voice and broadcasts them to the user.
[0066] The other steps and parameters are the same as those in the first to third embodiments.
[0067] The planning and guidance module uses the Yolact deep learning model and the bidirectional A* path planning algorithm. The input RGB-D video stream is used to output movement instructions, which realizes map acquisition based on the first-person perspective and real-time path planning. Traditional target navigation path planning algorithms are mostly based on preloaded global maps or real-time bird's-eye views for planning on the real ground. Although this method is conducive to improving the accuracy of later positioning, for guiding tasks in unstructured environments, it is unrealistic to have the blind walk around indoors to establish a global map for path planning. Moreover, the planning and guidance module of the present invention can also target dynamic objects, always follow the target position and make corresponding plans. At this point, the advantages of the path planning method based on the first perspective of the present invention are more prominent:
[0068] 1) Better robustness: Robustness can be improved by adjusting the system through real-time feedback;
[0069] 2) Better safety: Real-time perception and local planning enable safer navigation in unknown or dangerous environments;
[0070] 3) Better privacy: It reduces comprehensive scanning of the environment and data storage, which helps protect privacy.
[0071] This paper uses a public dataset of indoor multi-pedestrian scenes (H.Mu, G.Zhang, Z.Ma, M.Zhou and Z.Cao (2024) Dynamic Obstacle Avoidance System Based on Rapid Instance Segmentation Network. IEEE Transactions on Intelligent Transportation Systems 25(5): 4578-4592.) to train the Yolact deep learning model. The trained model can segment the ground, people, tables, and chairs. After obtaining the segmentation results, it is necessary to convert them into a map that can be used for path planning, that is, a map with a starting point, end point, and traversable area, which provides the conditions for the operation of the path planning algorithm.
[0072] A* is a classic heuristic path search method that not only estimates the cost from the current node to the next node, but also takes into account the cost to the target point, thereby finding the shortest path. It has the advantages of low complexity, easy implementation, low computational requirements, and can achieve global path planning. However, in actual operation, A* has limitations such as multiple consecutive turns, paths close to obstacles, and insufficient search efficiency (D. Zhang, C. Chen and G. Zhang (2024) AGV Path Planning Based on Improved A-star Algorithm. IEEE 7th Advanced Information Technology Electronic and Automation Control Conference (IAEAC) 1590-1595.). Therefore, the bidirectional A* path planning algorithm of the present invention adjusts and optimizes the traditional A* algorithm so that it can better serve the blind guide scene while further improving the efficiency of the search. Moreover, on the basis of bidirectional search, the traditional A* is modified and optimized from three aspects: neighborhood expansion mode, cost function modification, and path post-processing, making it more suitable for the actual needs of blind people walking.
[0073] Specific embodiment 5: This embodiment differs from any one of specific embodiments 1 to 4 in that the neighborhood expansion method of the bidirectional A* path planning algorithm is:
[0074] Mark the coordinates of the current position on the map as (x, y). Then the candidate neighboring points obtained by expanding the 4-neighborhood from the current position are (x, y+1), (x, y-1), (x+1, y), and (x-1, y), and expansion is prioritized in the y direction.
[0075] For any candidate neighbor point, continue to expand a (in this invention, the value of a is 10) pixels along the expansion direction of the candidate neighbor point. If an obstacle area is encountered during the expansion process, the candidate neighbor point is eliminated;
[0076] Similarly, each candidate neighbor point obtained by expanding the 4-neighborhood is processed separately, and then path planning is performed based on the remaining candidate neighbor points.
[0077] The other steps and parameters are the same as those in the first to fourth embodiments.
[0078] This implementation method makes the path safer and more convenient for blind people to understand by modifying the neighborhood expansion method. Figure 3 As shown. The neighborhood expansion methods of the A* algorithm include 4 neighborhoods, 8 neighborhoods, etc. Different neighborhood expansion methods are adapted to different scene needs. This embodiment adopts a 4-neighborhood expansion method, which limits the path turning angle to only ±90°, because for the blind, the compliance with the "turn left", "turn right" and "go forward" instructions is the highest. Assuming that the current position is (x, y), the candidate neighboring points obtained by the 4-neighborhood expansion are [(x, y+1), (x, y-1), (x+1, y), (x-1, y)]. Since the expansion rule is to expand in the y direction first, after the neighboring point (x±1, y) is obtained, the surrounding area of 10 pixels in the expansion direction is checked for the existence of obstacles. If the distance between this neighboring point and the obstacle is too close, its dual neighboring point is directly removed. Output is used as the result of this neighborhood expansion, so as to stay away from obstacles and ensure the safety of blind people walking. Figure 3 As shown in the figure, candidate neighbor point 3 fails the obstacle existence check in the extension direction, and the extension result outputs only neighbor points 1, 2, and 4.
[0079] Specific embodiment 6: This embodiment differs from any one of specific embodiments 1 to 5 in that the cost function of the bidirectional A* path planning algorithm is:
[0080] f(n)=g(n)+h(n)+t(n)
[0081] Where n represents the current node, f(n) represents the cost function, g(n) represents the sinking cost function, and t(n) represents the turning cost function;
[0082] If the current node is a node obtained in the forward search, and the current node is not collinear with the parent node or the grandparent node, then t(n) = 100;
[0083] If the current node is a node obtained in the reverse search, and the current node is not collinear with the parent node or the grandparent node, then t(n) = 50;
[0084] Otherwise, t(n)=0.
[0085] The other steps and parameters are the same as those in the first to fifth embodiments.
[0086] In the traditional A* algorithm, as each node is expanded, the node with the lowest expansion cost is selected until the target node is reached and the optimal path is obtained. During the search process, the total cost of each node is evaluated by adding g(n) and h(n). The cost function is defined as follows:
[0087] f(n)=g(n)+h(n)
[0088] Where n is the current node; the evaluation function f(n) represents the total cost associated with the current node; the sunk cost function g(n) expresses the cost of the number of steps from the starting point to the current node; and h(n) is a heuristic function that provides an estimate of the expected cost of reaching the end point from the current node (expressed as the Euclidean distance from the current node to the end point).
[0089] The traditional cost function f(n) greatly affects the search process of the A* algorithm and lacks consideration for cornering. In order to reduce the number of corners during the search process, the present invention improves the cost function f(n) by adding a corner cost function t(n) to the original one, that is, f(n) = g(n) + h(n) + t(n);
[0090] In a bidirectional A* search, the search starting from the starting point is called a forward search, and the search starting from the end point is called a reverse search. If the current node is an inflection point in the forward search (i.e., it is not collinear with the parent node or the grandparent node), then t(n) = 100; if the current node is an inflection point in the reverse search, then t(n) = 50; otherwise, if the current node is not an inflection point, t(n) = 0. This ensures that the blind always receive relatively concise and low-frequency guide instructions as they move forward. On the contrary, if we lower the t(n) of the forward search, the algorithm will tend to make the blind turn at the beginning, and then the change in the viewing angle of the head-mounted camera will replan the path, and the algorithm will still tend to make the blind turn, falling into a vicious cycle of going around in circles. The present invention can effectively avoid this vicious cycle by setting different turning cost function values in the bidirectional search.
[0091] Specific embodiment 7: This embodiment differs from any one of specific embodiments 1 to 6 in that the bidirectional A* path planning algorithm performs post-processing on the path after planning the path. The post-processing method is as follows:
[0092] Step 1: Put the starting point, end point and all intermediate nodes in the planned path into the list P0 in the order of the path. For the i-th node N i , if node N i-1 、N i and N i+1 Collinear, then N i Not an inflection point, otherwise N i is an inflection point; after traversing each node in the list P0, we get the inflection point list C consisting of the starting point, the end point and all the inflection points;
[0093] Step 2: Create an empty inflection point list T and record the kth inflection point in list C as c k , k∈[0,q], q+1 represents the total number of inflection points in list C, and the distance d(c k ,c k-1 );
[0094] Calculate w(k):
[0095] w(k)=d(c k ,c k-1 ) / n-1
[0096] Where n represents the number of nodes in list P0 including the starting point and the end point;
[0097] If w(k)≤0.2, then c k Put it into the list T, otherwise, do not put c k Put it into list T;
[0098] And perform step 3 for the inflection points in list T;
[0099] Step 3: For any inflection point c in the list T k , find the inflection point c in list C k Neighbor inflection point c k-1 and c k+1 , and c k The coordinates are marked as (x k ,y k ), c k-1 The coordinates are marked as (x k-1 ,y k-1 ), c k+1 The coordinates are marked as (x k+1 ,y k+1 );
[0100] If x k =x k-1 ≠x k+1 , then the inflection point c k The horizontal axis is optimized as If x k =x k+1 ≠x k-1 , then the inflection point c k The horizontal axis is optimized as Similarly, for the inflection point c k The vertical coordinate of the optimized The inflection point c k The corresponding optimized inflection point is recorded as
[0101] Use a straight line with c k-1 Connect, if with c k-1 If the line connecting the two does not pass through the obstacle, then use Replace the kth inflection point c in list C k Otherwise, do not calculate the inflection point c in list C. k Make changes;
[0102] After traversing each inflection point in list T, we get the optimized list C;
[0103] Step 4: Use straight lines to sequentially connect all points in the optimized list C to obtain a complete path.
[0104] The other steps and parameters are the same as those in the first to sixth embodiments.
[0105] It should be noted that when the inflection point c in the list T k Optimized to After that, when optimizing the next inflection point in list T, that is, the inflection point c in list T k+1 When optimizing, we will As the inflection point c k+1 The previous turning point.
[0106] In the path obtained after the search is completed, such as Figure 4a As shown in FIG, there are cases where there are continuous turns within a short distance. The present invention optimizes the entire path through a post-processing function to further reduce the number of turns, making the path and guidance instructions easier for the blind to understand and follow, as shown in FIG. Figure 4b The pseudo code of this embodiment is shown in Table 1:
[0107] Table 1 Path post-processing process
[0108]
[0109]
[0110] The movement instructions provided by the present invention for the blind are based on the relative position of the first turning point obtained after each planning, such as "turn left and walk 1.1 meters", "turn right and walk 1.1 meters", and "move forward 1.1 meters".
[0111] Specific embodiment eight: This embodiment differs from any one of specific embodiments one to seven in that the working process of the trajectory tracking module is as follows:
[0112] Step 1: Receive the target pixel coordinates (u g ,v g ) and depth information, and then use the camera internal parameters to convert the target pixel coordinates (u g ,v g ) is converted to world coordinates (X g ,Y g ,Z g ) and in world coordinates (X g ,Y g ,Z g ) as the origin and r as the radius to make a circle, and get the circular area;
[0113] Step 2: Use the self-positioning algorithm to process the image and IMU data to obtain the user's coordinates and orientation;
[0114] If the user's coordinates are within the circular area obtained in step 1, the user has reached the destination, and a navigation end reminder is sent to the multimodal interaction module, which is converted into voice and broadcast to the user by the multimodal interaction module;
[0115] If the user's coordinates are not within the circular area obtained in step 1, then according to the user's orientation and world coordinates (X g,Y g ,Z g ) outputs a movement instruction, wherein the movement instruction is to move right, left, forward or backward; and sends the movement instruction to the multimodal interaction module, which is converted into voice by the multimodal interaction module and broadcast to the user.
[0116] The other steps and parameters are the same as those in the first to seventh embodiments.
[0117] The trajectory tracking module takes the camera's RGB image and the three-axis acceleration and three-axis angular acceleration measured by the IMU as input, records the target's location, and outputs the historical trajectory, current position, and movement instructions. Throughout the dynamic process of guiding the blind, the target point may be lost at any time due to being too close for recognition, the user turning around, or the user lowering their head. In such cases, the planning and guidance module is unable to issue effective movement instructions due to the lack of its necessary conditions. The trajectory tracking module must combine the historical endpoint location with the user's current position to issue movement instructions and a reminder of the end of navigation, ensuring the consistency of instructions received by the user during movement.
[0118] The trajectory tracking module uses the VINS-fusion framework (Tong Qin, Shaozu Cao, Jie Pan, Shaojie Shen (2019) A General Optimization-based Framework for Global Pose Estimation with Multiple Sensors. ArXiv abs / 1901.03642.). During the guidance process, it is essential to track the positioning and motion state of the blind, which directly reflects the success of navigation and the safety of the blind. A tightly coupled nonlinear optimization method is used to obtain a high-precision visual-inertial odometry by fusing pre-integrated IMU measurements and feature observations. VINS can directly estimate motion state without relying on any prior map. The functional modules of VINS include five parts: data preprocessing, initialization, tightly coupled monocular VIO, relocalization, and global pose graph optimization.
[0119] Its workflow is a typical VSLAM framework process:
[0120] (1) When the system is just started, it is necessary to initialize the first few frames to obtain high-quality feature points for subsequent PnP pose solution;
[0121] (2) Each time an image frame comes in, the image needs to be pre-processed and the feature points are tracked by the KLT sparse optical flow algorithm. At the same time, new corner features are detected to maintain the minimum number of features in each image (100-300). The relative pose is solved using PnP based on the matched feature points.
[0122] (3) The backend uses visual reprojection to continuously optimize the key frames and map points filtered by the frontend to make the results more accurate;
[0123] (4) Loop detection determines whether there is a loop and reduces the global error (T. Qin, P. Li and S. Shen (2018) VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State Estimator. IEEE Transactions on Robotics 34 (4): 1004-1020.).
[0124] Because the Z-axis coordinate (i.e., height) of indoor objects is not very helpful for long-distance movement, the endpoint is projected vertically onto the ground, creating a circular area with a radius of r = 0.5m. Whether the target is within the circular area is used as the criterion for determining whether the target has been reached. If the target is not within the field of view, the user is guided to turn and search for the target based on its location and current posture until the target is recognizable. The system can then plan a new path and provide new guidance.
[0125] Experimental verification
[0126] 1. Experimental Setup
[0127] This paper sets up two sets of experiments, with finding the target object in an unstructured environment as the basic experimental task. Figure 7 The following is an experimental scene of blind guide movement. Figure 8 Nine experimental scenarios with similar spatial complexity were designed, and the calculation formula for the standard deviation σ of obstacle distribution is:
[0128]
[0129] Taking the centroid of the rectangular room as the origin, the distance from the center of the square projected by obstacle i (chair or box) on the ground to the origin is d i (i∈[1,16]);
[0130] The calculation formula of channel complexity C is:
[0131]
[0132] Where N is the number of channels, is the average channel width (unit: cm);
[0133] Experiment 1: Guide movement experiment Figure 7 As shown. Figure 8As shown, a total of nine rooms were built, unfamiliar to the subjects. The spatial complexity of the interior layout of the unfamiliar rooms (the standard deviation σ of the projection center of each obstacle to the center of the room, and the channel complexity C) was controlled to be within the same range. The subjects entered an unfamiliar room, starting from the doorway and ending at a randomly placed target object, and searched for the target object with the help of the system. To avoid performance bias caused by inference learning from repeated experiments, a control experiment was conducted one week later, with the room numbers in the order randomized. The control group subjects only relied on touch to find the target object in the room. The data we collected included:
[0134] 1) Movement time t (unit: seconds);
[0135] 2) Success score (100 points for touching the target, 60 points for approaching the target to within 0.5 meters, 0 points for not entering the 0.5-meter range of the target within 5 minutes, or 0 points for the subject's voluntary decision to stop);
[0136] 3) The actual path length L (in meters) from the start to the end (touching the target or entering the 0.5-meter area around the target);
[0137] 4) Number of collisions during the entire process (collision with walls or obstacles);
[0138] Experiment 2: Cognitive experiment. The subject entered a new room and stood at the door. After interacting with the spatial perception module for 1 minute to gain environmental cognition, the subject left the room and drew a cognitive sketch. In the control group, the subject gained environmental cognition through 1 minute of autonomous movement and exploration. After leaving the room, the subject drew a cognitive sketch. The spatial complexity of the interior layout of the new room was controlled within the same range. The cognitive sketches drawn by each subject were collected. The experimental process was as follows: Figure 9 shown.
[0139] 2. Data processing and discussion of results
[0140] In Experiment 1, the movement time t and path length L both demonstrate the performance of the system of the present invention in assisting users in performing tasks. We found that without the assistance of a guide system, subjects were easily disoriented and frightened during movement, thus actively stopping the experiment. Therefore, we made the control group's task easier than the experimental group. However, we controlled the difficulty within the same range, with the optimal path length ranging from 9 to 15 meters (the average length of the experimental group was 13.44 meters, and the average length of the control group was 12.43 meters), and the standard deviation of the obstacle distribution in the scene ranging from 75 to 150 (the average of the experimental group was 110.50, and the average of the control group was 99.22). We calculated the subjects' actual walking path length, the frequency of collisions (walls and obstacles), the ratio of the actual path length to the optimal path length, the success score, and the mean of each evaluation indicator.
[0141] The data of Experiment 1 are as follows Figure 10 As shown in the figure, the multiple relationship between the actual distance traveled and the optimal path length indicates that the blind person's guidance system significantly reduces unnecessary or wasted distance traveled, significantly improving pathfinding efficiency. It helps blind people move purposefully toward their destination, demonstrating significant guidance and targeting effectiveness. The obstacle avoidance system reduces the frequency of collisions during navigation, improving safety for blind people moving through unfamiliar rooms.
[0142] In Experiment 2, we compared the coordinates of obstacles, user positioning, and target positioning in the sketches drawn by the subjects with the actual information displayed in the room, and calculated the two-dimensional regression similarity R. 2 (Nobbir A, Miller HJ (2007) Time-space transformations of geographic space for exploring, analyzing and visualizing transportation systems. Journal of Transport Geography 15: 2-17.), quantitatively comparing the subjects' degree of cognition of the environment.
[0143] Two-dimensional correlation index calculation formula:
[0144]
[0145] R 2 The closer it is to 1, the higher the similarity is, and the closer it is to 0, the lower the similarity is. i ,y i ) is the design position of a point in the plane, They are x i and y i The true value in the sketch, and They are x i and y i Average value (in the real map).
[0146] The data of Experiment 2 are as follows Figure 11 Sketch similarity analysis shows that, overall, the spatial cognition module provides a more comprehensive understanding of the overall environment than autonomous exploration in the same timeframe. While autonomous exploration can reveal the composition of obstacles within a group, it's less useful for navigation. Blind people's habit of walking along the edges of obstacles or walls during movement also allows them to determine the exact number of chairs in a row.
[0147] The movement commands we designed for navigation not only serve as target guidance but also provide subjects with a reference standard to determine the degree of rotation, thereby strengthening their confidence in their actions and their trust in the system. Loss of turn angle and relative orientation was the main factor that caused subjects to stop the control experiment, and our guide system provides a basis for determining turn angles. Relative orientation, with the user as the center of a clock, is a guidance strategy that aligns with the living habits of blind people and is more effective in actual use than absolute directions of east, south, west, and north.
[0148] The above examples are merely illustrative of the calculation model and process of the present invention and are not intended to limit the embodiments of the present invention. Persons skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. This list of embodiments is not exhaustive; any obvious variations or modifications arising from the technical solution of the present invention remain within the scope of protection of the present invention.
Claims
1. A multimodal perception blind guidance system that takes into account both self-positioning and target guidance, characterized in that: The system includes a multimodal interaction module, a spatial mapping module, a planning guidance module, and a trajectory tracking module, wherein: In the awakening stage of the multimodal perception blind guide system: The user wakes up the system by inputting voice into the system's multimodal interaction module in the current environment, and then inputs questions about the entire environment information into the multimodal interaction module through voice. The multimodal interaction module is used to obtain the environment image and convert the user's input voice into text, and then send the text and image to the spatial mapping module; The spatial mapping module processes the input text and images and outputs a macro-description of the environment layout. The macro-description is sent to the multimodal interaction module, which converts the macro-description into speech and broadcasts it to the user. In the guidance stage of the multimodal perception guidance system: The user inputs voice into the system's multimodal interaction module to inquire about the location information of the target object; The multimodal interaction module is used to obtain the environment image and convert the user's input voice into text, and then send the text and image to the spatial mapping module; The spatial mapping module processes the input text and images, outputs a sentence containing the specific location of the target object, and sends the output sentence to the multimodal interaction module, which converts the output sentence into speech and broadcasts it to the user; After the user inputs a voice message to start navigation into the system's multimodal interaction module, if the image acquired by the multimodal interaction module contains the target object, the planning and guidance module processes the image acquired by the multimodal interaction module to obtain the location of the target object, then performs path planning and sends movement instructions to the multimodal interaction module based on the path planning results. The multimodal interaction module then converts the movement instructions into voice and broadcasts them to the user. If the image acquired by the multimodal interaction module does not contain the target object, the trajectory tracking module obtains the user's coordinates and orientation based on the image and IMU data, and outputs movement instructions or a reminder of the end of navigation based on the user's coordinates, orientation, and the most recently obtained target coordinates.
2. A multimodal perception blind guide system that takes into account both self-positioning and target guidance according to claim 1, characterized in that: The hardware part of the multimodal interaction module includes headphones with an integrated microphone and a camera.
3. The multimodal perception blind guide system that takes into account both self-positioning and target guidance according to claim 2, characterized in that: The software part of the spatial mapping module includes a large language model and a Yolact deep learning model; The large language model is fine-tuned by prompting, and the fine-tuned large language model is used to process text and images; The Yolact deep learning model is used to identify objects in images; The spatial mapping module outputs a macro-description text of the environment layout and a sentence containing the specific location of the target object based on the processing results of the large language model and the recognition results of the Yolact deep learning model.
4. The multimodal perception blind guidance system that takes into account both self-positioning and target guidance according to claim 3, characterized in that: The working process of the planning guidance module is as follows: The Yolact deep learning model is used to segment each object in the image. Then, based on the user's current location, the target location, and the locations of various objects in the current environment, a bidirectional A* path planning algorithm is used to plan a path from the user's current location to the target location. And according to the path planning results, movement instructions are sent to the multimodal interaction module, and the multimodal interaction module converts the movement instructions into voice and broadcasts them to the user.
5. The multimodal perception blind guidance system that takes into account both self-positioning and target guidance according to claim 4, characterized in that: The neighborhood expansion method of the bidirectional A* path planning algorithm is: Mark the coordinates of the current position on the map as (x, y). Then the candidate neighboring points obtained by expanding the 4-neighborhood from the current position are (x, y+1), (x, y-1), (x+1, y), and (x-1, y), and expansion is prioritized in the y direction. For any candidate neighbor point, continue to expand a pixels along the expansion direction of the candidate neighbor point. If an obstacle area is encountered during the expansion process, the candidate neighbor point will be eliminated. Similarly, each candidate neighbor point obtained by expanding the 4-neighborhood is processed separately, and then path planning is performed based on the remaining candidate neighbor points.
6. The multimodal perception blind guidance system that takes into account both self-positioning and target guidance according to claim 5, characterized in that: The cost function of the bidirectional A* path planning algorithm is: f(n)=g(n)+h(n)+t(n) Where n represents the current node, f(n) represents the cost function, g(n) represents the sinking cost function, and t(n) represents the turning cost function; If the current node is a node obtained in the forward search, and the current node is not collinear with the parent node or the grandparent node, then t(n) = 100; If the current node is a node obtained in the reverse search, and the current node is not collinear with the parent node or the grandparent node, then t(n) = 50; Otherwise, t(n)=0.
7. The multimodal perception blind guidance system that takes into account both self-positioning and target guidance according to claim 6, characterized in that: After the bidirectional A* path planning algorithm plans a path, it performs post-processing on the path. The post-processing method is as follows: Step 1: Put the starting point, end point and all intermediate nodes in the planned path into the list P0 in the order of the path. For the i-th node N i , if node N i-1 、N i and N i+1 Collinear, then N i Not an inflection point, otherwise N i is an inflection point; after traversing each node in the list P0, we get the inflection point list C consisting of the starting point, the end point and all the inflection points; Step 2: Create an empty inflection point list T and record the kth inflection point in list C as c k , k∈[0,q], q+1 represents the total number of inflection points in list C, and the distance d(c k ,c k-1 ); Calculate w(k): w(k)=d(c k ,c k-1 ) / n-1 Where n represents the number of nodes in list P0 including the starting point and the end point; If w(k)≤0.2, then c k Put it into the list T, otherwise, do not put c k Put it into list T; And perform step 3 for the inflection points in list T; Step 3: For any inflection point c in the list T k , find the inflection point c in list C k Neighbor inflection point c k-1 and c k+1 , and c k The coordinates are marked as (x k ,y k ), c k-1 The coordinates are marked as (x k-1 ,y k-1 ), c k+1 The coordinates are marked as (x k+1 ,y k+1 ); If x k =x k-1 ≠x k+1 , then the inflection point c k The horizontal axis is optimized as If x k =x k+1 ≠x k-1 , then the inflection point c k The horizontal axis is optimized as Similarly, for the inflection point c k The vertical coordinate of the optimized The inflection point c k The corresponding optimized inflection point is recorded as Use a straight line with c k-1 Connect, if with c k-1 If the line connecting the two does not pass through the obstacle, then use Replace the kth inflection point c in list C k Otherwise, do not calculate the inflection point c in list C. k Make changes; After traversing each inflection point in list T, we get the optimized list C; Step 4: Use straight lines to connect all the points in the optimized list C in sequence to obtain a complete path.
8. The multimodal perception blind guidance system that takes into account both self-positioning and target guidance according to claim 7, characterized in that: The working process of the trajectory tracking module is as follows: Step 1: Receive the target pixel coordinates (u g ,v g ) and depth information, and then use the camera internal parameters to convert the target pixel coordinates (u g ,v g ) is converted to world coordinates (X g ,Y g ,Z g ) and in world coordinates (X g ,Y g , Z g) is the origin and r is the radius to make a circle to obtain the circular area; Step 2: Use the self-positioning algorithm to process the image and IMU data to obtain the user's coordinates and orientation; If the user's coordinates are within the circular area obtained in step 1, the user has reached the destination, and a navigation end reminder is sent to the multimodal interaction module, which is converted into voice and broadcast to the user by the multimodal interaction module; If the user's coordinates are not within the circular area obtained in step 1, then according to the user's orientation and world coordinates (X g ,Y g ,Z g ) outputs a movement instruction, wherein the movement instruction is to move right, left, forward or backward; and sends the movement instruction to the multimodal interaction module, which is converted into voice by the multimodal interaction module and broadcast to the user.