Intelligent blind guiding method, system and device based on multi-modal data fusion
By using smart glasses and mobile devices to collaboratively process multimodal data, navigation commands are generated and auditory and tactile feedback is provided, solving the problems of large size and high power consumption of existing guide devices for the blind, and realizing lightweight, low-cost real-time navigation and safe navigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-13
AI Technical Summary
Existing electronic navigation devices for the visually impaired suffer from problems such as large size, high power consumption, high cost, inconvenient operation, and difficulty in achieving real-time response and safe navigation.
By collecting multimodal data in real time through smart glasses, generating micro-environmental data, and processing it in collaboration with mobile devices, navigation instructions are generated by combining local positioning and path data. Bone conduction headphones and vibration motors are used to provide auditory and tactile feedback to guide visually impaired people to their destination safely.
It achieves lightweight, low-power, and low-cost real-time navigation, enhancing the independence and safety of visually impaired people's travel, and providing rich navigation information feedback and dynamic obstacle avoidance capabilities.
Smart Images

Figure CN121655494A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and intelligent assisted navigation technology, and in particular to an intelligent guidance method, system and device based on multimodal data fusion. Background Technology
[0002] Independent and safe travel for visually impaired individuals is a matter of widespread social concern. Traditional assistive devices such as white canes and guide dogs have significant limitations. For example, white canes have a limited detection range and cannot detect obstacles above waist level or dynamic hazards at a distance; guide dogs are extremely expensive to train and are scarce, making them difficult to popularize.
[0003] In recent years, with the development of artificial intelligence technology, various electronic guide devices for the visually impaired have emerged. However, existing technical solutions generally suffer from the following problems: First, while systems relying on the fusion of multiple sensors such as lidar and ultrasound can achieve high perception accuracy, the large number of sensors and low integration result in large device size, high power consumption, and high cost, making them difficult to wear and limiting their widespread application in daily life. Second, while pure software solutions based on smartphone applications offer better portability, they require users to frequently hold their phones to scan the environment, which is inconvenient. Furthermore, limited by the phone's computing power, data processing latency is high, making it difficult to respond to sudden obstacles in real time and affecting safety.
[0004] Therefore, overcoming the shortcomings of the existing technology is an urgent problem to be solved in this technical field. Summary of the Invention
[0005] In response to the above-mentioned deficiencies or improvement needs of existing technologies, this invention proposes an intelligent guidance method, system, and device based on multimodal data fusion. By coordinating the processing of smart glasses and mobile terminals and fusing multimodal data to generate navigation instructions, it can effectively guide visually impaired individuals to arrive at and enter their destinations, thereby enhancing their independence in travel.
[0006] The embodiments of the present invention adopt the following technical solutions: In a first aspect, the present invention provides an intelligent guidance method for the blind based on multimodal data fusion. The method is executed collaboratively by smart glasses and a mobile terminal. Specifically, the smart glasses collect and process all raw data to obtain structured data corresponding to each raw data. The smart glasses integrate all structured data to obtain micro-environmental data that characterizes the spatial attributes and semantic features of each object, and then upload the micro-environmental data to the mobile device. The mobile device combines the aforementioned micro-environmental data, local positioning data, and macro-path data to generate a navigation environment situation map; When a user's voice command is received, the mobile device generates a hybrid navigation command based on the navigation environment situation map and feeds it back to the user to guide the user to arrive at and enter the destination.
[0007] Preferably, the raw data includes angular velocity, acceleration, and raw image; the structured data includes pose, geometric information, and semantic information; and the method further includes: The pose of the observation point in three-dimensional space is obtained by integrating the angular velocity and the acceleration. The original image is preprocessed to obtain the target image; Based on the visual features of objects in the target image, a depth map is obtained by performing depth estimation on the target image; and the geometric information of the depth map is extracted. Object detection and recognition are performed on the target image to obtain the two-dimensional bounding boxes of all objects and the semantic information corresponding to each object.
[0008] Preferably, the structured data includes pose, geometric information, and semantic information. The process of fusing all structured data to obtain micro-environmental data characterizing the spatial attributes and semantic features of each object, and uploading the micro-environmental data to the mobile terminal, includes: Synchronize the pose, geometric information, and semantic information in time; Based on the pose and geometric information after time synchronization, the depth values within the 2D bounding box are back-projected and aggregated with the pose after time synchronization to generate a 3D object containing 3D coordinates and dimensions. Based on the three-dimensional coordinates of the three-dimensional object, the three-dimensional object is associated with the semantic information after time synchronization to generate micro-environment data that characterizes the spatial attributes and semantic features of each object, and the micro-environment data is uploaded to the mobile terminal.
[0009] Preferably, the method further includes: When a user's voice command is received, the mobile device generates a macroscopic path from the starting position to the destination based on the navigation environment situation map; Based on the macro path, determine the path distance between the user's current location and the destination; When the path distance is less than the threshold, the smart glasses acquire all visual information within the user's field of vision; When the smart glasses detect a destination identifier in the visual information, they obtain the three-dimensional coordinates of the destination based on the depth information of the destination; and then upload the three-dimensional coordinates of the destination to the mobile device. The mobile device determines the direction and distance to the destination based on the three-dimensional coordinates of the destination, and generates guidance instructions to guide the user to enter the destination.
[0010] Preferably, the method further includes: As the user moves along the macroscopic path, the mobile device analyzes the microscopic environmental data to obtain obstacle information in the direction of travel; Based on the obstacle information and the destination, a detour path to avoid the obstacles is generated on the navigation environment situation map; The detour route is parsed into navigation instructions and fed back to the user through bone conduction headphones in the smart glasses to guide the user to the destination.
[0011] Preferably, the method further includes: The smart glasses acquire the user's original orientation at their current position on the macroscopic path; Obtain the target orientation of the macroscopic path at the current location; If the original orientation deviates from the target orientation, the target orientation is fed back to the user through the bone conduction headphones and / or the vibration motor in the smart glasses, so that the user can maintain the correct direction of travel.
[0012] Preferably, the method further includes: When a user's voice command is received, the mobile device generates a macroscopic path from the starting position to the destination based on the navigation environment situation map; The navigation instructions corresponding to the macro path are divided into multiple navigation node instructions; The navigation node instructions are sequentially fed back to the user through bone conduction headphones in the smart glasses to guide the user to the destination.
[0013] Preferably, the method further includes: The mobile device obtains the user's voice query command, determines the user's intent based on the voice query command, obtains the recognition result of the central area of the field of vision through smart glasses based on the user's intent, and then broadcasts the recognition result by the mobile device.
[0014] Secondly, the present invention provides an intelligent guidance system for the blind based on multimodal data fusion, the intelligent guidance system being used to implement the intelligent guidance method described in the first aspect, comprising: The smart glasses and a mobile device for interacting with the smart glasses; the smart glasses include a camera, an edge AI chip, an inertial measurement unit, bone conduction headphones, and a motor.
[0015] Thirdly, the present invention provides an intelligent guide device for the blind based on multimodal data fusion, specifically comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the processor to perform the intelligent guide method for the blind based on multimodal data fusion in the first aspect.
[0016] Fourthly, the present invention also provides a non-volatile computer storage medium storing computer-executable instructions that are executed by one or more processors to perform the method described in the first aspect.
[0017] Compared with existing technologies, the beneficial effects of this invention are as follows: Real-time collection of raw data via smart glasses generates micro-environmental data characterizing the spatial attributes and semantic features of each object; the mobile terminal receives the micro-environmental information uploaded by the smart glasses, integrates it with local positioning data and macro-path data, and dynamically generates path planning and navigation instructions. This method, through collaborative processing by smart glasses and the mobile terminal, and the fusion of multi-modal data to generate navigation instructions, can effectively guide visually impaired individuals to their destinations, enhancing their independence in travel. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0019] Figure 1 This is a flowchart illustrating an intelligent guide method for the blind based on multimodal data fusion provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the method of step 104 provided in an embodiment of the present invention; Figure 3 This is a flowchart illustrating another intelligent guidance method for the blind based on multimodal data fusion provided in an embodiment of the present invention; Figure 4 This is a structural diagram of the embedded system inside the smart glasses provided in an embodiment of the present invention; Figure 5 This is a structural diagram of the different methods executed based on different voice intents provided in the embodiments of the present invention; Figure 6 A schematic diagram of the structure of an intelligent guide device for the visually impaired based on multimodal data fusion, provided in an embodiment of the present invention; The accompanying figure is labeled as follows: 21: Processor; 22: Memory. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0021] Unless the context otherwise requires, throughout the specification and claims, the term "comprising" is interpreted as openly inclusive, meaning "including, but not limited to." In the description of the specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "example," "specific example," or "some examples" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics mentioned may be included in any suitable manner in any one or more embodiments or examples; that is, although they may be incorporated into embodiments or examples using the above terms for reasons such as order and position, it does not limit them to be incorporated in combination by a single embodiment or example.
[0022] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more. Furthermore, for example, the description may use the prefix "A" or "B" to describe the same type of nouns as two independent entities. In this case, the corresponding features defined with "A" and "B" are used only to distinguish between similar entities and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features.
[0023] In the description of this invention, the expression “A and / or B” (where A and B are used to formally represent specific features) will be used. The corresponding expression includes the following three combinations: only A, only B, and a combination of A and B.
[0024] As used in this invention, “about,” “approximately,” or “approximately” includes the stated value and the average value within an acceptable range of deviation from a particular value, wherein the acceptable range of deviation is determined by a person skilled in the art taking into account the measurement under discussion and the error associated with the measurement of the particular quantity (i.e., the limitations of the measurement system).
[0025] Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0026] Example 1: The intelligent guidance method for the blind provided in this embodiment is executed collaboratively by smart glasses and a mobile terminal within an intelligent guidance system. The intelligent guidance system employs a layered computing architecture combining real-time perception from the smart glasses and intelligent decision-making from the mobile terminal. The smart glasses perform tasks highly sensitive to latency, such as obstacle detection, depth estimation, and posture tracking, ensuring millisecond-level safety responses. The mobile terminal handles tasks requiring greater computing power, network connectivity, and complex logic, such as long-distance path planning, natural language understanding, accessing large cloud-based model services, and user interaction management.
[0027] This embodiment provides an intelligent guidance system for the blind based on multimodal data fusion. The system includes smart glasses and a mobile terminal for interacting with the smart glasses. The smart glasses include a camera, an edge artificial intelligence (AI) chip, an inertial measurement unit, bone conduction headphones, and a motor. The smart glasses are the core sensing and execution unit of the intelligent guidance system.
[0028] In one embodiment, the smart glasses are constructed entirely of lightweight materials (e.g., titanium alloy). Through a highly integrated circuit design, all components are seamlessly integrated into the shape of the smart glasses, keeping the total weight under 80 grams. The power management system of the smart glasses is meticulously optimized to ensure at least 4 hours of continuous high-intensity use on a single charge, achieving ultimate real-time performance, lightweight design, and reliability. All components include a camera, an edge AI chip, an inertial measurement unit, bone conduction headphones, and a high-precision linear vibration motor. The camera can be a miniature camera with a wide-angle field of view (≥120°) and high dynamic range. The wide angle ensures sufficient environmental coverage, avoiding blind spots; the high dynamic range ensures clear and usable image quality even under drastic changes in lighting (such as entering or exiting tunnels or walking against the light). The edge AI chip built into smart glasses can be a low-power, high-performance embedded neural processing unit (NPU) or branch processing unit (BPU) chip (such as the Horizon Robotics Journey series) designed specifically for edge AI applications. This edge AI chip can provide several TOPS (Tera Operations Per Second) of computing power with a power consumption of only a few watts, processing video streams captured by cameras in real time, running depth estimation, object detection, and 3D modeling algorithms to achieve low-latency environmental perception and local path planning. A high-precision nine-axis inertial measurement unit (IMU) is also integrated into the smart glasses. The IMU outputs attitude data at a frequency of several hundred hertz. The IMU data is used not only for motion estimation in Simultaneous Localization and Mapping (SLAM) to enhance the robustness of localization (especially in scenarios lacking visual features), but also to implement value-added safety functions such as fall detection. Bone conduction headphones are precisely encapsulated in the temples of smart glasses near the user's temporal bone. They transmit sound signals directly to the auditory nerve through minute vibrations, providing clear sound without disturbing ambient noise. High-precision linear vibration motors are built into the left and right temples of the smart glasses, generating vibration patterns of varying intensities and frequencies. For example, a short vibration can indicate a warning, while a sustained, increasing vibration can indicate approaching danger, providing the user with a rich tactile experience.
[0029] Two-dimensional images are acquired through the camera on the smart glasses, and angular velocities and angular inertia are obtained by combining the images with the IMU on the smart glasses. Based on the two-dimensional images, angular velocities, and angular inertia, a local three-dimensional point cloud map of the user's surroundings is constructed in real time using Visual Simultaneous Localization and Mapping (VSLAM) technology. The three-dimensional point cloud map not only contains geometric information of obstacles but also labels semantic information of targets, thus distinguishing between "steps," "pedestrians," and "door signs," providing rich data for subsequent navigation and interaction.
[0030] like Figure 1 As shown, Embodiment 1 of the present invention provides an intelligent guidance method for the blind based on multimodal data fusion, which specifically includes the following steps: Step 101: The smart glasses collect and process all raw data to obtain structured data corresponding to each raw data.
[0031] In one embodiment, the raw data consists of angular velocity and acceleration measured by the IMU in the smart glasses, and images captured by the camera in the smart glasses. The structured data corresponding to the angular velocity and acceleration is the pose. After processing, the raw images yield structured data including the geometric information of the depth map and the semantic information of the objects. The camera acquires images with maximum efficiency and adds a precise timestamp to each frame to obtain the raw images.
[0032] In one embodiment, after the smart glasses acquire angular velocity, acceleration, and raw images, they process these data separately through parallel tasks to obtain structured data for each raw data set. Specifically, integration is performed based on the angular velocity and acceleration to obtain the pose of the observation point in three-dimensional space; the raw images are preprocessed to obtain a target image; depth estimation is performed on the target image based on the visual features of objects in the target image to obtain a depth map; geometric information from the depth map is extracted; and object detection and recognition are performed on the target image to obtain two-dimensional bounding boxes for all objects and semantic information corresponding to each object.
[0033] In one embodiment, the smart glasses perform integral calculations based on the angular velocity and acceleration, running a visual-inertial SLAM algorithm at a high frequency to obtain the pose of the observation point in three-dimensional space and construct a sparse feature point map. Before processing the original image, the original image can be preprocessed with distortion correction and noise reduction to obtain a target image. Based on the visual features of objects in the target image, depth estimation is performed to obtain a depth map. Specifically, the relative distance between objects in the target image and the observation point is calculated using the UniDepth (Universal Monocular Metric Depth Estimation) model, and the relative distance is correlated with the pixel values in the target image to achieve pixel-level depth estimation and generate a depth map. The geometric information of the depth map is then extracted. The YOLOv12 (You Only Look Once version 12) model is used to perform object detection and recognition on the target image to obtain two-dimensional bounding boxes of all objects and semantic information corresponding to each object. The geometric information includes depth values and spatial positions.
[0034] Step 102: The smart glasses integrate all structured data to obtain micro-environmental data that characterizes the spatial attributes and semantic features of each object, and then upload the micro-environmental data to the mobile terminal.
[0035] In one embodiment, the smart glasses perform multimodal fusion of all structured data to obtain micro-environment data characterizing the spatial attributes and semantic features of each object. This micro-environment data is then packaged into a predefined, efficient binary format and transmitted to the mobile device via Bluetooth Low Energy channel at a low frequency (e.g., 5-10Hz). This transmission method avoids the high bandwidth and high power consumption issues associated with transmitting raw video streams. The micro-environment data includes pose, a list of detected objects, and the corresponding 3D information and Optical Character Recognition (OCR) text for each object.
[0036] In one embodiment, the smart glasses fuse pose, depth map geometric information, and target detection semantic information to obtain micro-environment data representing the spatial attributes and semantic features of each object. Specifically, the pose, geometric information, and semantic information are synchronized in time. Time synchronization refers to adjusting different types of data (i.e., pose, geometric information, and semantic information) to the same or similar timestamps to ensure that different types of data describe information at the same moment or within a very short time. Based on the time-synchronized pose and geometric information, the depth values within the two-dimensional bounding box are back-projected and aggregated with the pose to generate a three-dimensional object containing three-dimensional coordinates and dimensions. Based on the three-dimensional coordinates of the three-dimensional object, the three-dimensional object is associated with the time-synchronized semantic information to generate micro-environment data representing the spatial attributes and semantic features of each object. That is, on the three-dimensional map constructed by SLAM (i.e., the sparse feature point map in step 101), the spatial location is matched, and the three-dimensional point cloud clusters corresponding to each object are labeled with semantic tags (such as "pedestrian" and "step").
[0037] Step 103: The mobile device combines the micro-environment data, local positioning data, and macro-path data to generate a navigation environment situation map.
[0038] In one embodiment, the micro-environment data refers to the environmental data within the field of view of the smart glasses worn by the user. Local positioning data refers to the mobile device's own Global Positioning System (GPS) and compass data, which can represent the user's current location information. Macro-path data refers to the macro-path data from the cloud map service. The smart glasses upload the micro-environment data from step 102 to the mobile device, and the mobile device merges and aligns these data from different sources and coordinate systems to form a unified, global navigation environment situation map.
[0039] Step 104: When a user's voice command is received, the mobile terminal generates a hybrid navigation command based on the navigation environment situation map and feeds it back to the user to guide the user to arrive at and enter the destination.
[0040] In one embodiment, when the mobile device receives a user's voice command, it generates a hybrid navigation command based on the navigation environment situation map in step 103, and guides the user to the destination based on the hybrid navigation command. The hybrid navigation command means that the navigation command generated by the mobile device can be fed back to the user in multiple forms (e.g., voice from bone conduction headphones in smart glasses, vibration from a vibration motor, or broadcast by an app on the mobile device).
[0041] In this embodiment, raw data is collected in real time through smart glasses, and micro-environmental data representing the spatial attributes and semantic features of each object is generated. The mobile terminal receives the micro-environmental information uploaded by the smart glasses, and combines it with multi-modal data such as local positioning data and macro-path data to dynamically generate path planning and navigation instructions. This method, through collaborative processing by smart glasses and mobile terminal and the fusion of multi-modal data to generate navigation instructions, can effectively guide visually impaired people to arrive at and enter their destination, enhancing their independence in travel.
[0042] Traditional navigation systems mostly focus on outdoor route navigation, guiding users from their starting point to their destination. Navigation is considered complete when the destination enters the user's visual range. However, traditional navigation systems lack the ability to accurately identify and locate entrances and exits to destinations. For visually impaired individuals, true navigation completion means guiding them safely and accurately to their destination. This embodiment provides a full-process navigation method from "planning a long route" to "accurately completing the last step." For example, when a user is going to a destination or looking for a specific store, the app on the mobile device first uses a macro navigation mode to create a macro path for the user. When the user approaches the destination or arrives at the entrance of a specific store, the app on the mobile device switches from macro navigation mode to micro navigation mode to plan a path for the user to enter the destination.
[0043] like Figure 2 As shown, the embodiment provides a method for step 104, which specifically includes the following steps: Step 201: When a user's voice command is received, the mobile terminal generates a macroscopic path from the starting position to the destination based on the navigation environment situation map.
[0044] In one embodiment, for example, the voice command is to navigate to the nearest subway station. The mobile voice assistant recognizes the voice intent as navigation, and through point-of-interest search, determines the destination as the nearest subway station, and generates a macro-path from the user's current location to the subway station.
[0045] Step 202: Determine the path distance between the user's current location and the destination based on the macro path.
[0046] In one embodiment, path distance refers to the actual path length required to reach the destination from the user's current location along the planned navigation route.
[0047] Step 203: When the path distance is less than the threshold, the smart glasses acquire all visual information within the user's field of vision.
[0048] In one embodiment, the threshold can be set according to the actual situation. When the path distance is less than the threshold, the computing unit of the smart glasses activates a high-precision OCR model and scans all visual information within the user's field of vision using OCR technology. For example, if the user's destination is Starbucks, when the user enters a 50-meter radius around Starbucks, the smart glasses begin scanning all store signs within their field of vision.
[0049] Step 204: When the smart glasses detect a destination identifier in the visual information, the three-dimensional coordinates of the destination are obtained based on the depth information of the destination; the three-dimensional coordinates of the destination are uploaded to the mobile device.
[0050] In one embodiment, when the smart glasses detect a destination identifier in the visual information, the smart glasses analyze microscopic environmental data to obtain the depth information of the destination, and calculate the three-dimensional coordinates of the destination based on the depth information. For example, if OCR identifies and matches a Starbucks, the smart glasses combine the depth information of Starbucks to calculate the three-dimensional coordinates of Starbucks.
[0051] Step 205: The mobile terminal determines the direction and distance to enter the destination based on the three-dimensional coordinates of the destination, generates guidance instructions and feeds them back to the user through the bone conduction headphones in the smart glasses to guide the user to enter the destination.
[0052] In one embodiment, the mobile device determines the user's direction and distance from the current location to the destination based on the 3D coordinates of the destination uploaded by the smart glasses and a navigation environment map. This direction and distance are then converted into voice and fed back to the user via bone conduction headphones in the smart glasses. For example, the bone conduction headphones may contain voice prompts indicating the direction and distance (e.g., "The target is about 15 meters to your right; please continue along the current path," or "The target is now to your right; there are two steps at the entrance; please be careful") to guide the user until they are safely inside the store.
[0053] In one embodiment, considering the special circumstances of visually impaired individuals who have limited access to visual information and find it difficult to quickly understand complex route maps or simultaneously memorize multiple navigation commands like ordinary users, the system adopts a segmented guidance approach. Instead of presenting the entire macro-path to the user at once, the system breaks down the macro-path into a series of key navigation nodes (such as turning points and street crossings). When a user's voice command is received, the mobile terminal generates a macro-path from the starting point to the destination based on the navigation environment map; the navigation commands corresponding to the macro-path are divided into multiple navigation node commands; and these navigation node commands are sequentially fed back to the user through bone conduction headphones in the smart glasses to guide the user to the destination.
[0054] Before a user proceeds to the next navigation node, the intelligent navigation system activates a scenario-specific AI model to scan the environment, ensuring safe movement. Specifically, when the mobile device determines, via GPS and pedometer algorithms, that the user is approaching the next navigation node (e.g., a pedestrian crossing), it automatically sends a scene-switching command to the smart glasses. Upon receiving the command, the smart glasses' computing unit loads and activates the scenario-specific AI model for scene scanning. For example, at a pedestrian crossing, the intelligent navigation system activates a traffic light recognition model and a vehicle detection model, using these models to detect road conditions. The system then relays this information to the user via bone conduction headphones (e.g., "You have reached a pedestrian crossing; we are detecting road conditions for you," "The light is currently red," and "A vehicle is approaching from the left"). In addition, the smart glasses upload road condition information to the mobile device in real time. The decision engine on the mobile device combines the above road condition information to make a judgment. Only when the road condition is determined to be safe will it issue a clear passage instruction to guide the user to proceed (e.g., the light is green, the road is safe, and you can proceed). At the same time, the smart glasses can also emit a specific vibration pattern of "you can proceed" through the vibration motor to guide the user to proceed.
[0055] As users navigate along the macro-path, temporary or dynamic obstacles may appear, such as parked vehicles, construction barriers, temporary stalls, or gatherings of pedestrians. These obstacles may hinder the user's progress, so it is necessary to promptly replan detour routes for the user and provide clear turning and obstacle avoidance guidance to ensure a smooth and safe navigation process.
[0056] like Figure 3 As shown in the figure, this embodiment provides another intelligent guidance method for the blind based on multimodal data fusion, which specifically includes the following steps: Step 301: As the user travels along the macroscopic path, the mobile terminal analyzes the microscopic environmental data to obtain obstacle information in the direction of travel.
[0057] In one embodiment, as the user travels along the macroscopic path, the intelligent guidance system continuously uploads microscopic environmental data collected and processed in real time by the smart glasses, so that the mobile device can combine the microscopic environmental data to correct the macroscopic path and achieve true dynamic obstacle avoidance.
[0058] Furthermore, any static obstacle (such as a utility pole or trash can) or dynamic obstacle (such as a pedestrian or bicycle) that enters the user's safe distance (e.g., within 3 meters) will trigger a directional tactile vibration warning. At this time, the mobile device does not intervene; the smart glasses handle this high-frequency, low-latency obstacle avoidance entirely locally, ensuring the user's safety. The directional tactile vibration warning provides different information to the user through vibration direction and frequency. For example, if there is an obstacle on the left, the left temple of the smart glasses vibrates; if there is an obstacle on the right, the right temple vibrates; the continuously increasing vibration indicates that the user is getting closer to the obstacle.
[0059] Step 302: Based on the obstacle information and the destination, generate a detour path to avoid the obstacles on the navigation environment situation map.
[0060] In one embodiment, the mobile device continuously projects local obstacle information reported by the smart glasses onto the macroscopic path. If the macroscopic path is found to be blocked, the navigation engine activates a local replanning module. This local replanning module, based on a heuristic search algorithm or a dynamic window algorithm, quickly searches for the shortest and safest detour path within a small area on the navigation environment situation map. The microscopic environment data includes local obstacle information.
[0061] Step 303: The detour route is parsed into navigation instructions and fed back to the user through the bone conduction headphones in the smart glasses to guide the user to the destination.
[0062] In one embodiment, the mobile device parses the detour path in step 302 into navigation instructions and uploads them to the smart glasses. The smart glasses guide the user to bypass the current obstacle through bone conduction headphones, and then continue to travel on the macro path to reach the destination. For example, the voice instruction in the bone conduction headphones is "Please take 3 steps to the left front, bypass the obstacle, and then return to the original route".
[0063] As the user moves along the macroscopic path, the smart glasses collect data using a built-in IMU and visual odometry to prevent the user from deviating from the correct direction. Specifically, the smart glasses acquire the user's original orientation at their current position on the macroscopic path; and acquire the target orientation of the macroscopic path at that current position. If the original orientation deviates from the target orientation, the target orientation is fed back to the user via bone conduction headphones and / or vibration motors in the smart glasses to help the user maintain the correct direction of travel. For example, when the user deviates from the predetermined heading (e.g., turning more than 15 degrees), the bone conduction headphones emit a gentle prompt (such as "Please adjust slightly to the left"), while the temples in the corresponding direction emit a slight tactile vibration to help the user maintain the correct direction of travel. Providing auditory-tactile multimodal feedback through bone conduction headphones and directional vibration motors delivers rich information without hindering the user's reception of ambient sounds.
[0064] See Figure 4 This embodiment provides an embedded system within smart glasses. Specifically, the software architecture of the smart glasses' internal system includes a driver and data acquisition layer, a perception and computing core layer, and a decision and communication layer. The driver and data acquisition layer is responsible for directly interacting with hardware such as cameras and IMUs to capture raw data (i.e., angular velocity, acceleration, and images in step 101) with maximum efficiency, and to accurately timestamp each frame of data, providing a foundation for subsequent data fusion. The perception and computing core layer is the "brain" of the glasses. When the raw data stream enters, it triggers a parallel multi-task processing pipeline. Specifically, the visual front end performs preprocessing such as distortion correction and noise reduction on the images, and distributes the processed images to the SLAM thread, depth estimation thread, and object detection thread for subsequent processing. The SLAM thread runs the visual-inertial SLAM algorithm at a high frequency to calculate the precise pose of the glasses in three-dimensional space in real time and construct a sparse feature point map. The depth estimation thread inputs the processed image into the UniDepth model to generate a pixel-level dense depth map; the object detection thread inputs the processed image into the YOLOv12 model to identify all predefined targets and their 2D bounding boxes; the perception and computation core layer also includes a 3D reconstruction and semantic fusion thread, which fuses the pose of SLAM, the geometric information of the depth map, and the semantic information of object detection to obtain micro-environment data. The decision and communication layer is responsible for local rapid decision-making and structured data encapsulation and reporting. The local rapid decision-making layer performs rapid safety assessments based on the constructed semantic 3D map. For example, by analyzing the point cloud density and semantic labels in the user's direction of travel, it determines whether there is an immediate collision risk. Once detected, it bypasses the phone and directly triggers the vibration motor in the corresponding direction, achieving an instinctive warning with minimal latency. The structured data encapsulation and reporting layer packages the micro-environment data into a predefined, efficient binary format and sends it to the mobile app at a low frequency (e.g., 5-10Hz) via Bluetooth Low Energy channel, avoiding the high bandwidth and high power consumption problems caused by transmitting the original video stream.
[0065] The non-invasive auditory and tactile feedback provided in this embodiment effectively reduces the information load on a single sense, improves the recognizability and response speed of navigation commands, and enables users to accurately receive guidance information even in complex or noisy environments, thereby enhancing navigation reliability and user experience. Specifically, voice commands are transmitted through bone conduction headphones, enabling clear voice prompts without blocking the ear canal, ensuring that users always maintain awareness of surrounding sounds (such as vehicle horns and other people calling out), improving safety and situational awareness during navigation. Directional vibration motors on the temples of the glasses provide intuitive and rapid obstacle avoidance warnings. This intuitive feedback—vibrating to the left if there is an obstacle on the left, and to the right if there is an obstacle on the right—is more efficient than voice commands in noisy environments or emergencies.
[0066] In addition to navigation needs, users also have certain needs in information retrieval and emergency assistance. The mobile terminal provided in this embodiment integrates an advanced speech recognition application programming interface (API) and a natural language understanding module, and reserves an interface for connection with a cloud-based large language model (such as Tongyi Qianwen). When the mobile terminal receives a user's voice command, the speech recognition module converts the voice command into text, and combines it with image and / or text information captured by the smart glasses, sending it to the large model for analysis to determine the intent, categorizing it as navigation, environmental information query, or emergency assistance.
[0067] Specifically, the system parses user voice commands. If the command is intended for navigation, a hybrid navigation command is generated using steps 101-104. If the command is intended for environmental information query, descriptive voice feedback is generated by integrating the navigation environment situation map, environmental text information, and map service data. For complex environmental information queries (such as "What's good to eat near me?"), the mobile device not only calls the Point of Interest (POI) search function of the map service but also performs semantic processing on the results to present them in a more natural conversational manner. For visual question answering (such as "Can you tell me the color of this dress?"), the mobile device packages the real-time image transmitted by the glasses and the question text together and sends it to the large language model in the cloud via API. After receiving the model's answer, it is then read aloud to the user through a text-to-speech engine. If the command is intended for emergency assistance, the system automatically executes the emergency call process in response to the received emergency assistance signal. The mobile device provides a settings interface that allows users to customize the level of detail in the voice broadcast, the intensity of the vibration feedback, and emergency contacts. It can also manage emergency call logic. Upon receiving a user's request for help or a fall signal uploaded by the glasses, it will immediately execute the operations of making a phone call and sending a location SMS, providing the user with ultimate safety assurance. Through the above design, this embodiment provides an intelligent, reliable, and user-friendly intelligent guidance system for the visually impaired, using the smart glasses as a keen sentinel and the mobile terminal as the intelligent brain. The two work closely together to safeguard the independent travel of visually impaired users.
[0068] See Figure 5 When the user's voice command is "Go to Optics Valley World City", the voice assistant analyzes the information and opens outdoor navigation; when the user's voice command is "Where is Sunshine Barber Shop?", the voice assistant uses speech recognition to obtain the keyword "Sunshine Barber Shop", then calls the visual module to locate the shop and outputs the result (e.g., 10 meters ahead) to the user; when the user's voice command is "Contact volunteers", the voice assistant performs speech recognition, then calls the core processing module to make an emergency judgment and initiates a phone call to volunteers; when the user's voice command is "Open Baidu Maps", the AI assistant performs speech recognition and opens the map.
[0069] In this embodiment, the system employs an advanced deep learning model that exhibits excellent robustness to different lighting conditions and scenarios, and can recognize semantic information such as door signs and shop names, achieving a leap from physical obstacle avoidance to semantic navigation. Deploying the core perception algorithm on the edge AI chip of the glasses enables end-to-end low-latency processing, resulting in a response speed to dynamic obstacles far exceeding traditional solutions and significantly improving safety. The system requires only a single-lens camera as the main sensor, significantly reducing hardware costs. The highly integrated design makes the glasses lightweight and comfortable, suitable for extended wear.
[0070] For details on the methods for querying environmental information and providing emergency assistance, please refer to Examples 2 and 3.
[0071] Example 2: This embodiment demonstrates an intelligent guidance system for the visually impaired as an information probe for users. Through multimodal perception and real-time data processing technology, it can provide users with environmental information at different depths as needed. This environmental information includes basic safety information, on-demand queryable surrounding information, and immersive overall environmental overview information. Specifically, the mobile terminal acquires the user's voice query command, determines the user's intent based on the voice query command, and obtains the recognition result of the central visual area through smart glasses based on the user's intent. The mobile terminal then broadcasts the recognition result.
[0072] In one embodiment, basic safety information is enabled by default in any mode. The smart glasses can recognize objects, ground materials, and suspended objects in three-dimensional space. This highest priority safety information is immediately fed back through voice and touch, and the basic safety information is broadcast. This basic safety information includes recognition of ground conditions (e.g., "There is a sinking manhole cover ahead," "You are about to walk on the tactile paving") and recognition of obstacles in headspace (e.g., "Beware of low-hanging tree branches ahead").
[0073] In addition, the system supports on-demand queries for surrounding information. Specifically, users can make instant queries at any time while walking using voice commands (queries include objects, text, and public facilities). When a user wants to query an object (e.g., "What's in front of me?"), the smart glasses upload the recognition results of the central area of their field of vision (e.g., {Category: 'Bench'}) to the mobile device, which then announces: "There is a bench 2 meters ahead of you." When a user queries text (e.g., "Read this road sign"), the smart glasses run OCR, upload the recognized text to the mobile device, and the mobile device announces: "The road sign says 'Zhongshan Road'." When a user queries public facilities (e.g., "Are there any restrooms nearby?"), the mobile device's voice assistant interprets the intent and instructs the glasses' vision module to search for internationally recognized restroom signs within the field of vision. Simultaneously, the mobile device also utilizes the POI search function of the map service, combining information from both sources to provide the most reliable answer.
[0074] When a user arrives at an unfamiliar but relatively safe environment (such as a park or lobby), they can issue the command "Describe your surroundings." At this point, the system enters "Environment Overview" mode. The user can slowly rotate their body in place, and the system continuously scans the 360-degree environment, uploading all identified key objects and their relative positions to the mobile app. The app's scene-building engine then integrates this fragmented information into a coherent and orderly description, such as: "You are now in the center of a lobby. Directly in front of you is the service desk, to your right are several rows of seats, to your left is an escalator leading to the second floor, and behind you is the entrance door." This helps users quickly establish a spatial understanding of the new environment.
[0075] Example 3: This embodiment ensures that users can obtain assistance in the most efficient and reliable way in any emergency. The method provided in this embodiment helps users by identifying the emergency situation in their own state.
[0076] Users can dial pre-set emergency contact numbers with a single click using preset voice commands (such as "emergency call") or the "call volunteers" button on the app. Optionally, the IMU data in the glasses can be used to automatically send a distress message containing GPS location to emergency contacts in the event of an accident, based on a fall detection algorithm. Specifically, when a user feels unwell or is lost, they can proactively ask for help by saying a preset emergency wake-up word, such as "emergency call" or "help." The app's voice recognition module prioritizes these keywords. Once recognized, it skips the usual intent understanding and directly triggers the emergency call process. The system immediately begins dialing the user's pre-set emergency contact list until someone answers. To handle extreme situations where the user cannot speak or voice recognition fails, the smart glasses feature a separate, tactilely-sensible physical distress button on the temple. Specifically, the user presses and holds this button for 3 seconds. Upon detecting the long press signal, the glasses' embedded system immediately sends a highest-priority emergency distress command to the app via Bluetooth. Upon receiving this command, the app also immediately initiates the phone call process. During the user's movement, the glasses' IMU sensor continuously monitors the user's posture data; a fall detection algorithm runs within the edge computing module; this algorithm analyzes abnormal and drastic changes in acceleration and angular velocity to determine if the user has fallen. Once a suspected fall is detected, the system will first ask via voice: "Are you alright? Do you need help?" If no "cancel" voice command or physical button operation is received from the user within 15 seconds, the system will automatically trigger the emergency assistance process, making a phone call and simultaneously sending an emergency text message containing the user's last known GPS location to all emergency contacts.
[0077] Example 4: Based on the intelligent guidance method for the blind based on multimodal data fusion provided in the foregoing embodiments, the present invention also provides an intelligent guidance device for the blind based on multimodal data fusion that can be used to implement the above method, such as... Figure 6 The diagram shown is a schematic representation of the device architecture according to an embodiment of the present invention. The intelligent guide device for the visually impaired based on multimodal data fusion in this embodiment includes one or more processors 21 and a memory 22. Figure 6 Take a processor 21 as an example.
[0078] Processor 21 and memory 22 can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.
[0079] The memory 22, as a non-volatile computer-readable storage medium for a multimodal data fusion-based intelligent guidance method for the blind, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the multimodal data fusion-based intelligent guidance method in the aforementioned embodiments. The processor 21 executes various functional applications and data processing of the multimodal data fusion-based intelligent guidance device by running the non-volatile software programs, instructions, and modules stored in the memory 22, thereby realizing the multimodal data fusion-based intelligent guidance method of the aforementioned embodiments.
[0080] Memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 22 may include memory remotely located relative to processor 21, which can be connected to processor 21 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0081] The program instructions / modules are stored in memory 22. When executed by one or more processors 21, they execute the intelligent guidance method for the blind based on multimodal data fusion described in the foregoing embodiments, for example, executing the methods described above. Figures 1-3 The steps shown.
[0082] This invention also provides a non-volatile computer storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 6 One of the processors 21 can enable the above-described one or more processors to execute the intelligent guidance method for the blind based on multimodal data fusion in the foregoing embodiments, for example, to perform the above-described... Figures 1-3 The steps shown.
[0083] It is worth noting that the information interaction and execution process between the modules and units in the above-mentioned device and system are based on the same concept as the processing method embodiment of the present invention. For details, please refer to the description in the method embodiment of the present invention, and will not be repeated here.
[0084] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0085] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A smart guidance method for the blind based on multimodal data fusion, characterized in that, The method is executed collaboratively by smart glasses and a mobile device, including: The smart glasses collect and process all raw data to obtain structured data corresponding to each raw data. The smart glasses integrate all structured data to obtain micro-environmental data that characterizes the spatial attributes and semantic features of each object, and then upload the micro-environmental data to the mobile device. The mobile device combines the aforementioned micro-environmental data, local positioning data, and macro-path data to generate a navigation environment situation map; When a user's voice command is received, the mobile device generates a hybrid navigation command based on the navigation environment situation map and feeds it back to the user to guide the user to arrive at and enter the destination.
2. The intelligent guidance method for the blind based on multimodal data fusion according to claim 1, characterized in that, The raw data includes angular velocity, acceleration, and raw images; the structured data includes pose, geometric information, and semantic information; and the method further includes: The pose of the observation point in three-dimensional space is obtained by integrating the angular velocity and the acceleration. The original image is preprocessed to obtain the target image; Based on the visual features of objects in the target image, depth estimation is performed on the target image to obtain a depth map; and geometric information of the depth map is extracted. Object detection and recognition are performed on the target image to obtain the two-dimensional bounding boxes of all objects and the semantic information corresponding to each object.
3. The intelligent guidance method for the blind based on multimodal data fusion according to claim 1, characterized in that, The structured data includes pose, geometric information, and semantic information. The process of fusing all structured data to obtain micro-environmental data characterizing the spatial attributes and semantic features of each object, and uploading this micro-environmental data to the mobile device, includes: Synchronize the pose, geometric information, and semantic information in time; Based on the pose and geometric information after time synchronization, the depth values within the 2D bounding box are back-projected and aggregated with the pose after time synchronization to generate a 3D object containing 3D coordinates and dimensions. Based on the three-dimensional coordinates of the three-dimensional object, the three-dimensional object is associated with the semantic information after time synchronization to generate micro-environment data that characterizes the spatial attributes and semantic features of each object, and the micro-environment data is uploaded to the mobile terminal.
4. The intelligent guidance method for the blind based on multimodal data fusion according to claim 1, characterized in that, The method further includes: When a user's voice command is received, the mobile device generates a macroscopic path from the starting position to the destination based on the navigation environment situation map; Based on the macro path, determine the path distance between the user's current location and the destination; When the path distance is less than the threshold, the smart glasses acquire all visual information within the user's field of vision; When the smart glasses detect a destination identifier in the visual information, they obtain the three-dimensional coordinates of the destination based on the depth information of the destination; and then upload the three-dimensional coordinates of the destination to the mobile device. The mobile device determines the direction and distance to the destination based on the three-dimensional coordinates of the destination, and generates guidance instructions to guide the user to enter the destination.
5. The intelligent guidance method for the blind based on multimodal data fusion according to claim 4, characterized in that, The method further includes: As the user moves along the macroscopic path, the mobile device analyzes the microscopic environmental data to obtain obstacle information in the direction of travel; Based on the obstacle information and the destination, a detour path to avoid the obstacles is generated on the navigation environment situation map; The detour route is parsed into navigation instructions and fed back to the user through bone conduction headphones in the smart glasses to guide the user to the destination.
6. The intelligent guidance method for the blind based on multimodal data fusion according to claim 4, characterized in that, The method further includes: The smart glasses acquire the user's original orientation at their current position on the macroscopic path; Obtain the target orientation of the macroscopic path at the current location; If the original orientation deviates from the target orientation, the target orientation is fed back to the user through the bone conduction headphones and / or the vibration motor in the smart glasses, so that the user can maintain the correct direction of travel.
7. The intelligent guidance method for the blind based on multimodal data fusion according to claim 1, characterized in that, The method further includes: When a user's voice command is received, the mobile device generates a macroscopic path from the starting position to the destination based on the navigation environment situation map; The navigation instructions corresponding to the macro path are divided into multiple navigation node instructions; The navigation node instructions are sequentially fed back to the user through bone conduction headphones in the smart glasses to guide the user to the destination.
8. The intelligent guidance method for the blind based on multimodal data fusion according to any one of claims 1-7, characterized in that, The method further includes: The mobile device obtains the user's voice query command, determines the user's intent based on the voice query command, obtains the recognition result of the central area of the field of vision through smart glasses based on the user's intent, and then broadcasts the recognition result by the mobile device.
9. An intelligent guidance system for the blind based on multimodal data fusion, the intelligent guidance system being used to implement the intelligent guidance method as described in any one of claims 1-8, characterized in that, include: The smart glasses and a mobile device for interacting with the smart glasses; the smart glasses include a camera, an edge AI chip, an inertial measurement unit, bone conduction headphones, and a motor.
10. A smart guidance device for the blind based on multimodal data fusion, characterized in that, The device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the processor for performing the intelligent guide method for the blind based on multimodal data fusion as described in any one of claims 1-8.