A blind guiding robot interactive navigation system combining vision, inertial navigation and voice

By integrating visual, inertial navigation, and voice data, adaptive navigation and interaction with guide devices are achieved, solving the problem that existing guide devices cannot be personalized and context-aware in complex environments, and improving the travel safety and interactive experience of visually impaired people.

CN121346774BActive Publication Date: 2026-02-24SHANDONG SAIFEITE SAFETY ENG TECH DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511914576.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-02-24
Estimated Expiration
2045-12-18

AI Technical Summary

Technical Problem

Existing guide devices for the visually impaired struggle to provide personalized, context-aware, and adaptive guidance in complex environments, failing to effectively understand users' behavioral states and underlying intentions, leading to feelings of unease among visually impaired individuals during their travels.

Method used

By combining visual, inertial navigation, and voice data, and through multimodal perception data acquisition, context-user joint state feature extraction, user potential intent inference, and adaptive guidance strategy generation, adaptive navigation paths and interaction commands are generated, enabling proactive understanding and dynamic adjustment of user intent.

Benefits of technology

It improves the safety and interactive experience of visually impaired people when traveling in complex environments, reduces cognitive burden, enhances navigation autonomy and naturalness of interaction, provides rich multimodal feedback, and reduces anxiety and potential dangers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121346774B_ABST
    Figure CN121346774B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of blind guiding robot interactive navigation, and particularly discloses a blind guiding robot interactive navigation system combining vision, inertial navigation and voice, wherein the blind guiding robot interactive navigation system synchronously acquires environment vision, motion and sound data through a multi-modal perception data acquisition module, analyzes the data through a context-user joint state feature extraction module, combines a user potential intention deduction module, predicts the user intention, generates an adaptive guiding strategy, adjusts a navigation path according to the strategy through an adaptive guiding path optimization module, ensures safety and efficiency, and finally provides voice and tactile feedback through a situational multi-modal interactive instruction output module, thereby enhancing the interactive experience and navigation autonomy. The application solves the travel problem of the visually impaired in a complex environment, improves the individualization, situational perception ability and safety of blind guiding, and brings more natural and efficient navigation assistance to the visually impaired.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interactive navigation technology for guide robots, and more specifically, to an interactive navigation system for guide robots that combines vision, inertial navigation, and voice. Background Technology

[0002] Currently, visually impaired individuals face numerous challenges when traveling independently, especially in complex and ever-changing indoor and outdoor environments. Traditional guide tools and assistive devices often only provide basic obstacle avoidance or pre-defined route guidance, struggling to cope with sudden changes in circumstances or understand the subtle needs and emotions of users. For example, in crowded areas or when encountering environmental elements with specific contextual significance, such as warning signs on slippery surfaces, existing systems typically fail to provide appropriate and differentiated guidance, causing users to feel uneasy or unable to respond effectively. This lack of contextual awareness and personalized response capability is a core problem that current guide technology urgently needs to address.

[0003] Patent CN111142536B discloses an indoor guide robot for the blind, comprising: a robot terminal that collects laser ranging data, destination data, pose data, surrounding information data, and odometer data, and uploads the collected data to a guide decision-making layer. The robot responds to control information from the guide decision-making layer, moving to guide the blind person's movement. The guide decision-making layer includes four main functions: autonomous mapping, autonomous localization, intelligent guidance, and human-computer interaction. This invention introduces ROS and SLAM; it also incorporates voice interaction and a vibrating cane to improve the user experience. It assists the blind person to reach their destination faster and better through mapping, relocalization, path planning, and motion control; the integrated vibrating interactive handle and voice interaction function enhance the human-computer interaction experience.

[0004] The shortcomings of existing technologies lie in their relatively rudimentary and passive human-computer interaction, despite achieving SLAM-based indoor navigation, voice interaction, and vibration feedback. They primarily execute explicit user commands and announce road conditions, lacking effective understanding and adaptive responses to user behavior and potential intentions in complex environments. This invention addresses this issue by fusing visual, inertial navigation, and voice multimodal data. It not only perceives the physical environment but also extracts the user's involuntary physiological feedback and voice features, thereby deducing the user's potential intentions and dynamically generating adaptive guidance strategies and optimized navigation paths. This elevates the robot from a mere command-following tool into an intelligent partner capable of understanding user states and proactively adjusting interaction modes and information detail, thus enhancing the safety, personalized experience, and naturalness of human-computer interaction for the blind. Summary of the Invention

[0005] In view of this, in order to solve the problems mentioned in the background technology, a visual, inertial navigation and voice interactive navigation system for guide robots is proposed.

[0006] The objective of this invention can be achieved through the following technical solution: This invention provides an interactive navigation system for a guide robot that combines vision, inertial navigation, and voice, including: a multimodal perception data acquisition module, which synchronously acquires dynamic visual images, six-degree-of-freedom motion data, and environmental sound data, and marks the dynamic visual images, six-degree-of-freedom motion data, and environmental sound data with a unified timestamp to generate a synchronous multimodal perception data packet.

[0007] The context-user joint state feature extraction module parses synchronous multimodal perception data packets, extracts environmental semantic elements, involuntary physiological feedback features, and emotional speech features, and associates these elements with environmental semantic elements, involuntary physiological feedback features, and emotional speech features within a preset time window based on a unified timestamp to generate a context-user joint state feature set.

[0008] The user latent intent inference module predicts user latent intents and generates user latent intent labels based on the context-user joint state feature set and an intent inference model.

[0009] The adaptive guidance strategy generation module matches and activates adaptive guidance strategies from a preset strategy library based on the user's potential intent tags.

[0010] The adaptive guidance path optimization module obtains the robot's original navigation path and adjusts it according to the rules defined by the adaptive guidance strategy to generate an adaptive guidance path.

[0011] The contextualized multimodal interaction instruction output module executes an adaptive guidance path and, in conjunction with an adaptive guidance strategy, synthesizes contextualized multimodal interaction instructions.

[0012] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: (1) The present invention aims to improve the travel safety and interactive experience of visually impaired people in complex environments. By integrating multimodal perception data and intelligent inference, the system can obtain environmental semantics and the user's internal state in real time, thereby providing more accurate and proactive risk avoidance and situational adaptation guidance, which helps to reduce the user's anxiety and potential dangers.

[0013] (2) This invention realizes personalized and context-aware intelligent guidance for the blind. By deeply analyzing the visual, inertial navigation and voice data collected simultaneously, the system can identify the user's potential behavioral intentions and emotional expressions, and dynamically adjust the guidance strategy accordingly, so that the robot can be transformed from a single instruction executor into an intelligent partner that understands the user and actively cooperates, which helps to enhance the naturalness and effectiveness of the interaction.

[0014] (3) This invention helps reduce the cognitive burden on visually impaired individuals and enhance their navigation autonomy. The system can optimize navigation paths and information output methods according to the user's current intentions. For example, it can provide flexible collaborative suggestions during exploration, simplify instructions and shield secondary information when confused. This adaptive information presentation method makes it easier for users to understand and follow the guidance, while retaining a moderate degree of freedom of exploration.

[0015] (4) This invention provides a rich and intuitive interactive feedback through multimodal fusion output. By combining context-appropriate natural language instructions with the directional tactile vibration of the guide handle, it brings users a three-dimensional and clear guidance experience. This synergistic effect of touch and hearing helps to enhance the indicativeness and immersion of the guidance, thereby helping to build users' trust in the guide system. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the system module structure connection of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Please see Figure 1 This invention provides an interactive navigation system for a guide robot that combines vision, inertial navigation, and voice, including: a multimodal perception data acquisition module, a context-user joint state feature extraction module, a user potential intent inference module, an adaptive guidance strategy generation module, an adaptive guidance path optimization module, and a contextualized multimodal interaction command output module.

[0020] The multimodal perception data acquisition module is connected to the context-user joint state feature extraction module, the context-user joint state feature extraction module is connected to the user latent intention inference module, the user latent intention inference module is connected to the adaptive guidance strategy generation module, the adaptive guidance strategy generation module is connected to the adaptive guidance path optimization module, and the adaptive guidance path optimization module is connected to the contextualized multimodal interaction instruction output module.

[0021] The multimodal sensing data acquisition module synchronously acquires dynamic visual images, six-degree-of-freedom motion data, and ambient sound data, and marks the dynamic visual images, six-degree-of-freedom motion data, and ambient sound data with a unified timestamp to generate synchronous multimodal sensing data packets.

[0022] In a specific embodiment of the present invention, generating a synchronous multimodal perception data packet includes: activating the robot's vision sensor and acquiring dynamic visual images containing environmental objects and spatial relationships.

[0023] Simultaneously activate the inertial measurement unit built into the guide handle to capture the user's gait rhythm, body posture, and six degrees of freedom motion data of hand control during walking.

[0024] The system uses a microphone array to acquire user voice commands and background ambient sounds, and combines the user voice commands and background ambient sounds as ambient sound data.

[0025] The acquired dynamic visual images, six-degree-of-freedom motion data, and environmental sound data are labeled with unified timestamps and packaged to generate synchronous multimodal sensing data packets.

[0026] It should be noted that the goal of this step is to generate a synchronous multimodal sensing data package by simultaneously acquiring data from multiple sensors, providing a raw information foundation for subsequent analysis. The implementation process begins by starting and running the various sensors installed on the guide robot and its handle. Specifically, the central processing unit sends an activation command to the vision sensor, which is configured to continuously capture video streams of the environment in front of the robot, forming dynamic visual images. Simultaneously, a synchronization signal is sent to the inertial measurement unit built into the guide handle, which then begins recording the user's hand movements and posture changes at a frequency of 100 Hz, generating six degrees of freedom motion data. At the same time, the four-microphone circular array on the robot is also activated, continuously recording all sounds in the surrounding environment, including the user's voice commands and background ambient sounds. After data acquisition is complete, the system adds a unified and millisecond-accurate timestamp to each frame of dynamic visual image, each set of six-degree-of-freedom motion data, and each segment of ambient sound data. This timestamp originates from a unified system clock, ensuring strict time alignment of data from different sources. Finally, these dynamic visual images, six-degree-of-freedom motion data, and ambient sound data with a unified timestamp are integrated and packaged into a structured data set, namely the synchronous multimodal sensing data package. Its data structure is a composite data type containing a master timestamp field and three data payload fields that store dynamic visual images, six-degree-of-freedom motion data, and ambient sound data respectively, which serves as the input for subsequent steps.

[0027] The synchronous multimodal sensing data packet is a data structure that serves as a collection of time-synchronized visual, motion, and sound information. Its data structure is a composite data type containing a master timestamp field and three data payload fields. The three data payload fields store dynamic visual images, six-degree-of-freedom motion data, and ambient sound data, respectively. Dynamic visual images refer to a continuous video data stream used to reflect objects, layouts, and changes in the environment. Its data attributes are H.264 encoding format, a resolution of 1920x1080 pixels, and a frame rate of 30 frames per second (fps) RGB video stream. An inertial measurement unit (IMU) is an electronic device that measures and reports the velocity, orientation, and acceleration caused by gravity of an object. It integrates a three-axis accelerometer and a three-axis gyroscope, enabling it to capture motion and rotation in three-dimensional space. Six-degree-of-freedom motion data refers to a dataset describing the complete motion state of an object in three-dimensional space. Its data structure is a vector containing six floating-point values, representing translational motion along the X, Y, and Z axes and rotational motion (roll, pitch, yaw) around these three axes. The data sampling frequency is set to 100 Hz, based on the analysis of hand micro-movement frequencies of 200 users of different ages in daily walking scenarios. 100 Hz can fully capture intentional operations and unconscious tremors. A microphone array is a system composed of multiple microphones arranged in a specific geometric shape. Its function is to enhance the pickup of sound sources from specific directions and suppress background noise through multi-microphone signal processing technology. This solution uses a four-microphone ring array. Ambient sound data refers to unprocessed raw audio waveform data, recording all sound information present in the environment. Its data attributes are an audio stream in pulse-code modulation format with a sampling rate of 48 kHz and a bit depth of 16 bits.

[0028] For example, the system activates all sensors and begins generating a synchronous multimodal perception data packet. During the time interval from T0 to T0+10 seconds, the vision sensor acquires a 10-second dynamic visual image of the intersection's traffic lights, zebra crossing, and passing pedestrians and vehicles. The inertial measurement unit generates 1000 sets of six-DOF motion data within this 10-second period, showing a noticeable upward and paused wrist movement at the 3-second mark. The microphone array acquires a 10-second ambient sound data segment, clearly recording the noise of passing vehicles and a slight "hmm?" sound emitted by the user at the 3.5-second mark. The system timestamps the dynamic visual image, the 1000 sets of six-DOF motion data, and the ambient sound data with a synchronization time stamp from T0 to T0+10 seconds and encapsulates them together into a single synchronous multimodal perception data packet.

[0029] The context-user joint state feature extraction module parses synchronous multimodal perception data packets, extracts environmental semantic elements, involuntary physiological feedback features, and emotional speech features, and associates these elements with a preset time window based on a unified timestamp to generate a context-user joint state feature set.

[0030] In a specific embodiment of the present invention, the generation of context-user joint state feature set includes: extracting dynamic visual images from synchronous multimodal perception data packets and inputting them into an image analysis model trained on a public dataset to identify entities in the environment that have preset social functions or contextual meanings, i.e., environmental semantic elements.

[0031] It should be noted that the goal of this step is to perform deep analysis on the synchronous multimodal perception data packets generated in the previous step, extracting key information that reflects the current environment and user state, and integrating it into a context-user joint state feature set. The implementation process begins with receiving the synchronous multimodal perception data packets and processing the three types of data contained within them in parallel. First, dynamic visual images are extracted from the data packets and input into an image analysis model trained on a public dataset. This model is specifically designed to identify entities in the environment that are distinct from general obstacles and have specific social functions or contextual significance, i.e., environmental semantic elements.

[0032] Six-degree-of-freedom motion data is extracted from synchronous multimodal sensing data packets, and non-autonomous physiological feedback features are obtained by analyzing the six-degree-of-freedom motion data.

[0033] In a specific embodiment of the present invention, the step of obtaining non-voluntary physiological feedback features by parsing six-degree-of-freedom motion data includes: identifying the slowing or pausing characteristics of the user's gait by the change in acceleration data exceeding a preset threshold from the six-degree-of-freedom motion data.

[0034] The system identifies the user's body posture tilt characteristics by continuously deviating the angle value from the dynamic baseline by more than a preset angle after integrating the six-degree-of-freedom motion data through the gyroscope data.

[0035] Gait slowing or pausing characteristics and body posture tilting characteristics are used as involuntary physiological feedback characteristics.

[0036] It should be noted that, next, six-degree-of-freedom motion data is extracted from the synchronous multimodal sensing data packet. The system analyzes this time-series data to look for deviations in the user's gait and posture relative to their normal walking pattern. For example, the system detects a slowdown or pause in gait by calculating the instantaneous rate of change of acceleration data. Specifically, when the absolute value of this rate of change exceeds a preset impact threshold calculated based on the user's normal walking pattern within a short period of time, it is determined that a "slowdown or pause in gait" feature has occurred.

[0037] It should be further explained that, to determine body tilt, the system primarily analyzes the angular velocity data corresponding to the left and right tilt of the body from the gyroscope data. By integrating the angular velocity data over time, the system estimates the amount of angular change of the body over a period of time. When the calculated angular change consistently deviates from the dynamic baseline representing the user's normal posture over a relatively long period of time, and the deviation exceeds a preset micro-angle threshold, the system identifies a "body tilt" characteristic. The preset micro-angle threshold is set to 3 degrees based on expert experience. Through this method, involuntary physiological feedback characteristics are extracted.

[0038] Ambient sound data is extracted from synchronous multimodal sensing data packets. The signal processing capability of the microphone array is used to separate the user's voice from the background sound. Then, the audio emotion recognition model is used to analyze the data to identify non-instructional sighs, questions, or self-talk that can express emotions, i.e., emotional speech features.

[0039] It should be noted that, simultaneously, ambient sound data is extracted from the synchronous multimodal perception data packet. Utilizing the signal processing capabilities of the microphone array, the user's voice is first separated from the background sound. Then, the separated user voice is analyzed using an audio emotion recognition model. This model is trained based on audio features such as MFCC, fundamental frequency, energy, and speech rate to identify non-instructional, emotionally expressive sighs, questions, or self-talk—i.e., emotional speech features.

[0040] It should also be noted that, in this invention, the so-called non-command, emotionally expressive speech features refer to sounds uttered unconsciously or subconsciously by the user that do not contain explicit action commands such as "go forward" or "turn left," but reflect their current psychological state. For example, when a user faces a complex intersection, they might utter a rising "Hmm?" or "Uh..." This is not a command, but clearly expresses their "question" or "hesitation," and the system will identify and label it as a "questioning sound" feature. Similarly, when a user feels crowded or tired in a crowd, they might utter a soft sigh "Sigh..." This will be identified as expressing "fatigue" or "difficulty." Furthermore, when a user sees an interesting shop, they might mutter to themselves, "Oh, what's this?" This reveals a tendency towards "curiosity" or "exploration." The system captures and classifies these subtle sounds through an audio analysis model, aiming to quantify these valuable emotional cues into usable data, thereby more accurately inferring the user's true intentions.

[0041] By associating environmental semantic elements, involuntary physiological feedback features, and emotional speech features, a context-user joint state feature set is generated.

[0042] It should be noted that, finally, the system associates various features whose timestamp differences fall within a preset window (e.g., 500 milliseconds) based on a unified timestamp. For example, if an environmental semantic element is identified at time T1, and involuntary physiological feedback features and emotional speech features are detected within 500 milliseconds after T1, the system associates these three. All valid association entries generated within the current analysis period are collected together to form a context-user joint state feature set. Its data structure is a list, where each element is an association record containing a timestamp, an environmental semantic element identifier, a description of involuntary physiological feedback features, and a description of speech features.

[0043] The context-user joint state feature set is a structured dataset designed to integrate environmental and user-specific state information, providing comprehensive input for subsequent intent prediction. Its data structure is a list, where each element is an associated record containing a timestamp, environmental semantic element identifiers, involuntary physiological feedback feature descriptions, and emotional speech feature descriptions. Environmental semantic elements refer to entities that are distinct from general obstacles and possess specific social functions or contextual significance. The recognition capability here is based on an image analysis model trained on a public dataset containing 100,000 labeled socially functional entities. Involuntary physiological feedback features refer to motion data patterns representing the user's unconscious physiological responses. Their data attributes are descriptive labels, such as "gait slowed" or "body leaning to the left." These labels are generated based on a dynamic baseline established from the user's six degrees of freedom motion data over the first 60 seconds. A trigger occurs when real-time data deviates from this baseline by more than a preset threshold, which is set based on expert experience as a speed 30% below the baseline mean. Emotional speech features are non-instructional, emotionally expressive sighs, questions, or self-talk. Their data attributes are descriptive labels, such as "sigh" or "questioning tone." Their recognition is achieved through an audio emotion recognition model, which determines emotions by analyzing the prosodic features of the audio signal (such as fundamental frequency, energy, and speech rate). The model is based on training on a database of thousands of recordings containing five basic emotions (neutral, happy, sad, angry, and surprised).

[0044] For example, continuing the example from the previous section, the image analysis model identifies an environmental semantic element labeled "densely populated area" on the left side of the visual field at the 3-second mark of the dynamic visual image. The motion data analysis module detects a sudden drop in acceleration in the six-degree-of-freedom motion data at the 3-second mark and extracts the involuntary physiological feedback feature as "sudden slowing or hesitation in gait." The audio processing module separates and identifies the user's "hmm?" sound from the environmental sound data at the 3.5-second mark and extracts the emotional speech feature as "questioning sound." Since these three features are highly correlated in time, the system associates them to generate a context-user joint state feature set. This set contains a record with the content: {Time: "T0+3.5s", Environmental semantic element: "densely populated area", Involuntary physiological feedback feature: "sudden slowing or hesitation in gait", Emotional speech feature: "questioning sound"}.

[0045] The user potential intent inference module predicts user potential intents and generates user potential intent labels based on the context-user joint state feature set and an intent inference model.

[0046] In a specific embodiment of the present invention, the step of predicting the user's potential intent through the intent inference model and generating the user's potential intent label includes: constructing an intent inference model, wherein the intent inference model is used to associate the received environmental semantic elements with the corresponding non-autonomous physiological and speech features to a preset intent category.

[0047] It should be noted that the implementation process first requires a pre-defined intent inference model within the system. This intent-inference model is built upon a support vector machine (SVM) classifier, and its construction process includes: 1) defining feature engineering methods to construct the feature vectors required for training from the raw data; 2) determining the kernel function type and key parameters used in the SVM; 3) formulating training data rules for labeling intent categories such as "exploration" and "confusion"; and 4) clarifying the input and output data formats and inference process of the model. This model is based on supervised learning and is trained using training data containing various user behavior patterns in different contexts and their corresponding real intent annotations.

[0048] The input context-user joint state feature set, combined with user head orientation data obtained by visual sensors, when it is recognized that the duration of the user's head orientation toward a specific environmental semantic element exceeds a preset duration and the user's gait rhythm slows down, the intent inference model infers that the user's behavior pattern meets the triggering conditions for exploration intent.

[0049] When the environmental semantic elements are identified as complex intersections or densely populated areas, and the user's gait slows down or stops and questions are heard simultaneously, the intent inference model infers that the user's behavior pattern meets the triggering conditions for confused intent.

[0050] It should be noted that when the model receives a context-user joint state feature set as input, it checks whether the elements in the feature set meet a certain pre-set intention triggering condition. For example, combining user head orientation data obtained from visual sensors, if it identifies that the duration of the user's head facing a specific environmental semantic element exceeds a preset duration (e.g., 2 seconds), and simultaneously there is a non-voluntary physiological feedback feature of slowed gait, the model will associate this behavioral pattern with an exploration intention. Another rule is: when the environmental semantic element is identified as a complex intersection, and a non-voluntary physiological feedback feature of slowed or paused gait is detected simultaneously, accompanied by a questioning voice feature, the intention inference model infers that the behavioral pattern meets the triggering condition of a confused intention.

[0051] The user latent intent label is a string data type that visually represents the user's most likely current needs or behavioral tendencies. Its data attribute is a pre-defined intent category name, such as "exploration" or "confusion." The intent inference model is a pre-trained machine learning model that maps the context-user joint state feature set to the user latent intent label. This model is implemented using a support vector machine classifier, and its training dataset consists of 5000 sets of anonymized user behavior data. Each set of data contains context-user joint features and manually labeled real user intents. Each intent category defines a set of trigger conditions. These trigger conditions are the result of expert experience summarizing the combination of visual, inertial navigation, and speech features. For example, the trigger condition for the "exploration" intent is set as follows: simultaneously detecting "the user's head facing a specific environmental semantic element for more than 2 seconds" and "gait change rate less than 0.1 meters per second." The trigger condition for the "confusion" intent is set as follows: within a 500-millisecond window associated with the timestamp, simultaneously detecting environmental semantic elements as complex intersections or densely populated areas, non-voluntary physiological feedback features as a sudden slowing or hesitation in gait, and emotional speech features as a questioning tone.

[0052] For example, continuing the example from the previous step, the system receives a context-user joint state feature set, which includes the following records: {Time: "T0+3.5s", environmental semantic element: "crowded area", involuntary physiological feedback feature: "sudden slowing or hesitation in gait", emotional voice feature: "questioning sound"}. The intent inference model begins to analyze it. Assume there is a rule in the model that uses "crowded area" as a trigger condition, combined with additional visual sensor data—that is, the user's head frequently turns towards this area in the past 2 seconds—and identifies the combination of "sudden slowing or hesitation in gait" and "questioning sound" as highly associated with the "confused" intent. The model matches these conditions, calculating a probability score of 0.85 for the "confused" intent and 0.4 for the "exploratory" intent. Since the "confused" intent has the highest probability score, the model ultimately generates a unique, high-confidence user latent intent label representing the user's most likely need, with a value of "confused".

[0053] The adaptive guidance strategy generation module matches and activates an adaptive guidance strategy from a preset strategy library based on the user's potential intent tags.

[0054] In a specific embodiment of the present invention, the step of matching and activating an adaptive guidance strategy from a preset strategy library includes: the strategy library defines each guidance strategy in a parameterized form as corresponding to a potential user intent tag, and defines the interaction mode, information detail level, and language style.

[0055] It should be noted that the goal of this step is to dynamically match and generate a specific adaptive guidance strategy from a pre-defined strategy library based on the user's latent intent tags predicted in the previous step. The implementation process begins with the system receiving the user's latent intent tags output by the intent inference model. The system internally builds and maintains a strategy library, which is a key-value database where the "key" is the user's latent intent tag and the "value" is the corresponding adaptive guidance strategy parameter set. Each guidance strategy contains a set of clearly defined parameters that specify the mode of interaction between the robot and the user, the level of detail in the information provided, and the language style used in communication. For example, the "suggestive collaboration strategy" and the "high-priority simplification strategy" both have their corresponding parameterized definitions.

[0056] When the received user potential intent label is "exploration", a suggestive collaboration strategy is activated from the strategy library. This strategy changes the robot's role from a commander to a companion, uses suggestive language, and allows temporary deviations from the path.

[0057] When the received user potential intent label is confused, a high-priority simplification strategy is activated. This high-priority simplification strategy masks secondary environmental information and guides the user through complex areas with concise and clear instructions, generating an adaptive guidance strategy.

[0058] It's important to note that when the system receives a latent user intent tag such as "exploration" or "confusion," it immediately uses this tag as a search index to look up the strategy in the policy library. Once a matching entry is found, the system reads and activates the complete guidance strategy defined for that entry. For example, when the received latent user intent tag is "exploration," the system activates a guidance strategy called "suggestive collaboration." This strategy changes the robot's role from a command giver to a companion, sets the language style to suggestive, and allows temporary deviations at the path planning level. If the received latent user intent tag is "confusion," the system activates a high-priority simplification strategy. This strategy instructs subsequent modules to mask secondary environmental information, outputting only the most critical navigation commands, and correspondingly switches the language style to concise and clear commands. Finally, based on the search results, the system generates a complete adaptive guidance strategy containing specific definitions of interaction modes, information priorities, and language styles, and passes it to the subsequent path optimization and command generation modules.

[0059] The adaptive guidance strategy is a set of rules defining how the robot interacts with the user. Its data structure is a parameter set containing three fields: interaction mode, information detail level, and language style. The strategy library is a structured data storage system that pre-stores the correspondence between user latent intent tags and adaptive guidance strategies. Its data structure can be a key-value database, where the "key" is the string of the user's latent intent tag, and the "value" is the corresponding adaptive guidance strategy parameter set. The interaction mode defines the type of interaction between the navigation robot and the user. Its data attribute is a preset enumeration type, such as "commander mode" or "companion mode," based on the master-slave and collaborative models in human-computer interaction theory. The information detail level is a control parameter defining the number of environmental details included in the guidance information. Its data attribute is a three-level classification label: "complete," "core," and "simplified," based on the results of a survey of 100 visually impaired individuals' information needs in different scenarios. Language style refers to the tone and wording style used in navigation voice broadcasts. Its data attributes are preset enumeration types, such as "instructional", "suggestive" or "descriptive". This setting is designed to match the language acceptance preferences of users in different emotional states.

[0060] Continuing the example from the previous step, the system receives a user's latent intent label with the value "confused". The system immediately searches the policy library using "confused" as the keyword. A corresponding guidance policy exists in the policy library, named "High Priority Simplification". The system reads the detailed definition of this policy and generates an adaptive guidance policy with the following content: {Interaction Mode: "Instructor Mode", Information Level: "Simplified", Language Style: "Instructive"}. This adaptive guidance policy will be used to guide subsequent path planning and voice interaction, ensuring that the robot's guidance method can most effectively help users in a confused state.

[0061] The adaptive guidance path optimization module obtains the robot's original navigation path and adjusts it according to the rules defined by the adaptive guidance strategy to generate an adaptive guidance path.

[0062] In a specific embodiment of the present invention, adjusting the original navigation path of the robot according to the rules defined by the adaptive guidance strategy to generate an adaptive guidance path includes: obtaining the original navigation path planned by the robot navigation system based on map and obstacle avoidance.

[0063] It should be noted that the goal of this step is to adjust the original navigation path planned by the robot's navigation system based on the adaptive guidance strategy generated in the previous step, thereby generating an adaptive guidance path that better suits the user's current situation and intentions. The implementation process first retrieves the currently planned original navigation path from the robot's navigation system. This original navigation path is automatically planned by the robot's navigation system based on map information and obstacle avoidance algorithms, and is typically represented by a series of discrete waypoints or road segments. Next, the system calls upon the specific content defined by the adaptive guidance strategy output in the previous step, and determines how to adjust the original navigation path based on parameters such as "interaction mode," "information detail," and "language style" within the strategy.

[0064] If the adaptive guidance strategy is invoked, and the adaptive guidance strategy is advisory collaboration, then an adaptive guidance path containing a slow-down zone is generated around the target point of interest to the user in the original navigation path by interrupting and inserting a temporary exploration loop.

[0065] In a specific embodiment of the present invention, the slow-moving area is a circular area with a preset radius, centered on the target point of interest to the user.

[0066] It should be noted that the process of determining the target point of interest is as follows: the system combines the user's head orientation data obtained by the visual sensor with the gait features extracted from the six degrees of freedom motion data. When it is detected that the user's head is continuously facing a certain environmental semantic element for more than a preset time, and at the same time, the user's gait rhythm is slowed down, the environmental semantic element is determined as the target point of interest for the user.

[0067] It should also be noted that if the adaptive guidance strategy invoked is advisory collaboration, this strategy indicates that the user may be in an exploratory state. In this case, the system dynamically modifies the original navigation path, seamlessly integrating a buffer zone around the target point of interest into the path. Specifically, the system first identifies the critical path point P on the original navigation path that is closest to the target point, and uses this point as the entry point into the buffer zone. Then, the system temporarily interrupts the original navigation path at point P and generates a circular area with a preset radius R centered on P as the buffer zone. The preset radius R is set to 2 meters based on historical experience. When the robot reaches the boundary of this buffer zone, its standard path point tracking logic is paused, and the exploration logic within the buffer zone is executed instead: the robot's speed is reduced from the usual 0.8 m / s to 0.3 m / s, allowing for small-scale free movement within the area, while continuously adjusting its posture to face the target point, providing the user with opportunities for observation and perception. When the user issues a leave command or the dwell time exceeds a preset threshold, the exploration logic terminates. The preset threshold for the dwell time is set to 5 seconds based on historical experience. At this point, the system will replan a smooth connecting path from the robot's current position in the slow-moving area to the next critical path point (i.e., the path point after point P) on the original navigation path, and restore the normal navigation mode, thus allowing the robot to seamlessly return to the original navigation path and continue moving forward. In this way, the original single command path is optimized into an adaptive guided path that includes temporary exploration loops, ensuring the execution of the main task while meeting the user's immediate exploration needs.

[0068] If the adaptive guidance strategy is high-priority simplification, then the fine-tuning instructions for non-critical turning points in the original navigation path are filtered out, and only the guidance information of the start point, end point and critical turning points is retained to generate an adaptive guidance path.

[0069] It should be noted that if the adaptive guidance strategy is set to high-priority simplification, this indicates that the user may be confused and requires concise and clear guidance. In this case, the system simplifies the original navigation path, filtering out fine-tuning instructions for non-critical turning points. Specifically, critical turning points are identified by the geometric distance and angle changes between path points. If the angle change between three consecutive path points is less than a preset threshold, the intermediate path point is considered non-critical and removed. The preset threshold for angle change can be 15 degrees, which ensures that instructions are minimized without affecting the overall navigation direction determination. Only the starting point, ending point, and all critical turning points are retained in the original path, thus simplifying the path expression to the simplest form of adaptive guidance.

[0070] The original navigation path is a set of ordered spatial coordinates and motion commands calculated by the robot navigation system based on global map information and real-time sensor data. It guides the robot from its current position to the target position. Its data structure is an ordered list of path points, each containing three-dimensional coordinates (x, y, z) and orientation information (yaw angle). The adaptive guidance path, on the other hand, is a navigation path adjusted using an adaptive strategy based on the original navigation path and incorporating user intent and context. Its data type is either a corrected ordered list of path points or a complex data structure containing a list of path points and area constraint information.

[0071] For example, continuing the example from the previous section, the adaptive guidance strategy received by the system is {Interaction Mode: "Commander Mode", Information Level: "Simplified", Language Style: "Commanding"}, and the corresponding user's potential intent label is "Confusion". Assume the original navigation path contains 20 path points P1, P2, ..., P20 from point A to point Z. The system invokes the "High-Priority Simplification" strategy. By analyzing the angle changes between path points, it is found that the directional changes between points P3 and P4, P7 and P8, P10 and P11, and P12 and P13 are very small, with angles all less than 15 degrees, indicating that these are minor adjustments or redundant points in straight segments of the original path. Therefore, the system will delete these non-critical path points (P4, P8, P11, P12). The final generated adaptive guidance path will contain only the starting point P1, a few key turning points (such as P2, P5, P9, P14, P18), and the ending point P20, forming a path with fewer and simpler instructions, for example: P1->P2->P5->P9->P14->P18->P20. This simplified path will serve as the basis for the final navigation instructions.

[0072] The contextualized multimodal interaction instruction output module executes an adaptive guidance path and, in conjunction with an adaptive guidance strategy, synthesizes contextualized multimodal interaction instructions.

[0073] It should be noted that in this invention, the adaptive guidance strategy plays a central role in decision-making and transformation. Its key function is to transform the system's abstract understanding of the user's potential intentions into a set of specific, multi-dimensional robot action guidelines. First, it defines the role and style of human-computer interaction, determining whether the robot guides the user as a concise and clear instructor or as a gentle and suggestive companion, depending on whether the user is in an "exploration" or "confusion" state. Second, it directly guides the optimization of navigation paths. For example, when the user shows an exploratory intention, it dynamically inserts a "temporary exploration loop" to satisfy their curiosity into the original path; or when the user is confused, it filters out secondary path points to generate a clear path with the simplest instructions. Finally, the strategy also finely specifies the form of the final interactive instructions output to the user, including the level of detail in voice instructions, language style, and the matching tactile vibration pattern of the guide handle (such as short, forceful pulse vibrations or gentle directional vibrations). In summary, the adaptive guidance strategy transforms guide robots from rigid tools that rigidly execute preset routes into intelligent partners capable of sensing user status in real time, understanding user intentions, and dynamically adjusting their own behavior patterns. This significantly enhances the personalization, contextual awareness, and safety of navigation.

[0074] In a specific embodiment of the present invention, the synthesized contextualized multimodal interaction command includes: driving the robot to move along an adaptive guidance path.

[0075] It should be noted that the goal of this step is to ultimately execute the adaptive guidance path generated in the previous step, and, in conjunction with the adaptive guidance strategy, synthesize and output a set of contextualized multimodal interactive commands that include voice and haptic feedback. The implementation process begins with driving the robot's motion control system to ensure it moves precisely along the adaptive guidance path optimized in the previous step. The robot's wheeled chassis adjusts its speed and direction in real time according to the path point sequence to ensure smooth and accurate movement.

[0076] During the journey, the language style defined in the adaptive guidance strategy is invoked to transform the geometric information of the path points into natural language that conforms to the current context.

[0077] It's important to note that during the robot's movement, the system continuously invokes the generated adaptive guidance strategy. Specifically, it reads the "language style" parameter defined in the strategy. When the robot approaches a key waypoint in the adaptive guidance path, such as a corner or the entrance to a target area, the navigation module passes the geometric information of this waypoint (e.g., "turn left 90 degrees" or "5 meters ahead to the destination") to the interaction generation module. The interaction generation module then transforms this purely geometric information into natural language appropriate to the current context, based on the language style specified in the adaptive guidance strategy. For example, if the strategy defines the language style as "suggestive," the system might generate the voice prompt "There seems to be an interesting corner ahead; we can look to the left," instead of simply "Please turn left." If the language style is "instructive," it will generate the concise and clear "Turn left ahead."

[0078] Based on the interaction mode defined by the adaptive guidance strategy, natural language is combined with the directional vibration of the guide handle in a specific mode to generate contextualized multimodal interaction commands that include voice and tactile feedback and output them.

[0079] It's important to note that the system also combines the generated natural language with the directional vibrations of the guide handle according to the "interaction mode" defined in the adaptive guidance strategy. For example, in "companion mode," while issuing suggestive voice commands, the guide handle may generate a gentle and continuous guiding vibration, pointing in the suggested direction of movement, providing a mild prompt to the user. In "instructor mode," when issuing command-style voice commands, the guide handle will generate a short and powerful pulse vibration, clearly indicating the precise timing for turning. Ultimately, by organically combining path execution, context-appropriate natural language delivery, and the guide handle's specific vibration patterns, the system generates and outputs in real time a set of contextualized multimodal interaction commands that coordinate voice and tactile feedback and are highly matched to the user's current state and intentions (such as exploration or confusion). This set of commands is transmitted to the user through the robot's speaker and the guide handle's vibrator, completing the closed loop of the entire guide interaction.

[0080] Among them, contextualized multimodal interaction commands are a set of synchronously output commands containing information from multiple sensory channels. Their function is to provide users with a rich and context-appropriate guidance experience. Their data structure is a composite data packet containing voice command text and tactile vibration mode parameters. Directional vibration is a tactile feedback technology that generates directional perception in the user's hand by controlling the vibration timing and intensity of multiple vibration motors within the guide handle.

[0081] Continuing the example from the previous section, the system is executing the simplified adaptive guidance path generated for a user in a "confused" state. At this point, the adaptive guidance strategy is {Interaction Mode: "Commander Mode", Information Level: "Simplified", Language Style: "Imperative"}. As the robot approaches pathpoint P5, a critical right-turning point, the navigation module sends the "Turn Right" geometry to the interaction generation module. Because the language style is set to "Imperative", the module translates this information into concise natural language: "Turn Right". Simultaneously, because the interaction mode is "Commander Mode", the system sends commands to the guide handle.

[0082] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A navigation system for a blind robot that combines vision, inertial navigation, and voice, characterized in that, include: The multimodal perception data acquisition module simultaneously acquires dynamic visual images, six-degree-of-freedom motion data, and ambient sound data, and marks the dynamic visual images, six-degree-of-freedom motion data, and ambient sound data with a unified timestamp to generate a synchronous multimodal perception data packet; The context-user joint state feature extraction module parses synchronous multimodal perception data packets, extracts environmental semantic elements, involuntary physiological feedback features, and emotional speech features, and associates environmental semantic elements, involuntary physiological feedback features, and emotional speech features within a preset time window based on a unified timestamp to generate a context-user joint state feature set. The user potential intent inference module predicts user potential intents and generates user potential intent labels based on the context-user joint state feature set and intent inference model. The step of predicting users' potential intentions through an intent inference model and generating user potential intention tags includes: Construct an intent inference model, which is used to associate the received environmental semantic elements with the corresponding non-autonomous physiological and speech features to a preset intent category; The input context-user joint state feature set, combined with the user's head orientation data obtained by the visual sensor, when it is recognized that the duration of the user's head orientation toward a specific environmental semantic element exceeds a preset duration and the user's gait rhythm slows down, the intent inference model infers that the user's behavior pattern meets the triggering conditions for exploration intent. When the environmental semantic elements are identified as complex intersections or densely populated areas, and the user's gait slows down or stops and questions are heard simultaneously, the intent inference model infers that the user's behavior pattern meets the triggering conditions of confused intent. The adaptive guidance strategy generation module matches and activates adaptive guidance strategies from a preset strategy library based on the user's potential intent tags. The adaptive guidance path optimization module obtains the robot's original navigation path and adjusts it according to the rules defined by the adaptive guidance strategy to generate an adaptive guidance path. The contextualized multimodal interaction instruction output module executes an adaptive guidance path and, in conjunction with an adaptive guidance strategy, synthesizes contextualized multimodal interaction instructions.

2. The interactive navigation system for a blind guide robot combining vision, inertial navigation, and voice as described in claim 1, characterized in that: The generation of synchronous multimodal sensing data packets includes: Activate the robot's vision sensors to acquire dynamic visual images containing environmental objects and spatial relationships; Simultaneously activate the inertial measurement unit built into the guide handle to capture the user's gait rhythm, body posture, and six degrees of freedom motion data of hand control during walking; The microphone array is used to acquire user voice commands and background ambient sound, and the user voice commands and background ambient sound are used together as ambient sound data. The acquired dynamic visual images, six-degree-of-freedom motion data, and environmental sound data are labeled with unified timestamps and packaged to generate synchronous multimodal sensing data packets.

3. The interactive navigation system for a blind guide robot combining vision, inertial navigation, and voice as described in claim 1, characterized in that: The generated context-user joint state feature set includes: Dynamic visual images are extracted from synchronous multimodal perception data packets and input into an image analysis model trained on a public dataset to identify entities in the environment that have pre-defined social functions or contextual meanings, i.e., environmental semantic elements. Six-degree-of-freedom motion data is extracted from synchronous multimodal sensing data packets, and non-autonomous physiological feedback features are obtained by analyzing the six-degree-of-freedom motion data. Ambient sound data is extracted from synchronous multimodal sensing data packets. The signal processing capability of the microphone array is used to separate the user's voice from the background sound. Then, the audio emotion recognition model is used to analyze the data to identify non-instructional sighs, questions, or self-talk that can express emotions, i.e., emotional speech features. By associating environmental semantic elements, involuntary physiological feedback features, and emotional speech features, a context-user joint state feature set is generated.

4. The interactive navigation system for a blind guide robot combining vision, inertial navigation, and voice as described in claim 3, characterized in that: The process of analyzing six-degree-of-freedom motion data to obtain non-voluntary physiological feedback characteristics includes: The slowing or pausing characteristics of a user's gait are identified by changes in acceleration data exceeding a preset threshold from six degrees of freedom motion data. The user's body posture tilt characteristics are identified by the angle value obtained by integrating the gyroscope data from the six degrees of freedom motion data and continuously deviating from the dynamic baseline by more than a preset angle. Gait slowing or pausing characteristics and body posture tilting characteristics are used as involuntary physiological feedback characteristics.

5. The interactive navigation system for a blind guide robot combining vision, inertial navigation, and voice as described in claim 1, characterized in that: The step of matching and activating an adaptive guidance strategy from a preset strategy library includes: The strategy library defines each guidance strategy in a parameterized manner, corresponding to a potential user intent tag, and defines the interaction mode, information level of detail, and language style. When the received user potential intent label is "exploration", a suggestive collaborative strategy is activated from the strategy library. This suggestive collaborative strategy changes the robot's role from a commander to a companion, uses suggestive language, and allows temporary deviations from the path. When the received user potential intent label is confused, a high-priority simplification strategy is activated. This high-priority simplification strategy masks secondary environmental information and guides the user through complex areas with concise and clear instructions, generating an adaptive guidance strategy.

6. The interactive navigation system for a blind robot combining vision, inertial navigation, and voice as described in claim 5, characterized in that: The process of adjusting the robot's original navigation path according to the rules defined by the adaptive guidance strategy to generate an adaptive guidance path includes: Obtain the original navigation path planned by the robot navigation system based on map and obstacle avoidance; If the adaptive guidance strategy is invoked, and the adaptive guidance strategy is advisory collaboration, then an adaptive guidance path containing a slow-down zone is generated around the target point of interest to the user in the original navigation path by interrupting and inserting a temporary exploration loop. If the adaptive guidance strategy is high-priority simplification, then the fine-tuning instructions for non-critical turning points in the original navigation path are filtered out, and only the guidance information of the start point, end point and critical turning points is retained to generate an adaptive guidance path.

7. The interactive navigation system for a blind guide robot combining vision, inertial navigation, and voice as described in claim 6, characterized in that: The slow-moving area is a circular area with a preset radius, centered on the target point of interest to the user.

8. The interactive navigation system for a blind guide robot combining vision, inertial navigation, and voice as described in claim 1, characterized in that: The synthesized contextualized multimodal interaction instructions include: Drive the robot to move along an adaptive guidance path; During the journey, the language style defined in the adaptive guidance strategy is invoked to transform the geometric information of the path points into natural language that conforms to the current context; Based on the interaction mode defined by the adaptive guidance strategy, natural language is combined with the directional vibration of the guide handle in a specific mode to generate contextualized multimodal interaction commands that include voice and tactile feedback and output them.

Citation Information

Patent Citations

  • An indoor guide robot for the visually impaired

    CN111142536B

  • Indoor blind guiding robot

    CN111142536A

  • Blind person intelligent glasses based on binocular vision

    CN120045075A