Indoor barrier-free path dynamic planning and voice navigation system

Through multi-source perception fusion and dynamic path planning, the problem of insufficient recognition of dynamic environments by indoor navigation systems is solved, the reliability of passage and user interaction experience in complex environments are improved, and accurate recognition of targets such as humans and pets and the naturalness of voice navigation are achieved.

CN120820162AInactive Publication Date: 2025-10-21HANGZHOU YISHITONG TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511146428.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing indoor navigation systems lack the ability to perceive dynamic environmental changes in real time, and have difficulty identifying flexible obstructions and targets with thermal characteristics, leading to path misjudgment and traffic interruption. In addition, the voice navigation prompts lack environmental semantic understanding, making it difficult to meet users' requirements for natural interaction and accuracy.

Method used

A multi-source perception fusion module is used to fuse data from RGB-D cameras, millimeter-wave radars, and thermal infrared imaging modules. Target semantic recognition is performed through a multi-task neural network, dynamic path planning is performed using a finite state machine and the D*Lite algorithm, and real-time guidance is provided through a voice navigation module.

Benefits of technology

It significantly improves the ability to identify targets such as humans, pets, and flexible obstructions, reduces path misjudgments, improves traffic reliability and interactive experience, and realizes real-time assessment of complex environments and path guidance at the user perception level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120820162A_ABST
    Figure CN120820162A_ABST
Patent Text Reader

Abstract

The invention discloses an indoor barrier-free path dynamic planning and voice navigation system, and particularly relates to the technical field of path dynamic planning, and the system comprises a multi-source sensing fusion module, a motion mode modeling module, a path planning module and a voice navigation module. The system constructs a multi-modal intermediate representation tensor by fusing data of an RGB-D camera, a millimeter wave radar and a thermal infrared module, and realizes the recognition of the geometrical shape, the material type and the heat source state of a target based on a multi-task neural network model. In combination with track caching and state machine modeling methods, the system dynamically discriminates the motion mode of the obstacle target and constructs a state transition structure with time sequence consistency. The path planning module generates a path with fault tolerance and high accessibility according to the target semantic category and the motion state result in combination with map topology and navigation confidence constraints. And the voice navigation module generates a personalized voice instruction according to the path structure information and the current state of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of path dynamic planning, and more particularly to an indoor barrier-free path dynamic planning and voice navigation system. Background Art

[0002] With the growing demand for smart buildings and barrier-free environments, indoor navigation systems are increasingly being used in public places, medical institutions, supermarkets, and complex industrial parks to assist people with limited mobility, security patrols, and service robots in providing path guidance. However, existing indoor navigation systems primarily rely on fixed path planning and static obstacle modeling, lacking the ability to perceive dynamic environmental changes in real time and struggling to adapt to uncertainties such as movable obstacles, temporarily enclosed areas, and positioning errors in complex scenarios.

[0003] On the one hand, traditional navigation solutions often use single vision or lidar data as the basis for environmental perception, which makes it difficult to accurately identify flexible obstructions (such as plastic curtains, liquid pools) or targets with thermal characteristics (such as humans and pets), which can easily lead to path misjudgment or traffic interruption. On the other hand, existing path planning methods are often based on fixed maps and simplified traffic models. They are unable to dynamically adjust the path tolerance range based on the semantic attributes of the target and the movement trend. As a result, the system lacks the ability to effectively avoid and replan when faced with dynamic obstacles such as temporary stacking and service robot intersections. In addition, the current voice navigation prompts lack a deep understanding of the semantics of the environment, and the prompt information is too mechanical or generalized, making it difficult to meet the dual requirements of specific groups for natural voice interaction and guidance accuracy.

[0004] Therefore, there is an urgent need for an indoor navigation system with multi-source perception fusion, semantic recognition, dynamic path generation and voice interaction capabilities to achieve real-time assessment of traffic feasibility and path guidance at the user perception level in complex environments, thereby improving the system's traffic reliability and interactive experience. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an indoor barrier-free path dynamic planning and voice navigation system to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions: Indoor barrier-free path dynamic planning and voice navigation system, including the following modules: The multi-source perception fusion module is used to fuse the collected synchronous observation data, construct a unified multimodal intermediate representation tensor, and input it into a pre-trained multi-task neural network model to achieve target geometry recognition, material type judgment, and heat source state classification, thereby obtaining the target semantic category label. The synchronous observation data includes 2D color images and depth maps, point cloud data, and temperature distribution maps. The motion pattern modeling module is used to establish a short-term trajectory buffer of dynamic length based on the identified target semantic category label. The current target trajectory buffer length is adaptively adjusted based on four regulatory factors: the target semantic category label, expected residual time, local path feasibility redundancy, and user physiological response margin index. Within this short-term trajectory buffer of dynamic length, the target's average linear velocity, directional volatility index, and trajectory continuity parameter are extracted, and a finite state machine is constructed to achieve dynamic identification of the target's motion state and output target state information for path planning.

[0007] In a preferred embodiment, the multi-task neural network model includes a shared encoder and three semantic branch networks corresponding to geometric form, material type and heat source state respectively. The shared encoder adopts a residual convolution structure and outputs a shared feature map. The semantic branches all output category labels and confidences in a Softmax manner and are jointly trained through a weighted multi-task loss function.

[0008] In a preferred embodiment, the expected residual time is determined by three parts: the average residence time of targets of the same category in the area is used as a baseline value; a speed suppression factor with a value between zero and one is introduced according to the current moving speed of the target to perform monotonically reduce the baseline value; and a gain adjustment is performed on the reduced result according to the degree of trajectory loop of the target and the loop enhancement coefficient.

[0009] In a preferred embodiment, the trajectory loop degree refers to the spatial overlap between the target motion path and its historical trajectory.

[0010] In a preferred embodiment, the local path feasibility redundancy is determined by the following factors: the number of alternative paths that can be selected from the current position and the margin of the average width of the avoidance space along each pass path relative to the system preset standard detour width.

[0011] In a preferred embodiment, the user physiological response margin index is determined by taking the maximum acceptable delay time of the user during the navigation process as the upper limit, and estimating the time margin for completing prompts, turns or deceleration within a given lead time in combination with the user's current moving speed; wherein the maximum acceptable delay time is obtained based on historical interaction records.

[0012] In a preferred embodiment, the finite state machine includes four motion states, namely, a completely stationary state, a temporarily stationary but pushable state, a stable autonomous movement state, and a highly fluctuating or unpredictable movement state.

[0013] In a preferred embodiment, a path planning module is also included, which is used to combine the target semantic category label and the target motion state information, refer to the indoor map topology structure and the dynamic distribution state of obstacles, construct an enhanced path map and dynamically assign path nodes and path edge values, and generate a main path and alternative path set with fault tolerance and high confidence reachability based on the D*Lite algorithm.

[0014] In a preferred embodiment, a voice navigation module is further included, which is used to generate voice prompt content by combining the main path and the set of alternative paths output by the path planning module with the user's current state.

[0015] In a preferred embodiment, the voice navigation module dynamically generates three types of voice prompt content including path guidance prompts, environment avoidance prompts and system status feedback prompts based on the user's current location, facing direction, and historical interaction records.

[0016] Technical effects and advantages of the present invention: Compared with the existing technology, the present invention enables the system to simultaneously perceive the geometric shape, material type and heat source status of the target in complex indoor environments through the collaborative design of multi-source perception fusion and multi-task semantic recognition, significantly improving the ability to distinguish targets that were previously difficult to accurately identify, such as human bodies, pets, flexible obstructions and liquids, reducing the probability of path misjudgment and traffic interruption from the source, and solving the problem of insufficient semantic understanding caused by traditional solutions relying only on a single sensing channel, thereby improving traffic reliability and interactive experience.

[0017] In terms of temporal behavior understanding, the integrated modeling framework of "dynamic length short-term trajectory cache-finite state machine-temporal consistency constraint" proposed in this invention can adaptively adjust the observation window based on control factors such as target semantic category, expected residual time, local path feasibility redundancy and user physiological response margin, while ensuring response speed and suppressing misjudgment caused by short-term disturbances; combined with a four-state model including complete stillness, driven stillness, stable autonomous movement and high-fluctuation movement, as well as time threshold and fuzzy boundary mechanism, the system effectively avoids state "flickering switching" and obtains higher stability, interpretability and deployment reliability in continuous perception scenarios, thereby providing more reliable dynamic environment input for subsequent path planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1This is a schematic diagram of the structure of the indoor barrier-free path dynamic planning and voice navigation system of the present invention; Figure 2 This is a flow chart of the internal processing of the motion pattern modeling module of the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] Example 1 The present invention provides an indoor barrier-free path dynamic planning and voice navigation system, such as Figure 1 As shown, it includes the following modules: The multi-source perception fusion module is used for the initial perception and semantic recognition of target entities in the environment. When the target first enters the visual range, the system collects synchronous observation data based on the multi-source perception devices deployed in the environment and starts the initial perception processing process. The multi-source perception devices include RGB-D cameras, millimeter-wave radars, and thermal infrared imaging modules, which are used to obtain color texture, spatial contours, and thermal radiation characteristics, respectively. Among them, the RGB-D camera provides two-dimensional color images and depth maps to capture the boundary shape, surface texture density, and relative depth relationship of obstacles; the millimeter-wave radar provides point cloud data with anti-occlusion capabilities, enhancing the spatial configuration perception of targets in low-light or complex lighting environments; the thermal infrared module outputs a temperature distribution map to assist in determining whether the target is an entity with biological characteristics or an active heat source.

[0021] On edge computing nodes, the system performs simultaneous alignment and standardization preprocessing on these three types of raw data. This includes image brightness normalization, depth map hole completion, point cloud voxel downsampling, and thermal map temperature enhancement. After processing, the system concatenates the multimodal data along the channel dimension to form a unified intermediate representation tensor T ∈ ℝ^{H×W×C}, where H and W represent the spatial dimensions, and C represents the number of multimodal channels, including RGB image channels, depth map channels, point cloud projection density map channels, and thermal infrared image heat intensity channels.

[0022] This intermediate representation tensor is input into a pre-trained multi-task neural network model deployed on an edge computing node to perform parallel semantic recognition tasks. The pre-trained multi-task neural network model uses a shared feature encoder and three independent semantic task branches. The encoder uses a multi-layer residual convolutional network to extract the joint spatial features of the target structure and material, and outputs a shared feature map F_shared. The model then connects to three types of semantic branch networks: Geometric morphology branch: identifies whether the target has geometric properties such as regularity and symmetry; Material type branch: determines whether the target material is metal, fabric, glass, etc. Heat source status branch: Determines whether the target has a living heat source, an intermittent heat source, or no heat radiation.

[0023] Each of the three semantic branches outputs the label category and corresponding confidence level through the Softmax function. During the inference phase, the label with the highest confidence level is selected as the final recognition result. The pre-trained multi-task neural network model is trained using a multi-task learning mechanism, with a loss function calculated as: L_total = λ1·L_geom + λ2·L_mat + λ3·L_heat. L_geom, L_mat, and L_heat represent the cross-entropy loss functions for the three semantic tasks, respectively. λ1, λ2, and λ3 are weighting coefficients, dynamically set based on the sample distribution and training difficulty.

[0024] Based on the parallel reasoning results of the above model, the system preliminarily divides each target into the following six semantic categories: (1) Human body: with vertical structure, complex texture and stable heat source distribution; (2) Pets: small size, irregular symmetry, concentrated but unstable heat source; (3) Service robots or mobile devices: regular shape, uniform surface, and heat source mostly located in the chassis or motor part; (4) Furniture or movable objects: stable structure, lack of heat source, clear texture but no biological characteristics; (5) Temporary stacking objects: such as shelves and tool boxes, with chaotic shape, mixed texture and no stable thermal characteristics; (6) Liquid or flexible obstructions: such as stagnant water and plastic curtains, with blurred edges, strong reflection and discontinuous heat signals.

[0025] The motion pattern modeling module dynamically establishes a short-term trajectory cache for each target. This cache is a time-series data cache structure created for a specific object instance with a lifecycle. It temporarily stores the object's historical trajectory (position, velocity, etc.). It is typically created after the object is first recognized and continuously updated throughout its visible lifecycle. Each tracked obstacle target is stored separately and is typically deleted upon expiration (e.g., disappearance). Its primary use is in behavioral modeling and trajectory pattern identification.

[0026] like Figure 2 As shown in the figure, the current target's trajectory cache length is adaptively adjusted based on four control factors: target semantic category, expected residual time, local path feasibility redundancy, and user physiological response margin. Based on the trajectory cache update, the system continuously calculates motion parameters such as the target's average linear velocity, directional fluctuation, and trajectory continuity, and combines this with the target's semantic category to determine its motion state.

[0027] Specifically, after completing the preliminary semantic classification of the target, a short-term trajectory buffer with a dynamic length is established for the obstacle target. The trajectory buffer length of the current target is adaptively adjusted according to the following formula: ; 、 、 : It is a weight coefficient set by experience, used to balance the influence of residual time, bypassability and user margin. For example: = 0.5, = 2.0, = 1.0.

[0028] in: : The trajectory cache length of the current target, that is, the time span for which the system will retain the trajectory information of the target.

[0029] : The basic cache length is determined by the semantic category label of the target (e.g., "pets" is set to 3 seconds, "furniture" is set to 6 seconds, etc.).

[0030] : Expected residual time, which indicates how long the system predicts the target will stay at its current location. The specific calculation is: ; Wherein, μC is the average residence time of targets of the same category in the area; is the target moving speed; η is the speed suppression factor, with a value range of (0, 1), which is used to control the negative impact of linear speed on residual time; The trajectory loopback degree is defined as the degree of spatial overlap between the target's motion path and its historical trajectory within the current trajectory buffer.

[0031] : Local path feasibility redundancy. The larger the redundancy, the more sufficient the detour space is, and the cache can be shortened. It is calculated as: ; : Local path feasibility redundancy is a comprehensive indicator that describes the ability to bypass the target's current location. A larger value indicates greater local flexibility and suggests that the navigation system should reduce cache.

[0032] : The number of navigable branches in the path map within a set radius (such as 2 meters) centered on the current target location, that is, the number of alternative paths that can be selected from the current location.

[0033] :all The average width of the avoidance space in each traffic path.

[0034] : The system's preset standard detour width threshold is used to determine whether there is enough detour space in the current area. For example, it can be set to 0.8 meters (80 centimeters).

[0035] : User physiological response margin indicator, indicating how long in advance the system needs to observe the target to ensure navigation safety, calculated as: ; : User physiological response margin indicator. Indicates how much time the navigation system should reserve for the user to perceive and respond to obstacle changes in the current situation. The larger the value, the longer the trajectory cache should be. : Device compensation factor, set according to the type of assistive device used by the user (such as wheelchair, guide stick), reflecting its reaction speed and movement limitations.

[0036] : The maximum acceptable delay time during user navigation, estimated from historical interaction records, represents the upper limit of user tolerance for the navigation system's response speed.

[0037] : The user's current movement speed; this item is eliminated in the actual calculation and does not affect the final result.

[0038] The system immediately creates a The system continuously receives image sequences and point cloud data streams captured by RGB-D cameras and millimeter-wave radars deployed in the environment. It then uses a multi-target tracking algorithm based on Kalman prediction combined with deep appearance feature matching to correlate the spatial positions of targets in consecutive frames and construct trajectories.

[0039] For each obstacle target, within each sliding time window, the system calculates its average linear velocity, directional volatility index and trajectory continuity parameter in turn, and combines the actual effective length of the current trajectory buffer area to jointly construct a motion state judgment logic framework to achieve automatic recognition of the target's current motion mode.

[0040] Specifically, the average linear velocity is calculated as follows: ;in, represents the position vector of the target at time tᵢ; N is the total number of frames in the trajectory buffer, that is, the number of consecutive target states recorded in the buffer; is the average linear velocity; is the inter-frame time interval; It represents the Euclidean distance between the i-th and i+1-th frames, reflecting the displacement of the target in this time period.

[0041] The directional volatility indicator is calculated as follows: ; Represents the velocity vector of the target in segment i, that is: ; It is the angle function between two vectors. The return value is in radians or degrees (depending on the implementation setting), which represents the deflection angle between two adjacent velocity directions. It is a directional fluctuation index that reflects the stability of the target's moving direction. The larger the value, the more drastic the change in trajectory direction, such as sharp turns, swings, or jitters.

[0042] The calculation formula of trajectory continuity parameter is as follows: ; Indicates the actual cumulative length of the obstacle target's trajectory in the current buffer, which is the sum of the displacement distances between frames. The calculation method is: ; : represents the straight-line distance between the starting point and the ending point of the obstacle target in the buffer area, that is: ; Indicates the starting position of the obstacle target in the trajectory buffer; Indicates the end position of the obstacle target in the trajectory buffer. is the trajectory continuity parameter, reflecting whether the current trajectory has significant detours, repetitions or jumps. When it is close to 1, it means that the trajectory is close to a straight line and has good continuity; if it is far less than 1, it means that the trajectory is turning back, pausing or discontinuous.

[0043] After constructing the target trajectory buffer and extracting motion features, the system further introduces a finite state machine (FSM)-based modeling approach to dynamically determine the target's motion state. This state modeling abstracts the motion pattern into several explicit state nodes and constructs a state transition diagram based on the changes in observed indicators during trajectory evolution, enabling flow control and temporal consistency maintenance across multiple states.

[0044] Specifically, the system defines the following four basic motion states: (1) S1: completely stationary; (2) S2: temporarily stationary but capable of being pushed; (3) S3: stable autonomous movement; and (4) S4: highly volatile or unpredictable movement. Each state corresponds to a set of typical characteristic indicator intervals (such as a lower speed limit, an upper volatility limit, and a trajectory continuity threshold), which are used to trigger state transitions when the observed indicators meet the conditions.

[0045] In actual operation, the system constructs the current state vector V(t) for each sliding time window based on observation metrics (average linear velocity, directional volatility, and trajectory continuity) calculated from the trajectory cache. Combined with the previous state S(t–1), the system queries the state diagram to determine if legal transition conditions are met. If so, the system enters S(t); otherwise, it remains in the previous state. To enhance discrimination stability, the system introduces a time threshold constraint (e.g., 1.8 seconds) into the state diagram, requiring that state transitions be triggered only after the current state has persisted for a period exceeding this threshold, to prevent misjudgments caused by short-term disturbances.

[0046] Furthermore, the system introduces a "fuzzy boundary" mechanism to handle fuzzy discrimination bands between states. When an observation falls into the overlapping interval between two states, the system does not switch states directly. Instead, it enters an intermediate buffer state and continues observing for a complete sliding cycle. If the state criteria continue to meet a certain direction, the transition is confirmed; otherwise, the original state is maintained, thus avoiding the "flickering switching" phenomenon.

[0047] This motion pattern discrimination method based on state diagrams and time consistency constraints has higher stability, interpretability, and deployment reliability than traditional logic rules based on single-frame thresholds. It is particularly suitable for scenarios with continuous perception capabilities, and effectively improves the system's ability to model complex trajectory behaviors and suppress misjudgments.

[0048] The path planning module dynamically generates a feasible path with high confidence and fault tolerance based on the aforementioned semantic category classification and motion pattern recognition results, combined with the indoor map topology, obstacle distribution, and navigation location confidence constraints. This module leverages technologies such as enhanced path graph construction, dynamic cost assignment, and real-time search algorithms to achieve stable and controllable path generation in dynamic environments.

[0049] In each navigation cycle, the system first constructs an enhanced path map for the current environment based on the indoor topology map. This path map uses the building plan structure diagram as its topological basis and dynamically injects obstacle position information and behavior labels obtained from the multi-source perception fusion module and the motion pattern modeling module. In terms of constructing the obstacle impact area, the system sets different area attributes based on semantic categories and motion states: for example, the area where the target is a liquid or flexible obstruction and is in an unstable moving state will be set as a "high-risk soft exclusion zone", and the corresponding path node weight will be significantly increased; while the stationary furniture-type targets will only serve as a path cost increase area without blocking traffic, retaining the ability to circumvent flexibly.

[0050] Next, the system evaluates the user's current positioning accuracy and perception system confidence, combining it with trajectory prediction error to construct a path tolerance zone. If there is a significant deviation in the positioning error estimate or obstacle trajectory prediction interval, the system will establish redundant pass zones in the path map, prioritizing path segments with high branch density and ample local detour space. This ensures that the path can be quickly switched in the event of sudden occlusion or temporary environmental changes.

[0051] The system then assigns a multi-factor dynamic cost value to all node edges in the enhanced path graph. This cost value not only considers the geometric distance between nodes but also introduces the following three core control factors: (1) Obstacle Risk Factor: This factor combines semantic labels and behavioral state levels to quantify the degree to which the target hinders the feasibility of passage; (2) Path Substitutability Index: This factor evaluates the flexibility of passage based on the number of local path branches and the width of the obstacle avoidance space; and (3) Trajectory Prediction Confidence Penalty: This factor sets a path cost buffer in areas where there is high uncertainty about the target's future movement to improve the stability and accessibility of the path.

[0052] Based on the cost graph, the system uses the D* Lite path planning algorithm to search for a primary path. This D* Lite path planning algorithm features dynamic replanning capabilities, allowing for rapid adjustments to local path structures as environmental conditions change. During the search, the system uses the current user node as the starting point and, based on the terminal's target location and a defined heuristic distance function, generates a primary path and a set of alternative paths. Each path records the planning timestamp, accumulated cost, and key turning points.

[0053] Ultimately, the system encapsulates the path planning results into a structured path description object and outputs it to the voice navigation module and the visual interaction module. This path description object includes information such as a sequence of path nodes, path segment cost values, risk warning point markers, key path turning points, and their corresponding semantic labels (e.g., "need to avoid moving targets" or "path is narrow"). If the path contains ambiguous sections requiring user decision-making (e.g., areas with heavy service robot traffic or gaps between temporary storage objects), the system will append a "Pending Strategy" tag to the path description, detecting and adjusting the traffic strategy in real time during the navigation execution phase.

[0054] The voice navigation module converts the structured route description information output by the route planning module into user-friendly natural language instructions. Taking into account the user's current state, characteristics, and interaction preferences, it delivers real-time navigation instructions, guiding them through dynamic indoor environments. As the system's human-computer interaction terminal, this module not only provides route guidance but also handles environmental change feedback, abnormal status notifications, and strategy confirmation guidance, ultimately forming the final link in the communication process from route to user.

[0055] During each navigation cycle, the voice navigation module first receives the path description object output by the route planning module and parses the path node sequence, path segment cost values, key inflection point markers, risk warning points, and pending strategy labels. Based on this information, the system dynamically extracts the current valid path segment based on the user's current location, direction, movement speed, and historical interaction records, and creates a semantic navigation template for the current navigation window.

[0056] The semantic navigation template includes three types of voice prompt information: First, path guidance prompts inform users of travel directions and path structures. For example, "Turn left five meters ahead," "Go straight along the corridor for fifteen meters and you'll enter the elevator area," etc. These statements are generated based on the geometric relationships between path nodes and spatial orientation reasoning results, and simplified processing improves language clarity and response efficiency. Second, environmental avoidance prompts inform users of variable obstacles or risk areas ahead. Examples include: "A service robot is approaching you. Please keep to the right," "The area ahead is wet. Please watch your step." These statements are generated by the preceding module based on a combined judgment of the obstacle semantic category and motion pattern. The language intensity and urgency level are then adjusted based on the decreasing confidence level of the target's trajectory prediction. Thirdly, system status feedback prompts are used to inform users of positioning status, navigation status, or pending strategy confirmation. Examples include: "The current positioning error is large. Please wait for repositioning," "There is an unstable target ahead on the path. Do you want to detour?", "We are replanning your route. Please wait." These statements can enhance the interactive experience through voice template invocation, intonation adjustment, and synthesized speech speed control.

[0057] After generating the voice broadcast content, the system invokes a speech synthesis engine deployed on an edge computing node to generate a voice stream signal adapted to the current language environment and user preferences. The speech engine supports switching between multiple languages, including Mandarin, Cantonese, and English. It also supports speech rate adjustment, naturalization of voice, and keyword emphasis, ensuring clarity and listening comfort.

[0058] In addition, to deal with multi-source interference in the environment and delayed user response, the system has designed a "redundant confirmation broadcast mechanism" for navigation instructions. That is, when the user fails to complete the navigation status switch within a certain time period (such as not turning, not avoiding obstacles, etc.), the system will automatically give a second prompt without interfering with the user's movement, and try to switch the prompt semantic style (such as switching from "Please turn left" to "You are approaching the left channel entrance").

[0059] If the path description object contains the "Pending Strategy" tag (such as a temporary storage area, a service robot intersection area, etc.), the system will generate guiding voice clarification statements, such as "Please confirm whether there is a gap ahead for passage", "Do you see a robot approaching? If so, please avoid the right", etc., and record the user behavior results for subsequent model optimization and policy library updates.

[0060] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0061] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0062] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0063] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0064] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. Indoor barrier-free path dynamic planning and voice navigation system, characterized by: Includes the following modules: The multi-source perception fusion module is used to fuse the collected synchronous observation data, construct a unified multimodal intermediate representation tensor, and input it into a pre-trained multi-task neural network model to achieve target geometry recognition, material type judgment, and heat source state classification, thereby obtaining the target semantic category label. The synchronous observation data includes 2D color images and depth maps, point cloud data, and temperature distribution maps. The motion pattern modeling module is used to establish a short-term trajectory buffer of dynamic length based on the identified target semantic category label. The current target trajectory buffer length is adaptively adjusted based on four regulatory factors: the target semantic category label, expected residual time, local path feasibility redundancy, and user physiological response margin index. Within this short-term trajectory buffer of dynamic length, the target's average linear velocity, directional volatility index, and trajectory continuity parameter are extracted, and a finite state machine is constructed to achieve dynamic identification of the target's motion state and output target state information for path planning.

2. The indoor barrier-free path dynamic planning and voice navigation system according to claim 1, characterized in that: The multi-task neural network model includes a shared encoder and three semantic branch networks corresponding to geometric form, material type and heat source status respectively. The shared encoder adopts a residual convolution structure and outputs a shared feature map. The semantic branches all output category labels and confidences in a Softmax manner and are jointly trained through a weighted multi-task loss function.

3. The indoor barrier-free path dynamic planning and voice navigation system according to claim 1, characterized in that: The expected residual time is determined by three parts: the average residence time of targets of the same category in the area is used as a baseline value; a speed suppression factor between zero and one is introduced according to the current moving speed of the target to monotonically reduce the baseline value; and the gain of the reduced value is adjusted according to the target's trajectory loop degree and the loop enhancement coefficient.

4. The indoor barrier-free path dynamic planning and voice navigation system according to claim 3, characterized in that: The degree of trajectory looping refers to the spatial overlap between the target motion path and its historical trajectory.

5. The indoor barrier-free path dynamic planning and voice navigation system according to claim 1, characterized in that: The local path feasibility redundancy is determined by the following factors: the number of alternative paths that can be selected from the current position and the margin of the average width of the avoidance space along each pass path relative to the system's preset standard detour width.

6. The indoor barrier-free path dynamic planning and voice navigation system according to claim 1, characterized in that: The user physiological response margin index is determined by taking the maximum acceptable delay time of the user during the navigation process as the upper limit and combining the user's current movement speed to estimate the time margin for completing prompts, turns, or deceleration within a given lead time; among which, the maximum acceptable delay time is obtained based on historical interaction records.

7. The indoor barrier-free path dynamic planning and voice navigation system according to claim 1, characterized in that: The finite state machine contains four motion states: completely stationary, temporarily stationary but can be pushed, stable autonomous movement, and highly volatile or unpredictable movement.

8. The indoor barrier-free path dynamic planning and voice navigation system according to claim 1, characterized in that: It also includes a path planning module, which is used to combine the target semantic category label and target motion state information, refer to the indoor map topology structure and the dynamic distribution state of obstacles, build an enhanced path map and dynamically assign values ​​to path nodes and path edges, and generate a main path and alternative path set with fault tolerance and high confidence reachability based on the D*Lite algorithm.

9. The indoor barrier-free path dynamic planning and voice navigation system according to claim 8, characterized in that: It also includes a voice navigation module, which is used to generate voice prompt content by combining the main path and alternative path set output by the path planning module with the user's current status.

10. The indoor barrier-free path dynamic planning and voice navigation system according to claim 9, characterized in that: The voice navigation module dynamically generates three types of voice prompts, including path guidance prompts, environment avoidance prompts, and system status feedback prompts, based on the user's current location, facing direction, and historical interaction records.

Citation Information

Cited By

  • Multi-mode interactive parking space digital voice guiding system

    CN121034123A

  • Humanoid robot navigation method, device and equipment based on defect complementation

    CN121163530A