OVER-NAV vision-language combined navigation system and method thereof

The OVER-NAV vision-language joint navigation system achieves deep integration of visual information, natural language commands, and semantic maps, solving the positioning and understanding problems of traditional navigation systems in complex environments and providing a more accurate and intuitive navigation experience, especially for visually impaired people.

CN120800418APending Publication Date: 2025-10-17GUANGDONG UNIV OF PETROCHEMICAL TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510917359.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional navigation systems are inaccurate in complex environments, have difficulty understanding user intent, and lack the integration of information from different modalities, resulting in inaccurate and unintuitive navigation, especially inadequate services for visually impaired individuals.

Method used

The OVER-NAV vision-language joint navigation system is adopted. Through multimodal deep fusion technology, it combines visual information, natural language commands and semantic maps. It utilizes a topology-aware heterogeneous modal representation unified framework, a differential geometry-inspired dynamic attention guidance mechanism and a heterogeneous knowledge integration and collaborative optimization framework to achieve deep semantic fusion of visual, linguistic and map information.

Benefits of technology

It improves navigation accuracy by more than 35%, enhances the human-computer interaction experience, reduces response latency to 0.3 seconds, adapts to user language habits, maintains a recognition accuracy of over 90% in complex environments, achieves a navigation success rate of 82%, and supports the navigation needs of visually impaired individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120800418A_ABST
    Figure CN120800418A_ABST
Patent Text Reader

Abstract

The invention discloses an OVER-NAV vision-language combined navigation system and method, and belongs to the technical field of intelligent navigation.The system comprises a navigation server, a user handheld terminal and a navigation terminal.The navigation server comprises a vision module, a language obtaining module, a vision map module, a semantic map module and a multi-modal feature cross-modal interaction module; the core of the system is that a multi-modal feature cross-modal interaction module comprises a topology-aware heterogeneous modal representation unified framework, a differential geometry inspired dynamic attention guidance mechanism and a heterogeneous knowledge integration and collaborative optimization framework, and deep fusion of three modals of visual information, a language instruction and a semantic map is realized; according to the method, the navigation accuracy is remarkably improved, the natural language instruction understanding accuracy reaches 92%, the high recognition accuracy is still kept in illumination changing and noisy environments, and the method is particularly suitable for complex shopping mall environments, outdoor-indoor seamless transition navigation and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent navigation, in particular to a multimodal navigation system combining visual information and natural language instructions and a navigation method thereof. BACKGROUND

[0002] With the continuous development of artificial intelligence technology, navigation systems are developing towards being more intelligent and more natural. Traditional navigation systems mainly rely on single modal information such as GPS signals or visual maps, which often have problems such as inaccurate positioning and difficulty in understanding user intentions in complex environments. Especially for special groups such as visually impaired people, traditional navigation systems are difficult to provide intuitive and accurate navigation guidance.

[0003] In the prior art, visual navigation systems usually use feature point matching and SLAM algorithms to construct environment maps, but lack deep understanding of user language instructions; while language-based navigation systems can understand user instructions, they are difficult to accurately perceive the environment. In addition, the fusion of different modal information often stays at the level of simple superposition, lacking deep semantic fusion, making it difficult for the system to cope with complex and variable actual navigation scenarios.

[0004] Therefore, there is an urgent need for a navigation system that can deeply integrate visual information and language instructions and has intelligent decision-making capabilities to provide more accurate and natural navigation experiences. SUMMARY

[0005] The purpose of the present application is to provide an OVER-NAV visual-language joint navigation system and method, which realizes the seamless integration of visual information, natural language instructions and semantic maps through multimodal deep fusion technology, and provides more accurate and intelligent navigation services for users.

[0006] The present application proposes an OVER-NAV visual-language joint navigation system, which comprises: a navigation server comprising a visual module, a language acquisition module, a visual map module and a semantic map module; The visual module is used to obtain environment images, obstacles and road boundary information through a camera, a depth sensor and a LiDAR sensor, and send the obtained information to the visual map module. The language acquisition module is used to collect natural language instructions, extract language keywords from the natural language instructions, and send text information containing the language keywords to the semantic map module. The visual map module is used to process the information sent by the visual module using a SLAM algorithm to form a visual map, and convert the visual map into a point cloud format containing obstacle coordinate point information. The semantic map module is configured to associate the language keywords with the visual map, construct a semantic map describing instruction semantics, and update the semantic map in real time. The navigation server further comprises a multi-modal feature cross-modal interaction module configured to associate the visual map with the semantic map and generate navigation instructions. The multi-modal feature cross-modal interaction module comprises a topologically-aware heterogeneous modal representation unification framework, a differential geometry-inspired dynamic attention guidance mechanism, and a heterogeneous knowledge integration and collaborative optimization framework. A user handheld terminal is wirelessly connected to the navigation server and configured to send voice data and positioning signals to the navigation server and receive navigation instructions returned by the navigation server. A navigation terminal is wirelessly connected to the navigation server and configured to receive navigation instructions sent by the navigation server and guide the user to proceed.

[0007] Preferably, the topologically-aware heterogeneous modal representation unification framework comprises: A multi-representation space construction unit is configured to map visual data into a multi-level feature pyramid structure, map language data into a hierarchical semantic tree, and map map data into a graph structure representation. An inter-modal isomorphic mapping unit is configured to establish a visual-language mapping network, a language-map mapping network, and a visual-map mapping network to realize conversion between different modal representations. A multi-scale feature collaborative extraction unit is configured to extract and integrate feature information of different scales through a local feature extraction layer, a global context acquisition layer, a multi-layer feature fusion unit, and a boundary enhancement module.

[0008] Preferably, the differential geometry-inspired dynamic attention guidance mechanism comprises: A curvature-aware attention flow field construction unit is configured to calculate a local curvature index in a feature space, construct an attention flow field direction vector, and generate an attention distribution adaptive to local geometric characteristics. A multi-level attention allocation unit is configured to generate a global attention layer, a local attention layer, and a semantic-guided attention layer, and dynamically adjust a global-local balance coefficient according to task complexity. An adaptive attention update unit is configured to calculate the importance of each modal feature based on information entropy, adjust attention weights according to gradient information of a downstream task, and generate a final attention distribution by applying a temporal consistency constraint.

[0009] Preferably, the heterogeneous knowledge integration and collaborative optimization framework comprises: A multi-level knowledge representation and fusion unit is configured to maintain a modal-specific knowledge base, construct a cross-modal association graph, and form a unified knowledge representation. a dynamic knowledge flow and integration unit for controlling the flow direction and intensity of knowledge between different modalities, identifying and resolving conflicts between different modalities of knowledge, and evaluating the reliability and applicability of knowledge; a global consistency optimization unit for ensuring the internal consistency, environmental consistency and instruction consistency of the fused knowledge, and updating the unified knowledge representation through iterative optimization.

[0010] As a preferred embodiment, the visual module collects environmental information through multiple sensors and performs noise reduction, calibration and alignment processing; the language acquisition module analyzes the syntax and semantic structure of the instructions and determines the user's navigation intention; the visual map module uses the SLAM algorithm to construct an initial environmental map; the semantic map module generates a semantic enhanced map combining visual information and language instructions.

[0011] As a preferred embodiment, the navigation server further includes a path planning module for generating multiple candidate paths based on the fusion information generated by the multi-modal feature cross-modal interaction module, evaluating the safety, efficiency and user preference of the candidate paths, and determining the optimal path.

[0012] As a preferred embodiment, the navigation server further includes a user feedback processing module for receiving user confirmation or correction of navigation results, analyzing the semantics of user feedback, and updating the user preference model.

[0013] As a preferred embodiment, the navigation server further includes an environmental adaptation module for real-time updating of environmental maps and semantic information, adjusting system parameters to adapt to environmental changes, and recording navigation history to improve future navigation.

[0014] As a preferred embodiment, the navigation server, the user handheld terminal and the navigation terminal communicate information through a standardized tensor interface, a hierarchical caching mechanism, an asynchronous communication framework and adaptive compression encoding.

[0015] The OVER-NAV visual-language joint navigation method includes the following steps: Collecting environmental image, obstacle and road boundary information through the visual module, and sending the collected information to the visual map module; Collecting natural language instructions through the language acquisition module, extracting language keywords from the natural language instructions, and sending text information containing the language keywords to the semantic map module; Processing the information sent by the visual module to form a visual map using the SLAM algorithm through the visual map module, and converting the visual map into a point cloud format containing obstacle coordinate point information; Associating the language keywords with the visual map through the semantic map module, constructing a semantic map describing the semantics of the instructions, and performing real-time updating of the semantic map; associating the visual map with the semantic map through a multi-modal feature cross-modal interaction module, the multi-modal feature cross-modal interaction module comprising a topologically-aware heterogeneous modal representation unification framework, a differential geometry-inspired dynamic attention guidance mechanism, and a heterogeneous knowledge integration and collaborative optimization framework; generating navigation instructions and sending them to the user's handheld terminal and the navigation terminal; guiding the user to proceed through the navigation terminal.

[0016] The present application has the following beneficial effects: 1. Multi-modal deep fusion: Through the innovative topologically-aware representation framework and the differential geometry-inspired attention mechanism, deep semantic fusion of visual, language and map information is achieved, greatly improving the system's understanding ability of complex environments and instructions, with navigation accuracy improved by more than 35% compared to traditional systems.

[0017] 2. Enhanced human-computer interaction experience: The system's understanding accuracy of natural language instructions reaches 92%, response delay is reduced to 0.3 seconds, and it can adapt to the user's personal language habits, significantly improving the user experience.

[0018] 3. Complex environment adaptability: In scenes with drastic changes in lighting conditions, recognition accuracy remains above 90%; in noisy environments, instruction understanding accuracy remains at 85%; in complex environments with occlusions, navigation success rate reaches 82%.

[0019] 4. Special population barrier-free support: The system can understand the possibly inaccurate instructions issued by visually impaired people based on limited environmental perception and correctly interpret them, providing more comprehensive environmental descriptions for visually impaired users and enhancing their spatial perception ability.

[0020] 5. Robustness and adaptability: Through multi-modal redundancy design and graceful degradation mechanism, the system can maintain normal core functions when some functions fail, with fault detection delay not exceeding 500 milliseconds and slight faults recovered within 5 seconds. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is the overall architecture schematic diagram of the OVER-NAV visual-language joint navigation system of the present application; Figure 2 is the structure schematic diagram of the multi-modal feature cross-modal interaction module of the present application; Figure 3 is the workflow diagram of the topologically-aware heterogeneous modal representation unification framework of the present application; Figure 4 is the structure schematic diagram of the differential geometry-inspired dynamic attention guidance mechanism of the present application; Figure 5A workflow diagram of the isomerism knowledge integration and collaborative optimization framework of the present application; Figure 6 A flowchart of the OVER-NAV visual-linguistic joint navigation method of the present application. DETAILED DESCRIPTION

[0022] Reference will now be made to the drawings, and specific examples will be described in detail. Figures 1-6 Those skilled in the art will understand that these examples are only used to illustrate the present application, and do not limit the scope of the present application.

[0023] With reference to Figure 1 The OVER-NAV visual-linguistic joint navigation system provided by the present application includes a navigation server 1, a user handheld terminal 2 and a navigation terminal 3.

[0024] The navigation server 1 includes a visual module 11, a language acquisition module 12, a visual map module 13, a semantic map module 14 and a multi-modal feature cross-modal interaction module 15. The navigation server 1 can also include a path planning module 16, a user feedback processing module 17 and an environment adaptation module 18.

[0025] The visual module 11 is used to acquire environmental images, obstacles and road boundary information through a camera, a depth sensor and a LiDAR sensor, and send the acquired information to the visual map module 13. In an embodiment of the present application, the visual module 11 uses an RGB camera with a resolution of 1920x1080 to collect environmental images, a ToF depth camera to acquire obstacle information, and a 16-line LiDAR scanner to acquire road boundary information, with a sampling frequency of 10Hz. For example, in a shopping mall navigation scenario, the visual module 11 can simultaneously capture various environmental elements such as store signs, pedestrians, chairs, etc., providing comprehensive visual information for subsequent navigation.

[0026] The language acquisition module 12 is used to collect natural language instructions, extract language keywords from the natural language instructions, and send text information containing language keywords to the semantic map module 14. Preferably, the language acquisition module 12 uses a semantic analysis algorithm based on an attention mechanism to identify keywords such as location (starting point, ending point, passing point), direction, distance, time and condition in the instruction, with a vocabulary size of 50,000 words and an identification accuracy of 95%. In actual application, when the user says "Please take me to the coffee shop in front of me, avoiding crowded places", the language acquisition module 12 can accurately extract "coffee shop" as the destination and "avoid" and "crowded places" as the path constraint conditions.

[0027] The visual map module 13 is used to process the information sent by the visual module 11 to form a visual map using a SLAM algorithm, and convert the visual map into a point cloud format containing obstacle coordinate point information. In a preferred embodiment of the present application, the visual map module 13 uses the ORB-SLAM2 algorithm to construct the visual map, with a global map resolution of 0.5 meters / pixel and a local map resolution of 0.1 meters / pixel, and 1-3 topological nodes are set for every 10 square meters. This resolution setting can ensure navigation accuracy while controlling computing resource consumption, making it suitable for real-time navigation scenarios.

[0028] The semantic map module 14 is used to associate language keywords with the visual map, construct a semantic map describing the semantics of the instructions, and update the semantic map in real time. In addition, the semantic map module 14 can also add semantic labels to the map area, including roads, intersections, obstacles, markers, and functional areas, etc. For example, in a hospital navigation scenario, the semantic map module 14 can associate language keywords such as "radiology department" and "outpatient hall" with the corresponding areas in the visual map, making it easier for the system to understand the spatial meaning of the user's instruction "take me to the radiology department for examination".

[0029] The multi-modal feature cross-modal interaction module 15 is the core innovative module of the system, which is used to associate the visual map with the semantic map to generate navigation instructions. This module includes a topologically aware heterogeneous modal representation unification framework 151, a differential geometry inspired dynamic attention guidance mechanism 152, and a heterogeneous knowledge integration and collaborative optimization framework 153, which will be described in detail later.

[0030] The user's handheld terminal 2 is wirelessly connected to the navigation server 1, which is used to send voice data and positioning signals to the navigation server 1, and receive navigation instructions returned by the navigation server 1. In an embodiment of the present application, the user's handheld terminal 2 can be a smartphone or a dedicated navigation device, which communicates with the navigation server 1 through a WiFi6 or 5G network with a transmission delay of less than 50 milliseconds.

[0031] The navigation terminal 3 is wirelessly connected to the navigation server 1, which is used to receive the navigation instructions sent by the navigation server 1 and guide the user to proceed. Preferably, the navigation terminal 3 can be a smart glasses, a smart earphone, or a vibrating bracelet, etc. wearable device, which provides visual, auditory or tactile feedback to the user.

[0032] Reference Figure 2 The multi-modal feature cross-modal interaction module 15 is the core module of the present application, which includes a topologically aware heterogeneous modal representation unification framework 151, a differential geometry inspired dynamic attention guidance mechanism 152, and a heterogeneous knowledge integration and collaborative optimization framework 153.

[0033] Reference Figure 3The topological-aware heterogeneous modal representation unification framework 151 includes a multi-representation space construction unit 1511, an inter-modal isomorphism mapping unit 1512, and a multi-scale feature collaborative extraction unit 1513.

[0034] The multi-representation space construction unit 1511 is configured to map visual data into a multi-level feature pyramid structure, map language data into a hierarchical semantic tree, and map map data into a graph structure representation. Specifically, the visual data representation adopts a multi-level feature pyramid structure, which includes five scale levels, and each level is represented by a 32, 64, 128, 256, and 512-dimensional vector, respectively, to represent the features of an image region; the language data representation constructs a hierarchical semantic tree, which includes three layers of structure at the word level (up to 30 word units), the phrase level (up to 15 phrases), and the sentence level (overall semantics); and the map data representation adopts a graph structure representation, in which a node represents a location (including location coordinates and feature description), and an edge represents a connection relationship (including direction and distance information).

[0035] For example, when a user says "help me find a nearby children's clothing store" in a shopping mall environment, the system constructs the instruction as a semantic tree: the root node is the overall sentence semantics "find a location", the second layer nodes are "nearby" and "children's clothing store", and the third layer further expands the specific features of "children" and "clothing store". At the same time, the visual feature pyramid analyzes the environment image from coarse to fine, and the low-level captures the overall layout, and the high-level identifies specific store signs and product features.

[0036] The inter-modal isomorphism mapping unit 1512 is configured to establish a visual-language mapping network, a language-map mapping network, and a visual-map mapping network to realize the conversion between different modal representations. In actual applications, the mapping accuracy threshold is set to 0.85 to ensure the accuracy of inter-modal mapping; the feature dimension consistency constraint maps features of different dimensions to a unified 256-dimensional space through a projection matrix; and the topological preservation constraint maintains the neighborhood structure during the mapping process, with a similarity difference of no more than 0.2.

[0037] The inter-modal isomorphism mapping adopts the following mathematical expression: , Wherein: is a mapping function used to map source modal features and target modal features to a unified feature space; is a source modal feature, such as a feature vector from a visual module; is a target modal feature, such as a feature vector from a language module; represents a feature concatenation operation that connects two feature vectors in dimension; is a weight matrix used to learn the mapping relationship between different modalities; is an offset term used to adjust the offset of the mapping function; For the activation function, ReLU or tanh function is usually adopted to increase the non-linear ability of the mapping. The weight matrix The dimension is wherein , and are the source modality and target modality feature dimensions, respectively, is the feature dimension after mapping, which is uniformly set to 256.

[0038] In an actual navigation scene, such mapping can associate the language concept of "cafe" with the appearance features of the cafe in the visual image and the location information on the map. When the user asks "Is there a cafe nearby?", the system can identify the sign of the cafe in the visual scene and locate the corresponding position on the map.

[0039] The multi-scale feature collaborative extraction unit 1513 is used to extract and integrate feature information of different scales through a local feature extraction layer, a global context acquisition layer, a multi-layer feature fusion unit and a boundary enhancement module. The local feature extraction layer adopts a 3x3 convolution kernelx4 layers, and the feature map size is halved layer by layer; the global context acquisition layer adopts multi-scale dilated convolution with a dilated rate of 1, 2, 4 and 8; the multi-layer feature fusion unit adopts an adaptive feature fusion module to dynamically adjust the weight of different layer features according to the amount of information; the boundary enhancement module focuses on feature extraction of object boundaries and spatial transition areas to improve spatial positioning accuracy.

[0040] The multi-scale feature fusion adopts the following formula: , wherein: is the fused multi-scale feature, which is the weighted sum of the features of each layer; is the feature of the i-th layer, representing the feature extracted at different scale levels; is the weight of the i-th layer feature, determining the contribution of each layer feature to the final fusion result; is the number of feature layers, which is set to 5 in this embodiment. The weight is calculated by the following formula: , , wherein: is the information entropy of the i-th layer feature, measuring the amount of information contained in the feature of the i-th layer; is a learnable parameter, used to adjust the influence degree of information entropy on the weight; the denominator term is a normalization factor to ensure that the sum of all weights is 1. The higher the information entropy, the richer the information contained in the feature of the i-th layer, and the greater the weight.

[0041] ​In the shopping mall navigation scenario, when the system needs to recognize "elevator", the multi-scale feature fusion can simultaneously consider the overall contour of the elevator (from low-level features) and the detailed features such as buttons, digital display (from high-level features), greatly improving the recognition accuracy. Even in the case of dense flow and partial occlusion, the system can reliably identify the elevator position.

[0042] With reference to Figure 4 , the differential geometry inspired dynamic attention guiding mechanism 152 includes a curvature-aware attention flow field construction unit 1521, a multi-level attention distribution unit 1522, and an adaptive attention updating unit 1523.

[0043] The curvature-aware attention flow field construction unit 1521 is used to calculate the local curvature index in the feature space, construct the attention flow field direction vector, and generate the attention distribution that adapts to the local geometric characteristics. Specifically, the attention tensor has a shape of [batch size, source modality feature number, target modality feature number, attention dimension]; the flow field direction vector describes the flow direction of attention in the feature space, and has the same dimension as the feature dimension; the local curvature index describes the geometric characteristics of the local feature space, and is used to adjust the attention distribution strategy.

[0044] The local curvature is calculated using the following formula: , wherein: is the local curvature estimate value of point , which measures the bending degree of the feature space near the point; is a neighborhood point of , which is a point close to in the feature space; is the number of neighborhood points, which is usually 5-20 points; denotes the Euclidean distance from point to neighborhood point ; is the average distance from point to its neighborhood set , which is calculated as , wherein is the neighbor point of . Preferably, the value of k is set to 10, which ensures sufficient local geometric information capture capability.

[0045] In real navigation scenarios, when the user walks in a shopping mall, the system calculates the local curvature of the current visual feature space to determine the complexity of the environment. In an open hall area (low curvature area), the system pays more attention to distant signs and overall layout; while in a densely packed corridor (high curvature area), the system pays more attention to nearby detailed features and obstacles, thus providing more accurate navigation guidance.

[0046] The multi-level attention allocation unit 1522 is used to generate global attention layers, local attention layers, and semantic-guided attention layers, and dynamically adjust the global-local balance coefficient according to the task complexity. The global attention layer captures the overall correlation between modalities, with a dimension of [source modality number, target modality number]; the local attention layer focuses on the detail matching of the local area, with a dimension of [source feature number, target feature number]; the semantic-guided attention layer adjusts the attention allocation based on semantic correlation, with a dimension of [semantic category number, feature number].

[0047] The multi-level attention calculation formula is: , Wherein: is the final attention distribution, used to guide the system to pay attention to the importance of different features; is the global attention, focusing on the overall inter-modal relationship; is the local attention, focusing on local feature matching; is the semantic-guided attention, adjusting the attention allocation based on semantic correlation; is the global-local balance coefficient, controlling the proportion of global and local attention; represents the Hadamard product (element-level multiplication), which multiplies the elements at corresponding positions of two matrices. The global-local balance coefficient is initially set to 0.6 and dynamically adjusted according to the task complexity. The more complex the task, the smaller the value, the more inclined to focus on local details.

[0048] For example, when the user requests "take me to the radiology department" in the hospital lobby, the system first captures the overall relationship between the concept of "radiology department" and the hospital layout through global attention, determining the general direction. Subsequently, through local attention and semantic-guided attention, the system will focus on signs, door numbers, and other detailed features to ensure the accuracy of navigation. If the environment is simple and clear, the system will focus more on global attention; while in a complex and unclear environment, the system will automatically reduce the value and enhance the attention to local details.

[0049] The adaptive attention update unit 1523 is used to calculate the importance of each modality feature based on information entropy, adjust the attention weight according to the gradient information of the downstream task, and generate the final attention distribution by applying the temporal consistency constraint. In the preferred embodiment of the present application, the change rate of the attention distribution of adjacent time steps is controlled within 30%, ensuring the stability of the attention allocation.

[0050] Information entropy calculation formula: , Where: is the information entropy of the feature , quantifying the amount of information contained in the feature; is the probability of the feature value , usually obtained by normalizing the histogram of the feature value; is the number of possible feature values; denotes the logarithm function with base 2. The area with high feature entropy value contains more information and obtains higher attention weight.

[0051] Temporal consistency constraint formula: , Where: and are the attention distribution matrices of the current time step and the previous time step, respectively; denotes the Frobenius norm, calculated as the square root of the sum of the squares of all elements in the matrix, i.e. , where is the element of the matrix ; is the change rate threshold, set to 0.3, limiting the change amplitude of the attention distribution between adjacent time steps.

[0052] In actual navigation process, this temporal consistency constraint prevents the system attention from fluctuating sharply, providing a smooth navigation experience. For example, when the user is walking in the mall, even if there is a sudden short-term disturbance in the environment (such as people passing through), the system can maintain stable attention on the navigation target, avoiding unnecessary path adjustment or instruction change.

[0053] Referring to Figure 5 , the heterogeneous knowledge integration and collaborative optimization framework 153 includes a multi-level knowledge representation and fusion unit 1531, a dynamic knowledge flow and integration unit 1532, and a global consistency optimization unit 1533.

[0054] The multi-level knowledge representation and fusion unit 1531 is used to maintain modal-specific knowledge bases, construct cross-modal association graphs, and form unified knowledge representations. The modal-specific knowledge base contains entities (up to 1000), relationships (up to 50 categories), and attributes (up to 20 per entity); the cross-modal association graph represents the mapping relationship between different modal knowledge, containing nodes (knowledge entities) and edges (association types and strengths); the unified knowledge representation integrates the entity and relationship of multi-modal information.

[0055] Knowledge representation fusion formula: , Where: is the unified knowledge representation, which integrates the knowledge information of each modality; is the knowledge representation of the i-th modality, such as visual knowledge, language knowledge, or map knowledge; is the conversion function that maps the i-th modality knowledge to the unified representation space, usually implemented as a neural network; is the weight of the i-th modality, which determines the contribution of each modality knowledge to the unified representation; is the number of modalities, which is 3 in this system, corresponding to visual, language, and geographic modalities. The weight is dynamically adjusted according to the reliability of different modalities, for example, increasing the weight of the visual modality in a well-lit environment, and increasing the weight of the language modality in a complex instruction scenario. In the hospital navigation scenario, when the user requests "Take me to the CT room", the system will integrate three kinds of knowledge: language knowledge ("What is the CT room?"), visual knowledge (the appearance characteristics of the CT room), and map knowledge (the location and path of the CT room). In a noisy environment, the system will reduce the weight of the language modality; in a light-insufficient environment, it will reduce the weight of the visual modality; while the map knowledge usually maintains a relatively stable weight. Figure Three The dynamic knowledge flow and integration unit 1532 is used to control the flow direction and strength of knowledge between different modalities, identify and resolve conflicts between different modalities, and evaluate the reliability and applicability of knowledge. The knowledge flow threshold is set to 0.7 to control the selectivity of knowledge propagation; the conflict resolution strategy uses priority ordering (visual > map > language, but can be dynamically adjusted according to uncertainty); the knowledge integration temperature controls the smoothness of the fusion process, ranging from [0.1-1.0].

[0056] Knowledge flow intensity calculation formula:

[0057] , Where:

[0058] is the flow intensity from the i-th modality to the j-th modality; is the weight of the i-th modality; is the weight of the j-th modality; is the knowledge representation of the i-th modality; is the knowledge representation of the j-th modality. To the modal knowledge flow intensity, determine how much knowledge from the modal transferred to the modal ; For the modal and Similarity of knowledge, usually cosine similarity calculation; For temperature parameters, sensitivity of flow intensity; For sigmoid function, map input to [0,1] interval, form Temperature parameters The initial value is set to 0.5, which can be dynamically adjusted according to the difficulty of fusion task. The more complex the task is, The greater the value, the smoother the knowledge flow.

[0059] In actual navigation, this knowledge flow mechanism can handle inconsistencies between modalities. For example, when the visual system recognizes the "cafe" sign, but the map data shows that the location is "restaurant", the system will dynamically adjust the knowledge flow direction according to the reliability and recency of each modality, update the map knowledge or adjust the visual recognition result, to ensure the accuracy of navigation.

[0060] The global consistency optimization unit 1533 is used to ensure the internal consistency, environmental consistency and instruction consistency of the fused knowledge, and to update the unified knowledge representation through iterative optimization. Global consistency optimization uses a graph-based consistency scoring function: , Where: is the global consistency score of knowledge representation, which measures the overall consistency level of fused knowledge; is the internal consistency scoring function, which ensures that there is no contradiction within the knowledge representation; is the environmental consistency scoring function, which ensures that the knowledge is consistent with the current observed environment; is the instruction consistency scoring function, which ensures that the knowledge is consistent with the semantics of the user's instructions; , and are weight coefficients, which control the relative importance of the three consistencies. Preferably, , , , to ensure the balance of the three consistencies.

[0061] In complex mall navigation scenarios, global consistency optimization can handle multiple constraints. For example, when a user requests "take me to the nearest coffee shop, but avoid crowded areas," the system needs to balance three kinds of consistency: internal consistency (the path must be coherent and reasonable), environmental consistency (the path must be based on the actual observed environment), and instruction consistency (the path must satisfy the two possibly conflicting conditions of "nearest" and "avoid crowded areas"). Through iterative optimization, the system can find the best balance point and generate a navigation solution that satisfies multiple constraints.

[0062] The path planning module 16 is used to generate multiple candidate paths based on the fusion information generated by the multi-modal feature cross-modal interaction module 15, evaluate the safety, efficiency and user preference of the candidate paths, and determine the optimal path. In an embodiment of the present application, the number of candidate paths is 3-5 (in general cases) and at most 10 (in complex environments); the path evaluation dimensions include distance (30%), safety (30%), comfort (20%), and user preference (20%); the adjustment frequency is 5Hz when the environment changes rapidly and 1Hz when the environment is stable.

[0063] The user feedback processing module 17 is used to receive user confirmation or correction of the navigation result, analyze the semantics of the user feedback, and update the user preference model. Preferably, the user model dimensions include navigation preferences (10 categories), language expression habits (5 categories), and feedback modes (3 categories); the learning rate is initially 0.1 and decreases to 0.01 with the number of interactions; the memory capacity saves the last 100 interaction records, and important events are saved for a long time.

[0064] The environment adaptation module 18 is used to update the environment map and semantic information in real time, adjust the system parameters to adapt to environmental changes, and record the navigation history to improve future navigation. In an embodiment of the present application, the environment adaptation module 18 can identify and adapt to different scene types, including indoor (5 types), outdoor (8 types), and mixed (3 types) scenes; the parameter adjustment range is ±50% (based on the baseline configuration); the generalization ability evaluation index is that the first success rate in a new scene is not less than 70%.

[0065] The information transmission between the navigation server 1, the user handheld terminal 2 and the navigation terminal 3 is carried out through a standardized tensor interface, a hierarchical cache mechanism, an asynchronous communication framework and adaptive compression coding. The standardized tensor interface unifies the data transmitted between all modules into a standardized tensor format, including data itself, uncertainty estimation and timestamp information; the hierarchical cache mechanism sets up multi-level caches between modules, dynamically manages the caches according to the timeliness and importance of information, and reduces calculation redundancy; the asynchronous communication framework adopts an asynchronous communication mechanism based on a publish-subscribe mode, reduces the coupling degree between modules, and improves the parallel processing capability of the system; the adaptive compression coding dynamically adjusts the compression coding mode of information according to the characteristics of different modal information and the current system resource status, balances the communication efficiency and information fidelity.

[0066] With reference to Figure 6 The OVER-NAV visual-linguistic joint navigation method includes the following steps: S1: Collecting environment images, obstacles and road boundary information through the visual module 11, and sending the collected information to the visual map module 13.

[0067] S2: Collecting natural language instructions through the language acquisition module 12, extracting language keywords from the natural language instructions, and sending text information containing the language keywords to the semantic map module 14.

[0068] S3: Processing the information sent by the visual module 11 through the visual map module 13 to form a visual map using a SLAM algorithm, and converting the visual map into a point cloud format containing obstacle coordinate point information.

[0069] S4: Associating the language keywords with the visual map through the semantic map module 14, constructing a semantic map describing the semantics of the instructions, and updating the semantic map in real time.

[0070] S5: Associating the visual map with the semantic map through the multi-modal feature cross-modal interaction module 15 to generate navigation instructions.

[0071] S6: Sending the generated navigation instructions to the user handheld terminal 2 and the navigation terminal 3.

[0072] S7: Guiding the user to proceed through the navigation terminal 3.

[0073] In step S5, the working process of the multi-modal feature cross-modal interaction module 15 includes: S51: Mapping the visual data, language data and map data to a unified representation space through the topological perception heterogeneous modal representation unified framework 151, establishing a conversion relationship between different modal representations, and extracting multi-scale feature information.

[0074] S52: A dynamic attention guidance mechanism inspired by differential geometry152 is used to calculate the local curvature index in the feature space, generate a multi-level attention distribution, and dynamically adjust the attention weight based on information entropy and task gradient.

[0075] S53: Maintain the modality-specific knowledge base through the heterogeneous knowledge integration and collaborative optimization framework 153, control the flow of knowledge between different modalities, and ensure the global consistency of the fused knowledge.

[0076] The application of the system of the present invention in indoor navigation scenarios for visually impaired people has the following characteristics: The high-precision map resolution is set to 0.05 meters per pixel to ensure accurate positioning of obstacles; the voice interaction response time is less than 300 milliseconds, providing real-time feedback; the path planning adopts a safety-first strategy, and the avoidance distance is increased by 50% to ensure navigation safety.

[0077] For example, when a visually impaired user issues a command like "Take me to the conference room on the left," the system uses the language acquisition module 12 to extract the two keywords "left" and "conference room." The semantic map module 14 then associates these keywords with the visual map to determine the target location. The multimodal feature cross-modal interaction module 15 analyzes the user's current location, obstacle distribution, and target location to generate a safe path. The navigation terminal 3 then provides a voice prompt, "Go straight ahead for 5 meters, then turn right. Be careful, there's a chair ahead," to guide the user safely to their destination.

[0078] In a complex shopping mall environment, the system of the present invention sets the dynamic environment update frequency to 1 Hz, supports multi-layer maps of up to 10-story building structures, and adopts an enhanced crowd avoidance algorithm for predictive path planning.

[0079] When a user issues a command like "Take me to the coffee shop on the third floor," the system first determines the target location on the third floor and then plans a route that includes an elevator or escalator. During navigation, the system monitors crowd density in real time and dynamically adjusts the route to avoid crowded areas. Furthermore, the system provides rich descriptions of the surroundings, such as "There's a fountain 50 meters ahead. After passing the fountain, turn right to find the elevator," to enhance the user's spatial perception.

[0080] In the outdoor-indoor seamless transition scenario, the system integrates GPS, visual positioning and WiFi positioning technologies to enhance feature matching capabilities under multiple lighting conditions and establish correspondence between indoor and outdoor reference points to ensure navigation continuity.

[0081] For example, when a user enters a shopping mall from outdoors, the system will smoothly switch from GPS positioning to visual positioning, and the navigation instructions will be adjusted accordingly, from "200 meters north along the main road" to "turn left after entering the lobby, the elevator area is ahead", ensuring that the user receives continuous and consistent guidance throughout the navigation process.

[0082] Through the detailed description of the above embodiments, those skilled in the art should be able to understand the technical solutions and working principles of the OVER-NAV visual-linguistic joint navigation system and the method thereof, so as to realize the technical effects of the present application.

[0083] The above-described embodiments only express the specific implementation of the present application, which is described in a more specific and detailed manner, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that, for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application.

Claims

1. OVER-NAV vision-language joint navigation system, characterized by: include: Navigation server, including vision module, language acquisition module, visual map module and semantic map module; The visual module is used to obtain environmental images, obstacles and road boundary information through cameras, depth sensors and LiDAR sensors, and send the obtained information to the visual map module; The language acquisition module is used to collect natural language instructions, extract language keywords from the natural language instructions, and send text information containing the language keywords to the semantic map module; The visual map module is used to process the information sent by the visual module using the SLAM algorithm to form a visual map, and convert the visual map into a point cloud format containing obstacle coordinate point information; The semantic map module is used to associate the language keywords with the visual map, construct a semantic map describing the semantics of the instruction, and update the semantic map in real time; The navigation server further includes a multimodal feature cross-modal interaction module for associating the visual map with the semantic map to generate navigation instructions; The multimodal feature cross-modal interaction module includes a topology-aware unified framework for heterogeneous modal representation, a dynamic attention guidance mechanism inspired by differential geometry, and a heterogeneous knowledge integration and collaborative optimization framework; A user handheld terminal is wirelessly connected to the navigation server and is used to send voice data and positioning signals to the navigation server and receive navigation instructions returned by the navigation server; The navigation terminal is wirelessly connected to the navigation server and is used to receive navigation instructions sent by the navigation server and guide the user to move forward.

2. The OVER-NAV vision-language combined navigation system according to claim 1, characterized in that: The topology-aware unified framework for heterogeneous modal representation includes: A multi-representation space construction unit, used to map visual data into a multi-level feature pyramid structure, language data into a hierarchical semantic tree, and map data into a graph structure representation; Inter-modal isomorphic mapping unit, used to establish vision-language mapping network, language-map mapping network and vision-map mapping network to achieve conversion between different modal representations; The multi-scale feature collaborative extraction unit is used to extract and integrate feature information of different scales through the local feature extraction layer, the global context acquisition layer, the multi-layer feature fusion unit and the boundary enhancement module.

3. The OVER-NAV vision-language combined navigation system according to claim 1, characterized in that: The differential geometry-inspired dynamic attention guidance mechanism includes: A curvature-aware attention flow field construction unit is used to calculate the local curvature index in the feature space, construct the attention flow field direction vector, and generate an attention distribution that adapts to local geometric characteristics; A multi-level attention allocation unit, which generates global attention layers, local attention layers, and semantically guided attention layers, and dynamically adjusts the global-local balance coefficient according to task complexity; The adaptive attention update unit is used to calculate the importance of each modal feature based on information entropy, adjust the attention weight according to the gradient information of the downstream task, and apply temporal consistency constraints to generate the final attention distribution.

4. The OVER-NAV vision-language combined navigation system according to claim 1, characterized in that: The heterogeneous knowledge integration and collaborative optimization framework includes: A multi-level knowledge representation and fusion unit, used to maintain modality-specific knowledge bases, build cross-modality association graphs, and form a unified knowledge representation; Dynamic knowledge flow and integration unit, used to control the direction and intensity of knowledge flow between different modalities, identify and resolve conflicts between different modal knowledge, and evaluate the reliability and applicability of knowledge; The global consistency optimization unit is used to ensure the internal consistency, environmental consistency and instruction consistency of the fused knowledge, and to update the unified knowledge representation through iterative optimization.

5. The OVER-NAV vision-language combined navigation system according to claim 1, characterized in that: The vision module collects environmental information through multiple sensors and performs noise reduction, calibration and alignment processing; the language acquisition module analyzes the grammatical and semantic structure of the instructions and determines the user's navigation intention; The visual map module uses the SLAM algorithm to construct an initial environment map; the semantic map module combines visual information and language instructions to generate a semantically enhanced map.

6. The OVER-NAV vision-language combined navigation system according to claim 1, characterized in that: The navigation server also includes a path planning module for generating multiple candidate paths based on the fusion information generated by the multimodal feature cross-modal interaction module, evaluating the safety, efficiency and user preferences of the candidate paths, and determining the optimal path.

7. The OVER-NAV vision-language combined navigation system according to claim 1, characterized in that: The navigation server further comprises a user feedback processing module for receiving user confirmation or correction of navigation results, analyzing the semantics of user feedback, and updating the user preference model.

8. The OVER-NAV vision-language combined navigation system according to claim 1, characterized in that: The navigation server also includes an environment adaptation module for updating the environment map and semantic information in real time, adjusting system parameters to adapt to environmental changes, and recording navigation history to improve future navigation.

9. The OVER-NAV vision-language combined navigation system according to claim 1, characterized in that: Information is transmitted between the navigation server, the user handheld terminal and the navigation terminal via a standardized tensor interface, a layered cache mechanism, an asynchronous communication framework and adaptive compression coding.

10. OVER-NAV vision-language joint navigation method, characterized by: The following steps are involved: The visual module collects environmental images, obstacles, and road boundary information, and sends the collected information to the visual map module; Collecting natural language instructions through the language acquisition module, extracting language keywords from the natural language instructions, and sending text information containing the language keywords to the semantic map module; The visual map module processes the information sent by the visual module using the SLAM algorithm to form a visual map, and converts the visual map into a point cloud format containing obstacle coordinate point information; Associating the language keywords with the visual map through a semantic map module, constructing a semantic map that describes the semantics of the instruction, and updating the semantic map in real time; Associating the visual map with the semantic map through a multimodal feature cross-modal interaction module, wherein the multimodal feature cross-modal interaction module includes a topology-aware heterogeneous modal representation unified framework, a differential geometry-inspired dynamic attention guidance mechanism, and a heterogeneous knowledge integration and collaborative optimization framework; Generate navigation instructions and send them to the user's handheld terminal and navigation terminal; The user is guided forward by the navigation terminal.

Citation Information

Cited By

  • Underwater navigation method and system based on panoramic visual perception and language model reasoning

    CN121140810A

  • Underwater Navigation Method and System Based on Panoramic Visual Perception and Language Model Inference

    CN121140810B

  • Blind guiding robot interactive navigation system combining vision, inertial navigation and voice

    CN121346774A

  • Body type visual language navigation method based on modal quality dynamic regulation and control

    CN121655522A

  • Robot map asynchronous establishment method, electronic equipment and storage medium

    CN122170853A