A visual-language navigation method and system based on size model collaborative decision-making

CN122813831APending Publication Date: 2026-09-25TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610890614.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,该方案的协同框架仍局限于单一模型体系下的全局-局部模块划分,未解决大模型推理延迟与小模型泛化能力不足之间的核心矛盾;同时,其决策过程缺乏不确定度建模与自适应调整机制,无法根据环境动态变化对不同决策模式的权重进行动态优化,难以在保证实时性的同时,实现对复杂指令的深度理解与对动态环境的鲁棒适应

Benefits of technology

1、本发明中通过构建反应式小模型规划器与深思式大模型推理器并行决策的协同框架,结合全局-局部动态融合与因果链式思维两种决策模式,实现了对复杂自然语言指令的深度理解与实时环境应变。同时,基于不确定度的自适应融合机制动态调整两类决策权重,在保持系统实时性的同时提升了导航决策的鲁棒性、泛化性与可解释性,增强了移动机器人在动态复杂环境中的自主导航能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813831A_ABST
    Figure CN122813831A_ABST
Patent Text Reader

Abstract

The application relates to a visual-language navigation method and system based on a size model collaborative decision. In the method, a reactive small model planner is built, global-local dynamic fusion decisions are carried out by fusing panoramic images and natural language instructions, and a first decision probability is output; meanwhile, a visual language prompter and a deep thinking large model reasoner are built, a road point visual prompt is generated by combining panoramic images, road points and a navigation topology map, the visual prompt and the language instruction are integrated to build prompt content, a large model is input to complete causal chain thinking decision, and a second decision probability is output. Finally, the two decision probabilities are fused through an uncertainty adaptive collaborative strategy of commonality prediction to obtain a fusion decision probability, a target decision road point is determined according to the fusion decision probability, and navigation is performed, environment information is updated in real time throughout the whole process, and closed-loop control is formed until the navigation task is completed. Compared with the prior art, the application improves the robustness, generalization and explainability of navigation decisions while maintaining the real-time performance of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual-language navigation, and in particular to a visual-language navigation method and system based on size model collaborative decision-making. Background Technology

[0002] Vision-language navigation tasks for mobile robots aim to enable robots to understand and execute navigation commands given by humans in natural language, based on first-person visual observation information, in unknown or semi-unknown environments. This allows them to achieve tasks such as reaching the target location, path planning, and obstacle avoidance. These tasks typically require robots to possess cross-modal semantic understanding, environmental perception, and continuous decision-making and motion control capabilities, making it an important research direction in the fields of embodied intelligence and intelligent robotics.

[0003] In recent years, with the development of large-scale pre-trained models, large-scale vision-language models have demonstrated strong capabilities in general semantic understanding, commonsense reasoning, and cross-modal alignment, providing new technical pathways for complex instruction understanding and high-level decision-making. However, when existing large-scale vision-language models are directly applied in embodied navigation scenarios, they generally face problems such as high inference latency, insufficient response to real-time environmental changes, and limited coupling with underlying motion control, making it difficult to meet the real-time and stability requirements of mobile robots. Meanwhile, task-customized small-model architectures, trained end-to-end on vision-language navigation datasets, can achieve high execution efficiency and response speed, making them suitable for real-time planning and control. However, these methods typically rely on limited scenarios and data distributions, limiting their generalization ability in complex and variable real-world environments. Furthermore, their decision-making processes are mostly implicit mappings, lacking intermediate inference expressions and uncertainty modeling, making it difficult to provide interpretable decision-making basis.

[0004] For example, invention application CN118758310A discloses a multi-level cross-media fusion visual-language navigation method. This method predicts candidate waypoints using an observation-driven waypoint predictor, constructs a collaborative navigation planning module including a global navigation planning module and a local navigation planning module, and forms, maintains, and updates a global topology map to ultimately control the agent to reach the target location. However, the collaborative framework of this scheme is still limited to the global-local module partitioning under a single model system, failing to resolve the core contradiction between the inference latency of large models and the insufficient generalization ability of small models. Furthermore, its decision-making process lacks uncertainty modeling and adaptive adjustment mechanisms, making it unable to dynamically optimize the weights of different decision modes according to dynamic environmental changes. This makes it difficult to achieve a deep understanding of complex instructions and robust adaptation to dynamic environments while ensuring real-time performance.

[0005] In summary, current vision-language navigation methods suffer from difficulties in balancing real-time response efficiency with generalization and robustness in complex environments, and the decision-making process lacks interpretability and dynamic adaptive adjustment capabilities. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a vision-language navigation method and system based on uncertainty adaptive fusion.

[0007] The objective of this invention can be achieved through the following technical solutions: According to one aspect of the present invention, a visual-language navigation method based on size model collaborative decision-making is provided, the method steps including: It acquires preset natural language commands and collects current environmental observation information in real time; the environmental observation information includes panoramic images from a first-person perspective and radar depth information. Waypoints are predicted based on radar depth information, and a navigation topology map is constructed and dynamically updated in real time based on the waypoints. A reactive small model planner is constructed. Panoramic images and natural language instructions are input into the reactive small model planner, which performs global-local dynamic fusion decision-making and outputs the first decision probability. Construct a visual language prompter and a deep thinking large model inferencer; construct a panoramic environment observation waypoint visual prompt based on panoramic images, waypoints and navigation topology map; combine the visual prompt with natural language instructions, use the visual language prompter to construct prompt content, input the prompt content into the deep thinking large model inferencer, perform causal chain thinking decision-making, and output the second decision probability; Based on the uncertainty adaptive collaborative decision-making of common prediction, the first decision probability and the second decision probability are fused to obtain and output the fused decision probability; The target decision waypoint is selected from the waypoints based on the fusion decision probability. Navigation is completed based on the target decision waypoint. During the navigation process, environmental observation information is continuously updated to form a closed-loop control until the navigation ends.

[0008] As a preferred technical solution, the specific process of predicting waypoints based on radar depth information includes: The depth point cloud in the radar depth information is converted into a two-dimensional planar grid map, and K cluster centers are generated as candidate waypoints using the KMeans clustering method. The candidate waypoints are then filtered out, and points whose distance from obstacles is less than a first preset minimum distance or whose distance between two waypoints is less than a second preset minimum distance are removed, thus obtaining the waypoints.

[0009] As a preferred technical solution, the specific process of constructing and dynamically updating the navigation topology map in real time based on waypoints is as follows: The detected waypoints are mapped from the local coordinate system to the global coordinate system using the TF tree in the robot system ROS. Based on the spatial location and traversable connectivity of the waypoints in the global coordinate system, waypoint nodes and connecting edges between nodes are established to construct the initial navigation topology. As the robot moves, based on the coordinate transformation results refreshed in real time by the TF tree, the system adds detected valid waypoints to the navigation topology map, removes invalid waypoints outside the field of view, and synchronously updates the node connection relationships, thereby realizing the dynamic updating of the navigation topology map.

[0010] As a preferred technical solution, the specific process for global-local dynamic fusion decision-making is as follows: The panoramic image is segmented and corrected to obtain the corrected image; Visual information is extracted and encoded from the corrected image to obtain visual features; Encode the input natural language instructions to obtain language features; Collect historical memory image features, and based on visual features, linguistic features, and historical memory image features, perform cross-modal attention fusion from global and local perspectives by utilizing global and local branches respectively; In the global branch, the historical memory image features and the language features are aligned, and a global decision is output; in the local branch, the alignment of the visual features and the language features is emphasized, and a local decision is output. The global and local decisions are aggregated, and the aggregation result is processed using an activation function to obtain the first decision probability.

[0011] As a preferred technical solution, the specific process of segmenting and correcting panoramic images includes: The panoramic image is segmented in a counterclockwise direction at preset angle intervals, dividing the panoramic image into multiple surround view sub-images, and the panoramic distortion in the surround view sub-images is corrected by image geometric transformation. The specific process of extracting and encoding visual information from the corrected image to obtain encoded visual features includes: extracting and encoding the corresponding visual information from multiple surround view sub-images using a pre-trained CLIP-ViT model; encoding the orientation angle information of each sub-image using one-hot encoding; and fusing the visual information and the encoded orientation angles using a multilayer perceptron fusion function to obtain encoded visual features. The multilayer perceptron fusion function includes a linear mapping layer, a ReLU nonlinear mapping layer, and a LayerNorm operation.

[0012] As a preferred technical solution, when encoding the input natural language instructions, word segmentation is first performed to obtain a sequence containing L word segmentation units, where l is the position index of the word segmentation unit in the sequence, with a value range from 1 to L; for the word segmentation unit at the l-th position, the word vector representation corresponding to the word segmentation unit is first obtained through the word vector representation layer, and then the corresponding position vector representation is assigned to the position. The word vector representation and the position vector representation are added element by element to obtain the input representation vector corresponding to the l-th word segmentation unit; Subsequently, the sequence of input representation vectors corresponding to all word segmentation units is input into the RoberTA encoder. The RoberTA encoder performs contextual semantic encoding on the input sequence, and the output dimension is the language instruction contextual semantic feature matrix, which is the language feature after encoding.

[0013] As a preferred technical solution, the specific steps for constructing panoramic environment observation waypoint visual cues include: based on the navigation topology map and waypoints, selecting waypoints whose distance from the current location to the current position is less than a preset distance, mapping the selected waypoints to panoramic image coordinates and drawing them on the panoramic image to form panoramic environment observation waypoint visual cues.

[0014] As a preferred technical solution, before constructing the prompt content using the visual language prompter, historical navigation sequence information is also collected. In addition to visual prompts and natural language instructions, the input of the visual language prompter also includes preset task description information, historical navigation sequence information, and the output of the reactive small model planner.

[0015] As a preferred technical solution, the specific process of integrating the first decision probability and the second decision probability includes: Identify the decision uncertainty states corresponding to the first and second decision probabilities, adjust the fusion ratio of these two decision probabilities in real time based on the identified uncertainty states, and perform collaborative fusion processing on the first and second decision probabilities according to the adjusted fusion ratio to obtain the fused decision probability.

[0016] According to another aspect of the present invention, a visual-language navigation system based on size model collaborative decision-making is provided, the system comprising: The environmental perception and command input module is used to collect two types of environmental observation information in real time: panoramic images and radar depth information, and to input preset natural language commands. The waypoint mapping module is used to predict waypoints based on radar depth information and complete the real-time construction and dynamic updating of the navigation topology map; The small model decision module is used to perform global-local dynamic fusion decision-making and output the first decision probability. The large model reasoning prompt module is used to generate visual prompts for waypoints in the panoramic environment and construct complete prompt content. Based on the prompt content, it completes causal chain thinking decision-making and outputs the probability of the second decision. The adaptive collaborative fusion module is used to dynamically adjust the weights based on the uncertainty of the decision and complete the probabilistic fusion of the two decision paths. The navigation closed-loop execution module is used to determine the target waypoint based on the fusion decision probability and execute the navigation task, while updating environmental information in real time to form closed-loop control.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention constructs a collaborative framework for parallel decision-making between a reactive small-model planner and a thoughtful large-model inferencer, combining global-local dynamic fusion and causal chain thinking decision-making modes to achieve deep understanding of complex natural language instructions and real-time environmental adaptation. Simultaneously, an uncertainty-based adaptive fusion mechanism dynamically adjusts the weights of the two decision types, improving the robustness, generalization, and interpretability of navigation decisions while maintaining system real-time performance, thus enhancing the autonomous navigation capability of mobile robots in dynamic and complex environments.

[0018] 2. In this invention, radar depth point clouds are converted into grid maps and KMeans clustering is used to generate candidate waypoints. After distance constraint filtering, a set of waypoints with balanced distribution and high accessibility is obtained. At the same time, ROS TF trees are used to realize the unified mapping and real-time updating of waypoints in multiple coordinate systems, and a stable and consistent navigation topology map is constructed, which reduces planning failures caused by waypoint or map errors and improves the long-term reliability of the system.

[0019] 3. In this invention, distortion is eliminated and spatial information is preserved by segmenting and correcting panoramic images and performing orientation encoding, and visual features are extracted by combining a pre-trained model; at the same time, word vectors and position vectors are fused, and language instructions are deeply semantically encoded through the RoberTA encoder; thus improving the accuracy and richness of visual and language information representation respectively.

[0020] 4. In this invention, cross-modal attention fusion is performed on the encoded panoramic image and language commands using global and local branches respectively. This achieves multi-level alignment of historical memory, current local vision, and language commands, aggregating global and local information. This allows the planner to simultaneously understand the long-term task context and the immediate environment, improving the environmental adaptability of decision-making. By mapping and drawing neighboring waypoints onto the corrected panoramic image, an intuitive visual cue is constructed that integrates environmental observations and reachable waypoints. Simultaneously, multi-source information such as task descriptions, historical navigation sequences, and small model suggestions are integrated into the cue, collectively transforming abstract topological and historical information into concrete and structured cuees that are easy for the large model to understand, thereby improving the efficiency of causal reasoning and the coherence of decision-making.

[0021] 5. This invention constructs a complete navigation system that includes modules such as environmental perception, waypoint mapping, dual-model decision-making, adaptive fusion, and closed-loop execution. This system achieves closed-loop control from environmental perception to navigation actions, combining real-time response and deep reasoning capabilities. It improves the stability, generalization, and interpretability of visual-language navigation tasks in practical applications. Attached Figure Description

[0022] Figure 1 This is a schematic diagram illustrating the steps of a visual-language navigation method based on size model collaborative decision-making in this invention; Figure 2 This is a flowchart illustrating the specific implementation process of the visual-language navigation method in this embodiment; Figure 3 This is a schematic block diagram of the reactive small model decision maker in this invention; Figure 4 This is a schematic block diagram of the deep thinking large model inference engine in this invention; Figure 5 This is a flowchart illustrating the cloud-edge collaborative operation process of the physical robot in this embodiment. Figure 6a This is a schematic diagram of the mobile robot platform in the embodiment; Figure 6b This is a schematic diagram illustrating the actual deployment and application in the embodiment. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] Example In this embodiment, a visual-language navigation method and system based on size model collaborative decision-making is adopted, and the method flow is as follows: Figure 1 As shown, the specific steps include: It acquires preset natural language commands and collects current environmental observation information in real time; the environmental observation information includes panoramic images from a first-person perspective and radar depth information. Waypoints are predicted based on radar depth information, and a navigation topology map is constructed and dynamically updated in real time based on the waypoints. A reactive small model planner is constructed. Panoramic images and natural language commands are input into the reactive small model planner, which performs global-local dynamic fusion decision-making and outputs the first decision probability. Construct a visual language prompter and a deep thinking large model inferencer; construct a panoramic environment observation waypoint visual prompt based on panoramic images, waypoints and navigation topology map; combine the visual prompt with natural language instructions, use the visual language prompter to construct prompt content, input the prompt content into the deep thinking large model inferencer, perform causal chain thinking decision-making, and output the second decision probability; Based on the uncertainty adaptive collaborative decision-making of common prediction, the first decision probability and the second decision probability are fused to obtain and output the fused decision probability; The target decision waypoint is selected from the waypoints based on the fusion decision probability. Navigation is completed based on the target decision waypoint. During the navigation process, environmental observation information is continuously updated to form a closed-loop control until the navigation ends.

[0025] The robot receives first-person panoramic environmental imagery, radar depth information, and corresponding natural language navigation commands as input. The input includes first-person panoramic imagery from the robot's intelligent agent, containing RGB information. and radar depth information and the given natural language instructions ,in Indicates the first in the instruction Each token This is the instruction length.

[0026] A clustering-based panoramic candidate waypoint predictor model is constructed to predict potential navigation waypoints in the environment and to build and dynamically update the navigation waypoint connection topology map in real time. First, the depth point cloud is converted into a 2D planar grid map, and K-Means clustering is used to generate K cluster centers as candidate waypoints. To ensure the quality of the candidate waypoints, the clustering results are filtered, removing points that are too close to obstacles or too close to each other, thus ensuring the feasibility and balanced distribution of the candidate waypoints. These waypoints are then mapped to the local and global coordinate systems using a TF tree in the robot system's ROS, constructing and updating the global navigation topology map.

[0027] A reactive small model planner is constructed. By building a global-local dual-branch decision architecture, perception and action decision-making are collaboratively modeled to improve the generalization ability and response efficiency of the small model expert in complex and unknown scenarios. To improve the fine-grained encoding and directional expression capability of panoramic environment observation, it is first arranged in a counterclockwise direction at preset angle intervals. The panoramic image is segmented into segments. The system extracts panoramic sub-images and corrects panoramic distortion through image geometric transformation. Then, it extracts corresponding visual features from multiple panoramic sub-images using a pre-trained CLIP-ViT model, combined with one-hot encoding. The azimuth angle information of each sub-map is encoded and fused, and the calculation formula is as follows: in, This represents the image geometric transformation operator used for distortion correction. This represents a panorama cropping operator based on azimuth angle. Indicates the first The center azimuth angle corresponding to each sub-figure This represents a pre-trained visual encoding model. Represents the dimensions of visual features. This indicates a feature concatenation operation. This represents the fusion function of a multilayer perceptron (including linear mapping layers, ReLU nonlinear mapping layers, and LayerNorm operations). To integrate feature dimensions.

[0028] The input natural language commands are encoded using the RoberTA model, and the calculation formula is as follows: Simultaneously, for visual and language modalities, cross-modal attention fusion is performed from both local and global perspectives. Local branch It emphasizes the semantic alignment of current fine-grained panoramic image features and language instructions, while the global branch... Emphasizing coarse-grained long-term historical memory image features Semantic alignment with language instructions. Finally, a dynamic fusion decision module is used. To aggregate the decision outputs of local and global branches, and ultimately utilize... The decision probability is obtained using a function, and the calculation formula is shown below: in, Represents the learnable linear mapping parameters. This represents the total number of candidate points globally. This indicates the number of candidate points in the current step after fusing and filtering local and global candidate points.

[0029] Construct a deep-thinking large-scale model inference engine, based on panoramic environmental observation waypoint visual cues, and combined with a structured visual-language cues system and a causal chain thinking decision-making mechanism, to achieve interpretable multi-step reasoning and global decision-making; First, combining the obtained candidate waypoints and topology navigation connectivity map, nearby candidate points are mapped to panoramic image coordinates and directly plotted on the global image, forming panoramic environment observation waypoint visual cues. A structured visual-language cues system is then constructed, with inputs including task descriptions, instructions, visual observations, historical sequences, and expert suggestions from a small model. Through causal chain cues, a general visual-language large model is induced to output detailed thought processes, waypoint selection information, and corresponding confidence levels. And transform it into a sparse distribution. .

[0030] An uncertainty adaptive fusion decision-making module based on common prediction is constructed. During the inference stage, a high-confidence prediction set is dynamically constructed to adaptively quantify the decision uncertainty of the large and small models and realize the dynamic fusion of the decision results of the two models. Specifically, an uncertainty adaptive fusion decision-making module based on common prediction is constructed, and a confidence threshold is obtained using statistical methods during the calibration phase. During the inference phase, a reliable prediction set is constructed in real time. This quantifies the uncertainty of the small model expert and adaptively fuses the decision outputs from the small model expert and the large model expert based on the following formula: in, This represents the length of the constructed set of reliable predictions. Indicates the maximum length reduction factor. To preset the maximum length, To calculate the fusion decision probability, and These represent the output probabilities of the reactive small-model decision maker and the output probabilities of the thoughtful large-model inference maker, respectively.

[0031] The underlying motion control is performed by combining the SLAM online point target navigation algorithm and continuously updating environmental observation information during the navigation process to form closed-loop control until the decision model outputs the stop action.

[0032] Specifically, for the target decision waypoints obtained through collaborative decision-making using large and small models, the Simultaneous Localization and Mapping (SLAM) algorithm is invoked to perform target point navigation, and environmental observation information is updated in real time during the navigation process. The above decision-making and execution processes are carried out in a closed loop until the collaborative decision-making model outputs a stop action command, thereby completing the visual-language navigation task.

[0033] The specific implementation process of the method is as follows: Figure 2 As shown, the details are as follows: S1. The robot acquires first-person panoramic environmental image observation, depth radar information, and corresponding natural language navigation commands as input; The input includes a first-person panoramic view of the robot agent, containing RGB information. and radar depth information and the given natural language instructions ,in Indicates the first in the instruction Each token This is the length of the instruction. An example instruction might be: "Leave the bedroom, turn left, enter the kitchen, and stay next to the refrigerator." S2. Construct a cluster-based panoramic candidate waypoint predictor model to predict potential navigation waypoints in the environment and build and dynamically update the navigation waypoint connection topology map in real time. First, the depth point cloud is converted into a two-dimensional planar raster image, and K-Means clustering is used to generate K cluster centers as candidate waypoints. Let the robot at time [time value missing]... The acquired deep point cloud is ,in Represents the three-dimensional coordinates of the point cloud. This represents the total number of point clouds. The 3D point cloud is projected onto a horizontal plane along a robot height threshold to obtain a 2D point set. According to raster resolution Constructing a 2D Occupied Grid Map The element 1 indicates that the grid cell is occupied by an obstacle, and 0 indicates a feasible region. KMeans clustering is applied to the feasible regions to obtain an initial set of candidate points. Simultaneously, based on obstacle constraints and uniformity constraints, some infeasible waypoints and densely packed waypoints near obstacles are eliminated, resulting in the final candidate waypoint set. For globally accumulated candidate waypoints Build a navigation topology map Where, if the reachability constraint is satisfied: If there are no obstacles obstructing the connection between two points, then a topological connection is established. As environmental observations are updated, the collection... and Continuously updated dynamically.

[0034] S3. Construct a reactive small model planner. By building a global-local dual-branch decision architecture, perception and action decision-making are collaboratively modeled to improve the generalization ability and response efficiency of small model experts in complex and unknown scenarios. To improve the fine-grained coding and directional expression capabilities of panoramic environmental observation, the coding is first performed counterclockwise at preset angular intervals. The panoramic image is segmented into segments. The system extracts panoramic sub-images and corrects panoramic distortion through image geometric transformation. Then, it extracts corresponding visual features from multiple panoramic sub-images using a pre-trained CLIP-ViT model, combined with one-hot encoding. The azimuth angle information of each sub-map is encoded and fused, and the calculation formula is as follows: in, This represents the image geometric transformation operator used for distortion correction. This represents a panorama cropping operator based on azimuth angle. Indicates the first The center azimuth angle corresponding to each sub-figure This represents a pre-trained visual encoding model. Represents the dimensions of visual features. This indicates a feature concatenation operation. This represents the fusion function of a multilayer perceptron (including linear mapping layers, ReLU nonlinear mapping layers, and LayerNorm operations). To integrate feature dimensions.

[0035] The input natural language commands are encoded using the RoberTA model, and the calculation formula is as follows: Simultaneously, for visual and language modalities, cross-modal attention fusion is performed from both local and global perspectives. Local branch It emphasizes the semantic alignment of current fine-grained panoramic image features and language instructions, while the global branch... Emphasizing coarse-grained long-term historical memory image features Semantic alignment with language instructions is achieved. The cross-modal attention module employs cross-attention methods such as LXMERT (Learning Cross-Modality Encoder Representations from Transformers) and METER (Multimodal End-to-end Transformer) for computation, comprising multiple multi-head self-attention layers, multiple multi-head cross-attention layers, and multiple perceptron layers. Finally, a dynamic fusion decision module is utilized. To aggregate the decision outputs of local and global branches, and ultimately utilize... The decision probability is obtained using a function, and the calculation formula is shown below: in, Represents the learnable linear mapping parameters. This represents the total number of candidate points globally. This represents the number of candidate points in the current step after fusing and filtering local and global candidate points. The calculation process described above is as follows: Figure 3 As shown.

[0036] S4. Construct a deep-thinking large-scale model inference engine, based on panoramic environmental observation waypoint visual cues, and combined with a structured visual-language cues system and a causal chain thinking decision-making mechanism, to achieve interpretable multi-step reasoning and global decision-making. First, combining the candidate waypoints and topology navigation connection map obtained in S2, nearby candidate points are mapped to panoramic image coordinates and directly drawn on the global image to form panoramic environment observation waypoint visual cues. A structured visual-language cues system is then constructed, with inputs including task descriptions, instructions, visual observations, historical sequences, and expert suggestions from a small model. Through causal chain cues, the general visual-language large model is induced to output detailed thought processes, waypoint selection information, and corresponding confidence levels. And transform it into a sparse distribution. The above process is as follows: Figure 4 As shown.

[0037] S5. Construct an uncertainty adaptive fusion decision module based on common prediction, dynamically construct a high-confidence prediction set during the inference stage, adaptively quantify the decision uncertainty of the large and small models, and realize the dynamic fusion of the decision results of the two models. Construct an uncertainty adaptive fusion decision-making module based on common predictions. Specifically, assume... The input observation space consists of visual observations, voice commands, and navigation history. This represents the candidate set of possible outputs. For the action label space (i.e., all possible candidate waypoints), and This is the user-defined allowable error rate, which is the probability that the actual action will not be included in the prediction set.

[0038] During the calibration phase, 50% of the training set is randomly sampled as the calibration set, and the inconsistency score of each calibration sample is calculated. This score measures the degree of deviation of the sample from the training distribution; a higher score indicates a lower confidence level of the model in that action. Specifically, the formula for calculating the inconsistency score is as follows: During the inference phase, a reliable prediction set is constructed in real time. All candidate actions that satisfy the inconsistency score are included in the credible prediction set. : Size of the credible prediction set This naturally reflects the uncertainty of the model: a larger set indicates that more actions have a high probability, thus resulting in higher uncertainty; while if the prediction set only contains a single action (i.e., ... This means the model is very confident in its current decision, with low uncertainty. The decision outputs from both the small and large model experts are adaptively fused based on the following formula: in, This represents the length of the constructed set of reliable predictions. Indicates the maximum length reduction factor. To preset the maximum length, To calculate the fusion decision probability, and These represent the output probabilities of the reactive small model decision maker obtained in S3 and the output probabilities of the thoughtful large model inference maker obtained in S4, respectively.

[0039] S6. Combine the SLAM online point target navigation algorithm to perform low-level motion control and continuously update environmental observation information during navigation to form closed-loop control until the decision model outputs a stop action; For the target waypoints obtained through collaborative decision-making using large and small models, an online simultaneous localization and mapping (SLAM) algorithm is invoked to perform target point navigation, updating environmental observation information in real time during the navigation process. This algorithm benefits from continuous optimization and modular design of both the front-end and back-end, offering advantages in stable navigation capabilities and safe obstacle avoidance. The above decision-making and execution process is carried out in a closed-loop manner until the collaborative decision-making model outputs a stop action command, thus completing the visual-language navigation task. The flowchart of cloud-edge collaborative computing is shown below. Figure 5 As shown, the mobile robot platform is as follows Figure 6a As shown, the corresponding real-world application deployment illustration is as follows: Figure 6b As shown.

[0040] In summary, this method constructs a candidate waypoint generation mechanism based on two-dimensional grid modeling and clustering of deep point clouds, and combines obstacle distance constraints and waypoint spacing constraints to filter candidate waypoints, forming a spatially balanced and reachable set of navigation waypoints. On this basis, through unified mapping and topology connection updates in local and global coordinate systems, the reliability of candidate waypoint generation and the stability of navigation topology are effectively improved, thereby reducing the probability of path planning failure or frequent replanning in complex environments.

[0041] Furthermore, this method introduces a collaborative decision-making mechanism for large and small models based on uncertainty adaptive quantization. By evaluating the credibility and dynamically fusing the decision results of reactive small models and thoughtful large models, it ensures real-time navigation while taking into account high-level semantic understanding and reasoning capabilities. This effectively improves the robustness, generalization and interpretability of the decision-making process in visual-language navigation tasks, and enhances the autonomous navigation capabilities of mobile robots in complex and variable environments.

[0042] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A visual-language navigation method based on size-model collaborative decision-making, characterized in that, The method steps include: Acquire preset natural language commands and collect current environmental observation information in real time; the environmental observation information includes panoramic images from a first-person perspective and radar depth information; Waypoints are predicted based on radar depth information, and a navigation topology map is constructed and dynamically updated in real time based on the waypoints. A reactive small model planner is constructed. Panoramic images and natural language commands are input into the reactive small model planner, which performs global-local dynamic fusion decision-making and outputs the first decision probability. Construct a visual language prompter and a deep thinking large model inferencer; construct a panoramic environment observation waypoint visual prompt based on panoramic images, waypoints and navigation topology map; combine the visual prompt with natural language instructions, use the visual language prompter to construct prompt content, input the prompt content into the deep thinking large model inferencer, perform causal chain thinking decision-making, and output the second decision probability; Based on the uncertainty adaptive collaborative decision-making of common prediction, the first decision probability and the second decision probability are fused to obtain and output the fused decision probability; The target decision waypoint is selected from the waypoints based on the fusion decision probability. Navigation is completed based on the target decision waypoint. During the navigation process, environmental observation information is continuously updated to form a closed-loop control until the navigation ends.

2. The visual-language navigation method based on size model collaborative decision-making according to claim 1, characterized in that, The specific process of predicting waypoints based on radar depth information includes: The depth point cloud in the radar depth information is converted into a two-dimensional planar grid map, and K cluster centers are generated as candidate waypoints using the KMeans clustering method. The candidate waypoints are then filtered out, and points whose distance from obstacles is less than a first preset minimum distance or whose distance between two waypoints is less than a second preset minimum distance are removed, thus obtaining the waypoints.

3. The visual-language navigation method based on size model collaborative decision-making according to claim 1, characterized in that, The specific process of constructing and dynamically updating the navigation topology map in real time based on waypoints is as follows: The detected waypoints are mapped from the local coordinate system to the global coordinate system using the TF tree in the robot system ROS. Based on the spatial location and traversable connectivity of waypoints in the global coordinate system, waypoint nodes and connecting edges between nodes are established to construct an initial navigation topology. As the robot moves, based on the coordinate transformation results refreshed in real time by the TF tree, the system adds detected valid waypoints to the navigation topology map, removes invalid waypoints outside the field of view, and synchronously updates the node connection relationships, thereby realizing the dynamic updating of the navigation topology map.

4. The visual-language navigation method based on size model collaborative decision-making according to claim 1, characterized in that, The specific process of making global-local dynamic fusion decisions is as follows: The panoramic image is segmented and corrected to obtain the corrected image; Visual information is extracted and encoded from the corrected image to obtain visual features; Encode the input natural language instructions to obtain language features; Collect historical memory image features, and based on visual features, linguistic features, and historical memory image features, perform cross-modal attention fusion from global and local perspectives by utilizing global and local branches respectively; In the global branch, the historical memory image features and the language features are aligned, and a global decision is output; in the local branch, the alignment of the visual features and the language features is emphasized, and a local decision is output. The global and local decisions are aggregated, and the aggregation result is processed using an activation function to obtain the first decision probability.

5. The visual-language navigation method based on size model collaborative decision-making according to claim 4, characterized in that, The specific process of segmenting and correcting the panoramic image includes: segmenting the panoramic image in a counterclockwise direction at preset angle intervals, dividing the panoramic image into multiple surround view sub-images, and using image geometric transformation to correct the panoramic distortion in the surround view sub-images. The specific process of extracting and encoding visual information from the corrected image to obtain encoded visual features includes: extracting and encoding corresponding visual information from the multiple surround view sub-images using a pre-trained CLIP-ViT model, encoding the orientation angle information of each sub-image using one-hot encoding, and fusing the visual information and the encoded orientation angles using a multilayer perceptron fusion function to obtain encoded visual features; the multilayer perceptron fusion function includes a linear mapping layer, a ReLU nonlinear mapping layer, and a LayerNorm operation.

6. The visual-language navigation method based on size model collaborative decision-making according to claim 4, characterized in that, When encoding the input natural language instructions, word segmentation is first performed to obtain a sequence containing L word segmentation units, where l is the position index of the word segmentation unit in the sequence, with a value range from 1 to L; for the word segmentation unit at the l-th position, the word vector representation corresponding to the word segmentation unit is first obtained through the word vector representation layer, and then the corresponding position vector representation is assigned to the position. The word vector representation and the position vector representation are added element by element to obtain the input representation vector corresponding to the l-th word segmentation unit; Subsequently, the sequence of input representation vectors corresponding to all word segmentation units is input into the RoberTA encoder. The RoberTA encoder performs contextual semantic encoding on the input sequence, and the output dimension is the language instruction contextual semantic feature matrix, which is the language feature after encoding.

7. The visual-language navigation method based on size model collaborative decision-making according to claim 1, characterized in that, The specific steps for constructing panoramic environment observation waypoint visual cues include: based on the navigation topology map and waypoints, selecting waypoints whose distance from the current location to the current position is less than a preset distance, mapping the selected waypoints to panoramic image coordinates and drawing them on the panoramic image to form panoramic environment observation waypoint visual cues.

8. The visual-language navigation method based on size model collaborative decision-making according to claim 1, characterized in that, Before constructing the prompt content using the visual language prompter, historical navigation sequence information is also collected. In addition to visual prompts and natural language instructions, the input of the visual language prompter also includes preset task description information, historical navigation sequence information, and the output of the reactive small model planner.

9. A visual-language navigation method based on size model collaborative decision-making according to claim 1, characterized in that, The specific process of fusing the first decision probability and the second decision probability includes: Identify the decision uncertainty states corresponding to the first and second decision probabilities, adjust the fusion ratio of these two decision probabilities in real time based on the identified uncertainty states, and perform collaborative fusion processing on the first and second decision probabilities according to the adjusted fusion ratio to obtain the fused decision probability.

10. A visual-language navigation system based on size-model collaborative decision-making, characterized in that, The system operates using a visual-language navigation method based on size model collaborative decision-making as described in any one of claims 1-9, the system comprising: The environmental perception and command input module is used to collect two types of environmental observation information in real time: panoramic images and radar depth information, and to input preset natural language commands. The waypoint mapping module is used to predict waypoints based on radar depth information and complete the real-time construction and dynamic updating of the navigation topology map; The small model decision module is used to perform global-local dynamic fusion decision-making and output the first decision probability. The large model reasoning prompt module is used to generate visual prompts for waypoints in the panoramic environment and construct complete prompt content. Based on the prompt content, it completes causal chain thinking decision-making and outputs the probability of the second decision. The adaptive collaborative fusion module is used to dynamically adjust the weights based on the uncertainty of the decision and complete the probabilistic fusion of the two decision paths. The navigation closed-loop execution module is used to determine the target waypoint based on the fusion decision probability and execute the navigation task, while updating environmental information in real time to form closed-loop control.

Citation Information

Patent Citations

  • Multi-level cross-media fusion visual language navigation method

    CN118758310A