Method and device for improving reasoning ability of navigation robot based on thinking chain

Through a multi-step planning mechanism based on thinking chain, combined with navigation diagram structure and environmental observation information, multi-modal prompt words are generated and path selection is used for multi-modal big model, the problem of insufficient information integration and generalization capabilities in visual language navigation tasks is solved, and the reasoning and planning capabilities of navigation robots are improved.

CN120373472AActive Publication Date: 2025-07-25ZHEJIANG LAB

Patent Information

Application Number
CN202510858805.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The existing multimodal large models fail to fully integrate visual and language information in visual language navigation tasks, resulting in low efficiency in environmental perception and dynamic programming, and insufficient generalization ability in the case of scarcity of data and environmental differences.

Method used

Through a multi-step planning mechanism based on thinking chain, combining navigation diagram structure, historical paths and environmental observation information, multi-modal prompt words are generated and path selection and inspection are used for multi-modal large models to reduce information loss and improve the model's reasoning ability.

Benefits of technology

It significantly improves the inference ability and path planning accuracy of multimodal large models in visual language navigation tasks, and enhances the robot's performance ability in unseen environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373472A_ABST
    Figure CN120373472A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for improving the reasoning ability of a navigation robot based on a thinking chain. The method comprises the following steps: reading a graph structure of a navigation scene, extracting current observation information, forming a multi-modal cue word in combination with a historical path, prompting a multi-modal large model to generate a plurality of reference paths, then selecting an optimal path for the reference paths by utilizing a path selection large model, and further, aiming at different conditions of model selection options and output, selecting the optimal path by utilizing the path selection large model. And checking the options of the optimal path by using the option checking large model. The method specifically comprises the following steps: loading navigation chart structure data, and combining candidate viewpoints to generate action description; a plurality of paths are generated through a beam search method, then a path selection large model is prompted to evaluate the paths, an optimal decision is output, and for the condition that model selection options are different from output, options of the optimal path are corrected by using an option check large model. According to the method, the navigation information loss is remarkably reduced, and the multi-modal reasoning capability, the path planning accuracy and the generalization performance of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence applications and multimodal information fusion, and particularly relates to a method and device for improving the reasoning ability of a navigation robot based on a chain of thought. Background Art

[0002] Currently, artificial intelligence technology is developing rapidly, and the design of intelligent agents (Agents) is moving towards the general direction of multi-tasking and multimodality. In this trend, Embodied Intelligence, as a new intelligent paradigm that integrates perception, thinking, and action, has become a research hotspot. It particularly focuses on the interaction ability between the Agent and the environment, and completes complex tasks through dynamic perception and decision-making, such as application scenarios like smart homes, service robots, and autonomous driving. The core idea of this technology is to endow the Agent with human-like perception and action capabilities, enabling it to autonomously complete tasks in complex and open environments, overcoming the limitations of traditional artificial intelligence in dynamic environments.

[0003] At the same time, the rise of multimodal large models provides strong technical support for Embodied Intelligence. By integrating perceptual modalities such as vision, language, and sound, these models demonstrate remarkable cross-modal reasoning and task execution capabilities. For example, Vision-and-Language Navigation (VLN) is one of the important research directions of Embodied Intelligence, which requires the Agent to perform path planning and navigation in a visual environment under the guidance of language instructions, greatly testing the cross-modal reasoning ability and environmental adaptability of the model.

[0004] Before the advent of the large model era, the research on VLN mainly focused on aligning information from different modalities through representation learning and optimizing action selection through policy learning. These studies utilized pre-trained visual and language models (such as ResNet and BERT), combined with technologies such as cross-modal alignment and imitation learning, to address the performance gap between some known and unknown environments.

[0005] With the introduction of large language models (LLMs) and multimodal large models (MLLMs), the research on VLN has entered a new stage. These models have significantly improved the understanding ability and generalization ability of complex environments through multimodal pre-training. For example, a large model-based system can translate visual input into a text description and combine instructions for action prediction, greatly improving the accuracy of task planning and execution.

[0006] To further enhance the reasoning ability, researchers have proposed a CoT-based method to gradually complete complex tasks through explicit reasoning chains. As this method expands from unimodal to multimodal (such as the combination of vision and language), it enables the logical derivation and efficient execution of complex tasks. For example, models like EmbodiedGPT have introduced an embodied CoT reasoning mode, endowing the Agent with dynamic planning and real-time decision-making capabilities, opening up a new direction for multimodal embodied intelligence.

[0007] Despite significant technological progress, current research still faces the following key issues: 1. Deep integration of multimodal information: How to efficiently integrate information from modalities such as vision and language to achieve real-time environmental perception and dynamic planning. 2. Generalization ability: In the case of scarce data and environmental differences, how to enhance the performance of the Agent in unseen environments.

[0008] Therefore, the existing chain-of-thought methods have not fully unleashed the potential of multimodal large models in terms of the final effect, and there is still room for improvement in the reasoning effect. Summary of the Invention

[0009] The object of the present invention is to provide a method and device for enhancing the reasoning ability of a navigation robot based on the chain of thought in view of the deficiencies of the prior art.

[0010] The object of the present invention is achieved by the following technical solutions: A method for enhancing the reasoning ability of a navigation robot based on the chain of thought, comprising the following steps: (1) Load the navigation map of the current scene and obtain the graph structure; obtain the observation information of the current scene and generate a list of action prompts; and generate a trajectory text and a graph structure text according to the graph structure; (2) Concatenate the instruction, historical path, the plan planned in the previous step, the list of action prompts, the trajectory text, and the graph structure text into a first prompt word; (3) Subsequently, input the first prompt word and the current environmental view into a multimodal large model, and then output multiple thoughts and the corresponding next plan for each thought; (4) Concatenate the obtained multiple thoughts and the corresponding next plan for each thought and the first prompt word into a second prompt word and input it into an optimal path selection large model, and then output the optimal thought and the corresponding next plan for the optimal thought; (5) Concatenate the obtained optimal thought and the corresponding next plan for the optimal thought and the obtained list of action prompts into a third prompt word and input it into an option check large model, and then output the action prompt corresponding to the optimal thought and the motion description options; the robot moves to the next location according to the motion description options; (6) If the target location is not reached, repeat steps (1)-(5) until the robot reaches the target location.

[0011] Further, loading the navigation map of the current scene and obtaining the graph structure specifically includes: loading the navigation map of the current scene, acquiring the connectivity data in JSON format for each scene and reading it using the Dijkstra algorithm, filtering out all passable viewpoints, calculating the shortest path and the corresponding distance between each pair of viewpoints and storing them in the graph structure, taking each viewpoint as a node in the graph structure, taking the connection relationship between two viewpoints as an edge in the graph structure, and taking the distance between two viewpoints as the weight of the corresponding edge.

[0012] Further, obtaining the observation information of the current scene specifically includes: obtaining the instruction ID, scan ID, and the position information of the current viewpoint of the current scene; then obtaining all candidate viewpoints of the current viewpoint through the graph structure, continuously traversing each candidate viewpoint corresponding to the current viewpoint, generating the information of each candidate viewpoint and storing it in the candidate viewpoint list, where the information of the candidate viewpoint includes the feature information and position information of the candidate viewpoint; taking the observation information including the instruction ID, scan ID, the position information of the current viewpoint, and the candidate viewpoint list as the observation information of the current scene.

[0013] Further, generating an action prompt list according to the observation information specifically includes: according to the candidate viewpoint list in the observation information, comparing the horizontal viewpoint and rotation angle of the current viewpoint with the horizontal viewpoint and elevation angle of each candidate viewpoint to obtain the action prompt between the current viewpoint and each candidate viewpoint, and generating an action prompt list by associating each action prompt with a letter label; the action prompt includes turning left, going straight, turning right, going upstairs, and going downstairs.

[0014] Further, generating a trajectory text and a graph structure text according to the graph structure specifically includes: first constructing a trajectory text according to the shortest path between each pair of viewpoints in the graph structure; then splicing a graph structure text through the connection relationship between each viewpoint and other viewpoints in the graph structure.

[0015] Further, the multimodal large model is Qwen2-VL-7B; the optimal path selection large model is Qwen2.5-0.5B; the option check large model is Qwen2.5-0.5B.

[0016] The present invention further includes a device for improving the reasoning ability of a navigation robot based on a multimodal thinking chain and multi-step planning, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, it is used for the method of improving the reasoning ability of a navigation robot based on a multimodal thinking chain and multi-step planning as described above.

[0017] The present invention also includes a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the method for improving the reasoning ability of a navigation robot based on a multi-modal thought chain described above is implemented.

[0018] The beneficial effects of the present invention are as follows: 1) By combining the multi-modal thought chain and the multi-step planning mechanism, the present invention significantly improves the reasoning ability of the multi-modal large model (MLLM) in the visual language navigation (VLN) task; 2) By introducing the multi-modal thought chain, during the input process, by directly processing the environmental image rather than relying only on the language description, the information loss in the multi-modal conversion is reduced, enabling the model to more accurately understand the environment where the Agent is located, and improving the reliability and robustness of task execution; 3) With the multi-step planning mechanism as the core, the present invention regards the diverse thinking of the model as an opportunity to improve the reasoning ability. By reflecting on the self-generated reasoning paths, the model can significantly consider the reasons behind each path. Description of the Drawings

[0019] Figure 1 It is a flowchart of a method for improving the reasoning ability of a navigation robot based on the thought chain; Figure 2 It is a schematic diagram of constructing the first prompt word in Embodiment 1; Figure 3 It is a schematic diagram of the multi-modal large model, the optimal path selection large model, and the option check large model in Embodiment 1; Figure 4 It is a structural diagram of a device for improving the reasoning ability of a navigation robot based on the thought chain. Detailed Embodiments

[0020] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts are within the protection scope of the present invention.

[0021] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the" and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0022] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0023] Based on the multi-modal thought chain and multi-step planning mechanism, the present invention constructs a multi-step reasoning process by using the graph structure, historical path and environmental observation information in the navigation scenario, and at the same time integrates multi-modal inputs such as vision and language to achieve in-depth understanding and dynamic planning of the robot navigation task. In the method of the present invention, first, the scene data is preprocessed to obtain the navigation map and candidate view information, and multi-modal analysis is carried out in combination with environmental images, instruction texts, etc., and then action prompts and path decisions are continuously generated during the thought chain reasoning process; through the multi-step planning mechanism, multiple thoughts can be output simultaneously and the optimal path plan can be selected, so as to effectively filter and distinguish redundant or uncertain information in the multi-modal data. Since multi-modal features such as environmental images and language instructions are fully combined in each step, the present invention can explore the complementarity and relevance between modalities on the premise of reducing information loss, greatly improving the reasoning accuracy and generalization ability of the navigation robot, and providing a more flexible and efficient solution for the application of multi-modal large models in the embodied intelligence scenario.

[0024] Embodiment 1: As Figure 1 、 Figure 2 and Figure 3 shown, the present invention provides a method for improving the reasoning ability of a navigation robot based on a thought chain, including the following steps: (1) Load the navigation map of the current scene and obtain the graph structure; obtain the observation information of the current scene and generate a list of action prompts; and generate trajectory text and graph structure text according to the graph structure.

[0025] The step of loading the navigation map of the current scene and obtaining the graph structure is specifically as follows: load the navigation map of the current scene, obtain the connectivity data of each scene in JSON format and read it using the Dijkstra algorithm, filter out all passable viewpoints, calculate the shortest path and the corresponding distance between each pair of viewpoints and store them in the graph structure, take each viewpoint as a node in the graph structure, take the connection relationship between two viewpoints as an edge in the graph structure, and take the distance between two viewpoints as the weight of the corresponding edge. In this way, the shortest path planning information between the viewpoint nodes in each scene is obtained, providing structured data support for subsequent multi-modal navigation and decision-making.

[0026] The acquisition of the observation information of the current scene is specifically as follows: acquire the instruction ID (instr_id), scan ID (scan), and position information (position, heading, elevation) of the current perspective of the current scene; then obtain all candidate perspectives of the current perspective through the graph structure, continuously traverse each candidate perspective corresponding to the current perspective, generate the information of each candidate perspective and store it in the candidate perspective list (candidate), and the information of the candidate perspective includes the feature information and position information of the candidate perspective; take the observation information including the instruction ID, scan ID, position information of the current perspective, and the candidate perspective list as the observation information of the current scene.

[0027] The generation of the action prompt list according to the observation information is specifically as follows: according to the candidate perspective list in the observation information, compare the horizontal perspective and rotation perspective of the current perspective with the horizontal perspective and elevation angle of each candidate perspective to obtain the action prompt between the current perspective and each candidate perspective, and generate an action prompt list with each action prompt and a letter label; the action prompt includes turning left, going straight, turning right, going upstairs, and going downstairs.

[0028] For example, when the difference between the horizontal perspective of the candidate perspective and the horizontal perspective of the current perspective is positive, it can be determined that the action prompt of the candidate perspective is turning right; when the difference between the horizontal perspective of the candidate perspective and the horizontal perspective of the current perspective is negative, it is determined that the action prompt of the candidate perspective is turning left; when the difference between the horizontal perspective of the candidate perspective and the horizontal perspective of the current perspective is 0, it is determined that the action prompt of the candidate perspective is going straight; when the rotation perspective of the candidate perspective is greater than 0, it is determined that the action prompt of the candidate perspective is going upstairs; when the rotation perspective of the candidate perspective is less than 0, it is determined that the action prompt of the candidate perspective is going downstairs. Based on the direction descriptions obtained from the above judgments, when generating the action prompt words, only need to splice these direction descriptions (such as "turning left" or "going upstairs") with the relevant information of the target node to clearly and clearly express the operations required to move from the current perspective to the next target perspective.

[0029] The generation of the trajectory text and graph structure text according to the graph structure is specifically as follows: first construct the trajectory text according to the shortest path of each pair of perspectives in the graph structure; then splice the graph structure text through the connection relationship between each perspective and other perspectives in the graph structure.

[0030] (2) Splice the instruction, historical path, plan planned in the previous step, action prompt list, trajectory text, and graph structure text into the first prompt word, specifically as follows: First, generate a list of action options with alphabetical labels (e.g., "A. turn left to Place 2") for each candidate perspective, and insert the "stop" option at the front if applicable. Then, obtain the trajectory text and graph structure text for each example in the current batch, which describe the nodes that have been traversed in the navigation task, the connections between the nodes, and the unvisited ghost nodes. Next, depending on whether it is the first step (t == 0), append an empty or existing history to the first prompt. At the same time, the instruction, historical path, previously planned plan, and action prompt list are also integrated into the first prompt, along with a task description that includes the background setting, map description, execution requirements, and output format, so that the returned prompt contains a complete and coherent navigation context, providing the necessary environmental information and decision-making clues for subsequent multi-modal chain-of-thought reasoning.

[0031] (3) Subsequently, input the first prompt and the current environmental view into the multi-modal large model, and then output multiple thoughts and the corresponding next plan for each thought.

[0032] (4) Concatenate the obtained multiple thoughts, the corresponding next plan for each thought, and the first prompt to form a second prompt and input it into the optimal path selection large model, and then output the optimal thought and the corresponding next plan for the optimal thought; (5) Concatenate the obtained optimal thought, the corresponding next plan for the optimal thought, and the obtained action prompt list to form a third prompt and input it into the option checking large model, and then output the action prompt corresponding to the optimal thought and the motion description options; The robot moves to the next location according to the motion description options.

[0033] The multi-modal large model is Qwen2-VL-7B; the optimal path selection large model is Qwen2.5-0.5B; the option checking large model is Qwen2.5-0.5B.

[0034] (6) If the target location is not reached, repeat steps (1)-(5) until the robot reaches the target location.

[0035] Embodiment 2: This embodiment relates to a device for improving the reasoning ability of a navigation robot based on a chain of thought, including a memory and one or more processors. The memory stores executable code. When the one or more processors execute the executable code, it is used for the method of improving the reasoning ability of a navigation robot based on a chain of thought in Embodiment 1 above; The device embodiment can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer.

[0036] Such as Figure 4, at the hardware level, the knowledge distillation device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 shown method. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0037] Improvements to a technology can be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0038] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0039] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0040] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such a process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity, or device comprising the said element.

[0041] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0042] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0043] Embodiment 3: The embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for improving the reasoning ability of a navigation robot based on a chain of thought in Embodiment 1 above.

[0044] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for enhancing the reasoning ability of a navigation robot based on a chain of thought, characterized in that, It includes the following steps: (1) Load the navigation map of the current scene and obtain the graph structure; obtain the observation information of the current scene and generate a list of action prompts; and generate trajectory text and graph structure text according to the graph structure; (2) Concatenate the instruction, historical path, plan planned in the previous step, list of action prompts, trajectory text, and graph structure text into the first prompt; (3) Subsequently, input the first prompt and the current environmental view into the multi-modal large model, and then output multiple thoughts and the corresponding next plan for each thought; (4) Concatenate the obtained multiple thoughts, the corresponding next plan for each thought, and the first prompt into the second prompt and input it into the optimal path selection large model, and then output the optimal thought and the corresponding next plan for the optimal thought; (5) Concatenate the obtained optimal thought, the corresponding next plan for the optimal thought, and the obtained list of action prompts into the third prompt and input it into the option check large model, and then output the action prompt corresponding to the optimal thought and the motion description options; The robot moves to the next location according to the motion description options; (6) If the target location is not reached, repeat steps (1)-(5) until the robot reaches the target location.

2. A method for enhancing the reasoning ability of a navigation robot based on a chain of thought, characterized in that, The loading of the navigation map of the current scene and obtaining the graph structure are specifically as follows: Load the navigation map of the current scene, obtain the connectivity data in JSON format for each scene and read it using the Dijkstra algorithm, filter out all passable perspectives, calculate the shortest path and the corresponding distance between each pair of perspectives and store them in the graph structure, use each perspective as a node in the graph structure, use the connection relationship between two perspectives as an edge in the graph structure, and use the distance between two perspectives as the weight of the corresponding edge.

3. A method for enhancing the reasoning ability of a navigation robot based on a chain of thought according to claim 2, characterized in that The obtaining of the observation information of the current scene is specifically as follows: Obtain the instruction ID, scan ID, and position information of the current perspective in the current scene; subsequently, obtain all candidate perspectives of the current perspective through the graph structure, continuously traverse each candidate perspective corresponding to the current perspective, generate the information of each candidate perspective and store it in the candidate perspective list, and the information of the candidate perspective includes the feature information and position information of the candidate perspective; Use the observation information including the instruction ID, scan ID, position information of the current perspective, and the candidate perspective list as the observation information of the current scene.

4. A method for enhancing the reasoning ability of a navigation robot based on a chain of thought according to claim 3, characterized in that, The generation of the list of action prompts is specifically as follows: According to the list of candidate perspectives in the observation information, compare the horizontal perspective and rotation angle of the current perspective with the horizontal perspective and elevation angle of each candidate perspective to obtain the action prompt between the current perspective and each candidate perspective, and generate a list of action prompts by associating each action prompt with a letter label; The action prompts include turning left, going straight, turning right, going upstairs, and going downstairs.

5. A method for improving the reasoning ability of a navigation robot based on a chain of thought, as claimed in claim 4, wherein The generation of trajectory text and graph structure text according to the graph structure is specifically as follows: First, construct the trajectory text according to the shortest path between each pair of perspectives in the graph structure; subsequently, splice the graph structure text through the connection relationship between each perspective and other perspectives in the graph structure.

6. A method for improving the reasoning ability of a navigation robot based on a chain of thought, characterized in that, The multimodal large model is Qwen2-VL-7B; the optimal path selection large model is Qwen2.5-0.5B; the option checking large model is Qwen2.5-0.5B.

7. An apparatus for enhancing the reasoning ability of a navigation robot based on a chain of thought, characterized in that, It includes a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the method for improving the reasoning ability of a navigation robot based on the chain of thought according to any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, A program is stored thereon. When the program is executed by a processor, it implements the method for improving the reasoning ability of a navigation robot based on the chain of thought according to any one of claims 1-6.

Citation Information

Patent Citations

  • Visual language navigation method and system based on large language model

    CN118031964A

  • Visual semantic navigation frontier exploration method and device based on large language model

    CN118230272A

  • Robot positioning and navigation method and device based on large language model thinking chain

    CN119533474A

  • Robotic reasoning through planning with language models

    US20250018562A1

  • Neural Network Model Training Method, Electronic Device, Cloud, Cluster, and Medium

    US20250165782A1

Cited By

  • Low-altitude remote sensing image-oriented ground feature fine-grained attribute extraction method

    CN121661531A