Urban visual navigation method, device and equipment based on large language model and storage medium
By constructing a street view image fine-tuning dataset to fine-tune the large language model, and combining perception, reflection, and planning modules, the flexibility and interpretability issues of autonomous navigation in urban visual navigation are solved, and autonomous navigation in complex environments is achieved.
Patent Information
- Application Number
- CN202510493626.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies in urban visual navigation have problems such as insufficient flexibility, high training overhead, poor generalization performance and lack of interpretability, making it difficult to achieve autonomous navigation, especially when faced with unfamiliar environments or complex instructions.
By constructing a fine-tuning dataset of street view images based on the target city, the multimodal large language model is fine-tuned to generate a fine-tuned multimodal large language model. Combined with the perception, reflection and planning modules, autonomous navigation of the intelligent agent system is achieved.
It achieves autonomous navigation to the target location in complex urban environments, improves the success rate and explainability of navigation, has strong adaptability, and can handle long-distance dependencies and complex reasoning tasks.
Smart Images

Figure CN120593785A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual navigation technology, and in particular to a large language model-based urban visual navigation method, device, equipment and storage medium. Background Art
[0002] Visual-linguistic navigation aims to navigate to a target location using visual information (such as street view imagery) and verbal information (such as instructions and verbal descriptions of the target location). Urban visual navigation integrates cutting-edge technologies such as computer vision, deep learning, sensor fusion, and global navigation satellite systems to achieve precise perception and intelligent navigation of the urban environment. As a crucial component of modern urban transportation and intelligent systems, it is becoming a hot topic in research and application. With continuous technological advancements, urban visual navigation will not only provide safer and more efficient transportation solutions, but will also further promote the development of smart cities and improve people's daily lives.
[0003] Current research on vision-based navigation falls into several approaches: One approach uses visual sensors to construct maps and uses neural network training models to match language commands and execute navigation actions. However, this approach relies on pre-built map data and specified rules, lacking flexibility and making it difficult to adapt to unfamiliar environments or complex commands. Another approach uses deep reinforcement learning to train deep reinforcement learning models end-to-end for long-distance navigation in urban environments. However, this approach suffers from high training overhead, poor generalization performance, and lack of interpretability. It is particularly limited in handling long-distance dependencies, complex reasoning, and cross-modal understanding. These challenges often require intelligent agents to possess more advanced cognitive capabilities and more refined environmental perception.
[0004] In recent years, the emergence of large language models (LLMs) has brought new solutions to visual language navigation tasks. However, these solutions focus on enabling intelligent agents to learn to move according to specific natural language instructions, but cannot achieve autonomous navigation in urban scenarios.
[0005] Therefore, how to achieve autonomous navigation in urban scenarios through intelligent agents has become a technical problem that urgently needs to be solved in the industry. Summary of the Invention
[0006] In response to the above-mentioned deficiencies in the prior art, the present invention provides a method, apparatus, device and storage medium for urban visual navigation based on a large language model, which enables autonomous navigation in urban scenarios through an intelligent agent system.
[0007] In a first aspect, the present invention provides a method for urban visual navigation based on a large language model, the method comprising the following steps: Constructing a fine-tuning dataset based on multiple first street view images of the target city, and fine-tuning the multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model; wherein the annotation information of each first street view image includes landmark locations and landmark distance information corresponding to each first street view image; Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; Based on the natural language description of the target location, the intelligent agent system repeatedly executes the process of perception, reflection, planning and action until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target location.
[0008] According to the present invention, a large language model-based urban visual navigation method is provided. The intelligent agent system includes a perception module, a reflection module, and a planning module. Based on the natural language description of the target location, the intelligent agent system repeatedly performs the perception, reflection, planning, and action process until the target navigation task is completed, including: A natural language description of the target location and a second street view image are input, the perception module uses the fine-tuned large multimodal language model to identify landmarks in the second street view image, and infers the positional relationship between the current location and the target location based on the positional relationship between the target and the landmark; the inference process is implemented based on the law of cosines or the spatial cognition capability of the fine-tuned large multimodal language model; Based on the positional relationship between the current position and the target position, the reflection module corrects the reasoning error to obtain a relative relationship between the current position and the landmark position after reflection in the second street view image and an extracted experience memory; the extracted experience memory includes historical navigation data and an optimized navigation path; Based on the relative relationship between the reflected current location and the landmark location in the second street view image and the extracted experiential memory, the planning module decomposes the optimized navigation path into at least two sub-goals and generates specific action instructions for each sub-goal at the current moment; executing specific action instructions of each of the sub-goals at the current moment to drive the agent system to move to the next node in the optimized navigation path; The position of the next node is determined as the updated current position, and based on the updated current position and the natural language description of the target position, specific action instructions for each sub-target at the next moment are generated, and the specific action instructions for each sub-target at the next moment are executed to drive the intelligent body system to move toward the target position until the target navigation task is completed.
[0009] According to the urban visual navigation method based on a large language model provided by the present invention, the planning module includes a long-term planning module and a short-term decision-making module; the planning module decomposes the optimized navigation path into at least two sub-goals and generates specific action instructions for each sub-goal at the current moment, including: Decomposing the optimized navigation path into a plurality of sub-goals according to the road network structure and the natural language description of the target location by the long-term planning module; The short-term decision module generates specific action instructions for each sub-goal, and dynamically adjusts the action according to the road connection status to dynamically respond to environmental changes.
[0010] According to the present invention, a large language model-based urban visual navigation method is provided, wherein a fine-tuning dataset is constructed based on multiple first street view images of a target city, and the fine-tuning dataset is used to fine-tune a multimodal large language model to obtain a fine-tuned multimodal large language model, including: Determining, based on each of the first street view images, annotation information for each of the first street view images; wherein the location of a landmark corresponding to each of the first street view images is obtained by annotating a location rectangle in the image; and the landmark distance information corresponding to each of the first street view images represents a distance value from the first landmark in each of the first street view images to the current location, wherein the landmark distance information corresponding to each of the first street view images is a distance value in meters; the landmark includes a landmark building or a visibly distinguishable object in the environment; constructing the fine-tuning dataset based on the annotation information of each of the first street view images; the fine-tuning dataset is a training sample for the multimodal large language model; Based on the training samples of the multimodal large language model, the multimodal large language model is fine-tuned to obtain the fine-tuned multimodal large language model.
[0011] According to a large language model-based urban visual navigation method provided by the present invention, determining the labeling information of each first street view image based on each first street view image includes: For any of the first street view images, determining whether the first landmark exists in the first street view image by using a first prompt word; Marking a rectangular frame of the position of the first landmark in the first street view image using a second prompt word; estimating the distance between the first landmark and the current location using the third prompt word; Based on the position rectangle of each first landmark in each first street view image and the distance value between each first landmark and the current position, the labeling information of each first street view image is determined.
[0012] According to a large language model-based urban visual navigation method provided by the present invention, determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model includes: Determining the agent system based on the fine-tuned multimodal large language model and the deployment form of the agent; The deployment form of the intelligent agent includes an intelligent agent in a virtual simulation environment or a physical robot equipped with sensors.
[0013] According to a large language model-based urban visual navigation method provided by the present invention, the intelligent agent system also includes a reflection module, which includes long-term memory and working memory. The long-term memory includes episodic memory and semantic memory. The long-term memory is used to store historical navigation data and semantic experience. The working memory is used to perform prediction-reflection by retrieving episodic memory and semantic memory, combining current perception results with historical memory to correct reasoning errors and generate corrected target reasoning. The episodic memory is used to store historical navigation actions and corresponding perception results in natural language form. The semantic memory is used to summarize the episodic memory through the large language model to form a high-level navigation strategy.
[0014] In a second aspect, the present invention further provides a city visual navigation device based on a large language model, the device comprising the following modules: a model fine-tuning module, configured to construct a fine-tuning dataset based on a plurality of first street view images of a target city, and fine-tune the multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model; wherein the annotation information for each of the first street view images includes landmark locations and landmark distances corresponding to each of the first street view images; Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; An autonomous navigation module is used to repeatedly execute the process of perception, reflection, planning and action through the intelligent agent system based on a natural language description of the target location until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target location.
[0015] In a third aspect, the present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for urban visual navigation based on a large language model as described above is implemented.
[0016] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described urban visual navigation methods based on a large language model.
[0017] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described urban visual navigation methods based on a large language model.
[0018] The present invention provides a large language model-based urban visual navigation method, apparatus, device and storage medium. First, a fine-tuning dataset is constructed based on multiple first street view images of a target city, and the fine-tuning dataset is used to fine-tune the multimodal large language model to obtain a fine-tuned multimodal large language model, wherein the annotation information of each first street view image includes the landmark position and landmark distance information corresponding to each first street view image; then, based on the fine-tuned multimodal large language model, an intelligent agent system for urban visual navigation is determined; further, based on the natural language description of the target position, the intelligent agent system repeatedly executes the process of perception, reflection, planning and action until the target navigation task is completed, wherein the target position description includes the positional relationship between the target and the landmark, the positional relationship includes relative orientation and distance, and the target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target position.
[0019] In the present invention, the multimodal large language model is first fine-tuned based on the street view images of the target city. The fine-tuned multimodal large language model is more suitable for autonomous navigation in the target city scenario. Furthermore, by giving a natural language description of the target location, an intelligent agent system deployed with the fine-tuned multimodal large language model is used to achieve autonomous navigation to the target location in the urban environment, realizing autonomous navigation in the urban scenario through the intelligent agent system. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 This is one of the flow charts of the urban visual navigation method based on the large language model provided by the present invention.
[0022] Figure 2 It is a schematic diagram of the intelligent agent framework provided by the present invention.
[0023] Figure 3It is a schematic diagram of the effect of the target navigation task provided by the present invention.
[0024] Figure 4 It is a schematic diagram of the effect of the urban road network connection provided by the present invention.
[0025] Figure 5 This is the second flow chart of the urban visual navigation method based on the large language model provided by the present invention.
[0026] Figure 6 It is a structural diagram of the urban visual navigation device based on the large language model provided by the present invention.
[0027] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0029] In order to more clearly understand the various embodiments provided by the present invention, the technical content of the present invention is first introduced as follows: Existing related work has the following limitations: (1) Visual navigation methods based on data-driven and deep reinforcement learning rely on high-quality real navigation data, which is difficult to obtain and has high model training costs, resulting in insufficient generalization performance; (2) Existing visual navigation methods based on large language models only consider the agent's command execution capabilities and cannot perform navigation tasks normally in the absence of clear navigation instructions.
[0030] In view of the above shortcomings, the present invention provides an urban visual navigation method, device, equipment and storage medium based on a large language model, which realizes autonomous navigation to the target location in an urban environment through an intelligent agent system based on a given target location description and street view input.
[0031] The following combination Figure 1-Figure 7 The present invention describes the city visual navigation method, apparatus, device and storage medium based on a large language model.
[0032] Figure 1 This is one of the flow charts of the urban visual navigation method based on the large language model provided by the present invention. Figure 1 As shown, the method includes the following: Step 101: construct a fine-tuning dataset based on multiple first street view images of a target city, and use the fine-tuning dataset to fine-tune the multimodal large language model to obtain a fine-tuned multimodal large language model; wherein the annotation information of each first street view image includes the location and distance information of the landmarks corresponding to each first street view image; Specifically, first of all, it should be noted that the execution subject of the present invention is an electronic device, which is used to achieve autonomous navigation in urban scenarios.
[0033] In the method provided in this embodiment, the agent (i.e., the agent system) must first be deployed. It should be noted that the deployment of this agent system requires the selection of a specific target city, such as City A or City B. The deployment process of the agent (i.e., the agent system) is as follows: 1) Data Acquisition and Annotation: Collect street view images of the selected urban area. This step can be accessed through the commercial street view application programming interface (API). After obtaining the first street view image of the target city, the landmark locations and distance information must be annotated. Specifically, the annotation information includes the landmark locations and distance information corresponding to each first street view image. The landmark locations are annotated as rectangular boxes within the image, while the distance information is numerically annotated in meters. Landmarks are not limited to buildings; any object that is clearly distinguishable from the surrounding environment and visible in the majority of the selected area can serve as a landmark.
[0034] 2) Model Fine-tuning: A fine-tuning dataset is constructed using the annotated first street view images. Furthermore, based on the fine-tuning dataset, a chain of thought process is used to enhance the landmark recognition and reasoning capabilities of the multimodal large language model, resulting in a fine-tuned multimodal large language model. For example, after generating the fine-tuning dataset, the multimodal large language model is fine-tuned using Low-Rank Adaptation (LoRA).
[0035] LoRA is a technique for fine-tuning large pre-trained models. Its core idea is to simulate the amount of parameter changes through low-rank decomposition, thereby enabling indirect training of large models with a very small number of parameters. Specifically, LoRA adds low-rank matrices to specific layers of the model. These matrices are updated during the fine-tuning process, while the parameters of the original model remain frozen.
[0036] The multimodal large language model fine-tuned with street view images of the selected target city is more suitable for autonomous navigation in the target city.
[0037] Step 102: Determine an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; Specifically, based on a fine-tuned multimodal large language model, an intelligent agent system for urban visual navigation can be determined. For example, the trained multimodal large language model can be deployed on an intelligent agent, combined with a memory module and a planning module to form an intelligent agent system. The agent can be deployed as a robot in a real-world environment, for example.
[0038] Step 103: Based on the natural language description of the target location, the intelligent agent system repeatedly executes the process of perception, reflection, planning, and action until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current location of the intelligent agent system to the target location.
[0039] Specifically, after the model is deployed to create an intelligent agent system, the system can achieve autonomous navigation. The intelligent agent operation process mainly includes inputting a target location description, and the intelligent agent system repeatedly executes the process of perceiving street scene images, reflecting on historical memory, and planning and deciding navigation behaviors until the navigation target is reached. For example, the process of autonomous navigation using the intelligent agent system is as follows: A natural language representation of the target location is input to the agent system, which is then used to implement a target navigation task, which represents the navigation task from the agent system's current location to the target location. The natural language representation of the target location includes the positional relationship between the target and a landmark, where the positional relationship includes relative bearing and distance. For example, the natural language representation T of the target location could be "The target is 300 meters southeast of CCTV." Southeast indicates the relative bearing, and 300 meters indicates the distance.
[0040] Furthermore, the intelligent system repeats the process of perception, reflection, planning, and action until it navigates to the target location and completes the target navigation task.
[0041] Perception: Utilizes a fine-tuned multimodal large language model to identify landmarks in street view imagery and infers the relative positional relationship between the current location and the target location based on the target-landmark positional relationship. Specifically, the system receives a street view image input and a natural language description of the target location. The fine-tuned multimodal large language model is used to identify landmarks in the street view imagery and infers the relative positional relationship between the current location and the target location based on the target-landmark positional relationship.
[0042] Reflection: Combine the historical navigation data of the intelligent system's memory module with the current perception results (the relative position relationship between the current position and the target position) to correct reasoning errors and optimize path planning.
[0043] Planning: Decompose the optimized navigation path into multiple sub-goals based on the target location and road network structure, and generate specific movement instructions for each sub-goal through short-term decision-making.
[0044] Action: Execute the movement instructions to drive the agent to move towards the target location until the target navigation task is completed.
[0045] In summary, the intelligent agent system achieves autonomous navigation in an urban environment by repeatedly executing a closed-loop process of perception-reflection-planning-action, navigating from the current location to the target location.
[0046] The method provided in this embodiment first constructs a fine-tuning dataset based on multiple first street view images of a target city, and uses the fine-tuning dataset to fine-tune a multimodal large language model to obtain a fine-tuned multimodal large language model, wherein the annotation information of each first street view image includes the landmark location and landmark distance information corresponding to each first street view image; then, based on the fine-tuned multimodal large language model, an intelligent agent system for urban visual navigation is determined; further, based on a natural language description of the target location, the intelligent agent system repeatedly executes the process of perception, reflection, planning, and action until the target navigation task is completed, wherein the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current location of the intelligent agent system to the target location.
[0047] In the present invention, the multimodal large language model is first fine-tuned based on the street view images of the target city. The fine-tuned multimodal large language model is more suitable for autonomous navigation in the target city scenario. Furthermore, by giving a natural language description of the target location, an intelligent agent system deployed with the fine-tuned multimodal large language model is used to achieve autonomous navigation to the target location in the urban environment, realizing autonomous navigation in the urban scenario through the intelligent agent system.
[0048] According to the present invention, a large language model-based urban visual navigation method is provided. The intelligent agent system includes a perception module, a reflection module, and a planning module. Based on the natural language description of the target location, the intelligent agent system repeatedly executes the perception, reflection, planning, and action process until the target navigation task is completed, including: The system inputs a natural language description of the target location and a second street view image. The perception module uses a fine-tuned multimodal large language model to identify landmarks in the second street view image. The system then infers the positional relationship between the current location and the target location based on the positional relationship between the target and the landmarks. This inference process is based on the law of cosines or the spatial cognition capabilities of the fine-tuned multimodal large language model. Based on the positional relationship between the current location and the target location, the reflection module corrects the reasoning error to obtain the relative relationship between the current location and the landmark location after reflection in the second street view image and the extracted experience memory; the extracted experience memory includes historical navigation data and the optimized navigation path; Based on the relative relationship between the reflected current location and the landmark location in the second street view image and the extracted experience memory, the planning module decomposes the optimized navigation path into at least two sub-goals and generates specific action instructions for each sub-goal at the current moment; Execute the specific action instructions of each sub-goal at the current moment to drive the intelligent system to move to the next node in the optimized navigation path; The position of the next node is determined as the updated current position, and based on the natural language description of the updated current position and target position, the specific action instructions for each sub-target at the next moment are generated. The specific action instructions for each sub-target at the next moment are executed to drive the intelligent system to move towards the target position until the target navigation task is completed.
[0049] Specifically, in some embodiments, the agent system includes a perception module, a reflection module, and a planning module; Correspondingly, the specific implementation process of completing the target navigation task by the intelligent agent system in step 103 includes the following steps: First, a natural language description of the target location and a second street view image are input. The perception module uses a fine-tuned multimodal large language model to identify landmarks in the second street view image, and then infers the positional relationship between the current location and the target location based on the positional relationship between the target and the landmark.
[0050] Among them, for the fine-tuned model, when street view images are input, the intelligent agent can identify landmark buildings from the street view images and estimate their position and distance relationship. Then, it can combine the positional relationship between the target and the landmark (obtained from the target description) to infer the positional relationship between the current position and the target position. The reasoning process is based on the spatial cognition ability of the cosine theorem or a fine-tuned multimodal large language model. The reasoning results are convenient for providing support for subsequent navigation.
[0051] Furthermore, based on the positional relationship between the current position and the target position, the reflection module is used to correct the reasoning error, thereby obtaining the relative relationship between the current position and the landmark position after reflection in the second street view image and the extracted experience memory.
[0052] The reflection module includes long-term memory and working memory. Long-term memory consists of episodic memory and semantic memory. Episodic memory stores navigation data, while semantic memory preserves a summary of historical navigation experience. Working memory acts as a data buffer, processing visual perception results and retrieved memories. The retrieved experiential memory includes historical navigation data and optimized navigation paths.
[0053] Furthermore, based on the relative relationship between the reflected current position and the landmark position in the second street view image and the extracted experiential memory, the optimized navigation path is decomposed into at least two sub-goals through the planning module and specific action instructions for each sub-goal at the current moment are generated.
[0054] Furthermore, the specific action instructions for each sub-goal at the current moment are executed to drive the intelligent system to move to the next node in the optimized navigation path. Action means that the intelligent agent moves from one node to another, updates its position and explores the goal in the environment.
[0055] Furthermore, the next node's location is determined as the updated current location. Based on the updated current location and the natural language description of the target location, specific action instructions for each sub-goal at the next moment are generated. These specific action instructions for each sub-goal at the next moment are executed to drive the agent system toward the target location until the target navigation task is completed. In other words, by repeating the steps of perception-reflection-planning-action, the agent can autonomously navigate to the target location described in natural language in a complex urban environment.
[0056] For example, Figure 2 is a schematic diagram of the intelligent agent framework provided by the present invention, such as Figure 2 As shown, the intelligent agent includes a perception module, a reflection module and a planning module.
[0057] First, the street view image S at time t is input t The target description T is input into the perception module. For example, the target description T is “the target is 300 meters southeast of the CCTV station…”. The perception module first performs landmark recognition to obtain the position relationship S between the current position and the landmark position. lm , and then based on the position relationship S between the current position and the landmark position lm And the positional relationship R between the landmark position and the target position lm Perform orientation reasoning to obtain the position relationship R between the current position and the target position g t .
[0058] Furthermore, the working memory in the reflection module retrieves the historical navigation data stored in the long-term memory for prediction and reflection, and combines the road network connection E t After correction, the position relationship R between the current position and the target position after reflection is obtained l t Long-term memory includes episodic memory and semantic memory, which are used for storage and retrieval by working memory.
[0059] After that, the positional relationship R between the current position and the target position after reflection is calculated. l tAnd the extracted experience memory are input into the planning module, and the planned path P (P t-1 ) Generate P through long-term planning t , further combined with real-time decision-making (short-term decision-making) to generate specific action instructions α t , α t For example, "Behavior: Move West."
[0060] Figure 3 This is a schematic diagram of the effect of the target navigation task provided by the present invention, which realizes autonomous navigation from the current position to the target position Goal through the intelligent agent Agent. Figure 3 As shown, landmark1 represents landmark 1, landmark2 represents landmark 2, and the positional relationship between the current position and and the positional relationship between the target position and the landmark is inferred, and the positional relationship between the target position and the current position is obtained. Figure 3 Landmark 1 has been recognized. Figure 4 This is a schematic diagram of the effect of the urban road network connection provided by the present invention. Figure 4 (a), (b), (c), and (d) are schematic diagrams of road network connection effects in four different cities.
[0061] The method provided in this embodiment is based on a given target location description and street view input. The intelligent agent system can autonomously navigate to the target location described in natural language in a complex urban environment by repeatedly executing the steps of perception-reflection-planning-action. This realizes an intelligent agent that can autonomously navigate to the target location in an urban environment, and the success rate and interpretability of the intelligent agent in navigation tasks are greatly improved.
[0062] According to the present invention, a large language model-based urban visual navigation method includes a planning module comprising a long-term planning module and a short-term decision-making module. The planning module decomposes the optimized navigation path into at least two sub-goals and generates specific action instructions for each sub-goal at the current moment, including: Through the long-term planning module, the optimized navigation path is decomposed into multiple sub-goals according to the natural language description of the road network structure and the target location; Through the short-term decision-making module, specific action instructions for each sub-goal are generated, and the actions are dynamically adjusted according to the road connection status to dynamically respond to environmental changes.
[0063] Specifically, in some embodiments, the agent system planning module includes a long-term planning module and a short-term decision-making module. Rather than directly using visual perception information to determine next steps, the agent integrates a planning module that involves both long-term planning and short-term decision-making. Long-term planning is used to decompose the navigation path into multiple sub-goals based on the target location and road network structure; short-term decision-making is used to convert sub-goals into specific movement instructions and dynamically adjust actions based on road connectivity.
[0064] In this embodiment, the specific implementation process of decomposing the optimized navigation path into at least two sub-goals and generating specific action instructions for each sub-goal at the current moment by the planning module includes the following steps: First, the long-term planning module decomposes the optimized navigation path into multiple sub-goals based on the natural language description of the road network structure and target location. For example, long-term planning uses the target location obtained through the agent's reflection, retrieved historical memory, and the previous plan (retrieved navigation plan) as input. It updates the navigation plan by analyzing the execution of the previous plan and decomposing the possible paths into sub-goals. The agent system first analyzes the stage of execution of the original plan and then comprehensively considers goal inference, retrieved memory, and connection status to determine whether to update the plan. If an update to the previous plan is necessary, the agent system updates the previous plan by predicting possible routes to the target and decomposing the complete route into several sub-goals, such as {"move east" until [intersection] and then move north"}.
[0065] Furthermore, the short-term decision module generates specific action instructions for each sub-goal and dynamically adjusts actions based on road connectivity to respond to environmental changes. Specifically, after obtaining each sub-goal, the short-term decision maker converts the plan into a specific action based on road connectivity. Actions involve the agent moving from one node to another, updating its position, and exploring the environment.
[0066] In the method provided in this embodiment, the intelligent system integrates a planning module, which involves long-term planning and short-term decision-making. Long-term planning is used to decompose the navigation path into multiple sub-goals based on the target location and road network structure, and short-term decision-making is used to convert the sub-goals into specific movement instructions and dynamically adjust the actions according to the road connection status, so that the intelligent system can autonomously navigate to the target location described in natural language in a complex urban environment, and the navigation autonomy is good.
[0067] According to the present invention, a large language model-based urban visual navigation method is provided. A fine-tuning dataset is constructed based on multiple first street view images of a target city, and the fine-tuning dataset is used to fine-tune a multimodal large language model to obtain a fine-tuned multimodal large language model, including: Determining, based on each first street view image, annotation information for each first street view image; wherein the location of a landmark corresponding to each first street view image is obtained by annotating a location rectangle in the image; and the landmark distance information corresponding to each first street view image represents a distance value from the first landmark in each first street view image to the current location, wherein the landmark distance information corresponding to each first street view image is a distance value in meters; the landmark includes a landmark building or a visibly distinguishable object in the environment; Based on the annotation information of each first street view image, a fine-tuning dataset is constructed; the fine-tuning dataset is a training sample for the multimodal large language model; Based on the training samples of the multimodal large language model, the multimodal large language model is fine-tuned to obtain a fine-tuned multimodal large language model.
[0068] Specifically, in some embodiments, in step 101, a fine-tuning dataset is constructed based on multiple first street view images of the target city, and the fine-tuning dataset is used to fine-tune the multimodal large language model. The specific implementation process of obtaining the fine-tuned multimodal large language model includes the following steps: First, based on each first street view image, annotation information for each first street view image is determined. The landmark location corresponding to each first street view image is obtained by annotating the location rectangle in the image. The landmark distance information corresponding to each first street view image represents the distance from the first landmark in each first street view image to the current location. The landmark distance information corresponding to each first street view image is a distance value in meters. Landmarks include landmark buildings or visibly distinguishable objects in the environment.
[0069] Furthermore, based on the annotation information (landmark location and landmark distance information) of each first street view image, a fine-tuning dataset is constructed. The fine-tuning dataset serves as a training sample for the multimodal large language model, facilitating subsequent training of the multimodal large language model based on the fine-tuning dataset.
[0070] Furthermore, based on the training samples of the multimodal large language model, the multimodal large language model is fine-tuned to obtain a fine-tuned multimodal large language model.
[0071] In the fine-tuning of large multimodal language models, LoRA is widely used to integrate multi-modal information. Among them, LoRA is a fine-tuning technology that simulates the change of parameters through low-rank decomposition, and realizes indirect training of large models with extremely small number of parameters.
[0072] Key features of LoRA include: 1. Parameter efficiency: By introducing low-rank matrices, LoRA only needs to train a small number of parameters, greatly reducing computing resources and storage requirements.
[0073] 2. Avoid catastrophic forgetting: Since the original model parameters remain unchanged, LoRA can effectively avoid damage to the original task performance during fine-tuning.
[0074] 3. Flexibility: LoRA can be customized for different tasks or modalities and is suitable for a variety of application scenarios.
[0075] The multimodal large language model fine-tuned using LoRA can efficiently adapt to new tasks and data while maintaining the performance of the original model, and has important application value.
[0076] The method provided in this embodiment first determines, based on each first street view image, annotation information for each first street view image. The landmark location corresponding to each first street view image is obtained by annotating the landmark based on the location rectangle in the image. The landmark distance information corresponding to each first street view image represents the distance value from the first landmark in each first street view image to the current location. The landmark distance information corresponding to each first street view image is a distance value in meters. Landmarks include landmark buildings or visibly distinguishable objects in the environment. Then, based on the annotation information of each first street view image, a fine-tuning dataset is constructed. Furthermore, based on the training samples of the multimodal large language model (i.e., the annotation information of each first street view image), the multimodal large language model is fine-tuned to obtain a fine-tuned multimodal large language model. The fine-tuned multimodal large language model can better perform landmark recognition tasks, facilitate subsequent landmark-based location reasoning and navigation planning, and achieve target navigation tasks.
[0077] According to a large language model-based urban visual navigation method provided by the present invention, based on each first street view image, determining labeling information of each first street view image includes: For any first street view image, determining whether a first landmark exists in the first street view image by using a first prompt word; Marking a rectangular frame of a position of the first landmark in the first street view image using a second prompt word; Estimate the distance between the first landmark and the current location using the third prompt word; Based on the position rectangle of each first landmark in each first street view image and the distance value between each first landmark and the current position, the labeling information of each first street view image is determined.
[0078] Specifically, in some embodiments, the specific implementation process of determining the annotation information of each first street view image based on each first street view image is, for example, to construct a fine-tuning dataset based on certain rules. The rules are constructed in a chain-of-thought manner, and the implementation process includes the following steps: First, for any first street view image, whether a first landmark exists in the first street view image is determined by using a first prompt word, such as I. "Is landmark i in the image?"
[0079] Furthermore, the position rectangular box of the first landmark in the first street view image is marked with a second prompt word. The second prompt word is, for example, II. "Landmark i is in the image. What is its rectangular box in the image?" Furthermore, the distance between the first landmark and the current location is estimated using a third prompt word. The third prompt word may be, for example, III. "The location of landmark i in the image is xxx. Estimate the distance from the landmark to the current location." Furthermore, the position rectangular frame of each first landmark in each first street view image and the distance value between each first landmark and the current position are determined as the annotation information of each first street view image.
[0080] The method provided in this embodiment determines annotation information based on these three progressive prompt words (first prompt word, second prompt word, and third prompt word). It also constructs a fine-tuning dataset using a chain-of-thought approach, enabling the fine-tuned large language model to better predict and estimate landmark locations.
[0081] According to the present invention, a large language model-based urban visual navigation method is provided, which determines an intelligent agent system based on a fine-tuned multimodal large language model, including: Determine the agent system based on the fine-tuned multimodal large language model and the agent deployment form; Among them, the deployment forms of intelligent agents include intelligent agents in virtual simulation environments or physical robots equipped with sensors.
[0082] Specifically, the specific implementation process of determining the agent system based on the fine-tuned multimodal large language model in step 102 includes the following steps: First, we need to determine the deployment form of an agent. An agent is a software or hardware entity that can perceive its environment and make decisions to achieve specific goals. Depending on the application scenario and requirements, agents can be deployed in various forms. In this embodiment, the deployment forms of agents include agents in a virtual simulation environment or physical robots equipped with sensors.
[0083] Furthermore, based on the fine-tuned multimodal language model and the agent's deployment form, the agent system is determined. Specifically, the fine-tuned multimodal language model is deployed on a specific target software or hardware entity (such as an agent in a virtual simulation environment or a physical robot equipped with sensors). In other words, the trained model is deployed on the agent, combined with the reflection module and planning module, to form the agent system.
[0084] The process of deploying a model typically includes: 1. Model fine-tuning: Use LoRA technology to fine-tune the multimodal large model and save the fine-tuned model weights and configuration files.
[0085] 2. Environment Configuration: Ensure that the necessary dependent libraries, such as Transformers and Parameter-Efficient Fine-Tuning (PEFT), are installed in the deployment environment. Transformer is a deep neural network that incorporates fundamental technologies in natural language processing.
[0086] 3. Model loading: Load the fine-tuned model and use the corresponding processor (such as AnnotationProcessor) to preprocess the data.
[0087] 4. Interface design: Design the API interface, encapsulate the model's reasoning function, and provide a unified calling method.
[0088] 5. Deployment and testing: Deploy the model to the selected environment and perform functional and performance testing to ensure that the model operates normally.
[0089] The method provided in this embodiment determines the intelligent agent system based on the fine-tuned multimodal large language model and the deployment form of the intelligent agent, and deploys the fine-tuned multimodal large language model to the intelligent agent. Subsequently, the intelligent agent system can be used to achieve autonomous navigation in urban scenarios.
[0090] According to a large language model-based urban visual navigation method provided by the present invention, the intelligent agent system also includes a reflection module, which includes long-term memory and working memory, and the long-term memory includes episodic memory and semantic memory; the long-term memory is used to store historical navigation data and semantic experience, and the working memory is used to retrieve episodic memory and semantic memory, combine current perception results with historical memory for prediction-reflection to correct reasoning errors and generate corrected target reasoning; episodic memory is used to store historical navigation actions and corresponding perception results in natural language form, and semantic memory is used to summarize episodic memory through the large language model to form a high-level navigation strategy.
[0091] Specifically, in some embodiments, the intelligent agent system also includes a reflection module.
[0092] Among them, the reflection module includes two important components: long-term memory and working memory. Long-term memory includes two components: episodic memory and semantic memory.
[0093] Long-term memory is used to store historical navigation data and semantic experience. Episodic memory is used to store historical navigation actions and corresponding perception results in natural language. Semantic memory is used to summarize episodic memory using a large language model to form high-level navigation strategies. In other words, episodic memory stores navigation data, while semantic memory preserves the summary of historical navigation experience.
[0094] Working memory is used to retrieve episodic and semantic memories, combining current perception with historical memories for prediction and reflection to correct reasoning errors and generate corrected target inferences. Working memory can be understood as a data buffer, processing visual perception results and retrieved memories. During navigation, the intelligent system must maintain spatial awareness of its own position and the urban environment. This relies not only on current perception but also on historical trajectories, which is why we designed long-term memory. Because current perception is sometimes inadequate and past actions are not always accurate, a prediction-reflection mechanism is designed into working memory to make the intelligent system more robust.
[0095] Specifically, the technical details of each part are as follows: Episodic Memory: Episodic memory is navigation data expressed in natural language. When the agent makes a move, this action and visual perception are processed into a natural language sentence and stored. Using this stored navigation data, the agent can retrieve target inferences from past locations and detect whether connected nodes in the road network have been visited.
[0096] Semantic Memory: Episodic memory records navigational experiences, while the agent uses LLM to summarize and learn from episodic memory to form semantic memory. Semantic memory is a high-level cognitive function that helps the agent build an internal representation of the navigation map. Like humans, it can learn about the environment based on past experiences and learn more advanced navigation strategies, such as detours to reach the destination. These strategies can be restored to working memory and further facilitate the planning process. As the agent navigates, episodic and semantic memory are updated.
[0097] Working Memory: Working memory receives visual perception results and extracts relevant experience from long-term memory. It incorporates a prediction-reflection mechanism to address the problem of the agent losing its target direction when it cannot detect any landmarks in the street scene. The working memory output to the planning module integrates the target reasoning from the perception module and retrieved historical experience. This enables the agent to navigate complex environments, regardless of whether landmarks are visible or not, making navigation more flexible and robust.
[0098] In the method provided by this embodiment, during actual navigation, working memory receives visual perception results and extracts relevant experience from long-term memory. The output to the planning module is a combination of the target reasoning and retrieved historical experience from the perception module. This enables the intelligent system to handle complex environments, regardless of whether landmarks are visible, making intelligent navigation more flexible and robust.
[0099] Figure 5 This is the second flow chart of the urban visual navigation method based on the large language model provided by the present invention. Figure 5 As shown, the method includes: Data collection, model fine-tuning, model deployment, and agent execution.
[0100] The operation of the intelligent agent includes repeatedly executing the perception, reflection and planning steps until the target navigation task is completed.
[0101] The following describes the urban visual navigation device based on a large language model provided by the present invention. The urban visual navigation device based on a large language model described below and the urban visual navigation method based on a large language model described above can refer to each other.
[0102] Figure 6 This is a schematic diagram of the structure of the urban visual navigation device based on the large language model provided by the present invention. Figure 6 As shown, the urban visual navigation device 600 based on the large language model includes the following modules: A model fine-tuning module 610 is configured to construct a fine-tuning dataset based on a plurality of first street view images of a target city, and fine-tune the multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model; wherein the annotation information for each of the first street view images includes the location and distance information of landmarks corresponding to each of the first street view images; Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; The autonomous navigation module 620 is used to repeatedly execute the process of perception, reflection, planning and action through the intelligent agent system based on the natural language description of the target location until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative direction and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target location.
[0103] The device provided in this embodiment includes a model fine-tuning module 610 and an autonomous navigation module 620. First, the model fine-tuning module 610 constructs a fine-tuning dataset based on multiple first street view images of the target city, and uses the fine-tuning dataset to fine-tune the multimodal large language model to obtain a fine-tuned multimodal large language model, wherein the annotation information of each first street view image includes the landmark position and landmark distance information corresponding to each first street view image; then, the autonomous navigation module 620 determines the intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; further, based on the natural language description of the target position, the intelligent agent system repeatedly executes the perception, reflection, planning and action process until the target navigation task is completed, wherein the target position description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target position.
[0104] In the present invention, the multimodal large language model is first fine-tuned based on the street view images of the target city. The fine-tuned multimodal large language model is more suitable for autonomous navigation in the target city scenario. Furthermore, by giving a natural language description of the target location, an intelligent agent system deployed with the fine-tuned multimodal large language model is used to achieve autonomous navigation to the target location in the urban environment, realizing autonomous navigation in the urban scenario through the intelligent agent system.
[0105] According to the present invention, a large language model-based urban visual navigation device 600 includes an intelligent agent system comprising a perception module, a reflection module, and a planning module; and the autonomous navigation module 620 is specifically configured to: A natural language description of the target location and a second street view image are input, the perception module uses the fine-tuned large multimodal language model to identify landmarks in the second street view image, and infers the positional relationship between the current location and the target location based on the positional relationship between the target and the landmark; the inference process is implemented based on the law of cosines or the spatial cognition capability of the fine-tuned large multimodal language model; Based on the positional relationship between the current position and the target position, the reflection module corrects the reasoning error to obtain a relative relationship between the current position and the landmark position after reflection in the second street view image and an extracted experience memory; the extracted experience memory includes historical navigation data and an optimized navigation path; Based on the relative relationship between the reflected current location and the landmark location in the second street view image and the extracted experiential memory, the planning module decomposes the optimized navigation path into at least two sub-goals and generates specific action instructions for each sub-goal at the current moment; executing specific action instructions of each of the sub-goals at the current moment to drive the agent system to move to the next node in the optimized navigation path; The position of the next node is determined as the updated current position, and based on the updated current position and the natural language description of the target position, specific action instructions for each sub-target at the next moment are generated, and the specific action instructions for each sub-target at the next moment are executed to drive the intelligent body system to move toward the target position until the target navigation task is completed.
[0106] According to the large language model-based urban visual navigation device 600 provided by the present invention, the planning module includes a long-term planning module and a short-term decision-making module; the autonomous navigation module 620 is further configured to: Decomposing the optimized navigation path into a plurality of sub-goals according to the road network structure and the natural language description of the target location by the long-term planning module; The short-term decision module generates specific action instructions for each sub-goal, and dynamically adjusts the action according to the road connection status to dynamically respond to environmental changes.
[0107] According to the urban visual navigation device 600 based on a large language model provided by the present invention, the model fine-tuning module 610 is specifically used to: Determining, based on each of the first street view images, annotation information for each of the first street view images; wherein the location of a landmark corresponding to each of the first street view images is obtained by annotating a location rectangle in the image; and the landmark distance information corresponding to each of the first street view images represents a distance value from the first landmark in each of the first street view images to the current location, wherein the landmark distance information corresponding to each of the first street view images is a distance value in meters; the landmark includes a landmark building or a visibly distinguishable object in the environment; Constructing the fine-tuning dataset based on the annotation information of each of the first street view images; the fine-tuning dataset is a training sample for the multimodal large language model; Based on the training samples of the multimodal large language model, the multimodal large language model is fine-tuned to obtain the fine-tuned multimodal large language model.
[0108] According to the urban visual navigation device 600 based on a large language model provided by the present invention, the model fine-tuning module 610 is further used to: For any of the first street view images, determining whether the first landmark exists in the first street view image by using a first prompt word; Marking a rectangular frame of the position of the first landmark in the first street view image using a second prompt word; estimating the distance between the first landmark and the current location using the third prompt word; Based on the position rectangle of each first landmark in each first street view image and the distance value between each first landmark and the current position, the labeling information of each first street view image is determined.
[0109] According to the urban visual navigation device 600 based on a large language model provided by the present invention, the model fine-tuning module 610 is further used to: Determining the agent system based on the fine-tuned multimodal large language model and the deployment form of the agent; The deployment form of the intelligent agent includes an intelligent agent in a virtual simulation environment or a physical robot equipped with sensors.
[0110] According to an urban visual navigation device 600 based on a large language model provided by the present invention, the intelligent agent system also includes a reflection module, which includes long-term memory and working memory, and the long-term memory includes episodic memory and semantic memory; the long-term memory is used to store historical navigation data and semantic experience, and the working memory is used to perform prediction-reflection by retrieving episodic memory and semantic memory, combining current perception results with historical memory to correct reasoning errors and generate corrected target reasoning; the episodic memory is used to store historical navigation actions and corresponding perception results in natural language form, and the semantic memory is used to summarize the episodic memory through the large language model to form a high-level navigation strategy.
[0111] Figure 7 An example of a physical structure diagram of an electronic device is shown below. Figure 7 As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute the urban visual navigation method based on the large language model, which includes: Constructing a fine-tuning dataset based on multiple first street view images of the target city, and fine-tuning the multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model; wherein the annotation information of each first street view image includes landmark locations and landmark distance information corresponding to each first street view image; Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; Based on the natural language description of the target location, the intelligent agent system repeatedly executes the process of perception, reflection, planning and action until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target location.
[0112] Furthermore, the logic instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0113] In another aspect, the present invention further provides a computer program product, comprising a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the urban visual navigation method based on a large language model provided by the above methods, the method comprising: Constructing a fine-tuning dataset based on multiple first street view images of the target city, and fine-tuning the multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model; wherein the annotation information of each first street view image includes landmark locations and landmark distance information corresponding to each first street view image; Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; Based on the natural language description of the target location, the intelligent agent system repeatedly executes the process of perception, reflection, planning and action until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target location.
[0114] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the urban visual navigation method based on a large language model provided by the above methods, the method comprising: Constructing a fine-tuning dataset based on multiple first street view images of the target city, and fine-tuning the multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model; wherein the annotation information of each first street view image includes landmark locations and landmark distance information corresponding to each first street view image; Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; Based on the natural language description of the target location, the intelligent agent system repeatedly executes the process of perception, reflection, planning and action until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target location.
[0115] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0116] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A city visual navigation method based on a large language model, characterized in that: include: Constructing a fine-tuning dataset based on multiple first street view images of the target city, and fine-tuning the multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model; wherein the annotation information of each first street view image includes landmark locations and landmark distance information corresponding to each first street view image; Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; Based on the natural language description of the target location, the intelligent agent system repeatedly executes the process of perception, reflection, planning and action until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target location.
2. The urban visual navigation method based on a large language model according to claim 1 is characterized in that: The intelligent agent system includes a perception module, a reflection module, and a planning module. Based on the natural language description of the target location, the intelligent agent system repeatedly executes the process of perception, reflection, planning, and action until the target navigation task is completed, including: A natural language description of the target location and a second street view image are input, the perception module uses the fine-tuned large multimodal language model to identify landmarks in the second street view image, and infers the positional relationship between the current location and the target location based on the positional relationship between the target and the landmark; the inference process is implemented based on the law of cosines or the spatial cognition capability of the fine-tuned large multimodal language model; Based on the positional relationship between the current position and the target position, the reflection module corrects the reasoning error to obtain a relative relationship between the current position and the landmark position after reflection in the second street view image and an extracted experience memory; the extracted experience memory includes historical navigation data and an optimized navigation path; Based on the relative relationship between the reflected current location and the landmark location in the second street view image and the extracted experiential memory, the planning module decomposes the optimized navigation path into at least two sub-goals and generates specific action instructions for each sub-goal at the current moment; executing specific action instructions of each of the sub-goals at the current moment to drive the agent system to move to the next node in the optimized navigation path; The position of the next node is determined as the updated current position, and based on the updated current position and the natural language description of the target position, specific action instructions for each sub-target at the next moment are generated, and the specific action instructions for each sub-target at the next moment are executed to drive the intelligent body system to move toward the target position until the target navigation task is completed.
3. The urban visual navigation method based on a large language model according to claim 2 is characterized in that: The planning module includes a long-term planning module and a short-term decision-making module; Decomposing the optimized navigation path into at least two sub-goals and generating specific action instructions for each sub-goal at the current moment by the planning module includes: Decomposing the optimized navigation path into a plurality of sub-goals according to the road network structure and the natural language description of the target location by the long-term planning module; The short-term decision module generates specific action instructions for each sub-goal, and dynamically adjusts the action according to the road connection status to dynamically respond to environmental changes.
4. The urban visual navigation method based on a large language model according to claim 1 is characterized in that: The step of constructing a fine-tuning dataset based on a plurality of first street view images of a target city, and fine-tuning a multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model includes: Determining, based on each of the first street view images, annotation information for each of the first street view images; wherein the location of a landmark corresponding to each of the first street view images is obtained by annotating a location rectangle in the image; and the landmark distance information corresponding to each of the first street view images represents a distance value from the first landmark in each of the first street view images to the current location, wherein the landmark distance information corresponding to each of the first street view images is a distance value in meters; the landmark includes a landmark building or a visibly distinguishable object in the environment; constructing the fine-tuning dataset based on the annotation information of each of the first street view images; the fine-tuning dataset is a training sample for the multimodal large language model; Based on the training samples of the multimodal large language model, the multimodal large language model is fine-tuned to obtain the fine-tuned multimodal large language model.
5. The urban visual navigation method based on a large language model according to claim 4 is characterized in that: The determining, based on each of the first street view images, the labeling information of each of the first street view images includes: For any of the first street view images, determining whether the first landmark exists in the first street view image by using a first prompt word; Marking a rectangular frame of the position of the first landmark in the first street view image using a second prompt word; estimating the distance between the first landmark and the current location using the third prompt word; Based on the position rectangle of each first landmark in each first street view image and the distance value between each first landmark and the current position, the labeling information of each first street view image is determined.
6. The urban visual navigation method based on a large language model according to claim 1 is characterized in that: Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model includes: Determining the agent system based on the fine-tuned multimodal large language model and the deployment form of the agent; The deployment form of the intelligent agent includes an intelligent agent in a virtual simulation environment or a physical robot equipped with sensors.
7. The urban visual navigation method based on a large language model according to any one of claims 1 to 6, characterized in that: The intelligent agent system also includes a reflection module, which includes long-term memory and working memory, and the long-term memory includes episodic memory and semantic memory; the long-term memory is used to store historical navigation data and semantic experience, and the working memory is used to retrieve episodic memory and semantic memory, combine current perception results with historical memory for prediction-reflection to correct reasoning errors, and generate corrected target reasoning; the episodic memory is used to store historical navigation actions and corresponding perception results in natural language form, and the semantic memory is used to summarize episodic memory through a large language model to form a high-level navigation strategy.
8. A city visual navigation device based on a large language model, characterized in that: include: a model fine-tuning module, configured to construct a fine-tuning dataset based on a plurality of first street view images of a target city, and fine-tune the multimodal large language model using the fine-tuning dataset to obtain a fine-tuned multimodal large language model; wherein the annotation information for each of the first street view images includes landmark locations and landmark distances corresponding to each of the first street view images; Determining an intelligent agent system for urban visual navigation based on the fine-tuned multimodal large language model; An autonomous navigation module is used to repeatedly execute the process of perception, reflection, planning and action through the intelligent agent system based on a natural language description of the target location until the target navigation task is completed; the target location description includes the positional relationship between the target and the landmark, and the positional relationship includes relative orientation and distance. The target navigation task is used to represent the navigation task from the current position of the intelligent agent system to the target location.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the urban visual navigation method based on the large language model as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the urban visual navigation method based on a large language model as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
A low-altitude street view sign semantic navigation sample generation method, device and medium
CN122524148A