Navigation method, model training method, equipment and storage medium
Through the navigation method based on the visual language action model, the scene understanding ability is improved by using visual encoder and depth features, the problems of insufficient scenario understanding and poor performance in the prior art are solved, and more efficient action navigation is achieved.
Patent Information
- Application Number
- CN202411997347.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
When performing action navigation processing, the existing technology has low spatial reasoning capabilities and navigation performance due to insufficient scenario understanding and poor performance, which affects the accuracy of action navigation.
Using a navigation method based on the visual language action model, image features and depth features are extracted through a visual encoder, and combined with natural language navigation command information, action information is generated for action navigation. The model architecture includes a visual encoder, an alignment projection module and an action decoder. Through phased training, injecting depth information to improve scene comprehension capabilities.
It improves spatial reasoning capabilities and navigation performance, enhances the accuracy of action navigation, and can achieve better navigation effects in complex scenarios.
Smart Images

Figure CN119940365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of motion navigation technology, and in particular to a navigation method, a model training method, a device and a storage medium. Background Art
[0002] Vision-and-Language Navigation (VLN) has attracted much attention in recent years. VLN is a fundamental and challenging task in embedded artificial intelligence, requiring a robot to navigate in an unseen environment based on visual input and natural language instructions.
[0003] The latest VLN agents utilize a Vision Language Model (VLM) to model historical observations via visual tags generated by the VLM to reason about spatiotemporal relationships during navigation.
[0004] However, there are still problems with insufficient 3D scene understanding and poor performance when modeling historical observations to reason about spatial relationships during navigation via visual markers generated by VLM.
[0005] That is to say, when the relevant technology performs motion navigation processing, due to insufficient scene understanding and poor performance, it reduces the spatial reasoning ability and navigation performance, affecting the accuracy of motion navigation. Summary of the invention
[0006] The embodiments of the present application provide a navigation method, a model training method, a device and a storage medium, which can enhance spatial reasoning ability and navigation performance, and improve the accuracy of motion navigation.
[0007] The technical solution of the embodiment of the present application is implemented as follows:
[0008] In a first aspect, an embodiment of the present application provides a navigation method based on a visual language action model, the method comprising:
[0009] When receiving the natural language navigation instruction information, obtaining video data of a first time series; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information;
[0010] The action information of the second time series is obtained through a visual language action model based on natural language navigation instruction information, image information of the first time series and depth information of the first time series; wherein the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer; the visual language action model is obtained by phased training based on a training data set;
[0011] Action navigation is performed based on the action information of the second time series.
[0012] The embodiment of the present application proposes a navigation method based on a visual language action model, based on the received natural language navigation instruction information, combined with video data including image information and depth information, the corresponding action information is obtained through the visual language action model, and then action navigation is performed according to the action information. Among them, the model architecture and phased training strategy based on the visual language action model can respectively extract image features and depth features through the image coding layer and the depth coding layer in the visual encoder, so that in the process of reasoning action information, one-to-one corresponding image information and depth information can be introduced at the same time, solving the problems of insufficient scene understanding and poor performance, thereby improving spatial reasoning ability and navigation performance, and improving the accuracy of action navigation.
[0013] In a second aspect, an embodiment of the present application provides a method for training a visual language action model, the method comprising:
[0014] The initial model is trained in stages based on the training data set to obtain a visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer;
[0015] The initial model is trained in stages based on the training data set to obtain a visual language action model, including:
[0016] The parameters of the visual encoder and the action decoder are fixed, and the alignment projection module in the initial model is trained based on the training data set to obtain a model that completes the first stage of training;
[0017] Based on the model that has completed the first stage of training, the parameters of the image coding layer and the parameters of the action decoder in the visual encoder are fixed, and the depth coding layer, shared coding layer and alignment projection module are trained based on the training data set to obtain the model that has completed the second stage of training;
[0018] Based on the model that has completed the second stage of training, the parameters of the image encoding layer in the visual encoder are fixed, and the deep encoder, shared encoding layer, alignment projection module and action decoder are trained based on the training data set to obtain a visual language action model;
[0019] Among them, the training data set includes natural language command information, video data and actual action information; among them, the video data includes image information of a third time series and depth information of the third time series corresponding to the image information, and the actual action information includes actual action information of a fourth time series.
[0020] An embodiment of the present application proposes a training method for a visual language action model. During the model training process, the visual encoder, alignment projection module and action decoder in the visual language action model are trained in stages according to a training data set including natural language command information, image information, depth information and actual action information, and the depth information is injected into the model to obtain a model including a deep feature generation channel, so that the trained visual language action model has better scene understanding ability, and thus can obtain ideal spatial reasoning ability and navigation performance when the visual language action model is used for action navigation in the future.
[0021] In a third aspect, an embodiment of the present application provides a navigation device based on a visual language action model, and the navigation device based on the visual language action model includes:
[0022] An acquisition unit, configured to acquire video data of a first time series upon receiving natural language navigation instruction information; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information;
[0023] The acquisition unit is further used to obtain the action information of the second time series through a visual language action model based on the natural language navigation instruction information, the image information of the first time series and the depth information of the first time series; wherein the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer; the visual language action model is obtained by phased training based on a training data set;
[0024] The navigation unit is used to perform action navigation based on the action information of the second time series.
[0025] In a fourth aspect, an embodiment of the present application provides a training device for a visual language action model, and the training device for a visual language action model includes:
[0026] A training unit, used to train the initial model in stages based on the training data set to obtain a visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer;
[0027] The training unit is further used to fix the parameters of the visual encoder and the parameters of the action decoder, train the alignment projection module in the initial model based on the training data set, and obtain a model that completes the first stage of training; based on the model that completes the first stage of training, fix the parameters of the image coding layer in the visual encoder and the parameters of the action decoder, and train the depth coding layer, the shared coding layer, and the alignment projection module based on the training data set to obtain a model that completes the second stage of training; based on the model that completes the second stage of training, fix the parameters of the image coding layer in the visual encoder, and train the depth encoder, the shared coding layer, the alignment projection module, and the action decoder based on the training data set to obtain a visual language action model;
[0028] The training data set includes natural language instruction information, video data and actual action information; the video data includes image information of a third time series and depth information of a third time series, and the actual action information includes actual action information of a fourth time series.
[0029] In a fifth aspect, an embodiment of the present application provides a computer device, the computer device comprising: a processor and a memory; wherein:
[0030] The memory is used to store a computer program that can be run on the processor;
[0031] The processor is used to execute the method described in the first aspect or the second aspect when running the computer program.
[0032] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having a program stored thereon, and when the program is executed by a processor, the method described in the first aspect or the second aspect above is implemented.
[0033] The embodiment of the present application proposes a navigation method, a model training method, a device and a storage medium. On the one hand, based on the received natural language navigation instruction information, combined with the video data including image information and depth information, the corresponding action information is obtained through the visual language action model, and then the action navigation is performed according to the action information. Among them, based on the model architecture and the staged training strategy of the visual language action model, the image features and depth features can be extracted respectively through the image coding layer and the depth coding layer in the visual encoder, so that in the process of reasoning the action information, the one-to-one corresponding image information and depth information can be introduced at the same time, solving the problems of insufficient scene understanding and poor performance, thereby reducing the spatial reasoning ability and navigation performance, and improving the accuracy of action navigation. On the other hand, in the process of model training, the visual encoder, the alignment projection module and the action decoder in the visual language action model are respectively trained in stages according to the training data set including natural language instruction information, image information, depth information and actual action information, and the depth information is injected into the model to obtain a model including a deep feature generation channel, so that the trained visual language action model has a good scene understanding ability, and then the ideal spatial reasoning ability and navigation performance can be obtained when the visual language action model is used for action navigation in the future. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic diagram of the implementation flow of the navigation method based on the visual language action model proposed in the embodiment of the present application;
[0035] Figure 2 A schematic diagram of a visual language action model proposed in an embodiment of the present application;
[0036] Figure 3 A schematic diagram of implementing action navigation based on a visual language action model proposed in an embodiment of the present application;
[0037] Figure 4 A schematic diagram of implementing action navigation based on a visual language action model proposed in an embodiment of the present application;
[0038] Figure 5 A schematic diagram of implementing action navigation based on a visual language action model proposed in an embodiment of the present application;
[0039] Figure 6 A schematic diagram of implementing action navigation based on a visual language action model proposed in an embodiment of the present application;
[0040] Figure 7 A schematic diagram of navigation based on a visual language action model proposed in an embodiment of the present application;
[0041] Figure 8A schematic diagram of navigation based on a visual language action model proposed in an embodiment of the present application;
[0042] Fig. 9 A schematic diagram of the implementation flow of the training method for the visual language action model proposed in the embodiment of the present application;
[0043] Fig.10 A schematic diagram of an initial model proposed for an embodiment of the present application;
[0044] Fig.11 A schematic diagram of 4D spatiotemporal reasoning capability navigation in a visual language proposed in an embodiment of the present application;
[0045] Fig.12 A schematic diagram of 4D spatiotemporal reasoning capability navigation in a visual language proposed in an embodiment of the present application;
[0046] Fig.13 A schematic diagram of 4D spatiotemporal reasoning capability navigation in a visual language proposed in an embodiment of the present application;
[0047] Fig.14 A schematic diagram of 4D spatiotemporal reasoning capability navigation in a visual language proposed in an embodiment of the present application;
[0048] Fig.15 A schematic diagram of 4D spatiotemporal reasoning capability navigation in a visual language proposed in an embodiment of the present application;
[0049] Fig.16 A schematic diagram of implementing the action navigation proposed in the embodiment of the present application;
[0050] Fig.17 This is a schematic diagram of the evaluation of NaVid-4D proposed in the embodiment of the present application;
[0051] Fig.18 This is a schematic diagram of the experimental evaluation of NaVid-4D proposed in the embodiment of the present application;
[0052] Fig.19 A schematic diagram of the structure of a navigation device based on a visual language action model proposed in an embodiment of the present application;
[0053] Fig. 20 A schematic diagram of the structure of a training device for a visual language action model proposed in an embodiment of the present application;
[0054] Fig.21 A schematic diagram of the composition structure of the computer device proposed in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It is understood that the specific embodiments described herein are only used to explain the related applications, rather than to limit the applications. It should also be noted that, for ease of description, only the parts related to the related applications are shown in the drawings.
[0056] Vision-Language Navigation (VLN) has attracted much attention in recent years. In simulated environments, a common approach is to discretize the scene, where the robotic agent moves by teleporting between nodes on a predefined navigation graph by aligning language and visual observations to make decisions. However, these methods often perform poorly when VLN models trained in discrete space are directly transferred to continuous 3D real-world robotic applications. To address this issue, the Habitat simulator is utilized and the Vision-Language Navigation in Continuous Environments benchmark (VLN-CE) is proposed, which allows the robotic agent to freely navigate to any unobstructed space within the simulator. At each time step, the agent predicts actions based on visual observations and language instructions using direct low-level control prediction or selecting navigable subgoals estimated by a waypoint predictor. Recently, Large Language Models (LLMs) and Vision-Language Models (VLMs) trained on Internet-scale image, video, and text data have shown remarkable capabilities in multimodal reasoning and cross-domain generalization.
[0057] Vision and Language Navigation (VLN) is a fundamental and challenging task in embedded AI, requiring robots to navigate in unseen environments based on visual input and natural language instructions. Recent advances have transformed VLN from a discrete simulator setting to a more realistic continuous setting in the real world, enabling robots to navigate more naturally like humans, rather than just transitioning between waypoints on a predefined navigation map. Developing a truly general real-world VLN system places higher demands on 3D spatial intelligence, especially in reasoning about spatial relations (e.g., following instructions like “go to the farthest room”). Emerging large-scale vision-language models (VLMs) show great potential in addressing these challenges and shaping the future of VLN research due to their wide range of perception and advanced language understanding capabilities. State-of-the-art VLN agents leverage vision-language models (VLMs) to model historical observations via visual tags generated by the VLMs, thereby reasoning about spatiotemporal relations during navigation. While these methods benefit from the perceptual capabilities of current VLMs, they are still limited by a key drawback: the related VLMs are based on RGB visual ground truth models and do not explicitly incorporate depth information, which leads to insufficient 3D scene understanding and poor performance when the task requires reasoning about spatial relationships by modeling historical observations via VLM-generated visual landmarks during navigation.
[0058] Significant efforts have been made in improving the spatial reasoning capabilities of visual language models (VLMs). Pioneering research has mainly focused on incorporating 3D representations, such as multi-view images or point clouds, to inject spatial information into VLMs. However, the limited availability of multi-view image and point cloud data limits the effectiveness of these methods. Meanwhile, some methods attempt to enhance spatial reasoning capabilities without directly incorporating 3D representations. For example, ConceptGraph integrates scene graphs into VLMs to capture spatial relationships. SpatialVLM builds an Internet-scale 3D spatial reasoning dataset to train 2D VLMs, significantly improving their performance in spatial Visual Question-Answering (VQA) tasks. SpatialRGPT and SpatialBot further enhance the spatial reasoning capabilities of VLMs by integrating 3D inputs into their architectures. While these advances have demonstrated spatial intelligence in digital question-answering tasks, there is still a gap before they can be effectively applied to embodied visual language action (VLA) models.
[0059] That is to say, when performing motion navigation processing, the relevant technology has problems with insufficient scene understanding and poor performance, which in turn reduces the spatial reasoning ability and navigation performance, and affects the accuracy of motion navigation.
[0060] In order to solve the above problems, in an embodiment of the present application, the embodiment of the present application proposes a navigation method, a model training method, an apparatus and a storage medium, wherein the navigation device of the visual language action model obtains video data of a first time series when receiving natural language navigation instruction information; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information; through the visual language action model, based on the natural language navigation instruction information, the image information of the first time series and the depth information of the first time series, the action information of the second time series is obtained; wherein the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer; the visual language action model is obtained by phased training based on a training data set; and action navigation is performed based on the action information of the second time series. The training device of the visual language action model trains the initial model in stages based on the training data set to obtain the visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer; the initial model is trained in stages based on the training data set to obtain the visual language action model, including: fixing the parameters of the visual encoder and the parameters of the action decoder, training the alignment projection module in the initial model based on the training data set to obtain a model that has completed the first stage of training; based on the model that has completed the first stage of training, fixing the parameters of the image coding layer and the action decoder in the visual encoder The parameters of the decoder are determined based on the training data set, and the depth coding layer, shared coding layer and alignment projection module are trained to obtain a model that completes the second stage of training; based on the model that completes the second stage of training, the parameters of the image coding layer in the visual encoder are fixed, and the depth encoder, shared coding layer, alignment projection module and action decoder are trained based on the training data set to obtain a visual language action model; wherein the training data set includes natural language instruction information, video data and actual action information; wherein the video data includes image information of a third time series and depth information of a third time series corresponding to the image information, and the actual action information includes actual action information of a fourth time series. It can be seen that in an embodiment of the present application, on the one hand, based on the received natural language navigation instruction information, combined with the video data including image information and depth information, the corresponding action information is obtained through the visual language action model, and then action navigation is performed according to the action information. Among them, the model architecture based on the visual language action model and the staged training strategy can extract image features and depth features respectively through the image coding layer and depth coding layer in the visual encoder, so that in the process of reasoning action information, one-to-one corresponding image information and depth information can be introduced at the same time, solving the problems of insufficient scene understanding and poor performance, thereby improving spatial reasoning ability and navigation performance, and improving the accuracy of action navigation.On the other hand, during the model training process, the visual encoder, alignment projection module and action decoder in the visual language action model are trained in stages according to the training data set including natural language command information, image information, depth information and actual action information, and the depth information is injected into the model to obtain a model including a deep feature generation channel, so that the trained visual language action model has better scene understanding ability, and thus can obtain ideal spatial reasoning ability and navigation performance when using the visual language action model for action navigation in the future.
[0061] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0062] An embodiment of the present application provides a navigation method based on a visual language action model, wherein the navigation method based on a visual language action model can be applied to a navigation device or a computer device based on a visual language action model, and the present application does not specifically limit it. Below, taking a navigation device based on a visual language action model as an example, the navigation method based on a visual language action model proposed in the embodiment of the present application is exemplarily described.
[0063] Furthermore, in the embodiments of the application, Figure 1 The following is a schematic diagram of the implementation process of the navigation method based on the visual language action model proposed in the embodiment of the present application, as shown in FIG. Figure 1 As shown, the navigation method based on the visual language action model may include the following steps:
[0064] Step 101: upon receiving natural language navigation instruction information, obtaining video data of a first time series; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information.
[0065] In an embodiment of the present application, a navigation device based on a visual language action model may first receive natural language navigation instruction information, and then further acquire video data of a first time series in response to the received natural language navigation instruction information.
[0066] In an embodiment of the present application, the navigation device based on the visual language action model is a device or apparatus capable of navigating through the visual language action model, wherein the navigation device based on the visual language action model has spatial understanding ability and spatial reasoning ability.
[0067] For example, in some embodiments, the navigation device based on the visual language action model can be a robot, wherein the robot can implement action navigation processing through the pre-trained visual language action model. Of course, the navigation device based on the visual language action model can also be any other device with spatial understanding ability and spatial reasoning ability, which is not specifically limited in this application.
[0068] In an embodiment of the present application, natural language navigation instruction information can be used to indicate navigation information to a navigation device based on a visual language action model. Among them, natural language navigation instruction information can be understood as natural language instructions for navigation. Natural language instructions refer to instructions or commands expressed in languages used by humans in daily life (such as Chinese, English, etc.). These instructions can be understood by humans, and in some cases, can also be understood and executed by computers or artificial intelligence systems.
[0069] Exemplarily, in some embodiments, the natural language navigation instruction information received by the navigation device based on the visual language action model may include but is not limited to "walk forward three steps", "stop", "turn around and go back five meters, then turn right and walk to the end", etc.
[0070] Accordingly, in an embodiment of the present application, the navigation device based on the visual language action model also has the ability to parse natural language navigation instructions, that is, for the received natural language navigation instruction information, the navigation device based on the visual language action model can parse and obtain the effective information used for navigation. Among them, in order to understand and respond to the received natural language navigation instruction information, the navigation device based on the visual language action model usually needs to be supported by natural language processing (NLP) technology, and NLP technology includes but is not limited to: semantic understanding, syntactic analysis, named entity recognition, context understanding, etc.
[0071] Semantic understanding refers to analyzing the semantic content of instructions to understand their meaning and intent. Syntactic analysis refers to analyzing the syntactic structure of instructions to identify the subject, predicate, object and other components. Named entity recognition refers to identifying entity names in instructions, such as names of people, places, time, etc. Contextual understanding refers to understanding the context of the current instruction based on previous conversations or text content.
[0072] In an embodiment of the present application, the video data of the first time series acquired by the navigation device based on the visual language action model may include image information of the first time series, and may also include depth information of the first time series corresponding to the image information. The first time series may include a continuous time interval or a discrete time point. The time length corresponding to the first time series may be pre-set or randomly determined, and the present application does not specifically limit it.
[0073] It can be understood that, in the embodiments of the present application, at least one moment in the first time series can be understood as at least one time step, and the time length of the first time series can be understood as the number of time steps in the first time series.
[0074] Exemplarily, in some embodiments, assuming that the time length corresponding to the first time series is 3, that is, the number of corresponding time steps is 3, then the first time series may include three consecutive moments (t1, t2, t3).
[0075] Exemplarily, in some embodiments, assuming that natural language navigation instruction information is received at time t, correspondingly, the first time series includes n+1 moments from time tn to time t, the image information of the first time series includes n+1 image information corresponding to the n+1 moment, and the depth information of the first time series includes n+1 depth information corresponding to the n+1 moment. Wherein, n is an integer greater than or equal to 0, and t is an integer greater than n.
[0076] In an embodiment of the present application, the image information acquired by the navigation device based on the visual language action model may include an image sequence corresponding to the first time series.
[0077] Exemplarily, in some embodiments, assuming that the first time series is (t1, t2, t3), then the image information corresponding to the first time series may include an image sequence (x1, x2, x3), wherein image frame x1 corresponds to time t1, image frame x2 corresponds to time t2, and image frame x3 corresponds to time t3.
[0078] In an embodiment of the present application, the depth information corresponding to the image information acquired by the navigation device based on the visual language action model may include a depth map sequence corresponding to the first time series.
[0079] Exemplarily, in some embodiments, assuming that the first time series is (t1, t2, t3), then the depth information corresponding to the first time series may include a depth map sequence (d1, d2, d3), wherein the depth map d1 corresponds to time t1, the depth map d2 corresponds to time t2, and the depth map d3 corresponds to time t3.
[0080] In an embodiment of the present application, the image information and depth information included in the video data may be in one-to-one correspondence based on time, that is, for the tth moment, there may be one-to-one correspondence between the image information xt and the depth information dt, where t is an integer greater than 0.
[0081] Further, in an embodiment of the present application, the image information of the first time series may be obtained through a shooting module. The navigation device based on the visual language action model may be configured with a shooting module, and image acquisition may be performed through the configured shooting module, so that image information at any time may be obtained.
[0082] In an embodiment of the present application, after collecting image information through a shooting module, the navigation device based on the visual language action model can store the collected image information.
[0083] Furthermore, in an embodiment of the present application, the depth information corresponding to the image information in the first time series may be acquired through a depth detection module, or may be determined based on the collected image information, and the present application does not make any specific limitation thereto.
[0084] In an embodiment of the present application, the navigation device based on the visual language action model may also store the depth after acquiring the depth information corresponding to the image information.
[0085] Exemplarily, in some embodiments, the navigation device based on the visual language motion model may be configured with a depth detection module, and the depth information may be detected by the configured depth detection module, so that the depth information at any moment may be obtained.
[0086] For example, in some embodiments, after acquiring image information through the configured shooting module, the navigation device based on the visual language motion model can further estimate the depth information based on the image information to determine the corresponding depth information. For example, the navigation device based on the visual language motion model can obtain the corresponding depth information based on the image information estimation through a depth estimation algorithm or a depth estimation model.
[0087] That is, in an embodiment of the present application, a method for acquiring video data of a first time series may include: acquiring image information at time t through a shooting module; acquiring depth information at time t through a depth detection module; and simultaneously acquiring n pieces of image information and n pieces of depth information stored from time tn to time t-1, wherein n is an integer greater than or equal to 0, and t is an integer greater than n.
[0088] In an embodiment of the present application, another method for obtaining video data of a first time series may include: obtaining image information at moment t through a shooting module, and then performing depth estimation based on the image information at moment t to determine the depth information at moment t; and simultaneously obtaining n image information and n depth information stored from moment tn to moment t-1.
[0089] It can be understood that in the embodiments of the present application, the navigation device based on the visual language action model can store the image information and the corresponding depth information acquired at each historical moment, and then in the subsequent process of acquiring video data, on the one hand, it can acquire the image information and the corresponding depth information of the current moment in real time, and on the other hand, read the image information and the corresponding depth information of the stored historical moments, thereby forming the video data of the first time series.
[0090] That is to say, in the embodiment of the present application, assuming that the first time series includes multiple consecutive moments, the corresponding video data of the first time series may include video data at the current moment and video data at historical moments.
[0091] Step 102: Obtain action information of a second time series through a visual language action model based on natural language navigation instruction information, image information of the first time series, and depth information of the first time series; wherein the visual language action model includes a visual encoder, an alignment projection module, and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer, and a shared coding layer; and the visual language action model is obtained by phased training based on a training data set.
[0092] In an embodiment of the present application, when natural language navigation command information is received, after obtaining video data of a first time series, action information of a second time series can be further obtained through a visual language action model based on the natural language navigation command information, the image information of the first time series, and the depth information of the first time series.
[0093] It can be understood that in the embodiments of the present application, the visual language action model can be used to perform action navigation processing in combination with various types of data and information. For example, the visual language action model can perform robot navigation processing based on corresponding image information and depth information.
[0094] That is, in the embodiments of the present application, the visual language action model has the ability to accurately perform robot navigation processing by referring to image information and depth information.
[0095] Exemplarily, in some embodiments, the visual-language-motion model may be a VLM-based navigation model, such as NAVID-4D, wherein the visual-language-motion model can utilize RGB-D information to improve spatial reasoning ability and navigation performance.
[0096] Furthermore, in an embodiment of the present application, the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer.
[0097] For example, in some embodiments, Figure 2 This is a schematic diagram of the visual language action model proposed in the embodiment of the present application, such as Figure 2 As shown, the visual language action model can be composed of a visual encoder, an alignment projection module and an action decoder. Among them, the visual encoder can be used to extract feature information; the alignment projection module can perform alignment processing based on the extracted feature information and the acquired natural language navigation instruction information; the action decoder can determine the action information of the next time series based on the information token and the information token corresponding to the natural language navigation instruction information.
[0098] It should be noted that in an embodiment of the present application, the action decoder can be constructed based on a pre-trained large language model, and the codewords corresponding to the action signals are added to the code table of the large language model, so that it can output the robot action information in the form of text.
[0099] It should be noted that, in an embodiment of the present application, the modality alignment module can be used to map the visual representation to the input space of the pre-trained large language model, that is, the representation space shared with the language modality.
[0100] It can be understood that in an embodiment of the present application, the visual language action model can be based on the relevant visual language navigation, and further integrates a new model architecture that can use video data including image information and depth information (i.e., RGB-D video data) to realize navigation functions.
[0101] For example, in some embodiments, the visual language action model can be PHYSI-NAVID (a model combining physical operation (PhysX) and visual navigation (Navigation)) based on visual language navigation VLM, which can use the navigation agent information of RGB-D video to provide strong spatial information reasoning capabilities for decision-making. In the future, the core idea behind the visual language action model (PHYSI-NAVID) is to use a suitable depth data source and effectively integrate it, and integrate the depth information into the VLM framework, thereby improving the spatial understanding ability of the navigation agent.
[0102] Furthermore, in an embodiment of the present application, when obtaining action information of the second time series based on natural language navigation instruction information, image information of the first time series and depth information of the first time series, the image features and depth features corresponding to the video data of the first time series can be obtained through a visual encoder based on the image information of the first time series and the corresponding depth information of the first time series; then, image information tokens and depth information tokens can be obtained based on the image features, depth features and natural language navigation instruction information through an alignment projection module; and finally, the action information of the second time series can be obtained through an action decoder based on the image information tokens, depth information tokens and text information tokens corresponding to the natural language navigation instruction information.
[0103] For example, in some embodiments, Figure 3 This is a schematic diagram of implementing action navigation based on a visual language action model proposed in an embodiment of the present application, such as Figure 3 As shown, the image information of the first time series and the corresponding depth information of the first time series can be input into the visual encoder, and the image features and depth features corresponding to the video data of the first time series can be output respectively. Then, the image features and depth features can be input into the alignment projection module. At the same time, the natural language navigation instruction information can be input into the alignment projection module, and the image information token and the depth information token can be output respectively. Furthermore, the image information token, the depth information token and the text information token corresponding to the natural language navigation instruction information can be input into the action decoder, and finally the action information of the second time series can be output.
[0104] In the embodiment of the present application, the second time series may include a continuous time interval or a discrete time point. The time length corresponding to the second time series may be preset or randomly determined, and the present application does not specifically limit it.
[0105] It can be understood that, in the embodiment of the present application, at least one moment in the second time series can be understood as at least one time step, and the time length of the second time series can be understood as the number of time steps in the second time series.
[0106] In the embodiment of the present application, the time length of the second time series may be the same as or different from the time length of the first time series, and the present application does not make any specific limitation.
[0107] In the embodiment of the present application, assuming that the natural language navigation instruction information is received at the tth time, the action information of the second time series generated corresponding to the video data of the first time series may be the action information of multiple moments after the tth time. That is, the second time series may include one or more moments after the first time series.
[0108] Exemplarily, in some embodiments, it is assumed that the first time series includes n+1 moments from moment tn to moment t, and the second time series includes k moments from moment t+1 to moment t+k, where n and k are integers greater than or equal to 0.
[0109] It can be understood that, in the embodiment of the present application, the length n+1 of the first time series and the length k of the second time series may be the same or different, and the present application does not make any specific limitation.
[0110] Exemplarily, in some embodiments, assuming that the time length corresponding to the first time series is 3, that is, the corresponding number of time steps is 3, then the first time series may include three consecutive moments (t1, t2, t3), and assuming that the time length corresponding to the second time series is 2, that is, the corresponding number of time steps is 2, then the second time series may include two consecutive moments (t4, t5).
[0111] Exemplarily, in some embodiments, assuming that the time length corresponding to the first time series is 3, then the first time series may include three consecutive moments (t1, t2, t3), and assuming that the time length corresponding to the second time series is 3, then the second time series may include three consecutive moments (t4, t5, t6).
[0112] Furthermore, in an embodiment of the present application, when obtaining image features and depth features corresponding to video data of a first time series based on image information of the first time series and depth information of the first time series through a visual encoder, the encoded image information can be first obtained through an image coding layer based on the image information of the first time series; at the same time, the encoded depth information can be obtained through a depth coding layer based on the depth information of the first time series; finally, the image features and depth features can be obtained through a shared coding layer based on the encoded image information and the encoded image information.
[0113] It is understood that in an embodiment of the present application, the visual encoder can be a 3D-aware visual encoder, wherein a CLIP model based on ViT can be adopted and expanded to a 3D-aware encoder, i.e., a visual encoder. The encoder is responsible for learning the visual representation of RGB (image information) and depth information, and is initialized from the weights of the CLIP model pre-trained on a large-scale RGB image. In view of the different characteristics of RGB and depth information, different network parameters can be used in the shallow layer of the visual encoder to encode RGB and depth respectively, for example, RGB and depth are encoded respectively by an image coding layer and a depth coding layer. The two models share the same architecture and weight initialization with the RGB-based CLIP model. Afterwards, RGB and depth information have been represented in a relatively aligned visual space. Then, a shared layer, i.e., a shared coding layer pair, can be used for further encoding to obtain a set of visual markers. These visual markers encode visual information in a complementary manner, while RGB and depth markers synergistically encode texture and geometric features.
[0114] Further, in an embodiment of the present application, visual markers may be compressed to encode historical frames. For example, for each time step, each historical frame uses 4 content tokens and 1 context token, while the current frame uses 512 content tokens, 2 context tokens (including 257 RGB tokens and 257 depth tokens).
[0115] In an embodiment of the present application, a content token may include an image information token (RGB token) and a depth information token (depth token). A content token may be used to characterize key visual elements or features. For example, an RGB token may be used to represent color information in an image, and a depth token may be used to represent depth or distance information of an object in an image. A context token may be used to maintain or transmit information about the relationship between these elements or features. A context token may include an RGB token and / or a depth token.
[0116] It should be noted that the number of tokens mentioned here, such as 4 content tokens, 1 context token, 512 content tokens, etc., is only for illustrative purposes and may need to be adjusted according to specific circumstances in actual applications.
[0117] For example, in some embodiments, Figure 4 This is a schematic diagram of implementing action navigation based on a visual language action model proposed in an embodiment of the present application, such as Figure 4As shown, the visual encoder may include an image coding layer (RGB Layers), a depth coding layer (Depth Layers), and a shared coding layer (Shared Layers). The image information of the first time series may be input into the image coding layer (RGB Layers) to extract image features to obtain encoded image information; at the same time, the depth information of the first time series may be input into the depth coding layer (Depth Layers) to extract depth features to obtain encoded depth information; finally, the encoded image information and the encoded image information may be input into the shared coding layer (Shared Layers), and finally the image features and depth features may be output respectively.
[0118] It can be understood that in the embodiments of the present application, the image information and the depth information can be in one-to-one correspondence based on the first time series. Therefore, the image features and depth features finally obtained by the visual encoder are also in one-to-one correspondence based on the first time series.
[0119] Further, in an embodiment of the present application, the alignment projection module may include a modal alignment module and a projection module. The modal alignment module may be a Q-Former-based projector. The modal alignment module may be implemented by Q-Former, and the projection module may be implemented by Projector.
[0120] It is understandable that in an embodiment of the present application, a Q-Former-based method can be used to further convert visual markers into a more general space so that it is compatible with the processing of a pre-trained large language model in a subsequent action decoder. Among them, Q-Former is a Transformer module for processing image features and interacting with text features. It is a lightweight Transformer encoder that accepts features from an image encoder (such as ResNet, ViT, or SwinTransformer) and generates a series of query vectors. These query vectors can interact with the output of a text encoder (such as BERT or RoBERTa) to achieve cross-modal interaction between images and text. In the context of a large multimodal model, Projector is often used with modules such as Q-Former.
[0121] Furthermore, in an embodiment of the present application, when obtaining image information tokens and depth information tokens based on image features, depth features and natural language navigation instruction information through an alignment projection module, the aligned image features and aligned depth features can be obtained first through a modal alignment module based on image features, depth features and natural language navigation instructions; and then the image information tokens and depth information tokens can be obtained based on the aligned image features and aligned depth features through a projection module.
[0122] For example, in some embodiments, Figure 5 This is a schematic diagram of implementing action navigation based on a visual language action model proposed in an embodiment of the present application, such as Figure 5 As shown, the alignment and projection module includes a modal alignment module (Q-Former) and a projection module (Projector), wherein the image features, depth features and natural language navigation instruction information can be first input into the alignment module (Q-Former), and the alignment processing is performed through the interaction of the image features, depth features and text features to obtain the aligned image features and the aligned depth features, and then the aligned image features and the aligned depth features can be input into the projection module (Projector) to finally obtain the image information token and the depth information token.
[0123] Furthermore, in an embodiment of the present application, after obtaining the image information token and the depth information token, the text information token corresponding to the image information token, the depth information token and the natural language navigation instruction information can be input into the action decoder to finally obtain the action information of the second time series.
[0124] It can be understood that in an embodiment of the present application, an action decoder based on a pre-trained large language model, such as an action decoder based on a pre-trained large language model (LLM), can further add code words corresponding to action signals (such as text information tokens corresponding to natural language navigation instruction information) to the code table of the large language model, so that it can output the robot action information in the form of text.
[0125] That is to say, in an embodiment of the present application, through an action decoder based on a pre-trained large language model, the action information of the second time series finally outputted can be in the form of text, that is, text information. For example, the action information can be "stop", "turn around", "forward", "backward", etc.
[0126] Among them, in the field of artificial intelligence and machine learning, LLM is built based on deep learning technology and can understand and generate natural language text. In recent years, with the improvement of computing power and the increase of training data, large language models have made significant progress in the field of natural language processing, such as the GPT series (GPT-3, GPT-4, etc.) and BERT.
[0127] For example, in some embodiments, Figure 6 This is a schematic diagram of implementing action navigation based on a visual language action model proposed in an embodiment of the present application, such as Figure 6 As shown, the image information tokens (RGB tokens) and depth information tokens (Depth tokens) output by the alignment projection module, and the text information tokens (Texttokens) corresponding to the natural language navigation instruction information can be input as input information to the action decoder based on the pre-trained LLM, and then the action information corresponding to the second time series can be output.
[0128] It can be understood that in an embodiment of the present application, the action information corresponds to the second time series, wherein each moment (time step) in the second time series may correspond to action information, that is, the action information of the second time series may include action instructions corresponding to the moment (time step).
[0129] Exemplarily, in some embodiments, the second time series may include four consecutive moments (t1, t2, t3, t4), and the action information corresponding to the second time series may include forward corresponding to moment t1, forward corresponding to moment t2, left turn corresponding to moment t3, and stop corresponding to moment t4.
[0130] That is to say, in the embodiments of the present application, for different moments (time steps), the corresponding action instructions in the action information may be the same or different, and the present application does not make any specific limitation.
[0131] Further, in an embodiment of the present application, the visual language action model can be obtained by staged training based on a training data set. The staged training process may include: first fixing the parameters of the visual encoder and the parameters of the action decoder, training the alignment projection module in the initial model based on the training data set, and obtaining a model that completes the first stage of training; then based on the model that completes the first stage of training, fixing the parameters of the image coding layer in the visual encoder and the parameters of the action decoder, training the depth coding layer, shared coding layer and alignment projection module based on the training data set, and obtaining a model that completes the second stage of training; finally, based on the model that completes the second stage of training, fixing the parameters of the image coding layer in the visual encoder, training the depth encoder, shared coding layer, alignment projection module and action decoder based on the training data set, and obtaining a visual language action model.
[0132] That is to say, in an embodiment of the present application, the phased training may include at least three phases of training process. In the first training phase, the depth coding layer, shared coding layer and alignment projection module may be trained while other model data are fixed; the second training phase may further train the depth coding layer, shared coding layer and alignment projection module while other model data are fixed; the third training phase may train the depth encoder, shared coding layer, alignment projection module and action decoder while other model data are fixed, and finally the trained visual language action model may be obtained.
[0133] Further, in an embodiment of the present application, the training data set may include natural language instruction information, video data, and actual action information. Among them, the natural language instruction information can be understood as natural language instructions for navigation. Natural language instruction information refers to instructions or commands expressed in languages used by humans in daily life (such as Chinese, English, etc.). These instructions can be understood by humans, and in some cases, can also be understood and executed by computers or artificial intelligence systems. The video data includes image information of a third time series and depth information of a third time series corresponding to the image information, and the actual action information includes actual action information of a fourth time series.
[0134] Step 103: Perform action navigation based on the action information of the second time series.
[0135] In an embodiment of the present application, after obtaining the action information of the second time series based on natural language navigation instruction information, image information of the first time series and depth information of the first time series through a visual language action model, action navigation can be further performed based on the action information of the second time series.
[0136] It can be understood that in the embodiment of the present application, based on the image information of the first time series and the depth information corresponding to the image information, combined with the acquired natural language navigation instruction information, after the action information of the second time series is determined by the visual language action model, action navigation can be further performed according to the action information of the second time series. Among them, the action information of the second time series may include action instructions corresponding to the moment (time step), so in the process of action navigation, for each moment of the second time series, navigation is performed according to the corresponding action instruction.
[0137] Exemplarily, in some embodiments, assuming that the navigation device based on the visual language action model is a robot, the second time series may include five consecutive moments (t1, t2, t3, t4, t5), and the action information corresponding to the second time series may include forward corresponding to moment t1, right turn corresponding to moment t2, forward corresponding to moment t3, left turn corresponding to moment t4, and stop corresponding to moment t5. Accordingly, when the robot performs action navigation based on the action information of the second time series, it may move forward at moment t1, turn right at moment t2, move forward at moment t3, turn left at moment t4, and stop at moment t5.
[0138] In summary, the navigation method based on the visual language action model proposed in the above steps 101 to 103 uses a navigation agent based on a visual language model (VLM), that is, a visual language action model based on the VLM for action navigation processing. Among them, through the visual language action model, based on a given natural language instruction (natural language navigation instruction information), combined with video data including image information and depth information (RGB-D), spatial understanding and reasoning are performed, so that accurate action information can be generated, and action navigation is performed based on the action information.
[0139] Furthermore, the navigation method based on the visual language action model proposed in the embodiment of the present application has accurate spatial understanding and reasoning capabilities for the network data of the visual language action model because the visual language action model uses data from the simulation environment to learn the navigation strategy. Based on the visual language action model, without the need for pre-training the RGB-D base model, deep features can be directly injected into the VLM visual encoder, so that the visual language action model has better performance and achieves impressive VLN performance with spatial intelligence in the real world.
[0140] For example, in some embodiments, Figure 7 A schematic diagram of navigation based on a visual language action model proposed in an embodiment of the present application is shown in FIG. Figure 7As shown, based on the visual language action model proposed in the embodiment of the present application, the received natural language navigation instruction information, the acquired video data of the first time series including the image information of the first time series and the depth information of the first time series corresponding to the image information are input into the visual language action model, the action information of the time series can be outputted accordingly, and then the action navigation can be performed according to the action information based on the second time series.
[0141] For example, in some embodiments, Figure 8 A schematic diagram of navigation based on a visual language action model proposed in an embodiment of the present application is shown in FIG. Figure 8 As shown, the visual language action model proposed in the embodiment of the present application may include a visual encoder, an alignment projection module and an action decoder. Among them, the visual encoder may include an image coding layer (RGB Layers), a depth coding layer (Depth Layers), and a shared coding layer (Shared Layers), and the alignment projection module includes a modal alignment module (Q-Former) and a projection module (Projector). The image information of the first time series is input into the image coding layer (RGBLayers) to extract image features and obtain the encoded image information; at the same time, the depth information of the first time series is input into the depth coding layer (Depth Layers) to extract depth features and obtain the encoded depth information; then the encoded image information and the encoded image information after encoding are input into the shared coding layer (Shared Layers) to obtain image features and depth features respectively. Furthermore, the image features, depth features and natural language navigation instruction information are input into the alignment module (Q-Former) for alignment processing to obtain aligned image features and aligned depth features, and then the aligned image features and aligned depth features can be input into the projection module (Projector) to obtain image information tokens (RGB tokens) and depth information tokens (Depth tokens). Finally, the image information tokens (RGB tokens) and depth information tokens (Depth tokens) output by the alignment projection module, as well as the text information tokens (Text tokens) corresponding to the natural language navigation instruction information are input into the action decoder based on the pre-trained LLM, and the action information corresponding to the second time series is output.
[0142] The embodiment of the present application provides a navigation method based on a visual language action model, in which the navigation device of the visual language action model obtains video data of a first time series when receiving natural language navigation instruction information; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information; through the visual language action model, based on the natural language navigation instruction information, the image information of the first time series and the depth information of the first time series, the action information of the second time series is obtained; wherein the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer; the visual language action model is obtained by phased training based on a training data set; and action navigation is performed based on the action information of the second time series. It can be seen that in the embodiment of the present application, based on the received natural language navigation instruction information, combined with the video data including image information and depth information, the corresponding action information is obtained through the visual language action model, and then action navigation is performed according to the action information. Among them, the model architecture based on the visual language action model and the staged training strategy can extract image features and depth features respectively through the image coding layer and depth coding layer in the visual encoder, so that in the process of reasoning action information, one-to-one corresponding image information and depth information can be introduced at the same time, solving the problems of insufficient scene understanding and poor performance, thereby improving spatial reasoning ability and navigation performance, and improving the accuracy of action navigation.
[0143] An embodiment of the present application provides a method for training a visual language action model, wherein the method for training a visual language action model can be applied to a training device or computer device for a visual language action model, and the present application does not specifically limit it. Below, taking a training device for a visual language action model as an example, the method for training a visual language action model proposed in the embodiment of the present application is exemplarily described.
[0144] In the embodiments of the present application, the navigation device based on the visual language action model and the training device of the visual language action model can be the same device, or can be integrated in the same device or terminal, or can be different independently set devices, which is not specifically limited in the present application.
[0145] Furthermore, in the embodiments of the application, Fig. 9 This is a schematic diagram of the implementation process of the training method of the visual language action model proposed in the embodiment of the present application, such as Fig. 9 As shown, the training method of the visual language action model may include the following steps:
[0146] Step 201, training the initial model in stages based on the training data set to obtain a visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer.
[0147] In an embodiment of the present application, a training device for a visual language action model may use a training data set to perform phased training on an initial model, thereby obtaining a trained visual language action model.
[0148] For example, in some embodiments, Fig.10 A schematic diagram of the initial model proposed in the embodiment of the present application, such as Fig.10 As shown, the initial model can be composed of a visual encoder, an alignment projection module and an action decoder. The visual encoder can be used to extract feature information; the alignment projection module can perform alignment processing based on the extracted feature information and the acquired natural language navigation instruction information; the action decoder can determine the action information of the next time series based on the information token and the information token corresponding to the natural language navigation instruction information.
[0149] It should be noted that in an embodiment of the present application, the action decoder can be constructed based on a pre-trained large language model, and the codewords corresponding to the action signals are added to the code table of the large language model, so that it can output the robot action information in the form of text.
[0150] It should be noted that, in an embodiment of the present application, the modality alignment module can be used to map the visual representation to the input space of the pre-trained large language model, that is, the representation space shared with the language modality.
[0151] It can be understood that in an embodiment of the present application, the initial model can be based on the relevant visual language navigation, and further integrates a new model architecture that can use video data (i.e., RGB-D video data) including image information and depth information to achieve navigation functions.
[0152] Accordingly, in an embodiment of the present application, the visual language action model obtained based on the initial model training can be the PHYSI-NAVID (a model combining physical operation (PhysX) and visual navigation (Navigation)) of the VLM, which can use the navigation agent information of the RGB-D video to provide a strong spatial information reasoning capability for decision-making. In the future, the core idea behind the visual language action model (PHYSI-NAVID) is to use a suitable depth data source and effectively integrate it, and integrate the depth information into the VLM framework, thereby improving the spatial understanding ability of the navigation agent.
[0153] Furthermore, in an embodiment of the present application, in the process of performing phased training of the initial model based on a training data set, the parameters of the visual encoder and the parameters of the action decoder can be fixed first, and the alignment projection module in the initial model can be trained based on the training data set to obtain a model that has completed the first stage of training; then, based on the model that has completed the first stage of training, the parameters of the image coding layer in the visual encoder and the parameters of the action decoder are fixed, and the depth coding layer, shared coding layer and alignment projection module are trained based on the training data set to obtain a model that has completed the second stage of training; finally, based on the model that has completed the second stage of training, the parameters of the image coding layer in the visual encoder are fixed, and the depth encoder, shared coding layer, alignment projection module and action decoder are trained based on the training data set to obtain a visual language action model.
[0154] That is to say, in an embodiment of the present application, the visual language action model is obtained by staged training based on a training data set. Among them, the staged training may include at least three stages of training process. Among them, the first training stage can train the depth coding layer, shared coding layer and alignment projection module while fixing other model data; the second training stage can further train the depth coding layer, shared coding layer and alignment projection module while fixing other model data; the third training stage can train the depth encoder, shared coding layer, alignment projection module and action decoder while fixing other model data, and finally the trained visual language action model can be obtained.
[0155] Furthermore, in an embodiment of the present application, the training data set may include natural language instruction information, video data, and actual action information.
[0156] In the embodiments of the present application, natural language instruction information can be understood as natural language instructions for navigation. Among them, natural language instruction information refers to instructions or commands expressed in languages used by humans in daily life (such as Chinese, English, etc.). These instructions can be understood by humans, and in some cases, can also be understood and executed by computers or artificial intelligence systems.
[0157] For example, in some embodiments, the natural language instruction information may include but is not limited to "walk forward three steps", "stop", "turn around and walk back five meters, then turn right and walk to the end", etc.
[0158] It can be understood that, in the embodiment of the present application, the video data includes image information of a third time series and depth information of the third time series corresponding to the image information, and the actual action information includes actual action information of a fourth time series.
[0159] In the embodiment of the present application, the third time series may include a continuous time interval or a discrete time point. The time length corresponding to the third time series may be preset or randomly determined, and the present application does not specifically limit it.
[0160] It can be understood that, in the embodiment of the present application, at least one moment in the third time series can be understood as at least one time step, and the time length of the third time series can be understood as the number of time steps in the third time series.
[0161] Exemplarily, in some embodiments, assuming that the time length corresponding to the third time series is i+1, that is, the corresponding number of time steps is i, then the first time series may include i+1 consecutive moments (ts, ..., t(s+i)), where s and i are integers greater than 0.
[0162] Accordingly, in the embodiment of the present application, the image information of the third time series is the image sequence corresponding to the third time series. For example, assuming that the third time series is (ts, ..., t(s+i)), then the image information corresponding to the third time series may include an image sequence (xs, ..., x(s+i)), wherein the image frame xs corresponds to the time ts, and the image frame x(s+i) corresponds to the time t(s+i).
[0163] Accordingly, in an embodiment of the present application, the depth information of the third time series is a depth map sequence corresponding to the third time series. For example, assuming that the third time series is (ts, ..., t(s+i)), then the image information corresponding to the third time series may include a depth map sequence (ds, ..., d(s+i)), wherein the depth map ds corresponds to time ts, and the depth map d(s+i) corresponds to time t(s+i).
[0164] In an embodiment of the present application, the fourth time series may include a continuous time interval or a discrete time point. The time length corresponding to the fourth time series may be preset or randomly determined, and the present application does not specifically limit it.
[0165] It can be understood that, in the embodiment of the present application, at least one moment in the fourth time series can be understood as at least one time step, and the time length of the fourth time series can be understood as the number of time steps in the fourth time series.
[0166] In an embodiment of the present application, the time length of the fourth time series may be the same as or different from the time length of the third time series, which is not specifically limited in the present application. The fourth time series may include one or more moments after the third time series.
[0167] Exemplarily, in some embodiments, it is assumed that the third time series includes m+1 moments from moment tm to moment t, and the fourth time series includes p moments from moment t+1 to moment t+p, where m and p are integers greater than or equal to 0.
[0168] It can be understood that, in the embodiment of the present application, the length m+1 of the third time series and the length p of the fourth time series may be the same or different, and the present application does not make any specific limitation.
[0169] It can be understood that in the embodiments of the present application, the actual action information is obtained based on the natural language instruction information and corresponds to the fourth time series, wherein each moment (time step) in the fourth time series can correspond to action information, that is, the action information of the fourth time series can include action instructions corresponding to the moment (time step).
[0170] Exemplarily, in some embodiments, the fourth time series may include four consecutive moments (t1, t2, t3, t4), and the action information corresponding to the fourth time series may include a left turn corresponding to moment t1, forward corresponding to moment t2, forward corresponding to moment t3, and a stop corresponding to moment t4.
[0171] Further, in an embodiment of the present application, for the first training stage, when the parameters of the visual encoder and the parameters of the action decoder are fixed, the projection module in the initial model is trained based on the training data set to obtain a model that completes the first stage of training. The initial model can be used to obtain the first predicted action information of the fourth time series based on the first instruction information and the first video data in the training data set; wherein the first video data includes the first image information of the third time series and the first depth information of the third time series, and the training data set also includes the first actual action information of the fourth time series corresponding to the first instruction information; then the parameters of the visual encoder and the parameters of the action decoder can be fixed, and based on the first predicted action information of the fourth time series and the first actual action information of the fourth time series, the parameters of the alignment projection module can be updated to obtain a model that completes the first stage of training.
[0172] It can be understood that in the embodiment of the present application, in the first training stage, based on the first instruction information, after the action information corresponding to the first image information and the first depth information of the third time series, that is, the first predicted action information of the fourth time series, is derived through the initial model, the loss can be further calculated through the actual action information corresponding to the first image information and the first depth information of the third time series, that is, the first actual action information of the fourth time series, and then the model parameters can be corrected through the loss. Among them, the parameters of the alignment projection module can be corrected according to the loss, thereby completing the first stage of training.
[0173] That is, in the embodiment of the present application, in the first training stage, considering that both the visual encoder and the LLM-based action decoder have been pre-trained on a large scale, the projector (aligned projection module) can be pre-warmed while keeping the rest of the network frozen. At this stage, all layers of the visual encoder share RGB and depth information.
[0174] Further, in an embodiment of the present application, for the second training stage, based on the model that has completed the first stage training, the parameters of the image coding layer and the parameters of the action decoder in the visual encoder are fixed, and the depth coding layer, the shared coding layer and the alignment projection module are trained based on the training data set to obtain the model that has completed the second stage training. The second predicted action information of the fourth time series can be obtained based on the second instruction information and the second video data in the training data set by the model that has completed the first stage training; wherein the second video data includes the second image information of the third time series and the second depth information of the third time series, and the training data set also includes the second actual action information of the fourth time series corresponding to the second instruction information; the parameters of the image coding layer and the parameters of the action decoder in the visual encoder are fixed, and based on the second predicted action information of the fourth time series and the second actual action information of the fourth time series, the parameters of the depth coding layer, the parameters of the shared coding layer and the parameters of the alignment projection module are updated to obtain the model that has completed the second stage training.
[0175] It can be understood that in the embodiment of the present application, in the second training stage, based on the second instruction information, after the action information corresponding to the second image information and the second depth information of the third time series, that is, the second predicted action information of the fourth time series, is derived through the initial model, the loss can be further calculated through the actual action information corresponding to the second image information and the second depth information of the third time series, that is, the second actual action information of the fourth time series, and then the model parameters can be corrected through the loss. Among them, the parameters of the depth coding layer, the parameters of the shared coding layer, and the parameters of the alignment projection module can be corrected according to the loss, thereby completing the second stage training.
[0176] That is, in an embodiment of the present application, in the second training stage, on the basis of training the aligned projection module, the deep coding layer of the visual encoder and the subsequent shared encoder layer can be further trained, while keeping the shallow coding layer of the RGB and action decoders still frozen.
[0177] Further, in an embodiment of the present application, for the third training stage, based on the model that has completed the second stage training, the parameters of the image coding layer in the visual encoder are fixed, and the depth encoder, shared coding layer, alignment projection module and action decoder are trained based on the training data set to obtain the visual language action model. The third predicted action information of the fourth time series can be obtained based on the third instruction information and the third video data in the training data set by completing the second stage training model; wherein the third video data includes the third image information of the third time series and the third depth information of the third time series, and the training data set also includes the third actual action information of the fourth time series corresponding to the third instruction information; then the parameters of the image coding layer in the visual encoder are fixed, and based on the third predicted action information of the fourth time series and the third actual action information of the fourth time series, the parameters of the depth encoder, the parameters of the shared coding layer, the parameters of the alignment projection module and the parameters of the action decoder are updated to obtain the visual language action model.
[0178] It can be understood that in the embodiment of the present application, in the third training stage, based on the third instruction information, after the action information corresponding to the third image information and the third depth information of the third time series, that is, the third predicted action information of the fourth time series, is derived through the initial model, the loss can be further calculated through the actual action information corresponding to the third image information and the third depth information of the third time series, that is, the third actual action information of the fourth time series, and then the model parameters can be corrected through the loss. Among them, the parameters of the depth coding layer, the parameters of the shared coding layer, and the parameters of the alignment projection module can be corrected according to the loss, thereby completing the third stage of training.
[0179] That is, in an embodiment of the present application, in the third training stage, based on the training alignment projection module, the deep encoding layer of the visual encoder and the shared encoder layer, the action decoder is further unfrozen to enhance the instruction-aligned model output.
[0180] It can be understood that in the embodiments of the present application, model training data consisting of navigation strategy data and semantic understanding data are used in three training stages, wherein the present application does not specifically limit the loss functions used in the three training stages.
[0181] Furthermore, in an embodiment of the present application, after completing the model training of the three training stages, in the final stage, the DAgger algorithm may also be used to enhance the navigation strategy data.
[0182] In summary, through the training method of the visual language action model proposed in the above step 201, a model architecture including a depth information generation channel is proposed. At the same time, the video data in the training data set of the training model includes one-to-one corresponding image information and depth information, and the depth information can be injected into the conventional VLM. At the same time, using LLaMA-VID as the basic model for pre-training knowledge transfer, and integrating VLN-CE action planning and auxiliary tasks through a co-training framework, a visual language action model with relatively ideal performance can be finally obtained.
[0183] The embodiment of the present application provides a method for training a visual language action model, wherein the training device of the visual language action model performs phased training on an initial model based on a training data set to obtain a visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer; the initial model is trained in phases based on the training data set to obtain a visual language action model, including: fixing parameters of the visual encoder and parameters of the action decoder, training the alignment projection module in the initial model based on the training data set to obtain a model that has completed the first phase of training; based on the model that has completed the first phase of training, fixing the parameters of the visual encoder and the action decoder, and obtaining a model that has completed the first phase of training. The parameters of the image coding layer and the parameters of the action decoder are fixed, and the depth coding layer, the shared coding layer and the alignment projection module are trained based on the training data set to obtain a model that completes the second stage of training; based on the model that completes the second stage of training, the parameters of the image coding layer in the visual encoder are fixed, and the depth encoder, the shared coding layer, the alignment projection module and the action decoder are trained based on the training data set to obtain a visual language action model; wherein the training data set includes natural language instruction information, video data and actual action information; wherein the video data includes image information of a third time series and depth information of a third time series corresponding to the image information, and the actual action information includes actual action information of a fourth time series. It can be seen that in the embodiment of the present application, in the process of model training, the visual encoder, the alignment projection module and the action decoder in the visual language action model are respectively trained in stages according to the training data set including natural language instruction information, image information, depth information and actual action information, and the depth information is injected into the model to obtain a model including a deep feature generation channel, so that the trained visual language action model has a good scene understanding ability, and then when the visual language action model is used for action navigation in the future, it can obtain an ideal spatial reasoning ability and navigation performance.
[0184] Based on the above embodiments, another embodiment of the present application proposes a navigation method based on a visual language action model and a training method for a visual language action model. Among them, the visual language action model proposed in the embodiment of the present application can be NaVid-4D, which is a VLM-based navigation agent that can generate navigation actions end-to-end according to the instructions of a self-centered RGB-D video stream. For the training process of NaVid-4D, a new model paradigm and training strategy are proposed to explicitly encode and utilize depth information. Therefore, NaVid-4D particularly enhances the ability of spatial understanding and reasoning.
[0185] Understanding and reasoning about 4D spacetime is crucial for vision and language navigation (VLN). However, related research lacks in-depth exploration in this area, resulting in bottlenecks in the spatial perception and action accuracy of VLN agents. The visual language action model proposed in the embodiment of the present application, namely NaVid-4D, is the first to clearly demonstrate the capabilities of spatial intelligence in the real world.
[0186] NaVid-4D uses data from simulated environments to learn navigation strategies and uses network data to have accurate spatial understanding and reasoning capabilities. A method is proposed to inject depth features directly into the VLM visual encoder without the need for a pre-trained RGB-D base model. The depth information can be acquired or calculated and estimated. The use of the actual captured depth information is compared with the single estimated depth information. It is found that NaVid-4D is applicable in both cases, while the use of estimation is better and narrows the gap from simulation to reality.
[0187] A large number of experiments show that NaVid-4D proposed in the embodiments of the present application achieves state-of-the-art performance in a simulation environment and achieves impressive VLN performance with spatial intelligence in the real world.
[0188] The embodiments of the present application aim to bridge this gap by exploring the 3D spatial perception capabilities required for robots to successfully complete VLN tasks, thereby demonstrating spatial intelligence in the physical world.
[0189] In an embodiment of the present application, NaVid-4D is an end-to-end VLA model that requires an egocentric RGB-D video stream and natural language instructions as model input to generate navigation operations in a continuous environment, improve spatial understanding and reasoning capabilities in VLN tasks, and demonstrate spatial intelligence in the physical world.
[0190] NaVid-4D is built on top of a pre-trained VLM consisting of a visual base model and a large language model (LLM) to leverage the general knowledge gained from large-scale pre-training. It is extended to an end-to-end VLA model for robot navigation by learning navigation policies from data in simulation environments and further enhancing perception of Internet data. Compared to NaVid, which only supports RGB, NaVid-4D further explores three fundamental issues:
[0191] 1) Are explicit 3D features required in the VLN task?
[0192] 2) How to effectively integrate 3D features into the pre-trained base model?
[0193] 3) What is the impact of different sources of 3D information (captured vs. estimated) on the performance of the VLA model?
[0194] Depth information is necessary for digital tasks that require spatial understanding and reasoning, since embodying intelligent tasks requires actions in the physical world, which places higher demands on spatial perception. Therefore, it can be argued that explicitly utilizing depth information in VLA models is also crucial. So far, the biggest challenge has been the lack of sufficiently powerful RGB-D visual base models, since it is extremely expensive to obtain large-scale RGB-D image-text pairs for pre-training. To address this problem, this application proposes a new method to effectively extract deep features and inject them into an off-the-shelf RGB-based visual base model as an RGB-D encoder. Considering the domain shift between navigation data and jointly trained network data, as well as between simulated and real-world environments, the impact of captured depth and estimated depth on end-to-end performance is further compared. Among them, 1.8 million image-text pairs were collected in digital question-answering tasks and applied to the co-training of NaVid-4D to enhance its spatial perception capabilities.
[0195] Furthermore, in the embodiments of the present application, the contributions of the main work can be summarized into three aspects:
[0196] 1) We built NaVid-4D, an end-to-end VLAN model. Specifically, given natural language instructions, NaVid-4D only needs an egocentric RGB-D video stream as input information to perform spatial understanding and reasoning, thereby generating precise instructions after the robot operates.
[0197] 2) We propose a new model paradigm and its corresponding training strategy to effectively inject deep representations into existing RGB-based base models, thus avoiding the high cost of constructing RGB-D base models.
[0198] 3) We demonstrate the benefits of explicitly integrating VLN depth information in both simulated environments and simulated real-world generalization, and compare the impact of adopting captured or estimated depth information on the final end-to-end performance.
[0199] In an embodiment of the present application, NaVid-4D can understand the 3D world about the 1D temporal relationships between different tasks, and directly predict actions to follow given instructions. For example, Figures 11 to 15 This is a schematic diagram of the 4D spatiotemporal reasoning capability navigation in the visual language proposed in the embodiment of the present application, such as Fig.11 As shown in , for the reasoning about relative positions, given the natural language instruction is "enter the room from the door. Walk past the chair and wait behind the chair." Fig.12 As shown in , for the reasoning of distance comparison, the given natural language instruction is "go forward and wait next to the farthest chair". Fig.13 As shown in , for direction reasoning, the given natural language instruction is "turn left, walk to the chair, and wait." Fig.14 As shown in , for the reasoning of size comparison, the given natural language instruction is "move to the larger plant and stop". Fig.15 As shown in Figure 2, for reasoning about counting, the given natural language instruction is "walk to the first chair on your left and wait."
[0200] In an embodiment of the present application, at time t, a language instruction I is provided to an agent (such as a robot), which includes l words, and a video stream ot = {o0, o1, ..., ot}, which includes a sequence of RGB frames {x0, x1, ..., xt} and its corresponding depth map {d0, d1, ..., dt}. At time t, the agent, based on NaVid-4D proposed in an embodiment of the present application, needs to plan low-level operations {At, At+1, ..., At+k} for the next k steps based on the current observations or following the instruction I. Afterwards, At, At+1, when executing +k, the agent will go to the next step and receive a new observation ot+k. In this work, a VLA agent is constructed to generate at, at+1, , +k from a self-centered video stream Ot, and the instruction I is executed in an end-to-end manner at each time step. The role and method of explicitly utilizing depth information to release spatial intelligence in 4D modeling of VLN-CE are studied in depth.
[0201] NaVid-4D is an end-to-end Vision-Language-Action (VLA) model for VLN-CE, which is extended from a pre-trained Vision-Language Model (VLM). It receives language instructions and RGB-D video streams from a monoscopic centered camera as input to achieve end-to-end continuous navigation action generation. The input of NaVid-4D includes three fine-grained modalities, namely RGB (image information), depth (depth information), and language (natural language navigation instruction information). Although both RGB and depth information belong to visual modalities, the information they represent is both related and different. RGB is more suitable for expressing texture information, while depth maps are more suitable for conveying geometric information. Considering this property, a new model paradigm is designed to gradually adjust different modalities. Specifically, RGB and depth information are first learned to represent in a shared space, and then they are projected into a more general alignment space together with the language modality. It is worth mentioning that this model paradigm can be built on top of the classic pre-trained VLM to make full use of its general knowledge to learn navigation policies.
[0202] For example, Fig.16 This is a schematic diagram of the implementation of the action navigation proposed in the embodiment of the present application, such as Fig.16 As shown, the model paradigm of NaVid-4D includes a 3D perception visual encoder (visual encoder Vision Encoder), a Q-Former-based projector (aligned projection module, including Q-Former and Projector) and an action decoder based on a large language model (LLM) (action decoder LLM-based Action Decoder). Among them, under the framework of NaVid-4D, based on the natural language navigation instruction information "move to the third chair on the left and then stop", at each time step, NaVid-4D will take the self-centered RGB-D video stream and the given instruction as input, such as the image information of the first time series (..., x t-2 , x t-1 , x t ) and the corresponding first time series depth information (..., d t-2 , d t-1 , d t ), generating low-level executable actions for the next k steps in an end-to-end manner, such as forward, left, right, and stop action information at a time.
[0203] Among them, the 3D-aware visual encoder adopts the CLIP model based on ViT and expands it into a 3D-aware encoder in NaVid-4D. The encoder is responsible for learning the visual representation of RGB and depth information, and is initialized from the weights of the CLIP model pre-trained on large-scale RGB images. In view of the different characteristics of RGB and depth information, in this application, different network parameters are used in the shallow layer of the visual encoder to encode RGB and depth respectively, which are called image encoding layers (RGB Layers) and depth encoding layers (Depth Layers). These two models share the same architecture and weight initialization with the RGB-based CLIP model, but they each develop with the training strategy. After these layers, RGB and depth information have been represented in a relatively aligned visual space. Then, they can be further encoded with a shared encoding layer to obtain a set of visual tags. These tags encode visual information in a complementary way, while RGB and depth tags co-encode texture and geometric features. Compress the visual tags to encode historical frames. Experimentally, for each time step, each history frame adopts 4 content tokens and 1 context token, while the current frame adopts 512 content tokens, 2 context tokens (including 257 RGB tokens and 257 depth tokens).
[0204] Among them, the alignment projection module adopts a Q-Former based project to further transform the visual markers into a more general space, making it compatible with the processing of the pre-trained LLM.
[0205] Among them, the action decoder in NaVid-4D is extended from the open source LLM (i.e., Vicuna-7B) by adding a set of action tags to its vocabulary dictionary. These tags represent different low-level operations corresponding to "forward", "turn left", "turn right", and "stop". Unlike NaVi, which predicts the action type and corresponding parameters of the next step at the same time each time, NaVid-4D chooses to predict the next k steps (time steps, i.e., k moments) of discrete actions at each time of reasoning. Experiments have observed that this action modeling is comparable to the performance of NaVid, but provides better sampling efficiency. Among them, special tokens can be used to organize multiple model tokens, such as Fig.16 As shown, to facilitate training.
[0206] The training process of NaVid-4D consists of four different stages: 1) alignment projection module warm-up, 2) visual encoder training, 3) instruction adjustment, and 4) DAgger improvement. In the first training stage, the projector is warmed up while keeping the rest of the network frozen, considering that both the visual encoder and the LLM-based action decoder have been pre-trained on a large scale. In this stage, all layers of the visual encoder are shared through RGB and depth information. When entering the second training stage, the first twenty layers of the visual encoder are trained as deep encoding layers, as well as subsequent shared encoder layers, while keeping the shallow encoding layers of the RGB and action decoders still frozen. In the third stage, the action decoder is further unfrozen to enhance the model output for instruction alignment. In the first three stages, constructed training data consisting of navigation policy data and semantic understanding data is adopted, and in the last stage, the navigation policy data is also enhanced using the DAgger algorithm.
[0207] For the training dataset for model training, 320K visual language action samples were collected from the simulated environment Habitat to learn navigation strategies, and 10K navigation command reasoning samples were generated as NaVid to improve command understanding capabilities. In addition, a variety of different question-and-answer samples and text samples were used to give NaVid-4D a general multi-modal LLM function. In addition, 20K deep map understanding samples and 7.5K robot scene understanding samples from SpatialBot were integrated to further enhance the model's spatial understanding and reasoning capabilities. For example, these data can be mixed into a total of 1.84 million training samples for joint training, where all visual information is represented in RGB-D format.
[0208] In an embodiment of the present application, depth information is more readily available than other 3D representations, making it more feasible to meet the data volume and diversity requirements of the VLA model. In a simulation environment, depth information with ground-truth accuracy can be obtained, but this is almost impossible in the real world because the depth data collected in the real world inevitably contains hardware-limited noise. Due to the difference in depth distribution between simulated and real-world environments, previous work that directly used ground-truth depth in the simulator encountered a significant simulation-to-real gap. In addition, the Internet data used for joint training did not initially include depth information. This shows that using collected depth data can exploit depth with high accuracy in simulation, but will introduce distribution differences between simulation, real-world, and Internet data. This will have a negative impact on the performance of joint training using data from different sources. This application conducts extensive experiments to compare the two methods of obtaining depth, namely capture and estimation, in terms of final end-to-end performance, and further studies the best practices for injecting depth information into existing visual base models.
[0209] In an embodiment of the present application, NaVid-4D can be trained on 4 NVIDIA H800 graphics processors (Graphics Processing Unit, GPU) for 3 days in the first two training stages, and can be trained with 8 NVIDIAH800 GPUs for 2.5 days in the last two stages. In order to improve the obstacle avoidance performance, the A* algorithm can be used to refine the training trajectory, wherein the distance between the robot and the obstacle is considered in the cost function. Following NaVid, non-navigation video data is sampled at 1 frame per second (Frames Per Second, FPS), and all frames are reserved for navigation data. At each time step, the agent's action granularity is set to move forward 25 cm, rotate left / right 15 or decide to stop.
[0210] In an embodiment of the present application, this granularity ensures smooth and continuous motion in real-world experiments. Q-Former, BERT, and Vicuna-7B are initialized using the default pre-trained weights, and all layers of the visual encoder are initialized using the pre-trained weights of EVA-CLIP. In addition, the MLPlayers of the projector are initialized using the corresponding weights of the LLaMA VID trained in the first stage. For evaluation, ROS2 is used to transfer the observed images to a computer equipped with an NVIDIA GeForceRTX 3090, and regular expression matching is used to extract actions from the output of the model. In real-world experiments, Metric3Dv2 takes 1.3 seconds to estimate the depth, and the model of the present application takes 2.9 seconds to generate 4 actions at each time step.
[0211] In the examples of this application, experiments were conducted on NaVid-4D to study and evaluate the following basic questions:
[0212] (1) How does NaVid-4D perform against state-of-the-art VLN models?
[0213] (2) How beneficial is it for the VLN task to explicitly use depth information?
[0214] (3) What is the correct way to inject deep features into existing vision-based models?
[0215] (4) What impact do different depth acquisition methods have on the performance of NaVid-4D?
[0216] (5) Is it effective to predict multiple steps at each inference time to improve performance?
[0217] (6) Is NaVid-4D suitable for actual deployment?
[0218] Among them, for simulation experiments, experiments were conducted on the VLN CE benchmark, which provides 16,844 path instruction pairs in 90 visually realistic scenes in the Matterport3D dataset. For fair comparison, all methods are trained on a training split containing 10,819 R2R samples and evaluated on a test split of 1,839 R2R value samples. In order to reduce the comparison cost, only DAgger is used as shown in Table 1, and only SpatialQA data is used in Tables 1 and 5. Regarding the evaluation, the standard VLN evaluation metrics are followed to report the results, including success rate (SR), prediction success rate (OS), success rate weighted by path length (SPL), trajectory length (TL), and target navigation error (NE). If a stop decision is made within 3 meters of the target location, the event is considered successful.
[0219] Table 1
[0220] TL NE↓ OS↑ SR↑ SPL↑ Seq2Seq 9.30 7.77 37.0 25.0 22.0 CMA 8.64 7.37 40.0 32.0 30.0 WS-MGMap 6.28 9.55 10.8 5.00 4.43 NaVid 7.63 5.47 49.1 37.4 35.9 NaVid-4D 9.32 3.85 68.1 57.8 53.0
[0221] in, Fig.17 This is a schematic diagram of the evaluation of NaVid-4D proposed in the embodiment of the present application, as shown in Fig.17 As shown in Figure 1, for real-world experiments, in order to evaluate in a realistic environment, a Hexman Echo Plus was used as the robot base, equipped with a Kinect DK camera to capture RGBimes. To demonstrate the advantage of using depth information in normal navigation tasks, the following experiments were conducted in different realistic environments. In the following experiments, a STOP decision was considered successful if it occurred within 1.5 meters of the target location. Figure 11-Figure 15 The five reasoning task categories shown were experimented with to assess spatial reasoning ability, and an episode was considered successful if the goal was achieved.
[0222] In the embodiments of this application, problems (1) and (2) are studied by quantitatively comparing NaVid-4D with representative state-of-the-art baselines based on RGB and RGB-D as follows: Seq2Seq: Adopting a recurrent strategy to predict actions directly from RGB-D observations. CMA: Exploiting cross-modal attention between instructions and RGB-D observations for prediction. WS-MGMap: Exploiting multi-granularity graphs that combine object geometry, texture, and semantic information. NaVid: A video-based large visual language navigation model using only RGB input. The results in Table 1 show that NaVid-4D performs significantly better than all baselines in simulation, demonstrating its strong capabilities in this task. Compared with the previous NaVid, NaVid-4D outperforms its overall evaluation metrics, especially SR is improved by 20.4%. This clearly demonstrates the advantage of explicitly using depth information in the VLN model, which makes it easier to correctly find landmarks and follow instructions.
[0223] In embodiments of the present application, for problems (3), (4), and (5), ablation studies can be performed to compare different alternatives for model design and configuration.
[0224] Among them, for problem (3), different depth encoding strategies are compared in Table 2. In the experiments, using MLP layers to encode flat depth patches resulted in performance that was significantly worse than "with ViT". This result suggests that the ViT weights pre-trained on RGB can serve as a reasonable initialization for the deep encoder to promote alignment between RGB and depth information. Considering that RGB and depth are related but have different characteristics, non-shared weights can be adopted in the shallow layers of ViT to learn to present their features in a unified space. Therefore, in Table 3, an ablation study on the number of shared layers, that is, the location where the depth features are injected, is further performed. The results show that the best choice seems to share the first 20 layers of ViT, which means that the depth information should be injected at the midpoint of ViT.
[0225] Table 2
[0226] TL NE↓ OS↑ SR↑ SPL↑ w / o ViT 9.65 4.88 58.1 44.2 39.6 with ViT 9.81 4.62 60.9 47.3 41.9 Ours 8.84 4.32 61.0 51.4 47.6
[0227] Table 3
[0228] Num.of shared layers TL NE↓ OS↑ SR↑ SPL↑ 0(non-shared ViT) 9.81 4.62 60.9 47.3 41.9 10 9.55 4.34 63.9 51.0 45.8 20 (Recommended) 8.84 4.32 61.0 51.4 47.6 30 9.14 4.40 61.2 50.2 45.8
[0229] Among them, for problem (4), the benefits of explicitly using depth are rigorously evaluated and different methods for obtaining depth, i.e., captured vs. estimated, are compared in Table 4. The results significantly demonstrate the necessity of explicitly using thickness. Moreover, even if the depth captured in the simulator has ground truth accuracy, using estimated depth for training is disadvantageous. This is because the joint training data only includes estimated depth, and thus the difference between the depth captured in the navigation data and the depth estimated in the joint training data can negatively impact the performance. Furthermore, when using estimated depth during training, it is observed that estimated and captured depth have comparable performance in terms of evaluation performance. For real-world applications, using estimated values is also better because real-world captured depth images can be very noisy, especially when encountering transparent objects, and estimated depth can avoid these problems. Therefore, estimated depth is used in both simulated and real-world scenarios. In addition, ablation experiments are performed on the impact of using SpatialQA[6] data in Table 5. The results show that SpatialQA[6] data improves the performance of R2R by helping the model better extract and interpret depth information and its depth-related data.
[0230] Table 4
[0231] Train Test TL NE↓ OS↑ SR↑ SPL↑ w / o w / o 9.68 4.72 59.0 45.9 41.1 captured captured 9.03 4.34 60.5 51.1 46.5 captured estimated 8.84 4.32 61.0 51.4 47.6 estimated captured 9.28 4.36 63.5 51.7 47.0 estimated estimated 8.84 4.32 61.0 51.4 47.6
[0232] Table 5
[0233] TL NE↓ OS↑ SR↑ SPL↑ Ours 8.84 4.32 61.0 51.4 47.6 Ours (with SpatialQA) 9.25 4.16 64.4 53.8 49.3
[0234] Among them, for problem (5), the number of steps of the action chunk is ablated in Table 6. The results show that predicting 4 steps in each reasoning produces the highest performance. This is because predicting multiple steps makes the model learn knowledge for a longer time, while predicting too many steps makes the model more difficult to handle. In addition, chunking multiple steps can reduce the cost of reasoning, making 4-step action chunking a good balance between performance and efficiency.
[0235] Table 6
[0236] Num.of steps TL NE↓ OS↑ SR↑ SPL↑ 1 9.87 4.88 60.0 44.1 38.7 2 9.80 4.50 63.5 48.9 43.5 3(Recommended) 8.84 4.32 61.0 51.4 47.6 4 8.95 4.34 59.8 50.5 46.5
[0237] In the embodiments of this application, NaVid-4D is compared with the SOTA baseline NaVid on real-world tasks, including instruction tracking evaluation and spatial intelligence evaluation. For instruction tracking evaluation, Fig.18 This is a schematic diagram of the experimental evaluation of NaVid-4D proposed in the embodiment of the present application, as shown in Fig.18 As shown in the figure, the experimental results show that the proposed model is comparable to NaVid in real-world instruction tracking tasks. In addition, a case study of five types of spatial intelligence capabilities is conducted, and the results are shown in Fig.18As shown, the research results show that, with the support of estimated depth information, the performance of our model is significantly better than NaVid in these five task categories.
[0238] That said, experimental results show improvements in both simulated and real-world environments for standard navigation tasks as well as tasks requiring spatial reasoning. Currently, NaVid-4D still faces challenges in long-range tasks due to the lack of historical depth information input and the high GPU memory required to train ViT.
[0239] The embodiment of the present application provides a navigation method based on a visual language action model and a training method for a visual language action model. On the one hand, based on the received natural language navigation instruction information, combined with video data including image information and depth information, the corresponding action information is obtained through the visual language action model, and then action navigation is performed according to the action information. Among them, based on the model architecture and phased training strategy of the visual language action model, image features and depth features can be extracted respectively through the image coding layer and the depth coding layer in the visual encoder, so that in the process of reasoning action information, one-to-one corresponding image information and depth information can be introduced at the same time, solving the problems of insufficient scene understanding and poor performance, thereby reducing and improving spatial reasoning ability and navigation performance, and improving the accuracy of action navigation. On the other hand, in the process of model training, the visual encoder, alignment projection module and action decoder in the visual language action model are respectively trained in stages according to the training data set including natural language instruction information, image information, depth information and actual action information, and the depth information is injected into the model to obtain a model including a deep feature generation channel, so that the trained visual language action model has a good scene understanding ability, and then when the visual language action model is used for action navigation in the future, it can obtain ideal spatial reasoning ability and navigation performance.
[0240] Based on the above embodiment, in another embodiment of the present application, Fig.19 This is a schematic diagram of the structure of a navigation device based on a visual language action model proposed in an embodiment of the present application, such as Fig.19 As shown, the navigation device 110 based on the visual language action model proposed in the embodiment of the present application may include a first acquisition unit 1101, a navigation unit 1102,
[0241] The acquisition unit 1101 is used to acquire video data of a first time series when receiving natural language navigation instruction information; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information;
[0242] The acquisition unit 1101 is further used to obtain the action information of the second time series based on the natural language navigation instruction information, the image information of the first time series and the depth information of the first time series through the visual language action model; wherein the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer; the visual language action model is obtained by phased training based on the training data set;
[0243] The navigation unit 1102 is configured to perform action navigation based on the action information of the second time series.
[0244] Furthermore, in the embodiments of the present application, Fig. 20 A schematic diagram of the structure of the visual language action model training device proposed in the embodiment of the present application is shown in FIG. Fig. 20 As shown, the visual language action model training device 120 proposed in the embodiment of the present application may include a training unit 1201,
[0245] A training unit 1201 is used for training the initial model in stages based on the training data set to obtain a visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer;
[0246] The training unit 1201 is further used to fix the parameters of the visual encoder and the parameters of the action decoder, train the alignment projection module in the initial model based on the training data set, and obtain a model that completes the first stage of training; based on the model that completes the first stage of training, fix the parameters of the image coding layer in the visual encoder and the parameters of the action decoder, and train the depth coding layer, the shared coding layer, and the alignment projection module based on the training data set to obtain a model that completes the second stage of training; based on the model that completes the second stage of training, fix the parameters of the image coding layer in the visual encoder, and train the depth encoder, the shared coding layer, the alignment projection module, and the action decoder based on the training data set to obtain a visual language action model;
[0247] The training data set includes natural language instruction information, video data and actual action information; the video data includes image information of a third time series and depth information of a third time series, and the actual action information includes actual action information of a fourth time series.
[0248] In the embodiments of the present application, further, Fig.21 This is a schematic diagram of the structure of the computer device proposed in the embodiment of the present application, such as Fig.21As shown, the computer device 130 proposed in the embodiment of the present application includes a processor 1301 and a memory 1302 storing executable instructions of the processor 1301. Furthermore, the computer device 130 may also include a communication interface 1303 and a bus 1304 for connecting the processor 1301, the memory 1302 and the communication interface 1303.
[0249] In the embodiment of the present application, the above-mentioned processor 1301 can be at least one of an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a digital signal processor (Digital Signal Processor, DSP), a digital signal processing device (Digital Signal Processing Device, DSPD), a programmable logic device (ProgRAMmable Logic Device, PLD), a field programmable gate array (Field ProgRAMmable Gate Array, FPGA), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic device used to implement the above-mentioned processor function can also be other, and the embodiment of the present application is not specifically limited. The computer device 130 can also include a memory 1302, which can be connected to the processor 1301, wherein the memory 1302 is used to store executable program code, the program code includes computer operation instructions, and the memory 1302 may include a high-speed RAM memory, and may also include a non-volatile memory, for example, at least two disk memories.
[0250] In the embodiment of the present application, the bus 1304 is used to connect the communication interface 1303, the processor 1301 and the memory 1302, as well as the mutual communication between these devices.
[0251] In the embodiment of the present application, the memory 1302 is used to store instructions and data.
[0252] Furthermore, in an embodiment of the present application, the processor 1301 is configured to:
[0253] When receiving the natural language navigation instruction information, obtaining video data of a first time series; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information;
[0254] The action information of the second time series is obtained through a visual language action model based on natural language navigation instruction information, image information of the first time series and depth information of the first time series; wherein the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer; the visual language action model is obtained by phased training based on a training data set;
[0255] Action navigation is performed based on the action information of the second time series.
[0256] Furthermore, in an embodiment of the present application, the processor 1301 is further configured to:
[0257] The initial model is trained in stages based on the training data set to obtain a visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer;
[0258] The initial model is trained in stages based on the training data set to obtain a visual language action model, including:
[0259] The parameters of the visual encoder and the action decoder are fixed, and the alignment projection module in the initial model is trained based on the training data set to obtain a model that completes the first stage of training;
[0260] Based on the model that has completed the first stage of training, the parameters of the image coding layer and the parameters of the action decoder in the visual encoder are fixed, and the depth coding layer, shared coding layer and alignment projection module are trained based on the training data set to obtain the model that has completed the second stage of training;
[0261] Based on the model that has completed the second stage of training, the parameters of the image encoding layer in the visual encoder are fixed, and the deep encoder, shared encoding layer, alignment projection module and action decoder are trained based on the training data set to obtain a visual language action model;
[0262] Among them, the training data set includes natural language command information, video data and actual action information; among them, the video data includes image information of a third time series and depth information of the third time series corresponding to the image information, and the actual action information includes actual action information of a fourth time series.
[0263] In practical applications, the memory 1302 may be a volatile memory, such as a random access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above types of memories, and provide instructions and data to the processor 1301.
[0264] In addition, each functional module in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or software functional modules.
[0265] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method of this embodiment. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk, etc., which can store program code.
[0266] Specifically, the program instructions corresponding to the navigation method based on the visual language action model in this embodiment can be stored on a storage medium such as a CD, a hard disk, a USB flash drive, etc. When the program instructions corresponding to the navigation method based on the visual language action model in the storage medium are read or executed by an electronic device, the following steps are included:
[0267] When receiving the natural language navigation instruction information, obtaining video data of a first time series; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information;
[0268] The action information of the second time series is obtained through a visual language action model based on natural language navigation instruction information, image information of the first time series and depth information of the first time series; wherein the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer; the visual language action model is obtained by phased training based on a training data set;
[0269] Action navigation is performed based on the action information of the second time series.
[0270] Specifically, the program instructions corresponding to the training method of a visual language action model in this embodiment can be stored on a storage medium such as a CD, a hard disk, a USB flash drive, etc. When the program instructions corresponding to the training method of a visual language action model in the storage medium are read or executed by an electronic device, the following steps are included:
[0271] The initial model is trained in stages based on the training data set to obtain a visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer;
[0272] The initial model is trained in stages based on the training data set to obtain a visual language action model, including:
[0273] The parameters of the visual encoder and the action decoder are fixed, and the alignment projection module in the initial model is trained based on the training data set to obtain a model that completes the first stage of training;
[0274] Based on the model that has completed the first stage of training, the parameters of the image coding layer and the parameters of the action decoder in the visual encoder are fixed, and the depth coding layer, shared coding layer and alignment projection module are trained based on the training data set to obtain the model that has completed the second stage of training;
[0275] Based on the model that has completed the second stage of training, the parameters of the image encoding layer in the visual encoder are fixed, and the deep encoder, shared encoding layer, alignment projection module and action decoder are trained based on the training data set to obtain a visual language action model;
[0276] Among them, the training data set includes natural language command information, video data and actual action information; among them, the video data includes image information of a third time series and depth information of the third time series corresponding to the image information, and the actual action information includes actual action information of a fourth time series.
[0277] The embodiment of the present application also provides a computer program product.
[0278] In some embodiments, the computer program product may include a computer program or instructions.
[0279] In some embodiments, the computer program product can be applied to the computer device in the embodiments of the present application, and the computer program instructions enable the computer to execute the corresponding processes implemented by the computer device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0280] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0281] The present application is described with reference to implementation flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0282] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which is implemented in the implementation flow diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0283] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing the steps in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0284] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. A navigation method based on a visual language action model, characterized in that: The method comprises: When receiving natural language navigation instruction information, acquiring video data of a first time series; wherein the video data of the first time series includes image information of the first time series and depth information of the first time series corresponding to the image information; The action information of the second time series is obtained through a visual language action model based on the natural language navigation instruction information, the image information of the first time series and the depth information of the first time series; wherein the visual language action model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image encoding layer, a depth encoding layer and a shared encoding layer; the visual language action model is obtained by phased training based on a training data set; Perform action navigation based on the action information of the second time series.
2. The method according to claim 1, characterized in that In the case where the natural language navigation instruction information is received at the tth moment, the first time series includes n+1 moments from the tnth moment to the tth moment, the image information of the first time series includes n+1 image information corresponding to the n+1th moment, and the depth information of the first time series includes n+1 depth information corresponding to the n+1th moment; n is an integer greater than or equal to 0, and t is an integer greater than n; The second time series includes k moments from moment t+1 to moment t+k, where k is an integer greater than or equal to 0.
3. The method according to claim 1 or 2, characterized in that: The obtaining the action information of the second time series based on the natural language navigation instruction information, the image information of the first time series and the depth information of the first time series includes: Obtaining, by the visual encoder, image features and depth features corresponding to the video data of the first time series based on the image information of the first time series and the corresponding depth information of the first time series; Obtaining, by the alignment projection module, an image information token and a depth information token based on the image features, the depth features and the natural language navigation instruction information; The action information of the second time series is obtained through the action decoder based on the image information token, the depth information token and the text information token corresponding to the natural language navigation instruction information.
4. The method according to claim 3, characterized in that The obtaining, by the visual encoder, image features and depth features corresponding to the video data of the first time series based on the image information of the first time series and the depth information of the first time series, comprises: Obtaining encoded image information based on the image information of the first time series through the image encoding layer; Obtaining encoded depth information based on the depth information of the first time series through the depth coding layer; The image feature and the depth feature are obtained based on the encoded image information and the encoded image information through the shared coding layer.
5. The method according to claim 3 or 4, characterized in that: The alignment projection module includes a modality alignment module and a projection module, and the alignment projection module is used to obtain an image information token and a depth information token based on the image feature, the depth feature and the natural language navigation instruction information, including: Obtaining, by the modality alignment module, aligned image features and aligned depth features based on the image features, the depth features and the natural language navigation instructions; The image information token and the depth information token are obtained based on the aligned image features and the aligned depth features through the projection module.
6. The method according to claim 2, characterized in that The step of acquiring video data of a first time series includes: Acquire the image information at the tth moment through the shooting module; acquire the depth information at the tth moment through the depth detection module; acquire the n image information and n depth information stored from the tnth moment to the t-1th moment; or, The image information at the tth moment is acquired through the shooting module, and depth estimation is performed based on the image information at the tth moment to determine the depth information at the tth moment; and the n image information and n depth information stored from the tnth moment to the t-1th moment are acquired.
7. A method for training a visual language action model, characterized in that: The method comprises: The initial model is trained in stages based on the training data set to obtain a visual language action model; wherein the initial model includes a visual encoder, an alignment projection module and an action decoder, and the visual encoder includes an image coding layer, a depth coding layer and a shared coding layer; The initial model is trained in stages based on the training data set to obtain a visual language action model, including: Fixing the parameters of the visual encoder and the parameters of the action decoder, and training the alignment projection module in the initial model based on the training data set to obtain a model that completes the first stage of training; Based on the model that has completed the first stage of training, the parameters of the image coding layer and the parameters of the action decoder in the visual encoder are fixed, and the depth coding layer, the shared coding layer and the alignment projection module are trained based on the training data set to obtain a model that has completed the second stage of training; Based on the model that has completed the second stage of training, the parameters of the image coding layer in the visual encoder are fixed, and the depth encoder, the shared coding layer, the alignment projection module and the action decoder are trained based on the training data set to obtain the visual language action model; Wherein, the training data set includes natural language command information, video data and actual action information; wherein, the video data includes image information of a third time series and depth information of the third time series corresponding to the image information, and the actual action information includes actual action information of a fourth time series.
8. The method according to claim 7, characterized in that The fixing of the parameters of the visual encoder and the parameters of the action decoder, training the projection module in the initial model based on the training data set, and obtaining a model that completes the first stage of training, includes: Obtaining, by means of the initial model, first predicted action information of the fourth time series based on first instruction information and first video data in the training data set; wherein the first video data includes first image information of the third time series and first depth information of the third time series, and the training data set also includes first actual action information of the fourth time series corresponding to the first instruction information; Fixing the parameters of the visual encoder and the parameters of the action decoder, updating the parameters of the alignment projection module based on the first predicted action information of the fourth time series and the first actual action information of the fourth time series, and obtaining the model that completes the first stage of training; The method of fixing the parameters of the image coding layer and the action decoder in the visual encoder based on the model trained in the first stage, training the depth coding layer, the shared coding layer and the alignment projection module based on the training data set to obtain the model trained in the second stage includes: Obtaining, by the model that has completed the first stage of training, second predicted action information of the fourth time series based on the second instruction information and the second video data in the training data set; wherein the second video data includes the second image information of the third time series and the second depth information of the third time series, and the training data set also includes the second actual action information of the fourth time series corresponding to the second instruction information; Fixing the parameters of the image coding layer and the parameters of the action decoder in the visual encoder, updating the parameters of the depth coding layer, the parameters of the shared coding layer and the parameters of the alignment projection module based on the second predicted action information of the fourth time series and the second actual action information of the fourth time series, and obtaining the model that completes the second stage of training; The method of fixing the parameters of the image coding layer in the visual encoder based on the model that has completed the second stage of training, training the depth encoder, the shared coding layer, the alignment projection module and the action decoder based on the training data set to obtain the visual language action model includes: Obtaining, by means of the model that has completed the second-stage training, third predicted action information of the fourth time series based on the third instruction information and the third video data in the training data set; wherein the third video data includes third image information of the third time series and third depth information of the third time series, and the training data set also includes third actual action information of the fourth time series corresponding to the third instruction information; Fix the parameters of the image coding layer in the visual encoder, and based on the third predicted action information of the fourth time series and the third actual action information of the fourth time series, update the parameters of the depth encoder, the parameters of the shared coding layer, the parameters of the alignment projection module and the parameters of the action decoder to obtain the visual language action model.
9. A computer device, characterized in that: The computer device comprises: a processor and a memory; wherein, The memory is used to store a computer program that can be run on the processor; The processor is configured to execute the method according to any one of claims 1-6 or 7-8 when running the computer program.
10. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the method as described in any one of claims 1-6 or 7-8 is implemented.
Citation Information
Cited By
Robot navigation method, device and equipment and storage medium
CN121870775A
Robot navigation method, apparatus, device, and storage medium
CN121870775B
Visual language navigation method and system based on three-dimensional geometric perception and shape level representation
CN122392071A