Vision-language and exploration navigation method and system

Through the embodied navigation method of vision-language and exploration, multi-channel depth image video and real environment data are used, combined with paired sample structure strategies and boundary point dynamic selection strategies, the problem of insufficient perception and exploration capabilities of the agent in the dynamic environment is solved, and efficient and flexible navigation and exploration positioning are achieved.

CN119984294AActive Publication Date: 2025-05-13BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE

Patent Information

Application Number
CN202510475481.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing 3D vision-language models lack active perception and exploration capabilities in dynamic and partially observable environments, and embodied agents based on reinforcement learning have insufficient sample efficiency and generalization capabilities.

Method used

A embodied navigation method for vision-language and exploration is proposed. By acquiring multi-channel depth image videos, combining real environment data, using paired sample structure strategies and boundary point dynamic selection strategies, a large-scale trajectory data set is constructed, and online exploration is combined with dynamic spatial memory bank updates to achieve lifelong learning and exploration positioning.

Benefits of technology

It significantly improves exploration accuracy and efficiency, realizes efficient and flexible navigation in unknown real environments, and enhances the overall visual-language positioning and exploration capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119984294A_ABST
    Figure CN119984294A_ABST
Patent Text Reader

Abstract

The invention provides a vision-language and exploration navigation method and system, and the method comprises the steps: obtaining a multi-channel depth image video as trajectory data, combining real environment data, and collecting a vision-language and exploration large-scale trajectory data set according to a matching sample structure strategy and a boundary point dynamic selection strategy; setting a vision-language and exploration training strategy, and reading memorized query from the dynamic space memory bank to obtain a multi-channel depth image sequence; online exploration and dynamic space memory bank updating are combined, vision-language positioning and exploration are connected, three-dimensional world mobile understanding MTU 3D navigation is constructed, and lifelong learning and exploration positioning are achieved; and training the joint integrated optimization framework by using a large-scale trajectory data set, combining expert data with noisy navigation data to form a vision-language and exploration automatic trajectory hybrid model, and performing intelligent reasoning and navigation in a simulation environment and a real scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of spatial memory embodied intelligent navigation and exploration technology, and more specifically, to a vision-language and exploration embodied navigation method and system. Background Art

[0002] Embodied scene understanding requires not only recognition of observed objects, but also active exploration and reasoning in the three-dimensional physical world; multi-channel depth image RGB-D (color images and depth images collected in real time by a deep vision camera combined with information) Human search is driven by the seamless combination of common sense knowledge, spatial reasoning, and visual positioning; similarly, embodied agents must operate in a continuous closed-loop cycle of exploration, perception, reasoning, and action when navigating in a new environment; a key part of this process is understanding three-dimensional vision and language (3D-VL), which enables agents to think spatially and make informed decisions about where to explore; significant progress has been made in the field of 3D-VL in recent years; these models use three-dimensional reconstruction for visual positioning, question answering, dense annotation, and situational reasoning; recent methods, such as 3DVLP, PQ3D, and LEO, aim to handle multiple tasks in a single architecture through pre-training or joint unified training; however, existing 3D-VL models rely on static three-dimensional representations, assuming that a complete reconstruction of the environment is available in advance; although this is useful in offline visual-language localization, it is not feasible to use 3D-VL models in a single architecture. In the past, 3D-VL models have been proposed to be effective in positioning, but in the real world, embodied agents operate in partially observable and dynamic environments, and this assumption is not practical. In addition, these models usually lack active perception and exploration capabilities. In contrast, although embodied agents based on reinforcement learning can explore the environment, they often face problems such as low sample efficiency, poor generalization ability due to limited training data, and lack of clear spatial representation. Therefore, combining passive 3D-VL positioning with active exploration remains a key challenge in developing intelligent systems that can efficiently explore and understand the three-dimensional world. In the existing technical paradigm, although RL (reinforcement learning) can learn online and achieve exploration, it cannot achieve grounding and lifelong learning. Although 3D-VL can learn and achieve positioning throughout life, problems such as the inability to explore online remain to be solved. Therefore, it is necessary to propose a vision-language and exploration embodied navigation method and system to at least partially solve the problems existing in the existing technology. Summary of the invention

[0003] A series of simplified concepts are introduced in the content of the invention, which will be further described in detail in the specific implementation method section; the content of the invention of the present invention does not mean to attempt to limit the key features and essential technical features of the technical solution claimed for protection, nor does it mean to attempt to determine the scope of protection of the technical solution claimed for protection.

[0004] To at least partially solve the above problems, the present invention provides a visual-language and exploratory embodied navigation method, comprising: S10, obtain multi-channel depth image video as trajectory data, combine it with real environment data, and collect vision-language and explore large-scale trajectory data sets based on paired sample structure strategy and boundary point dynamic selection strategy; S20, setting the vision-language and exploration training strategy, reading the memory query from the dynamic spatial memory bank, and obtaining a multi-channel depth image sequence; S30, combines online exploration with dynamic spatial memory updates, connects visual-linguistic positioning and exploration, builds a three-dimensional world mobile understanding MTU 3D navigation, and realizes lifelong learning and exploration positioning; S40, builds a joint unified optimization framework, uses a large-scale trajectory dataset to train the joint unified optimization framework, combines expert data with noisy navigation data, forms a hybrid model of vision-language and exploration automatic trajectory, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios.

[0005] Preferably, S10 includes: S101, obtaining offline multi-channel deep image video in a multi-channel deep image video platform as trajectory data; according to the object query and the language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, using the paired samples for pre-training; collecting visual-language localization trajectory data; S102, combining visual-language positioning trajectory data, according to the exploration data structure, based on object query, boundary point query and target association to generate decisions, and the boundary points change dynamically during the exploration process, to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; setting a list of visited boundary points, and exploring only when there are better potential boundary points closer to the target, otherwise triggering an exception to prevent unnecessary exploration; collecting large-scale visual-language and exploration trajectory data sets; significantly improving exploration accuracy and efficiency.

[0006] Preferably, S20 includes: S201, setting the vision-language and exploration training strategy, and reading the memorized query from the dynamic spatial memory bank; S202, combining the language of the inquiry, through the trajectory planner, to conduct exploration, feedback to the world, and obtain a multi-channel depth image sequence.

[0007] Preferably, S30 includes: S301, taking the original multi-channel depth image frame as input, generating a single-frame local query and storing it in the global spatial memory, and performing dynamic spatial memory update; using the feature extraction and segmentation prior of the 2D basic model, online query representation learning, and generating a semantic spatial information query table that integrates semantic information and precise 3D spatial information; S302, by introducing a joint optimization framework, synchronous learning of object positioning and exploration is achieved; during the training process, the semantic spatial information query table is combined, and object queries retrieved from the spatial memory library and the boundaries of the explored map are processed as boundary queries at the same time, thereby achieving integrated end-to-end training of the two tasks and jointly exploring the positioning target; S303, by processing the multi-channel depth image sequence, setting a trajectory planner to guide navigation, and promoting the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception and action cycle to perform efficient and flexible navigation in an unknown real environment.

[0008] Preferably, S40 includes: S401, build a joint unified optimization framework and train it using a large-scale trajectory dataset; combine expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance training diversity; train the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; S402, based on the hybrid model of vision-language and exploration automatic trajectory, seamlessly migrates to simulated environments and real-world scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.

[0009] The present invention provides a visual-language and exploratory embodied navigation system, comprising: The trajectory boundary point dynamic selection subsystem obtains multi-channel depth image videos as trajectory data, combines real environment data, and collects vision-language and exploration large-scale trajectory data sets based on the paired sample structure strategy and boundary point dynamic selection strategy; The spatial memory query integration subsystem sets the visual-language and exploration training strategies, reads the memory query from the dynamic spatial memory library, and obtains a multi-channel depth image sequence; The 3D world mobile understanding navigation subsystem combines online exploration with dynamic spatial memory updates, connects visual-language positioning and exploration, builds 3D world mobile understanding MTU 3D navigation, and realizes lifelong learning and exploration positioning; The automatic trajectory hybrid reasoning subsystem builds a joint unified optimization framework, uses a large-scale trajectory dataset to train the joint unified optimization framework, combines expert data with noisy navigation data, forms a vision-language and exploration automatic trajectory hybrid model, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios.

[0010] Preferably, the trajectory boundary point dynamic selection subsystem includes: The visual-language trajectory collection subsystem obtains offline multi-channel deep image videos in the multi-channel deep image video platform as trajectory data; according to the object query and language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, the paired samples are used for pre-training; and the visual-language positioning trajectory data is collected; The dynamic strategy exploration trajectory collection subsystem combines visual-language positioning trajectory data, generates decisions based on object queries, boundary point queries and target associations according to the exploration data structure, and the boundary points change dynamically during the exploration process to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; set a list of visited boundary points, and explore only when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration; collect visual-language and exploration large-scale trajectory data sets; significantly improve exploration accuracy and efficiency.

[0011] Preferably, the spatial memory inquiry is combined with a subsystem, including: The dynamic spatial memory query subsystem sets the visual-language and exploration training strategies and reads the memory query from the dynamic spatial memory bank; The trajectory planning query feedback subsystem combines the query language with the trajectory planner to conduct exploration and feedback to the world to obtain a multi-channel depth image sequence.

[0012] Preferably, the three-dimensional world mobile understanding navigation subsystem includes: The spatial memory update query subsystem takes the original multi-channel depth image frame as input, generates a single-frame local query and stores it in the global spatial memory, and performs dynamic spatial memory update. With the help of feature extraction and segmentation priors of the 2D basic model, online query representation learning is performed to generate a semantic spatial information query table that integrates semantic information and precise 3D spatial information. The joint optimization framework synchronization subsystem realizes the synchronous learning of object positioning and exploration by introducing the joint optimization framework. During the training process, it combines the semantic spatial information query table and simultaneously processes the object query retrieved from the spatial memory library and the boundary of the explored map as the boundary query, thereby realizing the integrated end-to-end training of the two tasks and jointly exploring the positioning target. The guided navigation perception-action loop subsystem processes multi-channel depth image sequences, sets a trajectory planner to guide navigation, and promotes the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception-action loop for efficient and flexible navigation in unknown real environments.

[0013] Preferably, the automatic trajectory hybrid reasoning subsystem includes: The automatic trajectory hybrid model subsystem builds a joint unified optimization framework and trains the joint unified optimization framework using a large-scale trajectory dataset; combines expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance training diversity; trains the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; The hybrid model embodied navigation subsystem, based on the vision-language and exploration automatic trajectory hybrid model, seamlessly migrates to simulated environments and real-world scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.

[0014] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention provides a visual-language and exploration embodied navigation method and system, which obtains multi-channel depth image videos as trajectory data, combines real environment data, collects visual-language and exploration large-scale trajectory data sets according to paired sample structure strategy and boundary point dynamic selection strategy; sets visual-language and exploration training strategy, reads memorized queries from dynamic spatial memory library, and obtains multi-channel depth image sequence; combines online exploration with dynamic spatial memory library update, connects visual-language positioning and exploration, and constructs a three-dimensional world mobile understanding MTU 3D navigation, realizing lifelong learning and exploration positioning; constructing a joint unified optimization framework, using a large-scale trajectory data set to train the joint unified optimization framework, combining expert data with noisy navigation data, forming a visual-language and exploration automatic trajectory hybrid model, and performing intelligent reasoning and embodied navigation in simulated environments and real scenes; it can significantly improve the accuracy of tunnel lining crack detection; a joint unified framework that can simultaneously optimize visual-language positioning tasks and exploration tasks is proposed, thereby utilizing the complementary characteristics of the two to enhance the overall performance; a new visual-language and exploration training strategy is proposed, and embodied navigation pre-training is performed through a large-scale trajectory data set collected from real data and simulated data to achieve navigation in unknown real environments; a mobile understanding MTU 3D method (Move to Understand in a three-dimensional world) is constructed, connecting visual-language positioning and exploration to achieve efficient and flexible embodied navigation combining visual positioning and exploration in an embodied environment; and the efficiency of generating trajectory data for navigation pre-training is significantly improved; the present invention has important technical significance and significant effects; it has important technical significance and significant effects.

[0015] The present invention describes a visual-language and exploratory embodied navigation method and system. Other advantages, objectives and features of the present invention will be reflected in part through the following description, and in part will be understood by technical personnel in the field through research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a diagram of an embodiment of the visual-language and exploration embodied navigation system architecture described in the present invention.

[0017] Figure 2 This is a comparative example diagram of the prior art of the vision-language and exploration embodied navigation method and system described in the present invention. DETAILED DESCRIPTION

[0018] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments, so that those skilled in the art can implement it according to the instructions; as shown in the figure, the present invention provides a visual-language and exploratory embodied navigation method, including: S10, obtain multi-channel depth image video as trajectory data, combine it with real environment data, and collect vision-language and explore large-scale trajectory data sets based on paired sample structure strategy and boundary point dynamic selection strategy; S20, setting the vision-language and exploration training strategy, reading the memory query from the dynamic spatial memory bank, and obtaining a multi-channel depth image sequence; S30, combines online exploration with dynamic spatial memory updates, connects visual-linguistic positioning and exploration, builds a three-dimensional world mobile understanding MTU 3D navigation, and realizes lifelong learning and exploration positioning; S40, builds a joint unified optimization framework, uses a large-scale trajectory dataset to train the joint unified optimization framework, combines expert data with noisy navigation data, forms a hybrid model of vision-language and exploration automatic trajectory, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios.

[0019] The principle and effect of the above technical solution are as follows: The present invention provides an embodied navigation method of vision-language and exploration, including: obtaining multi-channel depth image videos as trajectory data, combining with real environment data, collecting large-scale trajectory data sets of vision-language and exploration according to paired sample structure strategy and boundary point dynamic selection strategy; setting vision-language and exploration training strategy, reading memorized queries from dynamic spatial memory library, and obtaining multi-channel depth image sequences; combining online exploration with dynamic spatial memory library update, connecting vision-language positioning and exploration, constructing three-dimensional world mobile understanding MTU3D navigation, and realizing lifelong learning and exploration positioning; constructing a joint unified optimization framework, using large-scale The simulated trajectory dataset is used to train a joint unified optimization framework, which combines expert data with noisy navigation data to form a visual-language and exploration automatic trajectory hybrid model, and performs intelligent reasoning and embodied navigation in simulated environments and real scenes; it can significantly improve the accuracy of tunnel lining crack detection; a joint unified framework that can simultaneously optimize visual-language localization tasks and exploration tasks is proposed, thereby utilizing the complementary characteristics of the two to enhance the overall performance; a new visual-language and exploration training strategy is proposed, which performs embodied navigation pre-training through large-scale trajectory datasets collected from real data and simulated data to achieve navigation in unknown real environments; a move to understand MTU 3D method (Move to Understand in a three-dimensional world) is constructed, which connects visual-language localization and exploration to achieve efficient and flexible embodied navigation combining visual localization and exploration in an embodied environment; and the efficiency of trajectory data generation for navigation pre-training is significantly improved; the present invention has important technical significance and significant effects.

[0020] In one embodiment, S10 includes: S101, obtaining offline multi-channel deep image video in a multi-channel deep image video platform as trajectory data; according to the object query and the language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, using the paired samples for pre-training; collecting visual-language localization trajectory data; S102, combining visual-language positioning trajectory data, according to the exploration data structure, based on object query, boundary point query and target association to generate decisions, and the boundary points change dynamically during the exploration process, to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; setting a list of visited boundary points, and exploring only when there are better potential boundary points closer to the target, otherwise triggering an exception to prevent unnecessary exploration; collecting large-scale visual-language and exploration trajectory data sets; significantly improving exploration accuracy and efficiency.

[0021] The principle and effect of the above technical solution are as follows: obtaining offline multi-channel depth image video in a multi-channel depth image video platform as trajectory data; according to the object query and language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, using paired samples for pre-training; collecting visual-language localization trajectory data; combining visual-language localization trajectory data, according to the exploration data structure, according to the object query, boundary point query and target association to generate a decision, and the boundary point changes dynamically during the exploration process, and a boundary point dynamic selection strategy is constructed; the boundary point changes dynamically during the exploration process Dynamic changes include: avoiding overfitting caused by using only the optimal boundary point through a dynamic boundary point selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; using the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment dataset to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; setting a list of visited boundary points, and exploring only when there are better potential boundary points closer to the target, otherwise triggering an exception to prevent unnecessary exploration; collecting large-scale vision-language and exploration trajectory datasets; and significantly improving exploration accuracy and efficiency.

[0022] In one embodiment, S20 includes: S201, setting the vision-language and exploration training strategy, and reading the memorized query from the dynamic spatial memory bank; S202, combining the language of the inquiry, through the trajectory planner, to conduct exploration, feedback to the world, and obtain a multi-channel depth image sequence.

[0023] The principle and effect of the above technical solution are: setting the visual-language and exploration training strategy, reading the memorized query from the dynamic spatial memory bank; combining the language of the query, through the trajectory planner, to explore and feed back into the world, and obtain a multi-channel depth image sequence; Setting up visual-language and exploration training strategies, reading memorized queries from a dynamic spatial memory library, combining the language of the query, exploring through a trajectory planner, feeding back into the world, and obtaining a multi-channel depth image sequence includes: visual-language and exploration training strategies, reading memorized queries from a dynamic spatial memory library, combining the language of the query (for example: a chair near me), obtaining object queries that need to be explored through spatial reasoning, and then obtaining their spatial position information (including xyz and pose), exploring through a trajectory planner (Trajectories Planner), feeding back into the world coordinate system, and obtaining a multi-channel depth image sequence.

[0024] In one embodiment, S30 includes: S301, taking the original multi-channel depth image frame as input, generating a single-frame local query and storing it in the global spatial memory, and performing dynamic spatial memory update; using the feature extraction and segmentation prior of the 2D basic model, online query representation learning, and generating a semantic spatial information query table that integrates semantic information and precise 3D spatial information; S302, by introducing a joint optimization framework, synchronous learning of object positioning and exploration is achieved; during the training process, the semantic spatial information query table is combined, and object queries retrieved from the spatial memory library and the boundaries of the explored map are processed as boundary queries at the same time, thereby achieving integrated end-to-end training of the two tasks and jointly exploring the positioning target; S303, by processing the multi-channel depth image sequence, setting a trajectory planner to guide navigation, and promoting the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception and action cycle to perform efficient and flexible navigation in an unknown real environment.

[0025] The principle and effect of the above technical solution are as follows: taking the original multi-channel depth image frame as input, generating a single-frame local query and storing it in the global spatial memory library, and dynamically updating the spatial memory library; with the help of feature extraction and segmentation priors of the 2D basic model, online query representation learning is performed to generate a semantic spatial information query table that integrates semantic information and precise 3D spatial information; by introducing a joint optimization framework, synchronous learning of object positioning and exploration is achieved; during the training process, the semantic spatial information query table is combined, and object queries retrieved from the spatial memory library and the boundaries of the explored map are processed as boundary queries at the same time, thereby achieving integrated end-to-end training of the two tasks and jointly exploring and locating the target; by processing a multi-channel depth image sequence, a trajectory planner is set to guide navigation, and the acquisition of a new multi-channel depth image sequence is promoted, thereby forming a continuous perception and action cycle for efficient and flexible navigation in an unknown real environment; Multi-channel depth images include color images and depth image information collected by the camera in real time; 2D basic models include: DINO (Distillation and NO labels) or SAM (Segment Anything Model); The joint unified optimization framework includes: combining the two previously different task training objectives of visual-language localization and exploration, integrating and co-optimizing the visual-language localization task and the exploration task, and utilizing the complementary characteristics of the two to enhance the spatial accuracy of the embodied navigation perception of the intelligent body in three-dimensional space and the level of intelligent decision-making and action in exploring the three-dimensional world, thus forming a joint unified optimization framework; By processing a multi-channel depth image sequence, setting a trajectory planner to guide navigation, and promoting the acquisition of new multi-channel depth image sequences, a continuous perception and action cycle is formed, and efficient and flexible navigation in an unknown real environment is performed, including: generating object queries by processing a multi-channel depth image sequence, and then storing these object queries in a dynamic spatial memory library; the spatial reasoning layer selects object queries that need to be explored from object queries and boundary queries, and obtains spatial position information of the selected object queries; setting a trajectory planner to guide navigation; passing the spatial position information of the selected object queries to the trajectory planner; the trajectory planner guides navigation and promotes the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception and action cycle, and efficient and flexible embodied navigation in an unknown real environment is performed.

[0026] like Figure 1As shown in the figure, unified grounding and exploration; partial multi-channel depth image sequences are obtained; online query feature learning is performed; online query feature learning includes: the obtained image and depth information as well as the position information are used to propose object queries through the feature extraction (query proposal) module of the 2D encoder and the 3D encoder, and the frontier queries can be obtained from the frontier mapping based on the position information; lower right corner: dynamic spatial memory bank, object queries and boundary queries are written through memory write; a joint unified optimization framework is constructed.

[0027] In one embodiment, S40 includes: S401, build a joint unified optimization framework and train it using a large-scale trajectory dataset; combine expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance training diversity; train the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; S402, based on the hybrid model of vision-language and exploration automatic trajectory, seamlessly migrates to simulated environments and real-world scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.

[0028] The principles and effects of the above technical solution are: constructing a joint unified optimization framework and using a large-scale trajectory dataset to train the joint unified optimization framework; combining expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance the diversity of training; training the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; based on the vision-language and exploration automatic trajectory hybrid model, seamlessly migrate to simulated environments and real-life scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.

[0029] The present invention provides a visual-language and exploratory embodied navigation system, comprising: The trajectory boundary point dynamic selection subsystem obtains multi-channel depth image videos as trajectory data, combines real environment data, and collects vision-language and exploration large-scale trajectory data sets based on the paired sample structure strategy and boundary point dynamic selection strategy; The spatial memory query integration subsystem sets the visual-language and exploration training strategies, reads the memory query from the dynamic spatial memory library, and obtains a multi-channel depth image sequence; The 3D world mobile understanding navigation subsystem combines online exploration with dynamic spatial memory updates, connects visual-language positioning and exploration, builds 3D world mobile understanding MTU 3D navigation, and realizes lifelong learning and exploration positioning; The automatic trajectory hybrid reasoning subsystem builds a joint unified optimization framework, uses a large-scale trajectory dataset to train the joint unified optimization framework, combines expert data with noisy navigation data, forms a vision-language and exploration automatic trajectory hybrid model, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios.

[0030] The principle and effect of the above technical solution are as follows: The present invention provides an embodied navigation system of vision-language and exploration, including: a trajectory boundary point dynamic selection subsystem, which obtains multi-channel depth image video as trajectory data, combines real environment data, and collects large-scale trajectory data sets of vision-language and exploration according to the paired sample structure strategy and the boundary point dynamic selection strategy; a spatial memory query combination subsystem, which sets the visual-language and exploration training strategy, reads the memory query from the dynamic spatial memory library, and obtains a multi-channel depth image sequence; a three-dimensional world mobile understanding navigation subsystem, which combines online exploration with dynamic spatial memory library updates, connects visual-language positioning and exploration, and constructs a three-dimensional world mobile understanding MTU. 3D navigation, realizing lifelong learning and exploration positioning; automatic trajectory hybrid reasoning subsystem, constructing a joint unified optimization framework, using a large-scale trajectory data set to train the joint unified optimization framework, combining expert data with noisy navigation data, forming a visual-language and exploration automatic trajectory hybrid model, and performing intelligent reasoning and embodied navigation in simulated environments and real scenes; it can significantly improve the accuracy of tunnel lining crack detection; a joint unified framework that can simultaneously optimize visual-language positioning tasks and exploration tasks is proposed, thereby utilizing the complementary characteristics of the two to enhance the overall performance; a new visual-language and exploration training strategy is proposed, and embodied navigation pre-training is performed through a large-scale trajectory data set collected from real data and simulated data to achieve navigation in unknown real environments; a mobile understanding in a three-dimensional world MTU 3D method (Move to Understand) is constructed, connecting visual-language positioning and exploration to achieve efficient and flexible embodied navigation combining visual positioning and exploration in an embodied environment; and the efficiency of generating trajectory data for navigation pre-training is significantly improved; the present invention has important technical significance and significant effects.

[0031] In one embodiment, the trajectory boundary point dynamic selection subsystem includes: The visual-language trajectory collection subsystem obtains offline multi-channel deep image videos in the multi-channel deep image video platform as trajectory data; according to the object query and language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, the paired samples are used for pre-training; and the visual-language positioning trajectory data is collected; The dynamic strategy exploration trajectory collection subsystem combines visual-language positioning trajectory data, generates decisions based on object queries, boundary point queries and target associations according to the exploration data structure, and the boundary points change dynamically during the exploration process to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; set a list of visited boundary points, and explore only when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration; collect visual-language and exploration large-scale trajectory data sets; significantly improve exploration accuracy and efficiency.

[0032] The principle and effect of the above technical solution are as follows: The trajectory boundary point dynamic selection subsystem includes: The visual-language trajectory collection subsystem obtains offline multi-channel deep image videos in the multi-channel deep image video platform as trajectory data; according to the object query and language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, the paired samples are used for pre-training; and the visual-language positioning trajectory data is collected; The dynamic strategy exploration trajectory collection subsystem combines visual-language positioning trajectory data, generates decisions based on object queries, boundary point queries and target associations according to the exploration data structure, and the boundary points change dynamically during the exploration process to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; set a list of visited boundary points, and explore only when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration; collect visual-language and exploration large-scale trajectory data sets; significantly improve exploration accuracy and efficiency.

[0033] In one embodiment, the spatial memory query is combined with a subsystem, including: The dynamic spatial memory query subsystem sets the visual-language and exploration training strategies and reads the memory query from the dynamic spatial memory bank; The trajectory planning query feedback subsystem combines the query language with the trajectory planner to conduct exploration and feedback to the world to obtain a multi-channel depth image sequence.

[0034] The principle and effect of the above technical solution are: a spatial memory query combined subsystem, including: a dynamic spatial memory query subsystem, setting a visual-language and exploration training strategy, and reading a memory query from a dynamic spatial memory library; a trajectory planning query feedback subsystem, combining the query language, through a trajectory planner, to explore, and feedback to the world to obtain a multi-channel depth image sequence; setting a visual-language and exploration training strategy, reading a memory query from a dynamic spatial memory library, combining the query language, through a trajectory planner, to explore, and feedback to the world to obtain a multi-channel depth image sequence including: a visual-language and exploration training strategy, reading a memory query from a dynamic spatial memory library, combining the query language (for example: a chair near me), through spatial reasoning (Spatial Reasoning), obtaining an object query that needs to be explored, and then obtaining its spatial position information (including xyz and pose), exploring through a trajectory planner (Trajectories Planner), and feeding back to the world coordinate system to obtain a multi-channel depth image sequence.

[0035] In one embodiment, a 3D world mobile understanding navigation subsystem includes: The spatial memory update query subsystem takes the original multi-channel depth image frame as input, generates a single-frame local query and stores it in the global spatial memory, and performs dynamic spatial memory update. With the help of feature extraction and segmentation priors of the 2D basic model, online query representation learning is performed to generate a semantic spatial information query table that integrates semantic information and precise 3D spatial information. The joint optimization framework synchronization subsystem realizes the synchronous learning of object positioning and exploration by introducing the joint optimization framework. During the training process, it combines the semantic spatial information query table and simultaneously processes the object query retrieved from the spatial memory library and the boundary of the explored map as the boundary query, thereby realizing the integrated end-to-end training of the two tasks and jointly exploring the positioning target. The guided navigation perception-action loop subsystem processes multi-channel depth image sequences, sets a trajectory planner to guide navigation, and promotes the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception-action loop for efficient and flexible navigation in unknown real environments.

[0036] The principle and effect of the above technical solution are as follows: a three-dimensional world mobile understanding navigation subsystem, including: a spatial memory update query subsystem, which takes the original multi-channel depth image frame as input, generates a single-frame local query and stores it in the global spatial memory, and performs dynamic spatial memory update; with the help of feature extraction and segmentation priors of the 2D basic model, online query representation learning, and generates a semantic spatial information query table that integrates semantic information and precise 3D spatial information; a joint optimization framework synchronization subsystem, which realizes synchronous learning of object positioning and exploration by introducing a joint optimization framework; during the training process, the semantic spatial information query table is combined, and the object query retrieved from the spatial memory and the boundary of the explored map are processed as boundary queries at the same time, thereby realizing integrated end-to-end training of the two tasks, and jointly exploring and locating the target; a guided navigation perception and action loop subsystem, which processes multi-channel depth image sequences, sets a trajectory planner to guide navigation, and promotes the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception and action cycle for efficient and flexible navigation in unknown real environments; Multi-channel depth images include color images and depth image information collected by the camera in real time; 2D basic models include: DINO (Distillation and NO labels) or SAM (Segment Anything Model); The joint unified optimization framework includes: combining the two previously different task training objectives of visual-language localization and exploration, integrating and co-optimizing the visual-language localization task and the exploration task, and utilizing the complementary characteristics of the two to enhance the spatial accuracy of the embodied navigation perception of the intelligent body in three-dimensional space and the level of intelligent decision-making and action in exploring the three-dimensional world, thus forming a joint unified optimization framework; By processing a multi-channel depth image sequence, setting a trajectory planner to guide navigation, and promoting the acquisition of new multi-channel depth image sequences, a continuous perception and action cycle is formed, and efficient and flexible navigation in an unknown real environment is performed, including: generating object queries by processing a multi-channel depth image sequence, and then storing these object queries in a dynamic spatial memory library; the spatial reasoning layer selects object queries that need to be explored from object queries and boundary queries, and obtains spatial position information of the selected object queries; setting a trajectory planner to guide navigation; passing the spatial position information of the selected object queries to the trajectory planner; the trajectory planner guides navigation and promotes the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception and action cycle, and efficient and flexible embodied navigation in an unknown real environment is performed.

[0037] like Figure 1As shown in the figure, unified grounding and exploration; partial multi-channel depth image sequences are obtained; online query feature learning is performed; online query feature learning includes: the obtained image and depth information as well as the position information are used to propose object queries through the feature extraction (query proposal) module of the 2D encoder and the 3D encoder, and the frontier queries can be obtained from the frontier mapping based on the position information; lower right corner: dynamic spatial memory bank, object queries and boundary queries are written through memory write; a joint unified optimization framework is constructed.

[0038] In one embodiment, the automatic trajectory hybrid reasoning subsystem includes: The automatic trajectory hybrid model subsystem builds a joint unified optimization framework and trains the joint unified optimization framework using a large-scale trajectory dataset; combines expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance training diversity; trains the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; The hybrid model embodied navigation subsystem, based on the vision-language and exploration automatic trajectory hybrid model, seamlessly migrates to simulated environments and real-world scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.

[0039] The principles and effects of the above technical solution are as follows: an automatic trajectory hybrid reasoning subsystem, including: an automatic trajectory hybrid model subsystem, which builds a joint unified optimization framework and uses a large-scale trajectory data set to train the joint unified optimization framework; combines expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance the diversity of training; trains the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; a hybrid model embodied navigation subsystem, which seamlessly migrates to simulated environments and real-world scenarios based on the vision-language and exploration automatic trajectory hybrid model to perform intelligent reasoning and embodied navigation in complex environments and noisy data.

[0040] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the illustrations shown and described herein.

Claims

1. A visual-language and exploratory embodied navigation method, characterized in that: include: S10, obtain multi-channel depth image video as trajectory data, combine it with real environment data, and collect vision-language and explore large-scale trajectory data sets based on paired sample structure strategy and boundary point dynamic selection strategy; S20, setting the vision-language and exploration training strategy, reading the memory query from the dynamic spatial memory bank, and obtaining a multi-channel depth image sequence; S30, combines online exploration with dynamic spatial memory updates, connects visual-linguistic positioning and exploration, builds MTU 3D navigation for mobile understanding of the three-dimensional world, and realizes lifelong learning and exploration positioning; S40, builds a joint unified optimization framework, uses a large-scale trajectory dataset to train the joint unified optimization framework, combines expert data with noisy navigation data, forms a hybrid model of vision-language and exploration automatic trajectory, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios.

2. A visual-language and exploratory embodied navigation method according to claim 1, characterized in that: S10 includes: S101, obtaining offline multi-channel deep image video in a multi-channel deep image video platform as trajectory data; according to the object query and the language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, using the paired samples for pre-training; collecting visual-language localization trajectory data; S102, combining visual-language positioning trajectory data, according to the exploration data structure, based on object query, boundary point query and target association to generate decisions, and the boundary points change dynamically during the exploration process, to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; setting a list of visited boundary points, and exploring only when there are better potential boundary points closer to the target, otherwise triggering an exception to prevent unnecessary exploration; collecting large-scale visual-language and exploration trajectory data sets; significantly improving exploration accuracy and efficiency.

3. A visual-language and exploratory embodied navigation method according to claim 1, characterized in that: The S20 includes: S201, setting the vision-language and exploration training strategy, and reading the memorized query from the dynamic spatial memory bank; S202, combining the language of the inquiry, through the trajectory planner, to conduct exploration, feedback to the world, and obtain a multi-channel depth image sequence.

4. The visual-language and exploratory embodied navigation method according to claim 1, characterized in that: S30 includes: S301, taking the original multi-channel depth image frame as input, generating a single-frame local query and storing it in the global spatial memory, and performing dynamic spatial memory update; using the feature extraction and segmentation prior of the 2D basic model, online query representation learning, and generating a semantic spatial information query table that integrates semantic information and precise 3D spatial information; S302, by introducing a joint optimization framework, synchronous learning of object positioning and exploration is achieved; during the training process, the semantic spatial information query table is combined, and object queries retrieved from the spatial memory library and the boundaries of the explored map are processed as boundary queries at the same time, thereby achieving integrated end-to-end training of the two tasks and jointly exploring the positioning target; S303, by processing the multi-channel depth image sequence, setting a trajectory planner to guide navigation, and promoting the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception and action cycle to perform efficient and flexible navigation in an unknown real environment.

5. The visual-language and exploratory embodied navigation method according to claim 1, characterized in that: S40 includes: S401, build a joint unified optimization framework and train it using a large-scale trajectory dataset; combine expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance training diversity; train the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; S402, based on the hybrid model of vision-language and exploration automatic trajectory, seamlessly migrates to simulated environments and real-world scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.

6. A visual-language and exploratory embodied navigation system, characterized in that: include: The trajectory boundary point dynamic selection subsystem obtains multi-channel depth image videos as trajectory data, combines real environment data, and collects vision-language and exploration large-scale trajectory data sets based on the paired sample structure strategy and boundary point dynamic selection strategy; The spatial memory query integration subsystem sets the visual-language and exploration training strategies, reads the memory query from the dynamic spatial memory library, and obtains a multi-channel depth image sequence; The 3D world mobile understanding navigation subsystem combines online exploration with dynamic spatial memory updates, connects visual-language positioning and exploration, builds 3D world mobile understanding MTU 3D navigation, and realizes lifelong learning and exploration positioning; The automatic trajectory hybrid reasoning subsystem builds a joint unified optimization framework, uses a large-scale trajectory dataset to train the joint unified optimization framework, combines expert data with noisy navigation data, forms a vision-language and exploration automatic trajectory hybrid model, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios.

7. A visual-language and exploratory embodied navigation system according to claim 6, characterized in that: The trajectory boundary point dynamic selection subsystem includes: The visual-language trajectory collection subsystem obtains offline multi-channel deep image videos in the multi-channel deep image video platform as trajectory data; according to the object query and language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, the paired samples are used for pre-training; and the visual-language positioning trajectory data is collected; The dynamic strategy exploration trajectory collection subsystem combines visual-language positioning trajectory data, generates decisions based on object queries, boundary point queries and target associations according to the exploration data structure, and the boundary points change dynamically during the exploration process to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; set a list of visited boundary points, and explore only when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration; collect visual-language and exploration large-scale trajectory data sets; significantly improve exploration accuracy and efficiency.

8. The visual-language and exploratory embodied navigation system according to claim 6, characterized in that: Spatial memory inquiry combined with subsystems, including: The dynamic spatial memory query subsystem sets the visual-language and exploration training strategies and reads the memory query from the dynamic spatial memory bank; The trajectory planning query feedback subsystem combines the query language with the trajectory planner to conduct exploration and feedback to the world to obtain a multi-channel depth image sequence.

9. The visual-language and exploratory embodied navigation system according to claim 6, characterized in that: The 3D world mobile understanding and navigation subsystem includes: The spatial memory update query subsystem takes the original multi-channel depth image frame as input, generates a single-frame local query and stores it in the global spatial memory, and performs dynamic spatial memory update. With the help of feature extraction and segmentation priors of the 2D basic model, online query representation learning is performed to generate a semantic spatial information query table that integrates semantic information and precise 3D spatial information. The joint optimization framework synchronization subsystem realizes the synchronous learning of object positioning and exploration by introducing the joint optimization framework. During the training process, it combines the semantic spatial information query table and simultaneously processes the object query retrieved from the spatial memory library and the boundary of the explored map as the boundary query, thereby realizing the integrated end-to-end training of the two tasks and jointly exploring the positioning target. The guided navigation perception-action loop subsystem processes multi-channel depth image sequences, sets a trajectory planner to guide navigation, and promotes the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception-action loop for efficient and flexible navigation in unknown real environments.

10. The visual-language and exploratory embodied navigation system according to claim 6, characterized in that: Automatic trajectory hybrid reasoning subsystem, including: The automatic trajectory hybrid model subsystem builds a joint unified optimization framework and trains the joint unified optimization framework using a large-scale trajectory dataset; combines expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance training diversity; trains the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; The hybrid model embodied navigation subsystem, based on the vision-language and exploration automatic trajectory hybrid model, seamlessly migrates to simulated environments and real-world scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.

Citation Information

Patent Citations

  • Visual language indoor navigation method and device, equipment and storage medium

    CN114897179A

  • Visual language navigation pre-training method based on prompt and automatic environment exploration

    CN114970457A

  • Large model driven body intelligent agent zero sample target navigation method

    CN118258396A

  • Modularized urban owned agent reasoning system, method and device and medium

    CN118798365A

  • Methods and system for goal-conditioned exploration for object goal navigation

    US20230080342A1

Cited By

  • Multi-modal data feature extraction and optimization method based on teleoperation robot task

    CN120387147A

  • Multimodal data feature extraction and optimization method based on teleoperation robot tasks

    CN120387147B

  • Planning method and device based on visual sequence, equipment and medium

    CN120783340A

  • Body navigation method, device and equipment based on brain inspiration structured spatial memory

    CN121252786A

  • Personal question and answer method based on observation-recording-decision-making mechanism

    CN121561066A