An Embodied Navigation Method and System for Vision-Language and Exploration
Through the embodied navigation method of vision-language and exploration, multi-channel depth image video and dynamic spatial memory library are used, combined with paired sample structure strategies and boundary point dynamic selection strategies, the problem of insufficient perception and exploration capabilities of embodied agents in the dynamic environment is solved, and efficient embodied navigation and exploration positioning is achieved.
Patent Information
- Application Number
- CN202510475481.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing 3D vision-language models lack active perception and exploration capabilities in dynamic and partially observable environments, and there are problems with the sample efficiency and generalization capabilities of embodied agents based on reinforcement learning.
A embodied navigation method for vision-language and exploration is proposed. By acquiring multi-channel depth image videos, combining real environment data, using paired sample structure strategies and boundary point dynamic selection strategies, a large-scale trajectory data set is constructed, and a dynamic spatial memory bank and a joint unified optimization framework is combined to realize online exploration and lifelong learning.
It significantly improves exploration accuracy and efficiency, can efficiently and flexibly navigate in unknown real environments, and improves the accuracy of tunnel lining crack detection.
Smart Images

Figure CN119984294B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of embodied intelligent navigation exploration in spatial memory, and more specifically, to an embodied navigation method and system for vision-language and exploration. Background Art
[0002] Embodied scene understanding not only requires identifying observed objects, but also demands active exploration and reasoning in the three-dimensional physical world; the search of humans with multi-channel depth images RGB-D (combined information of color pictures and depth pictures collected by a depth vision camera in real time) is driven by the seamless combination of common sense knowledge, spatial reasoning, and visual positioning; similarly, when an embodied agent navigates in a new environment, it must operate in a continuous closed-loop cycle of exploration, perception, reasoning, and action; a key part of this process is understanding three-dimensional vision and language (3D-VL), enabling the agent to think spatially and make informed decisions about exploration locations; in recent years, significant progress has been made in the field of 3D-VL; these models utilize three-dimensional reconstruction for visual positioning, question answering, dense annotation, and context reasoning; recent methods, such as 3DVLP, PQ3D, and LEO, aim to handle multiple tasks in a single architecture through pre-training or joint unified training; however, existing 3D-VL models rely on static three-dimensional representations, assuming that a complete reconstruction of the environment is available beforehand; while this is effective in offline vision-language positioning, in the real world, embodied agents operate in partially observable and dynamic environments, and this assumption is not practical; in addition, these models usually lack the ability of active perception and exploration; in contrast, embodied agents based on reinforcement learning can explore the environment, but often face problems such as low sample efficiency, poor generalization ability due to limited training data, and lack of explicit spatial representation; therefore, combining passive 3D-VL positioning with active exploration remains a key challenge in developing intelligent systems capable of efficiently exploring and understanding the three-dimensional world; in the existing technical paradigm, although RL (reinforcement learning) can learn online (Online) and achieve exploration (Exploration), it cannot achieve grounding and lifelong learning; problems such as that 3D-VL can achieve lifelong learning and grounding but cannot explore online remain to be solved; therefore, it is necessary to propose an embodied navigation method and system for vision-language and exploration to at least partially solve the problems existing in the prior art. Summary of the Invention
[0003] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further elaborated in detail in the Detailed Implementation section; the Summary of the Invention section of the present invention does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.
[0004] To at least partially solve the above problems, the present invention provides a visual-language and exploration embodied navigation method, including:
[0005] S10. Obtain a multi-channel depth image video as trajectory data, combine it with real environment data, and collect a large-scale visual-language and exploration trajectory dataset according to the paired sample structure strategy and the boundary point dynamic selection strategy;
[0006] S20. Set a visual-language and exploration training strategy, read the memory queries from the dynamic spatial memory bank, and obtain a multi-channel depth image sequence;
[0007] S30. Combine online exploration with the update of the dynamic spatial memory bank, connect visual-language localization and exploration, construct a three-dimensional world movement understanding MTU 3D navigation, and achieve lifelong learning and exploration localization;
[0008] S40. Construct a joint unified optimization framework, train the joint unified optimization framework using the large-scale trajectory dataset, combine expert data with noisy navigation data to form a visual-language and exploration automatic trajectory hybrid model, and perform intelligent reasoning and embodied navigation in a simulated environment and a real scenario.
[0009] Preferably, S10 includes:
[0010] S101. Obtain an offline multi-channel depth image video in a multi-channel depth image video platform as trajectory data; make corresponding decisions according to object queries and language descriptions, and associate each sample of the paired sample structure with the object query and the language description to generate a decision, and construct a paired sample structure strategy; perform pre-training using the paired samples; collect visual-language localization trajectory data;
[0011] S102. Combine the visual-language localization trajectory data, and according to the exploration data structure, generate decisions based on object queries, boundary point queries, and target associations, and the boundary points change dynamically during the exploration process, and construct a boundary point dynamic selection strategy; the boundary points changing dynamically during the exploration process include: avoiding overfitting caused by only using the optimal boundary points through a boundary point dynamic selection strategy including dynamically randomly selecting boundary points, optimally selecting boundary points, and randomly optimal hybrid selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment dataset to collect and generate trajectories from the simulated scanning environment; when generating trajectories, it is determined that the exploration is successful only when the target becomes visible and reachable; set a list of visited boundary points, and perform exploration only when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration; collect a large-scale visual-language and exploration trajectory dataset; significantly improve the exploration accuracy and exploration efficiency.
[0012] Preferably, S20 includes:
[0013] S201, Set the vision-language and exploration training strategy, and read the memory query from the dynamic spatial memory bank;
[0014] S202, Combine the query language, and through the trajectory planner, conduct exploration, feedback into the world, and obtain a multi-channel depth image sequence.
[0015] Preferably, S30 includes:
[0016] S301, Take the original multi-channel depth image frame as the input, generate a single-frame local query and store it in the global spatial memory bank to update the dynamic spatial memory bank; With the help of the feature extraction and segmentation prior of the 2D basic model, conduct online query representation learning, and generate a semantic space information query table that integrates semantic information and accurate 3D spatial information;
[0017] S302, By introducing a joint optimization framework, realize the synchronous learning of object localization and exploration; During the training process, combine the semantic space information query table, and simultaneously process the object query retrieved from the spatial memory bank and the boundary of the explored map as the boundary query, so as to realize the integrated end-to-end training of the two tasks and jointly unify the exploration and localization goals;
[0018] S303, By processing the multi-channel depth image sequence, set the trajectory planner to guide navigation, and promote the acquisition of a new multi-channel depth image sequence, so as to form a continuous perception and action cycle for efficient and flexible navigation in an unknown real environment.
[0019] Preferably, S40 includes:
[0020] S401, Build a joint unified optimization framework, and train the joint unified optimization framework with a large-scale trajectory dataset; Combine expert data and noisy navigation data to form an automatic trajectory mixing strategy to enhance the diversity of training; Train the joint unified optimization framework to form a vision-language and exploration automatic trajectory mixing model;
[0021] S402, According to the vision-language and exploration automatic trajectory mixing model, seamlessly migrate to the simulation environment and the real scenario for intelligent reasoning and embodied navigation in complex environments and noisy data.
[0022] The present invention provides an embodied navigation system for vision-language and exploration, including:
[0023] A dynamic selection subsystem for trajectory boundary points, which obtains a multi-channel depth image video as trajectory data, combines real environment data, and collects a large-scale vision-language and exploration trajectory dataset according to the paired sample structure strategy and the boundary point dynamic selection strategy;
[0024] Spatial memory query combination subsystem, set visual-language and exploration training strategies, read memory queries from the dynamic spatial memory bank, and obtain multi-channel depth image sequences;
[0025] Three-dimensional world movement understanding and navigation subsystem, combine online exploration with the update of the dynamic spatial memory bank, connect visual-language positioning and exploration, construct three-dimensional world movement understanding MTU 3D navigation, and achieve lifelong learning and exploration positioning;
[0026] Automatic trajectory hybrid inference subsystem, construct a joint unified optimization framework, train the joint unified optimization framework using a large-scale trajectory dataset, combine expert data with noisy navigation data to form a visual-language and exploration automatic trajectory hybrid model, and perform intelligent inference and embodied navigation in simulated environments and real-world scenarios.
[0027] Preferably, the trajectory boundary point dynamic selection subsystem includes:
[0028] Visual-language trajectory collection subsystem, obtain offline multi-channel depth image videos in the multi-channel depth image video platform as trajectory data; make corresponding decisions according to object queries and language descriptions, pair sample structures, each sample associates object queries with language descriptions to generate decisions, and construct a paired sample structure strategy; according to the paired sample structure, use paired samples for pre-training; collect visual-language positioning trajectory data;
[0029] Dynamic strategy exploration trajectory collection subsystem, combine visual-language positioning trajectory data, according to the exploration data structure, generate decisions according to object queries, boundary point queries and target associations, and the boundary points change dynamically during the exploration process, construct a boundary point dynamic selection strategy; the boundary points change dynamically during the exploration process includes: avoid overfitting caused by only using optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random-optimal hybrid selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment dataset to collect and generate trajectories from the simulated scanning environment; when generating trajectories, it is only determined that the exploration is successful when the target becomes visible and reachable; set a list of visited boundary points, and only explore when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration; collect visual-language and exploration large-scale trajectory datasets; significantly improve the exploration accuracy and exploration efficiency.
[0030] Preferably, the spatial memory query combination subsystem includes:
[0031] Dynamic spatial memory query subsystem, set visual-language and exploration training strategies, and read memory queries from the dynamic spatial memory bank;
[0032] The trajectory planning query feedback subsystem combines the query language, explores through the trajectory planner, feeds back into the world, and obtains a multi-channel depth image sequence.
[0033] Preferably, the three-dimensional world movement understanding and navigation subsystem includes:
[0034] The spatial memory bank update query subsystem takes the original multi-channel depth image frame as input, generates a single-frame local query and stores it in the global spatial memory bank for dynamic spatial memory bank update; with the help of the feature extraction and segmentation prior of the 2D basic model, online query representation learning is carried out to generate a semantic space information query table that fuses semantic information and precise 3D spatial information;
[0035] The joint optimization framework synchronization subsystem realizes the synchronous learning of object localization and exploration by introducing a joint optimization framework; during the training process, combined with the semantic space information query table, it simultaneously processes the object query retrieved from the spatial memory bank and the boundary of the explored map as a boundary query, so as to realize the integrated end-to-end training of the two tasks and jointly unify the exploration and localization goals;
[0036] The guided navigation sensory loop subsystem processes the multi-channel depth image sequence, sets the trajectory planner to guide navigation, and promotes the acquisition of a new multi-channel depth image sequence, thus forming a continuous perception and action loop for efficient and flexible navigation in an unknown real environment.
[0037] Preferably, the automatic trajectory hybrid inference subsystem includes:
[0038] The automatic trajectory hybrid model subsystem constructs a joint unified optimization framework, trains the joint unified optimization framework using a large-scale trajectory data set; combines expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance the diversity of training; trains the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model;
[0039] The hybrid model embodied navigation subsystem seamlessly migrates to the simulation environment and the real scene according to the vision-language and exploration automatic trajectory hybrid model for intelligent inference and embodied navigation in complex environments and noisy data.
[0040] Compared with the prior art, the present invention has at least the following beneficial effects:
[0041] An embodied navigation method and system for vision-language and exploration of the present invention acquires multi-channel depth image videos as trajectory data, combines real environment data, and collects a large-scale trajectory dataset for vision-language and exploration according to the paired sample structure strategy and the boundary point dynamic selection strategy; sets a vision-language and exploration training strategy, reads the queries memorized from the dynamic spatial memory bank to obtain multi-channel depth image sequences; combines online exploration with the update of the dynamic spatial memory bank, connects vision-language localization and exploration, constructs a 3D world movement understanding MTU 3D navigation to achieve lifelong learning and exploration localization; constructs a joint unified optimization framework, trains the joint unified optimization framework using the large-scale trajectory dataset, combines expert data with noisy navigation data to form an automatic trajectory hybrid model for vision-language and exploration, and performs intelligent reasoning and embodied navigation in simulated environments and real scenarios; can significantly improve the accuracy of tunnel lining crack detection; proposes a joint unified framework that can simultaneously optimize vision-language localization tasks and exploration tasks, thereby leveraging the complementary characteristics of the two to enhance the overall performance; proposes a brand-new vision-language and exploration training strategy, performs embodied navigation pre-training through a large-scale trajectory dataset collected from real data and simulated data to achieve navigation in unknown real environments; constructs a 3D world movement understanding MTU 3D method (Move to Understand, movement understanding in the 3D world), connects vision-language localization and exploration to achieve efficient and flexible embodied navigation that combines vision localization and exploration in an embodied environment; and significantly improves the generation efficiency of trajectory data for navigation pre-training; the present invention has important technical significance and remarkable effects; has important technical significance and remarkable effects.
[0042] An embodied navigation method and system for vision-language and exploration of the present invention, other advantages, objectives, and features of the present invention will be partially reflected by the following description, and will also be understood by those skilled in the art through the research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0044] Figure 1 It is a diagram of an embodiment of the architecture of an embodied navigation system for vision-language and exploration of the present invention.
[0045] Figure 2 It is a diagram of a comparison example of the prior art of an embodied navigation method and system for vision-language and exploration of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments, so that those skilled in the art can implement it with reference to the description; as shown in the figure, the present invention provides an embodied navigation method for vision-language and exploration, including:
[0047] S10, Obtain a multi-channel depth image video as trajectory data, combine it with real environment data, and collect a large-scale vision-language and exploration trajectory dataset according to the paired sample structure strategy and the boundary point dynamic selection strategy;
[0048] S20, Set the vision-language and exploration training strategy, read the memory queries from the dynamic spatial memory bank, and obtain a multi-channel depth image sequence;
[0049] S30, Combine online exploration with the update of the dynamic spatial memory bank, connect vision-language localization and exploration, construct a three-dimensional world movement understanding MTU 3D navigation, and achieve lifelong learning and exploration localization;
[0050] S40, Construct a joint unified optimization framework, train the joint unified optimization framework using the large-scale trajectory dataset, combine expert data with noisy navigation data, form a vision-language and exploration automatic trajectory hybrid model, and perform intelligent reasoning and embodied navigation in simulated environments and real scenarios.
[0051] The principles and effects of the above technical solution are as follows: The present invention provides an embodied navigation method for vision-language and exploration, including: obtaining a multi-channel depth image video as trajectory data, combining real environment data, and collecting a large-scale vision-language and exploration trajectory data set according to the paired sample structure strategy and the boundary point dynamic selection strategy; setting a vision-language and exploration training strategy, reading the queries memorized from the dynamic spatial memory bank to obtain a multi-channel depth image sequence; combining online exploration with the update of the dynamic spatial memory bank, connecting vision-language localization and exploration, constructing a three-dimensional world movement understanding MTU3D navigation to achieve lifelong learning and exploration localization; constructing a joint unified optimization framework, training the joint unified optimization framework using the large-scale trajectory data set, combining expert data with noisy navigation data to form a vision-language and exploration automatic trajectory hybrid model for intelligent reasoning and embodied navigation in simulated environments and real scenarios; being able to significantly improve the accuracy of tunnel lining crack detection; proposing a joint unified framework that can simultaneously optimize the vision-language localization task and the exploration task, thereby leveraging the complementary characteristics of the two to enhance the overall performance; proposing a brand-new vision-language and exploration training strategy, performing embodied navigation pre-training through a large-scale trajectory data set collected from real data and simulated data to achieve navigation in unknown real environments; constructing a three-dimensional world movement understanding MTU 3D method (Move to Understand, movement understanding in the three-dimensional world), connecting vision-language localization and exploration to achieve efficient and flexible embodied navigation that combines vision localization and exploration in an embodied environment; and significantly improving the generation efficiency of the trajectory data for navigation pre-training; the present invention has important technical significance and remarkable effects.
[0052] In one embodiment, S10 includes:
[0053] S101, obtaining an offline multi-channel depth image video in a multi-channel depth image video platform as trajectory data; making corresponding decisions according to object queries and language descriptions, associating each sample of the paired sample structure with the object query and the language description to generate a decision, constructing a paired sample structure strategy; using the paired samples for pre-training according to the paired sample structure; collecting vision-language localization trajectory data;
[0054] S102. Combine the visual-language localization trajectory data, construct a decision according to the exploration data structure, generate decisions based on object queries, boundary point queries, and target associations, where the boundary points change dynamically during the exploration process, and construct a dynamic boundary point selection strategy. The dynamic change of boundary points during the exploration process includes: By including a dynamic boundary point selection strategy that dynamically performs random boundary point selection, optimal boundary point selection, and random-optimal hybrid selection, avoid overfitting caused by only using optimal boundary points. Through the dynamic boundary point selection strategy, the 3D simulator uses the 3D environmental dataset to collect and generate trajectories from the simulated scanning environment. When generating trajectories, it is only determined that the exploration is successful when the target becomes visible and reachable. Set a list of visited boundary points, and only explore when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration. Collect visual-language and exploration large-scale trajectory datasets, significantly improving the exploration accuracy and exploration efficiency.
[0055] The principle and effect of the above technical solution are: Obtain the offline multi-channel depth image video in the multi-channel depth image video platform as trajectory data; Make corresponding decisions according to object queries and language descriptions, and pair the sample structure. Each sample associates the object query with the language description to generate a decision, and construct a paired sample structure strategy; According to the paired sample structure, use the paired samples for pre-training; Collect visual-language localization trajectory data; Combine the visual-language localization trajectory data, construct a decision according to the exploration data structure, generate decisions based on object queries, boundary point queries, and target associations, where the boundary points change dynamically during the exploration process, and construct a dynamic boundary point selection strategy. The dynamic change of boundary points during the exploration process includes: By including a dynamic boundary point selection strategy that dynamically performs random boundary point selection, optimal boundary point selection, and random-optimal hybrid selection, avoid overfitting caused by only using optimal boundary points. Through the dynamic boundary point selection strategy, the 3D simulator uses the 3D environmental dataset to collect and generate trajectories from the simulated scanning environment. When generating trajectories, it is only determined that the exploration is successful when the target becomes visible and reachable. Set a list of visited boundary points, and only explore when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration. Collect visual-language and exploration large-scale trajectory datasets, significantly improving the exploration accuracy and exploration efficiency.
[0056] In one embodiment, S20 includes:
[0057] S201. Set the visual-language and exploration training strategy, and read the remembered queries from the dynamic spatial memory bank.
[0058] S202. Combine the query language, perform exploration through the trajectory planner, feedback to the world, and obtain a multi-channel depth image sequence.
[0059] The principle and effect of the above technical solution are as follows: Set up a vision-language and exploration training strategy to read memory queries from the dynamic spatial memory bank; combine the query language, and through the trajectory planner, conduct exploration and feedback it into the world to obtain a multi-channel depth image sequence;
[0060] Setting up a vision-language and exploration training strategy, reading memory queries from the dynamic spatial memory bank, combining the query language (e.g., the chair near me), through spatial reasoning, obtaining the object query to be explored, and then obtaining its spatial position information (including xyz and pose), through the Trajectories Planner for exploration, and feedback it into the world coordinate system to obtain a multi-channel depth image sequence. The multi-channel depth image sequence includes: setting up a vision-language and exploration training strategy, reading memory queries from the dynamic spatial memory bank, combining the query language (e.g., the chair near me), through spatial reasoning, obtaining the object query to be explored, and then obtaining its spatial position information (including xyz and pose), through the Trajectories Planner for exploration, and feedback it into the world coordinate system to obtain a multi-channel depth image sequence.
[0061] In one embodiment, S30 includes:
[0062] S301: Taking the original multi-channel depth image frame as input, generating a single-frame local query and storing it in the global spatial memory bank to update the dynamic spatial memory bank; with the help of the feature extraction and segmentation prior of the 2D basic model, online query representation learning, generating a semantic space information query table that fuses semantic information and accurate 3D spatial information;
[0063] S302: By introducing a joint optimization framework, realizing the synchronous learning of object localization and exploration; during the training process, combining the semantic space information query table, simultaneously processing the object query retrieved from the spatial memory bank and the boundary of the explored map as a boundary query, so as to realize the integrated end-to-end training of the two tasks and jointly unify the exploration and localization objectives;
[0064] S303: By processing the multi-channel depth image sequence, setting up a trajectory planner to guide navigation and promoting the acquisition of a new multi-channel depth image sequence, thus forming a continuous perception and action cycle for efficient and flexible navigation in an unknown real environment.
[0065] The principles and effects of the above technical solutions are as follows: Taking the original multi-channel depth image frame as the input, generating a single-frame local query and storing it in the global spatial memory bank, and performing dynamic spatial memory bank update; leveraging the feature extraction and segmentation priors of the 2D basic model for online query representation learning to generate a semantic spatial information query table that integrates semantic information and accurate 3D spatial information; by introducing a joint optimization framework, realizing the synchronous learning of object localization and exploration; during the training process, combining the semantic spatial information query table to simultaneously process the object queries retrieved from the spatial memory bank and the boundaries of the explored map as boundary queries, thereby achieving integrated end-to-end training of the two tasks, jointly unifying the exploration and localization objectives; by processing the multi-channel depth image sequence, setting a trajectory planner to guide the navigation, and promoting the acquisition of new multi-channel depth image sequences, thus forming a continuous perception and action loop for efficient and flexible navigation in an unknown real environment;
[0066] The multi-channel depth image includes the color picture and depth picture information collected by the camera in real time; the 2D basic model includes: DINO (Distillation and NO labels) or SAM (Segment Anything Model);
[0067] The joint unified optimization framework includes: unifying the training objectives of the two previous different tasks of joint vision-language localization and exploration, integrating and co-optimizing the vision-language localization task and the exploration task, and utilizing their complementary characteristics to enhance the spatial accuracy of the agent's embodied navigation perception of the three-dimensional space and the intelligent decision-making action level of exploring the three-dimensional world, thus forming a joint unified optimization framework;
[0068] Efficient and flexible navigation in an unknown real environment by processing the multi-channel depth image sequence, setting a trajectory planner to guide the navigation, and promoting the acquisition of new multi-channel depth image sequences, thus forming a continuous perception and action loop includes: generating object queries by processing the multi-channel depth image sequence, and then storing these object queries in the dynamic spatial memory bank; the spatial reasoning layer selects the object queries to be explored from the object queries and boundary queries, and obtains the spatial position information of the selected object queries; setting a trajectory planner to guide the navigation; transmitting the spatial position information of the selected object queries to the trajectory planner; the trajectory planner guides the navigation and promotes the acquisition of new multi-channel depth image sequences, thus forming a continuous perception and action loop for efficient and flexible embodied navigation in an unknown real environment.
[0069] Such as Figure 1As shown, Unified Grounding and Exploration; obtaining partial RGB-D sequences; performing Online Query Representation Learning. Online Query Representation Learning includes: object queries are proposed through the feature extraction (query proposal) modules of the 2D encoder and 3D encoder for the obtained image, depth information, and position information. Frontier queries can be obtained from the frontier mapping based on the position information; bottom right: Dynamic Spatial Memory Bank, where object queries and frontier queries are written through Memory Write; constructing a unified optimization framework.
[0070] In one embodiment, S40 includes:
[0071] S401, constructing a unified optimization framework, training the unified optimization framework using a large-scale trajectory dataset; combining expert data with noisy navigation data to form an automatic trajectory mixing strategy to enhance the diversity of training; training the unified optimization framework to form a vision-language and exploration automatic trajectory mixing model;
[0072] S402, according to the vision-language and exploration automatic trajectory mixing model, seamlessly migrating to simulation environments and real-world scenarios for intelligent reasoning and embodied navigation in complex environments and with noisy data.
[0073] The principle and effect of the above technical solution are: constructing a unified optimization framework, training the unified optimization framework using a large-scale trajectory dataset; combining expert data with noisy navigation data to form an automatic trajectory mixing strategy to enhance the diversity of training; training the unified optimization framework to form a vision-language and exploration automatic trajectory mixing model; according to the vision-language and exploration automatic trajectory mixing model, seamlessly migrating to simulation environments and real-world scenarios for intelligent reasoning and embodied navigation in complex environments and with noisy data.
[0074] The present invention provides an embodied navigation system for vision-language and exploration, including:
[0075] A dynamic selection subsystem for trajectory boundary points, obtaining multi-channel depth image videos as trajectory data, combining with real environment data, and collecting a large-scale trajectory dataset for vision-language and exploration according to the paired sample structure strategy and the boundary point dynamic selection strategy;
[0076] The spatial memory query combination subsystem sets visual-language and exploration training strategies, reads the queries of memories from the dynamic spatial memory library, and obtains multi-channel depth image sequences;
[0077] The 3D world movement understanding and navigation subsystem combines online exploration with the update of the dynamic spatial memory library, connects visual-language positioning and exploration, constructs the 3D world movement understanding MTU 3D navigation, and realizes lifelong learning and exploration positioning;
[0078] The automatic trajectory hybrid reasoning subsystem constructs a joint unified optimization framework, trains the joint unified optimization framework using a large-scale trajectory dataset, combines expert data with noisy navigation data, forms a visual-language and exploration automatic trajectory hybrid model, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios.
[0079] The principles and effects of the above technical solutions are as follows: The present invention provides an embodied navigation system for visual-language and exploration, including: a trajectory boundary point dynamic selection subsystem that obtains multi-channel depth image videos as trajectory data, combines real environmental data, and collects a large-scale visual-language and exploration trajectory dataset according to the paired sample structure strategy and the boundary point dynamic selection strategy; a spatial memory query combination subsystem that sets visual-language and exploration training strategies, reads the queries of memories from the dynamic spatial memory library, and obtains multi-channel depth image sequences; a 3D world movement understanding and navigation subsystem that combines online exploration with the update of the dynamic spatial memory library, connects visual-language positioning and exploration, constructs the 3D world movement understanding MTU 3D navigation, and realizes lifelong learning and exploration positioning; an automatic trajectory hybrid reasoning subsystem that constructs a joint unified optimization framework, trains the joint unified optimization framework using a large-scale trajectory dataset, combines expert data with noisy navigation data, forms a visual-language and exploration automatic trajectory hybrid model, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios; it can significantly improve the accuracy of tunnel lining crack detection; a joint unified framework that can simultaneously optimize visual-language positioning tasks and exploration tasks is proposed, so as to utilize the complementary characteristics of the two to enhance the overall performance; a new visual-language and exploration training strategy is proposed, and embodied navigation pre-training is performed through a large-scale trajectory dataset collected from real data and simulated data to achieve navigation in unknown real environments; a 3D world movement understanding MTU 3D method (Move to Understand) is constructed, which connects visual-language positioning and exploration to achieve efficient and flexible embodied navigation that combines visual positioning and exploration in an embodied environment; and the generation efficiency of the trajectory data for navigation pre-training is significantly improved; the present invention has important technical significance and remarkable effects.
[0080] In one embodiment, the trajectory boundary point dynamic selection subsystem includes:
[0081] The vision-language trajectory collection subsystem obtains the offline multi-channel depth image video in the multi-channel depth image video platform as trajectory data; makes decisions corresponding to object queries and language descriptions, and for each sample in the paired sample structure, associates the object query with the language description to generate a decision, constructing a paired sample structure strategy; according to the paired sample structure, uses the paired samples for pre-training; collects vision-language localization trajectory data;
[0082] The dynamic policy exploration trajectory collection subsystem combines the vision-language localization trajectory data, and according to the exploration data structure, generates decisions based on object queries, boundary point queries, and target associations, and the boundary points change dynamically during the exploration process, constructing a dynamic boundary point selection strategy; the dynamic change of the boundary points during the exploration process includes: avoiding overfitting caused by only using the optimal boundary points through a dynamic boundary point selection strategy including randomly selecting boundary points dynamically, selecting the optimal boundary points, and a random-optimal hybrid selection; through the dynamic boundary point selection strategy, the 3D simulator uses the 3D environment dataset to collect and generate trajectories from the simulated scanning environment; when generating trajectories, it is determined that the exploration is successful only when the target becomes visible and reachable; sets a list of visited boundary points, and only explores when there are better potential boundary points closer to the target, otherwise triggers an exception to prevent unnecessary exploration; collects vision-language and exploration large-scale trajectory datasets; significantly improves the exploration accuracy and exploration efficiency.
[0083] The principle and effect of the above technical solution are as follows: The trajectory boundary point dynamic selection subsystem includes:
[0084] The vision-language trajectory collection subsystem obtains the offline multi-channel depth image video in the multi-channel depth image video platform as trajectory data; makes decisions corresponding to object queries and language descriptions, and for each sample in the paired sample structure, associates the object query with the language description to generate a decision, constructing a paired sample structure strategy; according to the paired sample structure, uses the paired samples for pre-training; collects vision-language localization trajectory data;
[0085] The dynamic policy exploration trajectory collection subsystem combines visual - language localization trajectory data, generates decisions according to the exploration data structure, object queries, boundary point queries, and target associations. The boundary points change dynamically during the exploration process, and a dynamic boundary point selection strategy is constructed. The dynamic change of boundary points during the exploration process includes: avoiding overfitting caused by only using optimal boundary points through a boundary point dynamic selection strategy that includes randomly selecting boundary points dynamically, selecting optimal boundary points, and a random - optimal hybrid selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment dataset to collect and generate trajectories from the simulated scanning environment; when generating trajectories, exploration is considered successful only when the target becomes visible and reachable; a list of visited boundary points is set, and exploration is only carried out when there are better potential boundary points closer to the target, otherwise an exception is triggered to prevent unnecessary exploration; collect visual - language and exploration large - scale trajectory datasets; significantly improve the exploration accuracy and exploration efficiency.
[0086] In one embodiment, the spatial memory query combination subsystem includes:
[0087] The dynamic spatial memory query subsystem sets visual - language and exploration training strategies and reads memory queries from the dynamic spatial memory library;
[0088] The trajectory planning query feedback subsystem combines the query language, explores through the trajectory planner, feeds back into the world, and obtains a multi - channel depth image sequence.
[0089] The principle and effect of the above - mentioned technical solution are as follows: The spatial memory query combination subsystem includes: the dynamic spatial memory query subsystem that sets visual - language and exploration training strategies and reads memory queries from the dynamic spatial memory library; the trajectory planning query feedback subsystem that combines the query language, explores through the trajectory planner, feeds back into the world, and obtains a multi - channel depth image sequence. Setting visual - language and exploration training strategies, reading memory queries from the dynamic spatial memory library, combining the query language, exploring through the trajectory planner, and feeding back into the world to obtain a multi - channel depth image sequence includes: visual - language and exploration training strategies, reading memory queries from the dynamic spatial memory library, combining the query language (such as: near my chair), obtaining the object query to be explored through spatial reasoning, and then obtaining its spatial position information (including xyz and pose), exploring through the Trajectories Planner, and feeding back into the world coordinate system to obtain a multi - channel depth image sequence.
[0090] In one embodiment, the three - dimensional world movement understanding and navigation subsystem includes:
[0091] The spatial memory bank update and query subsystem takes the original multi-channel depth image frames as input, generates single-frame local queries and stores them in the global spatial memory bank for dynamic spatial memory bank update; with the help of the feature extraction and segmentation priors of the 2D basic model, online query representation learning is carried out to generate a semantic spatial information query table that integrates semantic information and accurate 3D spatial information;
[0092] The joint optimization framework synchronization subsystem realizes the synchronous learning of object localization and exploration by introducing a joint optimization framework; during the training process, combined with the semantic spatial information query table, it simultaneously processes the object queries retrieved from the spatial memory bank and the boundaries of the explored map as boundary queries, so as to achieve the integrated end-to-end training of the two tasks and jointly unify the exploration and localization goals;
[0093] The guided navigation perception loop subsystem forms a continuous perception and action loop for efficient and flexible navigation in an unknown real environment by processing multi-channel depth image sequences, setting a trajectory planner to guide navigation, and promoting the acquisition of new multi-channel depth image sequences.
[0094] The principles and effects of the above technical solutions are as follows: The three-dimensional world movement understanding and navigation subsystem includes: the spatial memory bank update and query subsystem, which takes the original multi-channel depth image frames as input, generates single-frame local queries and stores them in the global spatial memory bank for dynamic spatial memory bank update; with the help of the feature extraction and segmentation priors of the 2D basic model, online query representation learning is carried out to generate a semantic spatial information query table that integrates semantic information and accurate 3D spatial information; the joint optimization framework synchronization subsystem realizes the synchronous learning of object localization and exploration by introducing a joint optimization framework; during the training process, combined with the semantic spatial information query table, it simultaneously processes the object queries retrieved from the spatial memory bank and the boundaries of the explored map as boundary queries, so as to achieve the integrated end-to-end training of the two tasks and jointly unify the exploration and localization goals; the guided navigation perception loop subsystem forms a continuous perception and action loop for efficient and flexible navigation in an unknown real environment by processing multi-channel depth image sequences, setting a trajectory planner to guide navigation, and promoting the acquisition of new multi-channel depth image sequences.
[0095] The multi-channel depth image includes the color picture and depth picture information collected by the camera in real time; the 2D basic model includes: DINO (Distillation and NO labels) or SAM (Segment Anything Model);
[0096] The unified optimization framework includes: unifying the training objectives of two previous different tasks, namely visual-language localization and exploration, integrating and co-optimizing the visual-language localization task and the exploration task, and leveraging their complementary characteristics to enhance the agent's embodied navigation perception of the three-dimensional space accuracy and the intelligent decision-making and action level in exploring the three-dimensional world, thereby forming a unified optimization framework;
[0097] By processing multi-channel depth image sequences, setting a trajectory planner to guide navigation, and facilitating the acquisition of new multi-channel depth image sequences, a continuous perception and action loop is formed for efficient and flexible navigation in an unknown real environment, including: generating object queries by processing multi-channel depth image sequences, and then storing these object queries in a dynamic spatial memory bank; the spatial reasoning layer selects the object queries to be explored from the object queries and boundary queries, and obtains the spatial position information of the selected object queries; setting a trajectory planner to guide navigation; transmitting the spatial position information of the selected object queries to the trajectory planner; the trajectory planner guides navigation and facilitates the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception and action loop for efficient and flexible embodied navigation in an unknown real environment.
[0098] Such as Figure 1 As shown, Unified Grounding and Exploration; obtaining partial multi-channel depth image sequences; performing online query feature learning. Online query feature learning includes: the obtained image, depth information, and position information are used to propose object queries through the feature extraction (query proposal) modules of the 2D encoder and 3D encoder, and frontier queries can be obtained from the frontier mapping based on the position information; in the lower right corner: Dynamic Spatial Memory Bank, writing object queries and frontier queries through Memory Write; constructing a unified optimization framework.
[0099] In one embodiment, the automatic trajectory hybrid inference subsystem includes:
[0100] The automatic trajectory hybrid model subsystem constructs a unified optimization framework, trains the unified optimization framework using a large-scale trajectory dataset; combines expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance the diversity of training; trains the unified optimization framework to form a visual-language and exploration automatic trajectory hybrid model;
[0101] The hybrid model embodied navigation subsystem, based on the visual-language and exploration automatic trajectory hybrid model, seamlessly migrates to simulated environments and real-world scenarios for intelligent reasoning and embodied navigation in complex environments and with noisy data.
[0102] The principle and effect of the above technical solution are as follows: The automatic trajectory hybrid reasoning subsystem includes: an automatic trajectory hybrid model subsystem that constructs a joint unified optimization framework and trains the joint unified optimization framework using a large-scale trajectory dataset; combines expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance the diversity of training; trains the joint unified optimization framework to form a visual-language and exploration automatic trajectory hybrid model; the hybrid model embodied navigation subsystem, based on the visual-language and exploration automatic trajectory hybrid model, seamlessly migrates to simulated environments and real-world scenarios for intelligent reasoning and embodied navigation in complex environments and with noisy data.
[0103] Although the embodiments of the present invention have been disclosed as above, it is not limited to only the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the illustrated and described examples here.
Claims
1. A visual-language and exploratory embodied navigation method, characterized in that: include: S10, obtain multi-channel depth image video as trajectory data, combine it with real environment data, and collect vision-language and explore large-scale trajectory data sets based on paired sample structure strategy and boundary point dynamic selection strategy; S20, setting the vision-language and exploration training strategy, reading the memory query from the dynamic spatial memory bank, and obtaining a multi-channel depth image sequence; S30, combines online exploration with dynamic spatial memory updates, connects visual-linguistic positioning and exploration, builds MTU 3D navigation for mobile understanding of the three-dimensional world, and realizes lifelong learning and exploration positioning; S40, build a joint unified optimization framework, use large-scale trajectory datasets to train the joint unified optimization framework, combine expert data with noisy navigation data, form a hybrid model of vision-language and exploration automatic trajectory, and perform intelligent reasoning and embodied navigation in simulated environments and real-world scenarios; S30 includes: S301, taking the original multi-channel depth image frame as input, generating a single-frame local query and storing it in the global spatial memory, and performing dynamic spatial memory update; using the feature extraction and segmentation prior of the 2D basic model, online query representation learning, and generating a semantic spatial information query table that integrates semantic information and precise 3D spatial information; S302, by introducing a joint optimization framework, synchronous learning of object positioning and exploration is achieved; During the training process, the semantic spatial information query table is combined to simultaneously process object queries retrieved from the spatial memory and the boundaries of the explored map as boundary queries, thereby achieving integrated end-to-end training of the two tasks and jointly exploring and locating the target; S303, by processing the multi-channel depth image sequence, setting a trajectory planner to guide navigation, and promoting the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception and action cycle to perform efficient and flexible navigation in an unknown real environment.
2. A visual-language and exploratory embodied navigation method according to claim 1, characterized in that: S10 includes: S101, obtaining offline multi-channel deep image video in a multi-channel deep image video platform as trajectory data; according to the object query and the language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, using the paired samples for pre-training; collecting visual-language localization trajectory data; S102, combining visual-language positioning trajectory data, according to the exploration data structure, based on object query, boundary point query and target association to generate decisions, and the boundary points change dynamically during the exploration process, to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; setting a list of visited boundary points, and exploring only when there are better potential boundary points closer to the target, otherwise triggering an exception to prevent unnecessary exploration; collecting large-scale visual-language and exploration trajectory data sets; significantly improving exploration accuracy and efficiency.
3. A visual-language and exploratory embodied navigation method according to claim 1, characterized in that: The S20 includes: S201, setting the vision-language and exploration training strategy, and reading the memorized query from the dynamic spatial memory bank; S202, combining the language of the inquiry, through the trajectory planner, to conduct exploration, feedback to the world, and obtain a multi-channel depth image sequence.
4. The visual-language and exploratory embodied navigation method according to claim 1, characterized in that: S40 includes: S401, build a joint unified optimization framework and train it using a large-scale trajectory dataset; combine expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance training diversity; train the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; S402, based on the hybrid model of vision-language and exploration automatic trajectory, seamlessly migrates to simulated environments and real-world scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.
5. A visual-language and exploratory embodied navigation system, characterized in that: include: The trajectory boundary point dynamic selection subsystem obtains multi-channel depth image videos as trajectory data, combines real environment data, and collects vision-language and exploration large-scale trajectory data sets based on the paired sample structure strategy and boundary point dynamic selection strategy; The spatial memory query integration subsystem sets the visual-language and exploration training strategies, reads the memory query from the dynamic spatial memory library, and obtains a multi-channel depth image sequence; The 3D world mobile understanding navigation subsystem combines online exploration with dynamic spatial memory updates, connects visual-language positioning and exploration, builds 3D world mobile understanding MTU 3D navigation, and realizes lifelong learning and exploration positioning; The automatic trajectory hybrid reasoning subsystem builds a joint unified optimization framework, which is trained using a large-scale trajectory dataset, combines expert data with noisy navigation data to form a vision-language and exploration automatic trajectory hybrid model, and performs intelligent reasoning and embodied navigation in simulated environments and real-world scenarios; The 3D world mobile understanding and navigation subsystem includes: The spatial memory update query subsystem takes the original multi-channel depth image frame as input, generates a single-frame local query and stores it in the global spatial memory, and performs dynamic spatial memory update. With the help of feature extraction and segmentation priors of the 2D basic model, online query representation learning is performed to generate a semantic spatial information query table that integrates semantic information and precise 3D spatial information. The joint optimization framework synchronization subsystem realizes the synchronous learning of object positioning and exploration by introducing the joint optimization framework. During the training process, it combines the semantic spatial information query table and simultaneously processes the object query retrieved from the spatial memory library and the boundary of the explored map as the boundary query, thereby realizing the integrated end-to-end training of the two tasks and jointly exploring the positioning target. The guided navigation perception-action loop subsystem processes multi-channel depth image sequences, sets a trajectory planner to guide navigation, and promotes the acquisition of new multi-channel depth image sequences, thereby forming a continuous perception-action loop for efficient and flexible navigation in unknown real environments.
6. A visual-language and exploratory embodied navigation system according to claim 5, characterized in that: The trajectory boundary point dynamic selection subsystem includes: The visual-language trajectory collection subsystem obtains offline multi-channel deep image videos in the multi-channel deep image video platform as trajectory data; according to the object query and language description corresponding decision, each sample of the paired sample structure associates the object query with the language description to generate a decision, and constructs a paired sample structure strategy; according to the paired sample structure, the paired samples are used for pre-training; and the visual-language positioning trajectory data is collected; The dynamic strategy exploration trajectory collection subsystem combines visual-language positioning trajectory data, generates decisions based on object queries, boundary point queries and target associations according to the exploration data structure, and the boundary points change dynamically during the exploration process to build a dynamic boundary point selection strategy; the dynamic changes of boundary points during the exploration process include: avoiding overfitting caused by using only the optimal boundary points through a boundary point dynamic selection strategy including dynamic random boundary point selection, optimal boundary point selection and random optimal mixed selection; through the boundary point dynamic selection strategy, the 3D simulator uses the 3D environment data set to collect and generate trajectories from the simulated scanning environment; when generating trajectories, the exploration is judged to be successful only when the target becomes visible and reachable; set a list of visited boundary points, and explore only when there are better potential boundary points closer to the target, otherwise trigger an exception to prevent unnecessary exploration; collect visual-language and exploration large-scale trajectory data sets; significantly improve exploration accuracy and efficiency.
7. The visual-language and exploratory embodied navigation system according to claim 5, characterized in that: Spatial memory inquiry combined with subsystems, including: The dynamic spatial memory query subsystem sets the visual-language and exploration training strategies and reads the memory query from the dynamic spatial memory bank; The trajectory planning query feedback subsystem combines the query language with the trajectory planner to conduct exploration and feedback to the world to obtain a multi-channel depth image sequence.
8. The visual-language and exploratory embodied navigation system according to claim 5, characterized in that: Automatic trajectory hybrid reasoning subsystem, including: The automatic trajectory hybrid model subsystem builds a joint unified optimization framework and trains the joint unified optimization framework using a large-scale trajectory dataset; combines expert data with noisy navigation data to form an automatic trajectory hybrid strategy to enhance training diversity; trains the joint unified optimization framework to form a vision-language and exploration automatic trajectory hybrid model; The hybrid model embodied navigation subsystem, based on the vision-language and exploration automatic trajectory hybrid model, seamlessly migrates to simulated environments and real-world scenarios to perform intelligent reasoning and embodied navigation in complex environments and noisy data.
Citation Information
Patent Citations
Modularized urban owned agent reasoning system, method and device and medium
CN118798365A