Robot zero-shot instance navigation method based on priori driving and multi-evidence fusion

CN122813832APending Publication Date: 2026-09-25NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610896550.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]针对现有机器人导航方法在新环境中泛化能力弱、依赖单一不稳定检测信号的问题,本发明提供一种基于先验驱动与多证据融合的机器人零样本实例导航方法,通过“目标+上下文”联合判断与自适应决策的机制,本发明能够在无需任何额外训练的情况下,显著提升机器人在全新未知环境中寻找新物体的成功率和效率

Benefits of technology

本发明通过构建完全基于预训练模型的零样本流程,实现了对新环境的开箱即用,无需任何额外训练。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813832A_ABST
    Figure CN122813832A_ABST
Patent Text Reader

Abstract

The application discloses a robot zero-sample instance navigation method based on prior driving and multi-evidence fusion, inputs natural language navigation instructions into a large language model to generate a context prior object list strongly related to a target object in the natural language navigation instructions; moves in an unknown environment according to an exploration strategy based on a front point to obtain visual observation information of the environment in real time; uses an open vocabulary detection model to detect the target object and the context prior object in the visual observation information in parallel; calculates direct semantic alignment scores and confidence evidence scores, fuses and calculates a comprehensive confidence score of the target object; compares the comprehensive confidence score with a dynamic confidence threshold to determine whether the target object exists in the current observation, and generates a robot navigation action based on the determination result. The application realizes efficient and robust navigation of a new object in an unknown environment without any scene-specific training, and significantly improves the success rate and efficiency of zero-sample instance navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot autonomous navigation technology, and in particular to a zero-shot instance navigation method for robots based on prior driving and multi-evidence fusion. Background Technology

[0002] Currently, enabling robots to navigate in an unfamiliar home based on verbal commands (such as "find the red cushion in the living room") is a challenge. Existing mainstream methods fall into two main categories: one is reinforcement learning methods that rely on extensive scene training, which perform poorly in new environments or when encountering new objects; the other is zero-shot detection methods using visual language models, but these methods rely too heavily on single visual recognition of the target object and are prone to failure when the target is occluded, in poor lighting, or when similar objects are present, leading to navigation interruptions. Therefore, a more robust and general robot navigation method is needed. Summary of the Invention

[0003] To address the issues of weak generalization ability and reliance on single, unstable detection signals in existing robot navigation methods in new environments, this invention provides a zero-shot instance navigation method for robots based on prior-driven and multi-evidence fusion. Through a mechanism of joint judgment and adaptive decision-making based on "target + context", this invention can significantly improve the success rate and efficiency of robots in finding new objects in completely new and unknown environments without any additional training.

[0004] This invention provides a zero-shot instance navigation method for robots based on prior-driven and multi-evidence fusion, comprising the following steps: Step S1: Input the natural language navigation instructions into the large language model to generate a list of contextual prior objects that are strongly related to the target object in the natural language navigation instructions; Step S2: Real-time acquisition of visual observation images containing RGB and depth information collected by the robot as it moves in the unknown environment based on the exploration strategy; Step S3: Using a zero-shot open vocabulary detection model, candidate target objects and context prior objects are detected in parallel in the visually observed image; Step S4: Combine the direct semantic alignment score of the candidate target object itself with the confidence evidence scores of all relevant context objects in its spatial neighborhood to calculate the comprehensive confidence score of each candidate target object. Step S5: Compare the comprehensive confidence score with a confidence threshold that decreases adaptively over time to determine whether the current candidate target object is the target object in the natural language navigation instruction, and control the robot to stop or continue exploring based on the determination result.

[0005] In one embodiment of the present invention, step S1 specifically includes: Step S101: Combine the natural language navigation instructions containing the description of the target object with the preset prompt word template to form a query for the large language model, so as to stimulate the common sense reasoning ability of the large language model to generate relevant context objects. Step S102: Parse and clean the text output of the large language model, extract a set of specific object names, and form an initial candidate context object set. Step S103: Calculate the semantic similarity between each candidate context object and the description of the target object using a pre-trained visual language model, and filter and sort the initial candidate context object set based on the semantic similarity score to generate the final context prior object list.

[0006] In one embodiment of the present invention, step S4 specifically includes: Step S401: For each detected candidate target object, calculate its direct semantic alignment score with the description of the natural language navigation instructions; Step S402: Delineate a spatial neighborhood centered on the candidate target object, and identify all detection instances belonging to the context prior object list located within the spatial neighborhood; Step S403: For each detected instance identified in the spatial neighborhood, calculate its matching score with the prior object list of its context and its direct semantic score with the description of the target object. Step S404: Aggregate the matching scores and direct semantic scores of all detected instances to obtain the confidence evidence score of the candidate target object; Step S405: The direct semantic alignment score and the confidence evidence score are fused using a nonlinear weighting function to obtain the comprehensive confidence score for each candidate target object.

[0007] In one embodiment of the present invention, step S5 specifically includes: The confidence threshold gradually decreases according to the time or steps already spent, following a predefined decay function. When the overall confidence score of any candidate target object exceeds the confidence threshold at the current moment, the candidate target object is determined to be the target object in the natural language navigation command, and the robot plans a path to move towards it and eventually stops. If no candidate target object triggers confirmation during the entire exploration process, the robot continues to execute the exploration strategy until the confidence threshold decays to the minimum value or the exploration ends.

[0008] In one embodiment of the present invention, in step S2, the exploration strategy is a hybrid strategy based on semantic value map and frontier detection; the semantic value map is constructed by calculating the global similarity between the current visual observation and the description of the natural language navigation instructions, and is used to guide the robot to the area more relevant to the description.

[0009] In one embodiment of the present invention, in step S3, each frame of visual observation image, along with the target object description and the context prior object list, is input into the open vocabulary detection model, and all detected target boxes, categories, and their corresponding confidence scores in the visual observation image are output.

[0010] The zero-shot instance navigation method for robots based on prior-driven and multi-evidence fusion according to the embodiments of the present invention has the following beneficial effects: This invention enables out-of-the-box use in new environments without any additional training by constructing a zero-shot process based entirely on a pre-trained model.

[0011] This invention significantly improves the robustness and success rate of navigation decisions in complex situations such as occlusion and changes in lighting by fusing direct evidence of the target with indirect evidence from the scene context.

[0012] This invention optimizes exploration efficiency by introducing a dynamic decision threshold that decays over time, thus avoiding early false alarms while ensuring that no target is missed in later stages.

[0013] This invention combines common sense from a large language model with open vocabulary detection, enabling robots to understand and execute navigation instructions for objects of unknown categories.

[0014] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0015] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A schematic diagram of the process for a zero-shot instance navigation method for robots based on prior driving and multi-evidence fusion provided by the present invention; Figure 2 This is a schematic diagram of the structure of a zero-shot instance navigation method for robots based on prior driving and multi-evidence fusion provided by the present invention. Detailed Implementation

[0016] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0017] In real-world scenarios, objects often appear in specific combinations (e.g., a remote control often appears near a sofa and a television). Based on this, this invention proposes a novel approach: instead of relying solely on the detection of the target itself, it simultaneously searches for the target and its common "contextual partners," fusing multiple pieces of evidence to make more reliable judgments, thereby improving navigation success rates in complex and unknown environments. Specifically, upon receiving a linguistic instruction (e.g., "find the lamp on the bedside table in the bedroom"), a large language model is first invoked to infer a list of contextual objects strongly related to the target based on common sense (e.g., "bed," "bedside table," "socket"). Subsequently, during exploration, the robot utilizes an open-vocabulary detection model to simultaneously search for the target object and these contextual objects. During decision-making, the system not only considers the confidence level of the target itself being detected but also calculates the presence, quantity, and quality of contextual objects around it, fusing these two pieces of evidence into a comprehensive score. Simultaneously, a judgment threshold that gradually decreases over time is adopted: a high threshold is set in the early stages of exploration to avoid misidentification; the standard is lowered in the later stages to prevent missing targets. Through this mechanism of joint judgment and adaptive decision-making based on "target + context", the present invention can significantly improve the success rate and efficiency of robots in finding new objects in completely new and unknown environments without any additional training.

[0018] like Figure 1 and Figure 2 As shown, the zero-shot instance navigation method for robots based on prior-driven and multi-evidence fusion includes the following steps: Step S1: Input the natural language navigation instructions into the large language model to generate a list of contextual prior objects that are strongly related to the target object in the natural language navigation instructions.

[0019] In an embodiment of the present invention, step S1 specifically includes: Step S101: Combine the natural language navigation instructions containing the description of the target object with the preset prompt word template to form a query for the large language model, so as to stimulate the common sense reasoning ability of the large language model to generate relevant context objects. Step S102: Parse and clean the text output of the large language model, extract a set of specific object names, and form an initial candidate context object set. Step S103: Calculate the semantic similarity between each candidate context object and the description of the target object using a pre-trained visual language model, and filter and sort the initial candidate context object set based on the semantic similarity score, removing irrelevant or overly general words (such as "things" or "objects"), and generating the final context prior object list.

[0020] The visual language model can be BLIP-2, whose text encoder is used to calculate the cosine similarity between the target description and the context object name.

[0021] In one specific embodiment, the extracted description of the target object is inserted into a preset prompt word template, such as "In an indoor environment, which other objects are most likely to appear next to a [target object]? Please list 5-8 of the most common related objects." This forms a query against a large language model. Based on its inherent common sense knowledge, the model generates a set of related object nouns; for example, for "remote control," it might output "sofa, coffee table, television, cushion, blanket." Subsequently, the system parses and deduplicates this list to form a structured contextual prior object list.

[0022] Step S2: Real-time acquisition of visual observation images containing RGB and depth information collected by the robot as it moves in the unknown environment based on the exploration strategy.

[0023] In step S2, the exploration strategy is a hybrid strategy based on semantic value map and frontier detection; the semantic value map is constructed by calculating the global similarity between the current visual observation and the description of the natural language navigation instructions, and is used to guide the robot to the area that is more relevant to the description.

[0024] Specifically, the robot's RGB-D camera continuously collects RGB images and depth information of the environment during exploration. Simultaneously, the robot employs a front-line-based exploration strategy to continuously explore the most valuable front-line points in the 2D semantic value map within the unknown indoor environment, and concurrently constructs a 2D occupancy grid map for real-time localization, obstacle avoidance, and 2D semantic value mapping.

[0025] Step S3: Using a zero-shot open vocabulary detection model, candidate target objects and context prior objects are detected in parallel in the visually observed image.

[0026] The zero-shot open vocabulary detection model is Grounding DINO, which takes a target description and a list of context objects as text input to achieve open vocabulary detection and segmentation of multiple objects in an image.

[0027] In step S3, each frame of visual observation image, along with the target object description and the context prior object list, is input into the pre-trained open vocabulary detection model, which outputs all detected target boxes, categories, and their corresponding confidence scores in the visual observation image.

[0028] Step S4: Combine the direct semantic alignment score of the candidate target object itself with the confidence evidence scores of all relevant context objects in its spatial neighborhood to calculate the comprehensive confidence score of each candidate target object.

[0029] In an embodiment of the present invention, step S4 specifically includes: Step S401: For each detected candidate target object, calculate its direct semantic alignment score with the description of the natural language navigation instructions; Step S402: Delineate a spatial neighborhood centered on the candidate target object, and identify all detection instances belonging to the context prior object list within that spatial neighborhood; Step S403: For each detected instance identified in the spatial neighborhood, calculate its matching score with the context prior object list and its direct semantic score with the target object description. Step S404: Aggregate the matching scores and direct semantic scores of all detected instances to obtain the confidence evidence score of the candidate target object; wherein, the confidence evidence score is determined by the semantic score and its spatial distance from the target box. The specific formula is as follows: , in, This is a detection instance for a candidate target object. Let C be the confidence evidence score for this detection instance, and C be the list of context prior objects. For a certain context prior object, For detection examples A detection instance within a spatial neighborhood, where r is the radius of the spatial neighborhood. For detection examples The direct semantic score of the target object description. For testing examples With context prior objects The direct semantic score.

[0030] Step S405: The direct semantic alignment score and the confidence evidence score are fused using a non-linear weighting function to obtain the comprehensive confidence score for each candidate target object. The formula for calculating the comprehensive confidence score is as follows: in, It is a nonlinear weighting function. For detection examples The confidence score of the evidence. For detection examples The direct semantic score of the target object description.

[0031] The nonlinear weighting function aims to emphasize that direct semantic alignment is a necessary condition, while contextual evidence is a sufficient condition, which can significantly enhance the overall confidence. The two work together to improve the robustness of the judgment.

[0032] The specific formula is as follows: in, Example of candidate target object detection The overall confidence score.

[0033] Specifically, for the current frame, all objects detected as "target objects" are first identified as candidate target objects, and their direct semantic alignment scores are directly calculated using BLIP-2. For each candidate bounding box, a spatial neighborhood, such as a circular region with radius R, is defined centered on it. Then, all detection instances located within this neighborhood and belonging to the context prior list are identified, and the aggregated score of all detection instances within this neighborhood is calculated as the confidence evidence score. Finally, the direct semantic alignment score and the confidence evidence score are non-linearly weighted and fused to obtain the comprehensive confidence score of the candidate target object. The core of this method is that the probability of a detected "target" being a true target is greatly increased if strongly related context objects frequently appear around it. The aggregation and fusion of contextual evidence aims to improve the system's robustness to occluded, rare viewpoints, or blurred targets. Even if the direct detection signal is weak, the system can still make a high-confidence judgment if there is strong contextual support.

[0034] Step S5: Compare the overall confidence score with a confidence threshold that decreases adaptively over time to determine whether the current candidate target object is the target object in the natural language navigation command, and control the robot to stop or continue exploring based on the determination result.

[0035] In an embodiment of the present invention, step S5 specifically includes: The confidence threshold gradually decreases according to the time or steps already spent, following a predefined decay function. When the overall confidence score of any candidate target object exceeds the confidence threshold at the current moment, the candidate target object is determined to be the target object in the natural language navigation command, and the robot plans a path to move toward it and eventually stops. If no candidate target object triggers confirmation during the entire exploration process, the robot continues to execute the exploration strategy until the confidence threshold decays to the minimum value or the exploration ends.

[0036] In step S5, a relatively high dynamic confidence threshold that decays over time is initialized. ,in, The final low dynamic confidence threshold, This serves as the initial high dynamic confidence threshold. The attenuation coefficient is... The exploration progress is measured as a percentage of the total budget (e.g., time elapsed as a percentage of total time). At the start of navigation, a high threshold is set to avoid false alarms. During exploration, this threshold decays exponentially based on the elapsed time or the number of exploration steps. This ensures the system can more actively identify targets in the later stages of exploration, preventing missed detections due to overly conservative approaches. For each frame's detection results, if the overall presence evidence score of a candidate target exceeds the current dynamic threshold, the target is considered found. The robot stops exploring and plans a path to move towards that target. If no candidate target exceeds the threshold, the robot continues to execute a frontier-based exploration strategy (selecting the frontier point with the highest semantic score) to acquire new observation information and repeat the exploration until the confidence threshold decays to its minimum value or the preset maximum exploration time is reached.

[0037] In the embodiments of the present invention, the decay function is an exponential decay function, which ensures that the threshold decreases monotonically over time, thereby maintaining caution in the early stage of exploration and improving the recall rate of the target in the later stage. The design of the dynamic decay threshold ensures that the system maintains caution in the early stage of exploration (high threshold to prevent false alarms) and improves sensitivity in the later stage of exploration (low threshold to prevent false negatives).

[0038] The entire method of this invention does not rely on any supervised training specific to the navigation task, target object category, or environment. Large-scale language models, open-vocabulary detection models, and visual language models are all pre-trained on public datasets and invoked in this method with zero samples. This method is applicable to various indoor service robot scenarios such as homes, offices, and warehouses, and can handle various household items specified by the user that were never seen during the training phase, achieving open-set object navigation.

[0039] This invention proposes a zero-shot instance navigation method for robots based on prior-driven and multi-evidence fusion. It operates in a zero-shot manner without any supervised training or fine-tuning specific to the navigation task or environment, and can be directly deployed in new, unknown environments with new target object descriptions. This invention utilizes the common-sense reasoning capabilities of a large language model to generate scene context priors for zero-shot object navigation; it employs a joint "target-context" detection and multi-evidence adaptive fusion mechanism to effectively overcome the vulnerability of single visual detection signals in complex environments; and it uses a dynamic decision threshold based on exploration progress adaptation to optimize the balance between success rate and efficiency during the search process. Through the organic combination of these technologies, this invention enables robots to stably and efficiently locate new objects specified by natural language in unknown indoor environments without any scene- or task-specific training.

[0040] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0041] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A zero-shot instance navigation method for robots based on prior-driven and multi-evidence fusion, characterized in that, Includes the following steps: Step S1: Input the natural language navigation instructions into the large language model to generate a list of contextual prior objects that are strongly related to the target object in the natural language navigation instructions; Step S2: Real-time acquisition of visual observation images containing RGB and depth information collected by the robot as it moves in the unknown environment based on the exploration strategy; Step S3: Using a zero-shot open vocabulary detection model, candidate target objects and context prior objects are detected in parallel in the visually observed image; Step S4: Combine the direct semantic alignment score of the candidate target object itself with the confidence evidence scores of all relevant context objects in its spatial neighborhood to calculate the comprehensive confidence score of each candidate target object. Step S5: Compare the comprehensive confidence score with a confidence threshold that decreases adaptively over time to determine whether the current candidate target object is the target object in the natural language navigation instruction, and control the robot to stop or continue exploring based on the determination result.

2. The method according to claim 1, characterized in that, Step S1 specifically includes: Step S101: Combine the natural language navigation instructions containing the description of the target object with the preset prompt word template to form a query for the large language model, so as to stimulate the common sense reasoning ability of the large language model to generate relevant context objects. Step S102: Parse and clean the text output of the large language model, extract a set of specific object names, and form an initial candidate context object set. Step S103: Calculate the semantic similarity between each candidate context object and the description of the target object using a pre-trained visual language model, and filter and sort the initial candidate context object set based on the semantic similarity score to generate the final context prior object list.

3. The method according to claim 1, characterized in that, Step S4 specifically includes: Step S401: For each detected candidate target object, calculate its direct semantic alignment score with the description of the natural language navigation instructions; Step S402: Delineate a spatial neighborhood centered on the candidate target object, and identify all detection instances belonging to the context prior object list located within the spatial neighborhood; Step S403: For each detected instance identified in the spatial neighborhood, calculate its matching score with the context prior object list and its direct semantic score with the target object description. Step S404: Aggregate the matching scores and direct semantic scores of all detected instances to obtain the confidence evidence score of the candidate target object; Step S405: The direct semantic alignment score and the confidence evidence score are fused using a nonlinear weighting function to obtain the comprehensive confidence score for each candidate target object.

4. The method according to claim 1, characterized in that, Step S5 specifically includes: The confidence threshold gradually decreases according to the time or steps already spent, following a predefined decay function. When the overall confidence score of any candidate target object exceeds the confidence threshold at the current moment, the candidate target object is determined to be the target object in the natural language navigation command, and the robot plans a path to move towards it and eventually stops. If no candidate target object triggers confirmation during the entire exploration process, the robot continues to execute the exploration strategy until the confidence threshold decays to the minimum value or the exploration ends.

5. The method according to claim 1, characterized in that, In step S2, the exploration strategy is a hybrid strategy based on semantic value map and frontier detection; the semantic value map is constructed by calculating the global similarity between the current visual observation and the description of the natural language navigation instructions, and is used to guide the robot to the area more relevant to the description.

6. The method according to claim 1, characterized in that, In step S3, each frame of visual observation image, along with the target object description and the context prior object list, is input into the open vocabulary detection model, which outputs all detected target boxes, categories, and their corresponding confidence scores in the visual observation image.