Zero-sample target navigation method and device, mobile robot and storage medium

By responding to natural language commands with a mobile robot, determining the target search space and generating navigation priorities, and collecting and observing images to identify target objects, the problem of unstable and inefficient navigation in complex environments is solved, and efficient and accurate target localization is achieved.

CN121804483APending Publication Date: 2026-04-07UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In complex or heavily occluded environments, existing zero-shot target navigation methods are easily affected by changes in lighting, object occlusion, or background interference, leading to unstable navigation decisions and low exploration efficiency.

Method used

By responding to natural language requests and commands through a mobile robot, the target search space is determined and environmental data is acquired. Navigation priorities are generated based on semantic relevance, and target objects are identified by collecting observation images, achieving efficient localization without task-specific training.

Benefits of technology

It improves the accuracy and efficiency of target object detection in unknown environments, solves the problems of navigation misjudgment and low exploration efficiency, and ensures that the robot can find target objects efficiently and accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121804483A_ABST
    Figure CN121804483A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a zero-sample target navigation method and device, a mobile robot and a storage medium, and relates to the technical field of agent navigation. The method comprises the steps of determining whether an environment map of a target search space is stored or not in response to a natural language request instruction of a target search task; if the environment map is not stored, acquiring environment data of the target search space, and determining an environment area based on the environment data; determining the semantic correlation between the target search task and each environment area, determining the navigation priority of each environment area, and generating a target moving path based on the navigation priority; when the vehicle travels to the target environment area based on the target moving path, collecting an observation image of the target environment area, and determining whether a target object exists in the target environment area based on the observation image; and if the target object exists, completing the target searching task. According to the scheme, the ability of the robot to accurately and efficiently position the target object in an unknown environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of agent navigation, and in particular to a zero-shot target navigation method and device, a mobile robot and a storage medium. BACKGROUND

[0002] In the field of agent navigation, zero-shot object navigation (ZSON) is considered as one of the key tasks to realize embodied AI. The core goal is to enable a mobile robot to autonomously understand semantic intent, explore the environment and successfully locate and reach the target object in a completely unfamiliar environment, only relying on natural language form task instructions (e.g., find a chair), without relying on training data and environment prior for the scene or target category.

[0003] At present, the user's input natural language instruction is usually semantically matched with the image collected in the environment based on a pre-trained visual-language model, and a navigation strategy is generated according to the matching result. This kind of method can realize basic target search function in part structured scene. However, in complex or severely occluded actual environment, since the matching process relies on image representation under limited visual angle, its semantic judgment is easily affected by light change, object occlusion or background interference, resulting in unstable navigation decision; at the same time, in order to cover potential target area, it is often necessary to perform traversal exploration on a large number of environment areas, and the overall efficiency is low. SUMMARY

[0004] The present application provides a zero-shot target navigation method, device, mobile robot and storage medium, to solve the problems of navigation misjudgment and low exploration efficiency caused by relying only on regional semantic matching, so as to improve the ability of the robot to accurately and efficiently locate the target object in an unknown environment based on natural language instructions without task-specific training.

[0005] According to an aspect of the present application, a zero-shot target navigation method is provided, which is executed by a mobile robot, and the method comprises:

[0006] In response to a natural language request instruction of a target search task, determining a target search space matched with the target search task, and determining whether an environment map of the target search space is stored;

[0007] In the case where it is determined that the environment map is not stored, acquiring environment data of the target search space, and determining at least one environment region based on the environment data;

[0008] determine semantic correlations between the target search task and each of the environment regions, determine navigation priorities of each of the environment regions based on the semantic correlations, and generate a target movement path based on the navigation priorities;

[0009] when traveling to a target environment region based on the target movement path, collect an observation image of the target environment region, and determine whether a target object exists in the target environment region based on the observation image; if the target object exists, the target search task is completed; wherein the target object is a search object of the target search task.

[0010] According to another aspect of the present application, a zero-sample target navigation device is provided, which is deployed in a mobile robot, and the device comprises:

[0011] a first determining module configured to determine a target search space matched with the target search task in response to a natural language request instruction of the target search task, and determine whether an environment map of the target search space is stored;

[0012] a second determining module configured to acquire environment data of the target search space and determine at least one environment region based on the environment data when it is determined that the environment map is not stored;

[0013] a movement path generating module configured to determine semantic correlations between the target search task and each of the environment regions, determine navigation priorities of each of the environment regions based on the semantic correlations, and generate a target movement path based on the navigation priorities;

[0014] a third determining module configured to collect an observation image of a target environment region when traveling to the target environment region based on the target movement path, and determine whether a target object exists in the target environment region based on the observation image; if the target object exists, the target search task is completed; wherein the target object is a search object of the target search task.

[0015] According to another aspect of the present application, a mobile robot is provided, which comprises:

[0016] at least one processor; and

[0017] a memory in communication connection with the at least one processor; wherein

[0018] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the zero-sample target navigation method according to any one of the embodiments of the present application.

[0019] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for causing a processor to implement the zero-shot goal navigation method according to any of the embodiments of the present application when executed.

[0020] According to another aspect of the present application, there is provided a computer program product comprising a computer program for implementing the zero-shot goal navigation method according to any of the embodiments of the present application when executed by a processor.

[0021] The technical solution of the embodiments of the present application can determine a target search space matched with a target search task in response to a natural language request instruction of the target search task, and determine whether an environment map of the target search space is stored; in the case that the environment map is not stored, environment data of the target search space is acquired, and at least one environment region is determined based on the environment data; semantic correlations between the target search task and each of the environment regions are determined, navigation priorities of each of the environment regions are determined based on the semantic correlations, and a target movement path is generated based on the navigation priorities; when traveling to a target environment region based on the target movement path, an observation image of the target environment region is collected, and whether a target object exists in the target environment region is determined based on the observation image; if the target object exists, the target search task is completed; wherein the target object is a search object of the target search task, and the above-mentioned technical solution can solve the problems of navigation misjudgment and low exploration efficiency caused by relying on only region-level semantic matching, thereby improving the ability of a robot to accurately and efficiently locate a target object in an unknown environment based on a natural language instruction without task-specific training.

[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0024] Figure 1 is a flowchart of a zero-shot goal navigation method according to an embodiment of the present application;

[0025] Figure 2is a flow chart of a zero-sample target navigation method according to Embodiment Two of the present application;

[0026] Figure 3 is a structural schematic diagram of a zero-sample target navigation device according to Embodiment Three of the present application;

[0027] Figure 4 is a structural schematic diagram of a mobile robot implementing a zero-sample target navigation method according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.

[0030] Embodiment One

[0031] Figure 1 is a flow chart of a zero-sample target navigation method according to Embodiment One of the present application. The present embodiment can be applicable to the case of searching for a zero-sample target object by a mobile robot. The method can be performed by a zero-sample target navigation device, which can be implemented in the form of hardware and / or software. The zero-sample target navigation device can be configured in a mobile robot, which can be an industrial robot, a service robot, a wheeled robot, a tracked robot, or the like. As shown in Figure 1 the method comprises:

[0032] Step 110, in response to a natural language request instruction of a target search task, determining a target search space matched with the target search task, and determining whether an environment map of the target search space is stored.

[0033] The natural language request instructions for the target search task can be navigation instructions issued by the user in natural language form, containing semantic information about the target object and / or target area, such as "Go to the kitchen to find the water glass" or "Where are the keys?". This embodiment does not limit these instructions.

[0034] The target search space refers to the potential environmental scope that is related to the semantics of the natural language request instruction of the target search task. For example, when the instruction mentions the kitchen, the target search space is the set of areas in the environment that are identified or predefined as the kitchen.

[0035] An environment map can be a pre-built and stored digital representation of the environment that includes geometric structures (e.g., occupancy grids) and / or semantic annotations (e.g., room type, location of key objects).

[0036] Optionally, in this embodiment, after receiving a natural language request instruction for a target search task, the mobile robot can parse the request instruction, determine the target search space, and further determine whether its internal storage module stores an environmental map of the target search space.

[0037] In one optional implementation of this embodiment, the mobile robot can perform semantic parsing on the input natural language request command to extract the target object name and target area keywords; further, based on an existing global semantic map or region naming database, it can match one or more candidate environment regions that are semantically consistent with the keywords to form a target search space; finally, it queries the system storage module to determine whether the environment map corresponding to the target search space already exists; if a map has been built, it can be loaded and used directly; if no map has been built, it can trigger subsequent online mapping or exploration processes.

[0038] For example, when a user issues a natural language command to the service robot, "Please go to the living room to find my remote control," the service robot first parses the command, identifying the target area keywords "living room" and the target object "remote control." Further, based on an existing semantic environment database, it matches and determines the physical area in the environment labeled as the living room as the target search space for this task. Further, it automatically queries the storage module to determine if a corresponding environmental map already exists for the living room area; if a map has already been created, it is directly loaded and used; if no map has been created, it is marked as not stored to trigger subsequent online mapping or exploration processes, thereby providing the necessary environmental context for zero-shot navigation.

[0039] Step 120: If it is determined that no environmental map is stored, obtain the environmental data of the target search space, and determine at least one environmental area based on the environmental data.

[0040] The environmental data can be raw perception information, such as point cloud, depth map, image or laser scanning data, etc., collected by sensors (e.g., lidar, depth camera or vision system) carried by the mobile robot in the target search space.

[0041] The environmental area can be a sub-area with semantic or geometric consistency obtained by performing semantic segmentation, clustering or spatial division on the environmental data, for example, a room, a functional area (e.g., dining table area, sofa area or reading area, etc.) or a navigable area defined by a key observation point.

[0042] Optionally, in the embodiment, when it is determined that the environmental map of the target search space has not been stored, the mobile robot collects raw environmental data in the space by a sensor (e.g., lidar), and synchronously constructs a geometric structure (e.g., occupancy grid map) based on the data. Further, the space is initially divided into areas; for example, the entire target search space can be divided into several environmental areas that do not overlap each other and have independent semantic or functional significance (e.g., a room is divided into an entrance area, a central area, a corner area or directly in the form of a room) according to spatial connectivity, wall segmentation or clustering algorithm.

[0043] In an example of the embodiment, when the robot receives the instruction "go to the balcony to find the mop" and finds that the balcony has not been mapped, it will enter the balcony area, obtain point cloud data by laser scanning, and construct a local two-dimensional occupancy grid map. Further, the entire balcony is initially divided into two independent environmental areas according to the spatial boundary: the balcony entrance area near the door and the drying area on the outer side; these two areas are used as the basic environmental areas for subsequent calculation of semantic correlation (e.g., the mop is more likely to be in the drying area) and generation of a navigation path in this task.

[0044] Step 130, determine the semantic correlation between the target search task and each environmental area, determine the navigation priority of each environmental area based on each semantic correlation, and generate a target movement path based on the navigation priority.

[0045] The semantic correlation can be the matching degree between the natural language request instruction of the target search task and each environmental area at the semantic level, which can be calculated by an optimized machine learning model.

[0046] The navigation priority can represent the order of visiting each environmental area, which is determined by the degree of semantic correlation; the higher the correlation, the higher the priority; the target movement path can be a continuous and passable robot motion trajectory generated according to the navigation priority and covering all or part of the environmental areas.

[0047] Optionally, in this embodiment, representative observation data (e.g., keyframe images or semantic summaries) can be extracted for each environmental region, and input along with natural language request instructions into a pre-trained visual-language model to calculate the semantic relevance score of that region. Further, regions can be sorted from high to low scores to form a navigation priority sequence. Finally, starting from the robot's current position, representative positions (e.g., geometric centers or entry points) of each environmental region are connected sequentially according to this priority sequence to generate a complete, obstacle-avoiding, and coherent movement path, which serves as the navigation instruction for the robot to perform its search task.

[0048] In an optional implementation of this embodiment, during the process of determining the semantic relevance between the target search task and each environmental region, and thus generating the target movement path, the semantic representation of each environmental region can be extracted. Specifically, if key observation frames (e.g., RGB images or point clouds with semantic labels) have been saved during the mapping stage, the most representative key frame is selected as the visual input for that region. If it is a newly segmented region in real time, image fragments covering the center or main field of view of that region are extracted from the currently collected observation data. Further, the natural language request command input by the user (e.g., find a red water glass) and the images corresponding to the above-mentioned regions are respectively fed into a pre-trained visual-language model with frozen parameters (e.g., CLIP). The model generates normalized text feature vectors and image feature vectors through its text encoder and image encoder, respectively, and calculates the cosine similarity between the two as the semantic relevance score of the environmental region. Furthermore, all environmental regions are sorted in descending order according to their semantic relevance scores to form a navigation priority sequence. Based on this, and combined with the robot's current pose, the representative positions of each environmental region in the priority sequence are used as waypoints. The underlying motion planning module is called to plan local passable paths segment by segment and then splice them into a continuous, collision-free, complete target movement path.

[0049] In this embodiment, the generated target movement path covers the entire environmental area and the access order strictly follows semantic priority, thereby ensuring that the robot prioritizes exploring the area most likely to contain the target object.

[0050] Step 140: When traveling to the target environment area based on the target movement path, collect observation images of the target environment area, and determine whether there is a target object in the target environment area based on the observation images; if there is a target object, the target search task is completed.

[0051] In this embodiment, the target object is the object to be searched in the target search task; the category of the target object is not pre-trained in the system and needs to be identified by zero-shot semantic understanding capability.

[0052] The target environment area can be any environment area determined through the above steps, and is not limited to it in this embodiment. The observed image of the target environment area can be one or more frames of images collected in real time by the mobile robot through its onboard visual sensors after entering the target environment area, used for fine-grained perception of the current scene.

[0053] In one optional implementation of this embodiment, when the mobile robot travels along the target movement path to the current target environment area, one or more frames of observation images are acquired. Further, a natural language request instruction (such as "find a blue backpack") is used as a text prompt and input into an open lexical object recognition model that supports zero-shot inference (e.g., a CLIP-based region proposal network). This model generates candidate object bounding boxes in the observation images and outputs a semantic matching score with the text prompt for each candidate box. If at least one candidate box has a score higher than a preset confidence threshold and its spatial scale and physical plausibility meet expectations, it is determined that a target object exists in the target environment area, and the task is successfully completed. Otherwise, it is determined that no target object was found, and the robot continues to the next environment area to perform the same verification process.

[0054] In another optional implementation of this embodiment, when the mobile robot travels along the target movement path to the target environment area with the highest current priority, it can perform short-distance adjustments or local circling movements within that area to acquire multiple frames of observation images covering different perspectives. Furthermore, each frame of observation image can be divided into several local image blocks, and the natural language request command and each image block are input into a pre-trained visual-language model to calculate the semantic matching score between each image block and the command. The maximum matching score among all image blocks is selected as the final verification score for the target environment area. If the verification score is greater than or equal to a preset confidence threshold, it is determined that a target object exists in the area, the target search task is immediately terminated, and a success status and target location information can be returned to the user. Otherwise, it is determined that the target object does not exist, and the robot continues along the target movement path to the next priority environment area, repeating the above acquisition and verification process until all areas have been traversed or the target object has been successfully found.

[0055] The technical solution of this embodiment involves a mobile robot responding to a natural language request command for a target search task, determining a target search space matching the target search task, and determining whether an environmental map of the target search space is stored. If no environmental map is stored, environmental data of the target search space is acquired, and at least one environmental region is determined based on the environmental data. The semantic relevance between the target search task and each environmental region is determined, and the navigation priority of each environmental region is determined based on the semantic relevance. A target movement path is generated based on the navigation priority. When the robot travels to the target environmental region based on the target movement path, an observation image of the target environmental region is acquired, and the presence of a target object in the target environmental region is determined based on the observation image. If a target object exists, the target search task is completed. Here, the target object is the search object of the target search task. This solution can solve the problems of navigation misjudgment and low exploration efficiency caused by relying solely on region-level semantic matching, thereby improving the robot's ability to accurately and efficiently locate target objects in unknown environments based on natural language commands without requiring task-specific training.

[0056] Example 2

[0057] Figure 2 This is a flowchart of a zero-sample target navigation method according to Embodiment 2 of the present invention. This embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the method includes:

[0058] Step 210: In response to the natural language request instruction of the target search task, determine the target search space that matches the target search task, and determine whether an environment map of the target search space is stored.

[0059] Optionally, in this embodiment, responding to a natural language request instruction for a target search task, determining a target search space matching the target search task, and determining whether an environmental map of the target search space is stored, may include: performing semantic parsing on the natural language request instruction to obtain a first semantic parsing result, and determining whether the first semantic parsing result contains spatial range information; if the first semantic parsing result contains spatial range information, extracting the spatial range information, and determining the target search space based on the spatial range information; if the first semantic parsing result does not contain spatial range information, using the physical area where the mobile robot is currently located or a preset exploration range as the target search space; and querying the mobile robot's map storage module to determine whether the map storage module stores an environmental map corresponding to the target search space.

[0060] The first semantic parsing result is the parsing result of the natural language request. In this embodiment, it is named the first semantic parsing result for ease of description, but it is not a limitation of this embodiment.

[0061] Optionally, in this embodiment, after the mobile robot interprets the natural language request instruction regarding the target search task, it can call a lightweight semantic parser (e.g., a fine-tuning module based on rule templates, named entity recognition models, or large language models) to parse the natural language request instruction and extract a structured first semantic parsing result. Further, it can determine whether the first semantic parsing result contains explicit spatial range information. If it does (e.g., parsing out the location entity "kitchen"), the location is used as a keyword to match the corresponding physical area coordinate range in the robot's semantic map index, and the target search space is determined accordingly. If it does not contain spatial range information (e.g., the natural language request instruction is only to find car keys), the room-level area where the mobile robot is currently located or the system's preset default exploration range (e.g., the known accessible area of ​​the entire house) is used as the target search space. Finally, the mobile robot can query the map storage module to check whether there is an environmental map that is consistent with or covers the target search space in terms of spatial range, in order to decide whether to skip the mapping stage and directly enter the navigation verification process.

[0062] The solution in this embodiment intelligently distinguishes whether spatial range information is included by analyzing the semantic parsing results of natural language commands, and dynamically determines the target search space accordingly. This avoids the problem of blindly expanding the search range when there are no regional limitations or ignoring the semantic context when there are clear location guidelines. At the same time, by querying the map storage module to determine whether the environment map exists, the overhead of repeated map building is effectively reduced, and the response efficiency and resource utilization in mixed known and unknown environments are improved.

[0063] Step 220: If it is determined that no environmental map is stored, obtain the environmental data of the target search space and determine at least one environmental area based on the environmental data.

[0064] Optionally, in this embodiment, if it is determined that no environmental map is stored, acquiring environmental data of the target search space and determining at least one environmental region based on the environmental data may include: collecting three-dimensional point cloud data of the target search space through a lidar sensor mounted on the mobile robot, and filtering each three-dimensional point cloud data to obtain a target point cloud; wherein, the target point cloud is a point cloud within the traversable height range of the mobile robot; projecting the target point cloud onto a two-dimensional plane to generate a top-down occupancy grid map, and marking the occupancy grid map with connected components to obtain continuous initial regions; determining the size information of each initial region, and if it is determined that the target size of the target initial region is greater than a preset size threshold, determining the local geometric features of the target initial region, and dividing the target initial region based on the local geometric features to obtain each divided region; and determining each divided region and each initial region as each environmental region.

[0065] The local geometric features include at least one of the following: passage, door, or wall.

[0066] In this embodiment, the occupied grid map can be a discretized map generated by orthogonally projecting the target point cloud onto a horizontal two-dimensional plane, where each grid can be marked as , free, occupied, or unknown, and can be used to describe the accessibility of the environment.

[0067] Step 230: Determine the semantic relevance between the target search task and each environmental region.

[0068] Optionally, in this embodiment, determining the semantic relevance between the target search task and each environmental region may include: acquiring at least one target observation image corresponding to each environmental region; inputting each target observation vector and natural language request into a pre-trained visual-language model to obtain a confidence score that each environmental region contains a target object; and determining the confidence score as the semantic relevance between the target search task and each environmental region.

[0069] The pre-trained visual-language model can be a multimodal model trained on large-scale image-text pairs (e.g., CLIP), whose visual and text encoders have frozen parameters and are used only for zero-shot inference.

[0070] The confidence score is a scalar value output by the model, reflecting the semantic probability that the corresponding environmental area contains the target object. The higher the value, the stronger the match. In this embodiment, the score is directly used as semantic relevance as the basis for subsequent navigation priority ranking.

[0071] In an optional implementation of this embodiment, the mobile robot can acquire target observation images corresponding to each environmental region. These target observation images can be keyframe images collected and cached within the corresponding region during the mobile robot's mapping or exploration process. Further, each target observation image and the user's input natural language request command are input into a pre-trained visual-language model. The visual encoder within this model processes the images to generate image feature vectors, and the text encoder processes the natural language request to generate text feature vectors. Based on the similarity between the two feature vectors (such as cosine similarity), a scalar value is output as the confidence score that the environmental region contains the target object. Finally, the confidence score corresponding to each environmental region is directly determined as the semantic relevance between the target search task and the environmental region.

[0072] For example, when a user gives the command "find a blue backpack" to the mobile robot, the mobile robot determines that the indoor environment has been pre-divided into three environmental areas: a living room area, a bedroom area, and an entryway area, and saves a key observation image for each area. Further, these three images, along with the text prompt "a blue backpack," are input into a pre-trained visual-language model. The model performs joint encoding on each image and text, and calculates three sets of image-text similarity scores, which are 0.28 (living room), 0.35 (bedroom), and 0.76 (entryway). These scores serve as the confidence scores for each area containing the target object and are directly used as the semantic relevance of the task to each area.

[0073] Step 240: Determine the navigation priority of each environmental region based on the semantic relevance, and generate the target movement path based on the navigation priority.

[0074] Optionally, in this embodiment, determining the navigation priority of each environmental region based on semantic relevance and generating a target movement path based on the navigation priority may include: sorting all environmental regions from high to low according to confidence scores to obtain a navigation priority sequence; and connecting all environmental regions sequentially according to the sorting order based on the navigation priority sequence to generate a continuous movement path that traverses all environmental regions as the target movement path.

[0075] The target movement path starts from the current position of the mobile robot, passes through representative positions of each environmental area in sequence, and is a continuous trajectory that covers the entire area and satisfies the feasibility of movement.

[0076] In one optional implementation of this embodiment, after determining the semantic relevance between the target search task and each environmental region, all environmental regions can be sorted in descending order according to their corresponding confidence scores to form a navigation priority sequence, wherein the higher the confidence score, the earlier the corresponding environmental region is in the sequence. Furthermore, the real-time pose of the mobile robot in the global coordinate system can be used as the starting point for path planning, and the representative position of the first environmental region in the navigation priority sequence can be used as the first target waypoint. Further, this starting point and the first waypoint are input into the underlying path planning module, and a graph search-based algorithm is run on the constructed grid map to generate a continuous and feasible local path that avoids obstacles. Using the previous waypoint as a new starting point, the representative positions of subsequent environmental regions in the navigation priority sequence are connected sequentially, and the local path planning process is repeated, generating traversable trajectories between adjacent waypoints segment by segment. All local path segments are spatially connected end-to-end, ultimately forming a complete movement path starting from the robot's current position, strictly following the semantic priority order, and covering all environmental regions. This movement path is output as the target movement path to the motion control module, which drives the mobile robot to sequentially visit each environmental region to perform the target search task.

[0077] For example, when the mobile robot is located in the center of the living room (current position), four environmental areas and their confidence scores are determined: kitchen (0.82), balcony (0.65), master bedroom (0.41), and bathroom (0.33). After sorting the scores from high to low, the navigation priority sequence is kitchen-balcony-master bedroom-bathroom. The path planning module takes the mobile robot's current position as the starting point and connects the representative positions of each area in sequence (e.g., inside the kitchen door, balcony entrance, master bedroom door, bathroom door). It calculates the obstacle avoidance path segment by segment on the occupied grid map and splices it to generate a continuous target movement path, guiding the robot to traverse all areas in this order.

[0078] Step 250: When traveling to the target environment area based on the target movement path, collect observation images of the target environment area, and determine whether there is a target object in the target environment area based on the observation images.

[0079] Optionally, in this embodiment, when the robot travels to the target environment area based on the target movement path, acquiring observation images of the target environment area and determining whether a target object exists in the target environment area based on the observation images may include: acquiring observation images of the target environment area from different angles using an image sensor when the mobile robot travels to the target environment area; inputting each observation image and the natural language request into a pre-trained visual-language model to obtain multiple matching scores that characterize the semantic matching degree between the natural language request instruction and the corresponding observation image; determining the maximum matching score among the multiple matching scores; when the maximum matching score is greater than or equal to a preset confidence threshold, determining the target observation image corresponding to the maximum matching score, and identifying the target object in the target observation image.

[0080] The target environment area can be any environmental area that the mobile robot has reached and is to be verified to contain the target object; the preset confidence threshold is a pre-set judgment boundary used to distinguish between the presence of a target and the absence of a target.

[0081] In one optional implementation of this embodiment, after the mobile robot travels to the target environment area, it can acquire observation images of the area from multiple perspectives using an image sensor; each observation image and the same natural language request are input into a pre-trained visual-language model, and the model calculates the matching score between each image and the instruction; the largest value among all matching scores is selected as the maximum matching score; when the maximum matching score is greater than or equal to a preset confidence threshold, the observation image corresponding to the score is determined as the target observation image, and the existence of the target object is located or confirmed in the image, thereby completing the target verification of the current area.

[0082] In one example of this embodiment, after the mobile robot moves to the study area, it slowly rotates around the desk and collects three observation images from different perspectives. Further, these three images, along with the user command "find a black mouse," are input into a pre-trained model, resulting in matching scores of 0.42, 0.68, and 0.53, respectively. The maximum matching score is 0.68, corresponding to the second image. Since this score is higher than the preset confidence threshold of 0.6, the second image is identified as the target observation image, and the black mouse located to the right of the keyboard is identified within it, confirming the existence of the target object.

[0083] In another optional implementation of this embodiment, after determining the observation image corresponding to the maximum matching score, the observation image can be further input into a target recognition model that supports open vocabulary or zero-shot capability. The model uses natural language requests as text prompts and directly generates the bounding box and category confidence of the target object in the image. If the model outputs at least one valid detection result (i.e., the bounding box area is greater than the preset minimum size and the positioning is reasonable), it is determined that the target object exists in the current environment area and the task is completed. If no valid detection result is output, it is determined that the target object does not exist and the process continues to the next environment area.

[0084] Step 260: After determining that there is no target object in the target environment area, travel along the target movement path to the next environment area, and repeat the steps of collecting and observing images and determining whether there is a target object in the next environment area until all environment areas along the target movement path have been traversed or the target object has been successfully found.

[0085] In one optional implementation of this embodiment, after verifying that no target object exists in the current target environment area, the mobile robot continues to travel along the generated target movement path to the next environment area in the path. Upon arrival, it acquires observation images of the area from multiple perspectives using an onboard image sensor, and inputs each image along with the original natural language request into a pre-trained visual-language model to calculate a matching score. The highest matching score is selected and compared with a preset confidence threshold to determine whether the target object exists. If it still does not exist, it continues to move to subsequent areas and repeats the same verification operation. This process continues until the mobile robot has completed its visit to all environment areas in the target movement path, or successfully confirmed the existence of the target object in a certain area, thereby terminating the search task.

[0086] For example, after verifying the remote control in the living room area, the mobile robot does not find it and then moves along the target movement path to the next high-priority area, the TV cabinet area. It collects three images from different angles in the TV cabinet area. After model matching, the maximum score is 0.71, which exceeds the threshold of 0.65. The target is determined to exist, and the task ends. If it is still not found, it continues to verify the sofa area and coffee table area in turn until the end of the path or the target is found.

[0087] The solution in this embodiment automatically moves to the next environmental area along the pre-planned path and repeats the target verification process after verification failure. This achieves orderly and closed-loop exploration of multiple candidate areas, avoiding premature termination of the task due to single misjudgment or partial occlusion. At the same time, using the completion of traversal or successful finding of the target as a clear termination condition ensures the integrity and robustness of the search process, improving the success rate and task reliability of zero-sample target navigation in complex environments without human intervention.

[0088] To better understand the zero-sample target navigation method involved in this embodiment, an example is used below to describe it, which mainly includes:

[0089] In the environmental modeling phase, the robot first collects 3D point cloud data using a LiDAR sensor. Through height thresholding and outlier removal, ground noise and invalid point clouds are eliminated, retaining only points within the passable height range. The processed point cloud is then projected onto a 2D plane, forming a top-down occupancy grid. Within this grid, connected component labeling methods are used to identify continuous passable areas, and excessively large areas are further subdivided through local geometric analysis (such as bottleneck detection of narrow passages and doorways), resulting in spatially semantically consistent partitioned units. These partitions are dynamically updated and merged during the exploration process, ensuring that the structural representation remains consistent with the robot's perception of the environment.

[0090] In the semantic reasoning phase, a two-tiered strategy of partition-level and pixel-level semantic reasoning is employed. Within each partition, the robot acquires several representative keyframes, automatically selected based on viewpoint coverage and image dissimilarity to ensure they reflect the overall characteristics of the partition. These keyframes are input into a pre-trained vision-language model, matched against task instructions (e.g., finding a refrigerator), and output a confidence score indicating that the partition contains the target. Simultaneously, pixel-level semantic estimation is continuously run during frame-by-frame navigation, generating local semantic score maps. These maps are then projected onto a global grid through viewpoint projection to provide more refined local semantic cues.

[0091] In the decision-making stage, a semantic utility-based fusion mechanism is proposed. By weighted combining partition-level scores and pixel-level scores, a comprehensive utility function is obtained: ;

[0092] in Represents candidate location units Overall utility, parameters Adaptively adjusts based on observation coverage in different regions. Within the reachable region set. In this scenario, the robot selects the position with the highest utility as its next target. .

[0093] In this way, the robot can prioritize entering high-probability areas while ensuring global semantic consistency, and avoid redundant exploration in low-yield partitions by combining local detailed information.

[0094] The solution of this invention introduces a spatial partitioning mechanism based on LiDAR, enabling the robot to navigate using structural units such as rooms or corridors. This avoids ineffective frame-by-frame traversal in large-scale environments, thereby reducing redundant paths and computational overhead. At the semantic reasoning level, this invention performs partition-level estimation through keyframe selection, combined with pixel-level semantic mapping, forming a multi-scale fusion reasoning that considers both global and local aspects. This allows the robot to maintain a high target detection rate even when the scene is occluded or visually blurred.

[0095] Example 3

[0096] Figure 3 This is a schematic diagram of a zero-sample target navigation device according to Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a first determining module 310, a second determining module 320, a moving path generating module 330, and a third determining module 340.

[0097] The first determining module 310 is used to determine the target search space that matches the target search task in response to the natural language request instruction of the target search task, and to determine whether an environmental map of the target search space is stored.

[0098] The second determining module 320 is used to obtain environmental data of the target search space when it is determined that the environmental map is not stored, and to determine at least one environmental area based on the environmental data.

[0099] The movement path generation module 330 is used to determine the semantic relevance between the target search task and each of the environmental regions, determine the navigation priority of each of the environmental regions based on the semantic relevance, and generate a target movement path based on the navigation priority.

[0100] The third determining module 340 is used to acquire observation images of the target environment area when traveling to the target environment area based on the target movement path, and determine whether a target object exists in the target environment area based on the observation images; if the target object exists, the target search task is completed; wherein, the target object is the search object of the target search task.

[0101] In this embodiment, the solution involves a first determining module responding to a natural language request command for a target search task, determining a target search space matching the target search task, and determining whether an environmental map of the target search space is stored. If the second determining module determines that the environmental map is not stored, it acquires environmental data of the target search space and determines at least one environmental region based on the environmental data. A movement path generation module determines the semantic relevance between the target search task and each of the environmental regions, determines the navigation priority of each environmental region based on the semantic relevance, and generates a target movement path based on the navigation priority. A third determining module, when navigating to the target environmental region based on the target movement path, acquires an observation image of the target environmental region and determines whether a target object exists in the target environmental region based on the observation image. If the target object exists, the target search task is completed. The target object is the object being searched for in the target search task. This solution addresses the navigation misjudgment and low exploration efficiency caused by relying solely on region-level semantic matching, thereby improving the robot's ability to accurately and efficiently locate target objects in unknown environments based on natural language commands without requiring task-specific training.

[0102] In an optional implementation of this embodiment, the first determining module 310 is specifically used to perform semantic parsing on the natural language request instruction to obtain a first semantic parsing result, and to determine whether the first semantic parsing result contains spatial range information;

[0103] If it is determined that the first semantic parsing result contains spatial range information, the spatial range information is extracted, and the target search space is determined based on the spatial range information;

[0104] If it is determined that the first semantic parsing result does not contain spatial range information, the physical area where the mobile robot is currently located or the preset exploration range shall be used as the target search space;

[0105] The map storage module of the mobile robot is queried to determine whether the map storage module stores an environmental map corresponding to the target search space.

[0106] In an optional implementation of this embodiment, the second determining module 320 is specifically used to collect three-dimensional point cloud data of the target search space by using a lidar sensor mounted on the mobile robot, and to filter each of the three-dimensional point cloud data to obtain a target point cloud; wherein, the target point cloud is a point cloud within the traversable height range of the mobile robot;

[0107] The target point cloud is projected onto a two-dimensional plane to generate a top-down occupied grid map, and the occupied grid map is labeled with connected components to obtain a continuous initial region.

[0108] The size information of each initial region is determined. If the target size of the initial region is greater than a preset size threshold, the local geometric features of the initial region are determined, and the initial region is divided based on the local geometric features to obtain each divided region. The local geometric features include at least one of the following: a passage, a door, or a wall.

[0109] Each of the defined regions and each of the defined initial regions are defined as the defined environmental regions.

[0110] In an optional implementation of this embodiment, the movement path generation module 330 is specifically used to acquire at least one target observation image corresponding to each of the environmental regions;

[0111] Each target observation vector and natural language request are input into a pre-trained visual-language model to obtain a confidence score that each environment region contains a target object.

[0112] The confidence score is determined as the semantic relevance between the target search task and each of the environmental regions.

[0113] In an optional implementation of this embodiment, the movement path generation module 330 is further configured to sort all environmental areas from high to low according to the confidence scores to obtain a navigation priority sequence;

[0114] Based on the navigation priority sequence, all environmental areas are connected sequentially in sorted order to generate a continuous movement path that traverses all environmental areas as the target movement path.

[0115] In an optional implementation of this embodiment, the third determining module 340 is specifically used to acquire observation images of the target environment area from different angles using an image sensor when the mobile robot travels to the target environment area;

[0116] Each of the observed images and the natural language request are respectively input into a pre-trained visual-language model to obtain multiple matching scores that characterize the semantic matching degree between the natural language request instruction and the corresponding observed image;

[0117] The maximum matching score among the multiple matching scores is determined. When the maximum matching score is greater than or equal to a preset confidence threshold, the target observation image corresponding to the maximum matching score is determined, and the target object is identified in the target observation image.

[0118] In an optional implementation of this embodiment, the zero-sample target navigation device further includes a traversal module, configured to, after determining that the target object does not exist in the target environment area, travel along the target movement path to the next environment area, and repeatedly execute the steps of acquiring and observing images and determining whether the target object exists in the next environment area, until all environment areas in the target movement path have been traversed or the target object has been successfully found.

[0119] The zero-sample target navigation device provided in the embodiments of the present invention can execute the zero-sample target navigation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0120] In the technical solutions of the embodiments of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of environmental data (such as point cloud data) involved all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0121] Example 4

[0122] Figure 4 A schematic diagram of a mobile robot 10, which can be used to implement embodiments of the present invention, is shown. The mobile robot is intended to represent various forms of digital computers, such as laptops, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The mobile robot can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0123] like Figure 4 As shown, the mobile robot 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from the storage unit 18. The RAM 13 can also store various programs and data required for the operation of the mobile robot 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0124] Multiple components in the mobile robot 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, optical disk, etc.; and a communication unit 19, such as a network card, modem, wireless transceiver, etc. The communication unit 19 allows the mobile robot 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0125] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods described above, such as zero-shot target navigation methods.

[0126] In some embodiments, the zero-sample target navigation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the mobile robot 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the zero-sample target navigation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the zero-sample target navigation method by any other suitable means (e.g., by means of firmware).

[0127] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0128] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0129] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on a mobile robot having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the mobile robot. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0131] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0132] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and Virtual Private Servers (VPS) in terms of management difficulty and weak business scalability.

[0133] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0134] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

[0135] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements a database detection method as provided in any embodiment of this application.

[0136] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LANs or WANs—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0137] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0138] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A zero-shot target navigation method, executed by a mobile robot, characterized in that, The method includes: In response to a natural language request instruction for a target search task, determine a target search space that matches the target search task, and determine whether an environmental map of the target search space is stored. If it is determined that the environmental map is not stored, environmental data of the target search space is obtained, and at least one environmental area is determined based on the environmental data; The semantic relevance between the target search task and each of the environmental regions is determined, the navigation priority of each of the environmental regions is determined based on the semantic relevance, and the target movement path is generated based on the navigation priority. When traveling to the target environment area based on the target movement path, an observation image of the target environment area is acquired, and the presence of a target object in the target environment area is determined based on the observation image; if the target object exists, the target search task is completed; wherein, the target object is the search object of the target search task.

2. The zero-sample target navigation method according to claim 1, characterized in that, The process of responding to a natural language request instruction for a target search task, determining a target search space matching the target search task, and determining whether an environmental map of the target search space is stored includes: The natural language request instruction is semantically parsed to obtain a first semantic parsing result, and it is determined whether the first semantic parsing result contains spatial range information; If it is determined that the first semantic parsing result contains spatial range information, the spatial range information is extracted, and the target search space is determined based on the spatial range information; If it is determined that the first semantic parsing result does not contain spatial range information, the physical area where the mobile robot is currently located or the preset exploration range shall be used as the target search space; The map storage module of the mobile robot is queried to determine whether the map storage module stores an environmental map corresponding to the target search space.

3. The zero-sample target navigation method according to claim 1, characterized in that, The step of acquiring environmental data of the target search space when it is determined that the environmental map is not stored, and determining at least one environmental region based on the environmental data, includes: The target point cloud is obtained by collecting three-dimensional point cloud data of the target search space by a lidar sensor mounted on the mobile robot and filtering each three-dimensional point cloud data; wherein, the target point cloud is the point cloud within the traversable height range of the mobile robot. The target point cloud is projected onto a two-dimensional plane to generate a top-down occupied grid map, and the occupied grid map is labeled with connected components to obtain a continuous initial region. The size information of each initial region is determined. If the target size of the initial region is greater than a preset size threshold, the local geometric features of the initial region are determined, and the initial region is divided based on the local geometric features to obtain each divided region. The local geometric features include at least one of the following: a passage, a door, or a wall. Each of the defined regions and each of the defined initial regions are defined as the defined environmental regions.

4. The zero-sample target navigation method according to claim 1, characterized in that, Determining the semantic relevance between the target search task and each of the environmental regions includes: At least one target observation image corresponding to each of the aforementioned environmental regions is acquired; Each target observation vector and natural language request are input into a pre-trained visual-language model to obtain a confidence score that each environment region contains a target object. The confidence score is determined as the semantic relevance between the target search task and each of the environmental regions.

5. The zero-sample target navigation method according to claim 4, characterized in that, The step of determining the navigation priority of each of the environmental regions based on the semantic relevance, and generating a target movement path based on the navigation priority, includes: All environmental regions are sorted from high to low according to the confidence scores to obtain a navigation priority sequence; Based on the navigation priority sequence, all environmental areas are connected sequentially in sorted order to generate a continuous movement path that traverses all environmental areas as the target movement path.

6. The zero-sample target navigation method according to claim 1, characterized in that, The step of acquiring observation images of the target environment area while traveling along the target movement path to the target environment area, and determining whether a target object exists in the target environment area based on the observation images, includes: When the mobile robot travels to the target environment area, it acquires observation images of the target environment area from different angles through an image sensor; Each of the observed images and the natural language request are respectively input into a pre-trained visual-language model to obtain multiple matching scores that characterize the semantic matching degree between the natural language request instruction and the corresponding observed image; The maximum matching score among the multiple matching scores is determined. When the maximum matching score is greater than or equal to a preset confidence threshold, the target observation image corresponding to the maximum matching score is determined, and the target object is identified in the target observation image.

7. The zero-sample target navigation method according to claim 1, characterized in that, The method further includes: After determining that the target object does not exist in the target environment area, the vehicle travels along the target movement path to the next environment area, and repeats the steps of acquiring and observing images and determining whether the target object exists in the next environment area until all environment areas in the target movement path have been traversed or the target object has been successfully found.

8. A zero-sample target navigation device, deployed on a mobile robot, characterized in that, include: The first determining module is used to respond to a natural language request instruction of the target search task, determine the target search space that matches the target search task, and determine whether an environmental map of the target search space is stored. The second determining module is used to acquire environmental data of the target search space when it is determined that the environmental map is not stored, and to determine at least one environmental area based on the environmental data. A movement path generation module is used to determine the semantic relevance between the target search task and each of the environmental regions, determine the navigation priority of each of the environmental regions based on the semantic relevance, and generate a target movement path based on the navigation priority. The third determining module is used to acquire observation images of the target environment area when traveling to the target environment area based on the target movement path, and determine whether a target object exists in the target environment area based on the observation images; if the target object exists, the target search task is completed; wherein, the target object is the search object of the target search task.

9. A mobile robot, characterized in that, The mobile robot includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the zero-sample target navigation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the zero-sample target navigation method of any one of claims 1-7.