Robot indoor object navigation method and system

By constructing an offline relational scene graph and using vision-language models and large language models, the robot can accurately identify and navigate to instance-level objects in complex and dynamic environments, solving the problem of existing technologies that cannot effectively navigate to everyday objects with unfixed positions, and achieving efficient instance-level object navigation.

CN120628100APending Publication Date: 2025-09-12BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510745747.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing object navigation methods have difficulty effectively navigating to everyday objects such as cups that are not fixed in position and are easily disturbed in complex and dynamic environments. They lack the ability to update the scene and cannot accurately identify and update the state of instance-level objects.

Method used

By constructing a semantic point cloud map to generate an offline hosting relationship scene graph, combined with visual-language models (such as CLIP and SBERT) and large language models (such as GPT-4o), dynamic objects and hosting relationships in the scene are identified and updated. The Markov decision process is used for navigation strategy, common sense reasoning is used to identify potential hosting objects, and the map is dynamically updated.

Benefits of technology

It improves the robot's navigation accuracy and flexibility in complex and dynamic environments, enables it to identify and process unknown objects, dynamically update the state of objects in the scene, and enhances the intelligence and adaptability of navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120628100A_ABST
    Figure CN120628100A_ABST
Patent Text Reader

Abstract

The invention discloses a robot indoor object navigation method and system, relates to the technical field of intelligent robots, and can realize robust instance level object navigation of a robot in an indoor daily environment. According to the technical scheme, the method comprises the following steps: generating an offline bearing relation scene graph from a semantic point cloud map; a navigation target is input, a target object on the map is obtained through matching, and the robot navigates to the target object according to a set navigation strategy. Map updating is conducted in the process that the robot navigates to the target, and if the target is not in the original place, the robot explores the possible target position according to a certain navigation strategy till the target is found.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent robots, and in particular to a method and system for indoor object navigation by robots. Background Art

[0002] With the rapid development of visual language models and large language models, the implementation of robot object navigation is gaining increasing attention. Object navigation requires not only that robots understand complex visual and language information but also that they can flexibly perform navigation tasks in dynamic environments. In an everyday environment, robots need to efficiently navigate to a designated target under various environmental conditions, whether the target is static furniture or a commonly used object with a constantly changing position (such as a cup). To accomplish this task, robots need to possess the following capabilities: first, they must be able to perceive and accurately represent the state of objects in the environment; second, they must be able to dynamically update the position and state of objects, especially when objects are frequently moving; and finally, they must possess efficient navigation capabilities, able to navigate strategically based on the current state of the object, its position changes, and the target location. The implementation of these capabilities is crucial for autonomous robot navigation in complex and dynamic environments.

[0003] Current object navigation methods can effectively navigate to static objects (such as sofas). However, they are usually limited to searching for objects at the semantic level and lack the ability to update the scene. Therefore, for everyday common objects like "cups on black tables", these objects often come in different colors and styles, may appear in different places such as kitchens and bedrooms, and their positions are not fixed. In addition, these objects are often carried by other objects, which means that the position of the carrying objects is also not fixed. Such navigation targets are highly dynamic and easily disturbed, so it is very challenging to navigate to them effectively and efficiently.

[0004] Therefore, it is necessary to design a new indoor object navigation method that can remember and update the latest state of the scene while accurately navigating to instance-level objects. Summary of the Invention

[0005] In view of this, the present invention provides a robot indoor object navigation method and system, which can realize robust instance-level object navigation of the robot in daily indoor environments.

[0006] To achieve the above object, the technical solution of the present invention includes the following steps:

[0007] S1: Generate offline hosting relationship scene graph from semantic point cloud map.

[0008] S2: Input the navigation target, match the target object on the map, and the robot navigates to the target object according to the set navigation strategy.

[0009] S3: The robot updates the map while navigating to the target. If the target is not at the original location, the robot explores possible target locations according to a certain navigation strategy until it finds the target.

[0010] Furthermore, S1: generating an offline hosting relationship scene graph from the semantic point cloud map, the specific steps are as follows:

[0011] S101: Build an open vocabulary instance map using pre-collected scene RGB-D data Add a description list cap generated by the Tokenize Anything model to each instance i and a text feature T_F encoded by the SBERT model i to enhance each instance.

[0012] S102: Construct a hosting relationship scene graph S_G, including a building and room layer, a hosting level layer, and an object layer.

[0013] Furthermore, S102: construct a hosting relationship scene graph S_G, including a building and room layer, a hosting level layer, and an object layer, specifically:

[0014] The building and room layers are constructed as follows: objects in the offline map are divided into different rooms using a priori-based approach. Then, the obtained room layers are merged to form the overall building layer.

[0015] The load-bearing layer is constructed as follows: Calculate each object O i Text feature T_F i SBERT-encoded text features for “furniture used to carry objects” The similarity between them is expressed as follows:

[0016]

[0017] Select the set of objects whose similarity score exceeds the specified threshold σ Its expression is as follows:

[0018]

[0019] Extract each mid-cap i The three most common descriptions are input into a large language model GPT-4o and are used to identify potential objects through specific prompts, denoted as

[0020] Finally, according to the criteria of whether the geometric size of the object exceeds a certain size and whether it is in contact with the ground, the final set of carrying objects is selected and recorded as Specifically:

[0021]

[0022] The object layer is constructed as follows: For any non-bearing layer object Based on O i The size and load-bearing objects The shortest distance between them and the spatial overlap relationship in the xyz direction are used to determine O i Whether it is the object O j Carrying capacity; h(O j ,O i ) is a comprehensive index including the above factors, where h(O j ,O i )=1 if all conditions are met; for any Definition by O j The set of objects C(O j )as follows:

[0023]

[0024] Furthermore, in S2, the navigation target is input and the target object on the map is matched, which specifically includes the following steps:

[0025] The input navigation target is a navigation instruction, specifically a text description or an image; the text description or image is encoded using the SBERT or CLIP model respectively.

[0026] The generated features are compared with the SBERT or CLIP features of each object in S_G using cosine similarity, where the object with the highest similarity score is selected as the target object. target .

[0027] Furthermore, the navigation strategy is specifically as follows:

[0028] The navigation strategy is modeled as a fixed-strategy Markov decision process, which includes:

[0029] State space S:

[0030] At the current step t: the robot's pose

[0031] Unexplored collection of host layer objects

[0032] The set of candidate target objects on the unexplored carrier layer

[0033] The flag indicating whether the target is found is F t∈{0,1}, where and Respectively represent L t , CR t and CT t The value range collection.

[0034] The state variables are:

[0035] S t =(L t ,CR t ,CT t ,F t )∈S

[0036] In the initial state S0=(L0,CR0,CT0,F0), L0 is the initial state of the robot, CT0=O target .

[0037] The action space A is:

[0038] A={Stop,Explore(cr),Goto(ct)|cr∈CR t ,ct∈CT t}

[0039] Among them, Stop means that the task has been completed or all the objects in the carrier layer have been explored, Explore(cr) and Goto(ct) respectively indicate the exploration of the carrier layer object cr∈CR t and navigate to the target object ct∈CT t location.

[0040] The robot is based on the current state S t and a specific policy π(·) to select the next action a in the action space A t ∈A.

[0041] Policy π(·): Given the current state S t =(L t ,CR t ,CT t ,F t ):

[0042] 1. If F t =1 or Then a t =Stop;

[0043] 2. If F t =0 and Prioritize the selection of a candidate object for operation. Specifically, let CT t ={O t1 ,…,O ti}; At the same time, store some additional variables: CT t With the target object O target SBERT similarity SS between t =ss t1 ,…,ss ti , current location L t With CT t The distance D t ={d t1 ,…,d ti}, and CT images observed by the robot camera t The average depth value of Any O tj ∈CT t Corresponding to ss tj d tj and The priority score is assessed in the following way:

[0044]

[0045] Among them, ss tj and P_R(O tj ) positive correlation, ss tj The larger it is, the higher the probability that the candidate is the target; and P_R(O tj ) is negatively correlated, based on the assumption that When it increases, the accuracy of the front-end detection model decreases; for ss tj The robot will navigate to the path with the maximum P_R(O tj ) and explore to find the object location of O target .

[0046] 3. If F t =0, and and At this time, LLM from CR t Select a bearing object cr k , and the robot performs action a t =Explore(cr k ); Specifically, extract CR t The LLM generates a description of each host object in the image and provides it as input together with the image or description of the target object; using the LLM's common sense understanding of object-host object relationships, the LLM identifies the host object that is most likely to contain the target object.

[0047] The state transfer process is as follows: If a t =Explore(cr) or a t= Goto(ct), where cr∈CR t ,ct∈CT t , during the robot's movement, let CR observed Represents the set of objects observed within a small radius r, in which there are no candidate targets. At the same time, let CT new Represents a set of new target candidates found on unexplored carrier objects; due to CT t Some candidate targets in CR may be observed Object load in CT t Will be updated to CT after removing these candidates t * Specifically, CT new The candidate targets in are those whose SBERT feature similarity with the target object exceeds the threshold σ1; in addition, the target object O target With CR observed The similarity between the objects carried in will not exceed σ1.

[0048] If a t =Explore(cr),CR t+1 and CT t+1 Updates as follows:

[0049] CR t+1 =CR t \(cr∪CR observed )

[0050] CT t+1 =CT t * ∪CT new

[0051] If a t =Goto(ct),CR t+1 and CT t+1 Updates as follows:

[0052] CR t+1 =CR t \(cr1∪CR observed ), ct∈C(cr1)

[0053] CT t+1 =CT t * ∪CT new {ct}

[0054] Calculate the target object O targetThe SBERT feature similarity between the object on cr or cr1; if the input command is an image, an LLM-based image comparison is also performed; if the combined score of the LLM image comparison and the SBERT text similarity exceeds the threshold σ2, then F is set t+1 = 1, and the task is marked as completed.

[0055] Furthermore, in S3, the specific map update method is:

[0056] During navigation, the robot periodically captures RGB images and depth images from the environment; the RGB images are processed by CropFormer, Tokenize Anything model, CLIP, and SBERT to obtain instance masks, descriptions, encoded CLIP features, and encoded SBERT features, respectively; for newly observed objects, the robot compares them with SBERT. G to identify the observed host object. Aspects of comparison include the size of objects, the distance between object center locations, and similarity scores based on CLIP and SBERT features.

[0057] For the currently observed instance, use h(·,·) to determine whether the instance is Carrying, where the newly observed carrying object set is O crd ;Will The object carried before and O crd Compare the objects; the comparison criteria include the size of the objects, the distance between the center positions, and the SBERT feature similarity score; after the comparison is completed, The objects on it will be updated based on the comparison results.

[0058] Another embodiment of the present invention further provides a robot indoor object navigation system, comprising: an offline map generation module, an object navigation module, and a map update module.

[0059] The offline graph generation module is used to generate an offline hosting relationship scene graph from a semantic point cloud map.

[0060] The object navigation module is used to receive the input navigation target, match the target object on the map, and the robot navigates to the target object according to the set navigation strategy.

[0061] The map update module is used to update the map while the robot is navigating to the target. If the target is not in the original location, the robot will explore the possible target locations according to a certain navigation strategy until it finds the target.

[0062] Furthermore, the offline graph generation module generates an offline hosting relationship scene graph from the semantic point cloud map in the following specific steps:

[0063] S101: Build an open vocabulary instance map using pre-collected scene RGB-D data Add a description list cap generated by the Tokenize Anything model to each instance i and a text feature T_F encoded by the SBERT model i to enhance each instance.

[0064] S102: Construct a hosting relationship scene graph S_G, including a building and room layer, a hosting level layer, and an object layer.

[0065] The building and room layers are constructed as follows: objects in the offline map are divided into different rooms using a priori-based approach. Then, the obtained room layers are merged to form the overall building layer.

[0066] The load-bearing layer is constructed as follows: Calculate each object O i Text feature T_F i SBERT-encoded text features for “furniture used to carry objects” The similarity between them is expressed as follows:

[0067]

[0068] Select the set of objects whose similarity score exceeds the specified threshold σ Its expression is as follows:

[0069]

[0070] Extract each mid-cap i The three most common descriptions are input into a large language model GPT-4o and are used to identify potential objects through specific prompts, denoted as

[0071] Finally, according to the criteria of whether the geometric size of the object exceeds a certain size and whether it is in contact with the ground, the final set of carrying objects is selected and recorded as Specifically:

[0072]

[0073] The object layer is constructed as follows: For any non-bearing layer object Based on O i The size and load-bearing objects The shortest distance between them and the spatial overlap relationship in the xyz direction are used to determine O i Whether it is the object Oj Carrying capacity; h(O j ,O i ) is a comprehensive index including the above factors, where h(O j ,O i )=1 if all conditions are met; for any Definition by O j The set of objects C(O j )as follows:

[0074]

[0075] Furthermore, the object navigation module specifically includes the following steps:

[0076] The input navigation target is a navigation instruction, specifically a text description or an image; the text description or image is encoded using the SBERT or CLIP model respectively.

[0077] The generated features are compared with the SBERT or CLIP features of each object in S_G using cosine similarity, where the object with the highest similarity score is selected as the target object. target .

[0078] Furthermore, the navigation strategy is specifically as follows:

[0079] The navigation strategy is modeled as a fixed-strategy Markov decision process, which includes:

[0080] State space S:

[0081] At the current step t: the robot's pose

[0082] Unexplored collection of host layer objects

[0083] The set of candidate target objects on the unexplored carrier layer

[0084] The flag indicating whether the target is found is F t ∈{0,1}, where and Respectively represent L t , CR t and CT t The value range set of ;

[0085] The state variables are:

[0086] S t =(L t ,CR t ,CT t ,F t)∈S

[0087] In the initial state S0=(L0,CR0,CT0,F0), L0 is the initial state of the robot, CT0=O target ;

[0088] The action space A is:

[0089] A={Stop,Explore(cr),Goto(ct)|cr∈CR t ,ct∈CT t}

[0090] Among them, Stop means that the task has been completed or all the objects in the carrier layer have been explored, and Explor(cr) and Goto(ct) respectively indicate the exploration of the carrier layer object cr∈CR t and navigate to the target object ct∈CT t location;

[0091] The robot is based on the current state S t and a specific policy π(·) to select the next action a in the action space A t ∈A;

[0092] Policy π(·): Given the current state S t =(L t ,CR t ,CT t ,F t ):

[0093] 1. If F t =1 or Then a t =Stop;

[0094] 2. If F t =0 and Prioritize the selection of a candidate object for operation. Specifically, let CT t ={O t1 ,…,O ti}; At the same time, store some additional variables: CT t With the target object O target SBERT similarity SS between t =ss t1 ,…,ss ti , current location L t With CT t The distance D t ={d t1 ,…,d ti}, and CT images observed by the robot camera t The average depth value of Any O tj ∈CT t Corresponding to ss tj d tj and The priority score is assessed in the following way:

[0095]

[0096] Among them, ss tj and P_R(O tj ) positive correlation, ss tj The larger it is, the higher the probability that the candidate is the target; and P_R(O tj ) is negatively correlated, based on the assumption that When it increases, the accuracy of the front-end detection model decreases; for ss tj The robot will navigate to the path with the maximum P_R(O tj ) and explore to find the object location of O target ;

[0097] 3. If F t =0, and and At this time, LLM from CR t Select a bearing object cr k , and the robot performs action a t =Explore(cr k ); Specifically, extract CR t The LLM generates a description of each object in the dataset and provides it, along with an image or description of the target object, as input to the LLM. Leveraging the LLM’s common-sense understanding of object-object relationships, the LLM identifies the object most likely to contain the target object.

[0098] The state transfer process is as follows: If a t =Explore(cr) or a t = Goto(ct), where cr∈CR t ,ct∈CT t , during the robot's movement, let CR observed Represents the set of objects observed within a small radius r, in which there are no candidate targets. At the same time, let CT new Represents a set of new target candidates found on unexplored carrier objects; due to CT t Some candidate targets in CR may be observed Object load in CT t Will be updated to CT after removing these candidatest * Specifically, CT new The candidate targets in are those whose SBERT feature similarity with the target object exceeds the threshold σ1; in addition, the target object O target With CR observed The similarity between the objects carried in will not exceed σ1;

[0099] If a t =explore(cr),CR t+1 and CT t+1 Updates as follows:

[0100] CR t+1 =CR t \(cr∪CR observed )

[0101] CT t+1 =CT t * ∪CT new

[0102] If a t =Goto(ct),CR t+1 and CT t+1 Updates as follows:

[0103] CR t+1 =CR t \(cr1∪CR observed ), ct∈C(cr1)

[0104] CT t+1 =CT t * ∪CT new {ct}

[0105] Calculate the target object O target The SBERT feature similarity between the object on cr or cr1; if the input command is an image, an LLM-based image comparison is also performed; if the combined score of the LLM image comparison and the SBERT text similarity exceeds the threshold σ2, then F is set t+1 =1, and the task is marked as completed;

[0106] Map update module, the specific map update method is:

[0107] During navigation, the robot periodically captures RGB images and depth images from the environment; the RGB images are processed by CropFormer, Tokenize Anything model, CLIP, and SBERT to obtain instance masks, descriptions, encoded CLIP features, and encoded SBERT features, respectively; for newly observed objects, the robot compares them with SBERT. G to identify the observed host object. The aspects of comparison include the size of objects, the distance between the center positions of objects, and the similarity scores based on CLIP and SBERT features;

[0108] For the currently observed instance, use h(·,·) to determine whether the instance is Carrying, where the newly observed carrying object set is O crd ;Will The object carried before and O crd Compare the objects; the comparison criteria include the size of the objects, the distance between the center positions, and the SBERT feature similarity score; after the comparison is completed, The objects on it will be updated based on the comparison results.

[0109] Beneficial effects:

[0110] 1. This paper proposes an instance navigation method based on open-set vocabulary, which identifies and updates dynamic objects and carrying relationships in the scene by combining visual-language models (such as CLIP and SBERT) and scene graphs. Based on the currently observed environmental information, the robot selects a suitable navigation target and explores by comparing the similarity between the object features and the target object, and adopting the common sense reasoning of LLM. Unlike traditional methods that mainly rely on fixed object categories or semantic labels, this method can process open-set vocabulary, support the recognition and navigation of unknown objects, and improve the adaptability of robots in complex and dynamic environments. The present invention can not only identify objects, but also model the carrying relationship between objects (i.e., the "carrier-carried" relationship). This function enables the robot to understand and dynamically update the state of objects in the scene, especially for everyday items that frequently change positions (such as cups, books, etc.), improving the accuracy and flexibility of navigation.

[0111] 2. By combining the commonsense reasoning capabilities of LLM with visual features (such as CLIP and SBERT), this method can better handle complex navigation tasks, such as inferring the possible placement of objects or handling dynamically changing target objects, effectively improving the intelligence and flexibility of navigation decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] Figure 1A block diagram of a robot indoor object navigation method and system principles provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0113] The present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0114] Example 1

[0115] An indoor robot object navigation algorithm provided by an embodiment of the present invention can realize robust instance-level object navigation of the robot in daily indoor environments.

[0116] The method framework is shown in Figure 1 .like Figure 1 As shown, the indoor object navigation algorithm of the present invention consists of three modules: generating an offline hosting relationship scene graph, object navigation, and map update. First, an offline hosting relationship scene graph is generated from a semantic point cloud map. Second, the navigation target is input and the target object on the map is matched. The robot navigates to the target and updates the changes in the scene along the way. If the target is not in the original location, the robot explores possible target locations according to a certain navigation strategy until the target is found. In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is further described in detail below with reference to specific implementation cases.

[0117] Step 1: The present invention first constructs an open vocabulary instance map M using pre-collected scene RGB-D data. The present invention adds a description list cap generated by the TokenizeAnything model to each instance. i and a text feature T_F encoded by the SBERT model i To enhance each instance. Next, a hosting relationship scene graph S_G is constructed.

[0118] Building and room layers: This paper uses a priori-based approach to divide objects in the offline map into different rooms. The resulting room layers are then merged to form the overall building layer.

[0119] Carrier level: The present invention calculates each object O i Text feature T_F i SBERT-encoded text features for “furniture used to carry objects” The similarity between them is expressed as follows.

[0120]

[0121] Next, the present invention selects a set of objects whose similarity scores exceed a specified threshold σ Its expression is as follows.

[0122]

[0123] Next, the present invention extracts each mid-cap i The three most common descriptions are fed into a large language model (GPT-4o is used in this test) and are used to identify potential objects through specific prompts, denoted as

[0124] Finally, the present invention selects the final set of load-bearing objects based on criteria such as whether the geometric size of the object exceeds a certain size and whether it is in contact with the ground, which is recorded as As shown in the formula below.

[0125]

[0126] Object layer: For any non-bearing layer object The present invention is based on i The size and load-bearing objects The shortest distance between them and the spatial overlap relationship in the xyz direction (exceeding a certain overlap rate) are used to judge O i Whether it is the object O j Carrying. j ,O i ) is defined as a comprehensive indicator including the above factors, where h(O j ,O i )=1 if all conditions are met. The present invention is defined by j The set of objects C(O j )as follows:

[0127]

[0128] The above completes the robot object navigation task in indoor environment.

[0129] Step 2: Navigation strategy for the object being moved. The input navigation command can be a text description or an image. The text description or image is encoded using the SBERT or CLIP model respectively. The generated features are then compared with the SBERT or CLIP features of each object in S_G using cosine similarity, similar to formula (1). The object with the highest similarity score is selected as the target object, O target .

[0130] The present invention models the exploration of a moved object as a Markov decision process with a fixed policy, as shown below.

[0131] State space S:

[0132] At the current step t, the present invention defines:

[0133] 1. Robot posture

[0134] 2. Unexplored collection of objects in the carrier layer

[0135] 3. The set of candidate target objects on the unexplored carrier layer

[0136] 4. Flag F indicating whether the target is found t ∈{0,1}. (where and Respectively represent L t , CR t and CT t The value range collection of .

[0137] The state variables are defined as follows:

[0138] S t =(L t ,CR t ,CT t ,F t )∈S(5)

[0139] In the initial state S0=(L0,CR0,CT0,F0), L0 is the initial state of the robot, CT0=O target .

[0140] Action space A:

[0141] A={Stop,Explore(cr),Goto(ct)|cr∈CR t ,ct∈CT t} (6)

[0142] Among them, Stop means that the task has been completed or all the objects in the carrier layer have been explored, Explore(cr) and Goto(ct) respectively indicate the exploration of the carrier layer object cr∈CR t and navigate to the target object ct∈CT t location.

[0143] The robot is based on the current state S t and a specific policy π(·) selects the next action a in (6) t ∈A.

[0144] Policy π(·): Given the current state S t =(L t ,CR t ,CTt ,F t ):

[0145] 1. If F t =1 or Then a t =Stop.

[0146] 2. If F t =0 and The present invention preferentially selects a candidate object to operate. Specifically, let CT t ={O t1 ,…,O ti}. At the same time, store some additional variables: CT t With the target object O target SBERT similarity SS between t =ss t1 ,…,ss ti , current location L t With CT t The distance D t ={d t1 ,…,d ti}, and CT images observed by the robot camera t The average depth value of Any O tj ∈CT t Corresponding to ss tj d tj and The priority score is evaluated in the following way.

[0147]

[0148] Among them, ss tj and P_R(O tj ) is positively correlated, the present invention assumes that ss tj The larger the value, the higher the probability that the candidate is the target. and P_R(O tj ) is negatively correlated, based on the assumption that As increases, the accuracy of the front-end detection model decreases. Considered as ss tj The robot will navigate to the path with the maximum P_R(O tj ) and explore to find the object location of O target .

[0149] 3. If F t =0, and and At this time, LLM will be tSelect a bearing object cr k , and the robot performs action a t =Explore(cr k Specifically, extract CR t The LLM then generates a description of each object in the dataset and provides it as input along with the image or description of the target object. Leveraging the LLM’s common sense understanding of object-object relationships (e.g., “a cup is unlikely to be placed on a toilet”), the LLM identifies the objects that are most likely to contain the target object.

[0150] State transition process: If a t =Explore(cr)(where cr∈CR t ) or a t = Goto(ct)(where ct∈CT t ), then during the robot's movement, let CR observed represents the set of objects observed within a small radius r, which have no candidate targets (based on the latest environmental observations); at the same time, let CT new represents the set of new target candidates found on unexplored carrier objects. t Some candidate targets in CR may be observed Object load in CT t Will be updated to CT after removing these candidates t * Specifically, CT new The candidate targets in are those whose SBERT feature similarity with the target object exceeds the threshold σ1. In addition, the target object O target With CR observed The similarity between the objects carried in will not exceed σ1.

[0151] 1. If a t =Explore(cr),CR t+1 and CT t+1 Updates as follows:

[0152] CR t+1 =CR t \(cr∪CR observed ) (8-a)

[0153] CT t+1 =CT t * ∪CT new (8-b)

[0154] 2. If a t =Goto(ct),CR t+1 and CTt+1 Updates as follows:

[0155] CR t+1 =CR t \(cr1∪CR observed ), ct∈C(cr1) (9-a)

[0156] CT t+1 =CT t * ∪CT new {ct} (9-b)

[0157] In either case, the target object O is calculated target The similarity between the SBERT features and the objects on cr or cr1. If the input command is an image, an LLM-based image comparison is also performed. If the combined score of the LLM image comparison and the SBERT text similarity exceeds the threshold σ2, then F is set t+1 = 1, and the task is marked as completed.

[0158] Step 3: Map Update. During navigation, the robot periodically captures RGB images and depth images from the environment. The RGB images are processed by CropFormer, TokenizeAnything model, CLIP, and SBERT to obtain instance masks, descriptions, encoded CLIP features, and encoded SBERT features, respectively. For newly observed objects, the robot compares them with SBERT. G to identify the observed host object. The main aspects of comparison include the size of objects, the distance between object center locations, and the similarity scores based on CLIP and SBERT features.

[0159] For the currently observed instances, use h(·,·) in formula (4) to determine whether they are Let the newly observed set of carrying objects be O crd Then, The object carried before and O crd Compare. The comparison criteria include the size of the object, the distance between the center positions and the SBERT feature similarity score. After the comparison is completed, The objects hosted on the are updated based on the comparison result: they are either added, removed, or remain unchanged.

[0160] Example 2:

[0161] An embodiment of the present invention further provides a robot indoor object navigation system, comprising: an offline map generation module, an object navigation module, and a map update module.

[0162] The offline graph generation module is used to generate an offline hosting relationship scene graph from a semantic point cloud map.

[0163] The object navigation module is used to receive the input navigation target, match the target object on the map, and the robot navigates to the target object according to the set navigation strategy.

[0164] The map update module is used to update the map while the robot is navigating to the target. If the target is not in the original location, the robot will explore the possible target locations according to a certain navigation strategy until it finds the target.

[0165] In the embodiment of the present invention, the offline graph generation module generates an offline hosting relationship scene graph from a semantic point cloud map in the following specific steps:

[0166] S101: Build an open vocabulary instance map using pre-collected scene RGB-D data Add a description list cap generated by the Tokenize Anything model to each instance i and a text feature T_F encoded by the SBERT model i to enhance each instance.

[0167] S102: Construct a hosting relationship scene graph S_G, including a building and room layer, a hosting level layer, and an object layer.

[0168] The building and room layers are constructed as follows: objects in the offline map are divided into different rooms using a priori-based approach. Then, the obtained room layers are merged to form the overall building layer.

[0169] The load-bearing layer is constructed as follows: Calculate each object O i Text feature T_F i SBERT-encoded text features for “furniture used to carry objects” The similarity between them is expressed as follows:

[0170]

[0171] Select the set of objects whose similarity score exceeds the specified threshold σ Its expression is as follows:

[0172]

[0173] Extract each mid-cap i The three most common descriptions are input into a large language model GPT-4o and are used to identify potential objects through specific prompts, denoted as

[0174] Finally, according to the criteria of whether the geometric size of the object exceeds a certain size and whether it is in contact with the ground, the final set of carrying objects is selected and recorded as Specifically:

[0175]

[0176] The object layer is constructed as follows: For any non-bearing layer object Based on O i The size and load-bearing objects The shortest distance between them and the spatial overlap relationship in the xyz direction are used to determine O i Whether it is the object O j Carrying capacity; h(O j ,O i ) is a comprehensive index including the above factors, where h(O j ,O i )=1 if all conditions are met; for any Definition by O j The set of objects C(O j )as follows:

[0177]

[0178] In the embodiment of the present invention, the object navigation module specifically includes the following steps:

[0179] The input navigation target is a navigation instruction, specifically a text description or an image; the text description or image is encoded using the SBERT or CLIP model respectively;

[0180] The generated features are compared with the SBERT or CLIP features of each object in S_G using cosine similarity, where the object with the highest similarity score is selected as the target object. target .

[0181] In the embodiment of the present invention, the navigation strategy is specifically:

[0182] The navigation strategy is modeled as a fixed-strategy Markov decision process, which includes:

[0183] State space S:

[0184] At the current step t: the robot's pose

[0185] Unexplored collection of host layer objects

[0186] The set of candidate target objects on the unexplored carrier layer

[0187] The flag indicating whether the target is found is F t ∈{0,1}, where and Respectively represent L t , CR t and CT t The value range set of ;

[0188] The state variables are:

[0189] S t =(L t ,CR t ,CT t ,F t )∈S

[0190] In the initial state S0=(L0,CR0,CT0,F0), L0 is the initial state of the robot, CT0=O target ;

[0191] The action space A is:

[0192] A={Stop,Explore(cr),Goto(ct)|cr∈CR t ,ct∈CT t}

[0193] Among them, Stop means that the task has been completed or all the objects in the carrier layer have been explored, Explore(cr) and Goto(ct) respectively indicate the exploration of the carrier layer object cr∈CR t and navigate to the target object ct∈CT t location.

[0194] The robot is based on the current state S t and a specific policy π(·) to select the next action a in the action space A t ∈A;

[0195] Policy π(·): Given the current state S t =(L t ,CT t ,CT t ,F t ):

[0196] 1. If F t =1 or Then a t =Stop;

[0197] 2. If F t =0 and Prioritize the selection of a candidate object for operation. Specifically, let CT t ={O t1 ,…,O ti}; At the same time, store some additional variables: CT t With the target object O target SBERT similarity SS between t =ss t1 ,…,ss ti , current location L t With CT t The distance D t ={d t1 ,…,d ti}, and CT images observed by the robot camera t The average depth value of Any O tj ∈CT t Corresponding to ss tj d tj and The priority score is assessed in the following way:

[0198]

[0199] Among them, ss tj and P_R(O tj ) positive correlation, ss tj The larger it is, the higher the probability that the candidate is the target; and P_R(O tj ) is negatively correlated, based on the assumption that When it increases, the accuracy of the front-end detection model decreases; for ss tj The robot will navigate to the path with the maximum P_R(O tj ) and explore to find the object location of O target .

[0200] 3. If F t =0, and and At this time, LLM from CR t Select a bearing object cr k , and the robot performs action a t =Explore(cr k ); Specifically, extract CR t The LLM generates a description of each host object in the image and provides it as input together with the image or description of the target object; using the LLM's common sense understanding of object-host object relationships, the LLM identifies the host object that is most likely to contain the target object.

[0201] The state transfer process is as follows: If a t =Explore(cr) or a t = Goto(ct), where cr∈CR t ,ct∈CT t , during the robot's movement, let CR observed Represents the set of objects observed within a small radius r, in which there are no candidate targets. At the same time, let CT new Represents a set of new target candidates found on unexplored carrier objects; due to CT t Some candidate targets in CR may be observed Object load in CT t Will be updated to CT after removing these candidates t * Specifically, CT new The candidate targets in are those whose SBERT feature similarity with the target object exceeds the threshold σ1; in addition, the target object O target With CR observed The similarity between the objects carried in will not exceed σ1.

[0202] If a t =Explore(cr),CR t+1 and CT t+1 Updates as follows:

[0203] CR t+1 =CR t \(cr∪CR observed )

[0204] CT t+1 =CT t *∪CT new

[0205] If a t =Goto(ct),CR t+1 and CT t+1 Updates as follows:

[0206] CR t+1 =CR t \(cr1∪CR observed ), ct∈C(cr1)

[0207] CT t+1 =CT t * ∪CT new {ct}

[0208] Calculate the target object Otarget The SBERT feature similarity between the object on cr or cr1; if the input command is an image, an LLM-based image comparison is also performed; if the combined score of the LLM image comparison and the SBERT text similarity exceeds the threshold σ2, then F is set t+1 = 1, and the task is marked as completed.

[0209] Map update module, the specific map update method is:

[0210] During navigation, the robot periodically captures RGB images and depth images from the environment; the RGB images are processed by CropFormer, Tokenize Anything model, CLIP, and SBERT to obtain instance masks, descriptions, encoded CLIP features, and encoded SBERT features, respectively; for newly observed objects, the robot compares them with SBERT. G to identify the observed host object. Aspects of comparison include the size of objects, the distance between object center locations, and similarity scores based on CLIP and SBERT features.

[0211] For the currently observed instance, use h(·,·) to determine whether the instance is Carrying, where the newly observed carrying object set is O crd ;Will The object carried before and O crd Compare the objects; the comparison criteria include the size of the objects, the distance between the center positions, and the SBERT feature similarity score; after the comparison is completed, The objects on it will be updated based on the comparison results.

[0212] Each module in this embodiment may be a program segment or a part of a code, and the above module, program segment, or part of a code includes one or more executable instructions for implementing a specified logical function.

[0213] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A robot indoor object navigation method, characterized in that: The steps include: S1: Generate offline hosting relationship scene graph from semantic point cloud map; S2: Input the navigation target, match the target object on the map, and the robot navigates to the target object according to the set navigation strategy; S3: The robot updates the map while navigating to the target. If the target is not at the original location, the robot explores possible target locations according to a certain navigation strategy until it finds the target.

2. The robot indoor object navigation method according to claim 1, characterized in that: S1: Generate an offline hosting relationship scene graph from a semantic point cloud map, specifically the following steps: S101: Build an open vocabulary instance map using pre-collected scene RGB-D data Add a description list cap generated by the Tokenize Anything model to each instance i and a text feature T_F encoded by the SBERT model i to enhance each instance; S102: Construct a hosting relationship scene graph S_G, including a building and room layer, a hosting level layer, and an object layer.

3. The robot indoor object navigation method according to claim 2, characterized in that: S102: Constructing a bearer relationship scene graph S_G, including a building and room layer, a bearer level layer, and an object layer, specifically: The building and room layers are constructed as follows: objects in the offline map are divided into different rooms using a priori-based approach, and then the obtained room layers are merged to form an overall building layer; The load-bearing layer is constructed as follows: Calculate each object O i Text feature T_F i SBERT-encoded text features for “furniture used to carry objects” The similarity between them is expressed as follows: Select the set of objects whose similarity score exceeds the specified threshold σ Its expression is as follows: Extract each mid-cap i The three most common descriptions are input into a large language model GPT-4o and are used to identify potential objects through specific prompts, denoted as Finally, according to the criteria of whether the geometric size of the object exceeds a certain size and whether it is in contact with the ground, the final set of carrying objects is selected and recorded as Specifically: The object layer is constructed as follows: for any non-bearing layer object Based on O i The size and load-bearing objects The shortest distance between them and the spatial overlap relationship in the xyz direction are used to determine O i Whether it is the object O j Carrying capacity; h(O j ,O i ) is a comprehensive index including the above factors, where h(O j ,O i )=1 if all conditions are met; for any Definition by O j The set of objects C(O j )as follows:

4. The robot indoor object navigation method according to claim 1, characterized in that: In S2, the navigation target is input and the target object on the map is matched. The specific steps include the following: The input navigation target is a navigation instruction, specifically a text description or an image; the text description or image is encoded using the SBERT or CLIP model respectively. The generated features are compared with the SBERT or CLIP features of each object in S_G using cosine similarity, where the object with the highest similarity score is selected as the target object. target .

5. The robot indoor object navigation method according to claim 1, characterized in that: The navigation strategy is specifically: The navigation strategy is modeled as a fixed-strategy Markov decision process, including: State space S: At the current step t: the robot's pose Unexplored collection of host layer objects The set of candidate target objects on the unexplored carrier layer The flag indicating whether the target is found is F t ∈{0,1}, where and Respectively represent L t , CR t and CT t The value range set of ; The state variables are: S t =(L t ,CR t ,CT t ,F t )∈S In the initial state S0=(L0,CR0,CT0,F0), L0 is the initial state of the robot, CT0=O target ; The action space A is: A={Stop,Explore(cr),Goto(ct)|cr∈CR t ,ct∈CT t } Among them, Stop means that the task has been completed or all the objects in the carrier layer have been explored, Explore(cr) and Goto(ct) respectively indicate the exploration of the carrier layer object cr∈CR t and navigate to the target object ct∈CT t location; The robot is based on the current state S t and a specific policy π(·) to select the next action a in the action space A t ∈A; Policy π(·): Given the current state S t =(L t ,CR t ,CT t ,F t ):

1. If F t =1 or Then a t =Stop; 2. If F t =0 and Prioritize the selection of a candidate object for operation. Specifically, let CT t ={O t1 ,…,O ti }; At the same time, store some additional variables: CT t With the target object O target SBERT similarity SS between t =ss t1 ,…,ss ti , current location L t With CT t The distance D t ={d t1 ,…,d ti }, and CT images observed by the robot camera t The average depth value of Any O tj ∈CT t Corresponding to ss tj d ij and The priority score is assessed in the following way: Among them, ss tj and P_R(O tj ) positive correlation, ss tj The larger it is, the higher the probability that the candidate is the target; and P_R(O tj ) is negatively correlated, based on the assumption that When it increases, the accuracy of the front-end detection model decreases; for ss tj The robot will navigate to the path with the maximum P_R(O tj ) and explore to find the object location of O target ; 3. If F t =0, and and At this time, LLM from CR t Select a bearing object cr k , and the robot performs action a t =Explore(cr k ); Specifically, extract CR t The LLM generates a description of each object in the dataset and provides it, along with an image or description of the target object, as input to the LLM. Leveraging the LLM’s common-sense understanding of object-object relationships, the LLM identifies the object most likely to contain the target object. The state transfer process is as follows: If a t =Explore(cr) or a t = Goto(ct), where cr∈CR t ,ct∈CT t , during the robot's movement, let CR observed Represents the set of objects observed within a small radius r, in which there are no candidate targets. At the same time, let CT new Represents a set of new target candidates found on unexplored carrier objects; due to CT t Some candidate targets in CR observed Object load in CT t Will be updated to CT after removing these candidates t * Specifically, CT new The candidate targets in are those whose SBERT feature similarity with the target object exceeds the threshold σ1; in addition, the target object O target With CR observed The similarity between the objects carried in will not exceed σ1; If a t =Explore(cr),CR t+1 and CT t+1 Updates as follows: CR t+1 =CR t \(cr∪CR observed ) CT t+1 =CT t * ∪CT new If a t =Goto(ct),CR t+1 and CT t+1 Updates as follows: CR t+1 =CR t \(cr1∪CR observed ),ct∈C(cr1) CT t+1 =CT t * ∪CT new {ct} Calculate the target object O target The SBERT feature similarity between the object on cr or cr1; if the input command is an image, an LLM-based image comparison is also performed; if the combined score of the LLM image comparison and the SBERT text similarity exceeds the threshold σ2, then F is set t+1 = 1, and the task is marked as completed.

6. The robot indoor object navigation method according to claim 1, characterized in that: In S3, the specific map update method is: During navigation, the robot periodically captures RGB images and depth images from the environment; the RGB images are processed by CropFormer, Tokenize Anything model, CLIP, and SBERT to obtain instance masks, descriptions, encoded CLIP features, and encoded SBERT features, respectively; for newly observed objects, the robot compares them with SBERT. G to identify the observed carrier object. The aspects of comparison include the size of objects, the distance between the center positions of objects, and the similarity scores based on CLIP and SBERT features; For the currently observed instance, use h(·,·) to determine whether the instance is Carrying, where the newly observed carrying object set is O crd ;Will The object carried before and O crd Compare objects using the size, distance between their centers, and SBERT feature similarity scores. After the comparison is completed, The objects on it will be updated based on the comparison results.

7. A robot indoor object navigation system, characterized in that: include: Offline map generation module, object navigation module and map update module; The offline graph generation module is used to generate an offline bearer relationship scene graph from a semantic point cloud map; The object navigation module is used to receive the input navigation target, match it to the target object on the map, and navigate the robot to the target object according to the set navigation strategy; The map update module is used to update the map when the robot navigates to the target. If the target is not at the original location, the robot explores possible target locations according to a certain navigation strategy until the target is found.

8. The robot indoor object navigation system according to claim 7, characterized in that: The offline graph generation module generates an offline hosting relationship scene graph from a semantic point cloud map in the following specific steps: S101: Build an open vocabulary instance map M using pre-collected scene RGB-D data; Add a description list cap generated by the Tokenize Anything model to each instance i and a text feature T_F encoded by the SBERT model i To enhance each instance; S102: Construct a hosting relationship scene graph S_G, including a building and room layer, a hosting level layer, and an object layer; The building and room layers are constructed as follows: objects in the offline map are divided into different rooms using a priori-based approach, and then the obtained room layers are merged to form an overall building layer; The load-bearing layer is constructed as follows: Calculate each object O i Text feature T_F i SBERT-encoded text features for “furniture used to carry objects” The similarity between them is expressed as follows: Select the set of objects whose similarity score exceeds the specified threshold σ Its expression is as follows: Extract each mid-cap i The three most common descriptions are input into a large language model GPT-4o and are used to identify potential objects through specific prompts, denoted as Finally, according to the criteria of whether the geometric size of the object exceeds a certain size and whether it is in contact with the ground, the final set of carrying objects is selected and recorded as Specifically: The object layer is constructed as follows: for any non-bearing layer object Based on O i The size and load-bearing objects The shortest distance between them and the spatial overlap relationship in the xyz direction are used to determine O i Whether it is the object O j Carrying capacity; h(O j ,O i ) is a comprehensive index including the above factors, where h(O j ,O i )=1 if all conditions are met; for any Definition by O j The set of objects C(O j )as follows:

9. The robot indoor object navigation system according to claim 7, wherein: The object navigation module specifically includes the following steps: The input navigation target is a navigation instruction, specifically a text description or an image; the text description or image is encoded using the SBERT or CLIP model respectively. The generated features are compared with the SBERT or CLIP features of each object in S_G using cosine similarity, where the object with the highest similarity score is selected as the target object. target .

10. The robot indoor object navigation system according to claim 9, characterized in that: The navigation strategy is specifically: The navigation strategy is modeled as a fixed-strategy Markov decision process, including: State space S: At the current step t: the robot's pose Unexplored collection of host layer objects The set of candidate target objects on the unexplored carrier layer The flag indicating whether the target is found is F t ∈{0,1}, where and Respectively represent L t , CR t and CT t The value range set of ; The state variables are: S t =(L t ,CR t ,CT t ,F t )∈S In the initial state S0=(L0,CR0,CT0,F0), L0 is the initial state of the robot, CT0=O target ; The action space A is: A={Stop,Explore(cr),Goto(ct)|cr∈CR t ,ct∈CT t } Among them, Stop means that the task has been completed or all the objects in the carrier layer have been explored, Explore(cr) and Goto(ct) respectively indicate the exploration of the carrier layer object cr∈CR t and navigate to the target object ct∈CT t location; The robot is based on the current state S t and a specific policy π(·) to select the next action a in the action space A t ∈A; Policy π(·): Given the current state S t =(L t ,CR t ,CT t ,F t ):

1. If F t =1 or Then a t =Stop; 2. If F t =0 and Prioritize the selection of a candidate object for operation. Specifically, let CT t ={O t1 ,…,O ti }; At the same time, store some additional variables: CT t With the target object O target SBERT similarity SS between t =ss t1 ,…,ss ti , current location L t With CT t The distance D t ={d t1 ,…,d ti }, and CT images observed by the robot camera t The average depth value of Any O tj ∈CT t Corresponding to ss tj d tj and The priority score is assessed in the following way: Among them, ss tj and P_R(O tj ) positive correlation, ss tj The larger it is, the higher the probability that the candidate is the target; and P_R(O tj ) is negatively correlated, based on the assumption that When it increases, the accuracy of the front-end detection model decreases; for ss tj The robot will navigate to the path with the maximum P_R(O tj ) and explore to find the object location of O target ; 3. If F t =0, and and At this time, LLM from CR t Select a bearing object cr k , and the robot performs action a t =Explore(cr k ); Specifically, extract CR t The LLM generates a description of each object in the dataset and provides it, along with an image or description of the target object, as input to the LLM. Leveraging the LLM’s common-sense understanding of object-object relationships, the LLM identifies the object most likely to contain the target object. The state transfer process is as follows: If a t =Explore(cr) or a t = Goto(ct), where cr∈CR t ,ct∈CT t , during the robot's movement, let CR observe d represents the set of objects observed within a small radius r, and there is no candidate target on the object in this set; at the same time, let CT new Represents a set of new target candidates found on unexplored carrier objects; due to CT t Some candidate targets in CR observed Object load in CT t Will be updated to CT after removing these candidates t * Specifically, CT new The candidate targets in are those whose SBERT feature similarity with the target object exceeds the threshold σ1; in addition, the target object O target With CR observed The similarity between the objects carried in will not exceed σ1; If a t =Explore(cr),CR t+1 and CT t+1 Updates as follows: CR t+1 =CRt \ (cr∪CR observed ) CT t+1 =CT t * ∪CT new If a t =Goto(ct),CR t+1 and CT t+1 Updates as follows: CR t+1 =CR t \(cr1∪CR observed ),ct∈C(cr1) CT t+1 =CT t * ∪CT new {ct} Calculate the target object O target The SBERT feature similarity between the object on cr or cr1; if the input command is an image, an LLM-based image comparison is also performed; if the combined score of the LLM image comparison and the SBERT text similarity exceeds the threshold σ2, then F is set t+1 =1, and the task is marked as completed; The map update module specifically updates the map in the following manner: During navigation, the robot periodically captures RGB images and depth images from the environment; the RGB images are processed by CropFormer, Tokenize Anything model, CLIP, and SBERT to obtain instance masks, descriptions, encoded CLIP features, and encoded SBERT features, respectively; for newly observed objects, the robot compares them with SBERT. G to identify the observed carrier object. The aspects of comparison include the size of objects, the distance between the center positions of objects, and the similarity scores based on CLIP and SBERT features; For the currently observed instance, use h(·,·) to determine whether the instance is Carrying, where the newly observed carrying object set is O crd ;Will The object carried before and O crd Compare the objects; the comparison criteria include the size of the objects, the distance between the center positions, and the SBERT feature similarity score; after the comparison is completed, The objects on it will be updated based on the comparison results.