Zero sample object navigation system and method using large language model

By integrating common-sense knowledge with large-scale language models and probabilistic soft logic, and combining scene understanding and common-sense reasoning, the generalization problem of object navigation in new environments is solved, achieving the effectiveness and adaptability of zero-shot object navigation.

CN121127346APending Publication Date: 2025-12-12SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480032159.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-03-19
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing object navigation technologies lack generalization ability when faced with new objects or new environments, and lack common sense reasoning ability, making it difficult to effectively navigate to a specific target object in an unfamiliar environment.

Method used

This method employs a large-scale language model combined with probabilistic soft logic (PSL) to integrate common-sense knowledge, achieving zero-shot object navigation through scene understanding and common-sense reasoning. It utilizes a pre-trained model for semantic scene understanding, combines cutting-edge detection techniques to infer the probability of a target object being in candidate rooms and objects, and updates the robot's position based on a map.

Benefits of technology

It enables effective navigation to target objects in unseen environments without additional training, improving the generalization ability of object navigation and enhancing adaptability in new environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121127346A_ABST
    Figure CN121127346A_ABST
Patent Text Reader

Abstract

A method includes determining a specified object to be positioned in an ambient environment. The method further includes causing the robot to capture an image of the ambient environment and a depth map. The method further includes predicting one or more rooms and one or more objects captured in the image using the scene understanding model. The method further includes updating a second map of the surroundings based on the predicted room, the predicted object, the depth map, and the location of the robot. The method further includes determining a likelihood that the specified object is within the candidate room and a likelihood that the specified object is near the candidate object using a pre-trained large language model. The method further includes moving the robot to a next location based on the likelihood and a second map for the robot to search for the specified object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to object navigation. More specifically, this disclosure relates to systems and methods for zero-shot object navigation using large language models. Background Technology

[0002] Object navigation refers to the task of an embodied agent navigating to a specific target object in an unknown environment. This task is fundamental to other navigation-based embodied tasks because it enables the agent to interact with the target object. Such object navigation tasks typically require large-scale training in a visual environment with labeled objects. Summary of the Invention

[0003] Technical solution

[0004] This disclosure provides systems and methods for zero-shot object navigation using large language models.

[0005] In a first embodiment, a method includes determining a designated object to be located in a surrounding environment, the surrounding environment including a plurality of candidate rooms and a plurality of candidate objects. The method further includes causing a robot to capture images and depth maps of the surrounding environment. The method further includes using a scene understanding model to predict one or more rooms and one or more objects captured in the images. The method further includes updating a second map of the surrounding environment based on one or more predicted rooms, one or more predicted objects, the depth map, and the robot's position. The method further includes using a pre-trained large language model to determine the probability that the designated object is in each of the candidate rooms and the probability that the designated object is near each of the candidate objects. Furthermore, the method includes causing the robot to move to a next location based on the determined probabilities and the second map of the surrounding environment so that the robot can search for the designated object.

[0006] In a second embodiment, an electronic device includes at least one processing device configured to determine a specified object to be located in a surrounding environment, the surrounding environment including a plurality of candidate rooms and a plurality of candidate objects. The at least one processing device is further configured to cause a robot to capture an image and a depth map of the surrounding environment. The at least one processing device is further configured to use a scene understanding model to predict one or more rooms and one or more objects captured in the image. The at least one processing device is further configured to update a second map of the surrounding environment based on one or more predicted rooms, one or more predicted objects, the depth map, and the robot's position. The at least one processing device is further configured to use a pre-trained large language model to determine the probability that the specified object is in each of the candidate rooms and the probability that the specified object is near each of the candidate objects. Furthermore, the at least one processing device is configured to cause the robot to move to a next location based on the determined probabilities and the second map of the surrounding environment so that the robot can search for the specified object.

[0007] In a third embodiment, a non-transitory machine-readable medium includes instructions that, when executed, cause at least one processor of an electronic device to determine a specified object to be located in a surrounding environment, the surrounding environment including a plurality of candidate rooms and a plurality of candidate objects. The non-transitory machine-readable medium also includes instructions that, when executed, cause the at least one processor to cause a robot to capture an image and a depth map of the surrounding environment. The non-transitory machine-readable medium further includes instructions that, when executed, cause the at least one processor to use a scene understanding model to predict one or more rooms and one or more objects captured in the image. The non-transitory machine-readable medium further includes instructions that, when executed, cause the at least one processor to update a second map of the surrounding environment based on one or more predicted rooms, one or more predicted objects, the depth map, and the robot's position. The non-transitory machine-readable medium further includes instructions that, when executed, cause the at least one processor to use a pre-trained large language model to determine the probability that the specified object is in each of the candidate rooms and the probability that the specified object is near each of the candidate objects. Furthermore, the non-transitory machine-readable medium contains instructions that, when executed, cause the at least one processor to move the robot to a next location based on a determined probability and the second map of the surrounding environment so that the robot can search for the designated object.

[0008] Other technical features will be readily apparent to those skilled in the art from the following figures, description and claims.

[0009] Before proceeding with the detailed description section below, it is helpful to define certain words and phrases used throughout this patent document. The terms “transmit,” “receive,” and “communicate,” and their derivatives, cover both direct and indirect communication. The terms “comprise,” “include,” and their derivatives, and their derivatives, mean unrestricted inclusion. The term “or” is inclusive, meaning “and / or.” The phrase “associated with,” and its derivatives, mean including, being included in, interconnected with, containing, being contained in, connected to or connected with, coupled to or coupled with, capable of communicating with, cooperating with, interleaving, juxtaposing, adjacent to, bound to or bound with, having, possessing the attributes of, being related to, etc.

[0010] Furthermore, the various functions described below can be implemented or supported by one or more computer programs, each computer program consisting of computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in appropriate computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium accessible by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, optical disc (CD), digital video disc (DVD), or any other type of storage. "Non-transitory" computer-readable medium excludes wired, wireless, optical, or other communication links that transmit transient electrical or other signals. Non-transitory computer-readable medium includes media where data can be permanently stored and media where data can be stored and later overwritten, such as rewritable optical discs or erasable memory devices.

[0011] As used herein, terms and phrases such as “have,” “may have,” “include,” or “may include” indicate the presence of a feature (such as a number, function, operation, or component, such as a part) without excluding the presence of other features. Furthermore, the phrases “A or B,” “at least one of A and / or B,” or “one or more of A and / or B” as used herein may include all possible combinations of A and B. For example, “A or B,” “at least one of A and B,” and “at least one of A or B” may indicate all of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. Furthermore, the terms “first” and “second” as used herein may modify various components regardless of their importance and do not limit these components. These terms are used only to distinguish one component from another. For example, a first user equipment and a second user equipment may refer to user equipments that are different from each other, regardless of the order or importance of the equipment. Without departing from the scope of this disclosure, a first component may be referred to as a second component, and vice versa.

[0012] It will be understood that when an element (such as the first element) is described as being "coupled" or "connected" (operably or communicatively) to another element (such as the second element), that element may be directly or indirectly coupled or connected to the other element via a third element. Conversely, it will be understood that when an element (such as the first element) is described as being "directly coupled" or "directly connected" to another element (such as the second element), no other element (such as a third element) lies between that element and the other element.

[0013] The phrase “configured (or set) as…” used in this document may be used interchangeably with phrases such as “suitable for…”, “capable of…”, “designed for…”, “adapted to…”, “made as…”, or “capable of…”, depending on the context. The phrase “configured (or set) as…” does not inherently mean “specifically designed in hardware as…”. Rather, the phrase “configured as…” can indicate that a device is capable of performing operations in conjunction with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” can refer to a general-purpose processor (such as a CPU or application processor) that performs the operations by executing one or more software programs stored in a memory device, or to a dedicated processor (such as an embedded processor) used to perform the operations.

[0014] The terms and phrases used herein are provided only to describe certain embodiments of this disclosure and are not intended to limit the scope of other embodiments of this disclosure. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” include plural references. All terms and phrases used herein, including technical and scientific terms and phrases, have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this disclosure pertain. It will be further understood that terms and phrases, such as those defined in common dictionaries, should be interpreted as having the same meaning as they have in the context of the relevant art and should not be interpreted in an idealized or overly formalized manner unless explicitly defined herein. In some cases, the terms and phrases defined herein may be interpreted as excluding embodiments of this disclosure.

[0015] Examples of "electronic devices" according to embodiments of this disclosure may include at least one of the following: smartphones, tablet PCs, mobile phones, video phones, e-book readers, desktop PCs, laptop computers, netbook computers, workstations, personal digital assistants (PDAs), portable multimedia players (PMPs), MP3 players, mobile medical devices, cameras, or wearable devices (such as smart glasses, head-mounted displays (HMDs), electronic clothing, electronic bracelets, electronic necklaces, electronic accessories, electronic tattoos, smart mirrors, or smartwatches). Other examples of electronic devices include smart home devices. Examples of smart home devices may include at least one of the following: a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a cleaner, an oven, a microwave oven, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), a smart speaker or loudspeaker with an integrated digital assistant (such as SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZONECHO), a game console (such as XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camera, or an electronic photo frame. Other examples of electronic devices include at least one of the following: various medical devices (such as various portable medical measurement devices (e.g., blood glucose measuring devices, heart rate measuring devices, or body temperature measuring devices), magnetic resonance angiography (MRA) devices, magnetic resonance imaging (MRI) devices, computed tomography (CT) devices, imaging devices, or ultrasound devices), navigation devices, global positioning system (GPS) receivers, event data loggers (EDR), flight data loggers (FDR), in-vehicle infotainment devices, marine electronic devices (such as marine navigation devices or gyrocompasses), avionics, security devices, in-vehicle control units, industrial or household robots, automated teller machines (ATMs), point-of-sale (POS) devices, or Internet of Things (IoT) devices (such as light bulbs, various sensors, electricity or gas meters, sprinklers, fire alarms, thermostats, streetlights, toasters, fitness equipment, hot water tanks, heaters, or boilers). Other examples of electronic devices include at least a portion of furniture or building / structural components, electronic whiteboards, electronic signature receiving devices, projectors, or various measuring devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). It should be noted that, according to various embodiments of this disclosure, the electronic device may be one or a combination of the devices listed above. According to certain embodiments of this disclosure, the electronic device may be a flexible electronic device. The electronic devices disclosed herein are not limited to the devices listed above and may include new electronic devices that emerge as technology develops.

[0016] In the following description, electronic devices will be described with reference to the accompanying drawings in accordance with various embodiments of this disclosure. The term "user" as used herein may refer to a person using the electronic device or another device (such as an artificial intelligence electronic device).

[0017] This patent document may provide definitions for certain other words and phrases throughout. Those skilled in the art will understand that, in many cases (if not most), such definitions apply to the past and future use of those defined words and phrases.

[0018] Nothing described in this application should be construed as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the subject matter of the patent protection is defined solely by the claims. Furthermore, unless expressed in the exact words “means for…”, no claim is intended to invoke 35 USC § 112(f). The applicant understands that any other terms used in the claims, including but not limited to “mechanism,” “module,” “device,” “unit,” “component,” “element,” “building block,” “device,” “machine,” “system,” “processor,” or “controller,” refer to structures known to those skilled in the art and are not intended to invoke 35 USC § 112(f). Attached Figure Description

[0019] To gain a more complete understanding of this disclosure and its advantages, the following description is now taken in conjunction with the accompanying drawings, in which the same reference numerals denote the same parts:

[0020] Figure 1 An example network configuration including electronic devices is shown according to this disclosure;

[0021] Figure 2 An example framework for zero-shot object navigation using a large language model is shown according to this disclosure;

[0022] Figure 3 An example environment in which an intelligent agent performs object navigation according to this disclosure is shown;

[0023] Figure 4 Example images captured by an agent during object navigation according to this disclosure are shown; and

[0024] Figure 5 An example method for zero-shot object navigation using a large language model, according to this disclosure, is shown. Detailed Implementation

[0025] The following discussion will be described with reference to the accompanying drawings. Figures 1 to 5Various embodiments of this disclosure are described. However, it should be understood that this disclosure is not limited to these embodiments, and all changes and / or equivalents or substitutions thereof are also within the scope of this disclosure.

[0026] As mentioned earlier, object navigation refers to the task of an embodied agent navigating to a specific target object in an unknown environment. This task can be fundamental to other navigation-based embodied tasks because it enables the agent to interact with the target object. Such object navigation tasks typically require large-scale training in visual environments with labeled objects.

[0027] While traditional object navigation techniques achieve good results when trained on specific datasets with limited target objects and similar environments, they can perform poorly when faced with new objects or environments due to distribution shifts. Real-world scenarios typically involve diverse objects and changing environments, making the collection of large amounts of annotated trajectory data both difficult and costly. Therefore, generalized zero-shot object navigation, where the navigation agent can adapt to new objects and environments without additional training, has become a key research direction.

[0028] To successfully navigate to a target object, an intelligent agent should possess both semantic scene understanding and common-sense reasoning capabilities. Semantic scene understanding involves recognizing objects existing in the environment, while common-sense reasoning involves logically inferring the location of the target object based on scene understanding. However, current zero-shot (i.e., unseen) object navigation methods do not effectively meet this requirement and often lack common-sense reasoning capabilities. Furthermore, some methods still require training in other goal-oriented navigation tasks and environments.

[0029] Some techniques enable zero-shot object navigation by transferring semantic scene understanding and commonsense reasoning knowledge from pre-trained models to open-world object navigation, without requiring any navigation experience or additional training in the visual environment. However, these large pre-trained models may not be able to directly and effectively generate navigation actions. Therefore, bridging the gap between pre-trained knowledge and navigation actions would be beneficial.

[0030] This disclosure provides several techniques for zero-shot object navigation using large language models (LLMs). As described in more detail below, the disclosed systems and methods provide a zero-shot object navigation framework that integrates common-sense knowledge into a frontier-based exploration (FBE) approach using probabilistic soft logic (PSL). PSL is a declarative template language that defines a special class of Markov random fields using first-order logic rules. PSL provides a simple framework for integrating common-sense knowledge from LLMs into exploration in a zero-shot manner. Unlike traditional techniques that rely on implicit training of common-sense knowledge using neural networks, the disclosed embodiments use soft logic predicates to represent knowledge in a continuous value space and then assign them to each frontier, thereby achieving more efficient exploration. In particular, the framework leverages a pre-trained model and can seamlessly generalize to unseen environments and new object types. Subsequently, the framework utilizes a pre-trained common-sense reasoning language model that infers the correspondence between rooms and objects using room and object information as context.

[0031] Note that although some embodiments discussed below are described in the context of the use of consumer electronic devices, such as home robots, this is merely an example, and it will be understood that the principles of this disclosure can be implemented in any number of other suitable contexts and can be used with any suitable device.

[0032] Figure 1 An example network configuration 100 including electronic devices is shown according to this disclosure. Figure 1 The embodiment of network configuration 100 shown is for illustrative purposes only. Other embodiments of network configuration 100 may be used without departing from the scope of this disclosure.

[0033] According to embodiments of this disclosure, electronic device 101 is included in network configuration 100. Electronic device 101 may include at least one of bus 110, processor 120, memory 130, input / output (I / O) interface 150, display 160, communication interface 170, or sensor 180. In some embodiments, electronic device 101 may not include at least one of these components, or at least one other component may be added. Bus 110 includes circuitry for connecting components 120-180 to each other and for transmitting communication (such as control messages and / or data) between components.

[0034] Processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In some embodiments, processor 120 includes one or more of a central processing unit (CPU), application processor (AP), communication processor (CP), or graphics processing unit (GPU). Processor 120 is capable of performing control and / or performing operations or data processing related to communication or other functions on at least one of the other components of electronic device 101. As described in more detail below, processor 120 can perform one or more operations for zero-sample object navigation using a large language model.

[0035] Memory 130 may include volatile and / or non-volatile memory. For example, memory 130 may be capable of storing commands or data associated with at least one other component of electronic device 101. According to embodiments of this disclosure, memory 130 may be capable of storing software and / or programs 140. Program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or application programs (or “applications”) 147. At least a portion of kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).

[0036] Kernel 141 can control or manage system resources (such as bus 110, processor 120, or memory 130) for performing operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). Kernel 141 provides an interface that allows middleware 143, API 145, or application 147 to access various components of electronic device 101 to control or manage system resources. Application 147 can support one or more functions for zero-sample object navigation using a large language model, as described below. These functions can be performed by a single application or by multiple applications, each performing one or more of these functions. For example, middleware 143 can act as a relay to allow API 145 or application 147 to communicate data with kernel 141. Multiple applications 147 can be provided. Middleware 143 can control work requests received from applications 147, such as by prioritizing the use of system resources of electronic device 101 (such as bus 110, processor 120, or memory 130) for at least one of the multiple applications 147. API 145 is an interface that allows application 147 to control functionality provided from kernel 141 or middleware 143. For example, API 145 includes at least one interface or function (such as a command) for file control, window control, image processing, or text control.

[0037] I / O interface 150 serves as an interface, for example, capable of transmitting commands or data input from a user or other external device to other components of electronic device 101. I / O interface 150 can also output commands or data received from other components of electronic device 101 to the user or other external device.

[0038] Display 160 includes, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot light-emitting diode (QLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. Display 160 can also be a depth-sensing display, such as a multi-focal display. Display 160 is capable of displaying various content (such as text, images, videos, icons, or symbols) to a user. Display 160 may include a touchscreen and can receive input such as touch, gestures, proximity, or hover input using an electronic pen or a user's body part.

[0039] Communication interface 170, for example, can establish communication between electronic device 101 and external electronic devices (such as first electronic device 102, second electronic device 104, or server 106). For example, communication interface 170 can connect to network 162 or 164 via wireless or wired communication to communicate with external electronic devices. Communication interface 170 can be a wired or wireless transceiver, or any other component for sending and receiving signals.

[0040] Wireless communication can use at least one of the following as a communication protocol: WiFi, Long Term Evolution (LTE), LTE-A Advanced, 5G, millimeter wave or 60 GHz wireless communication, wireless USB, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wi-Fi, or Global System for Mobile Communications (GSM). Wired connections may include at least one of the following: Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), RS-232 (Recommended Standard), or Common Old-Style Telephone Service (POTS). Network 162 or 164 includes at least one communication network, such as a computer network (e.g., a Local Area Network (LAN) or Wide Area Network (WAN)), the Internet, or a telephone network.

[0041] Electronic device 101 also includes one or more sensors 180 capable of measuring physical quantities or detecting the activation state of electronic device 101 and converting the measured or detected information into electrical signals. For example, one or more sensors 180 may include one or more cameras or other imaging sensors for capturing images of a scene. Sensor 180 may also include one or more buttons for touch input, gesture sensors, gyroscopes or gyro sensors, barometric pressure sensors, magnetic sensors or magnetometers, accelerometers or accelerometers, grip sensors, proximity sensors, color sensors (such as red-green-blue (RGB) sensors), biophysical sensors, temperature sensors, humidity sensors, illuminance sensors, ultraviolet (UV) sensors, electromyography (EMG) sensors, electroencephalography (EEG) sensors, electrocardiography (ECG) sensors, infrared (IR) sensors, ultrasound sensors, iris sensors, or fingerprint sensors. Sensor 180 may also include an inertial measurement unit, which may include one or more accelerometers, gyroscopes, and other components. Furthermore, sensor 180 may include control circuitry for controlling at least one of the sensors included herein. Any of these sensors 180 can be located within the electronic device 101.

[0042] In some embodiments, electronic device 101 may be a wearable device or a wearable device (such as an HMD) that can be mounted on an electronic device. For example, electronic device 101 may represent an AR wearable device, such as a head-mounted device or smart glasses with a display panel. In other embodiments, the first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device (such as an HMD) that can be mounted on an electronic device. In those other embodiments, when electronic device 101 is mounted in electronic device 102 (such as an HMD), electronic device 101 is able to communicate with electronic device 102 via communication interface 170. Electronic device 101 can connect directly to electronic device 102 to communicate with electronic device 102 without involving a separate network.

[0043] The first external electronic device 102, the second external electronic device 104, and the server 106 may each be the same as or a different type of device as electronic device 101. According to some embodiments of this disclosure, server 106 includes one or more servers. Furthermore, according to some embodiments of this disclosure, all or some of the operations performed on electronic device 101 may be performed on one or more other electronic devices (such as electronic devices 102 and 104 or server 106). Furthermore, according to some embodiments of this disclosure, when electronic device 101 is required to automatically or upon request perform certain functions or services, electronic device 101 can independently perform the function or service in its own name, or additionally, request another device (such as electronic devices 102 and 104 or server 106) to perform at least some of the associated functions. The other electronic device (such as electronic devices 102 and 104 or server 106) can perform the requested function or additional function and transmit the execution result to electronic device 101. Electronic device 101 can provide the requested function or service by processing the received result as is or additionally. For this purpose, cloud computing, distributed computing, or client-server computing technologies may be used, for example. Although Figure 1 The electronic device 101 is shown to include a communication interface 170 for communicating with an external electronic device 104 or a server 106 via a network 162 or 164; however, according to some embodiments of this disclosure, the electronic device 101 may operate independently without separate communication functionality.

[0044] Server 106 may include components 110-180 (or suitable subsets thereof) that are the same as or similar to electronic device 101. Server 106 is capable of supporting the driving of electronic device 101 by performing at least one of the operations (or functions) implemented on electronic device 101. For example, server 106 may include a processing module or processor that can support processor 120 implemented in electronic device 101. As described in more detail below, server 106 may perform one or more operations to support techniques for zero-shot object navigation using large language models.

[0045] although Figure 1 An example of a network configuration 100 including an electronic device 101 is shown, but other configurations are possible. Figure 1 Various changes can be made. For example, network configuration 100 can include any number of each component in any suitable arrangement. Typically, computing and communication systems have a wide variety of configurations, and Figure 1 This disclosure is not intended to limit the scope to any particular configuration. Furthermore, although... Figure 1 An operating environment in which the various features disclosed in this patent document can be used is shown, but these features can be used in any other suitable system.

[0046] Figure 2 An example framework 200 for zero-shot object navigation using a large language model, according to this disclosure, is shown. For ease of explanation, framework 200 is described as using the above-described... Figure 1 The network configuration 100 is implemented using one or more components, such as electronic device 101. However, this is merely an example, and the framework 200 can be implemented using any other suitable device (such as server 106) and any other suitable system.

[0047] As described in more detail below, electronic device 101 uses framework 200 to perform semantic scene understanding and commonsense reasoning in a zero-shot manner via one or more large pre-trained models. Electronic device 101 also combines state-of-the-art detection techniques with commonsense reasoning through PSL. To facilitate understanding of framework 200, further description of zero-shot object navigation and PSL used in conjunction with framework 200 may be helpful.

[0048] Zero-sample object navigation

[0049] In object navigation tasks, agents can be randomly placed in unseen environments. Figure 3 An example environment 300 in which an agent 302 performs object navigation according to this disclosure is shown. Here, the agent 302 is a robot, although the agent could be another electronic device capable of detecting the environment. The agent 302 is given a target object category (e.g., chair, fireplace, cabinet, etc.) and the purpose is to navigate to any object instance belonging to that category. At each step t, the agent 302 has observations including images and, sometimes, gesture readings. Figure 4 An example image 400 captured by agent 302 during object navigation is shown according to this disclosure. Based on this observation (including image 400), agent 302 is able to select an action in the action space, which may include a "stop" action to terminate the navigation process. Navigation is considered successful if agent 302 stops within a threshold distance of the target object and the object is visible without further movement.

[0050] As mentioned earlier, navigation learning in the real world is impractical due to its high cost, and most current methods struggle to generalize to new environments and objects. Therefore, the zero-shot navigation used in Framework 200 comprises three levels. The first level is the task level, where agent 302 can perform object-oriented tasks without object-oriented navigation training. The second level is the environment level, meaning that for a new set of environments, the model allows agent 302 to perform object navigation within that set of environments without training on any data from that environment. The third level is the object level, where the model can generalize to new target objects without further training.

[0051] Framework 200 uses a frontier-based probing principle, a heuristic probing method applicable to object navigation. For example... Figure 3 As shown, the leading edge 304 in environment 300 is defined as the boundary between the free region and the unseen region. The free region is defined as the region that agent 302 has seen and is not occupied by obstacles. Using leading edge-based detection, agent 302 is able to select the nearest leading edge 304 (with a distance threshold) after reaching it. The target object can be used as the next sub-target. The agent can navigate directly to the target object after detecting it. For example, such as... Figure 3 As shown, when the living room in environment 300 is detected, the agent 302 can prioritize detecting unseen areas in the living room to find the TV or sofa.

[0052] Probabilistic Soft Logic

[0053] PSL is a probabilistic programming language that uses a first-order logic-based syntax to define hinge-loss Markov random fields (HL-MRFs). Specifically, PSL models relational dependencies using weighted first-order logic clauses (called rules). For example, for the relational dependency between "a frontier near an object" and "selecting that frontier as a sub-objective," the following rule can be established:

[0054]

[0055] (1)

[0056]

[0057] In this rule, the predicate This indicates the relationship between the target object and one of the other objects. Is the frontier near the object? It is to select the leading edge of soft values, and This indicates the relative importance of the rule in the model.

[0058] In PSL inference, observed variables can be replaced by actual entities and values. , and This process is called grounding, and each specific instance of a rule is called a ground rule. PSL defines unobserved variables. The hinge loss potential function for each specific rule on the above is, for example, by the following formula:

[0059] (2)

[0060] here, It is a linear penalty function defined by PSL. This indicates the distance from the specified rule to its satisfaction. (Predicate) , The value is in the range [0, 1], and Optionally, the potential function can be squared.

[0061] Given all observed variables and unobserved variables PSL defines the HL-MRF over unobserved variables, which can be represented as follows:

[0062] (3)

[0063] (4)

[0064] here, Represents the number of potential functions, It is the first Each potential function It is aimed at The weight of the template rules.

[0065] The purpose of PSL inference is to find the method with the lowest penalty. The value of (i.e., finding the frontier score that best conforms to common sense). Here, the optimization of the distribution can be transformed into a convex optimization problem, which can be represented as follows:

[0066] (5)

[0067] like Figure 2As shown, electronic device 101 uses frame 200 to enable agent 302 to locate target object 201, which is a specified object among multiple candidate objects and candidate rooms within environment 300 surrounding agent 302. In this example, target object 201 may be a sofa. In some embodiments, target object 201 may be determined based on receiving a request to locate target object 201 within environment 300 (e.g., a request from a user). Once target object 201 is determined, electronic device 101 obtains at least one RGB image 202 and at least one depth map 204 of environment 300. RGB image 202 may represent... Figure 4 Image 400 (or represented thereto). When agent 302 is within (or near) environment 300, agent 302 can capture RGB image 202 and depth map 204. In some embodiments, electronic device 101 may represent agent 302 or a portion thereof. In other embodiments, electronic device 101 may represent a device different from agent 302. In some embodiments, electronic device 101 may, for example, cause agent 302 to capture RGB image 202 and depth map 204 by providing instructions to agent 302.

[0068] Scene understanding

[0069] To leverage a large language model for navigation inference, the RGB image 202 is converted into a semantic context in linguistic form. To achieve this, the electronic device 101 uses a scene understanding model 206 along with text prompts. In some embodiments, the scene understanding model 206 is a pre-trained, prompt-based, concretization-based language image model, such as GLIP. Unlike traditional object detection models and semantic segmentation models (such as Mask-RCNN) limited to fixed categories, the scene understanding model 206 constructs the detection task as a concretization problem by aligning proposed image regions with phrases in text prompts and predicting region-text alignment scores. Thanks to large-scale image-text training, the scene understanding model 206 is able to detect regular indoor elements in open-world settings. Therefore, it is more easily generalized to datasets with different environments and target objects to perform open-world object navigation.

[0070] Electronic Devices 101: Definition of Common Indoor Objects A collection, then using common objects and all possible target objects. The union of the elements is used to generate object hints for object materialization. 208. In some embodiments, object prompts 208 includes object names enclosed by periods ('.'). For example, if and If the object suggestion is 208, then it will be 'cabinet.chair.table'.

[0071] The object information in the current scene is a relatively low-level scene context. When people search for target objects in an unfamiliar environment, they typically consider higher-level context (e.g., "Which room should I go to?"). Therefore, to detect room information, room-specific cues are needed within an indoor environment. 210 Defines Common Rooms The electronic device 101 inputs object cues 208, room cues 210, and an RGB image 202 into a scene understanding model 206, which is capable of determining one or more predicted objects 212 and one or more predicted rooms 213, as well as bounding boxes from the current scene, such as by the following formula:

[0072] (6)

[0073] (7)

[0074] here, and These are the bounding boxes for the predicted object 212 and the predicted room 213. These cues can be easily extended to generalize to new test data to perform open-world semantic scene understanding.

[0075] Map building

[0076] Electronic device 101 uses depth map 204, agent position 214 (which may include agent pose information), and camera parameters as input to mapping operation 216, such as a simultaneous localization and mapping (SLAM) algorithm. Mapping operation 216 transforms the pixels of 2D RGB image 202 into 3D space, which is stored in 3D voxels. Mapping operation 216 also projects the 3D voxels along the height dimension into a 3D obstacle map, maintaining this 2D obstacle map during navigation. Furthermore, mapping operation 216 projects detected room and object locations into semantic map 218. For object detection, the center of the bounding box is projected onto the 2D location. For room detection, all pixels within the bounding box are projected onto the 2D map, and the projected location is recorded as the corresponding room.

[0077] Common sense reasoning

[0078] Intuitively, target object 201 should appear more frequently in certain rooms and near certain objects. For example, generally speaking, a sofa appears more frequently near a television in a living room than near a refrigerator in a kitchen. This common sense helps agent 302 search for target object 201. Therefore, after detecting room and object information in the current scene, electronic device 101 can use a pre-trained large language model 220 via text prompts 222 to perform common sense reasoning conditioned on target object 201 and semantic scene information. Here, large language model 220 can represent a publicly available large language model or any other suitable large language model. Text prompts 222 are input to large language model 220 and can include target object 201 as well as one or more candidate rooms and candidate objects in environment 300.

[0079] Specifically, for object-level and room-level inference, the large language model 220 is able to infer the target object 201—represented as —Is it possible to find every candidate object in object hint 208? Nearby, and whether target object 201 is possible in each room in room prompt 210. The prediction output of the large language model 220 can be a score for each (target, object) pair and (target, room) pair. , ∈ [0, 1]. For example, a text prompt 222 could be “Where might the sofa be near? Candidates: TV, table, bed, counter, …”. The predicted output from the large language model 220 could be a set of probability scores {0.6, 0.3, 0.4, 0.1, …}, where each probability score corresponds to one candidate object among the candidate objects. As another example, a text prompt 222 could be “Please provide a score indicating the probability of finding the sofa in the following rooms. Candidates: bedroom, kitchen, …”. The predicted output from the large language model 220 could be a set of probability scores {0.6, 0.1, …}, where each probability score corresponds to one candidate room among the candidate rooms. The techniques for obtaining probability scores from the large language model 220 and the syntax of the text prompt 222 can vary depending on the LLM.

[0080] Common sense-guided frontier exploration

[0081] In object navigation, efficient environmental probing can be crucial for finding target objects, as the object may not be visible from the agent's initial position. Traditional front-based probing can be used to explore the environment. However, selecting the nearest front as the sub-target to be probed may not be optimal in semantically rich environments and may violate common sense. For example, an agent might check the front behind a chaise lounge in a living room to search for a bed. Therefore, Framework 200 incorporates common-sense knowledge as well as front-based probing techniques. The goal is to make front-selection decisions more efficient. Not only based on the distance from agent 302 It is also based on objects surrounding the frontier 304 in environment 300. and room level information Such as through the following formula:

[0082] (8)

[0083] Intuitively, a front 304 is more likely to be chosen if it is close to an object in which the target object 201 is likely to be present, and if the front 304 is near or in a room in which the target object 201 is more likely to be present, and vice versa. However, these rules are not absolutely true or false. All fronts 304 that satisfy different numbers of rules can be compared. Moreover, the conditions in these rules are not absolutely true or false. For example, the target object 201 may be slightly or very likely to be present in the vicinity of a certain object.

[0084] Therefore, electronic device 101 performs planning operation 224, in which electronic device 101 combines common sense reasoning with front-based probing via PSL. In planning operation 224, electronic device 101 is capable of performing object reasoning, room reasoning, or both, to select a front 304 identified in semantic map 218. These will now be described.

[0085] Object reasoning

[0086] After the electronic device 101 detects objects and front edges 304 in the semantic map 218, the electronic device 101 can select these front edges 304 based on whether the front edge 304 is close to the object and whether the object is likely to appear around the target object 201. This selection can be based on the following PSL rules:

[0087]

[0088] (9)

[0089]

[0090] The rule encourages agent 302 to explore those frontiers 304 near some objects that may appear around target object 201 (i.e., explore near objects if the probability of that object being near target object 201 is greater than a predetermined threshold). The value is the co-occurrence score of (target, object) pairs predicted by the large language model 220. Furthermore, if, according to semantic map 218, the object is at the frontier 304... Within a meter range, then The value is the confidence score of the object prediction in the scene understanding model 206; otherwise... .

[0091] Furthermore, to suppress agent 302 from exploring those frontiers 304 near objects that are unlikely to be located around target object 201 (i.e., suppress agent 302 from exploring the vicinity of the object if the probability of the object being near target object 201 is below a predetermined threshold), corresponding negative PSL rules can be used, such as the following:

[0092]

[0093] (10)

[0094]

[0095] Room Deduction

[0096] Similar to object reasoning, room reasoning encourages agent 302 to detect the room in which the target object 201 might appear, or the front edge 304 of that room, and vice versa. Therefore, room reasoning can include two rules—a positive rule and a negative rule—such as the following:

[0097]

[0098] (11)

[0099]

[0100]

[0101] (12)

[0102]

[0103] As can be seen, equations (11) and (12) are similar to equations (9) and (10), but the terms 'Obj' and 'Object' are replaced with 'Room'.

[0104] Recent Frontiers

[0105] As described above, traditional front-edge-based detection techniques select the frontier that is at the shortest distance (above a threshold) to the agent. This encourages the agent to continue probing a region until there is nothing more to probe. Similarly, planning operation 224 can include a shortest-distance rule to encourage agent 302 to probe neighboring frontiers 304. An example shortest-distance rule is as follows:

[0106]

[0107] (13)

[0108] In some embodiments, a hard constraint on PSL can be implemented to limit the sum of the scores of all selected frontiers to one, such as by the following formula:

[0109] (14)

[0110] This constraint avoids degenerate solutions where all unobserved variables are equal to one and encourages competition among frontiers. All rules in the PSL model can use the same weights.

[0111] Once the electronic device 101 has selected the leading edge 304 to be detected for the agent 302, the electronic device 101 instructs the agent 302 to perform an action 226, such as moving the agent 302 to the position of the selected leading edge 304.

[0112] although Figures 2 to 4 An example and related details of a framework 200 for zero-shot object navigation using a large language model are shown, but further details are available. Figures 2 to 4 Various changes were made. For example, although frame 200 is described as involving a specific sequence of operations, regarding... Figure 2 The various operations described can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times). Furthermore, Figure 2 The specific operations shown are merely examples and can be performed using other techniques. Figure 2 Each operation shown.

[0113] Note that in Figures 2 to 4 Shown or about Figures 2 to 4 The described operations and functions can be implemented in any suitable manner in electronic devices 101, 102, 104, server 106, or other devices. For example, in some embodiments, they can be implemented or supported using one or more software applications or other software instructions executed by the processor 120 of electronic devices 101, 102, 104, server 106, or other devices. Figures 2 to 4 Shown or about Figures 2 to 4 The described operations and functions. In other embodiments, dedicated hardware components may be used to implement or support them. Figures 2 to 4 Shown or about Figures 2 to 4 At least some of the described operations and functions. Typically, this can be performed using any suitable hardware or any suitable combination of hardware and software / firmware instructions. Figures 2 to 4 Shown or about Figures 2 to 4 The described operations and functions. Furthermore, in Figures 2 to 4 Shown or about Figures 2 to 4 The described functions can be performed by a single device or by multiple devices.

[0114] Figure 5 An example method 500 for zero-shot object navigation using a large language model, according to this disclosure, is shown. For ease of illustration, Figure 5 The method 500 shown is described as using Figure 1 The electronic device 101 shown utilizes Figure 2 The framework 200 shown is executed. However, Figure 5 The method 500 shown can be used with any other suitable device or system.

[0115] like Figure 5 As shown, in step 501, a specified object to be located in the surrounding environment is determined. This surrounding environment includes multiple candidate rooms and multiple candidate objects. This may include, for example, electronic device 101 determining a target object 201 to be located within environment 300.

[0116] In step 503, the robot captures images and depth maps of the surrounding environment. This may include, for example, electronic devices 101 causing agent 302 to capture RGB images 202 and depth maps 204 of the surrounding environment 300.

[0117] In step 505, a scene understanding model is used to predict one or more rooms and one or more objects captured in the image. This may include, for example, the electronic device 101 using scene understanding model 206 to determine one or more predicted rooms 213 and one or more predicted objects 212.

[0118] In step 507, a second map of the surrounding environment is updated based on one or more predicted rooms, one or more predicted objects, a depth map, and the robot's position. This may include, for example, the electronic device 101 performing a mapping operation 216 to update the semantic map 218 of the environment 300 based on one or more predicted rooms 213, one or more predicted objects 212, a depth map 204, and the agent's position 214.

[0119] In step 509, a pre-trained large language model is used to determine the probability that the specified object is in each candidate room of the candidate rooms and the probability that the specified object is near each candidate object of the candidate objects. This may include, for example, the electronic device 101 using a large language model 220 to determine the probability that the target object 201 is in each candidate room of the candidate rooms and the probability that the specified object is near each candidate object of the candidate objects.

[0120] In step 511, based on the determined probabilities and a second map of the surrounding environment, the robot moves to the next location so that it can search for the specified object. This may include, for example, the electronic device 101 moving the agent 302 to the next location based on probabilities and a semantic map 218 determined from a large language model 220 so that the agent 302 can search for the target object 201.

[0121] although Figure 5 An example of method 500 for zero-shot object navigation using a large language model is shown, but it is possible to modify... Figure 5 Various changes were made. For example, although it is shown as a series of steps, Figure 5 The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0122] As described above, the disclosed embodiments introduce commonsense reasoning from large language models into object navigation tasks and propose a framework that leverages pre-trained visual and language models to perform zero-shot reasoning based on open-world object hierarchy and room hierarchy scene context. The disclosed embodiments also propose an exploration system that seamlessly integrates zero-shot commonsense reasoning with traditional exploration methods using probabilistic soft logic (PSL). This system is training-free and interpretable. The disclosed embodiments achieve improved results in zero-shot object navigation and significantly outperform conventional techniques on various object navigation datasets and benchmarks.

[0123] Although this disclosure has been described with reference to various exemplary embodiments, various changes and modifications will be made by those skilled in the art. This disclosure is intended to cover such changes and modifications that fall within the scope of the appended claims.

Claims

1. A method comprising: Identify a specific object to be located in the surrounding environment, which includes multiple candidate rooms and multiple candidate objects; The robot captures images and depth maps of the surrounding environment; The scene understanding model is used to predict one or more rooms and one or more objects captured in the image; A second map of the surrounding environment is updated based on one or more predicted rooms, one or more predicted objects, the depth map, and the robot's position; The probability of the specified object being in each of the candidate rooms and the probability of the specified object being near each of the candidate objects are determined using a pre-trained large language model. as well as Based on the determined probabilities and the second map of the surrounding environment, the robot moves to the next location so that it can search for the designated object.

2. The method of claim 1, wherein determining the designated object to be located in the surrounding environment comprises: The user receives a request to locate the specified object in the surrounding environment.

3. The method of claim 1, wherein determining the probability of the specified object being in each of the candidate rooms and the probability of the specified object being near each of the candidate objects using the pre-trained large language model comprises: Input a natural language query into the large language model, including the specified object and each candidate room in the candidate rooms; as well as The response is obtained from the large language model, and the response includes the probability score of each candidate room in the candidate rooms.

4. The method of claim 1, wherein moving the robot to the next location based on the determined probability and the second map of the surrounding environment comprises: Using a probabilistic soft logic algorithm and one or more of the determined probabilities, a frontier is selected from multiple frontiers identified in the second map.

5. The method of claim 1, wherein moving the robot to the next position comprises: If the probability of a first predicted object among the one or more predicted objects being near the specified object is greater than a threshold, then the robot is moved to an undetected location near the first predicted object.

6. The method of claim 5, wherein moving the robot to the next position further comprises: If the probability of the first predicted object being near the specified object is less than the threshold, then the robot is prevented from moving to the undetected location near the first predicted object.

7. The method of claim 1, wherein moving the robot to the next position comprises: If the probability that the specified object is in or near a first predicted room in one or more predicted rooms is greater than a threshold, then the robot is moved to an undetected location in or near the first predicted room.

8. An electronic device, comprising: At least one processing device, the at least one processing device being configured to: Identify a specific object to be located in the surrounding environment, which includes multiple candidate rooms and multiple candidate objects; The robot captures images and depth maps of the surrounding environment; The scene understanding model is used to predict one or more rooms and one or more objects captured in the image; A second map of the surrounding environment is updated based on one or more predicted rooms, one or more predicted objects, the depth map, and the robot's position; The probability of the specified object being in each of the candidate rooms and the probability of the specified object being near each of the candidate objects are determined using a pre-trained large language model. as well as Based on the determined probabilities and the second map of the surrounding environment, the robot moves to the next location so that it can search for the designated object.

9. The electronic device of claim 8, wherein, in order to determine the designated object to be located in the surrounding environment, the at least one processing device is configured to receive a request from a user to locate the designated object in the surrounding environment.

10. The electronic device of claim 8, wherein, in order to determine the probability of the specified object being in each of the candidate rooms and the probability of the specified object being near each of the candidate objects using the pre-trained large language model, the at least one processing device is configured to: Inputting a natural language query into the large language model includes the specified object and each candidate room in the candidate rooms; and The response is obtained from the large language model, and the response includes the probability score of each candidate room in the candidate rooms.

11. The electronic device of claim 8, wherein, in order to move the robot to the next position based on the determined probabilities and the second map of the surrounding environment, the at least one processing device is configured to select a frontier from a plurality of frontiers identified in the second map using a probabilistic soft logic algorithm and one or more of the determined probabilities.

12. The electronic device of claim 8, wherein, in order to move the robot to the next position, the at least one processing device is configured to: if the probability of a first predicted object among the one or more predicted objects being near the designated object is greater than a threshold, then move the robot to an undetected position near the first predicted object.

13. The electronic device of claim 12, wherein, in order to move the robot to the next position, the at least one processing device is further configured to: if the probability of the first predicted object being near the designated object is less than the threshold, then prevent the robot from moving to the undetected position near the first predicted object.

14. The electronic device of claim 8, wherein, in order to move the robot to the next location, the at least one processing device is configured to: if the probability that the designated object is in or near a first predicted room of one or more predicted rooms is greater than a threshold, then move the robot to an undetected location in or near the first predicted room.

15. A non-transitory machine-readable medium containing instructions that, when executed, cause at least one processor of an electronic device to: Identify a specific object to be located in the surrounding environment, which includes multiple candidate rooms and multiple candidate objects; The robot captures images and depth maps of the surrounding environment; The scene understanding model is used to predict one or more rooms and one or more objects captured in the image; A second map of the surrounding environment is updated based on one or more predicted rooms, one or more predicted objects, the depth map, and the robot's position; The probability of the specified object being in each of the candidate rooms and the probability of the specified object being near each of the candidate objects are determined using a pre-trained large language model. as well as Based on the determined probabilities and the second map of the surrounding environment, the robot is moved to the next location so that it can search for the designated object.