Multi-sensor perception model, and systems, devices, and methods thereof
A robotic device processes multiple sensor data streams in parallel using a perception algorithm to enhance navigation and interaction in dynamic environments by providing real-time, accurate environmental awareness, addressing the challenge of robust perception in complex settings.
Patent Information
- Application Number
- PCT/US2025/025617
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2025-04-21
- Publication Date
- 2025-10-23
AI Technical Summary
Robotic devices face challenges in navigating and interacting effectively in dynamic environments, such as hospitals, due to the need for robust perception capabilities that can handle multiple sensor data streams in real-time to ensure safety and efficiency.
A robotic device equipped with multiple sensors processes multiple data streams in parallel using a perception algorithm that includes feature extraction, transformer encoding, and decoding to generate accurate environmental representations, enabling efficient navigation and interaction.
The perception algorithm provides real-time, accurate, and comprehensive environmental awareness, allowing the robotic device to navigate around obstacles, interact with objects, and adapt its behavior based on social context, enhancing safety and functionality in hospital settings.
Smart Images

Figure US2025025617_23102025_PF_FP_ABST
Abstract
Description
MULTI-SENSOR PERCEPTION MODEL, AND SYSTEMS, DEVICES,AND METHODS THEREOFCross-Reference to Related Applications
[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 636,434, filed April 19, 2024, and entitled “Multi-Sensor Perception Model, and Systems, Devices, and Methods Thereof,” the entire disclosure of which is incorporated by reference herein.Technical Field
[0002] Embodiments described herein relate to a model for processing multiple streams of sensor data simultaneously for navigation and manipulation with a robotic device.Background
[0003] Robots may be deployed in dynamic environments (e.g., hospitals) and configured to assist with various logistical tasks. Effective navigation and interaction within complex hospital environments depends on robust perception capabilities to promote safety and efficiency.Summary
[0004] In some embodiments, a robotic device comprises a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; and one or more processors operatively coupled to the base. The one or more manipulating elements, and the set of sensors, the one or more processors are configured to receive, from the set of sensors, a plurality of data streams capturing information of an environment around the robotic device; input each data stream from the plurality of data streams into a respective feature extraction model from a plurality of feature extraction models to obtain a plurality of feature maps associated with the plurality of data streams; input the plurality of feature maps into a transformer encoder configured to generate a plurality of encoded outputs; input the plurality of encoded outputs into a plurality of decoders to generatea plurality of decoded outputs, each decoded output associated with a different representation of the environment; and determine a trajectory of the robotic device through the environment based on the different representations of the environment.
[0005] In some embodiments, a robotic device comprises a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; and one or more processors operatively coupled to the base, the one or more manipulating elements, and the set of sensors. The one or more processors are configured to receive, from each sensor of the set of sensors, information of an environment around the robotic device; input the information of the environment into an encoder model, the encoder model configured to generate one or more encoded outputs; input the one or more encoded outputs into a first decoder model configured to output object detection data, a second decoder model configured to output semantic segmentation data, and a third decoder model configured to output monocular depth data; and determine a trajectory of the robotic device based on the object detection data, the semantic segmentation data, and the monocular depth data.
[0006] In some embodiments, a robotic device comprises a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; a first processor operatively coupled to the base, the one or more manipulating elements, and the set of sensors and configured to receive sensor data from the set of sensors, the sensor data capturing information of an environment around the robotic device; and a second processor operatively coupled to the first processor. The second processor configured to receive, from the first processor, signals corresponding to the sensor data; input the sensor data into an encoder-decoder model configured to generate a plurality of outputs, each output of the plurality of outputs associated with a different representation of the environment, the first processor being configured to determine a plan for executing a task based on the plurality of outputs and to control the base and the one or more manipulating elements to execute the plan.
[0007] In some embodiments, a method comprises receiving, from a set of sensors, a plurality of data streams capturing information of an environment around a robotic device; inputting each data stream from the plurality of data streams into a respective feature extraction model from a plurality of feature extraction models to obtain a plurality of feature maps associated with the plurality of data streams; inputting the plurality of feature maps into a transformerencoder configured to generate a plurality of encoded outputs; inputting the plurality of encoded outputs into a plurality of decoders to generate a plurality of decoded outputs, each decoded output associated with a different representation of the environment; and determining a traj ectory of the robotic device through the environment based on the different representations of the environment.
[0008] In some embodiments, a method comprises receiving, from each sensor of the set of sensors, information of an environment around the robotic device; inputting the information of the environment into an encoder model, the encoder model configured to generate one or more encoded outputs; inputting the one or more encoded outputs into: a first decoder model configured to output object detection predictions, a second decoder model configured to output semantic segmentation data, and a third decoder model configured to output monocular depth data; and determining a trajectory of the robotic device based on the object detection predictions data, the semantic segmentation data, and the monocular depth data.
[0009] In some embodiments, a method comprises receiving, at a first processor, sensor data from a set of sensors of a robotic device, the sensor data capturing information of an environment around the robotic device; and receiving, at a second processor, signals corresponding to the sensor data from the first processor; inputting, using the second processor, the sensor data into an encoder-decoder model configured to generate a plurality of outputs, each output of the plurality of outputs associated with a different representation of the environment; and determining, using the first processor, a plan for executing a task based on the plurality of outputs and to control a base and one or more manipulating elements of the robotic device to execute the plan.Brief Description of the Drawings
[0010] FIG 1 is a schematic block diagram of a system including a robotic device, according to some embodiments.
[0011] FIG. 2 is a schematic block diagram of a configuration of a robotic device, according to some embodiments.
[0012] FIG. 3 is a schematic block diagram of a control unit of a robotic device, according to some embodiments.
[0013] FIG. 4 is a schematic illustration of a manipulating element of a robotic device, according to some embodiments.
[0014] FIG. 5 is a schematic illustration of a robotic device, according to some embodiments.
[0015] FIG. 6 is a schematic illustration of layers of a map of an environment generated by a robotic device, according to some embodiments.
[0016] FIG. 7 is a flow diagram of an example method of obtaining and perceiving information of an environment for learning and executing a skill, according to some embodiments.
[0017] FIG. 8 is a flow of information that a robotic device provides to and receives from a map maintained by the robotic device, according to some embodiments.
[0018] FIG. 9 is a flow chart of an example framework for training a student model based on a teacher model, according to some embodiments.
[0019] FIG. 10 is a flow chart of an algorithm for processing sensor data for navigation of a robotic device, according to some embodiments.
[0020] FIGS. 11A-11C show a flow chart of an algorithm for processing sensor data for navigation of a robotic device, according to some embodiments.
[0021] FIG. 12 is a schematic block diagram of flow of information between two processors of a robotic device to execute the perception algorithm, according to some embodiments.
[0022] FIGS. 13A-13B show results from performing rectification of a fisheye image to input into the perception algorithm.
[0023] FIGS. 14A-14C show example outputs of a perception algorithm for navigation of a robotic device.
[0024] FIG. 15 shows an example output of a perception algorithm for navigation of a robotic device.
[0025] FIG. 16 is a flow chart of an algorithm for processing sensor data for navigation of a robotic device, according to some embodiments.Detailed Description
[0026] Embodiments described herein may relate to a robotic device configured to execute a perception algorithm for processing and analyzing sensor data to provide the robotic device with semantic understanding of a surrounding environment (e.g., hospital). The perception algorithm may produce outputs that help the robotic device navigate through a dynamic environment and interact with objects in the environment. For example, the perception algorithm may produce outputs that help the robotic device execute various tasks such as maneuvering around moving objects, opening and moving through doors, riding elevators, etc. The perception algorithm may provide the robotic device environmental awareness by identifying environmental cues such as open doors and elevators, aiding the robotic device in decision-making and navigation.
[0027] The perception algorithm may be configured to receive multiple streams of sensor data and process the multiple streams of sensor data in parallel (i.e., simultaneously), thereby (1) speeding up an inference time of the algorithm and (2) providing the robotic device with an accurate and comprehensive representation of the surrounding environment at a given point in time. In some embodiments, the perception algorithm may be configured to receive different types of data in parallel (e.g., image data, infrared data, etc.). In some embodiments, the perception algorithm may continuously receive images (e.g., at 5 Hertz (Hz)-lOHz) from sensors (e.g., cameras) disposed on the robotic device (e.g., a head of the robotic device), with each sensor having a different field of view. The perception algorithm may output pixel-wise semantic segmentation, object detection, and monocular depth estimation for each image at a given time point. The semantic segmentation output may identify and label objects (e.g., people, wheelchairs, hospital beds, doors, etc.) The object detection output may detect dynamic objects in the scene and produce a 2D bounding box around the dynamic object, which can be used for facilitating trajectory planning around the dynamic object, for example. Monocular depth estimation may provide 3D spatial awareness, to help the robotic device ascertain a relative positioning of objects in the environment to the robotic device.
[0028] The perception algorithm may be trained with an apprenticeship architecture to provide accurate and robust performance. In some embodiments, the perception algorithm may be trained with cross-attention dropout, meaning that if one of the sensor streams is unavailable, the model may still produce reasonable outputs by taking advantage of sensor overlap. In someembodiments, if the robotic device is in a situation in which information from a sensor is important for the safety of the robotic device and / or a human in close proximity to the robotic device, the robotic device may be configured to produce a safety flag to escalate control of the robotic device to a user. In addition, the perception algorithm may halt navigation and / or operation (e.g., automatically or without a direct user input) of the manipulating element of the robotic device when people are detected in close proximity, mitigating potential hazards.
[0029] In some embodiments, the perception algorithm may be executed by a processor (e.g., a perception algorithm processor), and a separate processor (e.g., the main processor) may execute most or all other functions of the robotic device. In this way, the robotic device may free up computing power for the main processor to execute other tasks. For example, the robotic device may execute the perception algorithm on the perception algorithm processor to analyze and process large amounts of sensor data such that the main processor has available space to execute other important functions for the robotic device such as navigation, decision-making, learning, and / or planning.
[0030] The perception algorithm may be used to detect manipulation success. More specifically, the perception algorithm may provide a feedback mechanism to determine if manipulation by the manipulating element succeeded. The perception algorithm may aid the robotic device in making socially aware behavior changes. The model may provide 360° semantic information to the robot, and semantic scene understanding allows the robotic device to build a 3D representation of the environment in real-time or near-real time. This allows the robotic device to adapt its behavior, given the social context and / or a 3D representation of the environment. Overall, the perception algorithm may unlock a myriad of capabilities for the robotic device, enhancing navigation and manipulation functionality, safety, and social interactions within hospital settings.Systems and Devices
[0031] FIG. 1 is a schematic block diagram that illustrates a system 100, according to some embodiments. System 100 can be configured to obtain a representation of an environment and perceive the environment and / or learn and execute skills. For example, the system 100 may be configured to detect and / or identify obstacles in the environment such that the system 100 can learn and execute skills (e.g., manipulation skills) in an unstructured environment. System 100 can be implemented as a single device or may be implemented across multiple devices that areconnected to a network 105. For example, as depicted in FIG. 1, system 100 can include one or more compute devices, such as, for example, one or more robotic devices 102 and 110, a server 120, and additional compute device(s) 150. While four devices are shown, it should be understood that system 100 can include any number of compute devices, including compute devices not specifically shown in FIG. 1.
[0032] In some embodiments, system 100 includes a single robotic device, e.g., robotic device 102. Robotic device 102 can be configured to collect information about an environment (e.g., via one or more sensors) and to perceive information about the environment by executing one or more algorithms. In some embodiments, the robotic device 102 may be configured to receive and analyze multiple streams of data in parallel to reduce an amount of time for the robotic device to perform perception on the environment. Analysis may be on board the device or spread across the servers. For example, the robotic device may have one or more compute device(s) and / or processors on the device. In some embodiments, the sensor data and / or the analyzed data may be transmitted and / or stored in the server or compute devices. The robotic device 102 may be configured to interact with an environment including stationary and / or dynamic obstacles, meaning the robotic device 102 may be configured to detect and / or identify a class of object in the environment as well as detect and / or identify movement of objects in the environment.
[0033] In other embodiments, system 100 includes multiple robotic devices, e.g., robotic devices 102 and 110. Robotic device 102 can send and / or receive data to and / or from robotic device 110 via network 105. For example, robotic device 102 can send information that it perceives about an environment (e.g., a location of an object, an identity of an object, etc.) to robotic device 110, and can receive information about the environment from robotic device 110. Robotic devices 102 and 110 can also send and / or receive information to and / or from one another to learn and / or execute a skill based on the information about the respective environment of each robotic device 102. Robotic device 102 can be in a location that is the same as or different from robotic device 110. For example, robotic devices 102 and 110 can be located in the same room of a building (e.g., a hospital building). Alternatively, robotic device 102 can be located on a first floor of a building (e.g., a hospital building), and robotic device 110 can be located on a second floor of a building, and the two can communicate with one another to relay information about the different floors to one another (e.g., where objects are located on those floors, where a resource may be, etc.).
[0034] Compute device 150 can be any suitable processing device configured to run and / or execute certain functions. In a hospital setting, for example, a compute device 150 can be a diagnostic and / or treatment device that is capable of connecting to network 105 and communicating with other compute devices, including robotic device 102 and / or 110. Server 120 can be a dedicated server that manages robotic device 102 and / or 110. Server 120 can be in a location that is the same as or different from robotic device 102 and / or 110. For example, server 120 can be located in the same building as one or more robotic devices (e.g., a hospital building), and be managed by a local administrator (e.g., a hospital administrator). Alternatively, server 120 can be located at a remote location (e.g., a location associated with a manufacturer or provider of the robotic device).
[0035] Network 105 can be any type of network (e.g., a local area network (LAN), a wide area network (WAN), a virtual network, a telecommunications network) implemented as a wired network and / or wireless network and used to operatively couple compute devices, including robotic devices 102 and 110, server 120, and compute device(s) 150. As described in further detail herein, in some embodiments, for example, the compute devices are computers connected to each other via an Internet Service Provider (ISP) and the Internet (e.g., network 105). In some embodiments, a connection can be defined, via network 105, between any two compute devices. As shown in FIG. 1, for example, a connection can be defined between robotic device 102 and any one of robotic device 110, server 120, or additional compute device(s) 150. In some embodiments, the compute devices can communicate with each other (e.g., send data to and / or receive data from) and with the network 105 via intermediate networks and / or alternate networks (not shown in FIG. 1). Such intermediate networks and / or alternate networks can be of a same type and / or a different type of network as network 105. Each compute device can be any type of device configured to send data over the network 105 to send and / or receive data from one or more of the other compute devices.
[0036] In some embodiments, one or more robotic devices, e.g., robotic device 102 and / or 110, can be configured to communicate via network 105 with a server 120 and / or compute device 150. Server 120 can include component(s) that are remotely situated from the robotic devices and / or located on premises near the robotic devices. Compute device 150 can include component(s) that are remotely situated from the robotic devices, located on premises near the robotic devices, and / or integrated into a robotic device. Server 120 and / or compute device 150 can include a user interface that enables a user (e.g., a nearby user or a robot supervisor), tocontrol the operation of the robotic devices. The user can interrupt and / or modify the execution of one or more actions performed by the robotic devices. These actions can include, for example, navigation behaviors, manipulation behaviors, head behaviors, sounds / lights, and / or other components of a robotic device. In some embodiments, a robot supervisor can remotely monitor the robotic devices and control their operation for safety reasons. For example, the robot supervisor can command a robotic device to stop or modify an execution of an action to avoid endangering a human or causing damage to the robotic device or another object in an environment. The robotic device can solicit user intervention at points when the robotic device cannot confirm certain information about itself and / or the environment around itself. In some embodiments, if the robotic device cannot accurately perceive the surrounding environment (e.g., due to lack of sensor data, obstruction of a sensor, etc.) the robotic device can escalate control to the user to modify execution of the task and / or maneuver the robotic device to a safe position.
[0037] FIG. 2 schematically illustrates a robotic device 200, according to some embodiments. Robotic device 200 includes a control unit 202, a user interface 240, at least one manipulating element 250, and at least one sensor 270. Additionally, in some embodiments, robotic device 200 optionally includes at least one transport element 260. Control unit 202 includes a memory 220, a storage 230, a processor 204, a graphics processor 205, a system bus 206, and at least one input / output interface (“VO interface”) 208. Memory 220 can be, for example, a random access memory (RAM), a memory buffer, a hard drive, a database, an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), a readonly memory (ROM), and / or so forth. In some embodiments, memory 220 stores instructions that cause processor 204 to execute modules, processes, and / or functions associated with sensing or scanning an environment, learning a skill, and / or executing a skill. Storage 230 can be, for example, a hard drive, a database, a cloud storage, a network-attached storage device, or other data storage device. In some embodiments, storage 230 can store, for example, sensor data including state information regarding one or more components of robotic device 200 (e.g., manipulating element 250), learned models, marker location information, etc.
[0038] Processor 204 of control unit 202 can be any suitable processing device configured to run and / or execute functions associated with viewing an environment via the one or more sensors, processing one or more images obtained from the one or more sensors, detecting one or more objects in the environment, identifying the one or more objects in the environment,navigating the environment, learning a skill, and / or executing the skill based on the analysis of the one or more images of the environment. More specifically, processor 204 can be configured to execute modules, functions, and / or processes. In some embodiments, processor 204 can be a general purpose processor, a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a central processing unit (CPU), and / or the like.
[0039] Graphics processor 205 can be any suitable processing device configured to run and / or execute one or more display functions, e.g., functions associated with display device 242. In some embodiments, graphics processor 205 can be a low-powered graphics processing unit such as, for example, a dedicated graphics card or an integrated graphics processing unit.
[0040] System bus 206 can be any suitable component that enables processor 204, memory 220, storage 230, and / or other components of control unit 202 to communicate with each other. I / O interface(s) 208, connected to system bus 206, can be any suitable component that enables communication between internal components of control unit 202 (e.g., processor 204, memory 220, storage 230) and external input / output devices, such as user interface 240, manipulating element(s) 250, transport element(s) 260, and sensor(s) 270.
[0041] User interface 240 can include one or more components that are configured to receive inputs and send outputs to other devices and / or a user operating a device, e.g., a user operating robotic device 200. For example, user interface 240 can include a display device 242 (e.g., a display, a touch screen, etc.), an audio device 244 (e.g., a microphone, a speaker), and optionally one or more additional input / output device(s) (“I / O device(s)”) 246 configured for receiving an input and / or generating an output to a user.
[0042] Manipulating element(s) 250 can be any suitable component that is capable of manipulating and / or interacting with a stationary and / or moving object, including, for example, a human. In some embodiments, manipulating element(s) 250 can include a plurality of segments that are coupled to one another via joints that can provide for translation along and / or rotation about one or more axes. Manipulating element(s) 250 can optionally include an end effector that can engage with and / or otherwise interact with objects in an environment. For example, manipulating element can include a gripping mechanism that can releasably engage (e.g., grip) objects in the environment to pick up and / or transport the objects. Other examples of end effectors include, for example, vacuum engaging mechanism(s), magnetic engagingmechanism(s), suction mechanism(s), and / or combinations thereof. In some embodiments, one or more manipulating element(s) 250 can be retractable into a housing of the robotic device 200 when not in use to reduce one or more dimensions of the robotic device. In some embodiments, manipulating element(s) 250 can include a head or other humanoid component configured to interact with an environment and / or one or more objects within the environment, including humans. In some embodiments, manipulating element(s) 250 can include a transport element or base (e.g., transport element(s) 260). A detailed view of an example manipulating element is depicted in FIG. 4.
[0043] Transport element(s) 260 can be any suitable components configured for movement such as, for example, a wheel or a track. One or more transport element(s) 260 can be provided on a base portion of robotic device 200 to enable robotic device 200 to move around an environment. For example, robotic device 200 can include a plurality of wheels that enable it to navigate around a building, such as, for example, a hospital. Transport element(s) 260 can be designed and / or dimensioned to facilitate movement through tight and / or constrained spaces (e.g., small hallways and corridors, small rooms such as supply closets, etc.). In some embodiments, transport element(s) 260 can be rotatable about an axis and / or movable relative to one another (e.g., along a track). In some embodiments, one or more transport element(s) 260 can be retractable into a base of the robotic device 200 when not in use to reduce one or more dimensions of the robotic device. In some embodiments, transport element(s) 260 can be or form part of a manipulating element (e.g., manipulating element(s) 250).
[0044] Sensor(s) 270 can be any suitable component that enables robotic device 200 to capture information about the environment and / or objects in the environment around robotic device 200. Sensor(s) 270 can include, for example, image capture devices (e.g., cameras, such as a red-green-blue-depth (RGB-D) camera, fisheye cameras, or a webcam), audio devices (e.g., microphones), light sensors (e.g., light detection and ranging or lidar sensors, color detection sensors), proprioceptive sensors, position sensors, tactile sensors, force or torque sensors, temperature sensors, pressure sensors, motion sensors, sound detectors, etc. For example, sensor(s) 270 can include at least one image capture device such as a camera for capturing visual information about objects and the environment around robotic device 200. In some embodiments, sensor(s) 270 can include haptic sensors, e.g., sensors that can convey forces, vibrations, touch, and other non-visual information to robotic device 200.
[0045] In some embodiments, robotic device 200 can be have humanoid features, e.g., a head, a body, arms, legs, and / or a base. For example, robotic device 200 can include a face with eyes, a nose, a mouth, and other humanoid features. These humanoid feature can form and / or be part of one or more manipulating element(s). While not schematically depicted, robotic device 200 can also include actuators, motors, couplers, connectors, power sources (e.g., an onboard battery), and / or other components that link, actuate, and / or drive different portions of robotic device 200.
[0046] FIG. 3 is a block diagram that schematically illustrates a control unit 302, according to some embodiments. Control unit 302 can include similar components as control unit 202, and can be structurally and / or functionally similar to control unit 202. For example, control unit 302 includes a processor 304, a graphics processor 305, a memory 320, I / O interface(s) 308, a system bus 306, and a storage 330, which can be structurally and / or functionally similar to processor 204, memory 220, I / O interface(s) 208, system bus 206, and storage 230, respectively. Control unit 302 can be located on a robotic device and / or at a remote server that is connected to one or more robotic devices.
[0047] Memory 320 stores instructions that can cause processor 304 to execute modules, processes, and / or functions, illustrated as active sensing 322, learning 324, execution 328, and optionally data tracking and analytics 329. Active sensing 322, learning 324, execution 328, and data tracking and analytics 329 can be implemented as one or more programs and / or applications that are tied to hardware components (e.g., a sensor, a manipulating element, an I / O device, a processor, etc.). Active sensing 322, learning 324, execution 328, and data tracking and analytics 329 can be implemented by one robotic device or multiple robotic devices. For example, a robotic device can be configured to implement active sensing 322, learning 324, execution 328, and data tracking and analytics 329. As another example, a robotic device can be configured to implement active sensing 322, learning 324, and optionally execution 328. As another example, a robotic device can be configured to implement active sensing 322, learning 324, execution 328, and optionally data tracking and analytics 329. While not depicted, memory 320 can also store programs and / or applications associated with an operating system, and general robotic operations (e.g., power management, memory allocation, etc.).
[0048] In some embodiments, active sensing 322 can include active sensing or scanning of an environment, as described herein. In some embodiments, active sensing 322 can include active scanning of an environment and / or sensing or perceiving information associated with the environment, object(s) within the environment (e.g., including humans within the environment), and / or one or more conditions associated with a robotic device or system (e.g., by executing the perception algorithm).
[0049] Learning 324 can include modules, processes, and / or functions, that when implemented cause the robotic device to learn actions or skills. In some embodiments, skill model(s) 334 can be generated and / or updated based on learning 324. Execution 328 can cause the skill model(s) 334 to be implemented, thereby causing the robotic device to perform actions or skills. For example, learning from demonstration 326 can be configured to enable humans to demonstrate an action or a skill. Interactive learning 325 can be configured to receive inputs from a user that enable learning an action and / or updating a skill. Cache generation 327 can be configured to prompt a user to teach one or more pose that enable execution of a skill.
[0050] In some embodiments, execution 328 can be configured to cause the robotic device to execute an action and / or a skill learned from interactive learning 325, learning from demonstration 326, and / or cache generation 327. In some embodiments, execution 328 can include implementing one or more skill model(s) 334 to perform the action and / or the skill. Execution 328 can include modules, processes, and / or functions to determine joint configurations for the manipulating element (e.g., manipulating element 250) of the robotic device. In some embodiments, execution 328 can include modules, processes, and / or functions to implement one or more trajectories for the robotic device based on the learning 324.
[0051] In some embodiments, control unit 302 can also optionally include a data tracking & analytics element 329. Data tracking & analytics element 329 can be, for example, a computing element (e.g., a processor) configured to perform data tracking and / or analytics of the information collected by one or more robotic devices, e.g., information contained within state information 332 as further described below. For example, data tracking & analytics element 329 can be configured to manage inventory, e.g., tracking expiration dates, monitoring and recording the use of inventory items, ordering new inventory items, analyzing and recommending new inventory items, etc. In a hospital setting, data tracking & analytics element 329 can manage the use and / or maintenance of medical supplies and / or equipment.
[0052] Storage 330 stores information relating to an environment and / or objects within the environment, and learning and / or execution of skills (e.g., tasks and / or social behaviors). Storage 330 stores, for example, state information 332, skill model(s) 334, object information 340, machine learning libraries 342, and / or environmental constraints 354. Optionally, storage 330 can also store tracked information 356 and / or arbitration algorithm(s) 362.
[0053] State information 332 can include information regarding a state of a robotic device (e.g., robotic device 200) and / or an environment in which the robotic device is operating (e.g., a building, such as, for example, a hospital). In some embodiments, state information 332 can indicate a location of the robotic device within the environment, such as, for example, a room, a floor, an enclosed space, etc. For example, state information 332 can include a map of the environment, and indicate a location of the robotic device within that map. State information 332 can also include the location(s) of one or more objects (or markers representing and / or associated with objects) within the environment, e.g., within a map. Thus, state information 332 can identify a location of a robotic device relative to one or more objects. Objects can include any type of physical object that is located within the environment, including objects that define a space or an opening (e.g., surfaces or walls that define a doorway). Objects can be stationary or mobile. Examples of objects in an environment, such as, for example, a hospital, include equipment, supplies, instruments, tools, furniture, and / or humans (e.g., nurses, doctors, patients, etc.).
[0054] In some embodiments, state information 332 can include a representation or map of the environment along with static and / or dynamic information regarding objects (e.g., supplies, equipment, etc.) within the environment and social context information associated with humans and / or social settings within the environment, such as depicted in FIG. 6. Object information 340 can include information relating to physical object(s) in an environment. For example, object information 340 can include information identifying or quantifying different features of an object, such as, for example, location, color, shape, and surface features. Object information 340 can also identify codes, symbols, and other markers that are associated with a physical object, e.g., Quick Response or “QR” codes, barcodes, tags, etc. In some embodiments, object information 340 can include information characterizing an object within the environment, e.g., a doorway or hallway as being tight, a door handle as being a type of door handle, etc. Object information 340 can enable control unit 302 to identify physical object(s) in the environment.
[0055] Skill model(s) 334 are models that have been generated for performing different actions, and represent skills that have been learned (e.g., by implementing learning 324) by a robotic device. Optionally, sensory information can include transport element information. Transport element information can be associated with movements of a transport element (e.g., transport element(s) 260) of a robotic device as the robotic device undergoes a demonstration of a skill (e.g., navigation through a doorway, transport of an object, etc.). Transport element information can be recorded at specific points during a demonstration and / or execution of a skill, such as at keyframes associated with the skill, or alternatively, through a demonstration and / or execution of a skill.
[0056] Machine learning libraries 342 can include modules, processes, and / or functions relating to different algorithms for machine learning and / or model generation of different skills. In some embodiments, machine learning libraries can include methods such as Hidden Markov Models or “HMMs.” An example of an existing machine learning library in Python is scikit- learn. Storage 330 can also include additional software libraries relating to, for example, robotics simulation, motion planning and control, kinematics teaching and perception, etc.
[0057] Environmental constraints 354 can include information associated with objects and / or conditions within an environment that may restrict the operation of a robotic device within the environment. For example, environmental constraints 354 can include information associated with the size, configuration, and / or location of objects within an environment (e.g., supply bin, room, doorway, etc.), and / or information that indicates that certain areas (e.g., a room, a hallway, etc.) have restricted access. Environmental constraints 354 may affect the learning and / or execution of one or more actions within an environment. As such, an environmental constraint 354 may become part of each model for a skill that is executed within a context including the environmental constraint.
[0058] In some embodiments, an initial set of environmental constraints 354 (e.g., state information 331, object information 340, etc.) and / or skills (e.g., model(s) 334) can be provided to a robotic device, e.g., via a remote administrator or supervisor. The robotic device can adapt and / or add to its knowledge of environmental constraints 354 and / or skills 334 based on its own interactions, demonstrations, etc. with an environment or humans within the environment (e.g., patients, nurses, doctors, etc.) and / or via additional user input. Alternatively or additionally, a robot supervisor can update the robotic device’s knowledge of environmentalconstraints 354 and / or skills 334 based on new information collected by the robotic device or other robotic device(s) (e.g., other robotic device(s) within similar or the same environment, e.g., a hospital) and / or provided to the robotic supervisor by external parties (e.g., suppliers, administrators, manufacturers, etc.). Such updates can be periodically and / or continuously provided, as new information about an environment or skill is provided to the robotic device and / or robot supervisor.
[0059] Tracked information 356 includes information associated with a representation of an environment (e.g., map and / or representation 600) and / or information that is obtained from third-party systems (e.g., hospital electronic medical records, security system data, insurance data, vendor data, etc.) that is tracked and / or analyzed, e.g., by a data tracking & analytics element, such as data tracking & analytics element 329 or processor 302 executing data tracking & analytics 329. Examples of tracked information 356 include inventory data, supply chain data, point-of-use data, and / or patient data, as well as any aggregate data compiled from such data.
[0060] Arbitration algorithm(s) 362 include algorithms for arbitrating or selecting between different actions to execute (e.g., how to use different resources or components of a robotic device). Similar to I / O interface(s) 208, I / O interface(s) 308 can be any suitable component s) that enable communication between internal components of control unit 302 and external devices, such as a user interface, a manipulating element, a transport element, and / or compute device. I / O interface(s) 308 can include a network interface 360 that can connect control unit 302 to a network (e.g., network 105, as depicted in FIG. 1). Network interface 360 enables communications between control unit 302 (which can be located on a robotic device or another network device in communication with one or more robotic devices) and a remote device, such as a compute device that can be used by a robot supervisor to monitor and / or control one or more robotic devices. Network interface 360 can be configured to provide a wireless and / or wired connection to a network.
[0061] FIG. 4 schematically illustrates a manipulating element 450, according to some embodiments. Manipulating element 450 can form a part of a robotic device, such as, for example, robotic device 102 and / or 200. Manipulating element 450 can be implemented as an arm that includes two or more segments 452 coupled together via joints 454. Joints 454 can allow one or more degrees of freedom. For example, joints 454 can provide for translationalong and / or rotation about one or more axes. In an embodiment, manipulating element 450 can have seven degrees of freedom provided by joints 454. While four segments 452 and four joints 454 are depicted in FIG. 4, one of ordinary skill in the art would understand that a manipulating element can include a different number of segments and / or joints. A plurality of sensors 453, 455, 457, and 458 can be disposed on different components of manipulating element 450, e.g., segments 452, joints 454, and / or end effector 456. Sensors 453, 455, 457, and 458 can be configured to measure sensory information, including environmental information and / or manipulating element information. Examples of sensors include position encoders, torque and / or force sensors, touch and / or tactile sensors, image capture devices such as cameras, temperature sensors, pressure sensors, light sensors, etc.
[0062] Manipulating element 450 includes an end effector 456 that can be used to interact with objects in an environment. For example, end effector 456 can be used to engage with and / or manipulate different objects. Alternatively or additionally, end effector 456 can be used to interact with movable or dynamic objects, including, for example, humans. In some embodiments, end effector 456 can be a gripper that can releasably engage or grip one or more objects. For example, end effector 456 implemented as a gripper can pick up and move an object from a first location (e.g., a supply closet) to a second location (e.g., an office, a room, etc.).
[0063] FIG. 5 schematically illustrates a robotic device 500, according to some embodiments. Robotic device 500 includes a head 580, a body 588, and a base 586. Head 580 can be connected to body 588 via a segment 582 and one or more joints (not depicted). Segment 582 can be movable and / or flexible to enable head 580 to move relative to body 588. Head 580, segment 582, etc. can be examples of manipulating element(s), and include similar functionality and / or structure as other manipulating element(s) described herein.
[0064] Head 580 includes one or more image capture devices 572 and / or other sensors 570. Image capture device 572 and / or other sensors 570 (e.g., lidar sensors, motion sensors, etc.) can enable robotic device 500 to scan an environment and obtain a representation (e.g., a visual representation or other semantic representation) of the environment. In some embodiments, image capture device 572 can be a camera. In some embodiments, image capture device 572 can be movable such that it can be used to focus on different areas of the environment around robotic device 500. Image capture device 572 and / or other sensors 570 can collect and sendsensory information to a compute device or processor onboard robotic device 500, such as, for example, control unit 202 or 302. In some embodiments, head 580 of robotic device 500 can have a humanoid shape, and include one or more human features, e.g., eyes, nose, mouth, ears, etc. In such embodiments, image capture device 572 and / or other sensors 570 can be implemented as one or more human features. For example, image capture device 572 can be implemented as eyes on head 580.
[0065] In some embodiments, robotic device 500 can use image capture device 572 and / or other sensors 570 to scan an environment for information about objects in the environment, e.g., physical structures, devices, articles, humans, etc. Robotic device 500 can engage in active sensing, or robotic device 500 can initiate sensing or scanning in response to a trigger (e.g., an input from a user, a detected event or change in the environment). In some embodiments, robotic device 500 can engage in adaptive sensing where sensing can be performed based on stored knowledge and / or a user input. For example, robotic device 500 can identify an area in the environment to scan for an object based on prior information that it has on the object.
[0066] FIG. 6 shows a schematic diagram of a map 600 of the environment, according to some embodiments. The map 600 can include, for example, a navigation layer 610, a static semantic layer 620, a social layer 630, and a dynamic layer 640. Navigation layer 610 provides a general layout or map of the building, which may identify a number of floor(s) 612 with wall(s) 614, stair(s) 616, and other elements built into the building (e.g., hallways, openings, boundaries). Static semantic layer 620 identifies objects and / or spaces within the building, such as room(s) 622, object(s) 624, door(s) 626, etc. Static semantic layer 620 can identify which room(s) 622 or other spaces are accessible or not accessible to a robotic device. In some embodiments, static semantic layer 620 can provide a three dimensional map of the objects located within a building. Social layer 630 provides social context information 632. Social context information 632 includes information associated with humans within the building, such as past interactions between robotic device(s) and human(s). Social context information 632 can be used to track interactions between robotic device(s) and human(s), which can be used to generate and / or adapt existing models of skills involving one or more interactions between a robotic device and a human. For example, social context information 632 can indicate that a human is typically located at a particular location, such that a robotic device having knowledge of that information can adapt its execution of a skill that would require the robotic device to move near the location of the human. Dynamic layer 640 provides information on object(s) and other elements withina building that may move and / or change over time. For example, dynamic layer 640 can track movement(s) 644 and / or change(s) 646 associated with object(s) 642. In an embodiment, dynamic layer 640 can monitor the expiration date of an object 642 and identify when that object 642 has expired.
[0067] Map and / or representation 600 can be accessible to and / or managed by a control unit 602 (e.g., structurally and / or functionally similar to control unit 302), e.g., of a robotic device. Control unit 602 can include similar components as other control units described herein (e.g., control units 202 and / or 302). Control unit 602 can include a storage (similar to other storage elements described herein, such as, for example, storage 330) that stores state information 604, including map and / or representation 600 as well as information associated with one or more robotic devices (e.g., a configuration of an element of a robotic device, a location of the robotic device in a building, etc.). Control unit 602 can be located on a robotic device and / or at a remote server that is connected to one or more robotic devices. Robotic device(s) can be configured to update and maintain state information 604, including information associated with representation 600 of the environment, as the robotic device(s) collect information on their surrounding environment.
[0068] Information learned by a robotic device, e.g., from interactive learning 325, learning from demonstration 326, and / or cache generation 327 can feed into the different layers of the map 600 and be organized for future reference by the robotic device (and / or other robotic devices). For example, a robotic device may rely on information learned about different objects within an environment (e.g., a door) to decide how to arbitrate between different behaviors (e.g., waiting for a door to a tight doorway to be opened before going through the doorway, seeking assistance to open a door before going through a tight doorway).
[0069] FIG. 7 is a flow diagram of an example method of obtaining a representation of the environment and perceiving information in the environment for learning and / or executing a skill. As depicted in FIG. 7, a robotic device can scan an environment and obtain a representation of the environment, at 704. The robotic device can scan the environment using one or more sensors (e.g., sensor(s) 270 or 470, and / or image capture device(s) 572). The representation of the environment may include images and / or videos of one or more areas in the environment. In some embodiments, the robotic device can scan the environment using a movable camera, where the position and / or focus of the camera can be adjusted to capture areasin a scene of the environment. In some embodiments, the robotic device can use a plurality of sensors (e.g., standard RGB camera, RGB-D camera, fisheye camera, wide angle camera, night vision camera, infrared sensors, LiDAR sensors, etc.) to obtain the representation of the environment. In some embodiments, the robotic device may capture image data using one RGB-D camera and two fisheye camera. The sensors may capture representations of the environment in parallel and continuously (or semi-continuously) as the robotic device moves through the environment.
[0070] At 706, based on the data collected during the sensing, the robotic device can perform perception on the captured representation of the environment. For example, the one or more processors of the robotic device (e.g., processor(s) 204, 205) may be configured to execute a perception algorithm for processing and / or analyzing the representations the environment received from the sensors. In some embodiments, the data corresponding to the representation of the environment may be continuously (or semi-continuously) fed to the model such that the robotic device may perform perception as the robotic device navigates through the environment. In some embodiments, a subset of images may be fed into the perception algorithm (e.g., images can be dropped to accommodate latencies). The robotic device can use the perception algorithm to gain semantic understanding of the environment to avoid obstacles and / or interact with objects in the environment based on the results of the perception algorithm.
[0071] In some embodiments, performing perception may include performing semantic segmentation (e.g., pixel-wise semantic segmentation), object detection, and / or monocular depth estimation on images received from the sensors. In some embodiments, the robotic device may process and / or analyze the image data in parallel (i.e., simultaneously) to reduce a latency with which the robotic device can interact with the environment based on its perception of the environment. For example, each camera may capture an image at a first time point, and the images from each camera may be fed into the perception model concurrently. The model may then output the segmentation, object detection, and monocular depth estimation results for each camera concurrently. In some embodiments, the sensors may have overlapping fields of view such that the perception model may produce reasonable outputs when images from any one sensor are absent based on sensor to sensor overlap. In some embodiments, the robotic device may identify or classify the object (e.g., as a chair, door, person, bed, etc.). In some embodiments, the robotic device may use object detection to identify moving (i.e., dynamic)objects in the environment. The robotic device can then change its trajectory (e.g., of the manipulating element or other component) based on the movement of the object.
[0072] After obtaining a representation of the environment, at 704, and performing perception, at 706, the robotic device can initiate learning, at 702, or execution, at 712. In some embodiments, the robotic device may obtain representation of the environment and perform perception while the robot operates in learning mode 702 and / or execution mode 712. In some embodiments, when the robotic device enters learning mode 702 and / or execution mode 712, the robotic device may begin or continue scanning the environment and performing perception on the scanned environment. The robotic device can employ various methods to learn information about its environment, its current state, and / or one or more behaviors or skills. Additionally, the robotic device can employ various methods to execute one or more behaviors or skills. Further details related to the robotic device learning and executing skills is described in International Patent Publication No. WO 2020 / 047120, entitled, “Systems, apparatus, and methods for robotic learning and execution of skills”, filed August 28, 2019 and International Patent Publication No. WO 2022 / 170279, entitled, “Systems, apparatuses, and methods for robotic learning and execution of skills including navigation and manipulation functions”, filed February 8, 2022, the disclosure of each of which is incorporated by reference herein in its entirety.
[0073] In some embodiments, the information obtained during perception of the scanned environment can be used to generate a map maintained by the robotic device (e.g., map 600 shown in FIG. 6). FIG. 8 is a flow of information that the robotic device provides to and receives from a map 820 maintained by the robotic device. The map 820 can be stored and / or maintained by one or more robotic devices, such as any of the robotic devices as described herein. The map 820 can include one or more layer(s) 822 such as, for example, a navigation layer, a static layer, a dynamic layer, and a social layer.
[0074] A robotic device operating in the learning mode 802 can provide information that adds to and / or changes information within the map 820. For example, the robotic device can incorporate sensed information 804 (e.g., information collected by one or more sensor(s) of the robotic device) and / or derived information 806 (e.g., information derived by the robotic device based on, for example, analyzing sensed information 804) into the map 820. The robotic device can incorporate such information by adding the information to the map 820 and / or adaptingexisting information in the map 820. Additionally, the robotic device operating in the execution mode 812 can provide information (e.g., sensed information 814, derived information 816) that adds to and / or changes information within the map 820.
[0075] In some embodiments, the robotic device may segment the images received from the sensor(s) and then determine a location of the object and / or a class of the object based on the segmentation. The robotic device may store information including the location of the object and the class of the object in the map 820. In some embodiments, robotic device can identify an area in the environment to scan for an object based on prior information that the robotic device has on the object.
[0076] For example, the robotic device can scan a scene (e.g., an area of a room) and obtain a first representation of the scene. In the first representation, the robotic device can identify (e.g., based on image segmentation) that a first object is located in a first area and that a second object is located in a second area. The robotic device can store the locations of the first and second objects in the map 820 of the environment that it stored internally, such that the robotic device can use that information to locate the first and second objects when performing a future scan. When the robotic device returns to the scene and scans the scene a second time, the robotic device may obtain a different view of the scene. When performing this second scan, robotic device can obtain a second representation of the scene. To locate the first and second objects in the second representation, robotic device can refer to information that it had previously stored about the locations of those objects when it had obtained representation of the scene. Robotic device can take into account that its own location in the environment may have changed and recognize that first and second objects may be located in different areas of the second representation. The robotic device, by using previously stored information about the locations of the first and second objects, can automatically identify areas to scan closely (e.g., by zooming in, by slowly moving a camera through those areas) for certain objects.
[0077] As another example, the robotic device can build and / or load an obstacle model of a static local environment to use during planning. The obstacle model of the static local environment can be loaded quickly (e.g., have a small size) and can minimize robot scanning of the local environment. In some embodiments, the robotic device can use fiducial tags in the environment to build the obstacle map. For example, the robotic device may scan the environment and detect a fiducial tag on a wall. The robotic device can then scan a surroundingenvironment and generate one or more obstacles (e.g., a main wall obstacle on which the fiducial tag is located) from the fiducial, e.g., the coordinate associated with the fiducial tag in a map coordinate frame such that the six degrees of freedom (6DOF) location of the fiducial tag (e.g., X, Y, Z, roll, pitch, yaw). In some embodiments, the robotic device can execute a perception algorithm to determine one or more obstacles in the environment and store the obstacle based on the fiducial of a detected fiducial tag.
[0078] In some embodiments, a robotic device undergoing a demonstration for a skill (e.g., a demonstration of how to navigate through a doorway) can collect information during the demonstration that feeds into the layer(s) 822 of the map 820. The skill may be associated with a class of object (e.g., a door) that consistently exists at various locations throughout a building, e.g., as represented in the map 822. Characteristics and / or properties of the object can be recorded by various sensors on the robotic device and / or derived by the robotic device based on sensed information. For example, the robotic device can sense and / or perceive (e.g., via the perception model) a size of the doorway and determine that the doorway is a tight doorway. The robotic device can add this information regarding the doorway to the map 820, e.g., using one or more semantic labels, as raw or processed sensor information, as a derived rule, etc.
[0079] The map 820 can provide information that a robotic device can use while operating in a learning mode 802 or an execution mode 812. For example, a robotic device operating in the learning mode 802 can access the map 820 to obtain information regarding environmental constraint s), object(s), and / or social context(s) (e.g., similar to state information 332, object information 340, environmental constraint(s) 354, etc.). Such information can enable the robotic device to determine a location of object(s), identify characteristic(s) of object(s), analyze one or more environmental constraint s), select skill model(s) to use, prompt user(s) for input(s), etc., as described herein. Additionally or alternatively, a robotic device operating in the execution mode 812 can access the map 820 to obtain information regarding environmental constraint s), object(s), and / or social context(s), and use that information to evaluate an environment, arbitrate between different behaviors or skills, determining which skills to execute, and / or adapt behavior or skills to suit a particular environment.
[0080] In some embodiments, the map 820 can be centrally maintained for one or more robotic devices described herein. For example, the map 820 can be stored on a remote compute device (e.g., a server) and centrally hosted for a group of robotic devices that operate together, e.g., ina hospital. The map 820 can be updated as information (e.g., sensed information 804, 814 and derived information 806, 816) is received from the group of robotic devices as those robotic devices operate in learning mode 802 and / or execution mode 812, and be provided to each of the robotic devices as needed to execute actions, behavior, etc. As individual robotic devices within the group learn new information such that one or more layer(s) 822 of the map 820 are adapted, this information can be shared with the other robotic devices when they encounter a similar environment and / or execute a similar skill (e.g., action, behavior). In some embodiments, a local copy of the map 820 can be stored on each robotic device, which can be updated or synchronized (e.g., at predetermined intervals, during off hours or downtime) with a centrally maintained copy of the map 820. By regularly updating the map (e.g., with newly collected information, such as sensed information 804, 814 and derived information 806, 816) and / or sharing the map between robotic devices, each robotic device can have access to a more comprehensive map that provides it with more accurate information for interacting with and / or executing skills within its surrounding environment.Perception Algorithms1. Overview of Algorithm
[0081] In order to perceive and interact with a complex hospital environment including obstacles, the robotic device may receive a stream of images from the sensor(s) (e.g., cameras, light detection and ranging sensors (LiDARs), offline maps, etc.) and employ an algorithm (e.g., a perception algorithm) to process and analyze the sensor data (e.g., images, LiDARs, offline maps, etc.) live. The robotic device may employ a perception machine learning model, which processes inputs from more than one sensor stream (e.g., camera stream) simultaneously (or in parallel). The perception algorithm may include, for example, a neural network, a supervised learning model, nearest neighbor, support vector machine, deep learning model, etc. that is tailored for hospital environments. In some embodiments, the perception algorithm may include one or more convolutional neural networks (CNNs) and / or transformer architecture(s). In some embodiments, the model may be configured to receive a plurality of inputs. For example, the model may process data from 3 cameras including, for example, one RGB-depth (RGB-D) camera and two fisheye camera sensors. Parallel processing of the model inputs may reduce overall inference time, where inference time is defined as an amount of time it takes to produce model outputs, given the input sensor stream.
[0082] In some embodiments, the model may produce multiple outputs simultaneously. For example, the model may output (1) pixel-wise semantic segmentation, (2) object detection results, and / or (3) monocular depth estimation, enabling the robot to perceive its surroundings comprehensively. Semantic segmentation identifies and labels objects (e.g., people, wheelchairs, hospital beds, doors, etc.) Object detection detects dynamic and static objects in the scene in the form of 2D bounding boxes, thus facilitating trajectory planning around them. Monocular depth estimation provides 3D spatial awareness, which is crucial for navigating complex environments (e.g., a hospital environment).
[0083] The architecture of the model may maintain a consistent temporal relationship between inputs and outputs of the model. For example, each inference of the model at a point in time may correspond to the environment of the robot at that point in time. End-to-end learning may also be sped up as the model learns to optimize the processing of multiple inputs and the generation of multiple outputs. The model may combine or fuse inputs from a plurality of cameras and learn relationships between each of the cameras, for examplejoint representation and / or correlation between sensors. This model architecture also optimizes resource utilization by efficiently using available processing power.
[0084] The model may be trained using an apprenticeship learning (i.e., a knowledge distillation) paradigm, as shown in FIG. 9. As shown, image data may be input to both a teacher model, which is a large neural network, and a student model, which is a smaller neural network. Outputs from the teacher model (e.g., feature knowledge, response knowledge, and / or relation knowledge) may be fed into the student model to train the student model. For example, the outputs of the teacher model and the student model may be evaluated using a loss function (described in further detail with respect to FIG. 11 A-l 1C), and one or more outputs from the loss function indicating the difference in performance of the teacher model and / or the student model may then be fed into the student model. In some embodiments, after object detection, the model may automatically annotate the images (i.e., demarcate objects within a field of view or area of interest by adding a bounding box). In some embodiments, all detected objects in an image may be annotated automatically by the teacher model in a time between about 2 seconds to about 10 seconds, inclusive of all ranges and subranges therebetween. In comparison, manual annotation may take at least about 1 minute. The teacher- student learning architecture may provide cost-efficient training and / or significantly reduce annotation costs whilemaintaining performance. Additionally, the model may be trained with cross-attention dropout, as described in further detail below.2, Sensor Data Capture
[0085] The robotic device may include a plurality of sensors to provide a comprehensive view of the surrounding environment. In some embodiments, the perception algorithm may be configured to receive data from all of or any subset of the plurality of sensors of the robotic device. In some embodiments, the perception algorithm may be configured to receive any number of inputs (e.g., image or data stream from sensors) such as 1 input, 2 inputs, 3 inputs, 4 inputs, 5 inputs, 6 inputs, 7 inputs, 8 inputs, 9 inputs, 10 inputs, etc. In some embodiments, the perception algorithm can accept multiple images at one time. In some embodiments, the perception algorithm may be configured to detect images at, at least 5Hz. In some embodiments, each sensor may capture images at a frequency in a range of about 1Hz to about 30Hz, inclusive of all ranges and subranges therebetween, and each image may be fed into the perception algorithm. In some embodiments, each sensor may capture images at a frequency of about 5Hz to about 15Hz. In some embodiments, the perception algorithm may be configured to accept different types of sensor streams at one time (e.g., images, LiDAR, semantic maps, etc.).
[0086] In some embodiments, the model may wait for all of the inputs (e.g., images from all three of the cameras) to arrive before the model executes. In some embodiments, the model may execute before all of the inputs arrive. In some embodiments, a number of inputs and outputs may be adjusted. However, the size of the model and an amount of storage the model occupies is directly proportional to the number of inputs and outputs. The inference time is also directly proportional to the number of inputs and outputs. In some embodiments, the model having 3 inputs and 3 outputs enables the robot to comprehensively perceive the environment while minimizing the inference time.
[0087] The perception algorithm may be trained with cross-attention dropout, meaning that if one of the sensor streams is unavailable, the model may still produce reasonable outputs by taking advantage of sensor overlap. Cross-attention dropout ensures robustness of the model even in the event of sensor dropout. In some embodiments, the model may be trained with sensor dropout deliberately introduced such that the model learns to use sensor overlap to fill in the information if one of the sensors are unavailable.
[0088] In some embodiments, the perception algorithm may request data corresponding to each of the plurality of sensors depending on the perception task. In some embodiments, the perception algorithm may fuse the data from the plurality of sensors within a time slice. In some embodiments, the perception algorithm may check availability of the sensor to transmit the data stream before, during, or after requesting data from the corresponding sensor. In some embodiments, the perception algorithm may be configured to include a safety flag (e.g., “safety critical”) and define an escalation if one or more sensor stream having the safety flag are unavailable. This feature may be useful when the robotic device is in an environment in which perception from a particular sensor is critical to for the safety of a human in a vicinity of the robotic device and / or the robotic device itself.3, Model Architecture Design
[0089] FIG. 10 is a schematic block diagram depicting the model executed by the robotic device for processing sensor data to perceive its surrounding environment, according to an embodiment. This model architecture may use a large compute budget but allows overall lower latency and higher accuracy outputs. Methods to accommodate a large compute budget of a perception algorithm are described in further detail with regards to FIG. 12. The overall model architecture may include four main components: (1) a ResNet-50 CNN backbone to extract a compact feature representation, (2) an encoder-decoder transformer, (3) a feed forward network that makes object detection predictions and (4) two separate deconvolution heads for semantic segmentation and monocular depth prediction. As shown, the inputs to the model may include camera data, LiDAR data, and an offline map (e.g., similar to the map 820), but it can be appreciated that inputs can include more or less of the ones shown.
[0090] Each of the sensor inputs may then be fed into a respective CNN backbone. For example, an initial 3 channel image from a camera input may be fed into a conventional ResNet-50 CNN backbone to generate a lower-resolution activation map (also known as a feature map). Activation maps are generated for each of the sensor inputs separately. In other words, if the model intakes data from 3 cameras, 3 separate activation maps are generated corresponding to an image from each camera. The features from the activation maps may then be concatenated. The concatenated data is then input into an encoder / decoder layer. The encoder / decoder layer may include a transformer architecture and a feed forward network (FFN). Outputs from the transformer encoder are fed into one or more decoders, including, forexample, a semantic segmentation decoder and a depth decoder, and / or other decoders. The model may include one or more subsequent processing steps to transform the output(s) from the decoder to an image showing the semantic segmentation results, an image showing the monocular depth results, and / or any other output relevant to the robotic device.
[0091] FIGS. 11A-11C show a more detailed representation of a model architecture for processing sensor data to perceive its surrounding environment, according to an embodiment.. The model receives a first image (“Image A 10 A”) from the first camera of the robotic device and a second image (“Image B 10B”) from a second camera of the robotic device. Images 10A, 10B are each input into a first layer of a respective ResNet CNN Backbone 12A, 12B and an activation map (or feature map) for each image 14A, 14B is generated. In some embodiments, the activation map 14 A, 14B for each image may be input into a second ResNet CNN backbone 16 A, 16B. In some embodiments, a second activation map for each image 18 A, 18B may be generated. In some embodiments, a plurality of layers may be included between the activation map 14A, 14B and the second activation map 18A, 18B. In some embodiments, a number of layers may be determined based on experimentation.
[0092] The encoder / decoder layer 52, 54 includes a standard transformer architecture including a multi-head self-attention module and a feed forward network (FFN). Since the transformer architecture is permutation-invariant, positional encodings 20A, 20B may be added to the input of each attention layer, as shown in FIG. 1 IB. In order to achieve cross-attention across all input sensor streams, the learnt embeddings from all the sensor streams are fused together before moving them through the transformer encoder 52. Learnt embeddings are representations of input data (e.g., cameras, LiDARs, etc.) transformed into vectors of continuous numbers. Learn embeddings capture semantic relationships and properties of the input data and are optimized during model training process. As shown in FIGS. 1 IB-11C, the transformer decoder 54 has a first output 30A corresponding to the first image 10A and a second output 30B corresponding to the second image 10B.
[0093] Subsequent processing steps may be completed for image 10A and 10B. Posttransformer processing steps include a Bipartite Matching based 2D object detection head, a fully convolutional network based semantic segmentation head, and a fully convolutional network based monocular depth head. As shown in FIG. 11C, the Bipartite Matching based 2D object detection head may include a FNN 32A, 32B that outputs 2D bounding box predictionscorresponding to each object in the Images 10 A, 10B and a class score. For example, the images may be fed through the FNN 32, which may predict the normalized center coordinates of, the height, and the width of a 2D bounding box with respect to the input image 10 A. A linear layer may predict the class label (i.e., the categorization of the object) using softmax, for example. In some embodiments, the FNN 32 may be configured to determine a fixed-size set of N bounding boxes (i.e., the number of bounding boxes N is predetermined), where N is larger than the number of objects of interest in an image; therefore, an additional special class label 0 may be used to represent that no object is detected within a slot.
[0094] As shown in FIGS. 11A-11C, the semantic segmentation head intakes a fusion or combination of (1) a transformer decoder 54 output and (2) a CNN activation map output 14 A, 14B. For semantic segmentation, a locality of the input information should be preserved since the model performs a pixel-wise classification. Therefore, the model may preserve a locality of the input information into the semantic segmentation head. However, pooling layers of input CNN reduces a size of the extracted activation maps, and the transformer encoders and decoders 52, 54 maintain a size of the tensors provided by CNN activation maps (e.g., and therefore maintain the reduced size of the extracted activation maps from the images). Therefore, one or more up-sampling steps and / or deconvolutions 34A, 34B, 38 A, 38B may be used to ensure that the semantic segmentation and monocular depth estimation outputs match the input images sizes.
[0095] As pooling erases some information about the locality, skip connections may be employed in which relevant locality information is provided where it is useful or needed. For example, whenever an output activation map is up-sampled, the layers before the pooling layers (e.g., layer 14A and 14B), at which the activation map still has the original size is added pixelwise to the up-sampled activation map. In this way, the now up-sampled activation map 36A, 36B has both the locality detail from before pooling, but also the broad context information (e.g., details around segmentation and depth of the overall image) from after pooling and transformer encoder / decoder layers. This process allows a higher level of detail in the segmentation mask. The segmentation mask output may be a HxWxN sized tensor, where H and W are dimensions of the image and N is the number of classes of objects to be identified in the image. In order to determine a monocular depth output, the method may be similar to the semantic segmentation steps, except the output of the monocular depth head is a HxWxl sized tensor.
[0096] In some embodiments, one or more loss functions may be used to train the model. In some embodiments, three loss functions may be used to train the model: a first loss function may be used for evaluating object detection results, a second loss function may be used for evaluating semantic segmentation results, and a third loss function may be used for evaluating monocular depth perception results. In some embodiments, a Bipartite Matching loss function may be used for object detection. Unlike traditional object detectors that use anchor boxes and calculate individual losses for each anchor, a fixed number of bounding boxes for each image are used and a Bipartite Matching algorithm (Hungarian algorithm) may be employed to associate predicted bounding boxes with ground truth objects. This matching process minimizes the overall cost, considering both classification and bounding box regression errors. Once the matching is completed, the overall loss by aggregating the classification and bounding box regression losses for all matched pairs of predicted and ground truth boxes may be calculated.
[0097] In some embodiments, a Binary Cross Entropy with Logits Loss may be used for semantic segmentation. This loss function may be used for binary classification tasks by combining the sigmoid activation function and binary cross-entropy loss. Since the semantic segmentation head of the model described above produces a HxWxN tensor, where H and W are image dimensions and N is the number of classes, each of the N individual masks may be used as a one hot encoded binary segmentation mask. In some embodiments, a Mean Square Error (MSE) Loss function may be used for monocular depth prediction. The MSE Loss function computes the average squared difference between predictions of the model and the true target values.4, Division of Processing
[0098] FIG. 12 is a schematic block diagram of a robotic device 900 including a first processor904 and a second processor 905 that exchange information therebetween, according to an embodiment. The first processor 904 may be configured to execute all functions of the robotic device aside from the perception algorithm, whereas the second processor 905 may be configured to execute the perception algorithm. In some embodiments, the second processor905 may be configured to solely execute the perception algorithm. The first processor 904 may include a central processing unit, and the second processor 905 may include a graphicsprocessing unit such that the second processor 905 is equipped to process and analyze large amounts of image data.
[0099] The first processor 904 may receive data from the sensor(s) 970 (e.g., image capture device). In some embodiments, only the first processor 904 may receive data from the sensor(s), and the second processor 905 may only receive the sensor data 907 via the first processor 904. The first processor 904 may receive algorithm configurations 906 from the server 920, and the second processor 905 may receive (or may periodically request) algorithm configurations 906 from the first processor 904. After the second processor 905 executes the perception algorithm and produces algorithm outputs 908, these algorithm outputs 908 may be transmitted to the first processor 904 to inform other functions of the robotic device (e.g., navigation, manipulation, etc.) For example, camera images (e.g., from the sensor(s) 970) of a hospital bed moving closer to the robotic device may be streamed frame-by-frame from the first processor 904 to the second processor 905. Then the second processor 905 may detect objects including the hospital bed in each frame and send the results back to the first processor 904. The first processor 904 may use the object detection results to maneuver the robotic device out of the path of the hospital bed. In some embodiments, one or more sensors 970 may be coupled directly to the second processor 905, and the second processor 905 may run the perception algorithm and send the results to the first processor 904.
[0100] In some embodiments, the first processor 905 may send the algorithm outputs 908 to the first processor 904, and the first processor may send the algorithm outputs 908 to the server 920 for storage. In some embodiments, the first processor 904 may send updated information related to the perception algorithm to the second processor 905 such as new object classes. In some embodiments, the second processor 904 may receive this information from the server 920 and send the information to the second processor 905. For example, when the perception algorithm is to be updated with new information from server 920, the second processor 905 may request an updated configuration from the first processor 904 to provide the inputs for the updated perception algorithm.
[0101] With this design, no more than 25% of a single CPU core of the first processor 904 may be used while running the perception algorithm, leaving space for the robotic device 900 to execute other functions such as learning, planning, arbitration, etc. Implementing the perception algorithm on an independent processor minimizes changes to the operating systemor other installed software on the second processor 905. Furthermore, only the perception algorithm is to be loaded into the memory of the second processor 905, meaning there is no dynamic loading and unloading of different models onto the processor 905. This reduces resource usage in the second processor 905.
[0102] While two processors 904, 905 are depicted in FIG. 12, it can be appreciated that less or more processors can be used for this process. For example, a single processor can be used to run multiple processes including, for example, perception, in parallel. Alternatively, multiple processors can be used to run a single process, including, for example, perception.5, Results
[0103] In some embodiments, prior to the sensor data being fed into the perception algorithm, one or more pre-processing steps may be performed. For example, in embodiments in which fisheye images captured by a fisheye camera on the robotic device, a fisheye rectification step may occur, and the rectified fisheye image may be an input to the perception algorithm. Fisheye images may be captured because they provide a large field of view; however, fisheye images may be distorted especially toward outer edges of the image. Rectified fisheye images may improve performance of the perception algorithm in object detection and / or object tracking results. FIG. 13 A shows an original fisheye image on the right, and the rectified fisheye image on the left. FIG. 13B shows a rectified fisheye image on the left compared to the standard head camera image on the right. As shown in FIG. 13B, the rectified fisheye provides a larger field of view than the standard head camera image.
[0104] FIGS. 14A-14C show original images (top row), pixel-wise semantic segmentation results (center row), and monocular depth results (bottom row) from the perception algorithm. As shown, in FIG. 14A, the first fisheye camera images and results are shown in the left column, the standard head camera images results are shown in the middle column, and the second fisheye camera images and results are shown in the right column. The semantic segmentation head of the perception algorithm accurately segments different types of objects in the images. For example, the doors are shown in green or labeled D, and the hospital bed is shown in purple or labeled B (labels D and B are not outputs from the algorithm and added afterward for clarity). The monocular depth estimate images also accurately illustrate a gradient of depth through the hallway using a gradient in light intensity in the image.
[0105] As shown in FIG. 14B, the performance of the semantic segmentation head may be impacted by quality of the image captured. For example, the fisheye image (top) appears blurry, and therefore the perception algorithm does not produce as accurate and comprehensive segmentation results. However, the semantic segmentation head does successfully detect the person in the field of view of the robotic device.
[0106] FIG. 14C shows the output of the perception algorithm for a standard head camera image of a scene at a first time point Tl, a second time point T2, and a third time point T3. As shown, the semantic segmentation output (middle row) is able to accurately segment images when objects are stationary in the environment at Tl as well as when objects are dynamic as shown at T2 and T3. Additionally, the semantic segmentation outputs accurately group classes of outputs together such as humans, chairs, and desks. The monocular depth output of the perception algorithm (bottom row) also successfully indicates that the person is closer at T2 and further at T3.
[0107] FIG. 15 shows an example of the 2D bounding boxes with class labels with class scores overlayed. As shown, the perception algorithm correctly identifies the two trashcans and the two doors in the scene. The perception algorithm is able to classify both of the trashcans in this scene even though these are two different types of trashcans as well as both doors even though two different types of doors are present, demonstrating the robustness of the object identification outputs.
[0108] FIG. 16 is a flow chart of an example method 1000 of processing sensor information for navigation of a robotic device (e.g., robotic device 102, 200, 500, 900), according to an embodiment. In some embodiments, one or more steps of the method 1000 may be executed when the robotic device performs perception (e.g., as described at 706 in FIG. 7). The method 1000 may be similar to the algorithms described in FIGS. 10 and 11A-11C, and therefore, certain details are not described herein with respect to FIG. 16. In some embodiments, method 1000 may be executed by one or more processors (e.g., processors 204, 205, 904, 905) of the robotic device. In some embodiments, the method 1000 includes receiving one or more data streams capturing information of an environment, at 1002. In some embodiments, the processor may receive one or more data streams from one or more sensors of the robotic device. In some embodiments, a first processor of the robotic device may receive the one or more data streams from the one or more sensors and send signals corresponding to the data streams to a secondprocessor of the robotic device. In some embodiments, the one or more sensors may include image sensors of a robotic device (e.g., any of the robotic devices described herein). In some embodiments, the image sensors may include an RGB-D camera, a fisheye camera, a LiDAR sensor, or any sensor configured to capture information of the environment around the robotic device. In some embodiments, the one or more data streams may include a plurality of images captured over time. In some embodiments, the plurality of images may be captured at a predetermined frequency (e.g., 5 Hz to 10 Hz).
[0109] In some embodiments, the method 1000 may include executing an algorithm (e.g., a perception algorithm) based on the one or more data streams. The algorithm can be executed at a predetermined frequency. In some embodiments, the data streams may be continuously or periodically input into the algorithm and the algorithm may be periodically executed as the data streams are collected by the one or more sensors In some embodiments, the method 1000 may include waiting for each of the data streams captured to a time point or time window to be received from the one or more sensors before executing the algorithm. In some embodiments, the method 1000 includes inputting each data stream into a respective feature extraction model configured to output a plurality of feature maps associated with the data streams, at 1004. As described herein, a feature extraction model can include machine learning models, autoencoders, neural networks, dimensionality reduction models, or any type of model configured to extract features from raw or processed data. In some embodiments, the data streams may be captured from each sensor concurrently, and each data stream may be concurrently input into the feature extraction model(s). In some embodiments, the method 1000 may include waiting for all of the data streams to be received by the processor before inputting each data stream into a respective feature extraction model. In some embodiments, the method may include aligning the plurality of data streams in time and concurrently inputting at least a portion of the data streams into the feature extraction model(s). In some embodiments, additional inputs or information such as, for example, an offline map, fiducial marker coordinates, user inputs, etc., can be input into the algorithm (e.g., the feature extraction model).
[0110] In some embodiments, the feature extraction model(s) may be configured to output a plurality of feature maps (i.e., activation maps as described with respect to FIGS. 11A-11C) associated with the data streams. For example, the model may output a feature map or an activation map of an image (i.e., at a given time point) for each data stream. The algorithm mayinclude one or more convolutional neural networks (CNNs) and / or transformer architecture(s). In some embodiments, the feature extraction model may include a neural network. In some embodiments, the feature extraction model may include a CNN. In some embodiments, an image from each sensor may be input into a respective neural network, each respective neural network may be configured to output a feature map of the image. In some embodiments, the features of each feature map may be combined (i.e., concatenated) into a data structure. In some embodiments, the features of a feature map associated with a respective data stream may be concatenated. In some embodiments, the features of the feature maps across all the data streams may be concatenated. In some embodiments, the feature maps may be concatenated and transformed into an array or vector. In some embodiments, combining features of the feature maps may enable cross-attention across the plurality of data streams. In some embodiments, the feature maps may be tagged with positional encodings (e.g., 20 A, 20B as described in FIGS. 11 A). In some embodiments, the method 1000 may include inputting plurality of feature maps (e.g., with or without positional encodings and / or with or without concatenation) into a transformer encoder configured to generate a plurality of encoded outputs, at 1006. In some embodiments, the concatenated feature maps (e.g., concatenated by data stream or across data streams) may be input into the transformer encoder. In some embodiments, the feature maps associated with all data streams may be input into a single transformer encoder.[OHl] The method may include inputting the encoded outputs into one or more decoder model configured to generate one or more decoded outputs, at 1010. In some embodiments, the method 1000 may include inputting the plurality of encoded outputs into a plurality of decoders to generate a plurality of decoded outputs. In some embodiments, the encoded outputs may be input into a transformer decoder, and then the outputs from the transformer decoder may be input into a plurality of different decoder models. In some embodiments, the transformer decoder outputs may be input into a plurality of decoder models concurrently. In some embodiments, each decoder model of the plurality of decoder models may be configured to output a different representation of the environment. The representations of the environment may include, for example, analyzed or transformed images of the environment, images of the environment including one or more annotations, etc.
[0112] In some embodiments, the decoder model may include a first decoder model configured to output object detection data. For example, the first decoder model may include a feed forward network (FNN) configured to output at least one of object detection predictions and2D bounding box predictions for objects detected in the environment (e.g., “class score” and “bounding box” as shown in FIG. 11C). In some embodiments, the first decoder model may perform Bipartite Matching based 2D object detection. The object detection data may include predictions of a center point, a height, and a width of each object detected in an image. In some embodiments, the method 1000 may include determining movement of an object in the environment by comparing bounding boxes of the object across images captured at multiple time points. In some embodiments, the decoder models may include a second decoder model configured to output semantic segmentation data. In some embodiments, the second decoder model may include one or more deconvolutional layers (e.g., two deconvolutional layers) configured to up-sample the decoded output from the transformer decoder, and semantic segmentation may be performed on the up-sampled output. The semantic segmentation data may include identification, classification, and / or labeling of objects in the environment. In some embodiments, the decoder models may include a third decoder model configured to output depth perception data (e.g., monocular depth perception data). In some embodiments, the third decoder model may include one or more deconvolutional layers (e.g., two deconvolutional layers) configured to up-sample the decoded output from the transformer decoder, and monocular depth perception may be performed on the up-sampled output.
[0113] In some embodiments, the first decoder model, the second decoder model, and the third decoder model may be configured to perform object detection, the performing semantic segmentation, and the performing depth perception at the same time. In some embodiments, the perception algorithm may be executed at predetermined intervals to provide a representation of the environment as the robotic device moves through the therethrough. In some embodiments, object detection, semantic segmentation, and depth perception can be performed as the robotic device moves through the environment. In some embodiments, the method 1000 may include determining a trajectory of the robotic device based on the different representations of the environment (e.g., at least one of the object detection data, the semantic segmentation data, the depth perception data), at 1012. In some embodiments, the determining the trajectory may include determining a trajectory of the base or the manipulating element of the robotic device such that the robotic device can navigate through the environment.
[0114] In some embodiments, the robotic device may include a plurality of sensors, each sensor having a field of view that partially overlaps such that the algorithm can produce usable outputs when data from any one sensor from the one or more sensors is absent at a given time. In someembodiments, the algorithm is trained using an apprentice learning paradigm. In some embodiments, the method 1000 may include outputting a safety flag (e.g., to a user) if any one of the plurality of data streams are unavailable to prevent the robotic device from unsafely navigating. For example, if a data stream is unavailable for a predetermined period of time and / or if more than one data stream is unavailable, the method 1000 may include outputting a flag or indication that an analyzed representation of the environment cannot be generated. Therefore, the robotic device may be caused to stop navigation until the data streams are available. The decoded outputs from the method 1000 may enable the robotic device to safely interact with and / or navigate around objects in the environment. The method 1000 may enable the robotic device to generate plans and perform tasks based on periodically updated representations of the environment such that the robotic device can operate in a dynamic environment.6, Use Cases
[0115] The perception algorithm described herein can be used for comprehensive perception. The model may deliver rich information encompassing semantic understanding, object localization, and depth estimation. The model aids in identifying dynamic objects in the scene such as people, wheelchairs, hospital beds etc. and planning a safe trajectory around them.
[0116] The perception algorithm may be used for real-time safety measures. For example, the algorithm enables the robotic device to detect and respond to dynamic objects, ensuring safe navigation and operation in the presence of people. More specifically, the perception algorithm may include a safety feature that halts operation of the manipulating element of the robotic device when people are detected in close proximity, mitigating potential hazards.
[0117] The perception algorithm may be used to provide the robotic device environmental awareness: the model may identify environmental cues such as open doors and elevators, aiding the robot in decision-making and navigation. The algorithm may also aid in identifying the floor on which the robotic device is located by looking at visual cues inside an elevator, for example. The algorithm may communicate with the navigation system of the robotic device such that the robotic device is prepared for an arriving elevator when the robotic device is waiting in an elevator bay. This saves valuable time when the robotic device is boarding an elevator quickly, given the elevator door timing constraints.
[0118] The perception algorithm may be used to detect manipulation success. More specifically, the perception algorithm may provide a feedback mechanism to determine if manipulation by the manipulating element succeeded. For example, the algorithm can identify if an elevator door call button was successfully pressed or not, and if the robotic device should reattempt manipulation behavior or move on to the next behavior. Implementing this perception algorithm allows the robotic device to refine its estimates of locations of door access devices, call button panels, and elevator button panels instead of solely relying on pre-recorded information, which improves overall manipulation accuracy.
[0119] The perception algorithm may aid the robotic device in making socially aware behavior changes. The model provides 360° semantic information to the robot, and semantic scene understanding allows the robotic device to build 3D representation of the environment in realtime or near-real time. This allows the robotic device to adapt its behavior, given the social context. The perception algorithm may enable the robotic device to yield to wheelchairs, hospital beds carrying patients etc. The perception algorithm may also provide information to cue the robotic device to ask for help for certain tasks instead of causing inconvenience to people around the robot (e.g., ask for assistance in pressing the elevator floor button if the elevator is too busy). The perception algorithm also provides information so that the robotic device can park at appropriate locations to ensure the robot is not blocking any doorways, access to fire extinguishers etc., thereby enabling the robotic device to operate in busy settings without disruption.Enumerated Embodiments:
[0120] 1. A robotic device, comprising: a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; and one or more processors operatively coupled to the base, the one or more manipulating elements, and the set of sensors, the one or more processors configured to: receive, from the set of sensors, a plurality of data streams capturing information of an environment around the robotic device; input each data stream from the plurality of data streams into a respective feature extraction model from a plurality of feature extraction models to obtain a plurality of feature maps associated with the plurality of data streams; input the plurality of feature maps into a transformer encoder configured to generate a plurality of encoded outputs; input the plurality of encoded outputs into a plurality of decoders to generate a plurality of decodedoutputs, each decoded output associated with a different representation of the environment; and determine a trajectory of the robotic device through the environment based on the different representations of the environment.
[0121] 2. The robotic device of embodiment 1, wherein each data stream includes a plurality of images of the environment over time.
[0122] 3. The robotic device of embodiment 2, wherein the processor is configured to receive the plurality of images at a frequency between about 5 Hertz (Hz) and about 15 Hz.
[0123] 4. The robotic device of embodiment 1, wherein the different representations of the environment include a semantic segmentation data, a monocular depth perception data, and object detection data.
[0124] 5. The robotic device of embodiment 1, wherein the set of sensors includes at least one of a red-green-blue-depth (RGB-D) camera or a LiDAR sensor.
[0125] 6. The robotic device of embodiment 1, wherein the set of sensors are configured to have at least partially overlapping field of views such that at least a subset of the plurality of decoders produce usable outputs when data from any one sensor from the set of sensors is absent.
[0126] 7 The robotic device of embodiment 1, wherein the plurality of feature extractions models includes convolutional neural networks (CNN).
[0127] 8. The robotic device of embodiment 1, wherein the processor is further configured to associate, before inputting the plurality of feature maps into the transformer encoder, positional encodings with each feature map.
[0128] 9. The robotic device of embodiment 1, wherein the plurality of decoders includes a feed forward network configured to output at least one of object detection predictions and 2 dimensional (2D) bounding box predictions for objects detected in the environment.
[0129] 10. The robotic device of embodiment 9, wherein the processor is further configured to: identify moving objects in the environment based on at least one of the object detection predictions or the 2D bounding box predictions.
[0130] 11. The robotic device of embodiment 1, wherein the plurality of decoders includes at least one of: one or more deconvolutional layers configured to output a semantic segmentation and a monocular depth prediction of each data stream.
[0131] 12. The robotic device of embodiment 1, wherein the processor is further configured to: wait for all of the plurality of data streams to be received before inputting each data stream into the respective feature extraction model.
[0132] 13. The robotic device of embodiment 1, wherein the information of the environment includes a semantic map of the environment.
[0133] 14. The robotic device of embodiment 1, wherein the processor is configured to output a safety flag if any one of the plurality of data streams are unavailable to prevent the robotic device from navigating unsafely.
[0134] 15. A robotic device, comprising: a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; and one or more processors operatively coupled to the base, the one or more manipulating elements, and the set of sensors, the one or more processors configured to: receive, from each sensor of the set of sensors, information of an environment around the robotic device; input the information of the environment into an encoder model, the encoder model configured to generate one or more encoded outputs; input the one or more encoded outputs into: a first decoder model configured to output object detection data, a second decoder model configured to output semantic segmentation data, and a third decoder model configured to output monocular depth data; and determine a trajectory of the robotic device based on the object detection data, the semantic segmentation data, and the monocular depth data.
[0135] 16. The robotic device of embodiment 15, wherein the inputting information of the environment into the encoder model includes inputting information from each sensor from the set of sensors into a feature extraction model configured to output a plurality of feature maps.
[0136] 17. The robotic device of embodiment 16, wherein the feature extraction model includes a CNN.
[0137] 18. The robotic device of embodiment 16, wherein the encoder model includes a transformer encoder configured to receive the plurality of feature maps and output the one or more encoded outputs.
[0138] 19. The robotic device of embodiment 15, wherein the information of the environment includes a plurality of images of the environment over time.
[0139] 20. The robotic device of embodiment 15, wherein the set of sensors includes at least one of a red-green-blue-depth (RGB-D) camera or a LiDAR sensor.
[0140] 21 The robotic device of embodiment 15, wherein the object detection data includes at least one of object detection predictions or 2D bounding box predictions.
[0141] 22. The robotic device of embodiment 21, wherein the processor is further configured to: identify moving objects in the environment based on at least one of the object detection predictions or the 2D bounding box predictions.
[0142] 23. The robotic device of embodiment 15, wherein the processor is configured to output a safety flag if any one of the plurality of data streams are unavailable to prevent the robotic device from navigating unsafely.
[0143] 24. A robotic device, comprising: a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; a first processor operatively coupled to the base, the one or more manipulating elements, and the set of sensors, the first processor configured to receive sensor data from the set of sensors, the sensor data capturing information of an environment around the robotic device; and a second processor operatively coupled to the first processor, the second processor configured to: receive, from the first processor, signals corresponding to the sensor data; input the sensor data into an encoder-decoder model configured to generate a plurality of outputs, each output of the plurality of outputs associated with a different representation of the environment, the first processor being configured to determine a plan for executing a task based on the plurality of outputs and to control the base and the one or more manipulating elements to execute the plan.
[0144] 25. The robotic device of embodiment 24, wherein the set of sensors includes at least one of a red-green-blue-depth (RGB-D) camera or a LiDAR sensor.
[0145] 26. The robotic device of embodiment 24, wherein the sensor data includes a plurality of data streams, each data stream including a plurality of images of the surrounding environment over time.
[0146] 27. The robotic device of embodiment 24, wherein the encoder-decoder model includes a plurality of decoders configured to generate a plurality of decoded outputs, each decoded output associated with a different representation of the environment.
[0147] 28. The robotic device of embodiment 24, wherein the plurality of outputs includes object detection data including at least one of object detection predictions or 2D bounding box predictions.
[0148] 29. The robotic device of embodiment 28, wherein at least one of the first processor or the second processor is further configured to: identify moving objects in the environment based on at least one of the object detection predictions or the 2D bounding box predictions.
[0149] 30. The robotic device of embodiment 24, wherein the second processor is operatively coupled to a memory configured to store the encoder-decoder model thereon, the memory being separate from the first processor.
[0150] 31. The robotic device of embodiment 24, wherein the first processor is operatively coupled to a server, the first processor configured to receive algorithm configurations from the server and send signals corresponding to algorithm outputs to the server for storage.
[0151] 32. The robotic device of embodiment 24, wherein no more than 25% of a single processing core of the first processor is used while the encoder-decoder model is executed, thereby leaving processing power for the robotic device to execute other functions.
[0152] 33. A method, comprising: receiving, from a set of sensors, a plurality of data streams capturing information of an environment around a robotic device; inputting each data stream from the plurality of data streams into a respective feature extraction model from a plurality of feature extraction models to obtain a plurality of feature maps associated with the plurality of data streams; inputting the plurality of feature maps into a transformer encoder configured to generate a plurality of encoded outputs; inputting the plurality of encodedoutputs into a plurality of decoders to generate a plurality of decoded outputs, each decoded output associated with a different representation of the environment; and determining a trajectory of the robotic device through the environment based on the different representations of the environment.
[0153] 34. The method of embodiment 33, wherein each data stream includes a plurality of images of the environment over time.
[0154] 35. The method of embodiment 33, wherein the different representations of the environment include a semantic segmentation data, a monocular depth perception data, and object detection data.
[0155] 36. The method of embodiment 33, wherein the plurality of feature extractions models includes convolutional neural networks (CNN).
[0156] 37. The method of embodiment 33, wherein the plurality of decoders includes at least one of a feed forward network or one or more deconvolutional layers.
[0157] 38. The method of embodiment 33, wherein the method further includes: associating, before inputting the plurality of feature maps into the transformer encoder, positional encodings with each feature map.
[0158] 39. The method of embodiment 33, wherein the plurality of decoders includes a feed forward network configured to output at least one of object detection predictions and 2 dimensional (2D) bounding box predictions for objects detected in the environment.
[0159] 40. A method, comprising: receiving, from each sensor of the set of sensors, information of an environment around the robotic device; inputting the information of the environment into an encoder model, the encoder model configured to generate one or more encoded outputs; inputting the one or more encoded outputs into: a first decoder model configured to output object detection predictions, a second decoder model configured to output semantic segmentation data, and a third decoder model configured to output monocular depth data; and determining a trajectory of the robotic device based on the object detection predictions data, the semantic segmentation data, and the monocular depth data.
[0160] 41. The method of embodiment 40, wherein the inputting information of the environment into the encoder model includes inputting information from each sensor from the set of sensors into a feature extraction model configured to output a plurality of feature maps.
[0161] 42. The method of embodiment 41, wherein the encoder model includes a transformer encoder configured to receive the plurality of feature maps and output the one or more encoded outputs.
[0162] 43. The method of embodiment 40, wherein the information of the environment includes a plurality of images of the environment over time.
[0163] 44 The method of embodiment 40, wherein the object detection data includes at least one of object detection predictions or 2D bounding box predictions.
[0164] 45. The method of embodiment 33, wherein the method further includes: identifying moving objects in the environment based on at least one of the object detection predictions or the 2D bounding box predictions.
[0165] 46. A method, comprising: receiving, at a first processor, sensor data from a set of sensors of a robotic device, the sensor data capturing information of an environment around the robotic device; and receiving, at a second processor, signals corresponding to the sensor data from the first processor; inputting, via the second processor, the sensor data into an encoder-decoder model configured to generate a plurality of outputs, each output of the plurality of outputs associated with a different representation of the environment; and determining, via the first processor, a plan for executing a task based on the plurality of outputs and to control a base and one or more manipulating elements of the robotic device to execute the plan.
[0166] 47. The method of embodiment 46, wherein the sensor data includes a plurality of data streams, each data stream including a plurality of images of the surrounding environment over time.
[0167] 48. The method of embodiment 46, wherein the encoder-decoder model includes a plurality of decoders configured to generate a plurality of decoded outputs, each decoded output associated with a different representation of the environment.
[0168] 49. The method of claim 46, wherein the second processor is operatively coupled to a memory configured to store the encoder-decoder model thereon, the memory being separate from the first processor.
[0169] 50. The method of embodiment 46, wherein the first processor is operatively coupled to a server, the first processor configured to receive algorithm configurations from the server and send signals corresponding to algorithm outputs to the server for storage.
[0170] Various concepts may be embodied as one or more methods, of which at least one example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments. Put differently, it is to be understood that such features may not necessarily be limited to a particular order of execution, but rather, any number of threads, processes, services, servers, and / or the like that may execute serially, asynchronously, concurrently, in parallel, simultaneously, synchronously, and / or the like in a manner consistent with the disclosure. As such, some of these features may be mutually contradictory, in that they cannot be simultaneously present in a single embodiment. Similarly, some features are applicable to one aspect of the innovations, and inapplicable to others.
[0171] In addition, the disclosure may include other innovations not presently described. Applicant reserves all rights in such innovations, including the right to embodiment such innovations, file additional applications, continuations, continuations-in-part, divisionals, and / or the like thereof. As such, it should be understood that advantages, embodiments, examples, functional, features, logical, operational, organizational, structural, topological, and / or other aspects of the disclosure are not to be considered limitations on the disclosure as defined by the embodiments or limitations on equivalents to the embodiments. Depending on the particular desires and / or characteristics of an individual and / or enterprise user, database configuration and / or relational model, data type, data transmission and / or network framework, syntax structure, and / or the like, various embodiments of the technology disclosed herein may be implemented in a manner that enables a great deal of flexibility and customization as described herein.
[0172] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0173] As used herein, in particular embodiments, the terms “about” or “approximately” when preceding a numerical value indicates the value plus or minus a range of 10%. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the disclosure. That the upper and lower limits of these smaller ranges can independently be included in the smaller ranges is also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.
[0174] The phrase “and / or,” as used herein in the specification and in the embodiments, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0175] As used herein in the specification and in the embodiments, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the embodiments, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used hereinshall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the embodiments, shall have its ordinary meaning as used in the field of patent law.
[0176] As used herein in the specification and in the embodiments, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0177] In the embodiments, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.
[0178] While specific embodiments of the present disclosure have been outlined above, many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, the embodiments set forth herein are intended to be illustrative, not limiting. Various changes may be made without departing from the spirit and scope of the disclosure. Where methods and steps described above indicate certain events occurring in a certain order, those of ordinary skill in the art having the benefit of this disclosure would recognize that theordering of certain steps may be modified and such modification are in accordance with the variations of the invention. Additionally, certain of the steps may be performed concurrently in a parallel process when possible, as well as performed sequentially as described above. The embodiments have been particularly shown and described, but it will be understood that various changes in form and details may be made.
Claims
Claims:
1. A robotic device, comprising: a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; and one or more processors operatively coupled to the base, the one or more manipulating elements, and the set of sensors, the one or more processors configured to: receive, from the set of sensors, a plurality of data streams capturing information of an environment around the robotic device; input each data stream from the plurality of data streams into a respective feature extraction model from a plurality of feature extraction models to obtain a plurality of feature maps associated with the plurality of data streams; input the plurality of feature maps into a transformer encoder configured to generate a plurality of encoded outputs; input the plurality of encoded outputs into a plurality of decoders to generate a plurality of decoded outputs, each decoded output associated with a different representation of the environment; and determine a trajectory of the robotic device through the environment based on the different representations of the environment.
2. The robotic device of claim 1, wherein each data stream includes a plurality of images of the environment over time.
3. The robotic device of claim 2, wherein the processor is configured to receive the plurality of images at a frequency between about 5 Hertz (Hz) and about 15 Hz.
4. The robotic device of claim 1, wherein the different representations of the environment include a semantic segmentation data, a monocular depth perception data, and object detection data.
5. The robotic device of claim 1, wherein the set of sensors includes at least one of a red- green-blue-depth (RGB-D) camera or a LiDAR sensor.
6. The robotic device of claim 1, wherein the set of sensors are configured to have at least partially overlapping field of views such that at least a subset of the plurality of decoders produce usable outputs when data from any one sensor from the set of sensors is absent.
7. The robotic device of claim 1, wherein the plurality of feature extractions models includes convolutional neural networks (CNN).
8. The robotic device of claim 1, wherein the processor is further configured to: associate, before inputting the plurality of feature maps into the transformer encoder, positional encodings with each feature map.
9. The robotic device of claim 1, wherein the plurality of decoders includes a feed forward network configured to output at least one of object detection predictions and 2 dimensional (2D) bounding box predictions for objects detected in the environment.
10. The robotic device of claim 9, wherein the processor is further configured to: identify moving objects in the environment based on at least one of the object detection predictions or the 2D bounding box predictions.
11. The robotic device of claim 1, wherein the plurality of decoders includes at least one of: one or more deconvolutional layers configured to output a semantic segmentation and a monocular depth prediction of each data stream.
12. The robotic device of claim 1, wherein the processor is further configured to: wait for all of the plurality of data streams to be received before inputting each data stream into the respective feature extraction model.
13. The robotic device of claim 1, wherein the information of the environment includes a semantic map of the environment.
14. The robotic device of claim 1, wherein the processor is configured to output a safety flag if any one of the plurality of data streams are unavailable to prevent the robotic device from navigating unsafely.
15. A robotic device, comprising: a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; and one or more processors operatively coupled to the base, the one or more manipulating elements, and the set of sensors, the one or more processors configured to: receive, from each sensor of the set of sensors, information of an environment around the robotic device; input the information of the environment into an encoder model, the encoder model configured to generate one or more encoded outputs; input the one or more encoded outputs into: a first decoder model configured to output object detection data, a second decoder model configured to output semantic segmentation data, and a third decoder model configured to output monocular depth data; and determine a trajectory of the robotic device based on the object detection data, the semantic segmentation data, and the monocular depth data.
16. The robotic device of claim 15, wherein the inputting information of the environment into the encoder model includes inputting information from each sensor from the set of sensors into a feature extraction model configured to output a plurality of feature maps.
17. The robotic device of claim 16, wherein the feature extraction model includes a CNN.
18. The robotic device of claim 16, wherein the encoder model includes a transformer encoder configured to receive the plurality of feature maps and output the one or more encoded outputs.
19. The robotic device of claim 15, wherein the information of the environment includes a plurality of images of the environment over time.
20. The robotic device of claim 15, wherein the set of sensors includes at least one of a red-green-blue-depth (RGB-D) camera or a LiDAR sensor.21 The robotic device of claim 15, wherein the object detection data includes at least one of object detection predictions or 2D bounding box predictions.
22. The robotic device of claim 21, wherein the processor is further configured to: identify moving objects in the environment based on at least one of the object detection predictions or the 2D bounding box predictions.
23. The robotic device of claim 15, wherein the processor is configured to output a safety flag if any one of the plurality of data streams are unavailable to prevent the robotic device from navigating unsafely.
24. A robotic device, comprising: a base supported on a transport element; one or more manipulating elements coupled to the base and including an end effector; a set of sensors; a first processor operatively coupled to the base, the one or more manipulating elements, and the set of sensors, the first processor configured to receive sensor data from the set of sensors, the sensor data capturing information of an environment around the robotic device; and a second processor operatively coupled to the first processor, the second processor configured to: receive, from the first processor, signals corresponding to the sensor data; input the sensor data into an encoder-decoder model configured to generate a plurality of outputs, each output of the plurality of outputs associated with a different representation of the environment, the first processor being configured to determine a plan for executing a task based on the plurality of outputs and to control the base and the one or more manipulating elements to execute the plan.
25. The robotic device of claim 24, wherein the set of sensors includes at least one of a red-green-blue-depth (RGB-D) camera or a LiDAR sensor.
26. The robotic device of claim 24, wherein the sensor data includes a plurality of data streams, each data stream including a plurality of images of the surrounding environment over time.
27. The robotic device of claim 24, wherein the encoder-decoder model includes a plurality of decoders configured to generate a plurality of decoded outputs, each decoded output associated with a different representation of the environment.
28. The robotic device of claim 24, wherein the plurality of outputs includes object detection data including at least one of object detection predictions or 2D bounding box predictions.
29. The robotic device of claim 28, wherein at least one of the first processor or the second processor is further configured to: identify moving objects in the environment based on at least one of the object detection predictions or the 2D bounding box predictions.
30. The robotic device of claim 24, wherein the second processor is operatively coupled to a memory configured to store the encoder-decoder model thereon, the memory being separate from the first processor.
31. The robotic device of claim 24, wherein the first processor is operatively coupled to a server, the first processor configured to receive algorithm configurations from the server and send signals corresponding to algorithm outputs to the server for storage.
32. The robotic device of claim 24, wherein no more than 25% of a single processing core of the first processor is used while the encoder-decoder model is executed, thereby leaving processing power for the robotic device to execute other functions.
33. A method, comprising: receiving, from a set of sensors, a plurality of data streams capturing information of an environment around a robotic device;inputting each data stream from the plurality of data streams into a respective feature extraction model from a plurality of feature extraction models to obtain a plurality of feature maps associated with the plurality of data streams; inputting the plurality of feature maps into a transformer encoder configured to generate a plurality of encoded outputs; inputting the plurality of encoded outputs into a plurality of decoders to generate a plurality of decoded outputs, each decoded output associated with a different representation of the environment; and determining a trajectory of the robotic device through the environment based on the different representations of the environment.
34. The method of claim 33, wherein each data stream includes a plurality of images of the environment over time.
35. The method of claim 33, wherein the different representations of the environment include a semantic segmentation data, a monocular depth perception data, and object detection data.
36. The method of claim 33, wherein the plurality of feature extractions models includes convolutional neural networks (CNN).
37. The method of claim 33, wherein the plurality of decoders includes at least one of a feed forward network or one or more deconvolutional layers.
38. The method of claim 33, wherein the method further includes: associating, before inputting the plurality of feature maps into the transformer encoder, positional encodings with each feature map.
39. The method of claim 33, wherein the plurality of decoders includes a feed forward network configured to output at least one of object detection predictions and 2 dimensional (2D) bounding box predictions for objects detected in the environment.
40. A method, comprising:receiving, from each sensor of the set of sensors, information of an environment around the robotic device; inputting the information of the environment into an encoder model, the encoder model configured to generate one or more encoded outputs; inputting the one or more encoded outputs into: a first decoder model configured to output object detection predictions, a second decoder model configured to output semantic segmentation data, and a third decoder model configured to output monocular depth data; and determining a trajectory of the robotic device based on the object detection predictions data, the semantic segmentation data, and the monocular depth data.
41. The method of claim 40, wherein the inputting information of the environment into the encoder model includes inputting information from each sensor from the set of sensors into a feature extraction model configured to output a plurality of feature maps.
42. The method of claim 41, wherein the encoder model includes a transformer encoder configured to receive the plurality of feature maps and output the one or more encoded outputs.
43. The method of claim 40, wherein the information of the environment includes a plurality of images of the environment over time.44 The method of claim 40, wherein the object detection data includes at least one of object detection predictions or 2D bounding box predictions.
45. The method of claim 40, wherein the method further includes: identifying moving objects in the environment based on at least one of the object detection predictions or the 2D bounding box predictions.
46. A method, comprising: receiving, at a first processor, sensor data from a set of sensors of a robotic device, the sensor data capturing information of an environment around the robotic device; and receiving, at a second processor, signals corresponding to the sensor data from the first processor;inputting, using the second processor, the sensor data into an encoder-decoder model configured to generate a plurality of outputs, each output of the plurality of outputs associated with a different representation of the environment; and determining, using the first processor, a plan for executing a task based on the plurality of outputs and to control a base and one or more manipulating elements of the robotic device to execute the plan.
47. The method of claim 46, wherein the sensor data includes a plurality of data streams, each data stream including a plurality of images of the surrounding environment over time.
48. The method of claim 46, wherein the encoder-decoder model includes a plurality of decoders configured to generate a plurality of decoded outputs, each decoded output associated with a different representation of the environment.
49. The method of claim 46, wherein the second processor is operatively coupled to a memory configured to store the encoder-decoder model thereon, the memory being separate from the first processor.
50. The method of claim 46, wherein the second processor is operatively coupled to a memory configured to store the encoder-decoder model thereon, the memory being separate from the first processor.
Citation Information
Patent Citations
Systems, apparatus, and methods for robotic learning and execution of skills
WO2020047120A1
Systems, apparatuses, and methods for robotic learning and execution of skills including navigation and manipulation functions
WO2022170279A1
Image marking method, track planning method, marking model, device and system
CN114102575A
Sensor data fusion using cross-modal transformer
US11921824B1