Surface Touch Point Monitoring for Precision Cleaning

The key point representation of actors through sensor generation is expressed by the robot system accurately identifying and cleaning touched surfaces, solving the problem of low identification and cleaning efficiency in the prior art, and achieving efficient and accurate cleaning effects.

CN116669913BActive Publication Date: 2025-08-05GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180087649.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-22
Filing Date
2021-10-01
Publication Date
2025-08-05
Estimated Expiration
2041-10-01

AI Technical Summary

Technical Problem

In the prior art, when a robot system recognizes and cleanses the surface touched by the actor, it is difficult to efficiently and accurately identify and clean the touched surface, resulting in waste of resources and cleaning time too long.

Method used

The environment data is collected through sensors on the robot device, and a representation of the actor's key points is generated, whether a specific static feature lies within the threshold distance of the key points, and the map is updated to indicate that the part needs to be cleaned, and the robot device cleans the corresponding part according to the map.

Benefits of technology

It realizes that the robot system can accurately identify and clean the touched surfaces, reduce the use and time of cleaning resources, improve cleaning efficiency, avoid unnecessary cleaning, and adapt to dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116669913B_ABST
    Figure CN116669913B_ABST
Patent Text Reader

Abstract

A system includes a robotic device, a sensor disposed on the robotic device, and circuitry configured to perform operations. The operations include determining a map representing static features of an environment and receiving sensor data representing the environment from a sensor. The operations also include determining a representation of an agent within the environment based on the sensor data, wherein the representation includes key points representing corresponding body positions of the agent. The operations also include determining that a portion of a particular static feature is within a threshold distance of the particular key point, and based on this, updating the map to indicate that the portion is to be cleaned. The operations also include causing the robotic device to clean the portion of the particular static feature based on the updated map.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. patent application No. 17 / 131,252, filed on December 22, 2020, entitled “Monitoring of Surface TouchPoints for Precision Cleaning,” which is incorporated herein by reference as if fully set forth in this specification. Technical Field

[0003] This application relates to robots. Background Art

[0004] As technology advances, various types of robotic devices are being created to perform a variety of functions that can assist users. These devices can be used in applications such as material handling, transportation, welding, assembly, and distribution. Over time, the operation of these robotic systems has become more intelligent, efficient, and intuitive. As robotic systems become increasingly prevalent in many aspects of modern life, people expect them to be efficient. Consequently, the demand for efficient robotic systems is driving innovation in actuators, motion, sensing technologies, and component design and assembly. Summary of the Invention

[0005] The robot can be used to identify and clean surfaces that have been touched by an actor (e.g., a person, dog, cat, or another robot). The control system can receive sensor data representing an actor operating in the environment from sensors on the robot. The control system can determine the actor's representation by generating a plurality of key points representing the actor's corresponding body positions (such as fingers, hands, elbows, etc.) based on the sensor data. Based on the location of at least some of these key points, the control system can determine whether the actor has touched a feature in the environment. Specifically, the control system can determine whether any feature in the environment has entered a threshold distance of at least one specific key point among these key points (e.g., a hand key point). When a portion of an environmental feature is within the threshold distance of a specific key point, the portion can be designated for cleaning, and a map of the environment can be updated to indicate that the portion is to be cleaned. Once the actor moves to a different part of the environment, the robot can clean any portion of the environment that has been designated for cleaning based on the map. Once a portion has been cleaned, the map can be updated again to indicate that the portion has been cleaned.

[0006] In a first example embodiment, a system may include a robotic device, a sensor disposed on the robotic device, and circuitry configured to perform an operation. The operation may include determining a map representing a plurality of static features of an environment, and receiving sensor data representing the environment from a sensor. The operation may also include determining a representation of an actor in the environment based on the sensor data. The representation may include a plurality of key points representing corresponding body positions of the actor. The operation may additionally include determining that a portion of a particular static feature among the plurality of static features of the environment is within a threshold distance of a particular key point among the plurality of key points. The operation may also include updating a map to indicate that the portion of the particular static feature is to be cleaned based on determining that the particular key point is within a threshold distance of the portion of the particular static feature. The operation may also include causing the robotic device to clean the portion of the particular static feature based on the updated map.

[0007] In a second example embodiment, a computer-implemented method may include determining a map representing a plurality of static features of an environment and receiving sensor data representing the environment from a sensor disposed on a robotic device. The computer-implemented method may further include determining a representation of an actor in the environment based on the sensor data. The representation may include a plurality of key points representing corresponding body positions of the actor. The computer-implemented method may additionally include determining that a portion of a particular static feature among the plurality of static features of the environment is within a threshold distance of a particular key point among the plurality of key points. The computer-implemented method may further include updating the map to indicate that the portion of the particular static feature is to be cleaned based on determining that the particular key point is within the threshold distance of the portion of the particular static feature. The computer-implemented method may further include causing the robotic device to clean the portion of the particular static feature based on the updated map.

[0008] In a third example embodiment, a non-transitory computer-readable medium may have instructions stored thereon that, when executed by a computing device, cause the computing device to perform operations. The operations may include determining a map representing a plurality of static features of an environment, and receiving sensor data representing the environment from sensors disposed on a robotic device. The operations may further include determining a representation of an actor in the environment based on the sensor data. The representation may include a plurality of key points representing corresponding body positions of the actor. The operations may additionally include determining that a portion of a particular static feature among the plurality of static features of the environment is within a threshold distance of a particular key point among the plurality of key points. The operations may further include updating a map to indicate that the portion of the particular static feature is to be cleaned based on determining that the particular key point is within a threshold distance of the portion of the particular static feature. The operations may further include causing the robotic device to clean the portion of the particular static feature based on the updated map.

[0009] In a fourth example embodiment, a system may include components for determining a map representing a plurality of static features of an environment, and components for receiving sensor data representing the environment from sensors disposed on a robotic device. The system may also include components for determining a representation of an actor in the environment based on the sensor data. The representation may include a plurality of key points representing corresponding body positions of the actor. The system may additionally include components for determining that a portion of a particular static feature among the plurality of static features of the environment is within a threshold distance of a particular key point among the plurality of key points. The system may also include components for updating the map to indicate that the portion of the particular static feature is to be cleaned based on determining that the particular key point is within the threshold distance of the portion of the particular static feature. The system may also include components for causing the robotic device to clean the portion of the particular static feature based on the updated map.

[0010] In a fifth example embodiment, a system may include a sensor and circuitry configured to perform an operation. The operation may include determining a map representing multiple static features of an environment, and receiving sensor data representing the environment from a sensor. The operation may also include determining a representation of an actor in the environment based on the sensor data. The representation may include multiple key points representing corresponding body positions of the actor. The operation may additionally include determining that a portion of a specific static feature among the multiple static features of the environment is within a threshold distance of a specific key point among the multiple key points. The operation may further include updating the map to indicate that the portion of the specific static feature is determined to have been touched based on determining that the specific key point is within a threshold distance of the portion of the specific static feature. The operation may also include displaying a visual representation indicating that the portion of the specific static feature is determined to have been touched based on the updated map.

[0011] In a sixth example embodiment, a computer-implemented method may include determining a map representing multiple static features of an environment, and receiving sensor data representing the environment from a sensor. The computer-implemented method may also include determining a representation of an actor in the environment based on the sensor data. The representation may include multiple key points representing corresponding body positions of the actor. The computer-implemented method may additionally include determining that a portion of a specific static feature among the multiple static features of the environment is within a threshold distance of a specific key point among the multiple key points. The computer-implemented method may also include updating the map to indicate that the portion of the specific static feature is determined to have been touched based on determining that the specific key point is within the threshold distance of the portion of the specific static feature. The computer-implemented method may also include displaying a visual representation indicating that the portion of the specific static feature is determined to have been touched based on the updated map.

[0012] In a seventh example embodiment, a non-transitory computer-readable medium may store instructions that, when executed by a computing device, cause the computing device to perform operations. The operations may include determining a map representing multiple static features of an environment and receiving sensor data representing the environment from a sensor. The operations may also include determining a representation of an actor in the environment based on the sensor data. The representation may include multiple key points representing corresponding body positions of the actor. The operations may additionally include determining that a portion of a specific static feature among the multiple static features of the environment is within a threshold distance of a specific key point among the multiple key points. The operations may further include updating the map to indicate that the portion of the specific static feature is determined to have been touched based on determining that the specific key point is within a threshold distance of the portion of the specific static feature. The operations may also include displaying a visual representation indicating that the portion of the specific static feature is determined to have been touched based on the updated map.

[0013] In an eighth example embodiment, a system may include components for determining a map representing multiple static features of an environment, and components for receiving sensor data representing the environment from a sensor. The system may also include components for determining a representation of an actor in the environment based on the sensor data. The representation may include multiple key points representing corresponding body positions of the actor. The system may additionally include components for determining that a portion of a specific static feature among the multiple static features of the environment is within a threshold distance of a specific key point among the multiple key points. The system may also include components for updating the map to indicate that the portion of the specific static feature is determined to have been touched based on determining that the specific key point is within the threshold distance of the portion of the specific static feature. The system may also include components for displaying a visual representation indicating that the portion of the specific static feature is determined to have been touched based on the updated map.

[0014] These and other embodiments, aspects, advantages, and alternatives will become apparent to those skilled in the art upon reading the following detailed description and, where appropriate, referring to the accompanying drawings. Furthermore, this summary and the other descriptions and drawings provided herein are intended to illustrate the embodiments by way of example only, and as such, many variations are possible. For example, structural elements and process steps may be rearranged, combined, distributed, eliminated, or otherwise modified while remaining within the scope of the claimed embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A configuration of a robot system according to an example embodiment is illustrated.

[0016] Figure 2 A mobile robot according to an example embodiment is illustrated.

[0017] Figure 3 An exploded view of a mobile robot is illustrated according to an example embodiment.

[0018] Figure 4 A robotic arm is illustrated according to an example embodiment.

[0019] Figure 5A Actors interacting with environmental features according to example embodiments are illustrated.

[0020] Figure 5B Illustrated are key points on actors according to example embodiments.

[0021] Figure 5C Illustrated is a threshold distance around a key point according to an example embodiment.

[0022] Figure 5D Illustrated is a portion of an environment designated for cleaning according to an example embodiment.

[0023] Figure 6 A system according to an example embodiment is illustrated.

[0024] Figure 7 A flow chart according to an example embodiment is illustrated.

[0025] Figure 8 A flow chart according to an example embodiment is illustrated. DETAILED DESCRIPTION

[0026] Example methods, devices, and systems are described herein. It should be understood that the words "example" and "exemplary" as used herein mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "example," "exemplary," and / or "illustrative" is not necessarily to be construed as preferred or advantageous over other embodiments or features unless so stated. Therefore, other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein.

[0027] Therefore, the example embodiments described herein are not meant to be limiting.It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the accompanying drawings, may be arranged, substituted, combined, separated, and designed in a variety of different configurations.

[0028] Furthermore, unless the context suggests otherwise, the features illustrated in each figure may be used in combination with each other. Thus, the drawings should generally be considered as forming aspects of one or more overall embodiments, with it being understood that not all illustrated features are essential to every embodiment.

[0029] In addition, any enumeration of elements, blocks, or steps in this specification or claims is for clarity purposes. Therefore, such enumeration should not be interpreted as requiring or implying that these elements, blocks, or steps follow a specific arrangement or be performed in a specific order. Unless otherwise indicated, the drawings are not drawn to scale.

[0030] I. Overview

[0031] The robotic device can be configured to detect and / or clean surfaces in the environment that have been touched by an actor. Specifically, the robotic device can be mobile and therefore configured to traverse the environment. The robotic device can include one or more sensors thereon that are configured to generate sensor data representing static features (such as the ground, walls and / or furniture) and (non-static) actors (such as people, animals and / or other robotic devices). The one or more sensors can include, for example, a red, green, and blue (RGB) two-dimensional (2D) camera, a depth (three-dimensional (3D)) camera, a light detection and ranging (LIDAR) device, a radio detection and ranging (radar) device, a time of flight (ToF) camera, a thermal camera, and / or a near-infrared camera, among other possibilities. Sensor data can be collected as the robotic device performs various tasks in the environment, and / or as part of a dedicated data collection operation.

[0032] The control circuitry (e.g., forming part of a control system) can be configured to receive sensor data, process the sensor data, and issue commands to the robotic device to perform various operations within the environment. The control circuitry can be physically located on the robotic device and / or remote from the robotic device. Specifically, the control circuitry can be configured to generate a map representing at least static features of the environment. The map can be generated based on sensor data obtained from the robotic device, from one or more other robotic devices operating in the environment, and / or from one or more sensors distributed throughout the environment. A subset of these static features can be classified as normal touch / high touch features that are determined to interact with an actor at a relatively high frequency, and can therefore be monitored by the robotic device to determine whether these features have been touched.

[0033] In order to determine whether a static feature has been touched by an actor, the robotic device may receive sensor data representing an actor in the environment from one or more sensors. The control circuitry may be configured to determine a representation of the actor based on the sensor data. Specifically, the representation may include a plurality of key points representing the positions of the actor's body. For example, the key points may represent anatomical locations of the actor's body, such as joints, and may therefore represent ankles, knees, hips, shoulders, elbows, wrists, necks, and / or heads. In some cases, the key points may be interconnected to define a skeleton representing the actor's body. For example, the key points may be determined by a machine learning model such as a convolutional neural network (CNN).

[0034] The control circuit can also be configured to determine, for a particular key point among the plurality of key points, whether any portion of the static feature is within a threshold distance of the particular key point. To this end, the control circuit can be configured to transform the key point from the reference frame of the sensor data to the reference frame of the map. The particular key point may represent a subject position that is frequently used to interact with the environment, such as an arm, hand, and / or finger. In some cases, the control circuit can determine that a portion of a particular static feature, such as a desktop, is within a threshold distance of the particular key point. Based on this determination, the control circuit can be configured to update the map to indicate that the portion of the particular static feature is to be cleaned.

[0035] Specifically, the map can be updated to indicate that the portion of the particular static feature is determined / estimated to have been touched. Determining / estimating that the portion has been touched can include situations where the portion has actually been touched, as well as situations where the portion has not actually been touched but appears to have been touched based on the location of the particular key point. For example, a subset of the map representing the portion of the desktop that falls within a threshold distance of a particular key point can be marked as "touched" by updating one or more values (e.g., a touched / not touched flag) in a data structure representing the map.

[0036] In some cases, a map may represent an environment as a plurality of voxels. Each voxel may be marked as occupied by a static feature or unoccupied (e.g., representing an empty space). Thus, the control circuitry may be configured to identify occupied voxels that fall within a threshold distance of a particular keypoint and mark these voxels as to be cleaned (e.g., marking these voxels as touched). In other cases, a map may represent environmental features as a polygonal mesh comprising vertices, edges, and faces. Thus, the control circuitry may be configured to identify faces, vertices, and / or edges represented by the mesh that fall within a threshold distance of a particular keypoint and mark these faces, vertices, and / or edges for cleaning. Other arrangements of maps are also possible.

[0037] The sensor data may include multiple frames of sensor data representing the actor at different points in time. Thus, as the actor interacts with various features of the environment, interactions that are determined to involve the actor touching features of the environment may be detected by the control circuitry and indicated by updating the map. Thus, the map may store representations of portions of the environment that are determined to have been touched by the actor during a particular time period. When the actor leaves the relevant portion of the environment, the robotic device may be used to clean (e.g., decontaminate, disinfect, sterilize, and / or otherwise clean) the portion of the environment that is determined to have been touched. Such cleaning may include cleaning portions that have actually been touched, as well as portions that appear to have been touched but have not actually been touched. Once the robot has cleaned these portions, the map may be updated again to indicate that these portions have been cleaned and are therefore "untouched."

[0038] In some implementations, in addition to determining whether a particular static feature is within a threshold distance of a particular key point, the control circuitry may be configured to detect and / or verify a touch on a particular static feature based on a posture of the actor indicated by a plurality of key points and / or a motion pattern of the actor indicated by changes in the positioning of the plurality of key points over time. A machine learning model such as a feed-forward artificial neural network (FF-ANN) may be configured to determine that a particular arrangement of a plurality of key points indicates that the actor is engaging and / or making physical contact with the particular static feature. For example, an arrangement of a plurality of key points indicating that the actor is facing a table may indicate engagement with the table, while another arrangement of a plurality of key points indicating that the actor is walking along the edge of the table may indicate a lack of engagement and / or physical contact with the table. Another machine learning model such as a recurrent neural network (RNN) or a long short-term memory (LSTM) neural network may be configured to make similar determinations of engagement and / or physical contact with the table based on multiple sets of key points corresponding to different points in time.

[0039] In some cases, the sensor data used to identify touched features may additionally include sensor data collected by one or more other robots in the environment and / or sensors installed throughout the environment (e.g., security cameras). By including sensor data from multiple robots and / or installed sensors in the environment, more touch points can be detected and / or the detection of touch points can be more accurate due to the additional perspectives on the actor. For example, key points can be more accurately determined due to the more complete representation of the actor by the sensor data.

[0040] In some cases, due to the positioning of the robotic device relative to the actor when capturing sensor data, a particular key point on the actor may be obscured. For example, the robotic device may be located behind the actor and, therefore, may not be able to see the actor's hand located in front of the actor. Therefore, the control circuitry can be configured to estimate the location of one or more obscured key points based on the locations of other visible key points. Because the estimated location of a key point corresponding to an obscured subject position may not be as accurate as the estimated location of a key point when the subject position is unobstructed, the threshold distance for detecting a touch surface can be increased. For example, the threshold distance can range from a minimum threshold distance associated with a high confidence estimated location of a particular key point (e.g., when fully visible) to a maximum threshold distance associated with a low confidence estimated location of a particular key point (e.g., when fully obscured). Additionally or alternatively, when a particular key point is obscured, one or more other key points can be used to detect touched features. For example, when a hand is obscured, instead of using key points on an arm (e.g., an elbow), whether a particular feature has been touched can be determined based on whether these key points come within a different (e.g., greater) threshold distance of the particular feature.

[0041] The above-mentioned method for identifying and cleaning touched surfaces provides several advantages. Specifically, the robotic device may be able to accurately identify and clean parts of the environment that have actually and / or are likely to be touched, rather than indiscriminately cleaning all parts of the environment that may have been potentially touched. This enables the robotic device to reduce and / or minimize the use of cleaning products and reduce and / or minimize the total cleaning time. In addition, the robotic device may be able to reduce and / or minimize the time between the surface being touched and the surface being cleaned. That is, instead of performing general cleaning periodically (e.g., once a day, once a week, etc.), the robotic device is able to clean the surface when the actor who has touched the surface moves far enough away from the surface to allow the robot to operate safely in the relevant part of the environment (and before another actor touches the surface). In addition, since the robotic device may not need to comply with social distance rules, the robotic device can perform cleaning faster than humans and has fewer spatial restrictions.

[0042] II. Example Robotic System

[0043] Figure 1The diagram illustrates example configurations of robotic systems that can be used in conjunction with the implementations described herein. The robotic system 100 can be configured to operate autonomously, semi-autonomously, or using instructions provided by (multiple) users. The robotic system 100 can be implemented in various forms, such as a robotic arm, an industrial robot, or some other arrangement. Some example implementations involve a robotic system 100 that is engineered to be low-cost at scale and designed to support a variety of tasks. The robotic system 100 can be designed to be able to operate around people. The robotic system 100 can also be optimized for machine learning. Throughout the specification, the robotic system 100 may also be referred to as a robot, a robotic device, and / or a mobile robot, among other names.

[0044] like Figure 1 As shown, the robotic system 100 may include processor(s) 102, data storage device(s) 104, and controller(s) 108, which together may be part of a control system 118. The robotic system 100 may also include sensor(s) 112, power supply(s) 114, mechanical components 110, and electrical components 116. Nevertheless, the robotic system 100 is illustrated for illustrative purposes and may include more or fewer components. The various components of the robotic system 100 may be connected in any manner, including through wired or wireless connections. Furthermore, in some examples, the components of the robotic system 100 may be distributed across multiple physical entities rather than a single physical entity. Other example illustrations of the robotic system 100 may also exist.

[0045] The processor(s) 102 may operate as one or more general-purpose hardware processors or specialized hardware processors (e.g., digital signal processors, application-specific integrated circuits, etc.). The processor(s) 102 may be configured to execute computer-readable program instructions 106 and manipulate data 107, both of which are stored in the data storage device 104. The processor(s) 102 may also interact, directly or indirectly, with other components of the robotic system 100, such as the sensor(s) 112, the power supply(s) 114, the mechanical components 110, or the electrical components 116.

[0046] The data storage device 104 can be one or more types of hardware memory. For example, the data storage device 104 can include or take the form of one or more computer-readable storage media that can be read or accessed by the (multiple) processors 102. The one or more computer-readable storage media can include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or another type of memory or storage device, which can be integrated with the (multiple) processors 102 in whole or in part. In some implementations, the data storage device 104 can be a single physical device. In other implementations, the data storage device 104 can be implemented using two or more physical devices that can communicate with each other via wired or wireless communications. As previously described, the data storage device 104 can include computer-readable program instructions 106 and data 107. The data 107 can be any type of data, such as configuration data, sensor data, or diagnostic data, among other possibilities.

[0047] The controller 108 may include one or more electrical circuits, digital logic units, computer chips, or microprocessors configured to, among other tasks, interface between any combination of the mechanical components 110, the sensor(s) 112, the power supply(s) 114, the electrical components 116, the control system 118, or a user of the robotic system 100. In some implementations, the controller 108 may be a purpose-built embedded device configured to perform specific operations with one or more subsystems of the robotic system 100.

[0048] The control system 118 can monitor and physically change the operating conditions of the robotic system 100. Thus, the control system 118 can serve as a link between various parts of the robotic system 100, such as between the mechanical components 110 or the electrical components 116. In some cases, the control system 118 can serve as an interface between the robotic system 100 and another computing device. Furthermore, the control system 118 can serve as an interface between the robotic system 100 and a user. In some cases, the control system 118 can include various components for communicating with the robotic system 100, including joysticks, buttons, or ports. The example interfaces and communications mentioned above can be implemented via wired or wireless connections, or both. The control system 118 can also perform other operations for the robotic system 100.

[0049] During operation, the control system 118 can communicate with other systems of the robotic system 100 via wired and / or wireless connections, and can also be configured to communicate with one or more users of the robot. As one possible example, the control system 118 can receive input (e.g., from a user or from another robot) indicating an instruction to perform a requested task, such as picking up an object from one location and moving it to another location. Based on this input, the control system 118 can perform operations to cause the robotic system 100 to perform a series of movements to perform the requested task. As another example, the control system can receive input indicating an instruction to move to a requested location. In response, the control system 118 (possibly with the help of other components or systems) can determine the direction and speed at which the mobile robotic system 100 should move through the environment to reach the requested location.

[0050] The operations of the control system 118 may be performed by the processor(s) 102. Alternatively, the operations may be performed by the controller(s) 108, or a combination of the processor(s) 102 and the controller(s) 108. In some implementations, the control system 118 may be partially or fully located on a device external to the robotic system 100, and thus the robotic system 100 may be at least partially remotely controlled.

[0051] The mechanical assembly 110 represents the hardware of the robotic system 100 that enables the robotic system 100 to perform physical operations. As a few examples, the robotic system 100 may include one or more physical components, such as an arm, an end effector, a head, a neck, a torso, a base, and wheels. The physical components or other parts of the robotic system 100 may also include actuators arranged to move the physical components relative to each other. The robotic system 100 may also include one or more structural bodies for housing a control system 118 or other components, and may also include other types of mechanical assemblies. The specific mechanical assembly 110 used in a given robot may vary based on the design of the robot and may also vary based on the operations or tasks that the robot may be configured to perform.

[0052] In some examples, the mechanical assembly 110 may include one or more removable components. The robotic system 100 may be configured to add or remove such removable components, which may involve assistance from a user or another robot. For example, the robotic system 100 may be configured with a removable end effector or digit that can be replaced or changed as needed or desired. In some implementations, the robotic system 100 may include one or more removable or replaceable battery cells, control systems, power systems, buffers, or sensors. In some implementations, other types of removable components may be included.

[0053] The robotic system 100 may include sensor(s) 112 arranged to sense various aspects of the robotic system 100. The sensors 112 may include one or more force sensors, torque sensors, velocity sensors, acceleration sensors, positioning sensors, proximity sensors, motion sensors, position sensors, load sensors, temperature sensors, touch sensors, depth sensors, ultrasonic range sensors, infrared sensors, object sensors, or cameras, among other possibilities. In some examples, the robotic system 100 may be configured to receive sensor data from sensors that are physically separate from the robot (e.g., sensors located on other robots or within the environment in which the robot operates).

[0054] The sensor(s) 112 may provide sensor data to the processor(s) 102 (possibly via the data 107) to allow the robotic system 100 to interact with its environment and to monitor the operation of the robotic system 100. The sensor data may be used to control the control system 118's evaluation of various factors regarding the activation, movement, and deactivation of the mechanical and electrical components 110, 116. For example, the sensor(s) 112 may capture data corresponding to the topography of the environment or the location of nearby objects, which may aid in environmental recognition and navigation.

[0055] In some examples, the sensor(s) 112 may include RADAR (e.g., for long-range object detection, distance determination, or velocity determination), LIDAR (e.g., for short-range object detection, distance determination, or velocity determination), SONAR (e.g., for underwater object detection, distance determination, or velocity determination), The robotic system 100 may be configured to include a plurality of sensors 112 (e.g., for motion capture), one or more cameras (e.g., stereo cameras for 3D vision), a global positioning system (GPS) transceiver, or other sensors for capturing information about the environment in which the robotic system 100 is operating. The sensor(s) 112 may monitor the environment in real time and detect obstacles, terrain features, weather conditions, temperature, or other aspects of the environment. In another example, the sensor(s) 112 may capture data corresponding to one or more features of a target or identified object, such as the object's size, shape, outline, structure, or orientation.

[0056] Furthermore, the robotic system 100 may include sensor(s) 112 configured to receive information indicative of a state of the robotic system 100, including sensor(s) 112 that may monitor the state of various components of the robotic system 100. The sensor(s) 112 may measure system activity of the robotic system 100 and receive information based on the operation of various features of the robotic system 100, such as the operation of an extendable arm, an end effector, or other mechanical or electrical features of the robotic system 100. The data provided by the sensor(s) 112 may enable the control system 118 to determine errors in operation and to monitor the overall operation of the components of the robotic system 100.

[0057] As an example, the robotic system 100 can use force / torque sensors to measure the loads on various components of the robotic system 100. In some implementations, the robotic system 100 can include one or more force / torque sensors on an arm or end effector to measure the loads on actuators that move one or more components of the arm or end effector. In some examples, the robotic system 100 can include force / torque sensors at or near the wrist or end effector, but not at or near other joints of the robotic arm. In a further example, the robotic system 100 can use one or more positioning sensors to sense the positioning of actuators of the robotic system. For example, such positioning sensors can sense the extension, retraction, positioning, or rotational state of actuators on the arm or end effector.

[0058] As another example, the sensor(s) 112 may include one or more velocity or acceleration sensors. For example, the sensor(s) 112 may include an inertial measurement unit (IMU). The IMU may sense velocity and acceleration relative to a gravity vector in a world coordinate system. The velocity and acceleration sensed by the IMU may then be converted to velocity and acceleration of the robotic system 100 based on the position of the IMU in the robotic system 100 and the kinematics of the robotic system 100.

[0059] The robotic system 100 may include other types of sensors not explicitly discussed herein. Additionally or alternatively, the robotic system may use specific sensors for purposes not listed herein.

[0060] The robotic system 100 may also include one or more power sources 114 configured to provide power to various components of the robotic system 100. Among other possible power systems, the robotic system 100 may include a hydraulic system, an electrical system, a battery, or other types of power systems. As an example, the robotic system 100 may include one or more batteries configured to provide an electrical charge to the components of the robotic system 100. Some of the mechanical components 110 or the electrical components 116 may each be connected to a different power source, may be powered by the same power source, or may be powered by multiple power sources.

[0061] Any type of power source can be used to power the robotic system 100, such as electricity or a gasoline engine. Additionally or alternatively, the robotic system 100 can include a hydraulic system configured to provide power to the mechanical assembly 110 using fluid power. For example, the components of the robotic system 100 can operate based on hydraulic fluid transmitted to various hydraulic motors and hydraulic cylinders through the hydraulic system. The hydraulic system can transmit hydraulic power in the form of pressurized hydraulic fluid through pipes, flexible hoses, or other links between the components of the robotic system 100. The power source(s) 114 can be charged using various types of charging, such as a wired connection to an external power source, wireless charging, combustion, or other examples.

[0062] The electrical components 116 may include various mechanisms capable of processing, transmitting, or providing electrical charge or electrical signals. In possible examples, the electrical components 116 may include wiring, circuitry, or wireless communication transmitters and receivers to enable the operation of the robotic system 100. The electrical components 116 may interoperate with the mechanical components 110 to enable the robotic system 100 to perform various operations. For example, the electrical components 116 may be configured to provide power from the power source(s) 114 to the various mechanical components 110. Furthermore, the robotic system 100 may include an electric motor. Other examples of the electrical components 116 are also possible.

[0063] The robotic system 100 may include a main body that can be connected to or house accessories and components of the robotic system. As such, the structure of the main body may vary in examples and may further depend on the specific operation that a given robot may have been designed to perform. For example, a robot developed to carry heavy loads may have a wide main body so that the load can be placed. Similarly, a robot designed to operate in confined spaces may have a relatively tall, narrow main body. In addition, the main body or other components may be developed using various types of materials, such as metal or plastic. In other examples, the robot may have a main body with different structures or made of various types of materials.

[0064] The body or other components may include or carry sensor(s) 112. These sensors may be located at various locations on the robotic system 100, such as the body, head, neck, base, torso, arm, or end effector.

[0065] The robotic system 100 can be configured to carry a load, such as a type of cargo to be transported. In some examples, the robotic system 100 can place the load in a bin or other container attached to the robotic system 100. The load can also represent an external battery or other type of power source (e.g., a solar panel) that the robotic system 100 can utilize. Carrying a load represents one example use for which the robotic system 100 can be configured, but the robotic system 100 can also be configured to perform other operations.

[0066] As described above, the robotic system 100 can include various types of attachments, wheels, end effectors, gripping devices, and the like. In some examples, the robotic system 100 can include a mobile base having wheels, pedals, or some other form of mobility capability. Additionally, the robotic system 100 can include a robotic arm or some other form of robotic manipulator. In the case of a mobile base, the base can be considered one of the mechanical assemblies 110 and can include wheels driven by one or more actuators that allow movement of the robotic arm in addition to the rest of the body.

[0067] Figure 2 A mobile robot according to an example embodiment is illustrated. Figure 3 2 shows an exploded view of a mobile robot according to an exemplary embodiment. More specifically, robot 200 may include a mobile base 202, a mid-section 204, an arm 206, an end-of-arm system (EOAS) 208, a mast 210, a perception housing 212, and a perception kit 214. Robot 200 may also include a computing box 216 stored within mobile base 202.

[0068] The mobile base 202 includes two drive wheels at the front end of the robot 200 to provide mobility for the robot 200. The mobile base 202 also includes additional casters (not shown) to facilitate movement of the mobile base 202 on the ground. The mobile base 202 can have a modular architecture that allows the computing box 216 to be easily removed. The computing box 216 can be used as a removable control system for the robot 200 (rather than a mechanically integrated control system). After removing the external housing, the computing box 216 can be easily removed and / or replaced. The mobile base 202 can also be designed to allow additional modularity. For example, the mobile base 202 can also be designed so that the power system, batteries and / or external buffers can be easily removed and / or replaced.

[0069] The middle section 204 can be attached to the mobile base 202 at the front end of the mobile base 202. The middle section 204 includes a mounting post fixed to the mobile base 202. The middle section 204 also includes a rotational joint for the arm 206. More specifically, the middle section 204 includes the first two degrees of freedom of the arm 206 (the shoulder yaw J0 joint and the shoulder pitch J1 joint). The mounting post and the shoulder yaw J0 joint can form part of a stacking tower at the front of the mobile base 202. The mounting post and the shoulder yaw J0 joint can be coaxial. The length of the mounting post of the middle section 204 can be selected to provide the arm 206 with sufficient height to perform manipulation tasks at commonly encountered height levels (e.g., coffee table top and / or counter top level). The length of the mounting post of the middle section 204 can also allow the shoulder pitch J1 joint to rotate the arm 206 on the mobile base 202 without contacting the mobile base 202.

[0070] When connected to the middle section 204, the arm 206 can be a 7 degree of freedom (DOF) robotic arm. As described above, the first two degrees of freedom of the arm 206 can be included in the middle section 204. Figure 2 and 3 As shown, the remaining five degrees of freedom can be included in a separate part of the arm 206. The arm 206 can be made of a plastic integral link structure. The arm 206 can accommodate independent actuator modules, local motor drivers and through-hole cables.

[0071] EOAS 208 may be an end effector at the end of arm 206. EOAS 208 may allow robot 200 to manipulate objects in the environment. Figure 2 and 3 As shown, the EOAS 208 can be a gripper, such as an underactuated pinch gripper. The gripper can include one or more contact sensors, such as force / torque sensors, and / or non-contact sensors, such as one or more cameras, to facilitate object detection and gripper control. The EOAS 208 can also be a different type of gripper, such as a suction gripper, or a different type of tool, such as a drill or a brush. The EOAS 208 can also be interchangeable or include interchangeable components, such as gripper fingers.

[0072] The mast 210 can be a relatively long, narrow component between the shoulder pitch (J0) joint of the arm 206 and the sensing housing 212. The mast 210 can be part of a stacked tower in front of the mobile base 202. The mast 210 can be fixed relative to the mobile base 202. The mast 210 can be coaxial with the central portion 204. The length of the mast 210 can facilitate the sensing kit 214 sensing objects manipulated by the EOAS 208. The mast 210 can be of a length such that when the shoulder pitch (J1) joint rotates vertically upward, the highest point of the biceps of the arm 206 is approximately aligned with the top of the mast 210. The length of the mast 210 can be sufficient to prevent collision between the sensing housing 212 and the arm 206 when the shoulder pitch (J1) joint rotates vertically upward.

[0073] like Figure 2 and 3 As shown, mast 210 can include a 3D lidar sensor configured to collect depth information about the environment. The 3D lidar sensor can be coupled to a cutout portion of mast 210 and fixed at a downward angle. The lidar position can be optimized for positioning, navigation, and forward cliff detection.

[0074] The perception housing 212 may include at least one sensor that makes up the perception suite 214. The perception housing 212 may be connected to a pan / tilt controller to allow reorientation of the perception housing 212 (e.g., to view an object being manipulated by the EOAS 208). The perception housing 212 may be part of a stacked tower secured to the mobile base 202. The rear portion of the perception housing 212 may be coaxial with the mast 210.

[0075] The perception suite 214 may include a sensor suite configured to collect sensor data representing the environment of the robot 200. The perception suite 214 may include an infrared (IR) assisted stereo depth sensor. The perception suite 214 may additionally include a wide-angle red, green, and blue (RGB) camera for human-machine interaction and contextual information. The perception suite 214 may additionally include a high-resolution RGB camera for object classification. A facial halo surrounding the perception suite 214 may also be included to improve human-machine interaction and scene lighting. In some examples, the perception suite 214 may also include a projector configured to project images and / or video into the environment.

[0076] Figure 4A robotic arm according to an example embodiment is illustrated. The robotic arm includes seven degrees of freedom: shoulder yaw (J0), shoulder pitch (J1), biceps roll (J2), elbow pitch (J3), forearm roll (J4), wrist pitch (J5), and wrist roll (J6). Each joint can be coupled to one or more actuators. The actuators coupled to the joints can be operable to move a link along a kinematic chain (and any end effector attached to the robotic arm).

[0077] The shoulder deflection J0 joint allows the robot arm to rotate toward the front and back of the robot. One beneficial use of this motion is to allow the robot to pick up an object in front of the robot and quickly place it behind the robot (and vice versa). Another beneficial use of this motion is to quickly move the robot arm from a stowed configuration behind the robot to an active position in front of the robot (and vice versa).

[0078] The shoulder pitch J1 joint allows the robot to raise the robot arm (e.g., so that the biceps reach the level of the robot's sensor kit) and lower the robot arm (e.g., so that the biceps are just above the mobile base). This movement is beneficial in allowing the robot to effectively perform manipulation operations (e.g., top grasping and side grasping) at different target height levels in the environment. For example, the shoulder pitch J1 joint can be rotated to a vertical upward position to allow the robot to easily manipulate objects on a table in the environment. The shoulder pitch J1 joint can be rotated to a vertical downward position to allow the robot to easily manipulate objects on the ground in the environment.

[0079] The biceps flip J2 joint allows the robot to rotate the biceps to move the elbow and forearm relative to the biceps. This motion is particularly beneficial for providing the robot's perception suite with a clear view of the EOAS. By rotating the biceps flip J2 joint, the robot can extend the elbow and forearm to improve its view of objects held in the robot's gripper.

[0080] Moving down the kinematic chain, alternating pitch and roll joints (shoulder pitch J1, biceps roll J2, elbow pitch J3, forearm roll J4, wrist pitch J5, and wrist roll J6) are provided to improve maneuverability of the robot arm. The axes of the wrist pitch J5, wrist roll J6, and forearm roll J4 intersect to reduce arm motion required to reorient the object. The wrist roll J6 joint is provided in place of two pitch joints in the wrist to improve rotation of the object.

[0081] III. Example Touch Part Detection Process

[0082] Figure 5A 、 5B5C and 5D illustrate example processes for detecting a touched portion of an environmental feature, tracking the touched portion, and cleaning the touched portion. Specifically, Figure 5A The robot 200 is illustrated as obtaining sensor data via one or more sensors in the perception suite 214, the sensor data representing the area of the environment occupied by the actor 500, the table 502, and the cup 504, as indicated by field of view 524. The right hand of the actor 500 is shown moving from pose 526 (shown by a dashed line) to pose 528, for example, to reach out and grasp the cup 504.

[0083] When reaching out and / or grasping the cup 504, one or more body parts of the actor 500 may come into physical contact with an area of the table 502. Alternatively or additionally, the actor 500 may interact with the table 502 for reasons other than picking up an object from the table 502, such as by leaning on the table 502 for support, and may thus come into physical contact with the table 502. In some cases, such contact between the actor 500 and the table 502 may cause bacteria, viruses, sweat, and / or other unwanted matter that may have been transferred from a part of the actor 500 (e.g., the right hand) to stain, soil, and / or otherwise contaminate the table 502. Rather than having the robot 200 indiscriminately clean the entire table 502 and / or other features of the environment, the robot 200 may instead be configured to identify and clean specific parts of the table 502 that have been touched and / or appear to have been touched by the actor 500.

[0084] To this end, the control system of the robot 200 (e.g., the control system 118 or a similar component) can be configured to process the sensor data (e.g., represented by the field of view 524) to identify therein a plurality of key points representing corresponding body positions of the body of the actor 500. When the actor 500 is a human, the plurality of key points may represent predetermined locations on the human body, such as Figure 5B As shown. These predetermined subject positions can include a head keypoint 506, a neck keypoint 508, shoulder keypoints 510A and 510B, elbow keypoints 512A and 512B, hand keypoints 514A and 514B, a pelvis keypoint 516, a hip keypoint 518A and 518B, a knee keypoint 520A and 520B, and a foot keypoint 522A and 522B (i.e., keypoints 506-522B). Thus, at least a subset of keypoints 506-522B can include joints of the human body. Keypoints 506-522B can provide an accurate representation of the pose of actor 500 and can be generated and processed with less computational effort than some other representations, such as a polygonal mesh of actor 500.

[0085] In some implementations, some of the key points 506-522B may be omitted and / or other key points may be added. For example, the pelvis key point 516 may be omitted. In another example, each of the hand key points 514A and 514B may be further subdivided into five or more key points representing the fingers of each hand, among other possibilities. The key points 506-522B are shown interconnected to form a virtual human skeleton. Additionally, the key points 510B, 512B, and 514B are drawn with a different pattern than the other key points (and are interconnected with dashed lines) to indicate that when the robot 200 moves from Figure 5A and 5B When viewing actor 500 from the perspective shown, key points 510B, 512B, and 514B are not visible (ie, are hidden).

[0086] Alternatively, in some implementations, the plurality of key points may represent predetermined locations on the robot body, or predetermined locations on the body of an actor of a species other than humans (e.g., a dog or cat) that is also capable of interacting with features in the environment. Thus, the number and positioning of the key points may vary depending on the type of actor. Furthermore, the key points 506-522B may alternatively be referred to as nodes. Generally speaking, the plurality of key points may be predetermined body locations that the control system is configured to attempt to find or identify in the captured sensor data.

[0087] The control system can be configured to determine that a portion of a particular feature has been touched based on the relative positioning between (i) one or more of the key points 506-522B and (ii) a portion of the particular feature. Specifically, the control system can maintain a map of the environment in which at least static features of the environment (such as table 502) are represented. Static features can include landscape features, building fixtures, furniture, ground / floors, and other physical features. Static features can be fixed and / or repositionable, but can generally be expected to maintain a constant position in the environment for a predetermined period of time. In some implementations, static features can include any physical feature of the environment in addition to any actors present in the environment.

[0088] The map of the environment can be generated and / or updated based on sensor data collected by the robot 200 as it moves through the environment. For example, in addition to being used to detect the actor 500, the sensor data represented by the field of view 524 can be used to generate and / or update the map. Additionally or alternatively, the map can be generated and / or updated based on additional sensor data generated by other sensors. For example, the other sensors can be located on other robots operating in the environment and / or can be distributed at predetermined locations throughout the environment. Thus, in some cases, the map can be shared by multiple robots operating in the environment. In another example, the robot 200 can initially operate based on a predetermined map of the environment. As sensor data is generated over time by various sensors, the predetermined map can be updated based on the sensor data.

[0089] Based on the map and the plurality of key points, the control system can be configured to identify one or more portions of one or more features of the environment that have come within a threshold distance of one or more specific key points. Specifically, the respective locations of one or more specific key points among the key points 506-522B (e.g., hand key points 514A and 514B) can be transformed from the reference frame of the sensor (used to generate the sensor data based on which the key points 506-522B were identified) to the reference frame of the map of the environment representing the static features. Thus, the one or more specific key points can be represented as part of the map of the environment, and the distances between the one or more specific key points and the features of the environment can be directly compared.

[0090] Figure 5C The diagram illustrates the portion of determining whether an environmental feature is within a corresponding threshold distance of a particular key point and is therefore considered to have been touched. Specifically, Figure 5C A sphere 530 is shown centered on the right keypoint 514A. A radius 532 of the sphere 530 may define a threshold distance relative to the right keypoint 514A. Any portion of a static feature that falls within the volume of the sphere 530 may be considered to have been touched by the actor 500.

[0091] like Figure 5C As shown, a portion of table 502 intersects the volume of sphere 530. Therefore, Figure 5DThe diagram illustrates that portion 534 of workbench 502 has been designated for cleaning by robot 200. Portion 534 is circular in at least one dimension because it represents the geometric intersection of the tabletop of table 502 with the volume of sphere 530. Specifically, the map may be updated to indicate that portion 534 is to be cleaned by robot 200. For example, the values of one or more variables associated with portion 534 in the map (e.g., a clean / dirty flag, a touched / untouched flag, etc.) may be updated to indicate that portion 534 is to be cleaned. In some cases, one or more additional models and / or algorithms may be used to determine whether actor 500 intended to interact with table 502, or whether the intersection between sphere 530 and table 502 was the result of actor 500 passing near the table without engaging with table 502 and / or objects thereon. Such additional models and / or algorithms may be used to determine whether actor 500 intended to interact with table 502, or whether the intersection between sphere 530 and table 502 was the result of actor 500 passing near the table without engaging with table 502 and / or objects thereon. Figure 6 Figure and reference Figure 6 Discuss in more detail.

[0092] In one example, a map can represent an environment by a plurality of voxels. Each voxel can be marked as occupied or unoccupied by a static feature (e.g., representing an empty space). Each voxel can also be associated with a corresponding variable, the value of which can indicate the cleanliness of the corresponding part of the environment. In some cases, the variable can be a binary value and can therefore indicate whether the corresponding part of the environment is clean or dirty, and / or has been touched or not touched. In other cases, the variable can take more than two values and can therefore be configured to indicate various cleanliness states.

[0093] When at least a portion of an occupied voxel falls within a threshold distance of a particular keypoint (e.g., intersects with sphere 530), the occupied voxel can be marked as to be cleaned. When determining intersection with sphere 530, empty voxels can be ignored to reduce the amount of computation involved. In cases where the "cleanliness" variable can take on more than two values, the value of the variable can be based on the degree of intersection between the voxel and the sphere. For example, a voxel that is completely within the threshold distance can be marked as "very dirty," a voxel that is partially within the threshold distance can be marked as "moderately dirty," and a voxel that is completely outside the threshold distance can be marked as "clean."

[0094] In other examples, a map can represent the environment using a polygonal mesh comprising vertices, edges, and faces. Thus, for example, features of the environment can be represented as a combination of polygonal faces rather than voxels. Portions of the polygonal mesh can be marked for cleaning based on one or more vertices, edges, and / or faces being within a threshold distance. Other arrangements of the map are also possible.

[0095] In some cases, a map of an environment may associate each corresponding static feature represented by the map with an indication of the touch frequency of the corresponding feature. For example, the tabletop of table 502 may be classified as a high touch feature, as indicated by its shading pattern, while the legs of table 502 may be classified as a low touch feature, as indicated by its lack of a pattern. Other classifications and / or values, such as medium touch, are also possible. In some implementations, the touch frequency of a feature may be represented by associating each element of the map (e.g., voxel, polygonal face, etc.) with a corresponding touch frequency value. The classification of each static feature may be predefined (e.g., by a programmer), determined by a machine learning algorithm, and / or determined based on touch frequencies empirically observed by the robot 200.

[0096] Thus, in some cases, the control system may be configured to identify high-touch features that intersect sphere 530, but may not identify low-touch features that fall within sphere 530. In other cases, the control system may prioritize identifying and / or cleaning high-touch features over low-touch features. Thus, the amount of computation performed by the control system and the amount of cleaning performed by robot 200 may be reduced without significantly reducing the cleanliness of the environment. However, in further cases, high-touch features and low-touch features may be treated equally by the control system.

[0097] In some embodiments, the present invention provides the key point 506-522B of the invention, wherein the key point 506-522B is a key point ...

[0098] In some implementations, the value of the threshold distance associated with a particular keypoint can be based on a confidence value associated with the identification / determination of the particular keypoint. For example, a keypoint corresponding to a subject location that is explicitly represented (e.g., visible) in the sensor data and / or represented at a relatively high level of detail (e.g., resolution) can be associated with a relatively high confidence value. On the other hand, a keypoint corresponding to another subject location that is not represented in the sensor data and / or represented at a relatively low level of detail can be associated with a relatively low confidence level.

[0099] Thus, when a particular key point is identified with a relatively high confidence (e.g., when visible), the threshold distance for that particular key point may take a first value, and when the particular key point is identified with a relatively low confidence (e.g., when hidden), a second, smaller value may be taken. The threshold distance value may range from (i) a maximum threshold corresponding to the lowest allowable confidence value to (ii) a minimum threshold corresponding to the maximum confidence value (i.e., 100%). In some cases, the maximum and minimum thresholds may be predefined. Thus, to take into account key point identifications with lower confidence, the size of the sphere defined by the threshold distance may be increased, such that a larger portion of the environment is considered "touched" and more surface area is cleaned. Similarly, to take into account key point identifications with higher confidence, the size of the sphere defined by the threshold distance may be reduced, such that a smaller portion of the environment is considered "touched" and less surface area is cleaned.

[0100] As the actor 500 interacts with the environment, it can also be determined that the actor 500 has touched other portions of other static features within the environment. For example, based on additional sensor data (not shown) captured at different points in time, the above process can be used to determine that the actor 500 has touched portion 536 of the table 502, such as Figure 5D As shown. Therefore, the map can be further updated to indicate that portion 536 is also to be cleaned. Thus, the map can be used to represent the environment and track portions of environmental features that have (potentially and / or actually) been touched and are therefore designated for cleaning by the robot 200.

[0101] Once the actors 500 have evacuated the area around the table 502, thereby providing the robot 200 with sufficient space to move around the table 502, the robot 200 can be used to clean portions 534 and 536 of the table 502. To this end, the robot 200 can include an end effector 538 that is configured to allow for cleaning of portions of the environment. Although the end effector 538 is illustrated as a sponge, the end effector can additionally or alternatively include one or more brushes, sprayers (allowing for the application of cleaning agents), lights, vacuum cleaners, mops, sweepers, wipers, and / or other cleaning tools. In some cases, the robot 200 can be equipped with a plurality of different end effectors and can be configured to select a particular end effector from the plurality of different end effectors based on, for example, the type / classification of the surface to be cleaned, the degree of cleaning to be performed, and / or the amount of cleaning time allowed, among other considerations.

[0102] In some implementations, the environment map can indicate a degree of uncleanliness of a portion of the environment and / or a degree to which the portion of the environment is to be cleaned. Thus, updating the map to indicate that a particular portion of a particular static feature is to be cleaned can involve updating the map to indicate a degree of uncleanliness of a portion of the static feature and / or a degree to which the portion of the static feature is to be cleaned.

[0103] In one example, the map can be updated to indicate the degree of uncleanliness of portion 534 based on the distance between portion 534 and key point 514A. As key point 514A gets closer to portion 534 and / or its subregions, the degree of uncleanliness of portion 534 and / or its subregions increases. In another example, the map can be updated to indicate the degree of uncleanliness of portion 534 based on the number of times portion 534 was within a corresponding threshold distance of one or more specific key points among key points 506-522B before being cleaned. In other words, the degree of uncleanliness of portion 534 can be based on the number of times portion 534 is determined to have been touched, with the degree of uncleanliness increasing as the number of times portion 534 is determined to have been touched increases.

[0104] The map can be updated to track the degree of uncleanliness. For example, each section and / or sub-area thereof represented by the map can be associated with a "cleanliness" variable, which can range from 0 (clean) to 100 (very dirty), among other possibilities. The degree of cleanliness of section 534 by robot 200 can be based on the degree of uncleanliness of section 534, with the degree of cleanliness increasing as the degree of uncleanliness increases.

[0105] Once the robot 200 completes cleaning of portions 534 and 536, the corresponding portion of the map can be updated to indicate that portions 534 and 536 have been cleaned. In some cases, the respective times that portions 534 and 536 were touched and the respective times that portions 534 and 536 were cleaned can be stored and used to monitor the touch frequency and / or cleaning frequency associated with various features of the environment. For example, a second map for tracking the touch frequency of these surfaces can be updated to increase the touch count and cleaning count associated with table 502 by 2 and / or the touch count and cleaning count associated with each of portions 534 and 536 by 1. Thus, the relative touch frequencies of various features of the environment (or portions thereof) can be determined and used, for example, to determine which features of the environment are considered high touch and which are considered low touch.

[0106] In addition to or as an alternative to the cleaning portions 534 and 536, the robot 200 and / or a computing device associated therewith may be configured to provide one or more visual representations of information collected by the robot 200 and / or contained in the map (e.g., number of touches, number of cleanings, degree of touch, etc.). In one example, the robot 200 may be equipped with a projector that is configured to project the visual representation onto the environment. Thus, the robot 200 may use a projector to project visual representations of various aspects of the map onto the environment. Specifically, different portions of the visual representation of the map may be spatially aligned with corresponding portions of features of the environment. Thus, a given information unit may be visually overlaid on a portion of the environment described and / or quantified by the given information unit.

[0107] For example, a first portion of the visual representation may be spatially aligned with portion 534 of table 502 to visually indicate that portion 534 is to be cleaned, and a second portion of the visual representation may be spatially aligned with portion 536 of table 502 to visually indicate that portion 536 is to be cleaned. The shape and / or size of the first and second portions of the visual representation may be selected to match the physical shape and / or size of portions 534 and 536. The robot 200 may be configured to determine and apply a keystone correction to the visual representation based on the relative positioning of the projector and the portion of the environmental feature illuminated by it. This projection of visual representations of aspects of the map onto the environment may allow actors present in the environment to see aspects of the map overlaid on the environment. Thus, in some cases, the visual projection may be provided by the robot 200 based on and / or in response to detecting the presence of one or more actors in the environment.

[0108] In another example, the robot 200 and / or a computing device associated therewith can be configured to generate an image representing information collected by the robot 200 and / or contained in a map. The image can represent the environment from a particular perspective and can include spatially corresponding aspects of the map overlaid on the environment. In some implementations, the sensor data captured by the robot 200 can include a first image representing the environment. A second image can be generated by projecting aspects of the map onto the image space of the first image. For example, a visual representation of portion 534 and portion 536 can be projected onto the image space of the first image to indicate within the second image that portion 534 and portion 536 are to be cleaned.

[0109] The third image can be generated by inserting at least a portion of the second image representing relevant aspects of the map into the first image. Thus, for example, a third image can be generated by inserting at least a region of the second image representing portions 534 and 536 into the first image to indicate that portions 534 and 536 of table 502 are to be cleaned. The third image can be displayed to allow for visual presentation of various aspects of the map. For example, the third image can include a heat map indicating the degree of uncleanliness of a portion of the environment, a history of touches on the portion of the environment, and / or a history of cleanings of the portion of the environment, among other possibilities. Thus, an actor can use the third image to help clean the environment.

[0110] IV. Example Touch Part Detection System

[0111] Figure 6 A system 600 is illustrated that can be configured to identify / determine portions of an environment that are considered to have been touched by an actor. Specifically, the system 600 includes a key point extraction model 606, a key point proximity calculator 616, and a feature engagement model 624. Each of the model 606, the calculator 616, and the model 624 can be implemented using hardware components (e.g., dedicated circuitry), software instructions (e.g., configured to execute on general-purpose circuitry), and / or a combination thereof. In some implementations, the system 600 can form a subset of the control system 118.

[0112] The keypoint extraction model 606 can be configured to identify keypoints 608 based on at least the sensor data 602. The keypoints 608 can include a keypoint 610 and keypoints 612 to 614 (i.e., keypoints 610-614). Each of the keypoints 610-614 can be associated with a corresponding position (e.g., in the reference frame of the sensor data) and, in some cases, can be associated with a confidence value indicating the accuracy of the corresponding position. The sensor data 602 can be obtained by one or more sensors on the robot 200 (e.g., such as Figure 5A ), and the key point 608 may correspond to Figure 5B and 5C Key points 505-522B are shown in FIG.

[0113] In some implementations, the keypoint extraction model 606 can additionally be configured to determine keypoints 608 based on additional sensor data 604, which can be received from sensors on other robotic devices operating within the environment and / or sensors located at fixed locations throughout the environment. The additional sensor data 604 can represent the actor from a different perspective than the sensor data 602 and can therefore provide additional information about the location of the keypoints 608. The keypoint extraction model 606 can include a rule-based algorithm and / or a machine learning algorithm / model. For example, the keypoint extraction model 606 can include a feed-forward artificial neural network trained based on a plurality of training sensor data labeled with corresponding keypoints to identify the keypoints 608.

[0114] The key point proximity calculator 616 can be configured to identify a candidate touched portion 618 based on the environment map 630 and one or more key points 608. In some cases, the key point proximity calculator 616 can identify multiple candidate touched portions (e.g., multiple instances of the candidate touched portion 618), but for clarity of illustration, only a single candidate touched portion 618 is shown. The touch coordinates 622 of the candidate touched portion 618 can represent a portion of a static feature represented by the environment map 630 that falls within a corresponding threshold distance of the key point 620. Thus, the touch coordinates 622 can define an area and / or volume within the environment map 630 that is determined to have been touched. The key point 620 can represent a specific one of the key points 608 (e.g., key point 612) and can be associated with a body position (e.g., a finger, hand, arm, etc.) that is typically used to interact with and / or touch a surface.

[0115] Because some of the key points 610-614 may be associated with subject locations that may not be commonly used to touch and / or interact with environmental features, the key point proximity calculator 616 may, in some cases, ignore these key points when identifying candidate touched portions 618. However, in other cases, the key point proximity calculator 616 may be configured to consider each of the key points 610-614 when identifying candidate touched portions 618. The key point proximity calculator 616 may be configured to apply a corresponding threshold distance to each respective key point 610-614 depending on the subject location represented by the respective key point and / or a confidence value associated with the estimated location of the respective key point.

[0116] The keypoint proximity calculator 616 can be configured to transform the respective locations of one or more keypoints 608 (e.g., keypoint 612) from the reference frame of the sensor data to the reference frame of the environment map 630. Furthermore, the keypoint proximity calculator 616 can be configured to identify intersections between static features represented by the environment map 630 and respective threshold distance "spheres" associated with one or more of the keypoints 608. Thus, the candidate touched portions 618 can indicate areas of the environment that are likely to have been touched by an actor and, therefore, are candidates for cleaning by the robot 200.

[0117] In some implementations, the feature engagement model 624 can be omitted from the system 600. Thus, the candidate touched portion 618 can constitute the environment map update 626 because the touch coordinates 622 can identify an area to be cleaned in the environment map 630. In other implementations, the feature engagement model 624 can be configured to perform further processing to confirm and / or verify the accuracy of the candidate touched portion 618. Specifically, the feature engagement model 624 can be configured to determine whether the key point 608 indicates that the actor is engaged or disengaged with the environmental feature associated with the candidate touched portion 618, thereby serving as confirmation / validation of the output of the key point proximity calculator 616.

[0118] In one example, the feature engagement model 624 can be configured to determine the degree of engagement between (i) the actor represented by the key point 608 and (ii) the candidate touched portion 618 and / or underlying features of the environment. For example, the feature engagement model 624 can be trained to determine whether the actor is engaged with the candidate touched portion 618, whether it is neutral relative to the candidate touched portion 618, or whether it is disengaged from the candidate touched portion 618. The relative positioning of the key points 608 relative to each other and relative to the candidate touched portion 618 can provide clues indicating the degree of engagement. The feature engagement model 624 can be configured to determine the degree of engagement based on these clues, for example through a supervised training process. Therefore, the feature engagement model 624 can include one or more machine learning models that are trained to determine the degree of engagement between the actor and the portion 618.

[0119] In another example, the feature engagement model 624 can be configured to determine the degree of engagement further based on the actor's motion pattern over time. For example, the feature engagement model 624 can be configured to consider key points 608 associated with the most recent sensor data, as well as key points extracted from sensor data received at one or more earlier times, to determine the degree of engagement. Specifically, the actor's motion pattern over time can provide additional clues indicating the degree of engagement of the actor with the candidate touched portion 618. Therefore, the feature engagement model 624 can include one or more time-based machine learning models, such as a recurrent neural network or a long short-term memory (LSTM) neural network, trained to determine the degree of engagement based on a time series of key points.

[0120] When the feature engagement model 624 determines that the actor is engaged with the portion 618 and / or the underlying features of the environment (e.g., when the degree of engagement determined by the feature engagement model 624 exceeds a threshold engagement value), the candidate touched portion 618 can be used to generate an environment map update 626. That is, the environment map update 626 can indicate that the touch coordinates 622 are to be cleaned by the robot 200. The feature engagement model 624 can thus signal that the candidate touched portion 618 is a true class resulting from the actor's interaction with the candidate touched portion 618.

[0121] When the feature engagement model 624 determines that the actor is not engaged with the portion 618 and / or the underlying feature of the environment (e.g., when the degree of engagement determined by the feature engagement model 624 does not exceed a threshold engagement value), the candidate touched portion 618 can be discarded. That is, the environment map update 626 can not be generated based on the candidate touched portion 618. The feature engagement model 624 can therefore signal that the candidate touched portion 618 is a false positive class caused by the actor's body position entering within the corresponding threshold distance of the feature, but the actor did not actually interact with the candidate touched portion 618.

[0122] Feature engagement model 624 can thus help distinguish between (i) situations where an actor's body part is located close enough to a feature to cause a candidate touched portion 618 to be generated without actually interacting with the feature, and (ii) situations where an actor's body part is located close to and / or in contact with the feature as part of a physical interaction with the feature. Figure 5A (i) the actor 500 approaches the table 502 but does not interact with it, and (ii) the actor 500 picks up the cup 504 from the table 502 or otherwise interacts with the table 502.

[0123] Additionally, by using the keypoint proximity calculator 616 in conjunction with the feature engagement model 624, the computational cost associated with the system 600 and / or the amount of cleaning performed by the robot 200 can be reduced and / or minimized. Specifically, by executing the feature engagement model 624 relative to the candidate touched portions 618 rather than relative to any static features present in the environment, the amount of processing performed by the feature engagement model 624 can be reduced and / or minimized. Because features that are not within the corresponding threshold distance of one or more of the keypoints 608 are unlikely to have been touched by the actor, the degree of engagement between the actor and these features may be irrelevant to the purpose of cleaning the environment. Therefore, omitting the determination of the degree of engagement for features that are not near the actor can result in fewer executions of the feature engagement model 624. Similarly, by using the feature engagement model 624 to identify candidate touched portions for false positives, the robot 200 can focus on cleaning portions of the environment that are likely to be engaged by the actor and avoid cleaning portions of the environment that are likely not engaged by the actor.

[0124] In some implementations, the robot 200 may not be used to monitor and / or clean an environment. For example, sensor data for detecting touched surfaces can be received from sensors located at fixed locations within the environment, and cleaning can be performed by human actors. Specifically, the human actors can use the map and its visualization of the portion to be cleaned to perform a more targeted cleaning process. This approach can be used in environments where robotic equipment may not be suitable and / or operational, such as within a vehicle. Specifically, the system can be used to track surfaces touched by vehicle passengers, and the map can be used by the driver and / or operator of the vehicle to clean the surfaces touched by the passengers.

[0125] V. Additional Sample Operations

[0126] Figure 7 A flow chart illustrating operations related to identifying portions of an environment to be cleaned, tracking those portions, and cleaning those portions by a robot is shown. These operations may be performed by the robot system 100, the robot 200, and / or the system 600, among other possibilities. Figure 7 The embodiments of the present invention may be simplified by removing any one or more features shown therein. Furthermore, these embodiments may be combined with any of the features, aspects, and / or implementations described in any of the previous figures or otherwise herein.

[0127] Block 700 may include determining a map representing a plurality of static features of an environment.

[0128] Block 702 may include receiving sensor data representative of an environment from a sensor disposed on a robotic device.

[0129] Block 704 may include determining a representation of an actor in the environment based on the sensor data. The representation may include a plurality of key points representing corresponding body positions of the actor.

[0130] Block 706 may include determining that a portion of a particular static feature of the plurality of static features of the environment is within a threshold distance of a particular key point of the plurality of key points.

[0131] Block 708 may include updating the map to indicate that the portion of the particular static feature is to be cleaned based on determining that the particular keypoint is within the threshold distance of the portion of the particular static feature.

[0132] Block 710 may include causing the robotic device to clean the portion of the particular static feature based on the updated map.

[0133] In some embodiments, multiple key points may represent an actor's arms, and certain key points may represent an actor's hands.

[0134] In some embodiments, determining a representation of an actor may include determining a plurality of key points via a machine learning model trained to identify a plurality of key points within sensor data representing the actor, and mapping the positions of the plurality of key points from a reference frame of the sensor data to a reference frame of the map.

[0135] In some embodiments, the map may include a plurality of voxels. Determining that a particular keypoint in the plurality of keypoints is within a threshold distance of the portion of the particular static feature may include determining one or more voxels in the plurality of voxels that (i) represent the particular static feature and (ii) are within the threshold distance of the particular keypoint.

[0136] In some embodiments, based on the robotic device cleaning the portion of the particular static feature, the map may be updated to indicate that the portion of the particular static feature has been cleaned.

[0137] In some embodiments, a subset of multiple static features can be determined, wherein each corresponding static feature of the subset is associated with a touch frequency exceeding a threshold frequency. The subset can include a specific static feature. Based on determining the subset of multiple static features, it can be determined for each corresponding static feature of the subset whether one or more portions of the corresponding static feature are within a threshold distance of a specific key point. Based on determining whether one or more portions of the corresponding static feature are within the threshold distance of the specific key point, it can be determined that the portion of the specific static feature is within the threshold distance of the specific key point.

[0138] In some embodiments, determining the map may include receiving additional sensor data representing one or more of a plurality of static features of the environment from sensors as the robotic device moves within the environment, and generating the map based on the additional sensor data.

[0139] In some embodiments, determining the map may include, as the one or more other robotic devices move in the environment, receiving additional sensor data representing one or more of a plurality of static features of the environment from one or more other sensors disposed on the one or more other robotic devices, and generating the map based on the additional sensor data.

[0140] In some embodiments, based on the relative positioning of key points in a plurality of key points, it can be determined that an actor is engaged with a particular static feature. The map can be further updated based on determining that the actor is engaged with the particular static feature to indicate that the portion of the particular static feature is to be cleaned.

[0141] In some embodiments, the sensor data may include multiple sensor data frames representing an agent in the environment at multiple different points in time. Based on the multiple sensor data frames, multiple representations of the agent corresponding to the multiple different points in time may be determined. Based on the multiple representations of the agent, a motion pattern of the agent may be determined. It may be determined that the motion pattern of the agent indicates contact between the agent and a particular static feature. The map may be updated to indicate that the portion of the particular static feature is to be cleaned, further based on determining that the motion pattern of the agent indicates contact between the agent and the particular static feature.

[0142] In some embodiments, additional sensor data representing an actor within the environment may be received from additional sensors located within the environment. The additional sensor data may represent the actor from a perspective different from the sensor data. The representation of the actor may be further determined based on the additional sensor data.

[0143] In some embodiments, a subject position corresponding to a particular keypoint within the sensor data may be occluded. Determining a representation of the actor may include determining a location of the particular keypoint in the environment based on the unoccluded subject position of the actor.

[0144] In some embodiments, the threshold distance may have a first value when the subject position corresponding to the particular key point within the sensor data is not obscured, and may have a second value greater than the first value when the subject position corresponding to the particular key point within the sensor data is obscured.

[0145] In some embodiments, updating the map may include updating the map to indicate a touch time at which the portion of the particular static feature is indicated as to be cleaned. The map may also be configured to be updated to indicate a cleaning time at which the portion of the particular static feature was cleaned by the robotic device.

[0146] In some embodiments, a projector can be connected to the robotic device. The projector can be used to project a visual representation of the updated map onto the environment. A portion of the visual representation of the map can be spatially aligned with the portion of a particular static feature to visually indicate that the portion of the particular static feature is to be cleaned.

[0147] In some embodiments, sensor data representing an environment may include a first image representing the environment. A second image may be generated by projecting an updated map onto the image space of the first image to indicate within the second image the portion of a specific static feature to be cleaned. A third image may be generated by inserting at least a portion of the second image indicating the portion of the specific static feature to be cleaned into the first image. The third image may be displayed.

[0148] In some embodiments, updating the map to indicate that the portion of the particular static feature is to be cleaned may include updating the map to indicate a degree of uncleanliness of the portion of the static feature based on one or more of: (i) a distance between the portion of the particular static feature and a particular keypoint, or (ii) a number of times the portion of the particular static feature has been within a respective threshold distance of one or more of a plurality of keypoints. The degree of cleanliness of the portion of the particular static feature by the robotic device may be based on the degree of uncleanliness of the portion of the static feature.

[0149] Figure 8 A flowchart illustrating operations related to identifying portions of an environment determined to have been touched, tracking those portions, and displaying visual representations of those portions is shown. These operations may be performed by the robot system 100, the robot 200, and / or the system 600, among other possibilities. Figure 8 The embodiments of the present invention may be simplified by removing any one or more features shown therein. Furthermore, these embodiments may be combined with any of the features, aspects, and / or implementations described in any of the previous figures or otherwise herein.

[0150] Block 800 may include determining a map representing a plurality of static features of an environment.

[0151] Block 802 can include receiving sensor data representing an environment from a sensor. In some implementations, the sensor can be located on a robotic device.

[0152] Block 804 may include determining a representation of an actor in the environment based on the sensor data. The representation may include a plurality of key points representing corresponding body positions of the actor.

[0153] Block 806 may include determining that a portion of a particular static feature of the plurality of static features of the environment is within a threshold distance of a particular key point of the plurality of key points.

[0154] Block 808 may include updating the map to indicate that the portion of the particular static feature is determined to have been touched based on determining that the particular key point is within a threshold distance of the portion of the particular static feature. Determining that the portion of the particular static feature has been touched may include determining that the portion appears to have been touched based on the sensor data, and / or determining that the portion is actually physically touched.

[0155] Block 810 may include displaying a visual representation indicating that the portion of the particular static feature is determined to have been touched based on the updated map. The visual representation may be an image, a video, and / or a projection thereof on the environment, among other possibilities.

[0156] In some embodiments, multiple key points may represent an actor's arms, and certain key points may represent an actor's hands.

[0157] In some embodiments, determining a representation of an actor may include determining a plurality of key points via a machine learning model trained to identify a plurality of key points within sensor data representing the actor, and mapping the positions of the plurality of key points from a reference frame of the sensor data to a reference frame of the map.

[0158] In some embodiments, the map may include a plurality of voxels. Determining that a particular keypoint in the plurality of keypoints is within a threshold distance of the portion of the particular static feature may include determining one or more voxels in the plurality of voxels that (i) represent the particular static feature and (ii) are within the threshold distance of the particular keypoint.

[0159] In some embodiments, the robotic device may be caused to clean the portion of the particular static feature based on the updated map.

[0160] In some embodiments, based on the robotic device cleaning the portion of the particular static feature, the map may be updated to indicate that the portion of the particular static feature has been cleaned.

[0161] In some embodiments, an additional visual representation may be displayed to indicate that the portion of a particular static feature has been cleaned.

[0162] In some embodiments, a subset of multiple static features can be determined, wherein each corresponding static feature in the subset is associated with a touch frequency exceeding a threshold frequency. The subset can include a specific static feature. Based on determining the subset of multiple static features, it can be determined for each corresponding static feature in the subset whether one or more portions of the corresponding static feature are within a threshold distance of a specific key point. Based on determining whether one or more portions of the corresponding static feature are within the threshold distance of the specific key point, it can be determined that the portion of the specific static feature is within the threshold distance of the specific key point.

[0163] In some embodiments, determining the map may include receiving additional sensor data from a sensor representing one or more static features of a plurality of static features of the environment, and generating the map based on the additional sensor data.

[0164] In some embodiments, determining a map may include receiving additional sensor data representing one or more of a plurality of static features of the environment from one or more other sensors disposed on the one or more robotic devices as the one or more robotic devices move within the environment, and generating a map based on the additional sensor data.

[0165] In some embodiments, based on the relative positioning of key points in the plurality of key points, it can be determined that the actor is engaged with a particular static feature. The map can be further updated based on the determination that the actor is engaged with the particular static feature to indicate that the portion of the particular static feature is determined to be touched.

[0166] In some embodiments, the sensor data may include multiple sensor data frames representing an agent in the environment at multiple different points in time. Based on the multiple sensor data frames, multiple representations of the agent corresponding to the multiple different points in time may be determined. Based on the multiple representations of the agent, a motion pattern of the agent may be determined. It may be determined that the motion pattern of the agent indicates contact between the agent and a particular static feature. The map may be updated to indicate that the portion of the particular static feature is determined to be touched, further based on determining that the motion pattern of the agent indicates contact between the agent and the particular static feature.

[0167] In some embodiments, additional sensor data representing an actor within the environment may be received from additional sensors located within the environment. The additional sensor data may represent the actor from a perspective different from the sensor data. The representation of the actor may be further determined based on the additional sensor data.

[0168] In some embodiments, a subject position corresponding to a particular keypoint may be occluded within the sensor data.Determining a representation of the actor may include determining a location of the particular keypoint in the environment based on the unoccluded subject position of the actor.

[0169] In some embodiments, the threshold distance may have a first value when the subject position corresponding to the particular key point is not obscured within the sensor data, and may have a second value greater than the first value when the subject position corresponding to the particular key point is obscured within the sensor data.

[0170] In some embodiments, updating the map may include updating the map to indicate a touch time at which the portion of the particular static feature was determined to be touched.

[0171] In some embodiments, the map may be further configured to be updated to indicate the cleaning time at which the portion of the particular static feature was cleaned by the robotic device.

[0172] In some embodiments, the projector can be configured to project a visual representation onto the environment.A portion of the visual representation can be spatially aligned with the portion of the particular static feature to visually indicate that the portion of the particular static feature has been determined to have been touched.

[0173] In some embodiments, the sensor data representing the environment may include a first image representing the environment. A second image may be generated by projecting an updated map onto the image space of the first image to indicate within the second image the portion of the particular static feature determined to have been touched. A visual representation (e.g., a third image) may be generated by inserting at least a portion of the second image indicating the portion of the particular static feature determined to have been touched into the first image.

[0174] In some embodiments, updating the map to indicate that the portion of the particular static feature is determined to have been touched may include updating the map to indicate the extent to which the portion of the static feature has been touched based on one or more of: (i) a distance between the portion of the particular static feature and a particular keypoint, or (ii) a number of times the portion of the particular static feature has been within a respective threshold distance of one or more of a plurality of keypoints. The cleanliness of the portion of the particular static feature may be based on the extent to which the portion of the static feature has been touched.

[0175] VI. Conclusion

[0176] The present disclosure is not limited to the specific embodiments described in this application, which are intended to illustrate various aspects. It will be apparent to those skilled in the art that many modifications and variations can be made without departing from the scope thereof. In addition to the methods and apparatus described herein, functionally equivalent methods and apparatus within the scope of the present disclosure will be apparent to those skilled in the art based on the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.

[0177] The above detailed description describes various features and operations of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the accompanying drawings, similar symbols generally identify similar components unless the context indicates otherwise. The example embodiments described herein and in the accompanying drawings are not meant to be limiting. Other embodiments may be utilized and other changes may be made without departing from the scope of the subject matter presented herein. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the accompanying drawings, may be arranged, replaced, combined, separated, and designed in a variety of different configurations.

[0178] With respect to any or all message flow diagrams, scenarios, and process diagrams in the accompanying drawings, and as discussed herein, each step, block, and / or communication may represent processing of information and / or transmission of information, according to example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, the operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages may be performed out of the order shown or discussed, including substantially simultaneously or in reverse order, depending on the functionality involved. Additionally, more or fewer blocks and / or operations may be included with any message flow diagrams, scenarios, and process diagrams discussed herein. Figure 1 The present invention relates to a method for the use of message flow diagrams, scenarios and flowcharts, and these message flow diagrams, scenarios and flowcharts can be combined with each other in part or in whole.

[0179] The steps or blocks representing information processing may correspond to circuits that may be configured to perform the specific logical functions of the methods or techniques described herein. Alternatively or additionally, the blocks representing information processing may correspond to modules, fragments, or portions of program code (including associated data). The program code may include one or more instructions executable by a processor for implementing specific logical operations or actions in the methods or techniques. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device including random access memory (RAM), a disk drive, a solid-state drive, or another storage medium.

[0180] Computer-readable media can also include non-transitory computer-readable media, such as computer-readable media for short-term storage of data, such as register memory, processor cache, and RAM. Computer-readable media can also include non-transitory computer-readable media for long-term storage of program code and / or data. Thus, computer-readable media can include secondary or permanent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, solid-state drives, compact disk read-only memory (CD-ROM). Computer-readable media can also be any other volatile or non-volatile storage system. For example, computer-readable media can be considered to be computer-readable storage media, or tangible storage devices.

[0181] In addition, steps or blocks representing one or more information transfers may correspond to information transfers between software and / or hardware modules in the same physical device. However, other information transfers may be performed between software and / or hardware modules in different physical devices.

[0182] The specific arrangements shown in the accompanying drawings should not be considered restrictive. It should be understood that other embodiments may include more or fewer of each element shown in a given figure. In addition, some of the illustrated elements may be combined or omitted. In addition, example embodiments may include elements not shown in the accompanying drawings.

[0183] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes only and are not intended to be limiting, with the true scope being indicated by the appended claims.

Claims

1. A robotic system comprising: robotic equipment; a sensor disposed on the robotic device; and Circuitry configured to perform operations comprising: determining a map representing a plurality of static features of an environment; receiving sensor data representative of the environment from the sensor; determining a representation of an agent within the environment based on the sensor data, wherein the representation includes a plurality of key points representing corresponding body positions of the agent; determining that a portion of a particular static feature of the plurality of static features of the environment is within a threshold distance of a particular key point of the plurality of key points; Based on determining that the particular keypoint is within the threshold distance of the portion of the particular static feature, updating the map to indicate that the portion of the particular static feature is to be cleaned; and The robotic device is caused to clean the portion of the particular static feature based on the updated map.

2. The system according to claim 1, wherein: The plurality of key points represent arms of the actor, and wherein the specific key point represents a hand of the actor.

3. The system according to claim 1, wherein: Determining the representation of the actor includes: determining the plurality of key points via a machine learning model trained to identify the plurality of key points within sensor data representing the actor; and The positions of the plurality of key points are mapped from a reference system of the sensor data to a reference system of the map.

4. The system according to claim 1, wherein: The map includes a plurality of voxels, and wherein determining that the particular keypoint of the plurality of keypoints is within the indicated threshold distance of the portion of the particular static feature includes: One or more voxels of the plurality of voxels that (i) represent the particular static feature and (ii) are within the threshold distance of the particular keypoint are determined.

5. The system according to claim 1, wherein The operations further include: Based on the robotic device cleaning the portion of the particular static feature, the map is updated to indicate that the portion of the particular static feature has been cleaned.

6. The system according to claim 1, wherein: The operations further include: determining a subset of the plurality of static features, wherein each respective static feature in the subset is associated with a touch frequency exceeding a threshold frequency, and wherein the subset includes the particular static feature; Based on determining the subset of the plurality of static features, for each respective static feature in the subset, determining whether one or more portions of the respective static feature are within the threshold distance of the particular keypoint; and The portion of the particular static feature is determined to be within the threshold distance of the particular keypoint based on determining whether the one or more portions of the corresponding static feature are within the threshold distance of the particular keypoint.

7. The system according to claim 1, wherein: Determining the map includes: receiving, from the sensor, additional sensor data representative of one or more of the plurality of static features of the environment as the robotic device moves through the environment; and The map is generated based on the additional sensor data.

8. The system according to claim 1, wherein: Determining the map includes: receiving additional sensor data representative of one or more of the plurality of static features of the environment from one or more other sensors disposed on the one or more other robotic devices as the one or more other robotic devices move within the environment; and The map is generated based on the additional sensor data.

9. The system according to claim 1, wherein: The operations further include: The actor is determined to be engaged with the particular static feature based on relative positioning of keypoints in the plurality of keypoints, wherein additionally based on determining that the actor is engaged with the particular static feature, the map is updated to indicate that the portion of the particular static feature is to be cleaned.

10. The system according to claim 1, wherein: The sensor data includes a plurality of sensor data frames representing the actor within the environment at a plurality of different points in time, and wherein the operations further include: determining, based on the plurality of sensor data frames, a plurality of representations of the actor corresponding to the plurality of different points in time; determining a motion pattern of the agent based on the plurality of representations of the agent; and Determining that the movement pattern of the actor indicates contact between the actor and the specific static feature, wherein, additionally based on determining that the movement pattern of the actor indicates contact between the actor and the specific static feature, updating the map to indicate that the portion of the specific static feature is to be cleaned.

11. The system according to claim 1, wherein: The operations further include: receiving additional sensor data representing the actor within the environment from additional sensors located within the environment, wherein the additional sensor data represents the actor from a different perspective than the sensor data; and A representation of the actor is further determined based on the additional sensor data.

12. The system according to claim 1, wherein: The subject position corresponding to the particular keypoint is occluded within the sensor data, and wherein determining the representation of the actor comprises: The location of the specific key point in the environment is determined based on the unobstructed body position of the actor.

13. The system of claim 1, wherein: The threshold distance has a first value when the subject position corresponding to the particular key point is not obscured within the sensor data, and wherein the threshold distance has a second value greater than the first value when the subject position corresponding to the particular key point is obscured within the sensor data.

14. The system according to claim 1, wherein: Updates to the map include: The map is updated to indicate a touch time when the portion of the particular static feature is indicated to be cleaned, wherein the map is further configured to be updated to indicate a cleaning time when the portion of the particular static feature is cleaned by the robotic device.

15. The system of claim 1, further comprising a projector connected to the robotic device, wherein The operations further include: A visual representation of the updated map is projected onto the environment using the projector, wherein a portion of the visual representation of the map is spatially aligned with the portion of the particular static feature to visually indicate that the portion of the particular static feature is to be cleaned.

16. The system of claim 1, wherein: The sensor data representative of the environment includes a first image representative of the environment, and wherein the operations further comprise: generating a second image by projecting the updated map onto the image space of the first image to indicate within the second image the portion of the particular static feature to be cleaned; generating a third image by inserting at least a portion of the second image indicating the portion of the particular static feature to be cleaned into the first image; and The third image is displayed.

17. The system of claim 1, wherein: Updating the map to indicate the portion of the particular static feature to be cleaned includes: The map is updated to indicate a degree of uncleanliness of the portion of the static feature based on one or more of: (i) a distance between the portion of the specific static feature and the specific key point, or (ii) a number of times the portion of the specific static feature has been within a corresponding threshold distance of one or more specific key points among the multiple key points, wherein the degree of cleanliness of the portion of the specific static feature by the robotic device is based on the degree of uncleanliness of the portion of the static feature.

18. A computer-implemented method for controlling a robotic device, comprising: determining a map representing a plurality of static features of an environment; receiving sensor data representative of the environment from a sensor disposed on the robotic device; determining a representation of an agent within the environment based on the sensor data, wherein the representation includes a plurality of key points representing corresponding body positions of the agent; determining that a portion of a particular static feature of the plurality of static features of the environment is within a threshold distance of a particular key point of the plurality of key points; Based on determining that the particular keypoint is within the threshold distance of the portion of the particular static feature, updating the map to indicate that the portion of the particular static feature is to be cleaned; and The robotic device is caused to clean the portion of the particular static feature based on the updated map.

19. The computer-implemented method of claim 18, further comprising: Based on the robotic device cleaning the portion of the particular static feature, the map is updated to indicate that the portion of the particular static feature has been cleaned.

20. A non-transitory computer-readable medium having stored thereon instructions that, when executed by a computing device, cause the computing device to perform operations comprising: determining a map representing a plurality of static features of an environment; receiving sensor data representative of the environment from a sensor disposed on the robotic device; determining a representation of an agent within the environment based on the sensor data, wherein the representation includes a plurality of key points representing corresponding body positions of the agent; determining that a portion of a particular static feature of the plurality of static features of the environment is within a threshold distance of a particular key point of the plurality of key points; Based on determining that the particular keypoint is within the threshold distance of the portion of the particular static feature, updating the map to indicate that the portion of the particular static feature is to be cleaned; and The robotic device is caused to clean the portion of the particular static feature based on the updated map.

21. A computer-implemented method for identifying a touch, comprising: determining a map representing a plurality of static features of an environment; receiving sensor data representative of the environment from a sensor; determining a representation of an agent within the environment based on the sensor data, wherein the representation includes a plurality of key points representing corresponding body positions of the agent; determining that a portion of a particular static feature of the plurality of static features of the environment is within a threshold distance of a particular key point of the plurality of key points; Based on determining that the particular keypoint is within the threshold distance of the portion of the particular static feature, updating the map to indicate that the portion of the particular static feature is determined to have been touched; and Based on the updated map, an output is provided indicating that the portion of the particular static feature is determined to have been touched.

22. The computer-implemented method of claim 21 , wherein: Output provided includes: A visual representation is displayed indicating that the portion of the particular static feature is determined to have been touched.

23. The computer-implemented method of claim 21 , wherein: Output provided includes: A visual representation indicating that the portion of the particular static feature is determined to have been touched is projected onto the environment via a projector, wherein a portion of the visual representation is spatially aligned with the portion of the particular static feature to visually indicate that the portion of the particular static feature is determined to have been touched.

24. The computer-implemented method of claim 21 , wherein: The sensor data representative of the environment includes a first image representative of the environment, and wherein providing the output includes: generating a second image by projecting the updated map onto the image space of the first image to indicate within the second image the portion of the particular static feature determined to have been touched; generating a third image by inserting at least a portion of the second image indicating the portion of the specific static feature determined to have been touched into the first image; and The third image is displayed.

25. The computer-implemented method of claim 21 , wherein: The plurality of key points represent arms of the actor, and wherein the specific key point represents a hand of the actor.

26. The computer-implemented method of claim 21, wherein: Determining the representation of the actor includes: determining the plurality of key points by a machine learning model trained to identify the plurality of key points within sensor data representing the actor; and The positions of the plurality of key points are mapped from a reference system of the sensor data to a reference system of the map.

27. The computer-implemented method of claim 21, wherein: The map includes a plurality of voxels, and wherein determining that the particular keypoint of the plurality of keypoints is within the threshold distance of the portion of the particular static feature includes: One or more voxels of the plurality of voxels that (i) represent the particular static feature and (ii) are within the threshold distance of the particular keypoint are determined.

28. The computer-implemented method of claim 21 , further comprising: generating a request to clean the portion of the particular static feature based on the updated map; based on the portion of the particular static feature being cleaned, updating the map to indicate that the portion of the particular static feature has been cleaned; as well as An additional output is provided indicating that the portion of the particular static feature has been cleaned.

29. The computer-implemented method of claim 28, wherein: Updating the map to indicate that the portion of the particular static feature has been cleaned includes: The map is updated to indicate the cleaning time at which the portion of the particular static feature was cleaned.

30. The computer-implemented method of claim 21, wherein: Determining that the portion of the specific static feature is within the threshold distance of the specific key point includes: determining a subset of the plurality of static features, wherein each respective static feature in the subset is associated with a touch frequency exceeding a threshold frequency, and wherein the subset includes the particular static feature; Based on determining the subset of the plurality of static features, for each respective static feature in the subset, determining whether one or more portions of the respective static feature are within the threshold distance of the particular keypoint; and The portion of the particular static feature is determined to be within the threshold distance of the particular keypoint based on determining whether the one or more portions of the corresponding static feature are within the threshold distance of the particular keypoint.

31. The computer-implemented method of claim 21 , determining the map comprising: receiving, from the sensor, additional sensor data representative of one or more of the plurality of static features of the environment; as well as The map is generated based on the additional sensor data.

32. The computer-implemented method of claim 21 , further comprising: Based on the relative positioning of key points among the multiple key points, it is determined that the actor is engaged with the specific static feature, wherein, additionally based on the determination that the actor is engaged with the specific static feature, the map is updated to indicate that the portion of the specific static feature is determined to have been touched.

33. The computer-implemented method of claim 21 , wherein: The sensor data comprises a plurality of sensor data frames representing the actor within the environment at a plurality of different points in time, and wherein the method further comprises: determining, based on the plurality of sensor data frames, a plurality of representations of the actor corresponding to the plurality of different points in time; determining a motion pattern of the agent based on the plurality of representations of the agent; and Determining that the movement pattern of the actor indicates contact between the actor and the specific static feature, wherein, additionally based on determining that the movement pattern of the actor indicates contact between the actor and the specific static feature, updating the map to indicate that the portion of the specific static feature is determined to have been touched.

34. The computer-implemented method of claim 21 , wherein: The subject position corresponding to the particular keypoint is occluded within the sensor data, and wherein determining the representation of the actor comprises: The location of the specific key point in the environment is determined based on the unobstructed body position of the actor.

35. The computer-implemented method of claim 21 , wherein: The threshold distance has a first value when the subject position corresponding to the particular key point is not obscured within the sensor data, and wherein the threshold distance has a second value greater than the first value when the subject position corresponding to the particular key point is obscured within the sensor data.

36. The computer-implemented method of claim 21, wherein: Updates to the map include: The map is updated to indicate the touch time at which the portion of the particular static feature was determined to have been touched.

37. The computer-implemented method of claim 21, wherein: Updates to the map include: The map is updated to indicate an extent to which the portion of the particular static feature has been touched based on one or more of: (i) a distance between the portion of the particular static feature and the particular key point, or (ii) a number of times that the portion of the particular static feature has been within a corresponding threshold distance of one or more of the plurality of key points, wherein the extent to which the portion of the particular static feature is clean is based on the extent to which the portion of the static feature has been touched.

38. A system for identifying touch, comprising: sensor; and Circuitry configured to perform operations comprising: determining a map representing a plurality of static features of an environment; receiving sensor data representative of the environment from the sensor; determining a representation of an agent within the environment based on the sensor data, wherein the representation includes a plurality of key points representing corresponding body positions of the agent; determining that a portion of a particular static feature of the plurality of static features of the environment is within a threshold distance of a particular key point of the plurality of key points; Based on determining that the particular keypoint is within the threshold distance of the portion of the particular static feature, updating the map to indicate that the portion of the particular static feature is determined to have been touched; and Based on the updated map, an output is provided indicating that the portion of the particular static feature is determined to have been touched.

39. The system of claim 38, wherein: Providing said output includes: A visual representation is displayed indicating that the portion of the particular static feature is determined to have been touched.

40. A non-transitory computer-readable medium having stored thereon instructions that, when executed by a computing device, cause the computing device to perform operations comprising: determining a map representing a plurality of static features of an environment; receiving sensor data representative of the environment from a sensor; determining a representation of an agent within the environment based on the sensor data, wherein the representation includes a plurality of key points representing corresponding body positions of the agent; determining that a portion of a particular static feature of the plurality of static features of the environment is within a threshold distance of a particular key point of the plurality of key points; Based on determining that the particular keypoint is within the threshold distance of the portion of the particular static feature, updating the map to indicate that the portion of the particular static feature is determined to have been touched; and Based on the updated map, an output is provided indicating that the portion of the particular static feature is determined to have been touched.

Citation Information

Patent Citations

  • Robot control system

    DE102019122790A1

  • Mapping an environment using a state of a robotic device

    WO2020152436A1