Intelligent task assignment to mobile robots using augmented reality and 3D semantic maps
A 3D semantic map with augmented reality interface enhances mobile robot mapping and task performance by integrating user-provided semantic information, addressing limitations of 2D mapping and improving user interaction for efficient task execution.
Patent Information
- Application Number
- DE102025123767
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-12-24
AI Technical Summary
Mobile robots relying on conventional 2D mapping techniques for navigation and object identification have limited ability to accurately map environments and perform tasks efficiently, leading to inaccurate mapping and inefficient motion planning, with user interaction being limited to scheduling or manual control.
The use of a 3D semantic map that integrates semantic information generated from sensor data, combined with a graphical user interface for augmented reality on a personal electronic device, allows for intelligent task assignment and more precise user control of the robot's operations.
Enables more robust environmental mapping and intuitive user interaction, enhancing the robot's ability to identify objects and perform tasks efficiently by integrating user-provided semantic information into the 3D map for intelligent task planning.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AREA
[0001] The devices and methods disclosed in this document relate to mobile robots and, in particular, to the intelligent assignment of tasks to mobile robots using augmented reality and 3D semantic maps.
[0002] Unless otherwise stated herein, the materials described in this section shall not be deemed to represent prior art simply by virtue of their inclusion in this section.
[0003] Mobile robots that navigate their environment to perform a task, such as robotic vacuum cleaners, have become increasingly popular in recent years because they can conveniently and efficiently perform tasks like vacuuming floors autonomously. In many cases, these mobile robots navigate their environment using a 2D map, generated, for example, using LiDAR mapping techniques. However, due to this reliance on conventional 2D mapping techniques for navigation and understanding their surroundings, these mobile robots have a limited ability to identify and label objects in their environment, resulting in inaccurate mapping and inefficient motion planning. Such mobile robots are often operated using a software application, such as on a smartphone.However, user interaction with these devices is typically limited to setting a task schedule or manually controlling the device's movement. Therefore, there is a need for a system that allows for more robust environmental mapping and more precise and intuitive user interaction and control of the mapping and cleaning process. SUMMARY
[0004] A method for operating a mobile robot is disclosed. The method comprises storing a three-dimensional semantic map of the environment in memory, wherein the three-dimensional semantic map includes semantic information about the environment and is generated at least partially based on sensor data acquired by the mobile robot while navigating the environment. The method further comprises receiving an image of the environment captured by a camera of a personal electronic device. The method further comprises determining at least one first task to be performed by the mobile robot in the environment, at least partially based on the image and the three-dimensional semantic map.The procedure also involves operating the mobile robot to perform at least one initial task in the environment.
[0005] A further method for operating a mobile robot is also disclosed. This method involves storing a three-dimensional semantic map of the environment in memory, wherein the three-dimensional semantic map includes semantic information about the environment and is generated at least partially based on sensor data acquired by a mobile robot while navigating the environment. The method further involves receiving an image of the environment captured by a camera of a personal electronic device. The method further involves performing semantic segmentation of the image. The method further involves updating the three-dimensional semantic map of the environment based on the semantic segmentation of the image.The method also involves operating the mobile robot based on the updated three-dimensional semantic map.
[0006] A further method for operating a mobile robot is also disclosed. This method involves storing a three-dimensional semantic map of the environment in memory, wherein the three-dimensional semantic map comprises semantic information relating to the environment, and wherein the three-dimensional semantic map is generated at least partially based on sensor data acquired by a mobile robot while navigating the environment. The method further involves displaying a graphical user interface for augmented reality on a display of the personal electronic device, wherein the graphical user interface for augmented reality comprises a multitude of graphical elements superimposed on the environment, and the multitude of graphical elements includes graphical representations of the semantic information of the three-dimensional semantic map.The method further comprises determining at least one task to be performed by the mobile robot based on user input received via the graphical user interface for augmented reality. The method further comprises operating the mobile robot to perform the at least one task using the three-dimensional semantic map. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The above-mentioned designs and other features of the system and the procedures are explained in the following description in conjunction with the attached drawings. Fig. 1 summarizes the components and operations or operating procedures of a mobile robot system. Fig. Figure 2A shows an embodiment of a mobile robot of the mobile robot system. Fig. Figure 2B shows an example implementation of a cloud system for the mobile robot system. Fig. Figure 2C shows an embodiment of a mobile device of the mobile robot system. Fig. Figure 3 shows a flowchart for a procedure for operating a mobile robot. DETAILED DESCRIPTION
[0008] For a better understanding of the principles of the disclosure, reference is now made to the embodiments illustrated in the drawings and described in the following patent specification. It is understood that this is not intended to limit the scope of the disclosure. Furthermore, it is understood that the present disclosure includes all changes and modifications of the illustrated embodiments and further applications of the principles of the disclosure that would normally occur to a person skilled in the art in the field to which this disclosure belongs. Overview
[0009] With reference to Fig. The components and operations of a mobile robot system 10 are summarized. The mobile robot system 10 comprises at least one mobile robot 120 configured to perform a task in an environment. The mobile robot system 10 further comprises a mobile device 170 that can be used to operate and configure the mobile robot 120. Finally, the mobile robot system 10 also comprises a cloud system 150 that manages map data and task assignment for the mobile robot 120.
[0010] In general, the mobile robot 120 comprises a controller 121 configured to operate or control one or more sensors 126 and one or more actuators 128 to navigate autonomously in an environment to perform a task. In the embodiments described in detail herein, the mobile robot 120 is, in particular, a vacuuming robot or a mopping robot configured to navigate in the environment to clean a floor surface. However, it should be understood by a person skilled in the art that the systems and methods described herein can be applied to a variety of mobile robots that navigate autonomously in an environment to perform a task.
[0011] The mobile robot 120 is configured to use a 3D semantic map 164 of its environment for navigation and task performance. Specifically, when the mobile robot 120 is operating to perform tasks in its environment, the controller 121 controls the sensors 126 to detect the positions of walls, objects, or other obstacles. The data collected by the sensors 126 can include depth images, RGB images, and LiDAR point clouds. This sensor data is used to derive a comprehensive, high-resolution 3D semantic map 164 that includes both map data 165 and semantic information 166.
[0012] The map data 165 of the 3D semantic map 164 can take the form of a point cloud, a surface map, and / or a polygon mesh. The map data 165 includes at least three-dimensional geometric information (position, depth, and / or distance) of obstacles or other fixed objects and, in at least some embodiments, also photometric information (color, intensity, and / or brightness). The map data 165 is determined by the mobile robot 120 and / or the cloud system 150 based on sensor data, for example, using simultaneous localization and mapping (SLAM) and other image- or LiDAR-based odometry and mapping algorithms.
[0013] In contrast, the semantic information 166 of the 3D semantic map 164 includes additional information that helps the mobile robot 120 better understand its environment, such as space segmentation labels, object segmentation labels, cleanliness states (or similar states related to tasks performed by the mobile robot), and the like. Such labels can be in the form of text, color shades, a numerical value corresponding to a category or classification, or a unique identifier. Generally, the semantic information 166 is directly associated with specific parts of the map data 165, so that the semantic information 166 annotates the map data 165.The semantic information 166 is determined by the mobile robot 120 and / or the cloud system 150 on the basis of the sensor data and / or the map data 165, for example using one or more machine learning models, such as semantic segmentation models, and also on the basis of user input.
[0014] The 3D semantic map 164 is stored and managed by the cloud system 150. In some embodiments, the cloud system 150 receives sensor data and / or the map data 165 from the mobile robot 120 and processes the map data 165 to derive semantic information 166, for example, using a semantic segmentation model that extracts semantic information from each pixel in an image or from each point in a point cloud. In this way, the 3D semantic map 164 is enriched with additional semantic information that can be used when operating the mobile robot 120 to perform tasks in the environment. However, it should be noted that in some embodiments, the 3D semantic map 164 is stored and maintained locally using the mobile robot 120 or the mobile device 170.In some embodiments, the mobile robot 120 can maintain the 3D Semantic Map 164 with a lower level of detail for use in the event of a network failure.
[0015] With continued reference to Fig. 1. The mobile device 170 provides a mobile robot application used to operate and configure the mobile robot 120. Advantageously, the mobile robot application includes a graphical user interface (GUI) 40 for augmented reality (AR). AR technology is gaining increasing importance for improving human-robot interaction and robot interfaces by enabling more effective and intuitive communication between humans and machines. Using the camera and the graphical AR user interface of the mobile device 170, users can make informed and precise decisions for controlling the mobile robot 120. Graphical AR user interfaces have the potential to improve decision-making by aggregating and visualizing various data sources.By overlaying real-time data directly onto the real world instead of a computer screen, graphical AR user interfaces can revolutionize applications with visualized features such as semantic labels and safety heatmaps. Semantic labels can be in the form of words and / or phrases that can be encoded as plain text or translated into plain text. Alternatively, or additionally, semantic labels can be in the form of pixels shaded with a specific color indicating the semantic label, for example, red shading for "dirty." Interactive visualizations can allow users to examine and understand the robot's decision-making process as the physical environment and task requirements change.
[0016] For this purpose, the graphical AR user interface includes visualizations 50 in which semantic information 166 from the 3D semantic map 164 is directly overlaid onto real-time videos captured by a camera 182 of the mobile device 170. Furthermore, the mobile robot application allows the user to capture images of the environment with the camera 182 and enter custom semantic labels assigned to specific objects and areas within the images. The custom semantic labels can be entered using a user interface of the mobile device 170, such as an on-screen keyboard. Such images and user input 50 are processed by the cloud system 150 to update and enrich the 3D semantic map 164. Specifically, the custom semantic labels are assigned to specific 3D objects or 3D areas within the 3D semantic map 164.
[0017] In addition to maintaining the 3D semantic map 164, the cloud system 150, in some embodiments, also performs task assignment 20 and motion planning 30 on behalf of the mobile robot 120. Specifically, based on the captured images and / or the 3D semantic map 164, the cloud system 150 uses semantic segmentation to intelligently identify 70 and select 80 tasks to be performed by the mobile robot 120 in the environment. Specifically, the cloud system 150 processes the captured images and / or the 3D semantic map to determine the specific tasks that need to be performed. This could include detecting areas that need cleaning, detecting objects that need manipulation, defining an image target task, or other types of interactions. Once the tasks are identified, the system 10 prioritizes them based on predefined criteria, such as...Urgency, proximity, or other application-relevant factors. This step ensures that the robot focuses first on the most important tasks and completes them efficiently. In some embodiments of System 10, System 10 also suggests tasks to the user, who then makes the final decision. Mobile robot
[0018] Fig. Figure 2A shows an embodiment of the mobile robot 120. In the illustrated embodiment, the mobile robot 120 has, for example, a controller 121, one or more sensors 126, one or more actuators 128, and at least one network communication module 130. It is understood that the illustrated embodiment of the mobile robot 120 is only an embodiment and is merely representative of various types or configurations or arrangements of mobile robots that autonomously navigate or move within an environment to perform a task.
[0019] The controller 121 comprises a processor 122 and a memory 124, which stores an operating sequence 132 for operating the mobile robot. The processor 122 is configured to execute instructions for operating the mobile robot 120 to enable the features, functionalities, properties, and / or the like described herein. For this purpose, the processor 122 is functionally connected to the memory 124, the one or more sensors 126, and the one or more actuators 128. The processor 122 generally comprises one or more processors that can operate in parallel or otherwise coordinate with one another. It is known to those skilled in the art that a "processor" comprises any hardware system, hardware mechanism, or hardware component that processes data, signals, or other information.Therefore, the processor 122 can comprise a system with a central processing unit, graphics processing units, multiple processing units, dedicated circuits for achieving functionality, programmable logic, or other processing systems.
[0020] Memory 124 is configured to store data and program instructions which, when executed by processor 122, enable the mobile robot 120 to perform various operating sequences or operations described herein. Memory 124 can be any type of device capable of storing information accessible to processor 122, such as a memory card, ROM, RAM, hard disks, floppy disks, flash memory, or any other computer-readable medium serving as a data storage device, as is known to those skilled in the art. As explained in more detail below, processor 122 is configured to execute program instructions of the operating sequence 132 stored in memory 124 to navigate the environment and perform a task, such as cleaning a floor surface in the environment.In at least one embodiment, the operational sequence 132 uses the 3D Semantic Map 164, which virtually represents the environment, to assist in performing the task.
[0021] The one or more sensors 126 can comprise a variety of different sensors. In some embodiments, the sensors 126 include sensors configured to measure one or more accelerations, rotational speeds, and / or orientations of the mobile robot 120. In one embodiment, the sensors 126 include one or more accelerometers configured to measure linear accelerations of the mobile robot 120 along one or more axes (e.g., roll, pitch, and yaw axes), or one or more gyroscopes configured to measure rotational speeds of the mobile robot 120 along one or more axes (e.g., roll, pitch, and yaw axes), and / or an inertial measurement unit configured to measure all of the above quantities.
[0022] In at least some embodiments, the sensors 126 comprise a light sensor (e.g., LiDAR or another time-of-flight or structured light sensor) configured to emit measurement light (e.g., a laser) and receive the measurement light after it has been reflected in the environment. In time-of-flight-based embodiments, the processor 122 is configured to calculate travel times and / or return times for the measurement light. Based on the calculated travel times and / or return times, the processor 122 can, for example, generate the map data of the 3D semantic map, for example, in the form of a point cloud. In structured light-based embodiments, the processor 122 applies an algorithm to extract a 3D profile of surfaces onto which the structured light is projected (e.g., based on a stripe pattern generated on a surface).
[0023] In some embodiments, the sensors 126 include, as an alternative to or in addition to the light sensor, one or more cameras configured to capture a multitude of images of the environment as the mobile robot 120 navigates through it. The camera(s) generate image frames of the environment, each of which has a two-dimensional arrangement of pixels. Each pixel has corresponding photometric information (color, intensity, and / or brightness). In some embodiments, the camera(s) is / are configured to generate RGB-D images in which each pixel has corresponding photometric and geometric information (depth and / or distance).In such embodiments, the camera(s) can take the form of an RGB camera operating in conjunction with a LiDAR or IR sensor, in particular a LiDAR or IR camera configured to provide both photometric and geometric information. The LiDAR or IR camera can be separate from the RGB camera or directly integrated into it. Alternatively or additionally, the camera can have two RGB cameras configured to capture stereoscopic images from which depth and / or distance information can be derived. Based on RGB-D images captured while the mobile robot 120 navigates its environment, the map data for the 3D semantic map can be derived, for example, using visual and / or visual-inertial odometry methods such as simultaneous localization and mapping (SLAM) techniques.
[0024] The one or more actuators 128 comprise at least motors of a locomotion system that, for example, drive a set of wheels to cause the mobile robot 120 to move through the environment to perform the task. In addition, in some embodiments, the one or more actuators 128 comprise a vacuum suction system configured to vacuum a floor surface while the mobile robot 120 navigates the environment. Mobile robots 120 that perform other tasks in the environment may, of course, comprise different types of actuators 128 suitable for other tasks.
[0025] The Network Communications Module 130 may include one or more transceivers, modems, processors, memory, oscillators, antennas, or other hardware typically contained in a communications module to enable communication with various other devices, including at least the Cloud System 150 and / or the Mobile Device 170. Specifically, the Network Communications Module 130 generally includes a Wi-Fi module configured to enable communication with a Wi-Fi network and / or a Wi-Fi router (not shown). Additionally, the Network Communications Module 130 may include a Bluetooth® module configured to enable communication with the Mobile Device 170. Finally, the Network Communications Module 130 may include one or more cellular modems configured to communicate with wireless telephone networks.
[0026] The mobile robot 120 may also include a battery or other power source (not shown) configured to power the various components within the mobile robot 120. In one embodiment, the battery of the mobile robot 120 is a rechargeable battery configured to be recharged when the mobile robot 120 is connected to a base station configured for use with the mobile robot 120. Cloud system
[0027] As mentioned above, in at least some embodiments the mobile robot 120 can communicate with a cloud system 150. In particular, the cloud system 150 can be configured to store and maintain the 3D semantic map 164 and to perform intelligent task assignment and motion planning on behalf of the mobile robot 120. However, it should be noted that in some embodiments these functions can also be performed locally using the mobile robot 120 itself or by the mobile device 170.
[0028] Fig. Figure 2B shows an embodiment of the cloud system 150. The cloud system 150 comprises one or more cloud servers 152 and one or more cloud storage devices 162. The cloud servers 152 can include servers configured to perform a variety of functions for the cloud storage backend 150, including web servers or application servers, depending on the functions provided by the cloud system 150, but at least one or more database servers configured to manage map data received by the mobile robot 120 and stored in the cloud storage devices 162. Each cloud server 152, for example, includes a processor 154, memory 156, a user interface 158, and a network communication module 160.It is understood that the illustrated embodiment of the Cloud Server 152 is only one embodiment of a Cloud Server 152 and is merely representative of various types or configurations / arrangements of a personal computer, server or other data processing system that operates or is capable of operation in the manner described herein.
[0029] The processor 154 is configured to execute instructions to operate the cloud server 152 in order to enable the features, functionalities, properties, and / or the like described herein. For this purpose, the processor 154 is functionally connected to the memory 156, the user interface 158, and the network communication module 160. The processor 154 generally comprises one or more processors that can operate in parallel or otherwise coordinated with one another. It is known to those skilled in the art that a "processor" includes any hardware system, hardware mechanism, or hardware component that processes data, signals, or other information. Accordingly, the processor 154 may comprise a system with a central processing unit, graphics processing units, multiple processing units, dedicated circuits for achieving functionality, programmable logic, or other processing systems.
[0030] The cloud storage device 162 is configured to store map data received by the mobile robot 120. The cloud storage device 162 can be any type of long-term non-volatile storage device capable of storing information accessible to the processor 154, such as hard disks, solid-state drives, or any other computer-readable storage medium recognized by a person skilled in the art. Likewise, the memory 156 is configured to store program instructions which, when executed by the processor 154, enable the cloud server 152 to perform various operations described herein, including managing the map data stored in the cloud storage devices 162.The memory 156 can be any type of device or combination of devices that can store information which the processor 154 can access, such as memory cards, ROM, RAM, hard disks, floppy disks, flash memory or various other computer-readable media known to a specialist.
[0031] The cloud server 152 can be operated locally or remotely by an administrator. To enable local operation, the cloud server 152 can include the user interface 158. In at least one embodiment, the user interface 158 can suitably include an LCD display screen or the like, a mouse or other pointing device, a keyboard or other keypad or control panel, speakers, and a microphone, as is known to a person skilled in the art. Alternatively, in some embodiments, an administrator can operate the cloud server 152 remotely from another computing device that communicates with it via the network communication module 160 and has an analog user interface.
[0032] The network communication module 160 provides an interface that enables communication with any devices, including at least the mobile robot 120 and the mobile device 170. In particular, the network communication module 160 can include a LAN (Local Area Network) port that enables communication with any local computers located in the same or a nearby facility. Generally, the cloud server 152 communicates with remote computers over the internet via a separate modem and / or router on the local network. Alternatively, the network communication module 160 can also include a WAN (Wide Area Network) port that enables communication over the internet. In one embodiment, the network communication module 160 is equipped with a Wi-Fi transceiver or other wireless communication device.Therefore, it follows that communication with the Cloud Server 152 can take place via wired or wireless communication. This communication can be carried out using any of the various known communication protocols.
[0033] The cloud server 152 is configured to securely store and maintain the 3D semantic map 164 for the mobile robot 120 and to grant the mobile robot 120 and the mobile device 170 access to the 3D semantic map 164. In one embodiment, the cloud system 150 authenticates connections using a digitally signed key and various known network and computer security protocols. Furthermore, in some embodiments, the memory 156 stores program instructions for a task assignment program 168 to perform intelligent task assignment for the mobile robot 120 based on images received from the mobile device 170 and on the basis of the 3D semantic map 164. Mobile device
[0034] As mentioned above, the mobile robot 120 communicates with a mobile device 170, through which a user can manage and operate the mobile robot 120. The mobile device 170 can take the form of any personal electronic device, such as a mobile phone or a tablet computer. However, it should be noted that some of these functions can also be performed locally using the mobile robot 120 itself, for example, using a user interface integrated into the mobile robot 120.
[0035] Fig. Figure 2C shows an embodiment of the mobile device 170. The mobile device 170 comprises a processor 172, a memory 174, a display screen 176, at least one network communication module 178, and at least one camera 182. The processor 172 is configured to execute instructions for operating the mobile device 170 to enable the features, functions, characteristics, and / or the like described herein. For this purpose, the processor 172 is functionally connected to the memory 174, the display screen 176, and the network communication module 178. The processor 172 generally comprises one or more processors that can work together in parallel or otherwise. It is known to a person skilled in the art that a "processor" includes any hardware system, hardware mechanism, or hardware component that processes data, signals, or other information.Therefore, the processor 172 can comprise a system with a central processing unit, graphics processing units, multiple processing units, dedicated circuits for achieving functionality, programmable logic, or other processing systems.
[0036] Memory 174 is configured to store data and program instructions which, when executed by Processor 172, enable the mobile device 170 to perform various operations described herein. Memory 174 can be any type of device capable of storing information accessible to Processor 172, such as a memory card, ROM, RAM, hard disks, floppy disks, flash memory, or any other computer-readable storage medium used as a data storage device, as known to the average person skilled in the art. Among other things, Memory 174 stores an application 180 of the mobile robot. As further explained below, Processor 172 is configured to execute program instructions of the application 180 of the mobile robot in order to operate and configure the mobile robot 170.
[0037] The display screen 176 can comprise any of several known display types, such as LCD or OLED screens. In some embodiments, the display screen 176 can comprise touchscreens configured to receive touch input from a user. Alternatively or additionally, the mobile device 170 can include additional user interfaces, such as buttons, switches, a keyboard or other keypad, speakers, and a microphone.
[0038] The Network Communication Module 178 may include one or more transceivers, modems, processors, memory, oscillators, antennas, or other hardware typically included in a communication module to enable communication with various other devices, including at least the Cloud System 150 and / or the Mobile Robot 120. Specifically, the Network Communication Module 178 generally includes a Wi-Fi module configured to enable communication with a Wi-Fi network and / or a Wi-Fi router (not shown). Furthermore, the Network Communication Module 178 may include a Bluetooth® module (not shown) configured to enable communication with the Mobile Robot 120. Finally, the Network Communication Module 178 may include one or more cellular modems configured to communicate with wireless telephone networks.
[0039] The camera 182 is configured to capture a multitude of images of the environment as the mobile device 170 is moved through the environment by the user. The camera 182 is configured to generate image frames of the environment, each of which has a two-dimensional arrangement of pixels. Each pixel has at least corresponding photometric information (color, intensity, and / or brightness). In some embodiments, the camera 182 operates to generate RGB-D images in which each pixel has corresponding photometric information and geometric information (depth and / or distance). In such embodiments, the camera 182 can, for example, take the form of an RGB camera operating in conjunction with a LiDAR or IR sensor, in particular a LiDAR camera or IR camera, configured to provide both photometric and geometric information.The LiDAR or IR camera can be separate from the RGB camera or integrated directly into the RGB camera. Alternatively or additionally, the camera 129 can have two RGB cameras configured to capture stereoscopic images from which depth and / or distance information can be derived.
[0040] In some embodiments, the mobile device 170 may further comprise a plurality of sensors (not shown). In some embodiments, the sensors include sensors configured to measure one or more accelerations and / or rotational velocities of the mobile device 170. In one embodiment, the sensors include one or more accelerometers configured to measure linear accelerations of the mobile device 170 along one or more axes (e.g., roll, pitch, and yaw axes), and / or one or more gyroscopes configured to measure rotational velocities of the mobile device 170 along one or more axes (e.g., roll, pitch, and yaw axes). In some embodiments, the sensors include a GPS module configured to communicate with GPS satellites to generate GPS position data.The GPS module includes, for example, a GPS receiver, an amplifier and an antenna, as well as any other processors, memory, oscillators or other hardware that are typically included in a GPS module. Methods for providing intelligent and intuitive operating procedures for a mobile robot
[0041] The following describes various methods and processes for providing intelligent and intuitive operation of a mobile robot using a graphical user interface for augmented reality. In these descriptions, statements that a method, processor, and / or system performs a task or function refer to a controller or processor (e.g., processor 154 of cloud server 152, processor 122 of mobile robot 120, or processor 172 of mobile device 170) that executes programmed instructions stored in non-volatile, computer-readable storage media (e.g.,The data are stored in memory 156 of the cloud server 152, memory 124 of the mobile robot 120, or memory 174 of the mobile device 170, which are functionally connected to the controller or processor to manipulate data or to operate or control one or more components in the cloud server 152, the mobile robot 120, or the mobile device 170 in order to perform the task or function. Furthermore, the steps of the procedures can be performed in any possible chronological order, regardless of the order shown in the figures or the order in which the steps are described.
[0042] Furthermore, various graphical AR user interfaces for operating the mobile device 170 are described. In many cases, the graphical AR user interfaces comprise graphical elements overlaid on real-time images / videos captured by the camera 182. To provide these graphical AR user interfaces, the processor 172 executes instructions from the mobile robot's application 180 to render these graphical elements and controls the display screen 176 to overlay the graphical elements onto the real-time images / videos of the outside world. In many cases, the graphical elements are rendered in a position that depends on positional or orientation information received from any suitable combination of sensors (not shown) and the camera 182, thus simulating the presence of the graphical elements in the real-world environment.
[0043] Furthermore, various user interactions with the graphical AR user interfaces and their interactive graphical elements are described. To provide these user interactions, the processor 172 can render interactive graphical elements in the graphical AR user interface, receive user input from the user, for example by touching the display screen 176 or by manipulating another user interface of the mobile device 170, and execute instructions from the mobile robot's application 180 to perform certain operations in response to the user input.
[0044] Finally, various forms of motion tracking can be used to track the spatial positions and movements of the user or objects in the environment. To provide this tracking of spatial positions and movements, the processor 172 executes instructions from the mobile robot's application 180 to receive and process sensor data from any suitable combination of the sensors (not shown) and the camera 182 of the mobile device 170, and can optionally use visual and / or visual-inertial odometry methods such as simultaneous localization and mapping (SLAM) techniques.
[0045] Fig.Figure 3 shows a flowchart for a method 200 for operating a mobile robot using a graphical AR user interface of a mobile device. The method advantageously utilizes a 3D semantic map of the environment maintained by a cloud system. The graphical AR user interface advantageously allows the user to capture images of the environment and visualize the semantic information of the 3D semantic map in real time. The cloud system is advantageously designed to integrate new semantic information received via the graphical AR user interface into the 3D semantic map. The graphical AR user interface is also intended for assigning tasks to mobile robots, enabling the mobile robot to perform tasks based on user input via the graphical AR user interface.Finally, the cloud system uses semantic information from the 3D Semantic Map to automatically suggest tasks from images captured by the mobile device.
[0046] Method 200 begins by displaying a graphical user interface for augmented reality on a mobile device (block 205). Specifically, the processor 172 of the mobile device 170 is configured to execute program instructions from the application 180 of the mobile robot to provide a graphical AR user interface on the display screen 176. To this end, the processor 172 controls the camera 182 to capture a video of the environment, which is displayed in real time on the display screen 176 as the user moves the mobile device 170 to point the camera 182 at different areas within the environment. The graphical AR user interface comprises a variety of graphical elements that are overlaid or inserted into the real-time video of the environment and are designed to present information to the user on how to operate the mobile robot 120 in a more intelligent and intuitive way.Furthermore, at least some of the graphical elements in the graphical AR user interface may include interactive graphical elements, such as virtual buttons that can be selected by the user to perform various operations or processes, such as setting up, configuring and operating the mobile robot 120.
[0047] The graphical AR user interface allows the user to capture an image of the environment, for example, by interacting with a virtual image capture button within the graphical AR user interface or by manipulating a physical user interface of the mobile device 170. Specifically, in response to user input via a user interface of the mobile device 170, the processor 172 controls the camera 182 to capture an image of the environment. In at least some embodiments, in addition to capturing the image, the processor 172 also captures sensor data from other sensors of the mobile device 170, such as inertial data or audio data. The processor 172 stores the image and all other captured data in the memory 174 and / or controls the network communication module 178 to transmit the image and all other captured data to the cloud system 150.As explained in more detail below, such images are processed by the mobile robot system 10, in particular by the cloud system 150, to provide various functions including updating the 3D semantic map data 164 and intelligently assigning tasks to be performed by the mobile robot 120 in the environment.
[0048] Furthermore, the graphical AR user interface includes graphical representations of the semantic information from the 3D Semantic Map 164. For this purpose, the processor 172 controls the network communication module 178 to receive at least a portion of the 3D Semantic Map 164. Based on real-time images of the environment captured by the camera 182, as well as sensor data acquired by other sensors, such as inertial data, acceleration data, GPS position data, and the like, the processor 172 determines the position and orientation of the mobile device 170, for example, using SLAM techniques. Based on the determined position and orientation of the mobile device 170, the processor 172 determines a correspondence between surfaces and objects in the view of the camera 182 and surfaces and objects in the 3D Semantic Map 164.Based on this agreement, the processor 172 controls the display 176 to show graphical representations of the semantic information from the 3D semantic map 164, which are overlaid on the real-time video of the environment in the graphical AR user interface. The graphical representations can take various forms, such as floating text, highlighting of objects or surfaces, or displaying area boundaries, designed to provide the user with various information for operating the mobile robot 120 in a more intelligent and intuitive way.
[0049] In addition to visualizing semantic information currently contained in the 3D Semantic Map 164, the graphical AR user interface allows the user to input new semantic information related to the environment. In one embodiment, the user can input semantic information associated with a portion of the 3D Semantic Map 164 or a portion of an image captured by the camera 182. In response to user input received from the user via a user interface on the mobile device 170, the processor 172 stores additional user-provided semantic information and / or controls the network communication module 178 to transmit the user-provided semantic information to the cloud system 150. For example, the user can select an object in the environment and specify a text label indicating the object's name or type.In another example, the user can select existing semantic information, graphically represented in the AR graphical interface, and edit the semantic information. In yet another example, the user can interact with the AR graphical user interface to define a spatial area within the environment that the mobile robot 120 will use when performing a task, for example, by circling the spatial area of interest with their finger on the screen 176 of the mobile device 170. In one example, the user-defined spatial area is an area to be cleaned. In another example, the user-defined spatial area is a restricted area that the mobile robot 120 should not enter when performing tasks.
[0050] Finally, the graphical AR user interface allows the user to directly issue commands to the mobile robot 120. Specifically, in response to user input received from the user via the user interface of the mobile device 170, the processor 172 determines a user-defined task to be performed by the mobile robot 120 and controls the network communication module 178 to transmit a command to the mobile robot 120, which is configured to cause the mobile robot 120 to perform the user-defined task. For example, the user can select an object in the environment and instruct the mobile robot 120 to clean around the selected object or to avoid it.In another example, the user can define a trajectory through the environment, for instance, by swiping a finger on the screen 176 of the mobile device 170 along the trajectory that the mobile robot 120 is to follow. In this way, the user can instruct the mobile robot 120 to navigate along a user-defined trajectory, for example, to perform a task. In yet another example, the user can interact with the graphical AR user interface to define an area within the environment and instruct the mobile robot 120 to perform a task, such as cleaning, within the defined area.
[0051] Method 200 continues with the receipt of an image of an environment captured by a camera of the mobile device (Block 210). In particular, the processor 154 of the cloud system 150 controls the network communication module 160 to receive an image of the environment from the mobile device, which was captured by the camera 182 in the manner described above. In some embodiments, the processor 154 also receives additional sensor data from other sensors of the mobile device 170, such as inertial data or audio data. In some embodiments, the processor 154 also receives localization data from the mobile device 170, such as the position and orientation of the mobile device 170 within the environment at the time the image was captured.
[0052] The procedure continues next with the execution of semantic segmentation of the image (Block 215). In particular, the processor 154 of the cloud system 150 performs semantic segmentation on the received image to determine a variety of semantic labels that are assigned to specific areas or pixels within the received image. In at least some embodiments, the processor 154 further processes the image and / or the additional sensor data from other sensors prior to semantic segmentation, for example, to filter noise from the data and extract relevant features.
[0053] In some embodiments, the processor 154 executes a machine learning-based semantic segmentation model, such as Mask-RCNN, which extracts semantic information from each pixel in the received image. Such a machine learning-based semantic segmentation model is pre-trained using a large dataset of images of various indoor environments. In one embodiment, the semantic segmentation model is configured to output semantic labels that identify and characterize different objects and states within the image. For example, pixels in the image corresponding to the surface of a floor can be labeled as belonging to a floor. Similarly, pixels in the image corresponding to a piece of furniture, such as a sofa, can be labeled as belonging to a piece of furniture.
[0054] Method 200 further includes, in addition to receiving and processing the image, receiving semantic information provided by a user via the graphical user interface for augmented reality (Block 220). Specifically, the processor 154 of the cloud system 150 controls the network communication module 160 to receive semantic information provided by the user from the mobile device 170, either in conjunction with a portion of the 3D semantic map 164 or in conjunction with a portion of an image also received by the mobile device 170. As explained above, such semantic information is provided directly by the user via the graphical AR user interface of the mobile device 170.
[0055] Procedure 200 continues with the integration of new information into the 3D semantic map (Block 225). Specifically, the cloud system 150 stores the 3D semantic map 164 in the cloud storage devices 162. As mentioned above, the 3D semantic map 164 is derived from various sensor data from multiple sources, including at least sensor data acquired by the sensors 126 of the mobile robot 120, as well as images and other sensor data acquired by the camera 182 and other sensors of the mobile device 170. The 3D semantic map 164 includes map data, for example, in the form of a point cloud, a surface map, and / or a polygon mesh. The map data includes both photometric and three-dimensional geometric information representing the environment in which the mobile robot 120 is to operate. In addition to the map data, the 3D semantic map 164 includes semantic information.In some embodiments, the semantic information includes semantic labels that identify or characterize objects in the environment. Furthermore, in some embodiments, the semantic information includes semantic labels that identify or characterize specific spatial areas within the environment, such as room labels, cleaning areas, restricted areas, and the like. Finally, in some embodiments, the semantic information includes semantic labels that identify or characterize specific surfaces within the environment, such as labels that distinguish floors from walls, objects, and other obstacles, or that differentiate carpeted floors from hardwood floors, and the like.
[0056] When new map data and semantic information are received, the Cloud System 150 integrates this newly received data into the 3D Semantic Map 164. This integration helps the Mobile Robot 120 better understand the context and relationships between different elements in its environment. Furthermore, users generally capture images of their surroundings from a different perspective than the Mobile Robot 120, enabling more accurate and detailed mapping of the environment. Thus, over time, the 3D Semantic Map 164 begins to incorporate map data and semantic information that might be difficult for the Mobile Robot 120 to access on its own.
[0057] In particular, if an image of the environment is received from the mobile device 170, the processor 154 updates the map data of the 3D semantic map 164 based on the photometric and / or geometric information in the image. Furthermore, the processor 154 updates the semantic information of the 3D semantic map 164 based on the semantic segmentation labels derived from the image in block 215. Finally, if the mobile device 170 receives user-provided semantic information, the processor 154 updates the semantic information of the 3D semantic map 164 to include the user-provided semantic information.
[0058] In some cases, in order to perform these updates to the 3D semantic map 164, the processor 154 needs to know the position and orientation within the environment in which a received image was captured. In some embodiments, the mobile device 170 can provide the position and orientation within the environment in which the image was captured. In alternative embodiments, the processor 154 determines the position and orientation within the environment in which the image was captured based on the image itself. Specifically, to determine the position and orientation in which the image was captured, the processor 154 compares photometric and / or geometric information in the image with photometric and / or geometric information from the map data of the 3D semantic map 164.
[0059] In addition to integrating new information into the 3D semantic map, Method 200 also includes intelligent task assignment to the mobile robot. For this purpose, Method 200 continues with an assessment of the cleanliness state of the environment within the image (Block 230). Specifically, Processor 154 is configured to determine the cleanliness state of at least a portion of the environment within the image based on the received image. For example, the received image may include a portion of the environment that contains a table. Processor 154 may determine that dirt and debris have accumulated in areas below and around the table, resulting in a relatively poor cleanliness state. In some embodiments, Processor 154 determines the cleanliness state based on the semantic segmentation of the image.In some embodiments, the processor 154 can execute an additional machine learning model to process the image and / or the image's semantic segmentation labels to determine the cleanliness state of at least one part of the environment within the image. In some embodiments, the processor 154 updates the semantic information of the 3D semantic map 164 to update or reintegrate the cleanliness state of the at least one part of the environment within the image.
[0060] The procedure continues with the generation of a recommended task to be performed by a mobile robot in the environment (Block 235). Specifically, the processor 154 determines at least one recommended task to be performed by the mobile robot 120 in the environment based on the 3D semantic map 164, the received image, the semantic segmentation of the received image, and / or the cleanliness state of at least a part of the environment within the received image. To identify recommended tasks for any given image, the processor 154, in some embodiments, uses a combination of computer vision and machine learning techniques, in particular deep learning models such as convolutional neural networks and imitation learning.These models can be trained to recognize a variety of scenarios, objects, and environmental conditions in order to recommend the appropriate tasks.
[0061] In some embodiments, the processor 154 determines the at least one recommended task based on the cleanliness state of at least one part of the environment within the image. Specifically, in response to a particular part of the environment within the image being determined to have a cleanliness state below a threshold or otherwise not considered clean, the processor 154 determines a recommended task in which the mobile robot 120 navigates to and cleans that particular part of the environment. In one embodiment, the cleanliness state and the threshold state are characterized by the number of dirty objects per unit area or unit volume in the 3D semantic map 164.In another embodiment, the cleanliness state and the threshold state are characterized by a volume of dirty objects in a specific area relative to the total volume of the area.
[0062] In some embodiments, the processor 154 determines the at least one recommended task based on the semantic segmentation labels determined from the image, as well as on similar semantic information already stored in the 3D semantic map 164. In particular, in one embodiment, the processor 154 determines a recommended task in which the mobile robot 120 navigates to an area of the environment located near a specific type of object. For example, areas under and around tables may tend to accumulate dirt and debris more quickly than other areas. Consequently, the processor 154 can determine that cleaning tasks should be performed in an area under and around a table.
[0063] In some embodiments, the processor 154 can determine a recommended task in which the mobile robot 120 navigates to a specific area to acquire more sensor data in order to improve or update the 3D semantic map 164. For example, based on a received image, the processor 154 determines whether the existing information in the 3D semantic map 164 is incomplete or outdated and determines a recommended task to acquire further data in order to improve or update this information. For example, in one embodiment, the processor 154 determines that existing information in the 3D semantic map 164 is incomplete in response to the fact that points or areas are missing in a spatial volume of at least a predetermined size in the 3D semantic map 164.Furthermore, in one embodiment, the processor 154 determines that existing information in the 3D Semantic Map 164 is outdated in response to the fact that a spatial volume in the 3D Semantic Map 164 of at least a predetermined size has not been updated for at least a predetermined time interval.
[0064] In some embodiments, the processor 154 can similarly determine recommended tasks to be performed by the user before the mobile robot 120 performs any tasks. In particular, in one embodiment, based on the semantic segmentation of the image, the processor 154 can identify that a specific object should be removed, relocated, or moved from the environment before the mobile robot 120 performs a task in a particular area. In another embodiment, the processor 154 determines a recommended task in which the user is prompted to capture images of a specific area in the environment using the mobile device 170 in order to improve or update the 3D semantic map 164.
[0065] In some embodiments, once the recommended tasks have been identified, the processor 154 prioritizes them based on predefined criteria, such as urgency, proximity, or other factors relevant to the application. Specifically, the processor 154 assigns a priority to each specific or recommended task, which can take the form of a classification or a numerical value indicating the task's priority level. This step ensures that the robot focuses first on the most important tasks and performs them efficiently.
[0066] The procedure continues with a prompt to the user regarding the recommended task (Block 240). Specifically, the processor 154 controls the network communication module 160 to transmit the recommended tasks to the mobile device 170. The processor 172 of the mobile device 170 controls the display screen 176 to show a visual prompt to the user regarding the recommended tasks. In some embodiments, the prompt is an interactive prompt through which the user can authorize the execution of one or more of the recommended tasks. Furthermore, in some embodiments, the prompt is a suggestion that the user perform a task in the environment, such as removing or moving objects on the floor, before the mobile robot 120 performs a recommended task, such as cleaning the floor.
[0067] In the case of an interactive prompt, the processor 172 receives a user selection regarding the prompt via a user interface of the mobile device 170 and controls the network communication module 178 of the mobile device 170 to transmit the user selection to the cloud system 150 and / or the mobile robot 120. In some embodiments, the user selection is simply an affirmation or rejection indicating whether one or more recommended tasks should be performed. However, in other embodiments, the interactive prompt allows the user to change the priority of one or more recommended tasks or to add additional tasks to be performed, and the user selection includes such changes or additions.
[0068] If the user decides to proceed with one or more of the recommended tasks (Block 245), Procedure 200 continues by generating a motion plan for the mobile robot to perform the task (Block 250). Specifically, in response to a user selection received from the mobile device 170 indicating that a particular task is to be performed, Processor 154 determines a trajectory through the environment along which the mobile robot 120 is to navigate to perform the one or more tasks. Processor 154 determines the trajectory based on the 3D Semantic Map 164, taking into account the priority assigned to the recommended tasks if more than one of the recommended tasks is to be performed.The processor 154 generates the trajectory in a manner that takes into account the kinematics of the mobile robot 120, the environmental constraints, and the desired or prioritized task execution sequence. In some embodiments, the processor 154 determines the trajectory using a motion planning algorithm, such as A* or PRM, to generate an efficient and collision-free trajectory.
[0069] The procedure continues next with operating the mobile robot to perform the task and monitoring its progress (Block 255). Specifically, the processor 154 controls the network communication module 160 to transmit a command message to the mobile robot 120, configured to cause the mobile robot 120 to perform the one or more selected tasks authorized by the user. In at least one embodiment, the command message includes the specified trajectory along which the mobile robot 120 is to navigate through the environment to perform the one or more selected tasks. In some embodiments, the command message may include the latest version of the 3D semantic map 164 or otherwise cause the mobile robot 120 to retrieve the latest version of the 3D semantic map 164.
[0070] In response to receiving the command message, the controller 121 controls the actuators 128 of the mobile robot to cause the mobile robot to navigate through the environment along the specified trajectory and perform the one or more selected tasks. For this purpose, the controller 121 controls the sensors 126 to detect the positions of walls, objects, or other obstacles in the environment in order to locate the position of the mobile robot 120 within the environment and thus within the 3D semantic map 164. Once the mobile robot 120 is located, it uses the 3D semantic map 164 and the specified trajectory to perform the one or more selected tasks.
[0071] In at least some embodiments, the controller 121 controls the network communication module 130 to continuously and / or periodically transmit the current position of the mobile robot 120 within the environment and / or the current progress of each of the one or more selected tasks to the mobile device 170 or to the cloud system 150.
[0072] Finally, the process continues with the generation of AR graphic elements that visualize the 3D semantic map and / or the task being performed (Block 260). Specifically, the graphical AR user interface of the mobile device 170, as explained above, includes graphical representations of the semantic information in the 3D semantic map 164. Accordingly, as the 3D semantic map 164 is updated over time, the processor 172 controls the network communication module 178 to receive updated portions of the 3D semantic map 164. Based on the updated semantic information in the updated 3D semantic map, the processor 172 generates updated or new graphic elements that represent the new semantic information. The processor 172 controls the display screen 176 to display the updated or new graphic elements within the graphical AR user interface as needed, in the manner described above.
[0073] In some embodiments, the processor 172 generates graphical elements while the mobile robot 120 is performing a task in the environment. These elements represent the progress of the task or the motion planning (e.g., the trajectory) used or being used by the mobile robot 120 to perform the task. The processor 172 controls the display screen 176 to show these graphical elements within the graphical AR user interface as required. This allows the user to monitor the progress of tasks currently being performed or scheduled to be performed by the mobile robot 120.
[0074] Embodiments within the scope of the disclosure may also include non-volatile computer-readable storage media or machine-readable media on which computer-executable instructions (also referred to as program instructions) or data structures are stored or are stored. Such non-volatile computer-readable storage media or machine-readable media may be any available media accessible to a general-purpose or specialized computer. By way of example, and not as a limitation, such non-volatile computer-readable storage media or machine-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to receive, transmit, or store desired program code means in the form of computer-executable instructions or data structures.Combinations of the above-mentioned elements should also fall within the scope of non-volatile computer-readable storage media or machine-readable media.
[0075] Computer-executable instructions include, for example, instructions and data that cause a general-purpose computer, a specialized computer, or a specialized processing device to perform a particular function or group of functions. Computer-executable instructions also include program modules that are executed by computers in standalone or networked environments. In general, program modules include routines, programs, objects, components, and data structures, etc., that perform specific tasks or implement or realize certain abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of the program code means for performing steps of the procedures disclosed herein.The specific sequence of such executable instructions or associated data structures represents examples of corresponding actions for implementing / realizing the functions described in such steps.
[0076] Although the disclosure has been illustrated and described in detail in the drawings and the preceding description, it should nevertheless be regarded as illustrative and not limiting. It is understood that only preferred embodiments have been presented and that all changes, modifications, and further applications that are within the scope of the disclosure are to be protected.
Claims
[1] Method for operating a mobile robot comprising the method: Storing a three-dimensional semantic map of the environment in a memory, wherein the three-dimensional semantic map includes semantic information relating to the environment, and wherein the three-dimensional semantic map is generated at least partially based on sensor data acquired by the mobile robot while navigating the environment; Receiving an image of an environment captured by a camera of a personal electronic device; Determine at least one initial task to be performed by a mobile robot in the environment, at least partially based on the image and the three-dimensional semantic map; and Operating the mobile robot to perform at least one initial task in the environment. [2] Method according to claim 1, further comprising: Displaying a graphical user interface for augmented reality on a display of the personal electronic device, comprising a variety of graphical elements superimposed on the environment, wherein the variety of graphical elements includes graphical representations of the semantic information of the three-dimensional semantic map. [3] Method according to claim 2, wherein: The semantic information of the three-dimensional semantic map includes at least one semantic label of an object in the environment; and The multitude of graphical elements includes a graphical representation of at least one semantic label that is superimposed on the object in the environment. [4] Method according to claim 2, wherein: The semantic information of the three-dimensional semantic map must include at least one semantic label for a spatial area in the environment; and The multitude of graphic elements includes a graphic representation of at least one semantic label that is superimposed on the spatial area in the environment. [5] Method according to claim 4, wherein the spatial area is a user-defined spatial boundary used by the mobile robot to perform the at least one first task. [6] Method according to claim 4, wherein the spatial area is a surface of a floor corresponding to a specific room of the environment. [7] Method according to claim 4, wherein the room area is a surface of a floor which is to be cleaned during the at least one first task. [8] Method according to claim 2, wherein: The multitude of graphic elements includes a graphic representation of a trajectory of the mobile robot, along which the mobile robot navigates through the environment to perform at least one initial task. [9] Method according to claim 2, further comprising: Capturing an image of the environment with a camera of the personal electronic device in response to user input received via the graphical user interface for augmented reality. [10] Method according to claim 2, further comprising: Determine at least one second task to be performed by the mobile robot based on user input received via the graphical user interface for augmented reality. [11] Method according to claim 1, further comprising: Performing a semantic segmentation of the image; and Updating the three-dimensional semantic map of the environment based on the semantic segmentation of the image. [12] Method according to claim 1, wherein the updating of the three-dimensional semantic map is performed by a server that is remote from the mobile robot. [13] Method according to claim 1, further comprising: Receiving semantic information from a user regarding the environment within the image via a user interface of the personal electronic device; and Updating the three-dimensional semantic map of the environment based on semantic information provided by the user. [14] The method of claim 1, further comprising operating the mobile robot to perform the at least one first task: Determining a trajectory through the environment along which the mobile robot should navigate to perform at least one initial task, based on the three-dimensional semantic map of the environment; and Operating the mobile robot to navigate along the trajectory through the environment in order to perform at least one initial task in the environment. [15] The method of claim 1, further comprising determining the at least one first problem: Determining the cleanliness level of at least part of the environment within the image; and Determining the at least one initial task to be performed by the mobile robot in the environment, based on the cleanliness of at least one part of the environment. [16] Method according to claim 1, further comprising: Determine at least one third task to be performed by a user before the mobile robot performs at least one first task; Displaying a prompt to the user on a display of the personal electronic device, including at least one third task to be performed by the user. [17] Method according to claim 1, further comprising: Displaying a prompt to a user on a display of the personal electronic device, including at least one initial task to be performed by the mobile robot; and Receiving a user selection indicating whether at least one initial task is to be performed, via a user interface of the personal electronic device; wherein the mobile robot is operated to perform at least one initial task in the environment in response to the user selection indicating that at least one initial task is to be performed. [18] Method according to claim 1, wherein the at least one first object comprises cleaning a surface of a floor in a defined spatial area of the environment. [19] Method for operating a mobile robot comprising the method: Storing a three-dimensional semantic map of the environment in a memory, wherein the three-dimensional semantic map includes semantic information relating to the environment, and wherein the three-dimensional semantic map is generated at least partially based on sensor data acquired by a mobile robot while navigating the environment; Receiving an image of an environment captured by a camera of a personal electronic device; Performing a semantic segmentation of the image; Updating the three-dimensional semantic map of the environment based on the semantic segmentation of the image; and Operating the mobile robot based on the updated three-dimensional semantic map. [20] Method for operating a mobile robot comprising the method: Storing a three-dimensional semantic map of the environment in a memory, wherein the three-dimensional semantic map includes semantic information relating to the environment, and wherein the three-dimensional semantic map is generated at least partially based on sensor data acquired by a mobile robot while navigating the environment; Displaying a graphical user interface for augmented reality on a display of the personal electronic device, wherein the graphical user interface for augmented reality comprises a variety of graphical elements superimposed on the environment, the variety of graphical elements comprising graphical representations of the semantic information of the three-dimensional semantic map; Determine at least one task to be performed by the mobile robot based on user input received via the augmented reality graphical user interface; and Operating the mobile robot to perform at least one task using the three-dimensional semantic map.