Augmented reality coordination of human-robot interaction
Augmented reality systems enhance human-robot interaction by integrating robot data into the user's environment view, improving situational awareness and safety in collaborative environments.
Patent Information
- Application Number
- JP2024114564
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-03-05
- Filing Date
- 2024-07-18
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2039-03-05
AI Technical Summary
Current human-robot interaction systems face challenges in effective robot teleoperation and intent communication, particularly in collaborative environments, leading to reduced situational awareness and increased risk of collisions or unsafe operations due to the need for high operator skill and split attention between robot monitoring and environmental context.
The use of augmented reality (AR) technology to integrate robot data and user perspective, allowing for intuitive control and hazard avoidance through augmented environments, robots, and user interfaces, using ARHMDs to provide stereoscopic and context-aware feedback.
Enhances situational awareness and safety by integrating robot data directly into the user's environment view, reducing cognitive burden and minimizing distractions, thus improving task performance and collision avoidance.
Smart Images

Figure 0007792720000001 
Figure 0007792720000002 
Figure 0007792720000003
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Application No. 62 / 638,578, filed March 5, 2018, which is incorporated by reference herein in its entirety for all purposes. Statement Regarding Federally Sponsored Research
[0002] This invention was made with government support under Grant No. NNX16AR58G awarded by NASA. The government has certain rights in this invention.
[0003] Some embodiments relate to human-robot communication, and more particularly to novel techniques for using augmented reality to facilitate communication of intent and teleoperation of robots in collaborative human-robot environments. [Background technology]
[0004] Humans are increasingly working alongside robots. For example, aerial drones, rovers, robotic arms, and other robots are used in manufacturing environments alongside factory workers, in emergency situations alongside emergency responders, and in many other collaborative environments. Effective and safe deployment of robots in these environments can depend on several factors. For non-autonomous and semi-autonomous robots, one such factor is effective robot teleoperation. For example, some situations may depend on a human operator being able to precisely and predictably control one or more robots while simultaneously being able to monitor the robot's location, the robot's sensor feedback, the location of environmental objects, environmental feedback, and other inputs. For autonomous robots, another such factor is the effective communication of the robot's intent. For example, when a robot is making dynamic decisions about where and how it intends to move, those movements may be unpredictable to humans in the robot's environment. As a result, intuitively communicating the robot's intent can significantly improve safety and collaborative efficiency. Summary of the Invention [Means for solving the problem]
[0005] Systems and methods are described for novel techniques for using augmented reality to facilitate communication of intent and teleoperation of a robot in a human-robot collaborative environment. In some embodiments, a method for operating an augmented reality display is provided. The method can include receiving data collected by a robot being teleoperated by a user using the augmented reality display (e.g., wearing a head-mounted display). The data collected from the robot can include information about the local environment (e.g., temperature, noise, radiation levels, etc.) and / or one or more robot states or robot states. The augmented reality system can then localize the view of the augmented reality display to identify a user perspective of the environment currently within the line of sight of the user wearing the head-mounted display. The system can then link the user perspective of the environment currently within the user's line of sight to data collected by the robot. The augmented reality view in the augmented reality display can be updated to enhance the user perspective with robot data based on the data linked to the user perspective.
[0006] In some embodiments, one or more commands may be received from a user that instruct the robot to change one or more robot states. The commands can be analyzed to identify any hazards that, when executed, could damage the robot or violate safety rules. If a hazard is detected, a set of one or more modified commands can be generated to avoid the hazard. The augmented reality view in the augmented reality display can be updated to show the results of the robot executing the one or more modified commands. The set of one or more modified commands can then be sent to the robot for execution.
[0007] Embodiments of the present invention also include computer-readable storage media containing sets of instructions that cause one or more processors to perform the methods, method variations, and other operations described herein.
[0008] While multiple embodiments are disclosed, still other embodiments of the present invention will become apparent to those skilled in the art from the following detailed description, which shows and describes illustrative embodiments of the invention. As realized, the invention is capable of modification in various aspects, all without departing from the scope of the invention. Accordingly, the drawings and detailed description are to be regarded as illustrative in nature and not restrictive.
[0009] Embodiments of the present technology are described and explained using the accompanying drawings. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 illustrates an example of an environment in which some embodiments of the present technology may be utilized.
[0011] [Figure 2] FIG. 1 illustrates a set of components for a robot that may be used in accordance with one or more embodiments of the present technology.
[0012] [Figure 3] 1 is a flowchart illustrating a set of operations for operating an augmented reality interface in accordance with some embodiments of the present technology.
[0013] [Figure 4] 1 is a flowchart illustrating a set of operations for hazard detection within an augmented reality interface, in accordance with various embodiments of the present technology.
[0014] [Figure 5] FIG. 10 is a sequence diagram illustrating an exemplary set of communications between various components of a teleoperated robot, in accordance with some embodiments of the present technology.
[0015] [Figure 6A] FIG. 1 illustrates an example of an augmented reality of robotic teleoperation that may be used in accordance with one or more embodiments of the present technology. [Figure 6B] FIG. 1 illustrates an example of augmented reality remote control of a robot that may be used in accordance with one or more embodiments of the present technology. [Figure 6C] FIG. 1 illustrates an example of augmented reality remote control of a robot that may be used in accordance with one or more embodiments of the present technology.
[0016] [Figure 7A] FIG. 1 illustrates an example of augmented reality using a real-time virtual surrogate (RVS) of a virtual aerial robot sharing a physical environment with a physically embodied aerial robot, in accordance with some embodiments of the present technology. [Figure 7B] FIG. 1 illustrates an example of augmented reality using a real-time virtual surrogate (RVS) of a virtual aerial robot sharing a physical environment with a physically embodied aerial robot, in accordance with some embodiments of the present technology.
[0017] [Figure 8A] FIG. 10 illustrates an example of an augmented reality display of robot intentions to mediate collocated human-robot interaction by visually communicating the robot's behavioral intentions, in accordance with various embodiments of the present technology. [Figure 8B] FIG. 10 illustrates an example of an augmented reality display of robot intentions to mediate collocated human-robot interaction by visually communicating the robot's behavioral intentions, in accordance with various embodiments of the present technology. [Figure 8C] FIG. 10 illustrates an example of an augmented reality display of robot intentions to mediate collocated human-robot interaction by visually communicating the robot's behavioral intentions, in accordance with various embodiments of the present technology. [Figure 8D]FIG. 10 illustrates an example of an augmented reality display of robot intentions to mediate collocated human-robot interaction by visually communicating the robot's behavioral intentions, in accordance with various embodiments of the present technology.
[0018] [Figure 9] 10A-10C illustrate objective results showing that the augmented reality interface design of various embodiments of the present technology improved task performance in terms of accuracy and number of collisions while minimizing distractions in terms of number of gaze shifts and total time of distraction.
[0019] [Figure 10] FIG. 1 shows objective results showing that the RVS and WVS systems used in some embodiments showed improvement over baseline in all objective measures.
[0020] [Figure 11] 1 illustrates objective results showing that NavPoints, arrows, and gaze improved work performance by reducing inefficiency and wasted time.
[0021] [Figure 12] FIG. 1 is a block diagram illustrating an exemplary machine representing a computer systemization of a monitoring platform that may be used in various embodiments of the present technology. DETAILED DESCRIPTION OF THE INVENTION
[0022] The drawings are not necessarily drawn to scale. Similarly, some components and / or operations may be separated into various blocks or combined into a single block for purposes of illustrating some embodiments of the present technology. Moreover, while the present technology is susceptible to various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and are described in detail below. However, the intention is not to limit the present technology to the specific embodiments described. To the contrary, the present technology is intended to cover all modifications, equivalents, and alternatives falling within the scope of the technology as defined by the appended claims.
[0023] An increasing number of environments are emerging that involve coordination between human and robot activities. In some of these environments, humans and robots are performing completely separate tasks in overlapping spaces, so coordination can help avoid collisions, interference, injuries, and / or other undesirable outcomes. For example, it may be desirable for city pedestrians or factory workers to avoid collisions with autonomous vehicles and factory robots. In other of these environments, humans and robots are working together to perform specific tasks, so collaboration can help facilitate collaboration between robot and human activities. For example, coordination can facilitate the exchange of tools or parts between human and robot actors, notifications between human and robot actors of remaining and completed work, contribution of knowledge gained by one to the other, etc.
[0024] The embodiments described herein include novel techniques for improving collaboration between humans and robotic actors using augmented reality (AR). Various implementations include novel approaches to virtually augmenting the robotic actor, the environment in which the robot is operating, and / or the user interface for controlling the robot. For example, some implementations support the development of consumer see-through augmented reality head-mounted displays (ARHMDs). Leveraging recent advances, novel human-robot interactions are enabled. Generally, the approaches are categorized herein as "robot teleoperation" and "robot intent" embodiments. It will be appreciated that the categorization is for purposes of further clarity only, and that techniques described with respect to robot teleoperation embodiments can be applied to robot intent embodiments, and vice versa.
[0025] Before describing specific embodiments, it is helpful to provide an overview of some related art. Robot Interface Technology
[0026] While the embodiments described herein apply to any suitable type of robot (e.g., aerial robots, rovers, robotic arms, etc.), the discussion focuses on aerial robot interfaces to provide a useful exemplary context. Currently, most interfaces for aerial robots take one of two forms: a direct teleoperation interface, in which the robot is controlled by the user via a joystick (or similar device) with a video display, or a high-level supervisory interface, which allows the user to plot waypoints outlining a desired robot path.
[0027] Remote control interfaces for aerial robots often require users to possess a great deal of skill, being able to pilot the robot with potentially unfamiliar degrees of freedom while monitoring a live robot video feed. As an example of the skills required, the US Federal Aviation Administration (FAA) initially considered regulations requiring commercial aerial robot operators to hold a pilot's license, but even today such robots must be operated within line of sight.
[0028] Much research has attempted to improve the teleoperation paradigm. For example, certain interfaces provide the operator with a first-person view of the robot's video feed via display glasses. While this can be useful for certain tasks, it can also reduce overall situational awareness, as the operator loses all contextual knowledge of the surrounding environment outside the robot's direct field of view. Other interfaces combine live video displays with virtual map data, often blending teleoperation with a form of autonomous waypoint navigation. Still other approaches have aimed to develop control systems using multimodal interfaces, including exotic designs in which the user wears a head-mounted display (HMD) and the robot is controlled via a "floating head" metaphor, turning in sync with the user. Finally, other research has advanced the concept of "cognitive-primary control," which allows the user to "loosen" an aerial robot using gestures on a mobile touchscreen, supporting more precise and safer operation.
[0029] One limitation of many existing systems is that even though there are many deployments where co-located control could be useful (e.g., construction site monitoring, hilltop reconnaissance, factory logistics management, factory environmental surveys, etc.), or even when it is necessary for line-of-sight control, they often focus on remote teleoperation rather than co-located teleoperation. Additionally, current interfaces tend not to allow users to view information collected by the robot (e.g., live video feeds) while directly monitoring the robot itself within the operating environment. Instead, conventional interfaces typically present the robot video feed and other sensor information on a display (e.g., a mobile device) and require the user to choose between monitoring the robot or the robot's video feed, in a paradigm similar to someone texting while driving. Similar to texting while driving, monitoring a mobile device while teleoperating can distract the user, reducing situational awareness and leading to poor maneuvering and, potentially, a crash. Some of the enhancements described herein may be useful for: Augmented reality techniques enable new models of teleoperation interfaces that alleviate this problem, allowing users to combine the knowledge and benefits of both first-person and third-person views. Augmented reality technology
[0030] Augmented reality technology superimposes computer graphics onto a real-world environment in real time. AR interfaces have three main characteristics: (1) users can see real and virtual objects in a combined scene, (2) users have the impression that virtual objects are realistic representations and directly integrated into the real world, and (3) virtual objects can be interacted with in real time. This contrasts with purely virtual environments, or other parts of the mixed reality continuum, such as augmented reality, where real objects are mixed into the virtual environment.
[0031] Early AR systems were often custom-built in laboratories and were significantly limited in display fidelity, rendering speed, interaction support, and generalizability. However, recent advances in augmented reality head-mounted display (ARHMD) technology are creating an ecosystem of standardized consumer-grade see-through ARHMDs. For example, HoloLens and Meta2 ARHMDs both offer high-resolution stereoscopic virtual images displayed at a 60Hz refresh rate, built-in gesture tracking, depth sensing, and integration with standard development tools such as Unity and Unreal Engine. This advancement in hardware accessibility is creating new opportunities for exploring AR as an interaction medium to enhance human-robot interaction (HRI).
[0032] Various embodiments of the technology described herein explore the use of AR technology as an enabling technology for novel types of robot teleoperation and robot intent communication. Specifically, the embodiments described herein use AR to augment human-robot interaction with virtual imagery in three classifications: by (1) augmenting the environment, (2) augmenting the robot, and / or (3) augmenting the user interface. Such classifications are for further clarity only and do not limit the scope of any particular embodiment. Various examples are described with reference to ARHMDs, which enable several features such as stereoscopic information (e.g., including depth cues), a field of view comparable to a human user's normal field of view, hands-free operation, and an ultra-high level of immersion. Nevertheless, some or all of the embodiments described herein may be considered "windows on the world." This may similarly be implemented using any other suitable AR interface, such as a portable handheld tablet that acts as an "on the world" AR interface.
[0033] Augmented Environment: In this paradigm, the interface can display information about robot operation and data collected by the robot as a virtual image directly embedded in the context of the operating environment, using an environment-as-canvas metaphor. For example, objects the user / robot has inspected (or plans to inspect) may be highlighted, or information may be added to better illustrate the robot's field of view. This concept extends past work using mixed reality projection systems in three key ways. First, ARHMDs provide a fully three-dimensional "canvas" to utilize for virtual images, rather than the two-dimensional canvas from projector systems. Second, ARHMD environment augmentation does not risk occlusion, such as when the user or robot obstructs the projected light. Third, ARHMDs support stereoscopic environmental cues, which can more effectively leverage human depth perception, as opposed to the monocular cues in traditional projector systems.
[0034] Augmented Robot: In this archetype, a virtual image may be attached directly to the robot platform within a robot-as-canvas metaphor. This technique can provide context-sensitive cues to the operator in a more fluid manner than traditional interfaces. For example, rather than displaying a battery indicator on a 2D display and requiring the operator to look away from the robot to check its status, the virtual image instead provides an instruction icon directly above the robot in physical space, allowing the operator to maintain awareness of both the robot's location and status. This technique can also modify the robot's shape and / or function by creating new "virtual / physically embodied" cues; cues traditionally generated using the robot's physical aspects are generated using an indistinguishable virtual image.
[0035] For example, rather than directly modifying the robot platform to include traffic lights, an ARHMD interface may similarly overlay virtual traffic lights onto the robot in the same manner. Alternatively, virtual imagery may be used to give anthropomorphic or animalistic features to robots that lack this physical capability (e.g., adding a virtual body to a single manipulator or a virtual head to an aerial robot). Finally, virtual imagery may be used to make various aspects of the robot's form less salient or more prominent based on the user's role (e.g., an override switch may be hidden from a normal user but visible to an engineer). For example, robot form as a design variable that is fast, easy, and inexpensive to prototype and operate can enable several novel approaches, in contrast to traditional limitations on modifying robot form, such as prohibitive cost, time, and / or other constraints.
[0036] Augmented user interface: In this paradigm, virtual images are displayed directly in front of the user as an overlay to provide an interface to the physical world, inspired by "window-on-the-world" AR applications and head-up display technologies used in piloting. This interface-as-canvas metaphor may draw much inspiration from traditional 2D interface design, for example, maintaining a view of the robot while providing supplemental information about the robot's orientation, pose, GPS coordinates, or connection quality in the user's surroundings, thus maintaining situational awareness of the environment. Furthermore, the interface-as-canvas metaphor can uniquely provide egocentric cues either directly in front of the user's view or around them, compared to exocentric feedback provided by augmenting the environment or robot. For example, user interface augmentations may include a spatial minimap providing information about the robot's location or planned route to the user, robot status indicators (e.g., battery level, job progress, job queue), or a live video stream from the robot's camera.
[0037] These design paradigms can have advantages over traditional interfaces. For example, ARHMDs support stereoscopic cues that can more effectively leverage human depth perception, as opposed to monocular cues in traditional interfaces. Moreover, these paradigms enable interfaces that provide feedback directly in the context in which the robot is actually operating, reducing the need for context switching between monitoring the robot and behavioral data, thus helping to solve the gaze acquisition problem. Robot Teleoperation Embodiment
[0038] Various embodiments of the present technology explore the use of augmented reality technology to help mediate robotic teleoperation with novel forms of intuitive visual feedback. Human interaction with robots and semi-autonomous robots often involves some form of teleoperation. Generally, teleoperation is the remote electronic control of a robot or other type of machine. Teleoperation can include the manipulation of a user interface that allows a user to control some or all robot functions, such as the movement and actuation of indicators or sensors. Teleoperating a robot can often be a difficult task that requires significant user training and expertise, especially for platforms with many degrees of freedom (e.g., industrial manipulators and aerial robots). Users often struggle to synthesize the information the robot collects (e.g., camera streams) with situational knowledge of how the robot is moving in the environment.
[0039] Robotic teleoperation, in which a user manually controls a robot, typically requires a high degree of operator expertise and can impose a significant cognitive burden. However, it can also provide high precision and require little autonomy on the part of the robot. As a result, teleoperation remains the dominant paradigm for human-robot interaction in many domains, including the operation of surgical robots for medical purposes, robotic manipulators for space exploration, and aerial robots for disaster response. Even in future systems in which robots achieve greater degrees of autonomy than current human-robot teams, teleoperation may still play a role. For example, in the "shared control" and "user-directed guarded motion" paradigms, the robot allows the user to directly input teleoperation commands but uses these commands in an attempt to infer the user's intentions rather than precisely following the received input, especially if the received input could lead to unsafe operation.
[0040] A substantial body of research has investigated issues of human performance in various forms of robotic teleoperation interfaces and hybrid teleoperation / supervisory control systems. In particular, previous work has highlighted perspective-taking issues, the concept that poor awareness of the robot and its working environment can reduce situational awareness and therefore negatively impact operational effectiveness. This can be problematic both in remote exploration (as found in space exploration) and when operators and robots are co-located (as may occur in search and rescue or building inspection scenarios).
[0041] Current interface designs can exacerbate this problem, as live robot camera feeds are typically presented in one of two ways: directly through display glasses or on a traditional screen (e.g., a mobile device, tablet, or laptop computer). While video display glasses help users achieve an egocentric understanding of what the robot can see, they can also reduce overall situational awareness by removing a third-person perspective that can aid in understanding the operational situation, such as identifying obstacles and other surrounding objects that the robot cannot see directly. Meanwhile, routing robot camera feeds through traditional displays means that, at any given time, the operator can only view the video stream on their own display or on the robot within their physical space. As a result, the operator must constantly switch context between monitoring the robot's video feed and monitoring the robot, creating a split attention paradigm that bears similarities to texting while driving.
[0042] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the present technology. However, it will be apparent to one skilled in the art that embodiments of the present technology may be practiced without some of these specific details. The techniques presented herein may be implemented as dedicated hardware (e.g., circuits), software, and / or firmware. The electronic instructions may be embodied as programmable circuitry that is appropriately programmed with software, or as a combination of dedicated circuitry and programmable circuitry. Accordingly, embodiments may include a machine-readable medium having stored thereon instructions that can be used to program a computer (or other electronic device) to perform a process. Machine-readable media may include, but are not limited to, a floppy diskette, an optical disk, a compact disk read-only memory (CD-ROM), a magneto-optical disk, a ROM, a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, a flash memory, or any other type of medium / machine-readable medium suitable for storing electronic instructions.
[0043] Phrases such as "in some embodiments," "according to some embodiments," "in the illustrated embodiment," "in another embodiment," and the like generally mean that the particular feature, structure, or characteristic that follows the phrase is included in at least one implementation of the technology and may be included in more than one implementation. Additionally, such phrases do not necessarily refer to the same or different embodiments.
[0044] FIG. 1 illustrates an example of an environment 100 in which some embodiments of the present technology may be utilized. As illustrated in FIG. 1, the communication environment 100 may include one or more mobile robots 110, a user 120 with a controller that remotely controls the robot 110, and an augmented reality system (e.g., a head-mounted display 140) that provides the user with contextual information regarding the robot's operation or intended operation. Data from the robot and / or controller may be communicated or retrieved from a monitoring service 150. As such, in some embodiments, the robot 110, the augmented reality system, and the controller may include network communication components that enable these devices to communicate with remote servers or other portable electronic devices by transmitting and receiving wireless signals using licensed, semi-licensed, or unlicensed spectrum on a communication network. In some cases, the communication network may be comprised of multiple networks, even multiple heterogeneous networks, such as one or more border networks, voice networks, broadband networks, service provider networks, Internet Service Provider (ISP) networks, and / or public switched telephone networks (PSTNs), interconnected via gateways operable to facilitate communication between the various networks. The communication network may also include a third-party communication network, such as a Global System for Mobile Communications (GSM) mobile communication network, a Code Division / Time Division Multiple Access (CDMA / TDMA) mobile communication network, a third-generation or fourth-generation (3G / 4G) mobile communication network (e.g., General Packet Radio Service (GPRS / EGPRS), Enhanced Data Rates for GSM Evolution (EDGE), Universal Mobile Telecommunications System (UMTS), or Long Term Evolution (LTE) network), or other communication network.
[0045] The augmented reality system 140 can provide a direct line of sight to the robot 110 or can recreate the robot within the augmented reality system. The user 120 can develop commands to control the robot 110 using the controller 130. More specifically, according to various embodiments, rather than directly controlling the physical robot as in traditional teleoperation, teleoperation commands are intercepted and directed to a virtual surrogate. These commands can be processed by the augmented reality system and visualized on the head-mounted display 140 before being implemented on the actual robot 110. As such, the augmented reality display 142 may change over time. For example, as shown in FIG. 1, at time T1, the augmented reality display 142 may show only the robot and other physical items in a room, without any augmentations. The user can then see that the robot is moving around the room, as shown in FIG. 1. A user 120 may be requested to fly around the bull and various augmented reality information may be displayed to assist the user 120.
[0046] For example, in response to an initial command from a user, the augmented reality system may display, at time T2, a virtual line 147A directly beneath the robot connecting it to a double ring visualization lying on its back on the ground that serves as a depth indicator. At time T3, the display 142 may update to display a real-time virtual surrogate (RVS) showing where and how the robot will fly. Once the physical robot arrives, additional contextual information 149A-149B (e.g., the robot's battery life, height, sensor readings, etc.) may be displayed within the augmented reality display 142 at time T4.
[0047] 2 illustrates a set of components for a robot 200 that may be used in accordance with one or more embodiments of the present technology. As illustrated in FIG. 2, the robot 200 may include a power source 205 (e.g., a rechargeable battery), a memory 210 (e.g., volatile and / or non-volatile memory), an actuator 220, a sensor 225, a navigation system 230, a communication system 235, a human interface module 240, an inertial measurement unit (IMS) 245, a global positioning system (GPS) 250, a path estimator 255, a future state estimator 260, and / or a controller supervisor 265.
[0048] Additionally, robot 200 may include various processors, such as an application processor, various co-processors, and other dedicated processors for operating robot 200. In some embodiments, the processors include one or more dedicated or shared processors configured to perform signal processing (e.g., a baseband processor for cellular communications), implement / manage real-time wireless transmission operations (e.g., to a nearby augmented reality device such as a head-mounted display), or perform other calculations or decisions. The processors are communicatively coupled to memory 210 and may be configured to execute an operating system, a user interface, sensors 225, a navigation system 230, a communication system 235, a human interface module 240, an inertial measurement unit (IMS) 245, a global positioning system (GPS) 250, a path estimator 255, a future state estimator 260, and / or a controller supervisor 265, and / or other components. These processors, along with other components, may be powered by power source 205. Volatile and non-volatile memory found in various embodiments may include storage media for storing information such as processor-readable instructions, data structures, program modules, or other data. Some examples of information that may be stored include a basic input / output system (BIOS), an operating system, and applications.
[0049] Actuators 220 may be involved in moving and controlling various components or parts of robot 200. For example, in some embodiments, actuators 220 may include, but are not limited to, electric motors, linear actuators, piezoelectric actuators, servos, solenoids, stepper motors, etc. Sensors 225 may be used to detect conditions, events, or changes in the surrounding environment and generate corresponding signals that various components within the robot can act on. In some embodiments, sensors 225 may include one or more of the following: microphones, cameras, encoders, accelerometers, light sensors, motion sensors, radiation sensors, moisture sensors, chemical sensors, lidar, radar, etc. Some of these sensors may be used as parts of navigation system 230, which may be involved in determining a navigation path for robot 200 (e.g., taking into account obstacles detected by sensors 225). Navigation system 230 also includes, for example, a communication system 235, an IMU 24, and other related devices. 5, and / or GPS 250 may be used to receive input from external sources to determine optimal flight paths, detect and avoid objects, coordinate with other nearby robots using communication system 235, etc. For example, IMU 245 may determine the orientation and velocity of robot 200. Controller supervisor 265
[0050] The communication system 235 may include cellular and / or short-range communication components for transmitting and / or receiving information from other robots, augmented reality systems, controllers, data drops, etc. The human interface module 240 can receive input from and provide output to humans. The human interface module 240 can use a microphone to receive verbal commands or queries from nearby humans. For example, when a human passes by, the human can request the robot 200 to pause its movement using the human interface module 240 to allow the human to safely pass through. Upon detecting that the user has passed through (e.g., using the sensor 225), the robot 200 can resume the requested movement. As another example, the human interface module 240 can display one or more visual cues (e.g., projections, holograms, lights, etc.) to assist the user in the vicinity of the robot 200's intention. According to various embodiments, the human interface module can use an augmented reality display. The controller / supervisor 265 may delay the implementation of one or more commands to ensure the safety of the robot and nearby objects or humans.
[0051] The path estimator 255 receives the current state (e.g., speed, location, commands in queue, etc.) and can estimate a current navigation path the robot 200 can take. The future state estimator 260 can generate predictions of paths and activities the robot 200 is likely to take. Using the human interface module 240, indicators of those paths (e.g., probabilistic instructions) can be displayed (e.g., via an augmented reality display used by the human) to assist the human in understanding the robot 200's expected path, intentions, or activities. For example, this information can be broadcast using the communication system 235 to enable a nearby head-mounted display or other augmented reality system to create visualizations for the user.
[0052] FIG. 3 is a flowchart illustrating a set of operations 300 for operating an augmented reality interface in accordance with some embodiments of the present technology. These operations may be implemented or performed by various components of the augmented reality interface, such as, but not limited to, one or more processors, ASICs, displays, communication modules (e.g., Bluetooth, cellular, radio frequency, etc.), and / or other components. As shown in FIG. 3 , a receive operation 310 may receive robot data including robot state information (e.g., orientation, velocity, position, acceleration, etc.) and environmental data (e.g., collected by one or more sensors). This data may be received directly from the robot or indirectly via an intermediate source (e.g., monitoring service 150, relay, satellite, etc.). In some embodiments, the source of the data may be selected and / or switched based on task information, delay tolerance, communication network characteristics (e.g., delay), and / or other information. In some embodiments, some data may be obtained directly from the robot, while other data may be retrieved from information data drops that may be identified by the robot and / or the augmented reality system (e.g., based on the robot's location). Therefore, some data may be retrieved from other sources.
[0053] A transformation operation 320 can transform the data into a visual representation that is meaningful to humans. The human-meaningful representation may indicate the location of the robot in the environment, objects in the environment, and / or collected data. For example, the human-meaningful representation may include a virtual robot, waypoints, callouts, tables of data, etc. Once the data is converted into a human-meaningful visual representation, an update operation 330 may generate or update a rendering of the visualization via the augmented reality display.
[0054] The augmented reality system may be configured to monitor instructions from a user (e.g., to remotely control a robot). The system may constantly monitor for user instructions, which can take various forms, such as, but not limited to, voice commands, hand gestures, etc. If determine operation 340 determines that no user instructions have been received, determine operation 340 branches to receive operation 310 where further robot data is collected. If determine operation 340 determines that one or more user instructions have been received, determine operation 340 branches to update determine operation 350. Update determine operation 350 may analyze the commands to determine whether an update is needed. If update determine operation 350 determines that an update is needed, update determine operation 350 branches to update operation 360 where the head-mounted display is updated according to the user's instructions. For example, the display may be updated to provide contextual information, a projected robot path, hazard identification, zooming in or out of a particular area, etc.
[0055] If the update determination operation 350 determines that no update is needed or the update operation 360 is complete, these operations branch to a convert operation 370 where the instruction is converted into robot state information before being sent by a send operation 380. The robot can then implement the command based on the desired state information (e.g., using one or more controllers). The process then repeats.
[0056] 4 is a flowchart illustrating a set of operations 400 for hazard detection in an augmented reality interface, in accordance with various embodiments of the present technology. As shown in FIG. 4, a connect operation 410 connects a head-mounted display to a robot. A calibration operation 420 can then be initiated to ensure proper display of physical and virtual objects. For example, some augmented reality interfaces include a camera that replicates the line of sight (e.g., in the case of a closed system) and a camera that projects or renders contextual data onto the display. In either case, the augmented reality interface may need to identify the location and orientation of the camera to properly present a viewpoint to the user.
[0057] Once calibration is complete, a monitoring operation 430 can monitor user commands controlling the robot. Some embodiments include a correction function 440 that can identify user intent, potential dangers, or rule violations (e.g., staying below a maximum height, maximum speed, etc.) and automatically correct commands to prevent damage to the robot. This can, for example, keep the robot within desired location boundaries (e.g., identified by beacons or built into rules) or prevent the robot from colliding with walls, people, or other obstacles in the environment.
[0058] As shown in FIG. 4, hazard detection function 442 may first analyze the user's intent and command for hazards to execute. For example, the user's intent may be to scan radiation along a wall. However, the user command causes the robot to collide with the wall at a particular point or lose sight of the area. Determine operation 444 may determine whether a hazard or intent has been identified. In some embodiments, the identified intent or hazard may be immediately displayed on an augmented reality display. If determine operation 444 determines that a hazard or intent has been identified, determine operation 444 may determine whether the detected hazard The operation branches to a modify operation 446, which modifies the command based on the danger or intent. As a result, the robot operates more intuitively and automatically avoids any danger. If the determine operation 444 determines that no danger or intent has been identified, the determine operation 444 branches to an update operation 450, where the augmented reality visualization is updated. The update operation 450 may also be performed upon completion of the modify operation 446. This visualization gives the user time to reject any modifications or change the command before it is sent to the robot using a send operation 460.
[0059] FIG. 5 is a sequence diagram 500 illustrating an exemplary set of communications between various components of a teleoperated robot, in accordance with some embodiments of the present technology. As shown in FIG. 5, the robot 510 can connect to a monitoring service 520. The monitoring service 520 can verify validation credentials and send an acknowledgment back to the robot 510. The robot can collect status and environmental data and then report it to the monitoring service 520. The monitoring service can store the data. The head-mounted display can then connect to the monitoring service and request the data collected by the robot based on its current location. This data can then be sent back to the head-mounted display, which can use the data to generate an augmented reality visualization. For example, this may be useful for replaying a robot path or data collection activities. Similarly, such data drop functionality may be useful because different robots may have different sensors and capabilities (e.g., heat sensors, radiation sensors, motion sensors, etc.). As such, information collected from multiple robots can be overlaid or provided to provide a more coherent understanding of the environment.
[0060] 6A-6C show examples of augmented reality remote control of a robot that may be used in accordance with one or more embodiments of the present technology. In FIG. 6A, a frustum design 610 is shown augmenting the environment, giving the user a clear view of what actual objects are within the robot's field of view. FIG. 6B shows a callout design 620 that augments the robot like a thought bubble and attaches a panel with a live video feed above the robot. FIG. 6C shows a surrounding design 630 in which the live video feed provides a fixed window into the user's surroundings.
[0061] The frustum design 610 provides an example of augmenting the environment. This design has a spatial focus, as it provides a virtual image that displays the robot camera's frustum as a series of lines and points, similar to what appears to emanate from a virtual camera in a computer graphics and modeling application (e.g., Maya, Unity, etc.). The virtual frustum provides information about the robot's aspect ratio, orientation, and position while explicitly highlighting which objects in the environment are within the robot's field of view. Another advantage of this design is that it allows the robot sensor's frustum to be viewed even when there is essentially no real-time feedback and / or when taking aggregate measurements (e.g., LiDAR, air quality sensors, spectrometers, etc.).
[0062] Callout design 620 represents an example of augmenting the robot itself with virtual imagery. It uses metaphors inspired by callouts, speech bubbles, and thought bubbles, in line with previous work that renders a real robot camera feed within the context of a virtual representation of the robot's environment. Similarly, the robot's live camera feed is displayed on a panel with an orientation that corresponds to the orientation of the camera on the physical robot, allowing the information in the video to spatially resemble the corresponding physical environment. This feed also provides implicit information about the distance between the robot and the operator, as a perspective transformation is applied to the video callout panel based on the calculated offset between the user and the robot. As a result, the panel size is proportional to the operating distance, just as the perceived size of a real object is proportional to the viewer's distance. This design decision can potentially impair long-distance operation. While this may be advantageous, it can better support scalability and fan-out for operating multiple robots, and provides more realistic incorporation of information directly in situations where a robot is collecting data at any given time. Another variation on this design may instead use a billboard so that the callout panel always faces the operator. This ensures that the operator always has a direct view of the robot's video stream, but may remove potentially useful orientation cues. Another design variation may use orthographic projection rather than perspective, so that the video panel always remains a fixed size relative to the operator, regardless of the robot's distance (this creates a similar effect to the peripheral design described below).
[0063] Peripheral design 630 illustrates a potential method for an augmented reality user interface that provides contextual information about the robot's camera feed in an egocentric manner. This design displays a live robot video feed in a fixed window in the user's view. Designers can specify fixed parameters for the window's size and location (e.g., in front of the user or in their peripheral view) or provide support for dynamic interaction that allows users to customize the window's width, height, location, and even opacity. This work placed the window in the user's periphery, anchored in the upper right corner of the user's view, in a manner inspired by ambient and peripheral displays. Robot Teleoperation Embodiments Using Virtual Surrogates
[0064] Another set of embodiments addresses robot teleoperation by using AR to generate and utilize virtual surrogates. As used herein, a virtual surrogate is generally any suitable projection in an AR environment that represents a real robot. These embodiments demonstrate that AR can be used to provide a user with an immersive virtual robot surrogate to control, rather than directly controlling a physical robot. Controlling the surrogate allows the user to better predict how their actions will affect the system and predict the final pose and location of the physical robot that mimics the actions taken by the surrogate. While such a system could be useful for operating a wide variety of robots (e.g., manipulators, underwater robots, etc.), this work investigates interfaces for teleoperation of aerial robots. Two main designs explore potential tradeoffs for how such a surrogate interface affects the effectiveness of teleoperation.
[0065] Various embodiments may use a real-time virtual surrogate (RVS). This design presents a user with a virtual aerial robot that shares a physical environment with a physically embodied aerial robot. FIG. 7A shows an example of an RVS implementation with a virtual surrogate 710 and a physically embodied aerial robot 720. The appearance of the virtual surrogate is modeled after the physical robot 720 and may display a virtual “fishing line” 730 connecting the surrogate 710 to the physical robot 720. Additionally, a virtual line 740 may be rendered beneath the robot, connecting it to a double ring visualization 750 resting on the ground, which serves as a depth indicator (mimicking the drop shadow depth cues that have been shown to be effective in conveying depth in AR applications).
[0066] In this design, rather than directly controlling the physical robot as in traditional teleoperation, user teleoperation commands are intercepted and directed to a virtual surrogate 710. The virtual surrogate 710 is used as a setpoint or goal state for a planning algorithm running on the physical robot 720 that causes the physical robot 720 to constantly "chase" the surrogate 710 and stop only when the Euclidean distance between the virtual robot and the physical robot is zero. A simple PID controller is used. While the implementation selected a simple PID controller, any desired planning algorithm may be used. This design offers several adjustable parameters, including the virtual robot's control speed, the physical robot's 720 "chase speed" or delay time, and any additional constraints enforced by the planner (e.g., motion smoothness), but from an implementation perspective, this system only requires that the surrogate's 6DOF pose be transformed from the augmented reality coordinate system displaying the surrogate to the planning coordinate system as the desired goal state of the onboard robot autonomy. Overall, the design helps the user better understand how the physical robot 720 will respond in real time to the commands they issue and (given an appropriate "chase speed") gives the user an opportunity to evaluate and correct erroneous commands before the physical robot 720 actually executes them.
[0067] Some embodiments may use a waypoint virtual surrogate (WVS). This design extends the RVS model to provide more support for long-term planning, potentially at the expense of immediate precision control. Figure 7B shows an example of a WVS implementation. Similar to the RVS model, the teleoperator controls a virtual robot surrogate rather than directly manipulating the physical robot. However, in the WVS paradigm, the physical robot 720 remains in place while the user creates a plan by manipulating the virtual surrogate 710 to add / delete / edit virtual waypoints 760. At any point, the user can signal the physical robot 720 to begin executing a planned path defined by a series of 6DOF target poses specified by the waypoints 760. The user can edit the most recent waypoint, delete any planned waypoints, and add additional waypoints on the fly. This interface is inspired by recent research in high-level aerial robot interfaces and allows for the exploration of tradeoffs between RVS systems and more supervisory control strategies.
[0068] In addition to these two designs, a baseline teleoperation system can be implemented that allows the user to directly pilot a physical robot rather than controlling a virtual surrogate. This system can be based on modern aerial robot teleoperation interfaces, requiring the user to move the robot in 3D space using a joystick, but in the absence of user input, the robot autonomously maintains a stable hover.
[0069] The design used in various embodiments may be built on a back-end coordination system developed as a custom application within the Unity engine. The basis of this system is a virtual aerial robot object within a Unity application running on a Microsoft HoloLens. The application converts user input from an Xbox controller into a segmented list of desired poses for the virtual robot, which then navigates the scene according to the specifications of the waypoint list. With each iteration of the application engine's update loop, the virtual robot's 3D pose is sent from the HoloLens via UDP to an on-board system that controls the physical robot. The virtual robot's pose values are converted from Unity coordinates to real-world coordinates using a transformation matrix pre-computed during an initial calibration procedure that calculates the relative origin and basis vectors for each coordinate system. After this conversion, the pose values in the Unity scene correspond to the same locations in the user's real environment, allowing the physical robot to fly through the learning space in a manner consistent with the virtual drone.
[0070] In both RVS and WVS systems, the autonomy controlling the physical robot currently takes the form of a PID loop that uses the virtual robot's pose as a setpoint, but any robot planning algorithm can be used. The PID controller runs at 20 Hz and precisely controls the altitude, position, and orientation of the aerial robot. Because the current robotic platform (AscTec Hummingbird) lacks sufficient onboard sensing capabilities for accurate localization, the PID controller currently tracks the physical robot using motion-tracking cameras built into the environment. User control is implemented using an Xbox controller, with inputs aligned with the default AscTec controller, which is representative of modern teleoperation systems (Figures 8A-8D). The sensitivity of the Xbox controller was calibrated to be as similar as possible to commercial teleoperation systems and was kept constant across all interface designs. Robot Intention Embodiment
[0071] In addition to the embodiments described above, other embodiments include novel approaches to communicating robot behavioral intentions using augmented reality-based mediation of human-robot interactions, such as by generating and presenting physically embodied and / or other intuitive cues. Humans coordinate teamwork by communicating their intentions through social cues such as gestures and gaze behavior. When working with robots, especially autonomous robots (although this may also be the case for semi-autonomous or non-autonomous robots), similar coordination can benefit from communicating the robot's intentions. For example, when a robot is making dynamic decisions about where and how it intends to move, those movements may be unpredictable to humans in the robot's environment. As a result, intuitively communicating the robot's intentions can significantly improve safety and collaboration efficiency. However, these methods may not be possible with robots with constrained appearances that lack anthropomorphic or animal-like features, such as aerial or spherical robots.
[0072] Effective collaboration requires teammates to quickly and accurately communicate their intentions to build common ground, coordinate joint actions, and plan future activities. For example, previous research in social, cognitive, and behavioral sciences has found that collaborative activities fundamentally depend on mutual predictability—that is, the ability of each team member to quickly understand and predict their teammates' attitudes and behaviors. In collocated human-robot teams, poor communication of the robot's intentions and planned movements can lead to serious breakdowns that reduce safety, job performance, and perceptions of the robot's ease of use.
[0073] As a result, it can be difficult for users to understand when, where, and how their robot teammates will move. Providing support for this motion inference problem represents a major challenge to realizing safe and user-friendly robotic systems. In human-human teams, people use a variety of implicit and explicit cues, such as gaze, gestures, or other social behaviors, to communicate planned actions and movements, increasing team effectiveness and helping maintain trust. Research has demonstrated that robots can also use social cues to communicate both behavioral intentions and emotional states. However, it is not always clear how to apply these findings to robots lacking anthropomorphic and animalistic features, such as industrial robotic manipulators or aerial robotic drones.
[0074] Instead, alternative techniques can support movement inference, including generating legible movement trajectories, developing expressive motion primitives, verbalizing the robot's intentions using natural language, using projector-based or electronic display systems to provide additional information, and using optical signals as explicit directional cues. While such advances hold promise in increasing the safety and fluidity of interactions, various constraints arising from environmental, task, power, computational, and platform considerations may limit their feasibility or effectiveness in certain situations. For example, modifying a robot's behavior for readability or expressiveness may not always be possible in dynamic or cluttered environments, and natural language may be difficult to interpret in noisy environments (e.g., manufacturing warehouses or other public spaces). may not be a practical form of feedback for robotic platforms that generate a lot of noise (e.g., aerial robots), projections may be difficult to render on uneven surfaces, may not be noticeable in bright environments, and may be obstructed by the user or robot.
[0075] The problem of inferring robot movements can be seen as analogous to the "gulf of evaluation" problem that commonly arises between the representations provided by a system and the user's ability to interpret the system. This can be particularly challenging for robots with many degrees of freedom, such as aerial robots. Other issues can exacerbate this problem, including the robot's lack of ability to communicate intentions and goals using traditional methods, as well as the technological novelty / lack of a mature mental model for understanding the robot's behavior.
[0076] Previous research suggests that effective robot communication can improve users' perceptions of a robot's reliability, predictability, and transparency, increasing users' willingness to accept and use new robotic technologies in their work environments. Research has also shown that cues to a robot's intentions can help users anticipate and predict a robot's directional movements sooner, allowing users to respond more quickly in interactive tasks while increasing users' preference for working with robots. Legible movements that express a robot's intentions can further improve the fluidity and efficiency of interactions in human-robot collaborations.
[0077] Previous research in robot design has explored ways to effectively leverage users' prior experiences and mental models in human-human collaboration to bootstrap human-robot collaboration and instill social behaviors, such as gaze and gestures, that people typically use. Such behaviors have been explored for various robots using anthropomorphic and animal-like features. For aerial robots lacking such features, previous research has explored expressive flight patterns, demonstrating that specific behaviors based on biological motion and principles from film and animation can help compensate for the lack of developed mental models for unmanned flight locomotion. Research has also explored more explicit cues, such as the use of lights as indicators and mixed-reality projection systems, finding that the use of projected images can be advantageous for communicating spatial intent and instructions, such as informing a human collaborator of a robot's intended path.
[0078] While such research has shown promising benefits for improving human-robot interaction, including aerial robots, conventional methods are not without limitations. For example, the depicted flight motions may not be feasible in constrained environments or may be ineffective if the user only sees the robot and witnesses only part of the movement. Projection systems rely on environmental measurements and, in chaotic environments, face challenges at a distance and risk occlusion. By using AR as a means of communicating the robot's motion intentions, the embodiments described herein are not bound by these limitations.
[0079] Various embodiments for implementing novel robot intent techniques are described herein. FIG. 8 illustrates several examples of techniques for mediating collocated human-robot interactions by using augmented reality to visually communicate the robot's motion intent. For example, four different techniques for cueing the flight movements of an aerial robot are illustrated in FIGS. 8A-8D : the NavPoints technique, the Arrows technique, the Gaze technique, and the Utility technique, respectively. Together, these techniques represent a sampling of various design framework paradigms, offering potential tradeoffs in terms of conveyed information, information accuracy, generalizability, and potential for distraction / interface exaggeration.
[0080] These, along with the main design metaphors and interface techniques, are described with reference to aerial robots, similar design frameworks and methodologies, but can be applied in any suitable context for facilitating human-robot interaction (e.g., industrial robotic manipulators moving with high degrees of freedom).
[0081] The NavPoints design (shown in Figure 8A) is an example of an environment augmentation. This design has a spatial focus, as it provides a virtual image displaying the robot's planned flight path as a series of X lines and navigation waypoints, similar to those found in traditional waypoint delegation or monitoring interfaces. The lines sequentially connect the robot's current location to its future destinations. Destination waypoints are visualized as spheres, indicating the robot's exact destination in 3D space. Each destination sphere also renders a drop shadow on the ground directly below it, which has been shown to aid in user-generated depth estimation. Above each navigation point are two radial timers. The inner white timer indicates when the drone will arrive at that location, and the outer dark blue timer indicates when the robot will leave that location. The smaller spheres move along the lines, moving in the same direction and speed as the robot travels between destinations, providing clues for predicting the robot's future speed and direction. The information displayed by this design regarding speed and arrival / departure timing is thus explicitly displayed to the user.
[0082] The arrow design (shown in Figure 8B) offers an alternative example of how virtual imagery augments a shared environment. The NavPoints design provides users with a large amount of information, which may be distracting or confusing due to potential exaggeration. The arrow design takes a more minimal approach, focusing specifically on conveying temporal information, inspired by the common user experience with modern GPS systems. The virtual image consists of a blue arrowhead moving through 3D space on a precise path that will ultimately take the robot X seconds in the future. As the arrow moves, it leaves a line behind that traces the arrow's path back to the robot. This line allows users to explicitly see the path the arrow took, and the robot will eventually retrace that path. The line created by the arrow renders a drop shadow on the ground directly below it. The information displayed by this design regarding speed and arrival / departure timing must be inferred by the user by observing the arrow's individual movement through space.
[0083] The gaze design (shown in Figure 8C) represents one example of augmenting the robot itself with a virtual image. This design is inspired by previous research demonstrating the remarkable potential of gaze behavior to communicate intent, even in the case of aerial robots, as well as prior designs for robotic airships and research in robot teleoperation that explored the metaphor of treating an aerial robot as a "floating head." This design provides a virtual image that completely transforms the robot's shape by superimposing a white sphere of X meters diameter onto the pupil directly above the aerial robot, effectively transforming it from a multirotor into a "flying eye." While navigating between destinations, the eye model gazes at its current destination until it comes within a predetermined distance threshold of Y meters between itself and the current destination, at which point the eye turns and focuses on the robot's next destination. Such gaze shifts have been shown to be useful in predicting human behavioral intentions. These focus shifts preemptively reveal the robot's current destination to the user.
[0084] If the robot is to remain stationary at the destination for longer than Z seconds, the normally transparent lens over the pupil becomes opaque. When the currently stationary robot is within Z seconds of departure, the lens fades in and returns to transparent. This fade-in is done as a linear interpolation over the course of Z seconds. The effect of this lens fading in / out is as the robot becomes stationary and adjusts - The ciliary muscles shift as they contract and relax in human gaze as focus switches between near and far targets, informing the user of the length of time inspired by the focusing of a lens in a conventional camera. The size of the display was chosen to help the user more easily determine gaze direction at close and far distances from the robot. The back of the sphere directly behind the pupil is rendered flat to help infer eye rotation when the user is not facing the robot directly. Finally, the eye casts a drop shadow on the ground directly below it.
[0085] Another design (shown in Figure 8D) illustrates a potential way to extend user interfaces to provide contextual information in an egocentric manner. This design is inspired by peripheral utilities such as minimaps, radars, and off-screen indicators that often augment pilot interfaces, robot control interfaces, video game interfaces, and military applications. This design displays a 2D circular "radar" anchored in the lower left corner of the ARHMD display. The user is always displayed as a blue dot centered within the radar, while the robot is rendered as a red dot on the radar relative to the user's location. The size of the robot's radar dot is directly proportional to its current height. The radar's detection radius X can be customized by the interface designer or adjusted by the user. When the robot is within the user's field of view (FOV), it is superimposed with an aiming box; when not within the FOV, an off-screen indicator appears in the form of a directional arrow. This arrow is rendered along the side of the ARHMD display, pointing to the off-screen robot's location. Both the radar and the aiming box / off-screen indicator provide the user with a means to quickly position the robot relative to themselves.
[0086] Robot remote control experiment
[0087] We conducted a 4x1 between-participant experiment to evaluate how design affects user teleoperation of a collocated flying robot. The study involved participants operating a Parrot Bebop quadcopter and taking several photographs in a laboratory environment as an analog for aerial robot inspection and research tasks. The independent variable in this study corresponded to the type of teleoperation interface used by participants (four levels: frustum design, callout design, peripheral design, and baseline). In the baseline condition, participants still wore an ARHMD (to control for the possible effects of simply wearing an HMD) but did not view any augmented reality images. Instead, participants used the Freeflight Pro application, the official piloting interface provided by Parrot for the Bebop robot (the platform used in this experiment). The dependent variables included objective measures of task completion and subjective measures of operator comfort and confidence.
[0088] Our overall experimental design was inspired by the context in which unmanned aerial robots assist in environmental inspection and exploration within human environments, a practice already common among drone enthusiasts and likely soon to be found within domains including disaster response, operations aboard the International Space Station, and journalism. In the study, participants operated an aerial robot within a shared environment. The environment measured 5m x 5m x 3m and included a motion-tracking camera that was utilized to accurately track the robot and ensure the ARHMD visualization appeared in the appropriate position for the frustum and callout design (motion tracking was not required for the baseline condition or the peripheral design).
[0089] Two test targets decorated the walls of the experimental environment in the form of rectangular frames colored pink and purple with orange outlines. The larger pink target was 1.78 m x 1.0 m. The pink target is 1.35m x 0.76m and is 1.3m from the ground to its bottom edge. The smaller purple target is 1.35m x 0.76m and is 0.34m from the ground to its bottom edge. The aspect ratio of the targets matches perfectly with that of the robot's camera, allowing participants to take accurate photos of the targets. The pink target is large and high off the ground, making it relatively easy to capture a perfect image, while the purple target is more difficult due to its smaller size and shorter height above the ground (operating an aerial robot close to the ground is more difficult due to instability and updrafts being reflected off the ground).
[0090] To capture a perfect image of the purple target, the robot would need to fly closer to both the wall and the ground, increasing the likelihood of a collision due to operator error. Participants were tasked with piloting an aerial robot to take photos of the targets in a set order, inspecting the larger pink target first, followed by the smaller purple target. Participants were instructed to prioritize capturing images as quickly, accurately, and in as few total photos as possible. Regarding accuracy, participants were instructed to record photos that captured the entire pink or purple target area and to minimize any additional images within the photo (e.g., the orange target frame or other parts of the scene). Participants were able to determine when they had captured what they considered to be a suitable image so they could move on to the next target. Overall, this task mimicked an environmental inspection mission, requiring participants to operate a gaze robot to take off from a set starting position, inspect a series of targets in sequence, and land the robot within six minutes. If a participant crashed the robot, it was reset to its starting position and could then continue executing the task as long as time remained.
[0091] The experimental setup and implementation are described for context. The robotic platform used a Parrot Bebop quadcopter as the experimental aerial robot. The Bebop is a popular consumer "drone" with a digitally stabilized 14-megapixel HD camera and autonomous hovering capabilities, suitable for indoor and outdoor flight. The ARHMD platform included a Microsoft HoloLens as the ARHMD. HoloLens is a wireless, optical, see-through, stereoscopic augmented reality HMD with a 30° x 17.5° FOV, an inertial measurement unit, a depth sensor, an ambient light sensor, and multiple cameras and microphones supporting voice input, gesture recognition, and head tracking. HoloLens was selected due to its growing popularity, ease of access, ability to support hands-free AR, and high potential as a model for future consumer ARHMD systems. The teleoperation interface implementation included participants in a baseline condition wearing a HoloLens ARHMD but without augmented reality visualization. Instead, they controlled the robot via the "Free Flight Pro" application on an iPad®. The Free Flight application is the default control software for Bebop robots and is also developed by Parrot (the manufacturer of Bebop). It is a popular application (with an average rating of 3.8 from 24,654 reviews on the Android App Store) and represents the most modern aerial robot control interface in use today, aiming to provide users with an intuitive control scheme. The application provides touchscreen controls for robot takeoff / landing, positioning / orientation, and photo / video recording, all overlaid on a live video feed from the robot. In the absence of user input, the application ensures the robot continues to hover and automatically lands the robot if it detects low battery.
[0092] Unfortunately, limitations of the robot platform and the Freeflight application prevented the inventors from using it as a control input for other experimental conditions, nor could they stream the robot's video feed to any other device when Freeflight was connected to the robot. This precluded the use of the Freeflight interface in callout or peripheral designs. Instead, in the AR condition, participants received AR feedback while operating the robot using a wireless Xbox One controller. Button / joystick mapping was configured on the Xbox controller to match the touchscreen controls within the Freeflight Pro application, and the sensitivity of the calibrated Xbox controller made it as similar to the Freeflight controller as possible (Figure 7). As with the Freeflight Pro app, the robot continued to hover in place in the absence of user input.
[0093] Augmented reality visualizations of the frustum, callout, and peripheral designs were implemented using the Unity game engine and deployed as applications running on a Microsoft HoloLens. The frustum and callout designs required real-time understanding of the robot's location, so that virtual images could be accurately displayed either within the environment relative to the robot (frustum) or directly as an attachment to the robot itself (callout). To accomplish this, a motion tracking camera precisely localized the robot and provided the robot's location and orientation values to the HoloLens application. While the peripheral design did not rely on a motion tracking setup, both the callout and peripheral designs displayed a live video feed from the robot's camera as a virtual object in augmented reality. To accomplish this, the robot's video stream was wirelessly broadcast to the HoloLens by routing it through a desktop computer. This method resulted in an average frame rate of 15 frames per second (FPS), slightly slower than the ~30 FPS provided to participants in the baseline condition by the Freeflight Pro application.
[0094] Each of the three main designs tested had several parameters that could be adjusted by the designer or user, but each of these parameters was modified to control potential variance in the experiment. The frustum was displayed as a red line in the wireframe view, as opposed to a shaded or highlighted area to minimize potential occlusion of the environment or task target. Callouts were designed to emanate from the top of the robot, so that the video feed always appeared 13.5 cm above the center of the Bebop while the operator flew it throughout the environment. The peripheral design placed the camera feed window in the upper right corner of the user's view.
[0095] The experiment included 48 participants (28 men, 19 women, and 1 self-reported non-binary). Men and women were evenly distributed across conditions. Participants' mean age was 22.2 years (SD = 7.2), with a range of 18 to 58 years. On a 7-point scale, participants reported some prior familiarity with both aerial robots (M = 3.48, SD = 1.6) and ARHMDs (M = 3.38, SD = 1.75).
[0096] Procedurally, the experiment took approximately 30 minutes and consisted of five main phases: (1) introduction, (2) calibration, (3) training, (4) task, and (5) conclusion. First, participants were given a high-level overview of the experiment, signed a consent form, and were then guided to the experimental space. Participants then donned the HoloLens and, depending on their condition, were given either an iPad running the Freeflight Pro application or an Xbox controller. At this point, participants were assigned to either the conditions (baseline participants, no AR application for the frustum ... The appropriate HoloLens application was also launched based on the callout or surroundings. Participants then received control instructions (i.e., button map for Freeflight and Xbox controller) and had two minutes to practice piloting the robot. After the two minutes were up, the robot landed and was placed in a fixed starting position for all participants.
[0097] Participants then completed the main task of inspecting targets in sequence. Participants received six minutes to pilot the robot to take off from a fixed starting position, capture images of the targets, and then land, simulating an environmental inspection mission. If a participant crashed the robot, it was reset to its starting position, and the participant was allowed to continue the task if time remained. Once participants completed the task or the six minutes allotted for the task ran out, they were given a post-survey about their experience and then debriefed.
[0098] Objective, behavioral, and subjective measures were used to characterize the usability of the interface design. Several objective aspects of task accuracy were measured, including precision, measured by how well the participant's photograph captured the test target, each of which consisted of a visible, uniform grid of 297mm x 210mm rectangles, allowing accuracy to be measured by comparing the rectangles in the perfect photograph with the photograph captured by the participant; completion time, measured by total flight time (less time means more efficient performance); and operational errors, which were the number of times the participant crashed or otherwise landed the robot too early.
[0099] First- and third-person video was also recorded to analyze behavioral patterns of participants' movements. Two coders annotated the video data from each interaction based on when participants could and could not see the robot. Data was split evenly between the coders, with a 15% overlap in data coded by both. Inter-rater reliability analysis revealed substantial agreement between the raters (Cohen's kappa = .92). This coding allowed us to calculate distractive gaze shifts—the number of times participants were distracted by looking away from the robot during the task—and distraction time—the total time spent without looking at the robot. Both measures were relevant because many small gaze shifts can be just as detrimental as fewer but longer periods of distraction.
[0100] Several 7-point scales were constructed using Likert-type survey items to capture subjective participant responses. These scales measured how the interface design affected participants' comfort (3 items, Cronbach's α = .86), confidence (5 items, Cronbach's α = .95), and perceived task difficulty while operating the robot (5 items, Cronbach's α = .93). Participants also provided open-ended responses regarding their experience.
[0101] Participants rated perceived ease of use using the System Usability Scale (SUS), an industry-standard 10-item survey. A SUS score below 68 is considered below average, a score above 68 is considered above average, and a score above 80.3 is considered within the top 10 percentile.
[0102] Data from the objective and subjective measures were also analyzed using a one-way analysis of variance (ANOVA) with experimental condition (i.e., remote control interface) as a fixed effect. Post-hoc tests controlled for Type I error using Tukey's Honestly Significant Difference (HSD) for comparative effectiveness across each interface.
[0103] Figure 9 shows the objective results indicating that the augmented reality interface design improved task performance in terms of accuracy and number of collisions while minimizing distractions in terms of number of gaze shifts and total distraction time. (*), (**), and (***) denote comparisons with p<0.05, p<0.01, and p<0.001, respectively.
[0104] Objective Results - Objective task metrics were analyzed to confirm the usefulness of the design, which allowed participants to more effectively teleoperate the collocated aerial robot. Results found a significant main effect of design on task performance scores for accuracy, F(3,44) = 25.01, p < .0001. Tukey's HSD revealed that the frustum (M = 63.2%) and callout (M = 67.0%) interfaces significantly improved test performance over the baseline interface (M = 31.33%), with the peripheral design (M = 81.1%) demonstrating a further advantage by significantly outperforming both the frustum and callout (all post-hoc results, p < .0001). Results also found a significant main effect of design on task completion time, F(3,44) = 3.83, p = .016. Post-hoc comparisons with the baseline (M = 239.70 s) revealed that participants were able to complete the task significantly faster using the frustum (M = 140.69 s), p = .017, and periphery (M = 154.44 s), p = .050, but not the callout (M = 191.09 s), p = .434. When examining the occurrence of user collisions with the robot, results showed a significant effect of interface design on operation errors, F(3,44) = 9.24, p < .001, with each of the AR designs resulting in a significantly reduced number of collisions compared to the baseline (frustum: M = .250, p < .0001; callout: M = .667, p = .003; periphery: M = .584, p = .001; baseline: M = 2.17).
[0105] Behavioral metrics were also analyzed to understand user distraction, showing significant main effects for the number of distracting gaze shifts, F(3,44) = 40.28, p < .001, and total distraction time, F(3,44) = 48.72, p < .001. Post-hoc tests showed that all three AR designs significantly reduced both the number and length of distractions compared to baseline (p < .0001 for all comparisons).
[0106] Subjective Results - Participants rated several aspects of their experience teleoperating the aerial robot. Results showed a significant effect of interface design on users' comfort working with the robot, F(3,44) = 8.12, p < .001, with users being more comfortable using the frustum (M = 5.58), p = .002, callout (M = 5.87), p < .001, and periphery (M = 5.14), p = .019 designs than the current Freeflight interface (M = 3.64). Results also showed a significant effect of design on confidence in operating the robot, F(3,44) = 7.93, p < .001. Post-hoc comparisons to baseline (M = 3.0) revealed that users were significantly more confident using the frustum (M = 5.02) and callout (M = 5.0) designs, both p < .001, but not the periphery (M = 4.0), p = .179. Results also showed a significant effect of design on perceived task difficulty, F(3,44) = 4.17, p = .011, with participants finding the task significantly easier using the following designs than baseline (M = 2.5): frustum (M = 4.03), p = .027, callout (M = 4.13), p = .016, and periphery (M = 3.78), p = .047.
[0107] Perceived interface usability was assessed with the SUS. Results showed a significant effect of interface design on the total SUS score, F(3,44) = 7.38, p < .001. Tukey's HSD revealed that each of the AR designs had a significantly higher usability rating than the baseline (frustum: M = 80.21, p < .001). 001, Callout: M = 76.88, p = .003, Peripheral: M = 71.04, p = .041, Baseline: M = 55.0).
[0108] The results of the experiments described above generally demonstrated significant improvements over state-of-the-art interfaces representing popular designs currently in use. Each design enabled users to complete inspection tasks faster and more accurately than using the robot's default control interface, while also leading to safer operation with fewer collisions. Overall, users rated these designs as much more favorable in terms of usability and their own comfort and confidence when operating the robot. For example, while the interface design embodiments described herein enabled users to obtain live video feedback without taking their eyes off the robot, traditional designs tend to force users to context-switch to closely monitor the robot's camera feed at any time, sacrificing either situational awareness of the robot in its environment or their ability to do so.
[0109] In this experiment, as with real-world placement, users needed to understand both of these aspects to successfully complete their task. This was particularly important when participants attempted to inspect a small pink target close to the ground. The size and placement of this target required participants to navigate fairly close to both the wall and the ground, which challenged the robot's internal stabilization mechanism and potentially caused the robot to drift while hovering. Observation of the experimental recordings reveals that participants using the frustum, callout, and perimeter design were able to quickly notice the drift and re-stabilize the robot. However, participants in the baseline condition often stared at the tablet providing the robot's video feed rather than monitoring the robot itself, and therefore took much longer to realize the robot was drifting. By the time these participants noticed the drift, they often corrected too late, leading to collisions, landings, or overcorrections, giving participants the impression that the robot was difficult to control. Robot Teleoperation with a Virtual Surrogate Implementation
[0110] We conducted a 3 × 1 within-participant experiment to evaluate how the RVS and WVS designs affect user experience when teleoperating a co-located flying robot. The study protocol involved participants navigating an AscTec Hummingbird quadcopter throughout the experimental environment, visiting six points of interest from which they "collected data" while simultaneously completing a quiz mimicking the task of operating the robot while simultaneously analyzing the data collected by the robot on the fly. The independent variable in this study corresponded to the type of teleoperation interface participants used: a baseline teleoperation system in which Xbox controller input directly controlled the physical robot, a real-time virtual surrogate design, or a waypoint virtual surrogate system. Dependent variables included objective measures of mean completion time, response time, and interface usage, as well as subjective rankings directly comparing each interface, and overall ratings of perceived multitasking ability, stress, and ease of use.
[0111] The experiment, inspired by similar use cases for aerial robots in disaster response, construction, and space exploration, represents a scenario in which a user remotely controls a line-of-sight aerial robot to collect and analyze environmental data. The experimental environment in which participants controlled the aerial robot measured 6m x 10m x 6m and included several motion-tracking cameras used as part of the back-end system to locate the robot and ensure safe operation.
[0112] The experimental task required users to complete two subtasks: (1) visiting a set of "points of interest" (POIs) and Participants were instructed to (1) pilot the aerial robot to "collect data" and (2) periodically complete a quiz demonstrating the concept of analyzing the data the robot had just collected. POIs were designated by six stools placed within the environment. Participants were instructed to pilot the robot to visit each stool in a specific order and maintain a stable hover over the stool for 5 seconds to simulate environmental data collection. The stools were arranged in two rows of three. In each row, the distance between the first and second POI was 2 meters, while the second and third POIs were 4 meters apart. To examine the user experience of operating the robot across various distances, the environment was divided by tape into two zones: a user-permitted area and a user-restricted area. As a result, sometimes (when visiting stools within the user-permitted area) users were able to operate the robot from closer range, but other times (when visiting distant stools within the user-restricted area) users were forced to operate the robot from a greater distance.
[0113] In all conditions, while remotely controlling the robot, participants wore a Microsoft HoloLens ARHMD that used augmented reality imagery to present the order in which to visit POIs; the POI order was designated by a virtual number (1–6) that appeared above each stool. In all conditions, participants also received AR feedback in the form of a virtual progress bar that filled as the participant hovered over a POI, indicating the robot was "collecting data." If the participant prematurely left the POI (designated by a virtual cylinder outlining the POI area), their progress was lost and they had to reposition the robot so that it was within the POI radius, and "data collection" (i.e., the progress bar) would restart.
[0114] Upon completing each POI, participants were presented with a data analysis subtask and required to answer two multiple-choice quiz questions. These questions simulated the concept of participants analyzing the "data" the robot had just collected at the POI. The quiz questions were displayed on a smartphone attached to the user's wrist (the user's preferred wrist), again simulating the idea of using a robot in the field to collect and analyze data. Each quiz question presented the user with a sentence of approximately 25 characters and required the user to select an answer from four options corresponding to the number of vowels contained in the sentence. Participants were not forced to complete the quiz immediately; instead, they could continue piloting the robot to new POIs and complete the "data analysis" subtask whenever they liked. However, each completed POI added another two questions to the quiz queue. Successful completion of the entire task required participants to collect data from all POIs and answer all quiz questions.
[0115] In this experiment, an AscTec Hummingbird quadcopter was used as the aerial robot. The Hummingbird is a popular research "drone" suitable for indoor and outdoor flight. Various embodiments may use a Microsoft HoloLens as the ARHMD and the Unity game engine, which is used to develop and deploy applications. HoloLens is a wireless, optical, see-through, stereoscopic augmented reality HMD with a 30° x 17.5° FOV, an inertial measurement unit, a depth sensor, an ambient light sensor, and multiple cameras and microphones that support voice input, gesture recognition, and head tracking. HoloLens was selected due to its growing popularity, ease of access, ability to support hands-free AR, and potential as a model for future consumer ARHMD systems.
[0116] A total of 18 participants (11 males and 7 females) participated in the study. The population sample included both novice users and those with experience piloting aerial robots. Overall, 7 participants represent expert users recruited from a local "Drone Club." Eight participants reported some familiarity with aerial robots, while three participants had little or no experience operating flying robots. The mean age of participants was 20.7 years (SD = 3.59), with a range of 18 to 27 years.
[0117] The study took approximately 80 minutes and consisted of four main phases: (1) introduction, (2) prototype evaluation (which had six subphases and was repeated three times for each participant), (3) summary evaluation, and (4) conclusion. In the first phase, (1) participants signed a consent form, were guided to the experimental environment, and read an instruction sheet detailing the task and task rules. In the second phase, (2) participants completed the main experimental task (visiting points of interest (POIs) to "collect data" and answer quiz questions) using one of three interfaces (baseline, real-time virtual surrogate, or waypoint virtual surrogate). Each participant completed this phase three times, using each interface once, and the order of the interfaces was coordinated across participants to mitigate potential transition effects, such as learning or fatigue, that could occur due to the within-participant design.
[0118] This phase consisted of six subphases: (A) Participants first watched a short 60-second tutorial video presenting the interface design they would use, covering both the controls and what the visual feedback would look like. (B) Next, the ARHMD application was initiated, calibrated, and fitted to each participant, and the researcher verbally confirmed that participants could see the augmented reality images as intended. (C) Next, participants were given two minutes to test the interface and become familiar with the controller, augmented reality images, and robot. (D) Next, participants performed the main experimental task, in which they piloted the aerial robot to a series of POIs, hovering for 5 seconds at each, and then completing a quiz. The order of the POIs was kept constant to control for potential variability between trials. As described above, participants were free to complete the POIs and quizzes in any order they chose (e.g., participants could complete all POIs and then all quizzes, or they could complete quizzes simultaneously while piloting the robot), but participants could not start a quiz before completing the POI corresponding to that quiz. Participants were instructed that their goal was to complete all tasks (visit all POIs and complete all quizzes) as quickly as possible. (E) After completing all tasks, the experimenter administered a survey surveying participants regarding the interface design they had just used. After completing the survey, participants repeated this entire phase two more times. (4) After completing the main tasks using each interface design, participants completed the entire task one last time as part of a summary evaluation. During this phase, participants were free to switch between any interface design at any time as many times as they wished, allowing them to record objective data regarding their user preferences for the interface they chose to use. (5) After completing the summary evaluation, participants were given a final post-survey that collected data regarding their overall experience and subjective rankings comparing each interface.
[0119] Both objective and subjective measures were used to evaluate the designs. Several objective aspects of task performance were measured, including completion time, measured in seconds by the time elapsed from starting the task to completing the final data analysis quiz; response time, which was the average time elapsed in seconds from scanning a POI to completing the associated data analysis quiz for all six points of interest; and design usage, measured by the percentage of total task time participants used each interface design during a summary evaluation phase in which participants completed the task while freely switching between designs.
[0120] Data was also collected from several subjective measures. After using each interface, participants completed the industry-standard 10-item System Usability Survey. The User Interface Usability Scale (SUS) was used to assess perceived interface usability. SUS scores below 68 are considered below average, scores above 68 are considered above average, and scores above 80.3 are considered within the top 10 percentile. In addition to the SUS, several scales were constructed from 7-point Likert-style survey items to measure participants' perceptions and preferences. The scales ranked perceived ease of operation (2 items, Cronbach's a = .77), ease of accurate positioning (2 items, Cronbach's a = .77), multitasking ability (3 items, Cronbach's a = .94), and stress (5 items, Cronbach's a = .91). Following the summary evaluation phase, participants were asked to directly compare the three designs and rank them relative to each other (1 (best) to 3 (worst)). Participants rated the designs in terms of ease of learning and their willingness to use them in the future. Finally, qualitative feedback was obtained through open-ended questions administered to each participant as part of various surveys. The questions include (but are not limited to) "what made performing the job easier" These included "what made the task easier," "what made performing the task harder," and "how did this design impact your ability to control the drone." The objective measures, SUS, and constructed rating scales were analyzed using repeated measures analysis of variance with experimental condition (i.e., interface design) as a fixed effect, and the order of conditions was included as a covariate to control potential variance that may arise from the ordering of effects. Post-hoc tests used Tukey's Honestly Significant Difference (HSD) to control for Type I error when comparing the effectiveness between each interface. Participants' rankings of each interface were analyzed with the nonparametric Kruskal-Wallis test with experimental condition as a fixed effect. Post-hoc comparisons used Dunn's test to analyze specific design-sample pairs for random significance.
[0121] FIG. 10 shows the objective results, where the RVS and WVS systems demonstrated improvement over baseline in all objective measures.
[0122] Objective Results - Task performance metrics were analyzed to determine whether the AR surrogate design helped participants more effectively teleoperate a co-located aerial robot. A significant main effect of robot interface design was found for task completion time, F(2,45) = 13.65, p < .001. Using Tukey's HSD, it was found that the real-time virtual surrogate (M = 186.39 s), p = .001, and pass-through virtual surrogate (M = 184.39 s), p = .001, designs significantly improved completion time compared to the baseline interface (M = 260.11 s). A significant main effect of design was found for response time, F(2,45) = 8.43, p < .001. Post-hoc comparisons with baseline (M = 90.56 s) revealed that participants were significantly faster at responding to the data analysis quiz using the RVS (M = 47.44 s), p = .004, and WVS (M = 44.61 s), p = .002, designs. Finally, a significant main effect was found for the rate of design use during the final summary assessment task, F(2,51) = 34.92, p < .001. Tukey's HSD revealed that participants used the WVS (M = 81.94%) significantly more than the virtual surrogate (M = 18.06%) and baseline (M = 0%) designs (all comparisons at p < .001), with not a single participant ever using the baseline design at any point.
[0123] Subjective Results - Perceived interface usability was assessed using the SUS. A significant effect of interface design was found on the SUS total score, F(2,45) = 5.91, p = .005. Tukey's HSD showed that both AR designs performed significantly better than the baseline. The results revealed that participants rated the design as significantly easier to use than the baseline (RVS: M = 86, p = .008; WVS: M = 83.6, p = .022; baseline: M = 66.8). Participants rated the design in terms of ease of distal manipulation. A significant main effect was found for ease of distal manipulation, F(2,45) = 5.56, p = .007. Tukey's HSD confirmed that both the RVS (M = 6.25), p = .016, and the WVS (M = 6.25), p = .016, were rated significantly higher than the baseline (M = 4.61). Finally, there was a significant main effect for ease of design and precise positioning, F(2,45) = 4.9, p = .012. Post-hoc analyses revealed that RVS (M = 6.31), p = .012, was rated significantly higher than baseline (M = 4.72), whereas WVS (M = 5.89), p = .075, was found to be insignificant at the a = .05 level. Participants also rated their perceived ability to perform multitasks using each interface. A significant main effect of interface design on these ratings was found, F(2,45) = 22.93, p < .001. Tukey's HSD indicated that both RVS (M = 5.42), p < .001, and WVS (M = 6.10), p < .001, were rated significantly higher than baseline (M = 3.00). The post-hoc survey also collected data on perceived stress. A significant main effect of interface was found for stress ratings, F(2, 45) = 6.87, p = .003; Tukey's HSD indicated that the baseline (M = 3.27) design caused more perceived stress than either AR design (RVS: M = 2.14, p = .013; WVS: M = 1.99, p = .004). Finally, following the overview evaluation phase, participants directly compared the designs to each other, ranking them from 1 (best) to 3 (worst) on how easily they found the interface to learn and which design they would like to use in the future. No significant main effect of how easily the design was learned was found on participants' rankings, H = 5.07, p = .079. However, a significant main effect was found for which design participants would like to use in the future, H = 16.85, p < .001.Post hoc analyses using Dunn's test found that participants consistently ranked both RVS (M = 1.89), p = .017, and WVS (M = 1.5), p < .001, higher than baseline (M = 2.61), but no significant differences were observed when comparing relative rankings between RVS and WVS.
[0124] Overall, the surrogate AR design enabled users to complete tasks and respond to data analysis quizzes faster than the baseline teleoperation interface, which was modeled after existing systems commonly used today. Both the RVS and WVS designs provided a preview of the robot's actions, which has the advantage that if users are satisfied with the preview, it frees up time to monitor the robot while completing other simultaneous tasks. This perspective is also supported by our subjective results, in which the surrogate design outperformed the baseline in terms of perceived ability to perform multiple tasks and a desire to use the surrogate interface in the future. Overall, both novice and expert participants agreed that the surrogate design provided a preview of the robot's movements and found the preview very helpful when operating the robot. These findings help support the hypothesis that providing support for the goal / action / evaluation cycle can improve teleoperation. Robot Intention Experiment
[0125] The inventors conducted a 5x1 between-participant experiment to evaluate how the design affected user interaction with a flying robot within a shared workspace. The independent variable in this study was the type of AR feedback received by the user (five levels: baseline and the four designs described above). In the baseline condition, participants still wore the ARHMD but did not view any virtual imagery. Instead, participants in this condition were informed that the robot had a clear "front" that always indicated the direction of flight. This baseline behavior meant that the robot always oriented itself in the direction of movement and utilized the only physically embodied cues provided by the robot's default morphology. All conditions shared this baseline orientation behavior. Dependent variables included objective measures of task performance and efficiency, as well as subjective ratings of communication clarity and ease of use of the robot.
[0126] The overall experimental setup was inspired by a potential future in manufacturing, where unmanned aerial robots might assist with logistics management. In the study, participants worked with an aerial robot in a shared environment designed to mimic a small warehouse. The environment measured 20 ft x 35 ft x 20 ft and included motion-tracking cameras for precise robot navigation to ensure participant safety. Six workspaces were arranged in a three-by-two pattern within the physical space. Each workspace had at least 5 ft of open space around it and supported containers of colored beads. Each bead container held beads of only one color, corresponding to either green, black, yellow, white, blue, or red. While sharing the environment with the aerial robot, participants were tasked with collecting beads from these containers and tying them together to create bead strings. Participants were instructed that their goal was to make as many bead strings as possible in exactly eight minutes. Each completed string consisted of 25 beads. Individual instructions were provided for each string, describing the target color and the amount of beads to be used. For example, one string might require 10 blue beads, 5 red beads, and 10 green beads. Along with these instructions, participants were instructed about three additional rules. 1. Participants were only allowed to pick up one bead at a time and only allowed to attach beads to the string while at the work station. 2. Participants could collect the colors in any order, but once they selected a color, they had to stay in that color's location until they had strung all the beads of that color, as indicated by the string instructions (i.e., they could not mix colors). 3. The robot visited each workstation from time to time (ostensibly to monitor the supply of beads). If the robot flew into a workstation where a participant was working, the participant was asked to move at least 2 meters away from the workstation (i.e., back to the social distance informed by proxemics) and wait until the robot had left (i.e., giving the robot priority for the workstation) before continuing their work.
[0127] This task was designed to emulate an assembly task that might be found in a warehouse, with a shared resource (i.e., the workspace) between the user and the robot. Because the robot has priority in the workspace, the task required participants to understand and predict the robot's intentions to optimally plan their activities and maximize task efficiency.
[0128] The experiment used an AscTec Hummingbird robot as the unmanned flight platform (Figure 8). During the experiment, the robot autonomously flew to pre-programmed waypoints throughout the experimental environment using a PID controller that received input about the robot's position using a motion capture system. During the study, the researchers provided an emergency kill switch that could disarm the robot for safety, but this was never required. The experiment used a Microsoft HoloLens as the ARHMD. HoloLens is a wireless, optical, see-through, stereoscopic augmented reality HMD with a 30° x 17.5° FOV, an inertial measurement unit, a depth sensor, an ambient light sensor, and multiple cameras and microphones that support voice input, gesture recognition, and head tracking. HoloLens was chosen due to its growing popularity, ease of access, ability to support hands-free AR, and potential as a model for future consumer ARHMD systems.
[0129] A custom experimentation framework was developed to implement the designs, deploy them to HoloLens, and ensure that visualizations were properly synchronized with the robot's behavior. The four designs described above were prototyped using Unity, a popular game and animation engine for designing and developing virtual and AR applications. A waypoint system was also developed that allowed for the specification of a sequential list of desired robot destinations (i.e., the location and orientation of the target robot in six degrees of freedom space), the desired travel speed to each destination, and a wait time (possibly zero) at each destination. An invisible virtual drone object was added to the Unity scene that navigated the scene according to the waypoint list specification, and its movement controlled the flight of the physical robot.
[0130] With each iteration of the game engine's update loop, the virtual drone's 3D pose is sent from the HoloLens via UDP to the back-end system controlling the physical robot. The virtual drone's pose values are transformed from Unity coordinates to real-world coordinates using a transformation matrix pre-calculated in the initial calibration procedure. After this transformation, the pose values in the Unity scene correspond to the same location in the user's environment, allowing the Unity application to point to a waypoint delegation interface, allowing the physical robot to fly through the learning space in a manner consistent with the virtual drone. The waypoint list was populated by randomly selecting the workspace destination, wait time, and speed from the ranges listed above to ensure there were no flight path patterns for participants to learn. This list was entered once before the start of the study as a whole and then kept constant for all participants across all conditions to remove potential variations resulting from different robot trajectories. The total time to traverse the waypoint list took 8 minutes to complete, which was equal to the total duration of the task (selected due to battery / flight time capabilities).
[0131] Heuristic parameters were also adjusted during the design implementation. NavPoints offers the possibility of displaying the complete flight path of all waypoints, but early pilot testing revealed that this led to cases of distracting exaggeration. Instead, only a portion of the robot's path was shown, which was dynamically updated so that the visualization always displayed the destinations of three waypoints ahead of the robot's current location. The arrow design was implemented so that the blue arrow always remained 15 seconds ahead of the aerial robot (i.e., indicating the robot's location 15 seconds later). The heuristic parameters selected for eye movement in the gaze design were derived from research on human gaze behavior and were set at a distance threshold of 2 meters and a fade threshold of 3 seconds. The radius of the eye model was set to 0.6 meters to ensure that the hologram covered the entire robot's body. The "radar" in the utility design used a 15-meter detection radius to cover the entire area of the experimental environment.
[0132] A total of 60 participants (40 males, 20 females, balanced across conditions) participated in the study. The mean age of participants was 20.7 years (SD = 4.8), ranging from 18 to 45 years. Using a 7-point scale, participants reported some prior knowledge of aerial robots (M = 3.75, SD = 1.71) but little knowledge of ARHMDs (M = 2.65, SD = 1.83). The study took approximately 30 minutes and consisted of five phases: (1) introduction, (2) training, (3) calibration, (4) work, and (5) conclusion. (1) First, participants signed a consent form and were guided to the workspace. (2) Next, participants read identical instruction sheets detailing the work and work rules. Participants assigned to one of the four AR design conditions watched a corresponding 60-second tutorial video that provided brief instructions regarding the AR feedback they would receive based on the relative novelty of the ARHMD technology.
[0133] Only participants assigned to the baseline condition were verbally informed that the robot would always move toward the direction its marked "front" was facing. (3) Next The ARHMD application was initiated, calibrated, and adapted to each participant, including those assigned to the baseline condition (even if they did not receive AR feedback). (4) Next, participants performed the main task for 8 minutes, creating as many bead strings as possible while sharing the environment with the co-located aerial robot. (5) Once the 8-minute task was complete, participants were asked to stop and received a post-survey about their experience.
[0134] A combination of four objective and subjective measures was used to characterize the design's effectiveness. Objective work efficiency was measured by the total time participants spent waiting while interrupted by the robot, which could be avoided by understanding the robot's intentions and planning their own tasks (less time indicates better performance / efficiency). When calculating work efficiency, the time variance of users who avoided the workspace was removed. The interruption timer was started only when the robot was directly above the workspace. This allowed for consistent measurements across participants. To measure subjective participant perceptions and preferences, several scales were constructed from 7-point Likert-style questionnaire items. The scales ranked the clarity of the interface design (4 items, Cronbach's α = .85), perceptions of the robot as a teammate, both as an individual's work partner (4 items, Cronbach's α = .91), and as a potential work partner for others (2 items, Cronbach's α = .83), as well as the overall ease of use of the design (2 items, Cronbach's α = .73). Qualitative feedback was obtained through open-ended questions administered to each participant as part of an exit questionnaire to describe their experiences "working with the AR user interface," "working alongside the aerial robot," and "completing their task." Data were analyzed using a one-way analysis of variance (ANOVA) with experimental condition (i.e., interface design) as a fixed effect. Post-hoc tests used Dunnett's method to control for Type I error when evaluating the ARHMD design against the baseline condition, while Tukey's Honestly Significant Difference (HSD) tests compared validity across designs.
[0135] Figure 11 shows the objective results that NavPoints, arrows, and gaze improved task performance by reducing inefficiency and wasted time. The subjective results reveal that NavPoints outperformed other designs in terms of user preference and robot perception.
[0136] Objective Results - Task performance metrics were analyzed to confirm that the designs were useful for participants to quickly and accurately infer the robot's intentions and plan their activities more effectively. A significant main effect of ARHMD interface design was found for total time spent interrupting, F(4,55) = 12.56, p < .001. Comparing the performance of each design to the baseline using Dunnett's multiple comparisons test showed that using NavPoints (p < .001), arrows (p < .001), and gaze (p = .003), but not utilities (p = .104), significantly reduced the total time lost to interruption.
[0137] Subjective Results - Participants rated several aspects of the robot's communication of its behavioral intentions. A significant effect of design was found for perceived clarity of communication, F(4,55) = 11.04, p < .001. Post-hoc comparisons using Dunnett's test revealed that the NavPoints design was rated significantly higher than the baseline (p < .001), but no significant effects were found from the other designs. Participants' responses to the robot were analyzed in terms of how they viewed it as a collaborative partner within the work environment. A significant main effect of design margin (0.1 > p > 0.05) was found for participants' perception of the robot as a good work partner for themselves, F(4,55) = 2.48, p = .054. A significant main effect of design was also found for participants' perception of the robot as a good work partner for others, F(4,55) = 2.54, p = .049. Post-hoc comparisons revealed that NavPoints was the only design to significantly improve participants' perception of the robot as a good work partner for themselves (p = .03) and others (p = .029) over baseline. Finally, designs were compared to each other along the usability metric of how the displayed virtual images affected participants' understanding of the robot's movement intent. A significant main effect of design was found for perceived ease of use for understanding intent, F(3,44) = 25.32, p < .001. Post-hoc comparisons using Tukey's HSD found that NavPoints (M = 6.96), p < .001, arrows (M = 6.67), p < .001, and gazes (M = 5.83), p < .001, were ranked as significantly more useful than gazes (M = 4.21). NavPoints were also found to be significantly more useful than gazes, p = .012, and arrows were rated slightly more useful than gazes, p = .092.
[0138] The NavPoints, arrows, and gaze designs improved task performance by reducing inefficiencies, allowing participants to better predict the robot's intentions and plan their actions accordingly, reducing the amount of time they spent interrupted and unproductive. However, the utility model did not provide a similar improvement over the baseline condition. This may be due to the utility design emphasizing the robot's current positioning relative to the user, rather than displaying cues that help users predict the robot's future destination, as in other designs. Participant responses support this conclusion and reveal similarities between baseline and utility participants. Another possible cause of the utility design's poor performance may be the small scale of the task. In our experimental scenario, there was only one robot. This allowed participants to constantly face and listen to the robot while working in the workspace and navigating the environment. When a space is shared by two or more robots, it is unlikely that participants would be able to simultaneously track all robots using visual and audio alone. In this case, the utility design may scale appropriately and conservatively support tracking all nearby robots even better than some other designs. However, for the single robot in this experiment, Navpoints, Arrows, and Gaze all performed significantly better than the baseline in terms of reducing inefficiency. Participants often noted that explicitly visualizing the robot's movements made the task easier. Overview of an Exemplary Computer System
[0139] Aspects and implementations of the monitoring platform of the present disclosure have been described in the general context of various steps and operations. These various steps and operations may be performed by hardware components or may be embodied in computer-executable instructions that may be used to cause a general-purpose or special-purpose processor (e.g., within a computer, server, or other computing device) that is programmed with the instructions to perform the steps or operations. For example, the steps or operations may be performed by a combination of hardware, software, and / or firmware.
[0140] 12 is a block diagram illustrating an exemplary machine representing a computer implementation of a monitoring system. A system controller 1200 communicates with one or more users 1225, client / terminal devices 1220, user input devices 1205, peripheral devices 1210, optional co-processor devices (e.g., crypto-processor devices) 1215, and and a network 1230. A user can engage with the controller 1200 via a terminal device 1220 over the network 1230.
[0141] A computer can process information using a central processing unit (CPU) or processor. The processor may include a programmable general-purpose or special-purpose microprocessor, a programmable controller, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), an embedded component, a combination of such devices, etc. The processor executes program components in response to user- and / or system-generated requests. One or more of these components may be implemented in software, hardware, or both hardware and software. The processor passes instructions (e.g., operational instructions and data instructions) to enable various operations.
[0142] The controller 1200 may include, among other things, a clock 1265, a CPU 1270, memory such as read-only memory (ROM) 1285 and random access memory (RAM) 1280, and a coprocessor 1275. These controller components are connected to a system bus 1260, which may be connected to an interface bus 1235 via the system bus 1260. Additionally, a user input device 1205, a peripheral device 1210, a coprocessor device 1215, etc. may be connected to the system bus 1260 via the interface bus 1235. The interface bus 1235 may be connected to several interface adapters, such as a processor interface 1240, an input / output interface (I / O) 1245, a network interface 1250, and a storage interface 1255.
[0143] The processor interface 1240 can facilitate communication between the coprocessor device 1215 and the coprocessor 1275. In one implementation, the processor interface 1240 can expedite encryption and decryption of requests or data. The input / output interface (I / O) 1245 facilitates communication between the user input device 1205, peripheral device 1210, coprocessor device 1215, etc. and components of the controller 1200 using protocols such as protocols for handling audio, data, video interfaces, wireless transceivers, etc. (e.g., Bluetooth, IEEE 1394a-b, serial, Universal Serial Bus (USB), Digital Visual Interface (DVI), 802.11a / b / g / n / x, cellular, etc.). The network interface 1250 can communicate with the network 1230. Via the network 1230, the controller 1200 can access a remote terminal device 1220. The network interface 1250 can use a variety of wired and wireless connection protocols, such as a direct connection, Ethernet, or wireless connections such as IEEE 802.11a-x.
[0144] Examples of network 1230 include the Internet, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a wireless network (e.g., using the Wireless Application Protocol WAP), a secure custom connection, etc. Network interface 1250, in some embodiments, may include a firewall that can administer and / or manage permissions for accessing / proxying data within a computer network and track changes in trust levels between various machines and / or applications. A firewall may regulate the flow of traffic between a particular set of machines and applications, between machines, and / or between applications, e.g., sharing traffic and resources between these various entities. The firewall may be any number of modules having any combination of hardware and / or software components capable of enforcing a predetermined set of access rights to access and manipulate objects. A firewall may further manage and / or access access control lists that detail permissions, including, for example, access and manipulation rights of individuals, machines, and / or applications, and the circumstances under which the permissions are subject. Other network security functions performed by or included in the functionality of a firewall may be, for example, but not limited to, intrusion prevention, intrusion detection, next-generation firewall, personal firewall, etc., without departing from the novel technology of the present disclosure.
[0145] The storage interface 1255 can communicate with several storage devices, such as the storage device 1290, a removable disk device, etc. The storage interface 1255 can use various connection protocols, such as Serial Advanced Technology Attachment (SATA), IEEE 1394, Ethernet, and Universal Serial Bus (USB).
[0146] User input devices 1205 and peripheral devices 1210 may be connected to I / O interface 1245, and potentially other interfaces, buses, and / or components. User input devices 1205 may include card readers, fingerprint readers, joysticks, keyboards, microphones, mice, remote controls, retina readers, touch screens, sensors, etc. Peripheral devices 1210 may include antennas, audio devices (e.g., microphones, speakers, etc.), cameras, external processors, communication devices, radio frequency identifiers (RFIDs), scanners, printers, storage devices, transceivers, etc. Coprocessor devices 1215 may be connected to controller 1200 via interface bus 1235 and may include microcontrollers, processors, interfaces, or other devices.
[0147] Computer-executable instructions and data may be stored in memory accessible by the processor (e.g., registers, cache memory, random access memory, flash, etc.). These stored instruction codes (e.g., programs) may involve the processor components, motherboard, and / or other system components to perform desired operations. The controller 1200 may use various forms of memory, including on-chip CPU memory (e.g., registers), RAM 1280, ROM 1285, and storage devices 1290. The storage devices 1290 may use any number of tangible, non-transitory storage devices or systems, such as fixed or removable magnetic disk drives, optical drives, solid-state memory devices, and other processor-readable storage media. The computer-executable instructions stored in memory may include the monitoring service 150 having one or more program modules, such as routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. For example, the memory may include an operating system (OS) component 1295, modules, and other components, database tables, etc. These modules / components may be stored and accessed from storage devices, including external storage devices, accessible via an interface bus.
[0148] The database component can store programs that are executed by a processor to process the stored data. The database component may be implemented in the form of a relational, scalable, and secure database. Examples of such databases include DB2, MySQL, Oracle, Sybase, etc. Alternatively, the database may be an array, hash, list, stack, structured text file, or the like. It may be implemented using various standard data structures (e.g., XML), tables, etc. Such data structures may be stored in memory and / or in structured files.
[0149] The controller 1200 may be implemented in a distributed computing environment where tasks or modules are performed by remote processing devices linked through a communications network, such as a local area network (“LAN”), a wide area network (“WAN”), or the Internet. In a distributed computing environment, program modules or subroutines may be located in both local and remote memory storage devices. Distributed computing may be used to load balance and / or aggregate resources for processing. Alternatively, aspects of the controller 1200 may be electronically distributed over the Internet or other networks (including wireless networks). Those skilled in the relevant art will recognize that portions of the monitoring service may reside on server computers and corresponding portions may reside on client computers. Data structures and data transmissions specific to aspects of the controller 1200 are also encompassed within the scope of the present disclosure. conclusion
[0150] Unless the context clearly requires otherwise, throughout the description and claims, words like "comprises," "comprising," and the like should be construed in an inclusive sense, i.e., "including but not limited to," as opposed to an exclusive or exhaustive sense. As used herein, the terms "connected," "coupled," or any variation thereof, mean either a direct or indirect connection or coupling between two or more elements, and the coupling or connection between the elements may be physical, logical, or a combination thereof. Furthermore, the words "herein," "above," "below," and words of similar import, when used in this application, refer to this application as a whole and not to any particular portions of this application. Where the context permits, words in the above detailed description using singular or plural numbers may also include plural or singular, respectively. The word "or," in connection with a list of two or more items, covers the following interpretations of that word: any of the items in the list, all of the items in the list, and all of any combination of the items in the list.
[0151] The above-described detailed description of examples of the present technology is not intended to be exhaustive or to limit the present technology to the precise form disclosed above. While specific examples of the present technology are described above for illustrative purposes, those skilled in the art will recognize that various equivalent modifications are possible within the scope of the present technology. For example, while processes or blocks are presented in a given order, alternative implementations may perform routines having steps or use systems having blocks in a different order, and some processes or blocks may be deleted, moved, added, subdivided, combined, and / or modified to provide alternative or subcombinations. Each of these processes or blocks may be implemented in a variety of different ways. Also, while processes or blocks are sometimes shown as being performed sequentially, these processes or blocks may instead be performed or implemented in parallel, or may be performed at different times. Furthermore, any specific numbers referenced herein are merely examples, and alternative implementations may use different values or ranges.
[0152] The teachings of the technology provided herein may be applied to other systems, not necessarily the system described above. Elements and operations of the various examples described above may be combined to provide further implementations of the technology. Some alternative implementations of the technology may include additional elements as well as fewer elements than the implementations described above.
[0153] These and other changes can be made to the technology in light of the above detailed description. While the above description describes specific examples of the technology and describes the best mode contemplated, no matter how detailed the above appears in text, the technology can be practiced in many ways. System details, while still encompassed by the technology disclosed herein, may vary considerably in its specific implementation. As noted above, specific terms used when describing particular features or aspects of the technology should not be taken to mean that the terms are redefined herein to be limited to any specific characteristic, feature, or aspect of the technology to which they relate. In general, the terms used in the following claims should not be construed to limit the technology to the specific examples disclosed herein unless the above detailed description section explicitly defines such terms. Thus, the actual scope of the technology encompasses not only the disclosed examples but also all equivalent ways of practicing or implementing the technology under the scope of the claims.
[0154] To reduce the number of claims, certain aspects of the present technology are presented below in specific claim forms, but Applicant contemplates various aspects of the present technology in any number of claim forms. For example, while only one aspect of the present technology is recited as a computer-readable medium claim, other aspects may likewise be embodied as computer-readable medium claims or in other forms, such as means-plus-function claims. While any claim intended to be treated under 35 U.S.C. §112(f) begins with the words "means for," use of the term "for" in any other context does not seek treatment under 35 U.S.C. §112(f). Accordingly, Applicant reserves the right to seek additional claims to pursue such additional claim forms, either in the present application or in any pending application, after filing this application.
Claims
1. 1. A method of operating a robot, comprising: receiving data collected by the robot; the data collected from the robot includes information about (i) the robot's local environment, (ii) objects in the robot's local environment, and (iii) one or more robot states of the robot; receiving one or more commands from a user wearing a head-mounted display (HMD) instructing the robot to change the one or more robot states c; intercepting and directing the one or more commands to a virtual surrogate of the robot, wherein the virtual surrogate and the robot are visually communicated on an augmented or mixed reality view in the HMD, and the user controls the virtual surrogate rather than directly manipulating the robot; updating the augmented or mixed reality view in the HMD, analyzing the one or more commands to identify hazards that, when executed, may lead to at least one of damage to the robot, damage to the local environment, and violation of operating rules; automatically generating a set of one or more modified commands to avoid said hazard; updating the augmented or mixed reality view in the HMD to display the results of execution of one or more modified commands by the robot; sending the set of one or more modified commands to the robot to be executed, thereby updating the augmented or mixed reality view in the HMD; A method comprising:
2. The method of claim 1 , wherein the updated view displays waypoints of the robot's planned travel path, times of arrival at the waypoints, and times of departure from the waypoints.
3. Upon identifying the danger, the method further includes determining the user's intent by analyzing the one or more commands; The method of claim 2 , wherein generating the set of one or more modified commands maintains the user's intent.
4. Integrating the user's perspective with the data collected from the robot; visually providing to the user of the HMD information regarding future robot states, including position, velocity, time of arrival, and time of departure, resulting from one or more commands instructing the robot to change the one or more robot states; The method of claim 1 further comprising:
5. updating the augmented reality or mixed reality view in the HMD using augmented reality; The appearance of the robot; and The method of claim 1 , further comprising virtually augmenting at least one user interface for remotely controlling the robot.
6. 6. The method of claim 5, wherein updating the augmented reality or mixed reality view includes virtually augmenting the user interface with augmented reality, the user interface being virtually augmented to facilitate teleoperation of the robot using a peripheral design.
7. The virtual surrogate a real-time virtual surrogate of said robot, such that virtual teleoperation of the real-time virtual surrogate directs teleoperation of said robot; or a pass-point virtual surrogate for the robot, such that manipulation of a virtual pass-point using the pass-point virtual surrogate directs teleoperation of the robot; The method of claim 5 or 6, wherein the user interface is virtually augmented to include one or more of one or more non-virtual scene objects of an environmental context in which the robot is operating.
8. 10. The method of claim 1, further comprising tracking locations where the robot has previously traveled, and wherein updating the augmented or mixed reality view in the HMD comprises displaying data identifying locations where the robot has previously traveled.
9. 10. The method of claim 1, further comprising identifying a location of the robot, and wherein updating the augmented reality or composite view in the HMD includes a radar-like design to assist in remote operation of the robot.
10. The method of claim 1 , further comprising receiving robot state information and environmental data from data drops at locations identifiable by the robot.
11. 1. A system comprising a non-virtual robot that uses sensors or cameras to collect data as the non-virtual robot navigates a local environment, the system comprising: the data collected from the non-virtual robot includes information about the local environment of the non-virtual robot and one or more robot states of the non-virtual robot; an augmented reality system comprising an augmented reality display, a communications module, a processor, and a non-transitory computer-readable medium having instructions stored thereon; The instructions, when executed by the processor, receiving, via the communication module, at least a portion of the data collected by the non-virtual robot; receiving, via the communication module, one or more commands from a user instructing the non-virtual robot to change the one or more robot states; intercepting and directing the one or more commands to a virtual surrogate of the non-virtual robot, wherein the virtual surrogate and the non-virtual robot are visually communicated on the augmented reality display, and the user controls the virtual surrogate rather than directly manipulating the non-virtual robot; and updating an augmented reality view in the augmented reality display; analyzing the one or more commands to identify hazards that, when executed, may lead to at least one of damage to the non-virtual robot, damage to the local environment, and violation of behavior rules; automatically generating a set of one or more modified commands to avoid said hazard; updating the augmented reality view in the augmented reality display to display a result of execution of one or more modified commands by the non-virtual robot; instructing the communications module to transmit the set of one or more modified commands to the non-virtual robot to be executed, thereby updating an augmented reality view in the augmented reality display; The system causes the augmented reality system to perform the above.
12. further comprising one or more sensors associated with the augmented reality display; the one or more sensors are configured to monitor the one or more commands from the user teleoperating the non-virtual robot; and 12. The system of claim 11, wherein the communications module transmits the commands detected by the one or more sensors associated with the augmented reality display and transfers information about the one or more robot states that enables the augmented reality view to be updated.
13. one or more interfaces for receiving the one or more commands that change the one or more robot states of the non-virtual robot; a hazard detection module, under control of the processor, receiving the one or more commands that change the one or more robot states of the non-virtual robot; analyzing the commands to identify hazards that may damage the non-virtual robot, damage the local environment of the non-virtual robot, injure a co-located human, or violate operating rules when executed; and a hazard detection module that generates the set of one or more modified commands to avoid the hazard or violation of the operating rules; 1. An intent analyzer, comprising: Upon identifying the danger, determining the user's intent by analyzing the command; The system of claim 11 , further comprising: an intent analyzer configured to generate the set of one or more modified commands to maintain the user's intent.
14. When executed by the processor, the instructions further cause the augmented reality system to: updating the augmented reality view on the augmented reality display to display a result of execution of the one or more modified commands by the non-virtual robot; 14. The system of claim 13, wherein the communications module is used to transmit the set of one or more commands for the non-virtual robot to be executed and a robot state of the non-virtual robot to the augmented reality system, such that images in the augmented reality display can provide the user, or other users, with further insight into the non-virtual robot and the current robot state of the non-virtual robot.
15. the augmented reality display: context information, the expected robot path, Identified hazards, Zoom in or out on a specific area, one or more virtual objects in a combined real and virtual scene; one or more non-virtual objects in the combined real and virtual scene; and an actual camera feed of the non-virtual robot within the context of a virtual representation of the local environment of the non-virtual robot; The system of claim 11 , comprising at least one representation of:
16. the robot is a non-virtual robot, the non-virtual robot being remotely operated by the user using the augmented reality or mixed reality system to control the non-virtual robot; The method of claim 1 , wherein the virtual surrogate of the robot emulates the expected responses of the non-virtual robot to the one or more commands modified to avoid the hazard when executed.
17. The method of claim 16, wherein a violation of the operating rules causes the non-virtual robot to exceed a maximum height, fall below a minimum height, exceed a maximum speed, or fall below a minimum speed.
18. 20. The method of claim 17, further comprising broadcasting the future state of the non-virtual robot to a further augmented reality or mixed reality system before the one or more commands as modified by the modifying step are executed by the non-virtual robot.
19. The method of claim 1 , wherein the robot is an aerial robot.
20. The system of claim 11 , wherein the non-virtual robot is a non-virtual aerial robot.
21. The method of claim 16 , wherein the non-virtual robot and the virtual surrogate of the robot are aerial robots.
Citation Information
Patent Citations
Object testing method, apparatus and system
CN107004039A
Picking device
JP1995285622A
Input device and method using head motion
JP2010231290A
Robot system
JP2012171024A
Display control method, display control program, and information processing apparatus
JP2016184294A