System and method for detecting human gaze and gestures in unconstrained environments
The system addresses the limitation of existing human motion detection systems by using 360-degree imaging and depth sensing to detect human attention and identify objects in unrestricted environments, achieving accurate and efficient object recognition.
Patent Information
- Application Number
- JP2020557283
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-04-17
- Filing Date
- 2019-04-11
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2039-04-11
AI Technical Summary
Existing human motion detection systems, particularly those for gaze and gesture recognition, are limited to constrained environments and cannot accurately detect human attention or point to objects in unrestricted settings.
A system that uses a 360-degree image capture device and depth sensors to generate environment maps and determine the direction of human attention, projecting directional vectors onto the environment map to identify objects of interest and search for their identity in a knowledge base.
Enables accurate detection and identification of objects in unconstrained environments by combining 360-degree imaging, depth sensing, and machine learning to interpret human gestures and gaze, allowing for improved interaction with intelligent machines.
Smart Images

Figure 0007675970000001 
Figure 0007675970000002 
Figure 0007675970000003
Abstract
Description
[Technical field]
[0001] The present disclosure relates generally to detection of human activity, and more specifically, to human gaze or gestural activity in unconstrained environments by intelligent machines. [Background technology]
[0002] Machine detection of human motion has become widespread with the development of interactive computing, machine learning, and artificial intelligence. Desktop computers, smartphones, tablets, video game systems, and other devices with built-in or peripheral cameras often include the capability to detect human motion and even a user's gaze (e.g., the direction a user's eyes are looking).
[0003] Known gesture control and eye-tracking systems rely on narrow-field-of-view cameras and are designed for environments where the user's movements are constrained in some way (e.g., estimating the gaze of a driver or a person in front of a computer screen). These systems do not allow the user to move freely, and in most cases are unable to observe the same environment as the user does. In a more general environment, the user may teach the robot to engage (e.g., move or deliver) an object with other objects in the space, but these systems are not informative and cannot understand what the user is looking at or pointing at, since the camera does not capture the object of interest. Summary of the Invention
[0004] An embodiment of the present disclosure includes a system for detecting gaze and gestures in an unconstrained environment. The system may include an image capture device configured to generate a 360-degree image of the unconstrained environment including an individual and an object of interest. The system may further include a depth sensor configured to generate a three-dimensional depth map of the unconstrained environment and a knowledge base. The system may include a processor configured to generate an environment map from the 360-degree image and the depth map and determine a gaze direction of the individual in the 360-degree image. The processor may further be configured to generate a directional vector representing the gaze direction and project the directional vector from the individual onto the environment map. The processor may be configured to detect an intersection of the directional vector with the object of interest and use the data regarding the object of interest to search the knowledge base for an identity of the object of interest.
[0005] Additional features of the system may include a monocular panoramic camera as an image capture device. Alternatively, the 360-degree image may be stitched together from multiple images. The processor may be configured to determine the attention direction from at least one of the following: gaze, personal gesture direction, or a combination thereof. The depth sensor may be one of an infrared sensor, a neural network trained for depth estimation, or a simultaneous localization and mapping sensor.
[0006] In a further embodiment of the present disclosure, a method for locating and identifying an object is disclosed. The method may include capturing a 360-degree image of an unconstrained environment including an individual and an object of interest, and generating an environment map. A gaze direction may be determined from the individual in the captured 360-degree image, and a directional vector may be generated indicative of the gaze direction. The directional vector may be projected from the individual in the environment map, and an intersection of the directional vector with the object of interest may be detected.
[0007] Additional features of the system and method may include searching a knowledge base for the identity of the object of interest and initiating a learning process when the identity of the object of interest is not found in the knowledge base. The learning process may be a neural network. The environment map may be generated by acquiring depth information of the environment and combining the depth information with the 360-degree image. The depth information may be acquired from an infrared sensor, a simultaneous localization and mapping process, or a neural network. The environment map may be a saliency map and / or may include semantic parsing of the environment. The system and method may further estimate gaze and / or gesture direction of the individual, and the attention direction may be determined from the gaze and / or gesture. The gaze may be acquired from an infrared sensor.
[0008] Another embodiment of the present disclosure may include a system for locating and identifying an object. The system may include a monocular camera configured to capture a 360-degree image of an unconstrained environment including an individual and an object of interest. The infrared depth sensor may be configured to generate a 3D depth map of the unconstrained environment. The system may further include a knowledge base and a processor. The processor may be configured to generate a saliency map from the 360-degree image and the depth map to determine a direction of attention of the individual. The processor may be further configured to generate a three-dimensional directional vector representing the direction of attention and project the three-dimensional directional vector from the individual onto the environmental map. The processor may be further configured to detect an intersection of the three-dimensional directional vector with the object of interest and use the captured image data for the object of interest to search the knowledge base for an identity of the object of interest. [Brief description of the drawings]
[0009] Embodiments of devices, systems, and methods are illustrated in the figures in the accompanying drawings, which are intended to be illustrative and not limiting, and like references are intended to refer to like or corresponding parts.
[0010] [Figure 1] FIG. 1 is a conceptual diagram of an unconstrained environment illustrating a situation in which a system and method for detecting human gaze and / or gestural actions by an intelligent machine according to an embodiment of the present disclosure may be implemented. [Figure 2A] FIG. 1 is a conceptual diagram of an image capture device observing an individual's gestures, according to an embodiment of the present disclosure. [Figure 2B] FIG. 1 is a conceptual diagram of an image capture device observing an individual's gestures, according to an embodiment of the present disclosure. [Figure 3A] FIG. 1 is a conceptual diagram of a 360-degree panoramic image and the resulting saliency map generated by the system from 2D image and depth sensing data, according to an embodiment of the present disclosure. [Figure 3B] FIG. 1 is a conceptual diagram of a heat map, according to an embodiment of the present disclosure. [Figure 4] 1 illustrates a system for detecting human gaze and / or gestural movements, according to an embodiment of the present disclosure. [Diagram 5] 1 is a flowchart of a method for locating and identifying an object using gaze and gesture detection according to an embodiment of the present disclosure. [Figure 6] 11 is a further flow chart illustrating the refinement of gaze determination according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] The present disclosure describes a system and method for estimating and tracking human gaze and gestures to predict, locate and identify objects in unconstrained environments. The system may rely on advanced imaging techniques, machine learning and artificial intelligence to identify objects and their locations based on interpretation of human gestures, gaze, recognition voice and any combination thereof. Embodiments of the present disclosure are described in more detail below with reference to the accompanying drawings. However, the above may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein.
[0012] The present disclosure provides systems and methods for capturing an individual's surrounding environment in addition to a human's gaze or gestures to locate and identify objects. According to one embodiment, a 360-degree monocular camera may provide a two-dimensional ("2D") 360-degree image of the unconstrained environment. One or more infrared depth sensors may be used to create a depth map to enhance the captured scene to provide depth estimation parameters for individuals and objects in the unconstrained environment.
[0013] The system may process the 2D images and depth map to determine three-dimensional ("3D") vectors indicative of the gaze and / or gestures of the individual, such as the direction of the pointing finger. The system may generate a saliency map to define areas of the scene that are likely to be attended to. The saliency map may provide a technological advantage in predicting with greater precision the object or location at which the individual is looking or gesturing. The 3D vector may be projected onto the observed image of the environment to determine whether it intersects with any objects and locations identified in the environment surrounding the individual. Once the system has located and identified the object with sufficient certainty, the system may further interact with the object. If an object is located but unknown to the system, a machine learning process may be initiated to further train the system and expand the knowledge base.
[0014] FIG. 1 is a conceptual diagram illustrating an unconstrained environment 100 or a context in which a system and method for detecting human gaze and / or gestural actions by an intelligent machine is implemented according to an embodiment of the present disclosure. The gaze and gesture detection system may be implemented in a computing device such as a robot 105 equipped with an image capture device 110. The environment 100 may include one or more individuals 112, 113, 114, and a number of objects 115. An orientation individual 114 may look or gesturing specifically toward one of the objects 115. The image capture device 110 may capture the gaze direction 121 of the individual 114. The image capture device may also capture or detect a gesture 122 by the orientation individual 114. The gesture detection information may be used to supplement or confirm the direction of attention of the pointing individual 114. 1, the plurality of objects 115 in the unconstrained environment 100 may include a first window 117, a second window 119, and a painting 118 between the two windows 117 and 119. The painting 118 may be hung on a wall in the unconstrained environment 100 or may be placed on a stand or easel.
[0015] As used herein, the term "unconstrained environment" may include any space in which an individual's actions or behavior are not constrained by the spatial field of view of a camera or image capture device 110. In contrast to a driver constrained to a seat in a car, or a computer user constrained directly in front of the field of view of a computer camera, an individual in the illustrative unconstrained environment 100 may move freely throughout the area and still be within the field of view of the image capture device 110.
[0016] The computer or robot 105 may include a 360-degree monocular camera 111 as, or as a component of, the image capture device 110. The use of the 360-degree monocular camera 111 provides a substantially complete view of the unconstrained environment 100. Thus, regardless of the location of the individuals 112, 113, 114, or the objects 115, the individuals 112, 113, 114, or the objects 115 may be captured by the camera 111. Thus, the individuals 112, 113, 114 and the robot 105 may move within the environment 100 and still the camera 111 may capture the individuals 112, 113, 114 and the objects 115 within the unconstrained environment 100.
[0017] The system may further include additional environmental data gathering capabilities to gather depth information of the unconstrained environment 100. The system may generate a depth map by capturing environmental information related to a number of objects 115, individuals 112, 113, 114, and contours of the environment such as walls, pillars, and other structures. The robot 105 and image capture device 110 may include one or more infrared ("IR") sensors 120 that scan or image the environment 100. Additionally, the robot 105 and image capture device 110 may use the displacement of the IR signal to generate a depth image that may be used by the system to estimate the depth of objects or individuals in the environment from the image capture device 110.
[0018] If the robot 105, or other computing device, is a mobile robot, depth information of the environment may be obtained using simultaneous localization and mapping ("SLAM") techniques. SLAM techniques provide a process by which a mobile robot may build a map of an unknown environment as it moves through the environment. In one configuration, SLAM techniques may use features extracted from an imaging device to estimate the robot's location in an unconstrained environment. This may be done in the context of the robot 105 mapping the environment, such that the robot 105 knows the 3D relative positions of objects in the environment. This may include objects that the robot 105 has previously observed, but that are now occluded by other objects or scene structures.
[0019] According to one embodiment of the present disclosure, environmental depth information may be generated using an artificial neural network trained for depth estimation in an unconstrained environment, such as the unconstrained environment of FIG. 1. The artificial neural network may generate environmental depth information from or in addition to the methods and systems discussed above. In one configuration, the artificial neural network may be trained using object information. When based on a system with no prior experience of the objects or their environment, the artificial neural network may individually acquire unique defining features for each object or individual in the environment.
[0020] The artificial neural network may be trained for one or more specific tasks. Tasks may include, but are not limited to, person detection, human body pose estimation, face detection, object instance segmentation, recognition and pose estimation, metric estimation of scene depth, and segmentation of the scene into further attributes such as material properties and descriptive descriptions. The system may observe new objects or structures that the system cannot confidently recognize or segment. In such cases, the system may be able to use the extracted attributes and depth to obtain unique defining features for each object individually.
[0021] In addition to mapping contours and contours, objects, and individuals in the unconstrained environment 100, the system may further include the ability to capture or detect gaze information of the individual 114 performing the gesture. Gaze information may be obtained using an IR sensor, an RGB camera, and / or one or more optical tracking devices in communication with the robot 105. For example, an IR beam is transmitted by the IR sensor 120 and reflected from the eye of the individual 114 performing the gesture. The reflected IR beam may be detected by a camera or optical sensor. As described below, in further refinements, gaze data may be observed over a period of time using multiple images, scans, or the like to further acquire gaze direction or objects of interest in the field of view.
[0022] The image information is then analyzed by a processor to extract eye rotation from changes in reflection. The processor may use the corneal reflection and center of the pupil as features over time. Alternatively, reflections from the anterior cornea and posterior lens may be tracked as features. As part of the depth map, or captured and processed separately, the system may use IR images of the eyes of the individual 114 performing the gesture to determine the individual's gaze direction. As described below, the system may extrapolate the line of sight from the individual across the environmental image map, looking for the intersection of the projected line of sight and the object of interest.
[0023] In addition to observing and capturing the individual's position and eye position, the system may also detect gestures 122 (e.g., pointing poses) of the directional individual 114. The gestures 122 may be used to identify, aid in identifying, or confirm the direction of attention of the directional individual 114. As described below, the system may detect extended arms or hands. Based on the angle between the directional individual's 114 gesturing arm and body, the system may extrapolate a directional line or vector in the unconstrained environment 100.
[0024] 2A illustrates an example conceptual diagram 200 of an image capture device 110 observing the gaze 130 of a directional individual 114. In addition to detecting the gaze 130, the image capture device 110 may identify the pointing direction 124 or other directional gestures of the directional individual 114 to triangulate the location or object of interest of the directional individual 114. The system may analyze the image map from the IR sensor 120 or the 360-degree monocular camera 111. Based on the analysis, the system may apply a pattern or shape recognition process to identify the directional individual 114 and the pointing direction 124.
[0025] The direction individual 114, gaze 130, and gesture 122 may be used to extrapolate the gaze direction 126 from the gesturing individual 114. The system may detect both the gaze 130 and gesture 122 of the direction individual 114. The system may identify an extended arm or appendage that extends from a larger mass such as the body 128 of the direction individual 114. Data from the gaze direction and gesture direction may be combined. The system may use distances 136 and 137 from the image capture device 110 to the ends of the direction individual 114 and pointing direction 124 to determine a directional vector 132. As shown in the example of FIG. 2A, the directional vector 132 may be extrapolated from the eyes of the individual 114 to the direction of the gesture 122. The system may use an image map to determine that the directional vector 132 intersects with an object of interest such as a painting 118. An object of interest may be identified based on the intersection of the directional vector 132 with the object of interest.
[0026] FIG. 2B shows FIG. 2A as it may be interpreted by the system. As shown in FIG. 2B, the system may analyze an image map from an IR sensor 120' or a 360-degree monocular camera 111'. The directional vector 132' may be estimated, for example, by triangulating the body 128', line of sight 130', and gesture 122' of the individual 114' making the direction indication. The system may estimate the directional vector 132' using a distance 136' from the image capture device 110' to the eye of the individual 114' making the gesture and a distance 137' to the intersection of the line of sight 130' and the gesture 122'. The directional vector 132' may extend along the line of sight 130' to its intersection with the object of interest 118' (e.g., painting 118). While the examples shown in FIG. 2A and FIG. 2B may use both line of sight and gesture recognition, the system may use only line of sight to locate and identify the object of interest, or may use gesture information to supplement the accuracy of line of sight detection. Only gesture information may be used, for example, extrapolating a gesture vector 134' using only the gesture 122' and only in the gesture direction 122' from the individual 114', and still the gesture vector 134' intersects with the object of interest 118'.
[0027] The system may generate a saliency map to define areas of the scene that are likely to be attended to. The system may analyze the images and depth map to generate and project 3D vectors onto the saliency map that indicate gaze and / or gesture directions by the individuals. FIG. 3A is a conceptual diagram of a 360-degree panoramic image 300 and a saliency map 300' generated by the system based on 2D image and depth sensing data. The panoramic image 300 may be a 2D image. In this example, the panoramic image 300 depicts individuals 112, 113, 114, windows 117, 119, and painting 118 in a linear fashion as the image capture device scans the environment.
[0028] The saliency map 300' may be based on the 360-degree panoramic image 300 and an IR depth map. As mentioned above, the IR depth map may be acquired by an image capture device and an IR sensor. The saliency map 300' may be a topographically-distributed map indicating the saliency of an associated visual scene, such as a 360-degree image of an unconstrained environment. The image data and the depth sensing data may be combined to form the saliency map 300'. The saliency map 300' may identify areas of interest in the environment. The system may identify the individuals 112', 113', 114', the objects 117', 119' (e.g., windows 117 and 119), and the objects of interest 118' (e.g., painting 118) as areas of interest. For illustrative purposes, the objects or areas of interest in the saliency map 300' are depicted in opposite contrast in the panoramic image 300. The system may determine these objects and individuals as of interest due to the difference between the surrounding areas and known areas of non-interest.
[0029] For example, the system may be more interested in identifiable objects in the environment compared to walls, floors, and / or other locations that appear as large continuous areas in the image data. Areas of less interest, such as walls and / or floors, may be identified because they are continuous and homogenous pixel data that span large areas in the environment. Individuals and objects may be identified as objects of interest because their pixel data is different and contrasting compared to the pixel data of the surrounding areas. The system may analyze the saliency map 300' along with gaze and / or gesture information to identify objects or locations that are the focus of an individual's attention.
[0030] A 3D directional vector 138' may be generated by the system and projected onto the saliency map 300'. The directional vector 138' may be projected onto the rest of the observed scene, including any uncertainty boundaries. When the data is visualized into a 3D model, uncertainty due to noise and other factors may be present in the data. The uncertainty data may be provided along with other data to indicate the level of uncertainty (e.g., fidelity). The uncertainty boundary may be estimated from the field of view or the person's relative pose by a gaze detection system. For example, when a person is not looking in the direction of the camera, the uncertainty boundary may be large due to the lack of a direct line of sight between the camera and the individual's eyes. On the other hand, if the person is close to the camera and positioned parallel to the front, the uncertainty boundary may be smaller. The uncertainty boundary may effectively change the estimated 3D vector into a cone with the eyeball at the apex.
[0031] Gesture detection and related information may be used to narrow the uncertainty bounds. For example, gesture detection and information may be used to supplement analysis to locate and identify objects of interest when the individual making the direction gesture is not facing the camera and therefore the individual's eyes cannot be detected by the system. The system may be capable of generating a wide cone of attention based on the direction of the individual's head when the individual is not facing the camera. If a gesture is detected, the system may generate a gesture cone of attention based on the directionality of the detected gesture. The system may then determine the intersection or overlap of the gaze and gesture cones of attention. The intersection of the two cones of attention may provide a narrower 3D vector cone within which the system can locate and identify objects that may be of interest.
[0032] As discussed above, the system may use the depth map data to determine the distance from the image capture device to the gesturing individual 114'. Based on the determined distance, the system may project a directional vector 138' of the detected gaze or gesture direction from the location of the gesturing individual 114' to identify an intersection with the object of interest 118'. In the example of FIG. 3, the system may identify the object of interest 118' on the path of the directional vector 138'. Thus, the system may identify the object of interest 118' as an object of attention of the gesturing individual 114' based on the intersection with the directional vector 138'.
[0033] As previously described, the gesture information may be observed by the system and used by the system to locate and identify the object of interest 118'. The gesture vector 140' may be generated using information from the saliency map 300'. The gesture vector 140' may be based on the angle of the gesturing arm relative to the body of the individual 114' and may be extrapolated to a direction in the saliency map 300'. As shown in FIG. 3, the gesture vector 140' may still intersect with the same object of interest 118'. To improve accuracy and certainty, the system may use the gaze and gesture vector to triangulate or otherwise detect intersection with the object of interest.
[0034] As a further effort to reduce uncertainty, the system may refine gaze detection by generating a heat map of the direction individual's gaze, combining discrete gaze snapshots over a period of time. Not only are human eyes constantly moving, but the direction individual's attention may change. For example, the direction individual may momentarily gaze back to the robot, another individual, or another part of the object of interest. A single image capture of the direction individual's eyes may not be sufficient or accurate enough to identify the correct object of attention.
[0035] FIG. 3B is a conceptual diagram of a heat map 301 according to one embodiment of the present disclosure. Using a heat map 301 incorporating one or more images or scans of the directional individual's gaze over a period of time can provide an additional level of accuracy. The heat map 301 may combine data from a series of gaze images or scans acquired over a period of time t to form a gaze range 142'' as shown in FIG. 3B to identify areas the directional individual looks at most frequently. The areas 112'', 113'', 117'', 118'', 119'' shown in FIG. 3B indicate the density and / or frequency of image scans that may determine that the directional individual looks at that particular area. The system may estimate that during the scanning period t, the directional individual 114'' fixates on one area 118'' more frequently than other areas in the environment. According to the exemplary heat map 301, the most scanned gazes based on the density of scans may indicate the area of interest 118''.
[0036] Less frequent areas of interest may be detected in adjacent areas 117'', 119'', which represent windows 117 and 119, respectively. Areas 112'', 113'', which represent other individuals 112, 113 in the environment, may be shown as less frequent scanned data. The less frequent areas identified in the heat map may indicate that the directional individual 114'' momentarily looked at windows and other individuals in the environment during time period t. It should be understood that the representation of the directional individual 114'' in FIG. 3B is provided for illustrative purposes only to provide additional context and a frame of reference. It is unlikely that the directional individual would have gazed upon themselves during time period t when the system collected gaze data.
[0037] The heatmap 301 may be combined with a saliency map to determine the intersection of areas of interest in the heatmap with objects identified in the saliency map. The system may identify where objects in the environment are physically located (e.g., saliency map) and may also determine the direction in which an individual giving a direction is gazing or gesturing (e.g., heatmap). Combining the two sets of data may indicate the location of the object of interest with some level of certainty.
[0038] Once the system has identified an object of interest using any of the methods described above, the system may determine the identity of the identified object. The system may search a knowledge base (e.g., feature representations provided by a neural network) using the image and depth data acquired by the image capture device to associate the identified object of interest 118' with similar objects in the knowledge base. If a match is made with a certain degree of certainty, the system may associate the identified object of interest 118' with the matching object.
[0039] However, if the object of interest does not match, the system may begin a learning process to identify the identity of the object. For example, the system may indicate that the identified object of interest is unknown and prompt the directional individual or other operator to request information. The system's request may include an audio or visual prompt, or any other input request, which provides the system with information regarding the identity of the object of interest. The system may use machine learning techniques and user input to associate image data about the object with the identities of known objects, such that the object of interest (e.g., a painting) and observed objects with similar image data may become known to the system.
[0040] The system may use user input to provide identifying information to the knowledge base. For example, the system may use speech recognition and natural language understanding to obtain the name and other information of the object from the user input. For example, if the system cannot identify an object in the knowledge base, the system may output an audio or visual prompt for identifying information. In response to the prompt, the individual may provide an identification of the object by voice. The system may record an audio signal of the individual's identifying voice and may process the audio signal using a natural language processor to parse and convert it into machine understandable text. The system may then associate the correlating identifying information with the image data and store the association in the knowledge base for future recall. Additionally, the system may record additional information about the object, such as the object's location relative to the environment, the object's proximity to other identifiable objects, or the like. The additional data collected may assist the system in future queries or processing of the object in the environment.
[0041] The system may also have learned or be programmed with a library of commands, objects, and affordances for each object. The affordances may include various interactions that may occur between an individual or machine and other environments. As the system learns the identity of an object, it may associate the object with a set of affordances that provide additional context for how the system may locate, identify, and interact with a given object. For example, if the system has been trained to recognize a table, a specific affordance may be associated with that object for future interactions. Given the object identity of a table, the system may associate specific characteristics with the object. Examples of attributes may include, but are not limited to, a flat surface on which an item can be placed, a surface that is elevated above the floor, or the like. Thus, future interactions with the individual providing the direction may instruct the system to locate the object "on the table." The system may recognize the identity of the table, extract the relevant attributes, and perceive that the table includes a surface that is elevated above the ground.
[0042] Such affordances and attributes associated with the objects may be used to further filter the environment information such that the system first identifies the table and then locates the requested object in the small environment around the table (and its surface). The association of objects with attributes and features may be accomplished during training, semantic scene parsing, or other machine learning stages as described herein.
[0043] 4 illustrates a system 400 according to an embodiment of the present disclosure. The system 400 may include a dedicated computing device 405. As detailed above, the computing device 405 may be implemented as a mobile robot or may be a separate, integrated or distributed computing device, such as a special purpose computer designed and implemented to locate and identify objects of interest using human gaze and gesture detection. The computing device 405 may include data sources, servers, and / or client devices specifically designed for the system described herein.
[0044] The computing device 405 may include an interactive robot capable of interacting with humans. The computing device 405 may also be a device such as a laptop computer, a desktop computer, a personal digital assistant, a tablet, a mobile phone, a television, a set-top box, a wearable computer, and the like, capable of interacting with other devices or sensors via the network 402, but located in an unconstrained environment as previously described herein. In certain embodiments, the computing device 405 may be implemented using hardware or a combination of hardware and software. The computing device 405 may be a standalone device, a device integrated with other entities or devices, a platform distributed across multiple entities, or a virtual device running a virtual environment.
[0045] Network 402 may include data network(s) or internetwork(s) suitable for communicating data and control information between components of system 400. This may include public networks such as the Internet, private networks, public switched telephone networks, or communications networks such as cellular networks using third generation cellular technologies (e.g., 3G or IMT-2000), fourth generation cellular technologies (e.g., 4G, LTE, MT-Advanced, E-UTRA, etc.), WiMAX-Advanced (IEEE 802.16m), and / or other technologies, as well as any of a variety of enterprise, metropolitan area, campus, or other local area networks, or corporate networks with any switches, routers, hubs, gateways, and the like that may be used to transmit data between communicating components as described herein within computer system 400. Network 402 may include a combination of data networks and need not be strictly limited to public or private networks. The system may include external devices 404. External devices 404 may be computers or other external resources connected to computing device 405 via network 402.
[0046] Generally, computing device 405 may include a processor 406, memory 408, a network interface 410, data storage 412, and one or more input / output interfaces 414. Computing device 405 may further include or be in communication with wired or wireless peripherals 416 or other external input / output devices, such as a remote control, communication device, or the like, that may be connected to input / output interface 414.
[0047] Processor 406 may be any processor or processing circuit capable of processing instructions for execution within computer 405 or system 400. Processor 406 may include one or more single-threaded processors, multi-threaded processors, multi-core processors, or the like. Processor 406 may be capable of processing instructions stored in memory 408 or data storage 412 and providing the functionality described herein. Processor 406 may be a single processor or may include multiple processors that cooperate with system 400 and other processors to process instructions in parallel.
[0048] Memory 408 may store information within computer 405. Memory 408 may include volatile or non-volatile memory or other computer readable media, including, but not limited to, random access memory (RAM), flash memory, read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), registers, etc. Memory 408 may store program instructions, program data, executable files, and other software and data useful for controlling the operation of computer 405 and configuring computer 405 to perform functions for a user. Memory 408 may include many different levels and types of memory for different aspects of computer 405 operation. For example, a processor may include on-board memory and / or cache for fast access to certain data or instructions, and may also include a separate main memory or the like as needed to expand memory capacity. All of these types of memory may be part of memory 408 as contemplated herein.
[0049] Memory 408 may generally include a non-volatile computer-readable medium containing computer code that, when executed by computer 405, creates an execution environment for the computer program (e.g., code constituting a processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof) and performs some or all of the steps described or illustrated in the flowcharts and / or other algorithms described herein.
[0050] Although a single memory 408 is depicted, it will be understood that multiple memories may be usefully incorporated into the computer 405. For example, a first memory may provide non-volatile storage, such as a disk drive, for persistent or long-term storage of files and code, even when the computer 405 is powered off. A second memory, such as a random access memory, may provide a volatile (but faster) memory for storing instructions and data for an executing process. A third memory may be used to improve performance by providing faster memory physically adjacent to the processor 406 for use as registers, caches, etc.
[0051] Network interface 410 may include hardware and / or software that connects computing device 405 to communicate with other resources over network 402. This may include remote resources accessible over the Internet, or local resources available using a short-range communications protocol, such as using a physical connection (e.g., Ethernet), radio frequency communications (e.g., Wi-Fi), optical communications (e.g., fiber optics, infrared, or the like), ultrasonic communications, or a combination of these or other media that may be used to transport data between computing device 405 and other devices. Network interface 410 may include, for example, a router, a modem, a network card, an infrared transceiver, a radio frequency (RF) transceiver, a near-field communications interface, a radio frequency identification (RFID) tag reader, or any other data reading or writing resource, or the like.
[0052] More generally, the network interface 410 may include any combination of hardware and software suitable for connecting the components of the computing device 405 to other computing or communication resources. This may include, by way of non-limiting example, electronic devices for wired or wireless Ethernet connections operating in accordance with the IEEE 802.11 standard (or variants thereof) or other short-range or long-range wireless networking components or the like. This may include hardware for short-range communications such as Bluetooth or infrared transceivers that may be used to connect to other local devices or to connect to a local area network or the like that is in turn connected to a data network 402 such as the Internet. This may further include hardware / software for WiMAX connections or cellular network connections (e.g., using CDMA, GSM, LTE, or other suitable protocols or combinations of protocols). The network interface 410 may be included as part of the input / output interface 414, or vice versa.
[0053] Data storage 412 may be an internal memory storage device providing a computer-readable medium, such as a disk drive, optical drive, magnetic drive, flash drive, or other device capable of providing mass storage for computer 405. Data storage 412 may store computer-readable instructions, data structures, program modules, and other data for computer 405 or system 400 in a non-volatile form for relatively long-term persistent storage and subsequent retrieval and use. For example, data storage 412 may store an operating system, application programs, program data, databases, files, and other modules or other software objects, and the like.
[0054] The input / output interface 414 may support input and output to other devices that may be connected to the computer 405. This may include, for example, a serial port (e.g., an RS-232 port), a universal serial bus (USB) port, an optical port, an Ethernet port, a telephone port, an audio jack, a component audio / video input, an HDMI port, etc., any of which may be used to provide a wired connection to other local devices. It may also include an infrared interface, an RF interface, a magnetic card reader, or other input / output systems to wirelessly connect to communicate with other local devices. Although the network interface 410 for network communication is described separately from the input / output interface 414 for local device communication, it will be understood that the two interfaces may be the same or may share functionality, such as a USB port being used to connect to Wi-Fi accessories or an Ethernet connection being used to connect to storage devices connected to the local network.
[0055] Peripherals 416 may include devices used to provide information to or receive information from computer 405. This may include human input / output (I / O) devices such as a remote control, keyboard, mouse, mouse pad, trackball, joystick, microphone, foot pedals, touch screen, scanner, or other devices that may be used to provide input to computer 405. This may include displays, speakers, printers, projectors, headsets, or any other audiovisual devices for presenting information. Peripherals 416 may also include digital signal processors, actuators, or other devices that assist in the control or communication of other devices or components.
[0056] Other input / output devices suitable for use as peripherals 416 include haptic devices, three-dimensional rendering systems, augmented reality displays, and the like. In one embodiment, peripherals 416 may function as a network interface 410, such as a USB device configured to provide communication via short-range (e.g., Bluetooth, Wi-Fi, infrared, RF, or the like) or long-range (cellular data or WiMAX) communication protocols. In another embodiment, peripherals 416 may augment the operation of computer 405 with additional functionality or features, such as a Global Positioning System (GPS) device, security dongle, or other device. In another embodiment, peripherals 416 may include storage devices, such as flash cards, USB drives, or other solid-state devices, or optical drives, magnetic drives, disk drives, or other devices or combinations of devices suitable for mass storage. More generally, any device or combination of devices suitable for use with computer 405 may be used as peripherals 416 as contemplated herein.
[0057] The controller 418 may be responsible for controlling the operation or motion control of the computer 405. For example, the computer 405 may be implemented as a mobile robot in which the controller 418 may be responsible for monitoring and controlling all of the motion and motion functions of the computer 405. The controller 418 may be responsible for controlling motors, wheels, gears, arms or jointed limbs or fingers associated with the movement or operation of the computer's operational components. The controller may be coupled to the input / output interface 414 and peripherals 416, where the peripherals 416 are remote controllers capable of directing the operation of the computer.
[0058] The image capture device 420 may include cameras, sensors, or other image data gathering devices for use with the computing device 405's ability to capture and record image data related to the environment. For example, the image capture device 420 may include a 360-degree panoramic camera 422 suitable for capturing 2D images of the unconstrained environment. The image capture device 420 may further include a depth sensing device 424, such as a stereo camera, structured light camera, time-of-flight camera, or the like, that provides the computing device 405 with the ability to generate a depth map of the environment and generate a saliency map detailing individuals, gestures, and areas of interest within the environment along with the 2D panoramic image. The image capture device 420 may further include an additional gaze detection device, such as an optical observation device, or the like. Alternatively, the depth sensing device 424 may be capable of detecting gaze, for example using an IR camera and sensor, as described above.
[0059] The knowledge base 426 may serve as a learning center for the computer 405 as described above. The knowledge base 426 may be in the form of an artificial intelligence system, a deep neural network, a semantic scene parsing system, a simultaneous localization and mapping system, a computer vision system, or the like. The knowledge base may be coupled to the input / output interface 414 or to the network interface 410 for communication with external data sources via the network 402 for additional information regarding the identification and association of unknown objects. The knowledge base 426 may serve as a basis for identifying and associating objects of interest located near the system using image data of the environment. The knowledge base 426 may be coupled to the data storage 412 or memory 408 for storing and retrieving data of known and unknown identified objects and attributes associated with those objects.
[0060] Other hardware 428, such as a coprocessor, a digital signal processing system, a math coprocessor, a graphics engine, a video driver, a microphone, or speakers, may be incorporated into the calculator 405. Such other hardware 428 may include expansion input / output ports, additional memory, or additional drives (e.g., disk drives or other accessories).
[0061] The bus 430 or combination of buses may act as an electrical or electromechanical backbone interconnecting components of the computer 405, such as the processor 406, memory 408, network interface 410, data storage 412, input / output interface 414, controller 418, image capture device 420, knowledge base 426, and other hardware 428. As shown, the components of the computer 405 may be interconnected using a system bus 430 such that they are in a communicative relationship to share control, command, data, or power.
[0062] 5 is a flow chart illustrating a method 500 for locating and identifying an object using human gaze and gesture detection, according to an embodiment of the present disclosure. As previously discussed, the systems described herein may use environmental image data, human gaze information, and gesture information to locate and identify an object of interest to an individual. As shown in block 505, the system may receive a request to identify an object. For example, an individual may initiate a location and detection request using voice commands or haptic commands on a remote control or other interface in communication with the system.
[0063] When the system receives a voice or other request, it may parse the request using natural language processing, learned behaviors, or the like to further refine the localization and identification process. As previously described, the system may be trained or programmed with a saliency map and a library of available commands, objects, and affordances that refine the object localization and detection. The language of the request may provide the system with additional contextual information that allows the system to narrow the search area or environment. For example, if an individual providing direction makes a request to "locate objects on the table" and fixates on a box on the table, the system may parse the request and identify the table as a narrower environment in which the system may search for objects. The system may know that a table has an elevated flat surface on which objects may be placed. The system may exclude walls, floors, and other objects from the environment that do not share the known attributes of the table. The library of known and learned commands, objects, and affordances may provide a foundational knowledge base from which the system can begin the localization and identification process.
[0064] As shown in block 510, the system may capture images and depth information about the unconstrained environment. The system may capture 2D images of the unconstrained environment. The 2D environment images may be captured by a 360-degree panoramic camera, or a series of images taken by one or more cameras may be stitched together to form a 360-degree representation of the environment. The system may also capture depth images. The depth images may be obtained using an IR camera, sensor, or the like. The depth information of the environment may also be obtained using SLAM (e.g., for mobile robotic systems), or a deep neural network trained on depth estimation of unconstrained environments, such as indoor scenes. The depth information may be used to measure distances between the system, individuals, objects, and / or other definable environmental contours once the system detects them.
[0065] As shown in block 515, the system may construct or derive an environment map. The environment map may be based on 2D panoramic imagery and depth sensing data acquired by the system. Alternatively, the environment map may be derived from the system's memory if the environment has been previously mapped or uploaded to the system. The environment map may be generated as a saliency map that details the imagery and depth data as a more meaningful and easier representation of the environment for the system to interpret. Alternatively, the environment map may be generated by semantic parsing of the scene as objects and their identities (if known). The system may define areas of interest based on differentiated image data (e.g., identifying areas of differentiated pixels or image data compared to a typical homogenous pixel area).
[0066] Once the environment map has been constructed or derived, the system may locate and identify individuals as opposed to objects, and the individual's gaze, as shown in block 520. A bounding box may be generated around the identified individual's head and eyes. Within the bounding box, the individual's gaze may be detected using an IR sensor or optical / head tracker to form an estimated gaze vector for the purpose of locating and following the individual's eye movements.
[0067] The system may detect gaze by an individual, as shown in block 525. As a result of analyzing the environment map and identifying an individual, the direction of gaze of the individual making the direction indication may be detected. Gaze may be detected using an RGB camera, IR sensor, or the like to determine movement and the direction of gaze of the individual's eyes.
[0068] The system may then generate a directional vector, as shown in block 530. Analysis of the individual gaze and gestures may enable the system to project a directional vector onto the scene of the environment in an estimated direction of gaze and gestures by the individual across from the individual onto the environmental map. The directional vector may be estimated by extrapolation from the gaze and gestures by the individual in the environment.
[0069] As shown in block 535, the system may search for and locate objects of interest based on gaze and gestures by the individual. The system may analyze the extrapolated directional vector to determine whether the vector intersects with any objects or areas of interest identified in the saliency map. If the vector intersects with an area of interest in the saliency map, the system may determine that the area of interest is an object of interest for the individual.
[0070] The system may determine if the object is known, as shown in block 540. Once the object of interest has been located, the system may query a knowledge base, such as a neural network, an image database, or other data source, to compare known objects with associated image data as well as image data for the object.
[0071] If the object is known, the system may solicit further instructions from the individual, as shown in block 550. For example, the system may issue an audio or visual indication that the object has been located and identified, and the identity of the object. The system may request further instructions from the individual regarding the object, or regarding locating and identifying other objects. If the object is unknown to the system, as shown in block 545, the system may begin a learning process in which the system is trained by interacting with the individual and recording the image data and associated identity of the object in a knowledge base. The image data and associated identity may be recalled the next time the system is asked to identify an object with similar image data.
[0072] 6 shows a further flow chart for refining gaze determination according to an embodiment of the present disclosure. As shown in block 605 and as described above, the system may receive a request from an individual providing directions to locate and identify an object of interest. The request may be voice-based or initiated by a user interface or other controller located on the system or remotely. The system may programmatically or by advanced machine learning techniques parse complex, context-based requests to initiate search and location of the object of interest. The request may further include additional instructions to execute after the object of interest is identified. The system may again rely on programs or advanced knowledge bases to parse and execute the additional instructions.
[0073] As shown in block 610 and as described above, the system may capture sensor information about the unconstrained environment. To capture the environment, sensor data may be acquired by an RGB camera, an IR sensor, or the like. As shown in block 615, the system may use the captured sensor information to build or derive an environment map. The system may compare some or all of the newly acquired sensor data to other environment maps that the system may store in memory. If the environment is known to the system, a previously stored environment map may be retrieved. However, if the environment is unknown to the system, the system may register the environment map in memory for use in subsequent operations.
[0074] As shown in block 620 and described herein, the system may also locate and identify individuals in the environment based on the captured sensor data. Pattern recognition, facial recognition, or the like may be implemented to locate and identify the presence of individuals in the environment. If there is more than one individual in the environment, the system may implement additional processing to determine which individual is the individual providing the direction. To locate the object of interest of the individual providing the direction and identify where the individual is looking, the system may analyze data about the individuals to determine if one of the individuals is looking directly at the system, making a gesture, or has a particular characteristic that the system can be programmed to locate. For example, the individual providing the direction may be holding a remote or other controller that the system can identify, which may indicate that the holder is the individual providing the direction. If the system is unable to identify the individual providing the direction, the system may issue a prompt for further instructions.
[0075] As shown in block 625, the system may estimate the directional individual's activities, gaze, and / or gestures. As described above, the system may detect the directional individual's gaze by scanning eye movements and processing where the directional individual is looking. Gaze detection and tracking may be supplemented with other observable information to give the system additional contextual information for continuing the analysis. For example, the system may make observations to determine whether the individual is engaged in a particular activity, posture, or position to refine the environment in which the system is searching for the object of interest. In addition, as described above, the system may provide additional information by detecting the directional individual's gestures. From these observations, the system may form a set of parameters that the system uses to filter the environmental information into narrower ranges of interest to locate the object of interest.
[0076] As shown in block 630, the environment map, including image and depth information, may be used to generate a saliency map. The saliency map may be generated using the techniques described above, including semantic scene parsing, or other image processing techniques. Additional information acquired by the system may aid in the generation of the saliency map. Observations of the individual associated with block 625 may be used to further develop or constrain the saliency map. If the individual providing the direction is observed to further refine the environment, the system may only be interested in that restricted area. For example, if the individual providing the direction is gesturing towards a portion of the environment, the system may limit the saliency map to that area of the environment.
[0077] In addition, the system may be able to use the contextual information and affordances extracted from the original request in block 605 to further restrict the environment. For example, if the request included an identification of an area of the environment, such as a corner, or included the contextual command "open the window," the system may restrict the environment and saliency map to only objects on the wall. In this way, the system can generate a focal saliency map with a more limited region of possible objects or areas of interest.
[0078] As shown in step 635, the system may generate a heat map of attention based on the observed turning individual's activity and gaze. The heat map of attention may be generated by the system using multiple scans of the turning individual over a period of time. As mentioned above, due to the natural movement of the eyes and the tendency of the turning individual to gaze at multiple areas when making a request, the system may perform multiple scans to generate a heat map detailing the frequency and density of the turning individual's gaze locations. The area with the most concentrated gaze by the turning individual may represent the area or direction in which the turning individual is focusing. Although several discrete gaze directions may be detected, the fusion of several gaze scans may provide a significant number of scans directed to the area in which the target of attention is located.
[0079] As shown in block 640, the heat map of attention may be combined, overlaid, or overlapped with the saliency map to determine where the object of interest is. The system may determine the location of the object of interest by determining the intersection of the area of interest on the saliency map with the detected area of gaze concentration. If the system can determine such an intersection, the system may have located the object of interest and may then move on to identifying the object. As mentioned above, if the object of interest can be identified using any information the system has acquired along the way, the system may so indicate. If the object is unknown to the system, a learning process may begin.
[0080] Although the embodiments described herein include mobile robots that detect human gaze and gestures, those skilled in the art will recognize that the system is not limited to only such robots, but may be implemented in other types of computing devices, whether distributed or centralized, without departing from the subject matter of this disclosure.
[0081] Additionally, while the exemplary systems and methods described in this disclosure implement a 360-degree panoramic camera to capture 2D images of the unconstrained environment, one skilled in the art will recognize that other image capture devices may be used, for example, an RGB camera, a rotating camera, or multiple cameras with overlapping fields of view may be used to capture a series of environment images that the system's processor may stitch together to form a 360-degree representation of the environment.
[0082] Although embodiments of the present disclosure describe the use of IR cameras, sensors, and the like to generate depth map information, those skilled in the art will appreciate that other depth sensing techniques, including but not limited to SLAM or deep neural networks trained for depth estimation, may be used without departing from the subject matter of the present disclosure.
[0083] Reference to a singular item should be understood to include the plural item and vice versa, unless expressly stated otherwise or clear from the context. Grammatical conjunctions are intended to express any and all disjunctions and connections of clauses, sentences, words, and the like, unless expressly stated otherwise or clear from the context. Thus, the term "or" should be understood generally to mean "and / or" and the like.
[0084] As used herein, unless otherwise stated herein, reference to a range of values is not intended to be limiting, but instead refers individually to any and all values in that range, and each individual value within such range is incorporated herein as if it were individually set forth herein. "About," "approximately," "substantially," or the like, along with numerical or directional values, should be interpreted as indicating deviations that would be appreciated by one of ordinary skill in the art to operate satisfactorily for the intended purpose. Value ranges and / or numerical values are provided herein as examples only, and do not constitute limitations on the scope of the described embodiments. Any and all examples or exemplary language (such as, for example, or the like) provided herein are intended merely to further clarify the embodiments, and do not limit the scope of the embodiments. No language in this specification should be construed as indicating that any unclaimed element is essential to the implementation of the embodiment.
[0085] In the specification and the following claims, the terms "first," "second," "third," "front," "rear," and the like are used for convenience only and should not be construed as limiting terms, unless expressly stated otherwise.
[0086] It should be understood that the above-described methods and systems are intended to be illustrative and not limiting. Numerous variations, additions, omissions, and other modifications will be apparent to those skilled in the art. In addition, the order of the method steps and figures in the above description does not require the steps to be performed in the order recited, unless a particular order is expressly required or clear from the context. Thus, while specific embodiments have been shown and described, it will be apparent to those skilled in the art that various changes and modifications in form and details may be made without departing from the spirit and scope of the present disclosure, which are intended to form part of the present disclosure, which should be interpreted in the broadest sense permitted by law, as defined by the following claims. The invention disclosed in this specification includes the following aspects. [Aspect 1] 1. A system for locating and identifying an object, comprising: an image capture device configured to generate a 360 degree image of an unconstrained environment including an individual and an object of interest; a depth sensor configured to generate a depth map of an unconstrained environment; Knowledge base and 1. A processor comprising: generating an environment map from the 360 degree image and the depth map; determining a direction of gaze from the individual in the 360 degree image; generating a directional vector representing the gaze direction; projecting the directional vector from the individual onto the environment map; Detecting intersections of the directional vector with the object of interest; a processor configured to use data relating to the object of interest to search the knowledge base for an identity of the object of interest; A system comprising: [Aspect 2] 2. The system of claim 1, wherein the image capture device is a panoramic camera. [Aspect 3] The system of aspect 1, wherein the processor is further configured to determine the attention direction from at least one of a gaze direction, a gesture direction of the individual, or a combination thereof. Aspect 4 2. The system of claim 1, wherein the depth sensor is one of an infrared sensor, a neural network trained for depth estimation, or a simultaneous localization and mapping sensor. Aspect 5 1. A method for locating and identifying an object, comprising: capturing a 360 degree image of an unconstrained environment including an individual and an object of interest; generating an environment map; determining a direction of gaze from the individual in the captured 360 degree image; generating a directional vector representing the gaze direction; projecting the directional vector from the individual onto the environment map; detecting intersections of the directional vector with the object of interest; A method comprising: Aspect 6 searching a knowledge base for the identity of the object of interest; 6. The method of embodiment 5, further comprising initiating a learning process when the identity of the object of interest is not found in the knowledge base. Aspect 7 7. The method of embodiment 6, wherein the learning process comprises a neural network. Aspect 8 6. The method of claim 5, wherein the 360-degree image is captured by a monocular panoramic camera. Aspect 9 6. The method of claim 5, further comprising capturing a plurality of images, wherein the 360-degree image comprises a stitching together of the plurality of images. Aspect 10 Generating the environment map Obtaining depth information of an environment; combining depth information of the environment with the 360 degree image; 6. The method of embodiment 5, comprising: Aspect 11 11. The method of embodiment 10, wherein the depth information of the environment is obtained by an infrared sensor. Aspect 12 11. The method of claim 10, wherein the depth information of the environment is obtained from a concurrent process of localization and mapping. Aspect 13 11. The method of claim 10, wherein the depth information of the environment is obtained from a neural network. Aspect 14 6. The method of claim 5, wherein the environment map includes a saliency map. Aspect 15 6. The method of claim 5, wherein the environment map comprises a semantic parsing of the environment. Aspect 16 6. The method of claim 5, further comprising estimating a gaze of the individual, wherein the gaze direction is determined from the gaze. Aspect 17 17. The method of embodiment 16, wherein the line of sight is estimated using an infrared sensor. Aspect 18 6. The method of embodiment 5, further comprising estimating a gesture direction of the individual, wherein the gaze direction is determined from the gesture direction. Aspect 19 6. The method of claim 5, further comprising estimating gaze and gesture directions of the individual, the gaze direction being determined from the gaze and gesture directions. Aspect 20 1. A system for locating and identifying an object, comprising: a monocular camera configured to capture a 360 degree image of an unconstrained environment including an individual and an object of interest; an infrared depth sensor configured to generate a three-dimensional depth map of the unconstrained environment; Knowledge base and 1. A processor comprising: generating a saliency map from the 360 degree image and the depth map; determining a direction of attention from the individual; generating a three-dimensional directional vector representing the gaze direction; projecting the three-dimensional directional vectors from the individuals onto the saliency map; detecting an intersection of the three-dimensional directional vector with the object of interest; a processor configured to use captured image data relating to the object of interest to search the knowledge base for an identity of the object of interest; A system comprising:
Claims
1. 1. A system for locating and identifying an object, comprising: a monocular camera configured to capture a 360 degree image of an unconstrained environment including an individual and an object of interest; an infrared depth sensor configured to generate a three-dimensional depth map of the unconstrained environment; Knowledge base and 1. A processor comprising: generating a saliency map from the 360 degree image and the depth map; determining a direction of attention from the individual; generating a three-dimensional directional vector representing the gaze direction; projecting the three-dimensional directional vectors from the individuals onto the saliency map; detecting an intersection of the object of interest with the three-dimensional directional vector; a processor configured to search the knowledge base for an identity of the object of interest detected at an intersection of the object of interest with the three-dimensional directional vector; A system comprising:
2. The system of claim 1 , wherein the monocular camera is a panoramic camera.
3. The system of claim 1 , wherein the 360 degree image comprises multiple stitched images.
4. The system of claim 1 , wherein the processor is further configured to estimate a gesture direction of the individual, and the gaze direction is determined from the gesture direction.
5. The system of claim 1 , wherein the processor is further configured to estimate gaze and gesture directions of the individual, the attention direction being determined from the gaze and gesture directions.
6. The system of claim 1 , wherein the processor is further configured to initiate a learning process when the identity of the object of interest is not found in the knowledge base.
7. The system of claim 6 , wherein the learning process comprises a neural network.
8. The system of claim 1 , wherein the processor is further configured to estimate a gaze direction of the individual, and the gaze direction is determined from the gaze direction.
9. The system of claim 8 , wherein the line of sight is estimated using an infrared sensor.
Citation Information
Patent Citations
Method for sorting defect, device therefor and method for generating data for instruction
JP2000057349A
Image processing system, method of processing image, and program
JP2010268158A
Diagnosis support device, diagnosis support method, and diagnosis support program
JP2015116319A
Method, apparatus, and computer program product for personalized depth-of-field omnidirectional video
JP2017041242A
Machine learning device
JP2017224184A