System for performing eye detection and / or tracking

By using multiple imaging devices and actuator systems, the problem of low detection accuracy and load caused by high-resolution imaging devices in existing eye-tracking systems in wide-angle lens environments has been solved, achieving efficient and low-power eye-tracking results.

CN111798488BActive Publication Date: 2026-03-20TOBI TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-08
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing eye-tracking systems, due to the use of wide-angle lenses in environments such as vehicle passenger compartments, result in only a small portion of the image representing the user's eyes, limiting detection and tracking capabilities. Furthermore, using high-resolution imaging equipment increases processing load, latency, and power consumption.

Method used

The system employs multiple imaging devices and actuator systems. The first imaging device initially locates the user's face, and the actuator rotates the second imaging device to cover the user's face with its field of view (FOV). Combined with light source illumination, high-quality second image data is generated for precise eye tracking.

Benefits of technology

It improves the detection accuracy of eye position and gaze direction, reduces system power consumption and heat, lowers processing load, and is suitable for applications in enclosed or confined spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111798488B_ABST
    Figure CN111798488B_ABST
Patent Text Reader

Abstract

This disclosure describes, in part, systems and techniques for performing eye tracking. For example, a system can include a first imaging device that generates first image data. The system can then analyze the first image data to determine a position of a user’s face. Using the position, the system can cause an actuator to move from a first position to a second position in order to direct a second imaging device toward the user’s face. When in the second position, the second imaging device can generate second image data that represents at least the user’s face. The system can then analyze the second image data to determine a gaze direction of the user. In some instances, the first imaging device can include a first field of view (FOV) that is greater than a second FOV of the second imaging device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system for performing eye detection and / or tracking. Background Technology

[0002] Many systems use eye tracking to determine a user's gaze direction or other eye-related attributes. For example, in a vehicle, a system can use a camera installed in the vehicle to capture an image of a user inside the vehicle. The system can then analyze the image to at least determine the user's gaze direction while driving the vehicle. The vehicle control system can then use the gaze direction to determine whether the user is paying attention to the road while driving. If the user is not paying attention to the road, the vehicle control system can output an audible or other alert to the user.

[0003] In many cases, these systems encounter problems when performing eye tracking to determine a user's gaze direction or other eye-related attributes. For example, in large environments such as the passenger compartment of a vehicle, a camera might need a wide-angle lens to capture an image representing the environment. Thus, only a small portion of the image can represent the user's eyes, which can cause problems for systems using eye tracking to analyze the images. For instance, the system might be unable to recognize the user's eyes from the image. To compensate for this problem, some systems use high-resolution cameras to capture images. However, using high-resolution cameras can increase the processing load, potentially increasing latency, power consumption, and / or system heat generation. Summary of the Invention

[0004] According to a first aspect, a method according to claim 1 is provided.

[0005] Preferably, the method further includes determining a second position of the actuator based at least in part on a portion of the first image data.

[0006] According to a second aspect, a method according to claim 14 is provided.

[0007] Preferably, the method includes: generating first image data by using a first imaging device, generating first image data using a first frame rate and a first resolution; generating second image data by using a second imaging device, generating second image data using a second frame rate and a second resolution; and at least one of the first frame rates is different from the second frame rate, or the first resolution is different from the second resolution.

[0008] More preferably, it includes at least one of the following: the first frame rate is lower than the second frame rate; or the first resolution is lower than the second resolution.

[0009] Preferably, the method includes determining the second position of the actuator based at least in part on a position of the user’s face and a position of the second imaging device relative to a position of the first imaging device.

[0010] Preferably, in any aspect, the method further includes emitting light toward the user’s face using a light source associated with the second imaging device.

[0011] Preferably, in any aspect, the first image data further represents a first field of view, FOV, of the first imaging device; the second image data further represents a second FOV of the second imaging device; and the second FOV is smaller than the first FOV.

[0012] In yet another aspect, a system is provided that includes: a first imaging device; a second imaging device; an actuator to position the second imaging device; one or more processors; and one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations of any of the above aspects. BRIEF DESCRIPTION OF DRAWINGS

[0013] The detailed description is set forth with reference to the accompanying drawings. In the drawings, one or more leftmost digits of a reference number identify the figure in which the reference number first appears. The use of the same reference numbers in different figures indicates similar or identical items or features.

[0014] Figure 1 An example process for performing eye tracking using multiple imaging devices is shown.

[0015] Figure 2 An example of analyzing image data generated by multiple imaging devices to determine a user’s eye position and / or gaze direction is shown.

[0016] Figure 3 A block diagram of an example system for eye tracking using multiple imaging devices is shown.

[0017] Figure 4 An example diagram representing a system performing eye tracking is shown.

[0018] Figure 5 An example process for performing eye tracking using multiple imaging devices is shown.

[0019] Figure 6 An example process for determining when to adjust an actuator of an imaging device used for eye tracking is shown. DETAILED DESCRIPTION

[0020] As described above, conventional systems that perform eye tracking can use an imaging device, such as a camera, to generate image data representing an image. The system can then analyze the image data to determine a gaze direction or other eye-related attribute of a user. However, in many cases, such systems can have issues performing eye tracking and / or determining a gaze direction or other eye-related attribute of a user. For example, such as when the system is installed in a passenger cabin of a vehicle, the imaging device can need a wide-angle lens to capture an image representing a majority of the passenger cabin. As such, only a small portion of the image data can represent the eyes of a user, which can limit the system’s ability to accurately detect and track the position of the user’s eyes and / or the user’s gaze direction. To compensate for this issue, some conventional systems use high-resolution imaging devices. However, by using high-resolution imaging devices, such systems often suffer from a higher processing load to process the high-resolution image data for a wide field of view, which can increase latency, power consumption, and / or heat generated by the system.

[0021] This disclosure describes, in part, systems and techniques for performing eye tracking using multiple imaging devices. For example, a system can include at least a first imaging device, a second imaging device, and an actuator configured to rotate the second imaging device. The first imaging device can include a first field of view (FOV), and the second imaging device can include a second FOV. In some instances, the first FOV is different than the second FOV. For example, the first FOV of the first imaging device can be larger than the second FOV of the second imaging device, such that the second FOV includes only a portion of the first FOV. However, the system can use the actuator to rotate the second imaging device such that the second imaging device can scan substantially all of the first FOV.

[0022] To perform eye tracking, the system can generate image data (referred to as “first image data” in these examples) using the first imaging device. The system can then analyze the first image data using one or more algorithms associated with face detection. Based on this analysis, the system can determine a position of a face of a user (e.g., a direction from the first imaging device to the face of the user). For example, based on this analysis, the system can determine that a portion of the first image data represents the face of the user. Each portion of the first image data can be associated with a respective position (and / or a respective direction). As such, the system can determine the position (and / or direction) based on which portion of the first image data represents the face of the user. While this is just one example of using face detection to determine a position of a face of a user, in other examples, the system can analyze the first image data using any other algorithm and / or technique in order to determine a position of a face of a user.

[0023] The system can then use the position and / or orientation of the user's face to rotate the actuator from the first position to a second position. At the second position, a majority of the second FOV of the second imaging device can include the user's face. In other words, the system can use the position determined using the first image data to direct the second imaging device toward the user's face. Additionally, in some examples, the second imaging device can include a light source that emits light. In such examples, the light source can also be directed toward the user's face and can emit light in a limited area, such as an area proximate to the user's face. The system can then use the second imaging device to generate image data (referred to as "second image data" in these examples). In some instances, and because the second imaging device is directed toward the user's face, a greater portion of the second image data can represent the user's face as compared to the first image data. Moreover, in examples in which the light source emits light in a limited area, the light source can require less power as compared to illuminating a larger area (e.g., a passenger cabin or a field of view of the first imaging device).

[0024] The system can then analyze the second image data using one or more algorithms associated with eye tracking. Based on this analysis, the system can determine an eye position and / or, in the first aspect, a gaze direction of the user. For example, the system can analyze the second image data to identify a center of a pupil of the user's eye (e.g., an eye position). The system can additionally or alternatively analyze the second image data to determine a center of a corneal reflection created by light emitted by the light source. Using the position of the user's face, the center of the eye pupil, and / or the center of the corneal reflection, the system can determine a gaze direction of the user. While this is just one example of using eye tracking to determine a gaze direction of a user, in other examples, the system can analyze the second image data using any other algorithm and / or technique in order to determine a gaze direction of the user.

[0025] In some instances, the system can perform the above-described techniques in order to continuously or periodically track an eye position and / or gaze direction of a user. For example, as the position of the user's face changes, the system can continue to analyze first image data generated by the first imaging device to determine a new position of the user's face. The system can then cause the actuator to move from the second position to a third position. When at the third position, the second imaging device and / or the light source can be directed toward the user's face, which is now located at the new position. The system can then analyze second image data generated by the second imaging device to determine a new eye position and / or, in the first aspect, the system can determine a new gaze direction of the user.

[0026] In some instances, the system can output data representative of the position of the user's face, the user's eye position, and / or the first aspect of the user's gaze direction to one or more computing devices. For example, if the system is installed in or in communication with a vehicle, the system can output the data to one or more other computing devices installed in the vehicle (e.g., a vehicle driving system) and / or one or more remote systems via a network connection. The one or more computing devices and / or remote systems can then process the data. For example, the one or more computing devices can analyze the data in order to determine whether the user is paying attention to the road, whether the user is drowsy, whether the user sees objects in the vehicle's environment, etc. If the one or more computing devices determine that one of these or other applicable conditions exists, the one or more computing devices can cause an alert, such as a sound, a vibration, or a visible warning, to be output in order to alert the user.

[0027] In some instances, the system can be pre-installed in an environment, such as a passenger cabin of a vehicle. For example, a vehicle manufacturer can pre-install the system into a vehicle and then calibrate the imaging devices based on their locations in the vehicle. In other instances, the system can not be pre-installed in an environment. For example, the system can include a standalone or aftermarket system that can be installed within a variety of environments.

[0028] The second imaging device can be located at a known position relative to the first imaging device. As such, based on the position of the user's face relative to the first imaging device and the known position of the second imaging device relative to the first imaging device, the system can determine an angle from the second imaging device to the user's face. In some instances, the second imaging device can be positioned proximate to the first imaging device. For example, the second imaging device can be installed within a threshold distance to the first imaging device. The threshold distance can include, but is not limited to, 1 centimeter, 2 centimeters, 10 centimeters, etc. In other instances, the second imaging device can be spaced apart from the first imaging device. For example, the second imaging device can be spaced apart from the first imaging device by a distance greater than the threshold distance.

[0029] In some instances, when the system is installed in a vehicle, the first imaging device and / or the second imaging device can be installed in a front portion of the vehicle. For example, the first imaging device and / or the second imaging device can be installed in a dashboard of the vehicle, a rearview mirror of the vehicle, and / or any other location in, on, or proximate to which the first imaging device and / or the second imaging device can generate image data representative of the eyes of a user driving the vehicle.

[0030] In some instances, eye tracking is performed using a system that includes multiple imaging devices, which improves upon previous systems that performed eye tracking using only a single imaging device. For example, by directing a second imaging device toward the user's face, a larger portion of second image data generated by the second imaging device represents the user's face and / or eyes. This makes it easier for the system to analyze the second image data to determine eye position and / or the first aspect user's gaze direction than analyzing image data in which only a small portion of the image data represents the user's face and / or eyes. For another example, the system can be able to perform eye tracking using a low resolution imaging device, which consumes less power than a high resolution imaging device. This not only reduces the power consumed by the system, it also reduces the heat emitted by the system, which is important in enclosed or confined spaces, such as a vehicle's dashboard or rearview mirror.

[0031] Figure 1 An example process for performing eye detection and / or tracking using multiple imaging devices is shown. At 102, the system can determine a position 106 of a user's 108 face using first image data generated by a first imaging device 104. For example, the first imaging device 104 can generate first image data representing an environment 110. In the illustrated example, the environment 110 includes an interior cabin of a vehicle that includes at least the user 108. In other examples, the system can be used in an environment with any number of users. As shown, the first imaging device 104 includes a first FOV 114 that includes the user 108 and an additional user 112. To determine the position 106 of the user's 108 face, the system can analyze the first image data using one or more algorithms associated with face detection. Based on this analysis, the system can determine a position 106 of the user's 108 face relative to the position of the first imaging device 104 within the environment 110. Figure 1

[0032] In some instances, the system can determine the position 106 of the user's 108 face because the user 108 is the driver of the vehicle. In such instances, the system can determine that the user 108 is the driver based on the user's 108 relative position within the environment 110. For example, the user 108 can be in a position within the environment 110 that drivers typically occupy.

[0033] ​In some examples, the position 106 can represent a two-dimensional and / or three-dimensional position of the user's 108 face within the environment 110. Additionally or alternatively, in some examples, the position 106 can represent a direction 116 from the first imaging device 104 to the user's 108 face. In such examples, the direction 116 can include a two-dimensional vector and / or a three-dimensional vector. In some examples, as described herein, the second imaging device 120 is positioned proximate to the first imaging device 104 in order to minimize and / or eliminate parallax when adjusting the second imaging device 120.

[0034] At 118, the system can cause movement of an actuator associated with the second imaging device 120 based at least in part on the position 106. For example, as shown at 122, the system can cause the actuator associated with the second imaging device 120 to move from a first position to a second position. When the actuator is in the second position, the second imaging device 120 can be directed toward the user's 108 face. For example, a greater portion of a second FOV 124 of the second imaging device 120 can include the user's 108 face as compared to the first FOV 114 of the first imaging device 118. Additionally, in some examples, the second imaging device 120 can include a light source. In such examples, when the actuator is in the second position, the light source can be directed toward the user's face.

[0035] At 126, the system can determine a position of the user's 108 eyes using second image data generated by the second imaging device 120. For example, the second imaging device 120 can generate second image data, where the second image data represents at least the user's 108 face. The system can then analyze the second image data using one or more algorithms associated with eye tracking. Based on the analysis, the system can determine a position of the user's 108 eyes. In some examples, and based on the analysis concurrently, the system can further determine a gaze direction 128 of the user's 108.

[0036] As Figure 1 As further shown in the example of FIG. 1, the system can continuously or periodically perform the example process. For example, the system can determine a new position of the user's 108 face using first image data generated by the first imaging device 104. The system can then cause additional movement of the actuator associated with the second imaging device 120 based at least in part on the new position. Additionally, the system can determine a new position of the user's 108 eyes and / or a new gaze direction using second image data generated by the second imaging device 120. In other words, the system can continue to track the user's 108 eyes over time using the first imaging device 104 and the second imaging device 120.

[0037] Figure 2An example is shown in which image data generated by multiple imaging devices is analyzed to determine an eye position and / or gaze direction of a user 202. For example, a system can use a first imaging device to generate first image data. In Figure 2 In the example shown, the first image data represents at least one image 204 that depicts at least the user 202 and an additional user 206. The system can then analyze the first image data using one or more algorithms associated with face detection. Based on this analysis, the system can determine a position 208 of the face of the user 202.

[0038] The system can then cause an actuator associated with the second imaging device to move from a first position to a second position such that the second imaging device is directed toward the face of the user 202. While in the second position, the system can use the second imaging device to generate second image data. In Figure 2 In the example shown, the second image data represents only a portion of the first image data. For example, the second image data represents at least one image 210 that depicts the face of the user 202. As shown, a greater portion of the second image data represents the face of the user 202 as compared to the first image data.

[0039] The system can then analyze the second image data using one or more algorithms associated with eye tracking. Based on this analysis, the system can determine at least an eye portion 212 of the user 202 and / or a gaze direction of the user 202. In some instances, the system can then output data representing the position 208 of the face of the user 202, the eye position 212 of the user 202, and / or the gaze direction of the user 202.

[0040] Figure 3 A block diagram of an example system 302 that uses multiple imaging devices for eye tracking is shown. As shown, the system 302 includes at least a first imaging device 304 (which can represent and / or be similar to the first imaging device 104), a second imaging device 306 (which can represent and / or be similar to the second imaging device 120), and an actuator 308 configured to rotate the second imaging device 306. The first imaging device 304 can include a still image camera, a video camera, a digital camera, and / or any other type of device that generates first image data 310. In some instances, the first imaging device 304 can include a wide-angle lens that provides the first imaging device 304 with a wide FOV.

[0041] Additionally, the second imaging device 306 can include a still image camera, a video camera, a digital camera, and / or any other type of device that generates the second image data 312. In some instances, the second imaging device 306 includes a large focal length and / or a large depth of field lens that provides the second imaging device 306 with a smaller FOV compared to the first imaging device 304. However, the system 302 can use the actuator 308 to rotate the second imaging device 306 such that the second imaging device 306 can scan the entire FOV of the first imaging device 304.

[0042] For example, the actuator 308 can include any type of hardware device configured to rotate about one or more axes. The second imaging device 306 can be attached to the actuator 308 such that when the actuator rotates, the second imaging device 306 also rotates, thereby changing the viewing direction of the second imaging device 306. In some instances, the light source 314 can also be attached to the actuator 308 and / or to the second imaging device 306. In such instances, the actuator 308 can further rotate in order to change the direction in which the light source 314 emits light thereon. For example, the light source 314 can emit light in a similar direction as the direction in which the second imaging device 306 is generating the second image data 312 such that the light illuminates the FOV of the second imaging device 306. The light source 314 can include, but is not limited to, a light emitting diode, an infrared light source, and / or any other type of light source that emits visible and / or non-visible light.

[0043] In some instances, a first frame rate and / or a first resolution associated with the first imaging device 304 can be different than a second frame rate and / or a second resolution associated with the second imaging device 306. In some instances, a first frame rate and / or a first resolution associated with the first imaging device 304 can be the same as a second frame rate and / or a second resolution associated with the second imaging device 306.

[0044] As further shown in Figure 3 The system 302 can include a face detector component 316, a control component 318, and an eye tracking component 320, as further shown in FIG. 3. The face detector component 316 can be configured to analyze the first image data 310 in order to determine a location of a face of a user. For example, the face detector component 316 can analyze the first image data 310 using one or more algorithms associated with face detection. The one or more algorithms can include, but are not limited to, a neural network algorithm, a principal component analysis algorithm, an independent component analysis algorithm, a linear discriminant analysis algorithm, an evolutionary pursuit algorithm, an elastic bunch graph matching algorithm, and / or any other type of algorithm that the face detector component 316 can use to perform face detection on the first image data 310.

[0045] In some instances, the position can correspond to a direction from the first imaging device 304 to the user's face. For example, to determine the position of the face, the face detector component 316 analyzes the first image data 310 using one or more algorithms. Based on the analysis, the face detection component 316 can determine a direction from the first imaging device 310 to the user's face. The direction can correspond to a two-dimensional vector and / or a three-dimensional vector from the first imaging device 304 to the user's face. After determining the position of the user's face, the face detector component 316 can generate face position data 322 that represents the position of the user's face.

[0046] The control component 318 can be configured to use the face position data 322 to move the actuator 308 from a current position to a new position. When in the new position, the second imaging device 306 and / or the light source 314 can be directed toward the user's face. For example, when in the new position, a greater portion of the FOV of the second imaging device 306 can include the user's face. Further, a greater portion of the light emitted by the light source 314 can be directed toward the user's face, rather than other objects located within a similar environment as the user. In some instances, to move the actuator 308, the control component 318 can determine the new position based on the position (e.g., direction) represented by the position data 322.

[0047] For example, the control component 318 can use one or more algorithms to determine the new position of the actuator 308. In some instances, the one or more algorithms can determine the new position based on the position (and / or direction) represented by the face position data 322, the position of the second imaging device 306, and / or the distance between the first imaging device 304 and the second imaging device 306. For example, the control component 318 can determine two-dimensional coordinates that indicate a direction from the face to the first imaging device 304, determine a distance from the first imaging device 304 to the face, and convert the two-dimensional coordinates to a three-dimensional vector using the distance, where a particular distance along the vector gives a three-dimensional position of the face. The control component 318 can then use the three-dimensional vector, the position of the first imaging device 304, and the position of the second imaging device 306 to determine a direction vector between the second imaging device 306 and the face. Further, the control component 318 can convert the direction vector to polar coordinates for driving the actuator 308. While this is just one example for determining the new position, the control component 318 can use any other technique to determine the new position of the actuator 308.

[0048] After determining the position, the control component 318 can generate control data 324 that represents the new position of the actuator 308. In some instances, the control data 324 can represent polar coordinates for driving the actuator 308.

[0049] The actuator 308 and / or the second imaging device 306 can then move from the current position to the new position represented by the control data 324 using the control data 324. Additionally, the actuator 308 and / or the second imaging device 306 can generate position feedback data 326 representing the current position of the actuator 308 and / or the second imaging device 306. In some instances, the control component 318 uses the position feedback data 326 to determine when the second imaging device 306 is directed toward the user’s face.

[0050] In some instances, to determine when the second imaging device 306 is directed toward the user’s face, the control component 318 can analyze the second image data 312 using a similar process as described above with respect to the first image data 310. Based on the analysis, the control component 318 can determine a direction vector between the second imaging device 306 and the user’s face. The control component 318 can then use the direction vector to determine a polar coordinate. If the polar coordinate is the same (and / or within a threshold difference) as the polar coordinate determined using the first image data 310, the control component 318 can determine that the second imaging device 306 is directed toward the user’s face. However, if the polar coordinate is different (e.g., outside of the threshold difference), the control component 318 can use the polar coordinate to further drive the actuator 308. In other words, the control component 318 can use similar techniques as described above with respect to the first image data 310 in order to further direct the second imaging device 306 toward the user’s face.

[0051] The eye tracking component 320 can be configured to analyze the second image data 312 in order to determine an eye position and / or a gaze direction of the user. For example, the eye tracking component 320 can analyze the second image data 312 using one or more algorithms associated with eye tracking. The one or more algorithms can include, but are not limited to, neural network algorithms and / or any other type of algorithm associated with eye tracking. The eye position can represent a three-dimensional position of the eye relative to the second imaging device 306. Additionally, the gaze direction can represent a vector originating from the eye and expressed in a coordinate system associated with the second imaging device 306. After determining the eye position and / or the gaze direction of the user, the eye tracking component 320 can generate eye tracking data 328 representing the eye position and / or the gaze direction.

[0052] In some instances, the second imaging device 306 and / or the control component 318 can then determine the eye position and / or the gaze direction in a global coordinate system associated with the environment using the eye position and / or the gaze direction expressed in the coordinate system associated with the second imaging device 306, the position of the second imaging device 306, and / or the orientation of the second imaging device 306. For example, the second imaging device 306 and / or the control component 318 can determine the eye position and / or the gaze direction relative to the passenger cabin of the vehicle.

[0053] As Figure 3 As further shown in FIG. 3, system 302 includes a processor 330, a network interface 332, and a memory 334. As used herein, a processor such as processor 330 can include multiple processors and / or a processor having multiple cores. Further, the processor can include one or more cores of different types. For example, the processor can include an application processor unit, a graphics processing unit, and / or the like. In one instance, the processor can include a microcontroller and / or a microprocessor. Processor 330 can include a graphics processing unit (GPU), a microprocessor, a digital signal processor, or other processing units or components known in the art. Alternatively, or additionally, the functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc. Further, processor 330 can have its own local memory, which can also store program components, program data, and / or one or more operating systems.

[0054] Memory 334 can include volatile and nonvolatile memory, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program components, or other data. Memory 334 includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, RAID storage systems, or any other medium which can be used to store the desired information and which can be accessed by computing device. Memory 334 can be implemented as a computer-readable storage medium (“CRSM”) which can be any available physical medium that can be accessed by processor 330 to execute instructions stored on memory 334. In one basic example, the CRSM can include random access memory (“RAM”), and flash memory. In other examples, a CRSM can include, but is not limited to, read-only memory (“ROM”), electrically erasable programmable read-only memory (“EEPROM”), or any other tangible medium which can be used to store the desired information and which can be accessed by the processor.

[0055] Further, functional components can be stored in respective memory, or the same functionality can alternatively be implemented in hardware, firmware, application specific integrated circuits, field programmable gate arrays, or as a system on a chip (SoC). Further, while not shown, each respective memory discussed herein, such as memory 334, can include at least one operating system (OS) component configured to manage hardware resource devices, such as network interfaces, I / O devices of the respective device, etc., and provide various services to applications or components executing on the processor. Such OS components can implement a variant of the FreeBSD operating system published by the FreeBSD Project; other UNIX or UNIX-like variants; a variant of the Linux operating system published by Linus Torvalds; the FireOS operating system by Amazon.com, Inc. of Seattle, Washington; the Windows operating system by Microsoft Corporation of Redmond, Washington; the LynxOS by Lynx Software Technologies, Inc. of San Jose, California; the OSE by Enea AB of Sweden; the QNX RTOS by BlackBerry Limited; and the like.

[0056] Network interface 332 can enable system 302 to send data to and / or receive data from other electronic devices. Network interface 332 can include one or more network interface controllers (NICs) or other types of transceiver devices to send and receive data over a network. For example, network interface 332 can include a personal area network (PAN) component to support messages over one or more short-range wireless message channels. For example, the PAN component can support messages that comply with at least one of the following standards: IEEE 802.15.4 (ZigBee), IEEE 802.15.1 (Bluetooth), IEEE 802.11 (WiFi), or any other PAN message protocol. Further, network interface 332 can include a wide area network (WAN) component to support messages over a wide area network. Further, the network interface can support system 302 to communicate using a controller area network bus.

[0057] Operations and / or functions associated with and / or described with respect to components of system 302 can be performed with cloud-based computing resources. For example, a network-based system such as an elastic compute cloud system or similar system can be used to generate and / or present a virtual computing environment to perform some or all of the functions described herein. Further, or alternatively, one or more systems can be utilized that can be configured to perform operations without providing and / or managing servers, such as a Lambda system or similar system.

[0058] Although Figure 3 Examples of the face detector component 316, the control component 318, and the eye tracking component 320 are illustrated as including hardware components, in other examples, one or more of the face detector component 316, the control component 318, and the eye tracking component 320 can include software stored in the memory 334. Moreover, although Figure 3 Examples of the face detector component 316, the control component 318, and the eye tracking component 320 are illustrated as including hardware components, in other examples, one or more of the face detector component 316, the control component 318, and the eye tracking component 320 can include software stored in the memory 334. Moreover, although

[0059] Moreover, although Figure 3 Examples of the face detector component 316, the control component 318, and the eye tracking component 320 are illustrated as including hardware components, in other examples, one or more of the face detector component 316, the control component 318, and the eye tracking component 320 can include software stored in the memory 334. Moreover, although

[0060] As described herein, a machine learning model can include, but is not limited to, a neural network (e.g., a You Only Look Once (YOLO) neural network, VGG, DenseNet, PointNet, a convolutional neural network (CNN), a stacked autoencoder, a deep Boltzmann machine (DBM), a deep belief network (DBN)), a regression algorithm (e.g., ordinary least squares regression (OLSR), linear regression, logistic regression, stepwise regression, multivariate adaptive regression splines (MARS), locally estimated scatterplot smoothing (LOESS)), a Bayesian algorithm (e.g., Naive Bayes, Gaussian Naive Bayes, Multinomial Naive Bayes, Average One-Dependence Estimators (AODE), Bayesian Belief Network (BNN), Bayesian Network), a clustering algorithm (e.g., k-means, k-medians, expectation maximization (EM), hierarchical clustering), an association rule learning algorithm (e.g., perceptron, back-propagation, Hopfield network, radial basis function network (RBFN)), supervised learning, unsupervised learning, semi-supervised learning, etc. Additional or alternative examples of neural network architectures can include neural networks such as ResNet50, ResNetlOl, VGG, DenseNet, PointNet, etc. Although discussed in the context of neural networks, any type of machine learning can be used in accordance with the present disclosure. For example, machine learning algorithms can include, but are not limited to, regression algorithms, instance-based algorithms, Bayesian algorithms, association rule learning algorithms, deep learning algorithms, etc.

[0061] Figure 4An example of a system 302 that performs eye tracking is shown. As shown, a first imaging device 304 can generate first image data 310 and then send the first image data 310 to a face detector component 316. The face detector component 316 can then analyze the first image data 310 to determine a location of a user's face. Additionally, the face detection component 316 can generate face location data 322 that represents the location and send the face location data 322 to a control component 318.

[0062] The control component 318 can use the face location data 322 to generate control data 324, where the control data 324 represents a new location for the second imaging device 306. The control component 318 can then send the control data 324 to the second imaging device 306 (and / or an actuator 308). Based on receiving the control data 324, the actuator 308 of the second imaging device 306 can move from a current location to the new location. While the actuator is in the new location, the second imaging device 306 can send location feedback data 326 to the control component 318, where the location feedback data 326 indicates that the actuator 308 is in the new location. Additionally, the second imaging device 306 can generate second image data 312 that represents at least the user's face. The second imaging device 306 can then send the second image data 312 to the face detector component 316.

[0063] The face detector component 316 can analyze the second image data 312 to determine a user's eye location and / or gaze direction. The face detector component 316 can then generate eye tracking data 328 that represents the eye location and / or gaze direction and send the eye tracking data 328 to the control component 318. In some instances, the control component 318 uses the eye tracking data 328 along with the new face location data 322 to determine a new location for the second imaging device 306.

[0064] As Figure 4 As further shown in the example of FIG. 4, the control component 318 can send output data 402 to an external system 404. The output data 402 can include, but is not limited to, the face location data 322 and / or the eye tracking data 328. In some instances, the external system 404 can be included in a device similar to the system 302. For example, the system 302 can be installed in or on a vehicle and the external system 404 can include a vehicle control system. In some instances, the external system 404 can include a remote system (e.g., a fleet monitoring service that monitors drivers, a remote image analysis service, etc.). In such instances, the system 302 can send the output data 402 to the external system 404 over a network.

[0065] Figures 5-6An example process for performing eye tracking is shown. The processes described herein are illustrated as a collection of blocks in logical flow graphs, which represent a sequence of operations that can be implemented in hardware, software, or a combination thereof. In the context of software, the blocks represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processors, program the processors to perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the blocks are described is not intended to be limiting. Any number of the described blocks can be implemented in any order and / or in parallel, and not all blocks need to be executed.

[0066] Figure 5 An example process 500 for performing eye tracking using multiple imaging devices is shown. At 502, the system can generate, using a first imaging device, first image data representing at least a user. In some instances, the first imaging device can include a first FOV, a first resolution, and / or a first frame rate. In some instances, the system can be installed within a vehicle. For example, the first imaging device can be installed in a front portion of the vehicle, such as a dashboard. The first image data can then represent a passenger cabin of the vehicle that includes at least the user (e.g., a driver).

[0067] At 504, the system can analyze the first image data using one or more algorithms associated with face detection, and at 506, the system can determine a position of a face of the user. In some instances, the position can represent a direction from the first imaging device to the face of the user. For example, the position can represent a two-dimensional or three-dimensional vector that indicates a direction from the first imaging device to the face of the user. In some instances, the system determines the position based on which portion of the first image data represents the face of the user.

[0068] At 508, the system can cause, based at least in part on the position, an actuator associated with a second imaging device to move from a first position to a second position. For example, the actuator can rotate along one or more axes to move from the first position to the second position. When in the second position, the second imaging device can be directed toward the face of the user. Further, in some instances, when in the second position, a light source can be directed toward the face of the user.

[0069] At 510, the system can generate second image data representing at least the user's eyes using a second imaging device. For example, once the actuator is in the second position, the second imaging device can generate the second image data. In some instances, the second imaging device can include a second FOV, a second resolution, and / or a second frame rate. At least one of the second FOV can be different than the first FOV, the second resolution can be different than the first resolution, or the second frame rate can be different than the first frame rate. In some instances, such as when the system is installed in a vehicle, the second imaging device can also be installed in a front portion (e.g., dashboard) of the vehicle.

[0070] At 512, the system can analyze the second image data using one or more algorithms associated with eye tracking, and at 514, the system can determine at least one of an eye position of the user or a gaze direction of the user.

[0071] At 516, the system can output data representing at least one of a position of the user's face, an eye position of the user, or a gaze direction of the user. In some instances, the system can send the data to another system located in a similar device as the system. For example, if the system is installed in a vehicle, the system can send the data to a vehicle driving system. In some instances, the system can send the data to a remote system over a network connection. In some instances, the system can continue to perform the example process 500 in order to track the eyes of the user.

[0072] Figure 6 An example process 600 for determining when to adjust an actuator of an imaging device for eye tracking is shown. At 602, the system can cause an actuator associated with an imaging device to move to a first position, the first position associated with a first position within an environment. In some examples, the system can cause the actuator to move to the first position based on determining that a face of a user is located at the first position. In some instances, the first position can be based on a first direction from the other imaging device to the face of the user.

[0073] At 604, the system can determine a second position associated with the face of the user located within the environment. In some instances, the system can determine the second position by analyzing image data generated by the other imaging device using one or more algorithms associated with face detection. In some instances, the system can determine the second position based on receiving data from the electronic device indicating the second position. In any instance, the second position can be based on a second direction from the other imaging device to the face of the user.

[0074] At 606, the system can determine whether the second position is different from the first position. For example, the system can compare the second position and the first position. In some instances, comparing the second position and the first position can include comparing the second orientation and the first orientation. Based on the comparison, the system can determine whether the second position is different from the first position. In some instances, the system can determine that the second position is different from the first position based on a difference between the second orientation and the first orientation exceeding a threshold in any dimension. The threshold can include, but is not limited to, one degree, five degrees, ten degrees, etc.

[0075] If, at 606, the system determines that the second position is not different from the first position, at 608, the system can determine to leave the actuator in the first position. For example, if the system determines that the second position is not different from the first position, the system can determine that the imaging device is still directed toward the user’s face. As such, the system can determine not to move the actuator in order to change the orientation of the imaging device. The system can then analyze image data generated by the imaging device in order to determine the user’s eye position and / or gaze direction.

[0076] However, if, at 606, the system determines that the second position is different from the first position, at 610, the system can determine a second position for the actuator based at least in part on the second position. For example, if the system determines that the second position is different from the first position, the system can determine that the imaging device is no longer directed toward the user’s face. The system can then determine the second position such that the imaging device is directed toward the user’s face, which is now located at the second position. In some instances, the system determines the second position based at least in part on the second position, positions of other imaging devices within the environment, and / or positions of imaging devices within the environment.

[0077] At 612, the system can cause the actuator to move from the first position to the second position. As described above, when the actuator is in the second position, the imaging device can be directed toward the user’s face, which is located at the second position. Once the actuator is in the second position, the system can analyze image data generated by the imaging device in order to determine the user’s eye position and / or gaze direction. In some instances, the system can then proceed to the example process 600 in order to keep the imaging device directed toward the user’s face.

[0078] Referring back to Figures 1-3In a second aspect of the application, as an alternative or in addition to the gaze direction of the user 108, 202 being tracked respectively, the FOV 124 of the second imaging device 120, 306 is narrow enough and the resolution of the optical system of the second imaging device 120, 306 is high enough to provide sufficient image information from the eye region of the user 108, 202 to enable the diameter of one or more pupils of the user 108, 202 to be measured. Methods for analyzing images such as Figure 2 the image portion of the region 212 for the iris region and for identifying the pupil boundary.

[0079] In such embodiments, it can be beneficial for the light source 314 to emit infrared (IR) or near-IR light and for the second imaging device 120, 306 to be sensitive to the wavelengths of light emitted by the light source 314. Once the second imaging device 120, 306 and the light source 314 are directed towards the face of the user 108, 202, the continuous IR images provided by the second imaging device 120, 306 can provide a clear image of the user’s pupil even in low ambient light conditions, thereby enabling the pupil size and position to be tracked over time.

[0080] In embodiments, one or both of the first imaging device 104, 304 or the second imaging device 120, 306 is sensitive to visible light wavelengths. This can be achieved by providing a Bayer filter type image sensor within the required device, where the pixels are divided into respective RGB and IR sub-pixels, or alternatively, a white IR pixel sensor can be employed.

[0081] In any case, the visible spectrum image information received from one or both of the first imaging device 104, 304 or the second imaging device 120, 306 is used to measure the amount of light falling on the face of the user, and in particular the eye region 212.

[0082] In particular, where the first imaging device 104, 304 is configured to be sensitive to visible light, the device can be configured to take a plurality of images, each with an increased exposure time for any given exposure time of the second imaging device 120, 306. These techniques are typically used to produce high dynamic range (HDR) images, avoiding saturation in very unevenly illuminated scenes. Using such HDR component images enables unsaturated image information to be extracted from the face region of the user 108, 202, thereby avoiding errors in measuring variations in the level of light illuminating the face of the user.

[0083] Knowing the distance of the user's face from the imaging device providing the information, and the gain and exposure parameters of the imaging device, it is possible to determine the illumination of the user's face, and in particular their eye region 212, for any given captured image.

[0084] It is noted that in a variant of the above described implementation, instead of the imaging device providing the illumination information, a photosensor (not shown) with a lens can be used to measure the visible light intensity falling on the user's face including the eye region 212. The field of view of the photosensor should be small enough to be limited to the user's face and exclude any background. Thus, the photosensor can be mounted in the same housing as the second imaging device 120, 306 so that it also moves and continuously monitors the same portion of the environment 110 as the second imaging device 120, 306.

[0085] In any case, the illumination level of the user's face including the eye region 212 can then be correlated with the pupil size to determine the user's response to varying light levels incident on their face.

[0086] This allows a variant of the system 302, and in particular the eye tracking component 320, to track one or more of the following: the immediate response to any significant change in illumination of the user's face, the short-term adaptation of the pupil size to the illumination level, and the long-term adaptation of the pupil size to the illumination level.

[0087] Taking into account age, gender, ethnicity, and any other relevant information, the eye tracking component 320 itself, or indeed any other component such as the control component 318 with this data, can then determine whether any of these responses is within the nominal limits for a given user.

[0088] If not, the control component 318 or any component making such a determination can signal to an external system, such as the vehicle control system 404, that the user can not be in a suitable state to control the vehicle, and appropriate action can be taken, including limiting the speed of the vehicle and / or ensuring that the vehicle can be safely stopped.

[0089] In a variant of the above described implementation, the second imaging device 120, 306 can include a narrow field of view thermal camera focusing on only one eye, instead of or in addition to a near-infrared (NIR) camera.

[0090] In the above described implementation, the component determining whether the user's pupil response to light changes is within the nominal limits is continuously operating during the operation of the vehicle, and the changes in light level are typically caused by changes in ambient light level, including changes in road lighting.

[0091] In a variation of the described implementation, a visible light illuminator (not shown) can be provided and can be directed to illuminate at least the face of the user 108, 202. Thus, for example, the illuminator can be driven by the control assembly 318 or any other suitable assembly to emit a flash of known intensity before the car starts testing the pupil reaction time of the driver in order to detect alcohol intoxication. Moreover, if the pupil reaction of the driver is outside the nominal limits, the control assembly 318 or any assembly making such a determination can signal an external system such as the vehicle control system 404 that the user can not be in a suitable state to control the vehicle and appropriate action can be taken.

[0092] Another function of actuating such a visible light illuminator while the vehicle is driving or stationary can be to induce a blink reflex in any of the users 108, 202 or 112, 206 while deploying one or more respective airbags for that user or each user to minimize the chance of injury to the eyes by debris thrown by the airbags in the event of an accident.

[0093] In yet another variation of the above described implementation, the iris information can be extracted from the image data acquired by the second imaging device 120, 306 and provided to a biometric authentication unit, for example as described in US-2019-364229 (Reference File: FN-629) in order to authenticate the user and allow the user to control one or more functions of the vehicle or access personal information of the user, for example their age, etc. to help decide whether to allow them to control the vehicle.

[0094] While the foregoing application has been described with respect to particular examples, it is to be understood that the scope of the application is not limited to these particular examples. Since other modifications and changes varied to fit particular operating requirements and environments can be apparent to those skilled in the art, the application is not considered limited to the examples chosen for purposes of disclosure, and covers all changes and modifications that do not constitute departures from the true spirit and scope of the application.

[0095] While the present application describes implementations with specific structural features and / or method acts, it is to be understood that the claims are not necessarily limited to the specific features or acts described. Rather, the specific features and acts are merely illustrative of a few of the many ways to implement the claims.

[0096] Conclusion

[0097] While various examples and implementations have been described herein, examples and implementations can be combined, rearranged, and modified, to achieve other variations within the scope of the disclosure.

[0098] Although the embodiments have been described with reference to a particular feature or acts, it will be understood that the disclosure is not limited to the particular features or acts described and / or illustrated. To the contrary, the specific features and acts set forth herein are merely illustrative of specific ways to make and use embodiments of the claimed subject matter. Each of the claims is intended to encompass within its scope any and all equivalents thereof, as well as any and all combinations of elements from different claims.

Claims

1. A method, the method comprising: Repeat the following steps during consecutive image generation times: Using a first imaging device, first image data representing a user is generated, the first imaging device being coupled to the vehicle and configured to image the user when the user is in the driving position of the vehicle; It is determined that a portion of the first image data represents the user's face; At least in part based on said portion of the first image data, an actuator associated with the second imaging device is caused to move from a first position to a second position, thereby causing the second imaging device to rotate, wherein the second imaging device is coupled to the vehicle, and wherein the second imaging device consists of only a single imaging device; Using the second imaging device, generate second image data that at least represents the user's eye; The size of at least one pupil of the user is determined, at least in part, based on the second image data; Determine the level of visible light incident on the user; as well as Output third image data, which indicates the change in the size of the at least one pupil relative to the change in the visible light level from one image generation time to another when the vehicle is in motion.

2. The method according to claim 1, wherein the first imaging device consists of only a single imaging device.

3. The method according to claim 1, further comprising: When the vehicle is stationary, the step of generating second image data that at least represents the user's eyes is performed; A visible light source is actuated and directed toward the user to cause a significant change in the level of visible light incident on the user. as well as After the visible light source is activated, at least one fourth image data representing at least the user's eye is generated.

4. The method of claim 1, wherein the change in visible light level includes one or more of the following: an immediate user response to any significant change in the user's face; The size of at least one pupil adapts to the short-term illumination level; And the long-term adaptation of the size of the at least one pupil to the level of illumination, wherein the variation in the visible light level takes into account at least one or more of the user’s age, sex or ethnicity.

5. The method of claim 1, wherein at least one of the first imaging device or the second imaging device comprises a camera sensitive to at least visible wavelengths, and wherein the method further comprises using visible wavelength image data to determine the level of visible light incident on the user.

6. The method of claim 1, further comprising using a single-cell illumination sensor directed toward the user to determine the level of visible light incident on the user.

7. The method of claim 1, wherein the first image data comprises a plurality of image data, each image data being acquired at a continuously longer exposure time in the first time.

8. The method of claim 2, wherein the second imaging device comprises any one of: a camera, the camera being sensitive to at least infrared wavelengths; or a thermal camera, the thermal camera being configured to image only the user's monocular area.

9. The method according to claim 3, comprising: In response to determining that the vehicle may be involved in an accident, the visible light source is activated to induce at least one blink in the user when at least one vehicle airbag deploys.

10. The method according to claim 1, further comprising: The orientation of the user's face is determined, at least in part, based on said portion of the first image data. The movement of the actuator from the first position to the second position is at least partially based on the direction.

11. The method according to claim 1, further comprising: The user's gaze direction is determined, at least in part, based on the second image data. The data associated with the user's eye position includes data representing the user's gaze direction.

12. The method according to claim 1, wherein: Generating the first image data includes using the first imaging device, using a first frame rate and a first resolution to generate the first image data; Generating the second image data includes using the second imaging device, using a second frame rate and a second resolution to generate the second image data; and At least one of the first frame rates is different from the second frame rate or the first resolution is different from the second resolution.

13. The method according to claim 1, further comprising: Using the first imaging device, generate fourth image data representing the user; It is determined that a portion of the fourth image data represents the user's face; At least in part based on the portion of the fourth image data, the actuator associated with the second imaging device is caused to move from the second position to the third position; Using the second imaging device, fifth image data representing at least the user's eye is generated; The user's eye position is determined, at least in part, based on the fifth image data; and Output sixth image data associated with the user's eye position.

14. A method, the method comprising: Using a first imaging device with a first field of view, first image data at a first resolution is generated, the first image data representing the user; The location of the user's face is determined using at least the first image data; At least in part based on the position, an actuator associated with a second imaging device is moved from a first position to a second position, the second imaging device having a second field of view smaller than the first field of view, thereby causing the second imaging device to rotate, wherein the second imaging device consists of only a single imaging device; Using the second imaging device, second image data at a second resolution lower than the first resolution is generated, and the second image data at least represents the user's eye; At least the second image data is used to determine the user's gaze direction; as well as Output data representing the user's gaze direction.

15. The method of claim 14, wherein determining the location of the user's face comprises: The first image data is analyzed using one or more algorithms associated with face detection; Based at least in part on the analysis of the first image data, it is determined that a portion of the first image data represents the user's face; and The position of the user's face relative to the second imaging device is determined, at least in part, based on the portion of the first image data.

16. The method of claim 14, wherein determining the user's gaze direction comprises: The second image data is analyzed using one or more algorithms associated with eye tracking; The user's eye position is determined, at least in part, based on the analysis of the second image data; as well as The user's gaze direction is determined at least in part based on the eye position.

Citation Information

Patent Citations

  • Image acquisition system for off-axis eye images

    US10909363B2

  • Multispectral image processing system for face detection

    US20190364229A1

  • Apparatus for determining the alertness of a driver

    US6097295A