Neural network based estimation of head pose and gaze using photorealistic synthetic data
By training a neural network to separate head and eye movements, building a facial texture image database and performing 3D reconstruction, the accuracy problem of driver head posture and line of sight angle estimation is solved, and the preventive capability of the driver distraction monitoring system is improved.
Patent Information
- Application Number
- CN201980096086.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-05-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2039-05-13
AI Technical Summary
Existing technologies have difficulty accurately estimating a driver's head posture and viewing angle, resulting in insufficient accuracy of driver distraction monitoring systems and an inability to effectively prevent accidents.
By training a neural network and using multiple two-dimensional facial images to separate head and eye movements, a facial texture image database is constructed. The database is then used to train a second neural network to estimate head posture and gaze angle. Combined with image fusion and 3D reconstruction technology, realistic images are generated for training.
It achieves accurate estimation of the driver's head posture and sight angle, improves the accuracy of the driver distraction monitoring system, and can prevent accidents in a timely manner.
Smart Images

Figure CN114041175B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to training neural networks, and in particular, to training a neural network using a data set to estimate the head pose and gaze angle of a vehicle driver. Background Art
[0002] Driver distraction is an increasingly common cause of vehicle accidents, particularly with the increasing use of technologies such as mobile devices that divert drivers' attention from the road. Driver distraction monitoring and avoidance are key to ensuring a safe driving environment, not only for the distracted driver, but also for other nearby drivers who may be affected by the distracted driver. Vehicles capable of driver monitoring can take measures to prevent or assist in preventing accidents caused by driver distraction. For example, an alarm system can be activated to warn the driver of distraction, or automated functions such as braking and steering can be activated to control the vehicle until the driver is no longer distracted. To detect driver distraction, these alarm and preventative monitoring systems can use the driver's head pose and gaze angle to assess the current state. However, because head and eye movements are often independent of each other, accurate head pose and gaze estimation presents a significant challenge in computer vision technology. Summary of the Invention
[0003] According to one aspect of the present disclosure, a computer-implemented method for estimating head pose and gaze angle is provided, comprising: training a first neural network using a plurality of two-dimensional (2D) facial images to separate head and eye movements of the 2D face, wherein the training of the first neural network comprises: mapping a 2D face in the plurality of 2D facial images to a face position image, and constructing a facial texture image of the 2D face based on the face position image; storing an eye texture image including a gaze angle extracted from the facial texture image of the 2D face in a database; replacing an eye region of the facial texture image with the eye texture image including the gaze angle stored in the database to generate a modified facial texture image; reconstructing the modified facial texture image to generate a modified 2D facial image including a modified head pose and gaze angle as training data, and storing the training data in the database; and estimating the head pose and gaze angle by training a second neural network using the training data, wherein the training of the second neural network comprises: collecting the training data from the database, and simultaneously applying one or more transformations to the modified 2D facial image of the training data and the corresponding eye region of the modified 2D facial image.
[0004] Optionally, in any of the above aspects, the mapping also includes: mapping the 2D faces in the multiple 2D facial images to a position map using a face alignment method, wherein the facial position image aligns the 2D faces in the multiple 2D facial images to 3D coordinates of a reconstructed 3D (three-dimensional, 3D) model of the 2D faces in the multiple 2D facial images; and the constructing also includes: constructing the facial texture image for the 2D faces in the multiple 2D facial images based on the facial position image or the facial 3D deformable model to represent the texture of the aligned 2D faces.
[0005] Optionally, in any of the above aspects, the storing further includes: extracting the facial texture image from the 2D face in the multiple 2D facial images based on the facial position image; cropping the eye region from the facial texture image to create a cropped eye texture image based on the aligned 2D faces in the multiple 2D facial images based on landmarks; and storing the cropped eye texture image in the database.
[0006] Optionally, in any of the above aspects, the cropped eye UV texture image is marked as a difference between the head pose and the gaze angle of the 2D face in the plurality of 2D facial images.
[0007] Optionally, in any of the above aspects, the replacement also includes: selecting the eye area of the cropped eye texture image from the database based on the landmark; and replacing the eye area in the facial texture image with the cropped eye texture image in the database based on the alignment coordinates of the landmark to generate a modified facial texture map of the 2D face in the multiple 2D facial images.
[0008] Optionally, in any of the above aspects, the replacement also includes: applying image fusion to merge the cropped eye texture image selected from the database into the modified facial texture map of the 2D face in the multiple 2D facial images; and training a generative adversarial network (GAN) or using a method based on local gradient information to smooth the color and texture in the eye area of the modified facial texture image.
[0009] Optionally, in any of the above aspects, the computer-implemented method further includes: deforming the modified facial texture image of the 2D face onto a facial 3D morphable model (3DMM) to reconstruct a 3D facial model based on the modified facial texture image using a line of sight direction; applying a rotation matrix to the reconstructed 3D facial model to change the head posture, and changing the line of sight angle to be consistent with the head posture; projecting the 3D facial model after applying the rotation matrix into a 2D image space to generate the modified 2D facial image; and storing the modified 2D facial image in the database.
[0010] Optionally, in any of the above aspects, the gaze direction is calculated by adding the relative gaze direction stored in the cropped eye texture image selected from the database to the head posture.
[0011] Optionally, in any of the above aspects, the estimation further includes: collecting a 2D facial image of a vehicle driver with one or more head postures to generate a driver dataset; and applying the driver dataset to fine-tune the second neural network to estimate the driver's head posture and sight angle.
[0012] Optionally, in any of the above aspects, the 2D facial image of the driver is captured by a capture device and the 2D facial image is uploaded to a network for processing; and the processed 2D facial image of the driver is downloaded to the vehicle.
[0013] Optionally, in any of the above aspects, the first neural network is an encoder-decoder type neural network, used to map the 2D facial image to a corresponding position map.
[0014] Optionally, in any of the above aspects, in the facial position image, the RGB (red, green, blue, RGB) grayscale value of each pixel represents the 3D coordinate of the corresponding facial point in the reconstructed 3D model.
[0015] According to another aspect of the present disclosure, a device for estimating head pose and gaze angle is provided, comprising: a non-transitory memory comprising instructions; and one or more processors in communication with the memory, wherein the one or more processors execute the instructions to perform the following operations: training a first neural network using a plurality of two-dimensional (2D) facial images to separate the movements of the head and eyes of the 2D face, wherein the training the first network comprises: mapping a 2D face in the plurality of 2D facial images to a facial position image, and constructing a facial texture image of the 2D face based on the facial position image; extracting an eye texture image including the gaze angle from the facial texture image of the 2D face The method comprises the steps of: storing the eye texture image including the gaze angle in a database; replacing the eye region of the facial texture image with the eye texture image including the gaze angle stored in the database to generate a modified facial texture image; reconstructing the modified facial texture image to generate a modified 2D facial image including the modified head pose and gaze angle as training data, and storing the training data in the database; and estimating the head pose and gaze angle by training a second neural network using the training data, wherein the training the second neural network comprises: collecting the training data from the database, and simultaneously applying one or more transformations to the modified 2D facial image of the training data and the corresponding eye region of the modified 2D facial image.
[0016] According to another aspect of the present disclosure, a non-transitory computer-readable medium is provided, which stores computer instructions for estimating head posture and gaze angle, wherein when one or more processors execute the computer instructions, the one or more processors perform the following steps: training a first neural network using a plurality of two-dimensional (2D) facial images to separate the movements of the head and eyes of the 2D face, wherein the training of the first network includes: mapping a 2D face in the plurality of 2D facial images to a facial position image, and constructing a facial texture image of the 2D face based on the facial position image; storing an eye texture image including the gaze angle extracted from the facial texture image of the 2D face in a database; replacing an eye region of the facial texture image with the eye texture image including the gaze angle stored in the database to generate a modified facial texture image; reconstructing the modified facial texture image to generate a modified 2D facial image including a modified head pose and gaze angle as training data, and storing the training data in the database; and estimating the head pose and gaze angle by training a second neural network using the training data, wherein the training the second neural network comprises: collecting the training data from the database, and simultaneously applying one or more transformations to the modified 2D facial image of the training data and the corresponding eye region of the modified 2D facial image.
[0017] This summary briefly introduces a series of concepts that are further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. The claimed subject matter is not limited to implementations that solve any or all of the problems raised in the background technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The various aspects of the disclosure are illustrated by way of example and not limitation in the accompanying figures in which like reference numerals indicate elements thereof.
[0019] Figure 1A A system for estimating head pose and gaze estimation according to one embodiment of the present technology is shown;
[0020] Figure 1B Shown according to Figure 1A An exemplary head pose and gaze estimator for
[0021] Figure 2A An example flow chart for estimating head pose and gaze angle according to an embodiment of the present disclosure is shown;
[0022] Figure 2BExamples of origin and coincident head pose and gaze are shown;
[0023] Figure 2C and Figure 2D shows the training of a neural network for two-dimensional (2D) facial data;
[0024] Figure 3A and Figure 3B Shows an example of constructing an eye UV texture dataset;
[0025] Figure 4A and Figure 4B Shows an example of replacing the eye area in a facial UV texture image;
[0026] Figure 5A and Figure 5B An example flow chart of reconstructing 3D faces and generating training data is shown;
[0027] Figure 6 An example of a multimodal CNN (convolutional neural network) for estimating head pose and gaze angle is shown.
[0028] Figure 7A and Figure 7B shows a flowchart for fine-tuning a pre-trained model; and
[0029] Figure 8 A computing system is shown upon which embodiments of the invention may be implemented. DETAILED DESCRIPTION
[0030] The present disclosure will now be described with reference to the accompanying drawings, which generally relate to driver behavior detection.
[0031] A technique for estimating head pose and gaze is disclosed, in which head motion in two-dimensional (2D) facial images can be separated from gaze motion. Using a face alignment method, such as a deep neural network (DNN), head pose is aligned and the 2D facial image is mapped from the 2D image space to a new UV space. The new UV space is a 2D image plane parameterized from a 3D space, used to represent the three-dimensional (3D) geometry (UV position image) and the corresponding texture of the 2D facial image (UV texture image). The UV texture image can be used to crop the eye region (with different gaze angles) to create a dataset of eye UV texture images in a database. For any 2D facial image (e.g., a frontal face image), the eye region in its UV texture image can be replaced with any image in the eye UV texture dataset stored in the database. The facial image can then be reconstructed from the UV space to a 3D space. A rotation matrix is then applied to the new 3D face and projected back into 2D space to synthesize a large number of new realistic images with different head poses and gaze angles. The photorealistic images can be used to train a multimodal convolutional neural network (CNN) to simultaneously estimate head pose and gaze angle. This technique can also be applied to other facial attributes, such as expression or fatigue, to generate datasets on, but not limited to, yawning, closed eyes, and more.
[0032] It should be understood that embodiments of the present invention can be implemented in many different forms, and the scope of the claims should not be interpreted as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention thorough and complete, and will fully convey the embodiment concepts of the present invention to those skilled in the art. In fact, the present invention is disclosed to cover substitutes, modifications and equivalents of these embodiments included in the spirit and scope of the present invention as defined by the appended claims. In addition, in the following detailed description of the disclosed embodiments, many specific details are set forth in order to provide a thorough understanding. However, it will be clear to those skilled in the art that, without these specific details, the present embodiment of the present invention can be implemented.
[0033] Data-driven DNN technology has been one of the most significant advances of the past decade, particularly as it relates to computer vision. Large, accurately labeled datasets are crucial for DNN training. However, there are no readily available, large-scale datasets of head pose and gaze angle with sufficient data volume for such training. This is primarily due to the need for an experimental environment to collect and acquire the data. For example, a commonly used dataset is the Columbia Gaze Dataset, which was collected using a well-designed camera array and a chin-rest with many fixed head poses. While the Columbia Gaze Dataset is an excellent public dataset for algorithm research, it is still insufficient for training a stable gaze and head pose estimation network. An explanation of the Columbia Gaze Dataset is publicly available in "Gaze Locking: Passive Eye Contact Detection for Human-Object Interaction" (B.A. Smith et al.).
[0034] Another example is SmartEye Eye tracking systems are among the most advanced remote gaze analyzers available today. They can accurately estimate a person's head pose and gaze. However, they have several drawbacks, including complex calibration of the imaging system, limited availability of near-infrared (NIR) images for training and testing, and the sensitivity of head pose and gaze estimation to drift in imaging system parameters based on geometric calculations. Furthermore, the imaging system itself is expensive.
[0035] Due to these and other limitations of training datasets, head pose and gaze estimation tasks are usually considered to be two independent tasks in the field of computer vision.
[0036] Figure 1A A system for head pose and gaze estimation according to one embodiment of the present technology is shown. A head pose and gaze estimator 106 is shown mounted or otherwise included within a vehicle 101, which also includes a cabin in which a driver 102 may be seated. The head pose and gaze estimator 106, or one or more portions thereof, may be accessed by an in-cabin computer system and / or a mobile computing device, such as, but not limited to, a smartphone, tablet, notebook computer, laptop computer, and / or the like.
[0037] According to certain embodiments of the present technology, the head pose and gaze estimator 106 obtains current data of the driver 102 of the vehicle 101 from one or more sensors. In other embodiments, the head pose and gaze estimator 106 also obtains additional information related to features such as facial features of the driver 102, historical head pose and gaze information, and the like from one or more databases 140. The head pose and gaze estimator 106 analyzes the current data and / or additional information of the driver 102 of the vehicle 101 to identify the driver's head pose and gaze. As described below, such analysis can be performed using one or more computer-implemented neural networks and / or some other computer-implemented models.
[0038] like Figure 1A As shown, the driver distraction system 106 is communicatively coupled to a capture device 103 that can be used to capture current data of the driver of the vehicle 101. In one embodiment, the capture device 103 includes sensors and other devices for capturing current data of the driver 102 of the vehicle 101. The captured data can be processed by a processor 108 that includes hardware and / or software to detect and track driver movement, head posture, and gaze direction. As will be described in further detail below, reference Figure 1B , the capture device may additionally include one or more cameras, microphones, or other sensors to capture data.
[0039] In one embodiment, the capture device 103 may be configured as follows: Figure 1A It is shown outside the driver distraction system 106, or it can be included as part of the driver distraction system 106 depending on the specific implementation. Figure 1B Additional details of the driver distraction system 106 according to certain embodiments of the present technology are described.
[0040] Still refer to Figure 1ADriver distraction system 106 is also shown as being communicatively coupled to various types of vehicle-related sensors 105 included in vehicle 101. Such sensors 105 may include, but are not limited to, speedometers, global positioning system (GPS) receivers, and clocks. Driver distraction system 106 is also shown as being communicatively coupled to one or more communication networks 130, which provide access to one or more databases 140 and / or other types of data stores. Databases 140 and / or other types of data stores may store vehicle data for vehicle 101. Examples of such data include, but are not limited to, driving record data, driving performance data, driver's license type data, driver's facial features, driving head posture, driver's line of sight, etc. Such data may be stored in a local database or other data store located within vehicle 101. However, the data may be stored in one or more databases 140 or other data stores located remotely from vehicle 101. Thus, such databases 140 or other data stores may be communicatively coupled to the driver distraction system via one or more communication networks 130.
[0041] The communication network 130 may include a data network, a wireless network, a telephone network, or any combination thereof. It is contemplated that the data network may be any local area network (LAN), metropolitan area network (MAN), wide area network (WAN), a public data network (e.g., the Internet), a short-range wireless network, or any other suitable packet-switched network. In addition, the wireless network may be, for example, a cellular network, and may employ various technologies, including Enhanced Data rates for GSM Evolution (EDGE), general packet radio service (GPRS), global system for mobile communications (GSM), Internet protocol multimedia subsystem (IMS), universal mobile telecommunications system (UMTS), and any other suitable wireless medium, such as worldwide interoperability for microwave access (WiMAX), long term evolution (LTE) network, code division multiple access (CDMA), wideband code division multiple access (WCDMA), wireless fidelity (Wi-Fi), wireless LAN (WLAN), Internet Protocol (IP) data transmission, satellite, mobile ad-hoc network (MANET), etc., or any combination thereof. The communication network 130 can be connected to the communication device 102 ( Figure 1B ) etc. provide communication capabilities between the driver distraction system 106 and the database 140 and / or other data storage.
[0042] Although reference is made to vehicle 101 Figure 1A Although embodiments of the present invention are provided, it should be understood that the disclosed technology can be applied to a wide range of technical fields and is not limited to vehicles. For example, in addition to vehicles, the disclosed technology can be used in virtual or augmented reality devices or simulators, where head pose and gaze estimation, vehicle data, and / or scene information may be required.
[0043] Now refer to Figure 1B
[0046] Further details of the driver distraction system 106 according to certain embodiments of the present technology are described. The driver distraction system 106 includes a capture device 103, one or more processors 108, a vehicle system 104, a navigation system 107, a machine learning engine 109, an input / output (I / O) interface 114, a memory 116, visual / audio alerts 118, a communication device 120, and a database 140 (which may also be part of the driver distraction system).
[0044] The capture device 103 can be responsible for monitoring and identifying driver behavior based on captured driver motion and / or audio data using one or more capture devices located in the cab, such as sensors 103A, cameras 103B, or microphones 103C. In one embodiment, the capture device 103 is used to capture the motion of the driver's head and face, while in other embodiments, the motion of the driver's torso and / or the driver's limbs and hands is also captured. For example, the detection and tracking 108A, the head pose estimator 108B, and the gaze direction estimator 108C can monitor the driver's motion captured by the capture device 103 to detect specific gestures, such as head pose, or whether a person is looking in a specific direction.
[0045] Other embodiments include capturing audio data via microphone 103C, either together with or separately from the driver movement data. The captured audio may be, for example, an audio signal of driver 102 captured by microphone 103C. The audio may be analyzed to detect various functions that may vary depending on the driver's state. Examples of such audio features include the driver's voice, passenger voices, music, etc.
[0046] Although capture device 103 is depicted as a single device having multiple components, it should be understood that each component (e.g., sensor, camera, microphone, etc.) can be a separate component located in a different area of vehicle 101. For example, sensor 103A, camera 103B, microphone 103C, and depth sensor 103D can each be located in a different area of the cabin. In another example, a single component of capture device 103 can be part of another component or device. For example, camera 103B and visual / audio alert 118 can be part of a mobile phone or tablet (not shown) placed in the vehicle cabin, while sensor 103A and microphone 103C can be separately located in different locations in the vehicle cabin.
[0047] Detection and tracking 108A monitors facial features of driver 102 captured by capture device 103, and can then extract the facial features after the driver's face is detected. The term facial features includes, but is not limited to, points surrounding the eye, nose, and mouth areas, as well as points outlining portions of the detected facial contour of driver 102. Based on the monitored facial features, the initial position of one or more eye features of driver 102's eyeballs can be detected. The eye features can include an iris and a first corner of the eye and a second corner of the eye. Thus, for example, detecting the position of each of the one or more eye features includes detecting the position of the iris, detecting the position of the first corner of the eye, and detecting the position of the second corner of the eye.
[0048] The head pose estimator 108B uses the monitored facial features to estimate the head pose of the driver 102. As used herein, the term "head pose" describes an angle that refers to the relative orientation of the driver's head with respect to the plane of the capture device 103. In one embodiment, the head pose includes the yaw and pitch angles of the driver's head with respect to the plane of the capture device. In another embodiment, the head pose includes the yaw, pitch, and roll angles of the driver's head with respect to the plane of the capture device. Figure 5B Describe head posture in more detail.
[0049] The gaze direction estimator 108C estimates the gaze direction (and gaze angle) of the driver. In operation of the gaze direction estimator 108C, the capture device 103 may capture an image or set of images (e.g., of the driver of the vehicle). The capture device 103 may transmit the images to the gaze direction estimator 108C, where the gaze direction estimator 108C detects facial features from the images and tracks the driver's gaze (e.g., over time). Smart Eye The eye tracking system is such a gaze direction estimator.
[0050] In another embodiment, the gaze direction estimator 108C can detect the eyes from the captured image. For example, the gaze direction estimator 108C can rely on the center of the eye to determine the gaze direction. In short, it can be assumed that the driver is looking forward relative to the orientation of his head. In some embodiments, the gaze direction estimator 108C provides more accurate gaze tracking by detecting pupil or iris position or using a geometric model that is based on the estimated head pose and the detected iris and the position of each of the first and second eye corners. Pupil and / or iris tracking enables the gaze direction estimator 108C to detect gaze direction separately from the head pose. Drivers typically visually scan their surroundings with little or no head movement (e.g., glancing left or right (or up or down) to better see items or objects outside their direct line of sight). These visual scans are often directed at objects on or near the road (e.g., to view road signs, pedestrians near the road, etc.) and objects within the vehicle's cab (e.g., to view console readouts such as the vehicle speed, operate the radio or other dashboard devices, or view / operate a personal mobile device). In some cases, the driver may glance at some or all of these objects with minimal head movement (e.g., out of the corner of their eye). By tracking the pupil and / or iris, the gaze direction estimator 108C can detect upward, downward, and sideways gazes, which are not detectable in systems that track head position alone.
[0051] In one embodiment, based on the detected facial features, the gaze direction estimator 108C can cause the processor 108 to determine the gaze direction (e.g., the driver's gaze on the vehicle). In some embodiments, the gaze direction estimator 108C receives a series of images (and / or videos). The gaze direction estimator 108C can detect facial features in multiple images (e.g., a series of images or an image sequence). Thus, the gaze direction estimator 108C can track the gaze direction over time and store this information in, for example, the database 140.
[0052] In addition to the gesture and gaze detection described above, the processor 108 may also include an image corrector 108D, a video enhancer 108E, a video scene analyzer 108F, and / or other data processing and analysis to determine scene information captured by the capture device 103 .
[0053] Image corrector 108D receives the captured data and can perform corrections such as video stabilization. For example, bumps in the road may cause the data to be jittery, blurry, or distorted. The image corrector can stabilize the image to prevent horizontal and / or vertical jitter, and / or correct for translation, rotation, and / or scaling.
[0054] The video enhancer 108E can perform additional enhancement or processing in situations with poor lighting or high data compression. Video processing and enhancement may include, but is not limited to, gamma correction, dehazing, and / or deblurring. Other video processing enhancement algorithms may be used to reduce noise in low-light video inputs, followed by contrast enhancement techniques such as, but not limited to, tone mapping, histogram stretching and equalization, and gamma correction to restore visual information in low-light video.
[0055] The video scene analyzer 108F can identify the content of the video from the capture device 103. For example, the content of the video may include a scene or scene sequence from the forward-facing camera 103B in the vehicle. The analysis of the video may involve a variety of technologies, including but not limited to low-level content analysis such as feature extraction, structural analysis, object detection, tracking, and high-level semantic analysis such as scene analysis, event detection, and video mining. For example, by identifying the content of the incoming video signal, it can be determined whether the vehicle 101 is traveling along a highway or within a city, whether there are pedestrians, animals, or other objects / obstacles on the road, etc. By performing image processing (e.g., image correction, video enhancement, etc.) before or simultaneously with image analysis (e.g., video scene analysis), the image data can be prepared in an appropriate manner for the type of analysis being performed. For example, in image correction performed to reduce blur, video scene analysis can be performed more accurately by removing edge lines used for object recognition.
[0056] The vehicle system 104 may provide signals corresponding to any state of the vehicle, the vehicle environment, or output from any other information source connected to the vehicle. For example, vehicle data outputs may include analog signals (e.g., current speed), digital signals provided by individual information sources (e.g., a clock, a thermometer, a position sensor such as a Global Positioning System (GPS) sensor, etc.), and digital signals transmitted via a vehicle data network (e.g., an engine controller area network (CAN) bus through which engine-related information may be transmitted, a climate control CAN bus through which climate control-related information may be transmitted, and a multimedia data network through which multimedia data may be transmitted between multimedia components in the vehicle). For example, the vehicle system 104 may retrieve from the engine CAN bus the current speed of the vehicle estimated by wheel sensors, the power status of the vehicle via the vehicle's battery and / or power distribution system, the ignition status of the vehicle, and the like.
[0057] The navigation system 107 of the vehicle 101 can generate and / or receive navigation information, such as location information (e.g., via a GPS sensor and / or other sensors 105), route guidance, traffic information, point-of-interest (POI) identification, and / or provide other navigation services to the driver. In one embodiment, the navigation system or a portion of the navigation system is communicatively coupled to the vehicle 101 and is located remotely from the vehicle 101.
[0058] Various input / output devices can be used to present information to a user and / or other components or devices via the input / output interface 114. Examples of input devices include a keyboard, a microphone, touch functionality (e.g., a capacitor or other sensor for detecting physical touch), a camera (e.g., which can use visible or non-visible wavelengths such as infrared frequencies to identify motion as gestures that do not involve touch), etc. Examples of output devices include visual / audio alerts 118, such as displays, speakers, etc. In one embodiment, the I / O interface 114 receives driver motion data and / or audio data of the driver 102 from the capture device 103. The driver motion data can be related to the eyes and face of the driver 102, etc., which can be analyzed by the processor 108.
[0059] The data collected by driver distraction system 106 can be stored in database 140, memory 116, or any combination thereof. In one embodiment, the collected data comes from one or more sources external to vehicle 101. The stored information can be data related to driver distraction and safety, such as information captured by capture device 103. In one embodiment, the data stored in database 140 can be a collection of data collected for one or more drivers of vehicle 101. In one embodiment, the collected data is head posture data of the driver of vehicle 101. In another embodiment, the collected data is gaze direction data of the driver of vehicle 101. The collected data can also be used to generate datasets and information that can be used to train models for machine learning, such as machine learning engine 109.
[0060] In one embodiment, the memory 116 may store instructions executable by the processor 108 and the machine learning engine 109, as well as programs or applications (not shown) that may be loaded and executed by the processor 108. In one embodiment, the machine learning engine 109 includes executable code stored in the memory 116 that may be executed by the processor 108 and selects one or more machine learning models stored in the memory 116 (or database 140). The machine models may be developed using well-known and conventional machine learning and deep learning techniques, such as implementations of convolutional neural networks (CNNs), as described in more detail below.
[0061] Using all or part of the data collected and acquired from the various components, the driver distraction system 106 can calculate a level of driver distraction. The level of driver distraction can be based on a threshold level input into the system, or based on previously collected and acquired information (e.g., historical information), which is analyzed to determine when the driver is considered distracted. In one embodiment, a weight or score can represent the level of driver distraction and can be based on information obtained from observing the driver, vehicle, and / or the surrounding environment. These observations can be compared to threshold levels or previously collected and acquired information. For example, a route in bad weather, during rush hour, or at night may require a higher level of driver attention than a route in good weather, during off-peak hours, and during daytime surroundings. Within these sections, portions where lower levels of driver distraction are likely to occur are considered safe driving zones, while portions where higher levels of driver distraction are likely to occur are considered distracted driving zones. In another example, a driver may require a higher level of attention when driving along a winding road or highway than when driving along a straight road or a cul-de-sac. In this case, when driving along a winding road or a highway, the driver may be more distracted for part of the route, while when driving along a straight road or a cul-de-sac, the driver may be less distracted for part of the route.
[0062] Other examples include calculating a driver distraction score when the driver is looking forward (e.g., as determined based on an interior image) compared to when the driver is looking down or to the side. When the driver is considered to be looking forward, the associated score (and level of distraction) will be considered lower than when the driver is looking down or to the side. Many other factors may be considered when calculating the score, such as the noise level of the vehicle cabin (determined based on detected sound information, etc.), or the direction of sight that is obscured but detectable by the vehicle's sensors (e.g., determined from a vehicle proximity sensor, external images, etc.) toward a dangerous or unsafe object. It will be appreciated that other driver distraction scores may be calculated as long as any other suitable set of inputs is provided.
[0063] Figure 2A An example flow chart for estimating head pose and gaze angle according to an embodiment of the present disclosure is shown. In an embodiment, the flow chart may be a computer-implemented method that is at least partially performed by hardware and / or software components shown in the various figures and as described herein. In an embodiment, the disclosed process may be performed by Figure 1A and Figure 1B In one embodiment, a software component executed by one or more processors, such as processor 108 or processor 802, performs at least a portion of the process.
[0064] In process 200, the head pose and gaze angle of a person (e.g., a driver of a vehicle) are estimated. Steps 210-216 involve generating a dataset with accurate head pose and gaze angle labels, which is used in step 218 to train a multimodal CNN for estimating head pose and gaze angle. To compute head pose and gaze estimates, as Figure 2B As shown, when the person's head is facing the front of the capture device 103 (e.g., camera 103B), the yaw, pitch, and roll angles (α, β, γ) of the head pose are equal to (0°, 0°, 0°). In one embodiment, when the camera 103B captures a two-dimensional (2D) face image (i.e., a facial image) of the person, the gaze angle of the person's eyes is determined by the position of the center of the pupil relative to the corner of the eye.
[0065] Based on this assumption, different poses can be used to align a person's head so that the movement of the head (pose) can be separated from the movement of the eyes (gaze). In one embodiment, as described below, by changing the position of the pupil center (relative to the eye corner) in the aligned images, the gaze angle can also be changed when reconstructing the 2D image. When making these determinations, the origins of the head pose and gaze coordinates are coincident, as shown in Figure 2. Figure 2B According to this embodiment, origin sight line 201' and origin head pose 203' are shown as two separate dashed lines, where the sight line angle is (h, v)' and the head pose angle has three rotation angles: yaw, pitch, and roll (α1, β1, γ1)'. The coordinates of origin sight line 201' and origin head pose angle 203' (e.g., (h, v)' and (α1, β1, γ1)') are identical (or nearly identical) to the coincident origin head pose and sight line 205' (i.e., the head pose and sight line are shown as a single dashed line).
[0066] Based on the above assumptions, in step 210, an encoder-decoder deep neural network (DNN) face alignment method is trained using the 2D face image to generate an aligned UV position image and a facial UV texture image of the 2D face to separate the head and eye movements of the 2D face, as shown below. Figure 2C and 2D Then, in step 212 ( Figure 3A and Figure 3B ), extract the eye UV texture image including the sight angle from the 2D face UV texture image and store it in the database. In step 214 ( Figure 4A and Figure 4B), the eye region of the facial UV texture image is replaced with an eye UV texture image including a sight angle retrieved from a database. Replacing the eye region with the eye UV texture image generates a modified facial UV texture image. In step 216, the modified facial UV texture image is reconstructed to generate a modified 2D facial image including a modified head pose and sight angle as training data stored in the database ( Figure 5A and Figure 5B In step 218, the head pose and gaze angle of the person are estimated by training a CNN using the training data ( Figure 6 ). Each step is explained in detail below with reference to the corresponding figures.
[0067] Figure 2C and Figure 2D An example of training a neural network for two-dimensional (2D) face alignment is shown. The training of the neural network in the described embodiment is to Figure 2A A detailed description of the training step 210 is given in FIG. Figure 2C , the 2D facial image 202 is input into a deep neural network (DNN) 203, such as an encoder-decoder type DNN (or face alignment network), where the machine learning engine 109 aligns the 2D facial image into a facial UV position image (or position map) 204, and constructs a facial UV texture image (or texture map) 206. In other words, the facial UV position image 204 represents the complete 3D facial structure of the 2D image, recording the 3D positions of all points in the UV space while maintaining a close correspondence with the semantic meaning of each point in the UV space. As will be understood by those skilled in the art, UV space (or UV coordinates) is a 2D image plane parameterized from a 3D space that can be used to represent 3D geometry (i.e., the facial UV position image 204) and the corresponding texture of the face (i.e., the facial UV texture image 206), where "U" and "V" are the axes of the image plane (as "X", "Y", and "Z" are used as coordinates in 3D space). In one embodiment, the dataset used to train the DNN 203 is a public dataset, such as the 300W-LP (Large Pose) dataset. Although this example involves a neural network, it should be understood that other face alignment methods can be used to generate the facial UV position image (or position map) 204 and then construct the corresponding facial UV texture image 206.
[0068] In one example, the 2D image is Figure 2D, which is processed by an encoder-decoder type DNN 203. For head pose image 202A, encoder-decoder type DNN 203 aligns the head (face) and maps the facial image from 2D image space to a corresponding UV space. Thus, the 2D image plane is parameterized from 3D space to represent the corresponding 3D geometry (facial UV position image 204A) and facial texture (facial UV texture image 206A). Thus, the head and eye movements are separated, so that the head pose (facial UV position image) and the gaze direction (facial UV texture image) are separated from each other.
[0069] In one embodiment, the generation of the UV position image and the UV texture map is implemented according to "Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network" by Feng et al. However, it should be understood that any number of known face alignment techniques may be used to generate the facial UV position image and the facial UV texture image from the head pose image.
[0070] Figure 3A and Figure 3B An example of constructing an eye UV texture dataset is shown. In the described embodiment, the construction of the eye UV texture dataset is a detailed description of step 212 of FIG. 2 . Figure 3A and 3B , a head pose image (with a known head pose and gaze angle) is input into DNN203, wherein steps 202-206 may be repeated for each head pose image. For example, head pose image 302A is input into DNN203. As described above, head pose image 302A is input into DNN203, and an aligned facial UV position image and facial UV texture image are generated. Then, in step 304, facial UV texture image 304A may be extracted from the generated information. For example, after processing the input head pose image 302A, facial UV texture image 304A is extracted from DNN203. Using facial UV texture image 304A extracted from step 304, eye region 310 may be cropped from facial UV texture image 304A in step 306 to generate eye UV texture image 306A.
[0071] In one embodiment, the eye region is cropped from the extracted facial UV texture image 304A based on the aligned facial landmarks determined during the face alignment performed by the DNN 203 (described below). For example, for a 2D image (i.e., head pose image) with a known head pose with angle (α1, β1, γ1)′ and a line of sight with angle (h, v)′α, the above reference Figure 2B The difference between the head pose angle and the gaze angle is calculated as the eye UV texture image 306A with the eyes cropped. More specifically, if the 2D facial image has 3D head pose Euler angles Yaw, Pitch, and Roll such that H = (α1, β1, γ1)', and gaze Euler angles Yaw and Pitch such that G = (h, v)', the difference can be calculated as:
[0072]
[0073] Then, in step 308, an eye UV texture dataset can be constructed using each eye UV texture image 306A from each input 2D facial image for storage in a database (e.g., database 308A). The stored eye UV texture dataset can be retrieved for subsequent processing to replace the eye region of the facial UV texture image with the eye region in database 308. In one embodiment, the eye UV texture database is constructed using any 2D facial image dataset with known head pose and gaze angle, such as the Columbia gaze dataset.
[0074] Figure 4A and Figure 4B An example of replacing the eye region in the facial UV texture image is shown. The replacement of the eye region in the facial UV texture image of the depicted embodiment is a detailed description of step 214 of FIG. 2 . Before replacing the eye region in the facial UV texture with an eye UV texture image selected from the eye UV texture image dataset stored in the database 308A, a 2D facial image (e.g., a front-view facial image of (α2, β2, γ2)′=(0°, 0°, 0°)) with a known head pose (α2, β2, γ2)′ is input to the DNN 203 to generate the eye region in the facial UV texture image corresponding to the eye UV texture image. Figure 2C At steps 402 and 404 of steps 202-206 in FIG. 2 , the aligned facial UV position image 204A and facial UV texture image 206A are obtained.
[0075] Then, in step 406 (corresponding to Figure 3AIn step 306 of the above process, an eye region 310 is cropped from the aligned facial UV texture image 304A based on the UV space (coordinates) of the eye landmarks. During the generation of the UV position image 204A and the UV texture image 206A, the UV space of the eye landmarks used to crop the eye region 310 is directly determined by the DNN 203. In one embodiment, the UV space of the eye landmarks can be determined using any number of different well-known face alignment or facial landmark localization techniques, such as, but not limited to, regression techniques, active appearance models (AAMs), active shape models (ASMs), constrained local models (CLMs), mnemonic descent methods, and cascaded autoencoders, cascaded CNNs, generative adversarial networks (GANs), and the like.
[0076] Then, the eye region 310 cropped from the facial UV texture image in step 404 is replaced with the eye region 310 having the eye UV texture selected from the database 308A storing eye UV textures, as described above with reference to FIG. Figure 3A and 3B For example, in step 408 (corresponding to Figure 3A In step 308), the cropped eye region 306A is selected from the eye UV texture image in the database 308A and replaces the eye region of the facial UV texture image 410A, as shown in FIG. Figure 4B The resulting image is the modified facial UV texture map 412A. Figure 5A The output is processed in steps 502-506.
[0077] In another embodiment, Gaussian mixture image synthesis is used to replace the eye region 310 of the facial UV texture image 410A. It should be understood that any number of different eye region replacement techniques may be employed, as will be readily appreciated by those skilled in the art.
[0078] In some embodiments, due to different imaging conditions between the current input 2D facial image in step 402 and the eye UV texture image 306A selected from database 308A, replacing the eye region of the facial UV texture image with the eye UV texture image from database 308A may result in at least partial visual discontinuity in color distribution and / or texture. To eliminate this visual discontinuity, a gradient-based image fusion algorithm may be used to merge the selected eye UV texture image 306A into the aligned facial UV texture image 410A. For example, image fusion techniques may use gradient-based methods to preserve important local perceptual cues while avoiding traditional issues such as aliasing, ghosting, and halos. An example of such a technique, which can be used to merge (fuse) two images, is described in "Image Fusion for Context Enhancement and Video Surrealism" by Raskar et al. In another embodiment, a generative adversarial network (GAN) may be trained to smoothly modify local color / texture details in the eye region to perform the replacement step. That is, GAN can use unlabeled real data to improve the realism of images from the simulator while preserving the annotation information.
[0079] Figure 5A and Figure 5B An example flow chart of reconstructing a 3D face and generating training data is shown. Reconstructing a 3D face and generating training data in the described embodiment is a detailed description of step 216 of FIG. 2 . After outputting the modified UV texture image for the 2D face image ( Figure 4A In step 412), a dataset (e.g., a 2D photorealistic synthetic dataset) may be generated by reconstructing a 3D facial model, rotating the reconstructed 3D facial model, and projecting the rotated 3D facial model back into a 2D image space, as described below with reference to steps 502-510.
[0080] In step 506, step 504 (corresponding to Figure 4A The modified facial UV texture 412A of step 412) is warped onto the facial 3D morphable model (3DMM) of step 502 to reconstruct a 3D facial model 512 of the head of the person with the modified gaze direction, as shown in FIG. Figure 5BIn one embodiment, a 2D facial image such as 2D facial image 202A ( Figure 2D ) is fitted to a 3DMM, and this fitting is achieved by minimizing the difference between the image and model appearance. In a variation, regression-based 3DMM fitting can be applied, whereby the model parameters are estimated by regressing features at landmark locations. Examples of warping 2D facial images onto 3DMMs are described in “Face alignment across large poses: a 3D solution” by Zhu et al. and “Appearance-based gaze estimation in the wild” by Zhang et al.
[0081] Although the facial UV position image can be used to reconstruct the 3D facial model 512, other techniques can also be used. For example, a 3D dense face alignment (3DDFA) framework can also be used to reconstruct the 3D facial model 512. In another example, a face alignment technique can be used to reconstruct the 3D facial model 512. This technique is described in "How far are we from solving the 2D & 3D face alignment problem? (and a dataset of 230,000 3D facial landmarks)" by Bulat et al.
[0082] After reconstruction, the rotation matrix R is applied to the reconstructed 3D facial model 512 to change the person's head pose in step 508. Changing the head pose also changes the angle of view of the person's eyes because the reconstructed 3D facial model is considered to be a rigid object. After rotation, in step 510, the 3D facial model 512 is projected back into the 2D image space. It should be understood that other projection techniques can be applied in step 510. For example, the 3D facial model 512 can be projected directly onto a 2D plane and an artificial background added to it. In another example, a 3D image meshing method can be used to rotate the head and project the 3D head back into 2D. Such a technique is disclosed in "Face alignment across large poses: a 3D solution" by Zhu et al.
[0083] In one embodiment, the generated 2D facial image has a known head pose (α2+α3, β2+β3, γ2+γ3) based on the rotation matrix R, and the gaze direction (α2+α3, β2+β3, γ2+γ3) is obtained by adding the relative gaze direction of the eye UV texture image selected from the database 308A to the head pose. For example, when (α=0, β3=0, γ3=0), the gaze angle becomes (α2+θ, β2+Φ)', while the head pose remains unchanged.
[0084] refer to Figure 5B , if the three Euler angles (yaw, pitch, and roll) of the rotation applied to the reconstructed 3D facial model 512 are (α3, β3, γ3), the rotation matrix R can be written as:
[0085] R=R y (α3)R x (β3)R z (γ3)…(2),
[0086] in:
[0087]
[0088] As shown in the figure, the three matrices represent the basic rotation matrices around the X, Y, and Z axes. For any point V = [x, y, z] on the 3D face model T , after rotation, point V becomes:
[0089]
[0090] After rotation, the new head pose can be calculated as:
[0091] H ′ =(α2+α3,β2+β3,γ2+γ3)…(4),
[0092] And the sight angle can be calculated as:
[0093]
[0094] To project the rotated 3D facial model 512 back into 2D image space in step 510 , the camera intrinsic matrix is applied to the 3D facial model 512 according to the following:
[0095]
[0096] Where [x′, y′, z′] T is the new 3D space (coordinate) after rotation, f x and f y is the focal length expressed in pixels, (c x ,c y ) is the principal point usually located at the center of the image, s is the scaling factor, [u,v] T are the coordinates of the corresponding points on the 2D image.
[0097] In order to generate a 2D facial image dataset with head pose and gaze angle (e.g., a 2D realistic synthetic dataset), the rotation operation of step 508 is repeated for each 3D facial model reconstructed based on the 3DMM facial model and the modified facial UV texture. The resulting dataset can then be stored as a 2D realistic synthetic dataset, for example, in a database.
[0098] Figure 6 An example of a multimodal CNN for estimating head pose and gaze angle is shown. In the described embodiment, the CNN 600 for simultaneously estimating head pose and gaze angle is a Figure 2A A detailed description of step 218 will be used in the execution Figure 5ACNN 600 is trained end-to-end using the 2D realistic synthetic dataset generated during step 510 of . As shown, CNN 600 receives two inputs: a 2D facial image 602 and an eye region 610 of the 2D facial image 602. Each input is passed to a stack of convolutional layers 604A and 604B, respectively, where batch normalization and activation functions can be applied. For example, in one embodiment, each convolutional layer 604A and 604B is followed by a batch normalization (BN) layer (not shown). Each convolutional layer 604A and 604B simultaneously extracts deeper features based on the input features provided by the previous layer and automatically learns task-related feature representations of the 2D facial image 602 and the eye region 610. The features extracted from each input (i.e., head pose and gaze) are flattened and merged for input into a fully connected layer 606, and head pose angles 608A (e.g., yaw, roll, pitch) and gaze angles 608B (e.g., theta, phi) of a 2D facial image input from a 2D realistic synthetic dataset are estimated in layer 606. In one embodiment, processing the input in the manner described above may also be referred to as applying a transformation to the input data. More generally, when the input data is processed by CNN 600, transformations are applied to different layers of the network. These transformations can be linear transformations, statistical normalizations, or other mathematical functions. Non-limiting examples of transformations include mirroring, rotation, smoothing, contrast reduction, etc., which can be applied to the input data.
[0099] Figure 7A and Figure 7B A flow chart for fine-tuning a pre-trained model is shown. In one embodiment, reference Figure 7A, the driver 102 in vehicle 101 ( FIG. 1 ) can provide an image with the driver's head pose to fine-tune the CNN model. The dataset 702A (e.g., a 2D photorealistic synthetic dataset) generated according to the above steps is used to train a CNN, such as CNN 600, in a laboratory 702 (or any other location where processing equipment is located). The trained CNN model 702B can be stored, for example, in a vehicle memory 702F of vehicle 101. The driver 102 of vehicle 101 will have a captured 2D facial image 702C. In real-world applications such as real-time applications 702G, to preserve more details of the driver's facial texture, as much of the subject's frontal facial image as possible will be captured. The 2D facial image of the driver 102 can be captured by a capture device 103 within vehicle 101. Capture device 103 can be part of the vehicle, such as a sensor 103A, a camera 103B, or a depth sensor 103D, or can be located within the vehicle, such as a mobile phone or tablet. In one embodiment, the 2D facial image is captured while the driver 102 is driving vehicle 101. In another embodiment, a 2D facial image is captured when the vehicle 101 is parked or in a non-moving state. The captured 2D facial image of the driver can then be used to perform the aforementioned steps (e.g., Figures 2A-6 ) generates a driver dataset 702D (e.g., a driver 2D simulation dataset). In one embodiment, the driver dataset 702D can be stored in the vehicle memory 702F. The driver dataset 702D will be used to fine-tune the CNN model 702B to more accurately estimate the head pose and gaze angle of the driver 102 and used during real-time application 702G (e.g., during driving the vehicle 101).
[0100] In another embodiment, reference Figure 7B , the driver 102 of the vehicle 101 ( FIG1 ) can provide an image with the driver’s head pose to fine-tune the CNN model. Figure 7A In different embodiments, the driver 102 may capture the 2D facial image 704C before driving the vehicle 101. In real applications such as real-time applications 704G, in order to preserve more details of the driver's facial texture, as much of the subject's front-view facial image as possible is captured. In one embodiment, a capture device 103 such as a camera 103B, a depth sensor 103D, a mobile phone, or a tablet computer may be used to capture the 2D facial image 704C. The captured 2D facial image 704C may then be uploaded to the cloud 704, where it may be used to generate a 3D facial image. Figures 2A-6 In the steps of, (based on the different 2D facial images being uploaded) a dataset 704D (e.g., a 2D realistic synthetic dataset) with different postures and sight angles is generated for the driver 102. In the cloud 704, using Figures 2A-6A CNN model 704B, such as CNN 600, is pre-trained using a large dataset 704A (which may be stored in a database) of different objects generated in the steps of FIG. Subsequently, a driver dataset 704D is applied to fine-tune the pre-trained CNN model 704B to more accurately estimate the head pose and gaze angle of the driver 102. The fine-tuned CNN model 704E can be downloaded to the vehicle 101 and stored, for example, in the vehicle memory 704F and used during real-time application 704G (e.g., during driving the vehicle 101).
[0101] Figure 8 800 is a computer system that can implement embodiments of the present invention. The computer system 800 can be programmed (e.g., via computer program code or instructions) to use driver behavior detection as described herein to enhance driver safety, and includes a communication mechanism such as a bus 810 to pass information between other internal components of the computer system 800 and external components. In one embodiment, the computer system 800 is Figure 1A The computer system 800, or a portion thereof, constitutes a means for performing one or more steps for enhancing driver safety using the driver behavior detection method.
[0102] The bus 810 includes one or more parallel information conductors that enable rapid transfer of information between devices coupled to the bus 810. One or more processors 802 are coupled to the bus 810 for processing information.
[0103] One or more processors 802 perform a set of operations on information (or data) specified by a computer program code to enhance driver safety using driver behavior detection. Computer program code is a set of instructions or statements that provide instructions for the operation of a processor and / or computer system to perform a specified function. For example, the code can be written in a computer programming language that is compiled into a native instruction set of the processor. The code can also be written directly using a native instruction set (e.g., machine language). The set of operations includes importing information from bus 810 and placing information on bus 810. Each operation in the set of operations that can be performed by the processor is represented to the processor by information called an instruction (e.g., an operation code of one or more digits). The sequence of operations, such as the sequence of operation codes to be executed by the processor 802, constitutes processor instructions, also known as computer system instructions or simply computer instructions.
[0104] The computer system 800 also includes a memory 804 coupled to the bus 810. The memory 804, such as a random access memory (RAM) or any other dynamic storage device, stores information including processor instructions for enhancing driver safety using driver behavior detection. Dynamic memory allows the computer system 800 to change the information stored therein. RAM allows information units stored at locations called memory addresses to be stored and retrieved independently of information at adjacent addresses. The memory 804 is also used by the processor 802 to store temporary values while executing processor instructions. The computer system 800 also includes a read-only memory (ROM) 806 or any other static storage device coupled to the bus 810 for storing static information. A non-volatile (persistent) storage device 808, such as a magnetic disk, optical disk, or flash memory card, is also coupled to the bus 810 for storing information including instructions.
[0105] In one embodiment, information including instructions for enhancing distracted driver safety using a head pose and gaze estimator is provided to bus 810 for use by the processor from an external input device 812, such as a keyboard, microphone, infrared (IR) remote control, joystick, game pad, stylus, touch screen, head-mounted display, or sensor operated by a human user. A sensor detects conditions in its vicinity and converts these detections into physical representations compatible with measurable phenomena used to represent information in computer system 800. Other external devices coupled to bus 810 are primarily used for human interaction and include: a display device 814 for presenting text or images; a pointing device 816, such as a mouse, trackball, cursor direction keys, or motion sensor, for controlling the position of a small cursor image presented on display 814 and issuing commands associated with graphical elements presented on display 814; and one or more camera sensors 884 for capturing, recording, and storing one or more still and / or moving images (e.g., video, movie, etc.), which may also include audio recording.
[0106] In the illustrated embodiment, dedicated hardware, such as an application-specific integrated circuit (ASIC) 820, is coupled to bus 810. The dedicated hardware is used to quickly perform special-purpose operations not performed by processor 802.
[0107] Computer system 800 also includes a communication interface 870 coupled to bus 810. Communication interface 870 provides one-way or two-way communication coupling to various external devices operating with their respective processors. Typically, coupling is via a network link 878, which is connected to a local network 880 to which various external devices, such as servers or databases, can be connected. Alternatively, link 878 can be directly connected to an Internet service provider (ISP) 884 or a network 890, such as the Internet. Network link 878 can be wired or wireless. For example, communication interface 870 can be a parallel port, a serial port, or a universal serial bus (USB) port on a personal computer. In some embodiments, communication interface 870 is an integrated services digital network (ISDN) card, a digital subscriber line (DSL) card, or a telephone modem, providing an information communication connection to a corresponding type of telephone line. In some embodiments, the communication interface 870 is a cable modem that converts signals on the bus 810 into signals for communication connection via a coaxial cable or into optical signals for communication connection via a fiber optic cable. As another example, the communication interface 870 can be a local area network (LAN) card to provide a data communication connection to a compatible LAN such as Ethernet. A wireless link can also be implemented. For a wireless link, the communication interface 870 sends and / or receives electrical, acoustic or electromagnetic signals, including infrared and optical signals, which carry information streams such as digital data. For example, in a wireless handheld device such as a mobile phone, the communication interface 870 includes a radio frequency band electromagnetic transmitter and receiver called a wireless transceiver. In some embodiments, the communication interface 870 is capable of connecting to a communication network to enhance the safety of distracted drivers using a head posture and sight line estimator of a mobile device such as a mobile phone or tablet computer.
[0108] Network link 878 typically provides information using a transmission medium through one or more networks to other devices that use or process the information. For example, network link 878 may provide a connection through local network 880 to a host computer 882 or to equipment 884 operated by an ISP. ISP equipment 884, in turn, provides data communication services through the public, global packet-switched communications network now commonly referred to as the Internet 890.
[0109] A computer connected to the Internet, referred to as a server host 882, hosts a process that provides services in response to information received via the Internet. For example, the server host 882 hosts a process that provides information representing video data for presentation at the display 814. It is contemplated that the components of the system 800 can be deployed in various configurations in other computer systems (e.g., the host 882 and the server 882).
[0110] At least some embodiments of the present invention involve implementing some or all of the techniques described herein using computer system 800. According to one embodiment of the present invention, these techniques are performed by computer system 800 in response to processor 802 executing one or more sequences of one or more processor instructions contained in memory 804. These instructions, also known as computer instructions, software, and program code, may be read into memory 804 from another computer-readable medium, such as storage device 808 or network link 878. Execution of the sequences of instructions contained in memory 804 causes processor 802 to perform one or more of the method steps described herein.
[0111] It should be understood that the present invention can be embodied in many different forms and should not be construed as being limited to the embodiments described herein. On the contrary, providing these embodiments will make this subject matter thorough and complete, and will fully convey the present invention to those skilled in the art. In fact, this subject matter is intended to cover alternatives, modifications, and equivalents of these embodiments included within the spirit and scope of this subject matter disclosure as defined by the appended claims. In addition, in the following detailed description of this subject matter, many specific details are described in order to provide a thorough understanding of this subject matter. However, it will be clear to those skilled in the art that the subject matter of this claim can be practiced without such specific details. Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams and the combination of blocks in the flowchart illustrations and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable instruction execution device create a mechanism for implementing the functions / actions specified in the flowchart and / or block diagram.
[0112] Computer-readable non-transitory media include various types of computer-readable media, including magnetic storage media, optical storage media, and solid-state storage media, and specifically do not include signals. It should be understood that the software can be installed in the device and sold with the device. Alternatively, the software can be obtained and loaded into the device, including through optical disk media or any other means of obtaining the software from a network or distribution system, including, for example, obtaining the software from a server owned by the software creator or from a server not owned by the software creator but used by the software creator. For example, the software can be stored on a server for distribution via the Internet.
[0113] Computer-readable storage media, excluding propagating signals, are accessible by a computer and / or processor and include removable and / or non-removable volatile and non-volatile internal and / or external media. Various types of storage media can accommodate data storage in any suitable digital format for a computer. Those skilled in the art will appreciate that other types of computer-readable media, such as zip drives, solid-state drives, magnetic tapes, flash memory cards, flash drives, cartridges, etc., can be used to store computer-executable instructions for performing the novel methods (actions) of the disclosed architecture.
[0114] The terms used herein are for the purpose of describing particular aspects only and are not intended to limit the present invention. Unless the context clearly indicates otherwise, the singular forms "a," "an," and "the" as used herein include the plural forms thereof. It should be further understood that the term "comprising" as used in this specification is used to indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0115] The description of the present disclosure has been presented for purposes of illustration and description and is not intended to be exhaustive or limited to the disclosure in the form disclosed. Various modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The various aspects of the invention have been chosen and described in order to better explain the principles of the invention and its practical application, and to enable those skilled in the art to understand the various modifications of the invention as it may be adapted for the particular use contemplated.
[0116] For the purposes of this document, each process associated with the disclosed technology can be performed serially by one or more computing devices. Each step in a process can be performed by the same or different computing devices as used in other steps, and each step does not necessarily have to be performed by a single computing device.
[0117] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A computer-implemented method for estimating head pose and gaze angle, characterized in that: include: Training a first neural network using a plurality of two-dimensional (2D) facial images to separate head and eye movements of the 2D face, wherein training the first neural network comprises: mapping 2D faces in the plurality of 2D facial images to a face position image, and constructing a facial texture image of the 2D face based on the facial position image; storing an eye texture image including a sight angle extracted from the facial texture image of the 2D face in a database; replacing the eye region of the facial texture image with the eye texture image including the sight line angle stored in the database to generate a modified facial texture image; reconstructing the modified facial texture image to generate a modified 2D facial image including a modified head pose and gaze angle as training data, and storing the training data in the database; and The head pose and the gaze angle are estimated by training a second neural network using the training data, wherein the training the second neural network comprises: collecting the training data from the database, and concurrently applying one or more transformations to the modified 2D facial image of the training data and corresponding eye regions of the modified 2D facial image; deforming the modified facial texture image of the 2D face onto a 3D morphable model (3DMM) of the face to reconstruct a 3D facial model based on the modified facial texture image using a gaze direction; applying a rotation matrix to the reconstructed 3D facial model to change the head pose, and changing the gaze angle to be consistent with the head pose; Projecting the 3D facial model after applying the rotation matrix into a 2D image space to generate the modified 2D facial image; and The modified 2D facial image is stored in the database.
2. The computer-implemented method of claim 1, wherein: The mapping further comprises: mapping the 2D faces in the plurality of 2D facial images to a position map using a face alignment method, wherein the face position map aligns the 2D faces in the plurality of 2D facial images to 3D coordinates of a reconstructed 3D model for the 2D faces in the plurality of 2D facial images; and The constructing further includes: constructing the facial texture image for the 2D face in the plurality of 2D facial images based on the facial position image or the facial 3D deformable model, so as to represent the texture of the aligned 2D face.
3. The computer-implemented method of claim 1 , wherein: The storage further includes: extracting the facial texture image from the 2D face in the plurality of 2D facial images based on the facial position image; cropping the eye region from the facial texture image to create a cropped eye texture image from the aligned 2D faces in the plurality of 2D facial images based on the landmarks; and The cropped eye texture image is stored in the database.
4. The computer-implemented method of claim 3, wherein: The cropped eye texture image is labeled as a difference between the head pose and the gaze angle of the 2D face in the plurality of 2D facial images.
5. The computer-implemented method according to claim 3 or 4, wherein: The replacement also includes: selecting the eye region of the cropped eye texture image from the database based on the landmark; and Based on the aligned coordinates of the landmarks, the eye region in the facial texture image is replaced with the cropped eye texture image in the database to generate a modified facial texture map of the 2D face in the plurality of 2D facial images.
6. The computer-implemented method according to claim 3 or 4, wherein: The replacement also includes: applying image fusion to merge the cropped eye texture image selected from the database into the modified facial texture map of the 2D face in the plurality of 2D facial images; and A generative adversarial network (GAN) is trained or a method based on local gradient information is used to smooth the color and texture in the eye region of the modified facial texture image.
7. The computer-implemented method according to claim 3 or 4, characterized in that The gaze direction is calculated by adding the relative gaze direction stored in the cropped eye texture image selected from the database to the head pose.
8. The computer-implemented method according to any one of claims 1 to 4, wherein: The estimates also include: collecting 2D facial images of a vehicle driver with one or more head poses to generate a driver dataset; and The driver dataset is applied to fine-tune the second neural network to estimate the driver's head pose and gaze angle.
9. The computer-implemented method according to any one of claims 1 to 4, wherein: The 2D facial image of the driver is captured by a capture device and uploaded to a network for processing; and the processed 2D facial image of the driver is downloaded to a vehicle.
10. The computer-implemented method of claim 1, wherein: The first neural network is an encoder-decoder type neural network for mapping the 2D facial image to a corresponding position map.
11. The computer-implemented method according to claim 1 or 2, wherein: In the facial position image, the RGB grayscale value of each pixel represents the 3D coordinate of the corresponding facial point in the reconstructed 3D model.
12. A device for estimating head posture and gaze angle, characterized in that include: Non-transitory memory including instructions; as well as one or more processors in communication with the memory, wherein the one or more processors execute the instructions to: Training a first neural network using a plurality of two-dimensional (2D) facial images to separate head and eye movements of the 2D face, wherein training the first neural network comprises: mapping 2D faces in the plurality of 2D facial images to a face position image, and constructing a facial texture image of the 2D face based on the facial position image; storing an eye texture image including a sight angle extracted from the facial texture image of the 2D face in a database; replacing the eye region of the facial texture image with the eye texture image including the sight line angle stored in the database to generate a modified facial texture image; reconstructing the modified facial texture image to generate a modified 2D facial image including a modified head pose and gaze angle as training data, and storing the training data in the database; and The head pose and the gaze angle are estimated by training a second neural network using the training data, wherein the training the second neural network comprises: collecting the training data from the database, and concurrently applying one or more transformations to the modified 2D facial image of the training data and corresponding eye regions of the modified 2D facial image; The one or more processors further execute the instructions to: deforming the modified facial texture image of the 2D face onto a 3D morphable model (3DMM) of the face to reconstruct a 3D facial model based on the modified facial texture image using a gaze direction; applying a rotation matrix to the reconstructed 3D facial model to change the head pose, and changing the gaze angle to be consistent with the head pose; Projecting the 3D facial model after applying the rotation matrix into a 2D image space to generate the modified 2D facial image; and The modified 2D facial image is stored in the database.
13. The device according to claim 12, characterized in that The one or more processors further execute the instructions to: mapping the 2D faces in the plurality of 2D facial images to a position map using a face alignment method, wherein the face position map aligns the 2D faces in the plurality of 2D facial images to 3D coordinates of a reconstructed 3D model of the 2D faces in the plurality of 2D facial images; and Based on the facial position image or the facial 3D deformable model, the facial texture image is constructed for the 2D face in the plurality of 2D facial images to represent the texture of the aligned 2D face.
14. The device according to claim 12, characterized in that The one or more processors further execute the instructions to: extracting the facial texture image from the 2D face in the plurality of 2D facial images based on the facial position image; cropping the eye region from the facial texture image to create a cropped eye texture image from aligned 2D faces in the plurality of 2D facial images based on the landmarks; as well as The cropped eye texture image is stored in the database.
15. The device according to claim 14, characterized in that The cropped eye texture image is labeled as a difference between the head pose and the gaze angle of the 2D face in the plurality of 2D facial images.
16. The device according to claim 14 or 15, characterized in that The one or more processors further execute the instructions to: selecting the eye region of the cropped eye texture image from the database based on the landmark; and Based on the aligned coordinates of the landmarks, the eye region in the facial texture image is replaced with the cropped eye texture image in the database to generate a modified facial texture map of the 2D face in the plurality of 2D facial images.
17. The device according to claim 14 or 15, characterized in that The one or more processors further execute the instructions to: applying image fusion to merge the cropped eye texture image selected from the database into the modified facial texture map of the 2D face in the plurality of 2D facial images; and A generative adversarial network (GAN) is trained or a method based on local gradient information is used to smooth the color and texture in the eye region of the modified facial texture image.
18. The device according to claim 14 or 15, characterized in that The gaze direction is calculated by adding the relative gaze direction stored in the cropped eye texture image selected from the database to the head pose.
19. The apparatus according to any one of claims 12 to 15, characterized in that The one or more processors further execute the instructions to: collecting 2D facial images of vehicle drivers with one or more head poses to generate a driver dataset; as well as The driver dataset is applied to fine-tune the second neural network to estimate the driver's head pose and gaze angle.
20. The apparatus according to any one of claims 12 to 15, characterized in that The 2D facial image of the driver is captured by a capture device and uploaded to a network for processing; and the processed 2D facial image of the driver is downloaded to a vehicle.
21. The device according to claim 12, characterized in that The first neural network is an encoder-decoder type neural network for mapping the 2D facial image to a corresponding position map.
22. The device according to claim 12 or 13, characterized in that In the facial position image, the RGB grayscale value of each pixel represents the 3D coordinate of the corresponding facial point in the reconstructed 3D model.
23. A non-transitory computer readable medium storing computer instructions for estimating head pose and gaze angle, characterized in that: When one or more processors execute the computer instructions, the one or more processors are caused to perform the following steps: Training a first neural network using a plurality of two-dimensional (2D) facial images to separate head and eye movements of the 2D face, wherein training the first neural network comprises: mapping 2D faces in the plurality of 2D facial images to a face position image, and constructing a facial texture image of the 2D face based on the facial position image; storing an eye texture image including a sight angle extracted from the facial texture image of the 2D face in a database; replacing the eye region of the facial texture image with the eye texture image including the sight line angle stored in the database to generate a modified facial texture image; reconstructing the modified facial texture image to generate a modified 2D facial image including a modified head pose and gaze angle as training data, and storing the training data in the database; and The head pose and the gaze angle are estimated by training a second neural network using the training data, wherein the training the second neural network comprises: collecting the training data from the database, and concurrently applying one or more transformations to the modified 2D facial image of the training data and corresponding eye regions of the modified 2D facial image; deforming the modified facial texture image of the 2D face onto a 3D morphable model (3DMM) of the face to reconstruct a 3D facial model based on the modified facial texture image using a gaze direction; applying a rotation matrix to the reconstructed 3D facial model to change the head pose, and changing the gaze angle to be consistent with the head pose; Projecting the 3D facial model after applying the rotation matrix into a 2D image space to generate the modified 2D facial image; and The modified 2D facial image is stored in the database.
24. The non-transitory computer readable medium of claim 23, wherein: The mapping further comprises: mapping the 2D faces in the plurality of 2D facial images to a position map using a face alignment method, wherein the face position map aligns the 2D faces in the plurality of 2D facial images to 3D coordinates of a reconstructed 3D model for the 2D faces in the plurality of 2D facial images; and The constructing further includes: constructing the facial texture image for the 2D face in the plurality of 2D facial images based on the facial position image or the facial 3D deformable model, so as to represent the texture of the aligned 2D face.
25. The non-transitory computer readable medium of claim 23, wherein: The storage includes: extracting the facial texture image from the 2D face in the plurality of 2D facial images based on the facial position image; cropping the eye region from the facial texture image to create a cropped eye texture image from the aligned 2D faces in the plurality of 2D facial images based on the landmarks; and The cropped eye texture image is stored in the database.
26. The non-transitory computer readable medium of claim 25, wherein: The cropped eye texture image is labeled as a difference between the head pose and the gaze angle of the 2D face in the plurality of 2D facial images.
27. The non-transitory computer readable medium according to claim 25 or 26, wherein: The replacement also includes: selecting the eye region of the cropped eye texture image from the database based on the landmark; and Based on the aligned coordinates of the landmarks, the eye region in the facial texture image is replaced with the cropped eye texture image in the database to generate a modified facial texture map of the 2D face in the plurality of 2D facial images.
28. The non-transitory computer readable medium according to claim 25 or 26, wherein: The replacement also includes: applying image fusion to merge the cropped eye texture image selected from the database into the modified facial texture map of the 2D face in the plurality of 2D facial images; and A generative adversarial network (GAN) is trained or a method based on local gradient information is used to smooth the color and texture in the eye region of the modified facial texture image.
29. The non-transitory computer readable medium according to claim 25 or 26, wherein: The gaze direction is calculated by adding the relative gaze direction stored in the cropped eye texture image selected from the database to the head pose.
30. The non-transitory computer readable medium according to any one of claims 23 to 26, wherein: The estimates also include: collecting 2D facial images of a vehicle driver with one or more head poses to generate a driver dataset; and The driver dataset is applied to fine-tune the second neural network to estimate the driver's head pose and gaze angle.
31. The non-transitory computer readable medium according to any one of claims 23 to 26, wherein: The 2D facial image of the driver is captured by a capture device and uploaded to a network for processing; and the processed 2D facial image of the driver is downloaded to a vehicle.
32. The non-transitory computer readable medium of claim 23, wherein: The first neural network is an encoder-decoder type neural network for mapping the 2D facial image to a corresponding position map.
33. The non-transitory computer readable medium according to claim 23 or 24, wherein: In the facial position image, the RGB grayscale value of each pixel represents the 3D coordinate of the corresponding facial point in the reconstructed 3D model.
Citation Information
Patent Citations
Data enhancement method based on artificial face
CN108805094A
Unconstrained appearance-based gaze estimation
US20180181809A1