System and method for monitoring a human face by combining synthetic and real facial features

The system efficiently maps synthetic facial features to real facial features in driver monitoring systems by merging them using a processing unit with a learning engine, addressing the challenges of data annotation and reducing costs and time in data acquisition.

DE102024132874A1Pending Publication Date: 2025-06-05MERCEDES BENZ GROUP AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024132874
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-11-11
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing driver monitoring systems face challenges in efficiently mapping synthetic facial features to real facial features, particularly in complex regions like the eyes and jaw line, due to the time-consuming and costly process of collecting and labeling real data.

Method used

A system and method that merges synthetic facial features with real facial features by using a processing unit with a learning engine to receive image frames, extract real face landmarks, generate a face segmentation map, project synthetic face landmarks onto the map, and determine key points, thereby efficiently bridging the annotation gap between synthetic and real data.

Benefits of technology

This approach enables efficient and accurate monitoring of a driver's face, reducing the need for real data acquisition, saving time and cost, and improving the training of machine learning models by closing the annotation gap.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides a system and method for monitoring a human face. At 302, a face segmentation map is formed, at 304, face counters are drawn onto the face segmentation map. At 306, internal edges are removed from the face counters, and at 308, landmarks are projected onto the face counter. At 310, a principal line connecting the forehead and chin point is derived. At 312, a connecting line between the corners of the eyes is found, which is recursively split to obtain three evenly spaced points on the line. At 314, lines are then drawn through the evenly spaced points, and at 316, the points connecting the facial contours and the lines yield new keypoints. At 318, the center points for each pair of eye landmarks are found. At 320, lines are drawn through the center points parallel to the principal line to obtain new keypoints at 322.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELDThe present disclosure relates generally to the field of driver monitoring systems (DMS) in vehicles. More particularly, the present disclosure relates to a system and method for monitoring a vehicle user's face by merging synthetic facial features with real-world facial features.BACKGROUNDThe driver monitoring system (DMS) is a safety-critical component in modern cars. The DMS monitors the driver's attention and, if necessary, issues warnings, which is important to avoid accidents.Facial landmarks are used to recognize and track the eyes and mouth (face) of drivers and passengers. The recognition of facial features is of decisive importance for safety-critical application cases of artificial intelligence (AI) in the vehicle, since it can aid in the recognition of the face and the eyes of the driver and is also decisive in cases such as eye blinking.Facial landmarks are created by collecting and labeling real data. However, collecting and labeling the real data is very time consuming and expensive. To overcome these problems, synthetic data is created using computer graphics that are easily created and ideal. Game engines such as Unity and Ordinary can quickly generate a large number of photorealistic images. Synthetic data can provide not only accurate ground truth, but also various modalities, which is helpful in model training. However, a gap is created between the real data and the synthetic data, making it difficult to associate the synthetic data with the real data.Various works have been carried out to solve the above problems. For example, the non-patented literature "Face Alignment Refinement" describes an observation-based active matching algorithm to correct the inaccurate landmarks of a given contour from a face shape returned from the face alignment. It also includes a data controlled validation frame for refinement.The non-patent literature "Üb üb it till you make it: face analysis in the wild using synthetic data alon" describes systems and methods for carrying out face-related computer vision in free wild web using synthetic data only. Attempts are being made to bridge the domain gap between real and synthetic data with data mixing, domain adaptation and domain-adaptive training and to synthesize data with a minimum domain gap, so that models trained by machine learning on synthetic data can be generalized to real datasets in free wild-type.Although in the mentioned documents and conventional solutions synthetic data is used to bridge the gap between the real data and the synthetic data, there is still the possibility of providing an improved system and method that determines synthetic data and more efficiently maps it to real data associated with in-car scenarios.OBJECTS OF PRESENT DECISIONA general object of the present disclosure is to provide a system and a method that avoids the above-mentioned problems and enables efficient monitoring of a driver's face in a vehicle.Another object of the present disclosure is to provide a system and method for generating synthetic facial features that resemble real facial features.Another object of the present disclosure is to provide a system and method that merges synthetic facial features with real facial features, thereby enabling time-saving operation.Another object of the present disclosure is to reduce the need for real data acquisition and thus save time and cost.Another object of the present disclosure is to efficiently bridge annotations between the synthetic facial features and the real facial features, even for complex facial regions such as the eyes and the jaw line.Another object of the present disclosure is to efficiently train the involved learning machine by closing the gap between the synthetic facial features and the real facial features.Another object of the present disclosure is to monitor the face and detect changes in gaze, eye blinking, facial expression, and facial gestures.Another object of the present disclosure is to provide an efficient, accurate, time-saving, and low cost system and method for monitoring the face of the driver.SUMMARYAspects of the present disclosure generally relate to the field of driver monitoring systems (DMS) in vehicles. More particularly, the present disclosure relates to a system and method for monitoring a vehicle user's face by merging synthetic facial features with real-world facial features.In one aspect, the present disclosure relates to a system for mapping synthetic face landmarks to real face landmarks of a human subject in a vehicle. The system includes a processing unit including a learning engine, a processor, and a memory operatively coupled to the processor, the memory storing instructions that when executed by the processor. The processing unit is configured to: receive one or more image frames including a human person's face from an image acquisition unit coupled to the vehicle by communicating with the image acquisition unit; extract real face landmarks from the one or more received image frames; generate a face segmentation map associated with the human person's face in consideration of the real face landmarks; draw one or more face contours on the generated face segmentation map; project synthetic face landmarks onto the one or more face contours; determine a line connecting eye corners of both eyes of the human entity, and recursively split the determined line to obtain a plurality of equi-spaced points on the line; drawing one or more lines through each of the obtained points, wherein at least one of the obtained points connecting one of the one or more facial contours and one of the one or more lines provides new key points.In one embodiment, the processing unit may be configured to acquire a set of three-dimensional (3D) synthetic landmarks and convert the acquired 3D set of synthetic landmarks into two-dimensional (2D) annotated landmarks, wherein the converted 2D annotated landmarks are mapped to the facial contours.In an embodiment, the processing unit may generate, by the learning engine, the synthetic face landmarks associated with the human subject by considering the one or more received image frames and the associated human subject's face.In one embodiment, the synthetic landmarks may be fixed points whose positions depend on factors such as head orientation, gaze angle, and face construction.In one embodiment, the one or more drawn surface contours may include inner edges and outer edges, wherein the processing unit may be configured to eliminate the inner edges of the drawn surface contours.In another aspect, the present disclosure relates to a method for mapping synthetic facial landmarks to real facial landmarks of a human subject in a vehicle. The method comprises: (i) receiving, at a processing unit, one or more image frames comprising a human person's face by communicating with an image acquisition unit coupled to the vehicle; (ii) extracting real face landmarks from the one or more received image frames at the processing unit; (iii) generating, at the processing unit, a face segmentation map associated with the human person's face in consideration of the real face landmarks; (iv) drawing, at the processing unit, one or more face contours on the generated face segmentation map; (v) projecting synthetic face landmarks onto the one or more face contours in the processing unit; (vi) determining a line connecting the eye corners of both eyes of the human entity and recursively dividing the determined line to obtain a plurality of equi-spaced points on the line; and (vii) drawing one or more lines through each of the obtained points, wherein at least one of the obtained points connecting one of the one or more facial contours and one of the one or more lines provides new key points.In one embodiment, the method may include obtaining a set of three-dimensional (3D) synthetic landmarks and converting the obtained 3D set of synthetic landmarks to two-dimensional (2D) annotated landmarks, wherein the converted 2D annotated landmarks are mapped to the facial contours.In one embodiment, the method may include generating, by the learning engine, the synthetic face landmarks associated with the human subject in consideration of the one or more received image frames and the associated human subject's face.In one embodiment, the synthetic aspects may be fixed points whose positions depend on factors such as head orientation, gaze angle, and face construction.In another embodiment, the one or more drawn surface contours may include inner edges and outer edges, and the method may further include eliminating the inner edges of the drawn surface contours.BRIEF DESCRIPTION OF THE DRAWINGSThe accompanying drawings serve to further understand the present disclosure and form part of this description. The drawings illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. FIG. 1 illustrates an example network architecture of the proposed system to explain its operation according to an embodiment of the present disclosure. FIG. 2 shows an example block diagram of a processing unit of the proposed system according to an embodiment of the present disclosure. FIG. 3 is a flow chart illustrating the step-by-step operation of the proposed system according to an embodiment of the present disclosure. FIGS. 4A and 4B show example representations of synthetic images according to an embodiment of the present disclosure. FIG. 5 shows an exemplary illustration of the formation of eye contour and principal line according to an embodiment of the present disclosure. FIG. 6 is a flow chart illustrating the proposed method according to an embodiment of the present disclosure. FIG. 7 shows an example of a computer system in or with which embodiments of the present disclosure may be implemented.DETAILED DESCRIPTIONThe following is a detailed description of the embodiments of the disclosure illustrated in the accompanying drawings. The embodiments are so detailed as to clearly convey the disclosure. However, it is not intended to limit the predictable variations of embodiments with the particularity provided; on the contrary, it is intended to cover all modifications, equivalents and alternatives falling within the spirit and scope of the present disclosure as defined by the appended claims.The embodiments discussed herein relate generally to the field of driver monitoring systems (DMS) in vehicles. More particularly, the present disclosure relates to a system and method for monitoring a user's face, particularly a driver, of a vehicle by combining synthetic facial features with real facial features.Referring to FIG. 1, the proposed system 100 (interchangeably referred to herein as system 100) facilitates the efficient monitoring of a human face, i.e., a human person / user's face, in a vehicle 110 by merging real-world face landmarks (also referred to herein as real landmarks) of the user with corresponding synthetic face landmarks (also referred to herein as synthetic landmarks).In one embodiment, the system 100 includes an image capture unit 102 coupled to the vehicle 110 that can capture one or more images of a human person's face. In an embodiment, the image capturing unit 102 may include an image sensor and a camera that may be placed at a predefined position in the vehicle 110, such that the image capturing unit 102 may easily capture images of the human person's face, in particular the face of a driver driving the vehicle 110 or an authorized user connected to the vehicle 110, such as the owner of the vehicle 110.According to one embodiment, the system 100 includes a processing unit 106 that includes a learning engine 104. The processing unit 106 may receive the one or more frames from the image acquisition unit 102 by communicating with the image acquisition unit 102. In one embodiment, the processing unit 106 may extract real face landmarks from the one or more received image frames and generate a face segmentation map associated with the human being's face by considering the real face landmarks.In one embodiment, the processing unit 106 may draw one or more facial contours on the generated face segmentation map. In an example embodiment, the one or more drawn face contours may include inner edges and outer edges, and the processing unit 106 may eliminate the inner edges of the drawn face contours to improve accuracy.In another embodiment, the processing unit 106 may project synthetic face landmarks onto the one or more face contours. In one implementation, the processing unit 106 may generate, by the learning engine 104, the synthetic face landmarks associated with the human entity by considering the one or more received image frames and the associated human entity face. In an example embodiment, the processing unit 106 may acquire a set of three-dimensional (3D) synthetic landmarks and convert the acquired 3D set of synthetic landmarks into two-dimensional (2D) annotated landmarks, wherein the converted 2D annotated landmarks may be mapped to the facial contours. Moreover, the synthetic landmarks may be fixed points, so that the positions of the synthetic landmarks depend on factors such as head orientation, viewing angle, and face construction.In one embodiment, the processing unit 106 may determine a line connecting the eye angles of both eyes of the human being and recursively split the determined line to obtain a plurality of evenly spaced points on the line.In another embodiment, the processing unit 106 may draw one or more lines through each of the obtained points, wherein at least one of the obtained points connecting one of the one or more facial contours and one of the one or more lines may provide new key points, such as eye key points, forehead key points, jaw line key points, and other similar key points associated with the human being's face.The term "real facial features" refers to annotations of the human face in two-dimensional (2D) space, and the term "synthetic facial features" refers to computer-calculated three-dimensional (3D) features of the human face projected on 2D.In one embodiment, eliminating the annotation gap may make the synthetic data more useful during training. The system 100 trained on annotated synthetic data performs better than unbridged annotation data. This may also reduce the need to collect real data and save cost.In one embodiment, the processing unit 106 may be in communication with the image acquisition unit 102 and the learning engine 104 via a network 108. Further, the network 108 may be a wireless network, a wired network, or a combination thereof, which may be implemented as any of various types of networks, such as intranet, local area network (LAN), wide area network (WAN), Internet, and the like. Moreover, the network 108 may be either a dedicated network or a shared network. The shared network may represent a connection of various types of networks that may use a variety of protocols, such as hypertext transfer protocol (HTTP), transmission control protocol / internet protocol (TCP / IP), wireless application protocol (WAP), and the like.In one embodiment, the system 100 may be implemented using any one or a combination of hardware components and software components such as a cloud, a server 112, a computer system, a computing device, a network device, and the like. Further, the processing unit 106 may interact with the image capture unit 102 via a website or application that may be located in the proposed system 100. In one implementation, system 100 may be accessed via a website or application that may be configured with any operating system, including, but not limited to, Android™ iOS™ and the like.A block diagram 200 in FIG. 2 illustrates example functional units of the processing unit 106, which may include one or more processors 202, which may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logic circuits, and / or any devices that process data based on operating instructions. Among other capabilities, the one or more processor(s) 202 is / are configured to be able to fetch and execute computer readable instructions stored in a memory 204 of the processing unit 106. The memory 204 may store one or more computer readable instructions or routines that may be fetched and executed to create or share the data units via a network service. The memory 204 may include any nonvolatile memory device, for example, a volatile memory such as RAM or a nonvolatile memory such as EPROM, flash memory, and the like.In one embodiment, the processing unit 106 may also include one or more interfaces 206. The interface(s) 206 may include a variety of interfaces, e.g., interfaces for data input and output devices referred to as I / O devices, storage devices, and the like. The interface(s) 206 may facilitate communication of the processing unit 106 with various devices connected to the processing unit 106. The interface(s) 206 may also provide a communication path for one or more components of the processing unit 106. Examples of such components include processing engine(s) 208 and database 210.In one embodiment, the processing machine(s) 208 may be implemented as a combination of hardware and programming (e.g., programmable instructions) to implement one or more functions of the processing machine(s) 208. In the examples described herein, such combinations of hardware and programming may be implemented in various ways. For example, the programming for the processing machine(s) 208 may consist of processor-executable instructions stored on a non-transitory machine-readable storage medium, and the hardware for the processing machine(s) 208 may include a processing resource (e.g., one or more processors) for executing such instructions.In the present examples, the machine readable storage medium may store instructions that, when executed by the processing resource, implement the processing engine(s) 208. In such examples, the processing unit 106 may include the machine readable storage medium storing the instructions and the processing resource for executing the instructions, or the machine readable storage medium may be separate but accessible to the system 100 and the processing resource. In other examples, the processing engine(s) 208 may be implemented by electronic circuitry. Database 210 may include data that is either stored or generated as a result of functionalities implemented by one of the components of processing machine(s) 208. In one embodiment, the processing machine(s) 208 may include an extraction unit 212, a merging unit 214, and other unit(s) 216. The other unit(s) 216 may implement functionality(s) that supplement the applications / functions executed by the processing unit 106.According to an embodiment, the extraction unit 212 may extract real face landmarks from one or more image frames received from the image acquisition unit 102.In one embodiment, the extraction unit 212 may acquire a set of three-dimensional (3D) synthetic landmarks from the learning engine 104 and then convert the acquired 3D set of synthetic landmarks into two-dimensional (2D) annotated landmarks, wherein the converted 2D annotated landmarks may be mapped to the facial contours. In an example embodiment, the extraction unit 212 may generate, by the learning engine 104, the synthetic face landmarks associated with the human entity by considering the one or more received image frames and the associated face of the human entity. The synthetic landmarks may be fixed points whose positions may depend on factors such as head orientation, gaze angle, and face construction.According to another embodiment, the merging unit 214 may merge the real face landmarks and the synthetic face landmarks by properly annotation the synthetic face landmarks.Generally, the synthetic landmarks may differ in two important facial regions - jaw line and eyes. The synthetic landmarks are generated by the system and are therefore consistent so that each rule-based model once implemented can be scaled to a large set of images. Consequently, each matched landmark may have the same consistency throughout all data sets.In one embodiment, for a given set of synthetic 3D landmarks (L S) and the corresponding face segmentation map (S S) the system 100 may convert the 3D landmarks to annotated 2D landmarks (L H). In an exemplary embodiment, the L H annotations are made on visible face contours and the S S may provide accurate contours of the face.In addition, a binary mask may be derived from the face segmentation map. The system 100 may derive a contour (C) from the binary mask. In one embodiment, the system 100 could already assume that the features of a face (also referred to herein as facial features) are symmetrical in nature, and accordingly the system 100 may derive the contour (C).In one embodiment, the system 100 may use any mechanism known in the art for edge detection to derive the contour C from the binary mask. Take as an example that a point A on an image is a hidden point in 3D, but may appear on the face contour when labeled in 2D. In addition, in the face geometry there is a corresponding point B on the other side of the face which should not be obscured. In addition, a line L could be drawn to pass through the point A and the point B. The new key points may be where the line L meets the contour C.Referring to FIGS. 3 and 4A and 4B, in the flowchart 300, a face segmentation map 402 may be created in block 302. In one embodiment, segmentation data may be easily generated along with images and their corresponding annotations, and then a binary mask of the face (face mask 404) may be created using the segmentation data.In block 304, face counters 406 may be drawn on the face segmentation map. Further, in block 306, inner edges may be removed from the face counters 406, and then in block 308, landmarks may be projected onto the face counters 406. In a preferred embodiment, as shown in FIG. 4B, two sets of contours can be drawn on the face segmentation map - an inner contour 406-1 and an outer contour 406-2. Further, the outer contour 406- 2 may be eliminated by ANDing with the binary face mask 404.In one embodiment, the system 100 may derive a principal line connecting the forehead and the chin point at block 310, as shown in FIG. 5. In block 312, the system 100 may find a line connecting the eye angles and recursively share the line to obtain three equally spaced points on the line. Further, at block 314, one or more lines may be drawn through the evenly spaced points, such that at block 316, the points connecting the facial contours and the lines may form new key points.In another embodiment, after deriving the principal line connecting the forehead and chin points, the midpoints for each pair of eye marks may be found in block 318. In an exemplary embodiment, let the forehead point (x f, y f) and the chin point (x c,, y c). Then, let the slope (m) of a line between the end point (x f, y f) and the chin point be (x c,, y c) be calculated as -In addition, the center points of a line segment connecting two vertices may be recursively calculated using the formula -. Therefore, the corresponding line equation can be derived with y - y i= m (x - x i) intersecting the eye contour to form a line mask M l for each division point on the line segment. Moreover, new points can be obtained by M l. M c, where M c is the mask of the eye contour.In one embodiment, at block 320, one or more lines may be drawn through midpoints parallel to the principal line, and then at block 322, the points connecting the facial contours and the lines may form new key points.In an example embodiment, all key points may be derived using a facial feature database that can estimate 3D facial features in real-time, even on mobile computing devices. The facial landmark network database may employ machine learning (ML) to determine the 3D face surface.In one embodiment, the synthetic annotations are fixed points, the positions of which depend on head orientation, viewing angle, face construction, etc. In another embodiment, the real annotations may be based entirely on the viewing angle. Therefore, in many cases, there may be a large gap between annotations of the synthetic data and the real data.FIG. 6 shows the proposed method 600 (also referred to here as method 600) for assigning synthetic facial features to real facial features of a human person in a vehicle.In one embodiment, the method 600 includes, at block 602, receiving, at a processing unit, one or more images comprising a face of the human being by communication with an image capture unit connected to the vehicle.At block 604, the method 600 includes extracting real face landmarks from one or more received image frames at the processing unit.At block 606, the method 600 includes generating, at the processing unit, a face segmentation map associated with the human entity's face by considering the real face landmarks.At block 608, the method 600 includes the processing unit drawing one or more facial contours on the generated face segmentation map.At block 610, the method 600 includes projecting synthetic face landmarks onto the one or more face contours in the processing unit.At block 612, the method 600 includes determining a line connecting the eye angles of both eyes of the human being and recursively splitting the determined line to obtain a plurality of equally spaced points on the line.At block 614, the method 600 includes drawing one or more lines through each of the obtained points, where at least one of the obtained points connects one of the one or more facial contours and one of the one or more lines, thereby creating new key points.In another embodiment, the method 600 may include obtaining a set of three-dimensional (3D) synthetic landmarks, and further converting the obtained 3D set of synthetic landmarks to two-dimensional (2D) annotated landmarks, wherein the converted 2D annotated landmarks may be mapped to the facial contours.In another embodiment, the method 600 may include generating, by the learning engine, the synthetic face landmarks associated with the human subject by considering the one or more received image frames and the associated faces of the human subject. In an exemplary embodiment, the synthetic landmarks may be fixed points whose positions depend on factors such as head orientation, gaze angle, and face construction.In another embodiment, the one or more drawn surface contours may include inner edges and outer edges, wherein the method 600 may include eliminating the inner edges of the drawn surface contours.In one embodiment, there is a large gap between the annotations of the synthetic facial features and the real facial features relating to the user's eyes. The proposed system 100 and method 600 may minimize this gap and therefore accurately identify the eye key points.In one embodiment, the proposed system 100 and method 600 may track the human person's face in an efficient manner. Moreover, the proposed system 100 and method 600 may also be used to detect changes in gaze, blink, facial expression, or facial gesture, and accordingly may control operation of the vehicle 110, such as playing music, enabling / disabling various components of the vehicle 110, such as windows and climate control, and maneuvering the vehicle 110.As shown in FIG. 7, computer system 700 may include an external storage device 710, a bus 720, a main memory 730, a read only memory 740, a mass storage device 750, one or more communication ports 760, and a processor 770. One skilled in the art will understand that computer system 700 may include more than one processor and communication ports. Processor 770 may include various modules associated with embodiments of the present disclosure. The communication port / ports 760 may be an RS-232 port for use with a modem-based dial-up connection, a 10 / 100 Ethernet port, a gigabit or 10 gigabit port over copper or fiber, a serial port, a parallel port, or other existing or future ports. The communication ports 760 may be selected depending on the network, such as a local area network (LAN), a wide area network (WAN), or any network to which the computer system 700 is connected.In one embodiment, main memory 730 may be random access memory (RAM) or any other dynamic storage device well known in the art. The read-only memory 740 may be any static storage device, e.g., but not limited to a programmable read only memory (PROM) chip for storing static information, e.g., start or basic input / output system (BIOS) commands for the processor 770. Mass storage device 750 may be any current or future mass storage solution that may be used to store information and / or instructions. Example mass storage solutions include, but are not limited to, parallel advanced technology attachment (PATA) or serial advanced technology attachment (SATA) hard drives or solid state drives (internal or external, e.g., with universal serial bus (USB) and / or Firewire interfaces).In one embodiment, bus 720 may communicatively connect processor(s) 770 to the other memory, memory, and communication blocks. Bus 720 may be, for example, a peripheral component interconnect PCI) / PCI extended (PCI-X) bus, small computer system interface (SCSI), USB, or the like, to connect expansion cards, drives, and other subsystems, as well as other buses, such as a front side bus (FSB), that connects processor 770 to computer system 700.In another embodiment, operator and management interfaces, e.g., a screen, keyboard, and cursor control device, may also be connected to bus 720 to support direct operator interaction with computer system 700. Other operator and management interfaces may be provided via network connections connected via communication interface(s) 760. The above-described components are intended to be illustrative of various possibilities. The example computer system 700 described above is not intended to limit the scope of the present disclosure in any way.While the foregoing describes various embodiments of the disclosure, other and further embodiments of the disclosure may be developed without departing from the basic scope of the disclosure. The scope of the disclosure is defined by the following claims. The disclosure is not limited to the described embodiments, versions, or examples that are included to enable a person of ordinary skill in the art to make and use the disclosure when combined with information and knowledge available to the person of ordinary skill in the art.ADVANTAGES OF THE PRESENT DISCLOSUREThe present disclosure provides a system and method for facilitating efficient monitoring of a driver's face in a vehicle.The present disclosure provides a system and method for generating synthetic facial features that resemble real facial features.The present disclosure provides a system and method that merges synthetic facial features with real facial features, thereby allowing for a time-saving operation.The present disclosure provides a system and method that reduces the need to acquire real data and thus saves time and cost.The present disclosure provides a system and method that efficiently bridges annotations between the synthetic facial features and the real facial features, even for complex areas of the face, such as the eyes and the jaw line.The present disclosure provides a system and method for training the learning machine involved efficiently by closing the annotation gap between the synthetic facial features and the real facial features.The present disclosure provides a system and method for monitoring the face and detecting changes in eye impact, eye blinking, facial expression, and facial gestures.The present disclosure provides an efficient, accurate, time-saving, and low cost system and method for monitoring the face of the driver.

Claims

A system (100) for imaging synthetic face landmarks with real face landmarks of a human entity in a vehicle (110), the system (100) comprising: a processing unit (106) comprising a learning engine (104) and a processor and memory operatively coupled to the processor, the memory storing instructions that, when executed by the processor, are configured to: receive one or more image frames comprising a human person's face from an image acquisition unit (102) coupled to the vehicle (110) by communicating with the image acquisition unit (102); extract real face features from one or more received image frames; generate a face segmentation map associated with the human person's face; draw one or more face contours on the generated face segmentation map taking into account the real face features; A method of projecting synthetic facial features onto the one or more facial contours; determining a line connecting the eye angles of both eyes of the human being and recursively dividing the determined line to obtain a plurality of equally spaced points on the line; and drawing one or more lines through each of the obtained points, wherein at least one of the obtained points connecting one of the one or more facial contours and one of the one or more lines represent new key points.The system (100) of claim 1, wherein the processing unit (106) is configured to: acquire a set of three-dimensional (3D) synthetic landmarks; and convert the acquired 3D set of synthetic landmarks to two-dimensional (2D) annotated landmarks, wherein the converted 2D annotated landmarks are mapped to the facial contours.The system (100) of claim 1, wherein the processing unit (106), by the learning engine (104), generates the synthetic face landmarks associated with the human subject in consideration of the one or more received image frames and the associated human subject's face.The system (100) of claim 3, wherein the synthetic landmarks are fixed points whose positions depend on factors such as head orientation, gaze angle, and face construction.The system (100) of claim 1, wherein the one or more drawn surface contours include inner edges and outer edges, wherein the processing unit (106) is configured to eliminate the inner edges of the drawn surface contours.A method (600) of mapping synthetic face landmarks with real face landmarks of a human entity in a vehicle, the method (600) comprising: receiving (602), at a processing unit, one or more image frames comprising a human person's face by communication with an image capture unit coupled to the vehicle; extracting (604), at the processing unit, real face landmarks from the one or more received image frames; generating (606), at the processing unit, a face segmentation map associated with the human entity's face by accounting for the real face landmarks; drawing (608), at the processing unit, one or more face contours on the generated face segmentation map; projecting (610), at the processing unit, synthetic face landmarks onto the one or more face contours; determining (612) a line connecting the eye angles of both eyes of the human being and recursively splitting the determined line to obtain a plurality of equally sized points on the line; and drawing (614) one or more lines through each of the obtained points, wherein at least one of the obtained points connecting one of the one or more facial contours and one of the one or more lines provides new key points.The method (600) of claim 7, wherein the method (600) comprises: obtaining a set of three-dimensional (3D) synthetic landmarks; and converting the obtained 3D set of synthetic landmarks into two-dimensional (2D) annotated landmarks, wherein the converted annotated 2D landmarks are mapped to the facial contours.The method (600) of claim 7, wherein the method (600) comprises generating, by the learning engine, the synthetic face landmarks associated with the human being in consideration of the one or more received image frames and the associated human being's face.The method (600) of claim 8, wherein the synthetic face landmarks are fixed points whose positions depend on factors such as head orientation, gaze angle, and face construction.The method (600) of claim 7, wherein the one or more drawn surface contours comprise inner edges and outer edges, the method (600) further comprising eliminating the inner edges of the drawn surface contours.