Image processing method and device
By using prior information in the image processing method to track the facial area and set the scanning area and reset the window when the tracking fails, the problem of eye tracking failure in the illumination change environment is solved, and the detection speed and accuracy are improved.
Patent Information
- Application Number
- CN202010650136.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-26
- Filing Date
- 2020-07-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-07-08
AI Technical Summary
In an environment with sharp illumination, the stability of camera-based eye tracking technology is reduced, resulting in the user's eye tracking failure and the ability to quickly re-detect the position and coordinates of the eyes.
By acquiring an image frame, the facial area is tracked based on the prior information of the previous frame. If the tracking fails, the scanning area is set based on the second prior information, and the facial area is detected in the area. If the detection fails, the scan area is reset by sequentially enlarging the size of the window.
Improves the accuracy and speed of tracking the user's face or eyes in an illumination changing environment, ensuring the rapid and stable coordinate detection of the eyes.
Smart Images

Figure CN112434549B_ABST
Abstract
Description
[0001] This application claims the priority benefit of Korean Patent Application No. 10-2019-0104570 filed on August 26, 2019, in the Korean Intellectual Property Office, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0002] Methods and apparatuses consistent with example embodiments relate to an image processing method and an image processing apparatus. Background Art
[0003] Camera-based eye tracking technology can be applied to multiple fields (such as, for example, automatic stereoscopic three-dimensional (3D) ultra-multi-view displays and / or head-up displays (HUDs) based on viewpoint tracking). The performance of camera-based eye tracking technology depends on the image quality of the camera and / or the performance of the eye tracking method. In an environment with a sharp change in illumination (for example, in a driving environment), the stability of the operation of camera-based eye tracking technology is reduced due to backlight, strong sunlight, dark and low illumination environments, tunnel passage during driving, and / or driver movement. Considering the driving environment where an augmented reality (AR) 3D HUD can be used, there is a need for the following method: the method is responsive to the failure of tracking the user's eyes due to the influence of the user's movement or illumination, and can quickly re-detect the position of the eyes and obtain the coordinates of the eyes. Summary of the invention
[0004] One or more example embodiments may address at least the above problems and / or disadvantages and other disadvantages not described above. In addition, example embodiments are not required to overcome the disadvantages described above, and example embodiments may not overcome any of the problems described above.
[0005] One or more example embodiments provide a method in which, in response to a failure in tracking a user's face or eyes due to the influence of the user's motion or illumination, the position of the eyes can be quickly re-detected and the coordinates of the eyes can be obtained. Therefore, the accuracy and speed of tracking the user's face or eyes can be improved.
[0006] According to one aspect of an example embodiment, there is provided an image processing method, the image processing method comprising: acquiring an image frame; tracking a facial area of a user based on first prior information obtained from at least one previous frame of the image frame; setting a scanning area in the image frame based on second prior information obtained from the at least one previous frame based on determining that tracking of the facial area based on the first prior information has failed; and detecting the facial area in the image frame based on the scanning area.
[0007] The second prior information may include information related to at least one previous scanning area, the detection of the facial area is performed in the at least one previous frame based on the at least one previous scanning area, and the setting step may include resetting the scanning area to the area to which the at least one previous scanning area is expanded.
[0008] The resetting may include resetting the scanning area by sequentially enlarging a size of a window used to set the previous scanning area based on whether tracking of the face area in the previous scanning area has failed.
[0009] The step of resetting the scanning area by sequentially expanding the size of the window may include: sequentially expanding the size of the window used to set the previous scanning area based on the number of times the tracking of the facial area in the previous scanning area has failed; and resetting the scanning area based on the size of the sequentially expanded window.
[0010] Based on the number of times that tracking of the facial region has failed, the step of sequentially expanding the size of the window may include at least one of the following items: based on determining that tracking of the facial region has failed once, expanding the size of the window to the size of the first window, and the previous scanning area is expanded upward, downward, left and right to the first window; based on determining that tracking of the facial region has failed twice, expanding the size of the window to the size of the second window, and the scanning area of the first window is expanded to the second window to the left and right; and based on determining that tracking of the facial region has failed three times, expanding the size of the window to the size of the third window, and the scanning area of the second window is expanded to the third window upward and downward.
[0011] The image processing method may further include setting an initial scanning window corresponding to the scanning area based on the pupil center coordinates of the user accumulated in the at least one previous frame.
[0012] The image processing method may further include: selecting an initial scanning window corresponding to the scanning area from among a plurality of candidate windows based on a feature part of a face of a user included in the image frame; and setting the initial scanning area based on the initial scanning window.
[0013] The selecting may include selecting an initial scanning window from among the plurality of candidate windows based on statistical location coordinates of the user and a location of a camera used to capture the image frame.
[0014] The tracking may include: aligning a plurality of predetermined feature points at a plurality of feature parts included in the face area; and tracking the face of the user based on the aligned plurality of predetermined feature points.
[0015] The aligning may include mapping the plurality of predetermined feature points based on image information in the face region.
[0016] The aligning may include aligning the plurality of predetermined feature points at the plurality of feature parts included in the face region and an adjacent region of the face region.
[0017] The first prior information may include at least one of the following items: the user's pupil center coordinates accumulated in the at least one previous frame, the position coordinates of feature points corresponding to the user's face in the at least one previous frame, and the position coordinates of feature points corresponding to the user's eyes and nose in the at least one previous frame.
[0018] The step of tracking may include: generating a tracking map corresponding to the facial region based on the first prior information; and tracking the facial region of the user based on the tracking map.
[0019] The generating step may include generating a tracking map based on a movable range of the facial region in the image frame according to the first a priori information.
[0020] The image processing method may further include outputting information related to the detected facial region of the user.
[0021] The outputting may include outputting information related to at least one of positions of a pupil and a nose included in the scanning area, a viewpoint passing through the position of the pupil, and a facial expression of the user expressed in the scanning area.
[0022] The image frame may include at least one of a color image frame and an infrared image frame.
[0023] According to an aspect of example embodiments, there is provided a non-transitory computer-readable storage medium storing instructions which, when executed by at least one processor, cause the at least one processor to perform the image processing method.
[0024] According to one aspect of an example embodiment, there is provided an image processing device, comprising: a sensor configured to acquire an image frame; a processor configured to: track a facial region of a user based on first prior information obtained from at least one previous frame of the image frame, set a scanning region in the image frame based on second prior information obtained from the at least one previous frame based on determining that tracking of the facial region based on the first prior information has failed, and detect the facial region in the image frame based on the scanning region; and a display configured to output information related to the detected facial region of the user.
[0025] The second prior information may include information related to at least one previous scanning area, the detection of the facial area is performed in the at least one previous frame based on the at least one previous scanning area, and the processor may also be configured to reset the scanning area to the area to which the at least one previous scanning area is enlarged.
[0026] The processor may be further configured to reset the scanning area by sequentially enlarging a size of a window used to set the previous scanning area based on whether tracking of the face area in the previous scanning area fails.
[0027] The processor may be further configured to sequentially expand a size of a window used to set a previous scan area based on a number of times tracking of a face area in the previous scan area has failed, and reset the scan area based on the size of the sequentially expanded window.
[0028] The processor may also be configured to perform at least one of the following: based on determining that tracking of the facial area has failed once in a previous scanning area, expanding the size of the window to the size of the first window, and the previous scanning area is expanded upward, downward, left and right to the first window; based on determining that tracking of the facial area has failed twice, expanding the size of the window to the size of the second window, and expanding the scanning area of the first window to the left and right to the second window; and based on determining that tracking of the facial area has failed three times, expanding the size of the window to the size of the third window, and expanding the scanning area of the second window to the third window upward and downward.
[0029] The processor may be further configured to set an initial scanning window corresponding to the scanning area based on the pupil center coordinates of the user accumulated in the at least one previous frame.
[0030] The processor may be further configured to select an initial scanning window corresponding to the scanning area from among a plurality of candidate windows based on a feature portion of a face of the user included in the image frame, and set the initial scanning area based on the initial scanning window.
[0031] The processor may be further configured to select an initial scanning window from among the plurality of candidate windows based on statistical location coordinates of the user and a location of a camera used to capture the image frame.
[0032] The processor may be further configured to align a plurality of predetermined feature points at a plurality of feature locations included in the face region, and track the face of the user based on the aligned plurality of predetermined feature points.
[0033] The processor may be further configured to map the plurality of predetermined feature points based on image information in the facial region.
[0034] The processor may be further configured to: align the plurality of predetermined feature points at the plurality of feature parts included in the face region and an adjacent region of the face region.
[0035] The first prior information may include at least one of the following items: the user's pupil center coordinates accumulated in the at least one previous frame, the position coordinates of feature points corresponding to the user's face in the at least one previous frame, and the position coordinates of feature points corresponding to the user's eyes and nose in the at least one previous frame.
[0036] The processor may be further configured to generate a tracking map corresponding to the facial region based on the first a priori information, and track the facial region of the user based on the tracking map.
[0037] The processor may be further configured to generate a tracking map based on a movable range of the facial region in the image frame according to the first a priori information.
[0038] The display may be further configured to output information related to at least one of positions of a pupil and a nose included in the scanning area, a viewpoint through the position of the pupil, and a facial expression of the user expressed in the scanning area.
[0039] The image frame may include at least one of a color image frame and an infrared image frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The above and / or other aspects will become more apparent by describing certain example embodiments with reference to the accompanying drawings, in which:
[0041] Figure 1 illustrates a process for tracking and detecting facial regions according to an example embodiment;
[0042] Figure 2 , Figure 3 and Figure 4 is a flowchart illustrating an image processing method according to an example embodiment;
[0043] Figure 5 illustrates an example of setting a position of an initial scanning window according to an example embodiment;
[0044] Figure 6 An example of setting a scanning area according to an example embodiment is shown;
[0045] Figure 7 illustrates an example of tracking a user's facial area according to an example embodiment; and
[0046] Figure 8 is a block diagram illustrating an image processing apparatus according to an example embodiment. DETAILED DESCRIPTION
[0047] Hereinafter, example embodiments will be described in detail with reference to the accompanying drawings. However, the scope of the disclosure should not be interpreted as being limited to the example embodiments set forth herein. Throughout the present disclosure, the same reference numerals in the accompanying drawings represent the same elements.
[0048] Various modifications may be made to the exemplary embodiments. Here, the exemplary embodiments are not to be interpreted as limited to the disclosure, and should be understood to include all changes, equivalents, and substitutes within the scope of the disclosed ideas and technologies.
[0049] The terms used herein are only used to describe specific example embodiments and will not limit the example embodiments. Unless the context clearly indicates otherwise, as used herein, the singular form is also intended to include the plural form. It will also be understood that the terms "comprise" and / or "include" when used herein indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof.
[0050] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. A statement such as "at least one of...", when following a list of elements, modifies the entire list of elements without modifying the individual elements in the list. For example, the statement "at least one of a, b, and c" should be understood to include only a, only b, only c, both a and b, both a and c, both b and c, or all of a, b, and c.
[0051] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the example embodiments belong. It will also be understood that, unless expressly defined as such herein, terms (such as those defined in common dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense.
[0052] The example embodiments set forth below may be used in autostereoscopic three-dimensional (3D) monitors, autostereoscopic 3D tablets / smartphones, and 3D head-up displays (HUDs) to output eye coordinates by tracking the user's eyes using an infrared (IR) camera or an RGB camera. In addition, the example embodiments may be implemented in the form of a software algorithm in a chip of a monitor, or in the form of an application (app) on a tablet / smartphone, or as an eye tracking device. The example embodiments may be applied to, for example, autonomous vehicles, smart vehicles, smart phones, and mobile devices. Hereinafter, the example embodiments will be described in detail with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals are used for the same elements.
[0053] Figure 1 FIG. 1 shows a process for tracking and detecting facial regions according to an example embodiment. Figure 1, showing a process performed in an image processing device to detect a facial region by scanning an image frame 110 and to track eyes and a nose included in the facial region. However, in the present invention, the tracking object is not limited thereto. For example, the tracking object may be at least one of the body parts of the user associated with facial detection and / or facial tracking (e.g., as a non-limiting example, eyes, nose, eyebrows, mouth, and ears, etc.). In the following, for ease of description, the description is mainly based on eyes and nose as tracking objects. However, it should be understood that other tracking objects are also feasible.
[0054] Assume that in an application for tracking a face based primarily on pupils or eyes and noses, the eye detector may fail to track the eyes or eye regions in the image frame 110. When tracking of the user's eyes in the image frame 110 has failed, the image processing device may detect the user's eye and nose regions using a scanning window 115 as shown in the image 120. The scanning window 115 may correspond to a window for setting a scanning region in the image frame, which will be further described below.
[0055] The image processing device may align a plurality of predetermined feature points corresponding to the eye region and the nose region included in the scanning window 115 in the image 120 at positions corresponding to the face. The plurality of predetermined feature points may be, for example, feature points corresponding to key points representing features of the face such as the eyes and the nose. For example, the plurality of feature points may be indicated as dots (●) and / or asterisks (*) as shown in the images 120 and 130 or as various other marks. Figure 1 In the example embodiment of FIG. 1 , 11 feature points are aligned at corresponding positions on the face.
[0056] As shown in image 130 , the image processing apparatus may extract the user's pupil and / or the user's facial region by tracking the user's face based on a plurality of feature points aligned in image 120 .
[0057] When extraction of the user's pupil and / or extraction of the user's facial region from the image 130 has failed, the image processing apparatus may rescan the image frame 110 in operation 140 .
[0058] When rescanning is performed to detect a facial area in response to a facial tracking failure (tracking loss), the image processing device may limit the scanning area for detecting eyes or eye areas in the image frame based on prior information (e.g., a previously detected facial area or scanning area), rather than scanning the entire image frame, thereby increasing the rate (or speed) of detecting eye coordinates. In an example embodiment in which a driver in a vehicle is photographed by using a camera whose position in the vehicle is fixed and the driver's range of movement is within a limited range, the scanning area in the image frame may be limited, whereby the rate of eye or facial detection may be increased. In addition, according to an example embodiment, when detection of a facial area in a limited scanning area has failed, the size of the window for resetting the scanning area may be sequentially enlarged, which may reduce latency in detecting eye coordinates.
[0059] Figure 2 is a flowchart illustrating an image processing method according to an example embodiment. Figure 2 In operation 210, the image processing device may acquire an image frame. The image frame may include, for example, a color image frame and / or an IR image frame. The image frame may correspond to, for example, an image of the driver captured by an image sensor or camera provided in the vehicle.
[0060] In operation 220, the image processing apparatus may track the facial region of the user based on first a priori information obtained from at least one previous frame in the obtained image frame. If the image frame is acquired at a time point t, at least one previous frame may be acquired at a time point (e.g., t-1) or a time point (e.g., t-1, t-2, and t-3) earlier than the time point t at which the image frame is acquired. The first a priori information may include, for example, at least one of the following items: the pupil center coordinates of the user accumulated in at least one previous frame, the position coordinates of feature points corresponding to the user's face in at least one previous frame, and the position coordinates of feature points corresponding to the user's eyes and nose in at least one previous frame.
[0061] In operation 220, the image processing device may align a plurality of predetermined feature points at a plurality of feature parts of the facial region, and track the user's face based on the aligned plurality of feature points. The "plurality of feature parts" may be partial parts or regions included in the facial region of the image frame, and include, for example, eyes, nose, mouth, eyebrows, and glasses. In this example, the image processing device may align a plurality of feature points in the facial region and / or align a plurality of feature points at a plurality of feature parts included in a neighboring region of the facial region. The image processing device may move (or map) a plurality of predetermined feature points based on image information in the facial region and / or a neighboring region of the facial region.
[0062] In operation 220, the image processing apparatus may determine whether tracking of the user's facial region is successful or has failed based on the first a priori information. Figure 7 An example of tracking a facial region by an image processing device is described in detail.
[0063] In operation 230, in response to determining that tracking of the facial region has failed, the image processing apparatus may set a scanning region in the image frame based on second a priori information obtained from at least one previous frame. The second a priori information may include, for example, information related to at least one previous scanning region based on which detection of the facial region is performed in at least one previous frame. The at least one previous scanning region may be, for example, at least one initial scanning region set in an initial scanning window.
[0064] Hereinafter, for the convenience of description, a window for setting a scanning area in an image frame will be referred to as a "scanning window", and a scanning window initially used to set a scanning area in an image frame will be referred to as an "initial scanning window". A scanning area set by the initial scanning window will be referred to as an "initial scanning area". Figure 5 An example of setting the position and size of the initial scanning window is described in detail.
[0065] In one example, “setting a scan area” may include setting a scan area, setting an initial scan area, or adjusting or resetting a scan area.
[0066] In operation 230, the image processing apparatus may reset the scanning area to the area to which the previous scanning area (e.g., as a non-limiting example, the scanning area used for the most recent tracking) was enlarged. The image processing apparatus may reset the scanning area by sequentially enlarging the size of the window used to set the scanning area based on whether the tracking of the face area in the scanning area has failed. Figure 6 An example of setting or resetting a scanning area by an image processing apparatus is described in detail.
[0067] In operation 240, the image processing device may detect a facial region in the image frame based on the reset scanning area. In one example, the image processing device may output information related to the detected facial region. The information related to the facial region of the user may include, for example, the positions of the pupil and the nose included in the scanning area, the viewpoint through the position of the pupil, and the facial expression of the user expressed in the scanning area. The image processing device may output the information related to the facial region explicitly or implicitly. The expression "explicitly outputting information related to the facial region" means performing the following operations: the operation may include, for example, displaying the position of the pupil included in the facial region and / or the facial expression expressed in the facial region on the screen, and / or outputting information about the position of the pupil included in the facial region and / or the facial expression expressed in the facial region on the screen through audio. The expression "implicitly outputting information related to the facial region" means performing the following operations: the operation may include, for example, adjusting the image displayed on the HUD based on the position of the pupil included in the facial region or the viewpoint through the position of the pupil, or providing a service corresponding to the facial expression expressed in the facial region.
[0068] Figure 3 is a flowchart illustrating an image processing method according to an example embodiment. Figure 3 In operation 310, the image processing apparatus may acquire an nth image frame from a camera. The image frame may be, for example, an RGB color image frame or an IR image frame.
[0069] In operation 320, the image processing apparatus may determine whether eyes and a nose are detected in a previous (n-1)th image frame. The image processing apparatus may determine whether eyes and a nose are detected in an initial scanning area of a previous (n-1)th image frame. In response to determining that eyes and a nose are not detected in operation 320, in operation 370, the image processing apparatus may detect eyes or eyes and a nose by setting or adjusting a scanning area based on prior information. On the other hand, in response to determining that eyes and a nose are detected in operation 320, in operation 330, the image processing apparatus may align predetermined feature points (e.g., such as a plurality of feature points) at, for example, detected eyes or detected eyes and a nose. Figure 1 ). For example, in operation 330, the image processing device may align predetermined feature points at a plurality of feature parts of the detected eyes or the detected eyes and nose included in the scanning area or the adjacent area of the scanning area. As a non-limiting example, the predetermined feature points may include, for example, three feature points of each eye, one feature point between the eyes, one feature point at the tip of the nose, and three feature points of the mouth (or three feature points of the nose).
[0070] In operation 330, the image processing device may align a plurality of feature points at a plurality of feature parts included in a facial region and / or an adjacent region of the facial region. The image processing device may move (or map) a plurality of predetermined feature points to be aligned at a plurality of feature parts based on image information in the facial region and / or an adjacent region of the facial region. The image processing device may identify the positions of feature parts corresponding to the eyes and nose of the user from the facial region of the image frame based on various methods (such as, for example, a supervised descent method (SDM) for aligning feature points on the shape of an image using a descent vector learned from an initial shape configuration, an active shape model (ASM) for aligning feature points based on a principal component analysis (PCA) of a shape and a shape, an active appearance model (AAM), or a constrained local model (CLM)). The image processing device may move a plurality of predetermined feature points to be aligned at the positions of the identified plurality of feature parts. For example, when the image frame is an initial image frame, the plurality of feature points to be aligned may correspond to the average position of feature parts of a plurality of users. In addition, when the image frame is not an initial image frame, the plurality of feature points to be aligned may correspond to the plurality of feature points aligned based on a previous image frame.
[0071] In operation 340, the image processing apparatus may check an alignment result corresponding to a combination of a plurality of feature parts aligned with the feature point. The image processing apparatus may check an alignment result of a facial region corresponding to a combination of a plurality of feature parts (e.g., eyes and nose) in the scan region based on information in the scan region.
[0072] The image processing device may check whether a plurality of feature parts in the scanned area are classes corresponding to a combination of eyes and noses based on the image information in the scanned area. The image processing device may use a checker to check whether the scanned area is a face class, for example, based on a scale-invariant feature transform (SIFR) feature. Here, the "SIFR feature" may be obtained by the following two operations. The image processing device may extract candidate feature points having a local maximum or minimum brightness of the image in the scale space from the image data of the scanned area through an image pyramid, and select feature points to be used for image matching by filtering out feature points with low contrast. The image processing device may obtain a directional component based on a gradient of a neighboring area with respect to the selected feature point, and generate a descriptor by resetting an area of interest with respect to the obtained directional component and detecting the size of the feature point. Here, the descriptor may correspond to the SIFR feature. In addition, if the feature points corresponding to the key points of the eyes and nose of each face stored in the training image database (DB) are aligned in the face region of the training image frame, the "checker" may be a classifier trained using SIFT features extracted from the aligned feature points. The checker may check whether the face region where the feature points are aligned corresponds to a real face class based on the image information in the face region of the image frame. The checker may be, for example, a support vector machine classifier. The checker may also be referred to as a "face checker" because the checker checks the alignment for face regions.
[0073] In operation 350, the image processing apparatus may determine whether eyes are detected as a result of the checking in operation 340. In response to determining that eyes are not detected in operation 350, in operation 370, the image processing apparatus may detect eyes or eyes and nose by setting a scanning area based on prior information.
[0074] In operation 360 , in response to determining that the eye is detected, the image processing apparatus may output coordinates of the eye or coordinates of the pupil.
[0075] Figure 4 is a flowchart illustrating an image processing method according to an example embodiment. Figure 4 In operation 410, the image processing apparatus may determine whether a facial region including eyes and a nose is detected in an image frame (e.g., an nth frame). In response to determining that the eyes and the nose are detected, in operation 440, the image processing apparatus may align a plurality of predetermined feature points at a plurality of feature parts of the facial region (e.g., the eyes, the nose, a middle part between the eyes, and the pupil). In operation 450, the image processing apparatus may track the facial region of the user based on the aligned plurality of feature points.
[0076] On the other hand, in response to determining that the facial region is not detected, the image processing apparatus may set a scanning region in the image frame based on prior information obtained from at least one previous frame (e.g., the (n-1)th frame) of the image frame in operation 420. The scanning region may be set by a window for setting the scanning region.
[0077] In operation 430, the image processing apparatus may determine whether a facial region is detected in the set scanning area (i.e., whether detection of the facial region in the image frame (e.g., the nth frame) is successful by using the set scanning area). In response to determining that the detection of the facial region is successful, in operation 440, the image processing apparatus may align a plurality of predetermined feature points at a plurality of feature parts of the facial region. In operation 450, the image processing apparatus may track the facial region of the user based on the aligned plurality of feature points.
[0078] In response to determining that the detection of the facial region has failed, the image processing device may reset the scanning region by sequentially expanding the size of the window used to set the scanning region. In response to determining that the detection of the facial region has failed in operation 430, the image processing device may expand the size of the scanning window in operation 460. In operation 460, the image processing device may reset the scanning region by sequentially expanding the size of the window used to set the scanning region based on whether the tracking of the facial region in the scanning region has failed. In operation 420, the image processing device may reset the scanning region based on the size of the sequentially expanded window. The image processing device may reset the scanning region by repeating the sequential expansion of the size of the scanning window based on the number of times the tracking of the facial region in the scanning region has failed. For example, whenever the tracking of the facial region in the scanning region has failed, the image processing device may reset the scanning region by sequentially expanding the size of the scanning window. The number of iterations of expanding the size of the scanning window may be determined by the user as a predetermined number of times (e.g., three or four times). Repetition may be performed until the tracking of the facial region in the scanning region is successful. In other words, it is possible for the size of the scanning window to expand to the entire image frame. Referring to Figure 6 An example of setting or resetting a scanning area by an image processing apparatus is described in detail.
[0079] Figure 5 An example of setting the position of the initial scanning window according to an exemplary embodiment is shown. Figure 5 , showing an example of setting an initial scanning window 517 in an image frame 510.
[0080] The image processing apparatus may select an initial scanning window 517 corresponding to the scanning area from among the plurality of candidate windows 515 based on the characteristic parts of the user's face (e.g., eyes, nose, mouth, eyebrows, and glasses) included in the facial area of the image frame 510. The image processing apparatus may select the initial scanning window 517 from among the plurality of candidate windows 515 based on the statistical position coordinates of the user and the position of the camera used to capture the image frame 510. In the example of face detection in a driving environment, the statistical position coordinates of the user may correspond to, for example, position coordinates that average position coordinates to which a user sitting on a driver's seat of a vehicle may move on average. The image processing apparatus may set the initial scanning area based on the position of the initial scanning window 517.
[0081] In an exemplary embodiment, the image processing device may set an initial scanning window 517 for the image frame based on the pupil center coordinates of the user accumulated in at least one previous frame. The image processing device may set the initial scanning window 517, for example, in an area having a left margin and a right margin based on the pupil center coordinates of the user.
[0082] Figure 6 An example of setting a scanning area according to an exemplary embodiment is shown. Figure 6 , an initial scanning window 611 , an image 610 provided with a first window 615 , an image 630 provided with a second window 635 , and an image 650 provided with a third window 655 are shown.
[0083] The image processing apparatus may set the scanning area by sequentially expanding the size of a window for setting the scanning area based on whether tracking of the face area in the scanning area has failed.
[0084] As shown in image 610, if the tracking of the facial region through the scanning window has failed once, the image processing apparatus may expand the size of the scanning window to the size of the first window 615. For example, the size of the first window 615 may correspond to the size of the initial scanning area from Figure 6 The size of the initial scanning window 611 after it is expanded upward, downward, leftward and rightward. In detail, the size of the first window 615 may correspond to the size after the initial scanning area is expanded upward, downward, leftward and rightward by, for example, 5%.
[0085] like Figure 6 As shown in the image 630 of FIG. 1 , if the tracking of the facial region has failed twice, the image processing device may expand the size of the scanning window to the size of the second window 635. For example, the size of the second window 635 may correspond to the size of the second window 635 based on the facial region. Figure 6The size of the scanning area after the first window 615 is expanded to the left and right. In detail, the size of the second window 635 may correspond to the size of the first window 615 expanded to the left and right by "the distance from the middle part between the eyes to each eye in the initial scanning window 611 + a predetermined margin (e.g., 2 mm)". For example, the distance from the middle part between the eyes to the right eye (or left eye) may be 3.5 cm, and this distance may be obtained by averaging the distances from the middle part between the eyes of multiple users to the right eye (or left eye). In this example, the size of the second window 635 may correspond to the size of the first window 615 expanded to the left and right by 3.7 cm.
[0086] like Figure 6 As shown in the image 650 of FIG. 1 , if the tracking of the facial region has failed three times, the image processing device may expand the size of the scanning window to the size of the third window 655. For example, the size of the third window 655 may correspond to the size of the third window 655 based on the facial region. Figure 6 The size of the scanning area after the second window 635 is expanded upward and downward. In detail, the size of the third window 655 may correspond to the size of the second window 635 expanded upward and downward by "the distance from the middle part between the eyes to each eye in the initial scanning window 611 + a predetermined margin (e.g., 2 mm)". For example, the distance from the middle part between the eyes to the right eye (or left eye) may be 3.5 cm. In this example, the size of the third window 655 may correspond to the size of the second window 635 expanded upward and downward by 3.7 cm.
[0087] Although in Figure 6 615 corresponds to the size of the initial scanning window 611 after it is expanded upward, downward, leftward, and rightward, the size of the second window 635 corresponds to the size of the first window 615 after it is expanded leftward and rightward, and the size of the third window 655 corresponds to the size of the second window 635 after it is expanded upward and downward, but these are examples given for the purpose of illustration only, and the disclosure is not limited to the directions of expansion described herein. It should be understood that the window can be expanded in any direction of upward, downward, leftward, and rightward, or any combination thereof, or in any other direction.
[0088] Figure 7 An example of tracking a user's facial area according to an example embodiment is shown. Figure 7 , an image frame 710 and a tracking map 720 generated based on a previous frame of the image frame 710 are shown.
[0089] The image processing apparatus may generate a tracking map 720 corresponding to the facial region of the image frame 710 based on prior information (e.g., first prior information) obtained from at least one previous frame of the image frame 710. The first prior information may include, for example, at least one of the following items: pupil center coordinates of the user accumulated in at least one previous frame, position coordinates of feature points corresponding to the face of the user in at least one previous frame, and position coordinates of feature points corresponding to the eyes and nose of the user in at least one previous frame.
[0090] For example, the image processing apparatus may determine, based on the first a priori information, a moving range of a facial region that is movable in the image frame 710. The moving range of the facial region may include: a moving range of the facial region for a case where a driver sitting in a driver's seat moves his upper body or head left and right, and a moving range of the facial region for a case where the driver turns his upper body or head forward and backward.
[0091] The image processing device may generate a tracking map 720 based on the movement range of the facial region. The tracking map 720 may include coordinates corresponding to the maximum movement range of the facial region for the facial region to move in the upward, downward, leftward, and rightward directions. The image processing device may track the facial region of the user based on the tracking map 720.
[0092] Figure 8 is a block diagram illustrating an image processing apparatus according to an example embodiment. Figure 8 , the image processing device 800 may include a sensor 810, a processor 830, a memory 850, a communication interface 870, and a display 890. The sensor 810, the processor 830, the memory 850, the communication interface 870, and the display 890 may communicate with each other via a communication bus 805.
[0093] The sensor 810 may acquire an image frame. The sensor 810 may be, for example, an image sensor, a visual sensor, or an IR camera configured to capture an input image through IR radiation. The image frame may include, for example, a facial image of a user or an image of a user driving a vehicle.
[0094] The processor 830 may track the facial region of the user based on first a priori information obtained from at least one previous frame of the image frame. In response to determining that the tracking of the facial region based on the first a priori information has failed, the processor 830 may set a scanning region in the image frame based on second a priori information obtained from at least one previous frame. The processor 830 may detect the facial region in the image frame based on the scanning region. A single processor 830 or a plurality of processors 830 may be provided.
[0095] The processor 830 may execute Figures 1 to 7The processor 830 may execute a program and control the image processing apparatus 800. The program code to be executed by the processor 830 may be stored in the memory 850. The processor 830 may be configured as, for example, a central processing unit (CPU), a graphics processing unit (GPU), or a neural network processor (NPU).
[0096] The memory 850 may store an image frame acquired by the sensor 810, first and second a priori information obtained by the processor 830 from at least one previous frame of the image frame, a facial region of a user obtained by the processor 830, and / or information related to the facial region. In addition, the memory 850 may store information related to the facial region detected by the processor 830. The memory 850 may be a volatile memory or a nonvolatile memory.
[0097] The communication interface 870 may receive an image frame from outside the image processing apparatus 800. The communication interface 870 may output a facial region detected by the processor 830 and / or information related to the facial region of the user. The communication interface 870 may receive an image frame captured outside the image processing apparatus 800 or information of various sensors received from outside the image processing apparatus 800.
[0098] The display 890 may display the processing result obtained by the processor 830, for example, information related to the facial region of the user. For example, when the image processing apparatus 800 is embedded in a vehicle, the display 890 may be configured as a HUD.
[0099] The unit described herein can be implemented using hardware components, software components or their combination. For example, the processing device can be implemented using one or more general or special computers (such as, for example, with a processor, a controller, an arithmetic logic unit, a digital signal processor, a microcomputer, a field programmable array, a programmable logic unit, a microprocessor, any other device that can respond and execute instructions in a limited manner). The processing device can run an operating system (OS) and one or more software applications running on the OS. The processing device can also access, store, manipulate, process and generate data in response to the execution of software. For simple purposes, the description of the processing device is used as a singular, however, it will be understood by those skilled in the art that the processing device may include multiple processing elements and multiple types of processing elements. For example, the processing device may include multiple processors, or a processor and a controller. In addition, different processing configurations (such as, parallel processors) are feasible.
[0100] Software may include computer programs, code segments, instructions, or some combination thereof, to independently or collectively instruct or configure a processing device to operate as desired. Software and data may be permanently or semi-permanently implemented in any type of machine, component, physical or virtual device, computer storage medium or device, or in a propagating signal wave capable of providing instructions or data to or interpreted by a processing device. Software may also be distributed on networked computer systems so that the software is stored and executed in a distributed manner. In particular, software and data may be stored by one or more non-transitory computer-readable recording media.
[0101] The method according to the example embodiment described herein may be recorded in a non-transitory computer-readable medium, which includes program instructions for performing various operations implemented by a computer. The medium may also include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the medium may be program instructions specially designed and constructed for the purpose of implementation here, or may be known to a person of ordinary skill in the relevant field. Examples of non-transitory computer-readable media include: magnetic media (such as hard disks, floppy disks, and tapes); optical media (such as compact disks (CDs), read-only memories (ROMs), and digital versatile disks (DVDs)); magneto-optical media (such as floppy disks); and hardware devices (such as ROMs, random access memories (RAMs), flash memories) and the like that are specially configured to store and execute program instructions. Examples of program instructions include both machine codes such as those generated by a compiler and files containing high-level codes that can be executed by a computer using an interpreter. The above-mentioned device may be configured to act as one or more software modules to perform the operations of the above-mentioned example embodiments, or vice versa.
[0102] A number of examples have been described above. However, it should be understood that various modifications may be made. For example, suitable results may be achieved if the described techniques are performed in a different order, and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner and / or replaced or supplemented by other components or their equivalents. Therefore, other implementations are within the scope of the claims.
Claims
1. An image processing method, include: Get image frame; tracking a facial region of the user based on first a priori information obtained from at least one previous frame of the image frames; resetting the scanning area in the image frame based on second a priori information obtained from the at least one previous frame in response to determining that tracking of the facial area based on the first a priori information has failed; as well as The facial region is detected in the image frame based on the reset scanning region, and the steps of resetting the scanning region and detecting the facial region are repeatedly performed until the tracking of the facial region in the scanning region is successful. wherein the second a priori information includes information related to at least one initial scanning area, detection of the facial area is performed in the at least one previous frame based on the at least one initial scanning area, and the at least one initial scanning area is determined based on the statistical position coordinates of the user and the position of the camera used for shooting, The statistical position coordinates of the user correspond to the position coordinates obtained by averaging the position coordinates within a limited range to which the user sitting on the driver's seat moves on average.
2. The image processing method according to claim 1, in, The step of resetting the scanning area includes resetting the scanning area to the area to which the at least one initial scanning area is expanded.
3. The image processing method according to claim 2, in, The step of resetting the scanning area includes resetting the scanning area by sequentially enlarging a size of a window for setting the scanning area based on whether tracking of a face area in the scanning area has failed.
4. The image processing method according to claim 3, in, The steps of resetting the scanning area by sequentially enlarging the size of the window include: sequentially expanding the size of a window for setting the scan area based on the number of times tracking of a face area in the scan area has failed; and The scanning area is reset based on the size of the sequentially enlarged window.
5. The image processing method according to claim 4, in, Based on the number of times that tracking of the face region in the scan area has failed, the step of sequentially enlarging the size of the window includes at least one of the following: Based on determining that tracking of the facial region in the scan area has failed once, expanding the size of the window to the size of the first window, the initial scan area expanding upward, downward, leftward, and rightward to the first window; Based on determining that tracking of the facial region in the scan area has failed twice, expanding the size of the window to a size of a second window, the scan area of the first window expanding to the left and right to the second window; and Based on determining that tracking of the face region in the scan area has failed three times, the size of the window is expanded to the size of a third window, and the scan area based on the second window is expanded upward and downward to the third window.
6. The image processing method according to claim 1, further comprising: include: An initial scanning window corresponding to the scanning area is set based on the pupil center coordinates of the user accumulated in the at least one previous frame.
7. The image processing method according to claim 1, further comprising: include: selecting an initial scanning window corresponding to the scanning area from among a plurality of candidate windows based on a feature part of a user's face included in the image frame; as well as The initial scan area is set based on the initial scan window.
8. The image processing method according to claim 7, in, The step of selecting includes selecting an initial scanning window from among the plurality of candidate windows based on the statistical position coordinates of the user and the position of a camera used to capture the image frame.
9. The image processing method according to claim 1, in, The steps of tracking include: aligning a plurality of predetermined feature points at a plurality of feature locations included in the face region and / or an adjacent region of the face region; and The user's face is tracked based on the alignment of a plurality of predetermined feature points.
10. The image processing method according to claim 9, in, The step of aligning includes mapping the plurality of predetermined feature points based on image information in the face region and / or an area adjacent to the face region.
11. The image processing method according to claim 1, in, The first prior information includes at least one of the following items: the user's pupil center coordinates accumulated in the at least one previous frame, the position coordinates of feature points corresponding to the user's face in the at least one previous frame, and the position coordinates of feature points corresponding to the user's eyes and nose in the at least one previous frame.
12. The image processing method according to claim 1, in, The steps of tracking include: generating a tracking map corresponding to the facial region based on the first prior information; and The user's facial region is tracked based on the tracking map.
13. The image processing method according to claim 12, in, The generating step includes generating a tracking map based on a movable range of the face region in the image frame according to first a priori information.
14. The image processing method according to claim 1, further comprising: include: Information related to the detected facial region of the user is output.
15. The image processing method according to claim 14, in, The step of outputting includes outputting information related to at least one of positions of a pupil and a nose included in the scanning area, a viewpoint passing through the position of the pupil, and a facial expression of a user expressed in the scanning area.
16. The image processing method according to claim 1, in, The image frame includes at least one of a color image frame and an infrared image frame.
17. A non-transitory computer-readable storage medium storing instructions, which, when executed by at least one processor, cause the at least one processor to perform an image processing method, the image processing method include: Get image frame; tracking a facial region of the user based on first a priori information obtained from at least one previous frame of the image frames; resetting the scanning area in the image frame based on second a priori information obtained from the at least one previous frame in response to determining that tracking of the facial area based on the first a priori information has failed; as well as The facial region is detected in the image frame based on the reset scanning region, and the steps of resetting the scanning region and detecting the facial region are repeatedly performed until the tracking of the facial region in the scanning region is successful. wherein the second a priori information includes information related to at least one initial scanning area, detection of the facial area is performed in the at least one previous frame based on the at least one initial scanning area, and the at least one initial scanning area is determined based on the statistical position coordinates of the user and the position of the camera used for shooting, The statistical position coordinates of the user correspond to the position coordinates obtained by averaging the position coordinates within a limited range to which the user sitting on the driver's seat moves on average.
18. An image processing device, include: a sensor configured to acquire an image frame; The processor is configured as: tracking a facial region of the user based on first a priori information obtained from at least one previous frame of the image frame, Resetting the scanning area in the image frame based on second a priori information obtained from the at least one previous frame in response to determining that tracking of the facial area based on the first a priori information has failed, and Detecting a facial region in the image frame based on the reset scanning region, and repeatedly performing the steps of resetting the scanning region and detecting the facial region until tracking of the facial region in the scanning region is successful; as well as a display configured to output information related to the detected facial region of the user, wherein the second a priori information includes information related to at least one initial scanning area, detection of the facial area is performed in the at least one previous frame based on the at least one initial scanning area, and the at least one initial scanning area is determined based on the statistical position coordinates of the user and the position of the camera used for shooting, The statistical position coordinates of the user correspond to the position coordinates obtained by averaging the position coordinates within a limited range to which the user sitting on the driver's seat moves on average.
19. The image processing device according to claim 18, in, The processor is further configured to reset the scanning area to the area to which the at least one initial scanning area is enlarged.
20. The image processing device according to claim 19, in, The processor is further configured to reset the scanning area by sequentially enlarging a size of a window for setting the scanning area based on whether tracking of the face area in the scanning area fails.
21. The image processing device according to claim 20, in, The processor is further configured to sequentially expand a size of a window for setting the scanning area based on the number of times tracking of the face area in the scanning area has failed, and reset the scanning area based on the size of the sequentially expanded window.
22. The image processing device according to claim 21, in, The processor is further configured to perform at least one of the following: Based on determining that tracking of the facial region in the scan area has failed once, expanding the size of the window to the size of the first window, the initial scan area expanding upward, downward, leftward, and rightward to the first window; Based on determining that tracking of the facial region in the scan area has failed twice, expanding the size of the window to a size of a second window, the scan area of the first window expanding to the left and right to the second window; and Based on determining that tracking of the face region in the scan area has failed three times, the size of the window is expanded to the size of a third window, and the scan area based on the second window is expanded upward and downward to the third window.
23. The image processing device according to claim 18, in, The processor is further configured to set an initial scanning window corresponding to the scanning area based on the pupil center coordinates of the user accumulated in the at least one previous frame.
24. The image processing apparatus according to claim 18, in, The processor is further configured to select an initial scanning window corresponding to the scanning area from among a plurality of candidate windows based on a feature part of the user's face included in the image frame, and set the initial scanning area based on the initial scanning window.
25. The image processing device according to claim 24, in, The processor is further configured to select an initial scanning window from among the plurality of candidate windows based on the statistical location coordinates of the user and the location of a camera used to capture the image frame.
26. The image processing apparatus according to claim 18, in, The processor is further configured to align a plurality of predetermined feature points at a plurality of feature locations included in the face region and / or an adjacent region of the face region, and track the user's face based on the aligned plurality of predetermined feature points.
27. The image processing device according to claim 26, in, The processor is further configured to map the plurality of predetermined feature points based on image information in the facial region and / or an area adjacent to the facial region.
28. The image processing apparatus according to claim 18, in, The first prior information includes at least one of the following items: the user's pupil center coordinates accumulated in the at least one previous frame, the position coordinates of feature points corresponding to the user's face in the at least one previous frame, and the position coordinates of feature points corresponding to the user's eyes and nose in the at least one previous frame.
29. The image processing apparatus according to claim 18, in, The processor is further configured to generate a tracking map corresponding to the facial region based on the first a priori information, and track the facial region of the user based on the tracking map.
30. The image processing device according to claim 29, in, The processor is further configured to generate a tracking map based on a movable range of the facial region in the image frame according to the first a priori information.
31. The image processing apparatus according to claim 18, in, The display is further configured to output information related to at least one of positions of a pupil and a nose included in the scanning area, a viewpoint through the position of the pupil, and a facial expression of a user expressed in the scanning area.
32. The image processing apparatus according to claim 18, in, The image frame includes at least one of a color image frame and an infrared image frame.
Citation Information
Patent Citations
Powder metallurgy mixture, sintered body, and method for producing sintered body
KR1020190104570A
Face tracking for controlling imaging parameters
US20090303342A1