System for user interface using facial gesture recognition and face verification and method therefor
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- IND ACAD COOPERATION GRP OF BAEKSEOK UNIV
- Filing Date
- 2024-01-09
- Publication Date
- 2026-07-29
Smart Images

Figure 112024003284816-PAT00097_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a user interface system and method using face gesture recognition and face authentication, and more specifically, to a user interface system and method using face authentication and face gesture recognition that authenticates the user's face in the process of controlling a user interface by recognizing face gestures composed of combinations such as 3D face angles, binocular opening / closing states and the time of maintenance thereof, and provides a face gesture-based user interface only to authorized operators. Background Technology
[0002] Human-Computer Interaction (HCI) technology has undergone continuous technological innovation to enhance natural communication between humans and computers or machines through means that transcend the five human senses, moving beyond the Command Line Interface (CLI), which involves inputting characters using a keyboard, to the Graphic User Interface (GUI), which utilizes graphic input methods such as icons, buttons, windows, dialog boxes, scrolls, and drag-and-drop. Even recently, the most common HCI technology for controlling computers or digital devices remains the keyboard and mouse-based GUI, and the mouse maintains its status as the undisputed core pointing input device of the GUI. When a user manipulates the mouse to trigger user events such as left / right clicks, double clicks, drags, drops, and scrolling up / down at the cursor position, the operating system realizes user interaction on the computer by processing commands corresponding to those user events. However, driven by the rapid advancements in sensor technology, artificial intelligence, and CPUs, GPUs, and memory, new concepts of Natural User Interface (NUI) are being researched one after another. These methods move beyond the traditional approach using only keyboards and mice to operate digital devices using biosignals such as Electromyogram (EMG), Electrocardiogram (ECG / EKG), Electroencephalogram (EEG), Sphygmogram, Electrooculogram (EOG), and Facial EMG, in addition to the user's voice, gaze, facial expressions, gestures, and touch. The implementation of practical and highly sophisticated NUI is a long-cherished goal and an immediate challenge for the HCI field.User interfaces using NUI are evolving in a direction that eliminates the point of contact between humans and machines and autonomously understands human intent. NUI not only provides an intuitive and natural user experience but also removes the constraint that a keyboard and mouse must be installed for computer operation.
[0003] Among the HCI technologies that provide the foundation for this transition, gesture recognition technology—such as that involving hands, faces, and body movements—has been researched since the 1990s across various fields, including human-robot interaction, 3D game interfaces, virtual reality, interaction with home appliances or mobile devices, medical posture correction or movement measurement, sign language recognition, and drone control. NUI enables natural interactions that reflect users' cognitive or physical abilities or the situations they face, allowing for user-friendly processing of desired tasks by operating according to the context and purpose. Above all, there is a demand for advanced NUI technology capable of intuitively and rapidly processing tasks in environments of the materials, parts, and equipment (Sobujang) industry, where various equipment and a number of collaborators coexist, as well as in residential and commercial gathering places.
[0004] Recently, as Google’s MediaPipe has emerged as a central cross-platform framework in the field of gesture recognition, hand gesture interface technologies based on the MediaPipe Hands model are being proposed. The research team of this invention has also partially disclosed, through papers and patent applications, hand gesture interface technologies capable of detecting static and dynamic hand postures to enable operations such as dialing, clicking a mouse, and writing. However, hand gesture-based user interfaces have limitations; prolonged use can lead to accumulated fatigue or spasms, and gesture recognition is impossible when both hands are not free.
[0005] On the other hand, facial gesture-based user interfaces cause less fatigue than hand gestures during prolonged use, and above all, since the face and gaze move together, they are more intuitive and offer good user convenience. However, in environments such as the materials, parts, and equipment industry where many collaborators exist, or in densely populated residential or commercial settings, if registered operators with operational authority could be quickly identified and facial gesture-based user interface technology utilized without a separate medium, user convenience and security could be further enhanced. Prior art literature
[0006] Registered Patent No. 10-2522142 (April 11, 2023) The problem to be solved
[0007] The present invention aims to solve the aforementioned problem by providing a user interface system and method using face gesture recognition and face authentication, which recognizes face gestures composed of combinations such as 3D face angles of Pan angle, Tilt angle, and Roll angle, binocular opening / closing state, and the maintenance time thereof in a series of input frames to provide user interface functions, and in the process combines InsightFace-based face authentication technology to identify an authorized operator and provide control of the face gesture-based user interface only to that operator. means of solving the problem
[0008] A user interface system using face gesture recognition and face authentication according to the present invention comprises a face authentication unit that identifies the face features of a user granted operation authority among a plurality of users in a series of input image frames, and then selects a face-authenticated operator through this identification, and
[0009] It includes a user interface unit that allows facial gesture recognition of an operator selected through the facial authentication unit, displays the facial gesture recognition result on a screen when necessary to interact with the operator, and controls an interface to generate a user event based on the facial gesture recognition result and execute a corresponding command.
[0010] The face authentication unit comprises: a face authentication dataset generation module that generates a face authentication dataset by sampling a series of person image frames, including frontal, side, and tilted face compositions, at regular frame intervals; a face data learning module that generates a face recognition model through a supervised learning process that classifies faces using cosine similarity after extracting and normalizing multidimensional embedding vectors from face regions detected from the face authentication dataset; a face recognition module that detects face regions in a series of input image frames and performs face recognition when input into the face recognition model, and outputs the corresponding face class and recognition confidence; and an operator authentication module that authenticates the user of the face as the operator only when the recognized face class matches a registered operator and the recognition confidence is greater than or equal to a preset authentication threshold.
[0011] The above-mentioned face authentication unit detects a face area and performs face recognition to authenticate a user with a recognition confidence level higher than a preset level as the operator; however, if multiple people exist within the input image, immediately after one person is recognized, the unit repeats the process of recognizing other people while immediately erasing or ignoring the recognized person from the screen.
[0012] The user interface unit includes an estimation module that detects the movement of a user's face in a series of input video frames, estimates the pan and tilt angles of the face, and calculates the horizontal and vertical coordinates of a cursor position on the screen corresponding to each face angle; and a face gesture recognition module that detects the movement of the face and binoculars through face landmarks estimated in the input video frames, recognizes a face gesture composed of a combination of the roll angle of the face, the opening and closing state of the binoculars, and the holding time thereof, generates a user event corresponding to the face gesture in association with the cursor position calculated by the estimation module, and controls the interface to execute a command corresponding to the user event.
[0013] The above estimation module further includes a cursor stabilization unit that stabilizes the fine shaking of the cursor by applying adaptive moving average processing based on the cursor movement speed to the horizontal and vertical coordinates of the cursor on the screen.
[0014] The above-described user interface unit further includes a binocular opening / closing determination unit that, when determining the binocular simultaneous opening / closing state, repeatedly checks for discontinuities in the binocular simultaneous opening / closing and corrects the temporary opening / closing discrepancy state of the binoculars under specific conditions so that it is unified into either an open or closed state.
[0015] The user interface unit, as soon as the face gesture interface is activated, suppresses the movement of the user's face and ensures the user looks at the exact center of the screen; it then cumulatively averages the nose tip landmark coordinates estimated from video frames received during an initial predetermined time to establish them as the angle calculation reference point of the 3D face landmark coordinate system; subsequently, it obtains a 3D vector connecting the origin of the 3D face landmark coordinate system to the angle calculation reference point; and after obtaining a 3D vector connecting the nose tip landmark coordinates estimated after establishing the angle calculation reference point from the origin of the 3D face landmark coordinate system, it calculates the horizontal and vertical angles between the two 3D vectors and sets them as the face's Pan angle and Tilt angle, respectively; and when the face's Pan angle is within the range [(-)maximum Pan angle, (+)maximum Pan angle] with the screen width and the face's maximum Pan angle preset, it converts the face's Pan angle into a screen horizontal coordinate within the range [0, screen width] using trigonometric functions and proportional equations; and when the face's Tilt angle is within [(-)maximum When within the range of [Tilt Angle, (+)Maximum Tilt Angle], the Tilt Angle of the face is converted into a vertical screen coordinate within the range of [0, Screen Vertical Size] using trigonometric functions and a proportional equation.
[0016] The user interface unit calculates the Roll angle of the face as the angle between the vector connecting the right eye corner landmark coordinate of the user face to the left eye corner landmark coordinate and the horizontal coordinate axis of the 3D face landmark coordinate system, and if the absolute value of the Roll angle of the face is greater than or equal to a preset threshold Roll angle, it determines whether it is a positive (+) angle or a negative (-) angle and generates either a Scroll Down Event or a Scroll Up Event, and generates the other user event for an angle with the opposite sign.
[0017] The above user interface unit recognizes a downward scroll gesture and generates a corresponding downward scroll event when the Roll angle of the face exceeds a positive (+) threshold Roll angle due to tilting the head toward the right shoulder while the corners of the left and right eyes are horizontal, and when the Roll angle of the face rotates further toward the left shoulder than a negative (-) threshold Roll angle due to tilting the head toward the left shoulder, it recognizes an upward scroll gesture and generates a corresponding upward scroll event.
[0018] The user interface unit recognizes a facial gesture composed of a combination of the estimated open / closed state of both eyes and the duration thereof, using the uppermost and lowermost left eyelid landmark coordinates and the uppermost and lowermost right eyelid landmark coordinates of the user's face; wherein if the difference in the vertical axis coordinate values of the uppermost and lowermost left eyelid landmark coordinates of the user's face is below a specified threshold, the left eye is determined to be in a closed state, and otherwise, it is determined to be in an open state; and if the difference in the vertical axis coordinate values of the uppermost and lowermost right eyelid landmark coordinates of the user's face is below the specified threshold, the right eye is determined to be in a closed state, and otherwise, it is determined to be in an open state; wherein if the facial gesture is in a closed state of both eyes within a first fixed time range and then both eyes open, it is recognized as a left-click gesture and a corresponding left-click event is generated; wherein if the facial gesture is in a closed state of one eyeball within a first fixed time range and then that eyeball opens, it is recognized as a right-click gesture and a corresponding right-click event is generated; and if the facial gesture repeats the left-click gesture twice within a second fixed time range, it is recognized as a double-click gesture and accordingly A corresponding double-click event is generated, and if the cursor position moves after the right-click gesture is maintained for more than a third predetermined time, it is recognized as a drag gesture and a corresponding drag event is generated, and if both eyes are open during the drag event, it is recognized as a drop gesture and a corresponding drop event is generated, and if the face gesture is maintained in a closed state for more than a fourth predetermined time which is longer than the third predetermined time, and both eyes are open, if the face gesture interface is currently in a deactivated state, it is recognized as an interface activation gesture and a corresponding interface activation event is generated to activate the face gesture interface.Otherwise, conversely, if it is in an active state, it is recognized as an interface deactivation gesture, and the face gesture interface is terminated by generating a corresponding interface deactivation event.
[0019] To prevent frequent malfunctions when transitioning from a closed state to an open state, the above-mentioned eye opening / closing determination unit repeatedly performs a process of forcibly correcting and recognizing that both eyes are closed if, when recognizing eye opening / closing, it examines the two most recent frames based on the point in time when both eyes are recognized as open, and if the immediately preceding frame between the current frame (where both eyes are open) and the frame before the previous frame (where both eyes are closed) is recognized as having only one eye open, it recognizes that both eyes are closed.
[0020] The above 3D facial landmark coordinate system is the coordinate system of the 3D facial landmarks of the MediaPipe Face Mesh model.
[0021] A method using a user interface system utilizing facial gesture recognition and facial authentication comprises: (a) a step in which the user interface system identifies facial features of a user granted operation authority among a plurality of users in a series of input image frames and selects a facially authenticated operator therefrom; and (b) a step in which the user interface system allows facial gesture recognition of the selected operator, displays the facial gesture recognition result on a screen when necessary to interact with the operator, and controls an interface to generate a user event based on the facial gesture recognition result and execute a corresponding command.
[0022] The above step (a) includes the step of the user interface system generating a face authentication dataset by sampling a series of person image frames including frontal, side, and tilted face compositions at regular frame intervals; the step of the user interface system generating a face recognition model through a supervised learning process that classifies faces using cosine similarity after extracting and normalizing multidimensional embedding vectors from face regions detected from the face authentication dataset; the step of detecting face regions in a series of input image frames and inputting them into the face recognition model to perform face recognition and output the corresponding face class and recognition confidence; and the step of the user interface system authenticating the user of the face as the operator only if the recognized face class matches the registered operator and the recognition confidence is greater than or equal to a preset authentication threshold.
[0023] Step (a) above involves the user interface system detecting a face area and performing face recognition to authenticate a user with a recognition confidence level higher than a preset level as the operator, wherein if multiple people exist within the input image, the process of recognizing other people is repeated immediately after one person is recognized by removing or ignoring that recognized person from the screen.
[0024] The above step (b) includes: (b-1) a step in which the user interface system detects the movement of the user's face in a series of input image frames and estimates the Pan angle and Tilt angle of the face, and then calculates the horizontal and vertical coordinates of the cursor position on the screen corresponding to each face angle; and (b-2) a step in which the user interface system detects the movement of the face and binoculars through the face landmarks estimated in the input image frames, recognizes a face gesture composed of a combination of the face's Roll angle, the opening and closing state of the binoculars, and the holding time thereof, generates a user event corresponding to the face gesture in association with the cursor position, and controls the interface to execute a command corresponding to the user event.
[0025] The above step (b) further includes a cursor stabilization step in which the user interface system applies adaptive moving average processing based on the cursor movement speed to the horizontal and vertical coordinates of the cursor on the screen to stabilize the fine shaking of the cursor.
[0026] The above step (b) further includes a binocular opening / closing determination step, which, when the user interface system determines the binocular simultaneous opening / closing state, repeatedly checks for discontinuities in the binocular simultaneous opening / closing and corrects the temporary discrepancy in the opening / closing state of the binoculars under specific conditions so that it is unified into either an open or closed state.
[0027] Step (b) above involves, as soon as the face gesture interface is activated, suppressing the user's facial movement and ensuring they look at the center of the screen, accumulating and averaging the nose tip landmark coordinates estimated from video frames received during an initial predetermined time to establish them as the angle calculation reference point of the 3D face landmark coordinate system, obtaining a 3D vector connecting the origin of the 3D face landmark coordinate system to the angle calculation reference point, obtaining a 3D vector connecting the nose tip landmark coordinates estimated after establishing the angle calculation reference point from the origin of the 3D face landmark coordinate system, calculating the horizontal and vertical angles between the two 3D vectors to establish the face's Pan angle and Tilt angle, respectively, and, with the screen width and face's maximum Pan angle pre-set, when the face's Pan angle is within the range [(-)maximum Pan angle, (+)maximum Pan angle], converting the face's Pan angle into a screen horizontal coordinate within the range [0, screen width] using trigonometric functions and proportional equations, and with the screen height and face's maximum Tilt angle pre-set, when the face's Tilt angle When within the range [(-)maximum tilt angle, (+)maximum tilt angle], the tilt angle of the face is converted into a vertical screen coordinate within the range [0, screen vertical size] using trigonometric functions and a proportion.
[0028] The above step (b) includes a process in which the user interface system determines the Roll angle of the face as the angle between the vector connecting the right eye corner landmark coordinate of the user face to the left eye corner landmark coordinate and the horizontal coordinate axis of the 3D face landmark coordinate system, and if the absolute value of the Roll angle of the face is greater than or equal to a preset threshold Roll angle, it determines whether it is a positive (+) angle or a negative (-) angle and generates either a Scroll Down Event or a Scroll Up Event, and generates the other user event for an angle with the opposite sign.
[0029] The above step (b) includes a process in which, when the Roll angle of the face exceeds a positive (+) threshold Roll angle due to the user interface system tilting the head toward the right shoulder while the left and right corners of the eyes are horizontal, it is recognized as a downward scroll gesture and a corresponding downward scroll event is generated, and when the head is tilted toward the left shoulder and the Roll angle of the face rotates further toward the left shoulder than a negative (-) threshold Roll angle, it is recognized as an upward scroll gesture and a corresponding upward scroll event is generated.
[0030] The above step (b) involves the user interface system recognizing a facial gesture composed of a combination of the open / closed state of both eyes and the duration thereof, estimated using the uppermost and lowermost left eyelid landmark coordinates and the uppermost and lowermost right eyelid landmark coordinates of the user's face; wherein if the difference in the vertical axis coordinate values of the uppermost and lowermost left eyelid landmark coordinates of the user's face is below a specified threshold, the left eye is determined to be in a closed state, otherwise it is determined to be in an open state; and if the difference in the vertical axis coordinate values of the uppermost and lowermost right eyelid landmark coordinates of the user's face is below the specified threshold, the right eye is determined to be in a closed state, otherwise it is determined to be in an open state; wherein if the facial gesture is followed by the opening of both eyes after maintaining a closed state of both eyes within a first fixed time range, it is recognized as a left-click gesture and a corresponding left-click event is generated; if the facial gesture is followed by the opening of one eyeball after maintaining a closed state of that eyeball within a first fixed time range, it is recognized as a right-click gesture and a corresponding right-click event is generated; and if the facial gesture repeats the left-click gesture twice within a second fixed time range, a double It recognizes a click gesture and generates a corresponding double-click event; if the cursor position moves after the right-click gesture is maintained for more than a third predetermined time, it recognizes it as a drag gesture and generates a corresponding drag event; if both eyes are open during the drag event, it recognizes it as a drop gesture and generates a corresponding drop event; and if the face gesture is maintained in a closed state for more than a fourth predetermined time which is longer than the third predetermined time, and both eyes are open, if the face gesture interface is currently inactive, it recognizes it as an interface activation gesture and generates a corresponding interface activation event to activate the face gesture interface.Otherwise, if it is in an active state, it includes a process of terminating the face gesture interface by recognizing it as an interface deactivation gesture and generating a corresponding interface deactivation event.
[0031] The above-mentioned binocular opening / closing determination step includes a process to repeatedly perform a process of forcibly correcting the recognition to both eyes as closed in order to prevent frequent malfunctions when transitioning from a binocular closed state to a binocular open state. This is done by examining the two most recent frames based on the point in time when both eyes are recognized as open when recognizing binocular opening / closing, and if the immediately preceding frame between the current frame (where both eyes are open) and the frame before the previous frame (where both eyes are closed) is recognized as having only one eye open, the recognition is performed as if both eyes are closed.
[0032] The above 3D facial landmark coordinate system is the coordinate system of the 3D facial landmarks of the MediaPipe Face Mesh model. Effects of the invention
[0033] According to the present invention, in an industrial or living environment where multiple collaborators exist, a registered operator is identified among multiple people, and a face gesture-based user interface is provided without a separate medium with excellent intuitiveness and user friendliness, thereby having the effect of improving convenience and security. Brief explanation of the drawing
[0034] FIG. 1 is a diagram showing a MediaPipe Face Mesh 3D face landmark of a user interface system using face gesture recognition and face authentication according to one embodiment of the present invention. FIG. 2 is a configuration diagram of a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention. FIG. 3 shows seven landmarks used for determining a face gesture in a user interface system using face gesture recognition and face authentication according to one embodiment of the present invention. FIG. 4 is a drawing showing the Pan angle, Tilt angle, and Roll angle of a face according to one embodiment of the present invention. FIG. 5 is a diagram showing an example of a face pan angle display in a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention. FIG. 6 is a diagram showing an example of a face tilt angle display in a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention. FIG. 7 is a diagram showing the relationship between the face pan angle and the corresponding cursor position on the screen of a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention. FIG. 8 shows the face roll angle of a user interface system using face gesture recognition and face authentication according to one embodiment of the present invention. FIG. 9 is a diagram illustrating the process of improving the malfunction of simultaneous opening and closing of both eyes in a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention. FIG. 10 is a diagram illustrating face recognition of a user interface system using face gesture recognition and face authentication according to an embodiment of the invention. FIG. 11 is a schematic overall flowchart of a method using a user interface system utilizing face gesture recognition and face authentication according to one embodiment of the present invention. FIG. 12 is a diagram illustrating interface control operations through face authentication and gesture recognition of a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention. FIG. 13 is a flowchart of the interface operation of a method using a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention. FIG. 14 is a schematic flowchart of the operation of a user interface system using face gesture recognition and face authentication according to the present embodiment. FIG. 15 is a diagram showing a user scenario simulation scene of a user interface system using face gesture recognition and face authentication according to one embodiment of the present invention. Specific details for implementing the invention
[0035] The specific structural or functional descriptions presented in the embodiments of the present invention are merely illustrative for the purpose of explaining embodiments according to the concept of the present invention, and embodiments according to the concept of the present invention may be implemented in various forms. Furthermore, it should not be interpreted as being limited to the embodiments described in this specification, but should be understood to include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the present invention. Meanwhile, terms such as "first" and / or "second" in the present invention may be used to describe various components, but said components are not limited to said terms. For the purpose of distinguishing one component from other components, for example, within the scope of rights according to the concept of the present invention, a first component may be named a second component, and similarly, a second component may be named a first component.
[0036] A user interface system using face gesture recognition and face authentication according to one embodiment of the present invention proposes a user interface system and method using face authentication and face gesture recognition that recognizes face gestures composed of combinations such as 3D face angles, binocular opening / closing states, and maintenance times using a MediaPipe Face Mesh model, authenticates the user's face during the process of controlling the user interface, and provides a face gesture-based user interface only to authorized operators.
[0037] According to the present embodiment, first, a pre-registered operator is detected within the screen using a face recognition model generated by utilizing InsightFace. Subsequently, it is determined whether the operator is registered to allow face gesture recognition, and the 3D coordinates of 7 face landmarks are obtained using a MediaPipe Face Mesh model. Using these obtained 3D coordinates, Pan, Tilt, Roll angles, and whether both eyes are open or closed are calculated to control user events and cursors, and to display them on the screen. Through the simulation of various user scenarios such as web surfing, document work, and gaming, it was confirmed that, according to the present invention, an average face gesture recognition rate of 98.7% can be obtained.
[0038] Hereinafter, a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention will be described with reference to the attached drawings.
[0039] FIG. 1 is a diagram showing a MediaPipe Face Mesh 3D face landmark of a user interface system using face gesture recognition and face authentication according to one embodiment of the present invention.
[0040] Since there are many technologies available for facial landmark detectors, such as Dlib, InsightFace, Face-api, OpenVINO, OpenSeeFace, and StackML in addition to MediaPipe Face Mesh, one can choose any of these or develop one in-house. However, as Google’s MediaPipe has recently emerged as the center of cross-platform frameworks in the field of gesture recognition, in one embodiment of the present invention, facial gestures are to be recognized using 3D facial landmarks of the MediaPipe Face Mesh model.
[0041] Google's MediaPipe is a cross-platform / machine learning framework specialized for real-time media processing. Among MediaPipe's internal models, the MediaPipe Face Mesh model is a 3D face landmark detection solution that detects face regions in a series of input frames and then estimates 468 3D face landmarks within the detected face region, as shown in Fig. 1. The MediaPipe Face Mesh model consists of a Detector model that detects face regions in a series of input frames and a 3D Face Landmark model that predicts the approximate 3D surface within this face region. Estimating the 3D coordinates of the face across the entire space within the input frames has the advantage of obtaining global depth information. However, it places a burden on real-time processing due to excessive computation. Accordingly, the MediaPipe Face Mesh model first detects the face region in the input frames and then estimates the 3D coordinates within the face region. By reducing unnecessary computations and focusing on 3D landmark estimation, it improves the accuracy of estimating 3D coordinate values within the face region. The MediaPipe Face Mesh model uses a right-handed orthonormal metric 3D coordinate system as the 3D face landmark coordinate system. The estimated coordinate values of each 3D landmark consist of X, Y, and Z axis values, each ranging from [0,1], [0,1], to [-1,1]. The X and Y axis values are normalized to the position of each face landmark within the [0,1] range, based on the image width and height, respectively. The Z axis values represent the depth of the face landmark within the [-1,1] range, with the origin located at the center of the head. The magnitude of this depth is It uses the same scale as the axis, and the smaller the value, the closer the face landmark is to the camera.
[0042] FIG. 2 is a configuration diagram of a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention.
[0043] As illustrated in FIG. 2, a user interface system (10) using face gesture recognition and face authentication according to one embodiment of the present invention includes a face authentication unit (110) and a user interface unit (120).
[0044] The face authentication unit (110) identifies the facial features of a user granted operation authority among a number of users in a series of input image frames, and then selects a face-authenticated operator through this. To perform this function, the face authentication unit (110) includes a face authentication dataset generation module (111), a face data learning module (112), a face recognition module (113), and an operator authentication module (114).
[0045] The face authentication dataset generation module (111) generates a face authentication dataset by sampling a series of person image frames, including front, side, and tilted face compositions, at regular frame intervals.
[0046] The face data learning module (112) generates a face recognition model through a supervised learning process that extracts and normalizes multidimensional (e.g., 512 dimensions) embedding vectors from the face region detected from the face authentication dataset, and then classifies the face using cosine similarity.
[0047] The face authentication dataset is generated by storing frame data (i.e., sampling) at intervals of 5 frames from two 8-second videos of people at 30fps (frames per second), ensuring that there are 100 data points per person. A total of three face angles are used: frontal, profile, and tilted.
[0048] When training for face authentication, the face authentication dataset is not input directly; instead, face regions are detected using InsightFace’s RetinaFace, followed by embedding using InsightFace’s ArcFace. The embedding results in 512-dimensional face feature data, which is saved as a NumPy npy file. The ArcFace utilized in this embodiment extracts embedding features using the DCNN (Deep Convolutional Neural Network) embedding method. After normalizing the extracted embedding features and weights, the cosine similarity is obtained by taking the inner product of these two values, and this value is used to classify face classes. The embedding data is generated with dimensions of (n, 512), where n represents the number of data points.
[0049] Next, the saved npy file is received as input, trained using Keras, and a face recognition model is created. The created face recognition model is converted into an onnx file and saved.
[0050] The face recognition module (113) detects face regions in a series of input image frames and then inputs them into the face recognition model to perform face recognition and outputs the corresponding face class and recognition confidence level.
[0051] In this embodiment, the face data learning module (112) and the face recognition module (113) perform face recognition using the pre-trained buffalo_l model of InsightFace. There are 5 onnx files in the buffalo_l model, among which face detection is performed using RetinaFace and face recognition is performed using ArcFace.
[0052] At this time, RetinaFace finds the locations of the human face region and face feature points through the Box and 5 landmarks of FIG. 10. The face region found in this way is converted into a multidimensional embedding vector (i.e., embedding features or face features) through Arcface. Then, the face recognition module (113) performs face recognition using an onnx file to determine who the recognized face is, and the operator authentication module (114) finally authenticates the face as the face of a registered operator who has the authority to operate the equipment when the recognized face is a registered operator and the recognition reliability (or accuracy) is greater than or equal to a preset authentication threshold (appropriate level 0.95 or strict level 0.97).
[0053] For reference, InsightFace-based face authentication performs face authentication at intervals of 1 / 5 of the frames per second (fps), but in this embodiment, not only the Holistic model of MediaPipe but also the Face Mesh model is used, so it is performed at intervals of 1 / 6 of the frames per second (fps) to reduce the computational load.
[0054] Then, the face recognition module (113) detects face regions in a series of input image frames and inputs them into the face recognition model to perform face recognition and outputs the corresponding face class and recognition confidence level.
[0055] The operator authentication module (114) authenticates the user of the face as the operator only when the recognized face class matches the registered operator and the recognition reliability is greater than or equal to a preset authentication threshold.
[0056] The face authentication unit (110) according to the present embodiment detects a face area and performs face recognition to authenticate a user who is above a preset recognition reliability level as an operator. However, if there are multiple people in the input image, immediately after one person is recognized, the process of recognizing other people is repeated while the recognized person is immediately erased or ignored from the screen.
[0057] The user interface unit (120) allows the recognition of a facial gesture of an operator selected through a facial authentication unit, displays the facial gesture recognition result on a screen when necessary to interact with the operator, and controls the interface to generate a user event based on the facial gesture recognition result and execute a corresponding command. The user interface unit (120) for performing these functions includes an estimation module (121) and a facial gesture recognition module (122).
[0058] FIG. 3 shows seven landmarks used for determining a face gesture in a user interface system using face gesture recognition and face authentication according to one embodiment of the present invention.
[0059] The above estimation module (121) detects a face region in a series of camera input frames using a MediaPipe Face Mesh model, and then estimates the Pan angle and Tilt angle of the face using the 3D nose tip landmark coordinates among the 7 selected face landmarks (nose tip landmark, right eye corner landmark, left eye corner landmark, right uppermost eyelid landmark, right lowermost eyelid landmark, left uppermost eyelid landmark, left lowermost eyelid landmark) among the 3D face landmark coordinates of the face region, and calculates the horizontal and vertical coordinates of the cursor position on the screen corresponding to these angles.
[0060] To describe the operation process of the present embodiment in detail, the estimation module (121) detects a face region in each input frame using the MediaPipe Face Mech model in a series of camera input frames, and then estimates seven landmark coordinates of FIG. 3 selected from the 3D face landmark coordinates of FIG. 1 within the face region. If no face region is detected during this process, the process of receiving the next frame from the camera is repeated. If the seven 3D face landmark coordinates estimated in each frame of the series of camera input sequence are examined and it is determined that a preset face gesture interface enable event has been input, the Pan angle and Tilt angle of the face are calculated using the 3D nose tip landmark coordinates, and then the horizontal and vertical coordinates of the cursor position on the screen corresponding to these angles are converted. Otherwise, the process of receiving the next frame from the camera is repeated. For reference, the 3D face landmarks of the MediaPipe Face Mech model are assigned a total of 468 numbers from 0 to 467. As shown in Fig. 3, the nose tip landmark is number 1, the right outer corner of the eye landmark is number 33, the left outer corner of the eye landmark is number 263, the right uppermost eyelid landmark is number 159, the right lowermost eyelid landmark is number 145, the left uppermost eyelid landmark is number 386, and the left lowermost eyelid landmark is number 374.
[0061] The estimation module (121) further includes a cursor stabilization unit that stabilizes the fine shaking of the cursor by applying adaptive moving average processing according to the cursor movement speed to the horizontal and vertical coordinates of the cursor on the screen.
[0062] And the above-mentioned face gesture recognition module (122) further includes a two-eye opening / closing determination unit that, when determining the state of simultaneous opening / closing of both eyes, repeatedly checks the discontinuity of simultaneous opening / closing of both eyes and corrects the state of temporary discrepancy in opening / closing of both eyes under specific conditions so that it is unified into either an open or closed state.
[0063] According to the estimation module (121) of the user interface unit according to the present embodiment, the Pan direction of the face image is the direction in which the landmark moves along the X-axis in the direction of moving the head left and right, the Tilt direction of the face image is the direction in which the landmark moves along the Y-axis in the direction of moving the head up and down, and the Roll direction of the face image is the direction in which the angle of difference between the two eyes in the Y-axis occurs when performing the action of attaching one ear to the shoulder, and the cursor position is estimated through the 3D coordinates of the landmark.
[0064] The face gesture recognition module (122) recognizes a face gesture from face information consisting of the roll angle of the face estimated using the right eye corner landmark coordinates and the left eye corner landmark coordinates among the 3D face landmark coordinates of the MediaPipe Face Mesh model of the face area, the opening and closing state of both eyes estimated using the left top and bottom eyelid landmark coordinates and the right top and bottom eyelid landmark coordinates, and the holding time thereof, and generates a user event by associating it with the cursor position calculated by the estimation module (121), and controls the interface to execute a command corresponding to the user event.
[0065] The facial gesture recognition is explained as follows.
[0066] The information used for determining facial gestures is the opening and closing status of both eyes and the holding time, as well as the face roll angle, of the current frame and the previous frame. To this end, in this embodiment, the opening and closing status of both eyes and the face roll angle are calculated using the right eye corner landmark, left eye corner landmark, right eye corner landmark, right eye corner landmark, left eye corner landmark, left eye corner landmark, and left eye corner landmark among the seven selected face landmarks (nose tip landmark, right eye corner landmark, left eye corner landmark, right uppermost eyelid landmark, right lowermost eyelid landmark), left uppermost eyelid landmark, and left lowermost eyelid landmark.
[0067] FIG. 4 is a diagram showing the Pan angle, Tilt angle, and Roll angle of a face according to an embodiment of the present invention. As shown in FIG. 4, the Pan angle of the face is a face angle that varies as the head is turned left and right while the torso is the axis of rotation; in this embodiment, the relative change of this angle is converted to correspond to the horizontal coordinate axis of the cursor on the screen. The Tilt angle of the face is a face angle that varies as the head is moved up and down; the relative change of this angle is converted to correspond to the vertical coordinate axis of the cursor on the screen. The Roll angle of the face is a face angle that varies according to the movement of tilting one ear toward the left or right shoulder, and is used when generating an up / down scroll event.
[0068] FIGS. 5 and FIGS. 6 are drawings illustrating examples of face Pan angle and Tilt angle display of a user interface system using face gesture recognition and face authentication according to one embodiment of the present invention.
[0069] The left-right and up-down movement of the cursor position on the screen according to the present embodiment occurs by varying the Pan angle and Tilt angle of the face as shown in FIGS. 5 and 6, respectively. The Pan angle of the face is a horizontal rotation of the face, which moves the cursor in the horizontal direction of the screen. And, the Tilt angle of the face is a vertical rotation of the face, which moves the cursor in the vertical direction of the screen.
[0070] In one embodiment of the present invention, the Pan angle and Tilt angle of the face are calculated using the 3D nose tip landmark coordinates, which is landmark number 1 of the MediaPipe Face Mesh model in FIG. 1. First, as soon as the face gesture interface is activated, the movement of the user's face is suppressed for an initial predetermined time (e.g., 3 seconds), and the nose tip landmark coordinates of the input frames are cumulatively averaged to determine the 'angle calculation reference point' of the 3D face landmark coordinate system. The reason for initializing the angle calculation reference point with this cumulative average value is to stabilize the angle calculation reference point that serves as the standard when calculating the face angle. Subsequently, a 3D vector connecting the origin of the 3D face landmark coordinate system to the angle calculation reference point is obtained, and a 3D vector connecting the nose tip landmark coordinates estimated after determining the angle calculation reference point is obtained. Then, the horizontal and vertical angles between these two 3D vectors are calculated and determined as the Pan angle and Tilt angle of the face, respectively.
[0071] As explained earlier, the origin of the 3D face landmark coordinate system of the original MediaPipe Face Mesh model exists separately. However, for the sake of convenience of explanation, Figures 4 through 8 assume that the origin of the 3D face landmark coordinate system of the MediaPipe Face Mesh model is virtually moved to the center of the 3D head, and that there exists a 3D virtual coordinate system in which the Z-axis of that coordinate system is aligned with a 3D vector connecting the origin of the original MediaPipe Face Mesh model's 3D face landmark coordinate system to the angle calculation reference point. In this case, the vector representing the 3D vector connecting the origin of the MediaPipe Face Mesh model's 3D face landmark coordinate system to the nose tip landmark coordinate is a vector Let's say that. And the point marking the 3D nose tip landmark in that 3D virtual coordinate system is the coordinate Let's assume that. Fig. 5 is a 3D vector Projecting the 3D coordinates P and 3D coordinates onto the ZX plane of the 3D virtual coordinate system to obtain 2D vectors respectively and 2D coordinates This indicates the Pan angle of the face in Fig. 5 because the 3D vector connecting the origin of the 3D face landmark coordinate system to the angle calculation reference point was aligned with the Z-axis of this 3D virtual coordinate system. is a 2D vector using Equation (1) The angle between and the Z-axis can be easily calculated using the arccosine function (inverse cosine function). In Equation (1), is a 2D coordinate It is the value obtained by projecting it onto the Z-axis of that coordinate system, and is a vector It is the magnitude (norm). Fig. 6 is a 3D vector Projecting the 3D coordinates P and 3D coordinates onto the ZY plane of the 3D virtual coordinate system to obtain 2D vectors respectively and 2D coordinates It indicates. Likewise, the tilt angle of the face is using Equation (2) The angle between and the Z-axis can be easily calculated using the arc cosine function. In Equation (2), is a 2D coordinate It is the value obtained by projecting it onto the Z-axis of that coordinate system, and is a vector It is the size of.
[0072] (Mathematical Formula 1)
[0073]
[0074] (Mathematical Formula 2)
[0075]
[0076] Screen width in advance , screen vertical size , maximum pan angle of the face and maximum tilt angle If is determined, the pan angle of the face calculated earlier and Tilt angle Using , the horizontal coordinates of the cursor on the screen through Equation (3) and Equation (4) and vertical coordinates Calculates. Fig. 7 shows the pan angle of the face. Horizontal coordinates on the screen for This represents the process of converting it. In other words, the pan angle of the face. this When within the range, this Pan angle through Equation (3) [0, Horizontal screen coordinates within the range It can be converted into. And the tilt angle of the face this When within the range, this tilt angle through Equation (4) [0, ] Screen vertical coordinates within the range It can be converted into. Therefore, the 'angle calculation reference point' of the 3D face landmark coordinate system described earlier corresponds to being located exactly in the center across the vertical and vertical directions of the screen.
[0077] In equation (3) ... is a hatched area on the left and right sides of the screen area of FIG. 7, which is a marginal zone in the horizontal direction of the screen for ease of operation, and is of Equation (4). ... is the free area in the vertical direction of the screen. And in Equation (3) and Equation (4), respectively and is predetermined and It is determined according to. For example, and are each It is set to this, but if necessary, the Pan and Tilt directions can be set to different angles.
[0078] (Mathematical Formula 3)
[0079]
[0080] (Mathematical Formula 4)
[0081]
[0082] In addition, while a 3D nose tip landmark was used to determine the Pan and Tilt angles of the face in one embodiment of the present invention in the above process, it is a well known fact that there are many facial landmarks that can replace it while producing a similar effect. For example, in addition to the nose tip, there are facial landmarks on the vertical bisectors of the outer corners of the eyes or pupils, facial landmarks on the ridge of the nose, the philtrum, facial landmarks on the vertical bisectors of both corners of the mouth or the central axis of the lips, and chin landmarks. It is evident that using any one or more combinations of these alternative facial landmarks is within the scope of the present invention.
[0083] Meanwhile, through a user scenario simulation of a face gesture interface according to one embodiment of the present invention, it was found that unstable micro-tremors occur when moving the cursor. It was found that this phenomenon becomes more severe for inexperienced users and accelerates as tension increases for precise cursor operation. To solve this problem, in one embodiment of the present invention, the estimation module (121) further includes a cursor stabilization unit that stabilizes the micro-tremors of the cursor by applying adaptive moving average processing according to the cursor movement speed to the horizontal and vertical coordinates of the cursor on the screen, thereby suppressing the micro-tremors of the cursor and promoting the refinement and stabilization of cursor operation.
[0084] In one embodiment of the present invention, adaptive moving average processing is applied to the current cursor position throughout the entire process of calculating the cursor position to mitigate fine shaking. Accordingly, all cursor positions used as user commands are values to which accumulated moving average processing has been applied, except for the first three or four frames. This shall be referred to as the 'average cursor position'. In one embodiment of the present invention, as shown in Equation (5), the average cursor positions of the most recent three or four frames, including the cursor position of the current frame, are summed according to the speed of cursor movement, and the resulting average value is used as the actual cursor position of the current frame.
[0085] (Mathematical Formula 5)
[0086]
[0087] For example, the current cursor position that has not yet undergone averaging. w and previous average cursor position The pixel distance between When saying, pixel distance of less than 10 If so, the average value of the average cursor positions of the four most recent frames, including the current frame's cursor position, is the current frame's actual cursor position. Use as. Otherwise, the cursor position in the current frame. The average value of the average cursor positions of the most recent 3 frames, including the actual cursor position It is used as.
[0088] The number of frames used when calculating the average is the movement distance The reason for the variation depending on the frame rate is that increasing the number of frames used improves the fine shaking of the cursor, but causes a problem where the cursor's movement speed is delayed, resulting in reduced responsiveness. In other words, when fine movement of less than 10 pixels is required, the last 4 frames are used to focus more on improving shaking than movement speed, and when movement of 10 pixels or more is required, the last 3 frames are used to take movement speed into account more than shaking improvement.
[0089] In the above average processing steps, increasing the number of frames improves the stabilization of fine cursor jitter but degrades cursor movement responsiveness; therefore, this can be explained as variable depending on user preference. Additionally, the preset cursor distance threshold used to determine the speed of cursor movement. In the case of [this example], an example set to 10 pixels was described, but since this is a value obtained experimentally, it is a value that can be changed according to user preference.
[0090] Next, the user interface unit (120) of the present embodiment estimates the roll angle of the face using the right eye corner landmark coordinates and the left eye corner landmark coordinates of FIG. 3 among the 3D face landmark coordinates of the MediaPipe Face Mesh model of the face area. Then, it recognizes a face gesture from the face information composed of the binocular opening / closing state and the maintenance time thereof, which is estimated using the left uppermost and lowermost eyelid landmark coordinates and the right uppermost and lowermost eyelid landmark coordinates of FIG. 3, generates a user event related to the cursor position calculated by the estimation module (121), and controls the interface to execute a command corresponding to this user event.
[0091] The face information used for face gesture recognition in the user interface section (120) is the open / closed state of both eyes and the duration of the maintenance of the state and the face roll angle of the current frame and past frames. To this end, the open / closed state of both eyes and the face roll angle are calculated using a total of six landmarks consisting of the right eye corner landmark coordinates, the left eye corner landmark coordinates, the left uppermost and lowermost eyelid landmark coordinates, and the right uppermost and lowermost eyelid landmark coordinates of FIG. 3.
[0092] First, referring to Fig. 8, the process of calculating the face roll angle is explained as follows. The vector representing the 3D vector connecting the 3D right outer corner landmark coordinates of the MediaPipe Face Mesh model to the 3D left outer corner landmark coordinates, displayed in the previously described 3D virtual coordinate system, is a vector Let's assume that. Also, let 3D coordinates R and 3D coordinates L be the points marking the 3D right outer corner landmark coordinates and 3D left outer corner landmark coordinates in the 3D virtual coordinate system, respectively. Fig. 8 is a 3D vector , projecting 3D coordinates R and L onto the XY plane of this 3D virtual coordinate system to create a 2D vector , 2D coordinates and 2D coordinates It represents. Of course, here, the vector is a 2D coordinate 2D coordinates in It is a 2D vector connecting.
[0093] For reference, Figure 8 is illustrated based on the condition that no horizontal flip is applied. If a horizontal flip is applied to the face visible on the screen to create a mirror effect, the positions of the left and right eyes will be changed.
[0094] The roll angle of the face as the head is tilted toward the left and right shoulders is variable, and this angle is the vector in Fig. 8 As the angle formed by and the X-axis, it can be easily calculated using the arctangent function (inverse tangent function) as in Equation (6). In Equation (6), is a vector It is the X component of, and is a vector This represents the Y component of.
[0095] (Mathematical Formula 6)
[0096]
[0097] Afterwards, the roll angle of the face The absolute value of the specified critical Roll angle ( Determine if it is ) or more and trigger the vertical scroll event in Table 1. For example, with the corners of the left and right eyes horizontal, tilt the head toward the right shoulder in the Roll direction to determine the Roll angle of the face. go If this occurs, a Scroll Down Event is triggered. Then, tilt the head toward the left shoulder in the Roll direction, and the face's Roll angle go If rotated further toward the left shoulder, it triggers a Scroll Up Event. Of course, the roll angle of the face. It is a well-known fact that for the negative (-) and positive (+) directions, downward scroll events and upward scroll events can be generated in the opposite way.
[0098] The roll angle of the face in the above process It measures a positive (+) angle in the Y-axis direction relative to the X-axis according to the principle of the right-handed coordinate system. As explained earlier, if a horizontal flip is applied to the screen, the positions of the left and right eyes change, and the positive (+) direction of the X-axis also changes to the exact opposite.
[0099] In addition, while 3D left and right eye corner landmarks were used to determine the Roll angle of the face in one embodiment of the present invention in the above process, it is a well known fact that there are many face landmarks that can replace them while producing a similar effect. For example, in addition to the left and right eye corners, there are pairs (or pairs) of face landmarks that provide left-right horizontal symmetry in the face, such as the left and right pupils, left and right corners of the mouth, left and right ears, left and right cheekbones, and left and right nostrils. It is evident that using any one or more combinations of these pairs of face landmarks that provide left-right horizontal symmetry is within the scope of the present invention.
[0100] Next, a facial gesture composed of a combination of the estimated open / closed state of both eyes and their maintenance time using the top-left and bottom-right eyelid landmark coordinates of the user's face, wherein if the difference in the vertical axis coordinate values of the top-left and bottom-left eyelid landmark coordinates of the user's face is less than or equal to a specified threshold, the left eye is determined to be in a closed state, otherwise it is determined to be in an open state, and if the difference in the vertical axis coordinate values of the top-right and bottom-right eyelid landmark coordinates of the user's face is less than or equal to the aforementioned specified threshold, the right eye is determined to be in a closed state, otherwise it is determined to be in an open state.
[0101] In other words, a total of four landmarks—the coordinates of the top and bottom left eyelids and the top and bottom right eyelids—are used to determine whether both eyes are open or closed. For example, if the difference between the Y-axis coordinate values of the top and bottom left eyelids is less than or equal to a specified threshold, the left eye is determined to be in a closed state (i.e., closed, closed), and otherwise, the left eye is determined to be in an open state (i.e., open, floating). Similarly, if the difference between the Y-axis coordinate values of the top and bottom right eyelids is less than or equal to a specified threshold, the right eye is determined to be in a closed state, and otherwise, the right eye is determined to be in an open state.
[0102] A face gesture composed of a combination of the open / closed state of both eyes (left and right eyes) and the duration thereof is recognized, and the Left Click Event, Right Click Event, Double Click Event, Drag Event, Drop Event, Interface Enable Event, and Interface Disable Event of Table 1 are generated. The types of face gestures and corresponding user events of the face gesture interface according to an embodiment of the present invention are as shown in Table 1.
[0103] User's facial gestures Time and operating conditions User Events After maintaining a closed state of both eyes within a certain time range, both eyes open More than 0.2 seconds less than 3 seconds Left click After maintaining a closed state of one eyeball within a certain time range, the eyeball opens. More than 0.2 seconds less than 3 seconds Right-click Repeat the left-click gesture twice within a certain period of time. Continuous repetition interval less than 2 seconds Double click After holding the right-click gesture for a certain period of time, move the cursor. Over 1 second Drag Opening of both eyes during a drag event - drop Roll direction tilt on the right shoulder Critical Roll Angle: +35° Down scroll Roll direction tilt on the left shoulder Critical Roll Angle: -35° Scroll up After maintaining a closed state of both eyes for a certain period of time, both eyes open Exceeding 3 seconds Enable / Disable Interface
[0104] Referring to Table 1, if the face gesture is maintained in a closed state for both eyes within a first fixed time range (e.g., greater than 0.2 seconds and less than 3 seconds) and then both eyes are opened, it is recognized as a left-click gesture and a corresponding left-click event is generated, and if the face gesture is maintained in a closed state for one eye within a first fixed time range (greater than 0.2 seconds and less than 3 seconds) and then that eye is opened, it is recognized as a right-click gesture and a corresponding right-click event is generated.
[0105] Then, if the face gesture repeats the left-click gesture twice within a second set of times (continuous repetition interval of less than 2 seconds), it is recognized as a double-click gesture and a corresponding double-click event is generated; if the right-click gesture is maintained for more than a third set of times (1 second) and the cursor position is moved, it is recognized as a drag gesture and a corresponding drag event is generated; if both eyes are open during the drag event, it is recognized as a drop gesture and a corresponding drop event is generated; if both eyes are closed for more than a fourth set of times (3 seconds) which is longer than the third set of times (1 second) and both eyes are open, if the face gesture interface is currently in an inactive state, it is recognized as an interface activation gesture and an interface activation event is generated to activate the face gesture interface; otherwise, if it is in an active state, it is recognized as an interface deactivation gesture and an interface deactivation event is generated to terminate the face gesture interface.
[0106] Hereinafter, the process for improving the malfunction of simultaneous opening and closing of both eyes is described with reference to FIG. 9. FIG. 9 is a diagram illustrating the process for improving the malfunction of simultaneous opening and closing of both eyes in a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention.
[0107] During the experimental process of a face gesture interface according to one embodiment of the present invention, it was discovered that when both eyes are opened and closed simultaneously, cases often occur where the left and right eyeballs open and close with slight discrepancies due to individual differences, as shown in the test conditions of Fig. 9. In other words, when both eyes are opened and closed simultaneously, a problem arises in which the recognition rate of both eye opening and closing gestures is reduced because the left and right eyeballs open and close with slight discrepancies during the transition process.
[0108] To solve this problem, the user interface unit (120) of one embodiment of the present invention further includes a two-eye opening / closing determination unit that, when determining the two-eye simultaneous opening / closing state, repeatedly checks the discontinuity of the two-eye simultaneous opening / closing and modifies the temporary opening / closing discrepancy state of the two eyes under specific conditions so that it is unified into either an open or closed state.
[0109] Referring to FIG. 9, the binocular opening / closing determination unit of one embodiment of the present invention, since malfunctions occur more frequently when transitioning from a binocular closed state to a binocular open state, repeatedly performs a process of forcibly correcting and recognizing that both eyes are closed when recognizing binocular opening / closing, based on the point in time when both eyes are recognized as being open, by examining the two most recent frames and, between the current frame (frame t) where both eyes are open and the frame two prior to the previous frame (frame t-2) where both eyes are closed, if the immediately preceding frame (frame t-1) is recognized as having only one eyeball open, then both eyes are recognized as being closed. In other words, by investigating the discontinuity of simultaneous opening / closing of both eyes and resolving the problem of temporary opening / closing inconsistency of both eyes, the malfunction of the corresponding command is improved. As a result, it was confirmed that the recognition rate of the binocular opening / closing gesture increased significantly enough to offset the computational burden. Through this, it was possible to significantly reduce malfunctions in left-click, right-click, double-left-click, and interface activation / deactivation events using the binocular opening / closing mode. In the above process of improving the malfunction of simultaneous opening and closing of both eyes, if the previous frame (frame t-1) is recognized as having only one eye open, it is acceptable to forcibly correct and recognize it as having both eyes open.
[0110] FIG. 10 is a diagram illustrating face recognition of a user interface system using face gesture recognition and face authentication according to an embodiment of the invention.
[0111] InsightFace is an integrated library for 2D and 3D deep face analysis. First, the face authentication dataset is embeddinged using InsightFace’s RetinaFace and ArcFace, and then a face recognition model is created by training it via Keras. Real-time video is loaded into OpenCV and embedding it in the same manner as the face authentication dataset to transform facial features into 512 dimensions. These facial features are input into the face recognition model to determine the identity of the recognized face. At this stage, if the user is a registered operator, facial gesture recognition is permitted.
[0112] FIG. 11 is a schematic overall flowchart of a method using a user interface system utilizing face gesture recognition and face authentication according to one embodiment of the present invention.
[0113] As illustrated in FIG. 11, a user interface system using facial gesture recognition and facial authentication (hereinafter referred to as the "user interface system") identifies facial features of a user who is granted operating authority among a number of users in a series of input image frames, and then selects an operator whose face is authenticated through this (a).
[0114] Next, the user interface system allows the recognition of a selected operator's face gesture, displays the result of the face gesture recognition on the screen when necessary to interact with the operator, and controls the interface to generate a user event based on the result of the face gesture recognition and execute a corresponding command (b).
[0115] In step (a) of Fig. 11, the user interface system generates a face authentication dataset by sampling a series of person image frames, including front, side, and tilted face compositions, at regular frame intervals.
[0116] Next, the user interface system generates a face recognition model through a supervised learning process that extracts multidimensional (e.g., 512-dimensional) embedding vectors (i.e., face features) from face regions detected from a face authentication dataset, normalizes them, and then classifies faces using cosine similarity.
[0117] Next, the user interface system extracts facial feature points from the face region and performs face recognition using an embedding generated by Arcface. Then, when the recognized face is a registered operator and the recognition confidence (or accuracy) is 0.95 or higher, it finally authenticates the face as the face of a registered operator who has the authority to operate the equipment.
[0118] In addition, in step (a) of the method using a user interface system utilizing face gesture recognition and face authentication according to the present embodiment, the user interface system detects a face area and performs face recognition to authenticate a user with a recognition confidence level greater than or equal to a preset level as an operator, wherein if multiple people exist within the input image, the process of recognizing other people is repeated immediately after one person is recognized by immediately erasing or ignoring the recognized person from the screen.
[0119] In addition, in step (b) of the method using a user interface system using face gesture recognition and face authentication according to the present embodiment, the user interface system obtains Pan, Tilt, and Roll angles using the landmark 3D coordinates of the operator's face allowed in the image, detects the movement of the face image to estimate the cursor position, and controls user events including cursor movement, click, drag, and scroll on the GUI (Graphic User Interface) according to the Pan, Tilt, and Roll angles and whether both eyes are open or closed using the landmark 3D coordinates of the operator's face allowed in the image.
[0120] In this step (b), the user interface system estimates the cursor position through the 3D coordinates of the landmark by calculating the angle of difference between the two eyes in the Y-axis, such that the landmark moves along the X-axis in the direction of moving the head left and right in the Pan direction of the face image, the landmark moves along the Y-axis in the direction of moving the head up and down in the Tilt direction of the face image, and the angle of difference between the two eyes in the Roll direction of the face image when performing the action of trying to attach one ear to the shoulder.
[0121] And the user interface system of step (b) determines facial gestures based on the opening and closing of both eyes and the holding time and the face rotation (Roll) angle of the current and previous frames of the input video, and when determining the face rotation (Roll) angle and the opening and closing of both eyes using a landmark consisting of two points of the upper and lower eyelids that are in contact with the pupil for each eyeball of both eyes, only the Y-axis value among the 3D coordinate values of each point of the cursor is used, and when the difference between the Y-axis values of the upper and lower eyelids of the left and right eyes is less than or equal to a specified threshold, the corresponding eyeball is determined to be closed, and based on the opening and closing of both eyes, left-click, double-left-click, right-click, drag, drop, interface activation and deactivation events are generated.
[0122] FIG. 12 is a diagram illustrating interface control operations through face authentication and gesture recognition of a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention.
[0123] First, a user interface system using facial gesture recognition and facial authentication receives video input from a camera or external source (1).
[0124] Next, detect whether there is a person in the input image (2).
[0125] Afterwards, check whether there are multiple people in the input video to determine if they are registered operators (3-1).
[0126] Check the face of the registered operator among several people (4).
[0127] Afterwards, the Mediapipe face mesh allows for the inference (estimation) of 7 face landmarks, and the Mediapipe Face Mesh is used to infer and detect 7 face landmarks (5).
[0128] Next, facial gestures are recognized based on the inferred facial landmarks (6).
[0129] Displays the recognized gesture on the screen and controls the user interface according to the gesture (7).
[0130] Next, determine the conditions for the system to be terminated (8).
[0131] In this process, through steps ranging from camera video input to face authentication, face gesture recognition, and user interface control based on gestures, the user interface is controlled to generate user events based on the registered user's face and facial gestures, thereby executing corresponding commands.
[0132] FIG. 13 is a flowchart of the interface operation of a method using a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention.
[0133] As illustrated in FIG. 13, a user interface system using face gesture recognition and face authentication according to the present embodiment (hereinafter referred to as the "system") estimates the Pan angle and Tilt angle of the face in the face region detected in a series of camera input frames, and then calculates the horizontal coordinates and vertical coordinates of the cursor position on the screen corresponding to each face angle (a).
[0134] Next, the system recognizes a facial gesture from facial information consisting of the estimated face roll angle, binocular opening / closing state, and the holding time thereof in the face region, generates a user event related to the cursor position, and controls the interface to execute a command corresponding to the user event (b).
[0135] In step (a) of FIG. 13, the system estimates the Pan angle and Tilt angle of the user's face in a series of input frames for an initial predetermined time as soon as the face gesture interface mode is activated. At this time, the system determines the angle calculation reference point of the 3D face landmark coordinate system by accumulating and averaging the nose tip landmark coordinates of the input frames while the user suppresses facial movement and looks at the center of the screen. Subsequently, a 3D vector is obtained connecting the origin of the 3D face landmark coordinate system to the angle calculation reference point, and a 3D vector is obtained connecting the nose tip landmark coordinates estimated after determining the angle calculation reference point from the origin of the 3D face landmark coordinate system. Then, the horizontal and vertical angles between these two 3D vectors are calculated and determined as the Pan angle and Tilt angle of the face, respectively. In the process, the origin of the 3D face landmark coordinate system is virtually moved to the center of the 3D head, and assuming there is a 3D virtual coordinate system in which the Z-axis of the coordinate system is aligned with a 3D vector connecting the origin of the 3D face landmark coordinate system of the original MediaPipe Face Mesh model to an angle calculation reference point, the angle between the Z-axis and the 2D vector obtained by projecting the 3D virtual coordinate system onto the ZX plane with the 3D vector connecting the origin of the 3D face landmark coordinate system to the 3D nose tip landmark coordinate is defined as the Pan angle of the face. Then, the angle between the Z-axis and the 2D vector obtained by projecting the 3D virtual coordinate system onto the ZY plane with the 3D vector connecting the origin of the 3D face landmark coordinate system to the 3D nose tip landmark coordinate is defined as the Tilt angle of the face.
[0136] Additionally, since the Pan angle of the face varies as the head is turned left and right while the torso acts as the axis of rotation, the system converts this relative change in angle to correspond to the horizontal coordinate axis of the cursor on the screen. Similarly, since the Tilt angle of the face varies as the head moves up and down, the system converts this relative change in angle to correspond to the vertical coordinate axis of the cursor on the screen. Subsequently, adaptive moving average processing based on the cursor movement speed is applied to the aforementioned horizontal and vertical coordinates of the cursor on the screen to stabilize the fine shaking of the cursor.
[0137] Step (b) of FIG. 13 estimates the roll angle of the face using the right eye corner landmark coordinates and the left eye corner landmark coordinates of FIG. 3 among the 3D face landmark coordinates of the MediaPipe Face Mesh model of the face region. In this process, the angle between the X-axis and the 2D vector obtained by projecting the 3D vector connecting the right eye corner landmark coordinates to the left eye corner landmark coordinates onto the XY plane of a 3D virtual coordinate system is determined as the roll angle of the face. Then, the interface is controlled to recognize a face gesture from the face information composed of the binocular opening / closing state and the holding time thereof, estimated using the left uppermost and lowermost eyelid landmark coordinates and the right uppermost and lowermost eyelid landmark coordinates, generate a user event related to the cursor position, and execute a command corresponding to this user event.
[0138] Additionally, in step (b), the system estimates the face roll angle by projecting a 3D vector connecting the right outer corner landmark coordinates to the left outer corner landmark coordinates onto the XY plane of a 3D virtual coordinate system. Subsequently, the absolute value of the face roll angle is a specified threshold roll angle (e.g., Determines if it is ) or more and triggers an up / down scroll event. With the corners of the left and right eyes level, tilt the head toward the right shoulder in the Roll direction so that the face's Roll angle is If this occurs, a downward scroll event is triggered. Then, tilt the head toward the left shoulder in the Roll direction, the Roll angle of the face. go If rotated further toward the left shoulder, it triggers an upward scroll event.
[0139] And in step (b), the system determines the open / closed state of both eyes using a total of four landmarks: the top left and bottom left eyelid landmark coordinates and the top right and bottom right eyelid landmark coordinates. If the difference between the Y-axis coordinate values of the top left and bottom left eyelid landmark coordinates is less than or equal to a specified threshold, the left eye is determined to be in a closed state; otherwise, the left eye is determined to be in an open state. Similarly, if the difference between the Y-axis coordinate values of the top right and bottom right eyelid landmark coordinates is less than or equal to a specified threshold, the right eye is determined to be in a closed state; otherwise, the right eye is determined to be in an open state. A facial gesture composed of a combination of the open / closed state of both eyes (left and right eyes) and their maintenance time is recognized, and the left-click event, right-click event, double-click event, drag event, drop event, interface activation event, and interface deactivation event of Table 1 are generated.
[0140] And in step (b), when the system determines the simultaneous opening and closing state of both eyes, it repeatedly checks for discontinuities in the simultaneous opening and closing of both eyes and corrects the temporary discrepancy in the opening and closing state of both eyes under specific conditions so that it is unified into either an open or closed state. For example, when recognizing the opening and closing of both eyes, based on the point in time when both eyes are recognized as open, the system examines the two most recent frames. If the immediately preceding frame (frame t-1) between the current frame (frame t) where both eyes are open and the frame before the previous frame (frame t-2) where both eyes are closed is recognized as having only one eye open, the system repeatedly performs a process of forcibly correcting the recognition to recognize both eyes as closed. Through this, malfunctions occurring during left-click, right-click, double-left-click, and interface activation and deactivation events using the binocular opening and closing mode can be significantly reduced.
[0141] Hereinafter, a face gesture-based user interface system according to one embodiment of the present invention was designed and implemented using Python based on the object-oriented programming paradigm. Table 2 shows the program file names and dependency relationships of the face gesture-based user interface system implemented in this embodiment.
[0142] File A is the program's entry point and contains code that performs tasks corresponding to the top layer. File B is the implementation of the Utilized_face_mesh class, which inherits the MediaPipe Face Mesh model and modifies it according to requirements. It implements the Face_gesture_processor class, which handles face gestures, using six sub-dependency modules from E to J. File D is the implementation of the Painter class, which displays the cursor position on the screen and executes commands corresponding to the user event.
[0143] Nickname filename Subordinate dependency module A main.py B, C, D B utilized_face_mesh.py MediaPipe Face Mesh C face_gesture_processor.py E, F, G, H, I, J D painter.py - E face_info_per_frame.py - F face_angle_stabilizer.py E G cursor_position_calculator.py - H buffer_for_face_info.py E I buffer_for_cursor_position.py - J gesture_producer.py H, I
[0144] File E implements the Face_info_per_frame class, which calculates and manages face information such as the Pan angle, Tilt angle, Roll angle, binocular opening / closing state, and their duration for each frame. File F refers to the implementation of the Face_angle_stabilizer class, which stabilizes the Pan and Tilt angles of the face by summing and averaging the nose tip landmark coordinates of frames over a certain period (e.g., 3 seconds) immediately after activating the face interface to calculate stable angles, and using this averaged value as the reference point for angle calculation in the 3D face landmark coordinate system. File G implements the Cursor_position_calculator class, which calculates variable cursor positions when the Pan and Tilt angles of the face change across a sequence of input frames, and then stabilizes these cursor positions by applying an adaptive moving average to mitigate micro-jittering. File H implements the Buffer_for_face_info class to store frame-specific face information (Face_info_per_frame instances) within a limited size. The buffer is limited to a fixed size, and if new frames exceeding this limit are added, they are overwritten sequentially starting from the oldest frames. This ensures that face information corresponding to a fixed number of the most recent frames is maintained and available for use. File I implements the Buffer_for_cursor_position class, which stores the cursor position per frame within the limited size. File J implements the Gesture_producer class, which recognizes facial gestures from the face information stored in this buffer in H and I and triggers the corresponding user events.
[0145] A face gesture interface program according to one embodiment of the present invention was developed in an Anaconda virtual environment on the MacOS 13.3.1 (Ventura) operating system and utilizes Python 3.7.12, Numpy 1.21.6, OpenCV 4.7.0, and MediaPipe 0.9.0.1. The hardware development environment is a Macbook Air M1 with 16GB of memory, a CPU with a total of 8 cores consisting of 4 Apple Firestorm cores and 4 Apple Icestorm cores, and an Apple G13G GPU with 2.6 TFlops of floating-point operation (FP32) performance, and the monitor resolution is 1,440×900. As a result of performing a computer simulation as shown in FIG. 12 in this development environment, the frame rate per second was approximately 30 frames, which was sufficient for real-time processing.
[0146] In addition, to evaluate the usability and recognition performance of the face gesture interface according to one embodiment of the present invention in a Windows 11 environment, a computer simulation of one embodiment of the present invention was also performed using an NVIDIA RTX 3070 D6 (8GB) GPU, a LogiTech C920 HD Pro webcam, and Python 3.8 / MediaPipe 0.9.1.0 / OpenCV4.6.0 / CUDA 11.4 / cuDNN 8.2.2 in a Windows 11 Pro environment with an Intel Core i7-12700 (2.1GHz) CPU and DDR5 32GB RAM desktop.
[0147] FIG. 14 is a schematic flowchart of the operation of a user interface system using face gesture recognition and face authentication according to the present embodiment.
[0148] Import necessary modules and libraries step (1): Import necessary modules and libraries by importing libraries such as Python 3.7.12, Numpy 1.21.6, OpenCV 4.7.0, and MediaPipe 0.9.0.1.
[0149] Step of setting hyperparameters and creating instances (2): Various hyperparameters required for the operation of the MediaPipe Face Mesh model are set. Also, the width and height pixel sizes of the screen are set, and the maximum Pan angle and maximum Tilt angle are set. Afterwards, instances to be used are created. At this time, instances for loading the MediaPipe Face Mesh model, face region detection, 3D face landmark estimation, and face gesture recognition are created.
[0150] Camera input step (3): A series of input frames are received from the built-in camera of the Macbook Air M1 or from an external webcam.
[0151] Face presence determination step (4): In a series of camera input frames, the MediaPipe Face Mech model is used to detect a face region in each input frame and determine whether a face is present. If no face region is detected in this face determination step, the process of receiving the next frame from the camera is repeated.
[0152] Estimation of 3D face landmarks using the MediaPipe Face Mesh model (5): Seven selected landmark coordinates (nose tip landmark, right eye corner landmark, left eye corner landmark, right uppermost eyelid landmark, right lowermost eyelid landmark, left uppermost eyelid landmark, left lowermost eyelid landmark) within the detected face area are estimated using the MediaPipe Face Mesh model.
[0153] Step for estimating Pan angle and Tilt angle of face (6): When the 7 3D face landmark coordinates estimated in each frame of a series of camera input sequences are examined and it is determined that a preset face gesture interface activation event has been input, the Pan angle and Tilt angle of the face are estimated using the 3D nose tip landmark coordinates among the 7 3D face landmark coordinates.
[0154] To elaborate further on the process of estimating the Pan and Tilt angles, first, as soon as the face gesture interface mode is activated, the user is instructed to suppress facial movement and look at the center of the screen for an initial predetermined period. The nose tip landmark coordinates of the input frames are then summed and averaged to establish the angle calculation reference point of the 3D face landmark coordinate system. The reason for initializing the angle calculation reference point with this average value is to stabilize the reference point that serves as the standard when calculating the face angle. Subsequently, a 3D vector connecting the origin of the 3D face landmark coordinate system to the angle calculation reference point is calculated. Then, a 3D vector connecting the origin of the 3D face landmark coordinate system to the nose tip landmark coordinates input after this initialization is calculated. Finally, the horizontal and vertical angles between these two 3D vectors are calculated and determined as the face's Pan and Tilt angles.
[0155] In the process, the origin of the 3D face landmark coordinate system is virtually moved to the center of the 3D head, and assuming there is a 3D virtual coordinate system in which the Z-axis of the coordinate system is aligned with a 3D vector connecting the origin of the 3D face landmark coordinate system of the original MediaPipe Face Mesh model to an angle calculation reference point, the angle between the Z-axis and the 2D vector obtained by projecting the 3D virtual coordinate system onto the ZX plane with the 3D vector connecting the origin of the 3D face landmark coordinate system to the 3D nose tip landmark coordinate is defined as the Pan angle of the face. Then, the angle between the Z-axis and the 2D vector obtained by projecting the 3D virtual coordinate system onto the ZY plane with the 3D vector connecting the origin of the 3D face landmark coordinate system to the 3D nose tip landmark coordinate is defined as the Tilt angle of the face.
[0156] Cursor position calculation and cursor stabilization step (7): Since the Pan angle of the face is a face angle that varies as the head is turned left and right while the body is the axis of rotation, the relative change of this angle is converted to be variable corresponding to the horizontal coordinate axis of the cursor on the screen, and since the Tilt angle of the face is a face angle that varies as the head is moved up and down, the relative change of this angle is converted to be variable corresponding to the vertical coordinate axis of the cursor on the screen. Afterwards, adaptive moving average processing according to the cursor movement speed is applied to the horizontal and vertical coordinates of the cursor on the screen to stabilize the fine shaking of the cursor.
[0157] Step for estimating the Roll angle of the face (8): The Roll angle of the face is estimated using the right eye corner landmark coordinate and the left eye corner landmark coordinate among the 3D face landmark coordinates of the MediaPipe Face Mesh model of the face area. In this process, the angle between the X-axis and the 2D vector obtained by projecting the 3D vector connecting the right eye corner landmark coordinate and the left eye corner landmark coordinate onto the XY plane of the 3D virtual coordinate system is determined as the Roll angle of the face.
[0158] Step (9) for recognizing the opening and closing of both eyes and measuring the duration thereof: Among the seven 3D face landmarks, the coordinates of the upper left and lower left eyelid landmarks and the upper right and lower right eyelid landmarks are used to recognize the estimated opening and closing state of both eyes and to measure the duration thereof. In the process of recognizing the opening and closing of both eyes, if the difference between the Y-axis coordinate values of the upper left and lower left eyelid landmarks is less than or equal to a specified threshold, the left eye is determined to be in a closed state, and if not, the left eye is determined to be in an open state. Similarly, if the difference between the Y-axis coordinate values of the upper right and lower right eyelid landmarks is less than or equal to a specified threshold, the right eye is determined to be in a closed state, and if not, the right eye is determined to be in an open state. After recognizing the opening and closing state of the right eye and the left eye separately on a frame-by-frame basis, the duration thereof is measured using the frame per second.
[0159] Buffer storage step (10) of cursor position and face information: Face information consisting of the 3D face angles of the estimated face's Pan angle, Tilt angle and Roll angle, cursor position, binocular opening / closing state and the holding time thereof is stored in a buffer.
[0160] User event generation step for face gesture response (11): A face gesture composed of a combination of face information stored in the buffer is recognized, and a user event related to the cursor position is generated. A face gesture composed of a combination of the open / closed state of both eyes (left and right eyes) and the duration thereof is recognized, and the left-click event, right-click event, double-click event, drag event, drop event, interface activation event, and interface deactivation event of Table 1 are generated. Then, when determining the simultaneous open / closed state of both eyes, the discontinuity of the simultaneous open / closed state of both eyes is repeatedly checked, and under specific conditions, the temporary open / closed inconsistency state of both eyes is corrected so that it is unified into either an open or closed state. For example, to prevent frequent malfunctions during the transition from a closed eye state to an open eye state, when detecting eye opening / closing, the system repeatedly examines the two most recent frames based on the point in time when both eyes are recognized as open. If the immediately preceding frame (frame t-1) between the current frame (frame t) where both eyes are open and the frame before the previous one (frame t-2) where both eyes are closed is recognized as having only one eye open, the system forcibly corrects the recognition to recognize both eyes as closed. This process significantly reduces malfunctions associated with left-clicks, right-clicks, double-left-clicks, and interface activation / deactivation events that utilize the open / closed eye mode.
[0161] Then, it determines whether the absolute value of the face's roll angle is greater than or equal to a specified threshold roll angle (e.g., 35°) and triggers an up / down scroll event. With the corners of the left and right eyes level, tilting the head toward the right shoulder in the roll direction, the face's roll angle If this occurs, a downward scroll event is triggered. Then, tilt the head toward the left shoulder in the Roll direction, the Roll angle of the face. go If rotated further toward the left shoulder, it triggers an upward scroll event.
[0162] Interface control and screen display step (12): The interface is controlled to execute a command corresponding to the user event, and if necessary, face information and related action information are displayed on the manager screen. At this time, the information is not limited to visual display and can be displayed directly or indirectly through various expression media such as voice, sound, vibration, and lighting depending on the application field, target, or situation.
[0163] For example, FIG. 15 illustrates an administrator screen of an embodiment of the present invention, depicting a situation where the cursor is operated without using hands, and shows input frames flipped horizontally for ease of operation. FIG. 15 shows a user face within the administrator screen with a 3D face landmark-based face mesh applied, where the left eye and left eyebrow are displayed as a green mesh, and the right eye and right eyebrow are displayed as a red mesh. It also illustrates 3D face angle information displayed in the upper left corner of the administrator screen, consisting of the left-eye state and right-eye state in yellow, and the face's Pan angle, Tilt angle, and Roll angle in green. In this case, if the left-eye state is False, it means the eyeball is closed, and if it is True, it means the eyeball is open.
[0164] Termination determination step (13): While the face gesture interface according to the present invention is operating in an activated state, it is checked whether an interface deactivation event has occurred. If it is determined that an interface deactivation event has occurred, the face gesture interface mode is terminated. Otherwise, the process of returning to the camera input step (4) is repeatedly performed. At this time, the termination conditions of the face gesture interface mode of the present invention must include the occurrence of an interface deactivation event, but the face gesture interface mode may also be terminated by an operation such as pressing the 'ESC' key or clicking the 'Exit' button when the user desires.
[0165] To verify the usability and demonstrate application examples of a user interface system using facial gesture recognition and facial authentication according to one embodiment of the present invention, gesture recognition experiments and user scenario simulations were performed.
[0166] To measure the recognition rate of facial gestures, 5 experiment participants performed 30 times each of 8 gestures corresponding to left-click, double-click, right-click, drag, drop, up scroll, down scroll, and interface activation and deactivation events.
[0167] Table 3 shows the number of successful recognitions and the average recognition rate of eight facial gestures for each user before cursor stabilization with improved simultaneous opening and closing of both eyes and adaptive moving average processing was performed.
[0168] user Left Click Double Click Right Click Drag Drop Scroll Up Scroll Down Interface Enable / Disable total A 26 26 27 30 30 30 30 27 226 B 29 29 29 30 30 30 30 28 235 C 26 27 29 30 30 30 30 26 228 D 27 28 28 30 30 30 30 26 229 E 28 27 30 30 30 30 30 27 232 Avg(%) 90.6 91.3 95.3 100 100 100 100 89.3 95.8
[0169] As shown in Table 3, a usability verification experiment was conducted with a total of 5 participants, and the average recognition rate of facial gestures was 95.8%. The average recognition rates for each gesture were 90.6% for left-click, 91.3% for double-click, 95.3% for right-click, and 89.3% for interface activation and deactivation, while other gestures showed a recognition rate of 100%. For drag gestures, the user was notified of a minimum movement range and evaluated as recognized when moving beyond that range; for upward and downward scrolling, the user was notified of a minimum hold time and evaluated as recognized when the scroll command was executed for longer than that time.
[0170] Meanwhile, Table 4 shows the number of successful recognitions and the average recognition rate of eight facial gestures for each user after cursor stabilization was performed by applying improvement of malfunctions in simultaneous opening and closing of both eyes and adaptive moving average processing.
[0171] user Left Click Double Click Right Click Drag Drop Scroll Up Scroll Down Interface Enable / Disable total 1 30 30 29 29 30 30 30 30 238 2 30 30 26 28 30 30 30 28 232 3 30 30 30 30 30 30 30 28 238 4 30 30 30 30 30 30 30 26 236 5 30 30 27 30 30 30 30 28 237 6 30 30 29 30 30 30 30 28 237 7 30 30 30 30 30 30 30 30 240 8 30 29 30 30 30 30 30 26 235 8 30 30 30 29 30 30 30 29 238 10 30 30 30 30 30 30 30 28 238 Avg(%) 100 99.6 97.0 98.6 100 100 100 93.6 98.7
[0172] To measure the recognition rate of facial gestures, 10 participants were asked to perform facial gestures. Compared to before cursor stabilization was performed by applying adaptive moving average processing and improving the error rate of simultaneous opening and closing of both eyes, it was observed that the recognition rate of almost all facial gestures improved, and among them, the recognition rate of the left-click gesture improved the most, rising from 90.6% to 100%. The double-click improved from 91.6% to 99.6%, the right-click from 95.3% to 97.0%, and the interface activation and deactivation from 89.3% to 93.6%. Consequently, the overall average recognition rate of facial gestures was found to have increased from 95.8% to 98.7%.
[0173] FIG. 15 is a diagram showing a user scenario simulation scene of a user interface system using face gesture recognition and face authentication according to an embodiment of the present invention. FIG. 15 shows a series of camera input frames with a horizontal flip applied. When a horizontal flip is applied to the input frames, the positions of the left and right eyes are swapped, making it similar to looking at one's own face reflected in a mirror.
[0174] Figure 15 (a) shows a situation where the user closes both eyes to activate the interface. Checking the yellow text in the upper left corner of the administrator screen reveals that the user's eye opening / closing state is False, indicating that both eyes are closed. Figure 15 (b) shows a situation where the user performs block selection of a specific sentence via a drag event while web surfing. At this time, the user's right eye is in a closed (False) state, and it can be seen that some sentences in the upper right corner of the web page are selected in blue using the drag event. Figure 15 (c) shows a situation where the user scrolls down to view a document, and it can be confirmed that the rolling_angle in the green text in the upper left corner is 39.443°, which exceeds the threshold of 35°. Figure 15 (d) shows a situation where the user plays Slither.io, a simple game using a cursor, and it can be seen that the user's green worm is moving toward the cursor position calculated by the Pan angle and Tilt angle of the face.
[0175] The user interface system using face gesture recognition and face authentication according to the present embodiment estimates the user's 3D face landmarks from an input image using a MediaPipe Face Mesh model and extracts 3D coordinates from seven selected 3D face landmarks. Subsequently, the Pan angle and Tilt angle of the face are calculated from the 3D nose tip landmark coordinates to determine the position of the cursor for each frame, and the Roll angle of the face and the opening / closing state of both eyes are estimated using the remaining six 3D face landmarks to determine the face gesture. The user interface is then controlled through cursor movement and the occurrence of user events based on such face gestures. Furthermore, the face user interface proposed in this embodiment was able to improve malfunctions by applying adaptive moving average processing to the fine tremors of the cursor that occurred during the face gesture recognition process. Additionally, a decrease in the face gesture recognition rate occurs due to the left and right eyes opening / closing inconsistently with slight differences; however, malfunctions could be improved by repeatedly checking the discontinuity of simultaneous opening / closing of both eyes to block temporary discrepancies in opening / closing of both eyes.
[0176] As explained earlier, in the face gesture recognition rate experiment for usability verification, the average recognition rate of face gestures showed reliable performance of 98.7%, and the frames per second (fps) was 30 frames, confirming that there was no difficulty in real-time processing.
[0177] In the face gesture interface according to the present invention, the movement of the cursor and the occurrence of user events are controlled using face gestures, thereby enabling the interface to be controlled in a natural and intuitive manner. Furthermore, as the interface can be controlled without direct contact with the device, its potential for application is high in everyday life environments, the materials, parts, and equipment (SoBuJang) industry, and kiosk application fields.
[0178] The user interface unit proposed in this embodiment detects a person registered with a face recognition model, and if a registered person is found, obtains 3D coordinates from 7 designated face landmarks using the MediaPipe Face Mesh model.
[0179] Among the seven coordinates, the Pan and Tilt angles of the face are calculated using the 3D coordinates corresponding to the tip of the nose to determine the position of the cursor for each frame, and the Roll angle and whether both eyes are open or closed are calculated using the 3D coordinates of the remaining six landmarks (outer corners of both eyes, upper eyelids, and lower eyelids) to determine the facial gesture. Then, cursor movement and user events are controlled and displayed on the screen according to such facial gestures.
[0180] It is expected that users will be able to easily learn how to use it and conveniently, as it controls the cursor and performs various events such as clicking, dragging, and scrolling the web screen through simple actions, such as tilting the face toward the shoulder or closing and opening both eyes.
[0181] Furthermore, in a chaotic business environment for materials, parts, and equipment where various equipment is mixed, facial recognition allows for the identification of authorized operators among collaborators and enables the rapid execution of complex tasks with simple facial gestures.
Claims
Claim 1 A user interface system using face gesture recognition and face authentication, comprising: a face authentication unit that identifies the facial features of a user granted operation authority among a number of users in a series of input video frames and selects a face-authenticated operator through this; and a user interface unit that allows the recognition of the face gesture of the operator selected through the face authentication unit, displays the result of the face gesture recognition on a screen when necessary to interact with the operator, and generates a user event based on the result of the face gesture recognition to execute a corresponding command, wherein the user interface unit repeatedly performs a process of forcibly correcting and recognizing the two most recent frames based on the point of recognition when both eyes are open, in order to prevent frequent malfunctions when transitioning from a closed state to an open state, by examining the two most recent frames based on the point of recognition when both eyes are open, and if the immediately preceding frame between the current frame where both eyes are open and the frame before the previous frame where both eyes are closed is recognized as having only one eye open, then both eyes are closed. Claim 2 A user interface system using face gesture recognition and face authentication according to claim 1, wherein the face authentication unit comprises: a face authentication dataset generation module that generates a face authentication dataset by sampling a series of person image frames including front, side, and tilted face compositions at a fixed frame interval; a face data learning module that generates a face recognition model through a supervised learning process that classifies faces using cosine similarity after extracting and normalizing multidimensional embedding vectors from face regions detected from the face authentication dataset; a face recognition module that performs face recognition and outputs a corresponding face class and recognition reliability when face regions are detected in a series of input image frames and input into the face recognition model; and an operator authentication module that authenticates the user of the face as the operator only when the recognized face class matches a registered operator and the recognition reliability is greater than or equal to a preset authentication threshold. Claim 3 A user interface system using face gesture recognition and face authentication according to claim 1, wherein the face authentication unit detects a face region, performs face recognition, and authenticates a user with a recognition confidence level greater than or equal to a preset level as an operator, and when multiple people exist within an input image, immediately after one person is recognized, the recognized person is erased or ignored from the screen and the process of recognizing other people is repeated. Claim 4 A user interface system using face gesture recognition and face authentication according to claim 1, wherein the user interface unit includes an estimation module that detects the movement of a user face in a series of input image frames, estimates the Pan angle and Tilt angle of the face, and calculates the horizontal and vertical coordinates of a cursor position on the screen corresponding to each face angle; and a face gesture recognition module that detects the movement of the face and binoculars through face landmarks estimated in the input image frames, recognizes a face gesture composed of a combination of the Roll angle of the face, the opening and closing state of the binoculars, and the holding time thereof, generates a user event corresponding to the face gesture in association with the cursor position calculated by the estimation module, and controls the interface to execute a command corresponding to the user event. Claim 5 A user interface system using face gesture recognition and face authentication, characterized in that, in claim 4, the estimation module further includes a cursor stabilization unit that stabilizes the fine shaking of the cursor by applying adaptive moving average processing according to the cursor movement speed to the horizontal and vertical coordinates of the cursor on the screen. Claim 6 A user interface system using face gesture recognition and face authentication according to claim 1, wherein the user interface unit further includes a binocular opening / closing determination unit that, when determining the binocular simultaneous opening / closing state, repeatedly examines the discontinuity of the binocular simultaneous opening / closing and corrects the temporary opening / closing discrepancy state of the binoculars under specific conditions so that it is unified into either an open or closed state. Claim 7 In claim 1, the user interface unit, as soon as the face gesture interface is activated, suppresses the movement of the user's face and ensures that the user looks at the exact center of the screen, and sets the estimated nose tip landmark coordinates from video frames received during an initial predetermined time as the angle calculation reference point of the 3D face landmark coordinate system by accumulating and averaging them, then obtains a 3D vector connecting the origin of the 3D face landmark coordinate system to the angle calculation reference point, and then obtains a 3D vector connecting the nose tip landmark coordinates estimated after the angle calculation reference point is determined from the origin of the 3D face landmark coordinate system, and then calculates the horizontal and vertical angles between the two 3D vectors and sets them as the face's Pan angle and Tilt angle, respectively; and when the face's Pan angle is within the range [(-)maximum Pan angle, (+)maximum Pan angle] with the screen width and the face's maximum Pan angle preset, the face's Pan angle is converted into a screen horizontal coordinate within the range [0, screen width] using trigonometric functions and a proportional equation, and when the face's Tilt angle is within the range [0, screen width] with the screen height and the face's maximum Tilt angle preset A user interface system using face gesture recognition and face authentication, characterized by converting the tilt angle of the face into a screen vertical coordinate within the range [0, screen vertical size] using trigonometric functions and a proportional equation when the angle is within the range [(-)maximum tilt angle, (+)maximum tilt angle]. Claim 8 A user interface system using face gesture recognition and face authentication according to claim 1, wherein the user interface unit determines the Roll angle of the face as the angle between the vector connecting the right eye corner landmark coordinate of the user face to the left eye corner landmark coordinate and the horizontal coordinate axis of the 3D face landmark coordinate system, and if the absolute value of the Roll angle of the face is greater than or equal to a preset threshold Roll angle, determines whether it is a positive (+) angle or a negative (-) angle and generates either a Scroll Down Event or a Scroll Up Event, and generates the other user event for an angle with the opposite sign. Claim 9 A user interface system using face gesture recognition and face authentication, characterized in that, in claim 8, the user interface unit recognizes a downward scroll gesture and generates a corresponding downward scroll event when the Roll angle of the face exceeds a positive (+) threshold Roll angle by tilting the head toward the right shoulder while the left and right corners of the eyes are horizontal, and recognizes an upward scroll gesture and generates a corresponding upward scroll event when the Roll angle of the face rotates further toward the left shoulder than a negative (-) threshold Roll angle by tilting the head toward the left shoulder. Claim 10 In claim 1, the user interface unit recognizes a facial gesture composed of a combination of the open / closed state of both eyes and the maintenance time thereof, estimated using the uppermost and lowermost left eyelid landmark coordinates and the uppermost and lowermost right eyelid landmark coordinates of the user's face; wherein if the difference in the vertical axis coordinate values of the uppermost and lowermost left eyelid landmark coordinates of the user's face is less than or equal to a specified threshold, the left eye is determined to be in a closed state, and otherwise, it is determined to be in an open state; and if the difference in the vertical axis coordinate values of the uppermost and lowermost right eyelid landmark coordinates of the user's face is less than or equal to the specified threshold, the right eye is determined to be in a closed state, and otherwise, it is determined to be in an open state; wherein if the facial gesture is in a closed state of both eyes within a first fixed time range and then both eyes become open, it is recognized as a left-click gesture and a corresponding left-click event is generated; wherein if the facial gesture is in a closed state of one eyeball within a first fixed time range and then that eyeball becomes open, it is recognized as a right-click gesture and a corresponding right-click event is generated; and if the facial gesture repeats the left-click gesture twice within a second fixed time range, a double-click It recognizes it as a gesture and generates a corresponding double-click event; if the cursor position moves after the right-click gesture is maintained for more than a third predetermined time, it recognizes it as a drag gesture and generates a corresponding drag event; if both eyes are open during the drag event, it recognizes it as a drop gesture and generates a corresponding drop event; and if the face gesture is maintained in a closed state for more than a fourth predetermined time which is longer than the third predetermined time, and both eyes are open, and the face gesture interface is currently inactive, it recognizes it as an interface activation gesture and generates a corresponding interface activation event to activate the face gesture interface.A user interface system using face gesture recognition and face authentication, characterized by terminating the face gesture interface by recognizing, conversely, if the state is active, as an interface deactivation gesture and generating a corresponding interface deactivation event. Claim 11 delete Claim 12 A user interface system using facial gesture recognition and facial authentication, characterized in that, in either claim 7 or 8, the 3D facial landmark coordinate system is a coordinate system of 3D facial landmarks of a MediaPipe Face Mesh model. Claim 13 A method using a user interface system utilizing facial gesture recognition and facial authentication comprises: (a) a step in which the user interface system identifies facial features of a user granted operation authority among a plurality of users in a series of input image frames and selects an operator whose face is authenticated therefrom; and (b) a step in which the user interface system allows facial gesture recognition of the selected operator, displays the facial gesture recognition result on a screen when necessary to interact with the operator, and controls an interface to generate a user event based on the facial gesture recognition result and execute a corresponding command, wherein step (b) is characterized by repeatedly performing a process of forcibly correcting and recognizing that both eyes are closed when recognizing the opening and closing of both eyes, based on the point in time when both eyes are recognized as being open, and if the immediately preceding frame between the current frame where both eyes are open and the frame before the previous frame where both eyes are closed is recognized as having only one eye open, the user interface system utilizes a user interface system utilizing facial gesture recognition and facial authentication to prevent frequent malfunctions when transitioning from a closed state to an open state. Claim 14 In claim 13, the above step (a) comprises: a step in which the user interface system samples a series of person image frames including frontal, side, and tilted face compositions at regular frame intervals to generate a face authentication dataset; a step in which the user interface system generates a face recognition model through a supervised learning process that classifies faces using cosine similarity after extracting and normalizing multidimensional embedding vectors from face regions detected from the face authentication dataset; a step in which the user interface system detects face regions in a series of input image frames, performs face recognition when input into the face recognition model, and outputs the corresponding face class and recognition confidence; and a step in which the user interface system authenticates the user of the face as the operator only when the recognized face class matches a registered operator and the recognition confidence is greater than or equal to a preset authentication threshold. Claim 15 In claim 13, the above step (a) is characterized by the user interface system detecting a face region, performing face recognition to authenticate a user with a recognition confidence level greater than or equal to a preset level as the operator, and, in the case where multiple people exist within the input image, immediately after one person is recognized, repeating the process of recognizing other people while immediately erasing or ignoring the recognized person from the screen. Claim 16 In claim 13, the above step (b) comprises: (b-1) a step in which the user interface system detects the movement of the user face in a series of input image frames and estimates the Pan angle and Tilt angle of the face, and then calculates the horizontal and vertical coordinates of the cursor position on the screen corresponding to each face angle; and (b-2) a step in which the user interface system detects the movement of the face and binoculars through the face landmarks estimated in the input image frames and recognizes a face gesture composed of a combination of the Roll angle of the face, the opening and closing state of the binoculars, and the holding time thereof, and then generates a user event corresponding to the face gesture in association with the cursor position, and controls the interface to execute a command corresponding to the user event. Claim 17 In claim 13, the method using a user interface system utilizing face gesture recognition and face authentication is characterized in that step (b) further includes a cursor stabilization step in which the user interface system applies adaptive moving average processing according to the cursor movement speed to the horizontal and vertical coordinates of the cursor on the screen to stabilize the fine shaking of the cursor. Claim 18 A method using a user interface system using face gesture recognition and face authentication, wherein the above step (b) further includes a binocular opening / closing determination step in which, when the user interface system determines the binocular simultaneous opening / closing state, the discontinuity of the binocular simultaneous opening / closing is repeatedly examined and the temporary opening / closing discrepancy state of the binoculars is modified to unify the state of opening and closing of the binoculars under specific conditions into either an open or closed state. Claim 19 In claim 13, the above step (b) suppresses the movement of the user's face and directs it to look at the center of the screen as soon as the face gesture interface is activated, and sets the estimated nose tip landmark coordinates from video frames received during an initial predetermined time as the cumulative average to determine the angle calculation reference point of the 3D face landmark coordinate system, then obtains a 3D vector connecting the origin of the 3D face landmark coordinate system to the angle calculation reference point, and then obtains a 3D vector connecting the nose tip landmark coordinates estimated after determining the angle calculation reference point from the origin of the 3D face landmark coordinate system, and then calculates the horizontal and vertical angles between the two 3D vectors to determine the Pan angle and Tilt angle of the face, respectively; and when the Pan angle of the face is within the range [(-)max Pan angle, (+)max Pan angle] with the screen width and the maximum Pan angle of the face pre-set, the Pan angle of the face is converted into a screen horizontal coordinate within the range [0, screen width] using trigonometric functions and proportional equations, and with the screen height and the maximum Tilt angle of the face pre-set, the face's A method using a user interface system utilizing face gesture recognition and face authentication, characterized by converting the tilt angle of the face into a screen vertical coordinate within the range [0, screen vertical size] using trigonometric functions and a proportional equation when the tilt angle is within the range [(-)max tilt angle, (+)max tilt angle]. Claim 20 In claim 13, the above step (b) is characterized in that the user interface system determines the Roll angle of the face as the angle between the vector connecting the right eye corner landmark coordinate of the user face to the left eye corner landmark coordinate and the horizontal coordinate axis of the 3D face landmark coordinate system, and if the absolute value of the Roll angle of the face is greater than or equal to a preset threshold Roll angle, it determines whether it is a positive (+) angle or a negative (-) angle and generates either a Scroll Down Event or a Scroll Up Event, and generates the other user event for an angle with the opposite sign. Claim 21 In claim 20, the above step (b) is characterized by the user interface system recognizing a downward scroll gesture and generating a corresponding downward scroll event when the Roll angle of the face exceeds a positive (+) threshold Roll angle due to tilting the head toward the right shoulder while the left and right corners of the eyes are horizontal, and recognizing an upward scroll gesture and generating a corresponding upward scroll event when the Roll angle of the face rotates further toward the left shoulder than a negative (-) threshold Roll angle due to tilting the head toward the left shoulder. Claim 22 In Clause 13, the above step (b) recognizes a facial gesture composed of a combination of the open / closed state of both eyes and the maintenance time thereof, estimated by the user interface system using the top left and bottom right eyelid landmark coordinates and the top right and bottom right eyelid landmark coordinates of the user face, wherein if the difference in the vertical axis coordinate values of the top left and bottom left eyelid landmark coordinates of the user face is less than or equal to a specified threshold, the left eye is determined to be in a closed state, otherwise it is determined to be in an open state; and if the difference in the vertical axis coordinate values of the top right and bottom right eyelid landmark coordinates of the user face is less than or equal to the specified threshold, the right eye is determined to be in a closed state, otherwise it is determined to be in an open state; wherein if the facial gesture is in a closed state of both eyes within a first fixed time range and then both eyes are open, it is recognized as a left-click gesture and a corresponding left-click event is generated; and if the facial gesture is in a closed state of one eyeball within a first fixed time range and then that eyeball is open, it is recognized as a right-click gesture and a corresponding right-click event is generated; and if the facial gesture is in a left-click gesture within a second fixed time range If repeated twice, it is recognized as a double-click gesture and a corresponding double-click event is generated; if the cursor position moves after the right-click gesture is maintained for more than a third predetermined time, it is recognized as a drag gesture and a corresponding drag event is generated; if both eyes are open during the drag event, it is recognized as a drop gesture and a corresponding drop event is generated; and if the face gesture is maintained in a closed state for more than a fourth predetermined time which is longer than the third predetermined time, and both eyes are open, if the face gesture interface is currently inactive, it is recognized as an interface activation gesture and a corresponding interface activation event is generated to activate the face gesture interface.A method using a user interface system utilizing face gesture recognition and face authentication, characterized in that if, conversely, it is in an active state, it recognizes it as an interface deactivation gesture and terminates the face gesture interface by generating a corresponding interface deactivation event. Claim 23 delete Claim 24 A method using a user interface system utilizing facial gesture recognition and facial authentication, characterized in that, in either claim 19 or claim 20, the 3D facial landmark coordinate system is a coordinate system of 3D facial landmarks of a MediaPipe Face Mesh model.
Citation Information
Patent Citations
An adaptive gain adjustment method for mapping gesture motion to the interface
CN105912126B