Facial paralysis patient-oriented living body detection method and system and medium

Through the method of facial key point detection and dynamic threshold adjustment, the problem of liveness detection of patients with facial paralysis in the face recognition system is solved, and effective liveness identity verification of patients with facial paralysis is achieved.

CN120853233APending Publication Date: 2025-10-28JINAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510842080.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing facial recognition technology has difficulty in effectively detecting the live identity of patients with facial paralysis, resulting in them being unable to pass liveness detection in scenarios such as face-swiping shopping and government authentication.

Method used

By acquiring the user's detection video, performing facial key point detection, calculating the ratio of the width to height of the eyes and mouth, using the preset threshold to independently judge the blinking and mouth opening movements, and dynamically adjusting the detection threshold, liveness detection of patients with facial paralysis can be achieved.

Benefits of technology

It has improved the pass rate of liveness detection for patients with facial paralysis, enabling patients with injuries to only one eye or facial muscles to pass the liveness detection smoothly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853233A_ABST
    Figure CN120853233A_ABST
Patent Text Reader

Abstract

The invention discloses a living body detection method and system for a facial paralysis patient and a medium. The method comprises the following steps: acquiring a detection video of a user; face key point detection is carried out on each frame of image in the detection video, calculation is carried out based on a detection result, and an eye width-to-height ratio and a mouth width-to-height ratio corresponding to each frame of image are determined; performing independent blink detection on each eye based on a preset eye opening threshold and an eye width-to-height ratio corresponding to a plurality of frames of images in the video; performing mouth opening detection based on the mouth opening threshold and the mouth width-to-height ratio corresponding to each frame of image; the mouth opening threshold value is obtained by adjusting a preset mouth opening threshold value according to an average mouth width-to-height ratio of a plurality of previous continuous frame images in the detection video; and determining a living body detection result of the user according to the blink detection result and the mouth opening detection result. According to the embodiment of the invention, the living body detection passing rate of the facial paralysis patient can be effectively improved, and the method can be widely applied to the technical field of face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of facial recognition technology, and in particular to a liveness detection method, system, and medium for patients with facial paralysis. Background Technology

[0002] With the increasing prevalence of digital services in my country, facial recognition technology is relied upon as the core means of identity verification in various scenarios, such as shopping at street pharmacies using facial recognition, a series of authentications on government software, and online hospital report inquiries. Moreover, most facial recognition methods require users to perform standardized actions such as blinking and opening their mouths as the basis for determining liveness detection.

[0003] However, a significant portion of the population is being excluded from these identity verification methods—a large number of patients with limited facial muscle movement due to conditions such as facial paralysis, facial burns, and stroke (collectively referred to as facial paralysis patients in this article). Currently, situations exist where facial paralysis patients, unable to freely control their mouths or eyelids to perform standardized movements, experience failed medical insurance settlements and are forced to return to manual service windows; facial burn patients, due to eyelid deformation preventing them from closing, are identified as "not a human operator" by bank liveness detection systems, causing considerable inconvenience to their lives.

[0004] This shows that the pass rate of facial recognition technology that relies on standardized actions such as blinking and opening the mouth to detect the liveness of facial paralysis patients is very low. Summary of the Invention

[0005] In view of this, in order to solve one of the above problems, the purpose of this invention is to provide a liveness detection method, system and medium for patients with facial paralysis, which can effectively improve the pass rate of liveness detection for patients with facial paralysis.

[0006] On one hand, embodiments of the present invention provide a liveness detection method for patients with facial paralysis, including: Acquire the user's detection video; the detection video includes several frames of images; Facial landmark detection is performed on each frame of the image in the detection video, and calculations are performed based on the detection results to determine the width-to-height ratio of the eyes and the width-to-height ratio of the mouth for each frame of the image; the width-to-height ratio of the eyes includes the width-to-height ratio of the left eye and the width-to-height ratio of the right eye; Based on a preset eye-opening threshold, the width-to-height ratio of the left eye and the width-to-height ratio of the right eye corresponding to several frames of the image in the video are independently judged, and the blink detection result of the user is determined according to the judgment result; The mouth width-to-height ratio of each frame of the image is determined based on the mouth opening threshold to determine the user's mouth opening detection result; the mouth opening threshold is obtained by adjusting a preset mouth opening threshold based on the average mouth width-to-height ratio of the first few consecutive frames of the detected video. The liveness detection result of the user is determined based on the blink detection result and the mouth opening detection result.

[0007] Specifically, the aspect ratio of the eye corresponding to each frame of the image is determined in the following way: Facial key point detection is performed on the image, and the coordinates of the upper and lower edge points of the eyelids and the coordinates of the left and right edge points of the eyes are determined based on the detected eye key point coordinates. The vertical distance to the eye is determined based on the coordinates of the upper and lower edge points of the eyelid; The horizontal distance of the eye is determined based on the coordinates of the left and right edge points of the eye. The aspect ratio of the eye in the image is determined by calculating based on the first preset formula, the vertical distance of the eye, and the horizontal distance of the eye.

[0008] Specifically, the aspect ratio of the mouth corresponding to each frame of the image is determined in the following way: Facial key point detection is performed on the image, and the coordinates of the inner and outer edges of the lips and the lowest point of the chin are determined based on the detected mouth key point coordinates; the inner and outer edges of the lips include the two endpoints of the outer corner of the mouth and the midline of the upper and lower lips; The width of the outer edge of the lip is determined based on the coordinates of the two endpoints of the outer corner of the mouth; The height of the inner edge of the lips is determined based on the coordinates of the midline points of the upper and lower lips; The distance from the inner edge of the lip to the chin is determined based on the coordinates of the midline point of the lower lip and the coordinates of the lowest point of the chin. The width-to-height ratio of the mouth corresponding to the image is determined by calculating the outer edge width of the lips, the inner edge height of the lips, and the distance from the inner edge of the lips to the chin according to the second preset formula.

[0009] Specifically, the step of independently determining the aspect ratios of the left and right eyes corresponding to several frames of the image in the video based on a preset eye-opening threshold, and determining the user's blink detection result based on the determination result, includes: Based on a preset eye-opening threshold, the aspect ratio of the left eye corresponding to several frames in the video is determined to obtain the left eye determination result. Based on a preset eye-opening threshold, the aspect ratio of the right eye in several frames of the video is determined to obtain the right eye determination result. If the left eye assessment result and / or the right eye assessment result are both passed, the user's blink detection result is passed.

[0010] Specifically, the left-eye judgment result or the right-eye judgment result is obtained in the following way: When there are several consecutive frames of images in a series of images whose eye aspect ratios successively satisfy the following conditions: eye aspect ratio equal to preset eye-opening threshold, eye aspect ratio greater than preset eye-opening threshold, eye aspect ratio equal to preset eye-opening threshold, and eye aspect ratio less than preset eye-opening threshold, the judgment result is passed.

[0011] Specifically, adjusting the preset mouth opening threshold based on the average mouth width-to-height ratio of the preceding several consecutive frames in the detected video includes: If the mouth width-to-height ratio of the first few consecutive frames of the detection video does not meet the preset detection threshold, the average mouth width-to-height ratio of the first few consecutive frames of the detection video is matched with the preset detection threshold correspondence table, and the preset mouth opening threshold is adjusted according to the matching result.

[0012] Furthermore, if the user's liveness detection result is successful, the method further includes: A set of face images is obtained by capturing the face regions in the detected video using a classifier; The face image set is preprocessed to obtain a preprocessed image set; the image preprocessing includes grayscale conversion, histogram equalization, and normalization. The preprocessed image set is used to extract features based on the local binary model histogram algorithm to obtain a feature extraction set. The facial recognition result of the user is determined by comparing and analyzing the extracted feature set with the preset feature database.

[0013] On the other hand, embodiments of the present invention also provide a liveness detection system for patients with facial paralysis, comprising: The first module is used to acquire the user's detection video; the detection video includes several frames of images; The second module is used to perform facial key point detection on each frame of the image in the detection video and calculate based on the detection results to determine the width-to-height ratio of the eyes and the width-to-height ratio of the mouth corresponding to each frame of the image; the width-to-height ratio of the eyes includes the width-to-height ratio of the left eye and the width-to-height ratio of the right eye. The third module is used to independently determine the width-to-height ratio of the left eye and the width-to-height ratio of the right eye corresponding to several frames of the image in the video based on a preset eye-opening threshold, and determine the blink detection result of the user based on the judgment result; The fourth module is used to determine the mouth width-to-height ratio of each frame of the image based on the mouth opening threshold, and to determine the mouth opening detection result of the user; the mouth opening threshold is obtained by adjusting the preset mouth opening threshold according to the average mouth width-to-height ratio of the first few consecutive frames of the detection video. The fifth module is used to determine the user's liveness detection result based on the blink detection result and the mouth opening detection result.

[0014] On the other hand, embodiments of the present invention also provide a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the method described above.

[0015] On the other hand, embodiments of the present invention also provide a liveness detection system for patients with facial paralysis, including an image acquisition device and a computer device connected to the image acquisition device; wherein, The image acquisition device is used to acquire the user's detection video; The computer device includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described above.

[0016] Implementing the embodiments of the present invention has the following beneficial effects: This embodiment provides a liveness detection method, system, and medium for patients with facial paralysis. The method acquires a user's detection video, performs facial landmark detection on each frame of the video, and calculates the width-to-height ratio of the eyes and mouth for each frame based on the detection results. Then, it performs eye-opening and mouth-opening detection separately, and determines the user's liveness detection result based on the obtained results. Furthermore, this embodiment of the invention performs independent blink detection on each of the user's eyes, enabling liveness detection based on the blink detection results of each eye. This addresses situations where a single eye injury prevents complete binocular blink detection. Furthermore, the method of this invention allows facial paralysis patients to undergo liveness detection. On the other hand, the method adjusts the mouth opening detection threshold based on the average mouth width-to-height ratio of the first few consecutive frames obtained from video detection. This involves analyzing the user's average mouth movement amplitude over a certain period and automatically adjusting the movement threshold based on the user's actual facial muscle movement ability. This enables facial paralysis patients who cannot fully complete the mouth opening action due to facial muscle damage to pass liveness detection. In summary, this invention's liveness detection mechanism, based on dynamically adjusting the detection threshold and a blink detection mechanism capable of independently detecting each eye, effectively improves the liveness detection pass rate for facial paralysis patients. Attached Figure Description

[0017] Figure 1 This is a schematic flowchart of a liveness detection method for patients with facial paralysis provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of an image acquisition screen provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the acquisition of key eye point locations according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the location of key points on the mouth as provided in an embodiment of the present invention; Figure 5 This is a test diagram of a simulated facial paralysis patient undergoing blink detection, provided by an embodiment of the present invention; Figure 6 This is a test diagram provided by an embodiment of the present invention to simulate a facial paralysis patient performing a mouth-opening test; Figure 7 This is a schematic diagram of a face recognition test result provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of another face recognition test result provided by an embodiment of the present invention; Figure 9 This is a structural block diagram of a liveness detection system for patients with facial paralysis provided in an embodiment of the present invention; Figure 10 This is a structural block diagram of a liveness detection system for patients with facial paralysis provided in an embodiment of the present invention; Figure 11 This is a structural block diagram of a liveness detection device for patients with facial paralysis provided in an embodiment of the present invention. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adapted according to the understanding of those skilled in the art.

[0019] The following is a brief introduction to some of the technical terms used in the article: The House-Brackmann Facial Paralysis Scale is an internationally used assessment method, divided into six levels. Level I: Bilaterally symmetrical, normal facial muscle function; Level II: Mild impairment, statically symmetrical, can close eyes and move corners of the mouth with slight effort; Level III: Moderate impairment, differential muscle tone, unable to raise eyebrows, eyelids can close with effort, asymmetrical corners of the mouth; Level IV: Moderate to severe impairment, weakened or deformed muscle tone, unable to raise eyebrows, difficulty closing eyelids; Level V: Severe impairment, statically asymmetrical, no forehead movement, incomplete eyelid closure; Level VI: Complete facial paralysis, no tone, no synkinesis, etc.

[0020] Liveness detection: A technology used to verify whether a target is a real living person, widely used in facial recognition, identity authentication and other fields. It effectively distinguishes real faces from deception methods such as photos, videos or masks by analyzing the physiological characteristics (such as blinking and opening the mouth) or behavioral characteristics (such as head turning) of the face, thereby improving the security and reliability of the system.

[0021] The HBGS (House-Brackmann) system is a grading system that assesses the function of 10 facial expression muscles, including forehead wrinkles and eye fissures, ranging from normal (Grade I) to complete paralysis (Grade VI). The grading criteria range from normal to severe abnormality, with lower scores indicating better facial nerve function. It is a commonly used method for assessing the degree of facial nerve dysfunction.

[0022] Keypoint detection technology: This technology automatically identifies and locates key feature points of target objects (such as the human body, face, or hand) through algorithms. It is widely used in fields such as face recognition, motion capture, and pose estimation. The core methods include traditional image processing and deep learning-based solutions. It features high accuracy and real-time performance and is one of the important foundational technologies of computer vision and artificial intelligence.

[0023] MediaPipe Face Mesh: A facial landmark detection technology based on machine learning, capable of estimating 468 3D facial landmarks in real time, covering areas such as the eyes and nose. It employs a lightweight architecture and GPU acceleration, eliminating the need for a dedicated depth sensor, and provides a metric 3D space, making it suitable for applications such as facial expression analysis and virtual makeup.

[0024] LBPH (Local Binary Patterns Histograms) algorithm: an algorithm for image feature extraction and recognition. Based on local binary patterns, it divides an image into local regions, obtains binary codes by comparing the gray values ​​of pixels with those of their neighbors, and forms histograms to represent the features of these regions, thereby achieving image recognition. It is widely used in fields such as face recognition.

[0025] Confidence score: A numerical metric that measures the reliability of a model's predictions, typically ranging from 0 to 1 (or 0% to 100%). It reflects the model's certainty about the predicted answer: a higher value indicates a more reliable prediction, while a lower value suggests a potentially larger error. This metric is widely used in AI tasks such as classification and detection to help users assess the quality of results and assist in decision-making, and is one of the key parameters for evaluating model performance.

[0026] Eye Aspect Ratio (EAR): This index reflects the degree of eye opening and closing. It is calculated by measuring the length-to-width ratio of the eye area. It can be used for fatigue detection; the smaller the value, the more fatigued the eyes. It can also be used for blink detection; values ​​below the threshold can be used to determine blinking.

[0027] Mouth Aspect Ratio (MAR): A metric used to quantify the openness of the mouth. It is obtained by calculating the ratio of the width to the height of the mouth area, typically based on the coordinates of key points on the mouth's outline (such as the corners of the mouth and lips). This ratio effectively reflects the degree of mouth opening—a higher value indicates a more pronounced mouth opening.

[0028] like Figure 1 As shown in the figure, this embodiment of the invention provides a liveness detection method for patients with facial paralysis, which includes the following steps.

[0029] S100: Acquire the user's detection video; the detection video includes several frames of images.

[0030] The user's face is captured in real time by a camera, resulting in a captured image containing several frames; in this invention, the user group mainly discusses patients with facial paralysis.

[0031] Specifically, during the video acquisition process in step S100, the acquisition screen is displayed in real time, and a green rectangle is drawn around the detected face area. Simultaneously, the current acquisition progress (e.g., Collect: 10 / 100) is displayed on the screen. In one embodiment, the image acquisition screen is as follows: Figure 2 As shown.

[0032] S200: Perform facial landmark detection on each frame of the detected video and calculate the width-to-height ratio of the eyes and mouth corresponding to each frame based on the detection results; the width-to-height ratio of the eyes includes the width-to-height ratio of the left eye and the width-to-height ratio of the right eye.

[0033] Based on facial landmark detection technology, facial landmarks and their coordinates are extracted from the detected video (in this embodiment, the landmarks of the eyes and mouth are mainly referred to). Based on the extracted coordinates, the width-to-height ratio of the eyes (EAR) and the width-to-height ratio of the mouth (MAR) in each frame of the image are calculated.

[0034] S300: Based on a preset eye-opening threshold, the width-to-height ratios of the left and right eyes in several frames of the video are independently judged, and the blink detection result of the user is determined according to the judgment result.

[0035] Based on a preset eye-opening threshold, the width-to-height ratios of the left and right eyes and their corresponding eyes are independently detected for blinking. The user's blinking detection result is determined based on the judgment results of the left and right eyes.

[0036] S400: Determine the mouth width-to-height ratio of each frame image based on the mouth opening threshold to determine the user's mouth opening detection result; the mouth opening threshold is obtained by adjusting the preset mouth opening threshold based on the average mouth width-to-height ratio of the first few consecutive frames in the detection video.

[0037] The mouth width-to-height ratio of each frame of the image is judged based on the mouth opening threshold. When the mouth width-to-height ratio is higher than the mouth opening threshold, the mouth is considered to be in an open state. In this invention, the preset mouth opening threshold is adjusted based on the average mouth width-to-height ratio of the first few consecutive frames of the detected video. That is, the dynamic situation of the user's mouth over a certain period of time is analyzed to determine the actual mouth opening threshold used for judgment.

[0038] S500: Determines the user's liveness detection result based on blink detection results and mouth opening detection results.

[0039] The system analyzes the user's blink and mouth opening data. If both pass the liveness detection, the user is considered a "real person," meaning they have passed the liveness detection.

[0040] Specifically, in step S200, the aspect ratio of the eye corresponding to each frame image is determined in the following way: S210: Perform facial landmark detection on the image, and determine the coordinates of the upper and lower edge points of the eyelids and the coordinates of the left and right edge points of the eyes based on the detected eye landmark coordinates.

[0041] Using MediaPipe Face Mesh, a facial landmark detection technology, the coordinates of key points in the eyes are extracted, including the upper and lower edges of the eyelids and the left and right edges of the eyes, to form the eye contour, which is used to calculate the width-to-height ratio of the eyes.

[0042] In one embodiment, the locations of key eye points collected are as follows: Figure 3 As shown.

[0043] S220: Determine the vertical distance to the eye based on the coordinates of the upper and lower edge points of the eyelid.

[0044] The vertical distance to the eye is calculated based on the coordinates of the upper and lower edges of the eyelid, which is then used to calculate the width-to-height ratio of the eye.

[0045] In this embodiment, the vertical distances to the eyes are calculated based on the coordinates of the upper and lower edge points of the eyelids, and include vertical distance 1 and vertical distance 2.

[0046] S230: Determine the horizontal distance of the eye based on the coordinates of the left and right edge points of the eye.

[0047] The horizontal distance of the eye is calculated based on the coordinates of the left and right edge points of the eye, which is then used to calculate the width-to-height ratio of the eye.

[0048] S240: Calculate the width-to-height ratio of the eye in the image based on the first preset formula, the vertical distance of the eye and the horizontal distance of the eye.

[0049] The width-to-height ratio of the eyes is obtained by calculating the vertical distance (distance between the upper and lower eyelids) and the horizontal distance (width of the eyes). The calculation formula is shown in equation (1): (1) Specifically, in step S200, the aspect ratio of the mouth corresponding to each frame of the image is determined in the following way: S250: Perform facial landmark detection on the image, and determine the coordinates of the inner and outer edge points of the lips and the coordinates of the lowest point of the chin based on the detected mouth landmark coordinates; the inner and outer edge points of the lips include the two endpoints of the outer corner of the mouth and the midline point of the upper and lower lips.

[0050] In mouth opening detection, a facial mesh model based on MediaPipe is used, according to the predefined mouth key point index of MediaPipe, including the outer lip width (the two endpoints of the outer corner of the mouth), the inner lip height (the midline point of the upper and lower lips), and the chin reference point (the lowest point of the chin as a vertical reference).

[0051] In one embodiment, the collected key points of the mouth are as follows: Figure 4 As shown, S260: Determine the width of the outer edge of the lip based on the coordinates of the two endpoints of the outer corner of the mouth.

[0052] The width of the outer edge of the lip is calculated based on the coordinates of the two endpoints of the outer corner of the mouth, which is then used to calculate the mouth width-to-height ratio (MAR).

[0053] S270: Determine the height of the inner edge of the lips based on the coordinates of the midline points of the upper and lower lips.

[0054] The height of the inner edge of the lips is calculated based on the coordinates of the midline points of the upper and lower lips, which is then used to calculate the mouth width-to-height ratio (MAR).

[0055] S280: Determine the distance from the inner edge of the lip to the chin based on the coordinates of the midline point of the lower lip and the coordinates of the lowest point of the chin.

[0056] The distance from the inner edge of the lip to the chin is calculated based on the coordinates of the midline point of the lower lip and the coordinates of the lowest point of the chin (reference point), which is then used to calculate the mouth width-to-height ratio (MAR).

[0057] S290: Calculate the width-to-height ratio of the mouth corresponding to the image based on the second preset formula, the width of the outer edge of the lips, the height of the inner edge of the lips, and the distance from the inner edge of the lips to the chin.

[0058] The width-to-height ratio of the mouth is obtained by calculating the width and height of the mouth, as well as the vertical offset of the chin point; the calculation formula is shown in the following formula (2): (2) Specifically, in step S300, based on a preset eye-opening threshold, the aspect ratios of the left and right eyes corresponding to several frames of the image in the video are independently determined, and the blink detection result of the user is determined according to the determination result, including: S310: Based on a preset eye-opening threshold, determine the aspect ratio of the left eye corresponding to several frames in the video to obtain a left-eye determination result. Based on the preset eye-opening threshold, determine the aspect ratio of the right eye corresponding to several frames in the video to obtain a right-eye determination result.

[0059] The eye state is defined based on the relationship between the eye-opening threshold and the corresponding eye width-to-height ratio in the image. Blink detection is performed based on changes in the eye state to determine whether the eye has completed the blinking action.

[0060] Specifically, in step S310, the left-eye judgment result or the right-eye judgment result is obtained in the following way: When there are several consecutive frames of images in a series of images whose eye aspect ratios successively satisfy the following conditions: eye aspect ratio equal to preset eye-opening threshold, eye aspect ratio greater than preset eye-opening threshold, eye aspect ratio equal to preset eye-opening threshold, and eye aspect ratio less than preset eye-opening threshold, the judgment result is passed.

[0061] The system defines four eye states (open → closing → closed → opening) based on the relationship between the eye-opening threshold and the corresponding eye width-to-height ratio in the image. The completeness of the action is determined by these thresholds. Specifically, when the width-to-height ratio of any eye changes from above the eye-opening threshold to below the eye-closing threshold and then back above the eye-opening threshold, a blink-closing-opening cycle is considered complete, and the blink count for that eye is incremented by one.

[0062] In one embodiment, blink analysis is performed as follows: After testing by multiple people, the initial eye-opening threshold EAR was selected as 0.25.

[0063] Eyes open: The vertical spacing is small, and the horizontal spacing is large, i.e., EAR is greater than 0.25. When the EAR gradually decreases from greater than 0.25 to less than 0.25, it is considered that the user is closing their eyes. When the EAR gradually decreases further to less than 0.2125, it is considered that the user has closed their eyes.

[0064] Eyes closed: When the user's EAR gradually increases from a closed-eye state to EAR>0.225, the user is considered to be opening their eyes. When the EAR gradually increases again to EAR>0.25, the user is considered to have fully opened their eyes.

[0065] S320: If the left eye judgment result and / or the right eye judgment result are passed, the user's blink detection result is passed.

[0066] After a user completes the full blinking action of opening and closing one eye, the blink count for that eye increments by one, indicating that the blinking test has been passed. Therefore, if a facial paralysis patient can only perform blinking testing in one eye due to injury, they can still pass the blinking test by completing the blinking test in one eye, demonstrating the inclusivity of this detection mechanism for facial paralysis patients. Figure 5 The image shown is a test diagram simulating a blinking test in a patient with facial paralysis. Figure 5 Figures (1), (3), and (5) are the input detection images of the user, while Figures (2), (4), and (6) are the detection results corresponding to the detection images to their left. It can be seen that patients with facial paralysis can also pass the blink detection smoothly.

[0067] Specifically, in step S400, the preset mouth opening threshold is adjusted based on the average mouth width-to-height ratio of the first few consecutive frames in the detection video, including: If the mouth width-to-height ratio of the first few consecutive frames of the detection video does not meet the preset detection threshold, the average mouth width-to-height ratio of the first few consecutive frames of the detection video is matched with the preset detection threshold correspondence table, and the preset mouth opening threshold is adjusted according to the matching result.

[0068] The user's mouth movement ability is judged based on the average mouth width-to-height ratio of the first few consecutive images of the detection video. The average mouth width-to-height ratio is matched with a preset detection threshold correspondence table, and the mouth opening threshold is adaptively adjusted based on the matching result, so that even patients with facial paralysis and reduced mouth movement ability can pass the mouth opening test.

[0069] Optionally, the preset detection threshold correspondence table can be set according to the House-Brackmann facial paralysis classification (HBGS), as shown in Table 1.

[0070] Table 1

[0071] In one embodiment, when the mouth's width-to-height ratio exceeds a preset threshold (e.g., an initial threshold of 400), the mouth is considered to be in an open state. Normal users can easily reach this detection threshold. For patients with facial paralysis or other pathological conditions, if the user's mouth width-to-height ratio is detected to be between 250 and 200 for 15 consecutive frames, the user is determined to be a Grade III facial paralysis patient, and the detection threshold will be automatically set to 200. When the detected mouth width-to-height ratio exceeds the preset threshold, the mouth is considered to be in an open state, the mouth opening count is incremented by one, and the mouth opening detection is considered passed. Figure 6The image shown is a test diagram simulating a facial paralysis patient undergoing a mouth-opening test. It demonstrates that even patients with facial paralysis and mouth-related disabilities can successfully pass the mouth-opening test.

[0072] Furthermore, if the user's liveness detection result is successful, the method in this embodiment also includes a subsequent face recognition method for patients with facial paralysis. This method includes: S600: Capture the face regions in the detected video using a classifier to obtain a set of face images.

[0073] By detecting the user's face in the video frame by frame using a camera, after detecting the same face in several consecutive frames, the pixel coordinates of the face region are identified, and a separate face image is captured from each frame to obtain a set of face images.

[0074] S700: Perform image preprocessing on the face image set to obtain a preprocessed image set; the image preprocessing includes grayscale conversion, histogram equalization, and normalization.

[0075] The captured face image set is resized to a uniform size and converted to grayscale. Then, it undergoes histogram equalization to enhance image contrast and improve the accuracy of face detection.

[0076] S800: The preprocessed image set is subjected to feature extraction based on the local binary model histogram algorithm to obtain the feature extraction set.

[0077] The LBPH algorithm is used to extract local features from the preprocessed image that has been converted to grayscale, and the LBPH feature vector corresponding to the preprocessed image is obtained. The extracted feature vector is then used to construct a feature extraction set.

[0078] S900: Based on the comparative analysis of the extracted feature set and the preset feature database, determine the user's face recognition result.

[0079] The obtained feature extraction set is compared with a pre-trained preset feature database. By judging the difference in feature values ​​between the current image and the images in the database (in some embodiments, this can be reflected by outputting a confidence value), the user's face recognition result is identified and determined.

[0080] Specifically, the LBP feature vectors in the feature extraction set are compared with the feature vectors already stored in the preset database using Euclidean distance or other distance metrics to calculate the similarity between the two vectors, and finally the face recognition result is obtained.

[0081] The formula for calculating the three-dimensional Euclidean distance is shown in equation (3) below: (3) In the formula, The distance is the three-dimensional Euclidean distance. The coordinates of the LBP feature vector of the image to be tested. The coordinates of the LBP feature vectors that are already stored in the database.

[0082] Specifically, such as Figure 7 and Figure 8 The image shown is a schematic diagram illustrating the test results of face recognition using the method described in this example. Figure 5 This is a test image showing a successful facial recognition scan. Figure 6 These are test images showing failed face recognition. Figure 6 Figures (1) and (3) are the input user's detection images, and Figures (2) and (4) are the detection results corresponding to the detection images to their left. It can be seen that if the recognition passes, the user's ID information will be displayed, and if it fails, "unknown" and the confidence value will be displayed (Conf in the figure is the confidence value, and the output value in this case is 111.28). Among them, using a box in the shape of a face as the foreground image can indicate to the user to keep the face in the face area, so as to reduce the interference of the distance between the face of the object being detected and the camera on the recognition result.

[0083] Specifically, the pre-defined feature database can be trained in the following way: Real-time facial images of users are captured via camera and stored in a folder corresponding to the user's ID, providing a pre-training facial image dataset for subsequent facial model training. The facial training model transforms the collected facial data into a recognizable feature model, establishing a mapping relationship between facial features and user identity, and generating model files usable in the recognition stage.

[0084] Alternatively, the image processing method used during face training can be OpenCV, PIL, TensorFlow, or other image processing methods.

[0085] Specifically, the process of converting the collected face training image data into a dataset that can be used to identify face sample features includes: (1) Obtain a training face image dataset containing user information; (2) Perform feature extraction (such as global features, human eye features, nose features, mouth features, etc.) to obtain the feature value space of the training samples; (3) Construct the mapping relationship between facial features and user identity to obtain a facial sample feature dataset.

[0086] Implementing the embodiments of the present invention has the following beneficial effects: This embodiment provides a liveness detection method, system, and medium for patients with facial paralysis. The method acquires the user's detection video, performs facial key point detection on each frame of the detection video, and calculates the width-to-height ratio of the eyes and the width-to-height ratio of the mouth corresponding to each frame of the image based on the detection results. Then, it performs eye-opening detection and mouth-opening detection respectively, and determines the user's liveness detection result based on the obtained eye-opening and mouth-opening detection results. On the one hand, the method of this invention performs independent blink detection on each of the user's eyes, and performs liveness detection based on the blink detection results of each eye. This enables facial paralysis patients who cannot complete the blink detection of both eyes due to injury in one eye to still undergo liveness detection. On the other hand, the method of this invention adjusts the mouth opening detection threshold based on the average mouth width-to-height ratio of the first few consecutive frames of images obtained from video detection. That is, it analyzes the average mouth movement amplitude of the user within a certain period of time, and automatically adjusts the movement threshold according to the actual facial muscle movement ability of the user. This enables facial paralysis patients who cannot complete the mouth opening movement due to facial muscle damage to still pass the liveness detection. In summary, the mouth opening detection mechanism based on dynamically adjusting the detection threshold and the blink detection mechanism that can independently detect each eye to perform liveness detection on users can effectively improve the liveness detection pass rate of facial paralysis patients.

[0087] like Figure 9 As shown, this embodiment of the invention also provides a liveness detection system for patients with facial paralysis, comprising: The first module is used to acquire the user's detection video; the detection video includes several frames of images; The second module is used to detect facial key points in each frame of the video and calculate based on the detection results to determine the width-to-height ratio of the eyes and the width-to-height ratio of the mouth for each frame; the width-to-height ratio of the eyes includes the width-to-height ratio of the left eye and the width-to-height ratio of the right eye. The third module is used to independently judge the width-to-height ratio of the left eye and the width-to-height ratio of the right eye in several frames of the video based on a preset eye-opening threshold, and determine the user's blink detection result based on the judgment result. The fourth module is used to determine the mouth width-to-height ratio of each frame image based on the mouth opening threshold, and to determine the user's mouth opening detection result; the mouth opening threshold is obtained by adjusting the preset mouth opening threshold based on the average mouth width-to-height ratio of the first few consecutive frames in the detection video. The fifth module is used to determine the user's liveness detection results based on blink detection results and mouth opening detection results.

[0088] It is evident that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0089] like Figure 10 As shown, this embodiment of the invention also provides a liveness detection system for patients with facial paralysis, including an image acquisition device and a computer device connected to the image acquisition device; wherein, The image acquisition device is used to acquire the user's detection video; The computer device includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the steps of the liveness detection method for patients with facial paralysis as described in the above method embodiments.

[0090] Specifically, the image acquisition device is mainly implemented through a camera, and may specifically include at least one camera; while the computer device may be different types of electronic devices, including but not limited to desktop computers, laptops and other terminals.

[0091] It is evident that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0092] like Figure 11 As shown, this embodiment of the invention also provides a liveness detection device for patients with facial paralysis, comprising: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the steps of the liveness detection method for patients with facial paralysis as described in the above method embodiments.

[0093] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. The memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include remote memory located remotely relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0094] It is evident that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented in this device embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0095] Furthermore, embodiments of this application also disclose a computer program product or computer program stored in a computer-readable storage medium. A processor of a computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the methods described above.

[0096] This invention also provides a computer-readable storage medium storing a processor-executable program that, when executed by a processor, implements the above-described method. Similarly, the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0097] It is understood that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0098] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A live biopsy method for patients with facial paralysis, characterized in that, include: Obtain the user's test video; The detection video includes several frames of images; Facial landmark detection is performed on each frame of the image in the detection video, and calculations are performed based on the detection results to determine the width-to-height ratio of the eyes and the width-to-height ratio of the mouth corresponding to each frame of the image. The eye width-to-height ratio includes the width-to-height ratio of the left eye and the width-to-height ratio of the right eye; Based on a preset eye-opening threshold, the width-to-height ratio of the left eye and the width-to-height ratio of the right eye corresponding to several frames of the image in the video are independently judged, and the blink detection result of the user is determined according to the judgment result; The mouth width-to-height ratio of each frame of the image is determined based on the mouth opening threshold to determine the user's mouth opening detection result; The mouth opening threshold is obtained by adjusting the preset mouth opening threshold based on the average mouth width-to-height ratio of the first few consecutive frames in the detection video. The liveness detection result of the user is determined based on the blink detection result and the mouth opening detection result.

2. The method according to claim 1, characterized in that, The aspect ratio of the eye corresponding to each frame of the image is determined in the following way: Facial key point detection is performed on the image, and the coordinates of the upper and lower edge points of the eyelids and the coordinates of the left and right edge points of the eyes are determined based on the detected eye key point coordinates. The vertical distance to the eye is determined based on the coordinates of the upper and lower edge points of the eyelid; The horizontal distance of the eye is determined based on the coordinates of the left and right edge points of the eye. The aspect ratio of the eye in the image is determined by calculating based on the first preset formula, the vertical distance of the eye, and the horizontal distance of the eye.

3. The method according to claim 1, characterized in that, The aspect ratio of the mouth in each frame of the image is determined in the following way: Facial key point detection is performed on the image, and the coordinates of the inner and outer edges of the lips and the lowest point of the chin are determined based on the detected mouth key point coordinates; the inner and outer edges of the lips include the two endpoints of the outer corner of the mouth and the midline of the upper and lower lips; The width of the outer edge of the lip is determined based on the coordinates of the two endpoints of the outer corner of the mouth; The height of the inner edge of the lips is determined based on the coordinates of the midline points of the upper and lower lips; The distance from the inner edge of the lip to the chin is determined based on the coordinates of the midline point of the lower lip and the coordinates of the lowest point of the chin. The width-to-height ratio of the mouth corresponding to the image is determined by calculating the outer edge width of the lips, the inner edge height of the lips, and the distance from the inner edge of the lips to the chin according to the second preset formula.

4. The method according to claim 1, characterized in that, The step of independently determining the aspect ratios of the left and right eyes corresponding to several frames of the video based on a preset eye-opening threshold, and determining the user's blink detection result based on the determination results, includes: Based on a preset eye-opening threshold, the aspect ratio of the left eye corresponding to several frames in the video is determined to obtain the left eye determination result. Based on a preset eye-opening threshold, the aspect ratio of the right eye in several frames of the video is determined to obtain the right eye determination result. If the left eye assessment result and / or the right eye assessment result are both passed, the user's blink detection result is passed.

5. The method according to claim 4, characterized in that, The left-eye or right-eye judgment result is obtained in the following way: When there are several consecutive frames of images in a series of images whose eye aspect ratios successively satisfy the following conditions: eye aspect ratio equal to preset eye-opening threshold, eye aspect ratio greater than preset eye-opening threshold, eye aspect ratio equal to preset eye-opening threshold, and eye aspect ratio less than preset eye-opening threshold, the judgment result is passed.

6. The method according to claim 1, characterized in that, The step of adjusting the preset mouth opening threshold based on the average mouth width-to-height ratio of the preceding several consecutive frames in the detected video includes: If the mouth width-to-height ratio of the first few consecutive frames of the detection video does not meet the preset detection threshold, the average mouth width-to-height ratio of the first few consecutive frames of the detection video is matched with the preset detection threshold correspondence table, and the preset mouth opening threshold is adjusted according to the matching result.

7. The method according to claim 1, characterized in that, If the user's liveness detection result is successful, the method further includes: A set of face images is obtained by capturing the face regions in the detected video using a classifier; The face image set is preprocessed to obtain a preprocessed image set; the image preprocessing includes grayscale conversion, histogram equalization, and normalization. The preprocessed image set is used to extract features based on the local binary model histogram algorithm to obtain a feature extraction set. The facial recognition result of the user is determined by comparing and analyzing the extracted feature set with the preset feature database.

8. A liveness detection system for patients with facial paralysis, characterized in that, include: The first module is used to acquire the user's detection video; The detection video includes several frames of images; The second module is used to perform facial key point detection on each frame of the image in the detection video and calculate based on the detection results to determine the width-to-height ratio of the eyes and the width-to-height ratio of the mouth corresponding to each frame of the image. The eye width-to-height ratio includes the width-to-height ratio of the left eye and the width-to-height ratio of the right eye; The third module is used to independently determine the width-to-height ratio of the left eye and the width-to-height ratio of the right eye corresponding to several frames of the image in the video based on a preset eye-opening threshold, and determine the blink detection result of the user based on the judgment result; The fourth module is used to determine the mouth width-to-height ratio of each frame of the image based on the mouth opening threshold, and to determine the user's mouth opening detection result. The mouth opening threshold is obtained by adjusting the preset mouth opening threshold based on the average mouth width-to-height ratio of the first few consecutive frames in the detection video. The fifth module is used to determine the user's liveness detection result based on the blink detection result and the mouth opening detection result.

9. A computer-readable storage medium storing a processor-executable program, characterized in that, The processor-executable program, when executed by the processor, is used to perform the method as described in any one of claims 1 to 7.

10. A liveness detection system for patients with facial paralysis, characterized in that, It includes an image acquisition device and a computer device connected to the image acquisition device; wherein, The image acquisition device is used to acquire the user's detection video; The computer device includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • A Method and System for Identifying Proxy Attendance Based on Image Feature Comparison

    CN122416513A