A method for quantitatively assessing the facial condition of a parkinson's disease patient
By constructing a dataset for assessing the condition of masked faces, and utilizing various continuous facial expression videos and facial key point information, the problem of inaccurate assessment of masked faces in existing technologies has been solved, achieving higher assessment accuracy.
Patent Information
- Application Number
- CN202410659613.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2044-05-27
AI Technical Summary
Existing technologies for assessing mask-like facial features in Parkinson's disease patients suffer from several issues: reliance on localized facial expression analysis leads to missed key information, and head and camera movement affects assessment accuracy, resulting in inaccurate quantitative assessments.
A dataset for constructing a mask-face disease assessment model was developed by acquiring multiple continuous facial expression videos, detecting facial key point information, extracting temporal features of facial features, and fusing them to train the assessment model to improve assessment accuracy.
It improves the probability of capturing key facial expressions, reduces the interference of head movement and camera movement on facial expression changes, and improves the accuracy of assessing the severity of mask-like face syndrome.
Smart Images

Figure CN118522053B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Parkinson's disease medical technology, more specifically to the field of machine learning technology, and more specifically to a quantitative assessment method for mask-like facial features in Parkinson's disease patients. Background Technology
[0002] As Parkinson's disease progresses, subjects may exhibit reduced facial expressions, fixed gaze, and decreased blinking, resulting in stiff facial expressions, resembling a mask on the face, medically termed "masked face."
[0003] Traditional diagnosis of mask-like facial features in Parkinson's disease patients typically relies on experienced physicians observing changes in blink frequency and facial expressions during the patient's consultation, noting the presence of spontaneous smiling and lip separation, and using the MDS Unified-Parkinson's Disease Rating Scale (MDS-UPDRS) for quantitative assessment. However, the diagnostic results are highly subjective and influenced by the physician's experience.
[0004] Existing methods use mobile phone cameras to collect facial expression information and combine traditional machine learning or deep learning to extract, analyze and model the facial expressions of subjects. These methods can determine whether a subject has Parkinson's disease, but they cannot quantitatively assess the severity of the subject's mask-like face.
[0005] Relying solely on localized facial expressions (eyes or mouth) for diagnosis, or analyzing expressions for too short a time and missing crucial information, will lead to inaccurate quantitative assessments of facial expressions. This can result in misdiagnosis of healthy individuals and missed diagnoses of Parkinson's disease patients with mask-like facial features. For example, the blink rate of healthy individuals is 15-20 times per minute, while that of Parkinson's disease patients with mask-like facial features is 10 times per minute or less. However, healthy individuals may not blink for short periods, such as 5 seconds. Therefore, analyzing localized expressions in a short time can lead to numerous misdiagnoses of healthy individuals. Methods relying on optical flow information or frame difference motion information obtained by differentiating frames after face detection to assess facial movement and diagnose mask-like facial features are significantly affected by changes in optical flow or frame difference information caused by head movement or camera movement compared to subtle changes in facial expressions. This can severely interfere with the diagnosis and lead to inaccurate assessments.
[0006] Therefore, existing methods that rely solely on localized facial expressions of the subjects are prone to missing crucial facial information, leading to inaccurate quantitative assessments of facial expressions. Furthermore, current methods that depend on optical flow or frame difference information in video data to determine mask-like facial features are susceptible to the influence of head and neck movements, as well as camera movement, which can also result in inaccurate assessments.
[0007] It should be noted that the background information presented here is only for illustrating relevant information about the present invention to aid in understanding the technical solutions of the present invention, and does not imply that the relevant information is necessarily prior art. In the absence of evidence indicating that the relevant information was disclosed before the filing date of this invention, the relevant information should not be considered prior art. Summary of the Invention
[0008] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a quantitative assessment method for mask-like face symptoms in Parkinson's disease patients.
[0009] The objective of this invention is achieved through the following technical solution:
[0010] According to a first aspect of the present invention, a method for constructing a dataset for training a masked face disease assessment model is provided, comprising: step S1, acquiring facial expression videos of objects at each of multiple masked face disease levels, including multiple expression videos collected of the object when the object performs actions according to multiple preset continuous facial expression action requirements, wherein the object is a sample object; step S2, detecting facial key point information for each expression video for each object, extracting temporal features of facial features of the corresponding expression video based on the facial key point information, and fusing the temporal features of facial features of all expression videos of the object to obtain fused temporal features of facial features; step S3, constructing a dataset comprising multiple samples and corresponding labels, each sample being the fused temporal features of facial features of an object, and the label corresponding to the sample being the masked face disease level of the object corresponding to the sample.
[0011] In some embodiments of the present invention, the multiple continuous facial expression actions in step S1 include a first continuous facial expression action from calm to smiling and back to calm, a second continuous facial expression action from calm to surprised and back to calm, a third continuous facial expression action from calm to angry and back to calm, and a fourth continuous facial expression action of blinking. The multiple expression videos include: a first expression video captured on the subject when the subject performs the first continuous facial expression action multiple times within a predetermined time; a second expression video captured on the subject when the subject performs the second continuous facial expression action multiple times within a predetermined time; a third expression video captured on the subject when the subject performs the third continuous facial expression action multiple times within a predetermined time; and a fourth expression video captured on the subject when the subject performs the fourth continuous facial expression action multiple times within a predetermined time.
[0012] In some embodiments of the present invention, the temporal features of facial features of all expression videos of the object in step S2 include: temporal features of the corners of the mouth in a first type of expression video, including the variation features of the relative width of the left and right corners of the mouth; a first temporal feature of the lips in a second type of expression video, including the variation features of the relative height of the upper and lower lips; a temporal feature of the eyebrows and a second temporal feature of the lips in a third type of expression video, wherein the temporal feature of the eyebrows includes the variation features of the relative distance between the eyebrows, and the second temporal feature of the lips includes the variation features of the relative opening amplitude of the lips; and a temporal feature of the left eye, a temporal feature of the right eye, and a third temporal feature of the lips in a fourth type of expression video, wherein the temporal feature of the left eye includes the variation features of the relative distance between the upper and lower eyelids of the left eye, the temporal feature of the right eye includes the variation features of the relative distance between the upper and lower eyelids of the right eye, and the third temporal feature of the lips includes the variation features of the relative opening amplitude of the lips.
[0013] In some embodiments of the present invention, the facial key point information in step S2 includes facial key point information for each image frame of each expression video. The method of extracting the temporal features of facial features of the corresponding expression video includes: extracting multiple reference key point coordinates and multiple effective key point coordinates from the facial key point information of each image frame in the corresponding expression video; determining a reference distance based on the multiple reference key point coordinates; determining the relative distance features of facial features of each image frame based on the reference distance and the multiple effective key point coordinates; obtaining initial temporal features of facial features based on the relative distance features of facial features of all image frames in the corresponding expression video; performing linear interpolation on the initial temporal features of facial features to obtain the temporal features of facial features of the corresponding expression video in a predetermined dimension; wherein, the temporal features of facial features of all expression videos are stacked or spliced to obtain fused temporal features of facial features.
[0014] In some embodiments of the present invention, the plurality of reference key point coordinates include the key point coordinates of the left corner of the eye, the key point coordinates of the right corner of the eye, the key point coordinates of the root of the nose, and the key point coordinates of the tip of the nose. The method of determining the reference distance includes: using the interocular distance calculated based on the key point coordinates of the left corner of the eye and the key point coordinates of the right corner of the eye as the lateral reference distance; or using the nose length calculated based on the key point coordinates of the root of the nose and the key point coordinates of the tip of the nose as the longitudinal reference distance.
[0015] In some embodiments of the present invention, for the first type of facial expression video, the coordinates of multiple effective key points corresponding to each image frame include: multiple key point coordinates of the left corner of the mouth and multiple key point coordinates of the right corner of the mouth; for the first type of facial expression video, the relative distance features of facial features in each image frame include the relative width of the left and right corners of the mouth, and the method for determining the relative distance features of facial features includes: determining the width of the left and right corners of the mouth based on the coordinates of multiple key points of the left and right corners of the mouth and multiple key point coordinates of the right corner of the mouth, and obtaining the relative width of the left and right corners of the mouth based on the ratio of the width of the left and right corners of the mouth to the horizontal reference distance.
[0016] In some embodiments of the present invention, for the second type of facial expression video, the multiple effective key point coordinates corresponding to each image frame include multiple key point coordinates of the upper lip and multiple key point coordinates of the lower lip; for the second type of facial expression video, the relative distance features of facial features in each image frame include the relative height of the upper and lower lips, and the method for determining the relative distance features of facial features includes: determining the height of the upper and lower lips based on the multiple key point coordinates of the upper lip and multiple key point coordinates of the lower lip, and obtaining the relative height of the upper and lower lips based on the ratio of the height of the upper and lower lips to the longitudinal reference distance; wherein, the minimum value among all the relative heights of the upper and lower lips corresponding to all image frames is taken as the relative height of the upper and lower lips in the object's calm state.
[0017] In some embodiments of the present invention, for the third type of facial expression video, the coordinates of multiple effective key points corresponding to each image frame include: the coordinates of the key point of the left eyebrow tail, the coordinates of the key point of the right eyebrow tail, the coordinates of multiple key points of the upper lip, and the coordinates of multiple key points of the lower lip; for the third type of facial expression video, the relative distance features of the facial features in each image frame include the relative distance between the eyebrows and the relative opening of the lips, and the method for determining the relative distance features of the facial features includes: determining the distance between the eyebrows based on the coordinates of the key point of the left eyebrow tail and the key point of the right eyebrow tail, and obtaining the relative distance between the eyebrows based on the ratio of the distance between the eyebrows to the horizontal relative distance; determining the height of the upper and lower lips based on the coordinates of multiple key points of the upper lip and the multiple key points of the lower lip, and obtaining the relative height of the upper and lower lips in the corresponding image frame based on the ratio of the height of the upper and lower lips to the vertical reference distance; obtaining the relative height of the upper and lower lips in the calm state, and obtaining the relative opening of the lips based on the difference between the relative height of the upper and lower lips in the corresponding image frame and the relative height of the upper and lower lips in the calm state.
[0018] In some embodiments of the present invention, for the fourth type of facial expression video, the multiple effective key point coordinates corresponding to each image frame include: multiple key point coordinates of the left eye, multiple key point coordinates of the right eye, multiple key point coordinates of the upper lip, and multiple key point coordinates of the lower lip; for the fourth type of facial expression video, the relative distance features of facial features in each image frame include the relative distance between the upper and lower eyelids of the left eye, the relative distance between the upper and lower eyelids of the right eye, and the relative opening amplitude of the lips; the method for determining the relative distance features of facial features includes: determining the distance between the upper and lower eyelids of the left eye based on the multiple key point coordinates of the left eye, and determining the distance between the upper and lower eyelids of the left eye and the longitudinal reference... The ratio of distances is used to obtain the relative distance between the upper and lower eyelids of the left eye; the distance between the upper and lower eyelids of the right eye is determined based on the coordinates of multiple key points of the right eye, and the relative distance between the upper and lower eyelids of the right eye is obtained based on the ratio of the distance between the upper and lower eyelids of the right eye to the longitudinal reference distance; the height of the upper and lower lips is determined based on the coordinates of multiple key points of the upper and lower lips, and the relative height of the upper and lower lips in the corresponding image frame is obtained based on the ratio of the height of the upper and lower lips to the longitudinal reference distance; the relative height of the upper and lower lips in the calm state is obtained, and the relative opening amplitude of the lips is obtained based on the difference between the relative height of the upper and lower lips in the corresponding image frame and the relative height of the upper and lower lips in the calm state.
[0019] According to a second aspect of the present invention, a training method for a masked face disease assessment model is provided, comprising: obtaining a training set generated based on a dataset obtained by the method described in the first aspect of the present invention, which includes multiple samples and corresponding labels, wherein each sample is a fused temporal feature of facial features of an object, and the label corresponding to the sample is the masked face disease level of the corresponding object; training the model using the training set to assess the masked face disease level based on input samples, and updating the parameters of the model based on the loss calculated according to the difference between the assessed masked face disease level and the corresponding label, thereby obtaining a trained masked face disease assessment model.
[0020] According to a third aspect of the present invention, a method for quantitatively assessing mask face condition in Parkinson's disease patients is provided, comprising: acquiring fused temporal features of facial features of a subject to be assessed, wherein the subject to be assessed is a Parkinson's disease patient; and assessing the mask face condition level of the subject to be assessed based on the fused temporal features of facial features using a trained mask face condition assessment model obtained according to the method of the second aspect of the present invention.
[0021] According to a fourth aspect of the present invention, a quantitative assessment system for masked face syndrome in Parkinson's disease patients is provided, comprising: a data acquisition module for acquiring multiple facial images of the subject being assessed when performing actions according to preset requirements for multiple continuous facial expression movements, thereby obtaining multiple expression videos of the subject; a data processing module for detecting facial key point information for each expression video, extracting temporal features of facial features in the corresponding expression video based on the facial key point information, fusing the temporal features of facial features in all expression videos of the subject, thereby obtaining fused temporal features of facial features of the subject being assessed; and a trained masked face syndrome assessment model obtained according to the method described in the second aspect of the present invention for assessing the severity of masked face syndrome based on the fused temporal features of facial features of the subject being assessed.
[0022] According to a fifth aspect of the present invention, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions; wherein the one or more processors are configured to implement the steps of the method described in any one of the first, second, and / or third aspects of the present invention by executing the executable instructions.
[0023] Compared with the prior art, the advantages of the present invention are as follows:
[0024] The method of this invention first acquires facial expression videos of objects at each level, including multiple expression videos captured when the object performs actions according to preset continuous facial expression requirements, thus improving the probability of capturing key expressions. Second, it detects facial key point information for each object in each expression video, extracts temporal features of facial features from the corresponding expression video based on the facial key point information, and fuses the temporal features of facial features from all expression videos of the object to obtain fused temporal features of facial features. This solves the problem of inaccurate assessment of mask-like face syndrome caused by interference from head movement or camera movement on facial expression changes. Finally, the fused temporal features of facial features of one object are used as a sample in the dataset. The resulting dataset, containing multiple samples, has significantly improved sample quality, thereby improving the accuracy of the mask-like face syndrome assessment model for Parkinson's disease patients when training the model using this dataset. Attached Figure Description
[0025] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0026] Figure 1 This is a schematic diagram of a method for constructing a dataset for training a masked face disease assessment model according to an embodiment of the present invention;
[0027] Figure 2This is a schematic diagram illustrating three facial expression changes from calm to smiling and back to calm within 24 seconds, collected from a subject according to an embodiment of the present invention.
[0028] Figure 3 A schematic diagram illustrating three facial expression changes from calm to surprise and back to calm within 24 seconds, collected from a subject according to an embodiment of the present invention.
[0029] Figure 4 A schematic diagram illustrating three facial expression changes from calm to angry and back to calm within 24 seconds, based on an embodiment of the present invention.
[0030] Figure 5 This is a schematic diagram of facial expression changes of a subject performing multiple blinking actions within 15 seconds, collected according to an embodiment of the present invention.
[0031] Figure 6 This is a schematic diagram of facial landmark information detected by a facial landmark detection method according to an embodiment of the present invention.
[0032] Figure 7 A schematic diagram of the curve showing the relative width of the corners of the mouth changing over time in a video of a first facial expression, which is a change from calm to smiling and back to calm, according to an embodiment of the present invention.
[0033] Figure 8 A schematic diagram of the curve showing the change in the relative height of the upper and lower lips over time in a second facial expression video, which corresponds to the change in facial expression from calm to surprise and back to calm, according to an embodiment of the present invention.
[0034] Figure 9 A schematic diagram of the curve showing the relative distance between the eyebrows over time in a video of a third facial expression, which changes from calm to anger and back to calm, according to an embodiment of the present invention.
[0035] Figure 10 A schematic diagram of the relative lip opening amplitude over time in a video of an expression change from calm to anger and back to calm, according to an embodiment of the present invention.
[0036] Figure 11 A schematic diagram of the curve showing the change in the relative distance between the upper and lower eyelids of the left eye over time in a video of a fourth facial expression video corresponding to blinking, according to an embodiment of the present invention.
[0037] Figure 12 A schematic diagram of the curve showing the relative distance between the upper and lower eyelids of the right eye over time in a video of a fourth facial expression video corresponding to blinking, according to an embodiment of the present invention.
[0038] Figure 13This is a schematic diagram showing the curve of the relative opening of the lips over time in a video of a fourth facial expression corresponding to blinking, according to an embodiment of the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.
[0040] As mentioned in the background section, existing methods that rely solely on local facial expressions of the subject are prone to missing key facial information, leading to inaccurate quantitative assessments of the subject's facial expressions. Furthermore, existing methods that rely on optical flow or frame difference information in video data to determine mask-like facial conditions are affected by the subject's head and neck movements, as well as the movement of the phone's camera, which can also lead to inaccurate assessments of mask-like facial conditions.
[0041] To address the aforementioned problems, the inventors provide a quantitative assessment method for mask-like facial features in Parkinson's disease patients. First, by acquiring facial expression videos of individuals at each of the various mask-like facial feature severity levels, the method includes multiple expression videos captured when the individual performs actions according to preset continuous facial expression requirements. This solves the problem of missing key expressions when relying on local expression analysis, and improves the probability of capturing key expressions by acquiring multiple expression actions. Second, for each individual, facial keypoint information is detected for each expression video. Based on the facial keypoint information, temporal features of facial features corresponding to the expression video are extracted. These temporal features of facial features from all expression videos of the individual are then fused to obtain fused temporal features of facial features. This solves the problem of inaccurate mask-like facial feature assessment caused by interference from head movement or camera movement on facial expression changes. Finally, the fused temporal features of facial features from one individual are used as a sample in a dataset. The resulting dataset, containing multiple samples, has significantly improved sample quality, thereby improving the accuracy of the mask-like facial feature assessment model for Parkinson's disease patients when training the model using this dataset.
[0042] To better understand this invention, the following detailed description, in conjunction with the accompanying drawings and specific embodiments, covers four aspects: dataset, model structure, model training process, and application scenarios.
[0043] I. Dataset
[0044] According to one embodiment of the present invention, a method for constructing a dataset for training a mask-face disease assessment model is provided, see [link to relevant documentation]. Figure 1This is a flowchart illustrating the method for constructing a dataset for training a masked face disease assessment model. The method includes steps S1, S2, and S3:
[0045] In step S1, for multiple mask face disease levels, facial expression videos of the object under each level are acquired. These include multiple expression videos of the object when it performs actions according to multiple preset continuous facial expression action requirements, wherein the object is a sample object.
[0046] According to one embodiment of the present invention, the severity of mask-like facial expression is classified into five categories: normal, mild, moderate, and severe. The severity level in this invention is determined manually by senior physicians based on the patient's mask-like facial expression manifestations, according to the MDS-UPDRS mask-like facial expression scoring standard. Traditional diagnosis of mask-like facial expression symptoms in Parkinson's disease patients typically relies on experienced physicians observing the patient's state during and without conversation using the MDS-UPDRS scale. This includes observing the patient's blink frequency, the absence of facial expressions, spontaneous smiling, and lip separation, thereby quantitatively assessing the severity of facial expression symptoms in Parkinson's disease. The severity is categorized into five levels: 0, 1, 2, 3, and 4. 0: Normal, normal facial expressions. 1: Mild, mild mask-like facial expression, with only a decrease in blink frequency. 2: Mild, in addition to a decrease in blink frequency, there is also a decrease in lower facial expression, i.e., reduced movement around the mouth, such as a decrease in spontaneous smiling, but the lips do not open. 3: Moderate, mask-like facial expression present, with the lips sometimes opening even when the mouth is still. 4 indicates: severe, with a mask-like face, and the lips are mostly open when the mouth is still.
[0047] According to one embodiment of the present invention, the multiple continuous facial expression actions in step S1 include a first continuous facial expression action from calm to smiling and back to calm, a second continuous facial expression action from calm to surprised and back to calm, a third continuous facial expression action from calm to angry and back to calm, and a fourth continuous facial expression action of blinking. The multiple expression videos include: a first expression video captured on the subject when the subject performs the first continuous facial expression action multiple times within a predetermined time; a second expression video captured on the subject when the subject performs the second continuous facial expression action multiple times within a predetermined time; a third expression video captured on the subject when the subject performs the third continuous facial expression action multiple times within a predetermined time; and a fourth expression video captured on the subject when the subject performs the fourth continuous facial expression action multiple times within a predetermined time. The technical solution of this embodiment can achieve at least the following beneficial technical effects: by having the subject repeatedly perform facial expression changes from calm to smiling and back to calm, from calm to angry and back to calm, from calm to surprised and back to calm within a predetermined time period, as well as continuous blinking, such as a method of continuous blinking for 15 seconds, or repeatedly performing facial expression changes from calm to smiling and back to calm within 15 seconds, the probability of obtaining key expressions is increased by relying on multiple facial expression changes within a predetermined time period, thereby improving the quality of sample data and thus improving the accuracy of assessing the severity of mask-like facial expression in Parkinson's disease patients.
[0048] According to one embodiment of the present invention, the present invention assesses the severity of mask-like facial condition in subjects (also referred to as sample subjects in this invention) by detecting the flexibility of overall facial muscle movement. Firstly, four facial expressions capable of mobilizing overall facial muscle movement are designed, and subjects are asked to record these expressions using the front-facing camera of a mobile phone. During the recording phase, the subject's face is required to appear within a given frame. The predetermined time can be set to 15 seconds, 24 seconds, 30 seconds, or 35 seconds, etc., and the present invention is not limited to this. For example, with a predetermined time of 24 seconds, the first type of facial expression video requires the subject to complete three facial expression changes from calm to smiling and back to calm within 24 seconds. When smiling, the subject is required to raise the corners of their mouth and show their teeth as much as possible; a calm expression is the expression in a normal, relaxed state. See [link to relevant documentation]. Figure 2 The first type of video shows participants completing three facial expression changes from calm to smiling and back to calm within 24 seconds. The second type of video asks participants to complete three facial expression changes from calm to surprised and back to calm within 24 seconds. When surprised, participants are asked to open their mouths as wide as possible and widen their eyes. (See [link to video]). Figure 3The first type of video shows participants completing three facial expression changes from calm to surprise and back to calm within 24 seconds. The second type of video asks participants to complete three facial expression changes from calm to anger and back to calm within 24 seconds. When angry, participants are asked to lower their inner eyebrows, furrow their brows, and widen their eyes as much as possible. (See [link to video]). Figure 4 This is a diagram illustrating three facial expression changes from calm to angry and back to calm within 24 seconds, captured by the subject. A fourth type of facial expression video was captured, recorded at a predetermined 15-second interval, requiring the subject to blink as many times as possible within that 15-second timeframe. See [link / reference]. Figure 5 This is a diagram illustrating the facial expressions of a subject who blinks multiple times within 15 seconds.
[0049] In step S2, facial key point information for each type of expression video is detected for each object. Based on the facial key point information, the temporal features of facial features of the corresponding expression video are extracted. The temporal features of facial features of all expression videos of the object are fused to obtain fused temporal features of facial features.
[0050] According to an embodiment of the present invention, the temporal features of facial features of all expression videos of the object in step S2 include: temporal features of the corners of the mouth in a first type of expression video, including the variation features of the relative width of the left and right corners of the mouth; a first temporal feature of the lips in a second type of expression video, including the variation features of the relative height of the upper and lower lips; a temporal feature of the eyebrows and a second temporal feature of the lips in a third type of expression video, wherein the temporal feature of the eyebrows includes the variation features of the relative distance between the eyebrows, and the second temporal feature of the lips includes the variation features of the relative opening amplitude of the lips; and a temporal feature of the left eye, a temporal feature of the right eye, and a third temporal feature of the lips in a fourth type of expression video, wherein the temporal feature of the left eye includes the variation features of the relative distance between the upper and lower eyelids of the left eye, the temporal feature of the right eye includes the variation features of the relative distance between the upper and lower eyelids of the right eye, and the third temporal feature of the lips includes the variation features of the relative opening amplitude of the lips. The technical solution of this embodiment can achieve at least the following beneficial technical effects: by detecting facial key points to determine the relative distance changes of facial features and processing them into temporal features that reflect the continuous changes in the subject's facial expressions, the method solves the problem of interference caused by differences in facial features of different subjects, head movement or camera movement on facial expression changes, thereby improving the accuracy of the assessment of mask face condition.
[0051] According to an embodiment of the present invention, the facial key point information in step S2 includes facial key point information for each image frame of each expression video. The method for extracting the temporal features of facial features of the corresponding expression video includes: extracting multiple reference key point coordinates and multiple effective key point coordinates from the facial key point information of each image frame in the corresponding expression video; determining a reference distance based on the multiple reference key point coordinates; determining the relative distance features of facial features of each image frame based on the reference distance and the multiple effective key point coordinates; obtaining initial temporal features of facial features based on the relative distance features of facial features of all image frames in the corresponding expression video; performing linear interpolation on the initial temporal features of facial features to obtain the temporal features of facial features of the corresponding expression video in a predetermined dimension; wherein, the temporal features of facial features of all expression videos are stacked or spliced to obtain fused temporal features of facial features. The technical solution of this embodiment can achieve at least the following beneficial technical effects: select effective key points that can reflect changes in facial expressions, and select reference key points that are less affected by movement to determine the reference distance. Based on the relationship between the reference distance and the effective key points, determine the features of each frame (i.e., the relative distance of facial features), and perform linear interpolation to uniformly align them to the same dimension, thereby ensuring the quality and applicability of the dataset.
[0052] According to one embodiment of the present invention, a facial landmark detection method from the MediaPipe tool library is introduced to detect facial landmark information for each type of facial expression video. During the detection process, facial landmark information for each image frame of each type of facial expression video is detected in real time according to the camera's sampling frame rate, including the coordinates of 468 facial landmarks corresponding to each image frame and the corresponding sampling time. See also... Figure 6 This is a schematic diagram of facial landmark information detected by the facial landmark detection method. In the diagram, p corresponds to any i-th landmark. i This indicates that any two adjacent keypoints are connected. When eye keypoints are not detected: p 159 p 145 p 386 p 374 p 133 p 362 Key points about the nose: p 168 p2, key points of the mouth: p 61 p 291 p 12 p 15 When a subject's face is not fully visible in the video frame, they are reminded to look directly at the camera, correct their posture, and re-participate in the video recording. This ensures that the subject is fully visible in the video frame while preventing behaviors such as excessive head shaking, tilting, lowering, or tilting the head that could affect facial expression judgment. In the diagram, the coordinates of any i-th keypoint are represented by p. i (x i yi ) indicates that x i Represents the horizontal coordinate, y i This represents the vertical coordinate; therefore, for all reference keypoint coordinates corresponding to all facial expression videos, including p... 133 (x 133 y 133 ), p 362 (x 362 y 362 ), p 168 (x 168 y 168 p2(x2, y2); This refers to the coordinates of all valid keypoints corresponding to all facial expression videos, including: p 185 (x 185 y 185 ), p 409 (x 409 y 409 ), p 61 (x 61 y 61 ), p 291 (x 291 y 291 ), p 146 (x 146 y 146 ), p 375 (x 375 y 375 ), p 38 (x 38 y 38 ), p 86 (x 86 y 86 ), p 12 (x 12 y 12 ), p 15 (x 15 y 15 ), p 268 (x 268 y 268 ), p 316 (x 316 y 316 ), p 55 (x 55 y 55 ), p 285 (x 285 y 285 ), p 160 (x 160 y 160 ), p 144 (x 144 y 144 ), p 159(x 159 y 159 ), p 145 (x 145 y 145 ), p 158 (x 158 y 158 ), p 153 (x 153 y 153 ), p 385 (x 385 y 385 ), p 380 (x 380 y 380 ), p 386 (x 386 y 386 ), p 374 (x 374 y 374 ), p 387 (x 387 y 387 ), p 373 (x 373 y 373 ).
[0053] According to one embodiment of the present invention, the coordinates of multiple reference key points include the coordinates of the key point at the left corner of the eye, the key point at the right corner of the eye, the key point at the root of the nose, and the key point at the tip of the nose. The method for determining the reference distance includes: using the interocular distance calculated based on the coordinates of the key points at the left and right corners of the eye as the lateral reference distance; or using the nose length calculated based on the coordinates of the key points at the root of the nose and the tip of the nose as the longitudinal reference distance. The technical solution of this embodiment can achieve at least the following beneficial technical effects: by selecting the distance from the left corner of the eye to the right corner of the eye, which is less affected by movement, i.e., the interocular distance, as the lateral reference distance; and by selecting the distance between the root of the nose and the tip of the nose, which is less affected by movement, i.e., the nose length, as the longitudinal reference distance, the relative distance features of the facial features can be obtained more accurately based on the reference distance regardless of the subject's facial expression, avoiding the interference caused by differences in the subject's facial features and head or camera movement on facial expression changes, and improving the quality of extracted feature data.
[0054] According to one embodiment of the present invention, the coordinates of the key point at the left corner of the eye are p. 133 (x 133 y 133 The coordinates of the key point at the right corner of the eye are p. 362 (x 362 y 362 ), by calculating p 133 (x 133 y 133 ) and p362 (x 362 y 362 The Euclidean distance between the two eyes is w. r The interocular distance is used as a lateral reference distance, and the Euclidean distance is calculated as shown in the following formula (1).
[0055]
[0056] The key point coordinates of the nasal root are p 168 (x 168 y 168 The key point coordinates of the nose tip are p2(x2, y2). This is calculated by... 168 (x 168 y 168 The Euclidean distance between p1(x2, y2) and p2(x2, y2), which is the nose length h. r The nose length is used as the longitudinal reference distance. The method for calculating the nose length using Euclidean distance is shown in the following formula (2).
[0057]
[0058] It should be understood that this is for illustrative purposes only, and p can also be calculated. 133 (x 133 y 133 ) and p 362 (x 362 y 362 The Manhattan distance between the two eyes is used as the interocular distance, and p is calculated. 168 (x 168 y 168 The Manhattan distance between p1(x2, y2) and p2(x2, y2) can be used as the nose length, etc., and can be set according to specific needs. This invention does not impose any restrictions on this. The following uses the Euclidean distance calculation method to explain how to determine the relative distance characteristics of facial features:
[0059] By selecting key points that reflect changes in facial expressions and calculating the ratio of the distance between the selected key point coordinates to a reference distance at each sampling moment, the relative distance between corresponding key points can be avoided, thus avoiding the influence of individual differences in the facial features of the subject and camera and head movement on facial expression changes. Therefore, according to an embodiment of the present invention, for a first type of facial expression video, the coordinates of multiple effective key points corresponding to each image frame include: multiple key point coordinates of the left corner of the mouth and multiple key point coordinates of the right corner of the mouth; for the first type of facial expression video, the relative distance features of facial features in each image frame include the relative width of the left and right corners of the mouth, and the method for determining the relative distance features includes: determining the width of the left and right corners of the mouth based on the coordinates of multiple key points of the left and right corners of the mouth, and obtaining the relative width of the left and right corners of the mouth based on the ratio of the width of the left and right corners of the mouth to the lateral reference distance. The technical solution of this embodiment can at least achieve the following beneficial technical effects: calculating the ratio of the distance between the selected effective key point coordinates of the corners of the mouth in each image frame to a reference distance, as the relative distance between corresponding key points, can avoid the influence of individual differences in the facial features of the subject and camera and head movement on facial expression changes.
[0060] According to one embodiment of the present invention, for a video of an expression change from calm to smiling and back to calm, the change in the width of the left and right corners of the mouth reflects the movement of the perioral muscles during smiling. Therefore, the coordinates of several key points on the left corner of the mouth are selected as follows: p 185 (x 185 y 185 ), p 61 (x 61 y 61 ), p 146 (x 146 y 146 The coordinates of several key points selected at the right corner of the mouth are as follows: p 409 (x 409 y 409 ), p 291 (x 291 y 291 ), p 375 (x 375 y 375 The key point p is calculated using the same method as in the above embodiments. 185 (x 185 y 185 ) to key point p 409 (x 409 y 409 Euclidean distance d(p) between ) 185 p 409 ), key point p 61 (x 61 y 61 ) to key point p291 (x 291 y 291 Euclidean distance d(p) between ) 61 p 291 ), key point p 146 (x 146 y 146 ) to key point p 375 (x 375 y 375 Euclidean distance d(p) between ) 146 p 375 ), and the three Euclidean distances d(p) 185 p 409 ), d(p 61 p 291 ), d(p 146 p 375 The weighted average is then used to obtain the width of the left and right corners of the mouth. The relative width of the left and right corners of the mouth, mouth_width_smiling, is shown in the following formula (3):
[0061]
[0062] When performing a weighted average, the width of the left and right corners of the mouth is equal to That is, the three Euclidean distances d(p) 185 p 409 ), d(p 61 p 291 ), d(p 146 p 375 The corresponding weight parameters are all It should be understood that this is only for illustration, and it can also be represented by d(p) 185 p 409 ), d(p 61 p 291 ), d(p 146 p 375 Set different weight parameters, the sum of the three weight parameters should be 1, for example, the width of the left and right corners of the mouth should be equal to 0.5*d(p 185 p 409 )+0.3*d(p 61 p 291 )+0.2*d(p 146 p 375 The weight parameters are 0.5, 0.3, and 0.2 respectively, and this invention does not impose any restrictions on them.
[0063] For a video with a camera sampling frame rate of fps and a recording duration of t, the total number of image frames is t*fps. For example, the recording duration of the first expression video is 24 seconds, so the (1, 24fps) dimension of the relative width of the left and right corners of the mouth is used as the initial temporal feature of the corners of the mouth in the first expression video. Among them, if fps = 10, the dimension of the initial temporal feature of the corners of the mouth is (1, 24*10). In addition, since the left and right corners of the mouth do not move in the calm state, the relative width of the left and right corners of the mouth in the calm state is as shown in the following formula (4):
[0064]
[0065] Where min represents the minimum value, mouth_width_smiling i This represents the relative width of the left and right corners of the mouth in the i-th image frame, where i represents the image frame number. Since different brands and models of cameras have inconsistent sampling frame rates, this invention uses linear interpolation to align the initial temporal features of the mouth corners of the first expression video (1, 24fps) to (1, 1000) along a predetermined dimension, for example, a dimension of (1, 1000), thus obtaining the temporal features of the mouth corners, i.e., the variation features of the relative width of the left and right corners of the mouth. See also... Figure 7 This is a schematic diagram illustrating the change in the relative width of the corners of the mouth over time in a video showing the first facial expression, from calm to smiling and back to calm. In the diagram, the vertical axis represents the relative width of the left and right corners of the mouth (mouth_width_smiling), which is the width of the left and right corners of the mouth relative to the horizontal reference distance w. r The ratio, with the horizontal axis representing time and the unit being seconds.
[0066] By selecting key points that reflect changes in facial expressions and calculating the ratio of the distance between the selected key point coordinates to the reference distance at each sampling moment, the relative distance between the corresponding key points can be avoided, thus avoiding the influence of individual differences in the facial features of the subject and the effects of camera and head movement on facial expression changes. Therefore, according to an embodiment of the present invention, for the second type of facial expression video, the multiple effective key point coordinates corresponding to each image frame include multiple key point coordinates of the upper lip and multiple key point coordinates of the lower lip; for the second type of facial expression video, the relative distance features of facial features in each image frame include the relative height of the upper and lower lips, and the method for determining the relative distance features of facial features includes: determining the height of the upper and lower lips based on the multiple key point coordinates of the upper and lower lip, and obtaining the relative height of the upper and lower lips based on the ratio of the height of the upper and lower lips to the longitudinal reference distance; wherein, the minimum value among all the relative heights of the upper and lower lips corresponding to all image frames is taken as the relative height of the upper and lower lips in the calm state of the subject. The technical solution of this embodiment can achieve at least the following beneficial technical effects: calculating the ratio of the distance between the effective key points of the lips selected in each image frame to the reference distance, as the relative distance between the corresponding key points, can avoid the individual differences in the appearance of the lips of the object and the influence of camera and head movement on the change of expression.
[0067] According to one embodiment of the present invention, for a second type of facial expression video corresponding to a change in expression from calm to surprise and back to calm, the change in the height of the upper and lower lips reflects the movement of the perioral muscles during surprise. Therefore, the coordinates of multiple key points on the upper lip are selected: p 38 (x 38 y 38 ), p 12 (x 12 y 12 ), p 268 (x 268 y 268 ), and the coordinates of several key points on the lower lip: p 86 (x 86 y 86 ), p 15 (x 15 y 15 ), p 316 (x 316 y 316 The key point p is calculated using the same method as in the above embodiments. 38 (x 38 y 38 ) to key point p 86 (x 86 y 86 The distance d(p) between them 38 p 86 ), key point p 12(x 12 y 12 ) to key point p 15 (x 15 y 15 Euclidean distance d(p) between ) 12 p 15 ), key point p 268 (x 268 y 268 ) to key point p 316 (x 316 y 316 Euclidean distance d(p) between ) 268 p 316 ), and the three Euclidean distances d(p) 38 p 86 ), d(p 12 p 15 ), d(p 268 p 316 The weighted average of the upper and lower lip heights is obtained, which is the same as the first expression video processing method. To avoid the influence of differences in the facial features of the subjects and the camera and head movement, the relative height of the upper and lower lips, mouth_height_surprise, is shown in formula (5).
[0068]
[0069] Among them, the height of the upper and lower lips is equal to Therefore, when performing a weighted average, the three Euclidean distances d(p) 38 p 86 ), d(p 12 p 15 ), d(p 268 p 316 The corresponding weight parameters are all It should be understood that this is only for illustration, and it can also be represented by d(p) 38 p 86 ), d(p 12 p 15 ), d(p 268 p 316 Different weight parameters can be set, and the sum of the three weight parameters should be 1. Similarly, for example, the recording time of the second expression video is 24 seconds, so the feature value of the relative height of the upper and lower lips in the (1, 24fps) dimension is used as the initial lip first temporal feature of the second expression video. In addition, since the lips do not move in the calm state, the relative height of the upper and lower lips in the calm state, mouth_height_calm, can be calculated by formula (6).
[0070]
[0071] Where min represents the minimum value, mouth_height_surprise i This represents the relative height of the upper and lower lips in the i-th image frame, where i represents the frame number. Simultaneously, a linear interpolation method is used to align the initial temporal features of the lips (1, 24fps) in the second expression video to (1, 1000) along a predetermined dimension, for example, (1, 1000), resulting in the first temporal feature of the lips, i.e., the variation feature of the relative height of the upper and lower lips. See [link to documentation]. Figure 8 This is a schematic diagram illustrating the change in the relative height of the upper and lower lips over time in a video showing a second type of facial expression, from calm to surprise and back to calm. In the diagram, the vertical axis represents the relative height of the upper and lower lips (mouth_height_surprise), and the vertical reference distance (h) between the upper and lower lip heights. r The ratio, with the horizontal axis representing time and the unit being seconds.
[0072] By selecting key points that can reflect changes in facial expressions, and calculating the ratio between the coordinates of the selected key points at each sampling time to the reference distance, the relative distance between the corresponding key points can be used to avoid the influence of individual differences in the facial features of the subjects and the effects of camera and head movement on changes in facial expressions. Therefore, according to an embodiment of the present invention, for the third type of facial expression video, the multiple effective key point coordinates corresponding to each image frame include: key point coordinates of the left eyebrow tail, key point coordinates of the right eyebrow tail, multiple key point coordinates of the upper lip, and multiple key point coordinates of the lower lip; for the third type of facial expression video, the relative distance features of the facial features in each image frame include the relative distance between the eyebrows and the relative opening amplitude of the lips, and the method for determining the relative distance features of the facial features includes: determining the distance between the eyebrows based on the key point coordinates of the left eyebrow tail and the right eyebrow tail, and obtaining the relative distance between the eyebrows based on the ratio of the distance between the eyebrows to the horizontal relative distance; determining the height of the upper and lower lips based on the multiple key point coordinates of the upper lip and the multiple key point coordinates of the lower lip, and obtaining the relative height of the upper and lower lips in the corresponding image frame based on the ratio of the height of the upper and lower lips to the vertical reference distance; obtaining the relative height of the upper and lower lips in the calm state, and obtaining the relative opening amplitude of the lips based on the difference between the relative height of the upper and lower lips in the corresponding image frame and the relative height of the upper and lower lips in the calm state. The technical solution of this embodiment can achieve at least the following beneficial technical effects: by calculating the ratio of the distance between the effective key points of eyebrows and lips selected in each image frame to the reference distance, as the relative distance between the corresponding key points, the individual differences in the appearance of the subject's eyebrows and lips, as well as the influence of camera and head movement on facial expression changes, can be avoided.
[0073] According to one embodiment of the present invention, for a third facial expression video corresponding to a change in expression from calm to anger and back to calm, the change in distance from the left eyebrow tip to the right eyebrow tip reflects the movement of the periorbital muscles during anger. Therefore, the key point coordinate p of the left eyebrow tip is selected. 55 (x 55 y 55 The key point coordinates p of the right eyebrow tail 285 (x 285 y 285 ), and calculate key points p 55 (x 55 y 55 ) to key point p 285 (x 285 y 285 Euclidean distance d(p) between ) 55 p 285 ), d(p 55 p 285 The distance between the eyebrows is taken as the distance between the eyebrows. Finally, the relative distance between the eyebrows, eyebrow_distance, is shown in formula (7):
[0074]
[0075] Similarly, for example, if the third type of expression video is 24 seconds long, then the (1, 24fps) dimension of the relative distance between the eyebrows is used as the initial eyebrow temporal feature of the third type of expression video. Simultaneously, using linear interpolation, the initial eyebrow temporal feature (1, 24fps) of the third type of expression video is aligned to (1, 1000) along a predetermined dimension, for example, (1, 1000), to obtain the eyebrow temporal feature, i.e., the change feature of the relative distance between the eyebrows. See also... Figure 9 This is a schematic diagram illustrating the change in the relative distance between the eyebrows over time in a video showing a third facial expression, from calm to anger and back to calm. In the diagram, the vertical axis represents the relative distance between the eyebrows (eyebrow_distance), which is the distance between the eyebrows compared to the horizontal reference distance (w). r The ratio, with the horizontal axis representing time and the unit being seconds.
[0076] Simultaneously referring to the MDS-UPDRS diagnostic criteria regarding lip opening when the mouth is still, for the third type of facial expression video, the change in the distance between the upper and lower lips during anger compared to the calm state reflects the mouth opening state. According to one embodiment of the present invention, the relative height of the upper and lower lips is calculated in the same way as formula (5), since the relative lip opening amplitude of each image frame is the difference between the relative height of the upper and lower lips in the corresponding image frame and the relative height of the upper and lower lips in the calm state. Therefore, the relative lip opening amplitude lip_height_angry is calculated as shown in formula (8):
[0077]
[0078] in, This represents the relative height of the upper and lower lips. Similarly, for example, if the third expression video is 24 seconds long, then the (1, 24fps) dimension of the relative lip opening amplitude is used as the initial second temporal feature of the lips in the third expression video. Simultaneously, using linear interpolation, the initial second temporal feature of the lips in the third expression video (1, 24fps) is aligned to (1, 1000) along a predetermined dimension, for example, (1, 1000), to obtain the second temporal feature of the lips, i.e., the change feature of the relative lip opening amplitude. See also... Figure 10 This is a schematic diagram illustrating the change in the relative opening of the lips over time in a video depicting a third facial expression, from calm to anger and back to calm. In the diagram, the vertical axis represents the relative opening of the lips (lip_height_angry), and the horizontal axis represents time in seconds.
[0079] By selecting key points that reflect changes in facial expressions and calculating the ratio of the distance between the selected key point coordinates to a reference distance at each sampling moment, the relative distance between the corresponding key points can be avoided due to individual differences in the facial features of the subject and the influence of camera and head movement on facial expression changes. Therefore, according to an embodiment of the present invention, for the fourth type of facial expression video, the multiple effective key point coordinates corresponding to each image frame include: multiple key point coordinates of the left eye, multiple key point coordinates of the right eye, multiple key point coordinates of the upper lip, and multiple key point coordinates of the lower lip; for the fourth type of facial expression video, the relative distance features of facial features in each image frame include the relative distance between the upper and lower eyelids of the left eye, the relative distance between the upper and lower eyelids of the right eye, and the relative opening amplitude of the lips. The method for determining the relative distance features of facial features includes: determining the distance between the upper and lower eyelids of the left eye based on the multiple key point coordinates of the left eye, and determining the distance between the upper and lower eyelids of the left eye and the relative opening amplitude of the lips based on the distance between the upper and lower eyelids of the left eye and the reference distance. The relative distance between the upper and lower eyelids of the left eye is obtained by comparing the distances. The distance between the upper and lower eyelids of the right eye is determined based on the coordinates of multiple key points of the right eye. The relative distance between the upper and lower eyelids of the right eye is then obtained by comparing this distance with the longitudinal reference distance. The height of the upper and lower lips is determined based on the coordinates of multiple key points of the upper and lower lips. The relative height of the upper and lower lips in the corresponding image frame is obtained by comparing this height with the longitudinal reference distance. The relative height of the upper and lower lips in a calm state is obtained. The relative opening amplitude of the lips is obtained by comparing the relative height of the upper and lower lips in the corresponding image frame with the relative height of the upper and lower lips in a calm state. This embodiment achieves at least the following beneficial technical effects: by calculating the ratio of the distance between the effective key point coordinates of the eyes and lips selected in each image frame to the reference distance, as the relative distance between the corresponding key points, the individual differences in the appearance of the subject's eyes and lips, as well as the influence of camera and head movement on facial expression changes, can be avoided.
[0080] According to one embodiment of the present invention, for a second facial expression video where the expression change is blinking, the change in the height of the upper and lower eyelids of the left and right eyes reflects the movement of the muscles around the eyes during blinking. Therefore, for the left eye, the coordinates of multiple key points of the left eye are selected: p 160 (x 160 y 160 ), p 144 (x 144 y 144 ), p 159 (x 159 y 159 ), p 145 (x 145 y 145 ), p 158 (x 158 y 158 ), p 153 (x 153 y 153 ), and calculate key points p 160 (x 160 y 160 ) to key point p 144 (x 144 y 144 The Euclidean distance d(p) 160 p 144 ), key point p 159 (x 159 y 159 ) to key point p 145 (x 145 y 145 The Euclidean distance d(p) 159 p 145 ), key point p 158 (x 158 y 158 ) to key point p 153 (x 153 y 153 The Euclidean distance d(p) 158 p 153 For the three Euclidean distances d(p) 160 p 144 ), d(p 159 p 145 ), d(p 158 p 153 A weighted average is performed to obtain the distance between the upper and lower eyelids of the left eye. This distance is then compared with the longitudinal reference distance h. rThe ratio of the upper and lower eyelids of the left eye is used to obtain the relative distance between the upper and lower eyelids of the left eye, i.e., the relative distance between the upper and lower eyelids of the left eye, left_eye_height_blinking, is shown in the following formula (9):
[0081]
[0082] When performing a weighted average, the distance between the upper and lower eyelids of the left eye is equal to... That is, the three Euclidean distances d(p) 160 p 144 ), d(p 159 p 145 ), d(p 158 p 153 The corresponding weight parameters are all It should be understood that this is only for illustration, and it can also be represented by d(p) 160 p 144 ), d(p 159 p 145 ), d(p 158 p 153 Different weight parameters can be set, with the sum of the three weight parameters equal to 1. Similarly, for example, if the recording duration of the fourth expression video is 15 seconds, then the feature value of the relative distance between the upper and lower eyelids of the left eye in the (1, 15fps) dimension can be used as the initial left-eye temporal feature of the third expression video. Simultaneously, this invention uses linear interpolation to align the initial left-eye temporal feature (1, 24fps) of the fourth expression video to (1, 1000) along a predetermined dimension, for example, along (1, 1000), to obtain the left-eye temporal feature, i.e., the change feature of the relative distance between the upper and lower eyelids of the left eye. See also... Figure 11 This is a schematic diagram illustrating the change in the relative distance between the upper and lower eyelids of the left eye over time in a video showing the fourth facial expression corresponding to blinking. In the diagram, the vertical axis represents the relative distance between the upper and lower eyelids of the left eye (left_eye_height_blinking), and the horizontal axis represents time in seconds.
[0083] For the right eye, select the coordinates of multiple key points of the right eye: p 385 (x 385 y 385 ), p 380 (x 380 y 380 ), p 386 (x 386 y 386 ), p 374 (x 374 y 374 ), p 387 (x 387 y 387 ), p 373 (x373 y 373 ), and calculate key points p 385 (x 385 y 385 ) to key point p 380 (x 380 y 380 The Euclidean distance d(p) 385 p 380 ), key point p 386 (x 386 y 386 ) to key point p 374 (x 374 y 374 The Euclidean distance d(p) 386 p 374 ), key point p 387 (x 387 y 387 ) to key point p 373 (x 373 y 373 The Euclidean distance d(p) 387 p 373 For the three Euclidean distances d(p) 385 p 380 ), d(p 386 p 374 ), d(p 387 p 373 A weighted average is performed to obtain the distance between the upper and lower eyelids of the right eye. This distance is then compared with the longitudinal reference distance h. r The ratio of the upper and lower eyelids of the right eye is used to obtain the relative distance between the upper and lower eyelids of the right eye, i.e., the relative distance between the upper and lower eyelids of the right eye, right_eye_height_blinking, as shown in formula (10):
[0084]
[0085] When performing a weighted average, the distance between the upper and lower eyelids of the right eye is equal to... That is, the three Euclidean distances d(p) 385 p 380 ), d(p 386 p 374 ), d(p 387 p 373 The corresponding weight parameters are all It should be understood that this is only for illustration, and it can also be represented by d(p) 385 p 380 ), d(p 386 p 374 ), d(p 387 p 373Different weight parameters can be set, with the sum of the three weight parameters equal to 1. Similarly, for example, if the recording duration of the fourth expression video is 15 seconds, then the feature value of the relative distance between the upper and lower eyelids of the right eye in the (1, 15fps) dimension can be used as the initial right eye temporal feature of the third expression video. Then, a linear interpolation method is used to align the initial right eye temporal feature (1, 24fps) of the fourth expression video to (1, 1000) dimension using a predetermined dimension, for example, (1, 1000), to obtain the right eye temporal feature, i.e., the change feature of the relative distance between the upper and lower eyelids of the right eye. See [link to documentation]. Figure 12 This is a schematic diagram illustrating the change in the relative distance between the upper and lower eyelids of the right eye over time in a video showing the fourth facial expression corresponding to blinking. In the diagram, the vertical axis represents the relative distance between the upper and lower eyelids of the right eye (right_eye_height_blinking), and the horizontal axis represents time in seconds.
[0086] To determine whether there is involuntary mouth opening during blinking, the relative height of the upper and lower lips is calculated using the same method as formula (5) according to one embodiment of the present invention. Since the relative lip opening amplitude of each image frame is the difference between the relative height of the upper and lower lips in the corresponding image frame and the relative height of the upper and lower lips in the calm state, the relative lip opening amplitude lip_height_blinking is calculated as shown in formula (11):
[0087]
[0088] Similarly, for example, if the fourth expression video is 15 seconds long, then the (1, 15fps) dimension feature value of the relative opening amplitude of the lips is used as the initial third temporal feature of the lips in the third expression video. Simultaneously, this invention uses linear interpolation to align the initial third temporal feature of the lips in the fourth expression video (1, 24fps) dimension to (1, 1000) according to a predetermined dimension, for example, according to (1, 1000), to obtain the third temporal feature of the lips, i.e., the change feature of the relative opening amplitude of the lips. See also... Figure 13 This is a schematic diagram illustrating the change in the relative lip opening amplitude over time in a video showing the fourth facial expression corresponding to blinking. In the diagram, the vertical axis represents the relative lip opening amplitude (lip_height_blinking), and the horizontal axis represents time in seconds.
[0089] According to one embodiment of the present invention, during the fusion process, for each subject, the aforementioned temporal features of the corner of the mouth, the first temporal feature of the lips, the temporal features of the eyebrows, the second temporal feature of the lips, the temporal features of the left eye, the temporal features of the right eye, and the third temporal feature of the lips, a total of 7 groups, each with a dimension of (1, 1000), are vertically stacked to obtain a (7, 1000)-dimensional fused temporal feature of facial features that changes over time. Alternatively, the 7 groups of temporal features are horizontally spliced to obtain a (1, 7000)-dimensional fused temporal feature of facial features that changes over time.
[0090] In step S3, a dataset is constructed, comprising multiple samples and corresponding labels. Each sample represents the fused temporal features of facial features of an object, and the label corresponding to the sample indicates the severity level of the mask-like facial condition of the object. This dataset is then used as a training set to train a subsequent mask-like facial condition assessment model, enabling the assessment of the severity level of mask-like facial conditions in Parkinson's patients.
[0091] II. Model Structure
[0092] According to one embodiment of the present invention, the mask-face disease assessment model structure includes a multi-layer neural network, with each layer consisting of multiple neurons. The output of each layer in the multi-layer neural network becomes the input of the next layer, and the output of the last layer is used as the model output. If the sample is a (7, 1000)-dimensional fused temporal feature of facial features that varies over time and is obtained by vertically stacking, the mask-face disease assessment model can use LSTM, Convolutional Neural Network (CNN), ResNet, VGG, InceptionNet, AlexNet, and SENet with an attention mechanism. If the sample is a (1, 7000)-dimensional fused temporal feature of facial features that varies over time and is obtained by horizontally concatenating, the mask-face disease assessment model can use traditional machine learning classifiers such as MLP, SVM, Random Forest, XGBoost, and GBDT.
[0093] III. Model Training Process
[0094] According to an embodiment of the present invention, a training method for a mask face disease assessment model is provided, comprising: obtaining a training set generated based on a dataset obtained by the method described in the above embodiment, which includes multiple samples and corresponding labels, each sample being a fused temporal feature of facial features of an object, and the label corresponding to the sample being the mask face disease level of the corresponding object; training the model using the training set to assess the mask face disease level based on input samples, and updating the parameters of the model based on the loss calculated according to the difference between the assessed mask face disease level and the corresponding label, thereby obtaining the trained mask face disease assessment model.
[0095] According to one embodiment of the present invention, the cross-entropy loss function is used to calculate the difference between the model evaluation result and the label. The training process employs a five-fold cross-validation method, using the dataset obtained in the above embodiment to train the model to evaluate the severity of facial symptom severity based on temporal facial features, performing a five-class classification. The principle of the five-fold cross-validation method is to divide the obtained dataset into five parts, using one part as the validation set for each iteration, while the other four parts serve as the training set. This process is repeated five times, ensuring that each part of the data participates in validation without overlapping with the training set. This method is suitable for model training and validation when the amount of data is limited. In this invention, 80% of the patients' video data is used as the training set each time, and 20% of the patients' video data is used as the validation set, with cross-validation repeated five times until the data for each patient is validated.
[0096] IV. Application Scenarios
[0097] According to an embodiment of the present invention, a quantitative assessment method for mask face syndrome in Parkinson's disease patients is provided, comprising: acquiring fused temporal features of facial features of a subject to be assessed, wherein the subject to be assessed is a Parkinson's disease patient; and assessing the mask face syndrome level of the subject to be assessed based on the fused temporal features of facial features using a trained mask face syndrome assessment model obtained according to the method described in the above embodiment.
[0098] According to an embodiment of the present invention, a quantitative assessment system for masked face syndrome in Parkinson's disease patients is provided, comprising: a data acquisition module for acquiring multiple facial images of the subject being assessed when performing actions according to preset continuous facial expression requirements, thereby obtaining multiple expression videos of the subject; a data processing module for detecting facial key point information for each expression video, extracting temporal features of facial features in the corresponding expression video based on the facial key point information, fusing the temporal features of facial features in all expression videos of the subject to obtain fused temporal features of facial features of the subject being assessed; and a trained masked face syndrome assessment model obtained according to the method described in the above embodiment for assessing the severity of masked face syndrome based on the fused temporal features of facial features of the subject being assessed.
[0099] To verify the beneficial effects of this invention, the following training tests were conducted:
[0100] This invention collected fused temporal features of facial features from 120 subjects (i.e., sample subjects). Among them, 28 were normal individuals, 61 had mild mask-like face symptoms, 26 had mild mask-like face symptoms, 3 had moderate mask-like face symptoms, and 2 had severe mask-like face symptoms. Five-fold cross-validation was used to train and test the model for these patients. Furthermore, during the model training and testing phase, a Long Short-Term Memory (LSTM) network was selected as the classification network to train and test the features using five-class five-fold cross-validation. Each cross-validation training session consisted of 200 epochs. The mask-like face severity assessment results obtained through five-fold cross-validation are shown in Table 1 below.
[0101] Table 1: Assessment Results of Mask-like Face Disease Severity
[0102] Evaluation result = 0 24 3 0 0 0 Evaluation result = 1 4 58 1 0 0 Evaluation result = 2 0 0 25 1 0 Evaluation result = 3 0 0 0 2 0 Evaluation result = 4 0 0 0 0 2
[0103] As can be seen from the evaluation results in the table above, most of the evaluation results are the same as the labels. Therefore, the method of the present invention can achieve accurate evaluation of the sample objects, that is, the evaluation accuracy rate can reach 92.50%, which has high reliability in disease assessment.
[0104] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order, as long as the required function can be achieved.
[0105] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0106] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.
[0107] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for constructing a dataset for training a mask-face disease assessment model, characterized in that, include: Step S1: For multiple levels of mask-like facial expression disorder, acquire facial expression videos of the subject at each level. This includes multiple expression videos captured when the subject performs actions according to preset continuous facial expression commands. The subject is a sample subject. The multiple continuous facial expression commands include: a first type of continuous facial expression command from calm to smiling and back to calm; a second type of continuous facial expression command from calm to surprised and back to calm; a third type of continuous facial expression command from calm to angry and back to calm; and a fourth type of continuous facial expression command involving blinking. The multiple expression videos include: a first expression video captured when the subject performs the first type of continuous facial expression command multiple times within a predetermined time; a second expression video captured when the subject performs the second type of continuous facial expression command multiple times within a predetermined time; a third expression video captured when the subject performs the third type of continuous facial expression command multiple times within a predetermined time; and a fourth expression video captured when the subject performs the fourth type of continuous facial expression command multiple times within a predetermined time. Step S2: For each object, detect the facial key point information for each type of expression video. Extract the temporal features of the facial features of the corresponding expression video based on the facial key point information. Merge the temporal features of the facial features of all expression videos of the object to obtain the fused temporal features of the facial features. The temporal features of the facial features of all expression videos of the object include: the temporal features of the corners of the mouth in the first type of expression video, including the change features of the relative width of the left and right corners of the mouth; the first temporal features of the lips in the second type of expression video, including the change features of the relative height of the upper and lower lips; the temporal features of the eyebrows and the second temporal features of the lips in the third type of expression video, wherein the temporal features of the eyebrows include the change features of the relative distance between the eyebrows, and the second temporal features of the lips include the change features of the relative opening amplitude of the lips; the temporal features of the left eye, the right eye, and the third temporal features of the lips in the fourth type of expression video, wherein the temporal features of the left eye include the change features of the relative distance between the upper and lower eyelids of the left eye, the temporal features of the right eye include the change features of the relative distance between the upper and lower eyelids of the right eye, and the third temporal features of the lips include the change features of the relative opening amplitude of the lips. Step S3: Construct a dataset, which includes multiple samples and corresponding labels. Each sample is a fusion of the temporal features of facial features of an object, and the label corresponding to the sample is the severity level of the mask face disease of the object corresponding to the sample.
2. The method according to claim 1, characterized in that, The facial key point information in step S2 includes the facial key point information of each image frame of each type of expression video, and the method for extracting the temporal features of facial features of the corresponding expression video includes: Extract the coordinates of multiple reference key points and multiple valid key points from the facial key point information of each image frame in the corresponding expression video; The reference distance is determined based on the coordinates of multiple reference key points. The relative distance features of facial features in each image frame are determined based on the reference distance and the coordinates of multiple valid key points. The initial temporal features of facial features are obtained based on the relative distance features of facial features in all image frames in the corresponding expression video. Linear interpolation is performed on the initial facial temporal features to obtain the facial feature temporal features of the corresponding expression video under a predetermined dimension; In this process, the temporal features of facial features from all the facial expression videos are stacked or spliced together to obtain fused temporal features of facial features.
3. The method according to claim 2, characterized in that, The multiple reference key point coordinates include the key point coordinates of the left corner of the eye, the key point coordinates of the right corner of the eye, the key point coordinates of the root of the nose, and the key point coordinates of the tip of the nose. The method for determining the reference distance includes: The interocular distance calculated based on the coordinates of the key points at the left and right corners of the eye is used as the lateral reference distance; or The length of the nose, calculated based on the coordinates of the key points at the root of the nose and the tip of the nose, is used as the longitudinal reference distance.
4. The method according to claim 3, characterized in that, For the first type of facial expression video, the multiple valid key point coordinates corresponding to each image frame include: multiple key point coordinates of the left corner of the mouth and multiple key point coordinates of the right corner of the mouth; For the first type of facial expression video, the relative distance features of facial features in each image frame include the relative width of the left and right corners of the mouth. The method for determining the relative distance features of facial features includes: The widths of the left and right corners of the mouth are determined based on the coordinates of multiple key points at the left and right corners of the mouth. The relative widths of the left and right corners of the mouth are obtained by comparing the widths of the left and right corners of the mouth with the horizontal reference distance.
5. The method according to claim 3, characterized in that, For the second type of facial expression video, the coordinates of multiple valid key points corresponding to each image frame include the coordinates of multiple key points of the upper lip and the coordinates of multiple key points of the lower lip. For the second type of facial expression video, the relative distance features of facial features in each image frame include the relative height of the upper and lower lips. The methods for determining the relative distance features of facial features include: The height of the upper and lower lips is determined based on the coordinates of multiple key points of the upper lip and multiple key points of the lower lip. The relative height of the upper and lower lips is obtained based on the ratio of the height of the upper and lower lips to the longitudinal reference distance. Among them, the minimum value among all relative heights of the upper and lower lips corresponding to all image frames is taken as the relative height of the upper and lower lips in the object's calm state.
6. The method according to claim 5, characterized in that, For the third type of facial expression video, the coordinates of multiple valid key points corresponding to each image frame include: the key point coordinates of the left eyebrow tail, the key point coordinates of the right eyebrow tail, multiple key point coordinates of the upper lip, and multiple key point coordinates of the lower lip. For the third type of facial expression video, the relative distance features of facial features in each image frame include the relative distance between the eyebrows and the relative opening of the lips. The methods for determining the relative distance features of facial features include: The distance between the eyebrows is determined by the coordinates of the key points of the left eyebrow tail and the right eyebrow tail. The relative distance between the eyebrows is obtained by the ratio of the distance between the eyebrows to the horizontal relative distance. The height of the upper and lower lips is determined based on the coordinates of multiple key points of the upper lip and the coordinates of multiple key points of the lower lip. The relative height of the upper and lower lips in the corresponding image frame is obtained based on the ratio of the height of the upper and lower lips to the longitudinal reference distance. The relative height of the upper and lower lips in the calm state is obtained, and the relative opening amplitude of the lips is obtained based on the difference between the relative height of the upper and lower lips in the corresponding image frame and the relative height of the upper and lower lips in the calm state.
7. The method according to claim 5, characterized in that, For the fourth type of facial expression video, the multiple valid key point coordinates corresponding to each image frame include: multiple key point coordinates of the left eye, multiple key point coordinates of the right eye, multiple key point coordinates of the upper lip, and multiple key point coordinates of the lower lip. For the fourth type of facial expression video, the relative distance features of facial features in each image frame include the relative distance between the upper and lower eyelids of the left eye, the relative distance between the upper and lower eyelids of the right eye, and the relative opening amplitude of the lips. The methods for determining the relative distance features of facial features include: The distance between the upper and lower eyelids of the left eye is determined based on the coordinates of multiple key points of the left eye. The relative distance between the upper and lower eyelids of the left eye is obtained by comparing the distance between the upper and lower eyelids of the left eye with the longitudinal reference distance. The distance between the upper and lower eyelids of the right eye is determined based on the coordinates of multiple key points of the right eye. The relative distance between the upper and lower eyelids of the right eye is obtained by comparing the distance between the upper and lower eyelids of the right eye with the longitudinal reference distance. The height of the upper and lower lips is determined based on the coordinates of multiple key points of the upper lip and the coordinates of multiple key points of the lower lip. The relative height of the upper and lower lips in the corresponding image frame is obtained based on the ratio of the height of the upper and lower lips to the longitudinal reference distance. Obtain the relative height of the upper and lower lips in a calm state. Based on the difference between the relative height of the upper and lower lips in the corresponding image frame and the relative height of the upper and lower lips in a calm state, obtain the relative opening amplitude of the lips.
8. A training method for a mask-face disease assessment model, characterized in that, include: A training set is generated based on the dataset obtained by the method of any one of claims 1-7, which includes multiple samples and corresponding labels. Each sample is a fused temporal feature of facial features of an object, and the label corresponding to the sample is the mask face disease level of the corresponding object. The model is trained using the training set to assess the severity of masked face syndrome based on input samples. The parameters of the model are updated based on the loss calculated from the difference between the assessed severity of masked face syndrome and the corresponding label, resulting in a trained masked face syndrome assessment model.
9. A quantitative assessment method for mask-like facial features in Parkinson's disease patients, characterized in that, include: The fused temporal features of facial features of the subject to be evaluated are obtained; the subject to be evaluated is a Parkinson's disease patient. The mask face condition assessment model, trained according to the method of claim 8, is used to assess the mask face condition level of the subject based on the fused temporal features of the facial features.
10. A quantitative assessment system for mask-like facial features in Parkinson's disease patients, characterized in that, include: The data acquisition module is used to acquire multiple facial images of the object to be evaluated when it performs actions according to a variety of preset continuous facial expression requirements, and obtain multiple facial expression videos of the object. The data processing module is used to detect the facial key point information of each type of expression video, extract the temporal features of the facial features of the corresponding expression video based on the facial key point information, and fuse the temporal features of the facial features of all expression videos of the object to obtain the fused temporal features of the facial features of the object to be evaluated. The trained mask face condition assessment model obtained according to the method of claim 8 is used to assess the mask face condition level of the subject based on the fused temporal features of the facial features.
11. A computer-readable storage medium, characterized in that, It contains a computer program that can be executed by a processor to implement the steps of the method according to any one of claims 1 to 9.
12. An electronic device, characterized in that, include: One or more processors; as well as Memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 1 to 9 by executing the executable instructions.