Artificial intelligence automatic identification system based on facial expressions

Through an artificial intelligence automatic recognition system based on facial expressions, using indicators such as face forward proportion, neutral emotional proportion and head complexity, combined with machine learning technology, the problems of instability of the evaluation results of existing ASD children's identification tools and limitations of invasive detection tools are solved, and efficient and contactless screening for children with ASD is achieved.

CN119993474AActive Publication Date: 2025-05-13EAST CHINA NORMAL UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411842413.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-13
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

The existing children's identification tools for autism spectrum disorder (ASD) have problems with instability in assessment results, gender error and racial errors, and invasive detection tools are highly limited and have limited scope of application.

Method used

The artificial intelligence automatic recognition system based on facial expressions is adopted to guide the target object to generate specific emotions through the audio-visual stimulation module. The head detection module detects facial reaction information, including the face forward proportion, neutral emotion proportion and head complexity, and outputs classification labels in combination with the pre-trained classification model.

Benefits of technology

It has achieved unified and standardized identification in the early stages of children, and objective assessments without contact and stimulation are carried out based on machine learning technology, reducing time and labor costs, and achieving large-scale screening for children with ASD.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993474A_ABST
    Figure CN119993474A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an artificial intelligence automatic recognition system based on facial expressions. Comprising an audio-visual stimulation module used for outputting induction information used for guiding a target object to generate a specific emotion; the head detection module is connected with the information acquisition unit and is used for detecting head reaction information; and the label classification module is connected with the head detection module, and the label classification module is provided with a pre-trained classification model and is used for outputting classification labels according to the face forward proportion data, the neutral emotion proportion data and the head complexity data. Facial expressions are used as specific recognition indexes, unified and standardized recognition can be carried out in the early stage of children, meanwhile, based on the machine learning technology, the children can be objectively evaluated under the non-contact and non-stimulation conditions, the time cost and the labor cost are reduced to the maximum extent, and large-scale screening of ASD children is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence automatic recognition system based on facial expressions. Background Art

[0002] Specifically, there are two types of tools commonly used in identifying children with autism spectrum disorder (ASD), namely standardized identification tools and invasive identification tools. Standardized identification tools are in the form of questionnaires, interviews, observations, etc. However, differences in different identification standards can easily lead to disagreements in the evaluation results. In addition, since screening tools have a high degree of demand for subjective experience and judgment, gender errors and racial errors are more likely to occur in the evaluation process, resulting in a lack of stability in clinical diagnosis. In the early stages of children's development, the subtle differences shown by ASD children are difficult to be discovered and identified, resulting in a late diagnosis age.

[0003] Invasive detection tools are based on facial electromyography (FEMG) to identify muscle changes in ASD children and typically developing individuals by measuring electrical impulses from facial muscle contractions. However, this tool is only applicable to the corrugator and zygomatic major muscles and cannot define other facial movements, so its results are limited. In addition, since the electromyographic signal is weak, human hair and skin keratin will hinder its signal conduction. Therefore, before the actual measurement of wearing electrode patches, it is necessary to eliminate the influence of oils and other factors with alcohol. Overall, this process is invasive and not suitable for highly sensitive ASD people. Summary of the invention

[0004] The purpose of the present invention is to provide an artificial intelligence automatic recognition system based on facial expressions to solve the above technical problems;

[0005] The technical problem solved by the present invention can be achieved by adopting the following technical solutions:

[0006] The artificial intelligence automatic recognition system based on facial expression includes:

[0007] An audio-visual stimulation module outputs induced information to guide the target object to produce specific emotions;

[0008] A head detection module is connected to the information acquisition unit for acquiring head reaction information generated by the target object based on the induced information, and the head detection module is used to detect the head reaction information, including:

[0009] A face-forward ratio detection unit, connected to the information acquisition unit, for calculating the ratio of the face-forward time duration to the total time duration, and outputting face-forward ratio data;

[0010] A neutral emotion ratio detection module, connected to the information collection unit, is used to calculate the ratio of neutral emotion duration to total duration and output neutral emotion ratio data;

[0011] A head complexity detection unit, connected to the information collection unit, for calculating the head complexity of the target object and outputting head complexity data;

[0012] A label classification module is connected to the head detection module, and the label classification module has a pre-trained classification model, and the classification model is used to output a classification label according to the face-front ratio data, the neutral emotion ratio data and the head complexity data.

[0013] Preferably, the audio-visual stimulation module guides the target object to produce specific emotions including at least happiness, surprise and anger, the information collection unit is a camera set towards the face of the target object, and the head reaction information is a head reaction video taken by the camera.

[0014] Preferably, the face forward proportion detection unit has a head posture estimation model based on a convolutional neural network, the head posture estimation model traverses each video frame of the head reaction information, outputs a head posture feature corresponding to each video frame, and marks the corresponding picture according to the head posture feature, marking it as a forward-looking video frame and a non-forward-looking video frame;

[0015] The face-front proportion data is obtained based on statistics of the proportion of the forward-looking video frames in the total number of frames.

[0016] Preferably, the condition that needs to be satisfied by the forward-looking video frame is that the included angle between the pitch angle and the yaw angle of the head posture of the target object is less than or equal to a preset angle.

[0017] Preferably, the neutral emotion ratio detection module has an emotion detection model based on a convolutional neural network, and the emotion detection model is pre-trained with a data set containing facial expression information;

[0018] The emotion detection model traverses each video frame of the head reaction information, outputs the emotion label corresponding to each video frame, and marks the corresponding video frames according to the emotion label as neutral emotion video frames and other emotion video frames; based on the statistics of the proportion of the neutral emotion video frames in the total number of frames, the neutral emotion proportion data is obtained.

[0019] Preferably, the head complexity data output by the head complexity detection unit includes facial complexity, line of sight complexity and head posture complexity.

[0020] Preferably, the header complexity detection unit includes:

[0021] A head posture estimation model, connected to the information acquisition unit, for outputting the head posture of each video frame in the head reaction information;

[0022] A sight line detection model, connected to the information acquisition unit, for detecting a sight line vector of each video frame in the head reaction information;

[0023] A facial key point detection model is connected to the information acquisition unit and is used to detect facial key points of each video frame in the head reaction information.

[0024] Preferably, the header complexity detection unit further includes:

[0025] A head posture complexity calculation subunit, connected to the head posture estimation model, is used to calculate the angular velocity of the head swing according to the head posture, and obtain angular velocity sequence information as the head posture complexity;

[0026] A sight complexity calculation subunit, connected to the sight detection model, for calculating the Euclidean distance between the sight vectors of two adjacent video frames, and obtaining sight difference sequence information as the sight complexity;

[0027] A facial complexity calculation subunit, connected to the facial key point detection model, is used to calculate the Euclidean distance of the facial key point changes between two adjacent video frames, and obtain facial Euclidean distance sequence information as the facial complexity;

[0028] The sequence information processing subunit is connected to the head posture complexity calculation subunit, the line of sight complexity calculation subunit and the facial complexity calculation subunit, and is used to convert the angular velocity sequence information, the line of sight difference sequence information and the facial Euclidean distance sequence information into multi-scale entropy.

[0029] Preferably, the classification model adopts a machine learning model based on a gradient boosting framework and is trained using a binary classification task, and the classification labels output by the classification model include an autism label and a non-autism label.

[0030] Preferably, it also includes a model optimization unit, which is connected to the classification model and is used to update the classification model based on the error between the prediction result of the classification model and the actual label, and is used to construct a new round of model tree based on residual information and optimize the classification model by adjusting the leaf node value.

[0031] Beneficial effects of the present invention: Due to the adoption of the above technical scheme, the present invention uses facial expressions as specific recognition indicators, which can perform unified and standardized recognition in the early childhood. At the same time, based on machine learning technology, it can objectively evaluate children under non-contact and non-stimulation conditions, thereby minimizing time and labor costs and realizing large-scale screening of ASD children. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is an architecture diagram of an artificial intelligence automatic recognition system based on facial expressions in an embodiment of the present invention;

[0033] Figure 2 This is a diagram of the architecture of a header complexity detection unit in an embodiment of the present invention;

[0034] Figure 3 Schematic diagram of head posture characteristics of a character model in an embodiment of the present invention;

[0035] Figure 4 Schematic diagram of the ratio of the faces of typically developing children and autistic children facing forward in an embodiment of the present invention;

[0036] Figure 5 Schematic diagram of the difference in the distribution of neutral emotion ratios between typically developing children and autistic children in an embodiment of the present invention;

[0037] Figure 6a This is a diagram showing the difference in MSE values ​​of facial complexity between typically developing children and autistic children in an embodiment of the present invention;

[0038] Figure 6b This is a graph showing the difference in MSE values ​​of key eye points between children with typical development and children with autism in an embodiment of the present invention;

[0039] Figure 6c This is a graph showing the difference in MSE values ​​of key points of the nose of a typically developing child and autistic child in an embodiment of the present invention;

[0040] Figure 7 This is a graph showing the difference in the changing trend of multi-scale entropy of head movement between typically developed children and autistic children in an embodiment of the present invention;

[0041] Figure 8 This is a schematic diagram of automatic identification of training set data in an embodiment of the present invention;

[0042] Fig. 9 Schematic diagram of automatic identification of test set data in an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0045] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.

[0046] Artificial intelligence automatic recognition system based on facial expression, such as Figure 1 , Figure 2 As shown, including,

[0047] The audio-visual stimulation module 1 outputs the induced information for guiding the target object to produce specific emotions;

[0048] The head detection module 3 is connected to the information acquisition unit 2 for acquiring the head reaction information generated by the target object based on the induced information. The head detection module 3 is used to detect the head reaction information, including:

[0049] The face-forward ratio detection unit 31 is connected to the information collection unit 2 and is used to calculate the ratio of the face-forward time to the total time and output the face-forward ratio data;

[0050] The neutral emotion ratio detection module 32 is connected to the information collection unit 2 and is used to calculate the ratio of the neutral emotion duration to the total duration and output the neutral emotion ratio data;

[0051] The head complexity detection unit 33 is connected to the information collection unit 2 and is used to calculate the head complexity of the target object and output the head complexity data;

[0052] The label classification module 34 is connected to the head detection module 3. The label classification module 34 has a pre-trained classification model, and the classification model is used to output classification labels according to the face forward ratio data, the neutral emotion ratio data and the head complexity data.

[0053] Specifically, the working process of the artificial intelligence automatic recognition system based on facial expressions of the present invention mainly consists of two processes, namely, audio-visual stimulation and automatic detection.

[0054] First, the information collection unit 2, i.e., the camera, records the head reaction of the target object under specific audio-visual stimulation. The target object in the present invention is a child, and the stimulation is an animated video that can cause children to feel happy, surprised, or angry.

[0055] Secondly, it automatically detects various indicators of the individual head, mainly including facial emotions, head posture and head complexity. Among them, facial emotions use the proportion of neutral emotions as the estimation indicator, head posture uses the proportion of the face facing forward as the estimation indicator, and head complexity uses the complexity of the face, line of sight and head posture as the estimation indicator.

[0056] By developing an artificial intelligence automatic recognition system based on facial expressions, ASD children can be screened and warned on a large scale. The system uses facial expressions as specific recognition indicators, which can be uniformly and standardized in early childhood. At the same time, based on machine learning technology, it can objectively evaluate children without contact or stimulation, minimizing time and labor costs, and achieving large-scale screening of ASD children.

[0057] All codes in the present invention are in Python and implemented through the Pytorch framework, and the GPU is RTX3090.

[0058] In a preferred embodiment, the audio-visual stimulation module 1 guides the target object to produce specific emotions including at least happiness, surprise and anger, the information collection unit 2 is a camera set towards the face of the target object, and the head reaction information is a head reaction video taken by the camera.

[0059] In a preferred embodiment, the face forward proportion detection unit 31 has a head posture estimation model 331 based on a convolutional neural network. The head posture estimation model 331 traverses each video frame of the head reaction information, outputs the head posture feature corresponding to each video frame, and marks the corresponding picture according to the head posture feature, marking it as a forward-looking video frame and a non-forward-looking video frame;

[0060] The face forward ratio data is obtained based on the statistical proportion of the forward-looking video frames in the total number of frames.

[0061] Specifically, the face-front ratio detection unit 31 is used to perform face-front ratio data processing, and calculate the ratio of the duration of the child's face facing forward to the total duration in the video sequence.

[0062] First, the face-front ratio detection unit 31 traverses each frame of the video record, and outputs head posture features (yaw, pitch, roll) through the head posture estimation model 3313DDFA based on the convolutional neural network.

[0063] yaw is the yaw angle, the rotation angle around the vertical axis;

[0064] Pitch is the pitch angle, the rotation angle around the horizontal axis;

[0065] Roll is the roll angle, the rotation angle around the longitudinal axis.

[0066] The angle information of the head posture feature can accurately describe the orientation of the head in three-dimensional space. The model dataset is 300W-LP head posture data.

[0067] Specifically, face forward is defined as the angle between the pitch angle and the yaw angle of the head posture is less than or equal to 20 degrees, such as Figure 3 As shown, Figure 3 is a schematic diagram of head posture features. Figure 3 The character model 5 in the figure is only used as a reference. In actual scenes, the glasses are usually removed and only the head features are retained.

[0068] In a preferred embodiment, the condition that needs to be met for the video frame marked as looking forward is that the included angle of the pitch angle and the yaw angle of the head posture of the target object is less than or equal to a preset angle.

[0069] Specifically, the present invention defines the face facing forward as the angle between the pitch angle and the yaw angle of the head posture is less than or equal to 20 degrees. Based on this quantitative standard, the combination of the two angles is used to determine whether the face is facing forward.

[0070] For example, in a video frame, in the head posture features output by the 3DDFA model, if the angle between the pitch and yaw angles is calculated to be 15 degrees, then the face of this frame is judged to be facing forward; if the angle is 30 degrees, then the face is judged to be not facing forward.

[0071] The ratio of video frames with faces looking forward to those without looking forward in the total number of video frames is counted to obtain the face-looking forward ratio data.

[0072] In a preferred embodiment, the neutral emotion ratio detection module 32 has an emotion detection model based on a convolutional neural network, and the emotion detection model is pre-trained by a data set containing facial expression information;

[0073] The emotion detection model traverses each video frame of the head reaction information, outputs the emotion label corresponding to each video frame, and marks the corresponding video frames according to the emotion label, marking them as neutral emotion video frames and other emotion video frames; based on the statistical proportion of neutral emotion video frames in the total number of frames, the neutral emotion proportion data is obtained.

[0074] Specifically, the neutral emotion ratio detection module 32 is used to perform neutral emotion ratio data processing and calculate the ratio of the neutral emotion duration of children in the video sequence to the total duration.

[0075] The neutral emotion ratio detection module 32 traverses each frame of the video record, and based on the convolutional neural network, outputs an emotion label (such as Neutral, Happy, Sad, Surprise, Fear, Disgust, Anger) through the emotion detection model Affect Net, and its model data set is Affect Net Facial Expression data.

[0076] The neutral emotion ratio detection module 32 classifies the output frame labels into neutral emotions or other emotions, and counts the ratios of the two types of emotion video frames to the total number of video frames.

[0077] More specifically, for the emotion detection model Affect Net, a large amount of data is needed to learn the emotion categories corresponding to different facial expressions.

[0078] The present invention includes the data of facial expression information Affect Net Facial Expression data set to train the Affect Net model. The data set includes face images with different emotion labels (such as Neutral, Happy, Sad, etc.).

[0079] During the training process, the emotion detection model will learn the association between the facial features and emotion labels in these images, and gradually build a mapping relationship from facial features to emotion judgments. In this way, when the emotion labels are output for new images, that is, each frame in the video, it can judge the emotion label corresponding to the facial expression in this frame based on the previously learned mapping relationship.

[0080] The present invention uses the emotion detection model Affect Net to carry out the work. The emotion detection model is built based on a convolutional neural network and trained using Affect Net Facial Expression data. In the present invention, the emotion detection model analyzes and processes each frame of the video record one by one, and then outputs the corresponding emotion label.

[0081] Emotional labels include (Neutral, Happy, Sad, Surprise, Fear, Disgust, Anger), which means that you can determine the general emotional state of the child in each frame.

[0082] After obtaining the emotion label corresponding to each frame of the picture, the next step is to perform classification and statistics. These output frame labels are simply divided into two categories according to the emotion category: neutral emotion and other emotions. Among them, neutral emotion refers to the situation where the emotion label is "Neutral", while other emotions include non-neutral emotion types such as "Happy", "Sad", "Surprise", "Fear", "Disgust", and "Anger". After that, count the number of video frames belonging to neutral emotions and the number of video frames belonging to other emotions, and then calculate their proportions in the total number of frames in the entire video.

[0083] In a preferred embodiment, the head complexity data output by the head complexity detection unit 33 includes facial complexity, line of sight complexity and head posture complexity.

[0084] In a preferred embodiment, the header complexity detection unit 33 includes:

[0085] A head posture estimation model 331, connected to the information acquisition unit 2, for outputting the head posture of each video frame in the head reaction information;

[0086] The sight line detection model 332 is connected to the information acquisition unit 2 and is used to detect the sight line vector of each video frame in the head reaction information;

[0087] The facial key point detection model 333 is connected to the information acquisition unit 2 and is used to detect the facial key points of each video frame in the head reaction information.

[0088] In a preferred embodiment, the header complexity detection unit 33 further includes:

[0089] A head posture complexity calculation subunit 334, connected to the head posture estimation model 331, is used to calculate the angular velocity of the head swing according to the head posture, and obtain angular velocity sequence information as the head posture complexity;

[0090] The sight complexity calculation subunit 335 is connected to the sight detection model 332 and is used to calculate the Euclidean distance between the sight vectors of two adjacent video frames to obtain sight difference sequence information as sight complexity;

[0091] The facial complexity calculation subunit 336 is connected to the facial key point detection model 333 and is used to calculate the Euclidean distance of the facial key point changes between two adjacent video frames, and obtain the facial Euclidean distance sequence information as the facial complexity;

[0092] The sequence information processing subunit 337 is connected to the head posture complexity calculation subunit 334, the line of sight complexity calculation subunit 335 and the facial complexity calculation subunit 336, and is used to convert the angular velocity sequence information, the line of sight difference sequence information and the facial Euclidean distance sequence information into multi-scale entropy.

[0093] Specifically, the head complexity detection unit 33 processes and calculates the facial complexity, line of sight complexity and head posture complexity of the child in the video sequence. In this process, the facial complexity uses the Euclidean distance of the changes in key points of the face between two adjacent frames as an indicator, the line of sight complexity uses the Euclidean distance of the line of sight vectors between two adjacent frames as an indicator, and the head posture complexity uses the angular rate of head swing as an indicator.

[0094] The head complexity detection unit 33 traverses each frame of the video record, and outputs the head posture, 3D gaze vector and facial key points of each frame of the video through the head posture estimation model 331, the gaze detection model 332 and the facial key point detection model 333.

[0095] By calculating the angular rate of head swing, the difference in gaze vector and the Euclidean distance of changes in adjacent facial key points between each frame.

[0096] The sequence information processing subunit 337 processes the angular velocity sequence information, the line of sight difference sequence information and the facial Euclidean distance sequence information into a multi-scale entropy MSE.

[0097] Specifically, for the different data obtained in the previous step, the relevant change indicators between each frame are calculated respectively:

[0098] Head posture complexity index calculation: The angular rate of head swing is calculated to measure the head posture complexity. The angular rate describes the speed of head rotation, which is calculated by dividing the change in head posture angle between adjacent frames by the corresponding time interval.

[0099] Since the time interval between video frames is usually fixed, the angular velocity can be reflected by the angle change.

[0100] Calculation of line of sight complexity index: Calculate the Euclidean distance between the line of sight vectors of two adjacent frames. The Euclidean distance is used to measure the straight-line distance between two points in space;

[0101] The Euclidean distance between the sight vectors of two adjacent frames reflects the magnitude of the change in the sight direction at adjacent moments. If the Euclidean distance is large, it means that the sight direction has changed significantly and the sight change is complex; if the Euclidean distance is small, the sight direction is relatively stable and the complexity is low.

[0102] Facial complexity index calculation: The Euclidean distance of the key point changes between two adjacent frames of the face is used as a measurement index. Since the facial key points contain the coordinate information of the key parts of the face, calculating the Euclidean distance of the coordinate changes of these key points in two adjacent frames can reflect the degree of change of facial morphology at adjacent moments.

[0103] For example, blinking will change the coordinates of key points around the eyes, and smiling will change the key points around the mouth. These changes are quantified by Euclidean distance. The larger the distance, the more frequent and complex the facial morphology changes, and vice versa.

[0104] The angular velocity sequence information (reflecting the changes in head posture complexity), line of sight difference sequence information (reflecting the changes in line of sight complexity) and facial Euclidean distance sequence information (characterizing the changes in facial complexity) calculated previously are further processed and uniformly converted into multi-scale entropy MSE.

[0105] Multiscale entropy can comprehensively and quantitatively represent the overall complexity of children's faces, gazes, and head postures.

[0106] In a preferred embodiment, the classification model adopts a machine learning model based on a gradient boosting framework and is trained using a binary classification task. The classification labels output by the classification model include an autism label and a non-autism label.

[0107] The present invention adopts the LGB model as the classification model, namely the lightweight gradient boosting machine. The LGB model is a machine learning algorithm based on the gradient boosting framework, which is used to solve different types of prediction tasks such as classification and regression, and is particularly suitable for the binary classification task in the application scenario of the present invention: judging autism or non-autism.

[0108] The LGB model follows the basic idea of ​​gradient boosting. It constructs multiple weak learners (usually decision trees) in sequence, trains the next learner based on the prediction residual of the previous learner, and continuously iterates and optimizes, so that the prediction ability of the overall model is gradually improved. Finally, multiple weak learners are combined to form a strong learner to achieve more accurate fitting and prediction of the data.

[0109] Specifically, the present invention first initializes the classification model (LGB model) and inputs variable data such as the face forward ratio, neutral emotion ratio, and complexity into the model. In this process, the model will build an initial decision tree structure based on the set parameters, such as determining the general shape of the tree, node division rules, etc.

[0110] At the same time, optimization algorithms (such as gradient descent-related algorithms, etc.) are used to adjust the parameters in the model so that the loss function value of the model in the initial stage is minimized. The loss function is used to measure the difference between the model prediction result and the true label.

[0111] After the model has preliminary prediction results based on the initial parameters, it is necessary to compare these prediction results with the actual labels, that is, the errors between the known children and the actual situation of whether they have autism.

[0112] The error can be quantified in a variety of ways, such as directly calculating the difference between the predicted value and the true value, or measuring it based on the calculation results of the loss function.

[0113] Then, based on the calculated error information, the principle of gradient boosting is used to update the model parameters in the direction of error reduction.

[0114] For example, by adjusting the decision tree node division threshold, leaf node value and other parameters, the model can reduce errors in the next round of predictions and more accurately determine whether a child has autism.

[0115] By repeatedly calculating errors and updating parameters, the model gradually learns the hidden rules and feature relationships in the data, improving the accuracy of classification.

[0116] like Figures 4 to 7 As shown, based on different biomarkers, the differences between autistic children and normal children can be clearly seen, including Figure 4This is a diagram of the face-forward ratio. Children with autism (ASD) have a higher density when the face-forward ratio is around 0.7, while children with typical development have a higher density when the face-forward ratio is around 0.4.

[0117] Figure 5 The results show the differences in the distribution of neutral emotion ratios between typically developing children and children with autism, among which children with autism (ASD) are most concentrated when the neutral emotion ratio is about 0.2, while children with typically developing children are most concentrated when the neutral emotion ratio is about 0.6.

[0118] Figure 6a The blue curve (Typical) and the red curve (ASD) represent the overall difference of the MSE value of facial complexity. Figure 6b The blue curve (Typical) and the red curve (ASD) represent the MSE value difference of the eye key points of facial complexity. Figure 6c The blue curve (Typical) and the red curve (ASD) represent the MSE value differences of the nose key points of facial complexity.

[0119] Figure 7 The difference in the changing trends of head movement multi-scale entropy (MSE) between the two groups of data of typical development and autism spectrum disorder is shown.

[0120] based on Figures 4 to 7 The present invention allocates training set and test set in a ratio of 8:2. In the training set, the automatic recognition accuracy is 92%, the sensitivity is 93.8%, the specificity is 87.5%, the image threshold is 0.7, the true positive population / detected positive population is 0.94, the harmonic mean of precision and recall is 0.94, and the area under the ROC curve is 0.97. Figure 8 shown.

[0121] In the test set, the automatic recognition accuracy is 67%, the sensitivity is 50%, the specificity is 100%, the image threshold is 0.7, the true positive population / detected positive population is 1.0, the harmonic mean of precision and recall is 0.67, and the area under the ROC curve is 0.88. Fig. 9 shown.

[0122] In a preferred embodiment, it also includes a model optimization unit, which is connected to the classification model and is used to update the classification model based on the error between the prediction result of the classification model and the actual label, and is used to construct a new round of model tree based on residual information and optimize the classification model by adjusting the leaf node value.

[0123] Specifically, the variables of face-front ratio, neutral emotion ratio, and complexity situation are input into the model, and the classification model LGB model is used for binary classification task training to output autism or non-autism labels.

[0124] The execution code is as follows,

[0125] params = {'num_leaves': 3, #Number of leaf nodes

[0126] 'min_data_in_leaf':3,#Data in each leaf node

[0127] 'objective':'binary',#Task: Binary classification

[0128] 'max_depth':-1,#-1: No depth limit

[0129] "boosting_type":"gbdt",#

[0130] "metric":'auc',#metric

[0131] 'random_state':66,#random seed}

[0132] The present invention first initializes the model, inputs data and minimizes its loss, and divides its data into different nodes.

[0133] Then the error between the predicted result and the actual label is determined and the model is updated based on this.

[0134] A new model tree is constructed based on the residuals, and the leaf node values ​​are continuously adjusted to improve the model's ability to fit the data.

[0135] Specifically, the present invention constructs a new model tree based on the residual calculated in the previous step (i.e., the difference between the prediction result of the previous round of model and the true label).

[0136] In the process of building a new tree, the values ​​of leaf nodes will be continuously adjusted to optimize the classification model's ability to fit data samples under different feature combinations, so that the classification model can more accurately judge the category of the child based on input variables such as the proportion of the face facing forward, the proportion of neutral emotions, and the complexity, and more reliably output the label of autism or non-autism, thereby completing the task of auxiliary judgment of the child's autism condition.

[0137] It should be noted that the autism or non-autism label output by the system provided by the present invention is only auxiliary data for reference and needs to be input into other external medical diagnosis systems for use later. The output of the present invention is not a diagnosis result.

[0138] The above description is only a preferred embodiment of the present invention, and does not limit the implementation mode and protection scope of the present invention. For those skilled in the art, it should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.

Claims

1. An artificial intelligence automatic recognition system based on facial expressions, characterized in that: include, An audio-visual stimulation module (1) outputs induced information for guiding the target object to produce specific emotions; A head detection module (3) is connected to the information acquisition unit (2) for acquiring head reaction information generated by the target object based on the induced information, and the head detection module (3) is used to detect the head reaction information, including: A face-forward ratio detection unit (31), connected to the information collection unit (2), is used to calculate the ratio of the face-forward time duration to the total time duration, and output face-forward ratio data; A neutral emotion ratio detection module (32), connected to the information collection unit (2), is used to calculate the ratio of the neutral emotion duration to the total duration, and output neutral emotion ratio data; A head complexity detection unit (33), connected to the information collection unit (2), used to calculate the head complexity of the target object and output head complexity data; A label classification module (4) is connected to the head detection module (3), and the label classification module (4) is provided with a pre-trained classification model, and the classification model is used to output a classification label according to the face forward ratio data, the neutral emotion ratio data and the head complexity data.

2. The artificial intelligence automatic recognition system based on facial expression according to claim 1 is characterized in that: The audio-visual stimulation module (1) guides the target object to produce specific emotions, including at least happiness, surprise and anger; the information collection unit (2) is a camera arranged toward the face of the target object; and the head reaction information is a head reaction video captured by the camera.

3. The artificial intelligence automatic recognition system based on facial expression according to claim 1 is characterized in that: The face forward proportion detection unit (31) is provided with a head posture estimation model (331) based on a convolutional neural network, wherein the head posture estimation model (331) traverses each video frame of the head reaction information, outputs a head posture feature corresponding to each video frame, and marks the corresponding picture according to the head posture feature, marking it as a forward-looking video frame and a non-forward-looking video frame; The face-front proportion data is obtained based on statistics of the proportion of the forward-looking video frames in the total number of frames.

4. The artificial intelligence automatic recognition system based on facial expression according to claim 3 is characterized in that: The condition that needs to be satisfied by the forward-looking video frame is that the included angle of the pitch angle and the yaw angle of the head posture of the target object is less than or equal to a preset angle.

5. The artificial intelligence automatic recognition system based on facial expression according to claim 1 is characterized in that: The neutral emotion ratio detection module (32) is provided with an emotion detection model based on a convolutional neural network, and the emotion detection model is pre-trained with a data set containing facial expression information; The emotion detection model traverses each video frame of the head reaction information, outputs the emotion label corresponding to each video frame, and marks the corresponding video frames according to the emotion label as neutral emotion video frames and other emotion video frames; based on the statistics of the proportion of the neutral emotion video frames in the total number of frames, the neutral emotion proportion data is obtained.

6. The artificial intelligence automatic recognition system based on facial expression according to claim 1 is characterized in that: The head complexity data output by the head complexity detection unit (33) includes facial complexity, line of sight complexity and head posture complexity.

7. The artificial intelligence automatic recognition system based on facial expression according to claim 6 is characterized in that: The header complexity detection unit (33) comprises: A head posture estimation model (331), connected to the information acquisition unit (2), for outputting the head posture of each video frame in the head reaction information; A sight line detection model (332), connected to the information acquisition unit (2), for detecting a sight line vector of each video frame in the head reaction information; A facial key point detection model (333) is connected to the information acquisition unit (2) and is used to detect facial key points of each video frame in the head reaction information.

8. The artificial intelligence automatic recognition system based on facial expression according to claim 7 is characterized in that: The header complexity detection unit (33) further comprises: A head posture complexity calculation subunit (334), connected to the head posture estimation model (331), is used to calculate the angular velocity of the head swing according to the head posture, and obtain angular velocity sequence information as the head posture complexity; A sight line complexity calculation subunit (335), connected to the sight line detection model (332), is used to calculate the Euclidean distance between the sight line vectors of two adjacent video frames, and obtain sight line difference sequence information as the sight line complexity; A facial complexity calculation subunit (336), connected to the facial key point detection model (333), is used to calculate the Euclidean distance of the facial key point changes between two adjacent video frames, and obtain facial Euclidean distance sequence information as the facial complexity; The sequence information processing subunit (337) is connected to the head posture complexity calculation subunit (334), the line of sight complexity calculation subunit (335) and the facial complexity calculation subunit (336), and is used to convert the angular velocity sequence information, the line of sight difference sequence information and the facial Euclidean distance sequence information into multi-scale entropy.

9. The artificial intelligence automatic recognition system based on facial expression according to claim 1 is characterized in that: The classification model adopts a machine learning model based on a gradient boosting framework and is trained using a binary classification task. The classification labels output by the classification model include an autism label and a non-autism label.

10. The artificial intelligence automatic recognition system based on facial expression according to claim 1 is characterized in that: It also includes a model optimization unit, which is connected to the classification model and is used to update the classification model based on the error between the prediction result of the classification model and the actual label, and is used to construct a new round of model tree based on residual information and optimize the classification model by adjusting the leaf node value.

Citation Information

Patent Citations

  • Autism evaluation device and system based on parrot talk language paradigm behavior analysis

    CN110353703A

  • Early autism screening system based on human-computer interaction

    CN114974572A

  • Cognitive disorder recognition method based on facial information fusion

    CN119007273A

  • Methods, systems, and computer readable media for automated behavioral assessment

    US11158403B1

  • Facial expression detection for screening and treatment of affective disorders

    US20200151439A1