Autism spectrum disorder screening system and method based on eye movement and facial expression
By combining the screening system of eye movements and facial expressions and utilizing multi-scenario testing and training models, the problems of existing technologies relying on the subjective judgment of diagnosticians and the limitations of eye movement characteristics are solved, thus achieving efficient and accurate screening of autism spectrum disorders.
Patent Information
- Application Number
- CN202211107282.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2042-09-09
AI Technical Summary
Existing technologies are highly dependent on the subjective judgment of diagnosticians in the diagnosis of autism spectrum disorders, and the eye movement screening features are limited to eye gaze, resulting in high rates of missed diagnosis and misdiagnosis, low evaluation efficiency and poor accuracy.
A screening system based on eye movements and facial expressions is used. A multi-scenario test paradigm is presented through a display module. Eye movement information and facial videos are collected. Eye movement and expression features are extracted after pre-processing. Screening is performed using a trained screening model, and a comprehensive evaluation is performed combining eye movement and facial features.
It improves the accuracy and efficiency of autism spectrum disorder screening, reduces the rates of missed diagnosis and misdiagnosis, and achieves a more objective and comprehensive assessment.
Smart Images

Figure CN115429271B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the intersection of psychology, medicine, and artificial intelligence, and in particular to a system and method for screening autism spectrum disorder based on eye movements and facial expressions. Background Art
[0002] Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder that is hereditary and lifelong. Its etiology and course are complex, often manifesting at a young age. It is accompanied by difficulties with social interaction and judgment, and is associated with these difficulties. Unlike typically developing children (TD), children with ASD often exhibit issues such as eye avoidance, unusual repetitive movements, predilections, and stereotyped behaviors.
[0003] There are two main methods for diagnosing ASD: scales and instrumentation. Scales provide diagnosticians with professional, highly accurate judgments based on authoritative standards. Instrumentation uses specialized equipment to collect specific data. The data is then analyzed, analyzed, and visualized to explore differences between the subject and control groups, providing effective information for diagnosis and classification.
[0004] However, the use of traditional scales for autism screening places extremely high demands on diagnosticians. Diagnostics must have extensive clinical experience and an in-depth understanding of the developmental history of autism and related symptoms in order to make professional and highly accurate judgments. On the other hand, unlike data-driven methods, diagnostic results based on scales are often highly dependent on the subjective ideas of diagnosticians. Different experiences and different interpretations of scales may lead diagnosticians to give different results. Currently, the equipment for diagnosing ASD is mostly focused on the acquisition of brain imaging data (neuroimaging), posture control patterns, and eye movement data. In addition, studies on autism screening based on eye movements are limited to certain situations, and the extracted features are limited to features related to eye gaze. Research on autistic expressions mainly focuses on the emotion recognition ability of autistic patients (self and others), but their own emotional expressions are rarely paid attention to or reported. Summary of the Invention
[0005] The present application provides an autism spectrum disorder screening system and method based on eye movements and facial expressions to address the problems of missed diagnosis and misdiagnosis caused by defects such as the high reliance on diagnostic personnel in the diagnosis of autism spectrum disorder and the limitation of features extracted by diagnostic equipment to features related to eye gaze.
[0006] In a first aspect, the present application provides an autism spectrum disorder screening system based on eye movements and facial expressions, comprising:
[0007] a display module configured to display a test paradigm, the test paradigm comprising at least one scenario test task for testing different characteristics of a subject.
[0008] a collection module configured to collect eye movement information and facial video of the subject when watching the test paradigm, and send the eye movement information and the facial video to a preprocessing module.
[0009] the preprocessing module is configured to preprocess the facial video according to the eye movement information to obtain a face image frame containing a face frame of the subject corresponding to an eye movement entry in the eye movement information.
[0010] a feature extraction module configured to extract eye movement features and expression features of the subject from the eye movement information and the face image frame, the eye movement features comprising eye gaze features, eye physiological features and overall gaze features, and the expression features being features of an emotion proportion of the subject.
[0011] a screening module configured to input the eye movement features and the expression features into a trained screening model to obtain an autism spectrum disorder screening result of the subject.
[0012] In a second aspect, the present application provides an autism spectrum disorder screening method based on eye movement and facial expression, comprising:
[0013] displaying a test paradigm, the test paradigm comprising at least one scenario test task for testing different characteristics of a subject.
[0014] collecting eye movement information and facial video of the subject when watching the test paradigm, and sending the eye movement information and the facial video to a preprocessing module.
[0015] preprocessing the facial video according to the eye movement information to obtain a face image frame containing a face frame of the subject corresponding to an eye movement entry in the eye movement information.
[0016] extracting eye movement features and expression features of the subject from the eye movement information and the face image frame, the eye movement features comprising eye gaze features, eye physiological features and overall gaze features, and the expression features being features of an emotion proportion of the subject.
[0017] inputting the eye movement features and the expression features into a trained screening model to obtain an autism spectrum disorder screening result of the subject.
[0018] It can be seen from the above technical solutions that the application provides an autism spectrum disorder screening system and method based on eye movement and facial expression. First, a test paradigm with different cognitive tests is played to a subject. Eye information and facial video of the subject during watching the test paradigm are obtained. The facial video is preprocessed according to the eye information to obtain a face image frame containing a face frame of the subject corresponding to an eye movement item in the eye movement information. Eye movement features and expression features of the subject are extracted from the eye movement information and the face image frame. The eye movement features and the expression features are input into a trained screening model to obtain an autism spectrum disorder screening result of the subject output by the screening model. The test paradigm contains multiple scenarios, focuses on multiple behavior features of autism spectrum disorder patients, the screening model is obtained by training eye movement features and facial features of multiple sample subjects watching the test paradigm, a data set containing normal sample subjects and autism spectrum disorder sample subjects is collected, and autism spectrum disorder is screened from two aspects of eye movement and expression, which more objectively and comprehensively evaluates the subject, greatly reduces the misdiagnosis rate and the missed diagnosis rate, improves the screening accuracy and efficiency, and the screening method is simple and easy to implement. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the application, the drawings needed in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0020] Figure 1 A schematic block diagram of an exemplary autism spectrum disorder screening system based on eye movement and facial expression provided for the present embodiment is shown in the figure.
[0021] Figure 2 A schematic diagram of an exemplary face observation task provided for the present embodiment is shown in the figure.
[0022] Figure 3 A schematic diagram of an exemplary repetitive action preference test task provided for the present embodiment is shown in the figure.
[0023] Figure 4 A schematic diagram of another exemplary repetitive action preference test task provided for the present embodiment is shown in the figure.
[0024] Figure 5 A schematic diagram of an exemplary joint attention ability test task provided for the present embodiment is shown in the figure.
[0025] Figure 6 A schematic diagram of another exemplary joint attention ability test task provided for the present embodiment is shown in the figure.
[0026] Figure 7A schematic diagram of an exemplary dynamic social image and dynamic geometric image preference test task provided for this embodiment;
[0027] Figure 8 A schematic block diagram of an exemplary acquisition module provided in this embodiment;
[0028] Figure 9 This is a schematic diagram of an exemplary method of dividing a sub-scene into several regions of interest provided in this embodiment;
[0029] Figure 10 This is a schematic diagram of the key positions of the left and right eyes provided as an example in this embodiment;
[0030] Figure 11 A schematic diagram of an exemplary RMS-based key point mapping process provided in this embodiment;
[0031] Figure 12 A schematic diagram of an exemplary training emotion recognition model provided in this embodiment. DETAILED DESCRIPTION
[0032] In order to make the purpose and implementation of this application clearer, the exemplary implementation of this application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only part of the embodiments of this application, not all of the embodiments.
[0033] Existing methods for diagnosing ASD have extremely high requirements for diagnosticians, and the diagnostic results are highly dependent on the diagnosticians' subjective ideas. In addition, research on autism spectrum disorder screening based on eye movements is limited to certain situations, and the extracted features are limited to related features of eye gaze. These problems result in low efficiency and accuracy in the assessment of autism spectrum disorders and the inability to objectively and comprehensively assess the subjects. The present application provides an autism spectrum disorder screening system and method based on eye movements and facial expressions, which not only improves the efficiency and accuracy of autism spectrum disorder assessment, but also sets up test paradigms for multiple situations, screens for autism spectrum disorders from the two levels of eye movements and expressions, and more objectively and comprehensively assesses subjects, greatly reducing the missed diagnosis rate and misdiagnosis rate.
[0034] In the first aspect, the present application provides an autism spectrum disorder screening system based on eye movements and facial expressions, such as Figure 1 As shown, the autism spectrum disorder screening system includes a display module 11 , an acquisition module 12 , a preprocessing module 13 , a feature extraction module 14 and a screening module 15 .
[0035] The display module 11 is used to display a test paradigm, which includes at least one situational test task to test different characteristics of a subject.
[0036] In this embodiment, the display module 11 displays a test paradigm for the subject to watch. The test paradigm sets multiple scenario test tasks, which are presented to the subject in a preset order to test the subject's reaction to different scenario test tasks when watching the test paradigm, such as eye avoidance, unusual repetitive movements, preferences or behavioral stereotypes. The scenario test tasks include at least:
[0037] The face observation task is used to observe the differences in face observation patterns between ASD subjects (autism spectrum disorder subjects) and TD subjects (normal subjects). Figure 2 As shown, for example, multiple faces are displayed for viewing by a subject.
[0038] The repetitive motion preference test task is used to test the repetitive motion preference and compare the differences in the degree of preference for repetitive motion between ASD subjects and TD subjects. Figure 3 and 4 As shown, for example, the repetitive action preference test task can be a 2D or 3D animation, Figure 3 and Figure 4 There are two display contents for the animation screen. Figure 3 Two playback windows are shown, one displaying the five-pointed star and its circular orbit, and the other displaying the five-pointed star and its random orbit. When playing the animation, one window displays the five-pointed star rotating around the circular orbit, and the other window displays the five-pointed star moving in a random orbit. Figure 4 Two playback windows are shown, one playback window displays a triangular star in a rectangular frame and its elliptical orbit, and the other playback window displays a triangular star in a rectangular frame and its random orbit. When playing the animation, one window plays the triangular star rotating around the elliptical orbit, and the other window plays the triangular star moving in random orbit.
[0039] The joint attention ability test task is used to test the joint attention ability and measure the subject's ability to naturally look at and follow the gaze of others. Figure 5 and 6 As shown, exemplary, Figure 5 The animation content is shown as the image looking at the object. Figure 6 The display animation content is shown as an image looking at a person.
[0040] Dynamic social image and dynamic geometric image preference test tasks are used to test the preference for dynamic social images and dynamic geometric images, and to compare the differences in the degree of preference for social and geometric scenes between ASD subjects and TD subjects. Figure 7 As shown, figures and geometric patterns are displayed.
[0041] It should be noted that the display module 11 only needs to display the test paradigm to the subject for viewing, and can be a display, a projector, a computer or other equipment, and this application does not impose any restrictions on this.
[0042] The acquisition module 12 is used to acquire eye movement information and facial video of the subject when watching the test paradigm, and send the eye movement information and facial video to the preprocessing module 13.
[0043] In this embodiment, if Figure 8 As shown, the acquisition module 12 includes a facial acquisition unit 121 and an eye acquisition unit 122. Normal subjects and subjects with autism spectrum disorders exhibit different eye movement and facial expression data when viewing the test paradigm. This can be demonstrated through various variations in eye movement and facial expression characteristics. By using eye movement information and facial expressions in multiple contexts as analysis and prediction information, a more objective and comprehensive assessment of subjects can be achieved.
[0044] The facial acquisition unit 121 is configured to acquire a facial video of the subject while watching the test paradigm, and send the facial video to the pre-processing module 13 .
[0045] The eye acquisition unit 122 is used to acquire eye information of the subject when viewing the test paradigm, and send the eye information to the pre-processing module 13 .
[0046] The facial acquisition unit 121 and the eye acquisition unit 122 may be any device or equipment capable of acquiring eye movement data and facial videos, and may be used to acquire eye movement information and facial videos when the subject watches the test paradigm.
[0047] The face collection unit 121 includes a camera, the eye collection unit 122 includes an eye tracker, and the display module 11 is a portable display. The camera is placed beside the display to record the face information of the subject during the whole process of watching the test paradigm. The eye tracker is installed at the bottom of the display to capture the eye movement information of the subject. The eye tracker can be customized for each subject to accurately locate the position of the eyes and accurately calculate the direction of the eye movement, avoiding errors caused by the height and eye difference of the subject. Further, in order to reduce the complexity of the subject in the autism spectrum disorder test and avoid the subject's attention to things other than the test paradigm, a controller 16 is provided, which is connected with the display module 11, the face collection unit 121 and the eye collection unit 122, and is used to control the display of the display module 11 and the data collection of the face collection unit 121 and the eye collection unit 122. During the actual test of the subject, the controller controls the display module 11 to display the test paradigm composed of multiple scene test tasks and controls the collection module 12 to execute the data capture program. After the subject finishes watching, the program automatically exits and outputs the eye movement information file and the face video of the subject.
[0048] Further, the autism spectrum disorder screening system further comprises a monitoring module, which comprises a global camera and a monitor. The global camera is used to focus on and record the overall test situation of the subject in real time, so as to make subsequent evaluation and screening on the state of the subject and the availability of the data. The monitor is used to observe the picture recorded by the global camera in real time.
[0049] The preprocessing module 13 is used to preprocess the face video according to the eye movement information, so as to obtain a face image frame containing the face frame of the subject corresponding to the eye movement entry in the eye movement information.
[0050] The preprocessing module 13 receives the eye movement information and the face video from the collection module 12 and preprocesses the face video according to the eye movement information. Specifically, the face video of the subject is subjected to frame alignment operation. The face video is a dynamic image sequence, and the eye movement information is a series of eye movement entries. Therefore, the face video is read to associate the eye movement entry and the face image at the same time point. Frame alignment is a process of reading the image frame according to the eye movement entry in the eye movement information. Further, the image frame read from the face video is subjected to face matching operation. There may be a phenomenon that the read image frame contains not only the face of the subject (such as an observer in the corner of the test room, a teacher of the subject, etc.), but also some pixel differences similar to the face in the environment. Therefore, face matching operation is needed to ensure that the recognized and processed face is the subject himself.
[0051] The pre-processing module 13 pre-processes the facial video according to the eye movement information in the following manner:
[0052] Image information of each frame in the facial video is read to obtain a plurality of image frames and a plurality of frame number positions corresponding to the image frames.
[0053] The image frames are traversed.
[0054] If the frame number position of the image frame corresponds to the eye movement entry in the eye movement information, an image frame set is generated based on the image frame.
[0055] Face matching is performed on the image frames in the image frame set to obtain a face image frame containing the subject's face frame.
[0056] In this embodiment, the image information of each frame in the facial video is read from the beginning, and the corresponding frame number position of each frame is recorded. Based on the frame number column in the eye movement information file, it is determined whether the currently read facial video frame corresponds to an eye movement entry recorded in the eye movement information file. If not, the frame is directly discarded. If so, the frame is saved to a designated directory to obtain an image frame set. Face matching is then performed on the image frames in the image frame set to match the face corresponding to the subject.
[0057] The face matching of the image frames in the image frame set is performed in the following manner:
[0058] Performing face detection on the image frames in the image frame set, locating the face region, and obtaining a face image containing the face region;
[0059] Obtaining a description vector corresponding to a face in each face region, wherein the description vector is used to characterize features of the face;
[0060] Calculating the Euclidean distance between the descriptive vector corresponding to the face in each face region and the basic vector, wherein the basic vector is the descriptive vector corresponding to the face of the subject obtained by performing face detection on the face image of the subject in advance;
[0061] If the Euclidean distance is less than a preset threshold, the face corresponding to the Euclidean distance is successfully matched with the subject's face to obtain a face image frame containing the subject's face frame, and the subject's face frame contains the face area corresponding to the subject.
[0062] In the embodiment, the application uses Dlib to process images. Dlib is an open source library of machine learning, which is a C++ open source toolkit of machine learning algorithms and contains many algorithms of machine learning. When performing face matching on image frames in the image frame set, an image containing only one face of a subject is manually intercepted in advance. A face detector provided by Dlib is used to detect the face of the image containing only one face of the subject, and a face descriptor provided by Dlib is used to obtain a description vector of the face of the subject. The description vector is a 128-dimensional face description vector, which is a 128-dimensional feature vector of a face, and the vector is saved as a basic vector, denoted as Disp GT .
[0063] A face detector of Dlib is used to sequentially perform multi-face detection on image frames in the image frame set, and a plurality of face regions containing faces are detected. A face descriptor is used to obtain 128-dimensional face description vectors corresponding to all detected faces. Frame i (i = 1...n) represents the i-th image frame read, n represents the total number of image frames, Disp ij (i = 1...n, j = 1...m i ) represents a 128-dimensional face description vector of the j-th face in the i-th image frame, m i represents the number of faces detected in the i-th image frame.
[0064] The Euclidean distances between the face description vectors Disp ij and the true face description vectors Disp GT in the image frames are sequentially calculated and compared. The Euclidean distances are determined to be matched according to a preset face similarity threshold. For example, the face similarity threshold is set to 0.4.
[0065] If the Euclidean distance corresponding to a face in an image frame is less than the face similarity threshold, it is determined that the face matches the face of the subject, and a subject face frame is drawn in the face region corresponding to the face, that is, an external rectangular frame of the face. The subject face frame contains the face region corresponding to the face.
[0066] The feature extraction module 14 is configured to extract eye movement features and expression features of the subject from the eye movement information and the face image frames. The eye movement features include eye fixation features, eye physiological features, and overall fixation features. The expression features are features of the proportion of emotions of the subject.
[0067] In the embodiment, the feature extraction module 14 extracts eye movement features from the eye movement information and the face image frames by the following method:
[0068] inputting the eye movement information and the face image frame into a trained eye movement modality classification model to obtain eye movement features of the subject, the eye movement modality classification model being obtained by training a classifier based on eye movement information and face image frames of a plurality of sample subjects when the sample subjects view the test paradigm, and eye movement fixation features, eye movement physiological features, and overall fixation features of the plurality of sample subjects.
[0069] The training of the eye movement modality classification model is achieved by the following method:
[0070] dividing each sub-scene in different scenarios corresponding to the scenario test task in the test paradigm into a plurality of regions of interest.
[0071] obtaining eye movement information and face image frames of a plurality of sample subjects, the sample subjects including normal subjects and autism spectrum disorder subjects.
[0072] According to the eye movement information of the plurality of sample subjects and the regions of interest, calculating eye movement fixation features and overall fixation features of the plurality of sample subjects, the eye movement fixation features including total fixation point number, regional fixation point number, and inter-regional switching number, the total fixation point number being used to represent the fixation number of each sub-scene in different scenarios, the regional fixation point number being used to represent the fixation number of each region of interest in each sub-scene in different scenarios, the inter-regional switching number being used to represent the number of switching of fixation points between regions of interest in different scenarios, and the overall fixation features including fixation rate.
[0073] According to the eye movement information and the face image frames of the plurality of sample subjects, calculating eye movement physiological features of the plurality of sample subjects, the eye movement physiological features including eye width-height ratio, eyeball width-height ratio, and blink rate.
[0074] Based on the eye movement information and the face image frames of the plurality of sample subjects when the sample subjects view the test paradigm, and the eye movement fixation features, the eye movement physiological features, and the overall fixation features of the plurality of sample subjects, the eye movement modality classification model is trained.
[0075] In this embodiment, for a picture, the regions of interest of the ASD children and the TD children are different, for example, for a picture of a face, the ASD children pay more attention to the mouth region of the picture of the face, and the TD children pay more attention to the eye region. It can be understood that the test paradigm includes a plurality of scene test tasks, each scene includes a plurality of sub-scenes, and the subjects have different degrees of interest in different regions of the sub-scenes when watching different sub-scenes in the test paradigm, so the sub-scenes in the test paradigm can be divided into different regions for evaluating the degree of interest of the subjects in different regions, and then the relevant characteristics of the subjects are measured. As shown in Figure 9 To analyze the difference in observation mode between the ASD subjects and the TD subjects, a manual division method can be used to divide each sub-scene in different scenes corresponding to the scene test task into a plurality of regions of interest (ROI), including regions containing faces, objects, geometric figures, and other regions (background regions), and each region of interest is numbered, such as ROI_1, ROI_2, ROI_3, ROI_4, and the like. Figure 9 Exemplary regions of interest are shown in the division of sub-scenes in four different scenes. The attention of the subjects to different regions of interest in the sub-scenes of the above test paradigm can be measured by the number of gazes and the like. Generally, the more the number of gazes of the subjects to a certain region of interest, the more attention they pay to the region. Exemplary, the eye gaze features, eye physiological features, and overall gaze features can be used to reflect the attention.
[0076] In this embodiment, by obtaining the eye movement information and the face image frames of a plurality of sample subjects watching the test paradigm, the eye gaze features, the eye physiological features, and the overall gaze features are calculated according to the division of the regions of interest, and a total of 82-dimensional features are obtained, as shown in the following table:
[0077]
[0078] In this embodiment, the present application uses a logistic regression model (LR) to train a classifier, wherein, in order to simplify the training and prediction of the machine learning model, the features can be reduced in dimension, the main features can be retained, the amount of data can be greatly reduced, and the efficiency of training and prediction can be improved. The present application uses a cross-validated recursive feature elimination algorithm (RFECV) to perform feature dimensionality reduction, selects the best 75% of features (61-dimensional features) as the feature combination of eye movement modality, and uses the feature combination to train the classifier of the eye movement modality to obtain an eye movement modality classification model. The eye features currently extracted are limited to the relevant features of eye gaze. The present application proposes a feature extraction scheme based on multiple feature groups, which not only focuses on eye gaze features, but also focuses on important evaluation indicators such as eye physiological characteristics and overall gaze features in addition to eye gaze features, and more comprehensively evaluates the different characteristics of the subjects in multiple situations.
[0079] For the dominant eye feature, the specific calculation method is as follows:
[0080] Total number of fixations: Based on the eye movement information of the sample subjects, the number of fixations of the sample subjects in each sub-scene in different situations is counted, and a total of 12-dimensional features are extracted.
[0081] Regional fixation points: Based on the divided regions of interest, the region of interest to which each eye movement entry in the eye movement information belongs is determined, and cumulative calculations are performed to extract a total of 44-dimensional features.
[0082] Number of switching between regions: According to the order of the eye movement items in the eye movement information, the numbers of the regions of interest belonging to the current eye movement item are compared with those of the previous eye movement item. If they are different, the numbers are accumulated and calculated to extract a total of 11 dimensions of features.
[0083] For eye physiological characteristics, the eye aspect ratio, eyeball width-to-height ratio, and blink rate are calculated using the following methods:
[0084] Performing facial key point detection on the facial image frame of the sample subject to obtain a sample key point image, wherein the sample key point image includes facial key points of the sample subject.
[0085] Eye key points are extracted from the sample key point image, wherein the eye key points include the key point positions of the left eye and the key point positions of the right eye, and the key point positions include the eyelid positions, the eye corner positions, and the eyeball positions.
[0086] The eye aspect ratio and eyeball aspect ratio of the sample subject are calculated based on the eye key points, wherein the eye aspect ratio is the average of the eye aspect ratios of the left and right eyes, and the eyeball aspect ratio is the average of the eyeball aspect ratios of the left and right eyes.
[0087] According to the numerical value of the eye aspect ratio, a blink rate is determined.
[0088] Specifically, a key point detector provided by Dlib is used to detect key points of the matched face in the facial image frame of the sample subject, to obtain a sample key point image containing facial key points of the sample subject, and eye key points are extracted, as shown in Figure 10 Figure 10 In the formula, L i and R i (i = 1...6) represent the key positions of the left and right eyes, including the positions of the eyelids, canthi and eyeballs, L i (L xi , L yi ) and R i (R xi , R yi ) represent the coordinates of the corresponding positions, and the eye aspect ratio is calculated by the following formula:
[0089]
[0090]
[0091] In the formula, L whr and R whr are the eye aspect ratios of the left and right eyes, respectively. The eyeball aspect ratio is calculated by the following formula:
[0092]
[0093]
[0094] In the formula, LB whr and RB whr are the eyeball aspect ratios of the left and right eyes, respectively. The mean value of the eye aspect ratio and the mean value of the eyeball aspect ratio of the left and right eyes are taken as the eye aspect ratio and the eyeball aspect ratio of the current facial image frame, and finally the statistical quantity and expectation of the eye aspect ratio and the statistical quantity and expectation of the eyeball aspect ratio of all facial image frames of the sample subject are calculated, to obtain a total of 4-dimensional features. For the blink rate, the size of the eye aspect ratio is used as the standard to judge whether the eye blinks, and for example, the threshold is set to 0.5, if the reciprocal of the eye aspect ratio is less than 0.5, it is considered as blinking and is accumulated, to obtain a 1-dimensional feature.
[0095] For the overall fixation feature, the feature is used to evaluate the fixation of the sample subject in each sub-scene under different situations, and the specific calculation method is as follows:
[0096]
[0097] Where m represents the sample subject number, ranging from 1 to 66, and n represents the sub-scene number involved in the statistics, ranging from 1 to 10. represents the number of fixations of the mth sample subject in the nth sub-scene, It represents the comprehensive gaze rate of the mth sample subject in the nth sub-scene, that is, the ratio of the current user's gaze points to the average number of gaze points of all sample subjects in the same sub-scene. A total of 10 dimensions of gaze rate features are extracted for the sample subjects.
[0098] In this embodiment, compared with normally developing children, the expressions of children with autism spectrum disorders show stereotyped phenomena. Emotions, as a subjective feeling, are mainly conveyed through their external expression patterns - facial expressions. This application also uses the subject's expression as one of the information for analysis and prediction when screening for autism spectrum disorders, so as to evaluate the subject more objectively and comprehensively. The feature extraction module 14 extracts expression features from the eye movement information and the facial image frame in the following way:
[0099] Performing facial key point detection on the face area within the subject's face frame in the face image frame to obtain a key point image, wherein the key point image contains the subject's facial key points, and each of the subject's facial key points corresponds to a two-dimensional coordinate.
[0100] The face region is extracted from the key point image.
[0101] The width of the face area is adjusted to a preset width.
[0102] According to the preset width, the two-dimensional coordinates corresponding to the facial key points of the subject in the face area after width adjustment are obtained to obtain the facial features corresponding to the face image frame, where the facial features are one-dimensional information mapped from the two-dimensional coordinates.
[0103] Specifically, facial features are extracted from the face image frame to perform emotion recognition. Different from the previous method of dividing action units and calculating geometric distances, this application proposes a row-first mapping strategy (RMS) that maps the two-dimensional coordinates of key points into one-dimensional information and models the relative positions of all key points. Figure 11The figure shows the RMS-based keypoint mapping process. To protect the subject's personal information, this example uses publicly available images from the CK+ standard face database to demonstrate this process. First, using the keypoint detector provided by Dlib, keypoints are detected on the subject's face within the subject's face frame in the face image frame. This produces a keypoint image containing the subject's facial keypoints. The facial keypoints include 68 keypoints, each of which corresponds to a two-dimensional position coordinate to identify the keypoint's location. Figure 11 The rectangular frame outside the face shown in the figure represents the subject's face frame, marking the face position identified after face matching. The points on the face mark the locations of the detected facial landmarks. The face region is cropped from the keypoint image based on the acquired face positions. To preserve the relative position information of all keypoints, the cropped face region is treated as a pixel matrix, and a row-first strategy is used to map the two-dimensional coordinates into one-dimensional feature information.
[0104] Furthermore, in order to speed up the convergence of the model and avoid the influence of different distances between the person and the camera, the size of the face area is adjusted, and the cropped face area is adjusted based on the row. Figure 11 The coordinates of a key point in (L x , L y ), the width and height of the cropped face area are W and H respectively, the width of the face area is adjusted to a fixed value W′, and the coordinates of the key points after adjustment are Corresponding to Figure 11 Finally, the key points are completed according to the following formula The mapping:
[0105]
[0106] After completing the mapping of each facial key point, the facial features corresponding to the facial image frame are obtained, including 68-dimensional facial features.
[0107] The facial features and the eye physiological features are input into a trained emotion recognition model to obtain an emotion label of the face image frame, where the emotion label includes a basic emotion label and a neutral emotion label.
[0108] Among them, training the emotion recognition model is achieved through the following methods:
[0109] Obtain the model emotion set from the standard face database;
[0110] Filtering the model emotion set to obtain a filtered target emotion set, wherein the target emotion set includes basic emotions and neutral emotions;
[0111] Obtaining facial features corresponding to the facial images in the target emotion set and eye physiological features corresponding to the facial images in the target model dataset;
[0112] The emotion recognition model is trained based on the target emotion set, facial features corresponding to the face images in the target emotion set, and eye physiological features corresponding to the face images in the target model data set.
[0113] Specifically, a standard database is used as a data set to pre-train an emotion recognition model, and the emotion recognition model is used to perform emotion recognition on the faces matched in the face image frames of the subject. The standard database is the CK+ standard face database, and the CK+ standard face database is used as the data set of the model to train the entire emotion recognition model. Figure 12 As shown, a model emotion set was obtained from the CK+ standard face database. This model emotion set contains 593 expression sequences from 123 model subjects. Of these 593 expression sequences, 327 meet the definition of emotion prototypes, representing a total of seven emotions: six basic emotions (happiness, anger, fear, surprise, disgust, and sadness) and one neutral emotion (contempt). Each expression sequence progresses from a calm expression to a peak expression, meaning the order of the sequences represents the varying degrees of emotion. The model emotion set was preprocessed, filtering the set to select facial images representing the six basic emotions and a neutral emotion. A total of 2,822 facial images were selected for emotion recognition model training. In the stage of training the emotion recognition model, we first perform facial detection on each face image, including face detection and key point detection, and extract the 68-dimensional facial features of each face image and the eye aspect ratio and eyeball aspect ratio of each face image. In total, we extract the statistics and expectations of the eye aspect ratio of the face image and the statistics and expectations of the eyeball aspect ratio, a total of 4-dimensional features. Furthermore, in order to optimize the features, we use the cross-validated recursive feature elimination algorithm (RFECV) to reduce the feature dimensionality and select the optimal 61-dimensional features, of which 59 dimensions are selected for facial features and 2 dimensions are selected for eye features, which are denoted as M and M respectively. f 、M e Based on the selected 61-dimensional feature matrix M′=[M f , M e ], the classifier was trained using the logistic regression model (LR) and an emotion recognition model for emotion recognition was obtained.
[0114] The pre-trained emotion recognition model is used for emotion recognition on all face image frames of the subject in sequence, the facial features and the eye physiological features of the subject in the face image frames are combined with the features selected by the recursive feature elimination algorithm (RFECV) through cross validation, and the combined features are input into the emotion recognition model. The emotion recognition model outputs the emotion classification result of each face image frame to obtain an emotion label of the face image frame. The emotion label includes a basic emotion label and a neutral emotion label. The basic emotion label includes six basic emotions of happiness, anger, fear, surprise, disgust and sadness. The neutral emotion includes one neutral emotion of contempt. By labeling the face image frames of the subject, a relationship between emotion features and subject categories is established.
[0115] In this embodiment, when training the emotion recognition model, in addition to extracting facial features, physiological related features of eye width-height ratio and eyeball width-height ratio corresponding to each face image are extracted because eye state is a key factor for expression positioning, so as to improve the accuracy of emotion recognition.
[0116] After obtaining the emotion label of the face image frame, the emotion label is input into the trained expression modal classification model to obtain the expression features corresponding to the subject. The expression features are features of the emotion proportion of the subject. The expression modal classification model is obtained by training a classifier using the emotion labels corresponding to the face image frames of a plurality of sample subjects and the proportion of the emotion labels of the plurality of sample subjects as sample data.
[0117] The expression modal classification model is implemented in the following manner:
[0118] The emotion labels corresponding to the face image frames of the plurality of sample subjects and the proportion of the emotion labels of the plurality of sample subjects are obtained. Since the emotion state of the subject during the entire autism spectrum disorder test process needs to be obtained instead of a single emotion at a certain moment, the frequencies of the seven emotions in the emotion labels are counted respectively, and the proportions of the seven emotions are calculated, i.e., a 7-dimensional feature representing the emotion proportion of the subject. Based on the 7-dimensional feature representing the emotion proportion of the subject, a logistic regression model (Logistic Regression, LR) is used for classifier training of the expression modal to learn the correlation between the emotion label and the expression, and an expression modal classification model is obtained.
[0119] The screening module 15 is used for inputting the eye movement features and the expression features into the trained screening model to obtain the autism spectrum disorder screening result of the subject.
[0120] In this embodiment, the eye movement features and facial expression features output by the eye movement modality classification model and the facial expression modality classification model are input into a screening model, which then outputs a screening result indicating whether the subject has an autism spectrum disorder. For example, the screening result can be represented by 0 or 1, with 0 indicating no autism spectrum disorder and 1 indicating autism spectrum disorder.
[0121] Among them, the training screening model is achieved through the following methods:
[0122] Eye features of a plurality of sample subjects and facial expression features of a plurality of sample subjects are obtained, wherein the sample subjects include normal subjects and subjects with autism spectrum disorder.
[0123] The eye features of the multiple sample subjects and the expression features of the multiple sample subjects are fused into training samples.
[0124] The screening model is trained based on the training samples.
[0125] Specifically, the eye movement modality classification model and the expression modality classification model are fused by feature-level fusion, which directly combines the eye movement features and expression features output by the eye movement modality classification model and the expression modality classification model. That is, the 61-dimensional eye movement features (after dimensionality reduction) and the 7-dimensional expression features are combined into a 68-dimensional feature matrix, and the screening model is trained using a logistic regression model (LR).
[0126] In a second aspect, the present application provides a method for screening for autism spectrum disorders based on eye movements and facial expressions, comprising:
[0127] A test paradigm is displayed, which includes at least one situational test task to test different characteristics of the subject.
[0128] The eye movement information and facial video of the subject when watching the test paradigm are collected, and the eye movement information and the facial video are sent to a preprocessing module.
[0129] The facial video is preprocessed according to the eye movement information to obtain a facial image frame containing the subject's face corresponding to the eye movement entry in the eye movement information.
[0130] The eye movement features and expression features of the subject are extracted from the eye movement information and the facial image frame, wherein the eye movement features include eye gaze features, eye physiological features, and overall gaze features, and the expression features are features of the proportion of the subject's emotions.
[0131] The eye movement features and the facial expression features are input into a trained screening model to obtain autism spectrum disorder screening results of the subject.
[0132] The effects of the above method when applying the above system can be found in the description of the above system embodiment, which will not be repeated here.
[0133] As can be seen from the above technical solutions, this application provides an autism spectrum disorder screening system and method based on eye movements and facial expressions. A test paradigm including multiple scenarios is designed, focusing on various behavioral characteristics of patients with autism spectrum disorder, including eye movement characteristics and expression characteristics. Regarding eye movement characteristics, a feature extraction scheme based on multiple feature groups is proposed, focusing not only on eye gaze characteristics, but also on important evaluation indicators such as eye physiological characteristics and overall gaze characteristics in addition to eye gaze characteristics. Regarding expression characteristics, a new facial key feature extraction method - RMS is proposed, and an emotion recognition model for emotion recognition is trained. The screening model is obtained by feature fusion training of the eye movement characteristics and facial features of multiple sample subjects when viewing the test paradigm. A data set containing normal sample subjects and autism spectrum disorder sample subjects is collected, and autism spectrum disorder is screened from two levels of eye movement and expression. The subjects are evaluated more objectively and comprehensively, which greatly reduces the missed diagnosis rate and misdiagnosis rate, improves the screening accuracy and efficiency, and the screening method is simple and easy.
[0134] For ease of explanation, the above description has been presented in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations are possible. The above embodiments have been selected and described to better explain the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various variations of the embodiments suitable for specific use considerations.
Claims
1. Autism spectrum disorder screening system based on eye movements and facial expressions, characterized by: include: A display module, configured to display a test paradigm, wherein the test paradigm includes at least one situational test task to test different characteristics of a subject; an acquisition module, configured to acquire eye movement information and facial video of the subject while watching the test paradigm, and send the eye movement information and facial video to a preprocessing module; a preprocessing module, configured to preprocess the facial video according to the eye movement information to obtain a facial image frame containing the subject's face frame corresponding to an eye movement entry in the eye movement information; A feature extraction module is configured to extract the subject's eye movement features and expression features from the eye movement information and the facial image frame, wherein the eye movement features include eye gaze features, eye physiological features, and overall gaze features, and the expression features are features of the subject's emotional proportion. The feature extraction module extracts the eye movement features from the eye movement information and the facial image frame in the following manner: Inputting the eye movement information and the facial image frame into a trained eye movement modality classification model to obtain the subject's eye movement features, wherein the eye movement modality classification model is obtained by training a classifier using the eye movement information and facial image frames of multiple sample subjects when viewing the test paradigm, as well as the eye gaze features, eye physiological features, and overall gaze features of the multiple sample subjects as sample data; Training the eye movement modality classification model is achieved in the following way: Dividing each sub-scenario under different scenarios corresponding to the scenario test task in the test paradigm into a plurality of regions of interest; acquiring eye movement information and facial image frames of a plurality of sample subjects, wherein the sample subjects include normal subjects and subjects with autism spectrum disorder; Calculating eye gaze features and overall gaze features of the multiple sample subjects based on the eye movement information of the multiple sample subjects and the regions of interest, the eye gaze features including a total number of gaze points, a number of regional gaze points, and a number of inter-region switching times, the total number of gaze points being used to characterize the number of gazes of the subject on each sub-scene in different situations, the number of regional gaze points being used to characterize the number of gazes of the subject on each region of interest in each sub-scene in different situations, the number of inter-region switching times being used to characterize the number of times the subject's gaze point switches back and forth between the regions of interest in different situations, and the overall gaze features including a gaze rate; Calculating eye physiological characteristics of the multiple sample subjects based on the eye movement information and the facial image frames of the multiple sample subjects, the eye physiological characteristics including eye aspect ratio, eyeball width-to-height ratio, and blink rate; The eye movement modality classification model is trained based on the eye movement information and facial image frames of the multiple sample subjects when viewing the test paradigm, as well as the eye gaze features, eye physiological features, and overall gaze features corresponding to the multiple sample subjects; The feature extraction module extracts expression features from the eye movement information and the face image frame in the following manner: Performing facial key point detection on a facial region within a face frame of a subject in the facial image frame to obtain a key point image, wherein the key point image includes facial key points of the subject, and each facial key point of the subject corresponds to a two-dimensional coordinate; Extracting the face area from the key point image; Adjusting the width of the face area to a preset width; According to the preset width, obtaining two-dimensional coordinates corresponding to facial key points of the subject in the face area after the width is adjusted to obtain facial features corresponding to the face image frame, wherein the facial features are one-dimensional information mapped from the two-dimensional coordinates; Inputting the facial features and the eye physiological features into a trained emotion recognition model to obtain an emotion label for the face image frame, wherein the emotion label includes a basic emotion label and a neutral emotion label; Inputting the emotion label into a trained expression modality classification model to obtain expression features corresponding to the subject, wherein the expression modality classification model is obtained by training a classifier using emotion labels corresponding to facial image frames of multiple sample subjects and the proportion of emotion labels of the multiple sample subjects as sample data; The screening module is used to input the eye movement features and the facial expression features into a trained screening model to obtain the autism spectrum disorder screening result of the subject.
2. The autism spectrum disorder screening system based on eye movements and facial expressions according to claim 1, characterized in that: The situational test tasks include: a face observation task, a repetitive action preference test task, a joint attention ability test task, and a dynamic social image and dynamic geometric image preference test task.
3. The autism spectrum disorder screening system based on eye movements and facial expressions according to claim 1, characterized in that: The acquisition module includes: a facial acquisition unit, configured to acquire a facial video of the subject while watching the test paradigm, and send the facial video to the preprocessing module; An eye acquisition unit is used to acquire eye information of the subject when viewing the test paradigm, and send the eye information to the preprocessing module.
4. The autism spectrum disorder screening system based on eye movements and facial expressions according to claim 1, characterized in that: The preprocessing module preprocesses the facial video according to the eye movement information in the following manner: Reading image information of each frame in the facial video to obtain a plurality of image frames and a plurality of frame number positions corresponding to the image frames; traversing the image frames; If the frame number position of the image frame corresponds to the eye movement entry in the eye movement information, generating an image frame set based on the image frame; Face matching is performed on the image frames in the image frame set to obtain a face image frame containing the subject's face frame.
5. The autism spectrum disorder screening system based on eye movements and facial expressions according to claim 4, characterized in that: Performing face matching on the image frames in the image frame set is achieved by: Performing face detection on the image frames in the image frame set, locating the face region, and obtaining a face image containing the face region; Obtaining a description vector corresponding to a face in each face region, wherein the description vector is used to characterize features of the face; Calculating the Euclidean distance between the descriptive vector corresponding to the face in each face region and the basic vector, wherein the basic vector is the descriptive vector corresponding to the face of the subject obtained by performing face detection on the face image of the subject in advance; If the Euclidean distance is less than a preset threshold, the face corresponding to the Euclidean distance is successfully matched with the subject's face to obtain a face image frame containing the subject's face frame, and the subject's face frame contains the face area corresponding to the subject.
6. The autism spectrum disorder screening system based on eye movements and facial expressions according to claim 1, characterized in that: Calculation of eye aspect ratio, eyeball aspect ratio, and blink rate is achieved by: Performing facial key point detection on the facial image frame of the sample subject to obtain a sample key point image, wherein the sample key point image includes facial key points of the sample subject; Extracting eye key points from the sample key point image, wherein the eye key points include the key point positions of the left eye and the key point positions of the right eye, and the key point positions include the eyelid positions, the eye corner positions, and the eyeball positions; Calculating the eye aspect ratio and eyeball aspect ratio of the sample subject based on the eye key points, wherein the eye aspect ratio is the average of the eye aspect ratios of the left and right eyes, and the eyeball aspect ratio is the average of the eyeball aspect ratios of the left and right eyes; The blink rate is determined according to the value of the eye aspect ratio.
7. The autism spectrum disorder screening system based on eye movements and facial expressions according to claim 1, characterized in that: Training the emotion recognition model is achieved in the following ways: Obtain the model emotion set from the standard face database; Filtering the model emotion set to obtain a filtered target emotion set, wherein the target emotion set includes basic emotions and neutral emotions; Acquire facial features corresponding to the facial images of the target emotional set and eye physiological features corresponding to the facial images of the target emotional set; The emotion recognition model is trained based on the target emotion set, facial features corresponding to the facial images in the target emotion set, and eye physiological features corresponding to the facial images in the target emotion set.
8. The autism spectrum disorder screening system based on eye movements and facial expressions according to claim 1, characterized in that: Training the screening model is achieved by: Acquiring eye features and facial expression features of a plurality of sample subjects, wherein the sample subjects include normal subjects and subjects with autism spectrum disorder; fusing the eye features of the plurality of sample subjects and the expression features of the plurality of sample subjects into training samples; The screening model is trained based on the training samples.
9. Autism spectrum disorder screening method based on eye movements and facial expressions, characterized in that: include: displaying a test paradigm, the test paradigm comprising at least one situational test task to test different characteristics of the subject; collecting eye movement information and facial video of the subject while watching the test paradigm, and sending the eye movement information and facial video to a preprocessing module; pre-processing the facial video according to the eye movement information to obtain a facial image frame containing the subject's face corresponding to an eye movement entry in the eye movement information; Extracting eye movement features and expression features of the subject from the eye movement information and the facial image frame, wherein the eye movement features include eye gaze features, eye physiological features, and overall gaze features, and the expression features are features of the proportion of the subject's emotions; wherein extracting eye movement features from the eye movement information and the facial image frame is achieved by: Inputting the eye movement information and the facial image frame into a trained eye movement modality classification model to obtain the subject's eye movement features, wherein the eye movement modality classification model is obtained by training a classifier using the eye movement information and facial image frames of multiple sample subjects when viewing the test paradigm, as well as the eye gaze features, eye physiological features, and overall gaze features of the multiple sample subjects as sample data; Training the eye movement modality classification model is achieved in the following way: Dividing each sub-scenario under different scenarios corresponding to the scenario test task in the test paradigm into a plurality of regions of interest; acquiring eye movement information and facial image frames of a plurality of sample subjects, wherein the sample subjects include normal subjects and subjects with autism spectrum disorder; Calculating eye gaze features and overall gaze features of the multiple sample subjects based on the eye movement information of the multiple sample subjects and the regions of interest, the eye gaze features including a total number of gaze points, a number of regional gaze points, and a number of inter-region switching times, the total number of gaze points being used to characterize the number of gazes of the subject on each sub-scene in different situations, the number of regional gaze points being used to characterize the number of gazes of the subject on each region of interest in each sub-scene in different situations, the number of inter-region switching times being used to characterize the number of times the subject's gaze point switches back and forth between the regions of interest in different situations, and the overall gaze features including a gaze rate; Calculating eye physiological characteristics of the multiple sample subjects based on the eye movement information and the facial image frames of the multiple sample subjects, the eye physiological characteristics including eye aspect ratio, eyeball width-to-height ratio, and blink rate; The eye movement modality classification model is trained based on the eye movement information and facial image frames of the multiple sample subjects when viewing the test paradigm, as well as the eye gaze features, eye physiological features, and overall gaze features corresponding to the multiple sample subjects; Extracting expression features from the eye movement information and the facial image frame is achieved by: Performing facial key point detection on a facial region within a face frame of a subject in the facial image frame to obtain a key point image, wherein the key point image includes facial key points of the subject, and each facial key point of the subject corresponds to a two-dimensional coordinate; Extracting the face area from the key point image; Adjusting the width of the face area to a preset width; According to the preset width, obtaining two-dimensional coordinates corresponding to facial key points of the subject in the face area after the width is adjusted to obtain facial features corresponding to the face image frame, wherein the facial features are one-dimensional information mapped from the two-dimensional coordinates; Inputting the facial features and the eye physiological features into a trained emotion recognition model to obtain an emotion label for the face image frame, wherein the emotion label includes a basic emotion label and a neutral emotion label; Inputting the emotion label into a trained expression modality classification model to obtain expression features corresponding to the subject, wherein the expression modality classification model is obtained by training a classifier using emotion labels corresponding to facial image frames of multiple sample subjects and the proportion of emotion labels of the multiple sample subjects as sample data; The eye movement features and the facial expression features are input into a trained screening model to obtain autism spectrum disorder screening results of the subject.
Citation Information
Patent Citations
Machine learning-based method for evaluating and predicting ASD
CN105069304A
Face emotion recognition method and device, compute device and storage medium
CN109190487A
Interactive and adaptive learning, neurocognitive disorder diagnosis, and noncompliance detection systems using pupillary response and face tracking and emotion detection with associated methods
US20200178876A1