A mobile-based children's visual attention abnormality screening method based on multimodal data learning

By recording the head and face videos of children watching on smartphones, a multimodal data learning model is constructed, which solves the problem of medium and high-cost eye trackers for visual attention abnormalities screening, and achieves low-cost, comprehensive analysis of visual attention abnormalities, which is suitable for childhood disease screening.

CN115761908BActive Publication Date: 2025-08-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211164739.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-08-22
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

In the prior art, the research on visual attention abnormal screening is limited by high-cost eye trackers and ignores other modal data except eye tracking, such as facial expressions and head postures, resulting in limited development and promotion of visual attention abnormal screening.

Method used

Using a multimodal data learning method, children's head and face videos when watching videos are recorded through smartphone cameras, an eye movement estimation model is constructed, eye movement, facial expressions and head posture features are extracted, and LSTM network is used for feature fusion to realize visual attention abnormal screening.

Benefits of technology

It reduces the cost of screening for visual attention abnormalities, improves the abnormal recognition ability of the model, and can comprehensively analyze the observer's visual attention process under low-cost conditions. It is suitable for screening for diseases such as autism, attention deficits and hyperactivity disorder, and has important research significance and industrial application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761908B_ABST
    Figure CN115761908B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for screening abnormal visual attention of children on a mobile terminal based on multimodal data learning, and belongs to the field of computer vision. A calibration video and a test video are set, and the head and face videos of children watching the calibration video and the test video on a smart phone are recorded respectively; an eye movement estimation model is constructed, and the gaze point position is predicted frame by frame from the head and face video corresponding to the test video, and eye movement features are extracted from multiple angles such as the gaze point jump amplitude, gaze point jump angle, and area of ​​interest; facial expression features and head posture features are extracted from the head and face video corresponding to the test video respectively; and different modal features are fused using a long short-term memory network to achieve mapping from multimodal features to category labels. In the test phase, the head and face videos of the children to be classified are recorded when they are watching the mobile terminal video, and features such as eye movement, facial expression, and head posture are extracted and input into the trained model to determine whether they are abnormal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and relates to a method for identifying abnormal visual attention in children. The method collects head and facial videos of children watching videos on smartphones to estimate the eye movement position of each video frame, and then extracts multimodal features such as facial expressions and head posture to realize screening of abnormal visual perception in children. Background Art

[0002] Visual attention is a crucial regulatory mechanism in the human visual system, crucial for effectively processing complex and massive visual input. It rationally allocates limited information processing resources, enabling detailed analysis and processing of only a selected subregion of the visual system at the moment, while suppressing responses to areas outside of that region. To date, the vast majority of research on visual attention has focused on modeling common mechanisms across different observers, while research on the differentiation of visual attention processes has been limited.

[0003] To address this issue, scholars have recently begun exploring differences in visual attention across different categories of observers. Using comparative experiments, they have identified abnormalities in visual attention across certain demographics, providing objective evidence for disease screening and identification and promoting the development of intelligent diagnosis and treatment models. Jones and Klin, in their article "Attention to eyes is present but in decline in 2-6-month-old infants later diagnosed with autism," Nature, vol. 504, no. 7480, pp. 427-431, 2013, collected visual attention data from high-risk and low-risk children while watching videos of interactions between naturalistic caregivers. The data also measured the percentage of time spent fixating on the eyes, mouth, body, and objects. This study indicates that as young as 2-6 months, children with autism exhibit abnormal visual attention patterns that differ from typically developing children, primarily manifested by shorter fixations on social scenes and faces. In the article "Eye tracking during a continuous performance test: Utility for assessing ADHD patients," Journal of Attention Disorders, vol. 26, no. 2, pp. 245-255, 2022, Lev et al. compared the visual attention data of 35 adults with attention deficit hyperactivity disorder and a control group of observers under different tasks, and found that patients with attention deficit hyperactivity disorder spent significantly more time looking at irrelevant areas inside and outside the screen.

[0004] It should be noted that previous research on visual attention anomalies has primarily focused on eye movement analysis. For this purpose, current work has utilized eye trackers to record key eye movement data, such as the observer's visual attention allocation pattern. However, the high cost of eye trackers has limited the development and promotion of research on visual attention anomaly screening. Consequently, recent research has focused on eye movement estimation on mobile devices, such as smartphones. In their article, "Accelerating eye movement research via accurate and affordable smartphone eye tracking," published in Nature Communications, vol. 11, no. 1, pp. 1-12, 2020, Valliappa et al. constructed a deep learning-based mobile eye movement estimation model and compared it with a traditional head-mounted eye tracker in multiple visual attention tasks, such as visual search. This model validated the feasibility of using smartphone cameras for eye tracking in specific visual attention tasks. On the other hand, previous research on visual attention processes has often overlooked data from modalities other than eye movements, such as facial expressions and head posture. However, these data can also reflect anomalies in the visual attention process and play an important role in visual attention anomaly screening. Summary of the Invention

[0005] Technical problems to be solved

[0006] In order to avoid the shortcomings of the existing technology, the present invention provides a mobile terminal children's visual attention abnormality screening method based on multimodal data learning.

[0007] Technical Solution

[0008] A mobile terminal child visual attention abnormality screening method based on multimodal data learning is characterized by the following steps:

[0009] Step 1: Set up calibration and test videos, and record head and facial videos of m children with visual attention disorders and m children with normal development while watching the calibration and test videos on smartphones.

[0010] Step 2: Build an eye movement estimation model and use it to estimate the gaze point position of the head and face video frames corresponding to the test video, and extract eye movement features from the gaze point jump amplitude, gaze point jump angle and region of interest;

[0011] Step 3: Extract facial expression features and head movement features from the head and face videos corresponding to the test video. Then, use the long short-term memory (LSTM) network to fuse the different modal features and achieve the mapping from multimodal features to category labels.

[0012] Step 4: During the testing phase, record the head and face videos of the children to be classified while watching mobile videos, and extract multimodal features of eye movements, facial expressions, and head movements, and input them into the trained LSTM network model to determine whether there are any abnormalities.

[0013] A further technical solution of the present invention: Step 1 is as follows:

[0014] A smartphone application was used to play two calibration videos and one test video. The head and face videos of m children with visual attention disorders and m children with normal development were recorded while they watched calibration video 1, calibration video 2, and the test video.

[0015] The calibration video 1 is a 72-second video with a black background and 40 green dots flashing randomly, each in a random position, for 1.8 seconds, with the size varying from 18 to 50 pixels.

[0016] The calibration video 2 is a 72-second video with a black background and green dots moving in a zigzag pattern on the screen. Each dot takes 1.8 seconds, and the 40 dots move smoothly.

[0017] The test video is a 2-minute video composed of a short clip of a children's educational game and a video of geometric changes;

[0018] The mobile phone is placed horizontally, and the child sits about 30cm-40cm away from the phone to watch the mobile phone video. While the video is playing, the mobile phone camera is used to record the child's head and face video.

[0019] A further technical solution of the present invention: The eye movement estimation model constructed in step 2 is specifically as follows:

[0020] First, video frames are sampled from the head and face videos corresponding to the two calibration videos. At the same time, a single-order multi-layer detector is used to detect the face and left and right eyes in each video frame to obtain face images, left and right eye images and corresponding bounding boxes; next, the face and left and right eye images are converted to a fixed size, and the corresponding bounding boxes are normalized. The convolutional neural network is used to extract features from the face and left and right eye images respectively; the left and right eyes share the convolutional neural network weights, and the corresponding bounding boxes are then extracted using a fully connected layer to obtain the relative position features of the face and left and right eyes in the image; finally, the above extracted features are fused using a fully connected layer, and the final eye movement estimation result is obtained through the fully connected layer.

[0021] A further technical solution of the present invention: The facial expression feature extraction described in step 3 is specifically as follows: in each video frame, the LibFaceDetection technology is used to detect and crop the face; if no face is detected, the video frame is marked as no face is detected; if a face is detected, the cropped face data is input into a trained facial expression recognition network for classification and recognition; finally, the facial expression recognition results of all frames in each video are counted to generate a facial expression histogram.

[0022] A further technical solution of the present invention is as follows: the head motion features described in step 3 include head posture features and head motion distance features. The head posture feature extraction method uses a constrained local neural domain facial landmark detector to estimate different head posture angles, and then classifies the head posture by setting a threshold; the head posture angles include: head pitch angle p, head yaw angle y, and head roll angle r; based on the three head posture angles, the specific classification of head posture is as follows:

[0023]

[0024] The head movement distance feature extraction method first estimates the three-dimensional coordinates of the middle point between the two eyes to locate the head, then calculates the Euclidean distance between the head positioning points in each two adjacent frames in the three-dimensional space to measure the head movement distance, and finally quantifies the head movement distance according to five intervals.

[0025] A further technical solution of the present invention is as follows: In step 3, the mapping relationship between the multimodal feature sequence and the classification label is established using the LSTM network:

[0026] For the head and face videos of m children with visual attention abnormalities and m children with normal development when watching the test video, the head and face videos are sampled at a fixed frequency, and finally M frames of data are extracted from the total video of each observer for feature extraction. Then the multimodal feature sequence of each observer's video is F = {F1, F2, ..., F i ,…}, where F i Represents the multimodal features corresponding to the i-th paragraph;

[0027] F i Eye movement characteristics ET i , facial expression features FE i and head movement features HM i It consists of three parts, namely F i ={ET i ,FE i ,HM i}; where ET iET is the 5-dimensional gaze point jump amplitude histogram, 9-dimensional gaze point jump angle histogram and 4-dimensional gaze point semantic area percentage feature corresponding to the i-th video segment; i Corresponding to the 8-dimensional facial expression histogram, HM i Corresponding to the 9-dimensional head posture histogram and the 5-dimensional head movement distance histogram;

[0028] The multimodal feature sequences extracted from multiple continuous video segments are used as input, and the corresponding labels are used as output to train the LSTM network. The labels are 1 for children with abnormal visual attention and 0 for normal children. The mapping relationship between the multimodal feature sequences and the labels is obtained to realize the construction of the visual attention abnormality screening model.

[0029] A computer system, characterized in that it includes: one or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method.

[0030] A computer-readable storage medium is characterized by storing computer-executable instructions, which are used to implement the above method when executed.

[0031] Beneficial effects

[0032] The present invention provides a mobile terminal child visual attention abnormality screening method based on multimodal data learning, which has the following beneficial effects:

[0033] (1) The present invention focuses on the differentiation problem in visual attention. Different from the traditional analysis and research based on statistical theory, the present invention constructs a visual attention anomaly recognition model based on deep neural network from the perspective of learning, and realizes the direct use of mobile visual attention data for visual attention anomaly screening.

[0034] (2) Regarding the eye movement data, unlike traditional studies based on the comparison of low-level eye movement features, this paper introduces semantic perception differences into the model by setting regions of interest, and describes eye movement attributes by combining low-level gaze point features with high-level semantic features. On the other hand, in addition to eye movement data, this paper also focuses on other modal visual attention data such as facial expressions and head posture, and uses an LSTM network to fuse different modal features to more comprehensively analyze and reconstruct the observer's visual attention process, reduce the missed detection rate, and improve the model's anomaly recognition ability.

[0035] (3) Based on the research on visual attention using eye trackers, the present invention starts with the data collected from the mobile terminal to construct a low-cost and easy-to-promote visual attention abnormality screening model, which reflects abnormal behavior in an objective way and can help backward and remote areas promote large-scale screening of visual attention abnormality diseases. It has important research significance and industrial application value.

[0036] The present invention can be extended to various types of recognition and classification applications based on visual attention abnormalities, such as autism recognition, attention deficit and hyperactivity disorder recognition, etc., by changing the training samples. It can also be used as a feature and combined with other machine learning methods in applications such as target detection and recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like parts throughout the drawings.

[0038] Figure 1 It is the overall flow chart for realizing the present invention;

[0039] Figure 2 Schematic diagrams of the calibration video and test video of the mobile terminal in the present invention: (a) Schematic diagram of calibration video 1; (b) Schematic diagram of calibration video 2; (c) Schematic diagram of the test video;

[0040] Figure 3 This is a flow chart of eye movement estimation in the present invention;

[0041] Figure 4 This is a diagram of the eye movement estimation network model in the present invention;

[0042] Figure 5 Schematic diagram of the gaze point jump amplitude and jump angle in the present invention;

[0043] Figure 6 This is the facial expression recognition network model diagram in the present invention. DETAILED DESCRIPTION

[0044] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0045] The technical solution of the present invention is: first, set up a calibration video and a test video, and record the head and face videos of m children when they watch the calibration video and the test video on a smartphone respectively; secondly, build an eye movement estimation model, predict the gaze point position frame by frame from the head and face video corresponding to the test video, and extract eye movement features from multiple angles such as the gaze point jump amplitude, gaze point jump angle and area of ​​interest; next, extract facial expression features and head posture features from the head and face video corresponding to the test video respectively; then use the Long Short-Term Memory (LSTM) network to fuse different modal features to achieve mapping from multimodal features to category labels. Finally, in the test phase, record the head and face videos of the children to be classified when they watch the mobile video, and extract features such as eye movement, facial expression and head posture and input them into the trained model to determine whether they are abnormal. Its implementation steps include the following:

[0046] (1) Mobile visual attention data collection

[0047] Two calibration videos and one test video were played using a smartphone application. Meanwhile, the head and face videos of m children with visual attention abnormalities and m children with normal development were recorded while they watched calibration video 1, calibration video 2, and test video.

[0048] Calibration video 1: 72 seconds of video, black background, green dots, 40 dots flashing randomly, each in a random position, lit for 1.8 seconds, and the size varies from 18 to 50 pixels.

[0049] Calibration Video 2: 72 seconds of video with a black background. Green dots move in a zigzag pattern across the screen. Each dot takes 1.8 seconds, and about 40 dots move smoothly.

[0050] Test video: A 2-minute video composed of a short clip of a children's educational game and a video of geometric changes.

[0051] The mobile phone is placed horizontally, and the child sits about 30cm-40cm away from the phone to watch the mobile phone video. While the video is playing, the mobile phone camera is used to record the child's head and face video.

[0052] (2) Eye movement estimation and feature extraction

[0053] An eye movement estimation model is constructed to predict the gaze point position frame by frame; on the other hand, the eye movement features corresponding to each head and face video in the test video are extracted to prepare for the training of a multimodal visual attention abnormality screening model.

[0054] ①Eye movement estimation

[0055] To obtain more accurate eye movement estimation results, a general eye movement estimation model was first trained using head and face videos recorded by mobile phone cameras during the playback of calibration video 1 from the data of m typically developing children. The general eye movement estimation results were then personalized and calibrated using head and face videos recorded during the playback of calibration video 2. During model training, the center coordinates of the green origin in the corresponding video frame were used as the ground truth for the eye movement estimation.

[0056] First, video frames are sampled from the head and face videos corresponding to the two calibration videos at a frequency of 30 frames per second. At the same time, a single-shot multibox detector (SSD) is used to detect the face and left and right eyes in each video frame to obtain face images, left and right eye images, and corresponding bounding boxes. Next, the face and left and right eye images are converted to a fixed size, and the corresponding bounding boxes are normalized. A convolutional neural network is used to extract features from the face and left and right eye images respectively. The left and right eyes share the convolutional neural network weights, and then a fully connected layer is used to extract features from the corresponding bounding boxes to obtain the relative position features of the face and left and right eyes in the image. Finally, the fully connected layer is used to fuse the above-mentioned extracted features, and the final eye movement estimation result is obtained through the fully connected layer.

[0057] In eye movement estimation model training, a universal eye movement estimation model is first trained using data from different observers corresponding to calibration video 1. However, since the universal eye movement estimation model is trained from data from different observers, different observers have different inherent biases—inherent optical axis and visual axis biases (i.e., the person's kappa angle). Therefore, it is necessary to use the data from a specific observer in calibration video 2 as a calibration training set to calibrate the universal eye movement estimation model to account for these inherent biases. Specifically, the fused features extracted from the universal eye movement estimation model (i.e., the feature output of the penultimate layer of the universal eye movement model) are used as input. A support vector regression (SVR) algorithm is then used to obtain personalized calibration prediction results. The personalized SVR model is then trained using the data corresponding to the observer-specific calibration video 2, resulting in more accurate prediction results.

[0058] ② Eye movement feature extraction

[0059] The eye movement estimation model described above is used to estimate the gaze position of the head and facial video frames corresponding to the test video. Eye movement features are then extracted from three perspectives: gaze jump amplitude, gaze jump angle, and the number of gaze points in the semantic area. First, to extract the gaze jump amplitude feature, the Euclidean distance between the gaze points in two adjacent frames in two-dimensional space is calculated and divided into five intervals. Second, to extract the gaze jump angle feature, the angle between the line connecting the gaze jumps of two adjacent frames and the X-axis is calculated. The 360-degree angle is then evenly divided into eight intervals, and a 9-dimensional histogram feature is generated for no-gaze jumps.

[0060] Since both gaze point jump amplitude and angle are underlying eye movement features, to detect visual attention anomalies from a semantic perception perspective, we first annotate the video's regions of interest (ROIs)—the face, body, and other salient objects other than the human body. We then calculate the percentage of fixations on the face, body, other salient objects, and background in each video segment relative to the total number of fixations in that segment as semantic eye movement features. In summary, in the eye movement modality, each video segment corresponds to an 18-dimensional feature vector, 14 of which are underlying histogram features of gaze point jump amplitude and angle, and 4 correspond to the percentage of fixations in different semantic regions.

[0061] (3) Facial expression feature extraction

[0062] Extract the visual attention features under the facial expression modality corresponding to each head and face video in the test video to prepare for training the multimodal visual attention abnormality screening model.

[0063] First, face detection is performed using LibFaceDetection technology, which crops faces from the video. LibFaceDetection is faster than traditional multi-layer feedforward convolutional network algorithms and requires less computing power. If no face is detected, the video frame is classified as "no face detected." If a face is detected, the segmented region is used for subsequent facial expression feature extraction and classification.

[0064] Next, a facial expression recognition network was constructed using the MobileNetV3-Small network based on transfer learning to extract facial expression features from video data. MobileNetV3-Small is a network model designed specifically for mobile devices and embedded vision. It is lightweight, resulting in improved inference speed and accuracy. Transfer learning is a machine learning method that allows a model trained on one task to serve as the initial model for training a second, different task. During training, the main MobileNetV3-Small network was first trained using the ImageNet image recognition database to obtain the network's initial pre-trained parameters. Two fully connected layers were then sequentially connected to the output of the MobileNetV3-Small network, with a batch normalization (BN) layer added before each to prevent overfitting. The final layer was connected to a softmax layer to output the probabilities of seven facial expressions. After constructing the complete facial expression recognition network, the model was globally fine-tuned using the Extended Cohn-Kanade Dataset (CK+), a classic facial expression dataset, to improve recognition accuracy. This resulted in a facial expression recognition model based on the MobileNetV3-Small network.

[0065] Finally, facial expression features are extracted and a facial expression feature histogram is constructed. In each video frame, faces are detected and cropped using LibFaceDetection technology. If no face is detected, the video frame is marked as no face detected. If a face is detected, the cropped face data is input into a trained facial expression recognition network for classification and recognition (a total of seven facial expression categories: neutral, disgust, anger, sadness, fear, happiness, and surprise). Finally, the facial expression recognition results of all frames in each video are counted to generate an 8-dimensional facial expression histogram (neutral, disgust, anger, sadness, fear, happiness, surprise, and no face detected). In summary, in the facial expression modality, each video corresponds to an 8-dimensional facial expression histogram feature.

[0066] (4) Head motion feature extraction

[0067] Extract the visual attention features under the head movement mode corresponding to each head and face video in the test video to prepare for training the multimodal visual attention abnormality screening model.

[0068] Head movement is also an important attribute in the visual attention process, which can effectively reflect abnormal visual attention behavior. The present invention will describe the head movement characteristics from two perspectives: head posture and head movement distance. In order to extract the head posture features, a constrained local neural field (CLNF) facial landmark detector is first used to estimate different head posture angles, and then the head postures (lowering the head, raising the head, turning left and right, and tilting the head left and right) are classified by setting a threshold. The head posture angles include: head pitch angle p, head deflection angle y, and head roll angle r. According to the three head posture angles, the specific division of head posture is as follows:

[0069]

[0070] To extract head motion distance features, we first estimate the 3D coordinates of the midpoint between the two eyes to locate the head. We then calculate the Euclidean distance between the head positioning points in each adjacent frame to measure head motion distance. Finally, we quantify head motion distance according to five intervals. In summary, in the head motion mode, each video segment corresponds to a 14-dimensional histogram feature, with the three head posture angles corresponding to a 9-dimensional head posture histogram partition, and the head motion distance corresponding to five histogram interval partitions.

[0071] (5) Construction of a visual attention anomaly screening model based on multimodal feature fusion

[0072] Based on the above extraction of multimodal features such as eye movements, facial expressions, and head posture from the head and face video segments corresponding to the test video, the task of this step is to construct a multimodal feature sequence and then use the LSTM network to establish a mapping relationship between the multimodal feature sequence and the classification label.

[0073] First, the head and face videos of m children with visual attention disorders and m children with normal development were sampled at a frequency of 10 frames per second while they watched the test video. Finally, M frames of data were extracted from the total video of each observer. 50 frames were then used as a segment for feature extraction. The multimodal feature sequence of each observer's video was F = {F1, F2, ..., F i ,…}, where F i Represents the multimodal features corresponding to 50 frames of data in the i-th paragraph.

[0074] F i Eye-Tracking feature ET i , Facial Expression feature FE i and head movement characteristics HM i It consists of three parts, namely Fi ={ET i ,FE i ,HM i}. Among them ET i ET is the 5-dimensional gaze point jump amplitude histogram, 9-dimensional gaze point jump angle histogram and 4-dimensional gaze point semantic area (face, body, other salient objects and background) percentage features corresponding to the i-th video segment. i Corresponding to the 8-dimensional facial expression histogram, HM i This corresponds to a 9-dimensional head pose histogram and a 5-dimensional head motion distance histogram. In summary, each video segment corresponds to a 40-dimensional multimodal feature vector.

[0075] Finally, in order to train the visual attention abnormality screening model, the multimodal feature sequence extracted from the total video (multiple continuous video segments) is used as input, and the corresponding labels (children with visual attention abnormality are marked as 1, and normal children are marked as 0) are used as output to train the LSTM network, obtain the mapping relationship between the multimodal feature sequence and the label, and realize the construction of the visual attention abnormality screening model.

[0076] (6) Identification of visual attention abnormalities in children to be determined

[0077] For children of the to-be-determined category, we first record their head and facial videos when they watch calibration video 2 and the test video on a smartphone, and then use the head and facial videos corresponding to calibration video 2 to perform personalized fine-tuning on the eye movement estimation network.

[0078] The head and face video corresponding to the test video is sampled at a frequency of 10 frames per second and then segmented into video sequences of 50 frames. For each segment, the corresponding eye movement, facial expression, and head movement features are extracted to generate a 40-dimensional feature map. The features corresponding to multiple segments of the entire video are then combined to construct a multimodal feature sequence, which is then input into an LSTM network for anomaly detection.

[0079] In order to enable those skilled in the art to better understand the present invention, the present invention is described in detail below with reference to specific embodiments.

[0080] Example 1: Childhood Autism Screening Method Based on Multimodal Data Learning

[0081] Reference Figure 1 , the specific implementation steps of the present invention are as follows:

[0082] Step 1: Mobile visual attention data collection

[0083] Fifty autistic children and 50 normally developing children were recruited for the experiment.

[0084] Reference Figure 2, two calibration videos and one test video were played using a smartphone application, and the corresponding head and facial videos of the children were recorded while they watched calibration video 1, calibration video 2, and test video.

[0085] Calibration video 1: 72 seconds of video, black background, green dots, 40 dots flashing randomly, each in a random position, lit for 1.8 seconds, and the size varies from 18 to 50 pixels.

[0086] Calibration Video 2: 72 seconds of video with a black background. Green dots move in a zigzag pattern across the screen. Each dot takes 1.8 seconds, and about 40 dots move smoothly.

[0087] Test video: A 2-minute video composed of a short clip of a children's educational game and a video of geometric changes.

[0088] The mobile phone is placed horizontally, and the child sits about 30cm-40cm away from the phone to watch the mobile phone video. While the video is playing, the mobile phone camera is used to record the child's head and face video.

[0089] Step 2: Eye Movement Estimation and Feature Extraction

[0090] An eye movement estimation model is constructed to predict the gaze point position frame by frame; on the other hand, the eye movement features corresponding to each head and face video in the test video are extracted to prepare for the training of a multimodal visual attention abnormality screening model.

[0091] ①Eye movement estimation

[0092] To obtain more accurate eye movement estimation results, refer to Figure 3 First, a general eye movement estimation model is trained using head and face videos recorded by a mobile phone camera during the playback of calibration video 1 from a dataset of m typically developing children. The general eye movement estimation results are then personalized and calibrated using head and face videos recorded during the playback of calibration video 2. During model training, the center coordinates of the green origin in the corresponding video frame are used as the ground truth for the eye movement estimation.

[0093] Reference Figure 4First, video frames are sampled from the head and face videos corresponding to the two calibration videos at a frequency of 30 frames per second. At the same time, a single-shot multibox detector (SSD) is used to detect the face and left and right eyes in each video frame to obtain face images, left and right eye images, and corresponding bounding boxes. Next, the face and left and right eye images are converted to a fixed size, and the corresponding bounding boxes are normalized. The convolutional neural network is used to extract features from the face and left and right eye images respectively. The left and right eyes share the convolutional neural network weights, and then the fully connected layer is used to extract features from the corresponding bounding boxes to obtain the relative position features of the face and left and right eyes in the image. Finally, the fully connected layer is used to fuse the above-mentioned extracted features, and the final eye movement estimation result is obtained through the fully connected layer.

[0094] In eye movement estimation model training, a universal eye movement estimation model is first trained using data from different observers corresponding to calibration video 1. However, since the universal eye movement estimation model is trained from data from different observers, different observers have different inherent biases—inherent optical axis and visual axis biases (i.e., the person's kappa angle). Therefore, it is necessary to use the data from a specific observer in calibration video 2 as a calibration training set to calibrate the universal eye movement estimation model to account for these inherent biases. Specifically, the fused features extracted from the universal eye movement estimation model (i.e., the feature output of the penultimate layer of the universal eye movement model) are used as input. A support vector regression (SVR) algorithm is then used to obtain personalized calibration prediction results. The personalized SVR model is then trained using the data corresponding to the observer-specific calibration video 2, resulting in more accurate prediction results.

[0095] ② Eye movement feature extraction

[0096] The above eye movement estimation model is used to estimate the gaze point position of the head and face video frames corresponding to the test video, and then the eye movement features are extracted from three aspects: the gaze point jump amplitude, the gaze point jump angle and the number of gaze points in the semantic area. First, in order to extract the gaze point jump amplitude feature, refer to Figure 5 , calculate the Euclidean distance between the gaze points of each two adjacent frames in two-dimensional space and divide it into five intervals. Secondly, to extract the gaze point jump angle feature, calculate the angle between the line connecting the gaze point jumps of each two adjacent frames and the X-axis, and then evenly divide the 360 ​​degrees into 8 intervals and no jumps to generate a 9-dimensional histogram feature.

[0097] Since both gaze point jump amplitude and angle are underlying eye movement features, to detect visual attention anomalies from a semantic perception perspective, we first annotate the video's regions of interest (ROIs)—the face, body, and other salient objects other than the human body. We then calculate the percentage of fixations on the face, body, other salient objects, and background in each video segment relative to the total number of fixations in that segment as semantic eye movement features. In summary, in the eye movement modality, each video segment corresponds to an 18-dimensional feature vector, 14 of which are underlying histogram features of gaze point jump amplitude and angle, and 4 correspond to the percentage of fixations in different semantic regions.

[0098] Step 3: Facial expression feature extraction

[0099] First, the visual attention features of each head and face video under the facial expression modality corresponding to the test video are extracted to prepare for the training of the multimodal visual attention abnormality screening model.

[0100] First, face detection is performed using LibFaceDetection technology, which crops faces from the video. LibFaceDetection is faster than traditional multi-layer feedforward convolutional network algorithms and requires less computing power. If no face is detected, the video frame is classified as "no face detected." If a face is detected, the segmented region is used for subsequent facial expression feature extraction and classification.

[0101] Next, refer to Figure 6A facial expression recognition network was constructed using the MobileNetV3-Small network based on transfer learning to extract facial expression features from video data. MobileNetV3-Small is a network model designed specifically for mobile devices and embedded vision. It is lightweight, improving both inference speed and accuracy. Transfer learning is a machine learning method that allows a model trained on one task to serve as the initial model for training a second, different task. During training, the main MobileNetV3-Small network was first trained using the ImageNet image recognition database to obtain the network's initial pre-trained parameters. Two fully-connected layers were then sequentially connected to the output of the MobileNetV3-Small network, with a batch normalization (BN) layer added before each to prevent overfitting. The final layer was connected to a softmax layer to output the probabilities of seven facial expressions. After constructing the complete facial expression recognition network, the model was globally fine-tuned using the Extended Cohn-Kanade Dataset (CK+), a classic facial expression dataset, to improve recognition accuracy. The resulting facial expression recognition model based on the MobileNetV3-Small network was constructed.

[0102] Finally, facial expression features are extracted and a facial expression feature histogram is constructed. In each video frame, faces are detected and cropped using LibFaceDetection technology. If no face is detected, the video frame is marked as no face detected. If a face is detected, the cropped face data is input into a trained facial expression recognition network for classification and recognition (a total of seven facial expression categories: neutral, disgust, anger, sadness, fear, happiness, and surprise). Finally, the facial expression recognition results of all frames in each video are counted to generate an 8-dimensional facial expression histogram (neutral, disgust, anger, sadness, fear, happiness, surprise, and no face detected). In summary, in the facial expression modality, each video corresponds to an 8-dimensional facial expression histogram feature.

[0103] Step 4: Head motion feature extraction

[0104] Extract the visual attention features under the head movement mode corresponding to each head and face video in the test video to prepare for training the multimodal visual attention abnormality screening model.

[0105] Head movement is also an important attribute in the visual attention process, which can effectively reflect abnormal visual attention behavior. The present invention will describe the head movement characteristics from two perspectives: head posture and head movement distance. In order to extract the head posture features, a constrained local neural field (CLNF) facial landmark detector is first used to estimate different head posture angles, and then the head postures (lowering the head, raising the head, turning left and right, and tilting the head left and right) are classified by setting a threshold. The head posture angles include: head pitch angle p, head deflection angle y, and head roll angle r. According to the three head posture angles, the specific division of head posture is as follows:

[0106]

[0107] To extract head motion distance features, we first estimate the 3D coordinates of the midpoint between the two eyes to locate the head. We then calculate the Euclidean distance between the head positioning points in each adjacent frame to measure head motion distance. Finally, we quantify head motion distance according to five intervals. In summary, in the head motion mode, each video segment corresponds to a 14-dimensional histogram feature, with the three head posture angles corresponding to a 9-dimensional head posture histogram partition, and the head motion distance corresponding to five histogram interval partitions.

[0108] Step 5: Construction of a visual attention anomaly screening model based on multimodal feature fusion

[0109] Based on the above extraction of multimodal features such as eye movements, facial expressions, and head posture from the head and face video segments corresponding to the test video, the task of this step is to construct a multimodal feature sequence and then use the LSTM network to establish a mapping relationship between the multimodal feature sequence and the classification label.

[0110] First, the head and face videos of 50 autistic children and 50 normally developing children watching the test video were sampled at a frequency of 10 frames per second, and M frames of data were finally extracted from the total video of each observer. The 50 frames were then used as a segment for feature extraction, and the multimodal feature sequence of each observer's video was F = {F1, F2, ..., F i ,…}, where F i Represents the multimodal features corresponding to 50 frames of data in the i-th paragraph.

[0111] F i Eye-Tracking feature ET i , Facial Expression feature FE i and head movement characteristics HM i It consists of three parts, namely Fi ={ET i ,FE i ,HM i}. Among them ET i ET is the 5-dimensional gaze point jump amplitude histogram, 9-dimensional gaze point jump angle histogram and 4-dimensional gaze point semantic area (face, body, other salient objects and background) percentage features corresponding to the i-th video segment. i Corresponding to the 8-dimensional facial expression histogram, HM i This corresponds to a 9-dimensional head pose histogram and a 5-dimensional head motion distance histogram. In summary, each video segment corresponds to a 40-dimensional multimodal feature vector.

[0112] Finally, in order to train the visual attention anomaly screening model, the multimodal feature sequence extracted from the total video (multiple continuous video segments) is used as input, and the corresponding labels (autistic children are marked as 1 and normal children are marked as 0) are used as output to train the LSTM network, obtain the mapping relationship between the multimodal feature sequence and the label, and realize the construction of the visual attention anomaly screening model.

[0113] Step 6: Identification of children with autism

[0114] For children of the to-be-determined category, we first record their head and facial videos when they watch calibration video 2 and the test video on a smartphone, and then use the head and facial videos corresponding to calibration video 2 to perform personalized fine-tuning on the eye movement estimation network.

[0115] The head and face video corresponding to the test video is sampled at a frequency of 10 frames per second and then segmented into video sequences of 50 frames. For each segment, the corresponding eye movement, facial expression, and head movement features are extracted to generate a 40-dimensional feature corresponding to the segment. The features corresponding to multiple segments of the entire video are then combined to construct a multimodal feature sequence, which is finally input into an LSTM network for autism recognition (children with autism are marked as 1, and children without autism are marked as 0).

[0116] The method of the present invention is different from the traditional visual attention differentiation research based on statistical analysis. The goal of the present invention is to build a model for screening children's visual attention abnormalities, that is, to collect head and facial videos of children when watching videos on smartphones as model input, and directly judge whether there is any abnormality in the child's visual attention process based on the trained model, rather than simply comparing the differences in visual attention of different categories of people. Unlike traditional visual attention research based on eye trackers, the present invention will explore eye movement prediction models based on mobile terminals such as smartphones, realize visual attention abnormality screening based on mobile terminals, reduce the cost of visual attention abnormality screening, and help solve the problem of screening for diseases related to visual attention abnormalities in children in remote areas. On the basis of traditional visual attention research based on separate eye movement studies, the present invention will simultaneously extract eye movement, facial expression and head posture features, use data from different modalities of the same data source for fusion recognition, and more comprehensively analyze children's visual attention process, reduce the missed detection rate, and improve the model's abnormality recognition ability.

[0117] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present invention, and these modifications or replacements should all be included in the scope of protection of the present invention.

Claims

1. A mobile terminal children's visual attention abnormality screening method based on multimodal data learning, characterized by Here are the steps: Step 1: Set up the calibration video and test video and record them separately m Children with visual attention disorders m Head and face videos of typically developing children watching calibration and test videos on smartphones; Step 2: Construct an eye movement estimation model. Use the eye movement estimation model to estimate the gaze point position of the head and face video frames corresponding to the test video, and extract eye movement features from the gaze point jump amplitude, gaze point jump angle, and region of interest. The construction of the eye movement estimation model is as follows: First, video frames are sampled from the head and face videos corresponding to the two calibration videos. A single-stage multi-layer detector is used to detect the face and left and right eyes in each video frame, resulting in a face image, left and right eye images, and corresponding bounding boxes. Next, the face and left and right eye images are converted to a fixed size, the corresponding bounding boxes are normalized, and convolutional neural networks are used to extract features from the face and left and right eye images respectively. The left and right eyes share the convolutional neural network weights, and a fully connected layer is used to extract features from the corresponding bounding boxes to obtain the relative position characteristics of the face and the left and right eyes in the image. Finally, a fully connected layer is used to fuse the extracted features to obtain the final eye movement estimation result. Step 3: Extract facial expression features and head movement features from the head and face video corresponding to the test video; Then, the long short-term memory (LSTM) network is used to fuse the different modal features to achieve the mapping from multimodal features to category labels. The head motion features include head posture features and head motion distance features. The head posture feature extraction method uses a constrained local neural domain facial landmark detector to estimate different head posture angles, and then classifies the head posture by setting a threshold. The head posture angles include: head pitch angle , head tilt angle Head roll angle ; According to the three head posture angles, the specific classification of head posture is as follows: The head movement distance feature extraction method first estimates the three-dimensional coordinates of the middle point between the two eyes to locate the head, then calculates the Euclidean distance between the head positioning points in each two adjacent frames in the three-dimensional space to measure the head movement distance, and finally quantifies the head movement distance according to five intervals; Step 4: During the testing phase, record the head and face videos of the children to be classified while watching mobile videos, and extract multimodal features of eye movements, facial expressions, and head movements, and input them into the trained LSTM network model to determine whether there are any abnormalities.

2. The method for screening children's visual attention abnormalities on a mobile terminal based on multimodal data learning according to claim 1 is characterized by: Step 1 is as follows: Use the smartphone app to play two calibration videos and one test video, and record them separately. m Children with visual attention disorders m The corresponding head and face videos of the normally developing children when they watched calibration video 1, calibration video 2 and test video; The calibration video 1 is a 72-second video with a black background and 40 green dots flashing randomly, each in a random position, for 1.8 seconds, with the size varying from 18 to 50 pixels. The calibration video 2 is a 72-second video with a black background and green dots moving in a zigzag pattern on the screen. Each dot takes 1.8 seconds, and the 40 dots move smoothly. The test video is a 2-minute video composed of a short clip of a children's educational game and a video of geometric changes; The mobile phone is placed horizontally, and the child sits about 30cm-40cm away from the phone to watch the mobile phone video. While the video is playing, the mobile phone camera is used to record the child's head and face video.

3. The method for screening abnormal visual attention of children on a mobile terminal based on multimodal data learning according to claim 1 is characterized in that The facial expression feature extraction described in step 3 is specifically as follows: in each video frame, the LibFaceDetection technology is used to detect and crop the face. If no face is detected, the video frame is marked as no face is detected. If a face is detected, the cropped face data is input into the trained facial expression recognition network for classification and recognition; finally, the facial expression recognition results of all frames in each video are counted to generate a facial expression histogram.

4. The method for screening abnormal visual attention of children on a mobile terminal based on multimodal data learning according to claim 1 is characterized in that In step 3, the mapping relationship between the multimodal feature sequence and the classification label is established using the LSTM network as follows: against m Children with visual attention disorders m The head and face videos of normally developing children when they watched the test video were sampled at a fixed frequency, and finally the head and face videos were extracted from the total video of each observer. The feature extraction of frame data is performed, and the multimodal feature sequence of each observer's video is ,in Indicates the i Multimodal features corresponding to paragraphs; Eye movement characteristics , facial expression features and head movement characteristics It consists of three parts, namely ;in For this i 5-dimensional gaze point jump amplitude histogram, 9-dimensional gaze point jump angle histogram and 4-dimensional gaze point semantic area percentage features corresponding to each video segment; Corresponding to the 8-dimensional facial expression histogram, Corresponding to the 9-dimensional head posture histogram and the 5-dimensional head movement distance histogram; The multimodal feature sequences extracted from multiple continuous video segments are used as input, and the corresponding labels are used as output to train the LSTM network. The labels are 1 for children with abnormal visual attention and 0 for normal children. The mapping relationship between the multimodal feature sequences and the labels is obtained to realize the construction of the visual attention abnormality screening model.

5. A computer system, characterized in that include: One or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method of claim 1.

6. A computer-readable storage medium, characterized in that Computer-executable instructions are stored, and when the instructions are executed, they are used to implement the method of claim 1.

Citation Information

Patent Citations

  • Multi-feature fusion driver abnormal expression recognition method

    CN110334600A

  • Child ADHD screening evaluation system based on a multi-modal deep learning technology

    CN111528859A