System and method for autism identification with combined coarse and fine granularity
The system integrates coarse-grained and fine-grained feature extraction and fusion for enhanced autism identification, addressing reliability and accuracy issues in existing methods by leveraging facial expression analysis for improved diagnostic precision and early detection.
Patent Information
- Application Number
- JP2025127128
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-05
- Filing Date
- 2025-07-30
- Publication Date
- 2026-02-18
AI Technical Summary
Existing autism identification methods face challenges in reliability and accuracy due to subjective evaluations and limited multidimensional diagnostic information, hindering early detection and comprehensive understanding of facial behaviors in autism.
A system and method combining coarse-grained and fine-grained feature extraction and fusion, utilizing facial expression analysis to integrate multiple levels of cognitive impairment and functioning, including data acquisition, face detection, head pose estimation, and feature fusion through SEnet networks and LSTM neural networks for improved autism identification.
Enhances the reliability and accuracy of autism identification by providing a comprehensive explanation of facial behaviors, improving diagnostic precision and facilitating early intervention.
Smart Images

Figure 2026027197000001_ABST
Abstract
Description
[Technical Field]
[0001]
[0001] The present invention relates to the field of computer vision, specifically to coarse-grained and fine-grained The present invention relates to an autism identification system and method that combines [Background technology]
[0002]
[0002] Facial behavior analysis plays an important role in autism identification, mainly in the following areas: It appears in several directions. (1) Early Identification and Intervention: Autism typically manifests early in a child's life with specific social and communication challenges. Facial behavior analysis can help medical professionals and researchers detect early signs of autism. This can help identify and facilitate early intervention and treatment, improving treatment outcomes. do. (2) Objective evaluation: Facial behavior analysis is used to analyze characteristics of autistic patients, such as facial expressions, eye contact, and facial movements. This objectivity is not based on subjective evaluation. This helps reduce errors and increase diagnostic reliability and accuracy. (3) Complementary multidimensional diagnosis: Facial behavior analysis is typically used as an adjunct to other assessment tools in autism diagnosis. By combining it with other methods (e.g., behavioral observation, questionnaire survey, etc.), more comprehensive and multidimensional diagnostic information can be obtained. This can help the medical team better understand the patient's overall condition. , which helps to evaluate. (4) Research and development of new treatment methods: Facial behavior analysis can help researchers understand the pathogenesis and pathophysiology of autism. This study provides an important tool for studying facial behavioral changes and understanding treatment response. This will allow researchers to explore more deeply the neurobiological basis of autism and potentially develop new treatments. In short, facial behavior analysis can provide clues for the development of new methods for detecting potential autism. Not only can it help detect symptoms early, but it can also promote individualized and precise treatment.
[0003] With the development of deep learning theory, facial expression analysis technology has made significant progress. However, existing autism identification methods still face many challenges. [Summary of the Invention]
[0003]
[0004] In response to the above-mentioned deficiencies or needs for improvement of the current technology, the present invention provides a method for producing a coarse grain and a fine grain. A system and method for identifying autism that combines multiple levels of cognitive impairment and cognitive functioning to provide a comprehensive explanation of facial behavior in autism. This can effectively improve the reliability of autism identification.
[0005] In order to achieve the above object, one aspect of the present invention is to It collects slices while the user is watching a video, and each slice is sampled consecutively. a data acquisition module containing a plurality of images acquired by the data acquisition module; The multiple images in each slice are statistically analyzed based on a pre-defined statistical algorithm. a coarse-grained feature extraction module that calculates coarse-grained features for each slice based on the measurement results; Based on the coarse-grained features of the slices, we extract the slices from several different fine-grained feature extraction models. A fine-grained feature extraction module selects one input and acquires fine-grained features for each slice. Rules and The coarse-grained features and the fine-grained features of the slice are subjected to feature fusion to obtain a fusion feature, and the fusion feature is A coarse-grained and fine-grained combined feature fusion and identification module for autism identification based on provides an autism identification system.
[0006] Furthermore, the coarse-grained feature extraction module Face detection is performed on all images in the slice, and face coordinates, dimension information, and facial feature point positions are recorded. a face detection module for determining It invokes the head pose estimation algorithm and uses the face detection results to estimate the user's face in each image in each slice. a head pose estimation module for estimating a user's head pose; We calculate the user head pose variance for all images in each slice and calculate the user head pose variance. A head pose coarse-grained feature extraction module extracts coarse-grained head pose features for each slice. Ru and A face matching module performs face matching on all images in each slice using the face detection results. , The facial expression intensity estimation model is invoked to estimate the user's facial expression intensity in each image in each slice. ,The average facial expression intensity of all images in each slice is calculated. The facial expression intensity coarse-grained feature extraction model extracts the facial expression intensity coarse-grained features of each slice. Jules and Invokes a facial expression estimation model to estimate the user's facial expression category for each image in each slice. Then, we count the total number of times a specified facial expression appears in all images in each slice, and The total number of times the specified facial expression appears is binarized, and the facial expression category coarse-grained features of each slice are calculated. and a facial expression category coarse-grained feature extraction module for obtaining the facial expression category.
[0007] Furthermore, we can statisticize this slice based on the video corresponding to the slice. Determines the facial expression category of the specified facial expression.
[0008] Furthermore, the videos corresponding to the slices are videos that stimulate user joy. If the expression is a happy expression, make sure that the expression specified when calculating this slice is a happy expression. Determine.
[0009] Furthermore, the slices are divided into a plurality of different slices based on the coarse-grained features of the slices. Selecting the input to one of the fine-grained feature extraction models The total number of times a specified facial expression appears in all images in a slice exceeds a preset threshold. inputting the slice into a first fine-grained feature extraction model; The total number of times a specified facial expression appears in all images in a slice exceeds a preset threshold. If not, inputting this slice into a second fine-grained feature extraction model.
[0010] Furthermore, the first fine-grained feature extraction model and the second fine-grained feature extraction model are feature extraction models based on long short-term memory neural networks, The training sample sets of the first fine-grained feature extraction model and the second fine-grained feature extraction model are different. do.
[0011] Furthermore, the feature fusion of the coarse-grained features and fine-grained features of the slices is reducing the fine-grained features of the slice; The fine-grained features of the slices after dimensional reduction and the coarse-grained features of the slices are fed into the SEnet network. The attention weights are calculated based on the input, and the fine-grained features of the slices are then extracted based on the attention weights. This involves feature fusion of Rice's coarse-grained features.
[0012] Furthermore, performing autism identification based on the fusion features may include The combined features are input into a binary classifier to obtain an autism identification result.
[0013] Furthermore, collecting slices while the user is watching a video is ,collect multiple slices while a user is watching multiple videos, and each video is a slice To respond to Feature fusion of the coarse-grained features and fine-grained features of the slices is This involves feature fusion of coarse-grained features and fine-grained features.
[0014] Another aspect of the invention involves collecting slices while a user is watching a video; each slice containing a plurality of consecutively sampled images; The multiple images in each slice are statistically analyzed based on a pre-defined statistical algorithm. calculating coarse-grained features for each slice based on the measurement results; Based on the coarse-grained features of the slices, we extract the slices from several different fine-grained feature extraction models. Selecting one input to obtain fine-grained features for each slice; The coarse-grained features and the fine-grained features of the slice are subjected to feature fusion to obtain a fusion feature, and the fusion feature is applied to and performing autism identification based on the coarse-grained and fine-grained features. do.
[0015] Overall, the above technical solutions conceived by the present invention have advantages over the prior art. Combining coarse-grained feature extraction, fine-grained feature extraction, fusion and classification, it is possible to identify facial movements of autistic people. It can realize a comprehensive explanation of the reasons, and effectively improve the reliability of autism identification. Facial intensity estimation algorithm and facial expression for accurately extracting facial behavioral features of autistic patients Integrate the classification algorithm. In the classification process, first, calculate and analyze statistics to obtain coarse-grained classification. Next, we generate a time series progression analysis using a long short-term memory neural network. Finally, a feature-level attention network is used to capture the features and form a fine-grained analysis result. By using the adaptive fusion of all features, we obtained the final classification result. This can improve the accuracy of autism identification. [Brief explanation of the drawings]
[0004]
[0016] FIG. 1 shows an autism identification system that combines coarse-grained and fine-grained features according to an embodiment of the present invention. 1 is a schematic diagram of the operating principle of the system.
[0017] FIG. 2 shows an autism identification system that combines coarse-grained and fine-grained features according to an embodiment of the present invention. FIG. 1 is a detailed schematic diagram of the operating principle of the system.
[0018] FIG. 3 is a diagram illustrating a network structure of SENet according to an embodiment of the present invention. [Mode for Carrying Out the Invention]
[0005]
[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following drawings and The present invention will now be described in more detail with reference to examples. The embodiments are used only to illustrate the present invention and are not to be used to limit the present invention. It is understood that the technical features according to the various embodiments of the present invention described below are mutually exclusive. They may be combined with each other as long as no conflicts are created between them.
[0020] In describing the embodiments of the present invention, the terms "first" and "second" are used for the purpose of description. used only to indicate or imply the relative importance or number of technical features shown Therefore, the terms "first" and "second" are not specified. A feature may explicitly or implicitly include at least one feature. , means at least two, e.g., two, three, etc., unless otherwise specified.
[0021] Unless otherwise specified, "plurality" means two or more.
[0022] The terms "including" and "having" in the embodiments of the present invention, as well as Any variations of these are intended to cover their inclusion. For example, a series of steps or models Any process, method, device, article or apparatus that includes a module is not specifically listed as a step in the process. It need not be limited to processes or modules not specifically listed or those processes. may include other steps or modules inherent to a process, method, article, or apparatus. This is the intention.
[0023] The naming or numbering of steps described in the embodiments of the present invention is Steps in a method flow in chronological / logical first-come, first-served order, indicated by naming or number It does not mean that you must run a named or numbered flow. - Steps are in accordance with the technical purpose to be achieved if they achieve the same or similar technical effect. The execution order can be changed as needed.
[0024] References herein to "embodiments" mean that the invention is described in connection with the embodiments. It is believed that a particular feature, structure, or characteristic is included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily mean Not all of the above statements refer to the same embodiment, but rather to separate or alternative embodiments that are exclusive of other embodiments. Those skilled in the art will recognize that the embodiments described herein are not intended to be alternative embodiments. It is explicitly and implicitly understood that the embodiments of the present invention can be combined with the embodiments of the present invention.
[0025] The present invention provides a system and method for identifying autism that combines coarse and fine grained , each of which will be explained below.
[0026] In an embodiment of the present invention, a coarse-grained and fine-grained autism identification system is The operating principle is shown in FIGS.
[0027] In an embodiment of the present invention, the coarse-grained and fine-grained autism identification system is , It collects slices while the user is watching a video, and each slice is sampled consecutively. a data acquisition module containing a plurality of images acquired by the data acquisition module; The multiple images in each slice are statistically analyzed based on a pre-defined statistical algorithm. a coarse-grained feature extraction module that calculates coarse-grained features for each slice based on the measurement results; Based on the coarse-grained features of the slices, we extract the slices from several different fine-grained feature extraction models. A fine-grained feature extraction module selects one input and acquires fine-grained features for each slice. Rules and The coarse-grained features and the fine-grained features of the slice are subjected to feature fusion to obtain a fusion feature, and the fusion feature is and a feature fusion and identification module for autism identification based on the feature fusion and identification module.
[0028] The operating principles of each module will be specifically explained below. (1) Data collection module
[0029] The data collection module collects facial expression sequences as users watch the video. You can collect multiple videos and watch them consecutively, based on the start and end times of each video. The clipping continues based on the user's facial expression sequence, and the user's facial expression sequence after the clipping is treated as an image sequence, i.e., Each video is converted into "slices", one video corresponds to one slice, and each slice is a continuous It contains multiple images sampled at regular intervals.
[0030] The video that the user watches stimulates the user's positive emotions (e.g., happy facial expressions). The video may include various types of video, such as video recorded in a video stream or other types of video.
[0031] If a user watches multiple videos, the time during which the user watches the videos Collecting slices allows a user to collect multiple slices while watching multiple videos. Each video corresponds to a slice, and each slice is a continuous stream of video. It contains multiple images sampled periodically. (2) Coarse-grained feature extraction module
[0032] Perform face detection for each slice frame and obtain face coordinates, scale information, and The position of the facial feature points is determined. The detected face and feature point data are used to estimate the head pose. The algorithm is then invoked to estimate the detected head pose. Regarding the head pose variance, if the variance exceeds a preset threshold, the subject is notified. This indicates that there is a mismatch in the face detection results. Ensure that the facial eye coordinates are consistent, thereby ensuring that subsequent facial alignment elements are aligned. Eliminate interference with facial expression analysis. A frame-by-frame estimation of the subject's facial expression intensity is performed, and the intensity is calculated for all images in the slice. The average intensity of facial expressions is calculated, and if the average value exceeds a preset threshold, the subject's emotion is judged. This shows that the expression was activated. Perform frame-by-frame estimation of the subject's facial expression category and select the facial expression specified within the slice. The number of occurrences of emotions (e.g., happy expressions) is counted, and if the number of occurrences is greater than a preset threshold, If a category is higher than the threshold, it indicates that the subject activated the emotion in the corresponding category.
[0033] Furthermore, the coarse-grained feature extraction module: Face detection is performed on all images in the slice, and face coordinates, dimension information, and facial feature point positions are recorded. a face detection module for determining It invokes the head pose estimation algorithm and uses the face detection results to estimate the user's face in each image in each slice. a head pose estimation module for estimating a user's head pose; We calculate the user head pose variance for all images in each slice and calculate the user head pose variance. A head pose coarse-grained feature extraction module extracts coarse-grained head pose features for each slice. Ru and A face matching module performs face matching on all images in each slice using the face detection results. , The facial expression intensity estimation model is invoked to estimate the user's facial expression intensity in each image in each slice. ,The average facial expression intensity of all images in each slice is calculated. The facial expression intensity coarse-grained feature extraction model extracts the facial expression intensity coarse-grained features of each slice. Jules and Invokes a facial expression estimation model to estimate the user's facial expression category for each image in each slice. Then, we count the total number of times a specified facial expression appears in all images in each slice, and The total number of times the specified facial expression appears is binarized, and the facial expression category coarse-grained features of each slice are calculated. and a facial expression category coarse-grained feature extraction module for obtaining the facial expression category.
[0034] The following are coarse-grained feature extraction of head pose, coarse-grained feature extraction of facial expression intensity, and facial expression color. The specific principle of category coarse-grained feature extraction will now be described in detail.
[0035] Coarse-grained analysis is performed on the slice, and the input of the coarse-grained analysis is all the data in the slice. The facial expression intensity, facial expression category, and head pose estimation results for all pictures are shown in Fig. TIFF2026027197000002.tif591, where K represents the number of slices, and in one embodiment, the value of K is 6. Each slice corresponds to a 6-stage emotional stimulus video. , denoted by L. In one embodiment, the value of L is 200, i.e., one slice is 2 Contains 00 pictures. Given TIFF2026027197000003.tif440, the facial expression intensities predicted by six facial expression models are: The range of values for each element in TIFF2026027197000004.tif33 is TIFF2026027197000005.tif425, and the facial expression recognition model predicts the category. The predicted result is "joy" as 1, and the other facial expressions as 2. The facial expression is defined as 0. TIFF2026027197000006.tif312 contains the 3D Euler angles predicted by the head pose estimation algorithm, i.e., pitch, It consists of pitch, roll and yaw angles, and the values of each dimension are The range is The file is TIFF2026027197000007.tif424.
[0036] Coarse-grained analysis calculates different statistics for each of the above variables. For the k-th slice, the facial expression intensities of all pictures in the k-th slice are Average value of Statistic TIFF2026027197000008.tif33. TIFF2026027197000009.tif845In formula, TIFF2026027197000010.tif312 calculates the maximum value of the input data, i.e., the maximum value of the output of the six facial expression intensity estimation models. Take.
[0037] For the k-th slice, the index of all pictures in the k-th slice The number of times the specified facial expression occurred Statistic TIFF2026027197000011.tif33. TIFF2026027197000012.tif829Furthermore, the facial expression category specified when statistically analyzing this slice corresponds to the slice. This is determined based on the video.
[0038] The video corresponding to the slice is a video to stimulate the user's joy. If so, determine that the specified facial expression is a happy facial expression when calculating this slice. .
[0039] For the k-th slice, the beginning of all pictures in the k-th slice Range of change in body posture Statistic TIFF2026027197000013.tif33. TIFF2026027197000014.tif337In formula, TIFF2026027197000015.tif312 and min TIFF2026027197000016.tif33 represents calculating the maximum and minimum values of the input data, respectively.
[0040] The above statistical design discrimination rules were subjected to binarization to obtain binary features. Specifically, the average facial expression intensity of the k-th segment is a threshold. TIFF2026027197000017.tif33(threshold If the facial expression intensity is greater than 0.3, the facial expression intensity is considered to be 1 and activated. In all other cases, it is written as 0. The specific formula is as follows: TIFF2026027197000019.tif1021TIFF2026027197000020.tif33 represents the facial expression intensity coarse-grained feature of the k-th slice.
[0041] The number of occurrences of the joyful expression in the k-th slice is the threshold TIFF2026027197000021.tif34, and in one embodiment, TIFF2026027197000022.tif34 is taken at a value of 10 to suppress short-term noise that may be present. If the subject is considered to be expressing joy, it is marked as 1, otherwise it is marked as 0. The specific formula is as follows: TIFF2026027197000023.tif1022TIFF2026027197000024.tif33 represents the facial expression category coarse-grained features of the k-th slice.
[0042] Any Euler angle of the head pose of the k-th slice is thresholded. TIFF2026027197000025.tif34, and in one embodiment, If TIFF2026027197000026.tif49 takes π / 3, it is considered that attention is off and is written as 0. Otherwise This is denoted as 1. The specific formula is as follows: TIFF2026027197000027.tif1037TIFF2026027197000028.tif33 represent the head pose coarse-grained features of the k-th slice.
[0043] Finally, the coarse-grained features of all slices The resulting file is TIFF2026027197000029.tif33, which will be displayed as follows: TIFF2026027197000030.tif488 In this example, K=6 is used, so TIFF2026027197000031.tif33 outputs 18 feature values as features for coarse-grained analysis. (3) Fine-grained feature extraction module
[0044] The input for fine-grained analysis is slightly different from the input for coarse-grained analysis. Based on the judgment result, fine-grained analysis is performed on the slices to classify different slices, Different slices are input to different fine-grained feature extraction models.
[0045] Based on the coarse-grained characteristics of the slice, the slice is divided into a plurality of different fine-grained characteristics. Selecting it as input to one of the feature extraction models The total number of times a specified facial expression appears in all images in a slice exceeds a preset threshold. inputting the slice into a first fine-grained feature extraction model; The total number of times a specified facial expression appears in all images in a slice exceeds a preset threshold. If not, inputting this slice into a second fine-grained feature extraction model.
[0046] Videos that correspond to slices to stimulate users' emotions of joy If so, it is determined that the specified facial expression when calculating this slice is a happy facial expression. Then, we can divide the slices into two classes. The first class is the one with happy expressions. The second class is slices where the joy expression is not activated. The first slice is input to the first fine-grained feature extraction model, and the second slice is input to the second fine-grained feature extraction model. are input into the output model.
[0047] To account for the confusability of facial expressions in autism, in this example, a classifier is used to predict The observed probability distribution is then converted into fine-grained facial expression category features. Used as TIFF2026027197000032.tif33, TIFF2026027197000033.tif632, TIFF2026027197000034.tif441, where c represents the probability distribution dimension output by the classifier, and in one embodiment is 7 Since we use a facial expression classification model, c = 7. If the probability distribution has multiple peaks, the facial expressions are less likely to be confused. If there is no clear peak even if the facial expression tends to be in the category, it means that the confusability of that facial expression is high. Facial expression intensity characteristics TIFF2026027197000035.tif32 only keeps the prediction results of the joy expression intensity model to remove redundant information. TIFF2026027197000036.tif630, Given TIFF2026027197000037.tif441, the head pose features are consistent with the input of the coarse-grained analysis.
[0048] That is, facial expression category features for extracting fine-grained features and coarse-grained features The facial expression category features for extracting facial expression are different, and the facial expression intensity features for extracting fine-grained features are different. The facial expression intensity features for extracting features and coarse-grained features are different.
[0049] Input data for combining all features and extracting fine-grained features You get TIFF2026027197000038.tif518. TIFF2026027197000039.tif632TIFF2026027197000040.tif43 represent the input data for extracting fine-grained features of the k-th slice. where d=1+7+3=11.
[0050] The characteristics of these two slices are respectively compared to the two corresponding long and short-term memory The data was input to a neural network (LSTM) model to analyze the fine-grained features. The facial expression intensity estimation model outputs positive facial expression intensity features, and the facial expression The facial expression probability distribution output by the emotion category classifier is used as the facial expression category feature, and the head pose estimator The output 3D Euler angles represent the head pose features. The granular features are based on the results of coarse-grained facial expression category analysis and correspond to long-term short-term memory neural networks. Input the LSTM network and generate time series features. Extract TIFF2026027197000041.tif33, the specific formula is as follows: TIFF2026027197000042.tif840, TIFF2026027197000043.tif33 represents the fine-grained features of the k-th slice, TIFF2026027197000044.tif340 has the same network structure as the LSTM, but the trainable performance of the network is different. The parameters are obtained by training on different data. Specifically, coarse-grained analysis is The coarse-grained analysis determines the device data is for training LSTM1, and the coarse-grained analysis determines the Determine whether the data is for training LSTM2. Divide the data using a coarse grain. Then, by inputting different LSTMs, the difficulty of the task of extracting time series features is calculated. The coarse-grained analysis results are similar for autism and typical autism. This allows us to focus on the differences between the subjects and those with developmental disabilities. (4) Feature fusion and identification module
[0051] The feature fusion of the coarse-grained features and the fine-grained features of the slices is reducing the fine-grained features of the slice; The fine-grained features of the slice after dimension reduction and the coarse-grained features of the slice are used in the SEnet network ( The attention weights are calculated by inputting the information into the Squeeze-and-Excitation Networks (SQNs), and the attention weights are calculated based on the information. After the dimension reduction, the fine-grained features of the slices are merged with the coarse-grained features of the slices. .
[0052] Furthermore, when a user watches multiple videos, Collect multiple slices between each video, each corresponding to one slice, and Coarse-grained features and fine-grained features are fused.
[0053] Specifically, the extracted time series features are converted into fine-grained feature indices by linear layer projection. can be obtained, and the specific formula is as follows: In the TIFF2026027197000045.tif559 format, TIFF2026027197000046.tif37 reduces the feature dimension of the kth output signal to e dimensions. In one embodiment, e is set to 1.
[0054] The extracted fine-grained features are reduced in dimension by a linear projection layer, and then coarse-grained features are obtained. It facilitates the fusion of coarse-grained and fine-grained analysis results into a single feature vector. ,The simplified SEnet is input to obtain the feature weights, ,and feature recombination is realized. TIFF2026027197000047.tif435
[0055] TIFF2026027197000048.tif22 is the result of combining coarse-grained analysis and fine-grained analysis. SENet is a convolutional neural network. Extract channel attention from the network and map the features of different channels to attention weights. The purpose of this is to realize the disposal or retention of information based on the network configuration. The structure is shown in Figure 3. In the figure, global pooling is It is used to compress one feature plane into one feature value, and the first fully connected layer is It is used for compression (squeezing), compressing the original C feature channels to C / r dimensions, 2 The th fully connected layer further restores the feature dimension to C dimensions and excites useful features (Excit By compressing and restoring, the attention weights of useful features approach 1, and those of unnecessary features decrease. The attention weight approaches 0. By multiplying this attention weight with the corresponding feature channel, It is possible to retain useful features and suppress unnecessary features. It can be expressed as: TIFF2026027197000049.tif528In formula, TIFF2026027197000050.tif321 are the Sigmoid activation function and the ReLU activation function, respectively. TIFF2026027197000051.tif312 are the trainable parameters of the first and second fully connected layers, respectively. TIFF2026027197000052.tif36 is the attention weight for each channel.
[0056] The input features of an embodiment of the present invention are TIFF2026027197000053.tif3 has 16 dimensions, and each dimension can be considered as a feature channel. Direct access to the two fully connected layers of SEnet without requiring module global pooling. The attention weights can be calculated by the fusion weights. The final output fusion feature TIFF2026027197000054.tif33 is The file is TIFF2026027197000055.tif313. The output features were then subjected to a two-classifier to obtain the final evaluation results.
[0057] We conducted experimental verification of the autism identification system that combines the above coarse-grained and fine-grained methods. It was.
[0058] Below, we propose an intelligent identification method for autism that combines coarse-grained and fine-grained features. We first introduce the self-constructed dataset and experimental setup. ,The effectiveness of each module in the method is explained by ablation studies, and finally, the ,attention ,measurement ,is ,conducted. The interpretability of the method is further illustrated by visualization. (1) Data collection
[0059] To protect the privacy of autistic people, existing research has focused on autistic emotion data collection. Facial expression analysis, in particular, can reveal the identity of people with autism. and their parents may only use the data collected in general for academic research. However, the disclosure of facial data containing personal information is not permitted. Therefore, the present invention aims to The dataset is self-constructed and used to verify the effectiveness of the method described in this paper.
[0060] This dataset consisted of 81 subjects who participated with parental consent. We recruited 40 subjects (ages 4-16 years) as typical developmental subjects (mean age: 5.2 years, variance: 8 months), and 41 subjects with autism (mean age 5.0 years, variance: 14 months). The children with autism were recruited from special schools, and the entry criteria were as follows: (1) A double-blind diagnosis will be conducted by two pediatricians in charge of developmental behavior or deputy chief physicians. 2) The diagnostic criteria are based on the DSM-5 published by the American Psychiatric Association. (3) Ages 4 and up (4) Severe respiratory disease, schizophrenia, epilepsy, or other organic brain disease (5) The visual system is developing normally. (1) Age and autism group matched; (2) suspected or diagnosed; (3) No mental disorders and / or other developmental delays or learning disabilities; It's growing.
[0061] Stimulus materials were selected from movie segments that stimulated positive emotions in the subjects. Film segments are initially screened based on three criteria: ) Reasonable time: Fatigue caused by long viewing times affects the subjective experience of emotions. It is important to choose a segment of an appropriate length, taking into account your willpower and perseverance. (2) Understanding Significance: The test process requires obtaining emotional information from two groups of children in a short time. If the meaning of the ment is unclear, it may affect participants' emotional reaction times. Therefore, the content of the segments must be intuitive and easy for children to understand. (3) Effective emotional arousal: The effectiveness of emotional arousal is an important factor in assessing the quality of stimulus materials. Following the criteria, 20 movie segments that met the criteria were initially selected. Ming recruited 10 volunteers and asked them to measure the emotional encouragement of each segment through a self-report questionnaire. The questionnaire was designed by Gross and Levenson. Evaluate the degree of excitement of six emotions, including pleasure, joy, interest, sadness, disgust, and fear, referring to Each category is rated on a 9-point Liker scale ranging from 0 (none) to 8 (very strong). The higher the score, the better the movie. This indicates that the intensity of the emotion excited by the segment is high. Positive emotions, such as playing with your mother or a pet, or experiencing a funny or awkward moment. Select six high-scoring segments to awaken and adjust your self-assessment scores from low to high. The video resolution is 720 x 576 pixels (i.e. (i.e. 25.0 frames per second).
[0062] The experimental flow was as follows: the subject was asked to sit in a chair, and a message was displayed on the computer screen. Once the subject has calmed down, start the video playback program. The camera on the display will then open and record the subject's actions. If the video playback is stopped, take precautions and guidance to ensure the validity of the data. The data recorded by the camera will then be saved. (2) Details of the experimental setup
[0063] The accuracy of this method is estimated by three-fold cross-validation. All samples were divided into three groups on average, with 27 samples in each group. The number of autistic subjects and typically developing subjects in the group was as close as possible, and Two groups were used to train the model, and the other group was used to test the model. The experiment was repeated three times, and the average value was calculated for the three experimental results. We compare the average accuracy under different experimental settings under the study. (3) Ablation research
[0064] In order to clarify the specific role of each module in this method, performed ablation studies to compare experimental results under different settings. In experiment (1), ,This invention adopts the coarse-grained analysis module introduced in this ,paper to extract 18-dimensional coarse-grained features. The resulting data is input to a binary classifier to obtain the corresponding classification results. Using the fine-grained analysis module introduced in the text, we extract six-dimensional fine-grained features and then In experiment (3), the present invention uses coarse-grained and fine-grained classification. The 24-dimensional features extracted by granular analysis are input into the classifier to obtain more comprehensive classification results. Finally, in experiment (4), the present invention was further developed based on experiment (3) and proposed in section 5.3.3. The proposed attention module is added, and the weighted fusion features are input to the binary classifier for final classification. To ensure the reliability of the experimental results, the present invention adopts three-fold cross-validation. The average accuracy of three folds was calculated. For details of the specific experimental results, see Table 1.
[0065] Table 1 TIFF2026027197000056.tif37156
[0066] In experiment (3), the average classification accuracy reached 85.07%, which is higher than experiment (1). Compared to experiment (2), the results are improved by 4.82% and 3.57%, respectively. This significant improvement is due to the fact that We verified the complementarity of the coarse-grained analysis and the fine-grained analysis mentioned in Section .1. Specifically, The paper mainly focuses on the facial expression intensity and social congruence characteristics of autistic patients, and the fine-grained analysis The focus is on identifying characteristics of facial expressions that are difficult to understand.
[0067] The average accuracy of experiment (4) reached 88.74%, which is 3.6 times higher than experiment (3). This result indicates that the attention module plays a significant role in feature fusion. As mentioned above, partial features may not have significant differences, but attentional skills may be significantly different. The application of the mechanism effectively removes the interference of these insignificant features on the classification results. At the same time, when different features are combined, they repel each other and become redundant. However, attention mechanisms are able to solve these problems well. can be done.
[0068] An embodiment of the present invention collects slices while a user is watching a video, and each a slice comprises a plurality of consecutively sampled images; The multiple images in each slice are statistically analyzed based on a pre-defined statistical algorithm. calculating coarse-grained features for each slice based on the measurement results; Based on the coarse-grained features of the slices, we extract the slices from several different fine-grained feature extraction models. Selecting one input to obtain fine-grained features for each slice; The coarse-grained features and the fine-grained features of the slice are subjected to feature fusion to obtain a fusion feature, and the fusion feature is applied to and performing autism identification based on the coarse-grained and fine-grained methods of claim 1. A granularity-coupled autism identification method is further provided.
[0069] The autism identification method that combines coarse-grained and fine-grained methods is the autism identification system described above. The operating principle and technical effect are the same as those of the first embodiment, and therefore the description will not be repeated here.
[0070] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. All amendments, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of the present invention. It is easily understood by a person skilled in the art that the invention should be included within the scope of protection of the above.
Claims
1. A combined coarse-grained and fine-grained autism identification system, comprising: It collects slices while the user is watching a video, and each slice is sampled consecutively. a data acquisition module containing a plurality of images acquired by the data acquisition module; The multiple images in each slice are statistically analyzed based on a pre-defined statistical algorithm. a coarse-grained feature extraction module that calculates coarse-grained features for each slice based on the measurement results; Based on the coarse-grained features of the slices, we extract the slices from several different fine-grained feature extraction models. A fine-grained feature extraction module selects one input and acquires fine-grained features for each slice. Rules and The coarse-grained features and the fine-grained features of the slice are subjected to feature fusion to obtain a fusion feature, and the fusion feature is and a feature fusion and identification module for performing autism identification based on the coarse-grained feature fusion and identification module. Autism identification system combining high- and low-granularity.
2. The coarse-grained feature extraction module: Face detection is performed on all images in the slice, and face coordinates, dimension information, and facial feature point positions are recorded. a face detection module for determining It invokes the head pose estimation algorithm and uses the face detection results to estimate the user's face in each image in each slice. a head pose estimation module for estimating a user's head pose; We calculate the user head pose variance for all images in each slice and calculate the user head pose variance. A head pose coarse-grained feature extraction module extracts coarse-grained head pose features for each slice. Ru and A face matching module performs face matching on all images in each slice using the face detection results. 、 The facial expression intensity estimation model is invoked to estimate the user's facial expression intensity in each image in each slice. ,The average facial expression intensity of all images in each slice is calculated. The facial expression intensity coarse-grained feature extraction model extracts the facial expression intensity coarse-grained features of each slice. Jules and Invokes a facial expression estimation model to estimate the user's facial expression category for each image in each slice. Then, we count the total number of times a specified facial expression appears in all images in each slice, and The total number of times the specified facial expression appears is binarized, and the facial expression category coarse-grained features of each slice are calculated. and a facial expression category coarse-grained feature extraction module for acquiring the facial expression category. An autism identification system that combines coarse-grained and fine-grained methods as described in.
3. The facial expression specified when stating this slice based on the video corresponding to the slice.
3. The self-closing method according to claim 2, wherein the coarse-grained and fine-grained classes are combined to determine the category. Disease identification system.
4. If the video corresponding to a slice is a video that stimulates the user's joyful emotions, The facial expression designated when counting the chair is determined to be a happy facial expression. The coarse-grained and fine-grained autism identification system of claim 3.
5. Based on the coarse-grained features of the slices, the slices are extracted using a number of different fine-grained feature extraction models. Choosing to enter one of them is The total number of times a specified facial expression appears in all images in a slice exceeds a preset threshold. inputting the slice into a first fine-grained feature extraction model; The total number of times a specified facial expression appears in all images in a slice exceeds a preset threshold. If not, inputting the slice into a second fine-grained feature extraction model.
2. The autism identification system according to claim 1, wherein the coarse-grained and fine-grained features are combined.
6. The first fine-grained feature extraction model and the second fine-grained feature extraction model both have long and short time recording capabilities. a feature extraction model based on a memory neural network, and the training sample set of the second fine-grained feature extraction model are different.
6. An autism identification system that combines coarse-grained and fine-grained features as described in claim 5.
7. The feature fusion of the coarse-grained features and the fine-grained features of the slices is reducing the fine-grained features of the slice; The fine-grained features of the slice after dimensional reduction and the coarse-grained features of the slice are fed into the SEnet network. The attention weights are calculated based on the input, and the fine-grained features of the slices are then extracted based on the attention weights. and fusing the coarse-grained features of the rice grains. An autism identification system that combines granular and fine-grained.
8. The step of identifying autism based on the fusion feature includes inputting the fusion feature to a binary classifier. and obtaining an autism identification result. A fine-grained combined autism identification system.
9. Collecting slices while a user is watching a video allows a user to watch multiple videos. Collect multiple slices of the video while watching it, and each video corresponds to a slice. Feature fusion of the coarse-grained features and fine-grained features of the slices is 2. The method of claim 1, further comprising: feature fusion of coarse-grained features and fine-grained features. An autism identification system that combines coarse-grained and fine-grained techniques.
10. A method for identifying autism that combines coarse-grained and fine-grained features, comprising: It collects slices while the user is watching a video, and each slice is sampled consecutively. and The multiple images in each slice are statistically analyzed based on a pre-defined statistical algorithm. calculating coarse-grained features for each slice based on the measurement results; Based on the coarse-grained features of the slices, we extract the slices from several different fine-grained feature extraction models. Selecting one input to obtain fine-grained features for each slice; The coarse-grained features and the fine-grained features of the slice are subjected to feature fusion to obtain a fusion feature, and the fusion feature is applied to and performing autism identification based on the coarse-grained and fine-grained autonomic nervous system. Autism identification methods.
Citation Information
Patent Citations
Bidirectional LSTM micro-expression depression recognition method based on feature pyramid network
CN110472564A
Intelligent identification system for mental diseases
CN111012367A
Automatic detection method and automatic detection device for autism spectrum disorder based on video expression behavior analysis
CN111128368A
Expression recognition system and method based on enhancement CNN and cross-layer LSTM
CN111523461A
Video sequence facial expression recognition algorithm based on bidirectional time convolution network
CN113723204A