Autism classification method and device based on multi-normal-form multi-feature fusion
Through a multi-paradigm and multi-feature fusion autism screening method, a multi-layer perceptron mixture model (MLP-Mixer) is used to extract and fuse multi-dimensional eye movement features, which solves the problem of low autism screening accuracy in existing technologies and achieves more efficient autism screening and early diagnosis.
Patent Information
- Application Number
- CN202510795234.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-15
- Publication Date
- 2025-10-17
AI Technical Summary
Existing autism screening methods are mostly limited to a single feature under a single paradigm, and fail to fully explore the abnormal eye movements of autistic children under different gaze paradigms, resulting in limited sensitivity and specificity of the screening methods, making it difficult to meet the multidimensional needs of complex clinical environments.
A multi-paradigm and multi-feature fusion method is adopted. By designing free face views, emotion competition paradigm and social competition paradigm, combined with a multi-layer perceptron mixture model (MLP-Mixer), multi-dimensional eye movement features are extracted and fused to achieve the coordinated optimization of sensitivity and specificity of autism screening.
It improves the accuracy and reliability of autism screening, reduces the risk of misdiagnosis, increases the recognition rate of autistic children, and provides more reliable early diagnosis support.
Smart Images

Figure CN120805025A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present application relates to the technical field of autism screening, in particular to an autism classification method and device based on multi-paradigm multi-feature fusion. BACKGROUND
[0002] Autism (Autism Spectrum Disorder, ASD) is a complex and severe neurodevelopmental disorder, and its core symptoms include social interaction disorders, language development retardation, and repetitive stereotyped behaviors.
[0003] A large number of statistical analyses have confirmed that ASD children have obvious differences in eye movement characteristics from TD children under different eye movement research paradigms. In the static face fixation task and the static visual preference task, the ASD children have different fixation patterns on social pictures from the TD children. For example, in the static face fixation task, the ASD children have significantly less fixation time and fixation times on human faces and their core areas (left eye, right eye, nose, mouth, etc.) than the TD children, and more fixation on other areas; in the static visual preference task, the ASD children have a higher fixation percentage on non-social pictures than the TD children, and a lower fixation percentage on social scene pictures than the TD children, and less fixation on areas rich in social information such as characters in social scene pictures and more fixation on areas poor in social information such as backgrounds.
[0004] Although existing research has begun to use children's eye movement characteristics for ASD diagnosis, most researches are limited to single feature under a single paradigm or multiple features under a single paradigm, and fail to fully explore the abnormal performance of ASD children's eye movement under different fixation paradigms, resulting in limited sensitivity and specificity of existing screening methods, which is insufficient to cover the multi-dimensional needs of complex clinical environments. SUMMARY
[0005] The embodiment of the present application provides an autism classification method and device based on multi-paradigm multi-feature fusion, which at least solves the problem of low autism screening classification accuracy in related technologies.
[0006] According to an embodiment of the present application, an autism classification method based on multi-paradigm multi-feature fusion is provided, comprising:
[0007] Obtaining task execution data of a target object, wherein the task execution data includes face data and eye data of the target object when performing a preset task;
[0008] Extracting data features of the task execution data based on a preset feature extraction scheme, wherein the data features include task common features and task individual features;
[0009] The data features are fused and classified by a preset first model to determine the autism type of the target object.
[0010] In an example embodiment, the facial data and eye data of the target object performing the preset task include:
[0011] Eye gaze data and facial change data of the target object in a continuous language input process are acquired, wherein the continuous language input process includes continuously playing a dynamic video to the target object, and the dynamic video contains several times of first duration of voice natural pauses.
[0012] The eye gaze data and facial change data are analyzed to obtain eye gaze trajectory features and facial transition data.
[0013] In an example embodiment, the facial data and eye data of the target object performing the preset task include:
[0014] Eye switching data and facial switching data of the target object in a multi-emotion parallel stimulation process are acquired, wherein the multi-emotion parallel stimulation process includes showing parallel emotion images representing different emotions arranged according to a preset spatial position to the target object.
[0015] The eye switching data are analyzed to obtain eye gaze trajectory switching paths.
[0016] In an example embodiment, the facial data and eye data of the target object performing the preset task include:
[0017] Facial attention data and eye attention data of the target object in a mixed visual stimulation process are acquired, wherein the mixed visual stimulation process includes presenting several groups of visual stimulation image pairs to the target object in a first manner, and the first manner includes arranging the several groups of visual stimulation image pairs on the left and right sides respectively, wherein the left side is a combination of non-social class images, and the right side is a combination of social class images, and the images are exchanged left and right when a single group is presented for a second duration.
[0018] The facial attention data and eye attention data are analyzed to obtain attention resource allocation features.
[0019] In an example embodiment, the task common features include at least any one of an access proportion of a region of interest, a region with the most access times, an average pupil diameter, an average saccade amplitude, an average saccade speed, saccade times, blink times, gaze times, total blink duration, total gaze duration, maximum gaze duration.
[0020] According to another embodiment of the present application, there is provided an autism early screening device based on multi-paradigm and multi-feature fusion, comprising:
[0021] a data collection module configured to acquire task execution data of the target object, wherein the task execution data comprises facial data and eye data of the target object when performing a preset task;
[0022] a feature extraction module configured to extract data features of the task execution data based on a preset feature extraction scheme, wherein the data features comprise task common features and task individual features;
[0023] a fusion classification module configured to perform fusion classification processing on the data features by using a preset first model to determine the autism type of the target object.
[0024] In an example embodiment, the facial data and eye data of the target object when performing the preset task comprise:
[0025] acquiring eye fixation data and facial change data of the target object in a continuous language input process, wherein the continuous language input process comprises continuously playing a dynamic video to the target object, and the dynamic video comprises a plurality of first time length of voice natural pauses;
[0026] analyzing the eye fixation data and the facial change data to obtain eye fixation trajectory features and facial transition data.
[0027] In an example embodiment, the facial data and eye data of the target object when performing the preset task comprise:
[0028] acquiring eye switching data and facial switching data of the target object in a multi-emotion parallel stimulation process, wherein the multi-emotion parallel stimulation process comprises showing parallel emotion images representing different emotions arranged according to a preset spatial position to the target object;
[0029] analyzing the eye switching data to obtain an eye fixation trajectory switching path.
[0030] In an example embodiment, the facial data and eye data of the target object when performing the preset task comprise:
[0031] acquiring facial attention data and eye attention data of the target object in a mixed visual stimulation process, wherein the mixed visual stimulation process comprises presenting a plurality of sets of visual stimulation image pairs to the target object in a first manner, and the first manner comprises arranging the plurality of sets of visual stimulation image pairs on the left and right sides respectively, wherein the left side is a combination of non-social class images and the right side is a combination of social class images, and in the case of single set presentation reaching a second time length, the images are exchanged left and right;
[0032] The face attention data and the eye attention data are analyzed to obtain an attention resource allocation feature.
[0033] According to still another embodiment of the present application, a computer readable storage medium is also provided, in which a computer program is stored, wherein the computer program is configured to perform the steps of any of the above method embodiments when executed.
[0034] According to still another embodiment of the present application, an electronic device is also provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps of any of the above method embodiments.
[0035] According to the present application, since a plurality of task schemes are designed to screen and identify autism from multiple dimensions, the problem of low classification accuracy of autism screening can be solved, and the classification accuracy of autism screening is improved. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 is a flowchart of a multi-paradigm multi-feature fusion-based autism classification method according to an embodiment of the present application;
[0037] Figure 2 is a flowchart of a face free-view paradigm according to an embodiment of the present application;
[0038] Figure 3 is a flowchart of an emotional competition paradigm according to an embodiment of the present application;
[0039] Figure 4 is a flowchart of a social competition paradigm according to an embodiment of the present application;
[0040] Figure 5 is a structural block diagram of a multi-paradigm multi-feature fusion-based autism classification device according to an embodiment of the present application. DETAILED DESCRIPTION
[0041] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments.
[0042] Hereinafter, the terms "first", "second", and the like are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second", and the like can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0043] In addition, in the present application, the orientation terms such as "upper", "lower", "left", "right" and the like can include, but are not limited to, the orientation defined by the relative placement of the components in the drawings. It should be understood that these directional terms are relative concepts, and they are used for relative description and clarification, which can be changed accordingly according to the change of the placement of the components in the drawings.
[0044] In the present application, unless otherwise explicitly specified and limited, the term "connection" should be understood broadly, for example, "connection" can be fixed connection, or detachable connection, or integral; can be directly connected, or indirectly connected through intermediate medium. In addition, the term "coupling" can be an electrically connected manner for realizing signal transmission.
[0045] As used herein, "about", "approximately" or "nearly" includes the stated value and the average value within an acceptable deviation range of the specific value, wherein the acceptable deviation range is determined by the person of ordinary skill in the art considering the measurement being discussed and the error related to the measurement of the specific quantity (i.e., the limitation of the measurement system).
[0046] Autism Spectrum Disorder (ASD) is a complex and severe neurodevelopmental disorder, and its core symptoms include social interaction disorders, language development retardation, repetitive stereotyped behaviors, etc.
[0047] A large number of research results show that early identification, diagnostic evaluation and scientific intervention are very critical for autism. Early detection and intervention of autism can significantly improve the prognosis of patients. In medicine, based on the clinical manifestations and behavioral characteristics of autism patients, the diagnostic criteria for childhood autism are formulated at home and abroad, such as Clancy Behavior Scale (CBS), Modified Checklist for Autism in Toddlers (M-CHAT), Autism Behavior Checklist (ABC), Childhood Autism Rating Scale (CARS) and Autism Diagnostic Interview-Revised (ADI-R) and the like.
[0048] Traditional diagnostic methods require the collaborative participation of a multidisciplinary team (including psychiatrists, pediatricians, rehabilitation therapists, etc.) and parents to jointly complete a systematic assessment of autism in children. Because the assessment process involves multiple interviews, behavioral observations, and cross-scenario data collection, the diagnostic cycle is long and the operation is complex. In addition, the behavioral manifestations of autism vary greatly among different individuals, making it difficult to match universal judgment criteria, which may affect the accuracy of the diagnosis.
[0049] With advances in deep learning and computer vision technology, eye movement feature analysis based on a multi-task paradigm is gradually being used to assist in the early screening of autism spectrum disorder (ASD). For example, by designing standardized tasks (such as watching social scene videos and tracking moving objects), key feature data such as children's eye gaze area, eye movement amplitude, and pupil diameter changes are accurately recorded. Then, using machine learning algorithms, this data is compared with the feature data of typically developing children (TD) for learning and analysis, and a recognition model is established to develop an automated assessment tool that can be used in clinical practice.
[0050] A large number of statistical analyses have confirmed that the eye movement characteristics of ASD children are significantly different from those of TD children under different eye movement research paradigms. In the static face gaze task and the static visual preference task, the gaze patterns of ASD children on social pictures are different from those of TD children. For example, in the static face gaze task, ASD children gazed at the face and its core areas (left eye, right eye, nose, mouth, etc.) for a significantly shorter duration and number of times than TD children, but gazed at other areas more; in the static visual preference task, ASD children gazed at non-social pictures more than TD children, but gazed at social scene pictures less than TD children. In addition, in social scene pictures, they gazed less at areas rich in social information such as characters, and more at areas poor in social information such as backgrounds.
[0051] Although existing studies have begun to use children's eye movement characteristics to diagnose ASD, most studies are limited to a single feature under a single paradigm or multiple features under a single paradigm, and fail to fully explore the abnormal eye movement manifestations of ASD children under different gaze paradigms. As a result, the sensitivity and specificity of existing screening methods are limited, which is insufficient to cover the multidimensional needs of complex clinical environments.
[0052] The application is based on a multi-task paradigm and multi-feature fusion design, aiming to enhance the multi-paradigm multi-feature utilization capability of the existing ASD screening method. By designing multiple task paradigms to extract multi-dimensional eye movement features, using a multi-layer perceptron mixer model (Multi-Layer Perceptron-Mixer) for multi-source feature fusion and classification, the sensitivity and specificity of ASD screening are optimized, and the recognition rate of the model for ASD children is improved. This screening method effectively integrates the multi-dimensional eye movement features of children in different task scenarios, reduces the risk of misdiagnosis caused by single scene data bias, and provides more reliable technical support for early diagnosis and intervention.
[0053] In this embodiment, a autism classification method based on multi-paradigm multi-feature fusion is provided, Figure 1 is a flowchart of a autism classification method based on multi-paradigm multi-feature fusion according to an embodiment of the application, as Figure 1 shown, the flow includes the following steps:
[0054] Step S11, acquiring task execution data of a target object, wherein the task execution data includes facial data and eye data of the target object when performing a preset task;
[0055] In this embodiment, by fully mining the abnormal performance of ASD children's eye movement under different gaze tasks, the autism is screened and classified from multiple dimensions, and the screening and classification accuracy is improved.
[0056] Wherein, the target object includes children, adults and the like who need to be screened and classified for autism, and the facial data and eye data of the target object when performing a preset task include:
[0057] 1, free view paradigm of human face: acquiring eye gaze data and facial change data of the target object in a continuous language input process, wherein the continuous language input process includes continuously playing a dynamic video to the target object, and the dynamic video contains several times of first duration of voice natural pause; then analyzing the eye gaze data and facial change data to obtain eye gaze trajectory features and facial migration data.
[0058] As shown in detail, Figure 2 This paradigm mainly investigates the allocation strategy and dynamic adaptation ability of children's visual attention in real conversation situations; the flow is as Figure 2As shown, a 29-second dynamic video is continuously played in the center of the screen, presenting a scene in which a little girl looks directly at the camera and tells a story naturally. Four 0.5-second (i.e., the first time length) voice natural pauses are designed in the video to simulate the rhythm changes of real conversations. At the same time, the eye tracker collects the fixation trajectory features of the child during the continuous language input process, focusing on analyzing the key area fixation stability (i.e., the aforementioned eye fixation trajectory features) and the migration features of specific facial areas (i.e., the aforementioned facial migration data).
[0059] and / or,
[0060] 2. Emotional competition paradigm: obtaining eye switching data and facial switching data of a target object during a multi-emotion parallel stimulation process, wherein the multi-emotion parallel stimulation process includes showing the target object parallel emotion images representing different emotions arranged according to a preset spatial position; analyzing the eye switching data to obtain an eye fixation trajectory switching path.
[0061] This paradigm mainly investigates the visual selection preference and emotional information processing characteristics of children under multi-emotion parallel stimulation. The experimental procedure is as follows Figure 3 As shown, five groups of emotion images (i.e., the aforementioned parallel emotion images) are presented in the center of the screen each time, each group containing four equal-area regions (i.e., the aforementioned preset spatial positions) of upper left, upper right, lower left, and lower right, respectively displaying different facial emotions (joy, anger, sadness, and fear), and each group of images is presented for 5 seconds. The eye tracker records the eye fixation trajectory switching path (i.e., the aforementioned eye fixation trajectory switching path) of the child among the four quadrants. By analyzing the distribution of fixation hotspots in the competitive emotional stimulation, the preferential processing mode and visual selection strategy of children for conflicting emotional cues are revealed.
[0062] and / or,
[0063] 3. Social competition paradigm: obtaining facial attention data and eye attention data of a target object during a mixed visual stimulation process, wherein the mixed visual stimulation process includes presenting a plurality of pairs of visual stimulation images to the target object in a first manner, the first manner including arranging the plurality of pairs of visual stimulation images on the left and right sides, respectively, wherein the left side is a combination of non-social images and the right side is a combination of social images, and the images are exchanged left and right when a single group is presented for a second time length.
[0064] The facial attention data and eye attention data are analyzed to obtain attention resource allocation features.
[0065] This paradigm mainly investigates the attention resource allocation characteristics of children in a mixed visual stimulation environment. The experimental procedure is as follows Figure 4As shown, 8 sets of visual stimulus pairs are presented on the left and right sides of the screen simultaneously, the first four sets are left non-social images (abstract pattern combination) and right social images (human interaction scene), the last four sets of two types of images are exchanged in position, and the single set is presented for 5 seconds. By analyzing the distribution weight of the fixation point, the visual processing priority and selection characteristics of children to social and non-social information are revealed.
[0066] In step S12, data features of the task execution data are extracted based on a preset feature extraction scheme, and the data features include task common features and task individual features.
[0067] In this embodiment, on the basis of standardized data collection, based on the basic parameters of gaze coordinates, gaze times and pupil changes recorded by the eye tracker, feature extraction schemes are constructed for different data types and characteristics.
[0068] The task individual features include corresponding task data features in the aforementioned three task paradigms, such as eye gaze trajectory switching path, etc.; and the task common features include at least any one of the following: proportion of visited areas of interest, area with the most visits, average pupil diameter, average saccade amplitude, average saccade speed, saccade times, blink times, gaze times, total blink duration, total gaze duration, maximum gaze duration, etc.
[0069] The proportion of the visited areas of interest: according to the layout features of the visual elements of the aforementioned three types of experimental task paradigms, the screen to be displayed to the children is divided into areas of interest (AOI); then, according to the child's gaze point position collected in each frame, the spatial position relationship between the child's gaze point and the area of interest is matched frame by frame, and the cumulative number of gaze point landing points in different areas of interest is counted, and then the proportion of gaze points in each paradigm in each area of interest is calculated.
[0070] The calculation formula is shown in the following formula 1:
[0071]
[0072] In the formula, P ij represents the proportion of the number of gaze points in the i-th area of interest under the j-th paradigm to the total number of gaze points in all areas of interest under the paradigm; i and j are the index numbers of the AOI region and the experimental paradigm, respectively, T j is the total number of gaze points recorded in the j-th paradigm, x k represents the screen coordinates of the k-th gaze point, I(x k ∈AOI i ) is 1 when the coordinates x k are located in the AOI i , otherwise 0.
[0073] Most visited area: Specifically, the area of interest (AOI) number is divided, and the number of visits of children in each area of interest under each frame is counted to calculate the most visited area of interest. The formula is as follows:
[0074]
[0075] In the formula, M j represents the AOI number with the most visits in the paradigm. i represents the AOI number, T j is the total number of gaze points recorded in the jth paradigm, x k represents the screen coordinates of the kth gaze point, I(x k ∈AOU i ) is the indicator function, which is 1 when the current gaze point x k is located in the AOI i and the previous gaze point x k-1 is not in the AOI i-1 , otherwise it is 0.
[0076] Average pupil diameter: By tracking the pupil diameter data of children frame by frame, after excluding abnormal values caused by blinking or head movement, the average pupil diameter under each paradigm is calculated, and the formula is as follows:
[0077]
[0078] In the formula, D j represents the average pupil diameter of the jth paradigm, d k is the pupil diameter value of the kth valid measurement, T j is the total number of measurements under the paradigm, I(d k ∈V) is the indicator function, which is 1 when the measurement value d k is valid data (i.e. excluding abnormal values caused by blinking or head movement), otherwise it is 0, and N j is the number of valid measurements.
[0079] Average saccade amplitude: Detect the amplitude of each frame saccade when the gaze is in the AOI, and calculate the average amplitude of saccade under each paradigm after excluding invalid data, and the calculation formula is as follows:
[0080]
[0081] In the formula, A j represents the average saccade amplitude of the jth paradigm, s k is the saccade amplitude of the kth valid measurement, T j is the total number of saccades under the paradigm, I(s k ∈V) is the indicator function, which is 1 when the measurement value s k is valid data (i.e. excluding abnormal values caused by blinking or head movement), otherwise it is 0, and Nj is the number of valid measurements.
[0082] Average saccade velocity: The average saccade velocity of the child's gaze within the AOI throughout the paradigm is calculated as follows:
[0083]
[0084] where V j is the average saccade velocity of the jth paradigm, v k is the saccade velocity of the kth valid measurement, T j is the total number of saccades in the paradigm, i(v k ∈ V) is the indicator function, which takes the value 1 if the measured value v k is valid (i.e. excludes outliers due to blinks or head movements), and 0 otherwise, N j is the number of valid measurements.
[0085] Saccade count: The total number of saccades of the child's gaze within the AOI throughout the paradigm is calculated as follows:
[0086]
[0087] where C j is the number of valid saccades in the jth paradigm, T j is the total number of saccades in the paradigm, s k is the data record of the kth saccade, I(s k ∈ V) is the indicator function, which takes the value 1 if the measured value s k is valid (i.e. amplitude > 1° and velocity in the range 50-500° / s), and 0 otherwise.
[0088] Blink count: The total number of blinks of the child's gaze within the AOI throughout the paradigm is calculated as follows:
[0089]
[0090] where b j is the number of valid blinks in the jth paradigm, T j is the total number of saccades in the paradigm, b k is the data record of the kth blink, I(b k ∈ V) is the indicator function, which takes the value 1 if the measured value b k is valid (i.e. excludes invalid blinks due to head movements or device noise), and 0 otherwise.
[0091] Gaze count: The total number of gaze points of the child's gaze within the AOI throughout the paradigm is calculated as follows:
[0092]
[0093] where F j denotes the number of fixation points in AOI in the jth paradigm. T j is the total number of fixation points in the paradigm, x k is the coordinate of the kth fixation point, I(x k ∈ V) is an indicator function that takes the value 1 when x k is inside the AOI and 0 otherwise.
[0094] Total blink duration: the total blink duration is calculated as the sum of the blink durations of all fixation points before a blink in the jth paradigm. The formula is as follows:
[0095]
[0096] where B j denotes the total blink duration in AOI in the jth paradigm. T j is the number of blinks in the paradigm, t k is the duration of the kth blink, x k is the coordinate of the fixation point before the kth blink, I(x k ∈ V) is an indicator function that takes the value 1 when x k is inside the AOI and 0 otherwise.
[0097] Total fixation duration: the total fixation duration is calculated as the sum of the fixation durations of all fixation points in the jth paradigm. The formula is as follows:
[0098]
[0099] where E j denotes the total fixation duration in AOI in the jth paradigm. T j is the total number of fixation points in the paradigm, e k is the duration of the kth fixation point, x k is the coordinate of the kth fixation point, I(x k ∈ V) is an indicator function that takes the value 1 when x k is inside the AOI and 0 otherwise.
[0100] Maximum fixation duration: the maximum fixation duration is calculated as the maximum value of the durations of all fixation points in the jth paradigm. The formula is as follows:
[0101]
[0102] where G j denotes the maximum fixation duration in AOI in the jth paradigm. T j is the total number of fixation points in the paradigm, g k is the duration of the kth fixation point, x kis the coordinate of the kth gaze point, I(x k ∈V) is the indicator function, when x k The value is 1 if the coordinate is within the AOI, otherwise it is 0.
[0103] This patent designs separate features for the face free view paradigm, emotion competition paradigm, and social competition paradigm based on each paradigm's region of interest. For the face free view paradigm, as the first task paradigm, the first fixation duration is designed as a separate feature; for the emotion competition paradigm, the total dwell time is selected as a separate feature; and for the social competition paradigm, the number of re-looks is used as a separate feature:
[0104] First fixation duration: Calculate the duration of the first fixation point in each AOI under the face free view paradigm.
[0105] Total dwell time: Calculate the cumulative dwell time of all gazes within each AOI under the emotional competition paradigm. The calculation formula is as follows:
[0106]
[0107] Where S i Indicates the total stay time in the i-th AOI. j is the total number of fixations in the AOI, t k is the gaze time of the kth gaze point, x k is the coordinate of the kth gaze point, I(x k ∈V) is the indicator function, when x k The coordinate is 1 when it is within the AOI, and 0 otherwise, quantifying the intensity of children's attention to specific emotional stimuli.
[0108] Number of re-looks: Count the number of times children looked back in each AOI under the social competition paradigm. The calculation formula is as follows:
[0109]
[0110] Where R i Indicates the number of times the AOI is reviewed in this paradigm. i is the total number of fixations recorded in this paradigm, x k represents the screen coordinates of the kth gaze point, Represents the current gaze point x k Located in AOI i Inner, previous fixation point x k-1 Not in AOI i-1 And at the fixation point x k There was a fixation point x before m (m<k) located in AOI i It takes 1 when it is inside, and 0 otherwise.
[0111] Step S13, the data features are fused and classified by a preset first model to determine the autism type of the target object.
[0112] In this embodiment, in the feature fusion stage, the 11 common features extracted in the above three paradigm tasks and the respective paradigm individual features are integrated. Among them, the face free view paradigm, the emotion competition paradigm and the social competition paradigm are abbreviated as F face , F affect and F social respectively, and the fused feature Fus can be expressed as:
[0113]
[0114] In the classification stage, for the complex correlation between multi-paradigm and multi-feature, the application adopts a multi-layer perception hybrid model (MLP-Mixer, i.e., the aforementioned first model) for feature fusion and classification. This model can simultaneously analyze the interaction mode between global features and the detailed correlation of local features through a unique hybrid learning mechanism. Specifically, the MLP-Mixer processes data features in steps, not only retains the characteristics of different features, but also dynamically mines the potential relationship between them, and finally realizes the comprehensive modeling of multi-dimensional features. Compared with traditional neural network models, this model has better adaptability and higher training efficiency.
[0115] In summary, the application has the following beneficial effects
[0116] 1. Based on existing research, the application sets up three task paradigms: face free view paradigm, emotion competition paradigm and social competition paradigm, to fully tap the special response of ASD children in different scenarios. The core of the face free view paradigm is to investigate the spontaneous attention allocation pattern of ASD children to the core area of the face in a dynamic social scene, and to reveal the integration abnormalities of their attention system by analyzing the bias of their gaze distribution. The emotion competition paradigm presents four facial images with different emotional effects (joy, anger, sadness, and fear) side by side in the visual interface, and simultaneously captures the visual preference differentiation of ASD children under the stimulation of multiple emotions. The social competition paradigm presents a set of contrast stimuli on the left and right sides of the screen simultaneously, and cycles through 8 different content combinations, capturing the visual preference differentiation of children under the competition of social and non-social information in real time.
[0117] 2. The application uses a feature extraction module to preliminarily extract features, a feature fusion and classification module for multi-dimensional feature fusion and classification, to realize the collaborative optimization of ASD screening sensitivity and specificity, and to improve the recognition rate of the model for ASD children. First, from the eye movement data, eye movement features are preliminarily calculated and extracted. These features are obtained from three task paradigms of free view of human face, emotional competition and social competition, to ensure the authenticity and comprehensiveness of the data. Then, the extracted eye movement features are integrated and input into a multi-layer perception mixer hybrid model (MLP-Mixer) for multi-source feature fusion and classification. The MLP-Mixer effectively distinguishes the gaze preference of ASD children and TD children by modeling the distribution correlation and difference of eye movement features through its fully connected layer driven hybrid architecture. This invention realizes rapid screening of ASD children based on multi-task paradigm and multi-feature fusion, uses the lightweight architecture of MLP-Mixer, and provides an automated tool with screening efficiency and accuracy for primary medical scenarios.
[0118] 3. The application improves the accuracy and reliability of early screening of autism through multi-task paradigm design and multi-dimensional feature fusion technology. In view of the complexity of ASD children's behavior in different scenarios, a multi-dimensional task paradigm is designed to systematically capture diversified features such as gaze distribution, visual preference differentiation and social information processing. By fusing independent features of each paradigm and cross-paradigm general features, a multi-dimensional feature representation system is constructed, effectively covering the diversity of behavior differences in clinical environment, reducing the screening bias caused by single paradigm or single feature, and significantly improving the recognition ability of the model for abnormal behavior of ASD children.
[0119] 4. Based on the collaborative modeling of multi-source heterogeneous features, the application innovatively introduces a multi-feature extraction framework for multi-paradigm interaction, which fully explores the relevance and difference of behavior features in different paradigms through a dynamic interaction mechanism of independent features of each paradigm and cross-paradigm general features, retaining the independent characteristics of each task scenario and establishing common associations across paradigms. Compared with traditional single feature analysis methods, this method solves the problems of incomplete feature coverage in single paradigm and insufficient cross-scene feature adaptability through multi-paradigm multi-feature fusion, achieving collaborative optimization in screening efficiency and accuracy, and providing a screening tool with fast response and high reliability for primary medical scenarios.
[0120] 5. The application combined with multi-paradigm multi-feature fusion strategy shows good screening performance and stability.
[0121] The specific data is as follows:
[0122] The present application cites 20 ASD children and 20 TD children diagnosed by a certain people's hospital, aged 2-12 years old, and divides the training set (12 ASD and 12 TD) and the test set (8 ASD and 8 TD) according to the ratio of 6:4, and inputs them into the multi-layer perception hybrid model (MLP-Mixer) for feature fusion and classification.
[0123] The experimental results show that the method realizes balanced performance of classification accuracy 81.25%, sensitivity 87.5% and specificity 75% on the test set. The combination of independent feature extraction of each paradigm and cross-paradigm general feature modeling improves the systematic analysis ability of abnormal behaviors of ASD children. Compared with the prior art, the multi-dimensional fusion design of the present application effectively enhances the adaptability of the model to complex behavior patterns, and provides an intelligent solution with universality and practical value for early diagnosis of autism.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the method described in each embodiment of the present application.
[0125] In the present embodiment, an autism early screening device based on multi-paradigm multi-feature fusion is also provided, which is used to realize the above embodiments and preferred embodiments, which have been described and will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware, or a combination of software and hardware is also possible and is contemplated.
[0126] Figure 5 is a structural block diagram of an autism early screening device based on multi-paradigm multi-feature fusion according to an embodiment of the present application, as shown in Figure 5 , the device comprises:
[0127] The data acquisition module 51 is used to acquire the task execution data of the target object, wherein the task execution data includes the facial data and the eye data of the target object when performing the preset task;
[0128] The feature extraction module 52 is configured to extract data features of the task execution data based on a preset feature extraction scheme, the data features including task common features and task individual features.
[0129] The fusion classification module 53 is configured to perform fusion classification processing on the data features by using a preset first model to determine the autism type of the target object.
[0130] In an optional embodiment, the facial data and the eye data of the target object when performing the preset task include:
[0131] The eye gaze data and the facial change data of the target object in a continuous language input process are acquired, wherein the continuous language input process includes continuously playing a dynamic video to the target object, and the dynamic video contains a plurality of first-time-length voice natural pauses.
[0132] The eye gaze data and the facial change data are analyzed to obtain eye gaze trajectory features and facial transition data.
[0133] And / or,
[0134] The eye switching data and the facial switching data of the target object in a multi-emotion parallel stimulation process are acquired, wherein the multi-emotion parallel stimulation process includes showing parallel emotion images representing different emotions arranged according to a preset spatial position to the target object.
[0135] The eye switching data are analyzed to obtain an eye gaze trajectory switching path.
[0136] And / or,
[0137] The facial attention data and the eye attention data of the target object in a mixed visual stimulation process are acquired, wherein the mixed visual stimulation process includes presenting a plurality of sets of visual stimulation image pairs to the target object in a first manner, and the first manner includes arranging the plurality of sets of visual stimulation image pairs on the left and right sides respectively, wherein the left side is a combination of non-social class images, and the right side is a combination of social class images, and in the case of single set presentation reaching a second time length, the images are exchanged left and right.
[0138] The facial attention data and the eye attention data are analyzed to obtain attention resource allocation features.
[0139] It should be noted that each of the above modules can be implemented by software or hardware, and for the latter, the following implementation manners can be used, but are not limited thereto: all the above modules are located in the same processor; or the above modules are located in different processors in any combination.
[0140] The embodiment of the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is arranged to execute the steps in any of the method embodiments when running.
[0141] In an example embodiment, the computer readable storage medium can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media capable of storing a computer program.
[0142] The embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is arranged to execute the computer program to perform the steps in any of the method embodiments.
[0143] In an example embodiment, the electronic device can further comprise a transmission device and an input and output device, wherein the transmission device is connected to the processor, and the input and output device is connected to the processor.
[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0145] In the several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented by other ways. For example, the device embodiments described above are only illustrative, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0146] The units described as separate components can or can not be physically separate, and the components shown as units can be one physical unit or multiple physical units, that is, can be located in one place, or can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0147] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0148] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The software product is stored in a storage medium, including a plurality of instructions to make a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0149] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for autism classification based on multi-paradigm and multi-feature fusion, characterized by: include: Acquiring task execution data of the target object, wherein the task execution data includes facial data and eye data of the target object when performing a preset task; Extracting data features of the task execution data based on a preset feature extraction scheme, wherein the data features include common task features and individual task features; The data features are fused and classified using a preset first model to determine the autism type of the target object.
2. The autism classification method based on multi-paradigm and multi-feature fusion according to claim 1 is characterized in that: The facial data and eye data of the target object when performing the preset task include: Obtaining eye gaze data and facial change data of a target subject during a continuous speech input process, wherein the continuous speech input process includes continuously playing a dynamic video to the target subject, the dynamic video including a plurality of natural speech pauses of a first duration; The eye gaze data and the facial change data are analyzed to obtain eye gaze trajectory features and facial migration data.
3. The autism classification method based on multi-paradigm and multi-feature fusion according to claim 1, characterized in that: The facial data and eye data of the target object when performing the preset task include: Acquiring eye switching data and facial switching data of a target subject during a multi-emotion parallel stimulation process, wherein the multi-emotion parallel stimulation process includes presenting to the target subject parallel emotion images representing different emotions arranged in preset spatial positions; The eye switching data is analyzed to obtain an eye gaze trajectory switching path.
4. The autism classification method based on multi-paradigm and multi-feature fusion according to claim 1, characterized in that: The facial data and eye data of the target object when performing the preset task include: Obtaining facial attention data and eye attention data of a target subject during a mixed visual stimulation process, wherein the mixed visual stimulation process includes presenting a plurality of visual stimulation image pairs to the target subject in a first manner, wherein the first manner includes arranging the plurality of visual stimulation image pairs on left and right sides, with the left side comprising a non-social image combination and the right side comprising a social image combination, and swapping the left and right images when a single pair is presented for a second duration; The facial attention data and the eye attention data are parsed to obtain attention resource allocation features.
5. The autism classification method based on multi-paradigm and multi-feature fusion according to claim 1, characterized in that: The common characteristics of the tasks include at least one of the proportion of visited areas of interest, the area with the most visits, the average pupil diameter, the average saccade amplitude, the average saccade speed, the number of saccades, the number of blinks, the number of fixations, the total blink duration, the total fixation duration, and the maximum fixation duration.
6. An early screening device for autism based on multi-paradigm and multi-feature fusion, characterized in that: include: A data acquisition module is used to obtain task execution data of the target object, wherein the task execution data includes facial data and eye data of the target object when performing a preset task; A feature extraction module is used to extract data features of the task execution data based on a preset feature extraction scheme, wherein the data features include common features of tasks and individual features of tasks; The fusion classification module is used to perform fusion classification processing on the data features through a preset first model to determine the autism type of the target object.
7. The autism early screening device based on multi-paradigm and multi-feature fusion according to claim 6, characterized in that: The facial data and eye data of the target object when performing the preset task include: Obtaining eye gaze data and facial change data of a target subject during a continuous speech input process, wherein the continuous speech input process includes continuously playing a dynamic video to the target subject, the dynamic video including a plurality of natural speech pauses of a first duration; The eye gaze data and the facial change data are analyzed to obtain eye gaze trajectory features and facial migration data.
8. The autism early screening device based on multi-paradigm and multi-feature fusion according to claim 6, characterized in that: The facial data and eye data of the target object when performing the preset task include: Acquiring eye switching data and facial switching data of a target subject during a multi-emotion parallel stimulation process, wherein the multi-emotion parallel stimulation process includes presenting to the target subject parallel emotion images representing different emotions arranged in preset spatial positions; The eye switching data is analyzed to obtain an eye gaze trajectory switching path.
9. The autism early screening device based on multi-paradigm and multi-feature fusion according to claim 6, characterized in that: The facial data and eye data of the target object when performing the preset task include: Obtaining facial attention data and eye attention data of a target subject during a mixed visual stimulation process, wherein the mixed visual stimulation process includes presenting a plurality of visual stimulation image pairs to the target subject in a first manner, wherein the first manner includes arranging the plurality of visual stimulation image pairs on left and right sides, with the left side comprising a non-social image combination and the right side comprising a social image combination, and swapping the left and right images when a single pair is presented for a second duration; The facial attention data and the eye attention data are parsed to obtain attention resource allocation features.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when executed.