An autism spectrum grading method based on eye tracking, a terminal and a medium
Patent Information
- Application Number
- CN202610861865.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-04
AI Technical Summary
[0006]尽管眼动特征已被证实可作为ASD早期识别的客观生物标志物,但现有研究多将其用于患病与否的二分类筛查,极少将眼动特征与ASD等级量化关联,未能充分发挥眼动数据在等级细分评估中的客观支撑作用,导致等级评估仍缺乏可量化、可复现的生物学依据
1.提出了一种评估自闭症谱系障碍严重度的新方法,该方法整合了多模态眼动追踪生物标志物:注视热力图、注视轨迹图以及注视熵指标。该方法解决了传统二元ASD分类的关键局限性,能够实现精确的定量严重程度分级,并有助于早期临床筛查、精细化的患者分层以及个性化干预方案的制定。
Smart Images

Figure CN122692684A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autism spectrum disorder assessment technology, and in particular to an autism spectrum classification method, terminal and medium based on eye tracking. Background Technology
[0002] Autism is a congenital disorder of emotional contact; children with autism are born unable to form emotional connections like others. Currently, autism spectrum disorder is classified as a neurodevelopmental disorder, and it is one of the most congenital, lifelong, and easily inherited mental illnesses. Social communication impairments, repetitive and stereotyped behaviors, and restricted interests are the main characteristics of autism spectrum disorder. These not only directly affect the cognitive development, independent living skills, and quality of social integration of affected children, but also place a heavy burden of care and financial strain on families. Furthermore, it poses a continuous challenge to public service systems such as basic education, rehabilitation medicine, and social welfare. Currently, there is no cure; treatment mainly focuses on improving symptoms and enhancing function. The prognosis is influenced by a variety of factors. Therefore, early, continuous, and scientific intervention is crucial for the prognosis of autism. The earlier the intervention, the greater the potential for progress, the easier it is to integrate into social life, and the better the quality of life.
[0003] The Autism and Developmental Disorders Surveillance Network (ADDM Network), a leading global epidemiological surveillance system for autism, has released its latest data showing that, as of 2020, the prevalence of autism spectrum disorder (ASD) among 8-year-old children in the United States had climbed to 1 in 36. The prevalence is approximately 4% in boys and 1% in girls, with a significant and persistent gender difference. This figure represents a clear upward trend compared to previous monitoring results: the prevalence was 1 in 44 in 2018, 1 in 54 in 2016, and only 1 in 150 when the ADDM Network first published its data in 2000. It is noteworthy that this increase in prevalence stems not only from improved diagnostic techniques, increased public awareness, and optimized diagnostic criteria leading to higher case identification rates, but also reflects how ASD has evolved from a rare neurodevelopmental disorder into a prevalent developmental disorder among children and adolescents.
[0004] The diagnosis of autism is essentially a comprehensive clinical judgment. The core value of auxiliary assessment models is not to replace physician diagnosis, but to provide objective evidence for clinical decision-making through technological means, driving the assessment system from experience-driven to data-driven innovation. On the one hand, models can lower the assessment threshold; primary healthcare institutions only need to be equipped with simple eye-tracking devices to obtain standardized assessment reports through models, solving the assessment difficulties caused by the uneven distribution of high-quality medical resources. On the other hand, models can integrate multi-source data to construct a multimodal assessment framework, further improving the comprehensiveness and accuracy of assessments. This research aims to achieve more accurate early identification, a more efficient assessment process, and more personalized intervention programs, providing key technological support for improving the social functioning and social integration of children with autism.
[0005] Eye movement behavior, an objective biomarker detectable in the early stages of ASD identification, has shown great potential in the auxiliary diagnosis and severity assessment of autism spectrum disorders, as confirmed by numerous related studies. Because eye movement data acquisition is non-invasive, simple to operate, and does not require participants to possess complex language skills, increasing research indicates that eye-tracking analysis is a highly effective ASD screening method suitable for young children. For example, many studies have confirmed a close correlation between abnormal fixation preferences and ASD: children with ASD show significantly less time spent fixing their gaze on the facial eye area than typically developing children, while showing significantly higher attention to non-core social areas such as the mouth or objects.
[0006] Although eye movement features have been proven to serve as objective biomarkers for early identification of ASD, current research primarily uses them for binary screening to determine whether a child has the condition, rarely quantifying the correlation between eye movement features and ASD severity. This fails to fully leverage the objective supporting role of eye movement data in severity assessment, resulting in a lack of quantifiable and reproducible biological evidence for severity assessment. Furthermore, feature mining is often limited to a single dimension, focusing primarily on basic eye movement indicators such as fixation duration and saccade speed, without delving into the specific eye movement patterns of patients with different ASD severities, such as the degree of social cue avoidance and differences in joint attention-related eye movement trajectories. Moreover, a single eye movement modality cannot comprehensively cover the complex cognitive and behavioral characteristics related to ASD severity, easily leading to assessment bias. In addition, traditional multimodal fusion methods, such as fusing neuroimaging with behavioral data or fusing facial expressions with head movements and other facial features, have high requirements for equipment and environment and are prone to causing stress responses in children with autism. Summary of the Invention
[0007] Based on the limitations of existing technologies, this application proposes an eye-tracking-based autism spectrum classification method, terminal, and medium. It is based on the MAAT-Net multimodal fusion model of residual adaptive Transformer, which distinguishes normal children, children with mild to moderate autism, and children with severe autism by using the natural eye movement features of the target.
[0008] The technical solution of this application is as follows: On the one hand, the present invention provides a method for classifying autism spectrum based on eye-tracking, comprising the following steps: S1: Acquire the target's eye movement information and generate a gaze heatmap and gaze point trajectory map; S2: Obtain the clinical numerical characteristics of the target based on the target's eye movement information; S3: The target gaze heatmap, gaze trajectory map and clinical numerical features are fused using the multimodal fusion model MAAT-Net; The multimodal fusion model MAAT-Net extracts image features through an image feature extraction module, achieves semantic alignment of the feature space through multimodal feature projection and spatial alignment, fuses the cross-modal attention output features from two directions into the final interactive features through a bidirectional cross-modal attention interaction module, achieves feature fusion based on the spatiotemporal sequence aggregation of Transformer, and outputs the results through a deep classifier. S4: The output of the multimodal fusion model MAAT-Net predicts the probability of the severity of autism spectrum disorder.
[0009] Furthermore, in step S1, eye movement information is captured by an eye tracker after the target views the set image, and is stacked to form a gaze heatmap and a gaze point trajectory map.
[0010] Furthermore, in step S2, the clinical numerical features include one or more of the following: coordinate variance, spatial entropy, static fixation entropy, saccade fixation entropy, effective fixation / saccade count, average fixation duration, and mean duration of deduplicated eye movement events.
[0011] Furthermore, the multimodal fusion model MAAT-Net completes image feature extraction through an image feature extraction module. The image feature extraction module uses the lightweight backbone network EfficientNet-Lite0 as the feature extractor, which includes a shallow feature encoding layer, a core bottleneck layer, and a high-level feature aggregation layer. The shallow feature coding layer uses a standard 3×3 convolutional layer with a stride of 2, which includes a convolutional layer with 32 kernels, a batch normalization layer, and a ReLU6 activation function. The core bottleneck layer uses a 3×3 vertical convolutional layer with 32 groups to obtain Z1: (1) Then, the number of feature channels is expanded from 32 to 128 using a 1×1 convolution, as shown in the following formula: (2) The advanced feature aggregation layer increases the number of channels to 1280 dimensions through a 1×1 convolutional layer, and then uses an adaptive average pooling operator to compress the spatial dimension to 1×1, ultimately mapping the image modality into a fixed-length one-dimensional feature vector f. img ∈R 1280 ,in This represents the intermediate feature map processed by the lightweight convolutional module, and
[0012] (3) (4) Image features from various paradigms are projected onto a unified hidden layer dimension d through a linear layer. model Furthermore, learnable positional encodings are superimposed to capture structured information between tasks. This process is represented as follows: (5) Among them, E modal This is the modality-enhanced feature vector after fusing location information; Linear(·) represents the linear projection layer, whose core function is to map the original high-dimensional image features to a unified hidden layer dimension d. model ;f img ∈R 1280 For image modal one-dimensional feature vectors; PE∈R 3×64 This is a learnable location encoding matrix, where the dimension 3 corresponds to the three experimental paradigms, and 64 is the hidden layer dimension d. model The value of is used to encode sequence position information of different paradigms.
[0013] Furthermore, the multimodal fusion model MAAT-Net achieves semantic alignment of the feature space through multimodal feature projection and spatial alignment; Position encoding is achieved by fusing the projected features with element-wise addition. For trajectory modes, the enhanced features of the i-th paradigm are: (9) Where ET(i) represents the trajectory enhancement feature vector of the i-th paradigm, which integrates both trajectory content and location information; Wt represents the parameters of the linear projection layer, used to map the original trajectory features to a unified latent dimension dmodel; f (i) traj PE(i) is the trajectory image feature vector extracted from the i-th experimental paradigm; PE(i) is the learnable position code corresponding to the i-th paradigm, used to encode its position information in the evaluation sequence.
[0014] Furthermore, the multimodal fusion model MAAT-Net fuses the cross-modal attention output features from two directions into the final interaction features through a bidirectional cross-modal attention interaction module. The bidirectional cross-modal attention interaction module employs two independent multi-head attention modules. In each module, 64-dimensional features are projected into two 32-dimensional subspaces to compute attention in parallel. Each attention head in a multi-head attention model is calculated as follows: (10) The features after interaction are represented as follows: (11) Among them, F inter This represents the fused feature vector after bidirectional cross-modal interaction; MultiHeadAttn(⋅) represents the multi-head attention mechanism operator; Q traj、 Q heat Trajectory or heatmap features are used as query vectors in cross-modal attention computation, respectively; K heat K traj The heatmap or trajectory map features are used as the key vector V in cross-modal attention computation. heat V traj The heatmap and trajectory map features are used as value vectors in the cross-modal attention computation. The ⊕ symbol represents the feature concatenation operation, which fuses the cross-modal attention output features from the two directions into the final interactive feature.
[0015] Furthermore, the multimodal fusion model MAAT-Net achieves feature fusion based on the spatiotemporal sequence aggregation of Transformer and outputs the results through a deep classifier; Includes the following sub-steps: (1) Transformer sequence encoding: The input consists of multimodal features that have undergone interactive processing, and these features are then cascaded through formula (12) to form an enhanced feature sequence F. in ∈ R 9×dmodel : (12) A three-layer Transformer Encoder structure is adopted, and a multi-head self-attention mechanism is used to model long-range dependencies of heterogeneous features; each Transformer layer contains: Bullish Self-Attention Layer: (13) Feedforward network layer: (14) The feedforward network FFN(x) is: (15) (2) Learnable residual paths: The features of each mode are decomposed into trajectory features F. t Heatmap features F h sum numerical characteristics F n ; Perform element-wise mean pooling on all decomposed feature vectors within each mode to obtain the global feature g. t g h and g n The learnable residual weights α are activated using the sigmoid function; the residual enhancement features are calculated as follows: (16) in , , This is the global average of the original features.
[0016] (3) Modality-based adaptive weighting: The input is the concatenated and enhanced modal features: (17) Then, a small fully connected network, Modal Attention, is used to dynamically assign weights w=[w] to the three modalities. t ,w h ,w n ], (18) The final fusion feature f fused for: (19) fusion vector f fused The data is fed into a deep nonlinear classifier consisting of three linear layers, each containing a GELU activation function, batch normalization, and Dropout for regularization. The residual projection layer at the end of the classifier stacks the pre-classified deep features with the original fused features using learnable weights. The Softmax function generates the predicted probability distribution, and end-to-end training optimizes the cross-entropy loss function to output the predicted probability of the severity of autism spectrum disorder.
[0017] In another aspect, this application provides an electronic terminal, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is invoked and executed by the processor, it implements the steps in the method described above.
[0018] In another aspect, this application provides a computer-readable medium, characterized in that the computer-readable medium stores a computer program, which, when executed by a computer, implements the steps in the method described above.
[0019] In summary, the beneficial effects of this application are as follows: 1. A novel method for assessing the severity of autism spectrum disorder (ASD) is proposed, integrating multimodal eye-tracking biomarkers: fixation heatmaps, fixation trajectory maps, and fixation entropy indices. This method addresses key limitations of traditional binary ASD classification, enabling accurate quantitative severity grading and facilitating early clinical screening, refined patient stratification, and the development of personalized intervention programs.
[0020] 2. This innovative multimodal deep learning network integrates a bidirectional cross-modal attention mechanism with an adaptive residual fusion strategy. This framework achieves dynamic alignment and semantic complementarity of spatiotemporal features, and significantly improves the robustness of the model.
[0021] 3. This rapid, non-invasive eye-tracking experiment utilizes attention weighting analysis and visualization techniques to quantify the model's decision preferences: the typical developmental group relies on heatmaps, while the autism spectrum disorder group relies on trajectory maps. Furthermore, 3D coordination analysis of eye movements, head rotations, and gestures accurately identifies social cues, thereby detecting shared attention deficits and capturing key biomarkers, endowing the model with biological interpretability and enhancing its clinical reliability. Attached Figure Description
[0022] Figure 1 is a schematic diagram of the data collection process and joint attention task paradigm for multiple groups of subjects in one embodiment of the present invention. (a) Participants are screened, grouped, and calibrated with eye tracking based on the CARS scale to ensure data quality; (b) Paradigm design for the joint attention test, including stimulus presentation order and fatigue monitoring mechanism. P0 represents the buffer period; P1 represents the fixation indicator, P2 represents the fixation plus head rotation indicator, and P3 represents the fixation plus head rotation plus finger pointing to the red apple indicator; (c) Steps for extracting raw data and preprocessing data; (d) Subsequent model construction and classification.
[0023] Figure 2 is a heatmap and trajectory diagram of three representative samples of different grades selected from the participants in one embodiment of the present invention.
[0024] Figure 3 is an architecture diagram of a multimodal adaptive fusion autism spectrum disorder grading model in one embodiment of the present invention. (a) The overall model framework integrating eye-tracking trajectories, heatmaps, and clinical data. Multi-source features are dynamically fused through a cross-modal attention mechanism and a Transformer encoder, combined with modality adaptive attention; (b) An EfficientNet-Lite0 structure is used for preliminary feature extraction from heatmaps and trajectories; (c) A deep classifier is used to learn high-level semantic information from the fused features; (d) A modality adaptive attention structure is used to dynamically allocate modality weights to highlight key features.
[0025] Figure 4 shows the ROC curves and Acc and LOSS curves of eight eye-tracking feature extractors in one embodiment of the present invention. In the Acc and LOSS curves, the solid line represents the LOSS curve, and the dashed line represents the Acc curve. The same feature extractor is colored the same.
[0026] Figure 5 shows the horizontal distribution of modal weights based on the severity of ASD in one embodiment of the present invention. The violin plot illustrates the density and distribution of attention weights assigned to each modality by the model, and is presented in a classification manner according to the typical developmental group (level 0), the mild to moderate group (level 1), and the severe group (level 2).
[0027] Figure 6 shows the weight distribution of trajectory and heatmap attention mechanisms based on different paradigms in one embodiment of the present invention. The figure illustrates the attention allocation of the model under three different experimental paradigms (Paradigm 1-3). The left figure shows trajectory-based heatmap attention, and the right figure shows heatmap-based trajectory attention. The color intensity represents the weight magnitude, reflecting the degree of attention the model pays to features of different paradigms when fusing multi-source data.
[0028] Figure 7 This is a confusion matrix diagram obtained by performing five-fold cross-validation and LOSO validation on a tri-classification task and a binary classification task in one embodiment of the present invention, with recall, precision and overall accuracy.
[0029] Figure 8 This is an embodiment of the invention showing the ROC curves for ASD rank classification and TD and ASD classification tasks; the rank classification shows the model's classification performance for the three categories, and also provides the micro-average and macro-average curves, reflecting the model's discriminative ability on the overall dataset.
[0030] Figure 9 This is a ROC curve diagram from one embodiment of the present invention; Figure 10 It is the confusion matrix of five ablation models (excluding the complete model V6) in one embodiment of the present invention. Detailed Implementation
[0031] The specific embodiments of this application are described in detail below with reference to the accompanying drawings.
[0032] Before introducing the implementation scheme, the background of the present invention will be further explained.
[0033] The CARS scale is used to assess whether children aged 2 years and older have autism and the severity of autism. It covers 15 assessment items, each scored from 1 to 4 points based on the severity of symptoms. 1 point indicates age-appropriate behavior; 2 points indicate mild abnormality; 3 points indicate moderate abnormality; and 4 points indicate severe abnormality. The maximum score is 60. A total score below 30 suggests no autism; a total score of 30 to 36 with fewer than 5 items scoring below 3 suggests mild to moderate autism; a total score of 36 or higher with at least 5 items scoring above 3 suggests severe autism. The Centers for Disease Control and Prevention (CDC) currently lists the CARS scale as one of the recommended diagnostic tools for ASD.
[0034] Individuals with autism spectrum disorder (ASD) exhibit atypical patterns of social interaction, communication, and behavior, making it difficult to understand their social behaviors. The severity and symptoms of ASD vary from person to person, resulting in diverse behavioral manifestations. A recent study highlights the complexity of social interaction and the importance of face-to-face communication, noting that individuals with ASD are highly sensitive to social presence, which can affect their cognitive processing. Nonverbal cues such as eye contact and blinking are extensively studied because they reflect attention and engagement in social interactions. Atypical visual behaviors are most evident when studying fixation on social stimuli. Individuals with ASD often struggle to initiate and maintain eye contact and show reduced attention to facial areas. Manually measuring these cues is challenging, further underscoring the need for automated methods.
[0035] Joint attention refers to the ability to focus on the same object with another person. Research by Gideon Salter et al. shows that approximately 44% of infants begin exhibiting joint attention behaviors before 6 months of age, and 92% are able to do so by 9 months of age. In contrast, children with autism spectrum disorder (ASD) often struggle to capture and utilize social cues from others. Even when others show clear emotional or behavioral responses to environmental stimuli, ASD children may fail to notice and are unable to adjust their own behavior accordingly. This leads to their understanding of the environment relying more on direct personal experience than on indirect learning gained through social interaction. A core deficit in autism lies in the persistent impairment of social communication and interaction. Due to the lack of joint attention, ASD children cannot establish a shared focus of attention with others, resulting in a lack of common topics and a foundation for collaboration in social interactions, thus delaying the development of language and social skills. Therefore, this deficit is a core manifestation of social cognitive impairment in ASD children and a key objective in assessing the severity of autism and implementing intervention training.
[0036] Early eye-tracking studies for predicting autism spectrum disorder primarily employed traditional machine learning models such as logistic regression, support vector machines (SVM), and XGBoost. Recent research, through manual feature extraction, has revealed the potential of eye-tracking features as biomarkers to reflect visual attention and information processing deficits associated with ASD. However, the accuracy of machine learning classifiers varies significantly. As a crucial branch of artificial intelligence, deep learning has become a mainstream approach due to its ability to automatically learn features from raw eye-tracking data. For example, Ahmed et al. used deep convolutional neural networks (CNNs) for ASD classification, validating the effectiveness of ensemble learning. Tao et al. proposed the SP-ASDNET model, combining convolutional neural networks and long short-term memory networks to distinguish between ASD / TD groups by analyzing gaze paths on specific images. The CNN-RNN hybrid model achieved 17% higher accuracy in ASD classification than XGBoost, and Mujeeb et al. also demonstrated through eye-tracking research that deep learning outperforms machine learning in ASD screening.
[0037] Currently, technologies are available that utilize eye-tracking data from serious games for the verification and diagnosis of autism spectrum disorder (ASD). Eye-tracking technology has proven crucial for understanding the psychological functions within an individual's attention and cognitive systems. It is also a core biomarker for assisting in the identification of ASD. Non-invasive testing and treatment methods are also being strongly advocated. Currently, there are smartwatch-based educational learning systems for ASD patients that utilize artificial intelligence to enhance their educational and training effectiveness. Furthermore, in children aged 16 to 30 months referred to specialist clinics, eye-tracking-based social visual indicators can predict the clinical expert's diagnosis of autism. Therefore, based on previous research, a free-gazing paradigm specifically designed for ASD patients was adopted. This paradigm reduces the difficulty of subject cooperation by avoiding complex instructions, naturally capturing key eye-tracking features, thus demonstrating clinical value. Combined with simplified sequence segmentation, this method ensures data authenticity and provides high-quality training data for deep learning models, thereby promoting the clinical translation of ASD diagnostic tools.
[0038] A specific embodiment of the present invention provides a method for classifying autism spectrum based on eye-tracking, comprising the following steps: S1: Acquire the target's eye movement information and generate a gaze heatmap and gaze point trajectory map; S2: Obtain the clinical numerical characteristics of the target based on the target's eye movement information; S3: The target gaze heatmap, gaze trajectory map and clinical numerical features are fused using the multimodal fusion model MAAT-Net; The multimodal fusion model MAAT-Net extracts image features through an image feature extraction module, achieves semantic alignment of the feature space through multimodal feature projection and spatial alignment, fuses the cross-modal attention output features from two directions into the final interactive features through a bidirectional cross-modal attention interaction module, achieves feature fusion based on the spatiotemporal sequence aggregation of Transformer, and outputs the results through a deep classifier. S4: The output of the multimodal fusion model MAAT-Net predicts the probability of the severity of autism spectrum disorder.
[0039] In step S1, eye movement information is captured by an eye tracker after the target views the set image, and the eye movement information is stacked to form a gaze heatmap and a gaze point trajectory map.
[0040] In this embodiment, the target's eye movement information was obtained through experiments.
[0041] The subjects of the experiment were children with autism recruited from the Yangzhou Chuying Children's Development Center. A total of 74 children participated in the test, of whom 60 had completed the CARS scale assessment in advance. After data quality screening (effective eye movement data ratio ≥50%), 6 low-quality samples due to excessive head movement amplitude were removed, and 54 qualified samples were finally included. The clinical characteristics of the samples were as follows: 6 children without autism, 31 children with mild to moderate autism, and 17 children with severe autism; the age distribution was as follows: 13 preschool children aged 3-6 years and 41 school-age children aged 7-13 years.
[0042] To explore the social referencing and joint attention abilities of children with autism, a modified version of the NYSTRM study was developed, and a referential social stimulus video paradigm was designed. The specific design is as follows: The experimental stimulus videos focus on 'human-object' interactions as the core scenario, such as... Figure 1 As shown, the image includes a Chinese female model wearing a suit with her hair tied back, and two visual stimuli: a red apple image on the left and a green apple image on the right. The image size and position are fixed throughout the video to ensure stable spatial cues and exclude irrelevant variables, such as the distraction caused by the person changing clothes or the impact of changes in object position on spatial memory. This ensures that eye movement differences are only guided by social behavior or autistic traits.
[0043] The video is presented in a cyclical pattern, consisting of a 3-second behavioral stimulus followed by a 1-second fixation buffer. During the behavioral stimulus, the experimenter performs three referential social behaviors in a Latin square design sequence: glancing at the red apple with only eyes, turning the head to the right to look at the red apple, and turning the head to the red apple while pointing with the right hand. These three clear gaze and movement directions provide children with social reference cues, aiming to stimulate their collective attention. During the fixation buffer, a black plus sign appears in the center of the screen, connecting two consecutive stimuli. This eliminates visual persistence or cognitive interference from the previous stimulus and also allows participants time for cognitive recovery after completing their response, avoiding fatigue effects caused by rapid, continuous tasks.
[0044] The experiment consisted of 12 trials, with 3 trials per cue, for a total of 4 cues. To balance the direction of the cues and the left-right position of the apple, a Latin square design was used to present them in sequence. The specific playback order is shown in the squares in Fig. 1. The total duration of the experiment was 48 seconds. The duration of the experiment was designed to both conform to the characteristics of children's attention span, avoiding excessive rest periods that could lead to a loss of attention, and to clearly distinguish between spontaneous visual preferences and socially guided attention shifts through baseline-stimulus contrast.
[0045] The data collection and preprocessing process is as follows: The experimental setup included a Tobii Pro Fusion high-performance portable eye tracker to record binocular fixation as near-infrared light reflected from the cornea and pupil, and a laptop computer with Tobii Pro Lab software to present stimuli and collect eye movement data from the participants. The display screen resolution was set to 1920×1080 pixels at a frame rate of 30 frames per second; the eye tracker was positioned below the laptop screen; the sampling accuracy was 0.5°, and the eye tracker sampling frequency was 120Hz.
[0046] The eye-tracking experiments in this study were conducted entirely within the assessment room of the rehabilitation facility for children with autism. The testing environment was quiet, comfortable, well-lit, and enclosed. Only one tester and one subject were present during each experiment. Before the experiment began, a teacher from the rehabilitation facility escorted the subject into the eye-tracking assessment room. Upon arrival, the tester interacted with the subject to establish a good rapport, aiming for the best testing results.
[0047] During the experiment, children sat opposite a laptop and watched stimulating videos at a distance of approximately 50-60 cm from the screen. The screen height was at eye level, and the allowed head movement was 45cm × 25cm × 33cm. A natural viewing mode was used, and no additional task instructions were given during the experiment. The study used a cartoon duck with vocalizations included in the Tobii Studio software to attract the participants' gaze for 5-point calibration. Once calibration was successful, the formal eye-tracking experiment began.
[0048] After the experiment, children who completed the eye-tracking experiment were praised and rewarded with small snacks. Throughout the experiment, the tester was required to sit at a 90° angle to the subject, who was then tested individually. Children who experienced discomfort or did not wish to continue the test were allowed to leave and withdraw from the study.
[0049] The eye tracker simultaneously captures the subjects' eye movement information during the experiment, and finally exports the gaze heatmap and gaze point trajectory map through Tobii Pro Lab software as the core analysis carrier.
[0050] This study used Tobii Pro Lab software to standardize and preprocess the raw data. The core steps included invalid data removal and valid data rate quantification. First, noise data caused by blinking and large head movements were removed to ensure that the valid data rate was ≥90%. Then, the data quality was measured by the proportion of valid eye movement data points to the theoretical total number of data points. Samples with less than 50% valid eye movement data points were judged as substandard and removed.
[0051] The size of the training dataset is a key factor affecting the performance of deep learning models. Sufficient data can provide the model with rich feature information, helping it learn and summarize the inherent patterns of the samples. This study is limited by the difficulty in recruiting children with autism, the further reduction of samples for each subtype after classification by the CARS scale, and the strict restrictions on privacy protection and institutional data management, resulting in a limited size of the core eye-tracking dataset. At the same time, this dataset mainly collects data from children with autism, with relatively few typical children included, leading to a certain imbalance in the sample distribution within the dataset. To address the dual challenges of insufficient data size and uneven distribution, this study adopts a two-step strategy to optimize the dataset. First, the sample size is expanded based on the experimental design. Since the experimental paradigm is repeated four times according to the Latin square design and presented in groups, each sample in the paradigm combination of the four repeated experiments is considered an independent sample, increasing the sample size to four times the original size without introducing class bias. The sample and label information is shown in Table 1. During model training, the SMOTE algorithm is used to alleviate class imbalance. New feature vectors are synthesized only on the training set in each round by linear interpolation between minority class samples and their nearest neighbors, thus balancing the class distribution of the training set and guiding the model to learn a classification boundary with stronger generalization. The test set maintains its original distribution to ensure that the evaluation results can truly reflect the discriminative efficacy of the model in clinical scenarios.
[0052] To verify the robustness of the model and avoid overfitting, this study employed five-fold stratified cross-validation. Considering that each participant contributed sequence data across multiple experimental paradigms, the participant ID was used as the group label. This ensured that in each fold of the cross-validation, all samples from the same individual were strictly isolated in either the training or test set, effectively preventing data leakage due to the coupling of individual characteristics.
[0053]
[0054] Table 1 shows the distribution of the original and expanded samples, and the corresponding labels for each sample. Both the gaze heatmap and gaze trajectory map are stacked from eye movement information captured by the eye tracker after the sample viewed the stimulus image for 3 seconds, and both are output in JPG format. Figure 2 Heatmaps and trajectory maps were generated from viewing stimuli in three children with varying degrees of illness, illustrating their typical viewing patterns. The heatmaps, presented as visual heat distribution, visually represent the degree of focus on different areas of interest throughout the video viewing process, using a hot color scheme (red-yellow-green gradient). Red indicates areas of highest fixation density, followed by yellow, and green indicates areas of lower density, clearly showing the distribution intensity of visual attention. The trajectory maps, using different colors and employing fixation points connected by lines, visually represent changes in the child's fixation position. Furthermore, the size of the fixation point is directly proportional to the duration of the child's fixation.
[0055] By comparing the gaze heatmaps of children with typical developmental patterns and those with autistic tendencies, we can find that the heatmaps of children with autistic tendencies often exhibit characteristics such as "a high proportion of heatmaps for non-social objects during the resting period, a delayed and low proportion of heatmaps for target AOIs during the stimulus period, and dispersed heatmap transfer paths." Furthermore, the degree of these heatmap abnormalities is positively correlated with the severity of autistic social impairment. For example, the lower the proportion of gaze duration for target AOIs and the longer the gaze initiation latency, the more significant the child's combined attention deficit, potentially corresponding to a more severe autistic tendency. In conclusion, the gaze heatmap derived from this paradigm can serve as a supplementary behavioral indicator for assessing the severity of autism. Its visualized heatmap distribution characteristics not only intuitively reflect children's visual attention patterns but also provide objective and repeatable data support for early autism screening.
[0056] This embodiment proposes a multimodal fusion model, MAAT-Net, based on residual adaptive Transformer, which aims to integrate trajectory plots, heatmaps, and clinical numerical features for automatic assessment of ASD prevalence. Figure 3 (a) illustrates the complete process from feature extraction to feature processing and then to sample classification. This architecture captures spatial relationships between images through a bidirectional cross-modal attention mechanism, extracts deep semantic features using a Transformer encoder with residual paths, and finally achieves dynamic integration of different modalities through an adaptive weight module. This model achieves detailed classification of autism grading, realizing efficient classification tasks based on lightweight machine learning algorithms.
[0057] A. Image feature extraction: For image modalities, namely trajectory maps and heatmaps, the lightweight backbone network EfficientNet-Lite0 is used as the feature extractor. Figure 3 (b) illustrates its specific structure. This network, while maintaining computational efficiency, extracts high-dimensional semantic vectors through moving and reversing bottleneck convolutions, significantly reducing computational overhead while ensuring high-precision feature extraction. For each paradigm's image sequence, the extracted features are mapped to a vector f of length 1280. img ∈R 1280 While ensuring high-precision feature extraction, computational overhead is significantly reduced. The specific hierarchical structure is as follows: 1. Shallow Feature Coding Layer (Stem Stage) The network's first layer employs a standard 3×3 convolutional layer with a stride of 2, containing a convolutional layer with 32 kernels, a batch normalization layer, and a ReLU6 activation function. ReLU6 limits the dynamic range of activation values, enhancing the model's robustness on low-precision computing devices. The original image, with dimensions of 3 × 224 × 224, is first spatially downsampled to extract low-level visual features.
[0058] 2. Core Bottleneck Layer (Middle Stage) This section borrows design principles from MobileNetV2, employing a lightweight inverted bottleneck structure. Z1 is obtained by using 3×3 vertical convolutional layers with 32 groups, thus capturing spatial correlations without increasing the number of model parameters.
[0059] (1) Then, the number of feature channels is expanded from 32 to 128 using a 1×1 convolution, as shown in the following formula: (2) This "compression followed by expansion" structure can effectively perform linear combination of features in a higher-dimensional space.
[0060] 3. Advanced Feature Aggregation Layer (Head Stage) The number of channels is increased to 1280 dimensions using a 1×1 convolutional layer, as shown in Equation 3. This high-dimensional mapping aims to convert spatial features into semantic features, providing a rich expression space for subsequent multimodal fusion. An adaptive average pooling operator is used to compress the spatial dimension (H×W) to 1×1. Finally, the image modality is mapped to a fixed-length one-dimensional feature vector f. img ∈R 1280 As shown in Formula 4, where This represents the intermediate feature map processed by the lightweight convolutional module, and
[0061] (3) (4) Considering the temporal logic and behavioral correlations between different experimental paradigms, the image features of each paradigm are projected to a unified hidden layer dimension d through a linear layer. model Furthermore, learnable location codes are superimposed to capture structured information between tasks. This process is represented as: (5) Among them, E modal This is the modality-enhanced feature vector after fusing location information; Linear(·) represents the linear projection layer, whose core function is to map the original high-dimensional image features to a unified hidden layer dimension d. model ;f img ∈R 1280 The image modality one-dimensional feature vector mentioned above; PE∈R 3×64 This is a learnable location encoding matrix, where the dimension 3 corresponds to the three experimental paradigms, and 64 is the hidden layer dimension d.model The value of is used to encode sequence position information of different paradigms. This design ensures that the model can recognize the context of different visual tasks in the subsequent Transformer encoding stage.
[0062] B. Calculation of spatial, entropy, and dynamic eye-tracking numerical features: In addition to eye-tracking heatmaps and fixation trajectory maps, the following features were calculated and incorporated into the model training process using fixation location information at each time stamp of each sample. Specifically, in step S2, the clinical numerical features include coordinate variance, spatial entropy, static fixation entropy, saccade fixation entropy, number of effective fixations / saccades, average fixation duration, and mean duration of deduplicated eye-tracking events.
[0063] 1. Spatial distribution characteristics Coordinate variance (X / Y Variance): Reflects the spatial dispersion of the subject's fixation points, quantifying the stability of fixation points. It is calculated by plotting the fixation point sequence X = {x1,...,x...}. n The unbiased variance of} quantifies the spatial dispersion of the subject's fixation point in the horizontal and vertical directions, reflecting the stability of fixation.
[0064] (6) 2. Features based on information entropy To quantify the randomness of visual search based on entropy-driven uncertainty, and considering that entropy-based methods can detect heterogeneous sensory patterns in autism spectrum disorders, the visual stimulus region is divided into a grid consisting of M×N equal-area grids. For each grid, the fixation probability p within it is calculated. k Calculate the three entropy indices: Spatial Entropy: Based on information entropy theory, it quantifies the uncertainty and uniformity of the distribution of fixation points in visual space. A higher entropy value indicates a more dispersed fixation distribution. It reflects the global uniformity of the fixation point distribution in visual space.
[0065] Static Fixation Entropy: Focuses on the spatial distribution entropy of gaze points in a static visual scene, quantifying the uncertainty of gaze distribution without time-series dependence, highlighting the discrete characteristics of spatial location. The core difference between this feature and spatial entropy is that it eliminates the influence of the time dimension, reflecting only the static characteristics of spatial distribution.
[0066] Saccade Fixation Entropy: Saccade fixation entropy is a quantitative indicator characterizing the randomness of the distribution of saccade motion targets. Its core is a measure of the spatial distribution uncertainty of fixation points associated with saccades. During calculation, the set of start and end fixation points S of valid saccade events is first extracted. start With S end The result of merging is S saccade The total number of points q is counted; then, the gaze probability p of each grid is calculated by dividing the grid into M×N grids. k =q k / q.
[0067] The above entropy indices are all defined as follows: (7) 3. Visual processing and motion dynamics Effective Fixation / Saccade Count: Eye movement events were identified using the I-VT velocity threshold algorithm, with the criteria being fixation duration ≥ 60ms and saccade amplitude ≥ 1°, to characterize the activity and switching frequency of the subjects' visual exploration.
[0068] Mean Fixation Duration: The arithmetic mean of the durations of all valid fixation events is calculated as a core indicator reflecting the depth of visual information processing in the subjects.
[0069] Deduplicated Mean Event Duration: Valid events within the range of 10~500ms are extracted and deduplicated to obtain a unique duration sequence T′=[t1′,t2′,...,t...]. k ′](k≤m;Calculate the mean after deduplication, which is expressed as Mean in formula (8). deduplicated This is to eliminate statistical bias caused by equipment noise and redundant records.
[0070] (8) The aforementioned feature calculation methods all adhere to international standards for eye-tracking data analysis, taking into account both the physical meaning and repeatability of the features. All features have been optimized using clearly defined thresholds to eliminate the influence of device noise and outlier data, making them suitable for the differential representation of eye-tracking behavior patterns in a three-category autism classification task.
[0071] C. Feature Fusion Network: To address the heterogeneity between eye-tracking physiological characteristics and visual search patterns, this study designed and implemented a method that uses a cross-modal attention mechanism and a Transformer encoder to perform deep alignment and dynamic fusion of trajectory maps, heatmaps, and numerical features.
[0072] 1. Multimodal Feature Projection and Spatial Alignment The raw data of different modalities exist in completely heterogeneous feature spaces: trajectory images and heatmap images are extracted by convolutional neural networks to obtain a 1280-dimensional high-dimensional visual feature vector f. img ∈R 1280 Clinical numerical features only have a 9-dimensional numerical vector f. num ∈R 9 This dimensionality difference and distribution inconsistency directly hinders subsequent fusion operations. Therefore, dimensionality reduction compresses high-dimensional visual features into more discriminative low-dimensional representations, avoiding the "curse of dimensionality" in high-dimensional space and reducing computational complexity. Secondly, dimensionality expansion extends low-dimensional numerical features to a dimension matching the visual features, ensuring that numerical information is not dominated by visual features in subsequent attention calculations. Most importantly, the projection process itself achieves semantic alignment of the feature space, making features from different modalities comparable and operable in a unified latent space, laying the foundation for subsequent cross-modal interactions.
[0073] The three experimental paradigms—eye movement + head + hand movement, eye movement + head movement, and eye movement alone—constitute an inherently logical sequence of behaviors, reflecting a progressive process from complex behaviors to focused attention. Traditional methods often ignore this sequential association, simply splicing together the features of the three paradigms. This study innovatively introduces a learnable positional encoding PE∈R. 3 ×d model This design explicitly injects sequence information into the feature representation. On the one hand, positional encoding enables the model to distinguish the behavioral characteristics of the same subject in different paradigms, preserving the phased information of the evaluation process. On the other hand, the learnable positional encoding parameters are adaptively adjusted during training, better capturing the relative relationships and transition patterns between the three paradigms. Positional encoding is fused with the projected features through element-wise addition. For the trajectory modality, the enhanced feature for the i-th paradigm is: (9) Where ET(i) represents the trajectory enhancement feature vector of the i-th paradigm, which integrates both trajectory content and location information; Wt represents the parameters of the linear projection layer, used to map the original trajectory features to a unified latent dimension dmodel; f (i) trajPE(i) is the trajectory image feature vector extracted from the i-th experimental paradigm; PE(i) is the learnable positional code corresponding to the i-th paradigm, used to encode its positional information in the evaluation sequence. This design ensures that the features of each paradigm contain both its specific content information and its encoded positional information in the evaluation sequence.
[0074] 2. Bi-directional Cross-Modal Attention (Bi-CMA) To address the limitations of single-visual modality representation, this model constructs a bidirectional cross-modal interaction module. This module employs two independent multi-head attention modules. In each module, 64-dimensional features are projected into two 32-dimensional subspaces for parallel attention computation. This enables the model to simultaneously capture cross-modal relationships between trajectories and heatmaps, thereby achieving information complementarity.
[0075] This module supports bidirectional interactive retrieval of trajectory and heatmap features in the semantic space. A salient region of one modality is used to enhance the feature representation of another modality.
[0076] Trajectory features and heatmap features serve as each other's query vector and key-value vector, respectively. For example... Figure 3 As shown in (a), cross-modal attention (trajectory-focused heatmap) means that trajectory features actively focus on key spatial hotspot regions in heatmap features, thereby enabling trajectory features to integrate spatial distribution information from the heatmap. In contrast, cross-modal attention (heatmap-focused trajectory) means that heatmap features actively focus on the motion paths of key fixations in trajectory features, thereby enabling heatmap features to integrate temporal dynamic information from the trajectory.
[0077] Each attention head in a multi-head attention model is calculated as follows: (10) The features after interaction are represented as follows: (11) Among them, F inter This represents the fused feature vector after bidirectional cross-modal interaction; MultiHeadAttn(⋅) represents the multi-head attention mechanism operator, which is the core module for realizing cross-modal feature interaction; Q traj、 Q heat Trajectory or heatmap features are used as query vectors in cross-modal attention computation, respectively; K heat K traj分别 Use heatmap or trajectory map features as key vectors V in cross-modal attention computation. heat V traj分别The trajectory map features obtained from the heatmap are used as the value vector ⊕ in the cross-modal attention computation to represent the feature concatenation operation, which fuses the cross-modal attention output features from two directions into the final interactive features.
[0078] 3. Spatio-temporal Sequence Aggregation Based on Transformer It includes the following three sub-steps: (1) Transformer sequence encoding: The input consists of multimodal features that have undergone interactive processing, and these features are then cascaded through formula (12) to form an enhanced feature sequence F. in ∈ R 9×dmodel .
[0079] (12) The model adopts a three-layer Transformer Encoder structure. Figure 3 (a) is represented by blue squares, which use a multi-head self-attention mechanism to model the long-range dependencies of heterogeneous features under three experimental paradigms. Each Transformer layer contains: Bullish Self-Attention Layer: (13) Feedforward network layer: (14) The feedforward network FFN(x) is: (15) (2) Learnable residual paths: To prevent over-smoothing of features while preserving original observation information, learnable residual connections (such as...) were designed. Figure 3 (a) As shown by the dashed line). First, the features of each mode are decomposed into trajectory features F. t Heatmap features F h sum numerical characteristics F n Subsequently, element-wise mean pooling is performed on all decomposed feature vectors within each mode to eliminate local fluctuations, thereby obtaining the global feature g. t g h and g n The learnable residual weights α are activated using the sigmoid function. Finally, the residual enhancement features are calculated as follows: (16) in , , This is the global average of the original features.
[0080] (3) Modal-wise Adaptive Weighting: Considering the individual differences in the dominant modalities of different subjects in ASD diagnosis, an adaptive attention gating module was designed in the model.
[0081] The input is the concatenated and enhanced modal features: (17) Then, the module dynamically assigns weights w=[w] to the three modalities through a small fully connected network (Modal Attention). t ,w h ,w n ], (18) The final fusion feature f fused for: (19) D. Deep Classifier: fusion vector f fused It is input into a deep nonlinear classifier consisting of three linear layers, such as Figure 3 As shown in (c), each layer contains a GELU activation function, batch normalization, and Dropout (rate = 0.3) for regularization. The residual projection layer at the end of the classifier stacks the pre-classified deep features with the original fused features using learnable weights. The Softmax function generates the predicted probability distribution, and end-to-end training optimizes the cross-entropy loss function to output the predicted probability of autism spectrum disorder severity.
[0082] E. Evaluation Indicators: Five performance metrics were used to evaluate the task: accuracy, macro precision, macro recall, macro F1 score, and area under the curve (AUC). The F1 score represents the harmonic mean between precision and recall, while the AUC in the three-class classification task uses an OVR strategy, which calculates the overall AUC by weighting the sample size of each class.
[0083] Accuracy: Accuracy: Recall rate: F1 score: AUC: TP represents true positive, FP represents false positive, TN represents true negative, and FN represents false negative; the weighted average index is calculated based on the sample size N for each category. k The weighted calculation is based on the proportion of the total sample size N to adapt to imbalanced data scenarios.
[0084] The experiment and results are as follows: A. Implementation details: CUDA acceleration was employed; a fixed random seed of 42 was used to ensure experimental reproducibility, covering random, numpy, PyTorch, and CUDA random number generators. To avoid intra-group data leakage, a 5-fold hierarchical grouped K-fold cross-validation was used to partition the dataset, and the Synthetic Minority Oversampling Technique (SMOTE) was applied to the training set to alleviate class imbalance. For image preprocessing, all visual samples were resized to a uniform 224×224 pixels before feature extraction and normalized using a mean and standard deviation of 0.5.
[0085] The multimodal fusion model is configured with 64 core feature projection dimensions and a multi-head attention mechanism with two heads. The Transformer encoder consists of three stacked encoder layers, and the three fully connected layers of the deep classifier module have dropout rates of 0.3, 0.3, and 0.2, respectively. The residual connections are initialized with learnable weight parameters α of 0.5 and learnable scaling factors of 0.1, both of which are constrained to the range [0,1] by the Sigmoid function.
[0086] Model training was performed using the AdamW optimizer, with a weight decay factor of 1×10⁻⁶ for all trainable parameters. −2 The initial learning rate is 2×10 −4 The batch size was set to 8, and the total training epochs were 100. Gradient clipping with a maximum norm of 1.0 was implemented to prevent gradient explosion. The model's performance was primarily evaluated using weighted F1 scores, supplemented by overall accuracy and confusion matrix analysis to quantify the classification performance across all classes.
[0087] B. Eye-tracking feature extractor comparison experiment: To select the optimal eye-tracking feature extractor for the ASD-level tri-class classification task, a 5-fold hierarchical cross-validation experiment was conducted on eight mainstream pre-trained visual models. As shown in Table 2, EfficientNet-Lite0 exhibited the best performance, with an accuracy of 76.18%, precision of 77.56%, an F1 score of 0.7601, and an AUC of 0.7015, ranking first in all four core metrics and demonstrating optimal ability to capture eye-tracking data features and distinguish between categories. Following EfficientNet-Lite0, Vision Transformer, ResNet-50, and MobileNetV3-Large all achieved accuracy, F1, and AUC values between 64% and 70%. Vision Transformer, leveraging the global feature modeling capabilities of the Transformer architecture, demonstrated outstanding stability in precision, with a standard deviation of only 0.0507. VGG-16 and Swin... Transformer, InceptionV3, and DenseNet-121 performed poorly, with all metrics below 0.67. DenseNet-121 performed the worst, with an AUC of only 0.6018, making it difficult to effectively distinguish the differences in eye movement characteristics among different degrees of ASD.
[0088]
[0089] Table 2 Results of comparative experiments on feature extraction networks The results above show that EfficientNet-Lite0 and MobileNetV3-Large, as lightweight CNN models, have advantages in both performance and computational efficiency. While heavyweight CNN models such as ResNet-50 and DenseNet-121 possess deeper network structures and richer feature representation capabilities, they perform worse than lightweight models in this task and consume more computational resources. Vision Transformer and Swin Transformer rely on global feature modeling capabilities obtained through pre-training on large-scale natural images. However, the gaze trajectory and spatial distribution features of eye-tracking images are specific and differ significantly from the feature distribution of natural images, preventing them from fully leveraging their architectural advantages. Their performance is lower than EfficientNet-Lite0, and the accuracy of Swin Transformer is only 0.6618, failing to demonstrate the advantages of the cross-window attention mechanism. As a representative of traditional CNN models, VGG-16, despite its fully convolutional architecture being able to extract local features, lacks an adaptive feature adjustment mechanism and is subject to the risk of gradient vanishing. As a result, it performs poorly in eye-tracking feature extraction tasks, with all indicators at a below-average level, and its AUC value standard deviation reaches 0.1307, making it the least stable.
[0090] Depend on Figure 4 The ROC curves show that the curves of the eight pre-trained models are highly similar, with Micro-AUC values concentrated in a narrow range of 0.72 to 0.79. No single model's curve shows a significant lead or lag. Even the best-performing EfficientNet-Lite0's AUC value is only about 0.19 higher than the worst-performing DenseNet-121, and the curves of most models are very close. This indicates that when different pre-trained models are used as feature extractors, the eye-movement features captured by these models show little difference in their ability to distinguish categories in autistic children. This suggests that eye-movement features under this paradigm have strong consistency and captureability, and different network architectures do not produce a qualitative difference in their ability to extract core features. The training curves show that the loss curves of all models decrease rapidly in the early stages of training, and the final accuracy curves are very similar with small fluctuations. No single model shows a significant breakthrough or lag in accuracy. This further illustrates that the eye-movement features output by different feature extractors are of similar quality, and the models converge relatively quickly during training. The feature extraction stage is not the core factor determining the final performance; rather, the subsequent feature fusion strategies and classifier design are the key to improving model performance. The results of both figures jointly demonstrate that, within this experimental paradigm, the eye-movement features of autistic children exhibit strong uniformity and captureability, and the performance differences between different pre-trained models used as feature extractors are not significant. This implies that future research should focus on optimizing multimodal feature fusion strategies and classification modules, rather than simply pursuing more complex feature extraction networks.
[0091] Based on the above performance analysis and task requirements, EfficientNet-Lite0 was ultimately selected as the feature extractor for eye-tracking images.
[0092] C. Optimizer performance comparison experiment: To determine the optimal optimization strategy for the multimodal ASD hierarchical classification model, this study selected four mainstream optimizers—AdamW, Adam, SGD (with Nesterov momentum), and RMSprop—for comparative experiments. The experimental results are shown in Table 3. As can be seen from the table, AdamW performed best across all evaluation metrics, achieving an accuracy of 78.00%, which is 7.45%, 7.82%, and 20.73% higher than Adam, RMSprop, and SGD, respectively. This indicates that AdamW can more accurately complete the ASD hierarchical classification task and has the best optimization effect for multimodal feature fusion. The F1 score, as the harmonic mean of precision and recall, verifies AdamW's advantage in balancing the reduction of misclassification and the improvement of sample coverage, especially suitable for scenarios with imbalanced sample distribution, such as ASD hierarchical classification. AUC, as the core indicator for evaluating the discriminative ability of the classification model, reflects AdamW's stronger robustness on this task. Adam and RMSprop have similar performance and are in the second tier, while SGD's various metrics are significantly lower than other optimizers, showing a clear overall performance gap.
[0093]
[0094] Table 3 compares the accuracy, precision, recall, F1 score, and AUC score of four different optimizers. AdamW, an improved version of Adam, introduces a weight decay (1e-2) mechanism, effectively mitigating overfitting caused by parameter redundancy during multimodal feature fusion. It also retains Adam's adaptive learning rate characteristic, dynamically adjusting the learning rate based on gradient changes in the cross-modal attention module, balancing optimization efficiency and model generalization. While Adam performs slightly worse, its lack of an explicit weight decay mechanism makes it prone to parameter overfitting when fusing heterogeneous features such as trajectories, heatmaps, and numerical values, resulting in slightly lower performance than AdamW. RMSprop adjusts the learning rate through exponential moving average, making it more suitable for non-stationary target optimization scenarios such as recurrent neural networks; however, within the Transformer cross-modal attention framework of this study, its learning rate adjustment strategy cannot accurately match the gradient distribution of multimodal features, resulting in slightly inferior performance compared to Adam. Despite configuring SGD with 0.9 momentum and Nesterov acceleration, a fixed learning rate struggles to adapt to the complex optimization space of multimodal features, easily getting trapped in local optima and exhibiting slow convergence, ultimately leading to significantly lower performance across all metrics.
[0095] Experimental results show that AdamW is the optimal optimizer for the multimodal ASD hierarchical classification model in this study. Its weight decay mechanism and adaptive learning rate feature can effectively solve the overfitting problem in multimodal feature fusion while ensuring optimization efficiency. Therefore, AdamW was selected as the final optimizer when building the model.
[0096] D. Attention weight analysis of model decision-making: Figure 5 The violin plot reveals that the model employs drastically different decision-making strategies for different groups, a key finding of this study. The model's determination of the normal developmental group is highly dependent on the heatmap. As shown in the figure, the median weight of normal children (Level 0) on the heatmap is significantly higher than that of the mild-to-moderate group (Level 1) and the severe group (Level 2). This indicates that the visual attention of normal children exhibits high regularity and spatial consistency, allowing the model to efficiently determine ASD by identifying these typical attention areas. In contrast to the normal group, the model significantly shifts its attention to the trajectory plot when determining ASD patients. The weight distribution of the mild-to-moderate and severe groups on the trajectory plot is significantly higher than that of normal children. This suggests that the diagnostic signals for ASD are more hidden in the dynamic sequence of visual scans than in simple spatial distribution. Numerical indicators maintain a relatively stable low-weight distribution across groups, but exhibit greater volatility at Level 0, suggesting that certain physiological statistics play an auxiliary "corrective" role in excluding atypical samples.
[0097] The dynamic shift in weights reflects the fundamental differences in ASD behavioral manifestations. Normally developing children typically follow a highly consistent socialized visual paradigm when viewing social stimuli, a regularity most pronounced in static heatmaps. In contrast, the visual search patterns of ASD patients often exhibit fragmentation, atypicality, and disordered scanning logic. Simple spatial distribution is insufficient to capture these complex pathological features. Therefore, MAAT-Net's adaptive attention mechanism automatically increases the weight of trajectory maps when processing ASD samples, enhancing diagnostic sensitivity by capturing irregular visual scanning temporal logic. This attention shift from spatial distribution to dynamic logic demonstrates the unique advantage of multimodal fusion models in capturing the heterogeneity of ASD.
[0098] E. Distribution characteristics of cross-modal attention in trajectory-heatmap under multi-paradigm: Figure 6The diagram illustrates the cross-modal attention weight distribution between trajectory features and heatmap features across three experimental paradigms. The left figure shows the attention from trajectory to heatmap, where trajectory features are used as a query to focus on heatmap features; the right figure shows the attention from heatmap to trajectory, where heatmap features are used as a query to focus on trajectory features. Color scales are used to quantify the magnitude of the attention weights, ranging from 0.1 to 0.5.
[0099] Attention from trajectory to heatmap exhibits a clear intra-paradigm bias: trajectory features from Paradigm 1 and Paradigm 3 assign the highest weights to heatmap features within their respective paradigms (dark red, approximately 0.45–0.5). This pattern suggests that trajectory features tend to rely on intra-paradigm heatmap information to enhance their representational power. Conversely, attention from heatmap to trajectory shows a cross-paradigm bias: heatmap features from Paradigm 1 and Paradigm 2 prioritize trajectory features from other paradigms (dark red, ≈0.45–0.5). This implies that heatmap features (capturing spatial gaze density) require supplemental temporal trajectory information from other paradigms to improve their discriminative power, consistent with clinical observations that behavioral patterns of autism spectrum disorder are reflected in multiple experimental scenarios.
[0100] Overall, this asymmetric attention pattern suggests that the model dynamically adjusts its information fusion strategy based on modality features. Trajectory features leverage intra-paradigm consistency, while heatmap features utilize inter-paradigm complementarity. Furthermore, features from Paradigm 1 (YANSHEN+TOU+SHOU) are frequently assigned high weights, indicating that this paradigm contains the most clinically relevant behavioral biomarkers for stratifying autism spectrum disorder severity, thereby enhancing the model's interpretability and translational value.
[0101] F. Performance of MAAT-Net network: As shown in Table 4, the model demonstrates good discriminative ability across the three levels of ASD. The overall five-fold cross-validation accuracy reached 77.78%, the F1 score was 0.7769, the leave-one-out cross-validation accuracy was 81.48%, and the AUC value reached 0.8503, indicating good robustness in multi-class classification tasks. The model performed best in the mild to moderate category (Category 1), achieving an F1 score as high as 0.8254. While the F1 score for normal children (Category 0) was relatively lower, the AUC value reached 0.7821, showing that the model has strong potential discriminative power for this category.
[0102] Furthermore, the model's performance in a binary classification task distinguishing between ASD and TD was tested, as shown in Table 5. This multimodal fusion model was then transferred to a binary classification task involving ASD and typical development, demonstrating strong classification efficiency. The overall five-fold cross-validation accuracy reached 90.74%, and the LOSO accuracy was 92.59%, validating the model's generalization ability in binary classification scenarios. From a category-level perspective, the model's performance in recognizing the ASD category was particularly outstanding. In contrast, the model's performance in recognizing the TD category had room for improvement. The difference in performance between the two categories stemmed from the imbalanced distribution of the training data (192 ASD samples versus only 24 TD samples), leading to a higher fit of the majority class features during model learning. In summary, this multimodal fusion model exhibited excellent overall performance in the ASD binary classification task, and its ASD category recognition accuracy met practical application requirements, providing reliable technical support for the assisted diagnosis of ASD.
[0103] More specific classification results are available in Figure 7 The two confusion matrices are shown. It can be seen that the model performs best in classifying mild to moderate ASD in the three-class classification task, while it performs significantly better than TD in the two-class classification task. This may be related to the large difference in sample size. In addition, the high behavioral variability of TD subjects due to the lack of stable abnormal features is also one of the reasons for this result.
[0104]
[0105] Table 4 shows the accuracy, precision, recall, F1 score, and AUC of the model under five-fold cross-validation and LOSO validation at three autism levels.
[0106]
[0107] Table 5. Performance evaluation results of ASD and TD binary classification tasks. Num represents the number of samples in each class and the overall population. Accuracy, precision, recall, and F1 score are the core evaluation metrics for model classification performance, used to quantify the model's discriminative power against two classes in imbalanced sample scenarios. The area under the micro-mean ROC curve is 0.7986, effectively characterizing the model's overall performance across the entire cohort, aligning with practical clinical needs, namely prioritizing overall diagnostic accuracy. In contrast, the macro-mean ROC curve (dark pink dashed line) treats all ASD severity levels equally regardless of sample size. Its corresponding AUC value quantifies the model's balance performance across different classes, avoiding bias caused by class imbalance. The consistency between the two curves indicates that the model not only achieves satisfactory overall classification but also maintains stable performance at each ASD level, mitigating the impact of potential class imbalance and validating the reliability of the multimodal fusion framework.
[0108] G. Baseline Model Comparison: To validate the advantages of the proposed MAAT-Net model, ten baseline models were selected, including models using only unimodal data. Evaluation employed 5-fold hierarchical cross-validation and 50 bootstrap validations. All models were trained and validated using the same visual feature dataset to ensure the fairness and validity of the comparison.
[0109] For Table 6 and Figure 9A comprehensive analysis of the ROC curves yielded the following conclusions: 1) Simple fusion (direct splicing fusion) outperformed other baseline models, indicating that multimodal feature fusion significantly improves eye-tracking classification. 2) Traditional machine learning models are not suitable for fine-grained eye-tracking classification because they cannot explore cross-modal differences or capture the correlation and complex interactions of intrinsic features. 3) Classical convolutional neural network fusion models (VGG11, GoogLeNet, ResNet18) outperformed traditional machine learning, achieving a binary classification accuracy of nearly 80% under 5-fold cross-validation. However, due to the lack of modality-specific modeling and deep feature interactions, their performance and generalization stability were inferior to MAAT-Net. 4) MAAT-Net addressed nearly 20% of the shortcomings of Transformers models, such as feature smoothing, information decay, and uniform interactions, thus achieving deep fusion and adaptive modeling of eye-tracking data in autism spectrum disorders. 5) Models using single numerical features outperformed models relying solely on visual features, confirming that clinical entropy features can effectively extract key discriminative information from eye-tracking patterns. While a single visual modality provides complementary spatiotemporal information, multimodal fusion is crucial for capturing eye-tracking dynamics and ensuring accurate classification. 6) XGBoost is suitable for small to medium-sized datasets but fails to capture long-range dependencies in time series, resulting in poor performance. This highlights the high dimensionality, strong temporal correlation, and complex feature boundaries of eye-tracking data, while also showcasing the advantages of MAAT-Net's cross-modal attention mechanism, Transformer encoding, and residual connections. 7) Five-fold cross-validation is used to evaluate generalization ability, while BT validation enhances statistical reliability and reduces the tendency to underestimate the model's true performance with small sample sizes. The fundamental reason for the performance difference between the two is that five-fold cross-validation is susceptible to random sampling fluctuations, while BT validation reduces errors through resampling and better reflects the model's true upper limit. MAAT-Net demonstrates superior performance in both validation methods, indicating that its advantage stems from its inherent model design rather than sampling randomness.
[0110] The validation results in tri-class and bi-classification tasks fully demonstrate the effectiveness and superiority of the adopted multimodal attention fusion mechanism in extracting discriminative eye-tracking features and achieving accurate autism spectrum disorder classification.
[0111]
[0112] Table 6. Results of 5-fold cross-validation and Bootstrap validation for 10 baseline models and the MAAT-NET model in the autism spectrum disorder grading and classification task. H. Ablation test: To delve into the impact of each functional module in the feature fusion component on the performance of MAAT-Net, a series of ablation variant models were designed for quantitative evaluation. The model is deconstructed into four key mechanisms: bidirectional cross-modal attention, Transformer encoder, adaptive modal attention fusion, and two-layer residual architecture. Based on these mechanisms, the following seven sets of experiments were constructed: V0 (Baseline): The baseline model. All high-level interactions are removed, and only the raw features of the three modalities are linearly projected and concatenated before being classified by an MLP.
[0113] V1 (w / o Bi-CMA): Removes bidirectional cross-modal attention. Modal features do not interact before entering the Transformer, validating the importance of early alignment between modalities.
[0114] V2 (w / o TE): Remove the Transformer encoder. Pool the attention-enhanced features directly to verify the effectiveness of paradigmatic sequence modeling in extracting deep features.
[0115] V3 (w / o AMF): Removed modality adaptive fusion. Weighted fusion was replaced with equal-weighted average fusion to verify the importance of the model automatically adjusting the weights of different modalities (trajectory, heatmap, numerical).
[0116] V4 (w / o Feature-Res): Removes residual connections between the original features and Transformer features, verifying the effectiveness of preserving original feature information in preventing feature oversmoothing.
[0117] V5 (w / o Classifier-Res): Removes shortcut connections in deep classifiers and verifies the contribution of the residual term to gradient flow and feature reuse.
[0118] V6 (Full Model): A complete solution that includes all improved modules.
[0119] Table 7 lists the quantitative results of the ablation experiment, and the corresponding confusion matrix is as follows: Figure 10 As shown. By comparing the performance of different variants, the following conclusions are drawn: 1. The necessity of cross-modal attention and adaptive fusion: Comparing V1 and V6 reveals that removing bidirectional cross-modal attention reduces accuracy to 0.67. This indicates that without explicit feature alignment and interaction, the semantic gap between heterogeneous modalities will hinder model convergence and may even cause performance to drop below the baseline model V0.
[0120] Comparing V3 and V6 reveals that removing the adaptive weight module causes the accuracy to drop to 0.70. This confirms that modality dependence varies significantly across different autism spectrum severity levels, and static averaging fusion cannot adapt to such dynamic changes. The core role of the Transformer encoder: The core function of the Transformer encoder: Comparing V2 and V6 reveals the most significant performance degradation: accuracy drops from 0.78 to 0.65, while AUC plummets to 0.57. This strongly suggests that eye tracking and heatmaps are inherently time-series data. Failure to model long-range dependencies will result in the loss of crucial diagnostic information.
[0121] 3. Contribution of residual connections to training stability: Removing feature-level residuals (V4) or classifier residuals (V5) both led to performance degradation (accuracy decreased to 0.72 and 0.69, respectively). In particular, the results for V5 show that preserving direct connections of the original fused features in deep networks can effectively alleviate the gradient vanishing problem and enhance the model's generalization ability on small sample data.
[0122] In summary, the complete model V6 achieved the best performance across all metrics: Accuracy 0.78, F1-Score 0.77, demonstrating a good synergistic effect among the modules.
[0123]
[0124] Table 7 shows the model ablation experiments. V0-V7 are model variants with different combinations of core modules, which respectively verify the effectiveness of bidirectional cross-modal attention, Transformer encoder, modal adaptive attention fusion, and dual residual connection. Performance metrics include accuracy, precision, recall, F1 score, and AUC, all calculated based on 5-fold Stratified Group K-Fold cross-validation. The results are expressed as "mean ± standard deviation".
[0125] This embodiment designed an eye-tracking experimental paradigm incorporating three types of visual stimuli, recruiting participants with varying degrees of autism spectrum disorder (ASD) to collect effective eye-tracking data. Deep learning techniques were employed for feature extraction and classification modeling. Eye-tracking heatmaps, trajectory maps, and fixation entropy complement each other in terms of static attention, dynamic fixation sequences, and quantitative exploration, collectively constructing a comprehensive visual attention representation system for children with ASD. This system holds promise as an objective and quantifiable biomarker.
[0126] This embodiment proposes a multimodal adaptive fusion model, MAAT-Net, integrating eye-tracking trajectory maps, heatmaps, and fixation entropy indices. It integrates a Transformer encoder utilizing bidirectional cross-modal attention mechanisms, residual adaptive fusion strategies, and modal adaptive weighting to achieve dynamic alignment and semantic complementarity between spatial features (heatmaps) and temporal features (trajectory maps), effectively mitigating information loss in deep networks and efficiently modeling long-range dependencies between three experimental paradigms. Experimental results validate the model's excellent performance, achieving an accuracy of 77.78% in the ASD three-level classification task. When the model is extended to ASD and TD binary classification tasks, the overall accuracy reaches 90.74%. This study not only provides an efficient, non-invasive, and objective technical tool for assessing ASD severity but also enriches the biomarker library for ASD diagnosis by mining heterogeneous eye-tracking patterns in patients with different severity levels of ASD. The proposed framework breaks the reliance on subjective clinical scales, lowers the assessment threshold for primary healthcare institutions, and provides data support for the development of personalized intervention programs. This embodiment proposes a deep learning-based autism spectrum disorder classification scheme that utilizes 3D integrated visual task eye-tracking data to improve classification accuracy. Compared with traditional multimodal fusion methods, this scheme has the following advantages: portable eye-tracking devices enable convenient and low-cost data acquisition; non-invasive and contactless operation is suitable for the characteristics of children with autism and does not cause stress; the complementarity of multimodal data eliminates complex preprocessing steps, thereby promoting automated assessment.
[0127] Experimental results validated the effectiveness of the framework and the applicability of the workflow, confirming that it provides an efficient and non-invasive approach for classifying the severity of autism spectrum disorders. In five-fold cross-validation of a three-class classification task, the model achieved an overall accuracy of 77.78% and an F1 score of 0.7769, with an accuracy of 81.48% in leave-one-out validation. The model performed best in the mild to moderate ASD group, with an F1 score of 0.8254 in five-fold cross-validation, indicating that the residual adaptive fusion mechanism effectively captured the typical eye-tracking features of this group.
[0128] Bidirectional cross-modal attentional heatmap analysis revealed a significant interaction between trajectory features and heatmap features, reflecting distinct attentional preferences across different paradigms. This mechanism overcomes the limitations of simple feature splicing by extracting temporal patterns of attentional shifts. These patterns coincide with visual search abnormalities in individuals with autism spectrum disorder, further validating the value of eye-tracking patterns as stable biomarkers and the model's potential to simulate expert diagnosis.
[0129] Despite its innovative approach to multimodal fusion, this study has limitations. For example, differences in sample size between different ASD severity groups may affect model performance. Current attention mechanisms only focus on paradigm-level interactions; future research should explore more refined dynamic alignment mechanisms. As non-invasive biomarkers for the diagnosis of autism spectrum disorders, eye-tracking features demonstrate broad application prospects and clinical translational potential. They not only address the shortcomings of traditional methods (such as subjectivity and insufficient standardization of quantitative indicators) but also open new avenues for early screening and intervention.
[0130] Future work will expand the use of multi-center, large-scale eye-tracking datasets to enhance clinical extrapolation capabilities. A regression neural network based on the CARS scale will be constructed to develop refined assessment models for dynamic monitoring and personalized evaluation of patients with autism spectrum disorder. Through multi-center clinical validation and technological iteration, the system will be integrated into clinical workflows to improve the accuracy and standardization of autism spectrum disorder diagnosis and treatment.
[0131] This embodiment proposes a multimodal adaptive fusion model called MAAT-Net, which integrates eye-tracking maps, heatmaps, and gaze entropy. By utilizing bidirectional cross-modal attention, residual adaptive fusion, and a modally adaptive weighted Transformer encoder, the model achieves dynamic alignment of spatial and temporal features with semantic complementarity. This mitigates depth information loss and models long-range dependencies in the paradigm. Experiments validate its superior performance, achieving an accuracy of 77.78% in a three-class classification task for autism spectrum disorder and 90.74% in a two-class classification task for ASD and typical developmental disorder. This study provides an efficient, non-invasive, and objective tool for assessing the severity of autism spectrum disorder. It establishes a digital biomarker system based on eye-tracking behavior, enriching the biological evidence for diagnosis, reducing reliance on subjective clinical scales, lowering the assessment threshold at the primary care level, and providing reliable data support for early screening, tiered assessment, and personalized intervention for ASD.
[0132] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the inventive concept of this application, and these all fall within the protection scope of this application.
Claims
1. A method for classifying autism spectrum disorders based on eye-tracking, characterized in that, Includes the following steps: S1: Acquire the target's eye movement information and generate a gaze heatmap and gaze point trajectory map; S2: Obtain the clinical numerical characteristics of the target based on the target's eye movement information; S3: The target gaze heatmap, gaze trajectory map and clinical numerical features are fused using the multimodal fusion model MAAT-Net; The multimodal fusion model MAAT-Net extracts image features through an image feature extraction module, achieves semantic alignment of the feature space through multimodal feature projection and spatial alignment, fuses the cross-modal attention output features from two directions into the final interactive features through a bidirectional cross-modal attention interaction module, achieves feature fusion based on the spatiotemporal sequence aggregation of Transformer, and outputs the results through a deep classifier. S4: The output of the multimodal fusion model MAAT-Net predicts the probability of the severity of autism spectrum disorder.
2. The autism spectrum classification method based on eye-tracking according to claim 1, characterized in that, In step S1, eye movement information is captured by an eye tracker after the target views the set image, and the eye movement information is stacked to form a gaze heatmap and a gaze point trajectory map.
3. The autism spectrum classification method based on eye-tracking according to claim 1, characterized in that, In step S2, the clinical numerical features include one or more of the following: coordinate variance, spatial entropy, static fixation entropy, saccade fixation entropy, effective fixation / saccade count, average fixation duration, and mean duration of deduplicated eye movement events.
4. The autism spectrum classification method based on eye-tracking according to claim 1, characterized in that, The multimodal fusion model MAAT-Net extracts image features through an image feature extraction module. The image feature extraction module uses the lightweight backbone network EfficientNet-Lite0 as the feature extractor, which includes a shallow feature coding layer, a core bottleneck layer, and a high-level feature aggregation layer. The shallow feature coding layer uses a standard 3×3 convolutional layer with a stride of 2, which includes a convolutional layer with 32 kernels, a batch normalization layer, and a ReLU6 activation function. The core bottleneck layer uses a 3×3 vertical convolutional layer with 32 groups to obtain Z1: (1) Then, the number of feature channels is expanded from 32 to 128 using a 1×1 convolution, as shown in the following formula: (2) The advanced feature aggregation layer increases the number of channels to 1280 dimensions through a 1×1 convolutional layer, and then uses an adaptive average pooling operator to compress the spatial dimension to 1×1, ultimately mapping the image modality into a fixed-length one-dimensional feature vector f. img ∈R 1280 ,in This represents the intermediate feature map processed by the lightweight convolutional module, and ; (3) (4) Image features from various paradigms are projected onto a unified hidden layer dimension d through a linear layer. model Furthermore, learnable positional encodings are superimposed to capture structured information between tasks. This process is represented as follows: (5) Among them, E modal This is the modality-enhanced feature vector after fusing location information; Linear(·) represents the linear projection layer, whose core function is to map the original high-dimensional image features to a unified hidden layer dimension d. model ;f img ∈R 1280 For image modal one-dimensional feature vectors; PE∈R 3×64 This is a learnable location encoding matrix, where the dimension 3 corresponds to the three experimental paradigms, and 64 is the hidden layer dimension d. model The value of is used to encode sequence position information of different paradigms.
5. The autism spectrum classification method based on eye-tracking according to claim 1, characterized in that, The multimodal fusion model MAAT-Net achieves semantic alignment of the feature space through multimodal feature projection and spatial alignment. Position encoding is achieved by fusing the projected features with element-wise addition. For the trajectory mode, the enhanced features of the i-th paradigm are: (9) Where ET(i) represents the trajectory enhancement feature vector of the i-th paradigm, which integrates both trajectory content and location information; Wt represents the parameters of the linear projection layer, used to map the original trajectory features to a unified latent dimension dmodel; f (i) traj PE(i) is the trajectory image feature vector extracted from the i-th experimental paradigm; PE(i) is the learnable position code corresponding to the i-th paradigm, used to encode its position information in the evaluation sequence.
6. The autism spectrum classification method based on eye-tracking according to claim 1, characterized in that, The multimodal fusion model MAAT-Net fuses cross-modal attention output features from two directions into the final interactive features through a bidirectional cross-modal attention interaction module. The bidirectional cross-modal attention interaction module employs two independent multi-head attention modules. In each module, 64-dimensional features are projected into two 32-dimensional subspaces to compute attention in parallel. Each attention head in a multi-head attention model is calculated as follows: (10) The features after interaction are represented as follows: (11) Among them, F inter This represents the fused feature vector after bidirectional cross-modal interaction; MultiHeadAttn(⋅) represents the multi-head attention mechanism operator; Q traj、 Q heat Trajectory or heatmap features are used as query vectors in cross-modal attention computation, respectively; K heat K traj The heatmap or trajectory map features are used as the key vector V in cross-modal attention computation. heat V traj The heatmap and trajectory map features are used as value vectors in the cross-modal attention computation. The ⊕ symbol represents the feature concatenation operation, which fuses the cross-modal attention output features from the two directions into the final interactive feature.
7. The autism spectrum classification method based on eye-tracking according to claim 1, characterized in that, The multimodal fusion model MAAT-Net achieves feature fusion based on the spatiotemporal sequence aggregation of Transformer and outputs the results through a deep classifier. Includes the following sub-steps: (1) Transformer sequence encoding: The input consists of multimodal features that have undergone interactive processing, and these features are then cascaded through formula (12) to form an enhanced feature sequence F. in ∈ R 9×dmodel : (12) A three-layer Transformer Encoder structure is adopted, and a multi-head self-attention mechanism is used to model long-range dependencies of heterogeneous features; each Transformer layer contains: Bullish Self-Attention Layer: (13) Feedforward network layer: (14) The feedforward network FFN(x) is: (15) (2) Learnable residual paths: The features of each mode are decomposed into trajectory features F. t Heatmap features F h sum numerical characteristics F n ; Perform element-wise mean pooling on all decomposed feature vectors within each mode to obtain the global feature g. t g h and g n The learnable residual weights α are activated using the sigmoid function; the residual enhancement features are calculated as follows: (16) in , , This is the global average of the original features; (3) Modality-based adaptive weighting: The input is the concatenated and enhanced modal features: (17) Then, a small fully connected network, Modal Attention, is used to dynamically assign weights w=[w] to the three modalities. t ,w h ,w n ], (18) The final fusion feature f fused for: (19) fusion vector f fused The data is fed into a deep nonlinear classifier consisting of three linear layers, each containing a GELU activation function, batch normalization, and Dropout for regularization; the residual projection layer at the end of the classifier stacks the pre-classified deep features with the original fused features through learnable weights. The Softmax function generates the predicted probability distribution, and end-to-end training optimizes the cross-entropy loss function to output the predicted probability of the severity of autism spectrum disorder.
8. An electronic terminal, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is invoked and executed by the processor, it implements the steps of the method as described in any one of claims 1-7.
9. A computer-readable medium, characterized in that, The computer-readable medium stores a computer program that, when executed by a computer, implements the steps of the method as described in any one of claims 1-7.