Intelligent early-stage autism screening system based on multi-modal fusion
The multimodal fusion-based intelligent early autism screening system integrates acoustic, genetic, eye-tracking, and behavioral characteristics, solving the problems of time-consuming screening and reliance on subjective experience in existing technologies. It achieves early, efficient, and accurate autism screening, and is suitable for community hospitals and kindergartens.
Patent Information
- Application Number
- CN202511043631.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-07
AI Technical Summary
The existing autism screening and diagnosis system suffers from problems such as long processing time, reliance on subjective experience, high equipment costs, unreliable test results, and lack of multimodal data integration and analysis, which makes early diagnosis difficult and affects the effectiveness of intervention.
The early intelligent screening system for autism employs multimodal fusion, which includes data acquisition and preprocessing, feature extraction, monomodal analysis, and multimodal fusion. It integrates acoustic features, genetic information, eye movement trajectories, and behavioral characteristics to generate a comprehensive risk assessment report.
It improves the accuracy and speed of screening, enabling early diagnosis 2-3 years earlier, reducing reliance on professional resources, and is suitable for community hospitals and kindergartens. It can screen 30-40 children within 5-10 minutes, with an accuracy of 95.7%, sensitivity of 94.1%, and specificity of 96.8%.
Smart Images

Figure CN120913843A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to an early autism intelligent screening system based on multi-modal fusion, and belongs to the field of medical auxiliary diagnosis. BACKGROUND
[0002] Autism spectrum disorder, a disease with a growing number of patients worldwide, is a complex type of neurodevelopmental disease. Children with autism generally have difficulty in interpersonal communication and developing interests. In China, the average age of children diagnosed with autism is 4 to 5 years old, which is much later than the optimal intervention age of 1 to 3 years old, directly affecting the rehabilitation effect.
[0003] The current screening and diagnosis system faces some bottlenecks, such as traditional scales such as M-CHAT and CARS that rely on parents' subjective reports, which have relatively long time consumption and have certain difficulties in capturing biological characteristic indicators. The single biomarker detection scheme, such as gene sequencing, eye movement pattern analysis, voiceprint feature recognition or behavior action monitoring, all face problems and obstacles such as lack of sensitivity, limited specificity, and difficulty in dealing with the highly heterogeneous phenotype of autism spectrum disorder. As for advanced detection methods including functional magnetic resonance and electroencephalogram signals, the high cost of equipment, the complicated operation process and the high professional threshold required for result interpretation restrict the popularization and application in clinical practice.
[0004] The present study clearly points out that the current health assessment system for infants aged 1 to 2 years has many limitations. Doctors rely mainly on subjective experience and lack objective quantitative standards. The reliability of single physiological indicator detection results is not sufficient. The shortage of professional doctors and equipment resources leads to long detection waiting time. Various detection technologies are independent of each other and lack integrated analysis of multi-modal data. Therefore, there is an urgent need to develop a new intelligent monitoring system. SUMMARY
[0005] The purpose of the present application is to solve the above-mentioned problems in the background art and provide an early autism intelligent screening system based on multi-modal fusion.
[0006] The technical scheme adopted by the present application to achieve the above-mentioned purpose is as follows:
[0007] The early autism intelligent screening system based on multi-modal fusion comprises a data acquisition and preprocessing module for synchronously acquiring multiple modal original data and performing preliminary noise reduction and standardization processing.
[0008] A feature extraction module is used to extract specialized features from the original data of each modality.
[0009] A single-modal analysis module is configured to analyze each modal feature and output an independent risk probability;
[0010] A multi-modal fusion module is configured to integrate the results of each modal and generate a final risk assessment;
[0011] A result output module is configured to generate a visual report to support clinical decision-making;
[0012] The data acquisition and preprocessing module, the feature extraction module, the single-modal analysis module, the multi-modal fusion module, and the result output module are electrically connected in sequence. Compared with the prior art, the present application has the following advantages:
[0013] 1. The core breakthrough of the present application is that it breaks through the limitations of single-modal data and effectively integrates four types of heterogeneous data, namely acoustic features, genetic information, eye movement trajectories, and behavioral characteristics, through a multi-modal fusion method to improve screening accuracy.
[0014] 2. The present application also has obvious advantages in the time dimension of early screening. It can effectively evaluate children aged 1 to 2 years, which is 2 to 3 years earlier than traditional diagnostic methods, thereby enabling timely intervention.
[0015] 3. The present application can effectively improve screening speed. Compared with traditional methods, which require 3 to 6 hours to complete a comprehensive evaluation, the present application only needs 5 to 10 minutes to generate a comprehensive report, and a device can complete the screening of 30 to 40 children per day.
[0016] 4. The present application enables non-professionals to operate after simple training. This system is suitable for community hospitals, kindergartens, and other diverse scenarios, greatly reducing the dependence on professional resources, especially in areas lacking professional child psychiatrists, allowing more people to benefit from early intervention. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is the overall architecture diagram of the autism early intelligent screening system based on multi-modal fusion of the present application;
[0018] Figure 2 is the four-modal feature extraction and processing flowchart of the autism early intelligent screening system based on multi-modal fusion of the present application;
[0019] Figure 3 is the voiceprint analysis SE-ResNet50 model structure diagram of the autism early intelligent screening system based on multi-modal fusion of the present application;
[0020] Figure 4is a gene analysis CNN-LSTM hybrid model structure diagram of the early intelligence screening system for autism based on multi-modal fusion of the application;
[0021] Figure 5 is a multi-modal soft voting fusion mechanism schematic diagram of the early intelligence screening system for autism based on multi-modal fusion of the application;
[0022] Figure 6 is a normal child and ASD child MFCC feature comparison diagram of the early intelligence screening system for autism based on multi-modal fusion of the application;
[0023] Figure 7 is a cloud edge end collaborative architecture diagram of the early intelligence screening system for autism based on multi-modal fusion of the application;
[0024] Figure 8 is a system workflow diagram of the early intelligence screening system for autism based on multi-modal fusion of the application. DETAILED DESCRIPTION
[0025] The technical solutions in the application will be described clearly and completely below in conjunction with the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0026] Terminology explanation:
[0027] ASD - Autism Spectrum Disorder (Autism Spectrum Disorder);
[0028] MFCC - Mel-frequency Cepstral Coefficients (Mel-frequency Cepstral Coefficients);
[0029] SE-ResNet - Squeeze-and-Excitation Residual Network (Squeeze-and-Excitation Residual Network); CNN - Convolutional Neural Network (Convolutional Neural Network);
[0030] LSTM - Long Short-Term Memory (Long Short-Term Memory);
[0031] RNA - Ribonucleic Acid (Ribonucleic Acid);
[0032] VAD - Voice Activity Detection
[0033] Specific implementation one: as shown in the embodiment, the early autism intelligent screening system based on multi-modal fusion is recorded, including a data acquisition and preprocessing module, which is used for synchronously acquiring a plurality of modal original data, and performing preliminary noise reduction and standardization processing; Figures 1-8
[0034] The feature extraction module is used for extracting professional features from the original data of each mode;
[0035] The single modal analysis module is used for analyzing the features of each mode and outputting independent risk probability;
[0036] The multi-modal fusion module is used for integrating the results of each mode to generate the final risk assessment;
[0037] The result output module is used for generating visual report to support clinical decision;
[0038] The data acquisition and preprocessing module, the feature extraction module, the single modal analysis module, the multi-modal fusion module and the result output module are electrically connected in sequence.
[0039] Please refer to Figure 1 , the data acquisition module uses integrated design, the data acquisition and preprocessing module includes a voiceprint unit for capturing child voice; the voiceprint acquisition function is achieved by arranging three high-sensitivity microphones with physical noise reduction in a ring and cooperating with an ESP32 audio development board,
[0040] The gene unit is used for acquiring genetic samples; the gene acquisition function acquires samples through a saliva collector and completes the basic processing with a rapid RNA extraction kit,
[0041] The eye movement unit is used for recording eye movement trajectory under visual stimulation; a binocular 850nm infrared camera and a near-infrared eye movement tracker based on corneal reflection principle are used to accurately capture eye movement,
[0042] The behavior unit is used for capturing limb action video; a high-definition camera with resolution of 1080p and frame rate of 30fps is used to collect limb action video, and the above devices are designed with low power consumption and transmit data to processing terminal or cloud server through Bluetooth or WiFi.
[0043] The data acquisition and preprocessing module output content includes noise reduction voice, RNA sample, eye movement trajectory and skeleton point video.
[0044] Please refer to Figure 2 , the feature extraction module covers a multi-modal feature processing procedure. In analyzing the sound features, the system first uses intelligent noise reduction technology to improve the signal-to-noise ratio by more than 12 dB, then extracts 13-dimensional MFCC coefficients and their first and second-order differences to form a 39-dimensional feature vector, then uses short-time Fourier transform (window length 25 ms, frame shift 10 ms) to generate a spectrogram, obtains rhythm features such as speech rate and pause frequency, uses the LPC method to obtain the first three formant frequencies and bandwidths, and calculates quality features such as harmonic-to-noise ratio, signal-to-noise ratio, jitter, and perturbation, wherein the MFCC extraction formula is:
[0045]
[0046] 0≤m≤M, is an FFT coefficient, is a mel filter bank.
[0047] The genetic feature extraction extracts total RNA from the saliva sample, performs gene expression analysis using the BrainSpan Atlas dataset, converts the RNA sequence into k-mer frequency features with k=8, and then selects a key feature subset through the Chi-square method.
[0048] In analyzing the eye movement features, the system tracks the eye movement by locating the corneal reflection point and the pupil center. The system analyzes the position of the child's gaze point and the path of the rapid eye movement. The system pays special attention to the child's looking at the face and the direction of the gaze, which are social information.
[0049] In analyzing the behavior features, the system extracts joint angles, motion trajectories, and spatiotemporal features from the video stream based on the MediaPipe skeleton point tracking technology, forms quantitative indicators such as joint motion range, and adopts a 3D-CNN combined with a time series residual network model. The 3D convolution layer extracts spatiotemporal features (joint angles, motion trajectories, etc.) from the skeleton point video, the time series residual module enhances the long-term pattern recognition of stereotyped actions, and the time series pooling layer and the classifier output the behavior risk.
[0050] The single-modal analysis module includes a voiceprint analysis submodule for outputting a voice abnormality risk probability; please refer to Figure 3 In the single-modal analysis module, the voiceprint analysis uses an SE-ResNet50 deep convolutional neural network model based on a channel attention mechanism, the input layer receives a 224x224x1 MFCC image, is processed through a convolutional layer group containing multiple convolutional layers, batch normalization, and activation functions, is output by a fully connected layer to output the autism risk probability, and is connected through an SE module (global average pooling is used to compress the feature map, two fully connected networks are used to learn the channel relationship, and channel weight recalibration is used to enhance the key channel) and a residual connection to alleviate gradient disappearance.
[0051] The genetic analysis submodule is configured to output genetic risk probability; see Figure 4 The eye movement analysis submodule is configured to output visual behavior risk. The eye movement analysis submodule adopts a time sequence attention network, inputs eye movement trajectories and gaze time sequences of children under visual stimulation, processes time sequence characteristics of data through a GRU layer, captures key gaze area (such as facial, gaze direction and other social information area) patterns through a spatial attention module, identifies visual behavior features of important time points through a time sequence attention module, and finally outputs visual behavior risk probability through a full connection layer.
[0052] The eye movement analysis submodule is configured to output visual behavior risk. The eye movement analysis submodule adopts a time sequence attention network, inputs eye movement trajectories and gaze time sequences of children under visual stimulation, processes time sequence characteristics of data through a GRU layer, captures key gaze area (such as facial, gaze direction and other social information area) patterns through a spatial attention module, identifies visual behavior features of important time points through a time sequence attention module, and finally outputs visual behavior risk probability through a full connection layer.
[0053] The behavior analysis submodule is configured to output behavior risk. The behavior analysis submodule adopts a model combining 3D-CNN and time sequence residual network, extracts spatial and temporal features such as joint angle, motion trajectory and action frequency from limb motion video based on MediaPipe skeleton point tracking technology, processes spatial and temporal features through a 3D convolution layer, enhances learning of long-term behavior patterns such as stereotyped repetitive actions through a time sequence residual module, aggregates time dimension information through a time sequence pooling layer, and finally outputs behavior risk probability through a full connection classifier.
[0054] The behavior analysis submodule can extract spatial and temporal features, and the spatial and temporal features include joint angle and motion trajectory.
[0055] The multi-modal fusion module includes feature level fusion: voiceprint, gene, eye movement and behavior four modal features are normalized through Z-score standardization, mapped to a shared feature space using a four-layer autoencoder, complementary features with mutual information > 0.05 are screened and spliced to form a unified representation; Figure 5 Decision level fusion: based on a soft voting mechanism, risk probabilities independently output by each mode are integrated, and weights are adjusted through a dynamic weight formula, as follows:
[0056]
[0057]
[0058] wherein data quality score in the range of 0-1; represent the initial weights of each modality before fusion;
[0059] The soft voting ensemble weighted probability calculation formula is as follows:
[0060]
[0061] The base weight in the soft voting ensemble weighted probability calculation formula is the voiceprint feature weight 0.35, the genetic feature weight 0.30, the eye movement feature weight 0.20, and the behavior feature weight 0.15, and the final probability calculation formula is:
[0062]
[0063] wherein is the voiceprint model prediction probability, is the genetic model prediction probability, is the behavior model prediction probability, is the eye movement model prediction probability, is the voiceprint weight, is the genetic weight, is the behavior weight, is the eye movement weight.
[0064] When the data is missing, the single modality is used directly and the confidence is reduced, and the weight is adjusted according to the quality score, and the performance of the double modality system reaches the accuracy 95.7%, the sensitivity 94.1%, the specificity 96.8%, and the ROC-AUC 0.978.
[0065] Attention guided fusion: using a time-space attention mechanism (an attention module belonging to the prior art category, not an independent model), by dynamically capturing the sequence dependence in the time dimension and the feature association in the spatial dimension of multi-modal data, the weight distribution of high-value features such as social gaze area and stereotyped action patterns is strengthened, and the robustness of the model in the sparse annotation and environmental interference scene is improved.
[0066] The present application adopts a multi-modal integration method to fuse each feature result, including:
[0067] 1. Hard voting: determine the final classification according to the sample multi-model voting result;
[0068] 2. Soft voting: weighted average according to the probability value output by each model;
[0069] 3. Attention voting: dynamically adjust the weight of each model according to the importance of different features.
[0070] Through multi-modal fusion, the system can comprehensively analyze the multi-dimensional characteristics of autism, improving the accuracy and reliability of screening.
[0071] In the decision-level fusion, the risk probabilities independently output by each modality are integrated based on a soft voting mechanism, including voiceprint model, genetic model, eye movement model, and behavior model.
[0072] The result output module includes a risk assessment engine: integrating multi-modal fusion results, generating a comprehensive risk probability value in the range of 0-1 through normalization processing, and dividing risk levels according to threshold values;
[0073] Report generator: dynamically generating structured reports, including basic information of children, test time, overall risk level and confidence, independent results of each modality and contribution analysis.
[0074] Visualization system: provides multi-dimensional chart display, intuitively presenting key feature differences.
[0075] Decision support unit: based on risk level, matching standardized intervention suggestions.
[0076] The risk level greater than 85% is high risk, i.e. suggesting immediate professional diagnosis; 60%-85% is medium risk, i.e. suggesting close observation and review; less than 60% is low risk, i.e. suggesting normal monitoring.
[0077] Please refer to the self Figure 6 , the feature visualization part displayed therein covers acoustic features such as MFCC image and voiceprint graph, gene feature graph corresponding to key risk gene expression profile, and behavior feature graph showing key behavior characteristics and standard deviation relationship; in the report, not only basic content such as basic information of children and test time, but also comprehensive evaluation of overall risk level and confidence, sub-modal analysis of independent results of each modality and contribution, and intervention suggestions matching corresponding risk level are included.
[0078] Please refer to Figure 7The system is realized by using cloud edge-end collaborative architecture to realize flexible deployment; among them, the edge end is responsible for completing data collection and preprocessing, lightweight feature extraction, simple local reasoning and data quality evaluation, while the cloud provides GPU / TPU high-performance computing resources for processing complex case reasoning, optimizing the model based on new data, and managing the historical database, and the terminal device provides an intuitive operation interface to visualize the results, and collects expert opinions through the feedback channel to support system iteration. In the process of realizing the voiceprint analysis module, the data set used contains speech samples of children aged 1-6 (including 500 ASD children and 1000 normal children), which come from multiple scenes such as quiet clinics and homes, and contain different types of natural speech and directional task speech, recorded at a sampling rate of 44.1kHz and a quantization precision of 16bit;
[0079] In terms of feature extraction, it covers MFCC (with 13-dimensional basis plus 13-dimensional first-order difference plus 13-dimensional second-order difference), voiceprint graph generated by 256-point FFT (window length set to 25ms and frame shift to 10ms), rhythm features, formant features and quality features. Please refer to Figure 3 The SE-ResNet50 model is composed of 4 convolutional layers (the output channel numbers of the 4 convolutional layers are 64, 32, 128, and 64 respectively, and a LeakyReLU activation function is connected after each convolutional layer, and a max pooling layer is connected after every two convolutional layers) and 3 fully connected layers (the output dimensions are 128, 64, and 1 respectively, and a Dropout layer is connected after each fully connected layer to prevent overfitting); when training the model, data augmentation methods such as time stretching and pitch shifting are used, the batch size is set to 32, the initial learning rate is set to 0.0001 and the cosine annealing scheduling strategy is used, the Adam optimizer is used with a weight decay of 1e-5, the binary cross-entropy loss with class weights is used, the training number is 50 rounds and the early stopping strategy is used (patience is set to 10), and the final performance evaluation reaches an accuracy of 93.8%, a sensitivity of 91.2%, a specificity of 95.3%, and a ROC-AUC of 0.962.
[0080] Please refer to Figure 8In the system application process, the early preparation includes calibrating the equipment, setting a quiet and moderate light environment, and familiarizing the child with the equipment to reduce tension; information registration enters the basic information such as name and age, and the developmental history is filled in by the parents, and the past assessment records are entered; multi-modal data acquisition has realized the voiceprint data acquisition of 3-5 minute natural speech and directional speech tasks, and the gene acquisition of 2-3 ml saliva samples, and the design also includes 2-3 minutes of eye movement data (showing social stimulus materials to record eye movement) and 3-5 minutes of behavior data (record free activity and directional task limb movement) acquisition. Data processing is first processed locally, and then features are extracted, quality is evaluated, and then deep analysis is performed by the cloud, and finally multi-modal results are integrated; after the results are generated, a report containing risk assessment is generated, which is confirmed by the doctor, explained to the parents and provided with intervention suggestions (high risk referral to specialized institutions, medium risk 1-3 month review, low risk monitoring according to normal milestones).
[0081] It is apparent to those skilled in the art that the application is not limited to the details of the foregoing exemplary embodiments, and that the application can be implemented in other embodiments without departing from the spirit or essential characteristics of the application. Therefore, the embodiments should be considered in all respects as illustrative and not restrictive, the scope of the application being defined by the appended claims rather than by the foregoing description, and it is intended that all changes that come within the meaning and range of equivalency of the claims are embraced therein. No reference signs in the claims should be considered as limiting the scope of the claims to the features to which the reference signs are attached.
[0082] Furthermore, it should be understood that although the present specification is described in terms of embodiments, not every embodiment according to the present specification need necessarily include every independent technical feature mentioned in the specification, and that reference to a particular feature does not necessarily preclude its being combinable with other features. The description herein of any embodiment according to the present specification is specifically made subject to the existence of alternative embodiments, and that the descriptions and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. An autism early intelligent screening system based on multi-modal fusion, characterized in that: The data acquisition and preprocessing module is used for synchronously acquiring various modal original data, and performing preliminary noise reduction and standardization processing. The feature extraction module is used for extracting professionalized features from the original data of each modality. The single modality analysis module is used for analyzing the features of each modality and outputting independent risk probabilities. The multi-modality fusion module is used for integrating the results of each modality and generating a final risk assessment. The result output module is used for generating a visual report to support clinical decision-making. The data acquisition and preprocessing module, the feature extraction module, the single modality analysis module, the multi-modality fusion module and the result output module are electrically connected in sequence.
2. The multi-modal fusion based early autism intelligence screening system according to claim 1, wherein: The data acquisition and preprocessing module includes a voiceprint unit for capturing child speech, a gene unit for obtaining genetic samples, an eye movement unit for recording eye movement trajectories under visual stimulation, and a behavior unit for capturing limb movement videos. The output content of the data acquisition and preprocessing module includes noise-reduced speech, RNA samples, eye movement trajectories and skeletal point videos. The single modality analysis module includes a voiceprint analysis submodule for outputting speech anomaly risk probability, a gene analysis submodule for outputting genetic risk probability, an eye movement analysis submodule for outputting visual behavior risk, and a behavior analysis submodule for outputting behavior risk. The behavior analysis submodule can extract spatiotemporal features, including joint angles and movement trajectories.
3. The multi-modal fusion based early autism intelligence screening system according to claim 2, wherein: The multi-modality fusion module includes feature-level fusion: normalizing the voiceprint, gene, eye movement and behavior features through Z-score standardization, mapping them to a shared feature space using a 4-layer autoencoder, selecting complementary features with mutual information > 0.05 for splicing to form a unified representation; Decision-level fusion: integrating the risk probabilities independently output by each modality based on a soft voting mechanism, adjusting the weights through a dynamic weight formula, as follows: Soft voting integrated weighted probability calculation formula: Attention-guided fusion: using a time-space attention mechanism to dynamically capture the sequence dependence in the time dimension and feature association in the spatial dimension of multi-modal data, strengthening the weight distribution of high-value features such as social gaze areas and stereotyped action patterns, and improving the robustness of the model in sparse labeling and environmental interference scenarios.
4. The multi-modal fusion based early autism intelligence screening system according to claim 3, wherein: In the decision-level fusion, the risk probabilities independently output by each modality are integrated based on a soft voting mechanism, including voiceprint model, gene model, eye movement model and behavior model.
5. The multi-modal fusion based early autism intelligence screening system according to claim 3, wherein: In the soft voting integrated weighted probability calculation formula, the base weights are voiceprint feature weight 0.35, gene feature weight 0.30, eye movement feature weight 0.20 and behavior feature weight 0.15, and the final probability calculation formula is as follows: The result output module includes a risk assessment engine that integrates the multi-modality fusion results, generates a comprehensive risk probability value in the range of 0-1 through normalization processing, and divides the risk level according to the threshold; wherein a data quality score in the range of 0-1; represent initial weights of each modality before fusion. A report generator dynamically generates a structured report containing child basic information, test time, overall risk level and confidence, independent results of each modality and contribution analysis; ; A visualization system provides multi-dimensional chart display to intuitively present key feature differences.
6. The multi-modal fusion based early autism intelligence screening system according to claim 5, wherein: 7. The multi-modal fusion based early autism intelligence screening system according to claim 5, wherein: wherein is a voiceprint model prediction probability, is a genetic model prediction probability, is a behavior model prediction probability, is an eye movement model prediction probability, is a voiceprint weight, is a genetic weight, is a behavior weight, is an eye movement weight.
8. The multi-modal fusion based early autism intelligence screening system according to claim 5, wherein: Decision support unit: match standardized intervention recommendations based on risk level.
9. The multi-modal fusion based early autism intelligence screening system according to claim 8, wherein: Risk level greater than 85% is high risk; 60-85% is medium risk; less than 60% is low risk.