Autism evaluation system and method based on multi-modal time sequence data fusion

By adopting the dual timing alignment method of dynamic time alignment and LSTM fine-tuning in the autism assessment system, the cross-modal data timing misalignment problem is solved, and the reliability and robustness of the system are significantly improved.

CN120199486APending Publication Date: 2025-06-24HEFEI UNIV OF TECH
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510270983.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The problem of timing dislocation of cross-modal data in existing autism assessment systems leads to low assessment reliability.

Method used

The dual timing alignment method of dynamic time regularization (DTW) and LSTM fine-tuning is used to perform timing alignment of multimodal data to eliminate timing misalignment caused by differences in acquisition frequency.

Benefits of technology

The millisecond-level synchronization accuracy is achieved, eliminating small-range time deviations caused by delays in acquisition equipment or differences in physiological responses, and improving the robustness and reliability of the autism assessment system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199486A_ABST
    Figure CN120199486A_ABST
Patent Text Reader

Abstract

The invention provides an autism assessment system and method based on multi-modal time sequence data fusion, and relates to the technical field of children autism spectrum disorder assessment calculation processing. The invention innovatively provides a dual time sequence alignment method based on dynamic time warping (DTW) and LSTM prediction, the problem of time sequence dislocation of cross-modal data (heart rate / eye movement / limb movement) caused by acquisition frequency difference is solved, millisecond-level synchronization precision is realized, the technical problem of cross-modal data time sequence dislocation in an existing autism assessment system is solved, and the accuracy of time sequence alignment of the cross-modal data in the autism assessment system is improved. The small-range time deviation caused by acquisition equipment delay or physiological response difference is eliminated, and the robustness and reliability of the autism evaluation system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of child autism spectrum disorder assessment calculation and processing, and particularly relates to an autism assessment system and method based on multi-modal time-series data fusion. Background Art

[0002] Autism spectrum disorder is a rather common developmental disorder. Patients often cannot express themselves normally and participate in social activities in the early stage, and usually repetitive and restrictive behavioral actions will also occur. Currently, in the field of medical health, it is generally believed that the "cause of autism has not been clearly identified and there is no specific drug for treatment", however, the timing and method of later rehabilitation intervention will greatly affect the prognosis effect. The earlier it is discovered and intervened, the more significant the intervention effect will be, and the corresponding prognosis will also be better.

[0003] Traditional autism assessment techniques mainly rely on traditional behavioral observations and questionnaires, and usually require experts to evaluate according to specific behavioral criteria. This method has a large degree of subjectivity, and the assessment results are easily affected by observer bias. The booming development of modern information technology and artificial intelligence has opened up a broad development space and brought new hope for the diagnosis, rehabilitation, assistance and learning in the field of autism. Especially the in-depth application of artificial intelligence in multiple key fields, such as face recognition, expression analysis, speech emotion analysis, gesture recognition and motion analysis, etc., through machine learning technology, can analyze and understand the characteristics of autism from a more comprehensive and in-depth perspective, providing strong support for the development of related work. Existing autism assessment systems generally obtain multi-modal data such as behavioral data, electroencephalogram data, expression data and eye movement data of the object to be diagnosed; input the multi-modal data into a pre-trained multi-modal diagnosis model, use the multi-modal diagnosis model to extract single-modal features from the multi-modal data, and perform feature fusion diagnosis on the extracted single-modal features to obtain the autism diagnosis result of the object to be diagnosed. This method of feature fusion and analysis diagnosis of multi-modal data can achieve feature complementarity of different modalities and improve the accuracy and credibility of autism diagnosis results.

[0004] However, the above-mentioned autism assessment system does not consider the problem of time-series misalignment caused by differences in acquisition frequencies of cross-modal data, resulting in low reliability of the autism assessment system. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the prior art, the present invention provides an autism assessment system and method based on multi-modal time-series data fusion, and solves the technical problem of time-series misalignment of cross-modal data in the existing autism assessment system.

[0007] (2) Technical Solution

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0009] In a first aspect, the present invention provides an autism assessment system based on multi-modal time-series data fusion, including:

[0010] A data acquisition module for collecting multi-modal data of the person to be evaluated;

[0011] A data preprocessing module for preprocessing each modality in the multi-modal data of the person to be evaluated;

[0012] A time-series alignment module for performing time-series alignment on each modality in the processed multi-modal data through a dual time-series alignment method of dynamic time warping for rough alignment and LSTM fine-tuning to obtain aligned multi-modal data;

[0013] A single-modal feature extraction module for extracting individual single-modal features in the aligned multi-modal data by designing independent feature encoders for each modality;

[0014] An evaluation module for processing the individual single-modal features of the person to be evaluated through a pre-constructed autism assessment model to obtain an evaluation result.

[0015] Preferably, the multi-modal data includes: heart rate data, limb movement data, facial emotion data, eye movement data, and audio data.

[0016] Preferably, the preprocessing includes: general preprocessing and modality-specific preprocessing, where the general preprocessing includes timestamp alignment and synchronization processing, and missing value processing;

[0017] Among them, the timestamp alignment and synchronization processing process is as follows:

[0018] Align the timestamps of each modality data to the same benchmark;

[0019] Resample and interpolate the multi-modal data with timestamps aligned to the same benchmark to obtain multi-modal time-series data with strictly aligned time dimensions;

[0020] The modality-specific preprocessing refers to performing different standardization processes on data of different modalities.

[0021] Preferably, the performing time-series alignment on each modality in the processed multi-modal data through a dual time-series alignment method of dynamic time warping for rough alignment and LSTM fine-tuning includes:

[0022] Assume that for modalities m and n, their corresponding time series is: A m =(am1 , a m2 , …, a mM ) is the time series of modality m, B n = (b n1 , b n2 , …, b nN ) is the time series of modality n, DTW recurrence formula:

[0023]

[0024] Where:

[0025] dist(a i , b j ) is the distance metric between modality m and n at the i-th and j-th time steps, using Euclidean distance; D(i, j) is the cumulative distance of the corresponding alignment path;

[0026] After DTW coarse alignment, by taking the coarsely aligned time series as the input of the LSTM, the LSTM model learns how to further align the time steps according to the long-term dependencies between modalities.

[0027] Preferably, the pre-constructed autism assessment model includes a cross-modal dynamic fusion unit, a spatio-temporal joint unit, a pooling layer, and an inference layer;

[0028] Wherein,

[0029] The cross-modal dynamic fusion unit is used to process each single-modal feature to obtain cross-modal fusion time-series features;

[0030] The spatio-temporal joint unit is used to extract spatial data features and time-series data features from the cross-modal fusion time-series features to obtain features containing global time context information and features containing local spatial associations between modalities;

[0031] The pooling layer is used to pool the features containing global time context information and the features containing local spatial associations between modalities respectively, and then fuse them to obtain spatio-temporal features;

[0032] The inference layer is used to process the spatio-temporal features and output an evaluation result.

[0033] Preferably, the pre-constructed autism assessment model adopts two-stage optimization training during the training process, and the process is as follows:

[0034] In the early stage, the modality weights are fixed and converge stably, and in the later stage, adaptive λ adjustment based on feature entropy is enabled, specifically:

[0035] In the initial stage of training, the modality weights are fixed to a constant λ0, that is:

[0036] λm λ(t) = λ0 for t ≤ t transition

[0037] where λ m (t) represents the weight coefficient of mode m at time step t; λ0 is a constant value used to fix the weight of the mode; t transition is the time step of the transition stage, a hyperparameter set at the beginning of training;

[0038] As the training progresses and enters the later stage, the mode weight λ m (t) is dynamically adjusted based on the feature entropy H m (t). For mode m at time step t, the entropy value H m (t) can be calculated by the following formula:

[0039]

[0040] where: p m (t,i) is the probability distribution of the i-th feature of mode m at time step t; H m (t) is the entropy value of mode m at time step t; the following formula is used to update the mode weight:

[0041]

[0042] where: H max is the maximum value of the feature entropies of all modes, used to normalize the entropy value.

[0043] In a second aspect, the present invention provides an autism assessment method based on multi-modal time-series data fusion, including:

[0044] S1. Collect multi-modal data of the person to be evaluated;

[0045] S2. Preprocess each mode in the multi-modal data of the person to be evaluated;

[0046] S3. Perform time-series alignment on each mode in the processed multi-modal data through a dual time-series alignment method of dynamic time warping for rough alignment and LSTM fine-tuning to obtain the aligned multi-modal data;

[0047] S4. Extract individual single-modal features in the aligned multi-modal data by designing independent feature encoders for each mode;

[0048] S5. Process the individual single-modal features of the person to be evaluated through a pre-constructed autism assessment model to obtain an assessment result.

[0049] In a third aspect, the present invention provides a computer-readable storage medium storing a computer program for autism assessment based on multimodal temporal data fusion, wherein the computer program causes a computer to execute the autism assessment method based on multimodal temporal data fusion as described above.

[0050] In a fourth aspect, the present invention provides an electronic device, comprising:

[0051] One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the autism assessment method based on multimodal temporal data fusion as described above.

[0052] (III) Advantageous Effects

[0053] The present invention provides an autism assessment system and method based on multimodal temporal data fusion. Compared with the prior art, it has the following advantageous effects:

[0054] The present invention innovatively proposes a dual temporal alignment method based on dynamic time warping (DTW) and LSTM prediction to solve the problem of temporal misalignment of cross-modal data (heart rate / eye movement / limb movement) caused by differences in acquisition frequencies, achieving millisecond-level synchronization accuracy, thereby solving the technical problem of temporal misalignment of cross-modal data in existing autism assessment systems, eliminating small-range time deviations caused by acquisition device delays or physiological response differences, and improving the robustness and reliability of the autism assessment system. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0056] Figure 1 It is a block diagram of an autism assessment method based on multimodal temporal data fusion according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0058] Embodiments of this application provide an autism assessment system and method based on multimodal time-series data fusion, which solve the technical problem of cross-modal data time-series misalignment in existing autism assessment systems, eliminate small-scale time deviations caused by acquisition device delays or physiological response differences, and improve the robustness and reliability of autism assessment systems.

[0059] The technical solutions in the embodiments of this application to solve the above technical problems are generally as follows:

[0060] The autism assessment system and method based on multimodal time-series data fusion in the embodiments of this invention combine multimodal data acquisition and processing. By fusing various modal data such as limb movements, eye movements, heart rate, facial expressions, and audio, and combining advanced spatio-temporal modeling and dynamic fusion technologies, a more accurate and comprehensive assessment system is constructed. This system can dynamically perform weighted fusion on different modal data when facing diverse behaviors and physiological signals, thereby improving the accuracy of early autism diagnosis, identifying individual differences, and providing personalized support for subsequent interventions. At the same time, by introducing an adaptive spatio-temporal feature modeling method, the embodiments of this invention can effectively overcome the sensitivity of existing methods to modal noise and time-series deviations, and further improve the robustness and reliability of the system.

[0061] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0062] Embodiments of this invention provide an autism assessment system based on multimodal time-series data fusion, which includes:

[0063] A data acquisition module for acquiring multimodal data of the person to be evaluated;

[0064] A data preprocessing module for preprocessing each modality in the multimodal data of the person to be evaluated;

[0065] A time-series alignment module for performing time-series alignment on each modality in the processed multimodal data through a dual time-series alignment method of dynamic time warping for rough alignment and LSTM fine-tuning to obtain aligned multimodal data;

[0066] A single-modal feature extraction module for extracting individual single-modal features in the aligned multimodal data by designing independent feature encoders for each modality;

[0067] An assessment module for processing the individual single-modal features of the person to be evaluated through a pre-constructed autism assessment model to obtain an assessment result.

[0068] The embodiment of the present invention innovatively proposes a dual temporal alignment method based on Dynamic Time Warping (DTW) and LSTM prediction to solve the problem of temporal misalignment of cross-modal data (heart rate / eye movement / limb movement) caused by differences in acquisition frequencies, achieving millisecond-level synchronization accuracy, thereby solving the technical problem of temporal misalignment of cross-modal data in existing autism assessment systems, eliminating small-scale time deviations caused by acquisition device delays or physiological response differences, and improving the robustness and reliability of the autism assessment system.

[0069] The following is a detailed description of each module:

[0070] Data acquisition module:

[0071] This module is used to collect multi-modal data of the person to be evaluated. In the embodiment of the present invention, the interview and video datasets of the person to be evaluated are collected. The multi-modal data in this dataset includes: heart rate data, limb movement data, facial emotion data, eye movement data, audio data, etc.

[0072] Data preprocessing module:

[0073] Preprocessing includes general preprocessing and cross-modal preprocessing. Among them, general preprocessing includes timestamp alignment and synchronization processing, and missing value processing.

[0074] The process of timestamp alignment and synchronization is as follows:

[0075] Align the timestamps of each modal data to the same benchmark (referenced by the video acquisition device time);

[0076] Resample and interpolate the multi-modal data with timestamps aligned to the same benchmark to obtain multi-modal temporal data with strictly aligned time dimensions. The specific operation is as follows: perform linear interpolation on the low-sampling-rate modality (such as heart rate: 1Hz) to match the high-sampling-rate modality (such as eye movement: 30Hz), and output a unified time step (such as 30Hz).

[0077] The processing method for missing values is as follows:

[0078] Short-term missing (<1 second): Linear interpolation or forward filling.

[0079] Long-term missing (generally caused by sensor failure): Set to zero or mark the mask (subsequent models ignore invalid time periods through the attention mechanism).

[0080] Outlier filtering: Eliminate significantly abnormal data points based on medical prior knowledge (such as heart rate range 60 - 180 bpm).

[0081] Cross-modal preprocessing refers to performing different processing on data of different modalities, specifically as follows:

[0082] Heart rate: Perform Z-score normalization on the heart rate value

[0083]

[0084] Wherein:

[0085] x is a certain value in the heart rate data; μ is the mean of the heart rate data; σ is the standard deviation of the heart rate data, and z is the normalized heart rate data.

[0086] Limb movement data: Perform Min-Max normalization on the acceleration value to the interval [-1, 1] to avoid dimensional differences.

[0087]

[0088] Wherein:

[0089] x is a certain value in the acceleration data; x min is the minimum value of this data set; x max is the maximum value of this data set; x′ is the normalized limb movement data.

[0090] Facial emotion data: In the embodiments of the present invention, the OpenFace 2.0 toolkit is used to extract the AU (ActionUnit) intensity value (such as AU12 represents the upturned corners of the mouth) and the head pose (deflection / pitch / tilt) of each frame in the video.

[0091] Downsampling: Reduce the video frame rate (such as 30fps) to 10Hz aligned with the time axis to reduce redundant calculations.

[0092] Temporal smoothing: Perform moving average (window size 5 frames) on the AU intensity sequence to suppress instantaneous expression jitter.

[0093]

[0094] Wherein:

[0095] AU′ t is the AU intensity value of the t-th frame after smoothing; AU i is the AU intensity value of the i-th frame; N is the size of the sliding window (for example, 5 frames).

[0096] Moving average is a common smoothing technique used to remove high-frequency noise in time-series data. Here, the sliding window size is N (for example, 5 frames), which means averaging the AU intensity of the current frame and its N-1 frames before and after, thereby suppressing instantaneous expression fluctuations. Through this operation, the instantaneous changes in facial expressions are smoothed, and a more stable emotional trend is retained.

[0097] The processing process of eye movement data is as follows:

[0098] Track cleaning: Remove invalid fixation points (duration < 100ms) caused by blinking in eye movement data.

[0099] Gaze point analysis: Use pyGaze to extract gaze points from the original video and calculate the gaze duration and spatial distribution density.

[0100] Spatial normalization: Normalize screen coordinates to the [0,1] range to adapt to devices with different resolutions.

[0101] The audio data is processed as follows:

[0102] Denoising: Use spectral subtraction to remove ambient noise from audio data.

[0103] The denoised audio data is framed, specifically: framed with a frame length of 25ms and a step size of 10ms, and MFCC (Mel-frequency cepstral coefficients) and rhythmic features (fundamental frequency, energy) are extracted.

[0104] Timing alignment module:

[0105] This module is used to perform timing alignment on each modality in the processed multimodal data through a dual timing alignment method of dynamic time warping (DTW) coarse alignment and LSTM fine-tuning to obtain aligned multimodal data. In an embodiment of the present invention, in order to establish a cross-modal spatiotemporal consistency representation, the timing deviation caused by the acquisition device and physiological response delay is eliminated through an adaptive transformation layer to achieve millisecond-level (±50ms) precise alignment. A dual timing alignment method of dynamic time warping (DTW) coarse alignment and LSTM fine-tuning is used.

[0106] When performing time series alignment, the embodiment of the present invention first uses DTW for coarse alignment. For each pair of modalities m and n (for example, heart rate and body movement data), DTW calculates the minimum distance between them. Assume that for modalities m and n, the corresponding time series is: A m =(a m1 ,a m2 ,…,a mM ) is the time series of modality m (e.g., heart rate or eye movement data). n =(b n1 ,b n2 ,…,b nN ) is the time series of modality n (e.g., body movements or audio data). DTW recursive formula:

[0107]

[0108] in:

[0109] dist(a i ,bj ) is the distance metric between modalities m and n at time steps i and j, using Euclidean distance; D(i, j) is the cumulative distance of the corresponding alignment path.

[0110] After DTW rough alignment, LSTM is used to further fine-tune the time alignment to precisely adjust the time deviation between modalities. The core of LSTM is to learn the long-term dependencies in time series data based on its internal gating mechanism.

[0111] For each modality m and n (e.g., heart rate and limb movement), the LSTM model can further adjust the rough alignment result through its state transfer and memory mechanism to obtain the aligned multi-modal data.

[0112] The core formula of LSTM

[0113] f t = σ(W f (h t-1 , x t + b f )(forget gate)

[0114] i t = σ(W i [h t-1 , x t + b i )(input gate)

[0115]

[0116] o t = σ(W o [h t-1 , x t + b o )(output gate)

[0117] h t = o t * tanh(C t )(compute output)

[0118] Where:

[0119] f t 、i t 、o t are the forget gate, input gate, and output gate of LSTM respectively.

[0120] C t is the cell state, h t is the output of LSTM, representing the time series representation of modality data.

[0121] x t is the input data at the current moment, W f 、Wi , W C , W o are the weights of the LSTM, and b f , b i , b C , b o are the biases.

[0122] By taking the coarsely aligned time series as the input of the LSTM, the LSTM model learns how to further align time steps according to the long-term dependencies between modalities. The LSTM refines the time alignment between different modalities by adjusting the output h t to eliminate small-scale time deviations caused by acquisition device delays or physiological response differences.

[0123] Single-modal feature extraction module:

[0124] This module is used to extract individual single-modal features from the aligned multi-modal data by designing independent feature encoders for each modality. In the embodiments of the present invention, independent feature encoders are designed for each modality to extract highly discriminative temporal-spatial features, retain modality-specific information, and provide high-quality inputs for subsequent cross-modal fusion. The following is a detailed description of each feature encoder:

[0125] (1) The encoder structure for heart rate data (physiological information modality) is 1D CNN + BiLSTM:

[0126] Among them, the 1D CNN layer is used to capture local heart rate fluctuation patterns (such as short-term features of heart rate variability HRV).

[0127] The BiLSTM layer is used to model long-term temporal dependencies (such as the rising trend of heart rate under continuous stress).

[0128] (2) The encoder structure for limb movement data (behavior modality) is a spatio-temporal graph convolutional network (ST-GCN). Among them, the graph structure is defined as follows: taking human joints as nodes and bone connections as edges, a spatio-temporal graph is constructed. ST-GCN can simultaneously capture the spatial correlation of joint movements (such as hand-elbow coordinated movements) and the temporal evolution law (such as repetitive swinging arm movements).

[0129] (3) The encoder structure for facial emotion data (visual modality) is 3D CNN + Transformer:

[0130] Among them, 3D CNN is used to extract short-term temporal-spatial features from video segments (such as instantaneous changes in the upward curvature of the mouth).

[0131] The Transformer encoder is used to model long-range expression dynamics (such as the emotional transition from calm to anxious).

[0132] (4) The encoder structure for eye movement data (behavioral modality) is a graph neural network (GNN) + Temporal ConvNet:

[0133] Among them, the graph construction method is: converting the fixation point sequence into a graph structure (nodes = fixation points, edges = saccade paths).

[0134] The GNN layer is used to model the spatial correlation between fixation points (such as the transfer pattern of visual attention).

[0135] Temporal ConvNet is used to capture the temporal pattern of saccade behavior (such as the frequency of rapid eye movements).

[0136] (5) The encoder structure for audio data (speech modality) is Wav2Vec 2.0 + lightweight LSTM:

[0137] Wav2Vec 2.0 is used to extract context-related features of speech content (without relying on speech text).

[0138] The LSTM layer is used to model the temporal changes of prosodic features (such as intonation fluctuations and speech rate changes).

[0139] It should be noted that in the embodiments of the present invention, each encoder needs to be pre-trained before use, and the training process of these encoders is prior art and will not be elaborated here.

[0140] Evaluation module:

[0141] In this module, the single-modal features of the person to be evaluated are processed through a pre-constructed autism evaluation model to obtain an evaluation result. Among them, the evaluation result refers to identifying whether the person to be evaluated has autism. In the specific implementation process, the probability of whether the person to be evaluated has autism is obtained.

[0142] The training process of the pre-constructed autism evaluation model is as follows:

[0143] Step 1: Collect interview and video viewing datasets of autistic patients and normal children. The data collected in this dataset includes limb movement data, eye movement data, heart rate, facial videos, and interview audio data.

[0144] Step 2: Preprocess each modal data in the dataset. Its preprocessing process is similar to the data processing process in the above data preprocessing module and will not be elaborated here. In the embodiments of the present invention, performing data preprocessing work before starting model training can significantly improve the robustness of the model to multi-modal noise and provide high-quality input for subsequent dynamic fusion and spatio-temporal modeling.

[0145] Step 3: Perform temporal alignment on each modality in the preprocessed dataset through a dual temporal alignment method of dynamic time warping (DTW) coarse alignment and LSTM fine-tuning to obtain an aligned dataset. The alignment process in the above temporal alignment module is similar during this temporal alignment process and will not be elaborated here.

[0146] Step 4: Extract the individual unimodal features in the aligned dataset by designing independent feature encoders for each modality to obtain a unimodal feature set. The feature extraction process in the above unimodal feature extraction module is similar during this unimodal feature extraction process and will not be elaborated here.

[0147] Step 5: Train an autism assessment network with the unimodal feature set to obtain an autism assessment model.

[0148] The pre-constructed autism assessment model includes a cross-modal dynamic fusion unit, a spatio-temporal joint unit, a pooling layer, and an inference layer. The structure of the autism assessment model will be described in detail below:

[0149] The cross-modal dynamic fusion unit includes an independent linear transformation layer, a similarity calculation layer, a dynamic weight adjustment layer, a context fusion layer, and a feature aggregation layer designed for unimodal features.

[0150] The independent linear transformation layer is as follows: The features of each modality generate query, key, and value vectors through an independent linear transformation layer to capture modality-specific representations.

[0151] For the features of modality m where T is the time step and d m is the feature dimension of modality m, the linear transformation is as follows:

[0152] Q m = X m W q , K m = X m W k , V m = X m W v

[0153] where:

[0154] is the query vector; is the key vector; is the value vector; is the learned linear transformation matrix; d m represents the dimension of the query vector, d v represents the value dimension; d k represents the dimension of the key vector.

[0155] The similarity calculation layer is used to calculate the dot - product similarity between the queries of each modality and the keys of other modalities, and apply a scaling factor and a mask (excluding its own modality). To measure the similarity between different modalities, the embodiments of the present invention calculate the dot - product similarity between the queries and the keys, and apply a scaling factor and a mask (excluding its own modality). Suppose we want to calculate the similarity between modality m and modality n, the dot - product similarity formula is:

[0156]

[0157] where: S(Q m ,K n ) is the similarity between the query of modality m and the key of modality n; is the scaling factor, which is used to prevent the dot - product value from being too large.

[0158] In the dynamic weight adjustment layer, the Softmax normalization is used for the similarity scores to obtain the attention weights, which reflect the correlation strength between different modalities. The Softmax formula is as follows:

[0159]

[0160] where: α mn (t) represents the attention weight of modality m to modality n at time step t. The normalization operation ensures that the sum of the weights is 1.

[0161] The attention weights obtained by Softmax can dynamically adjust the weighting coefficients of each modality feature, so as to enhance the representation of relevant modalities and suppress the influence of irrelevant modalities.

[0162] In the context fusion layer, according to the attention weights, the values of other modalities are weighted and summed to generate the context vector C m :

[0163]

[0164] where: is the context vector of modality m at time step t; α mn (t) is the attention weight obtained by Softmax; V n (t) is the value vector of modality n at time step t; the context vector C m (t) represents the information fusion of modality m with other modalities at time step t.

[0165] After calculating the context vector, the context vector is fused with the original features through a residual connection to enhance the feature representation. This residual connection method helps to preserve the information of the original features while enhancing the interaction between modalities.

[0166] Finally, in the feature aggregation layer, the updated features of each modality are concatenated to form the temporal features of cross-modal fusion.

[0167]

[0168] Where: M is the total number of modalities; is the fused feature after all modalities are processed by the context fusion layer; the finally obtained X fused is the fused temporal feature containing multi-modal information and serves as the input to the subsequent unit.

[0169] In the spatio-temporal joint unit, a heterogeneous spatio-temporal network is constructed, and special network modules are designed to process time and space information respectively. The spatial modeling network (ST-GCN) is used to extract spatial data, and the temporal modeling network (Transformer) is used to process temporal data. The features containing the local spatial correlations between modalities after being processed by the spatial modeling network (ST-GCN) and the original input of the temporal modeling network (Transformer) are weighted and fused as the final feature input of the Transformer.

[0170] In the spatio-temporal joint unit, a hierarchical spatio-temporal model is adopted, decoupling time modeling and space modeling into two complementary sub-networks, and joint reasoning is achieved through an interaction mechanism. This unit includes a spatio-temporal feature decoupling layer, a global time modeling layer, and a local space modeling layer, and spatio-temporal interaction mechanisms are designed in the global time modeling layer and the local space modeling layer, where:

[0171] The global time modeling layer is used for the temporal self-attention network based on Transformer to model the long-term dependencies across time steps (such as emotion persistence, action periodicity).

[0172] The local space modeling layer is used for the improved ST-GCN (Spatio-Temporal Graph Convolutional Network) to model the spatial topological correlations between multi-modal features (such as the physiological coordination between heart rate and limb movements).

[0173] The input feature of the spatio-temporal joint unit is the multi-modal feature after cross-modal dynamic fusion (i.e., the temporal feature of cross-modal fusion), with the shape of (Batch, Seq_Len, Num_Modalities × d_model).

[0174] Batch: Batch size.

[0175] Seq_Len: The length of the time series, that is, the number of time steps.

[0176] Num_Modalities×d_model: The feature dimension after cross-modal fusion. Num Modalities represents the number of modalities (such as heart rate, facial expression, speech, etc.), and d_model is the dimension of the features of each modality.

[0177] In the spatio-temporal feature decoupling layer, the fused features are split into a time series (Seq_Len dimension) and spatial modalities (Num_Modalities dimension), and are respectively input into the spatio-temporal sub-networks. It should be noted that the decoupling operation here only re-partitions the organization way of the features, without changing the dimension or content of the features themselves. This way allows the time modeling network (Transformer) to focus on the dependencies in the time series, while the spatial modeling network (ST-GCN) processes the spatial relationships between different modalities.

[0178] The functions of the global time modeling layer (Transformer layer) include: capturing the evolution of global behavior patterns across time steps (such as the gradual change of emotions within 10 consecutive seconds).

[0179] The input shape is: (Batch, Seq_Len, Num_Modalities×d_model). At each time step, the feature values containing all modalities are used to describe the state at that time point, for capturing the dependencies and dynamic changes in time.

[0180] In the embodiments of the present invention, its key design is as follows:

[0181] Use multi-head self-attention in the time dimension (Seq_Len) to calculate the correlation weights of different time steps.

[0182] The role of the multi-head self-attention mechanism is to calculate the correlation weights between each time step. Specifically:

[0183] The self-attention mechanism can capture the long-term dependencies in the sequence. For example, the change of emotions or actions may require considering the data of the previous few time steps to judge the state at the current time point.

[0184] The multi-head mechanism captures different levels of correlations by calculating multiple attention heads in parallel, enhancing the model's ability to understand the relationships between different time steps.

[0185] In the time dimension (Seq_Len), this means that the model can adjust the representation of the current time step based on the data of all past time steps.

[0186] Introduce time position encoding to retain the temporal order information of actions / emotions.

[0187] Since Transformer does not have a recursive or convolutional structure, it does not inherently possess the ability to capture the temporal order in a sequence. Therefore, positional encoding needs to be introduced to provide the position information for each time step.

[0188] The formula for the fixed positional encoding (Sinusoidal Positional Encoding) used in the embodiments of the present invention is as follows:

[0189]

[0190] where pos is the current position, i is the dimension index of the positional encoding, and S model is the dimension of the model.

[0191] This encoding enables the model to understand the progression in time. For example, how the emotion changes from the first time step to the tenth time step.

[0192] Output: Features containing global time context information, with the shape of (Batch, Seq_Len, Num_Modalities × d_model).

[0193] Enhance the non - linear expression ability through a feed - forward network.

[0194] Local spatial modeling (improved ST - GCN layer)

[0195] Function: Model the local dynamic associations of multi - modal features in the spatial dimension (such as a sudden increase in heart rate occurring simultaneously with limb stiffness).

[0196] Input shape: (Batch, Num_Modalities, Seq_Len × d_model)

[0197] Key design:

[0198] Modal graph construction: Consider multi - modal features (such as heart rate, eye movement, etc.) as graph nodes, and define the dynamic association strength between modalities through a learnable adjacency matrix.

[0199] Spatio - temporal convolution kernel: Use separable spatio - temporal convolution to capture the spatial relationship between modalities (1×1 graph convolution) and temporal dynamics (1D temporal convolution) respectively.

[0200] Dynamic edge weights: Adjust the weights of the adjacency matrix according to the cross - modal attention scores to achieve adaptive spatial topology learning.

[0201] Output: Features containing local spatial associations between modalities, with the shape of (Batch, Num_Modalities, Seq_Len × d_model).

[0202] The spatio-temporal interaction mechanism refers to two-way feature interaction, specifically including:

[0203] From time to space: The time context features output by the Transformer are used as the initial node features of the ST-GCN to inject global temporal information.

[0204] From space to time: The modal correlation features output by the ST-GCN are used as the position bias term of the Transformer to enhance local spatial perception.

[0205] In the embodiment of the present invention, the pooling layer includes a time dimension pooling block, a space dimension pooling block, and a spatio-temporal feature fusion block. The input of the time dimension pooling block is the output of the global time modeling layer (B×T×d_t), where B represents the Batch size, T represents the number of time steps, and d_t represents the time feature dimension. The pooling method is to first perform adaptive max pooling and then attention-guided average pooling. Its output is B×k×d_t, where k is the top k response values of the key frames dynamically selected for each time channel during adaptive max pooling.

[0206] The calculation formula of attention-guided average pooling is as follows:

[0207]

[0208] α i = softmax(W q h i )

[0209] where T is the total number of time steps, representing the length of the input sequence (for example, collecting 30 seconds of data, 30 frames per second, then T = 900); h i represents the input feature vector at the i-th time step, such as features like encoded heart rate, facial expression, etc.; W q represents a learnable weight matrix used to project the input features into the Query Space; α i represents the attention weight at the i-th time step, generated by Softmax normalization, reflecting the importance of this time step; represents the output feature vector after pooling, aggregating the weighted information of all time steps.

[0210] The input of the space dimension pooling block is the output of the local space modeling layer (B×M×d_s), where M represents the number of modalities and d_s represents the space feature dimension. In the space dimension pooling block, first perform weighted aggregation based on the modal importance score through multi-modal graph pooling, and then perform dynamic edge pooling: retain the strongly associated modal pairs with weights > 0.7 in the adjacency matrix. Its output shape: B×d_s

[0211] The calculation formula of multi-modal graph pooling is as follows:

[0212]

[0213] Among them, M represents the total number of modalities (such as the number of modalities like heart rate, eye movement, facial expression, etc.); h m represents the feature vector of the m-th modality, the representation after being processed by the encoder (such as heart rate feature, eye movement feature, etc.); ∑ n≠m h n represents the sum of the feature vectors of all other modalities except the m-th modality; [h m ; ∑ n≠m h n represents the joint feature vector obtained by concatenating h m and ∑ n≠m h m ; W g represents a learnable weight matrix, which is used to map the concatenated features to a scalar importance score; σ is an activation function (Sigmoid), which represents compressing the weight score into the interval [0, 1], indicating the modality importance; s m represents the importance score of the m-th modality, and the larger the value, the more significant the contribution of this modality to the current task; represents the fused feature vector after pooling, which weighted-aggregates the features of all modalities.

[0214] In the spatio-temporal feature fusion block, its fusion strategy is as follows:

[0215]

[0216] Among them, h fusion the fused feature, of F skip represents a skip connection, and a 1×1 convolution is used in the skip connection to align the dimensions, and W t , W s represent learnable weight matrices.

[0217] The output shape of the spatio-temporal feature fusion block is spatio-temporal features of B×(k×d_t + d_s).

[0218] In the inference layer, a custom ASD classification head is mainly used to process spatio-temporal features to obtain the final inference. This classification head consists of a multi-layer perceptron (MLP) and a Sigmoid activation function. That is, the structure of the inference layer is: MLP(512→256→128→1)+Sigmoid activation. The MLP contains three fully connected layers. Among them, the input layer has 512 neurons, corresponding to the final feature dimension of the data; the first hidden layer has 256 neurons, and the second hidden layer has 128 neurons. The output of the hidden layer is transformed through the RelU activation function to learn the complex non-linear relationship between input features; the output layer consists of 1 neuron, and the output is transformed into the autism probability value p∈[0,1] predicted by the final model through the Sigmoid activation function.

[0219] During the training process of the autism assessment model, two-stage optimization training is adopted, and the process is as follows:

[0220] In the early stage, the modal weights (λ coefficients) are fixed to converge stably, and in the later stage, adaptive λ adjustment based on feature entropy is enabled.

[0221] Early stage (fixed modal weights):

[0222] In the initial stage of training, the modal weight (λ coefficient) is fixed to a constant λ0, that is:

[0223] λ m (t)=λ0 for t≤t transition

[0224] Among them, λ m (t) represents the weight coefficient of modality m at time step t; λ0 is a constant value used to fix the weight of the modality and ensure the stability of the initial training; t transition is the time step of the transition stage, which is a hyperparameter set in the initial stage of training.

[0225] In this stage, the weights of all modalities are equal, and the model focuses on finding the associations of basic features without overfitting to noise.

[0226] In the later stage, adaptive λ adjustment based on feature entropy:

[0227] As the training progresses and enters the later stage, the modal weight λ m (t) will no longer be a fixed value, but will be dynamically adjusted based on the feature entropy H m (t). The feature entropy measures the complexity and uncertainty of each modality. For modality m, at time step t, the entropy value H m (t) can be calculated by the following formula:

[0228]

[0229] where: p m (t, i) is the probability distribution of the i-th feature of mode m at time step t; H m (t) is the entropy value of mode m at time step t, which reflects the uncertainty or complexity of the mode; based on the entropy value, the weight of the mode is adaptively adjusted. The goal of adjusting the mode is to assign a higher weight to the more informative mode, and the following formula is used to update the mode weight:

[0230]

[0231] where: H max is the maximum value of the entropies of all mode features, used to normalize the entropy value.

[0232] In this way, the mode with a higher feature entropy will have a smaller weight, and vice versa, helping the model to focus on processing the more informative mode.

[0233] In the embodiment of the present invention, a dedicated loss function is designed for the training of the evaluation model, specifically:

[0234] During the training process, in combination with the loss functions of each mode, adjustment is performed through the weight coefficients. Specifically, assuming there are M modes, the loss function of each mode at time step t is where m′ represents the m′-th mode. Then, the total loss function can be expressed as:

[0235]

[0236] where:

[0237] is the loss function of the m′-th mode at time step t; is the weight coefficient of the m′-th mode at the current time step t; by adjusting the weight coefficient the contribution of different modes to the total loss can be controlled, thereby adjusting the attention degree of the model to each mode during the training process.

[0238] The embodiment of the present invention provides an autism evaluation method based on multi-modal time-series data fusion, as Figure 1 shown, including:

[0239] S1. Collect multi-modal data of the person to be evaluated;

[0240] S2. Preprocess each mode in the multi-modal data of the person to be evaluated;

[0241] S3. Use a dual temporal alignment method of dynamic time warping for rough alignment and LSTM fine-tuning to perform temporal alignment on each modality in the processed multi-modal data, obtaining the aligned multi-modal data;

[0242] S4. Extract individual single-modal features in the aligned multi-modal data by designing independent feature encoders for each modality;

[0243] S5. Process the individual single-modal features of the person to be evaluated through a pre-constructed autism assessment model to obtain an assessment result.

[0244] It can be understood that the autism assessment method based on multi-modal temporal data fusion provided by the embodiments of the present invention corresponds to the above-mentioned autism assessment system based on multi-modal temporal data fusion. For the explanations, examples, beneficial effects, etc. of the relevant content, reference can be made to the corresponding content in the autism assessment system based on multi-modal temporal data fusion, which will not be elaborated here.

[0245] The embodiments of the present invention also provide a computer-readable storage medium, which stores a computer program for autism assessment based on multi-modal temporal data fusion, wherein the computer program enables a computer to execute the autism assessment method based on multi-modal temporal data fusion as described above.

[0246] The embodiments of the present invention also provide an electronic device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the autism assessment method based on multi-modal temporal data fusion as described above.

[0247] In summary, compared with the prior art, the following beneficial effects are achieved:

[0248] 1. The embodiments of the present invention innovatively propose a dual temporal alignment method based on dynamic time warping (DTW) and LSTM prediction, which solves the problem of temporal misalignment of cross-modal data (heart rate / eye movement / limb movement) caused by differences in acquisition frequencies, achieving millisecond-level synchronization accuracy, thereby solving the technical problem of temporal misalignment of cross-modal data in existing autism assessment systems, and realizing the elimination of small-range time deviations caused by acquisition device delays or physiological response differences, improving the robustness and reliability of the autism assessment system.

[0249] 2. The embodiments of the present invention can dynamically perform weighted fusion on different modality data when facing diverse behaviors and physiological signals, thereby improving the accuracy of early diagnosis of autism, identifying individual differences, and providing personalized support for subsequent interventions.

[0250] 3. By introducing an adaptive spatio-temporal feature modeling method, the embodiments of the present invention can overcome the sensitivity of traditional methods to modal noise and temporal deviation, and further improve the robustness and reliability of the system.

[0251] 4. In the autism assessment model, a spatio-temporal separable convolutional structure (ST-CNN) is used to decouple and process the spatial feature extraction (2DCNN) and the time-dependent modeling (1D CNN) in parallel. Compared with the traditional 3D CNN, it can effectively reduce the computational complexity.

[0252] 5. The embodiments of the present invention develop a trajectory graph neural network (TGNN) for eye movement data, converting the sequence of fixation points into a spatio-temporal graph structure (nodes = fixation points, edges = saccade paths), and effectively capturing abnormal fixation patterns.

[0253] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0254] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An autism assessment system based on multimodal time series data fusion, characterized in that: include: A data collection module, used to collect multimodal data of the person to be evaluated; A data preprocessing module, used for preprocessing each mode in the multimodal data of the person to be evaluated; A timing alignment module is used to perform timing alignment on each mode in the processed multimodal data through a dual timing alignment method of dynamic time warping coarse alignment and LSTM fine-tuning to obtain aligned multimodal data; A unimodal feature extraction module is used to extract each unimodal feature in the aligned multimodal data by designing an independent feature encoder for each modality; The evaluation module is used to process each unimodal feature of the person to be evaluated through a pre-built autism evaluation model to obtain an evaluation result.

2. The autism assessment system based on multimodal time series data fusion as claimed in claim 1, characterized in that: The multimodal data includes: heart rate data, body movement data, facial emotion data, eye movement data and audio data.

3. The autism assessment system based on multimodal time series data fusion according to claim 1, characterized in that: The preprocessing includes: general preprocessing and sub-modal preprocessing, wherein the general preprocessing includes timestamp alignment and synchronization processing, and missing value processing; The timestamp alignment and synchronization process is as follows: Align the timestamps of each modality’s data to the same reference; Resample and interpolate the multimodal data whose timestamps are aligned to the same benchmark to obtain multimodal time series data with strict alignment in the time dimension; The sub-modal preprocessing refers to performing different standardization processes on data of different modalities.

4. The autism assessment system based on multimodal time series data fusion according to any one of claims 1 to 3, characterized in that: The dual time alignment method of dynamic time warping rough alignment and LSTM fine-tuning performs time alignment on each mode in the processed multimodal data, including: Assume that for modes m and n, the corresponding time series are: A m =(a m1 ,a m2 ,…,a mM ) is the time series of mode m, B n =(b n1 ,b n2 ,…,b nN ) is the time series of mode n, and the DTW recursive formula is: in: dist(a i ,b j ) is the distance metric between modes m and n at the i-th and j-th time steps, using the Euclidean distance; D(i,j) is the cumulative distance of the corresponding alignment path; After DTW coarse alignment, by using the coarsely aligned time series as the input of LSTM, the LSTM model learns how to further align time steps according to the long-term dependencies between modalities.

5. The autism assessment system based on multimodal time series data fusion according to any one of claims 1 to 3, characterized in that: The pre-built autism assessment model includes a cross-modal dynamic fusion unit, a spatiotemporal joint unit, a pooling layer, and an inference layer; in, The cross-modal dynamic fusion unit is used to process each single-modal feature to obtain a cross-modal fused temporal feature; The spatiotemporal joint unit is used to extract spatial data features and temporal data features from the cross-modal fused temporal features, and obtain features containing global temporal context information and features containing local spatial associations between modalities; The pooling layer is used to pool the features containing global temporal context information and the features containing local spatial associations between modalities respectively, and then fuse them to obtain spatiotemporal features; The inference layer is used to process the spatiotemporal features and output evaluation results.

6. The autism assessment system based on multimodal time series data fusion according to any one of claims 1 to 3, characterized in that: The pre-built autism assessment model adopts a two-stage optimization training during the training process, and the process is as follows: In the early stage, the fixed modal weights converge stably, and in the later stage, the adaptive λ adjustment based on characteristic entropy is enabled, specifically: In the early stage of training, the modal weight is fixed to a constant λ0, that is: λ m (t)=λ0 for t≤t transition Among them, λ m (t) represents the weight coefficient of mode m at time step t; λ0 is a constant value used to fix the weight of the mode; t transition is the time step of the transition phase, which is a hyperparameter set at the beginning of training; As training progresses, entering the later stages, the modal weight λ m (t) Based on the characteristic entropy H m (t) is dynamically adjusted. For mode m, at time step t, the entropy value H m (t) can be calculated by the following formula: Where: p m (t,i) is the probability distribution of the i-th feature of mode m at time step t; H m (t) is the entropy value of mode m at time step t; the following formula is used to update the modal weight: Where: H max It is the maximum value of all modal feature entropy and is used to normalize the entropy value.

7. A method for assessing autism based on multimodal time series data fusion, characterized in that: include: S1. Collect multimodal data of the person to be evaluated; S2, preprocessing each modality in the multimodal data of the person to be evaluated; S3, performing time alignment on each mode in the processed multimodal data through a dual time alignment method of dynamic time warping coarse alignment and LSTM fine tuning to obtain aligned multimodal data; S4, extracting each unimodal feature from the aligned multimodal data by designing an independent feature encoder for each modality; S5. The single-modal features of the person to be evaluated are processed by a pre-constructed autism assessment model to obtain an evaluation result.

8. A computer-readable storage medium, characterized in that: It stores a computer program for autism assessment based on multimodal time series data fusion, wherein the computer program enables a computer to execute the autism assessment method based on multimodal time series data fusion as described in claim 7.

9. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include a method for executing the autism assessment method based on multimodal time series data fusion as described in claim 7.

Citation Information

Cited By

  • Multi-project joint detection method for urinary system

    CN120446233A

  • Pilot attention state data acquisition and processing method and device and related equipment

    CN120616532A

  • Bimodal signal processing method for intelligent screening of mild cognitive impairment

    CN120763485A

  • Intelligent old-age care data system based on big data and data processing method

    CN120766953A

  • Multi-modal data integrated analysis system for screening children with autism

    CN121256711A