Autism symptom identification system based on multivariate emotion path analysis

Through the multivariate emotion path analysis system, combined with multimodal data acquisition, feature extraction and deep learning orthogonal fusion, the problem of modal isolation analysis in individual emotions cognitive analysis of autistic individuals is solved, and high-precision autism recognition and dynamic analysis of emotions cognition are achieved.

CN120345898APending Publication Date: 2025-07-22HUAZHONG NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510333746.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing technology has isolated analysis of multimodal data in the emotional cognitive analysis of autistic individuals, which leads to difficult to capture the relationship between modals, separation of external manifestations and internal physiological states, and lack of comprehensive analysis of emotional processes, resulting in insufficient identification classification accuracy.

Method used

The autism symptom recognition system using a multi-modal emotional path analysis is used to collect skin electrical activity, eye movement, brain activity and facial expression data through the multi-modal data acquisition module, and extract corresponding features in combination with the data feature extraction module. The emotional cognitive mode analysis module establishes a path relationship model, and uses the deep learning module to perform orthogonal fusion in reverse order to obtain complementary information between multi-modals.

Benefits of technology

It has improved the recognition accuracy of autistic children, realized multi-dimensional comprehensive analysis, comprehensively revealed the emotional information integration process, accurately captured the dynamic changes in emotional cognition, quantified the emotional cognitive characteristics of the autistic group, and improved the reliability and scientific nature of the recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345898A_ABST
    Figure CN120345898A_ABST
Patent Text Reader

Abstract

The invention discloses an autism symptom recognition system based on multivariate emotion path analysis, and the system comprises a multi-modal data collection module which is used for collecting the multi-modal data of a participant, and the multi-modal data comprises skin electrical activity data, eye movement fixation data, electroencephalogram activity data and facial expression data; the data feature extraction module is used for extracting multi-modal features from the multi-modal data; the emotion cognition mode analysis module is used for establishing a path relation model of the multi-modal data and the emotion cognition stage variables, establishing a path relation model of the emotion cognition stage variables and the emotion cognition ability, and calculating to obtain a path coefficient representing each modal; and the deep learning module is used for carrying out orthogonal fusion and autism identification on the mapping features of the multiple modes in an inverted order mode according to the path coefficients corresponding to the multiple modes. According to the method, multi-modal data are systematically collected, orthogonal fusion is carried out on multi-modal mapping features in a reverse order mode, and autism identification can be better carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of pattern recognition and special education, and more specifically, relates to an autism symptom recognition system for parsing multiple emotional paths. Background Art

[0002] Autism (Autism Spectrum Disorder), also known as autism, is a pervasive developmental disorder caused by nervous system disorders. Autism mostly occurs in infants and young children, and its main symptoms are: social communication disorder, speech development disorder, narrow range of interests, and stereotyped repetitive behaviors. As one of the core characteristics of the social communication disorder of autistic children, they have significant difficulties in emotional perception, emotional understanding, and emotional expression, showing deficiencies in emotional cognitive ability, which affects their social interaction and daily life.

[0003] With the rapid development of artificial intelligence technology, more and more intelligent means have been applied to the emotional cognitive analysis of autistic individuals. These means integrate computer vision, speech recognition, natural language processing, and physiological signal processing and recognition technologies to comprehensively analyze the emotions of autistic individuals in multiple modalities, so as to complete the intelligent evaluation of the emotions of autistic individuals. By analyzing the emotional cognitive pattern, an atypical emotional cognitive development pattern of autistic individuals can be found, and early identification and screening of mental diseases such as autism can be achieved.

[0004] Although significant progress has been made in the research in the field of autistic emotional cognition in the existing multi-modal analysis framework, however, due to the complex diversity of the emotional cognitive process and the wide heterogeneity of the manifestations of autistic individuals, this will bring many problems and challenges to the research on the emotions of autistic individuals. The isolated analysis of multi-modal data makes it difficult to capture the mutual relationship between modalities; the separation of external behavior and internal physiological state ignores the complexity of the emotions of autistic children; the lack of comprehensive analysis of the emotional process leads to one-sided research conclusions. The distinction of the emotional cognitive characteristics of autistic individuals is not significant, resulting in insufficient recognition and classification accuracy. Emotion is a dynamically developing process, so it is difficult to comprehensively reflect the true face of the emotional cognition of autistic children by only analyzing a certain stage or a certain modality. Summary of the Invention

[0005] Aiming at the above defects or improvement requirements of the existing technology, the present invention provides an autism symptom recognition system for parsing multiple emotional paths, aiming to systematically collect multi-modal data and orthogonally fuse the mapping features of multiple modalities in a reverse order, so as to better perform autism recognition.

[0006] To achieve the above object, according to one aspect of the present invention, there is provided an autism symptom recognition system for multi - emotion path analysis, including:

[0007] A multi - modal data acquisition module for acquiring four - modal data of a participant. The four - modal data include skin conductance activity data, eye - movement fixation data, electroencephalogram activity data, and facial expression data;

[0008] A data feature extraction module for respectively extracting four - modal features from the four - modal data. The four - modal features include skin conductance arousal features, eye - movement fixation features, electroencephalogram activation features, and facial action unit features;

[0009] An emotion cognitive pattern analysis module for defining four emotion cognitive stage variables, which are in one - to - one correspondence with the four - modal features. Respectively establish path relationship models between the four - modal features and the corresponding emotion cognitive stage variables, establish a path relationship model between the four emotion cognitive stage variables and emotion cognitive ability, and calculate path coefficients representing the correlation relationship between the four emotion cognitive stage variables and path coefficients representing the causal relationship between the four emotion cognitive stage variables and emotion cognitive ability;

[0010] A deep - learning module for extracting four - modal deep features from the four - modal features, performing standardization and linear transformation mapping on the four - modal deep features to obtain mapping features of the four modalities, sorting the path coefficients of the causal relationship between the four emotion cognitive stage variables and emotion cognitive ability from small to large, and orthogonally fusing the corresponding mapping features of the four modalities in reverse order to obtain a fused feature, and outputting an autism recognition result according to the fused feature.

[0011] Further, the process of orthogonally fusing the corresponding mapping features of the four modalities in reverse order includes:

[0012] According to the order of the absolute values of the path coefficients of the causal relationship between the four emotion cognitive stage variables and emotion cognitive ability from large to small, the mapping features of the corresponding modalities are called "rank 1 feature", "rank 2 feature", "rank 3 feature", and "rank 4 feature";

[0013] Perform the first - layer orthogonal fusion. The first - layer orthogonal fusion calculation formula: Where is the first - layer orthogonal fusion feature, n6 is the feature dimension, is the feature orthogonal splicing, is the rank 1 feature, is the rank 2 feature, R 12 is The upper - triangular matrix after splicing;

[0014] Perform the second - layer orthogonal fusion. The formula for the second - layer orthogonal fusion is: Among them, is the second - layer orthogonal fusion feature, n7 is the feature dimension, is the feature orthogonal splicing, is the ranked 3 feature, R 123 is the upper triangular matrix after splicing;

[0015] Perform the third - layer orthogonal fusion. The formula for the third - layer orthogonal fusion is: Among them, is the third - layer orthogonal fusion feature, n8 is the feature dimension, is the feature orthogonal splicing, is the ranked 4 feature, R 1234 is the upper triangular matrix after splicing.

[0016] Furthermore, the multiple emotion - recognition stage variables include an emotion - arousal variable, an emotion - perception variable, an emotion - understanding variable, and an emotion - expression variable.

[0017] Furthermore, the path - relationship model between the four modal features and the corresponding emotion - recognition stage variables is:

[0018] y i =λ i η i +ε i , i ∈ [1,4], y i is the i - th modal feature, i ∈ [1,4]; λ i represents the path coefficient between the i - th modal feature and the i - th emotion - recognition stage variable, η i is the i - th emotion - recognition stage variable, ε i is the measurement error, Cov() represents the covariance, and Var() represents the variance.

[0019] The path - relationship model of the path - relationship model between the four emotion - recognition stage variables and the emotion - recognition ability includes:

[0020] η i =μ pij η i +ν qi x o +δ i , i ∈ [1,4], j ∈ [1,4], where x o is the emotion - recognition ability, η j is the j - th emotion - recognition stage variable, μ pij represents the relationship between the i - th emotion - recognition stage variable and the j - th emotion - recognition stage variable, ν qiDenote the path coefficient between the i-th emotional cognitive stage variable and the emotional cognitive ability as δ i is the model residual vector, μ pij =Cov(η i ,η j ), i, j ∈ [1, 4], where i ≠ j,

[0021] Furthermore, the four modal deep features are extracted from the four modal features: the four modal features are respectively input into their respective self-attention models to extract four modal deep features.

[0022] Furthermore, the steps for extracting the galvanic skin response arousal feature include:

[0023] Remove the noise in the skin conductance activity data through four-layer wavelet transform;

[0024] Perform normalization processing on the denoised skin conductance activity data;

[0025] Calculate the average conductance level and average amplitude of the normalized skin conductance activity data, and use the average conductance level and average amplitude as the galvanic skin response arousal feature.

[0026] Furthermore, the steps for extracting the eye movement fixation feature include:

[0027] In the eye movement fixation data, record the eye movement fixation coordinates of the participant at time t as (x(t3), y(t3)). Divide it into n3 non-overlapping rectangular regions of interest according to the participant's fixation area. The i1-th rectangular region of interest is denoted as AOI(i1), AOI(i1) = [(x(i1), y(i1)), ((x(i1)+X(i1), y(i1)+Y(i1)))]. Among them, i1 ∈ (1, n3), n3 is the number of regions of interest, (X(i1), Y(i1)) is the size of the region of interest, and (x(i1), y(i1)) is the upper left coordinate of the region of interest. If x(i1) ≤ x(t3) ≤ x(i1)+X(i) and y(i1) ≤ y(t3) ≤ y(i1)+Y(i1), then it is determined that the participant's eye movement coordinates at time t3 fall into the region of interest AOI(i1), otherwise it is determined that the participant's eye movement coordinates at time t3 do not fall into the region of interest AOI(i1);

[0028] Calculate the number of fixations of the participant on the region of interest. When the eye movement fixation coordinates (x(t3), y(t3)) are within the region of interest, it is recorded as 1, otherwise it is 0, then Among them, is the number of fixations on the i1-th rectangular region of interest, t3 is the total time length. Add the number of fixations on the n3 regions of interest to obtain the final fixation feature FC of the participant. The calculation formula is:

[0029] Calculate the fixation time of the participant's gaze on the region of interest. Set each fixation time to a fixed duration Δt. Then, the fixation duration for the i1-th region of interest is: Among them, is the fixation time for the i1-th region of interest, t4 is the total time length. The sum of the fixation times for n3 regions of interest is the final fixation time feature FD of the participant. The calculation formula is:

[0030] Calculate the fixation range of the participant's gaze on the region of interest. The fixation range for the i1-th region of interest is Among them, and are the fixation ranges on the horizontal x and vertical y within the i1-th region of interest respectively. min(x(t5)|(x(t5),y(t5))) is the minimum value of the fixation coordinates at time t5, and max(x(t5)|(x(t5),y(t5))) is the maximum value of the fixation coordinates at time t5. The sum of the saccade ranges for n3 regions of interest is averaged to obtain the final fixation range feature FA of the participant. The calculation formula is:

[0031] Calculate the saccade amplitude of the participant's gaze on the region of interest. Then, the saccade amplitude for the i1-th region of interest is represented by where Among them, and are the coordinates of two adjacent fixation points. The sum of the saccade amplitudes for n3 regions of interest is the final saccade amplitude feature of the participant. The calculation formula is:

[0032] Furthermore, the steps for extracting electroencephalogram activation features include:

[0033] The electroencephalogram activity data is transformed from the time domain to the frequency domain through Fourier transform. The electroencephalogram activity data in the frequency domain is divided into 5 frequency bands according to the frequency bands to which they belong. The power spectral density of each frequency band is statistically calculated, and the power spectral densities of all frequency bands are used as the electroencephalogram activation features.

[0034] Furthermore, the steps for extracting facial action unit features include:

[0035] Extract facial action units from the facial expression data, record the frequencies and average intensities at which the facial action units appear, and use the frequencies and average intensities at which the facial action units appear as the facial action unit features.

[0036] Furthermore, a system for recognizing autism symptoms based on a multi - dimensional emotion path analysis further includes:

[0037] The paradigm material switching module is used to switch and display different emotion induction segments;

[0038] The data synchronization module is used to align the collected skin conductance data, eye movement fixation data, electroencephalogram data, and facial expression data based on timestamps before extracting four modal features;

[0039] The personal profile module is used to generate a personalized emotion recognition analysis report for each participant. The emotion recognition analysis report includes the autism recognition result of the participant based on multi-modal features, the path relationship model between the four modal features of the participant and the corresponding emotion recognition stage variables, and the path relationship model between the four emotion recognition stage variables of the participant and the emotion recognition ability.

[0040] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention has the following beneficial effects:

[0041] The present invention has the following advantages:

[0042] (1) It proposes to intelligently identify autism and its emotion recognition mode from multi-modal perspectives such as skin conductance, eye movement fixation, electroencephalogram, and facial expression, and based on the orthogonal fusion method of emotion path relationship, complementary information between modalities is obtained, improving the recognition accuracy of autistic children.

[0043] (2) It proposes a multi-dimensional comprehensive analysis strategy. According to different emotion expression methods, multi-modal data is divided into implicit states and explicit behaviors for joint analysis, more comprehensively revealing the emotion information integration process of participants, making emotion recognition analysis more reliable and scientific.

[0044] (3) It proposes a multi-stage emotion recognition analysis process to gradually analyze the emotion recognition stage, enabling researchers to more comprehensively and accurately capture the dynamic changes in the emotion recognition process of participants.

[0045] (4) It proposes to construct a path relationship model between different emotion recognition variables to quantify the emotion recognition characteristics of the autism group, thereby discovering the emotion recognition mode of the autism group. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a framework diagram of an autism symptom recognition system for multi-emotion path analysis according to an embodiment of the present invention;

[0047] Figure 2 is a schematic diagram of the control module of an autism symptom recognition system for multi-emotion path analysis according to an embodiment of the present invention;

[0048] Figure 3It is a schematic diagram of the path relationship model of variables in different emotional cognition stages in the embodiments of the present invention;

[0049] Figure 4 It is a schematic diagram of a reverse multi-modal deep fusion network based on the path relationship level in the embodiments of the present invention. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0051] In the description of the embodiments of the present application, the meaning of the term "a plurality" is two or more.

[0052] In the embodiments of the present invention, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or equipment.

[0053] In the embodiments of the present invention, the naming or numbering of steps does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can be changed in the execution order according to the technical objectives to be achieved, as long as the same or similar technical effects can be achieved.

[0054] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0055] As Figure 1 shown, a system for identifying autism symptoms by parsing multiple emotional paths in the embodiments of the present invention can comprehensively utilize the emotional cognition mode of autism and the results of multi-modal fusion data to perform intelligent screening and identification of autism, which can greatly improve the accuracy of identification. Specifically, it includes:

[0056] A multimodal data acquisition module for collecting four modal data of participants, and the four modal data include skin conductance activity data, eye movement fixation data, electroencephalogram activity data, and facial expression data.

[0057] A data feature extraction module for respectively extracting four modal features from the four modal data, and the four modal features include skin conductance arousal features, eye movement fixation features, electroencephalogram activation features, and facial action unit features.

[0058] An emotion recognition pattern analysis module for establishing a path relationship model between the four modal data and multiple emotion recognition stage variables, calculating the relationship between the four modal data and multiple emotion recognition stage variables, establishing a path relationship model between multiple emotion recognition stage variables and four modal emotion recognition abilities, calculating the relationship between multiple emotion recognition stage variables and four modal emotion recognition abilities, the path coefficients representing the correlation relationship between multiple emotion recognition stage variables, and the path coefficients representing the causal relationship between each modal emotion recognition stage variable and four modal emotion recognition abilities.

[0059] A deep learning module for extracting four modal deep features from the four modal features, performing standardization and linear transformation mapping on the four modal deep features to obtain four modal mapping features, sorting the path coefficients corresponding to the four modalities from small to large, and orthogonally fusing the four modal mapping features in reverse order to obtain a fused feature, and outputting an autism recognition result according to the fused feature.

[0060] Further, as Figure 2 shown, an autism symptom recognition system for multi - emotion path analysis according to an embodiment of the present invention may further include a multimodal scenario module, a control module, a paradigm material switching module, a multimodal data storage module, a data synchronization module, a result output module, and a personal profile module.

[0061] (1) A multimodal scenario module for integrally integrating multimodal acquisition devices required for experimental tasks.

[0062] (2) A control module for controlling all other modules.

[0063] (3) A paradigm material switching module for switching different emotion - inducing segments according to different requirements during the experiment and real - time feedback of the emotion - segment presentation screen to researchers; since it is necessary to collect multimodal data to comprehensively analyze the emotion recognition ability of participants, it is necessary to identify the starting point of time in the emotion - inducing segment to obtain relevant feature information of the multimodal data.

[0064] (4) Data synchronization module, which is used to align the collected multimodal data (skin conductance data, eye gaze data, electroencephalogram data, and facial expression data) based on timestamps, facilitating subsequent analysis of relevant emotional features.

[0065] (5) Multimodal data storage module, which is used to organize the collected multimodal data, establish a personalized file, and implement one file per person.

[0066] (6) Result output module, which is used to record and visually present the multimodal feature recognition results output by the deep learning module, so as to facilitate researchers to interpret and identify the recognition results of autistic children.

[0067] (7) Personal file module, which is used to generate a personalized emotional cognition analysis report for each participant, including the recognition effect of personal multimodal features, the analysis of the performance patterns in each emotional cognition stage, the generation map of the emotional cognition path, and the summary of autism risk-related features.

[0068] The working principles of each module are described in detail below.

[0069] The working principle of the multimodal data acquisition module is as follows.

[0070] The multimodal data collected by the multimodal data acquisition module includes implicit state data and explicit behavior data.

[0071] Implicit states include skin conductance data and electroencephalogram data, which are used to collect the internal physiological changes of participants under emotional induction.

[0072] Furthermore, the skin conductance data can be collected using one of Empatica E4 and Empatica Embrace2.

[0073] Furthermore, the electroencephalogram data can be collected using one of Emotiv Epoc+, BrainLink Pro, and BrainLink Lite.

[0074] Explicit behavior data includes eye gaze data and facial expression data, which are used to collect the non-physiological changes of the external behaviors of participants under emotional induction.

[0075] Furthermore, the eye gaze data can be collected using one of Tobii Pro Fusion, Tobii EyeTracker4, and Tobii Eye Tracker5.

[0076] Furthermore, the facial expression data can be collected using one of Logitech Brio 500 and Hikvision D5ACAM100D.

[0077] The working principle of the data feature extraction module is as follows.

[0078] The data feature extraction module is used for the preprocessing of multi-modal signals collected by participants during the experiment, including skin conductance activity data, eye movement fixation data, electroencephalogram activity data, and facial Action Units (AUs); and for extracting features of each modality according to emotion information, where the features of each modality include skin conductance arousal features, eye movement fixation features, electroencephalogram activation features, and facial Action Unit (AU) features.

[0079] Furthermore, the data feature extraction module includes: a skin conductance activity data processing sub-module, an eye movement fixation data processing sub-module, an electroencephalogram activity data processing sub-module, and a facial expression processing sub-module.

[0080] Furthermore, the skin conductance activity data processing sub-module is used to quantify the skin conductance arousal level of the participant. The collected skin conductance activity data is processed through four-layer wavelet transform to remove the noise influence in the skin conductance activity data. The wavelet transform is calculated by convolution operation and upsampling and downsampling. The wavelet convolution operation where F eda (n1) is the input original skin conductance activity data, is the wavelet basis function, and C(k1) is the result of the convolution operation. The sampling calculation result is i1 represents the number of layers of the wavelet transform, and t1 is the time length. In one embodiment, the value of t1 is the time length of 1 s, that is, the first layer is the second layer is A1(t1) = A2(t1) + B2(t1), the third layer is A2(t1) = A3(t1) + B3(t1), and the fourth layer is A3(t1) = A4(t1) + B4(t1).

[0081] Furthermore, the individual differences in the skin conductance activity data are eliminated through min-max data normalization processing, so that the value range of the skin conductance arousal level is [0, 1]. Specifically, for the skin conductance features of each participant where, F eda Max and F eda Min are the maximum and minimum values in the skin conductance activity data, respectively.

[0082] Furthermore, calculate the skin conductance level of the skin conductance activity data: where t2 is the time length, is the conductance level value at the j1-th second; average amplitude feature: where, n2 is the number of amplitudes, is the j2-th amplitude value, is the minimum value adjacent to the j2-th amplitude, representing the electrodermal arousal level.

[0083] Furthermore, the eye movement fixation data processing sub-module is used to quantify the participant's eye movement fixation level. The collected eye movement tracking data is used to determine the participant's eye movement fixation coordinates: (x(t3), y(t3)) through eye tracking technology. According to the participant's fixation area, it is divided into n3 non-overlapping rectangular regions of interest. The i1-th rectangular region of interest is denoted as AOI(i1), and AOI(i1) = [(x(i1), y(i1)), ((x(i1)+X(i1), y(i1)+Y(i1)))], where i1 ∈ (1, n3) is the region of interest, (X(i1), Y(i1)) is the size of the region of interest, and (x(i1), y(i1)) is the upper left coordinate of the region of interest. If x(i1) ≤ x(t3) ≤ x(i1)+X(i) and y(i1) ≤ y(t3) ≤ y(i1)+Y(i1), it is determined that the participant's eye movement coordinates at time t3 fall into the region of interest AOI(i1); otherwise, it is determined that the participant's eye movement coordinates at time t3 do not fall into the region of interest AOI(i1).

[0084] Furthermore, it is statistically determined whether the participant's fixation point falls into the fixation region of interest. When the eye movement fixation coordinates (x(t3), y(t3)) are within the region of interest, it is recorded as 1; otherwise, it is 0. The specific mathematical formula is: Then the number of fixations in the i1-th region of interest is: where is the number of fixations in the i1-th rectangular region of interest, and t3 is the total time length. When calculating the number of fixations for multiple regions of interest, the sum of the multiple numbers of fixations is the final fixation number feature of the participant. The calculation formula is: n3 is the number of regions of interest.

[0085] The fixation time is expressed as the time the participant stays in the region of interest. Each fixation time is set to a fixed duration Δt. Then the fixation duration in the i1-th region of interest is: where is the fixation time in the i1-th region of interest, and t4 is the total time length. When calculating the fixation time for multiple regions of interest, the sum of the multiple fixation times is the final fixation time feature of the participant. The calculation formula is: n3 is the number of regions of interest.

[0086] The fixation range is the spatial distribution of the participant in the region of interest, which is represented by the extreme value coordinates in the region of interest

[0087] where and They are the fixation ranges on the horizontal x and vertical y in the i1th region of interest respectively. min(x(t5)|(x(t5), y(t5))) is the minimum value of the fixation coordinates at time t5, and max(x(t5)|(x(t5), y(t5))) is the maximum value of the fixation coordinates at time t5. When calculating the fixation range for multiple regions of interest, the sum of multiple saccade ranges is averaged to obtain the final fixation range feature of the participant. The calculation formula is: n3 is the number of regions of interest.

[0088] The saccade amplitude is the moving distance of the participant from one fixation point to another. It is usually expressed in degrees. Among them, and are the coordinates of two adjacent fixation points. When calculating the saccade amplitude for multiple regions of interest, the sum of multiple saccade amplitudes is used as the final saccade amplitude feature of the participant. The calculation formula is: n3 is the number of regions of interest.

[0089] Furthermore, the electroencephalogram (EEG) activity data processing sub-module is used to quantify the EEG activation level of the participant. The received EEG activity data is preprocessed, such as filtering and noise reduction, through the EEGLAB toolbox. According to the relationship between the brain state and frequency characteristics, the EEG activity data is divided into five frequency bands (Delta band: 1 - 4 Hz, Theta band: 4 - 7 Hz, Alpha band: 7 - 13 Hz, Beta band: 13 - 31 Hz, and Gamma band: 31 - 50 Hz). The discrete short-time Fourier transform algorithm is used to project the time-series EEG activity data F eeg (u) onto the frequency domain. The calculation formula for the frequency-domain signal F(t6, f) is:

[0090]

[0091] Among them, f represents the frequency, T1(u - t6) is the time window; F eeg (t6, f) is the energy distribution of the source signal F eeg (u) after Fourier transform. As the window function moves along the time axis, a series of frequency-domain signals changing with time are obtained. Then, the power spectral density features PSD(i3, j3) in the five frequency bands are calculated. The calculation formula is:

[0092]

[0093] Among them, f represents the frequency, P(t6, f) represents the power spectral density, i3 (i3 = 1, 2…, n4) represents the serial numbers of all electrode channels, j3 (j3 = 1, 2…, 5) are the five EEG frequency bands, is the length of the time-frequency diagram in the time dimension. is the frequency change range of the frequency band.

[0094] Furthermore, the facial expression processing sub-module is used to quantify the level of the participant's Facial Action Unit (AU). For the received facial expression image, the face region of the participant is determined through a face detection algorithm, more accurate facial action units are extracted through the OpenFace detection algorithm, and the AU frequency characteristics that appear in different emotion inductions of the facial AU are recorded Among them, is the occurrence times of the i4th AU, and n5 is the number of action units. The average intensity feature of AU Among them, is the intensity of the i4th AU, and n5 is the number of action units. Both of them jointly reflect the dynamic changes and emotional states of facial expressions.

[0095] The working principle of the emotion cognitive pattern analysis module is as follows.

[0096] The emotion cognitive pattern analysis module is used to analyze the internal connection and mutual influence between the cognitive variables of the participant and the performance patterns in each emotion cognitive stage during the experiment, so as to ensure the comprehensiveness of emotion cognitive analysis, as Figure 3 shown;

[0097] Furthermore, four emotion cognitive stage variables are defined, and the four emotion cognitive stage variables correspond one-to-one with four modal features. According to the emotion construction process, it is subdivided into four key stages: emotion arousal, emotion perception, emotion understanding, and emotion expression. The emotion arousal variable corresponds to the skin conductance arousal feature, the emotion perception variable corresponds to the eye movement fixation feature, the emotion understanding variable corresponds to the electroencephalogram activation feature, and the emotion expression variable corresponds to the facial action unit feature. Emotion arousal refers to the intensity of the physiological or psychological reaction triggered by emotion. This psychological state can enable an individual to have a unique emotional experience and further stimulate the individual's behavior. Emotion perception is the individual's perception or awareness of their own and external emotional states. The emotion perception ability can enable the individual to keenly capture the emotional information in the scenario, thereby triggering corresponding emotional reactions. As an intermediate link in the emotion cognitive process, emotion perception connects emotion arousal and emotion understanding. Emotion understanding is the individual's ability to recognize, interpret, and understand external emotional states, which includes the understanding of the reasons behind emotions, emotional experiences, and the rationality of emotional expressions. Emotion expression refers to the relevant behavioral manifestations in the emotion cognitive process, which is the individual's ability to convey their inner emotional experience to others through language, facial expressions, behaviors, etc. Emotion expression is an important output of emotion cognition, which enables the individual to clearly recognize their emotional needs and emotional states, so as to better manage their emotions. The internal connection between the emotion cognitive variables and the emotion cognitive stages refers to the data quantification of the emotion cognitive pattern by constructing an emotion path model.

[0098] Furthermore, the emotion path relationship model is divided into two parts. One part is the measurement model, which is used to represent the observed variables (skin conductance arousal level y eda , eye movement fixation level y eye , y2 = y eye = {(FC), (FD), (FA), (SA), and electroencephalogram activation level y eeg , y3 = y eeg = {PSD(i3,j3) and facial action unit y exp , and the relationship between the endogenous latent variables (i.e., emotion cognitive stage variables: emotion arousal η1, emotion perception η2, emotion understanding η3, and emotion expression η4). The other part is the structural model, which is used to describe the path relationship and causal relationship between the endogenous latent variables (i.e., emotion cognitive stage variables: emotion arousal η1, emotion perception η2, emotion understanding η3, emotion expression η4) and the exogenous latent variable (i.e., emotion cognitive ability x o ).

[0099] Furthermore, the relationship calculation formula between measurement models is: y i = λ i η i + ε i , i ∈ [1,4]. Among them, y i is four kinds of observed variables, that is, four kinds of modal data; λ i is the factor loading matrix, representing the path coefficient between the observed variable and the endogenous latent variable, η i is four kinds of endogenous latent variables; ε i is the measurement error, which is calculated by maximizing the likelihood function of the observed data. The calculation formula is: Among them, N is the number of samples, α is the number of observed variables, Cov(ε i ) is the covariance formula. μ(ε i ) is the error mean.

[0100] Furthermore, the calculation formula of the factor loading matrix is: λ i takes values in the range [-1,1]. λ i > 0 shows a positive correlation, λ i < 0 shows a negative correlation, λ i = 0 shows no correlation, and the larger |λ i | is, the stronger the relationship is. Cov(y i ,η i ) is the covariance, and Var(y i ) and Var(η i ) are variances, and Var(η i ) = 1;

[0101] Furthermore, the relational calculation formula between structural models is: η i = μ pij η i + ν qi x o + δ i , where i ∈ [1, 4], j ∈ [1, 4]. Here, x o is an exogenous latent variable, and η i is an endogenous latent variable (emotion arousal η1, emotion perception η2, emotion understanding η3, and emotion expression η4); μ pij is the path coefficient between endogenous latent variables, specifically representing the relationship between the i-th and j-th endogenous latent variables; ν qi is the path coefficient between an endogenous latent variable and an exogenous latent variable, specifically representing the path coefficient between the i-th endogenous latent variable and the exogenous latent variable; δ i is the model residual vector, which is calculated by maximizing the likelihood function of the observed data. The calculation formula is: where M is the number of samples, β is the number of observed variables, Cov(δ i ) is the covariance formula, and μ(δ i ) is the residual mean.

[0102] Furthermore, the path coefficient between endogenous latent variables: μ pij = Cov(η i , η j ), where i, j ∈ [1, 4] and i ≠ j. Cov(η i , η j ) is the covariance. When μ pij > 0, it indicates a positive correlation; when μ pij < 0, it indicates a negative correlation; when μ pij = 0, it indicates no correlation; the larger μ pij is, the stronger the relationship. The path coefficient between an endogenous latent variable and an exogenous latent variable: ν qi ranges from [-1, 1]. When ν qi > 0, it indicates a positive correlation; when ν qi < 0, it indicates a negative correlation; when ν qi = 0, it indicates no correlation; the larger |ν qi | is, the stronger the relationship. Cov(η i , x o ) is the covariance, and Var(η i ) and Var(x o ) are variances, and Var(x o ) = 1.

[0103] Four ν will be calculated by the above method qi , where i ∈ [1, 4]. Since the four emotional cognitive stage variables correspond one-to-one with the four modal data, the four ν qi correspond one-to-one with the four modalities. ν q1 corresponds to the electrodermal arousal feature modality, ν q2 corresponds to the eye movement fixation feature modality, ν q3 corresponds to the electroencephalogram activation feature modality, and ν q4 corresponds to the facial action unit feature.

[0104] The analysis result of the emotional cognitive pattern analysis module can be used as one of the output results of the autism symptom recognition system, and is used to analyze the performance of autistic children in the emotional cognitive stages (emotional arousal, emotional perception, emotional understanding, and emotional expression). According to the analysis result of the emotional cognitive pattern, locate the specific disorders of autistic children and formulate personalized intervention plans. At the same time, combining multi-modal data such as skin conductance response, eye movement tracking, electroencephalogram activation level, and facial action units, an emotional cognitive path map is generated. These analysis results will be uniformly stored in the personalized file to quantify the abnormal emotional expressions of autistic children and provide a scientific basis for the early screening of autism. In addition, taking the path coefficients in the emotional cognitive model as the input indicators of the deep learning module effectively integrates the information of different data sources, extracts the causal relationships of different cognitive processes, makes up for the deficiency of the deep learning model in feature interpretability, and improves the accuracy and robustness of the autism symptom recognition system.

[0105] The working principle of the deep learning module is described below.

[0106] The deep learning module is used to perform deep feature extraction on the multi-modal features in the data processing modality and perform multi-modal fusion, capture more emotional representative features, obtain the information between multi-modal information, promote the complementarity between modalities and balance the output, and improve the recognition accuracy of autistic children, as Figure 4 shown.

[0107] Furthermore, the deep learning module is used to put the electrodermal arousal feature, eye movement fixation feature, electroencephalogram activation feature, and facial AU feature into the corresponding self-attention mechanism respectively. The self-attention mechanism can dynamically capture the dependencies between elements at different positions in the sequence and generate new feature representations based on these dependencies. Calculate the attention scores of each modality through scaled dot product, and use these scores to perform weighted summation on the vectors to obtain multi-modal deep features with more emotional representative information. In one embodiment, among the input features of the self-attention mechanism, the electrodermal arousal level y eda , the eye movement fixation level y eye , yeye ={(FC), (FD), (FA), (SA), and the electroencephalogram activation level y eeg , y eeg ={PSD(i3, j3) and facial action unit y exp , These four types of features are respectively input into four corresponding self-attention models to extract four corresponding deep emotion features F eda , F eye , F eeg , F exp , and the specific mathematical expression is: Among them, self-attention1(), self-attention2(), self-attention3(), and self-attention4() are four self-attention() with exactly the same network structure, and each parameter is obtained by training with each modal feature.

[0108] Furthermore, the feature fusion algorithm fuses the deep features extracted by the self-attention mechanism, including the deep galvanic skin response feature F eda , the deep eye movement feature F eye , the deep electroencephalogram feature F eeg , and the deep AU feature F exp According to the path coefficients in reverse order for fusion. By having this reverse-order fusion combination method, the potential relationships between different modalities can be effectively utilized, and by promoting the complementarity and redundancy extraction of each layer of modalities, their contribution values to the final output result can be correspondingly weighted.

[0109] Furthermore, the specific steps of the path coefficient reverse-order fusion modal method include: projecting the deep galvanic skin response feature F eda , the deep eye movement feature F eye , the deep electroencephalogram feature F eeg , and the deep AU feature F exp through the mapping layer projection, performing feature orthogonality through the orthogonalization layer, using the method of reverse sorting based on the path coefficient size in the fusion layer to determine the fusion order, and putting the fused features into the output layer for classification using a binary classifier to obtain the recognition result.

[0110] First, standardize the deep features of each modality so that their feature value means are 0 and variances are 1. The modal standardization feature calculation formula is: Among them, is the standardized feature (i.e., and ),F x is a multi-modal feature (four deep features, namely F eda 、F eye 、F eeg and F exp ), μ x and σ x are the mean and standard deviation of this modality respectively. Secondly, the standardized features of each modality are used as inputs and linearly transformed to map to a new space. The linear transformation formula is: where W eda 、W eye 、W eeg and W exp are mapping matrices, and Z eda 、Z eye 、Z eeg and Z exp are mapping features. Next, the mapping features of different modalities are made orthogonal to each other in the new space by using orthogonal triangular (Orthogonal Triangular Decomposition, QR) decomposition. The path factors of each emotional cognitive stage calculated by the emotional path model are sorted from small to large, and F min→max (|ν qi |) is used as the order of orthogonal fusion in the reverse ranking order, where |ν qi | is the absolute value of the path coefficient between the four endogenous latent variables and the exogenous latent variable, and F min→max () is to sort the magnitudes of the four path coefficients |ν qi | from small to large, and then from large to small according to |ν qi |. The mapping features of the corresponding modalities are abbreviated as "rank 1 feature", "rank 2 feature", "rank 3 feature" and "rank 4 feature", and the same applies hereinafter.

[0111] First, the first-layer orthogonal fusion is performed. The first-layer orthogonal fusion calculation formula: where, is the first-layer orthogonal fusion feature, n6 is the feature dimension, is the feature orthogonal splicing, is the rank 1 feature, is the rank 2 feature, and R 12 is the upper triangular matrix after splicing.

[0112] Secondly, the second-layer orthogonal fusion is performed. The second-layer orthogonal fusion calculation formula: where, is the second-layer orthogonal fusion feature, n7 is the feature dimension, is the feature orthogonal splicing, is the first-layer orthogonal fusion feature is the feature ranked 3, R 123 is the upper triangular matrix after splicing.

[0113] Then, the third-layer orthogonal fusion is performed. The calculation formula for the third-layer orthogonal fusion is: where, is the third-layer orthogonal fusion feature, n8 is the feature dimension, is the feature orthogonal splicing, is the second-layer orthogonal fusion feature, is the feature ranked 4, R 1234 is the upper triangular matrix after splicing.

[0114] The final fusion feature is fed into a binary classifier to obtain the final recognition result. The classifier includes a fully connected layer and a softmax operation, and this operation projects and generates a label sequence, and compares it with the true label. The loss function is where, T2 is the number of modalities, and represent the predicted label and the true label respectively. 0 represents autism, 1 represents non-autism, is the inner product of the i5-th modality mapping feature and the j5-th modality mapping feature, is the weight of linear correlation.

[0115] Through this fusion method of reverse ranking, complementary information of each modality can be better extracted, and the true contribution degree of each modality to the recognition task can be reflected.

[0116] The present invention proposes an autism symptom recognition system for multi-emotion path analysis, covering the entire process technical architecture from multi-modal data acquisition, processing, analysis to deep learning feature extraction and fusion classification. By integrating four modality features of galvanic skin response, eye movement, electroencephalogram and facial expression, and combining the self-attention mechanism and the path coefficient orthogonal fusion algorithm, the representation ability of multi-modal emotion features and the recognition accuracy of autism are improved, and the stability and credibility of the classification result are ensured. The result output module visually displays the analysis result through visualization, and the personal profile module generates a personalized analysis report, providing a scientific basis for researchers and clinicians.

[0117] The present invention can realize the intelligent screening and comprehensive analysis of the emotion recognition ability of autistic children, greatly improving the clinical application efficiency and research value.

[0118] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A system for identifying autism symptoms by analyzing multiple emotional paths, characterized in that Including: A multimodal data acquisition module for acquiring four modal data of a participant, where the four modal data include skin conductance activity data, eye gaze data, electroencephalogram activity data, and facial expression data; A data feature extraction module for respectively extracting four modal features from the four modal data, where the four modal features include skin conductance arousal features, eye gaze features, electroencephalogram activation features, and facial action unit features; An emotion recognition pattern analysis module for defining four emotion recognition stage variables, where the four emotion recognition stage variables correspond one-to-one with the four modal features, respectively establishing path relationship models between the four modal features and the corresponding emotion recognition stage variables, establishing a path relationship model between the four emotion recognition stage variables and emotion recognition ability, and calculating path coefficients representing the correlation relationship between the four emotion recognition stage variables, as well as path coefficients representing the causal relationship between the four emotion recognition stage variables and emotion recognition ability; A deep learning module for extracting four modal deep features from the four modal features, performing normalization and linear transformation mapping on the four modal deep features to obtain mapping features of the four modalities, sorting the path coefficients of the causal relationship between the four emotion recognition stage variables and emotion recognition ability from smallest to largest, and orthogonally fusing the corresponding mapping features of the four modalities in reverse order to obtain a fused feature, and outputting an autism recognition result according to the fused feature.

2. The autism symptom recognition system for multi - emotion path analysis according to claim 1, characterized in that, The process of orthogonally fusing the corresponding mapping features of the four modalities in reverse order includes: According to the order of the absolute values of the path coefficients of the causal relationship between the four emotion recognition stage variables and emotion recognition ability from largest to smallest, the mapping features of the corresponding modality are called "rank 1 feature", "rank 2 feature", "rank 3 feature", and "rank 4 feature"; Perform the first-layer orthogonal fusion. The calculation formula for the first-layer orthogonal fusion is: Among them, is the first-layer orthogonal fusion feature, n6 is the feature dimension, is the feature orthogonal splicing, is the ranked 1 feature, is the ranked 2 feature, R 12 is the upper triangular matrix after splicing; Perform the second-layer orthogonal fusion. The calculation formula for the second-layer orthogonal fusion is: Among them, is the second-layer orthogonal fusion feature, n7 is the feature dimension, is the feature orthogonal splicing, is the rank-3 feature, R 123 is the upper triangular matrix after splicing; Perform the third-layer orthogonal fusion. The calculation formula for the third-layer orthogonal fusion is: Among them, is the third-layer orthogonal fusion feature, n8 is the feature dimension, is the feature orthogonal splicing, is the rank-4 feature, R 1234 is the upper triangular matrix after splicing.

3. The autism symptom recognition system for multi - emotion path analysis according to claim 1, characterized in that, The multiple emotion recognition stage variables include an emotion arousal variable, an emotion perception variable, an emotion understanding variable, and an emotion expression variable.

4. The autism symptom recognition system for multi - emotion path analysis according to claim 3, characterized in that, The path relationship model between the four modal features and the corresponding emotion recognition stage variables is: y i = λ i η i + ε i , i ∈ [1, 4], y i is the i-th modal feature, i ∈ [1, 4]; λ i represents the path coefficient between the i-th modal feature and the i-th emotional recognition stage variable, η i is the i-th emotional recognition stage variable, ε i is the measurement error, Cov() represents covariance, and Var() represents variance. The path relationship model of the path relationship model between the four emotion recognition stage variables and emotion recognition ability includes: η i = μ pij η i + ν qi x o + δ i , i ∈ [1, 4], j ∈ [1, 4], where x o is the emotional cognitive ability, η j is the variable of the j-th emotional cognitive stage, μ pij represents the relationship between the i-th emotional cognitive stage variable and the j-th emotional cognitive stage variable, ν qi represents the path coefficient between the i-th emotional cognitive stage variable and the emotional cognitive ability, δ i is the model residual vector, μ pij = Cov(η i , η j ), i, j ∈ [1, 4], where i ≠ j, 5. The autism symptom recognition system for parsing multiple emotional paths according to claim 1, characterized in that, Extracting four modal deep features from the four modal features: respectively inputting the four modal features into their respective self-attention models to extract four modal deep features.

6. The autism symptom recognition system for multi - emotion path analysis according to claim 1, characterized in that, The steps for extracting skin conductance arousal features include: Removing noise in the skin conductance activity data through four-layer wavelet transform; Performing normalization processing on the denoised skin conductance activity data; Calculating the average conductance level and average amplitude of the normalized skin conductance activity data, and taking the average conductance level and average amplitude as skin conductance arousal features.

7. The autism symptom recognition system for multi - emotion path analysis according to claim 1, characterized in that, The steps for extracting eye gaze features include: In the eye movement fixation data, the eye movement fixation coordinates of the participant at time t are denoted as (x(t3), y(t3)). According to the fixation area of the participant, it is divided into n3 non-overlapping rectangular regions of interest. The i1-th rectangular region of interest is denoted as AOI(i1), and AOI(i1) = [(x(i1), y(i1)), ((x(i1)+X(i1), y(i1)+Y(i1)))], where i1 ∈ (1, n3), n3 is the number of regions of interest, (X(i1), Y(i1)) is the size of the region of interest, and (x(i1), y(i1)) is the coordinate of the upper left region of interest. If x(i1) ≤ x(t3) ≤ x(i1)+X(i) and y(i1) ≤ y(t3) ≤ y(i1)+Y(i1), then it is determined that the eye movement coordinates of the participant at time t3 fall into the region of interest AOI(i1); otherwise, it is determined that the eye movement coordinates of the participant at time t3 do not fall into the region of interest AOI(i1). Calculate the number of fixations of the participant's gaze on the region of interest. If the eye movement fixation coordinates (x(t3), y(t3)) are within the region of interest, it is recorded as 1; otherwise, it is 0. Then where is the number of fixations on the i1-th rectangular region of interest, t3 is the total time length, and the sum of the number of fixations on n3 regions of interest is the final fixation count feature FC of the participant. The calculation formula is: Calculate the fixation time of the participant's gaze on the region of interest, and set each fixation time to a fixed duration Δt. Then, the fixation duration of the i1-th region of interest is as follows: where is the fixation time of the i1-th region of interest, t4 is the total time length, and the sum of the fixation times of n3 regions of interest is the final fixation time feature FD of the participant. The calculation formula is: Calculate the fixation range of the participant's fixation on the region of interest. The fixation range of the i1-th region of interest is where and are the fixation ranges on the horizontal x and vertical y in the i1-th region of interest respectively. min(x(t5)|(x(t5),y(t5))) is the minimum value of the fixation coordinates at time t5, and max(x(t5)|(x(t5),y(t5))) is the maximum value of the fixation coordinates at time t5. The average of the saccade ranges of n3 regions of interest is taken as the final fixation range feature FA of the participant, and the calculation formula is: Calculate the saccade amplitude of the participant's gaze on the region of interest. The saccade amplitude of the i1-th region of interest is represented by where Among them, and are the coordinates of two adjacent fixation points. The sum of the saccade amplitudes of n3 regions of interest is the final saccade amplitude feature of the participant. The calculation formula is:

8. The autism symptom recognition system for multi - emotion path analysis according to claim 1, characterized in that The extraction of electroencephalogram activation features includes the steps of: The electroencephalogram activity data is transformed from the time domain to the frequency domain through Fourier transform. The electroencephalogram activity data in the frequency domain is divided into 5 frequency bands according to the frequency bands to which the frequencies belong. The power spectral density of each frequency band is statistically calculated, and the power spectral densities of all frequency bands are used as the electroencephalogram activation features.

9. The autism symptom recognition system for multi - emotion path analysis according to claim 1, characterized in that, The extraction of facial action unit features includes the steps of: Facial action units are extracted from the facial expression data, and the frequencies and average intensities of the occurrences of the facial action units are recorded. The frequencies and average intensities of the occurrences of the facial action units are used as the facial action unit features.

10. The autism symptom recognition system for multi - emotion path analysis according to claim 1, characterized in that, It also includes: A paradigm material switching module for switching and displaying different emotion induction segments; A data synchronization module for aligning the collected skin conductance activity data, eye movement fixation data, electroencephalogram activity data, and facial expression data based on timestamps before extracting the four modal features; A personal profile module for generating a personalized emotion cognition analysis report for each participant. The emotion cognition analysis report includes the autism recognition result of the participant based on multi-modal features, the path relationship model between the four modal features of the participant and the corresponding emotion cognition stage variables, and the path relationship model between the four emotion cognition stage variables of the participant and the emotion cognition ability.

Citation Information

Cited By

  • Line inspection channel crowd behavior identification method and system based on multi-modal data fusion

    CN120853224A

  • A line inspection channel crowd behavior recognition method and system based on multi-modal data fusion

    CN120853224B