Autism spectrum disorder diagnosis and treatment method and system

By combining feature extraction and cross-fusion of EEG signal data and eye movement thermal images, and using an autism spectrum disorder classification model for diagnosis, the problems of high subjectivity and misdiagnosis rate in the existing autism spectrum disorder diagnosis technology are solved, and more accurate diagnostic results are achieved.

CN119418908BActive Publication Date: 2025-09-12SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411509986.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-26
Publication Date
2025-09-12
Estimated Expiration
2044-10-26

AI Technical Summary

Technical Problem

Existing technologies for diagnosing autism spectrum disorders rely on behavioral observations and questionnaires, which are highly subjective, have high misdiagnosis rates, and have limitations when processing EEG and eye movement data.

Method used

By obtaining the EEG signal data and eye movement thermal images of the person to be diagnosed, feature extraction is performed using preset extraction conditions, cross-fusion is performed by combining EEG feature vectors and eye movement features, and classification prediction is performed using the autism spectrum disorder classification model to obtain the diagnosis result.

Benefits of technology

The accuracy of autism spectrum disorder diagnosis is improved, the problems of strong subjectivity and high misdiagnosis rate in existing technologies are solved, and a more accurate diagnosis is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418908B_ABST
    Figure CN119418908B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of medical diagnosis technology, and provides a diagnostic processing method and system for autism spectrum disorder, the method comprising: obtaining EEG signal data and eye movement thermal images of a person to be diagnosed in multiple diagnostic tasks; performing feature extraction on the EEG signal data and eye movement thermal images of the multiple diagnostic tasks according to preset extraction conditions, to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks; cross-fusing the EEG feature vectors and eye movement features of the multiple diagnostic tasks, to obtain multimodal fusion attention features of the multiple diagnostic tasks; classifying and predicting the multiple multimodal fusion attention features, to obtain a diagnostic result for the person to be diagnosed, thereby solving the problems of the prior art in the diagnosis of people with autism spectrum disorder, which are highly subjective, have a high misdiagnosis rate, and have limitations in data processing due to reliance on behavioral observation and questionnaire surveys.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of medical diagnosis technology, and in particular relates to a diagnosis and treatment method and system for autism spectrum disorder. Background Art

[0002] Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder. Early intervention and diagnosis can significantly improve patients' quality of life. Traditional diagnostic methods for ASD rely primarily on behavioral observation and questionnaires, which are subject to high subjectivity and misdiagnosis rates.

[0003] However, in recent years, with the development of technology, ASD can be diagnosed by analyzing EEG data or eye movement data. However, existing methods have certain limitations in processing these EEG data or eye movement data. Summary of the Invention

[0004] The embodiments of the present application provide a method and system for diagnosing and treating autism spectrum disorder, which can solve the problems of the existing technology in the diagnosis of people with autism spectrum disorder, such as strong subjectivity and high misdiagnosis rate due to reliance on behavioral observation and questionnaire surveys, as well as limitations in processing multimodal data such as electroencephalogram (EEG) or eye movement data.

[0005] In a first aspect, embodiments of the present application provide a method for diagnosing and treating autism spectrum disorder, the method comprising:

[0006] Acquiring EEG signal data and eye movement thermal images of a person to be diagnosed during multiple diagnostic tasks; wherein the diagnostic tasks are used to represent an active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed;

[0007] According to preset extraction conditions, feature extraction is performed on the EEG signal data and the eye movement thermal images of the multiple diagnostic tasks to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks;

[0008] Cross-fusing the EEG feature vectors and the eye movement features of the plurality of diagnostic tasks to obtain multimodal fusion attention features of the plurality of diagnostic tasks;

[0009] Classify and predict the multiple multimodal fusion attention features to obtain a diagnosis result for the person to be diagnosed.

[0010] In a possible implementation of the first aspect, the performing feature extraction on the EEG signal data and the eye movement thermal image of the multiple diagnostic tasks according to preset extraction conditions to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks includes:

[0011] According to preset eye movement feature extraction conditions, deep semantic feature extraction is performed on the eye movement thermal image using an eye movement thermal map reconstruction model to obtain the eye movement features corresponding to the eye movement thermal image;

[0012] According to the preset EEG feature extraction conditions, the EEG signal data is subjected to multi-scale extraction by an EEG feature extraction model to obtain the EEG feature vector corresponding to the EEG signal data.

[0013] In a possible implementation of the first aspect, performing deep semantic feature extraction on the eye movement thermal image using an eye movement heat map reconstruction model according to preset eye movement feature extraction conditions to obtain the eye movement features corresponding to the eye movement heat map includes:

[0014] According to the preset eye movement feature extraction conditions, the encoder in the eye movement heat map reconstruction model performs deep semantic feature extraction on the eye movement heat map to obtain initial eye movement features corresponding to the eye movement heat map;

[0015] The initial eye movement features are discretized and encoded through the discrete semantic information dictionary structure in the eye movement heat map reconstruction model to obtain the eye movement features corresponding to the eye movement heat map.

[0016] In a possible implementation of the first aspect, after obtaining the eye movement feature corresponding to the eye movement thermal image, the method further includes:

[0017] The eye movement features are sampled and reconstructed by a decoder in the eye movement heat map reconstruction model to obtain a reconstructed eye movement heat map corresponding to the eye movement heat map;

[0018] Calculating the loss between the eye movement thermographic image and the corresponding reconstructed eye movement thermographic image based on a preset loss function;

[0019] According to the loss, the gap between the initial eye movement feature and the eye movement feature is reduced by a random gradient method to optimize the eye movement heat map reconstruction model.

[0020] In a possible implementation of the first aspect, performing multi-scale extraction on the EEG signal data using an EEG feature extraction model according to preset EEG feature extraction conditions to obtain the EEG feature vector corresponding to the EEG signal data includes:

[0021] According to the preset EEG feature extraction conditions, the EEG signal data is subjected to sliding window processing to obtain the multi-band brain topography corresponding to the EEG signal data. Figure 3 dimensional data;

[0022] The multi-band brain topography Figure 3 dimensional data into a set of two-dimensional segments centered on the electrode;

[0023] Performing multi-scale extraction of the spatial position relationship between all electrodes in the two-dimensional segment set through a self-attention network in an EEG feature extraction model, obtaining an initial EEG feature vector corresponding to the EEG signal data;

[0024] The initial EEG feature vector is normalized to obtain the EEG feature vector corresponding to the EEG signal data.

[0025] In a possible implementation of the first aspect, classifying and predicting the plurality of multimodal fusion attention features to obtain a diagnosis result for the person to be diagnosed includes:

[0026] Using an autism spectrum disorder classification model, classifying and predicting the plurality of multimodal fusion attention features of the person to be diagnosed obtained in the proactive joint attention test, and obtaining a plurality of proactive prediction results corresponding to the plurality of multimodal fusion attention features in the proactive joint attention test;

[0027] The diagnosis result of the person to be diagnosed is determined based on the multiple initiative prediction results and the preset initiative characteristic level index.

[0028] In a possible implementation of the first aspect, the autism spectrum disorder classification model is obtained by:

[0029] Obtaining a training set; wherein the training set includes a plurality of sample data and a plurality of prediction results of the sample data;

[0030] Building an initial autism spectrum disorder classification model based on graph attention network;

[0031] The initial autism spectrum disorder classification model is trained using the plurality of sample data and the prediction results of the plurality of sample data to obtain the autism spectrum disorder classification model.

[0032] In a possible implementation of the first aspect, constructing an initial autism spectrum disorder classification model based on a graph attention network includes:

[0033] Using each electrode in the EEG signal data as a node of the graph attention network;

[0034] Using the multimodal fusion attention feature as a node input feature of the graph attention network and using the adjacency matrix corresponding to the EEG signal data as an edge input feature of the graph attention network; wherein the adjacency matrix is ​​generated by using the phase lag index between every two electrodes in the EEG signal data as the functional connectivity relationship;

[0035] Obtaining an input layer of the graph attention network according to the nodes of the graph attention network and the edges of the graph attention network;

[0036] Using a preset single-layer feedforward neural network as the graph attention layer of the graph attention network to convert the multimodal fusion attention features into high-level multimodal fusion attention features;

[0037] Using a preset two-layer fully connected neural network as the output layer of the graph attention network, so as to obtain the diagnosis result according to the high-level multimodal fusion attention feature;

[0038] The initial autism spectrum disorder classification model is obtained according to the input layer, the graph attention layer and the output layer of the graph attention network.

[0039] In a possible implementation of the first aspect, obtaining EEG signal data and eye movement thermal images of the person to be diagnosed in multiple diagnostic tasks includes:

[0040] collecting EEG signals of the person to be diagnosed in the plurality of diagnostic tasks by a high-density EEG device to obtain a plurality of initial EEG signal data of the person to be diagnosed;

[0041] capturing eye movement behavior information of the person to be diagnosed during the multiple diagnostic tasks using a high-frequency eye tracking device to obtain multiple initial eye movement thermal images of the person to be diagnosed;

[0042] performing a first preprocessing operation on the plurality of initial EEG signal data by using a bandpass filter and an independent component analysis method to obtain a plurality of EEG signal data corresponding to the plurality of initial EEG signal data;

[0043] A second preprocessing operation is performed on the multiple initial eye movement thermal images through a Kalman filter to obtain multiple eye movement thermal images corresponding to the multiple initial eye movement thermal images.

[0044] In a second aspect, an embodiment of the present application provides a diagnostic and processing device for autism spectrum disorder, the device comprising:

[0045] a data acquisition module for acquiring EEG signal data and eye movement thermal images of a person to be diagnosed during a plurality of diagnostic tasks; wherein the diagnostic tasks are used to represent an active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed;

[0046] a feature extraction module, configured to perform feature extraction on the EEG signal data and the eye movement thermal images of the plurality of diagnostic tasks according to preset extraction conditions, to obtain EEG feature vectors and eye movement features of the plurality of diagnostic tasks;

[0047] a feature fusion module, configured to cross-fuse the EEG feature vectors and the eye movement features of the plurality of diagnostic tasks to obtain multimodal fused attention features of the plurality of diagnostic tasks;

[0048] The classification prediction module is used to perform classification prediction on the multiple multimodal fusion attention features to obtain the diagnosis result of the person to be diagnosed.

[0049] In a third aspect, an embodiment of the present application provides a diagnosis and processing system for autism spectrum disorder, the system comprising: an electroencephalogram (EEG) data acquisition device, an eye movement data acquisition device, a data processing device, and a model reasoning device; wherein,

[0050] The EEG data acquisition device includes a high-density EEG device, an electrode cap, and a first data acquisition card; the EEG data acquisition device is used to collect initial EEG signal data of the person to be diagnosed in multiple diagnostic tasks in real time through the high-density EEG device when the electrode cap is attached to the scalp of the person to be diagnosed, and is used to convert the initial EEG signal data into digital signals through the first data acquisition card and transmit them to the data processing device; the diagnostic tasks are used to represent the active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed;

[0051] The eye movement data acquisition device includes a high-frequency eye tracking device and a second data acquisition card; the eye movement data acquisition device is used to collect multiple initial eye movement thermal images of the person to be diagnosed in multiple diagnostic tasks in real time through the high-frequency eye tracking device, and is used to transmit the multiple initial eye movement thermal images to the data processing device through the second data acquisition card;

[0052] The data processing device is used to receive and store the multiple initial EEG signal data transmitted from the EEG data acquisition device and the multiple initial eye movement thermal images transmitted from the eye movement data acquisition device; and is used to perform a first preprocessing operation on the multiple initial EEG signal data to obtain multiple EEG signal data; perform a second preprocessing operation on the multiple initial eye movement thermal images to obtain multiple eye movement thermal images; and is used to perform feature extraction on the EEG signal data and the eye movement thermal images of the multiple diagnostic tasks according to preset extraction conditions to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks; and is used to cross-fuse the EEG feature vectors and the eye movement features of the multiple diagnostic tasks to obtain multimodal fusion attention features of the multiple diagnostic tasks; and is used to transmit the multiple multimodal fusion attention features to the model inference device;

[0053] The model inference device is used to classify and predict the multiple multimodal fusion attention features through an autism spectrum disorder classification model to obtain a diagnosis result for the person to be diagnosed.

[0054] In a possible implementation of the third aspect, the system includes: a user interface device; wherein,

[0055] The user interface device is used to display a user interface for the operator to input relevant information and start the diagnostic task, and to display the collected EEG signal data and the eye movement thermal image in real time, as well as the diagnostic results.

[0056] In a fourth aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for diagnosing and treating autism spectrum disorders described in any one of the above items is implemented.

[0057] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any of the above-mentioned methods for diagnosing and treating autism spectrum disorders.

[0058] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute any one of the above-mentioned methods for diagnosing and treating autism spectrum disorders.

[0059] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0060] An embodiment of the present application provides a diagnostic and processing method for autism spectrum disorder, by obtaining EEG signal data and eye movement thermal images of a person to be diagnosed in multiple diagnostic tasks; wherein the diagnostic tasks are used to characterize an active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed; according to preset extraction conditions, feature extraction is performed on the EEG signal data and eye movement thermal images of the multiple diagnostic tasks to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks; the EEG feature vectors and eye movement features of the multiple diagnostic tasks are cross-fused to obtain multimodal fusion attention features of the multiple diagnostic tasks; the multiple multimodal fusion attention features are classified and predicted to obtain a diagnostic result for the person to be diagnosed. By cross-fusing the EEG feature vectors extracted from EEG signal data and the eye movement features extracted from the eye movement thermal image, a multimodal fusion attention feature is obtained; and the multimodal fusion attention feature is classified and predicted to improve the accuracy of autism spectrum disorder diagnosis for the diagnosed person, thereby solving the problems of the existing technology in the diagnosis of autism spectrum disorder people due to reliance on behavioral observation and questionnaire surveys, which are highly subjective, have a high misdiagnosis rate, and have limitations in processing multimodal data such as EEG or eye movement data. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0062] Figure 1 This is a flow chart of a method for diagnosing and treating autism spectrum disorders provided in one embodiment of the present application;

[0063] Figure 2 is a schematic diagram of a proactive joint attention test provided by an embodiment of the present application;

[0064] Figure 3 This is a schematic diagram of generating multimodal fusion attention features provided by an embodiment of the present application;

[0065] Figure 4 This is a schematic diagram of a classification model for autism spectrum disorders provided in one embodiment of the present application;

[0066] Figure 5 This is a structural diagram of a graph attention network provided by an embodiment of the present application;

[0067] Figure 6This is a schematic structural diagram of a diagnostic and processing device for autism spectrum disorder provided in one embodiment of the present application;

[0068] Figure 7 This is a schematic structural diagram of a diagnosis and processing system for autism spectrum disorder provided in one embodiment of the present application;

[0069] Figure 8 is a structural diagram of a diagnosis and processing system for autism spectrum disorder provided by another embodiment of the present application;

[0070] Figure 9 This is a structural diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0071] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0072] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0073] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0074] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0075] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0076] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0077] See also Figure 1 , Figure 1 : This is a flow chart of a method for diagnosing and treating autism spectrum disorder provided in one embodiment of the present application. The method includes:

[0078] S11. Obtaining EEG signal data and eye movement thermal images of the person to be diagnosed during multiple diagnostic tasks; wherein the diagnostic tasks are used to represent an active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed;

[0079] S12. Perform feature extraction on the EEG signal data and eye movement thermal images of the multiple diagnostic tasks according to preset extraction conditions to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks;

[0080] S13, cross-fusing the EEG feature vectors and eye movement features of multiple diagnostic tasks to obtain multimodal fusion attention features of multiple diagnostic tasks;

[0081] S14. Classify and predict multiple multimodal fusion attention features to obtain a diagnosis result for the person to be diagnosed.

[0082] It should be noted that, in this embodiment, the execution entity may be a terminal device such as a server, and there is no specific limitation on this.

[0083] In step S11, the person to be diagnosed is a person who needs to be diagnosed with autism spectrum disorder, and is diagnosed as to whether he or she is a patient with autism spectrum disorder. The EEG signal data is the EEG signal of the person to be diagnosed obtained during the active joint attention test, that is, the electrical activity of different areas of the brain of the person to be diagnosed. The eye movement thermal image is an image generated by a heat map of the eye movement behavior information such as the gaze point position, movement trajectory and gaze time of the eyeball of the person to be diagnosed captured during the active joint attention test. The eye movement thermal image can intuitively display the eye movement behavior information of the eyeball of the person to be diagnosed, which is helpful for analyzing the distribution of visual attention during the test.

[0084] The diagnostic task is an active joint attention test performed when diagnosing autism spectrum disorder on the person to be diagnosed. In the diagnosis process of autism spectrum disorder, the joint attention test is an important assessment link. Joint attention refers to the ability of an individual to pay attention to the same object or event together with others. The joint attention test is an active joint attention test. Among them, the active joint attention test assesses the ability of the person to be diagnosed to actively guide others to pay attention to a certain object or event. In this embodiment, when diagnosing autism spectrum disorder on the person to be diagnosed, the active joint attention test is mainly performed. By analyzing the EEG signal data and eye movement thermal images obtained when the person to be diagnosed is subjected to the active joint attention test, medical staff or evaluators can have a more comprehensive understanding of the social interaction ability of the person to be diagnosed, thereby obtaining a more accurate diagnosis result, and judging whether the person to be diagnosed is a patient with autism spectrum disorder, so as to provide a more comprehensive and detailed intervention plan.

[0085] It should be noted that in this embodiment, the diagnostic task is a proactive joint attention test, which can include multiple subtests. Each test requires a certain amount of time, and during each subtest, multiple EEG signal data and multiple eye movement thermographic images are collected according to preset parameters.

[0086] Specifically, if Figure 2 As shown, Figure 2 This is a schematic diagram of a proactive joint attention test provided in one embodiment of the present application. Figure 2 In the , IJA stands for Active Joint Attention Test. Figure 2 The process of the active joint attention test is as follows: first, a cheese stimulus is displayed as a clue in the center of the display of the test device to inform the person to be diagnosed to look at the cheese clue stimulus during the test. Afterwards, the eye contact between the operator and the person to be diagnosed in the video is formed on the starting interface, and the duration is 1000ms. Afterwards, the person to be diagnosed looks at the cheese target stimulus to the left / right / upper side at will, and the operator follows the gaze direction of the person to be diagnosed and looks at the same target stimulus, and the duration is up to 4000ms. If the person to be diagnosed does not have active joint attention, it will automatically jump to the reward page, and the duration is 1000ms. Finally, the eye contact condition is restarted again after a time interval of 2000ms. Generally speaking, the active joint attention test includes 60 sub-tests, and the experimental duration is about 30-40 minutes. In this embodiment, there is no specific limitation on the number of sub-tests, test duration, and acquisition duration in the active joint attention test.

[0087] In step S12, the preset extraction conditions are pre-set conditions for extracting features from the EEG signal data and the eye movement thermal image. The EEG feature vector is a feature vector obtained after feature extraction of the EEG signal data according to the preset extraction conditions, mainly including extracted time domain features (such as mean potential, standard deviation) and frequency domain features (such as power spectrum density, frequency band energy). The eye movement feature is a feature obtained after feature extraction of the eye movement thermal image according to the preset extraction conditions, mainly including features such as gaze point distribution, gaze duration, and scanning speed.

[0088] In step S13, cross fusion is the process of integrating data features from different modalities, that is, integrating EEG feature vectors and eye movement features from the same diagnostic task according to certain rules. Multimodal fusion attention features are fusion feature vectors obtained by introducing attention mechanism on the basis of integrating EEG feature vectors and eye movement features according to certain rules. Figure 3 As shown, Figure 3 This is a schematic diagram of a multimodal fusion attention feature generation method provided in one embodiment of the present application. By cross-fusion, the useful information from EEG feature vectors and eye movement features is effectively combined to generate a multimodal fusion attention feature, thereby achieving better performance in subsequent diagnostic processes.

[0089] In step S14, a deep learning classifier, such as an autism spectrum disorder classification model, can be used to classify and predict the multiple multimodal fusion attention features to obtain a diagnosis result for the person to be diagnosed. The diagnosis result can indicate whether the person to be diagnosed has an autism spectrum disorder.

[0090] It can be understood that the embodiments of the present application perform feature extraction, cross-fusion, classification prediction and other operations on the EEG signal data and eye movement thermal images obtained during the diagnosis of autism spectrum disorders, thereby achieving a comprehensive assessment of whether the person being diagnosed suffers from autism spectrum disorders, thereby helping to improve the accuracy and reliability of autism spectrum disorder diagnosis.

[0091] It can be understood that the embodiment of the present application provides a diagnostic and processing method for autism spectrum disorder, by obtaining EEG signal data and eye movement thermal images of the person to be diagnosed in multiple diagnostic tasks; wherein the diagnostic tasks are used to characterize the active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed; according to preset extraction conditions, feature extraction is performed on the EEG signal data and eye movement thermal images of multiple diagnostic tasks respectively to obtain EEG feature vectors and eye movement features of multiple diagnostic tasks; the EEG feature vectors and eye movement features of multiple diagnostic tasks are cross-fused respectively to obtain multimodal fusion attention features of multiple diagnostic tasks; multiple multimodal fusion attention features are classified and predicted to obtain the diagnostic results of the person to be diagnosed. By cross-fusing the EEG feature vectors extracted from EEG signal data and the eye movement features extracted from the eye movement thermal image, a multimodal fusion attention feature is obtained; and the multimodal fusion attention feature is classified and predicted to improve the accuracy of autism spectrum disorder diagnosis for the diagnosed person, thereby solving the problems of the existing technology in the diagnosis of autism spectrum disorder people due to reliance on behavioral observation and questionnaire surveys, which are highly subjective, have a high misdiagnosis rate, and have limitations in processing multimodal data such as EEG or eye movement data.

[0092] In one possible implementation, based on preset extraction conditions, feature extraction is performed on the EEG signal data and eye movement thermal images of multiple diagnostic tasks to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks, including:

[0093] According to the preset eye movement feature extraction conditions, the eye movement thermal image is reconstructed using the eye movement thermal image reconstruction model to extract deep semantic features and obtain the eye movement features corresponding to the eye movement thermal image.

[0094] According to the preset EEG feature extraction conditions, the EEG signal data is subjected to multi-scale extraction through the EEG feature extraction model to obtain the EEG feature vector corresponding to the EEG signal data.

[0095] It should be noted that the preset eye movement feature extraction conditions refer to a series of rules or standards that are pre-set based on the requirements of the diagnostic task and the characteristics of the eye movement thermal image to guide the extraction of eye movement features. The eye movement thermal map reconstruction model is a model used to process eye movement thermal images. It can be used to extract eye movement features and reconstruct eye movement thermal images to achieve a deep understanding and representation of eye movement information data. Deep semantic feature extraction refers to the use of deep learning methods to automatically learn and extract features with rich semantic information from data. In the process of processing eye movement thermal images, deep semantic feature extraction can capture information related to the human visual attention mechanism in eye movement thermal images, such as gaze point distribution, gaze duration, and scan speed, that is, eye movement features. These features can not only reflect the surface form of eye movement behavior, but also reveal the cognitive process and psychological state behind it.

[0096] Preset EEG feature extraction conditions refer to a series of rules and standards that are pre-set based on the diagnostic task and the characteristics of the EEG signal data to guide EEG feature extraction. The EEG feature extraction model is a model used to extract useful information from EEG signal data. It can be used to analyze the complex characteristics of EEG signals and extract EEG feature vectors, thereby achieving in-depth analysis and understanding of EEG signal data. Multi-scale extraction refers to the use of different scales or resolutions to analyze data and extract useful features at each scale. In the process of processing EEG signal data, multi-scale extraction can capture the changing patterns of EEG signals at different time scales and different spatial scales, thereby extracting richer and more comprehensive feature information, such as time domain features (such as mean potential, standard deviation) and frequency domain features (such as power spectral density, frequency band energy), that is, EEG feature vectors. These vectors can better understand the complexity and diversity of EEG signals, providing strong support for subsequent diagnostic processes such as classification and prediction.

[0097] In one possible implementation, based on preset eye movement feature extraction conditions, deep semantic feature extraction is performed on the eye movement thermal image using an eye movement heat map reconstruction model to obtain eye movement features corresponding to the eye movement thermal image, including:

[0098] According to the preset eye movement feature extraction conditions, the encoder in the eye movement heat map reconstruction model performs deep semantic feature extraction on the eye movement heat map to obtain the initial eye movement features corresponding to the eye movement heat map;

[0099] The initial eye movement features are discretized and encoded through the discrete semantic information dictionary structure in the eye movement heat map reconstruction model to obtain the eye movement features corresponding to the eye movement heat map.

[0100] It should be noted that if Figure 3 As shown, Figure 3This is a schematic diagram of a multimodal fusion attention feature generation provided by an embodiment of the present application. Figure 3 In [1], the eye movement heat map reconstruction model consists of an encoder, a discrete semantic information dictionary structure, and a decoder. The purpose is to obtain the key semantic information in the eye movement heat map, namely the eye movement features, through the training of the encoder-decoder model.

[0101] Specifically, if Figure 3 In the eye movement heat map reconstruction model, the encoder is composed of a convolutional neural network, and the deep semantic information of the eye movement heat map, that is, the initial eye movement features, can be obtained through multiple convolutional layers. The input of the encoder is the eye movement heat map, and the output is the initial eye movement features. The initial eye movement features can be called intermediate results. The content of the intermediate results contains the semantic information of the image, including the attention information of the eye movements of the person to be diagnosed. Unlike the ordinary encoding-decoding model, the ordinary encoding-decoding model directly sends the intermediate results output by the encoder to the decoder to realize the reconstruction task, while the eye movement heat map reconstruction model in the embodiment of the present application processes the intermediate results and introduces a discrete semantic information dictionary structure, namely the codebook structure. The codebook structure is a discrete semantic information dictionary that stores multiple token vectors with the same length as the number of hidden layer channels. The token is a discrete semantic information. Before the intermediate result is sent to the decoder, it is discretized and encoded through the codebook structure to obtain the eye movement features. Assume that the intermediate result output by the encoder (i.e., the initial eye movement feature) is regarded as h*w vectors of length n, and the codebook structure stores multiple token vectors of length n. The discretization encoding process is to traverse each vector of the intermediate result in the codebook structure to find the closest token vector; replace the token vector with the vector at the original position. After processing the vector of the intermediate result, the intermediate result is replaced by a discrete semantic information matrix composed of tokens, i.e., the eye movement feature. Afterwards, the eye movement feature will be used as the attention prior distribution and cross-fused with the EEG feature vector extracted by the EEG feature extraction model. It should be noted that in this embodiment, discrete semantic information is selected to match the diffuse feature map of the EEG signal data centered on the electrode. The electrodes distributed on the scalp are discrete information, and the discrete semantic information can be used as the attention information prior of the discrete electrodes.

[0102] In one possible implementation, after obtaining the eye movement features corresponding to the eye movement thermal image, the method further includes:

[0103] The decoder in the eye movement heat map reconstruction model samples and reconstructs the eye movement features to obtain a reconstructed eye movement heat map corresponding to the eye movement heat map;

[0104] Based on a preset loss function, the loss between the eye movement thermal image and the corresponding reconstructed eye movement thermal image is calculated;

[0105] According to the loss, the gap between the initial eye movement features and the eye movement features is reduced by stochastic gradient method to optimize the eye movement heat map reconstruction model.

[0106] Continuing with the above example, Figure 3 In the eye movement heatmap reconstruction model, the decoder is also composed of a convolutional neural network. It takes the discretized eye movement features as input, restores the image through a deconvolution upsampling process, and reconstructs the eye movement heatmap based on the semantic information of the discretely encoded eye movement features to obtain a reconstructed eye movement heatmap. Then, based on a preset loss function, the loss between the eye movement heatmap and the reconstructed eye movement heatmap is calculated. Based on this loss, a stochastic gradient method is used to reduce the gap between the initial eye movement features and the eye movement features, thereby optimizing the eye movement heatmap reconstruction model and further improving its performance and stability.

[0107] It should be noted that the preset loss function is a pre-set function used to measure the difference between the reconstructed image and the actual image in the eye movement heat map reconstruction model. The loss of the preset loss function is defined as the following formula:

[0108]

[0109] Wherein, L is the loss calculated by the preset loss function; log p(x||Zq(x)) is used to optimize the encoder and decoder in the eye movement heat map reconstruction model; ||sg[Ze(x)] is used to optimize the codebook; is a regularization term used to constrain the encoder training; x is the input eye movement thermal image; β is a constant, which can be 0.25 in this embodiment; sg is the gradient stop, whose calculation remains unchanged during forward propagation and its partial derivative is 0 during back propagation; e is the defined embedding space, e∈R F×D In the token vector, F is the number of token vectors, and D is the dimension of the token vector. Ze(x) is the initial eye movement feature extracted from the input eye movement thermal image; and Zq(x) is the quantized eye movement feature. For each eye movement thermal image, the encoder in the eye movement thermal map reconstruction model extracts its initial eye movement feature Ze(x). For each vector, the encoder finds the index of the closest vector in the codebook, a discrete semantic information dictionary structure. The closest vector is obtained based on the index, resulting in the quantized eye movement feature Zq(x). Zq(x) is then fed into the decoder, which outputs the reconstructed eye movement thermal image.

[0110] It should be noted that the stochastic gradient method is an optimization algorithm used to optimize parameters in a model. There are many types of stochastic gradient methods, such as standard gradient descent, stochastic gradient descent, etc. In this embodiment, the specific type of the stochastic preset gradient method is not limited.

[0111] In one possible implementation, based on preset EEG feature extraction conditions, the EEG signal data is subjected to multi-scale extraction by an EEG feature extraction model to obtain an EEG feature vector corresponding to the EEG signal data, including:

[0112] According to the preset EEG feature extraction conditions, the EEG signal data is processed by sliding window to obtain the multi-band brain topography corresponding to the EEG signal data. Figure 3 dimensional data;

[0113] Multi-band brain topography Figure 3 dimensional data into a set of two-dimensional segments centered on the electrode;

[0114] The self-attention network in the EEG feature extraction model performs multi-scale extraction of the spatial position relationship between all electrodes in the two-dimensional segment set to obtain the initial EEG feature vector corresponding to the EEG signal data;

[0115] The initial EEG feature vector is normalized to obtain the EEG feature vector corresponding to the EEG signal data.

[0116] It should be noted that if Figure 3 As shown, Figure 3 This is a schematic diagram of a multimodal fusion attention feature generation provided by an embodiment of the present application. In this embodiment, Figure 3 In the EEG feature extraction model, the encoder consists of a self-attention encoder and a convolutional encoder. The input is the EEG signal data of each subtest in the diagnostic task. By performing sliding window processing on the EEG signal data, the multi-band brain topography is obtained. Figure 3 dimensional data. Among them, sliding window processing is a commonly used analysis method for time series data. By sliding a window of fixed length on the data and processing the data in each window, the original time series data is converted into a series of feature sequences. In this embodiment, sliding window processing of EEG signal data helps to capture the characteristics of EEG activity in different time periods. Multi-band brain topography is a method of graphically displaying the distribution of EEG signals on the scalp surface. Multi-band brain topography Figure 3 Dimensional data refers to the distribution and changes of EEG signals in three-dimensional space at different frequency bands.

[0117] Two-dimensional segment collection refers to the multi-band brain topography Figure 3The 2D image set centered on the electrode is obtained by simplifying the 2D data. The activity of each electrode in a specific frequency band and time period is extracted separately to form a 2D image segment. These segments together form a 2D segment set, which is used to analyze the spatial position relationship of all electrodes.

[0118] The self-attention network is a deep learning architecture suitable for processing dependencies between elements in sequential data. It captures the global dependencies between elements by calculating the attention weight of each element in the sequence on other elements. Normalization is used to scale the numerical range of data to a specific range.

[0119] During feature extraction from EEG signal data, the self-attention network analyzes the spatial relationships between all electrodes in a two-dimensional segment set, thereby extracting features of inter-electrode interactions and coordinated activity, known as the initial EEG feature vector. This initial EEG feature vector is then normalized to produce a processed EEG feature vector. Normalization helps eliminate dimensional differences and numerical ranges between multi-scale features, making the EEG feature extraction model easier to learn and generalize.

[0120] Specifically, in this embodiment, the 128-channel EEG signal data of the β [14-30 Hz] frequency band of each diagnostic task is subjected to sliding windowing to fully exploit the frequency information of EEG and obtain brain topography of a total of 5 frequency bands. Figure 3 In order to make full use of the spatial prior information of the electrodes, the electrodes are used as the center to divide the segments of different scales into discrete semantic information inputs of the self-attention network in the self-attention encoder, and the electrode positions are used as the position codes of the segments to be input into the self-attention network. Figure 3 dimensional data is converted into a set of two-dimensional segments centered on the electrodes. Then, the self-attention network in the EEG feature extraction model performs multi-scale extraction of the spatial relationship between all electrodes in the two-dimensional segment set to obtain an initial EEG feature vector. This initial EEG feature vector is normalized to obtain the final EEG feature vector. Unlike convolutional neural networks, the self-attention network extracts global features. Its application to EEG signal data fully extracts the spatial relationship between all electrodes. The self-attention encoder consists of alternating layers of multi-head self-attention (MSA) and multi-layer perceptron (MLP). Layer normalization (LN) is applied before the MLP for normalization, and a residual connection network is applied after the MLP.

[0121] The specific composition of the self-attention encoder is shown in the following formula:

[0122]

[0123] In the above formula, Attention(Q,K,V) represents the self-attention encoder; Q is the query matrix; K is the matrix of the content key you want to focus on, and the K matrix contains the vector key; V and K are equal, and V is the probability distribution converted after similarity calculation using the Q matrix and the K matrix; softmax() is the activation function; d k is the dimension of the vector key; QK T It is a dot product operation used to calculate the attention weight of the Q matrix on the V matrix.

[0124] In order to better integrate the initial EEG feature vectors learned under multi-scale receptive fields by the convolution encoder, in this embodiment, the multi-scale initial EEG feature vectors are normalized to the same dimension to obtain the EEG feature vectors.

[0125] Finally, the EEG feature vector is cross-fused with the eye movement features obtained by the eye movement heat map reconstruction model to obtain a multimodal fusion attention feature with attention prior.

[0126] In one possible implementation, multiple multimodal fusion attention features are classified and predicted to obtain a diagnosis result for the person to be diagnosed, including:

[0127] Using the autism spectrum disorder classification model, multiple multimodal fusion attention features of the diagnosed person obtained in the proactive joint attention test are classified and predicted, and multiple proactive prediction results corresponding to the multiple multimodal fusion attention features in the proactive joint attention test are obtained;

[0128] The diagnosis result of the person to be diagnosed is determined based on multiple initiative prediction results and a preset initiative characteristic level index.

[0129] It should be noted that if Figure 4 As shown, Figure 4 This is a structural diagram of an autism spectrum disorder classification model provided in one embodiment of the present application. Figure 4 In this example, ω represents the proactive prediction result corresponding to the proactive joint attention test. The autism spectrum disorder classification model is a trained deep learning model that takes multimodal fusion attention features as input and outputs a diagnostic result corresponding to these multimodal fusion attention features, thereby helping the operator determine whether the patient to be diagnosed has an autism spectrum disorder.

[0130] Classification prediction refers to the use of an autism spectrum disorder classification model to predict the multimodal fusion attention features obtained from the proactive joint attention test. The prediction results, which correlate the individual's performance on the proactive joint attention test, are considered proactive predictions. The preset proactive feature level index is a pre-set weighting threshold for proactive prediction results in the proactive attention test for the diagnosis of autism spectrum disorder, and is used to translate proactive test results into more specific diagnostic evidence.

[0131] Based on the initiative prediction results and the preset initiative characteristic level index, the diagnosis result of the person to be diagnosed can be determined, that is, the diagnosis result can include whether the person to be diagnosed suffers from autism spectrum disorder, and can also include the severity of the autism spectrum disorder, possible manifestations and other results.

[0132] Specifically, in this embodiment, Figure 4 As shown, each subtest generates a prediction result for the person being diagnosed, which can be expressed as ASD or TD. ASD indicates autism spectrum disorder, with a corresponding prediction result of -1; TD indicates normal, with a corresponding prediction result of 0. The prediction result is multiplied by the preset proactive feature level index, and the prediction results of all 2n subtests are accumulated to obtain the final classification result. If the final result is a negative number, it indicates that the person being diagnosed has autism spectrum disorder (ASD); if the final result is a positive number, it indicates that the person being diagnosed is normal.

[0133] In one possible implementation, the autism spectrum disorder classification model is obtained by:

[0134] Obtaining a training set; wherein the training set includes a plurality of sample data and prediction results of the plurality of sample data;

[0135] Building an initial autism spectrum disorder classification model based on graph attention network;

[0136] An initial autism spectrum disorder classification model is trained using the plurality of sample data and the prediction results of the plurality of sample data to obtain an autism spectrum disorder classification model.

[0137] It should be noted that the Graph Attention Network is a neural network model based on graph-structured data. It can use graph-structured data to capture the complex relationships between nodes through an attention mechanism, thereby extracting features that are useful for autism spectrum disorder classification.

[0138] In this embodiment, it is necessary to perform model training on the constructed initial autism spectrum disorder classification model to obtain the autism spectrum disorder classification model. The initial autism spectrum disorder classification model is a newly constructed learning model that has not been model trained. The training set is a subset of data used to train the initial model. These data contain known input features and expected output labels, that is, sample data and the predicted results of the sample data. In this embodiment, the sample data is a multimodal fusion attention feature obtained after feature extraction and cross-fusion of EEG signal data and eye movement thermal images. The training set is usually data from people who have been diagnosed with autism spectrum disorder and data from people who have not been diagnosed with autism spectrum disorder. The predicted result of the sample data is the predicted result output by the model, that is, whether the person is a person with autism spectrum disorder.

[0139] It can be understood that by training the initial autism spectrum disorder classification model, the model can accurately classify and predict new data, accurately predict whether the person to be diagnosed is a person with autism spectrum disorder, and thus realize the diagnosis of people with autism spectrum disorder.

[0140] It should be noted that in order to further evaluate the feasibility of the autism spectrum disorder classification model for assisting clinical diagnosis, the accuracy of the model was compared with the results of the clinical gold scale (the clinical gold scale includes the Autism Parent Interview Scale (ADI-R) and the Autism Diagnostic Observation Scale (ADOS), and the feasibility of the model for assisting clinical use was assessed at three levels. Among them, the first level result was excellent, that is, the accuracy of the model was higher than the clinical accuracy of both ADI-R and ADOS; the second level result was good, that is, the accuracy of the model was only higher than the clinical accuracy of either ADI-R or ADOS; the third level result was unacceptable, that is, the accuracy of the model was lower than the clinical accuracy of both ADI-R and ADOS.

[0141] In one possible implementation, an initial autism spectrum disorder classification model is constructed based on a graph attention network, including:

[0142] Each electrode in the EEG signal data is used as a node in the graph attention network;

[0143] The multimodal fusion attention features are used as node input features of the graph attention network, and the adjacency matrix corresponding to the EEG signal data is used as the edge input feature of the graph attention network. The adjacency matrix is ​​generated by taking the phase lag index between each two electrodes in the EEG signal data as the functional connectivity relationship.

[0144] According to the nodes and edges of the graph attention network, the input layer of the graph attention network is obtained;

[0145] The preset single-layer feedforward neural network is used as the graph attention layer of the graph attention network to convert the multimodal fusion attention features into high-level multimodal fusion attention features;

[0146] A preset two-layer fully connected neural network is used as the output layer of the graph attention network to obtain diagnostic results based on high-level multimodal fusion attention features;

[0147] According to the input layer, graph attention layer and output layer of the graph attention network, the initial autism spectrum disorder classification model is obtained.

[0148] It should be noted that if Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of an autism spectrum disorder classification model provided in one embodiment of the present application. In this embodiment, a graph attention network is used to construct an initial autism spectrum disorder classification model. A graph attention network is a neural network model based on graph-structured data. A graph attention network can utilize graph-structured data and, through an attention mechanism, capture the complex relationships between nodes, thereby extracting features useful for autism spectrum disorder classification.

[0149] like Figure 4 As shown in Figure 2, the specific process of constructing the initial autism spectrum disorder classification model is as follows: each electrode in the EEG signal data is used as a node of the graph attention network; the multimodal fusion attention feature is used as the node input feature of the graph attention network, that is, Figure 4 The phase lag index between each two electrodes in the EEG signal data is used as the functional connectivity relationship (edge ​​of the graph attention network) to generate an adjacency matrix, and the adjacency matrix is ​​used as the edge input feature of the graph attention network, that is, Figure 4 The EEG network in the graph attention network is then used as the input layer of the graph attention network. A preset single-layer feedforward neural network is used as the graph attention layer of the graph attention network to convert the multimodal fusion attention features into high-level multimodal fusion attention features. A preset two-layer fully connected neural network is used as the output layer of the graph attention network to obtain diagnostic results based on the high-level multimodal fusion attention features. The above input layer, graph attention layer, and output layer constitute the initial autism spectrum disorder classification model.

[0150] The adjacency matrix is ​​used to represent the connectivity between nodes in a graph attention network. In EEG data processing, the phase lag index between each pair of electrodes is used as the functional connectivity to construct connectivity between nodes. The phase lag index is a metric used to quantify the stability of the phase difference between two signals. In EEG data processing, the phase lag index is used to assess signal synchronization or functional connectivity between different electrodes.

[0151] The input layer is the first layer to receive feature vectors, including node and edge input features, namely, multimodal fusion attention features and an adjacency matrix. The graph attention layer is the core component of the graph attention network. It uses an attention mechanism to update node feature representations. Each node updates its own features based on the features of its neighbors. Simultaneously, the attention mechanism dynamically adjusts the importance of different neighboring nodes to the current node. A feedforward neural network is a basic neural network architecture in which information propagates only forward from the input layer to the output layer, without loops or feedback connections. The pre-set single-layer feedforward neural network is a pre-set neural network architecture used to convert multimodal fusion attention features into high-level multimodal fusion attention features. The output layer is the final layer in the graph attention network that produces the final prediction or classification result. A fully connected neural network is also a basic neural network architecture in which every node in each layer is connected to every node in the next layer. The pre-set two-layer fully connected neural network is a pre-set neural network architecture used to generate diagnostic results for autism spectrum disorder based on high-level multimodal fusion attention features.

[0152] Specifically, if Figure 5 As shown, Figure 5 This is a schematic diagram of a graph attention network structure provided by an embodiment of the present application. In the structure of the graph attention network, the nodes are represented as Where N is the number of nodes. In order to transform the input multimodal fusion attention features into higher-level features, a learnable linear transformation is required. To this end, a shared linear transformation parameterized by the weight matrix is ​​performed on each node. Then, a shared attention mechanism is performed on the node to calculate the attention coefficient, which is expressed as follows:

[0153]

[0154] Among them, e ij is the attention coefficient, i.e., the importance of the feature vector of node j to the feature vector of node i; a() is the function for calculating the relevance between nodes i and j; W1 is the weight parameter for the feature transformation of the nodes in this layer; is the feature vector of node i; is the feature vector of node j; j∈N i , N i are all first-order neighbors of node i.

[0155] In order to make the attention coefficient e ij To make it easier to compare different nodes, we use an activation function to normalize them, i.e. softmax. j () function, as described below:

[0156]

[0157] in, The weight coefficient calculated for the attention mechanism; e ik is the attention coefficient that represents the importance of the feature vector of node k to the feature vector of node i. Figure 5 As shown in (5-a), α ij Represents the weight coefficient calculated by the attention mechanism, softmax j represents softmax j ()function, Represents a node, represents the feature vector of node i, represents the feature vector of node j, W1 is the weight parameter of the feature transformation of the node in this layer, and the weight coefficient calculated by the attention mechanism can be expressed as:

[0158]

[0159] in, is the feature vector of node k, k∈N i , N i are all the first-order neighbors of node i; LeakyReLU() represents the activation function; T represents transposition, || represents concatenation operation; W2 is the learning parameter.

[0160] The aggregation process of the graph attention network is as follows Figure 5 As shown in (5-b), the weight coefficient calculated by the attention mechanism is obtained After that, the normalized attention coefficient is used to perform the connection aggregation operation to calculate the final output features of each node, namely:

[0161]

[0162] in, represents the concatenation result, i.e., the high-level multimodal fusion attention feature; σ() represents the activation function; is the weight coefficient calculated by the attention mechanism; W2 is the learning parameter.

[0163] like Figure 5 In (5-b), is the eigenvector of node 1, is the eigenvector of node 2, is the eigenvector of node 3, The eigenvector of node 4, is the eigenvector of node 5, is the eigenvector of node 6, is the weight coefficient of node 1 calculated by the attention mechanism, is the weight coefficient between node 1 and node 2 calculated by the attention mechanism, The weight coefficient between nodes 1 and 3 calculated by the attention mechanism, The weight coefficient between nodes 1 and 4 calculated by the attention mechanism, The weight coefficient between nodes 1 and 5 calculated by the attention mechanism, The weight coefficient between nodes 1 and 6 calculated by the attention mechanism; concat / avg indicates a connection aggregation operation; Represents the high-level multimodal fusion attention features output by node 1 after processing through the graph attention network.

[0164] In the last part of the graph attention network, a two-layer fully connected neural network is used to perform binary classification on the autism spectrum disorder group and the normal control group to obtain the final diagnosis results.

[0165] In one possible implementation, obtaining EEG signal data and eye movement thermal images of a person to be diagnosed during multiple diagnostic tasks includes:

[0166] The EEG signals of the person to be diagnosed in multiple diagnostic tasks are collected by high-density EEG equipment to obtain multiple initial EEG signal data of the person to be diagnosed;

[0167] The eye movement behavior information of the person to be diagnosed during multiple diagnostic tasks is captured by a high-frequency eye tracking device to obtain multiple initial eye movement thermal images of the person to be diagnosed;

[0168] Performing a first preprocessing operation on the plurality of initial EEG signal data by using a bandpass filter and an independent component analysis method to obtain a plurality of EEG signal data corresponding to the plurality of initial EEG signal data;

[0169] A second preprocessing operation is performed on the multiple initial eye movement thermal images through a Kalman filter to obtain multiple eye movement thermal images corresponding to the multiple initial eye movement thermal images.

[0170] It should be noted that a high-density EEG device is a medical device used to record and analyze brain electrical signals. It is usually composed of multiple electrodes. These electrodes are placed on the scalp of the person to be diagnosed to capture the weak electrical signals generated by the activity of the brain neurons of the person to be diagnosed, and can provide more detailed and comprehensive brain electrical signal information. A high-frequency eye tracking device is a high-precision instrument used to track and record eye movement trajectories in real time. In this embodiment, the high-frequency eye tracking device is used to collect multiple eye movement thermal images of the person to be diagnosed during the diagnostic task in real time, that is, the eye movement information of the person to be diagnosed, mainly including eye movement information such as frequency information, time length information, position (up, down, left, right, etc.) of gaze.

[0171] Initial EEG signal data is the raw EEG signal collected directly from the patient's scalp using high-density EEG equipment during the active joint attention test. It represents the electrical activity in different regions of the patient's brain. This data can contain rich neural activity information, but it can also contain some noise. Initial eye movement thermal images are generated using a heat map, capturing eye movement information such as the position, movement trajectory, and duration of the patient's gaze during the active joint attention test using high-frequency eye tracking equipment. This image can be used to visually display the patient's eye movement behavior, helping to analyze the allocation of visual attention during the test.

[0172] A bandpass filter is a signal processing tool used to remove unwanted frequency components from a signal while retaining the frequency components of interest. Independent component analysis is a statistical method used to decompose multichannel signals into independent components. The first preprocessing operation is a series of preprocessing methods performed on the initial EEG signal data. This first preprocessing operation includes, but is not limited to, denoising through a bandpass filter and removing artifacts using independent component analysis.

[0173] Specifically, during the first preprocessing operation on the initial EEG signal data, the low-frequency drift and high-frequency noise in the initial EEG signal data are removed by a bandpass filter to obtain an intermediate-frequency signal related to brain activity; and the independent component analysis method is used to remove artifacts such as blinking and electromyographic artifacts and separate independent components related to brain activity, and finally the EEG signal data is obtained.

[0174] A Kalman filter is a recursive filter that can estimate the state of a dynamic system from a series of data in the presence of measurement noise. The second preprocessing operation is a series of preprocessing methods performed on the initial eye movement thermal image. This second preprocessing operation includes, but is not limited to, smoothing the eye movement trajectory through a Kalman filter and removing outliers.

[0175] It is understood that by performing a first preprocessing operation on the initial EEG signal data to obtain EEG signal data, the signal-to-noise ratio of the EEG signal data is improved, making it more suitable for subsequent feature extraction. By performing a second preprocessing operation on the initial eye movement thermal image to obtain an eye movement thermal image, the quality of the eye movement thermal image is improved so that it can accurately reflect the eye movements of the person to be diagnosed.

[0176] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0177] Corresponding to a method for diagnosing and treating autism spectrum disorder in the above embodiment, Figure 6 A schematic structural diagram of a diagnosis and treatment device for autism spectrum disorder provided in one embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0178] Reference Figure 6 The autism spectrum disorder diagnosis and processing device 3 of this embodiment includes:

[0179] The data acquisition module 31 is used to acquire EEG signal data and eye movement thermal images of the person to be diagnosed during multiple diagnostic tasks; wherein the diagnostic tasks are used to represent the active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed;

[0180] A feature extraction module 32 is used to extract features from the EEG signal data and eye movement thermal images of multiple diagnostic tasks according to preset extraction conditions, thereby obtaining EEG feature vectors and eye movement features of the multiple diagnostic tasks;

[0181] A feature fusion module 33 is used to cross-fuse the EEG feature vectors and eye movement features of multiple diagnostic tasks to obtain multimodal fusion attention features of multiple diagnostic tasks;

[0182] The classification prediction module 34 is used to perform classification prediction on multiple multimodal fusion attention features to obtain the diagnosis result of the person to be diagnosed.

[0183] It can be understood that the diagnosis and processing device 3 for autism spectrum disorder provided in the embodiment of the present application obtains the EEG signal data and eye movement thermal images of the person to be diagnosed in multiple diagnostic tasks through the data acquisition module 31; wherein the diagnostic task is used to characterize the active joint attention test for diagnosing autism spectrum disorder of the person to be diagnosed; then, the feature extraction module 32 performs feature extraction on the EEG signal data and eye movement thermal images of the multiple diagnostic tasks according to preset extraction conditions to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks; then, the feature fusion module 33 cross-fuses the EEG feature vectors and eye movement features of the multiple diagnostic tasks to obtain multimodal fused attention features of the multiple diagnostic tasks; finally, the classification prediction module 34 performs classification prediction on the multiple multimodal fused attention features to obtain the diagnosis result of the person to be diagnosed. By cross-fusing the EEG feature vectors extracted from EEG signal data and the eye movement features extracted from the eye movement thermal image, a multimodal fusion attention feature is obtained; and the multimodal fusion attention feature is classified and predicted to improve the accuracy of autism spectrum disorder diagnosis for the diagnosed person, thereby solving the problems of the existing technology in the diagnosis of autism spectrum disorder people due to reliance on behavioral observation and questionnaire surveys, which are highly subjective, have a high misdiagnosis rate, and have limitations in processing multimodal data such as EEG or eye movement data.

[0184] It should be noted that the information interaction, execution process, etc. between the modules in the above-mentioned autism spectrum disorder diagnosis and processing device 3 are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0185] The present application also provides a diagnosis and processing system for autism spectrum disorder, such as Figure 7 As shown, Figure 7 This is a schematic diagram of the structure of a diagnosis and processing system for autism spectrum disorder provided by one embodiment of the present application. Figure 7 The autism spectrum disorder diagnosis and processing system 4 of this embodiment includes: an EEG data acquisition device 41, an eye movement data acquisition device 42, a data processing device 43 and a model reasoning device 44; wherein,

[0186] The EEG data acquisition device 41 includes a high-density EEG device, an electrode cap, and a first data acquisition card. The EEG data acquisition device 41 is used to collect initial EEG signal data of the person to be diagnosed in multiple diagnostic tasks in real time through the high-density EEG device when the electrode cap is attached to the scalp of the person to be diagnosed, and is used to convert the initial EEG signal data into digital signals through the first data acquisition card and transmit them to the data processing device 43. The diagnostic tasks are used to represent the active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed.

[0187] The eye movement data acquisition device 42 includes a high-frequency eye tracking device and a second data acquisition card. The eye movement data acquisition device 42 is used to acquire multiple initial eye movement thermal images of the person to be diagnosed during multiple diagnostic tasks in real time using the high-frequency eye tracking device, and is used to transmit the multiple initial eye movement thermal images to the data processing device 43 via the second data acquisition card.

[0188] The data processing device 43 is used to receive and store a plurality of initial EEG signal data transmitted from the EEG data acquisition device 41 and a plurality of initial eye movement thermal images transmitted from the eye movement data acquisition device 42; and is used to perform a first preprocessing operation on the plurality of initial EEG signal data to obtain a plurality of EEG signal data; and to perform a second preprocessing operation on the plurality of initial eye movement thermal images to obtain a plurality of eye movement thermal images; and is used to perform feature extraction on the EEG signal data and eye movement thermal images of a plurality of diagnostic tasks according to preset extraction conditions to obtain EEG feature vectors and eye movement features of the plurality of diagnostic tasks; and is used to cross-fuse the EEG feature vectors and eye movement features of the plurality of diagnostic tasks to obtain multimodal fusion attention features of the plurality of diagnostic tasks; and is used to transmit the plurality of multimodal fusion attention features to the model inference device 44;

[0189] The model inference device 44 is used to classify and predict multiple multimodal fusion attention features through an autism spectrum disorder classification model to obtain a diagnosis result for the person to be diagnosed.

[0190] It should be noted that the diagnosis and processing system 4 for autism spectrum disorders includes an EEG data acquisition device 41 , an eye movement data acquisition device 42 , a data processing device 43 , and a model reasoning device 44 ;

[0191] The EEG data acquisition device 41 includes a high-density EEG device, an electrode cap, and a first data acquisition card. The high-density EEG device is a medical device used to record and analyze brain electrical signals, typically consisting of a plurality of electrodes. These electrodes are placed on the scalp of the person to be diagnosed to capture the weak electrical signals generated by the activity of the brain neurons of the person to be diagnosed, and can provide more detailed and comprehensive information on the brain's electrical signals. The electrode cap is a key component used in conjunction with the high-density EEG device, and is covered with electrode contacts connected to the high-density EEG device. When performing EEG data acquisition, the electrode cap is attached to the scalp of the person to be diagnosed to ensure good contact between the electrodes and the scalp, thereby accurately capturing the brain's electrical signals. The first data acquisition card is a hardware device used to convert the initial EEG signal data collected by the high-density EEG device and the electrode cap into digital signals, and transmit them to the data processing device 43 for subsequent processing and analysis.

[0192] Specifically, when the electrode cap is attached to the scalp of the person being diagnosed, the high-density EEG device collects the person's EEG signals during multiple diagnostic tasks in real time, obtaining initial EEG signal data. The initial EEG signal data is converted into digital signals by the first data acquisition card and transmitted to the data processing device 43. The EEG data acquisition device 41, which includes the high-density EEG device, the electrode cap, and the data acquisition card, achieves comprehensive and accurate collection and recording of the person's EEG signals.

[0193] The eye movement data acquisition device 42 includes a high-frequency eye tracking device and a second data acquisition card. The high-frequency eye tracking device is a high-precision instrument for real-time tracking and recording of eye movement trajectories. In this embodiment, the high-frequency eye tracking device is used to collect in real time a plurality of initial eye movement thermal images of the person to be diagnosed in the diagnostic task, that is, the original eye movement information of the person to be diagnosed, mainly including eye movement information such as frequency information, time length information, position (up, down, left, right, etc.) information of gaze. The second data acquisition card is a hardware device for transmitting the initial eye movement thermal images collected by the high-frequency eye tracking device to the data processing device 43 for subsequent processing and analysis. Among them, the initial eye movement thermal images are originally collected, and can intuitively display images of eye movement information such as the gaze pattern and distribution of the eyeballs of the person to be diagnosed when performing the active joint attention test.

[0194] The data processing device 43 is used to receive and store multiple initial EEG signal data transmitted from the EEG data acquisition device 41 and multiple initial eye movement thermal images transmitted from the eye movement data acquisition device 42. It performs a first preprocessing operation on the initial EEG signal data to obtain EEG signal data, and a second preprocessing operation on the initial eye movement thermal images to obtain eye movement thermal images. It then performs a series of operations on the EEG signal data and the eye movement thermal images, including feature extraction and feature fusion, to obtain multimodal fused attention features, which are then transmitted to the model inference device 44. The first preprocessing operation includes, but is not limited to, a series of preprocessing methods performed on the initial EEG signal data, such as denoising via a bandpass filter and artifact removal using independent component analysis. The second preprocessing operation includes, but is not limited to, smoothing eye movement trajectories using a Kalman filter and removing outliers. By efficiently and accurately processing and analyzing multimodal data, the data processing device 43 provides strong support for the subsequent classification and prediction by the model inference device 44.

[0195] Model inference device 44 is used to receive the multimodal fusion attention features sent by data processing device 43 and classify and predict them using the autism spectrum disorder classification model, thereby obtaining a diagnosis result for the patient. The autism spectrum disorder classification model is a trained model that can classify and predict the multimodal fusion attention features of the patient.

[0196] It should be noted that the diagnostic task is an active joint attention test conducted when the person to be diagnosed is diagnosed for autism spectrum disorder. In the diagnosis process of autism spectrum disorder, the joint attention test is an important assessment link. Joint attention refers to the ability of an individual to pay attention to the same object or event together with others. Joint attention is an active joint attention test. Among them, the active joint attention test assesses the ability of the person to be diagnosed to actively guide others to pay attention to a certain object or event. In this embodiment, when the person to be diagnosed is diagnosed for autism spectrum disorder, the active joint attention test is mainly conducted. By analyzing the EEG signal data and eye movement thermal images of the person to be diagnosed obtained when the person to be diagnosed is subjected to the active joint attention test, medical staff or evaluators can have a more comprehensive understanding of the social interaction ability of the person to be diagnosed, thereby obtaining more accurate diagnostic results, so as to provide a more comprehensive and detailed intervention plan.

[0197] It can be understood that the diagnosis and processing system 4 for autism spectrum disorder in this embodiment first obtains the EEG signal data and eye movement thermal images of the person to be diagnosed in multiple diagnostic tasks through the EEG data acquisition device 41 and the eye movement data acquisition device 42 respectively; wherein the diagnostic task is used to characterize the active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed; then, according to preset extraction conditions, the data processing device 43 performs feature extraction on the EEG signal data and eye movement thermal images of multiple diagnostic tasks respectively to obtain EEG feature vectors and eye movement features of multiple diagnostic tasks; the EEG feature vectors and eye movement features of multiple diagnostic tasks are cross-fused respectively to obtain multimodal fusion attention features of multiple diagnostic tasks; finally, the model inference device 44 performs classification and prediction on multiple multimodal fusion attention features to obtain the diagnosis result of the person to be diagnosed. By cross-fusing the EEG feature vectors extracted from EEG signal data and the eye movement features extracted from the eye movement thermal image, a multimodal fusion attention feature is obtained; and the multimodal fusion attention feature is classified and predicted to improve the accuracy of autism spectrum disorder diagnosis for the diagnosed person, thereby solving the problems of the existing technology in the diagnosis of autism spectrum disorder people due to reliance on behavioral observation and questionnaire surveys, which are highly subjective, have a high misdiagnosis rate, and have limitations in processing multimodal data such as EEG or eye movement data.

[0198] The present application also provides a diagnosis and processing system for autism spectrum disorder, such as Figure 8 As shown, Figure 8 This is a structural diagram of a diagnosis and processing system for autism spectrum disorder provided by another embodiment of the present application. Figure 8 , the autism spectrum disorder diagnosis and processing system 4 of this embodiment further includes: a user interface device 45;

[0199] The user interface device 45 is used to display a user interface for the operator to input relevant information and start the diagnostic task, and to display the collected EEG signal data and eye movement thermal images in real time, as well as the diagnostic results.

[0200] It should be noted that user interface device 45 is a device capable of displaying a graphical user interface. This user interface device 45 allows an operator to interact with the system. By providing a user interface, the operator can input relevant information and initiate diagnostic tests for the patient being diagnosed. It can also display the collected EEG signal data and eye movement thermal images in real time, as well as the final diagnostic results. This user interface device 45 can be, but is not limited to, a traditional computer monitor with a mouse and keyboard, a touch screen display, a tablet computer, a smartphone, or other device with both display and input functions.

[0201] It should be noted that the user interface device 45 can also display detailed diagnostic analysis information, such as feature weights, classification confidence, etc.

[0202] It should be noted that the autism spectrum disorder diagnosis and processing system 4 can also save the diagnosis results and generate relevant diagnostic reports. Specifically, the diagnosis results are saved, the information is analyzed, and a detailed diagnostic report is generated. The diagnostic report can be exported to PDF format for archiving and further analysis.

[0203] It should be noted that the autism spectrum disorder diagnosis and processing system 4 can also adjust and optimize the autism spectrum disorder classification model based on the diagnosis results and clinical feedback, and improve the training set of the autism spectrum disorder classification model by collecting more sample data, thereby improving the generalization ability of the model.

[0204] It should be noted that each device in the autism spectrum disorder diagnosis and treatment system 4 is regularly inspected to check its operating status and perform necessary maintenance. The system's data processing and model inference software are upgraded based on the latest research results and clinical needs.

[0205] The present application also provides an electronic device, such as Figure 9 As shown, Figure 9 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. Figure 9 The electronic device 5 of this embodiment includes: a memory 51, a processor 52, and a computer program stored in the memory 51 and executable on the processor 52. When the processor 52 executes the computer program, the steps of any one of the above-mentioned embodiments of the method for diagnosing and treating autism spectrum disorders are implemented.

[0206] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0207] The embodiment of the present application further provides a computer program product, which, when executed on a mobile terminal, enables the mobile terminal to implement the steps in the above-mentioned various method embodiments.

[0208] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can at least include: any entity or device capable of carrying computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. Examples include USB flash drives, mobile hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.

[0209] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0210] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0211] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0212] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0213] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for diagnosing and treating autism spectrum disorder, characterized in that: include: Acquiring EEG signal data and eye movement thermal images of a person to be diagnosed during multiple diagnostic tasks; wherein the diagnostic tasks are used to represent an active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed; According to preset extraction conditions, feature extraction is performed on the EEG signal data and the eye movement thermal images of the plurality of diagnostic tasks to obtain EEG feature vectors and eye movement features of the plurality of diagnostic tasks; wherein, according to the preset EEG feature extraction conditions, multi-scale extraction is performed on the EEG signal data by an EEG feature extraction model to obtain the EEG feature vector corresponding to the EEG signal data, including: According to the preset EEG feature extraction conditions, the EEG signal data is subjected to sliding window processing to obtain multi-band brain topography three-dimensional data corresponding to the EEG signal data; converting the three-dimensional data of the multi-band brain topography into a set of two-dimensional segments centered on the electrode; Performing multi-scale extraction of the spatial position relationship between all electrodes in the two-dimensional segment set through a self-attention network in an EEG feature extraction model, obtaining an initial EEG feature vector corresponding to the EEG signal data; Normalizing the initial EEG feature vector to obtain the EEG feature vector corresponding to the EEG signal data; Cross-fusing the EEG feature vectors and the eye movement features of the plurality of diagnostic tasks to obtain multimodal fusion attention features of the plurality of diagnostic tasks; Using an autism spectrum disorder classification model, a plurality of the multimodal fusion attention features are classified and predicted to obtain a diagnosis result for the person to be diagnosed, wherein the autism spectrum disorder classification model is obtained by training an initial autism spectrum disorder classification model constructed based on a graph attention network; Build an initial autism spectrum disorder classification model based on graph attention network, including: Using each electrode in the EEG signal data as a node of the graph attention network; Using the multimodal fusion attention feature as a node input feature of the graph attention network and using the adjacency matrix corresponding to the EEG signal data as an edge input feature of the graph attention network; wherein the adjacency matrix is ​​generated by using the phase lag index between every two electrodes in the EEG signal data as the functional connectivity relationship; Obtaining an input layer of the graph attention network according to the nodes of the graph attention network and the edges of the graph attention network; Using a preset single-layer feedforward neural network as the graph attention layer of the graph attention network to convert the multimodal fusion attention features into high-level multimodal fusion attention features; Using a preset two-layer fully connected neural network as the output layer of the graph attention network, so as to obtain the diagnosis result according to the high-level multimodal fusion attention feature; The initial autism spectrum disorder classification model is obtained according to the input layer, the graph attention layer and the output layer of the graph attention network.

2. The method for diagnosing and treating autism spectrum disorder according to claim 1, wherein: The step of extracting features from the EEG signal data and the eye movement thermal images of the plurality of diagnostic tasks according to the preset extraction conditions to obtain eye movement features of the plurality of diagnostic tasks includes: According to preset eye movement feature extraction conditions, deep semantic feature extraction is performed on the eye movement thermal image through an eye movement heat map reconstruction model to obtain the eye movement features corresponding to the eye movement heat image.

3. The method for diagnosing and treating autism spectrum disorder according to claim 2, wherein: The step of performing deep semantic feature extraction on the eye movement thermal image using an eye movement thermal image reconstruction model according to preset eye movement feature extraction conditions to obtain the eye movement features corresponding to the eye movement thermal image includes: According to the preset eye movement feature extraction conditions, the encoder in the eye movement heat map reconstruction model performs deep semantic feature extraction on the eye movement heat map to obtain initial eye movement features corresponding to the eye movement heat map; The initial eye movement features are discretized and encoded through the discrete semantic information dictionary structure in the eye movement heat map reconstruction model to obtain the eye movement features corresponding to the eye movement heat map.

4. The method for diagnosing and treating autism spectrum disorder according to claim 3, wherein: After obtaining the eye movement features corresponding to the eye movement thermal image, the method further includes: The eye movement features are sampled and reconstructed by a decoder in the eye movement heat map reconstruction model to obtain a reconstructed eye movement heat map corresponding to the eye movement heat map; Calculating the loss between the eye movement thermographic image and the corresponding reconstructed eye movement thermographic image based on a preset loss function; According to the loss, the gap between the initial eye movement feature and the eye movement feature is reduced by a random gradient method to optimize the eye movement heat map reconstruction model.

5. The method for diagnosing and treating autism spectrum disorder according to claim 1, wherein: The classifying and predicting the plurality of multimodal fusion attention features to obtain the diagnosis result of the person to be diagnosed includes: Using an autism spectrum disorder classification model, classifying and predicting the plurality of multimodal fusion attention features of the person to be diagnosed obtained in the proactive joint attention test, and obtaining a plurality of proactive prediction results corresponding to the plurality of multimodal fusion attention features in the proactive joint attention test; The diagnosis result of the person to be diagnosed is determined based on the multiple initiative prediction results and the preset initiative characteristic level index.

6. The method for diagnosing and treating autism spectrum disorder according to claim 5, wherein: The autism spectrum disorder classification model is obtained by the following method: Obtaining a training set; wherein the training set includes a plurality of sample data and a plurality of prediction results of the sample data; Building an initial autism spectrum disorder classification model based on graph attention network; The initial autism spectrum disorder classification model is trained using the plurality of sample data and the prediction results of the plurality of sample data to obtain the autism spectrum disorder classification model.

7. The method for diagnosing and treating autism spectrum disorder according to claim 1, wherein: The step of obtaining EEG signal data and eye movement thermal images of the person to be diagnosed during multiple diagnostic tasks includes: collecting EEG signals of the person to be diagnosed in the plurality of diagnostic tasks by a high-density EEG device to obtain a plurality of initial EEG signal data of the person to be diagnosed; capturing eye movement behavior information of the person to be diagnosed during the multiple diagnostic tasks using a high-frequency eye tracking device to obtain multiple initial eye movement thermal images of the person to be diagnosed; performing a first preprocessing operation on the plurality of initial EEG signal data by using a bandpass filter and an independent component analysis method to obtain a plurality of EEG signal data corresponding to the plurality of initial EEG signal data; A second preprocessing operation is performed on the multiple initial eye movement thermal images through a Kalman filter to obtain multiple eye movement thermal images corresponding to the multiple initial eye movement thermal images.

8. A diagnosis and treatment system for autism spectrum disorder, characterized in that: include: EEG data acquisition device, eye movement data acquisition device, data processing device and model reasoning device; wherein, The EEG data acquisition device includes a high-density EEG device, an electrode cap, and a first data acquisition card; the EEG data acquisition device is used to collect initial EEG signal data of the person to be diagnosed in multiple diagnostic tasks in real time through the high-density EEG device when the electrode cap is attached to the scalp of the person to be diagnosed, and is used to convert the initial EEG signal data into digital signals through the first data acquisition card and transmit them to the data processing device; the diagnostic tasks are used to represent the active joint attention test for diagnosing autism spectrum disorder in the person to be diagnosed; The eye movement data acquisition device includes a high-frequency eye tracking device and a second data acquisition card; the eye movement data acquisition device is used to collect multiple initial eye movement thermal images of the person to be diagnosed in multiple diagnostic tasks in real time through the high-frequency eye tracking device, and is used to transmit the multiple initial eye movement thermal images to the data processing device through the second data acquisition card; The data processing device is used to receive and store the multiple initial EEG signal data transmitted from the EEG data acquisition device and the multiple initial eye movement thermal images transmitted from the eye movement data acquisition device; and is used to perform a first preprocessing operation on the multiple initial EEG signal data to obtain multiple EEG signal data; perform a second preprocessing operation on the multiple initial eye movement thermal images to obtain multiple eye movement thermal images; and is used to perform feature extraction on the EEG signal data and the eye movement thermal images of the multiple diagnostic tasks according to preset extraction conditions to obtain EEG feature vectors and eye movement features of the multiple diagnostic tasks; and is used to cross-fuse the EEG feature vectors and the eye movement features of the multiple diagnostic tasks to obtain multimodal fusion attention features of the multiple diagnostic tasks; and is used to transmit the multiple multimodal fusion attention features to the model inference device; The data processing device is further configured to: perform sliding window processing on the EEG signal data according to the EEG feature extraction conditions to obtain multi-band brain topography three-dimensional data corresponding to the EEG signal data; convert the multi-band brain topography three-dimensional data into a two-dimensional segment set centered on the electrode; perform multi-scale extraction of the spatial position relationship between all electrodes in the two-dimensional segment set using a self-attention network in an EEG feature extraction model to obtain an initial EEG feature vector corresponding to the EEG signal data; and perform normalization processing on the initial EEG feature vector to obtain the EEG feature vector corresponding to the EEG signal data. The model inference device is used to classify and predict the plurality of multimodal fusion attention features using an autism spectrum disorder classification model to obtain a diagnosis result for the person to be diagnosed, wherein the autism spectrum disorder classification model is obtained by training an initial autism spectrum disorder classification model constructed based on a graph attention network; The model inference device is also used to: use each electrode in the EEG signal data as a node of the graph attention network; use the multimodal fusion attention feature as the node input feature of the graph attention network and use the adjacency matrix corresponding to the EEG signal data as the edge input feature of the graph attention network; wherein the adjacency matrix is ​​generated by using the phase lag index between every two electrodes in the EEG signal data as the functional connectivity relationship; obtain the input layer of the graph attention network based on the nodes of the graph attention network and the edges of the graph attention network; use a preset single-layer feedforward neural network as the graph attention layer of the graph attention network to convert the multimodal fusion attention feature into a high-level multimodal fusion attention feature; use a preset two-layer fully connected neural network as the output layer of the graph attention network to obtain the diagnosis result based on the high-level multimodal fusion attention feature; obtain the initial autism spectrum disorder classification model based on the input layer, the graph attention layer and the output layer of the graph attention network.

Citation Information

Patent Citations

  • Real-time attention assessment method and system fusing multiple physiological modalities

    CN113729710A

  • Multi-modal data driven autism detection system and device and storage medium

    CN114974571A