Attention deficit hyperactivity disorder (ADHD) evaluation system and method based on deep learning
Through deep learning technology, multi-dimensional evaluation of ADHD is achieved, which solves the subjectivity and accuracy of existing evaluation methods, and provides an efficient and safe evaluation system suitable for medical, home and school environments.
Patent Information
- Application Number
- CN202510157663.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing ADHD evaluation methods are subjective, low in accuracy, long time, and single function, and cannot fully and accurately evaluate the patient's condition.
A multimodal data fusion system based on deep learning is adopted to analyze data through high-resolution video, audio and physiological signal acquisition, and a multi-stream fusion convolutional neural network is used for data analysis, combining attention mechanism and dynamic time window technology to achieve multi-dimensional evaluation, and optimize system performance through feedback adjustment.
Improves the objectivity, accuracy and efficiency of ADHD assessments, provides detailed visual reports and personalized assessment results, ensuring data security and privacy protection, and is suitable for medical institutions, homes and school environments.
Smart Images

Figure CN120496815A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and healthcare, specifically a deep learning-based system and method for assessing attention deficit hyperactivity disorder (ADHD). Traditional assessment methods are highly subjective and inaccurate. This invention utilizes deep learning technology to achieve intelligent and objective assessment, improving accuracy and efficiency. This technology belongs to the field of intelligent ADHD assessment technology. Background Art
[0002] Attention Deficit Hyperactivity Disorder (ADHD) is a common neurodevelopmental disorder characterized by inattention, hyperactivity, and impulsivity. Epidemiological studies have shown that ADHD is more common in children and may persist into adulthood. Key symptoms include difficulty concentrating on tasks, easy distraction, restlessness, hyperactivity, and impulsive behavior.
[0003] Existing ADHD assessment methods primarily rely on physicians' subjective judgment and questionnaires, which are highly subjective and inaccurate. Furthermore, the assessment process is time-consuming and inefficient. Existing ADHD assessment systems often have limited functionality and data sources, making them incapable of comprehensively and accurately assessing a patient's condition. Summary of the Invention
[0004] The development of an AI-based intelligent ADHD assessment system is of great significance. AI technology can integrate and analyze multimodal data, improving the accuracy and objectivity of assessments. Furthermore, deep learning algorithms can rapidly process large amounts of data, enhancing assessment efficiency. Recent research data demonstrates that deep learning-based assessment systems offer significant advantages in accuracy and efficiency.
[0005] This paper provides a deep learning-based attention deficit hyperactivity disorder (ADHD) assessment system and method. The system comprises the following main functional modules: a data acquisition module, a deep learning module, an assessment module, a result output module, a feedback adjustment module, and a storage module. Through multimodal data acquisition, deep learning analysis, and intelligent assessment, the system achieves objective and intelligent assessment of ADHD.
[0006] The data acquisition module is responsible for collecting multimodal data, including high-resolution video, high-quality audio, and a variety of physiological signal data. This module uses a high-frame-rate camera to capture the subject's facial expressions, eye movements, and body movements; a highly sensitive microphone to record speech; and a wearable device to collect physiological signals such as heart rate, electrodermal conduction, and electroencephalogram (EEG). This multimodal data collection provides a rich source of information for subsequent in-depth analysis.
[0007] The deep learning module is the core of the system, and it adopts an innovative multi-stream fusion convolutional neural network structure. The network contains three parallel sub-networks: visual stream, audio stream, and physiological signal stream. Each sub-network extracts features from data of a specific modality. The visual stream uses a 3D convolutional neural network to extract spatiotemporal features, the audio stream uses a 1D convolutional neural network to analyze speech features, and the physiological signal stream uses a long short-term memory network (LSTM) to capture temporal patterns. The outputs of the three sub-networks are fused through an attention mechanism to achieve effective integration of multimodal data. The network is trained in an end-to-end manner, using a large-scale dataset with ADHD diagnostic labels and optimizing network parameters through a backpropagation algorithm.
[0008] The assessment module, based on the output of the deep learning module, provides a multidimensional assessment of ADHD. This module employs a hierarchical assessment strategy, first assessing the three primary dimensions of attention, impulsivity, and hyperactivity individually, and then synthesizing these dimensions to produce an overall assessment. Dynamic time windowing technology is incorporated into the assessment process to capture behavioral characteristics across different timescales, improving the accuracy and stability of the assessment.
[0009] The results output module uses visualization technology to present assessment results in an intuitive and easy-to-understand manner. The module generates a comprehensive report containing scores, confidence levels, and detailed analysis, and provides an interactive visualization interface that allows doctors and parents to deeply explore the assessment data.
[0010] The feedback adjustment module uses a continuous learning mechanism to continuously optimize system performance based on expert feedback and long-term follow-up data. This module utilizes an incremental learning algorithm to gradually adapt to new data distributions while maintaining existing knowledge, improving the system's generalization and adaptability. The storage module utilizes a distributed storage architecture to securely and efficiently manage large-scale multimodal data and assessment results. This module implements encrypted data storage and access control, ensuring patient privacy and data security.
[0011] The innovations of the present invention are mainly reflected in the following aspects:
[0012] 1. Multimodal data fusion: Through an innovative deep learning network structure, video, audio, and physiological signal data are effectively integrated to improve the comprehensiveness and accuracy of the assessment.
[0013] 2. Personalized evaluation model: The introduction of attention mechanism and dynamic time window technology enables accurate capture and analysis of individual behavioral characteristics.
[0014] 3. Continuous learning capability: Through the feedback adjustment module, the system can continuously optimize and adapt, improving the effectiveness of long-term use.
[0015] 4. Explainable design: The result output module provides detailed evaluation basis and visual analysis, enhancing the credibility and practicality of the system.
[0016] 5. Security and privacy protection: Advanced encryption and access control technologies are used to ensure the security of sensitive medical data.
[0017] This invention significantly improves the objectivity, accuracy, and efficiency of ADHD assessments through deep learning technology and multimodal data analysis, providing strong technical support for clinical diagnosis and scientific research. The system's modular design and scalability lend it broad application prospects, extending beyond healthcare institutions to include homes and schools, offering new possibilities for early screening and intervention for ADHD.
[0018] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory and do not limit the present disclosure. For better understanding and implementation, the present disclosure is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a structural block diagram of an ADHD assessment system based on deep learning according to an embodiment of the present disclosure; Figure 2 A flowchart of an ADHD assessment method according to an embodiment of the present disclosure is shown; Figure 3 A schematic diagram of a deep learning model training process illustrating an embodiment of the present disclosure; Figure 4 A schematic diagram of a multimodal data fusion process illustrating an embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0021] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0022] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the term "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0023] Example 1: System hardware composition
[0024] The deep learning-based ADHD assessment system of this invention utilizes advanced hardware to achieve high-precision, multimodal data acquisition. The system's primary hardware components include a high-resolution camera, a high-sensitivity microphone array, a multifunctional wearable device, and a high-performance computing server.
[0025] First, the system utilizes a 4K UHD camera equipped with a Sony IMX586 CMOS image sensor, boasting 48 megapixels and a frame rate of up to 60 fps. This camera offers excellent low-light performance and high dynamic range, enabling it to capture subtle facial expressions and body movements in a variety of lighting conditions. The camera is also equipped with advanced optical image stabilization technology, ensuring stable and clear images even when the subject is moving.
[0026] Secondly, the system utilizes a microphone array consisting of eight Shure SM58 high-fidelity microphones. Each microphone has a frequency response range of 50Hz-15kHz and a sensitivity of -54.5dBV / Pa. The microphone array utilizes beamforming technology to effectively suppress background noise and accurately capture the subject's voice information.
[0027] Furthermore, the system is equipped with an Empatica E4 wearable device for collecting physiological signal data. The E4 features a photoplethysmography (PPG) sensor, a 3-axis accelerometer, a electrodermal activity (EDA) sensor, and a skin temperature sensor. The PPG sensor has a sampling rate of 64Hz and can measure heart rate and blood oxygen saturation; the accelerometer has a sampling rate of 32Hz and is used to record movement information; the EDA sensor has a sampling rate of 4Hz and is used to measure skin conductivity; and the temperature sensor has a sampling rate of 4Hz and an accuracy of ±0.2°C.
[0028] Finally, the system uses the NVIDIA DGX A100 AI computing server for data processing and model training. The DGX A100 is equipped with eight NVIDIA A100 Tensor Core GPUs, each with 40GB of HBM2 memory, for a total of 320GB of GPU memory. The server is also equipped with two AMD EPYC 774264-core CPUs, for a total of 128 CPU cores, and 1TB of DDR4-3200 memory. This powerful computing power ensures the system can process large amounts of multimodal data in real time and quickly train complex deep learning models.
[0029] according to Figure 1 As shown in the system structure block diagram, the data acquisition module (S101) is responsible for collecting the above-mentioned multimodal data. Specifically, the camera collects video data, the microphone array collects audio data, and the wearable device collects physiological signal data. This data is then transmitted to the deep learning module (S102) for analysis. The evaluation module (S103) evaluates ADHD based on the analysis results, and the result output module (S104) presents the evaluation results in a visual format. The feedback adjustment module (S105) continuously optimizes system performance based on the evaluation results and expert feedback, while the storage module (S106) is responsible for securely storing all data and results.
[0030] Example 2: Deep Learning Model Network Structure
[0031] This paper uses an innovative multimodal deep learning model to accurately assess ADHD. This model integrates convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and attention mechanisms to effectively process multimodal data such as video, audio, and physiological signals.
[0032] The overall architecture of the model is as follows:
[0033] 1. Video processing branch: 3D convolutional neural network (3D-CNN) is used to extract spatiotemporal features.
[0034] 2. Audio processing branch: Use 1D convolutional neural network (1D-CNN) to extract audio features.
[0035] 3. Physiological signal processing branch: Bidirectional long short-term memory network (Bi-LSTM) is used to process time series data.
[0036] 4. Multimodal fusion layer: Use the attention mechanism to perform weighted fusion of different modal features.
[0037] 5. Classification layer: A fully connected layer is used for the final ADHD assessment.
[0038] The specific network structure is as follows:
[0039] 1. Video processing branch (3D-CNN):
[0040] – Input: 224x224x16 video clip (16 frames)
[0041] –5 3D convolutional layers, each followed by batch normalization and ReLU activation function
[0042] ––3 3D max pooling layers
[0043] 1 global average pooling layer
[0044] – Output: 2048-dimensional feature vector
[0045] 1. Audio processing branch (1D-CNN):
[0046] – Input: 16000x1 audio signal (1 second)
[0047] – 4 1D convolutional layers, each followed by batch normalization and ReLU activation function
[0048] –4 1D max pooling layers
[0049] – 1 global average pooling layer
[0050] – Output: 1024-dimensional feature vector
[0051] 2. Physiological signal processing branch (Bi-LSTM):
[0052] – Input: 128x4 physiological signal sequence (128 time steps, 4 types of signals) – 2 layers of bidirectional LSTM, 128 hidden units per layer
[0053] – Output: 256-dimensional feature vector
[0054] 3. Multimodal fusion layer (attention mechanism):
[0055] – Input: Feature vectors from three branches
[0056] – Self-attention layer: calculates the importance weight of each modality feature
[0057] – Weighted summation: Fusion of features from different modalities
[0058] – Output: 3328-dimensional fused feature vector
[0059] 4. Classification layer:
[0060] – 2 fully connected layers with 1024 and 256 neurons respectively, using ReLU activation function
[0061] –Dropout layer with a dropout rate of 0.5
[0062] – Output layer: 3 neurons, corresponding to the three subtypes of ADHD (attention-deficit, hyperactivity-impulsive, and mixed)
[0063] The mathematical representation of the model is as follows:
[0064] For the video processing branch, the output of 3D-CNN can be expressed as:
[0065] F v =CNN3D(V)
[0066] Where V is the input video, F v are the extracted video features.
[0067] For the audio processing branch, the output of 1D-CNN can be expressed as:
[0068] F a =CNN1D(A)
[0069] Where A is the input audio, F a is the extracted audio feature.
[0070] For the physiological signal processing branch, the output of Bi-LSTM can be expressed as:
[0071] F p =BiLSTM(P)
[0072] Where P is the input physiological signal sequence, F p is the extracted physiological signal feature.
[0073] In the multimodal fusion layer, the attention mechanism can be expressed as:
[0074]
[0075] where α i is the attention weight of each modality, F fused It is the fusion feature.
[0076] Finally, the output of the classification layer can be expressed as:
[0077]
[0078] where y is the final ADHD assessment result.
[0079] During training, the model uses the cross entropy loss function and the Adam optimizer, with a learning rate of 0.0001 and a batch size of 32. In addition to Dropout, L2 regularization is also used to prevent overfitting, and the weight decay coefficient is set to 0.0001.
[0080] according to Figure 3 The deep learning model training process shown in FIG3 first prepares the data (S301), including data cleaning, enhancement and standardization. Then the model described above is constructed (S302) and the training process is started (S303). The model performance is verified on the validation set (S304), and the hyperparameters are adjusted according to the verification results (S305) until the model performance meets the requirements and the training is completed (S306). This multimodal deep learning model can make full use of the information in the video, audio and physiological signals to achieve a comprehensive and accurate assessment of ADHD. Through the attention mechanism, the model can adaptively adjust the importance of different modal data, thereby improving the robustness and generalization ability of the evaluation. Example 3: Multimodal data processing process
[0081] The present invention uses an advanced multimodal data processing pipeline to ensure high-quality feature extraction from video, audio, and physiological signals. This pipeline involves complex signal processing and feature engineering techniques, laying a solid foundation for subsequent deep learning analysis.
[0082] 1. Video Processing
[0083] Video processing mainly involves facial expression recognition and body movement analysis. The processing flow is as follows:
[0084] a) Frame extraction: Frames are extracted from 4K videos at 30fps using the OpenCV library.
[0085] b) Face detection: We use the RetinaFace algorithm for face detection, which achieves an average accuracy of 96.5% on the WIDER FACE dataset. c) Facial landmark detection: PFLD (Practical Facial Landmark Detection) algorithm is used to locate 68 facial key points with an average error of less than 3 pixels. d) Expression recognition: Based on the extracted facial features, the VGGFace2 pre-trained model is used for transfer learning to recognize seven basic expressions (neutral, happy, sad, surprised, fear, disgust, and angry) with an accuracy rate of 93.7%. e) Pose Estimation: The OpenPose algorithm is used for full-body pose estimation, which can identify 18 key points and achieve a mean average precision (mAP) of 75.6%. f) Action classification: Based on posture sequences, a temporal convolutional network (TCN) is used for action classification, which can identify 10 common ADHD-related actions, such as fidgeting and pacing, with an accuracy rate of 89.3%. Video feature extraction can be expressed as: F v =f action (f pose (V))+f expression (f face (V)) Among them, f face 、f expression 、f pose and f action They represent face detection, expression recognition, posture estimation and action classification functions respectively. 2. Audio Processing Audio processing mainly focuses on the extraction of speech features and non-language sound features: a) Preprocessing: Use librosa library for audio resampling (16kHz) and segmentation (20ms frame length, 10ms step size). b) Voice Activity Detection (VAD): Using WebRTC's VAD algorithm, the accuracy rate reaches 98.2%. c) Fundamental frequency extraction: The improved YIN algorithm is used to extract the fundamental frequency with an average error of less than 2Hz. d) Mel-frequency cepstral coefficient (MFCC) extraction: Calculate 13-dimensional MFCC features and their first-order and second-order differences, a total of 39 dimensions. e) Pitch, loudness, and timbre features: Extracts perception-based features including pitch (F0), loudness (loudness contour), and timbre (spectral centroid, spectral flatness). f) Speech emotion recognition: Based on the above features, support vector machine (SVM) is used for speech emotion classification, which can identify 6 emotional states with an accuracy rate of 85.6%. Audio feature extraction can be expressed as: F a =[f MFCC (A),f pitch (A),f loudness (A),f timbre (A),f emotion (A)] Among them, f MFCC 、f pitch 、f loudness 、f timbre and f emotion They represent MFCC, pitch, loudness, timbre feature extraction and emotion recognition functions respectively. 3. Physiological signal processing Physiological signal processing involves analysis of heart rate variability (HRV), galvanic skin response (EDA), body temperature, and motion data: a) Heart rate variability analysis: Uses the Pan-Tompkins algorithm for R-wave detection with an accuracy of 99.8%. Calculate time domain features (SDNN, RMSSD, pNN50) and frequency domain features (LF, HF, LF / HF ratio). HRV feature extraction function: f HRV (P PPG ) b) Galvanic Skin Response Analysis: Separate skin conductance level (SCL) and skin conductance response (SCR) using the cvxEDA algorithm. Extracted features include SCL mean, SCR amplitude, SCR frequency, etc. EDA feature extraction function: f EDA (P EDA ) c) Body temperature analysis: Calculate average temperature, temperature change rate, and temperature fluctuation. Body temperature feature extraction function: f temp (P temp ) d) Sports data analysis: Denoise the acceleration data using wavelet transform. Calculate characteristics such as activity intensity, activity frequency, and posture changes. Motion feature extraction function: f motion (P acc ) Physiological signal feature extraction can be expressed as: F p =[f HRV (P PPG ),f EDA (P EDA ),f temp (P temp ),f motion (P acc )] according to Figure 4In the multimodal data fusion process shown, the system first collects multimodal data (S401) and then performs the aforementioned preliminary processing (S402). The feature extraction step (S403) corresponds to the detailed feature extraction process described above. Data fusion (S404) uses an attention mechanism to perform a weighted fusion of features from different modalities. Fusion result analysis (S405) includes feature importance analysis and correlation analysis. Finally, the system outputs the fusion results (S406), providing comprehensive data support for ADHD assessment. Example 4: Evaluation rules and scoring method The present invention uses an innovative ADHD assessment rule and scoring method that combines DSM-5 diagnostic criteria with a data-driven machine learning approach. The assessment process considers multiple dimensions, including attention, impulsivity, hyperactivity, and executive function, and is standardized according to age and gender. 1. Evaluation dimensions and characteristics a) Attention dimension: Sustained attention: Video analysis measures gaze duration and distraction frequency. Selective attention: Conversational response speed and accuracy based on audio analysis. Distraction: Heart rate variability indicators using physiological signal analysis. b) Impulsiveness dimension: Behavioral impulsivity: Frequency of sudden movements in video analysis. Cognitive impulsivity: Speech interruption and overlap frequencies in audio analysis. Emotional impulsivity: An index of mood swings based on facial expressions and vocal emotion. c) Hyperactivity dimension: Total activity: The intensity and frequency of activity from accelerometer data. Restless behavior: The frequency of behaviors such as fidgeting and pacing in the video analysis. Excessive speech: Speech rate and duration in audio analysis. d) Executive function dimension: Working Memory: A memory score based on task completion. Inhibitory control: an indicator of performance on the Go / No-Go task. Cognitive flexibility: task switching efficiency and error rates. 2. Scoring criteria and threshold setting For each dimension, we use the Z-score standardization method to convert the raw scores into standard scores: Where X is the raw score, μ is the mean score of the age- and gender-matched population, and σ is the standard deviation. The threshold is set as follows: Z ≥ 2.0: severe symptoms (3 points) 1.5≤Z<2.0: Moderate symptoms (2 points) 1.0≤Z<1.5: Mild symptoms (1 point) Z < 1.0: Normal range (0 points) 3. ADHD subtype classification Based on the diagnostic criteria of DSM-5, we defined the classification rules for ADHD subtypes: a) Inattention: Attention dimension score ≥ 6, and impulsivity and hyperactivity dimension scores < 6 b) Hyperactive-Impulsive Type: Impulsiveness or Hyperactivity score ≥ 6, and Attention score < 6 c) Mixed type: Attention dimension score ≥ 6, and impulsivity or hyperactivity dimension score ≥ 6 d) Non-ADHD: All dimensions score < 6 4. Evaluate Credibility Calculation In order to improve the reliability of the evaluation results, we introduced the evaluation credibility indicator: C=α·C data +β·C model +γ·C consistency in: ·C data : Data quality indicators, considering signal-to-noise ratio, data integrity, etc. ·C model : Model confidence, based on the probability value of softmax output. ·C consistency : Consistency index for multiple evaluations. α, β, γ: weight coefficients, satisfying α + β + γ = 1. 5. Evaluation result output The evaluation report output by the system includes the following contents: ADHD subtype diagnosis results Standardized scores and symptom severity for each dimension Assess credibility Detailed analysis of behavioral and physiological indicators Targeted intervention recommendations according to Figure 2 In the ADHD assessment process shown in Figure 2, after completing data collection (S202) and preprocessing (S203), the system uses a deep learning model for analysis (S204). The assessment results are calculated (S205) based on the aforementioned scoring criteria and classification rules. Finally, the system generates a detailed assessment report, concluding the assessment process (S206). Through this multi-dimensional, data-driven assessment method, the present invention provides more objective and comprehensive ADHD assessment results than traditional questionnaires. The system's scoring criteria take age and gender factors into account, improving assessment accuracy. Furthermore, the inclusion of assessment credibility makes the results more reliable, providing strong support for clinical diagnosis and intervention. Example 5: Practical application of the system and effect verification The deep learning-based ADHD assessment system presented in this paper has demonstrated excellent performance and efficiency in practical applications. The following details the system's usage and validation results. 1. System application scenarios This system is mainly used in the following scenarios: a) Child Health Centers: Screening children aged 5-12 for ADHD as part of routine physicals. b) School Counseling Rooms: Helping teachers identify students at risk for ADHD and providing early intervention recommendations. c) Mental Health Clinics: Assisting clinicians in diagnosing ADHD and providing objective, quantitative assessment data. d) Home Environments: Providing parents with daily ADHD symptom monitoring tools through portable devices. 2. Usage according to Figure 2 The ADHD assessment method flow shown is as follows: S201. Start evaluation: The user starts the evaluation process through the system interface. S202. Data Collection: a) Video Data: Use a high-definition camera to record the subject's performance during a standard task (such as a continuous performance test) for 15 minutes. b) Audio Data: Use a microphone array to collect the subject's voice data during the task. c) Physiological Signals: The subject wears an E4 wristband to collect physiological data such as heart rate and electrodermal activity. S203. Data preprocessing: The system automatically preprocesses the collected multimodal data, including noise removal and data standardization. S204. Deep Learning Analysis: The preprocessed data is fed into a pretrained deep learning model for analysis. The model outputs quantitative results for various behavioral and physiological indicators. S205. Obtain evaluation results: The system calculates the scores of each dimension of ADHD and the overall evaluation results based on the model output and the preset scoring criteria. S206. End of assessment: The system generates an assessment report, including ADHD risk level, detailed symptom analysis and intervention recommendations. 3. Effect verification experiment To verify the effectiveness of the system, we conducted large-scale clinical validation experiments.
[0086] Experimental design: • Sample size: 1,000 children aged 5-12 years, including 200 clinically diagnosed ADHD patients and 800 normal controls.
[0087] • Assessment methods: Each child received (1) traditional questionnaire assessment, (2) face-to-face consultation with a professional doctor, and (3) this system assessment.
[0088] • Evaluators: 5 child psychiatrists with experience in ADHD diagnosis.
[0089] • Blind method: The doctor is unaware of the subject's clinical diagnosis and the evaluation results of this system.
[0090] Experimental results: a) Comparison of diagnostic accuracy: • Traditional questionnaire: 82.5% • Doctor’s face-to-face consultation: 90.3% • This system: 94.7% b) Evaluation time comparison: • Traditional questionnaire: average 45 minutes • Doctor consultation: average 60 minutes • This system: average 20 minutes c) ADHD subtype classification accuracy: • Attention deficit disorder: 93.2% • Hyperactive-Impulsive: 92.8% • Mixed: 95.6% d) Consistency of symptom severity assessment (correlation coefficient with clinical scores): • Attention dimension: r = 0.89 • Impulsiveness dimension: r = 0.87 • Hyperactivity dimension: r = 0.91 e) False positive rate and false negative rate: • False positive rate: 3.2% • False negative rate: 2.1% f) Assessment credibility: The average assessment credibility is 0.92 (out of a maximum score of 1), with 95% of the assessment results having a credibility higher than 0.85.
[0091] g) User satisfaction survey: • 93% of parents think the system is easy to use • 97% of doctors believe that the information provided by the system is helpful for diagnosis • 89% of teachers said the system helped them better understand student behavior System Advantage Analysis Based on experimental results, this system has the following advantages over traditional methods: a) Higher Accuracy: The system's 94.7% diagnostic accuracy significantly surpasses traditional questionnaires and in-person doctor consultations. This is due to the deep learning model's comprehensive analysis of multimodal data, which can capture subtle behavioral patterns that are difficult for humans to perceive.
[0092] b) Higher efficiency: The average assessment time is only 20 minutes, which greatly improves screening efficiency. This makes large-scale ADHD screening possible, facilitating early detection and intervention.
[0093] c) Objective quantification: The system provides quantitative evaluation results based on objective data, avoiding the influence of human subjective factors and improving the consistency and repeatability of the evaluation.
[0094] d) Rich information: In addition to diagnostic results, the system also provides detailed behavioral and physiological indicator analysis, providing strong support for clinical diagnosis and personalized treatment plan formulation.
[0095] e) Continuous Monitoring: The system can easily conduct repeated assessments, which is beneficial for monitoring changes in ADHD symptoms and treatment effects.
[0096] f) Strong adaptability: Through continuous data collection and model updates, the system can continuously improve its performance and adapt to different populations and cultural backgrounds.
[0097] Innovative algorithm examples This system introduces several innovative algorithms in ADHD assessment, one of which is: Multi-modal Attention Fusion (MAF) This algorithm aims to adaptively adjust the importance of different modal data in the final decision. Its core idea is to achieve dynamic fusion between modalities by learning the attention weight of each modality.
[0098] The mathematical expression of the algorithm is as follows: Given video features ,Audio features , and physiological signal characteristics , the steps of the MAF algorithm are as follows: Calculate the self-attention for each modality:
[0099]
[0100] Calculate cross-attention between modalities:
[0101] Fusion of self-attention and cross-attention:
[0102] Calculate modal fusion weights:
[0103] Get the final fusion features:
[0104] in, , , , , , , is the learnable parameter matrix, is the bias vector, is the scaling factor of the attention mechanism.
[0105] This algorithm can effectively capture the complex interactions between different modalities of data and adaptively adjust the importance of each modality in different situations, thereby improving the accuracy and robustness of ADHD assessment.
[0106] In practical applications, the MAF algorithm enables the system to flexibly adjust its reliance on video, audio, and physiological signals based on the characteristics of different subjects. For example, for children who are not good at speaking, the system will automatically increase the weight of video and physiological signals; while for children with limited mobility, the system will rely more on audio and physiological signals.
[0107] Through the detailed examples described above, the present invention demonstrates its superior performance and wide applicability in practical applications. The system not only significantly improves the accuracy and efficiency of ADHD assessments but also provides rich, objective data support for clinical practice. This deep learning-based intelligent assessment system represents a major breakthrough in ADHD diagnosis and is expected to be widely adopted in the future, providing timely and accurate diagnosis and intervention support for more children at risk of ADHD.
Claims
1. A deep learning-based attention deficit hyperactivity disorder (ADHD) assessment system and method, characterized in that: It includes a data acquisition module, a deep learning module, an evaluation module, a result output module, a feedback adjustment module and a storage module; among them, the data acquisition module is responsible for collecting high-resolution video data, high-quality audio data and a variety of physiological signal data, using a high-frame rate camera to capture the subject's facial expressions, eye movements and body movements, using a high-sensitivity microphone to record voice, and collecting physiological signals such as heart rate, skin electricity, and brain electricity through wearable devices.
2. The deep learning-based attention deficit hyperactivity disorder assessment system according to claim 1, characterized in that: The deep learning module adopts an innovative multi-stream fusion convolutional neural network structure, which includes three parallel sub-networks: visual stream, audio stream, and physiological signal stream. The visual stream uses a 3D convolutional neural network to extract spatiotemporal features, the audio stream uses a 1D convolutional neural network to analyze speech features, and the physiological signal stream captures timing patterns through a long short-term memory network. The outputs of the three sub-networks are fused through an attention mechanism.
3. The evaluation module in the deep learning-based attention deficit hyperactivity disorder evaluation system according to claim 2, characterized in that: A multi-dimensional assessment of ADHD is achieved based on the output of the deep learning module. A hierarchical assessment strategy is adopted to first evaluate the three main dimensions of attention, impulsivity, and hyperactivity separately, and then comprehensively obtain the overall assessment results. Dynamic time window technology is introduced into the assessment process to capture behavioral characteristics at different time scales.
4. The evaluation module according to claim 3, characterized in that For the attention dimension, sustained attention is calculated through video analysis to calculate the gaze time and distraction frequency, selective attention is based on the conversation response speed and accuracy of audio analysis, and divided attention uses the heart rate variability indicator of physiological signal analysis.
5. The result output module in the deep learning-based attention deficit hyperactivity disorder assessment system according to claim 4, characterized in that: Visualization technology is used to present assessment results in an intuitive and easy-to-understand manner, generating comprehensive reports containing scores, confidence levels, and detailed analysis. An interactive visualization interface is provided, allowing doctors and parents to explore the assessment data in depth.
6. The feedback adjustment module in the deep learning-based attention deficit hyperactivity disorder assessment system according to claim 5, characterized in that: Through a continuous learning mechanism, the system performance is continuously optimized based on expert feedback and long-term follow-up data. Using an incremental learning algorithm, the system gradually adapts to the new data distribution while maintaining the original knowledge, thereby improving the system's generalization ability and adaptability.
7. The storage module in the deep learning-based attention deficit hyperactivity disorder assessment system according to claim 6, characterized in that: A distributed storage architecture is used to safely and efficiently manage large-scale multimodal data and evaluation results, implement data encryption storage and access control, and ensure patient privacy and data security.
8. The deep learning-based attention deficit hyperactivity disorder (ADHD) assessment system and method according to claim 7, characterized in that: It includes data acquisition, deep learning analysis, evaluation, result output, feedback adjustment and storage steps. The data acquisition module collects multimodal data and transmits it to the deep learning module for analysis. The evaluation module evaluates ADHD based on the analysis results. The result output module presents the evaluation results. The feedback adjustment module optimizes system performance. The storage module stores data and results.
9. The method for assessing attention deficit hyperactivity disorder based on deep learning according to claim 8, characterized in that: It uses innovative ADHD assessment rules and scoring methods, combining DSM-5 diagnostic criteria and data-driven machine learning methods, taking into account multiple dimensions of attention, impulsivity, hyperactivity, and executive function, and standardizing them according to age and gender.
10. The deep learning-based attention deficit hyperactivity disorder assessment method according to claim 9, characterized in that: Evaluation credibility indicators are introduced into the evaluation process. Through calculation of data quality indicators, model confidence and multiple evaluation consistency indicators, the reliability of evaluation results is improved, providing strong support for clinical diagnosis and intervention.