Multi-modal fusion-based depression classification method and system

Through a multimodal fusion depression classification method, combined with electrocardiogram, electroencephalogram, facial expressions and gastrointestinal environmental signals, a deep learning model is used to achieve efficient and accurate identification and dynamic monitoring of depression status, solving the accuracy and stability problems of single-modality diagnostic methods and providing new objective diagnostic support.

CN120656721APending Publication Date: 2025-09-16SHANDONG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510778226.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the existing technology, the single-modality physiological signal diagnosis method is easily interfered with by environmental pressure and task execution status in the diagnosis of depression, resulting in unstable recognition accuracy and difficulty in distinguishing the physiological differences between patients with depression and normal people. It also lacks effective objective physiological indicator support, resulting in a long diagnosis cycle and difficulty in achieving early screening and dynamic monitoring.

Method used

A multimodal fusion depression classification method is adopted. By synchronously collecting electrocardiogram signals, electroencephalogram signals, facial expression videos and gastrointestinal environment exhaled signals, combined with multimodal signal preprocessing, multi-domain feature extraction, multimodal fusion and deep learning classification models, efficient and accurate identification and dynamic monitoring of depression status can be achieved.

Benefits of technology

It makes full use of the complementary information of each modality, effectively suppresses single signal noise and interference, improves the stability and accuracy of depression identification, and achieves a comprehensive and stable reflection of the depression state. It is suitable for long-term clinical and family monitoring, and improves the accuracy, stability and real-time performance of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656721A_ABST
    Figure CN120656721A_ABST
Patent Text Reader

Abstract

The invention discloses a depression classification method and system based on multi-modal fusion, and relates to the technical field of multi-modal data processing and intelligent identification. Comprising a data acquisition module used for acquiring an electrocardiosignal, an electroencephalogram signal, a facial expression video and a gastrointestinal environment expiration signal; the pre-processing module is used for performing pre-processing operation on the four modal signals; the data fusion module is used for performing high-order feature extraction and dimensionality reduction on the four modal signals by using different deep learning sub-networks, considering the real-time performance and the mutual relation between different modals, and performing feature fusion on the four modal signals after dimensionality reduction based on an attention soft fusion strategy; and the classification detection module is used for performing classification detection on the comprehensive features by using a deep learning classification model. According to the method, real-time signals of four modes are collected, and efficient and accurate recognition and dynamic monitoring of the depression state are achieved in combination with multi-mode signal preprocessing, multi-domain feature extraction, multi-mode fusion and a deep learning classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data processing and intelligent recognition, and in particular to a depression classification method and system based on multimodal fusion. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Depression is a common mental disorder worldwide, characterized by persistent low mood, loss of interest, and cognitive impairment, severely impacting patients' daily lives and social interactions. Currently, the diagnosis of depression relies primarily on clinical interviews and questionnaires, lacking effective objective physiological indicators. The diagnostic cycle is long, making early screening and ongoing monitoring difficult.

[0004] With the continuous development of wearable devices and intelligent sensing technology, auxiliary diagnosis based on physiological signals has gradually become a research hotspot. Currently, electrocardiogram (ECG) and electroencephalogram (EEG) as important physiological signals can reflect the functional status of the autonomic nervous system and central nervous system and are widely used in the study of biomarkers for depression. However, diagnostic methods that rely solely on ECG or EEG signals have certain limitations: these physiological signals are easily interfered with by factors such as environmental stress and task execution status, resulting in unstable recognition accuracy and difficulty in distinguishing the physiological differences between patients with depression and normal people under stress.

[0005] At the same time, facial expressions and body posture, as important external manifestations of emotional state, can reflect an individual's psychological state and emotional characteristics. Depressed individuals often exhibit facial expressions characterized by reduced expressiveness, dull eyes, and drooping mouth corners. These behavioral characteristics clearly differ from the stress-related tension and anxiety experienced by healthy individuals. Integrating computer vision technology, real-time camera-based capture and analysis of localized facial expressions and movements of the eyes and mouth provide a crucial behavioral characteristic dimension for depression identification.

[0006] Furthermore, a growing body of research indicates that the gastrointestinal environment plays a key role in mood regulation and mental health. The gastrointestinal microbiota influences neurotransmitters and immune responses through the gut-brain axis, thereby influencing the development and manifestation of depression.

[0007] Currently, multimodal depression detection primarily relies on physiological signals like electrocardiograms and electroencephalograms (EEGs). Comprehensive diagnostic solutions that integrate behavioral characteristics like facial expressions with gastrointestinal biomarkers are lacking. In addition to issues such as signal noise interference, limited recognition accuracy, and the inability to achieve real-time monitoring, single-modal data is prone to misjudgments and incorrect identifications, making it difficult to fully and consistently reflect the complex pathological and psychological state of patients with depression, limiting diagnostic effectiveness. Summary of the Invention

[0008] In response to the shortcomings of the existing technology, the purpose of the present invention is to provide a depression classification method and system based on multimodal fusion. By synchronously collecting real-time signals of four modalities, combining multimodal signal preprocessing, multi-domain feature extraction, multimodal fusion and deep learning classification models, efficient and accurate identification and dynamic monitoring of depression status can be achieved.

[0009] In order to achieve the above object, the present invention is implemented through the following technical solutions:

[0010] The first aspect of the present invention provides a depression classification system based on multimodal fusion, comprising:

[0011] The data acquisition module is used to obtain four modal signals, namely electrocardiogram (ECG) signals, electroencephalogram (EEG) signals, facial expression videos, and gastrointestinal environment exhalation signals;

[0012] A preprocessing module is used to perform preprocessing operations on the four modal signals;

[0013] The data fusion module is used to extract high-order features and reduce the dimensionality of the four modal signals using different deep learning sub-networks. Considering the real-time performance and the relationship between different modalities, the four modal signals after dimensionality reduction are fused based on the soft fusion strategy of attention to obtain comprehensive features.

[0014] The classification detection module is used to use the deep learning classification model to perform classification detection on the comprehensive features and obtain the classification results.

[0015] A second aspect of the present invention provides a depression classification method based on multimodal fusion, comprising the following steps:

[0016] Acquire four modal signals, namely, electrocardiogram (ECG) signals, electroencephalogram (EEG) signals, facial expression videos, and gastrointestinal environment exhalation signals;

[0017] Perform preprocessing operations on the four modal signals;

[0018] Different deep learning sub-networks are used to extract high-order features and reduce the dimensionality of the four modal signals. Considering the real-time performance and the relationship between different modalities, the four modal signals after dimensionality reduction are fused based on the soft fusion strategy of attention to obtain comprehensive features.

[0019] The deep learning classification model is used to classify and detect the comprehensive features to obtain the classification results.

[0020] A third aspect of the present invention provides a computer-readable storage medium storing a computer program, which is suitable for being loaded by a processor and executing the steps in the depression classification method based on multimodal fusion as described in the second aspect of the present invention.

[0021] A fourth aspect of the present invention provides a computer device, comprising:

[0022] a processor adapted to execute a computer program;

[0023] A computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the depression classification method based on multimodal fusion as described in the second aspect of the present invention is implemented.

[0024] One or more of the above technical solutions have the following beneficial effects:

[0025] This paper discloses a depression classification method and system based on multimodal fusion. This method integrates multimodal signals from four dimensions: electrocardiogram (ECG), electroencephalogram (EEG), facial expression, and gastrointestinal environment. By combining multimodal signal preprocessing, multidomain feature extraction, multimodal fusion, and a deep learning classification model, it enables efficient and accurate identification and dynamic monitoring of depression. This method fully utilizes the complementary information of each modality, effectively suppressing noise and interference from a single signal, and improving the stability and accuracy of depression identification.

[0026] This invention introduces breath biomarker detection based on stable isotope labeling, enabling comprehensive reflection of the complex physiological and behavioral characteristics of patients with depression. The integration of a dynamic attention mechanism and modality loss compensation enhances system robustness. The lightweight model design meets the requirements of real-time edge inference, making it suitable for long-term clinical and home monitoring, significantly improving detection accuracy, stability, and real-time performance.

[0027] This invention combines multimodal real-time signals, including electrocardiogram (ECG), electroencephalogram (EEG), facial expressions (eye and mouth movements captured by a camera), and gastrointestinal environment (signals acquired through air puff detection), using a multi-layered deep learning network to identify and dynamically monitor depression. This approach has significant clinical application prospects and scientific research value, overcoming the shortcomings of existing technologies and providing new technical support for the objective diagnosis and early intervention of depression.

[0028] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0030] Figure 1 This is a framework diagram of a depression classification system based on multimodal fusion in Example 1 of the present invention;

[0031] Figure 2 This is a flow chart of a depression classification method based on multimodal fusion in Example 2 of the present invention. DETAILED DESCRIPTION

[0032] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0033] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations;

[0034] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0035] Example 1:

[0036] The first embodiment of the present invention provides a depression classification system based on multimodal fusion, such as Figure 1 As shown, it includes a data acquisition module, a preprocessing module, a data fusion module and a classification detection module.

[0037] The data acquisition module is used to obtain four modal signals, namely electrocardiogram (ECG) signals, electroencephalogram (EEG) signals, facial expression videos, and gastrointestinal environment exhalation signals.

[0038] The preprocessing module is used to perform preprocessing operations on the four modal signals.

[0039] The data fusion module is used to use different deep learning sub-networks to extract high-order features and reduce the dimensionality of the four modal signals respectively. Considering the real-time performance and the relationship between different modalities, the four modal signals after dimensionality reduction are fused based on the soft fusion strategy of attention to obtain comprehensive features.

[0040] The classification detection module is used to use the deep learning classification model to perform classification detection on the comprehensive features and obtain the classification results.

[0041] The data acquisition module includes an ECG signal acquisition module, an EEG signal acquisition module, a facial expression feature acquisition module, and a gastrointestinal environment signal acquisition module, respectively used to collect the subject's ECG signal, EEG signal, facial expression video, and gastrointestinal environment exhaled gas signals. Before the test, the subject prepared according to the standard, such as remaining in a resting state and fasting for several hours to prevent interference from gastrointestinal fermentation.

[0042] ECG signal acquisition module: uses two-lead electrodes, sets the sampling frequency to 250Hz, and collects ECG signals in real time.

[0043] EEG signal acquisition module: It uses the Fp1 and Fp2 leads covering the frontal lobe and sets the sampling frequency to 250 Hz to obtain EEG activity data.

[0044] Facial expression feature acquisition module: Use a 1080p resolution, 30fps camera to focus on the eye and mouth areas to capture facial movement videos.

[0045] Gastrointestinal environment signal acquisition module: equipped with an exhalation bag with a one-way valve and a variety of real-time gas sensors, it uses non-dispersive infrared spectroscopy (NDIR), semiconductor gas sensors, proton transfer reaction mass spectrometry (PTR-MS) and other technologies to achieve quantitative detection of various gas components such as 13CO2, H2, CH4 and volatile organic compounds (VOCs).

[0046] Specifically, during the test, the subject orally ingests pre-prepared stable isotope markers to activate corresponding metabolic reactions and produce measurable exhaled biomarker signals.

[0047] It should be noted that the markers used in the present invention are all safe and stable isotope tracers that have been used in medical testing. The dosage used is far below the human body's tolerance threshold, has no toxic side effects, and the consent of the subjects has been obtained.

[0048] Specifically include:

[0049] 13C-labeled urea: Used to assess the activity of Helicobacter pylori in the stomach or the status of intestinal urea-degrading bacteria. Approximately 15 to 30 minutes after a subject ingests a dose of 13C urea, 13CO2 in exhaled breath is detected. Subjects positive for Helicobacter pylori rapidly decompose urea in their stomachs to produce 13CO2 and NH3, significantly increasing the concentration of exhaled 13CO2. By analyzing the carbon isotope ratio in exhaled breath using non-intrusive infrared spectroscopy (NDIR) or gas chromatography-mass spectrometry (GC-MS), the status of the gastrointestinal flora can be non-invasively determined.

[0050] 13C-labeled glucose: used to assess the body's sugar metabolism and pancreatic islet function. After the subjects orally take a certain dose of 13C-glucose on an empty stomach, the curve of the change of 13CO_2 concentration in their exhaled breath over time is recorded. Healthy individuals will gradually experience a 13CO_2 peak about 30 to 60 minutes after intake, and then fall back; if there are metabolic abnormalities (such as insulin resistance), the 13CO_2 generation rate will be reduced or delayed. The present invention quantifies the glucose oxidation rate of the subjects by comparing their exhaled curve characteristics (peak size, peak time, cumulative recovery rate, etc.) to assist in judging the possible tendency of metabolic syndrome in patients with depression.

[0051] 13C-labeled tryptophan: Used to detect the activity of the tryptophan-kynurenine metabolic pathway, which is closely related to depression. Subjects ingest an appropriate amount of L-[1-13C]tryptophan (e.g., 150 mg) and then collect breath samples or continuously monitor the exhaled 13CO_2 ratio at regular intervals (e.g., 15 minutes) over the next three hours. Depressed patients often experience increased IDO enzyme activity due to immune activation, which promotes the breakdown of tryptophan via the kynurenine pathway and the release of labeled 13CO_2. The detected 13CO_2 cumulative recovery rate (CRR) and peak concentration are significantly higher than those of normal subjects, providing an objective metabolic marker for depression.

[0052] Regarding the ingestion of the above-mentioned isotope markers, the 13C breath detection scheme of the present invention adopts a flexible design, which supports the detection of a single marker, and can also select a combination of all or some markers to meet different research and clinical needs.

[0053] The following are two typical combination examples:

[0054] A. Single marker detection mode:

[0055] Depending on the target, any of the following markers can be used alone for detection without other interference:

[0056] 13C-labeled urea (100 mg): Exhaled breath was collected 15–30 minutes after oral administration on an empty stomach for detection of Helicobacter pylori activity.

[0057] 13C-labeled glucose (75 mg): Sampled 30–60 minutes after oral administration to assess glucose metabolism and insulin sensitivity.

[0058] 13C-labeled tryptophan (150 mg): Samples were collected every 15 minutes until 180 minutes for measurement of tryptophan-kynurenine pathway activity.

[0059] B. Multi-marker combined detection mode:

[0060] To fully obtain information on multiple metabolic pathways, subjects can ingest multiple 13C-labeled substances in stages, with sufficient time intervals between each stage to avoid interference from metabolic residues. The specific process is as follows:

[0061] Phase I: 13C-labeled urea: The subjects took 100 mg orally on an empty stomach, and samples were collected 15-30 minutes later (for Helicobacter pylori testing).

[0062] Interval ≥ 90 minutes: Wait until 13CO2 returns to baseline before entering the next stage.

[0063] Phase II: 13C-labeled glucose: Continuous sampling 30-60 minutes after 75 mg ingestion (glucose metabolism assessment).

[0064] Again, the interval is ≥90 minutes.

[0065] Phase III: 13C-labeled tryptophan: samples were collected every 15 minutes until 180 minutes after oral administration of 150 mg (immunometabolism pathway determination).

[0066] The above design ensures that each metabolic signal does not interfere with each other, and the detection species and detection time period can be flexibly selected according to actual needs. In some other embodiments, other marker combination intake schemes can be designed.

[0067] Afterwards, a hydrogen or methane breath test is performed to assess small intestinal flora overgrowth (SIBO) or intestinal fermentation. After the subject ingests a certain amount of lactulose or glucose, the breathalyzer is used to monitor the changes in H_2 and CH_4 concentrations in the exhaled air every 10-15 minutes. If there is abnormal fermentation in the small intestine, the subject will experience a significant increase in hydrogen peak within 30-60 minutes. The incidence of SIBO may be higher in patients with depression, and the excess hydrogen and methane produced by symptomatic intestinal flora disorders can serve as another evidence for the present invention to obtain information on gut-brain axis imbalance. Real-time gas sensors (such as electrochemical sensors or semiconductor gas sensors) can be used to continuously monitor the content of these small molecule gases to achieve non-invasive and continuous intestinal metabolic assessment.

[0068] After ingestion of the aforementioned markers, the system enters the monitoring phase, utilizing specialized markers to acquire gaseous indicators reflecting metabolic characteristics associated with depression. The airflow signal acquisition device includes an exhaled breath collection device with a one-way valve to ensure the collection of deep exhaled breath (alveolar air). Subjects follow instructions to blow into the sensor at various time points, or continue breathing normally through an oral and nasal mask, allowing the system to continuously sample exhaled air.

[0069] Breath tests using stable isotope-labeled substrates reveal that when subjects ingest 13C-labeled compounds, their metabolites (especially gases produced by gastrointestinal microbiota or neurotransmitter metabolic pathways) carry 13C and are excreted in the exhaled breath, potentially serving as objective markers of depression. Studies have found significant differences in the exhaled breath composition of patients with depression compared with healthy controls. For example, concentrations of volatile organic compounds (VOCs) with mass-to-charge ratios of m / z = 88, 89, and 90 are significantly lower in patients with MDD, and their temporal patterns of change are abnormal. These differences can be exploited to discriminate between patients with depression, with AUCs for classifying subjects from healthy controls reaching 0.80–0.94. Furthermore, 13C-breath tests can quantitatively assess the activity of metabolic pathways associated with depression. A typical example is the 13C-tryptophan breath test: after oral administration of 13C-labeled tryptophan, monitoring the time-dependent changes in the exhaled 13CO_2 / 12CO_2 ratio reflects the extent of tryptophan breakdown via the kynurenine pathway. Due to factors such as inflammation, patients with depression often experience accelerated tryptophan metabolism via the IDO pathway, resulting in higher 13CO2 production than healthy individuals. Experimental results showed that this marker was significantly elevated in the MDD group compared with the control group (p < 0.01), demonstrating that the 13C-tryptophan breath test can serve as a novel biomarker for identifying depression subtypes. Similarly, the 13C-urea breath test utilizes the urease of Helicobacter pylori to break down 13C-urea to produce 13CO2, a method commonly used to detect gastrointestinal dysbiosis. Studies have shown that H. pylori infection is associated with host mood and neurotransmitter metabolism, which can indirectly influence depressive symptoms. The 13C-glucose breath test is used to assess glucose metabolism and insulin sensitivity. After oral administration of 13C-labeled glucose, the rate of exhaled 13CO2 production is monitored to infer systemic energy metabolism. The present invention innovatively integrates the above-mentioned breath biomarker detection into depression screening. By analyzing gases or VOC components such as 13CO_2, H_2, CH_4 in exhaled breath, it captures the impact of the intestinal flora-brain axis and neural metabolism on the depressive state, providing a new biological pathway evidence for the objective diagnosis of depression.

[0070] The sensing part adopts corresponding technology according to the different target analytes: for example, for 13CO_2, non-dispersive infrared spectroscopy (NDIR) or laser isotope analyzer can be used to measure the 13CO_2 / 12CO_2 ratio in real time; for H_2 / CH_4, semiconductor gas sensor array or micro-chromatograph can be used for detection; for complex VOC spectra, electronic nose technologies such as proton transfer reaction mass spectrometry (PTR-MS) or surface acoustic wave sensors can be used for analysis. This embodiment preferably uses a real-time and portable sensing means, such as an infrared / laser-based isotope CO_2 analyzer that can give a 13CO_2 concentration reading within seconds; or an integrated gas-sensitive array chip that can achieve a second-level response to a variety of VOCs. This enables the changing trends of breath biomarkers to be captured in real time: the system will store the gas concentration data at consecutive time points and transmit it to the signal processing unit to prepare for subsequent feature extraction.

[0071] This example innovatively introduces a breath detection method based on stable isotope labeling. By ingesting 13C-labeled metabolic substrates, exhaled gas indicators related to gastrointestinal flora and neurotransmitter metabolism, such as 13CO_2, H_2, CH_4, and volatile organic compounds (VOCs), are obtained. This marker reflects the activity of metabolic pathways related to depression and the regulation mechanism of the gut-brain axis, providing an important biological basis for the diagnosis of depression. The multimodal fusion process adopts time synchronization and spatial alignment strategies, combined with a dynamic weight adjustment mechanism and modality loss compensation, to achieve effective fusion and robust classification of heterogeneous signals.

[0072] The system uses multiple sensors to synchronously collect real-time data from four modalities, including the subject's physiological signals (such as electrocardiogram (ECG) and electroencephalogram (EEG), behavioral characteristics (such as video capture of facial expressions), and gastrointestinal environmental biomarkers (obtained through air blowing tests), and performs intelligent fusion analysis on the information of each modality to efficiently and objectively identify the state of depression.

[0073] The preprocessing module includes a denoising module, a feature extraction module, and a data organization module. Denoising, feature extraction, and data organization are performed on each modal signal. After each modal signal is acquired, the preprocessing phase begins to improve cleanliness and feature quality.

[0074] The denoising process of the denoising module includes bandpass filtering (0.5–40 Hz) and 50 Hz notch filtering of ECG signals, bandpass filtering and independent component analysis to remove artifacts of EEG signals, key point detection and action unit recognition of facial videos, smoothing filtering and baseline correction of exhaled gas signals.

[0075] The features extracted by the feature extraction module include but are not limited to:

[0076] ECG signal features: Heart rate variability (HRV) time domain features, standard deviation (SDNN), root mean square (RMSSD) of adjacent RR differences, frequency domain features (low-frequency power and high-frequency power (calculated via FFT), and sample entropy (SampEn) for nonlinear complexity analysis. EEG signal features: Nonlinear indices such as frequency band power, approximate entropy, and Lyapunov exponent. Facial signal features: Facial expressions are calculated by calculating blink frequency, eyelid closure, and mouth corner movement amplitude through changes in key points. Exhaled gas concentration features: Exhaled gas concentration features include peak concentration (Cpeak), time to peak (tpeak), area under the curve (AUC), and cumulative recovery rate.

[0077] The data sorting module is used to synchronize and normalize the features using a sliding time window strategy. In this embodiment, the window length is set to 60 seconds and the step length is 5 seconds.

[0078] In a specific embodiment, the preprocessing process is described according to signals of different modalities:

[0079] ECG signal preprocessing process.

[0080] Denoising: The ECG signal is bandpass filtered (e.g., 0.5–40 Hz) to remove baseline drift and power-frequency noise, and then filtered at 50 Hz to remove power-frequency interference. The filter design uses a finite impulse response (FIR) bandpass filter to ensure linear phase characteristics and maintain ECG waveform distortion.

[0081] Feature extraction process: R wave detection uses the Pan-Tompkins algorithm, which extracts peak positions based on differentiation, squaring, and moving integrals, identifies continuous RR intervals, and removes abnormal peaks and artifacts. Based on the RR interval sequence, the time domain features of heart rate variability (HRV) are calculated, including the mean RR interval:

[0082]

[0083] in, represents the mean RR interval; RR i represents the time interval between two consecutive R wave peaks detected for the i-th time; N represents the number of valid RR intervals in the sampling time window; i is the summation variable, which represents the index number of the current RR interval.

[0084] Standard Deviation SDNN:

[0085]

[0086] And the root mean square RMSSD of the adjacent RR differences:

[0087]

[0088] The frequency domain features are calculated by fast Fourier transform (FFT) to calculate the power spectral density, which is divided into low frequency (LF, 0.04 to 0.15 Hz) and high frequency (HF, 0.15 to 0.4 Hz) power. The calculation formula is:

[0089]

[0090] Where LF represents low frequency power and HF represents high frequency power. P(f) represents the power spectral density function at frequency f (unit: ms) 2 / Hz), which is obtained by fast Fourier transform (FFT) of the RR interval sequence; f is the frequency variable, the unit is Hz;

[0091] In addition, the sample entropy is used to measure the complexity of the heart rate time series and provide a quantitative indicator for nonlinear characteristics.

[0092] EEG signal preprocessing process.

[0093] After the EEG signal was filtered in the same frequency band and subjected to independent component analysis (ICA) to remove eye movement and myoelectric artifacts, short-time Fourier transform (STFT) and wavelet transform were used to obtain multi-band time-frequency features. The power of each frequency band (δ, θ, α, β, and γ bands) was calculated as an energy index, which is calculated as:

[0094]

[0095] Among them, P band Indicates the total power of the target frequency band (Powerinband), in μV 2 or μV 2 / Hz; STFT(f) represents the short-time Fourier transform (SFT) amplitude of the signal at frequency f; |STFT(f)| 2 represents the power spectral density at frequency f; f low ,f high are the lower and upper frequencies of the frequency band, respectively, in Hz. For example, the θ band can be [4,8] Hz, and the α band can be [8,13] Hz. df represents the frequency increment, which is used for integration.

[0096] The spectral entropy is calculated by normalizing the spectral probability distribution:

[0097] H=-∑ j p j log p j .

[0098] Where pj represents the normalized power value of the jth frequency band. Nonlinear dynamic features include approximate entropy (ApEn) and maximum Lyapunov exponent, which are used to evaluate the complexity and chaotic properties of EEG signals.

[0099] Facial expression feature preprocessing process.

[0100] Facial expression features are identified using a high-precision keypoint detection algorithm to locate 68 standard facial points, focusing on extracting motion data from the eyelid and mouth regions. Action unit (AU) recognition, based on the open-source tool OpenFace, calculates muscle activation probabilities and quantifies expression intensity. Specific features include blink frequency (the number of eye closure events per unit time), eyelid closure (the change in distance between keypoints on the upper and lower eyelids), mouth corner movement amplitude (vertical displacement of keypoints), and movement duration, reflecting emotional expression and depression-related behavioral characteristics.

[0101] Preprocessing process of gastrointestinal environmental signals.

[0102] The gastrointestinal environment blowing signal is filtered by moving average to remove short-term noise, and the filtering window is about 5 seconds. The peak concentration and peak occurrence time of gas components such as 13CO2, H2, and CH4 are calculated, and the area under the curve AUC is calculated as follows:

[0103]

[0104] Wherein, t0 represents the integration starting time point, which is the starting time (unit: minute) when the exhalation signal is recorded after the marker is taken in. n Indicates the integration termination time point, that is, the time point when the last breath sample was collected (unit: minute).

[0105] The cumulative recovery rate (CRR) of the exhaled breath marker is defined as:

[0106]

[0107] Where C(t) is the gas concentration at time t, C baseline is the fasting baseline concentration, and Dose is the oral marker dose. Principal component analysis (PCA) was used to extract multidimensional ion peak intensities of volatile organic compounds (VOCs). Key features were extracted for metabolic abnormality discrimination, providing a comprehensive evaluation of intestinal metabolism.

[0108] The methods for extracting features of expired gas in the gastrointestinal environment include:

[0109] The 13CO2 / 12CO2 ratio is measured using a non-dispersive infrared (NDIR) sensor and the real-time isotope ratio R(t) is calculated:

[0110]

[0111] Calculate the time series exhaled gas concentration C(t) and smooth it using a moving average filter:

[0112] in, is the exhaled gas concentration value at time t after smoothing; C(tk) represents the original exhaled gas concentration value at the first k sampling points; L represents the length of the sliding average window, that is, the number of data points used to calculate the average value (for example, 5 means using the last 5 time points for smoothing); k is the summation index variable, which represents the time interval from the current time t; the value range is 0 to N-1, where N is the total number of sampling points.

[0113] Peak concentration C peak :

[0114]

[0115] Peak time t peak :

[0116]

[0117] In the data sorting module, features from different sources are aligned according to the time axis. In this embodiment, after preprocessing and feature extraction of each modal signal, a unified feature representation is constructed through a multimodal feature fusion strategy, and input into a deep learning model for depression status discrimination. The fusion process focuses on time synchronization and spatial alignment. Therefore, the data sorting module divides the data of each modality with a fixed sliding window, and fuses the ECG, EEG, expression and exhalation features of the same time segment accordingly. And the features are mapped to a common space through appropriate transformations to ensure semantic alignment. On the one hand, time synchronization ensures that events of different signals are comparable when fused; on the other hand, spatial alignment converts the features of each modality into a fusible format through normalization or shared embedding. For example, the two-dimensional matrix features are reshaped into a multi-channel tensor input into the fusion network, or a graph embedding-based method is used to maintain the consistency of the feature space of each modality.

[0118] The specific sliding time window strategy is:

[0119] a) Each modal feature sequence is segmented into a fixed-length 60-second time window with a step size of 5 seconds to ensure high temporal resolution and continuous coverage of the data.

[0120] b) For the characteristic sequence within each time window, calculate the first-order difference sequence and the second-order difference sequence to detect the signal fluctuation rate and acceleration:

[0121] Δx(t)=x(t)-x(t-1).

[0122] Among them, x(t) represents the value of the feature sequence at the current time point t; x(t-1) represents the value of the feature at the previous sampling time point; Δx(t) represents the feature change rate, which is the increment of the feature change at the current time point; its positive or negative value reflects the rising or falling trend of the signal.

[0123] c) Define a fluctuation threshold θ. When |Δx(t)|>θ or the second-order difference exceeds the threshold, dynamically adjust the start and end boundaries of the window to cover the abnormal signal segment and improve the event capture rate.

[0124] d) The features within the window are calculated by weighted averaging, with the weights based on the modal signal-to-noise ratio (SNR). i To distribute, the eigenvector calculation formula is

[0125]

[0126] Among them, F window is the feature in the window, M is the number of modalities, f i* is the i*th modal eigenvector.

[0127] e) Use the Kalman filter to recursively update the window boundaries and feature estimates, smooth out noise interference, and ensure the stability of time series features.

[0128] All features extracted by the feature extraction module are segmented and aggregated according to a 60-second sliding time window (step size 5 seconds) and uniformly normalized to ensure that the features of each modality are synchronized in time and consistent in scale, laying the foundation for subsequent fusion.

[0129] In the data fusion module, a multimodal convolutional neural network combined with a cross-modal attention mechanism is used for feature fusion to achieve modal dynamic weight allocation and joint representation.

[0130] In a specific embodiment, the data fusion module includes an encoding module, a soft fusion module and a feature compensation module.

[0131] The encoding module is used to encode the features of each modality separately using deep neural network sub-modules, which include one-dimensional convolutional network, two-dimensional convolutional network, three-dimensional convolutional network and long short-term memory network (LSTM).

[0132] Specifically, the pre-processed features of each modality first enter their respective deep learning sub-networks for high-order feature extraction and dimensionality reduction. In this embodiment, a multimodal convolutional neural network (MM-CNN) architecture is adopted: a separate convolution / pooling branch is designed for each modality to learn the local spatiotemporal features of the modality. For example, the ECG data uses a one-dimensional convolutional network to extract the timing pattern, the EEG data is input into a two-dimensional convolutional network to process the time-frequency graph, the facial expression uses a three-dimensional convolution or a temporal convolutional network to capture dynamic expression changes, and the inflated gas concentration sequence is extracted through a one-dimensional convolution or a long short-term memory network (LSTM) to extract dynamic features. After multi-layer nonlinear mapping of each branch, each modal branch outputs a compressed feature representation vector x before the fusion layer. i* Among them, x i* ∈Rd, i*=1,2,...,M, where i* is the i*th modality, M is the number of modalities, and d is the unified feature dimension. The soft fusion module is used to dynamically calculate the weights of each modality through a multi-head self-attention mechanism and fuse the signals of each modality.

[0133] Specifically, the soft fusion module of this embodiment introduces a dynamic weighting mechanism, leveraging an attention network to adaptively assign weights to features of different modalities, thereby emphasizing the extraction of key information. For example, when facial expression features are unusual (e.g., expressionless or reduced blinking), the attention mechanism will increase the weight of the behavioral modality in the comprehensive judgment criteria; conversely, when physiological signal changes are more indicative (e.g., reduced heart rate variability or increased theta wave power), the model will adaptively focus on ECG / EEG features.

[0134] This dynamic fusion strategy effectively combines the advantages of various heterogeneous modalities and improves the robustness of depression detection. In particular, the system remains robust when some modalities are missing or the signal quality is poor. On the one hand, the deep model has been trained for modality loss and can still complete the judgment based on the remaining modalities. On the other hand, the fusion architecture designs a modality gating mechanism to i* Calculating the effectiveness score γ i ∈[0,1]. If γ i <τ, it is considered an abnormal or failed mode, and its corresponding attention weight is reset to zero. The weights of the remaining modes are renormalized to maintain a weighted sum of 1 to ensure the robustness of the fusion process. Wherein, τ is the set modality judgment threshold. In addition, this embodiment also uses a cross-modal completion strategy to compensate for modality loss. For example, the missing modal features are reconstructed through a shared latent space: after projecting each modality into a common subspace, the data of the missing modality can be reconstructed using information from other modalities, thereby compensating for the loss of sensor data when necessary.

[0135] More specifically, the fusion module aims to combine representations from different modalities into a multi-channel tensor or joint vector. A simple implementation currently involves directly concatenating the modal vectors. More advanced implementations include tensor product fusion, which computes the outer product between modalities to generate a higher-order tensor containing the interaction information between them.

[0136] However, the above existing methods have the disadvantages of poor real-time performance and do not consider the dynamic changes in the relationship between modalities. Therefore, considering the real-time and complexity, this embodiment prefers soft fusion based on attention. Therefore, the soft fusion module of this embodiment uses a cross-modal attention mechanism to model the correlation between modalities, thereby dynamically adjusting the fusion representation. The formula is:

[0137]

[0138] Among them, Attention represents the attention mechanism, d k is the scaling dimension hyperparameter of each attention head, the query matrix Q, key matrix K, and value matrix V are the modal feature matrices after linear transformation. This mechanism automatically weights and fuses the features of each modality and outputs the joint features.

[0139] To improve the model's ability to detect abnormal signals (such as sudden changes in patterns and inter-modal asynchrony), auxiliary and contrastive loss terms are added during training. These guide the model in identifying which modal combinations are most representative of depressive states and strengthen the learning of modal synergy features. For example, the model is guided to learn the synchronization between decreased HRV and apathy, or the implicit correspondence between high EEG theta and increased gastrointestinal fluctuations. The output of the soft fusion module is the comprehensive feature, which is then connected to fully connected neurons or other discriminant heads for final depression risk prediction.

[0140] In this embodiment, the feature compensation module of the data fusion module is used to compensate for the missing features of each modality caused by insufficient signal quality or task-driven selective acquisition. The specific steps include:

[0141] 1. Calculate the effectiveness index γ for each modal feature of the input i ∈[0,1], realized through the gating network, if γ i <τ, it is judged as a missing or abnormal mode, where τ is the preset threshold.

[0142] 2. Reset the weights of the missing modalities to zero and renormalize the weights of the remaining modalities to ensure that the sum of the weights is 1. The formula is:

[0143]

[0144] Among them, w' i is the residual modal weight, and M is the number of modes.

[0145] 3. Use a cross-modal latent space mapping model to predict the probability distribution of missing modal features based on non-missing modal features to achieve feature compensation. A cross-modal latent space mapping model can be a generative network based on a variational autoencoder (VAE). The compensation process formula is as follows:

[0146]

[0147] Among them, f obs is the currently observed non-missing modal feature combination, is the estimated missing feature, G θ is the trained generative model parameter, θ represents its parameter set, f miss is the true feature representation vector of the missing mode, p(f miss |f obs ) is the probability distribution of missing modal features.

[0148] 4. The feature output after comprehensive compensation in the fusion layer ensures that even if some modalities are missing, it can still output stable and highly accurate depression classification results.

[0149] The classification detection module includes a lightweight designed and deployed classification network. The comprehensive features obtained by the soft fusion module are input into the classification network to output the depression classification results.

[0150] Specifically, the comprehensive features are input into a pre-trained deep learning classification model to output the depression detection results. The classification model can be a binary classification model to determine the presence or absence of depressive symptoms, or a multi-classification depression model that classifies the depression according to severity. The classification model can be a convolutional neural network, a long short-term memory network (LSTM), a Transformer, etc., or a combination thereof. In this embodiment, in order to take into account the characteristics of multimodal time series data and the requirements of real-time reasoning, a convolution-Transformer hybrid model is selected.

[0151] Specifically, after the fusion representation, the comprehensive features are obtained, and the Transformer layer is used to capture the long-range dependencies between the features of different modalities (the attention weight can explain the importance of each modality), and finally the discrimination results are obtained through several fully connected layers. Label supervised learning of depression and healthy samples is used in the model training stage, and the generalization performance can be improved through transfer learning (for example, using the pre-trained weights of the public emotion database). After the training, the model is pruned and optimized to adapt to the resource-constrained environment of the edge device. Specifically, this embodiment adopts a variety of model compression and acceleration methods: first, redundant parameters are pruned, and convolution kernels and connection weights that have little impact on the output are deleted, thereby greatly reducing the model size (studies have shown that structured pruning can reduce the model size by 75% while the performance remains basically unchanged); secondly, quantization technology is applied to compress the model weights from 32-bit floating points to 8-bit integers, etc., thereby reducing memory usage and speeding up reasoning; thirdly, distillation training is introduced, and a high-precision teacher model guides a smaller student model to learn the distribution of key features with almost no loss of accuracy. The model is compressed to a fraction of its original size. The optimized model can be easily exported to the ONNX open neural network exchange format for cross-platform deployment. The optimized model structure consists of a convolutional-Transformer hybrid network with a fully connected output layer. The final output of the model is a depression risk score, denoted as: s∈[0,1]. Here, s represents the probability that the current sample is in a depressive state. If s>τ*, a high-risk depression state can be determined based on a threshold. τ* is the depression threshold. This score can be used as a numerical basis for real-time risk assessment or converted into a classification judgment result using a threshold, supporting clinical risk screening or dynamic trend monitoring.

[0152] With the help of efficient inference engines such as ONNX Runtime, the model can be loaded and run on mobile terminals, wearable devices, or embedded hardware. The resulting model has a small number of parameters and fast computational speed, enabling real-time depression detection inference with low latency on edge devices. For example, the pruned and quantized model size is only tens of megabytes, enabling millisecond-level inference response on smartphone SoCs, thus meeting the requirements of real-time monitoring.

[0153] The system of this embodiment also includes a feedback monitoring module for continuous data collection, online processing and immediate feedback, dynamic assessment of depression risk based on a sliding window, and continuous monitoring.

[0154] In a specific embodiment, the data acquisition module feeds each data stream into an edge computing unit (such as a portable detector or smartphone) in real time, triggering signal preprocessing and feature extraction in parallel. At the end of each sliding time window, the new multimodal feature vector is immediately generated through the fusion network and classification model to generate a prediction result of the current depression state. Due to the use of a sliding window update mechanism, the system can output a new assessment result every few seconds, achieving nearly continuous psychological state monitoring. For example, with a 5-second step and a 60-second window length, the data within the last minute is rolled out every 5 seconds to output a depression risk score. In this way, when the subject's physiological or behavioral state changes, the system will reflect the fluctuation in the risk score within seconds, providing a timely warning. The entire signal processing and inference process is completed on the local device, without the need to upload data to the cloud, thereby minimizing latency and improving data privacy and security. After testing, running the method of the present invention on a typical ARM architecture mobile processor, the latency of a single prediction can be controlled to hundreds of milliseconds, meeting the requirements of real-time applications. The system also includes a friendly user interface for feedback of detection results. For example, when a high-risk depressive state is detected in multiple windows continuously, the user or medical staff will be reminded by a prompt tone or mobile App notification. Through this low-latency closed-loop feedback, the present invention can be used for long-term continuous monitoring of changes in the state of patients with depression, and assist in clinical decision-making and personal psychological management. For example, the patient uses the device to monitor at home every day for a period of time. The system automatically generates a depression state curve and abnormal event records, and remotely transmits them to the doctor for reference, thereby achieving early warning and efficacy tracking of depression. In summary, the present invention constructs a complete real-time system architecture from signal acquisition, fusion analysis to result feedback, which can run efficiently on edge devices and provide users with objective depression detection services anytime, anywhere.

[0155] The specific implementation of the feedback monitoring module functions includes:

[0156] S1: Calculate the depression risk score of the fused features within the current sliding window at intervals (5 seconds). The risk score is represented by the probability value output by the classification model.

[0157] S2: Use the Cumulative Sum Control Chart (CUSUM) algorithm to detect mutation points in the risk score sequence. CUSUM statistic S t The calculation formula is:

[0158] S t =max(0,S t-1 +(x t -μ0-Q)).

[0159] Among them, x tis the current score, μ0 is the expected mean, and Q is the reference value or detection sensitivity coefficient, which is used to set the "tolerance" of the detection system to small changes. The initial value S0 = 0; when S t When >h(threshold), an abnormal alarm is triggered.

[0160] S3: Combine the sliding window and mutation detection results to dynamically adjust the window boundary to capture abnormal signal fluctuations and improve sensitivity.

[0161] S4: Apply Bayesian filtering to the system based on the sliding window mechanism, and evaluate the depression state of the input multimodal fusion features every fixed time step (such as 5 seconds). Smooth the time series composed of a series of continuously generated depression risk score values, calculate the posterior risk probability, and reduce the impact of noise.

[0162] S5: The early warning information is pushed to mobile terminals and local display devices in real time through a secure and encrypted wireless communication module, enabling remote and on-site simultaneous monitoring.

[0163] This embodiment synchronously collects multimodal physiological and behavioral signals, combines them with stable isotope markers to assist in the detection of gastrointestinal metabolic status, uses a sliding time window strategy to achieve signal time alignment and dynamic feature extraction, adopts a deep neural network to encode the features of each modality separately, and realizes the dynamic allocation and fusion of modal weights based on a multi-head self-attention mechanism. In response to the problem of missing modal signals, this embodiment designs a gated weight adjustment and cross-modal latent space feature compensation mechanism to ensure that the system can still stably and accurately identify the depressive state in the case of partial multimodal loss. Through a continuous rolling risk assessment and mutation detection algorithm based on a sliding window, near-real-time dynamic monitoring and early warning feedback of the depressive state are achieved. This embodiment takes into account the temporal continuity of the signal and the comprehensive utilization of multimodal information, significantly improving the accuracy and reliability of depression classification, and has strong clinical application value and promotion potential.

[0164] Example 2:

[0165] The second embodiment of the present invention provides a depression classification method based on multimodal fusion, such as Figure 2 As shown, the following steps are included:

[0166] Acquire four modal signals, namely, electrocardiogram (ECG) signals, electroencephalogram (EEG) signals, facial expression videos, and gastrointestinal environment exhalation signals;

[0167] Perform preprocessing operations on the four modal signals;

[0168] Different deep learning sub-networks are used to extract high-order features and reduce the dimensionality of the four modal signals. Considering the real-time performance and the relationship between different modalities, the four modal signals after dimensionality reduction are fused based on the soft fusion strategy of attention to obtain comprehensive features.

[0169] The deep learning classification model is used to classify and detect the comprehensive features to obtain the classification results.

[0170] Example 3:

[0171] A third embodiment of the present invention provides a computer-readable storage medium storing a computer program. The computer program is suitable for being loaded by a processor and executing the steps of the depression classification method based on multimodal fusion as described in the second embodiment of the present invention.

[0172] Example 4:

[0173] A fourth embodiment of the present invention provides a computer device, comprising:

[0174] a processor adapted to execute a computer program;

[0175] A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of the depression classification method based on multimodal fusion as described in the first embodiment of the present invention are implemented.

[0176] The steps involved in the above embodiments 2, 3 and 4 correspond to those in the method embodiment 1. For the specific implementation methods, please refer to the relevant description part of the embodiment 1.

[0177] Those skilled in the art will appreciate that the units and algorithmic steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0178] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0179] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any technical object of a person skilled in the art that can be easily conceived of within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A depression classification system based on multimodal fusion, characterized by: include: The data acquisition module is used to obtain four modal signals, namely electrocardiogram (ECG) signals, electroencephalogram (EEG) signals, facial expression videos, and gastrointestinal environment exhalation signals; A preprocessing module is used to perform preprocessing operations on the four modal signals; The data fusion module is used to extract high-order features and reduce the dimensionality of the four modal signals using different deep learning sub-networks. Considering the real-time performance and the relationship between different modalities, the four modal signals after dimensionality reduction are fused based on the soft fusion strategy of attention to obtain comprehensive features. The classification detection module is used to use the deep learning classification model to perform classification detection on the comprehensive features and obtain the classification results.

2. The depression classification system based on multimodal fusion according to claim 1, characterized in that: The data acquisition module includes an ECG signal acquisition module, an EEG signal acquisition module, a facial expression feature acquisition module and a gastrointestinal environment signal acquisition module, which are used to collect the subject's ECG signal, EEG signal, facial expression video and gastrointestinal environment exhaled gas signal respectively.

3. The depression classification system based on multimodal fusion according to claim 2, characterized in that: The specific steps of the gastrointestinal environment signal acquisition module for gastrointestinal environment signals are as follows: The subjects took the stable isotope label orally; A hydrogen or methane breath test is then performed to assess for small intestinal bacterial overgrowth or intestinal fermentation.

4. The depression classification system based on multimodal fusion according to claim 1, characterized in that: The preprocessing module includes a denoising module, a feature extraction module and a data sorting module, which perform denoising, feature extraction and data sorting operations on each modal signal collected respectively.

5. The depression classification system based on multimodal fusion according to claim 4, characterized in that: The features extracted by the feature extraction module include: ECG features of ECG signals: time domain features of heart rate variability, standard deviation, root mean square of adjacent RR differences, frequency domain features of low-frequency power and high-frequency power, and sample entropy; EEG features of EEG signals: frequency band power, approximate entropy and Lyapunov exponent; Facial features of facial signals: Facial expressions are calculated by changing key points to calculate blink frequency, eyelid closure, and mouth corner movement amplitude; Exhaled breath characteristics of gastrointestinal environmental signals: peak concentration, time to peak, area under the curve, and cumulative recovery.

6. The depression classification system based on multimodal fusion according to claim 1, characterized in that: The data fusion module includes an encoding module, a soft fusion module and a feature compensation module. The encoding module is used to encode the features of each modality separately using a deep neural network sub-module. The soft fusion module is used to dynamically calculate the weights of each modality through a multi-head self-attention mechanism and fuse the signals of each modality. The feature compensation module is used to compensate for the missing features of each modality.

7. The depression classification system based on multimodal fusion according to claim 1, characterized in that: It also includes a feedback monitoring module for continuous data collection, online processing and instant feedback, dynamically assessing depression risks based on a sliding window, and achieving continuous monitoring.

8. A depression classification method based on multimodal fusion, characterized in that: The following steps are involved: Acquire four modal signals, namely, electrocardiogram (ECG) signals, electroencephalogram (EEG) signals, facial expression videos, and gastrointestinal environment exhalation signals; Perform preprocessing operations on the four modal signals; Different deep learning sub-networks are used to extract high-order features and reduce the dimensionality of the four modal signals. Considering the real-time performance and the relationship between different modalities, the four modal signals after dimensionality reduction are fused based on the soft fusion strategy of attention to obtain comprehensive features. The deep learning classification model is used to classify and detect the comprehensive features to obtain the classification results.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the depression classification method based on multimodal fusion according to claim 8.

10. A computer device, characterized in that: a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the depression classification method based on multimodal fusion according to claim 8 is implemented.

Citation Information

Cited By

  • Mild depression dynamic prediction system based on multi-modal time series data deep learning

    CN121506501A

  • Dynamic prediction system for mild depression based on deep learning of multi-modal time series data

    CN121506501B

  • Method, system and equipment for detecting exhaled nitric oxide based on multi-source data fusion

    CN122123676A