A method and device for early screening of parkinson's disease based on multi-modal physiological signals

CN117442159BActive Publication Date: 2026-09-08HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311310817.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-10
Publication Date
2026-09-08
Estimated Expiration
2043-10-10

AI Technical Summary

Technical Problem

[0008]1)传统的帕金森检测方法依赖于临床医生的主观判断,存在主观性和不稳定性,并且需要大量的时间和成本

Benefits of technology

[0021] Compared with the prior art, the advantages of the present invention are that the provided method for early screening of Parkinson's disease based on multimodal physiological signals, based on multimodal multiscale deep learning, combined with ECG and respiratory signals, realizes automated detection of Parkinson's disease and improves the accuracy and robustness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117442159B_ABST
    Figure CN117442159B_ABST
Patent Text Reader

Abstract

The application discloses a Parkinson early screening method and device based on multi-modal physiological signals. The method comprises the following steps: obtaining a target's respiratory signal and electrocardiogram signal; inputting the respiratory signal and electrocardiogram signal into a trained deep learning model to obtain a Parkinson screening result. The deep learning model comprises a first time feature extraction module, a second time feature extraction module, a first global feature extraction module, a second global feature extraction module and a classification module. The first time feature extraction module extracts different scale features of the respiratory signal, the second time feature extraction module extracts different scale features of the electrocardiogram signal and performs feature fusion, the first global feature extraction module is used for extracting global features of the respiratory signal, the second global feature extraction module is used for extracting global features of the electrocardiogram signal, and the classification module is used for outputting the Parkinson screening result. The application improves the reliability and robustness of Parkinson detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical physiological signal analysis technology, and more specifically, to a method and device for early screening of Parkinson's disease based on multimodal physiological signals. Background Technology

[0002] Parkinson's disease (PD) is a neurodegenerative disorder. Characteristic features of PD include the loss of neurons in specific regions of the substantia nigra and widespread intracellular protein accumulation. While PD is caused by the progressive degeneration and death of dopamine-producing neurons in the substantia nigra of the brain, the combination of these two major neuropathologies is specific for a definitive diagnosis of idiopathic PD.

[0003] From the early stages of the disease, various other cell types throughout the central and peripheral autonomic nervous systems are also affected. For example, neurons in the substantia nigra of the brain are lost. These neurons are responsible for producing a chemical called dopamine. Dopamine helps transmit information from the substantia nigra to other parts of the body, thus controlling movement. This degenerative process leads to decreased dopamine levels, resulting in motor symptoms commonly associated with Parkinson's disease (PD), including tremor, rigidity, and bradykinesia.

[0004] In addition to motor symptoms, most Parkinson's patients also experience non-motor symptoms. These non-motor symptoms involve a variety of functions, including sleep-wake cycle regulation disorders, cognitive impairment (including frontal lobe executive dysfunction, memory retrieval deficits, dementia, and hallucinations), mood and affective disturbances, autonomic dysfunction, and sensory symptoms and pain. Some of these symptoms can appear years or even decades before the onset of classic motor symptoms. Non-motor symptoms are increasingly prevalent throughout the disease process and are a major determinant of quality of life and overall disability progression.

[0005] Currently, there are multiple methods for diagnosing Parkinson's disease (PD). Regarding motor symptoms, clinicians can use the Movement Disorders Association Unified Parkinson's Disease Rating Scale (MDS-UPDRS) to assess PD progression and analyze gait differences between PD patients and healthy individuals to evaluate disease severity and progression. However, diagnosis based on motor symptoms may not detect PD early and is time-consuming. Furthermore, PD-related motor symptoms typically appear several years after onset, making them unsuitable for early screening. In addition, because the assessment of motor symptoms is subjective, multiple consultations with doctors or specialists are required before a definitive diagnosis is obtained.

[0006] Besides directly understanding the patient's medical history, imaging techniques can diagnose Parkinson's disease by providing objective markers of neurodegenerative changes. Examples include neuroimaging techniques such as magnetic resonance imaging (MRI), functional magnetic resonance imaging (fMRI), single-photon emission computed tomography (SPECT), and transcranial ultrasound (TCS). DaTscan is an imaging technique that diagnoses PD by measuring dopamine transporter levels in the brain. Due to the various biochemical changes associated with PD, biochemical markers including cerebrospinal fluid and α-synuclein are also used in the diagnosis of PD.

[0007] Analysis reveals the following shortcomings in existing technologies:

[0008] 1) Traditional Parkinson's disease detection methods rely on the subjective judgment of clinicians, which is subjective and unstable, and requires a lot of time and cost.

[0009] 2) Existing automated Parkinson's disease detection methods typically use only a single biological signal, such as ECG (electrocardiogram) or EEG (electroencephalogram) signals, which do not provide sufficient information when used alone.

[0010] 3) Existing automated Parkinson's disease detection methods typically use a single scale or feature extraction method, which fails to capture key feature information.

[0011] 4) Existing automated Parkinson's disease detection methods typically use shallow machine learning models, which cannot handle large amounts of complex biosignal data, thus limiting the accuracy and robustness of the detection. Summary of the Invention

[0012] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and device for early screening of Parkinson's disease based on multimodal physiological signals.

[0013] According to a first aspect of the present invention, a method for early screening of Parkinson's disease based on multimodal physiological signals is provided. The method includes the following steps:

[0014] Acquire the target's respiratory and electrocardiogram signals;

[0015] The respiratory and electrocardiogram signals are input into a trained deep learning model to obtain Parkinson's screening results;

[0016] The deep learning model includes a first temporal feature extraction module, a second temporal feature extraction module, a first global feature extraction module, a second global feature extraction module, and a classification module. The first temporal feature extraction module is used to extract features of different scales of respiratory signals. The second temporal feature extraction module is used to extract features of different scales of electrocardiogram (ECG) signals and fuse the different features. The first global feature extraction module is used to extract global features of respiratory signals using a self-attention mechanism. The second global feature extraction module is used to extract global features of ECG signals using a self-attention mechanism. The classification module is used to combine the features of respiratory signals and ECG signals to output Parkinson's screening results.

[0017] According to a second aspect of the present invention, a Parkinson's disease early screening device based on multimodal physiological signals is provided. The device comprises:

[0018] Signal acquisition unit: used to acquire the target's respiratory signals and electrocardiogram signals;

[0019] Prediction unit: used to input the respiratory signal and electrocardiogram signal into a trained deep learning model to obtain Parkinson's screening results;

[0020] The deep learning model includes a first temporal feature extraction module, a second temporal feature extraction module, a first global feature extraction module, a second global feature extraction module, and a classification module. The first temporal feature extraction module is used to extract features of different scales of respiratory signals. The second temporal feature extraction module is used to extract features of different scales of electrocardiogram (ECG) signals and fuse the different features. The first global feature extraction module is used to extract global features of respiratory signals using a self-attention mechanism. The second global feature extraction module is used to extract global features of ECG signals using a self-attention mechanism. The classification module is used to combine the features of respiratory signals and ECG signals to output Parkinson's screening results.

[0021] Compared with the prior art, the advantages of the present invention are that the provided method for early screening of Parkinson's disease based on multimodal physiological signals, based on multimodal multiscale deep learning, combined with ECG and respiratory signals, realizes automated detection of Parkinson's disease and improves the accuracy and robustness of detection.

[0022] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.

[0024] Figure 1This is a flowchart of an early Parkinson's disease screening method based on multimodal physiological signals according to an embodiment of the present invention;

[0025] Figure 2 This is a schematic diagram of a multimodal, multi-scale time-global fusion network according to an embodiment of the present invention;

[0026] Figure 3 This is a detailed network architecture diagram of a respiratory signal time feature extraction module according to an embodiment of the present invention;

[0027] Figure 4 This is a schematic diagram of a deep cyclic cell based on an SRU according to an embodiment of the present invention;

[0028] Figure 5 This is a specific network architecture of an electrocardiogram signal time feature extraction module according to an embodiment of the present invention;

[0029] Figure 6 This is a structural comparison diagram of a bottleneck layer, an inverted bottleneck layer, and a nonlinear feature fusion module according to an embodiment of the present invention;

[0030] Figure 7 This is a network structure diagram of a global feature extraction module according to an embodiment of the present invention;

[0031] Figure 8 This is a network structure diagram of a Parkinson's disease classification module according to an embodiment of the present invention;

[0032] Figure 9 This is a specific network structure diagram of a dual auxiliary task according to an embodiment of the present invention;

[0033] In the attached diagram, ECG Encoder, Breathing Encoder, QEEG Predictor, Breathing Signal, SimilarityMatching, Block, Conv Block, and Attention Weights are used to represent the signals generated by the electrocardiogram (ECG), breathing signal, and similarity matching, respectively. Detailed Implementation

[0034] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0035] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.

[0036] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0037] In all the examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0038] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0039] See Figure 1 As shown, the provided method for early Parkinson's disease screening based on multimodal physiological signals includes the following steps:

[0040] Step S110: Construct a multimodal, multi-scale deep learning model.

[0041] Parkinson's disease is a chronic neurodegenerative brain disorder that affects a person's ability to perform daily activities. Considering that electrocardiogram (ECG) changes may precede the onset of Parkinson's disease, and that respiratory symptoms often appear several years earlier than motor symptoms, respiratory and ECG signals can be used as biomarkers for detecting Parkinson's disease. Therefore, this invention proposes an early detection method for Parkinson's disease based on multimodal, multi-scale deep learning.

[0042] Combination Figure 2 As shown, the provided deep learning model mainly includes a multimodal feature extraction module and a Parkinson's disease classification module, forming a multimodal backbone network. The multimodal feature extraction module includes time feature extraction modules for two physiological signals and a global feature extraction module, namely, an electrocardiogram (ECG) time feature extraction module (or ECG signal feature extraction module), a respiratory signal feature extraction module, a respiratory signal global feature extraction module, and an ECG global feature extraction module. In one embodiment, the respiratory signal time feature extraction module uses a residual network to downsample and extract different features of the respiratory signal. The ECG time feature extraction module extracts different features of the ECG signal in parallel and utilizes a nonlinear feature fusion module to effectively fuse different features. Furthermore, to address the sparse supervision of Parkinson's disease labels, the deep learning model also introduces a dual auxiliary task: using two signals to predict the quantitative electroencephalogram (EEG) during the subject's sleep, providing additional labels for the auxiliary task and regularizing the model during training.

[0043] Specifically, for Figure 2The deep learning model architecture first takes the subject's nighttime respiratory signals and nighttime electrocardiogram (ECG) signals as input, extracts temporal and global features from these signals, and then performs classification. The respiratory signal temporal feature extraction module extracts features through residual information structures, while the ECG temporal feature extraction module performs convolutional parallel extraction on different features of the ECG signal and utilizes a nonlinear feature fusion module to perform nonlinear efficient fusion of different ECG signal features. The latter two modules then extract global features from their respective signals through a self-attention mechanism. Furthermore, to address the sparse supervision of Parkinson's disease labels, auxiliary tasks are introduced during the feature extraction process for the two signals: predicting the subject's quantitative electroencephalogram (qEEG) during sleep using the two signals. This provides additional labels for the auxiliary tasks and regularizes the model during training.

[0044] The following sections will describe in detail the implementation of the model data processing, respiratory signal time feature extraction module, electrocardiogram time feature extraction module, respiratory signal global feature extraction module, electrocardiogram global feature extraction module, Parkinson's disease classification module, and dual auxiliary task.

[0045] (1) Model data processing

[0046] In one embodiment, the data is processed to filter out nights shorter than 2 hours. Additionally, nights with distorted or absent respiratory and electrocardiogram (ECG) signals are also filtered out. In data processing, ECGs are downsampled from 512Hz to 128Hz and bandpass filtered to remove signals outside the 0.5 to 10Hz range to reduce noise. Since excessively large variances in some features can dominate the objective function, preventing the parameter estimator from correctly learning other features, z-score regularization is also applied. Respiratory signals are downsampled to 10Hz and z-score regularized to eliminate the impact of magnitude on data analysis, thus ensuring that information-rich features are extracted in subsequent feature extraction.

[0047] Table 1 shows the representation of commonly used symbols in the implementation process.

[0048] Table 1: Commonly Used Symbols in the M2TGNet Implementation Process

[0049]

[0050] (2) Respiratory signal time feature extraction module

[0051] In one embodiment, see Figure 3As shown, the respiratory signal temporal feature extraction module first uses a 3x1 convolution kernel to increase the number of channels in the respiratory signal, providing conditions for extracting more features. Then, the BN() function is used to normalize the module's input, followed by the linear rectified function ReLU(). Subsequently, eight layers of conventional residual blocks are used to downsample and extract different features of the respiratory signal. For example, the representation of each conventional residual block is as follows:

[0052] M (out) =BN(Conv (k) (Relu(BN(Conv (k) (M (in) ))))) (1)

[0053] Among them, M (out) It is a characteristic signal of a conventional residual block, M (in) This is the input of a regular residual block, and k represents the convolution kernel. Using the Batch Normalization (BN) function and subsequent activation functions such as ReLU() can accelerate the convergence speed of the model, reduce the risk of overfitting, and improve the model's generalization ability.

[0054] The respiratory signal temporal feature extraction module extracts respiratory signal temporal features at different scales by continuously increasing the number of channels in the conventional residual blocks, thereby enhancing the nonlinear fitting capability. Furthermore, to address gradient vanishing, a Simple Recurrent Unit (SRU) is added after each of the last three conventional residual blocks. For example, see... Figure 4 As shown, the calculation process of SRU is expressed as follows:

[0055]

[0056] Where f t To forget the door, z t To reset the door, c t For internal storage units, the final output state is h. t . Given the input sequence, W f W z W and b are parameter matrices in SRU. f and b z This is the bias unit vector.

[0057] SRU has the same parallelism as convolutional and feedforward networks. SRU replaces the use of convolution. Like QRNN and KNN, it has more recurrent connections. This allows the respiratory signal temporal feature extraction module to retain good modeling capabilities, extracting abnormal times and respiratory conditions from the signal, and also makes training faster.

[0058] (3) Electrocardiogram time feature extraction module

[0059] Because electrocardiogram (ECG) signals have a higher sampling rate, a greater amount of information is collected during nighttime sleep. Therefore, the ECG time feature extraction module adopts a different model design than the respiratory signal time feature extraction module. The ECG time feature extraction module also uses "depth convolution" because this effectively reduces FLOPS in the network.

[0060] See Figure 5 As shown, in one embodiment, for the ECG time feature extraction module, the number of blocks at each depth is adjusted to (3, 3, 9, 18), allowing the model to learn shallow time features in the early stages and deepen progressively in later stages. During this process, after each depth convolution is completed, the model is downsampled to an appropriate feature map size, reducing the feature signal length to 1 / 16 of its original size. Furthermore, each depth will reduce the input size to C×1. E The ECG data is processed by increasing the number of channels. This increases the complexity and expressive power of the ECG time feature extraction module, enabling it to more accurately recognize input features. Within each block, the ECG signal is first processed in parallel through three convolutions with batch normalization (BN) operations, resulting in feature maps for multiple receptive fields. The kernel sizes of the three branches are 3x1, 5x1, and 11x1, respectively. Each branch processes ECG signals of the same size into feature signals of the same size, extracting features at multiple time scales, thereby improving the module's generalization ability.

[0061]

[0062] in, Indicates having C in Input channels and C out k×1 convolution kernel operation on the output channel. μ k , σ k γ k ,β k Indicates following C (k) The cumulative mean, standard deviation, learning scaling factor, and bias in the BN layer, where {} represents the concatenation operation of the three branches. This represents the input and output of the electrocardiogram signal during feature mapping processing. BN() represents the BN operation, and the implementation formula of BN() is expressed as:

[0063]

[0064] in,

[0065] The three sets of feature signals with different receptive field characteristics obtained in this process are combined and input into the Nonlinear Feature Fusion Module to perform nonlinear and effective fusion of different features of the electrocardiogram signal to obtain the time characteristics of the electrocardiogram signal.

[0066] Combination Figure 6 As shown, the nonlinear feature fusion module is an improvement on the bottleneck layer and reverse bottleneck layer structures. The bottleneck layer initially emerged to reduce the number of parameters and computational cost, and to provide a more intuitive training and feature extraction process after dimensionality reduction. It consists of three convolutional parts: two point convolutions responsible for increasing and decreasing the dimensionality of the data features, and a depthwise convolution that performs the actual convolution operation. Figure 6 As shown in (a). Correspondingly, the subsequent inverse bottleneck layer structure makes the data dimension of the depthwise convolutional processing layer four times the input dimension, as shown in (a). Figure 6 As shown in (b), such operations can be used to obtain more valuable model performance improvements while ignoring computational costs.

[0067] Compared to Figure 6 (a) and Figure 6 (b) shows the structure. The nonlinear feature fusion module of this invention has undergone special optimization, shifting the functional layer implementation downwards and retaining only the nonlinear function between the dimensionality increase and decrease operations. See [link to relevant documentation]. Figure 6 As shown in (c), based on different feature sets in a coarse-joined form, a series of nonlinear operations interspersed within and varying the feature dimensions are used to obtain fused features with stronger semantic representation capabilities. Specifically, the three sets of feature signals with different receptive fields are merged and enter the nonlinear feature fusion module. First, they pass through a LayerNorm layer (layer normalization layer) for normalization along the feature dimensions. The formula for the LayerNorm layer is defined as follows:

[0068] M (out) =LayerNorm(M (in) ,ε,ω,δ) (5)

[0069] in The multi-scale concatenated features represent the input, ε is a constant added for numerical stability, ω is a learnable scaling parameter, and δ is a learnable offset parameter. Therefore, the LayerNorm process can be represented as:

[0070]

[0071] Among them, Var(M (in) E(M) is the feature variance input to the nonlinear feature fusion module.(in) ) represents the mean of the input, M (out) This represents the output after LayerNorm processing.

[0072] Then, the features processed by the LayerNorm layer are multiplied by four through a 1x1 convolution, followed by a ReLU activation function to deepen the non-linearity of the features, and then compressed by four through another 1x1 convolution. This operation mainly uses a learnable one-dimensional vector to scale the values ​​on each channel, thereby automatically adjusting the features and dynamically helping the model improve training quality. Finally, the features that have undergone sufficient non-linear operations are processed by a third 1x1 convolution, performing a final fusion along the feature dimension, reducing the feature size from the initial... It became Finally, the time feature tensor of the electrocardiogram signal is obtained through another LayerNorm operation.

[0073] (4) Global feature extraction module for electrocardiogram and global feature extraction module for respiratory signal.

[0074] After the time feature modules of the two signals extract their respective time features, this invention utilizes the similarity between the attention mechanism and the human observation mechanism in the global feature extraction module of the two signals. When extracting features, it first tends to focus on the important local information of the object, and then combines the information of different regions to form an overall impression of the observed object. Furthermore, the different levels of shallow feature information contained in the feature pool make the global feature extraction module more targeted when performing calculations, thereby establishing a more reliable global dependency relationship.

[0075] See Figure 7 As shown, for the global feature extraction module, the Queries, Keys, and Values ​​that implement the attention mechanism first select the feature map processed by 1x1 convolution as the mapping for Keys, aiming to minimize the difference between the key information in the attention operation process and the input data; and then use the feature map extracted after 1x1 convolution as the mapping for Queries, so that more abstract features can be considered in the process of calculating the update weights with Keys; finally, the feature map with more abstract global information is mapped to Values, so that it contains shallow global features at the starting point of the update process, providing a foundation for the entire global feature extraction process. Then, it is input into the normalization layer to generate the attention score for each signal feature, and then the time average of two signal features is calculated. The corresponding attention scores are then weighted to obtain the global Parkinson's feature G(M). (out) )∈R d×1 , where d is the fixed dimension of the global feature. Its calculation process is expressed as:

[0076]

[0077] Where Q, K, and V represent matrix-form Queries, Keys, and Values, respectively, d k Indicates the dimension of the vector Keys.

[0078] (5) Parkinson's Disease Classification Module

[0079] See Figure 8 As shown, in one embodiment, the Parkinson's disease classification module includes three fully connected layers and one sigmoid layer. The classifier outputs a PD diagnostic score, which is a number between 0 and 1. If the score exceeds 0.5, the person is considered to have Parkinson's disease.

[0080] (6) Dual auxiliary tasks

[0081] To predict whether a person has Parkinson's disease, approximately 10 hours of nighttime breathing and electrocardiogram (ECG) signals are required. This invention also introduces a dual auxiliary task to predict quantitative electroencephalograms (qEEGs) during the subject's sleep. This dual auxiliary task can provide additional labeling to help the model regularize during training. qEEG prediction was chosen as the auxiliary task because EEG is related to Parkinson's disease, breathing, and ECG.

[0082] Specifically, to generate qEEG tags, the true time-series EEG signal was first converted to the frequency domain using short-time Fourier transform and Welch periodogram methods. Time-series EEG signals were extracted from the C4-M1 channels, a common method in sleep research. The EEG spectrum was then decomposed into Δ (0.5-4Hz), θ (4-8Hz), θ (8-13Hz), and β (13-30Hz) frequency bands, and the power was normalized to obtain the relative power per second for each frequency band.

[0083] The dual-assisted task uses the encoded signals as input to predict the relative power of each EEG signal band, consisting of three 1D deconvolutional blocks (upsampling the extracted respiratory features to the same temporal resolution as the qEEG signal) and two fully connected layers. Each 1D deconvolutional block contains three deconvolutional layers, followed by batch normalization, rectified linear unit activation, and residual connections. The invention also utilizes skip connections, following a UNet architecture, by connecting the output of the SRU layer in the respiratory signal temporal feature extraction module and the deep convolutional features of the last two layers in the ECG temporal feature extraction module to the deconvolutional layers in the qEEG predictor. Figure 9 It is the specific network structure for dual auxiliary tasks, in which, Figure 9 (a) is the process corresponding to the respiratory signal in a dual-assistance task. Figure 9(b) is the process of the corresponding electrocardiogram signal in the dual auxiliary task.

[0084] Step S120: Train a deep learning model using a training set, which reflects the correspondence between respiratory signals, electrocardiogram signals and Parkinson's disease prediction labels.

[0085] For example, with the optimization objective of minimizing a predetermined loss function, a deep learning model is trained using a training set to obtain the model's optimized parameters. The training set reflects the correspondence between respiratory signals, electrocardiogram signals, and the presence or absence of Parkinson's disease. The loss function can be either the cross-entropy loss function or the squared loss function.

[0086] Taking the introduction of a dual-assistance task as an example, the overall loss function for training a deep learning model consists of two parts. The first part, Parkinson's disease classification, uses the cross-entropy loss function to measure the difference between the true probability distribution and the predicted probability distribution, expressed as:

[0087]

[0088] in, Y represents the loss in Parkinson's disease classification. PB This represents the true classification result. This indicates the predicted classification result.

[0089] The second part, the dual-assisted task using EEG signals, employed a mean squared error loss function to measure the squared average difference between the predicted output and the true label, expressed as:

[0090]

[0091]

[0092] in, The squared value of the average difference between the EEG signal predicted by the auxiliary task in the respiratory signal time feature extraction module and the actual EEG signal; X is the squared value of the average difference between the EEG signal predicted by the auxiliary task in the ECG time feature extraction module and the actual EEG signal, where n represents the length of the EEG signal, t∈(1,T) represents the time of the signal, and T represents the total duration. qeeg Represents the actual electroencephalogram (EEG) signal. These represent the EEG signals predicted by the auxiliary task in the respiratory signal time feature extraction module and the EEG signals predicted by the auxiliary task in the electrocardiogram time feature extraction module, respectively.

[0093] Therefore, in one embodiment, the overall loss function is expressed as follows:

[0094]

[0095] Step S130: The trained deep learning model is used to predict the Parkinson's diagnosis result of the target in real time.

[0096] Once the deep learning model is trained, it can be used as a Parkinson's disease detection model for actual Parkinson's diagnosis. For example, respiratory signals and electrocardiograms of the target over a period of time can be collected and input into the trained model to obtain clinical indicators of whether the target has Parkinson's disease.

[0097] It should be noted that, without departing from the spirit and scope of this invention, those skilled in the art can make appropriate changes or modifications to the above embodiments. For example, the size of the convolution kernels used in the respiratory signal time feature extraction module and the electrocardiogram time feature extraction module can be changed. The nonlinear operations in the electrocardiogram time feature extraction module and the respiratory signal time feature extraction module can be replaced by other nonlinear operations, such as replacing the Layer Norm operation with the BatchNormal operation, and replacing the ReLU activation function with the GELU activation function. The multi-scale used in the electrocardiogram time feature extraction module can be obtained using more or fewer convolutions. The functional layer in the nonlinear feature extraction module used in the electrocardiogram time feature extraction module can be moved to the top layer, i.e., feature fusion is performed first, followed by nonlinear transformation. In addition, the signals used can be replaced with other physiological signals related to Parkinson's disease. The number of skipped connections in the dual-assisted task can be increased or decreased. The signals used for assistance in the dual-assisted task can be replaced.

[0098] Accordingly, the present invention also provides an early Parkinson's disease screening device based on multimodal physiological signals, used to achieve one or more of the above aspects. For example, the device includes: a signal acquisition unit for acquiring the respiratory signal and electrocardiogram (ECG) signal of the target; and a prediction unit for inputting the respiratory signal and ECG signal into a trained deep learning model to obtain Parkinson's disease screening results. The deep learning model includes a first temporal feature extraction module, a second temporal feature extraction module, a first global feature extraction module, a second global feature extraction module, and a classification module. The first temporal feature extraction module extracts different scale features of the respiratory signal; the second temporal feature extraction module extracts different scale features of the ECG signal and fuses these different features; the first global feature extraction module extracts global features of the respiratory signal using a self-attention mechanism; the second global feature extraction module extracts global features of the ECG signal using a self-attention mechanism; and the classification module combines the features of the respiratory signal and ECG signal to output the Parkinson's disease screening results. The functions of the signal acquisition unit and the prediction unit can be implemented using a general-purpose processor, a dedicated processor, or an FPGA, etc.

[0099] To further verify the effectiveness of this invention, experiments were conducted. Experimental results show that the model provided by this invention is reliable and stable, and effectively improves the model's learning ability, resulting in a more robust and effective learning model. In the experimental verification, a 4-fold cross-validation method was used to avoid overfitting. The dataset was divided into training, validation, and test sets in a ratio of 13:5:3. The results were evaluated using the area under the curve (AUC), with test AUCs of 78.5%, 84.4%, 81.4%, and 76%, respectively, and an average AUC of 80.1%.

[0100] In summary, this invention provides an early detection method for Parkinson's disease based on multimodal, multi-scale deep learning. This method combines ECG and respiratory signals, using a multimodal, multi-scale deep learning model to extract and classify features from ECG and respiratory signal data during the subject's nighttime sleep. This captures richer biosignal information and key features, improving accuracy and robustness. Compared to existing technologies, this invention has the following advantages:

[0101] 1) The model provided by this invention can combine long-term data from two types of nighttime sleep physiological signals to provide sufficient information, and perform multimodal and multi-scale processing to extract and classify signal features from a large amount of complex biological signal data, thereby ensuring the accuracy and robustness of Parkinson's disease detection.

[0102] 2) The model provided by this invention can utilize auxiliary tasks in long-term one-dimensional nighttime sleep physiological signal data, avoiding the problem of unsatisfactory results caused by sparse supervision.

[0103] 3) The model provided by this invention can extract useful information from long-term one-dimensional nighttime sleep physiological signal data for classification, and the classification effect is better than that of training with only respiratory signals, providing an effective solution for long-term sequence processing.

[0104] 4) This invention follows the UNet network architecture design and provides a design idea for dual auxiliary tasks to solve the data sparsity problem.

[0105] 5) This invention provides a more convenient and reliable method for detecting Parkinson's disease, offering clinicians a better tool for early screening.

[0106] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0107] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0108] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0109] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, Python, etc., and conventional procedural programming languages ​​such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0110] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0111] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0112] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.

[0114] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.

Claims

1. A Parkinson's disease early screening device based on multimodal physiological signals, comprising: Signal acquisition unit: used to acquire the target's respiratory signals and electrocardiogram signals; Prediction unit: used to input the respiratory signal and electrocardiogram signal into a trained deep learning model to obtain Parkinson's screening results; The deep learning model includes a first temporal feature extraction module, a second temporal feature extraction module, a first global feature extraction module, a second global feature extraction module, and a classification module. The first temporal feature extraction module is used to extract features of different scales of respiratory signals. The second temporal feature extraction module is used to extract features of different scales of electrocardiogram (ECG) signals and fuse the different features. The first global feature extraction module is used to extract global features of respiratory signals using a self-attention mechanism. The second global feature extraction module is used to extract global features of ECG signals using a self-attention mechanism. The classification module is used to combine the features of respiratory signals and ECG signals to output Parkinson's screening results.

2. The apparatus according to claim 1, characterized in that, The first-time feature extraction module sequentially includes convolutional layers, batch normalization (BN) layers, activation layers, and multiple residual blocks, with a simple recurrent unit connected after the last set number of residual blocks.

3. The apparatus according to claim 2, characterized in that, The residual blocks are set to 8 layers, and a simple loop unit is connected after each of the last three layers of residual blocks. Each layer of residual blocks is represented as follows: in, It is a feature signal extracted from the residual block. It is the input of the residual block, and Conv indicates convolution processing. k This represents the convolution kernel.

4. The apparatus according to claim 1, characterized in that, The second time feature extraction module includes multiple depth blocks. For each depth block, there are multiple parallel branches and a nonlinear feature fusion module. The multiple parallel branches use convolution processing with BN operation to obtain feature maps of multiple receptive fields for the input electrocardiogram signal. Each branch processes the corresponding electrocardiogram signal into a feature signal of the same size to extract features at multiple time scales. The nonlinear feature fusion module is used to fuse features at different time scales.

5. The apparatus according to claim 1, characterized in that, For the first global feature extraction module and the second global feature extraction module, the process of implementing the attention mechanism is as follows: The convolutional feature map is used as the mapping for Keys, and the convolutional and extracted feature map is used as the mapping for Queries to obtain abstract global information. The feature map of the obtained abstract global information is mapped to Values, and then input into the normalization layer to generate the attention score of the corresponding signal feature; Calculate the time average of the signal features and weight them as global features using the corresponding attention scores.

6. The apparatus according to claim 1, characterized in that, The classification module includes multiple fully connected layers and one Sigmoid layer. The score output by the classifier is a number between 0 and 1, which is then compared with a set threshold to determine whether the target has Parkinson's disease.

7. The apparatus according to claim 1, characterized in that, During the training of the deep learning model, a dual auxiliary task is introduced to predict the quantitative electroencephalogram (EEG) during the subject's sleep. The overall loss function for training the deep learning model is set as follows: in: in, This represents the total loss value. This indicates a loss in the Parkinson's disease classification. This indicates the true classification results of Parkinson's disease. This indicates the classification results for predicting Parkinson's disease. The squared value of the average difference between the EEG signal predicted by the auxiliary task in the first-time feature extraction module and the actual EEG signal; It is the squared value of the average difference between the EEG signal predicted by the auxiliary task in the second temporal feature extraction module and the actual EEG signal. n Indicates the length of the electroencephalogram (EEG) signal. This indicates the signal duration, where T represents the total duration. Represents the actual electroencephalogram (EEG) signal. , These represent the EEG signals predicted by the auxiliary task in the first time feature extraction module and the EEG signals predicted by the auxiliary task in the second time feature extraction module, respectively.

8. The apparatus according to claim 4, characterized in that, The nonlinear feature fusion module sequentially includes a first LayerNorm layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a second LayerNorm layer. The first LayerNorm layer is used to perform layer normalization along the feature dimension, the first convolutional layer is used to dilate the number of channels, the second convolutional layer is used to scale the value on each channel using a learnable one-dimensional vector, the third convolutional layer is used to perform feature fusion along the feature dimension, and the second LayerNorm layer is used to obtain the time feature tensor of the electrocardiogram signal through layer normalization.

9. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by the processor, it performs the following steps: Acquire the target's respiratory and electrocardiogram signals; The respiratory and electrocardiogram signals are input into a trained deep learning model to obtain Parkinson's screening results; The deep learning model includes a first temporal feature extraction module, a second temporal feature extraction module, a first global feature extraction module, a second global feature extraction module, and a classification module. The first temporal feature extraction module is used to extract features of different scales of respiratory signals. The second temporal feature extraction module is used to extract features of different scales of electrocardiogram (ECG) signals and fuse the different features. The first global feature extraction module is used to extract global features of respiratory signals using a self-attention mechanism. The second global feature extraction module is used to extract global features of ECG signals using a self-attention mechanism. The classification module is used to combine the features of respiratory signals and ECG signals to output Parkinson's screening results.

Citation Information

Patent Citations

  • Multi-modal data acquisition device based on neural feedback

    CN113397502A

  • Electrocardiogram anomaly detection method and system based on multi-scale signal recovery

    CN116541791A