Brain-computer interface (BCI) based brain-inspired classification model construction method
Patent Information
- Application Number
- CN202411226542.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-09-03
AI Technical Summary
[0006]鉴于上述的分析,本发明实施例旨在提供一种面向BCG的类脑分类模型构建方法,用以解决现有方法中BCG数据集样本数量不足、以及BCG心拍分类精确度低的技术问题
[0048]与现有技术相比,本发明至少可实现如下有益效果之一:
Smart Images

Figure CN119167192B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent medical signal processing technology, and in particular to a method for constructing a brain-like classification model for BCG (Ballistocardiogram, cardiac impact signal). Background Technology
[0002] ECG (Electrocardiogram) is an electrical conduction signal that records the heart's electrical activity. By placing electrodes on the skin, it measures changes in the heart's voltage at different points in time. ECG is widely used to diagnose heart conditions such as arrhythmias, myocardial infarction, and myocardial ischemia. Doctors can assess the heart's health by analyzing the morphology and intervals of the ECG waveform, aiding in disease diagnosis, treatment planning, and monitoring treatment effectiveness. ECG can also be used to screen high-risk individuals, detect potential heart problems early, and prevent cardiovascular disease.
[0003] Unlike conventional cardiac systolic angiography (BCG), which records the mechanical motion of the heart, BCG measures the chest vibrations produced during a heartbeat to reflect the heart's contraction and relaxation. BCG provides information on cardiac contractility, heart rate, and heart volume, which is crucial for assessing cardiac function and its impact on heart disease. By using BCG, doctors can gain a more detailed understanding of a patient's heart condition, develop targeted and personalized treatment plans, improve treatment outcomes, and reduce the risk of heart attacks and complications.
[0004] Unlike ECG, which requires contact with the skin via electrodes and a conductive medium (such as gel), BCG measures the mechanical signal generated by the vibration of the chest during a heartbeat. BCG testing does not require direct contact with the skin and can be performed for extended periods without disturbing the subject. This non-invasive nature makes it particularly suitable for long-term monitoring and sleep tracking. Contactless wearable BCG testing devices have undergone rapid development and are now readily applicable in everyday environments, significantly improving testing efficiency while reducing costs. This makes abnormal heart rate detection more convenient, efficient, and accurate. BCG testing technology is not only used for heart rate monitoring, sleep structure analysis, and cardiac function monitoring and evaluation, but also shows potential in the diagnosis of cardiovascular diseases and postoperative recovery evaluation.
[0005] However, compared to the ample ECG datasets available, the number of publicly available BCG datasets from medical institutions is relatively insufficient to meet the enormous demand for various BCG analyses and deep learning applications. Furthermore, the accuracy of BCG in classifying abnormal heartbeats is also relatively low. Summary of the Invention
[0006] Based on the above analysis, the embodiments of the present invention aim to provide a method for constructing a brain-like classification model for BCG, in order to solve the technical problems of insufficient sample size in the BCG dataset and low accuracy of BCG heartbeat classification in existing methods.
[0007] This invention provides a method for constructing a brain-like classification model for BCG, comprising the following steps:
[0008] Construct the first training dataset, which includes the preprocessed ECG and the corresponding BCG;
[0009] The ECG-BCG mapping model is trained using the first training dataset until the loss function converges to obtain the trained ECG-BCG mapping model.
[0010] Based on the trained ECG-BCG mapping model, a second training dataset for the brain-like classification model is constructed; wherein, the second training dataset includes BCG and the corresponding ECG heartbeat type;
[0011] The brain-like classification model is trained using the second training dataset, and the trained brain-like classification model is obtained after reaching the maximum number of iterations.
[0012] Furthermore, an ECG-BCG mapping model is constructed based on the encoder-decoder structure of a U-shaped convolutional network, including an encoder module and a decoder module;
[0013] The encoder module is used to downsample the received preprocessed ECG to extract local features and output a feature map matrix representing the local features.
[0014] The decoder module restores the spatial dimension through a transposed convolutional layer, and fuses the local features extracted by the encoder module with the global features upsampled by the decoder module through skip connections to restore the original dimensional space, and outputs the BCG corresponding to the preprocessed ECG.
[0015] The feature map matrix is represented as a three-dimensional tensor, with the three dimensions being batch size, number of channels, and sequence length, respectively.
[0016] Furthermore, the encoder module is a downsampling structure, comprising five downsampling layers, wherein:
[0017] The first downsampling layer consists of a convolutional layer with a stride of 1, layer normalization, and PReLU activation. The convolutional layer increases the number of channels in the input feature map from 1 to 32 before outputting it to the second downsampling layer.
[0018] The second to fifth downsampling layers consist of two convolutional layers with a stride of 1 and a max pooling layer, respectively. After each convolutional layer, layer normalization and PReLU activation are performed in sequence. The first convolutional layer is used to double the number of channels in the feature map input to the current layer. The second convolutional layer keeps the number of channels unchanged and halves the sequence length dimension of the feature map input to the current layer through the max pooling layer, thus outputting the feature map with the changed number of channels and sequence length dimension.
[0019] Layer normalization is used to accelerate convergence, and PReLU activation introduces nonlinearity to enhance the encoder module's sensitivity to subtle changes in local features.
[0020] Furthermore, the decoder module is an upsampling structure, comprising six upsampling layers, as follows:
[0021] The first upsampling layer consists of two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed in sequence. The two convolutional operations with a stride of 1 keep the number of channels unchanged. The transposed convolutional layer is used to halve the number of channels and double the sequence length, thereby achieving spatial expansion and upsampling of features.
[0022] The second to fourth upsampling layers each consist of two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed sequentially. The first convolutional layer with a stride of 1 halves the number of channels in the feature map input to the current layer. The second convolutional layer with a stride of 1 keeps the number of channels unchanged, and then uses a transposed convolutional layer to halves the number of channels and doubles the sequence length. The input of the second upsampling layer and the output of the fourth downsampling layer, the input of the third upsampling layer and the output of the third downsampling layer, and the input of the fourth upsampling layer and the output of the second downsampling layer are sequentially connected in a skip connection to fuse the downsampled local features and the upsampled global features. The second to fourth upsampling layers output the feature map with the changed number of channels and sequence length to the next layer for upsampling.
[0023] The fifth upsampling layer includes two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed in sequence. The first convolutional layer with a stride of 1 halves the number of channels to refine the feature representation. The second convolutional layer with a stride of 1 keeps the number of channels unchanged and further extracts and fuses the downsampling and upsampling features.
[0024] The sixth upsampling layer uses a single convolutional layer to reduce the number of channels from 16 to 1, mapping the decoded feature map back to the dimension of the feature map input to the encoder module, and outputting the ECG-mapped BCG.
[0025] Further, the first training dataset is obtained by preprocessing the original signal sequences of ECG and corresponding BCG data obtained from the first public dataset. The steps include:
[0026] Starting from the first R peak of the original signal sequence of the ECG, and starting from the J peak in the original signal sequence of the BCG corresponding to the timestamp of the first R peak, non-overlapping segments are performed with a preset sequence length.
[0027] Data points before the first R peak in ECG and the first J peak in BCG, as well as data points that are less than the preset sequence length after segmentation, are discarded.
[0028] Normalize the amplitudes of the segmented ECG and BCG;
[0029] The normalized ECG and its corresponding BCG constitute each sample in the first training dataset.
[0030] Further, the first training dataset is loaded into the ECG-BCG mapping model for training;
[0031] A fixed learning rate is used, and forward propagation is performed using the MSE (mean squared error) loss function, while backpropagation is performed using the Adam optimizer. The parameters of the ECG-BCG mapping model are continuously adjusted to minimize the loss function until the loss function converges to obtain the trained ECG-BCG mapping model.
[0032] Furthermore, the brain-like classification model is constructed based on a spiking neural network, and includes, in sequence, a static convolutional layer, a channel attention layer, and a fully connected classification module;
[0033] The static convolutional layer includes a one-dimensional convolutional layer Conv1d, a LIFNode brain-like activation layer, and a pooling layer. The one-dimensional convolutional layer Conv1d is used for channel expansion to extract features. After the input feature map is convolved by Conv1d and activated by LIFNode brain-like activation, the number of channels is expanded from 1 to 4, and the sequence length is reduced from 304 to 284. Then, the pooling layer downsamples the feature map with expanded channels and reduced sequence length, and outputs it as X.
[0034] The channel attention layer includes a channel weight branch and a pass-in input branch;
[0035] The channel weight branch is combined with a custom convolution using the AdaptiveMaxPool1d function to output attention values for different channels, denoted as A.
[0036] The input branch first performs a multiplication operation between X and A, then passes through a convolutional layer and a pooling layer for downsampling, and outputs a feature map of a three-dimensional tensor to the fully connected classification module for classification.
[0037] The fully connected classification module includes a Flatten layer and a linear layer. The feature map of the input three-dimensional tensor is flattened into a two-dimensional tensor by the Flatten function, and then passed through the linear layer to obtain the classification prediction probability of the heartbeat type; the heartbeat classification result is obtained based on the classification prediction probability.
[0038] Further, ECG data is obtained from the second public dataset and preprocessed, then input into the trained ECG-BCG mapping model to obtain the corresponding BCG data, and the heartbeat type of the corresponding ECG in the second public dataset is used as a label to form the second training dataset.
[0039] The second training dataset is loaded into the brain-like classification model for training until the maximum number of iterations is reached.
[0040] Furthermore, the computation mechanism of the LIFNode neuromorphic activation layer is as follows:
[0041]
[0042] Where T is the preset brain-like time step, V[T] and V[T-1] are the brain-like neuron membrane potentials at brain-like time steps T and T-1, respectively, Y[T] is the feature map output by Conv1d convolution, and τ m V is the preset membrane time constant. reset To reset the potential;
[0043] If V[T] reaches the preset discharge threshold V threshold If the membrane potential returns to zero, V[T] is input to the substitution function of the step function for back propagation; the substitution function is as follows:
[0044]
[0045] Otherwise, if V[T] does not reach the preset discharge threshold V threshold Then the membrane potential of brain-like neurons remains V[T].
[0046] Furthermore, the first publicly available dataset is the NIH-published BALLISTOCARDIOGRAPHYDATASET; where ECG and BCG are in a one-to-one correspondence based on timestamps;
[0047] The second publicly available dataset is the electrocardiogram (ECG) signal dataset released by MIT-BIH, which includes ECG signals and corresponding heartbeat types.
[0048] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0049] 1. By using the ECG-BCG mapping model, the number of BCG samples was effectively increased, solving the problem of insufficient sample quantity in the BCG dataset and improving the utilization rate of BCG samples.
[0050] 2. An ECG-BCG mapping model was established using the encoder-decoder structure of a U-shaped convolutional network, which improved the model's ability to extract local and global features of ECG and BCG signals and enhanced the richness of feature representation.
[0051] 3. A brain-like classification model based on a spiking neural network improves the accuracy of heartbeat classification through the LIFNode brain-like activation module and channel attention mechanism;
[0052] 4. By integrating the NIH's BED-BASED BALLISTOCARDIOGRAPHY DATASET and the MIT-BIH's ECG dataset, a comprehensive and accurately labeled training dataset was constructed, providing high-quality data support for the training of brain-like classification models and achieving effective integration of public datasets.
[0053] 5. Compared with existing abnormal heartbeat detection methods, this invention makes better use of the mechanical information contained in BCG, and can more effectively detect cardiac mechanical motion. The heart-brain classification model has broad application prospects in clinical diagnosis, long-term monitoring and personalized treatment.
[0054] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0055] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0056] Figure 1 A flowchart illustrating a brain-like classification model construction method for BCG;
[0057] Figure 2 A schematic diagram showing the fixed time delay of the RR peak of ECG and the JJ peak of BCG collected simultaneously;
[0058] Figure 3The diagram shows the structure of the ECG-BCG mapping model based on the encoder-encoder structure of a U-shaped convolutional network.
[0059] Figure 4 This is a structural diagram of a brain-like classification model based on a spiking neural network. Detailed Implementation
[0060] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0061] The existing BCG dataset has a relatively insufficient sample size. To leverage the advantages of BCG signals in detecting heart rate abnormalities, which contain more information such as cardiac contractility than ECG signals, a lightweight intelligent algorithm that can be deployed on BCG devices is designed to achieve rapid and effective disease detection.
[0062] BCG signals provide a direct measurement of cardiac mechanical activity, including cardiac contractility and workload, which ECG cannot provide. Since BCG is non-invasive and wearable, it is ideally suited for deploying lightweight diagnostic algorithms on a chip. Brain-like neural networks offer the advantage of being lightweight compared to traditional convolutional networks. The combination of brain-like neural networks and BCG represents a cutting-edge innovation.
[0063] To address the aforementioned technical problems of insufficient existing cardiac impact signal datasets and low accuracy in cardiac beat classification based on cardiac impact signals, this invention proposes a brain-like classification model construction method for BCG.
[0064] First, given the relative insufficiency of existing cardiac impact signal datasets, based on the isoperiodicity of ECG and cardiac impact signals, such as... Figure 2 The image shows ECG and BCG data acquired simultaneously. The RR interval between ECG-R peaks and the JJ interval between JJ peaks of the cardiac impact signal have a fixed time delay. This invention proposes an ECG-BCG mapping model based on a U-shaped convolutional network encoder-decoder structure. This model accurately performs ECG-BCG mapping, effectively expands the BCG sample data, and solves the problem of insufficient sample quantity in existing cardiac impact signal datasets.
[0065] Secondly, given the excellent characteristics of BCG—low detection cost and rich detail—this invention utilizes BCG for abnormal heartbeat detection and classification. This invention proposes a brain-like classification model based on a brain-like neural network. Using this brain-like classification model, more accurate heartbeat classification and disease type identification can be easily achieved.
[0066] To achieve the above objectives, a specific embodiment of the present invention discloses a method for constructing a brain-like classification model for BCG, such as... Figure 1 As shown, it includes the following steps:
[0067] Step S1: Construct the first training dataset, which includes the preprocessed ECG and the corresponding BCG.
[0068] Step S2: Train the ECG-BCG mapping model using the first training dataset until the loss function converges to obtain the trained ECG-BCG mapping model.
[0069] Step S3: Based on the trained ECG-BCG mapping model, construct the second training dataset for the brain-like classification model; wherein, the second training dataset includes BCG and the corresponding ECG heartbeat type;
[0070] Step S4: Use the second training dataset to train the brain-like classification model, and obtain the trained brain-like classification model after reaching the maximum number of iterations.
[0071] Step S1 includes steps S11-S12.
[0072] The first training dataset is obtained by preprocessing the raw signal sequences of ECG and corresponding BCG data from the first publicly available dataset. The steps include:
[0073] Starting from the first R peak of the original signal sequence of the ECG, and starting from the J peak in the original signal sequence of the BCG corresponding to the timestamp of the first R peak, non-overlapping segments are performed with a preset sequence length.
[0074] Data points before the first R peak in ECG and the first J peak in BCG, as well as data points that are less than the preset sequence length after segmentation, are discarded.
[0075] Normalize the amplitudes of the segmented ECG and BCG;
[0076] The normalized ECG and its corresponding BCG constitute each sample in the first training dataset.
[0077] The original ECG signal sequence is used to find the first R peak of the signal sequentially from the beginning of the signal, which is taken as the starting point.
[0078] Step S11: Obtain the original signal sequence, including ECG and BCG synchronized with it, from the first public dataset.
[0079] The original signal sequence, including ECG and corresponding BCG, was obtained from the first public dataset and preprocessed.
[0080] like Figure 2 The image shows the BCG and ECG signals acquired simultaneously.
[0081] The original signal sequence, including ECG and corresponding BCG, was read from the first publicly available dataset and synchronized with each other using timestamps.
[0082] For example, the original signal sequence is an ECG in mat format and its corresponding BCG.
[0083] In this invention, the first public dataset is the NIH-published BED-BASEDBALLISTOCARDIOGRAPHY DATASET; wherein ECG and BCG are in a one-to-one correspondence based on timestamps;
[0084] The first publicly available dataset used in this invention comes from NIH (National Library of Medicine, National Center for Biotechnology Information). The name of this publicly available signal dataset is: BED-BASED BALLISTOCARDIOGRAPHY DATASET. The length of the published original signal sequence, ECG and BCG are both 107,616 data points.
[0085] In the publicly available signal dataset, there is a one-to-one correspondence between the original ECG and BCG signal sequences, with one ECG corresponding to one BCG. There are 354 samples, totaling 708 original signal sequence data.
[0086] Step S12: Preprocess the ECG and corresponding BCG of the original signal sequence to construct the first training dataset.
[0087] (1) Perform non-overlapping segmentation on the original ECG and corresponding BCG signal sequences.
[0088] The original signal sequence of ECG and BCG is divided into non-overlapping segments using a loop function, starting from the first R peak of ECG and starting from the J peak of the first R peak of the corresponding ECG for BCG, with an interval length of 304 data points. The segments are then saved as a two-dimensional tensor format, represented as [number of samples, sequence length].
[0089] The original signal sequence has 107,616 data points in the ECG and BCG segments. It is divided into 354 data segments with a non-overlapping interval of 304 data points each. Each segment contains 304 data points.
[0090] For example, the loop function uses a for loop; it is saved as a numpy tensor in Python 3.8 format with a size of [354, 304].
[0091] According to matrix storage, both BCG and ECG are stored as 354 signals with a length of 304 data points, represented as [354, 304].
[0092] For the same patient, ECG and BCG are collected at the same time and are matched one-to-one based on timestamps.
[0093] (2) Normalize the BCG and ECG signal data after non-overlapping segmentation to remove the original bias of the data values:
[0094] Normalization involves subtracting the mean (i.e., the mean amplitude of the signal) from each data point and then dividing by the standard deviation, thus converting the data into a format with a mean and unit variance. This facilitates subsequent data processing and analysis.
[0095]
[0096] Wherein, the current value is the amplitude of the signal tensor after non-overlapping segmentation; the average value is the sum of the amplitudes of all data points in the 304-point signal dataset divided by the number of data points; and the standard deviation is also the standard deviation of the values of the 304 points.
[0097] from Figure 2 It can be seen that without normalization, the vertical axis magnitudes of BCG and ECG are different, making subsequent calculations difficult. After normalization, all non-overlapping segmented ECGs and their corresponding BCGs are within the range [0, 1]. This simplifies subsequent operations.
[0098] (3) Add a channel to the normalized two-dimensional signal tensor to expand it into a three-dimensional tensor and construct the first training dataset.
[0099] For example, the reshape function based on Python 3.8 is used to expand the processed two-dimensional signal tensor from [number of samples, sequence length] to a three-dimensional tensor, represented as [number of samples, number of channels, sequence length], that is, from [354, 304] to [354, 1, 304], adding one channel. This prepares for the channel dimension transformation of the subsequent ECG-BCG mapping function and the brain-like classification model, and is adapted to and compatible with the subsequent ECG-ECG network structure and the subsequent ECG-BCG mapping model.
[0100] After preprocessing, each signal sample data is a preprocessed ECG, and the corresponding preprocessed BCG is used as a sample to form the first training dataset, which contains a total of 354 samples.
[0101] (4) Randomly sample the first training dataset and divide it into training set, validation set and test set according to a preset ratio;
[0102] For example, the obtained sample dataset is randomly divided into 354 samples in a preset ratio of 8:1:1, resulting in 283 samples forming the training set, 36 samples forming the validation set, and 35 samples forming the test set.
[0103] Step S2 includes steps S21-S22.
[0104] Step S21: Construct the ECG-BCG mapping model.
[0105] (1) Construct an ECG-BCG mapping model based on a U-shaped convolutional network encoder-decoder structure. The model structure diagram is as follows: Figure 3 As shown.
[0106] An ECG-BCG mapping model is constructed based on the encoder-decoder structure of a U-shaped convolutional network, including an encoder module and a decoder module;
[0107] The encoder module is used to downsample the received preprocessed ECG to extract local features and output a feature map matrix representing the local features.
[0108] The decoder module restores the spatial dimension through a transposed convolutional layer, and fuses the local features extracted by the encoder module with the global features upsampled by the decoder module through skip connections to restore the original dimensional space, and outputs the BCG corresponding to the preprocessed ECG.
[0109] The feature map matrix is represented as a three-dimensional tensor, with the three dimensions being batch size, number of channels, and sequence length, respectively.
[0110] The encoder module extracts local features from the input feature map, while the decoder module fuses the local features extracted by downsampling in the encoder module with the global features extracted by upsampling in the decoder module. The specific parameters for each layer of the ECG-BCG mapping model are shown in Table 1.
[0111] Table 1: Structural Parameters of the ECG-BCG Mapping Model
[0112]
[0113]
[0114]
[0115] like Figure 3 As shown:
[0116] The encoder module includes the first through fifth downsampling layers:
[0117] Within each layer, the number of samples remains constant.
[0118] (1) First downsampling layer: The input feature map dimension is [16, 1, 304], and the output feature map dimension is [16, 32, 304].
[0119] The first downsampling layer is a single convolutional layer, with no change in the sample number dimension and sequence length dimension. In the convolution, the number of channels is defined to change from 1 to 32.
[0120] (2) Second downsampling layer: The input is the output feature map dimension of the first downsampling layer, which is [16,32,304], and the output feature map dimension is [16,64,152].
[0121] Multiplying the channel number dimension by 2 gives a channel number of 64.
[0122] After convolution, the sequence length S l The change in dimensions is shown in formula (2):
[0123]
[0124] Wherein, input_shape is the sequence length of the input feature map of this layer, kernel_size is the convolution kernel size, p is the padding, and strikes is the stride.
[0125] The input sequence length for the second downsampling layer is 304. After convolution, the sequence length dimension is calculated using formula (2) as follows:
[0126]
[0127] In the max pooling layer, the sequence length dimension is divided by 2; the final sequence length dimension is 304 ÷ 2 = 152.
[0128] The final output feature map dimension of the second downsampling layer is [16, 64, 152].
[0129] (3) Third downsampling layer: The input is the output feature map dimension of the second downsampling layer, which is [16, 64, 152], and the output feature map dimension is [16, 128, 76].
[0130] Multiplying the channel number dimension by 2 gives a channel number of 128.
[0131] The sequence length dimension is calculated using formula (2) to obtain 152, and then the sequence length dimension of the max pooling layer is divided by 2; the final sequence length dimension is 152 ÷ 2 = 76;
[0132] The final output feature map dimension of the third downsampling layer is [16, 128, 76].
[0133] (4) Fourth downsampling layer: The input is the output feature map dimension of the third downsampling layer, which is [16,128,76], and the output feature map dimension is [16,256,38].
[0134] Multiplying the channel number dimension by 2 gives a channel number of 256.
[0135] The sequence length dimension is calculated using formula (2) to obtain 76, and then the sequence length dimension of the max pooling layer is divided by 2; the final sequence length dimension is 76 ÷ 2 = 38;
[0136] The final output feature map dimension of the fourth downsampling layer is [16, 256, 38].
[0137] (5) Fifth downsampling layer: The input is the output feature map dimension of the fourth downsampling layer, which is [16,256,38], and the output feature map dimension is [16,512,19].
[0138] Multiplying the channel number dimension by 2 gives a channel number of 512.
[0139] The sequence length dimension is calculated using formula (2) to obtain 38, and then the sequence length dimension of the max pooling layer is divided by 2; the final sequence length dimension is 38 ÷ 2 = 19;
[0140] The final output feature map dimension of the fifth downsampling layer is [16, 512, 19].
[0141] The decoder module includes the first through sixth upsampling layers:
[0142] (1) First upsampling layer: The input is the output feature map dimension of the fifth downsampling layer, which is [16,512,19], and the output feature map dimension is [16,256,38].
[0143] The channel number dimension is divided by 2 when transferred to the convolutional layer, resulting in a channel number of 256.
[0144] The sequence length dimension is calculated to be 19 using formula (2), and then the transposed convolutional layer sequence length is multiplied by 2, resulting in a final sequence length dimension of 38.
[0145] The final output feature map dimension of the first upsampling layer is [16, 256, 38].
[0146] (2) Second upsampling layer: The input is the output feature map dimension of the fourth downsampling layer [16,256,38] and the output feature map dimension of the first upsampling layer [16,256,38]. The jump concatenation is performed to obtain the feature map dimension [16,512,38], and the output feature map dimension is [16,128,76].
[0147] [16,256,38] and [16,256,38] are cascaded to obtain [16,512,38], which is used as the input of the second upsampling layer;
[0148] After dividing the number of channels in the convolutional layer by 2, and then dividing the number of channels in the transposed convolutional layer by 2 again, we get 128 channels.
[0149] The sequence length dimension is calculated to be 38 using formula (2), and then the sequence length dimension is multiplied by 2 after passing through the transposed convolutional layer to obtain 76;
[0150] The final output feature map dimension of the second upsampling layer is [16, 128, 76].
[0151] (3) Third upsampling layer: The input is the output feature map dimension of the third downsampling layer [16,128,76] and the output feature map dimension of the second upsampling layer [16,128,76]. The jump concatenation is performed to obtain the feature map dimension [16,256,76], and the output feature map dimension is [16,64,152].
[0152] [16,128,76] and [16,128,76] are cascaded to obtain [16,256,76], which is used as the input of the third upsampling layer;
[0153] After dividing the number of channels in the convolutional layer by 2, and then dividing the number of channels in the transposed convolutional layer by 2 again, we get 64 channels.
[0154] The sequence length dimension is calculated to be 76 using formula (2), and then the sequence length dimension is multiplied by 2 after passing through the transposed convolutional layer to obtain 152;
[0155] The final output feature map dimension of the third upsampling layer is [16, 64, 152].
[0156] (4) Fourth upsampling layer: The input is the output feature map dimension of the second downsampling layer [16,64,152] and the output feature map dimension of the third upsampling layer [16,64,152]. The skip concatenation is performed to obtain the feature map dimension [16,128,152], and the output feature map dimension is [16,32,304].
[0157] [16,64,152] and [16,64,152] are cascaded to obtain [16,128,152], which is used as the input of the third upsampling layer;
[0158] After dividing the number of channels in the convolutional layer by 2, and then dividing the number of channels in the transposed convolutional layer by 2 again, we get 32 channels.
[0159] The sequence length dimension is calculated to be 152 using formula (2), and then the sequence length dimension is multiplied by 2 after passing through the transposed convolutional layer to obtain 304;
[0160] The final output feature map dimension of the fourth upsampling layer is [16, 32, 304].
[0161] (5) Fifth upsampling layer: The input is the output feature map dimension of the fourth upsampling layer [16,32,304], and the output feature map dimension is [16,16,304].
[0162] Dividing the number of convolution channels by 2 yields 16;
[0163] The final output feature map of the fifth upsampling layer has dimensions of [16, 16, 304].
[0164] (6) Sixth upsampling layer: The input is the output feature map dimension of the fifth upsampling layer [16,16,304], and the output feature map dimension is [16,1,304].
[0165] The sixth upsampling layer is a single convolutional layer, reducing the number of channels from 16 to 1.
[0166] In this invention, the number of channels can be changed through convolution. There are several ways to change the dimension of the number of channels: increasing the dimension from 1 to 32, keeping the number of channels unchanged, multiplying the number of channels by 2, dividing the number of channels by 2, and reducing the dimension from 16 to 1.
[0167] (2) Construct an encoder module for extracting local features, which stacks 5 layers of convolutional neural network.
[0168] The encoder module is a downsampling structure, comprising five downsampling layers, wherein:
[0169] The first downsampling layer consists of a convolutional layer with a stride of 1, layer normalization, and PReLU activation. The convolutional layer increases the number of channels in the input feature map from 1 to 32 before outputting it to the second downsampling layer.
[0170] The second to fifth downsampling layers consist of two convolutional layers with a stride of 1 and a max pooling layer, respectively. After each convolutional layer, layer normalization and PReLU activation are performed in sequence. The first convolutional layer is used to double the number of channels in the feature map input to the current layer. The second convolutional layer keeps the number of channels unchanged and halves the sequence length dimension of the feature map input to the current layer through the max pooling layer, thus outputting the feature map with the changed number of channels and sequence length dimension.
[0171] Layer normalization is used to accelerate convergence, and PReLU activation introduces nonlinearity to enhance the encoder module's sensitivity to subtle changes in local features.
[0172] As shown in Table 1, the first downsampling layer includes a convolutional combination, which includes a convolutional layer, layer normalization, and PReLU activation. The input sample feature map dimension [16, 1, 304] is processed by one convolution operation, one layer normalization operation, and one PReLU activation operation, and the output feature map dimension is [16, 32, 304], which is used as the input of the second downsampling layer.
[0173] The second downsampling layer consists of two convolutional combinations and a max pooling layer. The kernel size of each convolution is 31×1 (stride is 1, padding is 15). The feature map dimension of the output signal of the second downsampling layer is [16, 64, 152].
[0174] In the second to fifth downsampling layers, the input feature map first undergoes two one-dimensional convolutional combination operations to extract features. The convolutional kernel size is 31×1 (stride is 1, padding is 15), and then it is input into the max pooling layer (pooling window size is 2×2) for downsampling.
[0175] For example, the batch size of the data input to the ECG-BCG mapping model is 16, i.e., B=16, and the sequence length is 304 data points, i.e., L=304. The dimensions of the input feature maps of the first to fifth downsampling layers of the encoder module are [16,1,304], [16,32,304], [16,64,152], [16,128,76], [16,256,38], respectively. The dimension of the final output feature map of the fifth encoding layer is [16,512,19].
[0176] In [16,1,304], 16 represents the batch size, 1 represents the number of channels, and 304 represents the signal sequence length, with the unit being the number of data points.
[0177] The first to fifth downsampling layers in this step extract local signal features more accurately, mainly through convolution operations, to extract local features of ECG and corresponding BCG.
[0178] The first to fifth downsampling layers extract local signal features mainly through the following aspects:
[0179] ① Convolution kernel size: The encoder module uses a 31×1 convolution kernel size, which means that the convolution operation is performed only within a very small local window of the ECG or BCG, which can capture local signal changes and features;
[0180] ② Convolutional layer stacking: The encoder is composed of multiple convolutional layers stacked together, each layer further refines the feature extraction based on the previous layer; each convolutional layer focuses on extracting features of a local region of the input signal;
[0181] ③Stride and padding: The stride in the convolution operation is 1 and the padding is 15, which means that the convolution kernel only slides one unit on the input signal. At the same time, the padding ensures that the convolution operation can cover the boundary area of the input signal, enhancing the ability to extract local features.
[0182] ④ Downsampling: Downsampling is performed through a max pooling layer, which reduces the dimension of the feature map while increasing the number of channels; this operation helps to reduce spatial redundancy of data while preserving important local features.
[0183] ⑤ Feature map dimension change: The dimension of the encoder's output feature map gradually decreases, indicating that the model gradually refines from a large range of local features to a smaller range of local features during the feature extraction process.
[0184] ⑥ PReLU activation function: The use of the PReLU activation function provides a non-linear transformation for the extraction of local features, enhancing the model's ability to express local features;
[0185] The encoder module can effectively extract local features from cardiac impact signals and corresponding ECGs, providing a foundation for subsequent ECG-BCG mapping and abnormal heartbeat classification, which is crucial for understanding the mechanical activity of the heart and diagnosing heart disease.
[0186] (3) Construct a decoder module that integrates local and global features, and stack 6 layers of convolutional neural network in sequence.
[0187] The decoder module is an upsampling structure, comprising six upsampling layers, as follows:
[0188] The first upsampling layer consists of two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed in sequence. The two convolutional operations with a stride of 1 keep the number of channels unchanged. The transposed convolutional layer is used to halve the number of channels and double the sequence length, thereby achieving spatial expansion and upsampling of features.
[0189] The second to fourth upsampling layers each consist of two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed sequentially. The first convolutional layer with a stride of 1 halves the number of channels in the feature map input to the current layer. The second convolutional layer with a stride of 1 keeps the number of channels unchanged, and then uses a transposed convolutional layer to halves the number of channels and doubles the sequence length. The input of the second upsampling layer and the output of the fourth downsampling layer, the input of the third upsampling layer and the output of the third downsampling layer, and the input of the fourth upsampling layer and the output of the second downsampling layer are sequentially connected in a skip connection to fuse the downsampled local features and the upsampled global features. The second to fourth upsampling layers output the feature map with the changed number of channels and sequence length to the next layer for upsampling.
[0190] The fifth upsampling layer includes two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed in sequence. The first convolutional layer with a stride of 1 halves the number of channels to refine the feature representation. The second convolutional layer with a stride of 1 keeps the number of channels unchanged and further extracts and fuses the downsampling and upsampling features.
[0191] The sixth upsampling layer uses a single convolutional layer to reduce the number of channels from 16 to 1, mapping the decoded feature map back to the dimension of the feature map input to the encoder module, and outputting the ECG-mapped BCG.
[0192] The first upsampling layer is a deconvolutional layer, which consists of two convolutional combinations and a transposed convolutional layer. The output feature map of the fifth layer is encoded and used as the input feature map of the first layer for decoding. After two one-dimensional convolutional combination operations, the kernel size is 31×1 (stride is 1 and padding is 15). Then, the transposed convolutional layer is used to obtain the first upsampling result.
[0193] The local feature map dimension [16,512,19] output by the fifth downsampling layer is input to the first upsampling layer. The feature map dimension [16,256,38] output by the first upsampling layer is used as the input to the second upsampling layer.
[0194] The second to fourth upsampling layers are feature fusion deconvolution layers;
[0195] The second to fourth upsampling layers consist of two convolutional combinations and one transposed convolutional layer, respectively. The first convolutional combination has a 31×1 kernel, a stride of 1, padding of 15, and channels divided by 2 (the number of channels is half that of the first decoding layer). The second convolutional combination has the same convolutional layer as the first convolutional combination. The transposed convolutional layer has half the number of channels of the convolutional combination and a dimension of 2.
[0196] The fifth upsampling layer is a combination of two convolutions; the output feature map of the fourth upsampling layer is used as the input of the fifth upsampling layer; the first convolution combination has a kernel of 31×1, a stride of 1, padding of 15, and the number of channels is half that of the fourth upsampling layer; the convolution parameters of the second convolution combination are exactly the same as those of the first convolution combination.
[0197] The sixth upsampling layer is a single convolutional layer with a kernel size of 31×1 (stride of 1 and padding of 15). The number of channels is changed from 16 to 1. It restores the multidimensional features to a feature map sequence of the same length as the input of the encoder module. This output feature map sequence is the preprocessed BCG corresponding to the ECG mapping.
[0198] The input feature map dimensions of each of the six upsampling layers in the decoder module are [16,512,19], [16,256,38], [16,128,76], [16,64,152], [16,32,304], and [16,16,304], respectively. After a single convolution in the last sixth upsampling layer, the final output ECG feature map dimension is [16,1,304].
[0199] Through layer-by-layer upsampling and transposed convolution operations in the decoder module, skip connections are used to directly fuse the local features extracted from the encoder with the global features in the decoder. Finally, the decoder maps the local features of the encoder back to the original signal space dimension, outputting a preprocessed BCG corresponding to the ECG.
[0200] Step S22: Train the ECG-BCG mapping model using the first training dataset until the loss function converges to obtain the trained ECG-BCG mapping model.
[0201] Load the first training dataset into the ECG-BCG mapping model for training;
[0202] A fixed learning rate is used, and forward propagation is performed using the MSE (mean squared error) loss function, while backpropagation is performed using the Adam optimizer. The parameters of the ECG-BCG mapping model are continuously adjusted to minimize the loss function until the loss function converges to obtain the trained ECG-BCG mapping model.
[0203] (1) Construct the loss function, determine the optimizer and hyperparameters, input the training set and validation set, and train the ECG-BCG model parameters;
[0204] (2) Forward propagation is performed using the MSE (mean squared error) loss function, as shown in formula (3).
[0205]
[0206] N is the total number of points contained in each cardiac impact signal sequence after non-overlapping segmentation, i.e., the length of the signal sequence, which is 304 in this embodiment. i This is the normalized amplitude value of the cardiac impact signal. The cardiac impact signal value is obtained after ECG-BCG mapping.
[0207] (3) Backpropagation is performed using the Adam optimizer, with an initial learning rate set to 1e. -4 The value is 0.0001, with a fixed learning rate, a batch size of 16, and a weight decay of 1e. -5 That is, 0.00001, with a preset maximum training rounds of 150;
[0208] For example, the software framework used in this invention to train the ECG-BCG mapping model is PyTorch 1.7.1 based on Python 3.8, the operating system is Ubuntu 20.0.4, the CUDA version is 11.3, and the hardware computing platform is a server equipped with an Nvidia RTX A40 graphics card.
[0209] (4) Train the network model and optimize the hyperparameters on the training set and validation set using the backpropagation algorithm until the loss function converges, thus completing the training.
[0210] First, train the software for 10 epochs using the training set, then validate it for 1 epoch using the validation set; then train it for another 10 epochs using the training set, then validate it for 1 epoch using the validation set.
[0211] The process continues until the loss function converges, then validates the model using a validation set to obtain the trained ECG-BCG mapping model. The model weight parameters are then saved. Finally, the weight parameters of the previously trained mapping model are loaded onto the constructed ECG-BCG mapping model.
[0212] Step S3, specifically.
[0213] ECG data is obtained from the second public dataset and preprocessed. It is then input into the trained ECG-BCG mapping model to obtain the corresponding BCG data. The heartbeat type of the corresponding ECG in the second public dataset is used as a label to form the second training dataset.
[0214] ECG signals and corresponding heart rate types were obtained from the second public dataset.
[0215] The second publicly available dataset is the electrocardiogram (ECG) signal dataset released by MIT-BIH, which includes ECG signals and corresponding heartbeat types.
[0216] (1) Obtain the original ECG format and corresponding label (i.e. heartbeat type), and perform data preprocessing;
[0217] For example, obtain the ECG in .dat format and its corresponding heart rate type, where the heart rate type is the type of heart rate abnormality disease.
[0218] In this embodiment, the second public dataset comes from the ECG signal dataset published by MIT-BIH, and the initial signal sequence length is 650,000.
[0219] Each ECG sequence in the public dataset corresponds to a specific type of heart rate abnormality.
[0220] The ECG data in this public dataset corresponds to the heart rate type, which is labeled by the doctor.
[0221] Using a loop function, the ECG is divided into a non-overlapping sequence of 304 data points, with the position of the first R peak as the reference.
[0222] For example, the ECG segments after non-overlapping segmentation are saved as a two-dimensional tensor format based on Python 3.8, with a feature map size of [2138, 304].
[0223] Normalize the ECG data after non-overlapping segmentation.
[0224] (2) After normalizing the non-overlapping segmented ECG, input it into the trained ECG-BCG mapping model to obtain BCG of the same size. Map the BCG and the corresponding ECG heartbeat type as samples of the second training dataset to form the second training dataset; divide it into training set, validation set and test set according to the preset ratio.
[0225] This step effectively increased the number of BCG samples.
[0226] Step S4 includes steps S41-S42.
[0227] Step S41: Construct a brain-like classification model based on a spiking neural network.
[0228] The brain-like classification model is built on a spiking neural network and includes, in sequence, a static convolutional layer, a channel attention layer, and a fully connected classification module.
[0229] The static convolutional layer includes a one-dimensional convolutional layer Conv1d, a LIFNode brain-like activation layer, and a pooling layer. The one-dimensional convolutional layer Conv1d is used for channel expansion to extract features. After the input feature map is convolved by Conv1d and activated by LIFNode brain-like activation, the number of channels is expanded from 1 to 4, and the sequence length is reduced from 304 to 284. Then, the pooling layer downsamples the feature map with expanded channels and reduced sequence length, and outputs it as X.
[0230] The channel attention layer includes a channel weight branch and a pass-in input branch;
[0231] The channel weight branch is combined with a custom convolution using the AdaptiveMaxPool1d function to output attention values for different channels, denoted as A.
[0232] The input branch first performs a multiplication operation X×A, then passes through a convolutional layer and a pooling layer for downsampling, and outputs a feature map of a three-dimensional tensor to the fully connected classification module for classification.
[0233] The fully connected classification module includes a Flatten layer and a linear layer. The feature map of the input three-dimensional tensor is flattened into a two-dimensional tensor by the Flatten function, and then passed through the linear layer to obtain the classification prediction probability of the heartbeat type; the heartbeat classification result is obtained based on the classification prediction probability.
[0234] After passing through the linear layer, the output result is the classification prediction probability result at a certain time. The entire brain-like classification model network runs the brain-like classification model training process from 0 to T at a preset brain-like time step T (T is an integer).
[0235] From 0 to T-1 times; each linear layer outputs a classification prediction probability result;
[0236] Then, the average of the classification prediction probabilities of the T brain-like time steps is taken, and this average is used as the classification prediction probability of the heartbeat type.
[0237] Table 2: Structural Parameters of Brain-Inspired Classification Models
[0238]
[0239]
[0240] like Figure 4 As shown:
[0241] (1) The first static convolutional layer is a one-dimensional convolutional layer Conv1d. The input data is the output feature map of the sixth upsampling layer [16,1,304]. After convolution, the input data is expanded in channels. The convolutional kernel size is 21×1, the stride is 1, the padding is 10, which is expanded from 1 to 4. The sequence length dimension is reduced from 304 to 284. The output feature map dimension is [16,4,284].
[0242] After passing through the LIFNode brain-like activation function layer, it is then downsampled through a one-dimensional pooling layer with a pooling window size of 2×1 and the sequence length dimension divided by 2; the final output feature map dimension of the static convolutional layer is [16,4,142].
[0243] For example, suppose the batch size of the input model is 16, i.e., B=16, and the sequence length is 300, i.e., L=300. Then the dimension of the input feature map of each layer of the static convolutional layer is [16,1,304], [2,4,284], and the dimension of the final output feature map after the pooling layer is [16,4,142].
[0244] The LIFNode brain-like activation layer, also known as the brain-like activation module, has a calculation mechanism as shown in formula (4):
[0245]
[0246] Where T is the preset brain-like time step, V[T] and V[T-1] are the brain-like neuron membrane potentials at brain-like time steps T and T-1, respectively, Y[T] is the feature map output by Conv1d convolution, and τ m V is the preset membrane time constant. reset To reset the potential;
[0247] If V[T] reaches the preset discharge threshold V threshold If the membrane potential returns to zero, V[T] is input to the substitution function of the step function for back propagation; the substitution function is shown in formula (5):
[0248]
[0249] Otherwise, if V[T] does not reach the preset discharge threshold V threshold Then the membrane potential of brain-like neurons remains V[T].
[0250] At T=0, brain-like neurons do not fire, they only charge;
[0251] When T>0, brain-like neurons both fire and charge.
[0252] For example, the membrane time constant τ m The default value is 2, and the reset potential V is... reset The preset value is 0, and the discharge threshold V is... threshold The default value is 1.0.
[0253] For the membrane potential V[T], the actual discharge process can be represented as a discrete step function, as shown in equation (6):
[0254] H(V(T)-V threshold )Formula (6)
[0255] Where H is the Heavydide step function, which is not differentiable. In order to calculate the backpropagation gradient, a substitute function σ(V[T]) is introduced.
[0256] The arctangent function arctan in the substitution function σ(V[T]) provides a smooth approximation when V[T] is close to the preset discharge threshold V. threshold The output of the substitution function is close to 1 or 0, and the substitution function is differentiable near the preset discharge threshold.
[0257] During the computation of the LIFNode brain-like activation layer, the dimension of the input feature map remains unchanged.
[0258] LIFNode represents the size of the output of the brain-like activation layer.
[0259] LIFNode is a computational model that simulates the firing and quiescent behavior of biological neurons, mimicking their firing and quiescent mechanisms. It reflects the working mechanism of real neurons to a certain extent and exhibits brain-like characteristics.
[0260] The LIFNode activation function module is used to increase the model's sensitivity to the temporal dynamics of ECG signals and improve its ability to recognize cardiac activity patterns. The brain-like design of the LIFNode activation layer avoids the neuronal saturation problem that may occur with traditional activation functions (such as Sigmoid or Tanh), thus maintaining the network's activity during training. Using the arctangent function as a replacement for the step function provides a differentiable path for backpropagation, allowing the network to be trained using optimization algorithms such as gradient descent.
[0261] (2) The channel attention layer contains two branches: the channel weight branch and the input propagation branch.
[0262] (2.1) The first branch is the channel weight branch. The input data is the output of the static convolutional layer. It is first combined with the AdaptiveMaxPool1d function and convolution to obtain the feature map dimension [16,4,142], which is set as A;
[0263] The convolution combination is a custom convolution combination structure, consisting of four layers: convolution (channel number 4→1) + ReLU activation function + convolution (channel number 1→4) + Sigmoid activation. The custom convolution combination does not change the number of channels, and the final result is still [16, 4, 142].
[0264] Custom convolutional combinations can be viewed as an implementation of an attention mechanism. Through this structure, the network can learn the importance or weights of different channels, thus focusing on features more relevant to the current task.
[0265] Convolution (4 channels → 1): Compresses the 4 channels of information at each position into 1 channel. This can be regarded as a selection process, in which the network learns which channels are important.
[0266] Calculate the sequence length dimension using formula (2)
[0267] ReLU activation function: Introduces non-linearity, enabling the model to learn and simulate more complex feature transformations;
[0268] Convolution (channel number 1→4): The compressed information is expanded back to 4 channels, but the weights of each channel have been adjusted.
[0269] Sigmoid activation: ensures the output is between 0 and 1, which can be seen as a normalization of the channel weights;
[0270] Custom convolutional combinations do not change the data dimension; the number of channels in the output is the same as the number of channels in the input, thus maintaining the channel dimension of the original data and facilitating fusion with the original data or the output of other branches.
[0271] The final output feature map dimension of the channel weight branch is [16, 4, 142].
[0272] (2.2) The second branch is the input passing branch, which includes weight multiplication, convolutional layer and pooling layer in sequence.
[0273] The original input and the extracted weights are multiplied together, X×A;
[0274] The result is then passed through a convolutional layer with a kernel size of 21×1 (stride of 1, padding of 10) and a pooling layer for downsampling (pooling window size of 2×1), and the output is [16,8,61], which is then input into the fully connected classification module for classification.
[0275] The output feature map of the channel weight branch has dimensions [16, 4, 142], which is used as the input of the pass-through input branch.
[0276] The dimension of the output feature map of the first layer of the input branch remains unchanged, still being [16,4,142];
[0277] The second layer, Conv1d, multiplies the number of channels by 2, reduces the sequence length dimension to 122, and outputs a feature map with dimensions of [16, 8, 122].
[0278] The third pooling layer has the same number of channels, but the sequence length dimension is halved, and the output feature map dimension is [16, 8, 61].
[0279] (3) The fully connected classification module includes a Flatten function and a linear layer. The input [16,8,61] is first flattened into one dimension [16,488] by the Flatten function, and then enters the linear layer to get [16,5], which gives the heartbeat classification probability. The heartbeat classification result is obtained based on the heartbeat classification probability.
[0280] Step S42: Train the brain-like classification model using the second training dataset.
[0281] The second training dataset is loaded into the brain-like classification model for training until the maximum number of iterations is reached.
[0282] Construct the loss function, determine the optimizer and hyperparameters, input the training set and validation set, and train the network model parameters;
[0283] Forward propagation is performed using the MSE (mean squared error) loss function:
[0284] Backpropagation is performed using the Adam optimizer, with an initial learning rate set to 1e. -4 The learning rate is 0.0001, the batch size is 16, the time step of the spiking neural network is 7, the time constant is 2, the pulse voltage threshold is 1, the reset potential after discharge is 0, and the maximum total training iterations are 150.
[0285] Train until the maximum number of iterations to obtain a trained brain-like classification model. Save the parameters of the brain-like classification model that achieve the highest classification accuracy in the validation set as the parameters of the trained brain-like classification model.
[0286] For example, the software framework used in the experimental training network of this invention is PyTorch 1.7.1 based on Python 3.6, the operating system is Ubuntu 20.0.4, the CUDA version is 11.3, and the hardware computing platform is a server equipped with an Nvidia RTX A40 graphics card.
[0287] The pre-trained weight parameters were loaded onto the constructed brain-like classification model, and the pre-processed BCG signal was input to obtain five predicted types of heart rate diseases.
[0288] Heart rate type refers to the category of heart rate abnormality, with 5 categories, using one-hot encoding (0-00001, 1-00010, 2-00100, 3-01000, 4-10000), as follows:
[0289] 0 - Normal;
[0290] Class 1-S disorders (premature atrial contractions, abnormal premature atrial contractions, junctional premature contractions, supraventricular premature contractions);
[0291] Class 2-V disorders (premature ventricular contractions, ventricular escape beats);
[0292] Class 3-F disease (ventricular fusion heartbeat);
[0293] Class 4-Q disorders (paced heartbeat, paced fusion heartbeat, unclassified heartbeat).
[0294] In summary, the brain-like classification model construction method for BCG according to the embodiments of the present invention has the following beneficial effects:
[0295] 1. By using the ECG-BCG mapping model, the number of BCG samples was effectively increased, solving the problem of insufficient sample quantity in the BCG dataset and improving the utilization rate of BCG samples.
[0296] 2. An ECG-BCG mapping model was established using the encoder-decoder structure of a U-shaped convolutional network, which improved the model's ability to extract local and global features of ECG and BCG signals and enhanced the richness of feature representation.
[0297] 3. A brain-like classification model based on a spiking neural network improves the accuracy of heartbeat classification through the LIFNode brain-like activation module and channel attention mechanism;
[0298] 4. By integrating the NIH's BED-BASED BALLISTOCARDIOGRAPHY DATASET and the MIT-BIH's ECG dataset, a comprehensive and accurately labeled training dataset was constructed, providing high-quality data support for the training of brain-like classification models and achieving effective integration of public datasets.
[0299] 5. Compared with existing abnormal heartbeat detection methods, this invention makes better use of the mechanical information contained in BCG, and can more effectively detect cardiac mechanical motion. The heart-brain classification model has broad application prospects in clinical diagnosis, long-term monitoring and personalized treatment.
[0300] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for constructing a brain-like classification model oriented towards BCG, characterized in that, Includes the following steps: Construct the first training dataset, which includes the preprocessed ECG and the corresponding BCG; The ECG-BCG mapping model is trained using the first training dataset until the loss function converges to obtain the trained ECG-BCG mapping model. An ECG-BCG mapping model is constructed based on the encoder-decoder structure of a U-shaped convolutional network, including an encoder module and a decoder module; The encoder module is used to downsample the received preprocessed ECG to extract local features and output a feature map matrix representing the local features. The decoder module restores the spatial dimension through a transposed convolutional layer, and fuses the local features extracted by the encoder module with the global features upsampled by the decoder module through skip connections to restore the original dimensional space, and outputs the BCG corresponding to the preprocessed ECG. The feature map matrix is represented as a three-dimensional tensor, with the three dimensions being batch size, number of channels, and sequence length, respectively. Based on the trained ECG-BCG mapping model, a second training dataset for the brain-like classification model is constructed; wherein, the second training dataset includes BCG and the corresponding ECG heartbeat type; The brain-like classification model is trained using the second training dataset, and the trained brain-like classification model is obtained after reaching the maximum number of iterations. The brain-like classification model is built on a spiking neural network and includes, in sequence, a static convolutional layer, a channel attention layer, and a fully connected classification module. The static convolutional layer includes a one-dimensional convolutional layer Conv1d, a LIFNode brain-like activation layer, and a pooling layer. The one-dimensional convolutional layer Conv1d is used for channel expansion to extract features. After the input feature map is convolved by Conv1d and activated by LIFNode brain-like activation, the number of channels is expanded from 1 to 4, and the sequence length is reduced from 304 to 284. Then, the pooling layer downsamples the feature map with expanded channels and reduced sequence length, and outputs it as X. The channel attention layer includes a channel weight branch and a pass-in input branch; The channel weight branch is combined with a custom convolution by the AdaptiveMaxPool1d function to output the attention value of different channels, which is set as A; The input branch first performs a multiplication operation between X and A, then passes through a convolutional layer and a pooling layer for downsampling, and outputs a feature map of a three-dimensional tensor to the fully connected classification module for classification. The fully connected classification module includes a Flatten layer and a linear layer. The feature map of the input three-dimensional tensor is flattened into a two-dimensional tensor by the Flatten function, and then passed through the linear layer to obtain the classification prediction probability of the heartbeat type; the heartbeat classification result is obtained based on the classification prediction probability.
2. The method according to claim 1, characterized in that, The encoder module is a downsampling structure, comprising five downsampling layers, wherein: The first downsampling layer consists of a convolutional layer with a stride of 1, layer normalization, and PReLU activation. The convolutional layer increases the number of channels in the input feature map from 1 to 32 before outputting it to the second downsampling layer. The second to fifth downsampling layers consist of two convolutional layers with a stride of 1 and a max pooling layer, respectively. After each convolutional layer, layer normalization and PReLU activation are performed in sequence. The first convolutional layer is used to double the number of channels in the feature map input to the current layer. The second convolutional layer keeps the number of channels unchanged and halves the sequence length dimension of the feature map input to the current layer through the max pooling layer, thus outputting the feature map with the changed number of channels and sequence length dimension. Layer normalization is used to accelerate convergence, and PReLU activation introduces nonlinearity to enhance the encoder module's sensitivity to subtle changes in local features.
3. The method according to claim 2, characterized in that, The decoder module is an upsampling structure, comprising six upsampling layers, as follows: The first upsampling layer consists of two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed in sequence. The two convolutional operations with a stride of 1 keep the number of channels unchanged. The transposed convolutional layer is used to halve the number of channels and double the sequence length, thereby achieving spatial expansion and upsampling of features. The second to fourth upsampling layers each consist of two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed sequentially. The first convolutional layer with a stride of 1 halves the number of channels in the feature map input to the current layer. The second convolutional layer with a stride of 1 keeps the number of channels unchanged, and then uses a transposed convolutional layer to halves the number of channels and doubles the sequence length. The input of the second upsampling layer and the output of the fourth downsampling layer, the input of the third upsampling layer and the output of the third downsampling layer, and the input of the fourth upsampling layer and the output of the second downsampling layer are sequentially connected in a skip connection to fuse the downsampled local features and the upsampled global features. The second to fourth upsampling layers respectively output the feature maps with altered channel number and sequence length to the next layer for upsampling; The fifth upsampling layer includes two convolutional layers with a stride of 1 and a transposed convolutional layer. After each convolutional layer with a stride of 1, layer normalization and PReLU activation are performed in sequence. The first convolutional layer with a stride of 1 halves the number of channels to refine the feature representation. The second convolutional layer with a stride of 1 keeps the number of channels unchanged and further extracts and fuses the downsampling and upsampling features. The sixth upsampling layer uses a single convolutional layer to reduce the number of channels from 16 to 1, mapping the decoded feature map back to the dimension of the feature map input to the encoder module, and outputting the ECG-mapped BCG.
4. The method according to claim 1, characterized in that, The first training dataset is obtained by preprocessing the raw signal sequences of ECG and corresponding BCG data from the first publicly available dataset. The steps include: Starting from the first R peak of the original signal sequence of the ECG, and starting from the J peak in the original signal sequence of the BCG corresponding to the timestamp of the first R peak, non-overlapping segments are performed with a preset sequence length. Data points before the first R peak in ECG and the first J peak in BCG, as well as data points that are less than the preset sequence length after segmentation, are discarded. Normalize the amplitudes of the segmented ECG and BCG; The normalized ECG and its corresponding BCG constitute each sample in the first training dataset.
5. The method according to claim 4, characterized in that, Load the first training dataset into the ECG-BCG mapping model for training; A fixed learning rate is used, and forward propagation is performed using the MSE (mean squared error) loss function, while backpropagation is performed using the Adam optimizer. The parameters of the ECG-BCG mapping model are continuously adjusted to minimize the loss function until the loss function converges to obtain the trained ECG-BCG mapping model.
6. The method according to claim 1, characterized in that, ECG data is obtained from the second public dataset and preprocessed. It is then input into the trained ECG-BCG mapping model to obtain the corresponding BCG data. The heartbeat type of the corresponding ECG in the second public dataset is used as a label to form the second training dataset. The second training dataset is loaded into the brain-like classification model for training until the maximum number of iterations is reached.
7. The method according to claim 1, characterized in that, The computation mechanism of the LIFNode neuromorphic activation layer is as follows: Where T is the preset brain-like time step, , These represent the neuronal membrane potentials at neuronal time steps T and T-1, respectively. The feature map output by Conv1d convolution. The preset membrane time constant, To reset the potential; if Reaching the preset discharge threshold Then the membrane potential returns to zero, and The substitution function, input to the step function, is backpropagated; the substitution function is as follows: Otherwise if The preset discharge threshold was not reached. Then the membrane potential of brain-like neurons remains .
8. The method according to claim 4 or 6, characterized in that, The first publicly available dataset is the NIH-published BED-BASEDBALLISTOCARDIOGRAPHY DATASET; where ECG and BCG are matched one-to-one based on timestamps; The second publicly available dataset is the ECG signal dataset released by MIT-BIH, which includes ECG signals and corresponding heartbeat types.
Citation Information
Patent Citations
LSTM-based automatic multi-classification identification method for ballistocardiogram signals
CN110427924A
Intelligent electrocardiosignal generation method based on multi-scale feature fusion of ballistocardiogram signals
CN115935295A