Multi-sensor monitoring system and method based on time alignment
By using a time-aligned multi-sensor monitoring system, combined with multimodal fusion analysis and deep learning networks, the problem of insufficient accuracy and precision in sleep monitoring by single-source monitoring systems is solved. This achieves high-precision alignment and fusion of multiple signals, improving the accuracy and interpretability of atrial fibrillation monitoring.
Patent Information
- Application Number
- CN202610173353.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-06
- Publication Date
- 2026-03-17
AI Technical Summary
Existing single-source monitoring systems have limited accuracy and precision in specific scenarios (such as sleep monitoring), making it difficult to effectively monitor the probability of atrial fibrillation.
A time-aligned multi-sensor monitoring system is adopted. The sensor acquisition module acquires electrocardiogram signals, photoplethysmography pulse wave signals, heart sound signals and respiratory impedance signals. These signals are aligned by the time alignment module. The multimodal fusion analysis module extracts local morphological features and dynamic weights. 1D-CNN and Bi-LSTM networks are used to capture long-term temporal dependencies and generate a 2D vector of atrial fibrillation probability.
It achieves high-precision alignment and fusion of multiple signals, improving the accuracy and interpretability of atrial fibrillation monitoring. It is applicable to various scenarios from static sleep monitoring to daily activities, assisting medical staff in making quick judgments.
Smart Images

Figure CN121667653A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-sensor monitoring technology, and specifically to a time-aligned multi-sensor monitoring system and method. Background Technology
[0002] Atrial fibrillation (AF) is a common cardiac arrhythmia where the heart loses its normal contractile function and exhibits irregular, flaring movements. Stroke is a complication of AF. Patients with AF are five times more likely to have a stroke than the general population, and approximately 20% of strokes are directly caused by AF. Several monitoring systems have been designed to better monitor AF.
[0003] For example, Chinese patent CN114469133B discloses a non-disruptive atrial fibrillation (AF) monitoring method, which includes the following steps: Step 1: First, the BCG signal is preprocessed by separating it from the original signal using denoising technology; Step 2: Then, the BCG signal is segmented; Step 3: Finally, the segmented segments are used to extract AF features through a deep learning model, and a softmax classifier is used to achieve non-disruptive AF monitoring. This method can achieve non-disruptive monitoring of AF during sleep based on BCG signals; and the proposed MS-DenseNet network structure can improve the accuracy of AF detection through multi-scale feature learning.
[0004] However, it has value in specific scenarios (such as sleep monitoring), but it is limited by its reliance on a single information source, and its accuracy and precision need to be improved.
[0005] Based on this, the present invention designs a time-aligned multi-sensor monitoring system and method to solve the above problems. Summary of the Invention
[0006] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a multi-sensor monitoring system and method based on time alignment.
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] A time-aligned multi-sensor monitoring system includes:
[0009] Sensor acquisition module: used to acquire ECG signals, photoplethysmography pulse wave signals, heart sound signals, and respiratory impedance signals containing master clock reception timestamps;
[0010] Time alignment module: used to time-align ECG signals, photoplethysmography (PPG) signals, heart sound signals, and respiratory impedance signals containing master clock receive timestamps, forming time-aligned ECG signals, PPG signals, heart sound signals, and respiratory impedance signals;
[0011] Multimodal fusion analysis module: Extracts local morphological features from time-aligned ECG signals, photoplethysmography (PPG) signals, heart sounds, and respiratory impedance signals to form feature maps; extracts dynamic weights between different signals and between different time points / features within the same signal; and then fuses these weights with the feature maps to form a weighted feature map; captures long-range temporal dependencies in the weighted feature maps to form a high-level semantic feature sequence of multi-signal, long-time-series contextual information; and generates a 2D vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence.
[0012] Results output module: Generates analysis reports including disease diagnosis and decision-making basis based on the analysis and processing results.
[0013] Furthermore, each sensor acquisition module is connected to a local high-precision clock synchronized with the main control unit. After the data acquired by each sensor acquisition module completes analog-to-digital conversion, the local high-precision clock immediately generates a master clock receiving timestamp for the sampled data.
[0014] Furthermore, the local high-precision clock has a time synchronization accuracy of less than 50ms.
[0015] Furthermore, the multimodal fusion analysis module includes a feature extraction unit, a cross-channel attention mechanism unit, a temporal modeling unit, and a comprehensive analysis unit;
[0016] Feature extraction unit: used to extract local morphological features within each signal in the weighted time-aligned ECG signal, photoplethysmography pulse wave signal, heart sound signal, and respiratory impedance signal, forming a feature map;
[0017] Cross-channel attention mechanism unit: used to extract dynamic weights between different signals and between different time points / features within the same signal, and then fuse them with the feature map to form a weighted feature map;
[0018] Temporal modeling unit: used to capture long-range temporal dependencies in the weighted feature map, forming a high-level semantic feature sequence with multiple signals and long temporal context information;
[0019] Comprehensive analysis unit: Generates a 2D vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence.
[0020] A monitoring method employing a time-aligned multi-sensor monitoring system includes the following steps:
[0021] Step 1: Acquire the ECG signal, photoplethysmography pulse wave signal, heart sound signal, and respiratory impedance signal containing the master clock reception timestamp;
[0022] Step 2: Time-align the acquisition of ECG signals, photoplethysmography (PPG) signals, heart sound signals, and respiratory impedance signals containing the master clock receiving timestamp to form time-aligned ECG signals, PPG signals, heart sound signals, and respiratory impedance signals;
[0023] Step 3: Extract local morphological features from time-aligned ECG, photoplethysmography (PPG), heart sound, and respiratory impedance signals to form feature maps. Extract dynamic weights between different signals and between different time points / features within the same signal, and then fuse them with the feature maps to form a weighted feature map. Capture the long-range temporal dependencies of the weighted feature maps to form a high-level semantic feature sequence with multiple signals and long-term temporal contextual information. Generate a 2D vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence.
[0024] Step 4: Determine if the probability of atrial fibrillation is greater than the threshold. If it is, generate an analysis report containing the probability value of atrial fibrillation and a suspected atrial fibrillation. If it is not, it is normal and monitoring continues.
[0025] Furthermore, after the ECG, photoplethysmography (PPG), heart sound, and respiratory impedance signals acquired by each sensor acquisition module are converted from analog to digital, the local high-precision clock immediately generates a master clock receiving timestamp for the sampled data. At the same time, the master clock receiving timestamp is bound to the ECG, PPG, heart sound, and respiratory impedance signals one by one and transmitted to the time alignment module.
[0026] Furthermore, the specific steps for step 2 are as follows:
[0027] Step 21: Receive the ECG signal, photoplethysmography pulse wave signal, heart sound signal, and respiratory impedance signal containing the master clock reception timestamp, and store them in the streaming alignment buffer;
[0028] Step 22: Based on the preset common timeline, for the target time point, extract at least two original data points from the streaming alignment buffer whose master clock received timestamps are closest to the target time point;
[0029] Step 23: Based on the values of the original data points and their master clock receiving timestamps, interpolation is used to calculate the estimated values of each sensor at the target time point;
[0030] Step 24: Combine and output the estimated values of all sensor acquisition modules at the target time point to form a fused data frame with acquisition time alignment.
[0031] Furthermore, step 3 is performed as follows:
[0032] Step 31: Extract local morphological features from the time-aligned electrocardiogram (ECG), photoplethysmography (PPG), heart sound, and respiratory impedance signals using a 1D-CNN network to form ECG feature maps, PPG feature maps, heart sound feature maps, and respiratory impedance signal feature maps, respectively.
[0033] Step 32: Extract the dynamic weights between aligned ECG signals, photoplethysmography (PPG) signals, heart sound signals, and respiratory impedance signals, as well as between different time points / features within the same signal, using a cross-channel attention mechanism. Then, fuse these weights with the feature maps to form weighted ECG signal feature maps, weighted PPG signal feature maps, weighted heart sound signal feature maps, and weighted respiratory impedance signal feature maps.
[0034] Step 33: Based on the Bi-LSTM network, capture the long-range temporal dependencies of the weighted ECG signal feature map, the weighted photoplethysmography pulse wave signal feature map, the weighted heart sound signal feature map, and the weighted respiratory impedance signal feature map to form a high-level semantic feature sequence with multiple signals and long-term contextual information.
[0035] Step 34: Generate a 2D vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence.
[0036] Furthermore, the specific steps in step 31 are as follows:
[0037] Step 311: Weights of the 32 convolutional kernels of length 64 in 1D Random initialization learning was performed to form 32 different convolutional kernels;
[0038] Step 312: Take a convolution kernel and align it with the 64 data points of the aligned ECG signal starting from the beginning point to obtain the starting point time step. Window data;
[0039] Step 313: Calculate the 64 weights of the convolution kernel and the corresponding starting point and time step of the 64 data points. dot product of window data , forming the first The initial feature sequence of each convolution ;
[0040] Step 314: Place the first Each convolution start point time step initial feature sequence The input is fed into a non-linear activation function for activation processing, and then the activation sequence is pooled according to the set... Each convolution start point time step The pooling window slides over the active data, retrieving the first data at a time. Each convolution start point time step The maximum value within the pooled window data;
[0041] Step 315: Next, move the convolution kernel one step to the right, repeat step 314, calculate the maximum value within the next pooling window, and after sliding the convolution kernel and the aligned ECG signal upwards, group all the dot product calculation results into a set in sequence. Each convolution start point time step ECG signal characteristics ;
[0042] Step 316: The remaining 31 convolutional kernels in the CNN layer simultaneously perform operations 312-315, resulting in 31 sets of ECG signal feature maps, thus obtaining 32 ECG signal feature maps. , ;
[0043] Step 317: Repeat steps 312-316 independently for the photoplethysmography (PPG) signal, heart sound signal, and respiratory impedance signal to obtain 32 PPG signal feature maps respectively. 32 heart sound signal feature maps and 32 respiratory impedance signal feature maps .
[0044] Furthermore, step 32 consists of the following specific steps:
[0045] Step 321: Extract the ECG signal feature map from Step 31. Photoplethysmography (PPG) signal characteristics Heart sound signal characteristic diagram and respiratory impedance signal characteristic map Concatenate along the feature dimensions to obtain a two-dimensional feature tensor. , Then Convert to feature sequence ;
[0046] Step 322: Use three independent, trainable weight matrices , , For the vector at each time step Perform a linear transformation to obtain the query vector sequence Q, the key vector sequence K, and the value vector sequence V;
[0047] Step 323: Query the key values in the global query vector and key vector series K. Calculate similarity Use the Softmax function to combine all similarities Convert to time steps Attention weights, generating attention weight sequences Ensure that the sum of all attention weights is 1;
[0048] Step 324: Convert each time step Attention weights are applied to the value vectors in the value vector sequence V to obtain the context vector. ;
[0049] Step 325: Assign attention weights to the sequence With two-dimensional feature tensor Multiplication yields a weighted fusion tensor Then, for the weighted fusion tensor The original channels were split into four weighted feature maps. , , and .
[0050] Beneficial effects: 1. This invention provides a unified time base for all sensors through a local high-precision clock, ensuring data alignment from the source; dynamically estimates and compensates for clock drift, achieving low-latency, high-precision alignment of continuous data streams; firstly, 1D-CNN extracts local morphological features (such as P waves and QRS complexes in ECG), then an attention mechanism adaptively fuses multi-channel information, followed by Bi-LSTM capturing long-term temporal patterns such as heart rate variability, forming a hierarchical feature understanding from micro to macro; finally, through a temporal attention mechanism, the model can dynamically focus on the key time period most relevant to atrial fibrillation detection (such as the arrhythmia episode), improving the interpretability and accuracy of the judgment, assisting medical staff in making quick judgments, and is applicable to various scenarios from static sleep monitoring to daily activity monitoring, with a wide range of applications. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0052] Figure 1 This is a flowchart of the time-aligned multi-sensor monitoring method of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0054] The present invention will be further described below with reference to embodiments.
[0055] Example 1: A time-aligned multi-sensor monitoring system, comprising the following steps:
[0056] Sensor acquisition module: used to acquire ECG signals, photoplethysmography pulse wave signals, heart sound signals, and respiratory impedance signals containing master clock reception timestamps;
[0057] Time alignment module: used to time-align ECG signals, photoplethysmography (PPG) signals, heart sound signals, and respiratory impedance signals containing master clock receive timestamps, forming time-aligned ECG signals, PPG signals, heart sound signals, and respiratory impedance signals;
[0058] Multimodal fusion analysis module: Extracts local morphological features from time-aligned ECG signals, photoplethysmography (PPG) signals, heart sounds, and respiratory impedance signals to form feature maps; extracts dynamic weights between different signals and between different time points / features within the same signal; and then fuses these weights with the feature maps to form a weighted feature map; captures long-range temporal dependencies in the weighted feature maps to form a high-level semantic feature sequence of multi-signal, long-time-series contextual information; and generates a 2D vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence.
[0059] Results output module: Generates analysis reports including disease diagnosis and decision-making basis based on the analysis and processing results.
[0060] Each sensor acquisition module is connected to a local high-precision clock that synchronizes with the main control unit;
[0061] After the data collected by each sensor acquisition module completes analog-to-digital conversion, the local high-precision clock immediately generates a master clock receiving timestamp for the sampled data;
[0062] The local high-precision clock has a time synchronization accuracy of less than 50ms.
[0063] The multimodal fusion analysis module includes a feature extraction unit, a cross-channel attention mechanism unit, a temporal modeling unit, and a comprehensive analysis unit;
[0064] Feature extraction unit: used to extract local morphological features within each signal in the weighted time-aligned ECG signal, photoplethysmography pulse wave signal, heart sound signal, and respiratory impedance signal, forming a feature map;
[0065] Cross-channel attention mechanism unit: used to extract dynamic weights between different signals and between different time points / features within the same signal, and then fuse them with the feature map to form a weighted feature map;
[0066] Temporal modeling unit: used to capture long-range temporal dependencies in the weighted feature map, forming a high-level semantic feature sequence with multiple signals and long temporal context information;
[0067] Comprehensive analysis unit: Generates a 2D vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence.
[0068] Example 2, please refer to Figure 1 The monitoring method employs a time-aligned multi-sensor monitoring system, with the following specific steps:
[0069] Step 1: Acquire the ECG signal, photoplethysmography pulse wave signal, heart sound signal, and respiratory impedance signal containing the master clock reception timestamp;
[0070] Step 2: Time-align the acquisition of ECG signals, photoplethysmography (PPG) signals, heart sound signals, and respiratory impedance signals containing the master clock receiving timestamp to form time-aligned ECG signals, PPG signals, heart sound signals, and respiratory impedance signals;
[0071] Step 3: Extract local morphological features from time-aligned ECG, photoplethysmography (PPG), heart sound, and respiratory impedance signals to form feature maps. Extract dynamic weights between different signals and between different time points / features within the same signal, and then fuse them with the feature maps to form a weighted feature map. Capture the long-range temporal dependencies of the weighted feature maps to form a high-level semantic feature sequence with multiple signals and long-term temporal contextual information. Generate a 2D vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence.
[0072] Step 4: Determine if the probability of atrial fibrillation is greater than the threshold. If it is, generate an analysis report containing the probability value of atrial fibrillation and a suspected atrial fibrillation. If it is not, it is normal and monitoring continues.
[0073] Step 1 Specific operation: After the ECG signal, photoplethysmography (PPG) signal, heart sound signal and respiratory impedance signal collected by each sensor acquisition module are converted from analog to digital, the local high-precision clock immediately generates a master clock receiving timestamp for the sampled data. At the same time, the master clock receiving timestamp is bound to the ECG signal, PPG signal, heart sound signal and respiratory impedance signal one by one and transmitted to the time alignment module.
[0074] The specific steps for step 2 are as follows:
[0075] Step 21: Receive the ECG signal, photoplethysmography pulse wave signal, heart sound signal, and respiratory impedance signal containing the master clock reception timestamp, and store them in the streaming alignment buffer;
[0076] Step 22: Based on the preset common timeline, for the target time point, extract at least two original data points from the streaming alignment buffer whose master clock received timestamps are closest to the target time point;
[0077] Step 23: Based on the values of the original data points and their master clock receiving timestamps, interpolation is used to calculate the estimated values of each sensor at the target time point;
[0078] Step 24: Combine and output the estimated values of all sensor acquisition modules at the target time point to form a fused data frame with acquisition time alignment;
[0079] The specific steps for interpolation calculation are as follows:
[0080] 1. Calculate the prior error
[0081]
[0082]
[0083]
[0084] For the first The master clock receives the timestamps of each data packet; For the first The slave clock timestamp of each data packet; For the first A vector of slave clock timestamps for each data packet; for Transpose matrix; For the first The estimated update state of each data packet;
[0085] 2. Calculate the gain vector
[0086]
[0087] For the first Update the covariance matrix values of each data packet;
[0088] It is a forgetting factor. Used to gradually reduce the impact of old data.
[0089] 3. Update the state estimate :
[0090]
[0091] 4. Update the covariance matrix values :
[0092]
[0093] It is the identity matrix;
[0094]
[0095] 5. Based on the updated state estimate :
[0096]
[0097]
[0098] For the first The clock drift rate of each data packet, ideally 1. This indicates that the slave clock is faster than the master clock; Indicates that the clock is slow; It is the first The clock offset value for each data packet.
[0099] Step 3 is performed as follows:
[0100] Step 31: Extract local morphological features from the time-aligned electrocardiogram (ECG), photoplethysmography (PPG), heart sound, and respiratory impedance signals using a 1D-CNN network to form ECG feature maps, PPG feature maps, heart sound feature maps, and respiratory impedance signal feature maps, respectively.
[0101] Step 32: Extract the dynamic weights between aligned ECG signals, photoplethysmography (PPG) signals, heart sound signals, and respiratory impedance signals, as well as between different time points / features within the same signal, using a cross-channel attention mechanism. Then, fuse these weights with the feature maps to form weighted ECG signal feature maps, weighted PPG signal feature maps, weighted heart sound signal feature maps, and weighted respiratory impedance signal feature maps.
[0102] Step 33: Based on the Bi-LSTM network, capture the long-range temporal dependencies of the weighted ECG signal feature map, the weighted photoplethysmography pulse wave signal feature map, the weighted heart sound signal feature map, and the weighted respiratory impedance signal feature map to form a high-level semantic feature sequence with multiple signals and long-term contextual information.
[0103] Step 34: Generate a 2D vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence.
[0104] The specific steps for step 31 are as follows:
[0105] Step 311: Weights of the 32 convolutional kernels of length 64 in 1D Random initialization learning was performed to form 32 different convolutional kernels;
[0106] No. The weights of each convolutional kernel The calculation is as follows:
[0107]
[0108] Step 312: Take a convolution kernel and align it with the 64 data points of the aligned ECG signal starting from the beginning point to obtain the starting point time step. Window data;
[0109] Window data The calculation is as follows:
[0110]
[0111] The starting point time step, Starting point time step Window data.
[0112] Step 313: Calculate the 64 weights of the convolution kernel and the corresponding starting point and time step of the 64 data points. dot product of window data , forming the first The initial feature sequence of each convolution ;
[0113]
[0114] For window data The i-th element in;
[0115]
[0116]
[0117] It is the length of the sequence after convolution. The length of the aligned electrocardiogram signal;
[0118] Step 314: Place the first Each convolution start point time step initial feature sequence The input is fed into a non-linear activation function for activation processing, and then the activation sequence is pooled according to the set... Each convolution start point time step The pooling window slides over the active data, retrieving the first data at a time. Each convolution start point time step The maximum value within the pooled window data;
[0119]
[0120]
[0121]
[0122]
[0123] For the first Each convolution start point time step The activated sequence; For the first Each convolution start point time step Pooled window data; For the first Each convolution start point time step The pooling output result, The length of the pooled sequence;
[0124] Step 315: Next, move the convolution kernel one step to the right, repeat step 314, calculate the maximum value within the next pooling window, and after sliding the convolution kernel and the aligned ECG signal upwards, group all the dot product calculation results into a set in sequence. Each convolution start point time step ECG signal characteristics ;
[0125] No. Each convolution start point time step ECG signal characteristics :
[0126]
[0127] Step 316: The remaining 31 convolutional kernels in the CNN layer simultaneously perform operations 312-315, resulting in 31 sets of ECG signal feature maps, thus obtaining 32 ECG signal feature maps. , ;
[0128] Step 317: Repeat steps 312-316 independently for the photoplethysmography (PPG) signal, heart sound signal, and respiratory impedance signal to obtain 32 PPG signal feature maps respectively. 32 heart sound signal feature maps and 32 respiratory impedance signal feature maps .
[0129] The non-linear activation function is the ReLU function;
[0130] The pooling window size is set to 2;
[0131] ECG signal characteristics include P wave shape characteristics, QRS complex shape characteristics, and T wave shape characteristics.
[0132] The photoplethysmography (PPG) signal feature map includes the main wave pulse wave feature and the dicrotic notch pulse wave feature.
[0133] The heart sound signal feature map includes S1 heart sound spectrum features, S2 heart sound spectrum features, and noise spectrum features.
[0134] The respiratory impedance signal feature map includes inspiratory cycle pattern features and expiratory cycle pattern features.
[0135] Step 32 is as follows:
[0136] Step 321: Extract the ECG signal feature map from Step 31. Photoplethysmography (PPG) signal characteristics Heart sound signal characteristic diagram and respiratory impedance signal characteristic map Concatenate along the feature dimensions to obtain a two-dimensional feature tensor. , Then Convert to feature sequence ;
[0137]
[0138] It contains all the characteristic information of four signals at time step t: electrocardiogram signal, photoplethysmography pulse wave signal, heart sound signal, and respiratory impedance signal;
[0139] Step 322: Use three independent, trainable weight matrices , , For the vector at each time step Perform a linear transformation to obtain the query vector sequence Q, the key vector sequence K, and the value vector sequence V;
[0140]
[0141]
[0142]
[0143]
[0144]
[0145]
[0146] Step 323: Query the key values in the global query vector and key vector series K. Calculate similarity Use the Softmax function to combine all similarities Convert to time steps Attention weights, generating attention weight sequences Ensure that the sum of all attention weights is 1;
[0147]
[0148]
[0149]
[0150] For each time step Attention weights;
[0151] This is the global query vector;
[0152] Step 324: Convert each time step Attention weights are applied to the value vectors in the value vector sequence V to obtain the context vector. ;
[0153]
[0154] , The cross-channel attention machine dimension;
[0155] Step 325: Assign attention weights to the sequence With two-dimensional feature tensor Multiplication yields a weighted fusion tensor Then, for the weighted fusion tensor The original channels were split into four weighted feature maps. , , and ;
[0156]
[0157]
[0158]
[0159]
[0160]
[0161] The specific steps for step 33 are as follows:
[0162] Step 331: Four weighted feature maps , , and Perform matrix transpose to obtain four sets of sequences. , , and ;
[0163] Step 332: Concatenate the four transposed sequences into a unified, comprehensive time-series feature sequence containing all signal information. ;
[0164]
[0165] Every time step A corresponding 128-dimensional feature vector fully encompasses the time steps. The weighted characteristics of ECG, pulse wave, heart sounds, and respiratory impedance, input In a two-layer Bi-LSTM network;
[0166] Step 333: Bi-LSTM forward LSTM layer: from arrive Sequential processing The forward hidden state sequence is obtained. , , ..., Bi-LSTM forward LSTM layer: from arrive Sequential processing The forward hidden state sequence is obtained. , , ..., );
[0167] Step 334: Hide the first layer of Bi-LSTM forward and backward As the input to the second-layer Bi-LSTM at time step t, the second-layer Bi-LSTM repeats step 333.
[0168] Step 335: Take the complete hidden state output of the last layer of the Bi-LSTM network at all time steps as the high-level semantic feature sequence. ;
[0169]
[0170] , It is twice the dimension of the hidden layer of Bi-LSTM. .
[0171] The specific steps for step 34 are as follows:
[0172] Step 341: Transform the high-level semantic feature sequence They are aggregated into a fixed-length global feature vector Z;
[0173] Introduce a trainable time-series query vector and weight matrix This is used to calculate the importance score for each time step t.
[0174]
[0175] Calculate the unnormalized attention score at each time step:
[0176]
[0177] The attention weights are obtained by performing Softmax normalization on the scores at all time steps.
[0178]
[0179] Weight Indicates the first The degree of contribution of the hidden state at each time step to the final classification.
[0180] It can dynamically focus on the key time periods most relevant to the diagnosis of atrial fibrillation (such as the episode of arrhythmia or the period when P waves are missing).
[0181] Step 342: Global Feature Vector Z and Context Vector The fusion feature vector is obtained by splicing and merging the features. ;
[0182] Step 343: Fuse feature vectors Perform full-face layer classification to generate a 2D vector containing [normal probability, atrial fibrillation probability].
[0183] By providing a unified time base for all sensors through a local high-precision clock, data alignment is ensured from the source. Clock drift is dynamically estimated and compensated for, achieving low-latency and high-precision alignment of continuous data streams. Local morphological features (such as P waves and QRS complexes in ECG) are first extracted by 1D-CNN, and then multi-channel information is adaptively fused by an attention mechanism. Next, long-term temporal patterns such as heart rate variability are captured by Bi-LSTM, forming a hierarchical feature understanding from micro to macro. Finally, the temporal attention mechanism enables the model to dynamically focus on the key time period most relevant to atrial fibrillation detection (such as the arrhythmia episode), improving the interpretability and accuracy of the judgment and assisting medical staff in making quick judgments. It is applicable to a variety of scenarios, from static sleep monitoring to daily activity monitoring, and has a wide range of applications.
[0184] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A time-aligned multi-sensor monitoring system based on: Comprise: Sensor acquisition module: for acquiring electrocardio signal, photoplethysmogram signal, heart sound signal and respiratory impedance signal containing master clock receiving timestamp; Time alignment module: for time alignment of electrocardio signal, photoplethysmogram signal, heart sound signal and respiratory impedance signal containing master clock receiving timestamp, forming time-aligned electrocardio signal, photoplethysmogram signal, heart sound signal and respiratory impedance signal; Multimodal fusion analysis module: extracting local morphological features in time-aligned electrocardio signal, photoplethysmogram signal, heart sound signal and respiratory impedance signal to form feature map, extracting dynamic weight between different signals and between different time points / features in the same signal, and then fusing with feature map to form weighted feature map; capture long-range temporal dependence of weighted feature map, form advanced semantic feature sequence of multi-signal, long-time sequence context information, and generate a 2-dimensional vector containing [normal probability, atrial fibrillation probability] according to the advanced semantic feature sequence; Result output module: generating analysis report including disease diagnosis and decision basis according to analysis processing result.
2. The time-alignment based multi-sensor monitoring system of claim 1, wherein, Each sensor acquisition module is connected with the local high-precision clock of the master control unit, and the data collected by each sensor acquisition module is immediately given a master clock receiving timestamp by the local high-precision clock after analog-to-digital conversion.
3. The time-alignment based multi-sensor monitoring system of claim 2, wherein, The time synchronization accuracy of the local high-precision clock is less than 50 ms.
4. The time-alignment based multi-sensor monitoring system of claim 3, wherein, The multimodal fusion analysis module comprises a feature extraction unit, a cross-channel attention mechanism unit, a time sequence modeling unit and a comprehensive analysis unit; Feature extraction unit: for extracting local morphological features inside each signal of time-aligned electrocardio signal, photoplethysmogram signal, heart sound signal and respiratory impedance signal, forming feature map; Cross-channel attention mechanism unit: for extracting dynamic weight between different signals and between different time points / features in the same signal, and then fusing with feature map to form weighted feature map; Time sequence modeling unit: for capturing long-range temporal dependence of weighted feature map, forming advanced semantic feature sequence of multi-signal, long-time sequence context information; Comprehensive analysis unit: generating a 2-dimensional vector containing [normal probability, atrial fibrillation probability] according to the advanced semantic feature sequence.
5. A monitoring method employing the time-alignment based multi-sensor monitoring system of claim 4, characterized by, Comprise the following steps: Step 1: acquiring electrocardio signal, photoplethysmogram signal, heart sound signal and respiratory impedance signal containing master clock receiving timestamp; Step 2: time alignment of electrocardio signal, photoplethysmogram signal, heart sound signal and respiratory impedance signal containing master clock receiving timestamp, forming time-aligned electrocardio signal, photoplethysmogram signal, heart sound signal and respiratory impedance signal; Step 3: Extract local morphological features in the time-aligned electrocardiogram signal, photoplethysmogram signal, heart sound signal, and respiratory impedance signal to form a feature map, extract dynamic weights between different signals and between different time points / features within the same signal, and then fuse the feature map to form a weighted feature map; capture long-range temporal dependencies of the weighted feature map to form a high-level semantic feature sequence of multi-signal and long-time context information, and generate a 2-dimensional vector containing [normal probability, atrial fibrillation probability] based on the high-level semantic feature sequence; Step 4: Determine whether the atrial fibrillation probability is greater than a threshold value. If the determination is yes, generate a suspected atrial fibrillation and an analysis report containing the atrial fibrillation probability value. If the determination is no, it is normal and continue monitoring.
6. The monitoring method according to claim 5, characterized in that, Step 1: The electrocardiogram signal, photoplethysmogram signal, heart sound signal, and respiratory impedance signal collected by each sensor acquisition module are immediately assigned a master clock reception timestamp by a local high-precision clock after analog-to-digital conversion, and the master clock reception timestamp is bound to the electrocardiogram signal, photoplethysmogram signal, heart sound signal, and respiratory impedance signal one by one and transmitted to the time alignment module.
7. The monitoring method according to claim 6, characterized in that, The specific operation of step 2 is as follows: Step 21: Receive the electrocardiogram signal, photoplethysmogram signal, heart sound signal, and respiratory impedance signal containing the master clock reception timestamp and store it in the streaming alignment buffer; Step 22: Based on the pre-set common time axis, for the target time point, take out at least two original data points with the closest master clock reception timestamp to the target time point from the streaming alignment buffer; Step 23: Calculate the estimated value of each sensor at the target time point based on the value of the original data point and its master clock reception timestamp through interpolation calculation; Step 24: Combine and output the estimated values of all sensor acquisition modules at the target time point to form a fusion data frame of acquisition time alignment.
8. The monitoring method according to claim 7, characterized in that, The specific operation of step 3 is as follows: Step 31: Extract local morphological features in the time-aligned electrocardiogram signal, photoplethysmogram signal, heart sound signal, and respiratory impedance signal through a 1D-CNN network to form an electrocardiogram signal feature map, photoplethysmogram signal feature map, heart sound signal feature map, and respiratory impedance signal feature map; Step 32: Extract dynamic weights between the aligned electrocardiogram signal, photoplethysmogram signal, heart sound signal, and respiratory impedance signal, and between different time points / features within the same signal through a cross-channel attention mechanism, and then fuse the feature map to form a weighted electrocardiogram signal feature map, weighted photoplethysmogram signal feature map, weighted heart sound signal feature map, and weighted respiratory impedance signal feature map; Step 33: Capture long-range temporal dependencies of the weighted electrocardiogram signal feature map, weighted photoplethysmogram signal feature map, weighted heart sound signal feature map, and weighted respiratory impedance signal feature map based on a Bi-LSTM network to form a high-level semantic feature sequence of multi-signal and long-time context information; Step 34: Generate a 2-dimensional vector containing [normal probability, atrial fibrillation probability] from the sequence of advanced semantic features.
9. The monitoring method according to claim 8, characterized in that, The specific operation of step 31 is as follows: Step 311: weights of 32 convolutional kernels of 1D of length 64 Randomly initialized learning is performed, shape 32 different convolutional kernels; Step 312: taking a convolution kernel, aligning the convolution kernel with the 64 data points of the starting point of the aligned electrocardio signal to obtain window data of the starting point time step ; Step 313: calculate the dot product of the 64 weights of the convolution kernel and the starting point time steps of the corresponding 64 data points of the window data , forming the initial feature sequence of the first convolution ; Step 314: input the initial feature sequence of the first convolution start point time step into a nonlinear activation function for activation processing, and then perform a pooling operation on the activation sequence, according to a set first convolution start point time step pooling window sliding on the activated data, and taking out the maximum value in the first convolution start point time step pooling window data each time Step 315: Then the convolution kernel moves right by one step, and step 314 is repeated to calculate the maximum value in the next pooling window data. After the convolution kernel is slid on the aligned electrocardio signal, all dot product calculation results are sequentially composed into a set of electrocardio signal feature maps of the convolution starting point time steps . Step 316: The remaining 31 convolution kernels in the CNN layer simultaneously perform the operations of steps 312-315 to obtain 31 groups of electrocardiosignal feature maps, thereby obtaining 32 electrocardiosignal feature maps , ; Step 317: Steps 312-316 are repeated independently for the photoplethysmographic pulse wave signal, the heart sound signal, and the respiratory impedance signal to obtain 32 photoplethysmographic pulse wave signal feature maps , 32 heart sound signal feature maps , and 32 respiratory impedance signal feature maps .
10. The monitoring method according to claim 9, characterized in that, The specific steps of step 32 are as follows: Step 321: the electrocardiosignal feature map output by step 31 , the photoplethysmogram signal feature map , the heart sound signal feature map , and the respiratory impedance signal feature map are spliced in the feature dimension to obtain a two-dimensional feature tensor , and then converted into a feature sequence ; Step 322: Use three independent, trainable weight matrices , , For the vector at each time step Perform a linear transformation to obtain the query vector sequence Q, the key vector sequence K, and the value vector sequence V; Step 323: Query the key value in the series K of key vectors by the global query vector Compute similarity Convert all similarities to attention weights at time step using a Softmax function, generating a sequence of attention weights ensuring that the sum of all attention weights is 1; Step 324: apply the attention weight for each time step to the value vector in the sequence of value vectors V, resulting in a context vector ; Step 325: multiply the attention weight sequence with the two-dimensional feature tensor to obtain a weighted fusion tensor Then, the weighted fusion tensor is split according to the original channel order to obtain four weighted feature maps , , and .
Citation Information
Patent Citations
A non-intrusive atrial fibrillation monitoring method
CN114469133B
Peripheral artery blood pressure waveform reconstruction system
CN113143230A
Blood pressure prediction method and device based on multi-wavelength photoelectric volume pulse wave fusion
CN118436326A
Efficient processing method for real-time acquisition and synchronization of multi-source heterogeneous data
CN120234530A
Information interaction control method and system of vital sign monitor
CN120565086A