Electrocardiogram analysis method and device based on deep learning model and medium
The deep learning model is used to preprocess and feature fusion of ECG signals, and combined with clinical information to correct abnormal detection results, the problems of misdiagnosis and poor robustness of traditional ECG analysis are solved, and efficient and accurate ECG automated analysis is achieved.
Patent Information
- Application Number
- CN202510398056.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Traditional electrocardiogram analysis relies on manual interpretation to be very different, time-consuming and easy to misdiagnose and misdiagnosis. The existing algorithms are poorly robust in low-quality signals, making it difficult to integrate clinical information, resulting in false positive or false negative results and misjudgment.
The deep learning model is used to pre-process the ECG signal, extract temporal and spatial features, combine cross-modal feature fusion, perform multi-anomaly detection and timing prediction tasks, and correct abnormality detection results through clinical information.
It improves the availability of low-quality electrocardiogram signals, enhances the ability to identify complex pathological patterns, reduces misjudgment, realizes end-to-end automated analysis, shortens analysis time and ensures consistency of results.
Smart Images

Figure CN120241091A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning, and specifically to an electrocardiogram analysis method, device, and medium based on a deep learning model. Background Art
[0002] In traditional solutions, the analysis of electrocardiograms mainly relies on doctors' visual interpretation of waveforms and empirical judgments. For example, abnormalities such as ST-segment changes and T-wave inversions may be judged as myocardial ischemia or electrolyte disorders. However, different doctors may have differences in the interpretation of the same electrocardiogram, which not only takes a long time but may also lead to misdiagnosis or missed diagnosis in complex or atypical cases.
[0003] Based on this, with the development of computer vision and machine learning technologies, they have been gradually introduced into electrocardiogram analysis, aiming to improve the automation level and diagnostic efficiency.
[0004] However, it still has the following problems: 1. Electrocardiogram signals are easily affected by noise such as electromyogram interference and baseline drift. Existing algorithms have poor robustness in low-quality signals (such as ECGs collected by wearable devices) and are prone to producing false positive or false negative results. 2. There is insufficient recognition of complex pathological patterns, and it is difficult to integrate clinical information such as the patient's medical history and physical signs, and misjudgment may still occur. Summary of the Invention
[0005] To solve the above problems, this application proposes an electrocardiogram analysis method based on a deep learning model, including: Obtain the electrocardiogram corresponding to the patient, and preprocess the electrocardiogram to suppress noise and obtain an electrocardiogram signal; Input the electrocardiogram signal into a pre-trained deep learning model, extract the time features and spatial features corresponding to the electrocardiogram signal through the deep learning model, and perform cross-modal feature fusion; According to the fused features, synchronously execute multiple anomaly detection tasks and synchronously execute a time series prediction task to obtain corresponding anomaly detection results and time series prediction results respectively; Obtain the clinical information corresponding to the patient, and for the anomaly detection result, correct the anomaly detection result according to the clinical information and the time series prediction result; Output the analysis result corresponding to the electrocardiogram according to the corrected anomaly detection result.
[0006] In one example, preprocessing the electrocardiogram to suppress noise and obtain an electrocardiogram signal specifically includes: Perform a fast Fourier transform on the electrocardiogram, extract the corresponding spectral features, and perform noise type detection based on the state of the corresponding frequency in the spectral features; and extract morphological features based on the electrocardiogram and perform noise type detection based on the morphological features; Based on the identified noise type, adopt the corresponding preprocessing method to suppress noise and obtain an electrocardiogram signal.
[0007] In one example, based on the identified noise type, adopting the corresponding preprocessing method to suppress noise and obtain an electrocardiogram signal specifically includes: For baseline drift, eliminate fluctuations by dynamically fitting the signal trend line; and / or, for power frequency interference, eliminate phase distortion through zero-phase filtering; and / or, for electromyogram interference, separate high-frequency noise through wavelet transform; and / or, for electrode contact interference, reconstruct the damaged signal through the spatio-temporal correlation between leads; Determine that the electrocardiogram signal after noise suppression meets the quality requirements according to the preset quality assessment index.
[0008] In one example, the model architecture of the deep learning model includes: an input layer, a spatio-temporal joint encoder, a feature pyramid, and an output layer; The input layer is used to input the electrocardiogram signal; The spatio-temporal joint encoder includes a time channel branch, a space channel branch, and a cross-modal attention fusion layer; Among them, the time channel branch includes a bidirectional LSTM structure to extract time features; The space channel branch includes a 3D-CNN structure to extract space features; The cross-modal attention fusion layer fuses the time features and the space features through a cross-modal attention mechanism to obtain fused features; The feature pyramid includes a one-dimensional convolutional layer, a batch normalization layer, and an activation function layer; The output layer includes multiple anomaly classification heads and a single time series prediction head; The multiple anomaly classification heads are independent branches and respectively output their corresponding anomaly detection results; The time series prediction head includes a causal convolutional layer and a self-attention layer and outputs a time series prediction result.
[0009] In one example, the training process of the deep learning model includes: training the spatio-temporal joint encoder, the anomaly classification heads, and the time series prediction head in stages; In the first stage, train the spatio-temporal joint encoder. During the training process, use a symmetric encoder-decoder structure and perform adversarial masking training by randomly masking segments of the electrocardiogram signal; In the second stage, the anomaly classification head is trained. During the training process, the parameters of the spatio-temporal joint encoder are frozen, and the anomaly judgment training is carried out by gradually unfreezing the feature pyramid and setting dynamic sample weights. In the third stage, the time series prediction head is trained. During the training process, the time series is decoupled for training, and uncertainty calibration is introduced for prediction ability training. In the fourth stage, the deep learning model is adversarially adjusted. During the training process, adversarial samples are generated, defense strategies are set, and stability dynamic evaluation is carried out. According to the evaluation results, the model parameters of the deep learning model are finely tuned. Among them, the second stage and the third stage perform two-stage cyclic training.
[0010] In one example, the clinical information corresponding to the patient is obtained. For the anomaly detection result, according to the clinical information and the time series prediction result, the anomaly detection result is corrected, specifically including: A correction rule library established in advance based on a clinical knowledge graph; the correction rule library includes multiple correction rules, and the correction rules include trigger conditions, correction actions, and corresponding weights. Obtain the clinical information corresponding to the patient, generate corresponding structured features according to the dimensional information in the clinical information, and generate a risk index corresponding to the patient according to the structured features. According to the structured features, the anomaly detection result, and the trigger conditions in the correction rules, determine the specified correction rules that are hit, and determine the specified dimensional information and the specified anomaly detection result corresponding to the specified correction rules. According to the historical clinical information, determine the likelihood ratio corresponding to the specified dimensional information, and use Bayes' formula to correct the initial anomaly probability of the specified anomaly detection result according to the likelihood ratio to obtain the corrected anomaly probability. According to the risk index, adjust the confidence threshold corresponding to the specified anomaly detection result, and correct the anomaly detection result according to the corrected anomaly probability and the adjusted confidence threshold.
[0011] In one example, the method further includes: Determine the predicted anomaly probability in the time series prediction result. If the predicted anomaly probability conforms to the anomaly detection result, output the corrected anomaly detection result. If the predicted anomaly probability does not conform to the anomaly detection result, determine the difference between the predicted anomaly probability and the corrected anomaly probability. If the difference is lower than a preset difference, the predicted anomaly probability and the corrected anomaly probability are weighted and summed to obtain a final anomaly probability, and the anomaly detection result is corrected according to the final anomaly probability and the adjusted confidence threshold; If the difference is higher than the preset difference, manual review and alarm are performed.
[0012] In one example, according to the corrected anomaly detection result, the analysis result corresponding to the electrocardiogram is output, specifically including: Determine other dimensional information in each dimensional information of the clinical information except the specified dimensional information; Encrypt the other dimensional information to obtain encrypted dimensional information; According to the encrypted dimensional information, the specified dimensional information, and the corrected anomaly detection result, output the analysis result corresponding to the electrocardiogram.
[0013] On the other hand, the present application also proposes an electrocardiogram analysis device based on a deep learning model, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: perform the electrocardiogram analysis method based on a deep learning model described in any of the above examples.
[0014] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set to: perform the electrocardiogram analysis method based on a deep learning model described in any of the above examples.
[0015] The electrocardiogram analysis method based on a deep learning model proposed by the present application can bring the following beneficial effects: 1. By the preprocessing module suppressing noises such as electromyogram interference and baseline drift, the signal quality can be improved, the usability of low-quality electrocardiogram signals (for example, data collected by wearable devices) can be enhanced, and the probabilities of false positives and false negatives caused by noises can be reduced.
[0016] 2. By the deep learning model simultaneously extracting the temporal dynamic features and spatial distribution features of the electrocardiogram signal, and combining the cross-modal fusion technology, the representation ability for complex pathological patterns such as ST-segment changes and T-wave inversions can be improved, and the pathological recognition ability can be enhanced.
[0017] 3. Execute multiple anomaly detection tasks and time series prediction tasks synchronously, utilize the correlation between tasks to constrain model learning, avoid local feature biases caused by a single task, improve the coverage of atypical cases, and collaboratively optimize the comprehensiveness of diagnosis.
[0018] 4. Dynamically calibrate the anomaly detection results of the deep learning model by integrating clinical prior knowledge such as the patient's medical history and physical signs, and combining the time series prediction results, so as to reduce misjudgments caused by the lack of context information.
[0019] 5. Form a complete closed loop from signal processing to feature extraction, multi-task analysis, and result correction, reduce the manual intervention link, significantly shorten the electrocardiogram analysis time and ensure result consistency, and achieve end-to-end automation to improve efficiency. Description of the Drawings
[0020] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic flowchart of the electrocardiogram analysis method based on a deep learning model in an embodiment of the present application; Figure 2 It is an electrocardiogram containing baseline drift in one case of an embodiment of the present application; Figure 3 It is an electrocardiogram containing power frequency interference in one case of an embodiment of the present application; Figure 4 It is an electrocardiogram containing several interferences in one case of an embodiment of the present application; Figure 5 It is a schematic architecture diagram of a deep learning model in one case of an embodiment of the present application; Figure 6 It is a schematic architecture diagram of a spatio-temporal joint encoder in one case of an embodiment of the present application; Figure 7 It is a schematic diagram of the data flow of the time series prediction head in one case of an embodiment of the present application; Figure 8 It is a schematic flowchart of anomaly detection result correction in one case of an embodiment of the present application; Figure 9 It is a schematic diagram of an electrocardiogram analysis device based on a deep learning model in an embodiment of the present application. Detailed Embodiments
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0022] The following will detail the technical solutions provided by each embodiment of this application in conjunction with the drawings.
[0023] As Figure 1 shown, the embodiment of this application provides an electrocardiogram analysis method based on a deep learning model, including: S101: Obtain the electrocardiogram corresponding to the patient, and preprocess the electrocardiogram to suppress noise and obtain an electrocardiogram signal.
[0024] Here, the user undergoing electrocardiogram detection is referred to as a patient. For some users, they do not have diseases themselves, but undergo electrocardiogram detection through regular physical examinations and other means. For the convenience of description, they are collectively referred to as patients here.
[0025] The method of obtaining the electrocardiogram can be to associate with the corresponding medical system and obtain the corresponding electrocardiogram through the interface of the medical system. During the process of obtaining the electrocardiogram, it is necessary to ensure the consent of the patient and protect the patient's privacy. For example, the electrocardiogram analysis is carried out in a trusted environment such as a trusted device and a trusted network, and only doctors or operators with permissions are allowed to observe.
[0026] Generally speaking, the directly obtained electrocardiogram may have noise due to reasons such as equipment and the patient's own state. Therefore, noise suppression is required before an electrocardiogram signal for analysis can be obtained.
[0027] Generally speaking, the reasons for generating noise may be various. Therefore, it is necessary to first analyze to determine the type of generated noise, and then adopt the corresponding method to eliminate the noise.
[0028] S102: Input the electrocardiogram signal into a pre-trained deep learning model, and extract the time features, spatial features corresponding to the electrocardiogram signal through the deep learning model, and perform cross-modal feature fusion.
[0029] The deep learning model performs machine learning based on a deep neural network (DNN). It includes two or more hidden layers, can learn complex features in the data, and make predictions and classifications through these features.
[0030] The deep learning model can extract features from electrocardiogram (ECG) signals. To extract the features of ECG signals more accurately, the time features and spatial features can be extracted synchronously in two dimensions: the time dimension and the spatial dimension, and then the features are fused to facilitate the subsequent analysis of ECG signals.
[0031] S103: According to the fused features, multiple anomaly detection tasks are synchronously executed, and a time series prediction task is synchronously executed to obtain corresponding anomaly detection results and time series prediction results respectively.
[0032] The anomaly detection task refers to detecting whether the ECG signal of the patient is abnormal in multiple directions. Multiple anomaly detection tasks can include: arrhythmia classification task (which includes five classification results: normal, atrial fibrillation, ventricular tachycardia, conduction block, others), ST-T segment abnormality detection task (which is a binary classification, including normal and abnormal), QRS complex morphology abnormality detection task (which is a binary classification, including normal and abnormal), heart rate variability analysis task, etc.
[0033] The time series prediction task refers to predicting the ECG signal within a certain period in the future according to the trend of the current ECG signal, so as to give an early warning and correct the anomaly detection results based on the prediction.
[0034] S104: Obtain the clinical information corresponding to the patient. For the anomaly detection results, according to the clinical information and the time series prediction results, the anomaly detection results are corrected.
[0035] If only the anomaly detection results of the deep learning model are used for the final judgment, the dimensions considered are too few to adapt to the different physical conditions of different patients. By combining the clinical information and the time series prediction results, personalized correction of the patient can be carried out, making the final anomaly detection results closer to the personalized state of the patient.
[0036] S105: According to the corrected anomaly detection results, output the analysis results corresponding to the electrocardiogram.
[0037] After obtaining the anomaly detection results, it can include: the anomaly probabilities corresponding to each anomaly detection task, judgment results (whether there is an anomaly), etc. After combining these contents, the analysis results corresponding to the electrocardiogram are obtained.
[0038] The analysis results of the electrocardiogram can be directly used for the judgment of the patient's condition, or can be used as auxiliary judgment results and sent to the attending doctor of the patient. The doctor makes a judgment and analysis based on the analysis results combined with artificial experience to obtain the final confirmed result of the patient's condition.
[0039] 1. The preprocessing module can suppress noise such as electromyographic interference and baseline drift, which can improve signal quality, enhance the availability of low-quality ECG signals (for example, data collected by wearable devices), and reduce the probability of false positives and false negatives caused by noise.
[0040] 2. Through the deep learning model, the temporal dynamic characteristics and spatial distribution characteristics of the ECG signal are extracted simultaneously, and combined with cross-modal fusion technology, the ability to characterize complex pathological patterns such as ST segment changes and T wave inversion is improved, and the pathology recognition ability is enhanced.
[0041] 3. Simultaneously execute multiple anomaly detection tasks and time series prediction tasks, use the correlation between tasks to constrain model learning, avoid local feature deviations caused by a single task, improve the coverage of atypical cases, and collaboratively optimize the comprehensiveness of diagnosis.
[0042] 4. By integrating clinical prior knowledge such as patient medical history and physical signs, and combining the time series prediction results, the abnormal detection results of the deep learning model are dynamically calibrated to reduce misjudgments caused by lack of contextual information.
[0043] 5. A complete closed loop is formed from signal processing to feature extraction, multi-task analysis, and result correction, which reduces manual intervention, significantly shortens ECG analysis time and ensures result consistency, achieving end-to-end automation and improving efficiency.
[0044] In one embodiment, when performing noise suppression and judging the noise type, the noise type can be judged from two dimensions: spectrum characteristics and morphological characteristics. The noise types include baseline drift, power frequency interference, myoelectric interference, electrode contact interference, and the like.
[0045] Perform fast Fourier transform on the electrocardiogram to extract the corresponding spectrum features, and perform noise type detection based on the state of the corresponding frequency in the spectrum features; perform morphological feature extraction based on the electrocardiogram, and perform noise type detection based on the morphological features.
[0046] For baseline drift, it is mainly manifested as excessive energy in the low frequency band (usually less than 2Hz), while for power frequency interference, it can be determined by detecting peaks near 50Hz and 60Hz, so the two can be judged by spectral characteristics.
[0047] To elaborate, for baseline drift, the sources of low-frequency components include respiratory movement, electrode movement, and skin electrolysis effect. Respiratory movement refers to the change in the impedance of the electrode-skin contact caused by the rise and fall of the chest (usually corresponding to 0.1 Hz ~ 0.5 Hz); electrode movement refers to the slow drift caused by the change of the patient's body position (usually less than 1 Hz); and skin electrolysis effect refers to the slow potential change caused by the drying of the conductive paste (usually corresponding to 0.05 Hz ~ 0.2 Hz).
[0048] As shown Figure 2 in the figure, an exemplary electrocardiogram with baseline drift is given, and the manifestations in its spectral characteristics may include: a wide energy band appears in the range of 0 Hz to 2 Hz, and energy aggregation appears in the low-frequency band; the main energy of the QRS wave (fully called the Q wave - R wave - S wave group) is concentrated in 5 - 15 Hz, without overlap with the baseline drift frequency, and separation from the QRS wave appears.
[0049] At this time, quantitative judgment can be carried out. When the proportion of the energy in the low-frequency band (0 Hz to 2 Hz) in the total energy reaches 30%, it can be determined that significant baseline drift exists.
[0050] Regarding power frequency interference, the operating frequency of the AC power supply (in different regions, its operating frequency may correspond to 50 Hz or 60 Hz) will generate an electromagnetic field, which is coupled to the electrocardiogram signal through the electrode wire.
[0051] As shown Figure 3 in the figure, an exemplary electrocardiogram with power frequency interference is given, and the manifestations in its spectral characteristics may include: on the spectrogram, a sharp peak centered on 50 Hz or 60 Hz appears, and a narrow-band sharp peak appears; there may be harmonic interference such as second harmonic corresponding to 100 Hz or 120 Hz, third harmonic corresponding to 150 Hz or 180 Hz, etc. When the power supply is unstable, the sharp peak will expand to the surrounding area, and bandwidth expansion appears. For example, 50 Hz expands to 49 Hz to 51 Hz around.
[0052] At this time, the identification of the sharp peak can be carried out to determine whether power frequency interference exists.
[0053] Regarding electromyogram interference, its local slope mutation point number can be calculated to identify the electromyogram noise density.
[0054] As shown Figure 4 in the figure, an exemplary electrocardiogram with electromyogram interference is given, and the manifestations in its morphological characteristics (also called time domain characteristics) may include: dense irregular sharp peaks are superimposed on the isoelectric line, and high-frequency burrs appear; different from the regularity of the QRS wave, the electromyogram noise appears randomly without periodicity.
[0055] At this time, local slope and density statistics can be carried out to judge whether electromyogram interference exists. For example, for the local slope, several sampling points are set, and then the difference values of each sampling point are calculated to obtain the corresponding slope. If the absolute value of the slope of consecutive multiple sampling points exceeds the threshold, it is considered that there is an electromyogram interference segment. For density statistics, the density of sampling points whose absolute value of the slope exceeds the threshold within a unit time can be statistically calculated. If the density is high (for example, more than 20%), it is determined as severe electromyogram noise.
[0056] For electrode contact interference, it mainly corresponds to the interference caused by long-term signal loss. Its manifestations in morphology can include: within a certain duration (for example, more than 1 second), the lead suddenly becomes a straight line, resulting in signal loss; the signal rapidly deviates by more than a certain value (for example, 2 mV) within a few seconds; the affected lead loses spatial correlation with other leads, resulting in inconsistency between leads.
[0057] At this time, for the inspection of lead consistency, the correlation coefficient of the QRS wave appearance time of each lead can be calculated. Usually, when it is higher than 0.9, it is normal, while the correlation coefficient of abnormal leads usually drops sharply below 0.3. For the detection of amplitude mutation, its change amplitude within a certain time (for example, 100 ms) can be detected. If the change amplitude exceeds 1 mV, it is determined as poor contact.
[0058] In this way, through corresponding methods, the detection of each noise type is completed. Then, according to the detected noise type, the corresponding preprocessing method is adopted for noise suppression.
[0059] Specifically, for baseline drift, the fluctuation is eliminated by dynamically fitting the signal trend line; for power frequency interference, the phase distortion is eliminated by zero-phase filtering; for electromyogram interference, the high-frequency noise is separated by wavelet transform; for electrode contact interference, the damaged signal is reconstructed through the spatio-temporal correlation between leads.
[0060] To expand, for baseline drift, noise suppression can be performed through a high-pass filter with a fixed cut-off frequency. However, this method is prone to distorting the ST segment morphology.
[0061] Based on this, an adaptive window segmentation method can be adopted, and the window length is dynamically adjusted according to the heart rate. Among them, the faster the heart rate, the shorter the window. Then, within each window, a polynomial fitting method (for example, using a polynomial of order 3 - 5. A 3rd-order polynomial can be used for gentle drift to avoid overfitting physiological fluctuations. A 5th-order polynomial can be used for severe mutations to better fit the rapidly changing non-linear trend. Other types use a 4th-order polynomial to balance flexibility and stability) is used to approximate the baseline trend. The optimal coefficients are found through the least squares method to minimize the sum of the squared residuals between the fitting curve and the original signal, thereby fitting and generating the corresponding fitting curve.
[0062] Subtract the fitting curve of this polynomial from the original electrocardiogram signal to achieve noise suppression. Of course, during this process, physiological fluctuation protection can also be carried out. Detect the fluctuations synchronized with the thoracic impedance signal and consider them as respiration-related fluctuations, and this physiological oscillation can be retained.
[0063] After experimental verification, this noise suppression method can make the maximum offset error of the ST segment less than 0.05 mV, meeting the clinical diagnosis requirements. And the processing time is less than 10 ms / lead, ensuring real-time performance.
[0064] For power frequency interference, a filter can be used for filtering. However, it is prone to introducing phase shift, resulting in deviation in the measurement of the QRS wave width. Moreover, it is difficult for a fixed center frequency to adapt to power grid fluctuations.
[0065] Based on this, zero-phase filtering is adopted. By means of bidirectional filtering (forward + reverse filtering), the filtering operation is made symmetric on the time axis, thereby canceling the phase shift. Among them, the forward filtering passes the original signal through the filter in sequence according to the time order. At this time, the signal is filtered, but a phase delay is generated, and the waveform shifts to the right. Reverse the time axis of the result after forward filtering and perform reverse filtering, passing through the filter again. At this time, a reverse phase delay is generated, which cancels the phase delay in the forward filtering, making the total phase delay 0. Multiply the amplitude-frequency responses of the two filterings, and the stopband attenuation can be doubled, and the transition band becomes steeper.
[0066] Its corresponding mathematical expression can be: y = fftshift( filtfilt(b, a, x) ), where y is the electrocardiogram signal after noise suppression, b is the filter numerator coefficient vector, a is the filter denominator coefficient vector, filtfilt is the forward filtering, and fftshift is the reverse filtering.
[0067] And adaptive tracking is adopted to monitor the energy peak value in the 50Hz - 60Hz frequency band in real time, dynamically adjust the notch center frequency, and perform automatic bandwidth contraction. For example, when the interference is weak, the bandwidth is 2Hz, and when the interference is strong, it expands to 5Hz.
[0068] After experimental verification, this noise suppression method can reduce the measurement error of the QRS wave width from ±8ms to ±2ms, and make the retention rate of the T wave amplitude exceed 98%, avoiding misdiagnosis of hypokalemia.
[0069] For electromyogram interference, since the db4 wavelet (the full name is Daubechies 4th order wavelet) has a short time domain range, is suitable for capturing the transient characteristics of the electrocardiogram signal, and the 2nd order vanishing moment can effectively represent the local polynomial trend of the signal, forming a contrast with the random spikes of the electromyogram noise, and is close to the steep rising edge shape of the QRS wave, which is conducive to distinguishing the true electrocardiogram characteristics from the noise, so the db4 wavelet is selected and hierarchical processing is carried out through the db4 wavelet. For example, assuming that the sampling rate of the electrocardiogram signal is 500Hz and its effective frequency range is 0Hz - 250Hz, it is divided into D1 - D5 and A5 layers. The corresponding frequency ranges and main components of each layer are shown in Table 1 below: Table 1 Wavelet decomposition level table
[0070] Based on this, for the high-frequency layer (D1~D2), which is the energy concentration area of EMG noise, strong suppression is required. For the middle-frequency layer (D3), considering both noise residue and physiological details, fine adjustment is needed. For the low-frequency layer (D4~A5), the main component of ECG should be retained to avoid over-smoothing.
[0071] Then, hierarchical noise suppression is performed, and different thresholds are applied to different levels.
[0072] The formula for the threshold function is: ; where is a non-linear filtering function used to screen the coefficients after wavelet decomposition. It determines which wavelet coefficients are regarded as noise and need to be suppressed, and which are regarded as real signals and need to be retained. c is the wavelet coefficient, representing the detail component of the signal in a certain frequency band, T is the threshold used to control the intensity of noise suppression. Generally, large coefficients correspond to real signals and small coefficients correspond to noise. Coefficients lower than T are set to zero or reduced, so as to suppress noise while retaining important signal features. sign(c) is used to retain the sign of the coefficient. is used to shrink the coefficients exceeding the threshold. Compared with the hard threshold that directly truncates the coefficients lower than the threshold, it retains mutations but may introduce artifacts. And compared with the soft threshold that linearly attenuates the coefficients higher than the threshold, it is smooth but may weaken the signal. By adopting this compromise scheme and adjusting the "hardness and softness" of the threshold shrinkage, a balance is achieved between retaining signal features and suppressing noise.
[0073] Among them, the adaptive calculation formula for the threshold T is: ; where is the threshold corresponding to the j-th layer, is the noise standard deviation of the wavelet coefficients in the j-th layer, which is used to reflect the noise intensity. N is the signal length (i.e., the number of sampling points), compensating for the influence of the data volume on the extreme value, and controlling the rate of change of the threshold with the increase of the data volume through lnN.
[0074] In this way, the greater the noise, the higher the threshold, realizing noise intensity compensation. And the longer the signal, the higher the probability of extreme noise values, so the threshold needs to be increased to avoid misjudgment, realizing data volume compensation.
[0075] Next, the estimation formula for the noise standard deviation is: ; where is the median absolute deviation of the detail coefficients in the j-th layer, and 0.6745 is the quantile scaling factor of the standard normal distribution, converting MAD into a standard deviation estimate.
[0076] Among them, since the layer corresponds to the highest frequency band, mainly containing noise and high-frequency harmonics of the QRS wave. And in most ECG signals, the noise in the high-frequency band dominates the real signal. Therefore The MAD of the layer can be approximated as the noise intensity. At this time, select the layer to calculate the noise standard deviation. The process mainly includes: calculating the absolute value of the layer coefficient , finding the median , calculating , and scaling to obtain the noise standard deviation .
[0077] Through experimental verification, this noise suppression method can increase the signal-to-noise ratio of EMG noise by about 15 dB and increase the P-wave detection rate from 70% to 92%.
[0078] For electrode contact interference, motion artifact processing can be used for morphological repair.
[0079] Its repair strategies mainly include spatial repair and temporal repair.
[0080] For spatial repair, the electrocardiogram (ECG) signal is essentially the projection of cardiac electrical activity at different positions on the body surface, and there is a spatial vector relationship between the 12 leads. When a certain lead is damaged due to motion artifact or electrode detachment, the spatial correlation of other leads can be used for repair.
[0081] Use historical data to establish the conversion relationship between leads. This conversion relationship is a multiple linear regression function. By solving the coefficients in it, the functional relationship between different leads is established. Among them, the coefficients can be updated regularly.
[0082] In this way, when an abnormal signal of a certain lead is detected, the corresponding lead signal can be obtained by solving through this functional relationship.
[0083] For temporal repair, utilize the temporal autocorrelation of the ECG signal to predict the current moment signal through historical data. When the actual signal deviates significantly from the predicted value, the predicted value is used to replace the abnormal segment.
[0084] A corresponding long short-term memory (LSTM) network can be pre-trained, which includes an input layer, a hidden layer, and an output layer. The input layer is used to input ECG image segments within a certain past time (for example, 200 ms). The hidden layer includes 2 layers of LSTM modules, and each layer includes 64 units. The output layer is used to output the voltage prediction value at the current moment.
[0085] When the actual signal deviates from the predicted value by a certain degree, the predicted value can be used to replace the actual signal.
[0086] Through experimental verification, this noise suppression method can reduce the misdiagnosis rate of false ventricular tachycardia caused by motion artifacts by 65% and keep the diagnostic integrity rate above 80% when the lead is detached.
[0087] Based on this, after noise suppression, it is also possible to determine whether the electrocardiogram after noise suppression meets the quality requirements according to the preset quality assessment indicators.
[0088] Among them, the quality assessment indicators can include the signal-to-noise ratio (SNR), waveform integrity, and the quality requirements can include: the SNR needs to be higher than 20 dB, the QRS wave loss rate in waveform integrity is lower than 0.1%, and the correlation between the ST segment distortion degree and the original signal is higher than 0.98.
[0089] In one embodiment, as Figure 5 shown, the model architecture of the deep learning model includes: an input layer, a spatio-temporal joint encoder, a feature pyramid, and an output layer.
[0090] Among them, the input layer is used to input the electrocardiogram signal. After noise suppression of the electrocardiogram, it is input into the input layer. Of course, the processing module corresponding to noise suppression can also be added to the deep learning model. The electrocardiogram without noise suppression is input into the input layer, and the electrocardiogram signal is obtained after noise suppression by this processing module, and then the electrocardiogram signal is input into the spatio-temporal joint encoder.
[0091] Generally speaking, the sampling rate of the electrocardiogram signal can be set to 100 Hz.
[0092] As Figure 6 shown, the spatio-temporal joint encoder includes a time channel branch, a space channel branch, and a cross-modal attention fusion layer.
[0093] The time channel branch includes a bidirectional LSTM structure to extract time features.
[0094] Specifically, the time channel branch adopts a BiLSTM (the full name is Bi-directional Long Short-Term Memory) channel, including 3 bidirectional LSTM layers and 1 time attention layer for outputting time features. The input dimension is set to [Batch × 3000 (time sampling points) × 12 (leads)], and each bidirectional LSTM layer is set with 256 hidden units to process the temporal evolution law of the entire electrocardiogram sequence.
[0095] After the first bidirectional LSTM layer and the second bidirectional LSTM layer, a regularization (Dropout) layer is set, and its Dropout rate can be set to 0.3 to prevent overfitting. The third bidirectional LSTM layer is connected to a multi-head time attention layer. For example, it is set to 8-head self-attention to automatically focus on key time nodes (such as the peak position of the R wave), so as to output time features.
[0096] The spatial channel branch includes a 3D-CNN structure to extract spatial features.
[0097] Specifically, the spatial channel branch includes a 3D convolutional layer, a pooling layer, and a depthwise separable convolutional layer for outputting spatial features. The input dimension is reshaped into a three-dimensional tensor of [Batch × 3000 (time sampling points) × 12 (leads) × 1 (channel)], and the output dimension is aligned with the time channel branch.
[0098] In the 3D convolutional layer, the convolutional kernel is set to 3×5×3, and a batch normalization layer (BN layer) and an activation function layer (ReLU layer) are set. The pooling layer uses a max pooling layer (MaxPool layer), and the size of the pooling unit is 1×2×1. In the depthwise separable convolutional layer (DepthwiseConv3D layer), the convolutional kernel is set to 3×3×3, and a batch normalization layer (BN layer), an activation function layer (ReLU layer), and a space-to-depth layer (SpaceToDepth layer) are set. The channel attention weighting layer performs channel attention calculation. The dimension alignment layer includes a 1×1 convolutional layer and an interpolation alignment layer. The number of channels is expanded to 512 through a 1×1 convolution, and the time dimension is interpolated to 3000 points to be aligned with the time channel branch.
[0099] The cross-modal attention fusion layer fuses the time features and spatial features through a cross-modal attention mechanism to obtain fused features.
[0100] Specifically, the cross-modal attention fusion layer realizes spatio-temporal feature interaction through cross attention. The time features output by the time channel branch are used as the query vector (Query, abbreviated as Q), and the spatial features output by the spatial channel branch are used as the key-value pair (including the key vector Key and the value vector Value, abbreviated as K and V respectively). After calculating the attention weights, the fused features are generated by weighting. Finally, the original time features are retained through a residual connection to form a joint feature vector with a dimension of 512.
[0101] The feature pyramid includes a one-dimensional convolutional layer, a batch normalization layer, and an activation function layer.
[0102] The dimension of the fused features input to the feature pyramid is [Batch × 3000 × 512]. Its structure includes layers P5 to P3. Each layer contains a one-dimensional convolutional layer (Conv1D), a batch normalization layer (BN layer), and an activation function layer (ReLU layer) for hierarchical downsampling. The 3000 dimension is sequentially reduced to 1500, 750, and 375, and the 512 number of channels is sequentially reduced to 256, 128, and 64. In each layer from P5 to P3, max pooling processing is performed using max pooling units of different sizes in parallel.
[0103] The generated hierarchical features include: [Batch×1500×256], [Batch×750×128], and [Batch×375×64]. Among them, 1500 points correspond to high-level features that capture the overall rhythm pattern of the electrocardiogram, 750 points correspond to middle-level features that identify local morphological changes such as the ST segment / T wave, and 375 points correspond to low-level features that analyze the fine structure of the QRS complex.
[0104] Among them, a one-dimensional dilated convolutional layer (DilatedConv1D layer) can also be added to the P3 layer, and its dilation rate is set to 4 to capture long-range dependencies.
[0105] Set the feature enhancement path. After the high-level features are upsampled by bilinear interpolation, they are concatenated with the next-level features, and the global semantics and local details are fused through a 3×1 convolution to achieve top-down propagation. The middle-level features of the original spatial channel branch are retained at each level, and after aligning the dimensions through a 1×1 convolution, they are added to the hierarchical features to achieve lateral connection.
[0106] The output layer includes multiple anomaly classification heads and a single time series prediction head, and all output heads are in a parallel relationship.
[0107] Specifically, the multiple anomaly classification heads are independent branches that respectively output their corresponding anomaly detection results. For example, there are four types of anomaly detection tasks, including: the arrhythmia classification task (which includes five classification results: normal, atrial fibrillation, ventricular tachycardia, conduction block, other), the ST-T segment anomaly detection task (which is a binary classification, including normal and abnormal), the QRS complex morphology anomaly detection task (which is a binary classification, including normal and abnormal), and the heart rate variability analysis task (which can be set as a regression task). Each anomaly classification head is used to perform the corresponding anomaly detection task.
[0108] For the arrhythmia classification task, global average pooling (GlobalAveragePool) is adopted, the activation function uses Softmax, and the loss function is set to categorical focal loss during training. Based on the features corresponding to the 1500 points of the highest level, the overall representation is obtained through global average pooling, and the 5-class probability distribution is output after passing through a 128-dimensional fully connected layer. Among them, global features such as P wave absence and irregular RR interval need to be concerned.
[0109] For the ST-T segment abnormality detection task, global average pooling (GlobalAveragePool) is adopted, the activation function is Sigmoid, the loss function is set to binary focal loss during training, the features corresponding to the 750 points in the middle layer are selected, and the ST segment region (80 ms - 120 ms after the J point) features are extracted by sliding the window in the time dimension. After weighting by spatial attention, it is judged whether the ST segment of each lead is elevated or depressed.
[0110] For the QRS complex morphology abnormality detection task, global average pooling (GlobalAveragePool) is adopted, the activation function is Sigmoid, the loss function is set to weighted binary cross-entropy loss (Weighted BCE Loss) during training. Utilizing the high time resolution of the features corresponding to the 375 points in the bottom layer, the starting and ending points of the QRS are located, and abnormalities such as bundle branch block are detected by combining morphological template matching.
[0111] For the heart rate variability analysis task, global average pooling (GlobalAveragePool) is adopted, the activation function is Linear, the loss function is set to Huber loss (Huber Loss) during training. Time domain (SDNN), frequency domain (LF, HF power) features are jointly extracted from the original RR interval sequence and the LSTM hidden state, and the HRV index is predicted through a regression network.
[0112] The time series prediction head includes a causal convolutional layer and a self-attention layer, and outputs the time series prediction result.
[0113] As Figure 7 shown, its input includes the bottom layer features of the feature pyramid (features of the P3 layer), the patient state memory (a dynamic memory unit maintained through the time channel branch, recording the feature evolution in the past period of time, which is a hidden state dimension of 128), the patient's metadata (encoding the patient's age, gender, medical history, etc. into a 32-dimensional vector), and the timestamp encoding (mapping the prediction time point to a 16-dimensional sine position encoding).
[0114] The above input features are concatenated and dimensionally increased, the merged dimension is Dense(256), and normalization is performed through LayerNorm.
[0115] Build a causal convolution with an increasing dilation rate of 4 layers, where the dilation values are 1, 2, 4, and 8 in sequence. In the structure of each layer, it includes a causal one-dimensional convolutional layer (CausalConv1D), a batch normalization layer (BatchNorm), an activation function layer (ReLU), and a regularization layer (Dropout). Among them, the convolutional kernel size of the causal one-dimensional convolutional layer is set to 5 (kernel_size = 5), and the interval between each point in the convolutional kernel is 1 unit (dilation_rate = 2). The Dropout rate of the regularization layer is set to 0.2.
[0116] Set up a self-attention mechanism. Through 4-head attention, process 256-dimensional time series features, and ensure that only historical information is focused on through an attention mask. The output is extended to 512 dimensions through a feed-forward network.
[0117] In the prediction output layer, multi-time point parallel prediction is performed. The predicted target time is divided into multiple windows (for example, one window every 30 minutes), and each window is processed independently.
[0118] In one embodiment, when training the deep learning model, it can be set to perform staged training on the spatio-temporal joint encoder, the anomaly classification head, and the time series prediction head.
[0119] Specifically, in the first stage, the spatio-temporal joint encoder is trained. During the training process, a symmetric encoder-decoder structure is used, and adversarial masking training is performed by randomly masking segments of the electrocardiogram signal.
[0120] The data source can be obtained from a clinical database. After preprocessing it, resample it to 300 Hz and perform lead-level normalization.
[0121] Set up a masking strategy. For example, each sample is masked with 3 - 5 segments, and the length of each segment is 200 ms - 500 ms. And enhance the masking of the key area. For example, the additional masking probability of the QRS complex region is increased by 30%.
[0122] At this time, only the spatio-temporal joint encoder is trained. Set the parameters of the optimizer as the learning rate lr = 5e-4, the decay rate betas = (0.9, 0.98), the decay weight weight_decay = 0.05, the batch size is 128, the learning rate strategy is to linearly warm up to 1e-3 in the first 5 rounds, and the learning rate is adjusted through cosine annealing decay (cycle = 50 rounds), and the lowest learning rate is 1e-6.
[0123] During training, the masking pattern is dynamically generated online for each batch. Adopt a weighted multi-feature mean absolute error loss function, set weights for the features corresponding to each lead, and calculate the loss by weighted summation of the mean absolute error of each feature.
[0124] In the second stage, the anomaly classification head is trained. During the training process, the parameters of the spatio-temporal joint encoder are frozen. The anomaly judgment training is carried out by gradually unfreezing the parameters of the feature pyramid and setting dynamic sample weights.
[0125] Set the ratio of the training set to the validation set to 8:1. In the training set, it is necessary to ensure that the number under each classification exceeds a certain value.
[0126] Set the augmentation strategy. Noise templates can be collected from real devices used to generate electrocardiograms for dynamic noise introduction. Lead replacement is performed by randomly swapping the aVR and aVL leads.
[0127] During training, unfreeze the parameters of the bottom layer of the feature pyramid. Set the parameters of the optimizer as learning rate 3e-4, decay rate betas=(0.95, 0.99), minimum value eps=1e-6. Adopt the Lookahead(k=5) optimization strategy, where the step parameter k=5 and the batch size is 64, ensuring that there are at least a certain number of anomaly types.
[0128] The regularization methods include label smoothing with smoothing=0.1, and also hierarchical Dropout with parameters set as 0.1 for the input layer and 0.3 for the intermediate layer.
[0129] Set dynamic sample weights to automatically allocate attention during training according to the historical performance of the task. After each round or every two rounds of training, determine the historical performance based on the validation set. Tasks with better historical performance will receive higher weights in the next round of training. At the same time, set the maximum and minimum values of the weights, for example, set them as 2.0 and 0.5 respectively.
[0130] Set the main task and auxiliary tasks during training. The main task can be the arrhythmia classification task. Set weights for each anomaly classification head, where the weights of the anomaly classification heads corresponding to the main task are higher. Set the weighted sum loss function corresponding to the output of each anomaly classification head, and set the early stopping conditions, which can include that the main task validation F1 parameter has not improved for multiple consecutive rounds, or the loss of the auxiliary task fluctuates above a certain value for multiple consecutive rounds.
[0131] In the third stage, the time series prediction head is trained. During the training process, decouple the training of the time series and introduce uncertainty calibration for prediction ability training.
[0132] The input sequence can be continuous electrocardiogram signals of a certain duration (for example, 6 hours), and set the step size of the sliding window (for example, 10 minutes). Perform positive sample augmentation. For samples with upcoming anomalies, extend the context forward for a certain duration (for example, 3 hours), and oversample rare events (such as precursors of ventricular fibrillation) by 500%.
[0133] Set the thawing strategy to gradually thaw the encoder during the training process. Increase the thawing ratio gradually in each training round. The initial thawing ratio can be set to 20%, and the thawing ratio increases by 1.6% for each additional round. The total number of training rounds can be set to 50, and the upper limit of the thawing ratio is set to 100%.
[0134] In each round, select the corresponding parameters for gradient calculation according to the thawing ratio.
[0135] Set the parameters of the optimizer to a learning rate of 1e-4 and a scheduling decay schedule_decay = 0.004.
[0136] Among them, the second stage and the third stage perform two-stage cyclic training. For example, in the morning every day, perform diagnostic mode training, freeze the time series prediction head, and only update the parameters of the anomaly classification head. In the afternoon, perform prediction mode training, freeze the anomaly classification head, and only update the parameters of the time series prediction head. In the fourth stage, perform adversarial adjustment on the deep learning model. During the training process, generate adversarial samples, set the defense strategy, and perform stability dynamic evaluation. According to the evaluation results, fine-tune the model parameters of the deep learning model.
[0137] The attack strategy is set to Projected Gradient Descent (PGD) attack. The parameters can be set as follows: the maximum norm constraint eps is 0.08, the step size alpha for each iteration is 0.02, the number of iteration steps steps is 7, and the boolean value targeted is set to False, indicating a non-targeted attack. Set the lead weights, for example, set to [1.0, 1.0, 1.0, 2.0, 2.0, 2.0, 1.5, 1.5, 1.5, 1.0, 1.0, 1.0]. The higher the weight, the more attention it receives, and the key attack is on leads V1 - V3 that are sensitive to ST segment changes.
[0138] In the defense training strategy, each batch contains 50% clean samples and 50% adversarial samples, and the labels of the adversarial samples remain the original values. Set the gradient penalty loss, and the hyperparameter λ is set to 10, which is calculated only on the adversarial samples.
[0139] Perform stability monitoring. Randomly mark 1 - 3 leads in the lead loss test to check the prediction fluctuations. The allowed fluctuation range is set to a decrease in AUROC below 0.05. In the temporal consistency check, for adjacent 10 - minute segments of the same patient, the difference in prediction results should be less than 15%.
[0140] Among the fine-tuning parameters, the learning rate is set to 1e-6 to stabilize the model at an extremely low learning rate, the number of training epochs is set to 15, and during regularization enhancement, the weight decay is set to 0.1, and the probability of all Dropout layers is increased by 0.1.
[0141] In one embodiment, as Figure 8 shown, when correcting the anomaly detection results, a correction rule library pre-established based on clinical information and a clinical knowledge graph can be determined; the correction rule library includes multiple correction rules, and the correction rules include trigger conditions, correction actions, and corresponding weights.
[0142] Among them, the clinical knowledge graph can be established through corresponding medical dictionaries, expert experience, medical journals, etc. This connection process can be to collect corresponding data from the above sources and then input it into a large model, and the clinical knowledge graph is output through the large model.
[0143] The clinical knowledge graph stores nodes and edges. The nodes represent different knowledge. For example, it includes clinical information and electrocardiogram information in various dimensions, and the edges represent the influence relationship between the nodes.
[0144] The correction rule library filters the nodes and edges related to the electrocardiogram signal in the clinical knowledge graph, and further combines expert experience and the large model to form corresponding correction rules. For example, if the following knowledge is obtained from the clinical knowledge graph: The prediction of atrial fibrillation in the elderly over 65 years old is at high risk; the use of digoxin may lead to false positives in ST segment abnormalities, etc. Then the trigger conditions can include: the age is greater than 65 years old, and the judgment probability of atrial fibrillation being positive in the current anomaly judgment result; the patient uses β-blockers, and the QRS complex is widened in the anomaly judgment result.
[0145] At the same time, corresponding correction actions are set. The correction actions can include: changing the judgment threshold of the anomaly, modifying the judgment probability of the anomaly, adding or reducing anomaly judgment conditions, etc. For example, for "The prediction of atrial fibrillation in the elderly over 65 years old is at high risk", the correction action can be: The positive judgment threshold of atrial fibrillation is lowered by 20%; for "The use of digoxin may lead to false positives in ST segment abnormalities", the correction action can be: The judgment of ventricular tachycardia requires an additional judgment condition of HR less than 50.
[0146] The corresponding weight can be determined based on the evidence level of the knowledge. The evidence level is mainly divided into five levels, from high to low: Level I, Level II, Level III, Level IV, and Level V. The higher the level, the higher the reliability of the knowledge. The way to judge reliability can be based on expert experience, large model judgment, judging the probability distribution in the collected historical records, etc., or a combination of multiple methods can be used for judgment. The higher the evidence level, the more reliable the correction rule. For example, based on comprehensive judgment, the corresponding weight for "the prediction of atrial fibrillation in the elderly over 65 years old is high risk" is set to 0.9, and the corresponding weight for "the use of digoxin may cause false positives of ST segment abnormalities" is set to 0.7.
[0147] Based on this, by assigning corresponding IDs to each rule, a correction rule composed of rule ID + trigger condition + correction action + corresponding weight is obtained. For example, based on the knowledge that "the prediction of atrial fibrillation in the elderly over 65 years old is high risk", a correction rule with rule ID R001 is generated. In this correction rule, the trigger condition is: the age is greater than 65 years old, and the judgment probability of atrial fibrillation in the current abnormal judgment result increases; the correction action is to lower the atrial fibrillation judgment threshold by 20%, and the corresponding weight is 0.9.
[0148] In addition, for the clinical knowledge graph and the correction rule library, it can be updated regularly to adapt to the latest research developments in medicine.
[0149] Obtain the corresponding clinical information of the patient, generate corresponding structured features according to the dimensional information in the clinical information, and generate the corresponding risk index of the patient according to the structured features.
[0150] The clinical information includes multiple dimensional information, such as age, gender, medical history, medications taken, etc. Here, each dimensional information is quantitatively encoded, thereby converted into the corresponding label vector to generate the corresponding structured features for subsequent analysis.
[0151] Among them, the types of quantitative encoding can include continuous data binning, categorical data encoding, and feature engineering quantization. Continuous data binning corresponds to the dimensions that can be directly quantified. For example, for the age dimension, it is segmented. 20 to 30 years old is young, 31 to 50 years old is middle-aged, and over 51 years old is elderly. Categorical data encoding corresponds to the dimensions that are difficult to directly quantify. For example, for the gender dimension, male is classified and encoded as 1, and female is classified and encoded as 0. Feature engineering quantization is to convert abstract features into computable indicators. For example, for medical history, medications taken, etc., corresponding levels are set according to the impact of each medical history and medications taken on electrocardiogram, and then corresponding quantization is carried out.
[0152] The risk index can be set through preset rules. Several conditions are preset in advance, and weights are set for each condition. Weighted summation is performed based on the conditions met by the patient. For example, weighted summation is performed based on age ( +1 point for every 10 years), diabetes ( +2 points), coronary heart disease ( +1.5 points), etc.
[0153] Based on the structured features, anomaly detection results, and trigger conditions in the correction rules, determine the specified correction rules that are hit, and determine the specified dimension information and specified anomaly detection results corresponding to the specified correction rules. If a certain dimension information hits the trigger condition of a certain correction rule, for the convenience of description, this dimension information and correction rule are called the specified dimension information and specified correction rule, respectively, and the corresponding anomaly detection result in the specified correction rule is called the specified anomaly detection result.
[0154] Based on the historical clinical information, determine the likelihood ratio corresponding to the specified dimension information, and through Bayes' formula, correct the initial anomaly probability of the specified anomaly detection result according to the likelihood ratio to obtain the corrected anomaly probability.
[0155] The likelihood ratio (LR) is a composite index that reflects both sensitivity and specificity, and is used to evaluate the authenticity of a diagnostic test. In this application, it is mainly used to verify the authenticity of the trigger condition. For example, for elderly people over 65 years old, according to the historical clinical information of the hospital, count the ratio of their occurrence frequencies in the abnormal group and the normal group. Assume that the probability of elderly patients in the atrial fibrillation positive group is 2.3 times that of the negative group, then the likelihood ratio LR = 2.3.
[0156] Bayes' formula can include: , where is the corrected anomaly probability, is the initial anomaly probability, is the total likelihood ratio, which is obtained by multiplying the likelihood ratios of all specified correction conditions that are hit.
[0157] In this way, the anomaly probability output by the deep learning model can be personalized corrected through the clinical information of this patient. For example, assume that the patient simultaneously hits the trigger conditions of the following two correction rules: the age is greater than 65 years old, and the judgment probability of atrial fibrillation in the current anomaly judgment result; the patient uses β-blockers, and the QRS complex is widened in the anomaly judgment result. And it is calculated that the corresponding likelihood ratios LR are 2.3 and 0.8 respectively, then the total likelihood ratio = 2.3 * 0.8 = 1.84. Then assume that the initial anomaly probability of atrial fibrillation positive output by the deep learning model for this patient is 72%, then the corrected anomaly probability = (0.72 * 1.84) / (1 + 0.72 * (1.84 - 1)) = 81.3%.
[0158] Meanwhile, according to the risk index, the confidence threshold corresponding to the specified anomaly detection result is adjusted, and the anomaly detection result is corrected according to the corrected anomaly probability and the adjusted confidence threshold. Among them, the confidence threshold means that when the anomaly probability judged by the deep learning model is higher than this confidence threshold, it is considered that there is an anomaly.
[0159] For high-risk patients, their characteristics usually include multiple underlying diseases, frequent use of corresponding drugs, poor physiological reserve, etc., and there is a high probability of false positives. Therefore, a higher confidence threshold can be set for them to improve specificity.
[0160] Set the confidence threshold adjustment formula as: ; where is the adjusted confidence threshold, is the confidence threshold before adjustment, is the risk index. At this time, assume that due to old age and hypertension, the calculated risk index of this patient is 6.8, and assume that the original confidence threshold is 0.5. Then the adjusted confidence threshold = 0.5 * (1 + 0.05 * 6.8) = 0.67. Since the corrected anomaly probability calculated above is 81.3%, which is higher than the adjusted confidence threshold of 0.67, it can be judged that it is an anomaly, atrial fibrillation positive.
[0161] Furthermore, after obtaining the anomaly detection result, the predicted anomaly probability in the time series prediction result can also be determined. This predicted anomaly probability is usually the predicted anomaly probability within a certain future time, and the highest point or average value within a period of time can be selected as the predicted anomaly probability.
[0162] If the predicted anomaly probability conforms to the anomaly detection result, the corrected anomaly detection result is output. For example, assume that the predicted anomaly probability is 80%, which is higher than the adjusted confidence threshold of 0.67 and conforms to the judgment of atrial fibrillation positive with the corrected anomaly probability, then keep this judgment unchanged.
[0163] If the predicted anomaly probability does not conform to the anomaly detection result, the difference between the predicted anomaly probability and the corrected anomaly probability is determined. For example, assume that the predicted anomaly probability is 65%, which is lower than the adjusted confidence threshold of 0.67 and does not conform to the judgment of atrial fibrillation positive. At this time, calculate the difference between the predicted anomaly probability and the corrected anomaly probability as 16.3%.
[0164] If the difference is lower than the preset difference, the predicted anomaly probability and the corrected anomaly probability are weighted and summed to obtain the final anomaly probability, and the anomaly detection result is corrected according to the final anomaly probability and the adjusted confidence threshold. For example, if the preset difference is set to 30% and the difference is 16.3%, which is lower than the preset difference, and the weight of the corrected anomaly probability is 0.7 and the weight of the predicted anomaly probability is 0.3, the final anomaly probability = 0.7 * 81.3% + 0.3 * 65% = 76.41%. This final anomaly probability is still higher than the adjusted confidence threshold of 0.67, so the final anomaly detection result is still atrial fibrillation positive.
[0165] If the difference is higher than the preset difference, it is considered that the difference is large, and manual review and alarm are performed to force manual review and additional inspections to prevent missed diagnosis and misdiagnosis.
[0166] In one embodiment, when outputting the analysis result corresponding to the electrocardiogram of the patient, to protect the privacy of the user, it is possible to determine, among the various dimensional information of the clinical information, the other dimensional information except the specified dimensional information, and encrypt the other dimensional information to obtain the encrypted dimensional information.
[0167] According to the encrypted dimensional information, the specified dimensional information, and the corrected anomaly detection result, the analysis result corresponding to the electrocardiogram is output.
[0168] Among them, the encryption method can adopt a symmetric encryption algorithm. For example, the Data Encryption Standard (DES) algorithm can encrypt and decrypt through a secret key.
[0169] The reason for not encrypting the specified dimensional information is that if manual review by a doctor is still required, quick review can be performed through the specified dimensional information.
[0170] Of course, selective encryption can be performed according to the output target during output. For example, if the output target is the attending doctor of the patient, encryption may not be performed; if it is someone else, encryption is required to protect the privacy of the patient. If someone else has the corresponding permission and has the secret key, they can also decrypt the encrypted dimensional information to obtain the corresponding information.
[0171] As Figure 9 shown, the embodiment of the present application also provides an electrocardiogram analysis device based on a deep learning model, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: the electrocardiogram analysis method based on a deep learning model described in any of the foregoing embodiments.
[0172] An embodiment of the present application further provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are configured to: the electrocardiogram analysis method based on a deep learning model described in any of the foregoing embodiments.
[0173] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0174] The device and medium provided by the embodiments of the present application correspond one by one to the method. Therefore, the device and medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.
[0175] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An electrocardiogram analysis method based on a deep learning model, characterized in that, Including: Obtain the electrocardiogram (ECG) corresponding to the patient, and preprocess the ECG to suppress noise and obtain an electrocardiogram signal; Input the electrocardiogram signal into a pre-trained deep learning model, and extract the time features and spatial features corresponding to the electrocardiogram signal through the deep learning model, and perform cross-modal feature fusion; According to the fused features, synchronously execute multiple anomaly detection tasks and synchronously execute a time series prediction task to obtain corresponding anomaly detection results and time series prediction results respectively; Obtain the clinical information corresponding to the patient, and correct the anomaly detection results according to the clinical information and the time series prediction results for the anomaly detection results; Output the analysis result corresponding to the electrocardiogram according to the corrected anomaly detection result.
2. The electrocardiogram analysis method based on a deep learning model according to claim 1, wherein Preprocess the electrocardiogram to suppress noise and obtain an electrocardiogram signal, specifically including: Perform a fast Fourier transform on the electrocardiogram, extract the corresponding spectral features, and detect the noise type according to the state of the corresponding frequency in the spectral features; and extract morphological features according to the electrocardiogram, and detect the noise type according to the morphological features; Based on the identified noise type, adopt the corresponding preprocessing method to suppress noise and obtain an electrocardiogram signal.
3. The electrocardiogram analysis method based on a deep learning model according to claim 2, wherein Based on the identified noise type, adopt the corresponding preprocessing method to suppress noise and obtain an electrocardiogram signal, specifically including: For baseline drift, eliminate fluctuations by dynamically fitting the signal trend line; and / or, for power frequency interference, eliminate phase distortion through zero-phase filtering; and / or, for electromyogram interference, separate high-frequency noise through wavelet transform; and / or, for electrode contact interference, reconstruct the damaged signal through the spatio-temporal correlation between leads; Determine that the electrocardiogram signal after noise suppression meets the quality requirements according to the preset quality evaluation index.
4. The electrocardiogram analysis method based on a deep learning model according to claim 1, characterized in that, The model architecture of the deep learning model includes: an input layer, a spatio-temporal joint encoder, a feature pyramid, and an output layer; The input layer is used to input the electrocardiogram signal; The spatio-temporal joint encoder includes a time channel branch, a space channel branch, and a cross-modal attention fusion layer; Among them, the time channel branch includes a bidirectional LSTM structure to extract time features; The space channel branch includes a 3D-CNN structure to extract spatial features; The cross-modal attention fusion layer fuses the time features and the spatial features through a cross-modal attention mechanism to obtain fused features; The feature pyramid includes a one-dimensional convolutional layer, a batch normalization layer, and an activation function layer; The output layer includes multiple anomaly classification heads and a single time series prediction head; The multiple anomaly classification heads are independent branches and respectively output their corresponding anomaly detection results; The time series prediction head includes a causal convolutional layer and a self-attention layer, and outputs a time series prediction result.
5. The electrocardiogram analysis method based on a deep learning model according to claim 4, characterized in that, The training process of the deep learning model includes: training the spatio-temporal joint encoder, the anomaly classification head, and the time series prediction head in stages; In the first stage, train the spatio-temporal joint encoder. During the training process, use a symmetric encoder-decoder structure and perform adversarial masking training by randomly masking segments of the electrocardiogram signal; In the second stage, the anomaly classification head is trained. During the training process, the parameters of the spatio-temporal joint encoder are frozen, and the feature pyramid is gradually unfrozen, and dynamic sample weights are set for anomaly judgment training; In the third stage, the time series prediction head is trained. During the training process, the time series is decoupled and trained, and uncertainty calibration is introduced for prediction ability training; In the fourth stage, the deep learning model is adversarially adjusted. During the training process, adversarial samples are generated, defense strategies are set, and stability dynamic evaluation is performed. According to the evaluation results, the model parameters of the deep learning model are fine-tuned; Among them, the second stage and the third stage perform two-stage cyclic training.
6. The electrocardiogram analysis method based on a deep learning model according to claim 1, wherein Obtain the clinical information corresponding to the patient. For the anomaly detection result, according to the clinical information and the time series prediction result, the anomaly detection result is corrected, specifically including: A correction rule library established in advance based on a clinical knowledge graph; multiple correction rules are included in the correction rule library, and each correction rule includes a trigger condition, a correction action, and a corresponding weight; Obtain the clinical information corresponding to the patient, generate corresponding structured features according to each dimension information in the clinical information, and generate a risk index corresponding to the patient according to the structured features; According to the structured features, the anomaly detection result, and the trigger condition in the correction rule, determine the specified correction rule that is hit, and determine the specified dimension information and the specified anomaly detection result corresponding to the specified correction rule; According to the historical clinical information, determine the likelihood ratio corresponding to the specified dimension information, and use Bayes' formula to correct the initial anomaly probability of the specified anomaly detection result according to the likelihood ratio to obtain a corrected anomaly probability; According to the risk index, adjust the confidence threshold corresponding to the specified anomaly detection result, and correct the anomaly detection result according to the corrected anomaly probability and the adjusted confidence threshold.
7. The electrocardiogram analysis method based on a deep learning model according to claim 6, wherein The method further includes: Determine the predicted anomaly probability in the time series prediction result; If the predicted anomaly probability conforms to the anomaly detection result, output the corrected anomaly detection result; If the predicted anomaly probability does not conform to the anomaly detection result, determine the difference between the predicted anomaly probability and the corrected anomaly probability; If the difference is lower than the preset difference, perform weighted summation of the predicted anomaly probability and the corrected anomaly probability to obtain a final anomaly probability, and correct the anomaly detection result according to the final anomaly probability and the adjusted confidence threshold; If the difference is higher than the preset difference, perform manual review and alarm.
8. The electrocardiogram analysis method based on a deep learning model according to claim 6, characterized in that According to the corrected anomaly detection result, output the analysis result corresponding to the electrocardiogram, specifically including: Determine the other dimension information in each dimension information of the clinical information except the specified dimension information; Encrypt the other dimension information to obtain encrypted dimension information; According to the encrypted dimension information, the specified dimension information, and the corrected anomaly detection result, output the analysis result corresponding to the electrocardiogram.
9. An electrocardiogram analysis device based on a deep learning model, characterized in that, Includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: the electrocardiogram analysis method based on a deep learning model according to any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: the electrocardiogram analysis method based on a deep learning model according to any one of claims 1 to 8.
Citation Information
Patent Citations
Electrocardiosignal analysis method based on deep learning
CN109864714A
Arrhythmia classification algorithm of C-LSTM for physiological parameter monitoring
CN113397555A
Remote sensing few-sample target detection method based on condition prompt and causal learning
CN118334519A
Multi-modal fusion algorithm for electrocardiosignal anomaly detection
CN118520279A
Cell classification using central enhancement of feature map
CN119156650A
Cited By
Exercise load electrocardio test system for multi-mode signal processing and intelligent analysis
CN121400843A
Defibrillator human body shaking detection method and system fusing electrocardio and transthoracic impedance
CN121401608A
Method and system for detecting human body shaking of defibrillator fusing electrocardio and thoracic impedance
CN121401608B