Dynamic electrocardiogram heart beat classification method, device and equipment and storage medium

Through the multi-band coupling matrix and dynamic attention mechanism, the problems of low recognition of small-amplitude characteristic waves and multi-perspective information heterogeneity in the existing electrocardiogram heartbeat classification are solved, achieving high-precision heartbeat classification and improved model robustness.

CN120744581AActive Publication Date: 2025-10-03GUANGDONG UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510877536.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-03
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In existing electrocardiogram (ECG) beat classification methods, the second-order coupling features of small-amplitude characteristic waves have low recognition accuracy, lack an adaptive mechanism to deal with the heterogeneity of temporal information and non-stationary signal-to-noise ratio of multi-perspective ECG features, and are unable to adaptively adjust the optimal recognition range of different beats.

Method used

A multi-band coupling matrix is ​​used to enhance the second-order coupling feature recognition of the characteristic wave. A multi-perspective feature fusion module based on channel-time decoupling cross-attention and a dynamic attention module based on ECG heartbeat multi-neighborhood level-by-level information fusion are constructed. The multi-perspective cross-attention mechanism and the dynamic attention mechanism are used to adaptively extract and fuse multi-perspective heartbeat information.

Benefits of technology

It significantly improves the accuracy and generalization performance of heartbeat classification, improves the recognition of small-amplitude characteristic waves, enhances the robustness and generalization ability of the model, and reduces the risk of overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744581A_ABST
    Figure CN120744581A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic electrocardiogram heart beat classification method and device, equipment and a storage medium. Relates to the technical field of electrocardiogram cardiac beat classification. The method comprises the following steps: preprocessing a data set to obtain multi-view data of ECG heart beats, and performing feature extraction to obtain multi-view features; the multi-view feature embedding of a plurality of heart beats is grouped according to the neighborhood size increasing from 0 to 1, the features of each view in each group are unified into the same length on a time axis and are connected on a channel axis to form a multi-range group, and a multi-view cross attention mechanism is used for each range group to obtain a multi-view fusion feature; discriminative information is extracted from multiple groups of features in different ranges, and attention enhancement features are obtained; and inputting the attention enhancement feature and the R-R interval into a classification head together to obtain a classification prediction result. The inherent diversity of the ECG signals is effectively handled, and the generalization performance of the classification model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electrocardiogram (ECG) beat classification, and in particular to a method, apparatus, device, and storage medium for dynamic electrocardiogram (DECT) beat classification. Background Art

[0002] The electrocardiogram (ECG) is a non-invasive technique for assessing cardiac health and is crucial for detecting arrhythmias in clinical practice. However, manual ECG diagnosis by experienced clinicians is time-consuming and inefficient, especially during long-term monitoring. Consequently, extensive research is devoted to developing computer-assisted diagnosis and automated classification methods to improve the efficiency of ECG analysis.

[0003] Existing ECG heartbeat classification methods can be roughly divided into two categories based on their input heartbeat number. The first category focuses on extracting discriminative features from single heartbeat waveforms and multi-beat RR interval sequences. For example: The paper "J.Niu, Y.Tang, Z.Sun, and W.Zhang, "Inter-patient ECG classification with symbolic representations and multi-perspective convolutional neural networks," IEEE J.Biomed.Health Inform., vol.24, no.5, pp.1321-1332, May 2020." improves the robustness of heartbeat morphology representation by converting heartbeat morphology into symbolic space through a discretization strategy, and maps the RR interval information into the heartbeat embedding matrix, thereby enhancing the generalization ability of the heartbeat classification model. The paper "X. Li, F. Zhang, Z. Sun, D. Li, X. Kong, and Y. Zhang, "Automatic heartbeat classification using S-shaped reconstruction and a squeeze and excitation residual network," Comput. Biol. Med., vol. 140, 2022, Art. no. 105108." segments the heartbeat signal and stacks it into a two-dimensional matrix. It then uses two-dimensional convolution to expand the receptive field, enhancing the model's ability to perceive contextual information in the ECG waveform. However, the discriminative information contained in a single heartbeat waveform and RR interval is limited. In clinical practice, many abnormal heartbeats, such as supraventricular ectopy and ventricular fusion, require the integration of contextual information from multiple heartbeat waveforms for accurate identification. Therefore, the second type of method classifies the target heartbeat based on ECG data from multiple adjacent heartbeats. For example, the paper "X. Zhai and C. Tin, "Automated ECG classification using dual heartbeat coupling based on convolutional neural network," IEEE Access, vol. 6, pp. 27465-27472, 2018." calculates the outer product matrix of adjacent heartbeats, highlights the second-order coupling relationship between characteristic waves, and uses CNN to extract the contextual information of the outer product matrix, significantly improving the accuracy of heartbeat classification.The paper "F.Li, Y.Xu, Z.Chen, and Z.Liu, "Automated heartbeat classification using 3-D inputs based on convolutional neural network with multi-fields ofview," IEEE Access, vol. 7, pp. 76295-76304, 2019." uses a two-dimensional convolutional network to extract multi-view information from a matrix consisting of ECG heartbeat morphology, RR intervals, and correlation coefficients of adjacent heartbeats, significantly improving the recognition accuracy of supraventricular ectopic beats and ventricular ectopic beats. The paper "S. Mousavi and F. Afghah, "Inter- and intra-patient ECG heartbeat classification for arrhythmia detection: A sequence-to-sequence deep learning approach," in Proc. IEEE Int. Conf. Acoust., Speech Signal Process., 2019, pp. 1308-1312," constructs a sequence-to-sequence prediction model based on a bidirectional recurrent neural network (BiRNN) to extract potential contextual information from multi-beat ECGs, effectively improving the recognition accuracy of supraventricular ectopic beats and ventricular fusion beats. However, for the task of modeling long multi-beat long-time series, BiRNN has difficulty learning long-range dependencies due to vanishing gradients. Therefore, the paper "Y.Xia, Y.Xiong, and K.Wang, "A transformer model blended with CNN and denoising autoencoder for inter-patient ECG arrhythmia classification," Biomed.Signal Process.Control, vol.86, 2023, Art.no.105271." uses a CNN-Transformer hybrid architecture to represent the local-global correlation of multi-heart beat sequences step by step, significantly improving the accuracy of heart beat classification.On this basis, the paper "H.Peng, X.Chang, Z.Yao, D.Shi, and Y.Chen, "A deep learning framework for ECG denoising and classification," Biomed.Signal Process.Control, vol.94, Aug.2024, Art.no.106441." proposed a wavelet CNN-Transformer encoder, which decouples the potential temporal features of ECG through a priori orthogonal transformation, effectively reducing the deep learning model's dependence on training samples.

[0004] Currently, the shortcomings of the existing technology are:

[0005] First, existing methods for representing heartbeats using second-order coupling matrices suffer from low coupling values ​​between small-amplitude feature waves, resulting in low recognition of their corresponding second-order coupling features. Second, multi-view ECG features contain heterogeneous temporal information and a non-stationary signal-to-noise ratio. Existing technologies lack adaptive fusion mechanisms to address these differences. Third, the optimal recognition ranges for different heartbeats vary. Existing technologies lack dynamic attention mechanisms, making it impossible to adaptively adjust the scope of contextual perception. Summary of the Invention

[0006] The present application provides a dynamic electrocardiogram (DECT) beat classification method, apparatus, device and storage medium to overcome three key limitations of existing methods: (1) the low recognition of second-order coupling features between small-amplitude feature waves in the ECG characterization method based on the outer product matrix of adjacent beat sequences; (2) the lack of an effective mechanism to deal with the heterogeneity of temporal information and the non-stationarity of signal-to-noise ratio contained in multi-view ECG features; and (3) the lack of a dynamic attention mechanism to adaptively adjust the optimal recognition range of different beats to achieve high-precision beat classification.

[0007] In a first aspect, the present application provides a method for classifying heart beats of a dynamic electrocardiogram, comprising:

[0008] Preprocessing the data set to obtain multi-view data of ECG heartbeats; wherein the multi-view data of ECG heartbeats includes heartbeat waveforms, low-frequency coupling matrices, high-frequency coupling matrices, and RR intervals;

[0009] Extracting features from the multi-view data of the ECG heartbeat to obtain multi-view features;

[0010] The multi-view feature embeddings of multiple heartbeats are grouped according to the neighborhood size increasing from 0 to 1. The features of each view in each group are unified to the same length on the time axis and connected on the channel axis to form multi-range groups. The multi-view cross-attention mechanism is used for each range group to obtain multi-view fusion features.

[0011] Based on the multi-view fusion features, discriminative information is extracted from multiple groups of features in different ranges to obtain attention enhancement features;

[0012] The attention enhancement feature is input into the classification head together with the RR interval to obtain the classification prediction result.

[0013] In one possible design, the data set is preprocessed to obtain multi-view ECG heartbeat data in the following ways:

[0014] Acquiring electrocardiogram data, and performing baseline drift removal and signal normalization on the electrocardiogram data to obtain a preprocessed ECG signal;

[0015] Based on the preprocessed ECG signal, for each heartbeat, with the R peak position as a reference point, the ECG signal is segmented into a plurality of time sample segments, each segment containing data t1 milliseconds before and t2 milliseconds after the R peak as a heartbeat segment;

[0016] Two band-pass filters are applied to the preprocessed ECG signal to obtain a low-frequency signal below the first frequency and a high-frequency signal from the first frequency to the second frequency. The outer product operation is performed on the two adjacent segments before and after each heartbeat to obtain the low-frequency coupling matrix and the high-frequency coupling matrix.

[0017] The RR interval is calculated based on the R peak position recorded in the ECG dataset.

[0018] In one possible design, feature extraction is performed on the multi-view data of the ECG heartbeat to obtain multi-view features, including:

[0019] Establishing a multi-view feature extractor; wherein the multi-view feature extractor includes feature extractors of different viewpoints;

[0020] The multi-perspective data generated by the target heartbeat and its surrounding heartbeats within a maximum neighborhood size l are input into feature extractors of different perspectives to obtain features of each perspective and obtain the multi-perspective features; wherein the feature extractors of different perspectives include a first network and two second networks, the first network extracts the morphological features of the heartbeat from the heartbeat segment as features of one perspective, and the two second networks extract the second-order features of the heartbeat from the low-frequency coupling matrix and the high-frequency coupling matrix as features of one perspective respectively.

[0021] In one possible design, the first network includes three identical first convolutional blocks, each of which includes a convolutional layer, a batch normalization layer, an activation layer, a maximum pooling layer, and a squeeze activation module connected in sequence;

[0022] The second network includes three second convolution blocks, each of which includes a second convolution layer, a second batch normalization layer, and a second maximum pooling layer connected in sequence, and the three second convolution blocks have different kernel sizes.

[0023] In one possible design, a multi-view cross-attention mechanism is used for each range group to obtain multi-view fusion features, including:

[0024] The multi-view features in each range group are spliced ​​together, and then a linear layer is used to generate the fusion feature Z f , the calculation process is:

[0025] Z f =Linear(Cat(G)) (1)

[0026] Where, G=[Z mor ,Z low ,Z high ], represents the multi-view features within each range group, Z mor ,Z low ,Z high They represent heartbeat waveform features, low-frequency coupling matrix features, and high-frequency coupling matrix features, respectively. Cat represents the concatenation operation, and Linear represents the fully connected layer.

[0027] Based on the fusion feature Z f , the first channel weight and the second channel weight are generated by the following formula:

[0028] U f =MLP(AvgPool(Z f ))+MLP(MaxPool(Z f )) (2)

[0029] U k =MLP(AvgPool(Z k ))+MLP(MaxPool(Z k ))

[0030] Where U f represents the first channel weight, U k Represents the second channel weight, Z k Represents the input features, k∈{mor,high,low}, represents three perspectives, mor, high, low represent the heartbeat waveform, low-frequency coupling matrix, and high-frequency coupling matrix respectively, MLP represents multi-layer perceptron, AvgPool represents global average pooling, and MaxPool represents global maximum pooling;

[0031] Based on the first channel weight and the second channel weight, the first feature and the second feature are calculated using the following formula:

[0032]

[0033] Where, P f represents the first feature, Softmax represents the normalized exponential function, T represents the matrix transpose, P k Indicates the second feature;

[0034] Based on the first and second features, the third channel weight and the fourth channel weight are calculated using the following formula:

[0035]

[0036] Where W f and W k Represent the third channel weight and the fourth channel weight respectively;

[0037] Based on the third channel weight and the fourth channel weight, the feature embedding of each view is calculated by the following formula:

[0038]

[0039] Where, Represents the feature embedding of view k;

[0040] Based on the embedding of each view feature, the multi-view fusion feature is calculated by the following formula:

[0041]

[0042] Where, represents the multi-view fusion feature, They represent the fused heart waveform features, the fused low-frequency coupling matrix features, and the fused high-frequency coupling matrix features respectively.

[0043] In one possible design, based on the multi-view fusion features, discriminative information is extracted from multiple groups of features in different ranges to obtain attention enhancement features, including:

[0044] The multi-view fusion feature of the i-th range group is expressed as:

[0045]

[0046] Where, represents the multi-view fusion features of the first range group, represents the multi-view fusion features of the i-th range group, MLP represents the multi-layer perceptron operation, and l+1 represents the total number of groups;

[0047] Connect the multi-view fusion features of all groups along the time axis to form a unified feature representation H:

[0048]

[0049] Where, Cat temp Represents the splicing operation along the time axis, represents the multi-view fusion features of the second range group, Represents the multi-view fusion features of the last range group;

[0050] Specifically, we unify the feature representation Input into a multi-strategy attention mechanism, the time feature is regarded as a series of timestamps, where T = L (l + 1) represents the total number of timestamps, and the linear transformation is used to transform Generate query tokens And convert the unified feature representation H into a key Sum represents a set of real numbers, D represents the number of feature channels for each perspective, and L represents the number of time axis features for each group;

[0051] For the t-th query term q in the query Q t , the attention operation is performed through the following formula:

[0052]

[0053] Where, α t,i Represents the query word q t With the i-th keyword k i The attention weight between k represents the characteristic dimension of the bond, e represents the natural constant, V t Indicates that you want to t The value of the attention operation, Attn(q t ,K t ,V t ) represents q t The result after attention operation, α t,j Indicates q t With key K t The attention score between the j-th time axis features, v i Indicates V t The i-th time axis feature;

[0054] For query q t , derive the corresponding keys and values ​​through different mapping strategies to obtain different categories f, calculate the derived keys and values ​​through formula (10), and calculate the compressed keys and compressed values ​​through formula (11):

[0055]

[0056] Where, and Represents the derived keys and values, It is a learnable multilayer perceptron that maps keys or values ​​within a group to a single compressed key. or compressed value k (i-1)×L+1:i×L represents the i-th group of time axis features of K, v (i-1)×L+1:i×L represents the i-th group of time axis features of V, f K (q t ,K t ,V t ) represents the derivation strategy of K, f V (q t ,K t ,V t ) represents the derivation strategy of V;

[0057] Based on the derived keys and values, fine-grained keys and fine-grained values ​​are calculated. The calculation formula for the fine-grained key is as follows:

[0058]

[0059] Where [·] represents the index operator used to access vector elements, Rank(·) represents the ranking position in descending order, Rank = 1 corresponds to the highest score, Cat represents the concatenation operation, Indicates q t With compression key The attention score between the l+1 time axis feature vectors in, Indicates the first n subscripts of the l+1 attention scores ranked from large to small, express The ranking number of the attention score of the i-th time axis feature vector, n means selecting the top n feature vectors, Represents a fine-grained key;

[0060] Based on fine-grained keys and fine-grained values The attention enhancement feature o is calculated by the following formula t :

[0061]

[0062] Where g cmp ∈[0,1] and g slc ∈[0,1] is the gate score of the corresponding strategy, Indicates q tThe attention results obtained by the compressed attention strategy, Indicates q t Attention results obtained by selecting an attention strategy;

[0063] Determine g using the following formula cmp and g slc :

[0064] g cmp =Sigmoid(MLP(Q)),g slc =Sigmoid(MLP(Q)) (16)

[0065] Where Sigmoid represents the Sigmoid activation function.

[0066] In one possible design, the classification head includes seven sequentially connected modules, wherein the first module is used to perform an average pooling operation on the attention enhancement features in the time dimension; the second module is used to splice the pooled features and the RR interval; the third module is used to perform a linear transformation on the spliced ​​features; the fourth module introduces nonlinear characteristics through the Relu activation function; the fifth module is used to randomly discard some neurons; the sixth module is used to perform a linear transformation again; the seventh module is used to convert the output into a probability distribution through a normalized exponential function, obtain the final classification result and output it.

[0067] In a second aspect, the present application provides a device for classifying heart beats in a dynamic electrocardiogram, the device comprising:

[0068] A data preprocessing module is configured to preprocess the data set to obtain multi-view data of ECG heartbeats; wherein the multi-view data of ECG heartbeats includes heartbeat waveforms, low-frequency coupling matrices, high-frequency coupling matrices, and RR intervals;

[0069] a multi-view feature extraction module configured to extract features from the multi-view data of the ECG heartbeat to obtain multi-view features;

[0070] The multi-view feature fusion module is configured to group the multi-view feature embeddings of multiple heartbeats into groups according to the neighborhood size increasing from 0 to 1. The features of each view in each group are unified to the same length on the time axis and connected on the channel axis to form multi-range groups. The multi-view cross-attention mechanism is used for each range group to obtain the multi-view fusion features.

[0071] a multi-range group attention module configured to extract discriminative information from the multi-group features of different ranges based on the multi-view fusion features to obtain attention enhancement features;

[0072] The classification module is configured to input the attention enhancement feature and the RR interval into the classification head to obtain a classification prediction result.

[0073] In a third aspect, an embodiment of the present application provides an electronic device comprising: at least one processor and a memory; the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the dynamic electrocardiogram beat classification method described in the first aspect and various possible designs of the first aspect.

[0074] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the dynamic electrocardiogram beat classification method described in the first aspect and various possible designs of the first aspect is implemented.

[0075] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the dynamic electrocardiogram beat classification method described in the first aspect and various possible designs of the first aspect.

[0076] The method, device, equipment, and storage medium for classifying dynamic electrocardiogram beats provided in this application have at least the following beneficial effects:

[0077] This application uses a multi-band coupling matrix to enhance the recognition of the second-order coupling features of the characteristic wave, and constructs a multi-perspective feature fusion module based on channel-time decoupling cross-attention and a dynamic attention module based on the step-by-step information fusion of ECG heartbeat multi-neighborhood ranges. It can adaptively extract and fuse multi-perspective heartbeat information in different ranges, thereby effectively addressing the inherent diversity of ECG signals and improving the generalization performance of the classification model. Specifically, compared with the existing technology, the advantages of this application are reflected in the following three points:

[0078] First, in the existing ECG characterization method based on the outer product matrix of adjacent heartbeat sequences, the coupling value between small-amplitude characteristic waves (QSPT wave groups) is much smaller than the coupling value between large-amplitude characteristic waves (R waves), making them difficult to identify. To address this technical defect, this application proposes a heartbeat characterization method based on a multi-band coupling matrix. This method significantly improves the recognition of the second-order coupling features of the characteristic waves by extracting coupling matrices from low-frequency band signals below 17 Hz and high-frequency band signals from 17 Hz to 100 Hz, highlighting richer heartbeat discrimination information.

[0079] Second, to address the technical deficiencies of existing technologies, which lack effective mechanisms to address the temporal information heterogeneity and signal-to-noise ratio non-stationarity inherent in multi-view ECG features, this paper proposes a Multi-View Feature Fusion (MVCAF) module based on channel-temporal decoupled cross-attention. By combining a two-level fusion strategy with a channel-temporal decoupled cross-attention mechanism, this module achieves adaptive fusion of multi-view features, significantly enhancing the discriminative power of feature representations and effectively improving the generalization performance of the model.

[0080] Third, to address the technical limitations of existing technologies, which lack a dynamic attention mechanism and are unable to adaptively adjust the optimal recognition range for different heartbeats, this paper proposes a dynamic attention module (MRGA) based on the step-by-step information fusion of ECG heartbeat multi-neighborhood ranges. This module uses a progressive fusion strategy for heartbeat multi-neighborhood ranges and a dynamic attention mechanism to selectively extract key features from local details and multi-beat contextual information, thereby reducing the risk of model overfitting and improving model robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0082] Figure 1 A simplified flowchart of a method for classifying heart beats in a dynamic electrocardiogram provided in an embodiment of the present application;

[0083] Figure 2 A specific flow chart of a method for classifying heart beats in a dynamic electrocardiogram provided in an embodiment of the present application;

[0084] Figure 3 A structural diagram of an ECG classification model based on multi-view and multi-range attention provided in an embodiment of the present application;

[0085] Figure 4 Structural diagrams of a one-dimensional convolutional encoder and a two-dimensional convolutional encoder provided in an embodiment of the present application; wherein (a) is a one-dimensional convolutional encoder; (b) is a two-dimensional convolutional encoder;

[0086] Figure 5 A structural diagram of the multi-view cross-fusion module provided in an embodiment of the present application; wherein, (a) the overall structure; (b) the cross-attention fusion process;

[0087] Figure 6 A structural diagram of a classification head provided in an embodiment of the present application;

[0088] Figure 7 This is a structural diagram of the dynamic electrocardiogram beat classification device provided in an embodiment of the present application.

[0089] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0090] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0091] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of information such as financial data or user data involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0092] It should be noted that in the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned. They should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.

[0093] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0094] The present application embodiment provides a method for classifying heart beats of a dynamic electrocardiogram, such as Figure 1 The figure shows a simplified flowchart of the Holter ECG heartbeat classification method provided in an embodiment of the present application. The main steps of this Holter ECG heartbeat classification method include: preprocessing the data set, including baseline drift removal and signal normalization; acquiring multi-view ECG heartbeat data, including heartbeat waveforms, low-frequency coupling matrices, high-frequency coupling matrices, and RR intervals; using a classification model to process the multi-view ECG heartbeat data; and finally, classifying and predicting the target heartbeat into five arrhythmia categories.

[0095] Specifically, if Figure 2 As shown in FIG, it is a specific flow chart of the dynamic electrocardiogram heart beat classification method provided in an embodiment of the present application. The dynamic electrocardiogram heart beat classification method includes the following steps S100-S500.

[0096] S100: Preprocessing the data set to obtain multi-view data of ECG heartbeats; wherein the multi-view data of ECG heartbeats includes heartbeat waveforms, low-frequency coupling matrices, high-frequency coupling matrices, and RR intervals.

[0097] In some embodiments, step S100 may be implemented through the following steps S101 - S103 .

[0098] S101: Prepare dataset.

[0099] The MIT-BIH Arrhythmia Database was used as the experimental dataset. It contains 48 ECG recordings from 47 subjects, sampled at 360 Hz and processed through a 0.1 to 100 Hz bandpass filter. Each recording contains two ECG lead signals and lasts approximately 30 minutes. Cardiologists provided reference annotations for each heartbeat in the database. Heartbeats were detected using lead II ECG signals. Arrhythmias were strictly classified into five main categories according to the American Association for Medical Instrumentation (AAMI) classification: normal (N), supraventricular ectopic beat (SVEB), ventricular ectopic beat (VEB), fusion beat (F), and unknown beat (Q).

[0100] S102: Data preprocessing.

[0101] First, two nested median filters were used to estimate baseline drift, which was then subtracted from the raw signal. The time window sizes for the two median filters were 200 ms and 600 ms, respectively. The ECG signal was then normalized by dividing the entire signal by the mean amplitude of the R-peak in the corresponding recording.

[0102] S103: Acquire multi-view data.

[0103] For each heartbeat, the ECG signal is segmented into segments of 216 time samples with the R peak position as the reference point. Each segment contains data 250 milliseconds before and 350 milliseconds after the R peak, which serves as the first input data of the model.

[0104] The preprocessed signal is filtered using two bandpass filters, respectively, to obtain a low-frequency signal below 17 Hz and a high-frequency signal between 17 Hz and 100 Hz. The outer product of the two adjacent segments before and after each heartbeat is performed to generate the low-frequency and high-frequency coupling matrices, which serve as the other two inputs to the model. The two adjacent segments before and after each heartbeat are obtained by taking the larger of the two R-R intervals before and after the current heartbeat as the window unit. The first segment contains the ECG signal segment two window units forward and one window unit backward relative to the R-peak position of the current heartbeat, while the second segment contains the ECG signal segment one window unit forward and two window units backward relative to the R-peak position of the current heartbeat. After calculating the outer product, the coupling matrix is ​​scaled to a uniform size of 73×73 using bilinear interpolation.

[0105] The RR interval is calculated based on the R peak position recorded in the ECG data set and used as the fourth input data of the model.

[0106] S200: Extract features from the multi-view data of the ECG heartbeat to obtain multi-view features.

[0107] S300: The multi-view feature embeddings of multiple heartbeats are grouped according to the neighborhood size increasing from 0 to 1. The features of each view in each group are unified to the same length on the time axis and connected on the channel axis to form multi-range groups. The multi-view cross-attention mechanism is used for each range group to obtain multi-view fusion features.

[0108] S400: Based on multi-view fusion features, discriminative information is extracted from multiple groups of features in different ranges to obtain attention enhancement features.

[0109] S500: Input the attention enhancement feature and the RR interval into the classification head to obtain the classification prediction result.

[0110] It should be noted that the above steps S200-S500 can be implemented by constructing an ECG heart beat classification model. The overall architecture of the ECG heart beat classification model constructed in this embodiment is as follows: Figure 3 As shown, the ECG heartbeat classification model mainly consists of three stages: multi-view feature extraction stage, multi-view feature fusion stage, and multi-range group attention stage (corresponding to steps S200-S400 respectively).

[0111] In the multi-view feature extraction stage, the neighborhood size s is used to represent the number of adjacent heartbeats on each side of the target heartbeat. iThe multi-view data generated by the heartbeats and their surrounding heartbeats within a maximum neighborhood size l is input into feature extractors at different viewpoints to obtain features for each viewpoint. The classification model uses a 1-D CNN network to extract heartbeat morphological features from heartbeat segments and two 2-D CNN networks with the same structure to extract second-order heartbeat features from the coupling matrix.

[0112] In the multi-view feature fusion stage, the multi-view feature embeddings of multiple heartbeats are grouped according to the neighborhood size from 0 to l. The features of each view in each group are unified to the same length L on the time axis and then connected on the channel axis to form a multi-range group G s Then, a multi-view cross attention fusion (MVCAF) module is used for each group to obtain the fused multi-view features.

[0113] In the multi-range group attention stage, a multi-range group attention (MRGA) module is used to extract discriminative information from multi-group features of different ranges and output attention-enhanced features O.

[0114] Finally, the attention enhancement feature and the RR interval feature R are input into the classification head together to obtain the final classification prediction result (corresponding to step S500).

[0115] In some embodiments, the ECG heartbeat classification model includes a multi-view feature extractor, a multi-view cross-attention fusion module, a multi-range group attention module (MRGA module) and a classification head, wherein the multi-view feature extractor, the multi-view cross-attention fusion module, the multi-range group attention module and the classification head are respectively configured to perform steps S200-S500 as above.

[0116] The multi-view feature extractor includes feature extractors of different viewpoints; wherein, the feature extractors of different viewpoints include a first network and two second networks, the first network extracts the morphological features of the heartbeat from the heartbeat segment as the features of one viewpoint, and the two second networks extract the second-order features of the heartbeat from the low-frequency coupling matrix and the high-frequency coupling matrix as the features of one viewpoint respectively.

[0117] For example, the first network can be selected as a 1-D CNN network, whose structure is as follows Figure 4As shown in (a), the 1-D CNN network consists of three repeated convolutional blocks, each of which contains a 15×15 convolutional layer 1-D Conv, a batch normalization layer BatchNorm, a Relu activation layer Relu, a 3×2 maximum pooling layer MaxPool, and a squeeze activation module (Squeeze-and-excitation, SE) at the end.

[0118] The second network can be selected as a 2D CNN network, whose structure is as follows Figure 4 As shown in (b), the encoder of a two-dimensional convolutional neural network (2DCNN) consists of three convolutional blocks. Each convolutional block contains a 2-D Conv layer, a Batch Normalization layer (BatchNorm), and a MaxPool layer. The kernel sizes of these three convolutional layers are 8×8, 10×10, and 5×5, respectively.

[0119] The structure of the multi-view cross attention fusion module (MVCAF module) is as follows Figure 5 As shown in (a), Figure 5 (b) shows the cross-attention fusion process.

[0120] The MVCAF module uses a multi-level feature fusion mechanism to integrate multi-view information layer by layer. First, the multi-view features in each group are spliced ​​together, and then the spliced ​​features are processed using a linear layer to generate the fused feature Z. f .

[0121] Z f =Linear(Cat(G)) (1) Where G = [Z mor ,Z low ,Z high ], represents the multi-view features within each range group, Z mor ,Z low ,Z high They represent heart beat waveform features, low-frequency coupling matrix features, and high-frequency coupling matrix features respectively. Cat represents the concatenation operation, and Linear represents the fully connected layer.

[0122] By integrating multi-view information from a global perspective, the fused features enhance the complementary interaction between multi-view features. Among the three branches of MVCAF, the cross-attention fusion (CAF) module is responsible for coordinating the interaction between the fused features and the features of each input branch.

[0123] Taking into account the heterogeneity of different dimensions, the CAF module performs cross-attention operations on the channel dimension and the time dimension to exploit the potential dependencies of multi-view data in multiple dimensions. CAF first performs maximum pooling and average pooling respectively, and obtains the fusion feature Z f And input feature Z k Generate channel weights and

[0124]

[0125] Where U f represents the first channel weight, U k Represents the second channel weight, Z k Represents the input features, k∈{mor,high,low}, represents three perspectives, mor, high, low represent the heartbeat waveform, low-frequency coupling matrix, and high-frequency coupling matrix respectively, MLP represents multi-layer perceptron, AvgPool represents global average pooling, and MaxPool represents global maximum pooling.

[0126] Max pooling and average pooling respectively collect important clues about different object features to infer finer attention. By multiplying the channel weights, we get the first cross matrix of shape L×L Then, a softmax operation is applied to the cross matrix and its transpose respectively to obtain the attention scores of the two inputs. The feature obtained by the first cross attention multiplication is calculated as follows:

[0127]

[0128] Where, P f represents the first feature, Softmax represents the normalized exponential function, T represents the matrix transpose, P k Indicates the second feature.

[0129] Then, the two features are transposed, and CAF performs max pooling and average pooling again to obtain channel weights. and

[0130]

[0131] Where W f and W k Represent the third channel weight and the fourth channel weight respectively.

[0132] By multiplying the third channel weight and the fourth channel weight, a second cross matrix with a shape of D×D is obtained Features computed by the second cross-attention multiplication

[0133]

[0134] Where, represents the feature embedding of view k.

[0135] MVCAF performs a similar fusion operation on the three input features to obtain the final multi-view fused feature embedding.

[0136]

[0137] Where, represents the multi-view fusion feature, They represent the fused heart waveform features, the fused low-frequency coupling matrix features, and the fused high-frequency coupling matrix features respectively.

[0138] The structure of the MRGA module is as follows Figure 2 First, MRGA adopts a hierarchical progressive fusion strategy to integrate multi-range temporal features. To represent the fusion of the i-th group features of the previous group, the fusion of the i-th group can be defined as

[0139]

[0140] Where, represents the multi-view fusion features of the first range group, represents the multi-view fusion features of the i-th range group, MLP represents the multi-layer perceptron operation, and l+1 represents the total number of groups.

[0141] Then, the feature embeddings of all groups are concatenated along the time axis to form a unified feature representation H.

[0142]

[0143] Where, Cat temp Represents the splicing operation along the time axis, represents the multi-view fusion features of the second range group, Represents the multi-view fusion features of the last range group.

[0144] Specifically, this embodiment unifies the features Input into a multi-strategy attention mechanism, the time feature is regarded as a series of tokens, where T = L (l + 1) represents the total number of tokens. MRGA transforms the last range group into Generate query tokens And convert the uniform feature H into a key Sum

[0145] For the t-th query token q in the query Q t , the attention operation is defined as

[0146]

[0147] Where, α t,i Represents the query word q t With the i-th keyword k i The attention weight between k represents the characteristic dimension of the bond, e represents the natural constant, V t Indicates that you want to t The value of the attention operation, Attn(q t ,K t ,V t ) represents q t The result of the attention operation, α t,j Indicates q t With key K t The attention score between the j-th time axis features, v i Indicates V t The i-th time axis feature of .

[0148] With query q t The corresponding key tokens and value tokens are derived through different mapping strategies to obtain different categories f.

[0149]

[0150] Where, and Represents the derived keys and values, It is a learnable multilayer perceptron that maps keys or values ​​within a group to a single compressed key. or compressed value k (i-1)×L+1:i×L represents the i-th group of time axis features of K, v (i-1)×L+1:i×L represents the i-th group of time axis features of V, f K (q t ,K t ,V t ) represents the derivation strategy of K, f V (q t ,K t ,V t ) represents the derivation strategy of V.

[0151] For compressed value representation The formula is similar. The compressed representation can capture coarser-grained high-level semantic information and reduce the computational burden of the attention mechanism.

[0152] The token selection strategy identifies and retains the most relevant group tokens by compressed key values. We can retain these tokens in the top n groups sorted by group importance score, which is

[0153]

[0154] Where [·] represents the index operator used to access vector elements, Rank(·) represents the ranking position in descending order, Rank = 1 corresponds to the highest score, Cat represents the concatenation operation, Indicates q t With compression key The attention score between the l+1 time axis feature vectors in, Indicates the first n subscripts of the l+1 attention scores ranked from large to small, express The ranking number of the attention score of the i-th time axis feature vector, n means selecting the top n feature vectors, Represents a fine-grained key.

[0155] For fine-grained values There is a similar formula. Setting n = 1 means focusing on the group with the highest score.

[0156] The final output of the NGSA module is O=o 1:L It is a dynamic combination of the above two mapping strategies, specifically expressed as

[0157]

[0158] Where g cmp ∈[0,1] and g slc ∈[0,1] are the gate scores of the corresponding strategies, which are derived from the query Q through MLP and sigmoid activation function. Indicates q t The attention results obtained by the compressed attention strategy, Indicates q t Attention results obtained after selecting the attention strategy.

[0159] g cmp =Sigmoid(MLP(Q)),g slc =Sigmoid(MLP(Q)) (16)

[0160] Where Sigmoid represents the Sigmoid activation function.

[0161] like Figure 6 As shown in the figure, the classification head consists of seven sequentially connected modules. The first module is Temporal Average Pooling, which is used to perform average pooling operation on the input feature O in the time dimension; the second module is Concatenate, which concatenates the pooled features and the RR feature R; the third module is Linear, which performs linear transformation on the concatenated features; the fourth module is Relu, which introduces nonlinear characteristics through the Relu activation function; the fifth module is Dropout, which can prevent overfitting and randomly discard some neurons; the sixth module is Linear, which performs linear transformation again; the seventh module is Softmax, which converts the output into a probability distribution, obtains the final classification result and outputs it.

[0162] In summary, the core advantages of the Holter ECG heartbeat classification method provided in the embodiments of the present application are:

[0163] (1) Multi-band second-order coupling feature extraction method for adjacent heartbeat sequences. Given that the frequency domain distribution of the QRS complex and the PT wave is significantly different, and the energy of the PT wave is mainly concentrated in the frequency band below 17 Hz, the present invention extracts coupling matrices from the low-frequency band signal below 17 Hz and the high-frequency band signal between 17 Hz and 100 Hz, respectively, to enhance the recognition of the second-order coupling features between different characteristic waves and highlight richer heartbeat discrimination information.

[0164] (2) Multi-view feature fusion module based on channel-time decoupling cross attention. Figure 5 As shown in the figure, the multi-view feature adaptive fusion (MVCAF) module uses a two-level feature fusion mechanism to perform multi-level integration of multi-view input features. This module adaptively extracts the most discriminative feature information from each view by mining the potential correlation between multi-view features. Specifically, the input of the MVCAF module is the features of the three views. First, the features of the three views are spliced, and high-level fusion features are generated through a linear projection layer. Subsequently, the MVCAF module dynamically adjusts the weight contribution of each view feature based on the guidance of the high-level fusion features and the cross attention fusion (CAF) sub-module. Finally, the MVCAF module outputs a fusion feature spliced ​​from the three view features after weight adjustment.

[0165] (3) Dynamic attention module based on step-by-step information fusion of ECG heart beat multi-neighborhood range. Figure 3As shown in the multi-range group attention area in the figure, the multi-range group attention (MRGA) module is used to adaptively aggregate key information from different neighborhood ranges. This module dynamically adjusts the importance weights of each neighborhood range and uses the compression attention mechanism and the selection attention mechanism to achieve the adaptive aggregation. Specifically, the input of the MRGA module includes features of multiple time range groups. First, the MRGA module uses a hierarchical progressive fusion strategy to integrate the features of each time range group. Compared with simple feature stacking input, this fusion strategy can more effectively model the dynamic characteristics of multi-range time series, thereby improving the representation ability of time series features. Subsequently, the MRGA module performs cross-attention operations of the compression strategy and the selection strategy, interacting the large-range group features with the features of each range group respectively, and can efficiently and dynamically capture discriminative information under related ranges. Finally, the MRGA module outputs features after gated weighting processing.

[0166] The present application also provides a device for classifying heart beats of a dynamic electrocardiogram. Figure 7 As shown, the dynamic electrocardiogram heart beat classification device includes:

[0167] The data preprocessing module 701 is configured to preprocess the data set to obtain multi-view data of ECG heartbeats, wherein the multi-view data of ECG heartbeats includes heartbeat waveforms, low-frequency coupling matrices, high-frequency coupling matrices, and RR intervals;

[0168] A multi-view feature extraction module 702 is configured to extract features from the multi-view data of the ECG heartbeat to obtain multi-view features;

[0169] The multi-view feature fusion module 703 is configured to group the multi-view feature embeddings of multiple heartbeats into groups according to the neighborhood size increasing from 0 to 1, unify the features of each view in each group to the same length on the time axis, and connect them on the channel axis to form multiple range groups. The multi-view cross-attention mechanism is applied to each range group to obtain the multi-view fusion feature.

[0170] The multi-range group attention module 704 is configured to extract discriminative information from the multi-group features of different ranges based on the multi-view fusion features to obtain attention enhancement features;

[0171] The classification module 705 is configured to input the attention enhancement feature and the RR interval into the classification head to obtain a classification prediction result.

[0172] An embodiment of the present application provides an electronic device, which may include a processor and a memory, wherein the processor and the memory can communicate with each other; illustratively, the processor and the memory communicate with each other via a communication bus.

[0173] The processor executes the computer-executable instructions stored in the memory, so that the processor implements the solutions in the above embodiments. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0174] The communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. System buses can be categorized as address buses, data buses, and control buses. Transceivers enable communication between the database access device and other computers (e.g., clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) or non-volatile memory.

[0175] The electronic device provided in the embodiment of the present application may be the terminal device of the above embodiment.

[0176] An embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on a computer, the computer executes the technical solution of the dynamic electrocardiogram beat classification method of the above embodiment.

[0177] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when at least one processor executes the computer program, it can implement the technical solution of the dynamic electrocardiogram beat classification method in the above embodiment.

[0178] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.

[0179] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these modules may be selected to implement the solution of this embodiment based on actual needs.

[0180] In addition, the functional modules in the various embodiments of the present application may be integrated into a single processing unit, or each module may exist physically separately, or two or more modules may be integrated into a single unit. The above-mentioned modules may be implemented in the form of hardware or hardware plus software functional units.

[0181] The above-mentioned integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some steps of the methods of various embodiments of the present application.

[0182] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), or application-specific integrated circuits (ASICs). A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.

[0183] The memory may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk.

[0184] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, and control buses.

[0185] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0186] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic control unit or a main control device.

[0187] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for classifying heart beats of a dynamic electrocardiogram, characterized in that: The method comprises: Preprocessing the data set to obtain multi-view data of ECG heartbeats; wherein the multi-view data of ECG heartbeats includes heartbeat waveforms, low-frequency coupling matrices, high-frequency coupling matrices, and RR intervals; Extracting features from the multi-view data of the ECG heartbeat to obtain multi-view features; The multi-view feature embeddings of multiple heartbeats are grouped according to the neighborhood size increasing from 0 to 1. The features of each view in each group are unified to the same length on the time axis and connected on the channel axis to form multi-range groups. The multi-view cross-attention mechanism is used for each range group to obtain multi-view fusion features. Based on the multi-view fusion features, discriminative information is extracted from multiple groups of features in different ranges to obtain attention enhancement features; The attention enhancement feature is input into the classification head together with the RR interval to obtain the classification prediction result.

2. The method for classifying heart beats of a dynamic electrocardiogram according to claim 1, wherein: Methods for preprocessing the data set to obtain multi-view ECG heartbeat data include: Acquiring electrocardiogram data, and performing baseline drift removal and signal normalization on the electrocardiogram data to obtain a preprocessed ECG signal; Based on the preprocessed ECG signal, for each heartbeat, with the R peak position as a reference point, the ECG signal is segmented into a plurality of time sample segments, each segment containing data t1 milliseconds before and t2 milliseconds after the R peak as a heartbeat segment; Two band-pass filters are applied to the preprocessed ECG signal to obtain a low-frequency signal below the first frequency and a high-frequency signal from the first frequency to the second frequency. The outer product operation is performed on the two adjacent segments before and after each heartbeat to obtain the low-frequency coupling matrix and the high-frequency coupling matrix. The RR interval is calculated based on the R peak position recorded in the ECG dataset.

3. The method for classifying heart beats of a dynamic electrocardiogram according to claim 1, wherein: Feature extraction is performed on the multi-view data of the ECG heartbeat to obtain multi-view features, including: Establishing a multi-view feature extractor; wherein the multi-view feature extractor includes feature extractors of different viewpoints; The multi-perspective data generated by the target heartbeat and its surrounding heartbeats within a maximum neighborhood size l are input into feature extractors of different perspectives to obtain features of each perspective and obtain the multi-perspective features; wherein the feature extractors of different perspectives include a first network and two second networks, the first network extracts the morphological features of the heartbeat from the heartbeat segment as features of one perspective, and the two second networks extract the second-order features of the heartbeat from the low-frequency coupling matrix and the high-frequency coupling matrix as features of one perspective respectively.

4. The method for classifying dynamic electrocardiogram beats according to claim 3, wherein: The first network includes three identical first convolutional blocks, each of which includes a convolutional layer, a batch normalization layer, an activation layer, a maximum pooling layer, and a squeeze activation module connected in sequence; The second network includes three second convolution blocks, each of which includes a second convolution layer, a second batch normalization layer, and a second maximum pooling layer connected in sequence, and the three second convolution blocks have different kernel sizes.

5. The method for classifying dynamic electrocardiogram beats according to claim 1, wherein: A multi-view cross-attention mechanism is used for each range group to obtain multi-view fusion features, including: The multi-view features in each range group are spliced ​​together, and then a linear layer is used to generate the fusion feature Z f , the calculation process is: Z f =Linear(Cat(G)) (1) Where, G=[Z mor ,Z low ,Z high ], represents the multi-view features within each range group, Z mor ,Z low ,Z high They represent heartbeat waveform features, low-frequency coupling matrix features, and high-frequency coupling matrix features, respectively. Cat represents the concatenation operation, and Linear represents the fully connected layer. Based on the fusion feature Z f , the first channel weight and the second channel weight are generated by the following formula: Where U f represents the first channel weight, U k Represents the second channel weight, Z k Represents the input features, k∈{mor,high,low}, represents three perspectives, mor, high, low represent the heartbeat waveform, low-frequency coupling matrix, and high-frequency coupling matrix respectively, MLP represents multi-layer perceptron, AvgPool represents global average pooling, and MaxPool represents global maximum pooling; Based on the first channel weight and the second channel weight, the first feature and the second feature are calculated using the following formula: Where, P f represents the first feature, Softmax represents the normalized exponential function, T represents the matrix transpose, P k Indicates the second feature; Based on the first and second features, the third channel weight and the fourth channel weight are calculated using the following formula: Where W f and W k Represent the third channel weight and the fourth channel weight respectively; Based on the third channel weight and the fourth channel weight, the feature embedding of each view is calculated by the following formula: Where, Represents the feature embedding of view k; Based on the embedding of each view feature, the multi-view fusion feature is calculated by the following formula: Where, represents the multi-view fusion feature, They represent the fused heart waveform features, the fused low-frequency coupling matrix features, and the fused high-frequency coupling matrix features respectively.

6. The method for classifying heart beats of a dynamic electrocardiogram according to claim 1, wherein: Based on the multi-view fusion features, discriminative information is extracted from multiple groups of features in different ranges to obtain attention enhancement features, including: The multi-view fusion feature of the i-th range group is expressed as: Where, represents the multi-view fusion features of the first range group, represents the multi-view fusion features of the i+1th range group, MLP represents the multi-layer perceptron operation, and l+1 represents the total number of groups; Connect the multi-view fusion features of all groups along the time axis to form a unified feature representation H: Where, Cat temp Represents the splicing operation along the time axis, represents the multi-view fusion features of the second range group, Represents the multi-view fusion features of the last range group; Specifically, we unify the feature representation Input into a multi-strategy attention mechanism, the time feature is regarded as a series of timestamps, where T = L (l + 1) represents the total number of timestamps, and the linear transformation is used to transform Generate query tokens And convert the unified feature representation H into a key Sum represents a set of real numbers, D represents the number of feature channels for each perspective, and L represents the number of time axis features for each group; For the t-th query term q in the query Q t , the attention operation is performed through the following formula: Where, α t,i Represents the query word q t With the i-th keyword k i The attention weight between k represents the characteristic dimension of the bond, e represents the natural constant, V t Indicates that you want to t The value of the attention operation, Attn(q t ,K t ,V t ) represents q t The result after attention operation, α t,j Indicates q t With key K t The attention score between the j-th time axis features, v i Indicates V t The i-th time axis feature; For query q t , derive the corresponding keys and values ​​through different mapping strategies to obtain different categories f, calculate the derived keys and values ​​through formula (10), and calculate the compressed keys and compressed values ​​through formula (11): Where, and Represents the derived keys and values, It is a learnable multilayer perceptron that maps keys or values ​​within a group to a single compressed key. or compressed value k (i-1)×L+1:i×L represents the i-th group of time axis features of K, v (i-1)×L+1:i×L represents the i-th group of time axis features of V, f K (q t ,K t ,V t ) represents the derivation strategy of K, f V (q t ,K t ,V t ) represents the derivation strategy of V; Based on the derived keys and values, fine-grained keys and fine-grained values ​​are calculated. The calculation formula for the fine-grained key is as follows: Where [·] represents the index operator used to access vector elements, Rank(·) represents the ranking position in descending order, Rank = 1 corresponds to the highest score, Cat represents the concatenation operation, Indicates q t With compression key The attention score between the l+1 time axis feature vectors in, Indicates the first n subscripts of the l+1 attention scores ranked from large to small, express The ranking number of the attention score of the i-th time axis feature vector, n means selecting the top n feature vectors, Represents a fine-grained key; Based on fine-grained keys and fine-grained values The attention enhancement feature o is calculated by the following formula t : Where g cmp ∈[0,1] and g slc ∈[0,1] is the gate score of the corresponding strategy, Indicates q t The attention results obtained by the compressed attention strategy, Indicates q t Attention results obtained by selecting an attention strategy; Determine g using the following formula cmp and g slc : g cmp =Sigmoid(MLP(Q)),g slc =Sigmoid(MLP(Q)) (16) Where Sigmoid represents the Sigmoid activation function.

7. The method for classifying heart beats of a dynamic electrocardiogram according to claim 1, wherein: The classification head includes seven sequentially connected modules, wherein the first module is used to perform an average pooling operation on the attention enhancement features in the time dimension; the second module is used to splice the pooled features and the RR interval; the third module is used to perform a linear transformation on the spliced ​​features; the fourth module introduces nonlinear characteristics through the ReLU activation function; the fifth module is used to randomly discard some neurons; the sixth module is used to perform a linear transformation again; the seventh module is used to convert the output into a probability distribution through a normalized exponential function, obtain the final classification result and output it.

8. A dynamic electrocardiogram beat classification device, characterized in that: The device comprises: A data preprocessing module is configured to preprocess the data set to obtain multi-view data of ECG heartbeats; wherein the multi-view data of ECG heartbeats includes heartbeat waveforms, low-frequency coupling matrices, high-frequency coupling matrices, and RR intervals; a multi-view feature extraction module configured to extract features from the multi-view data of the ECG heartbeat to obtain multi-view features; The multi-view feature fusion module is configured to group the multi-view feature embeddings of multiple heartbeats into groups according to the neighborhood size increasing from 0 to 1. The features of each view in each group are unified to the same length on the time axis and connected on the channel axis to form multi-range groups. The multi-view cross-attention mechanism is used for each range group to obtain the multi-view fusion features. a multi-range group attention module configured to extract discriminative information from the multi-group features of different ranges based on the multi-view fusion features to obtain attention enhancement features; The classification module is configured to input the attention enhancement feature and the RR interval into the classification head to obtain a classification prediction result.

9. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the Holter electrocardiogram beat classification method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the Holter electrocardiogram beat classification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Heart beat classification method, heart beat classification device and electronic equipment

    CN109948396A

  • Electrocardiogram cardiac beat classification method and system, medium, equipment and terminal

    CN115462797A

  • Multi-lead electrocardiogram classification and identification method based on convolution and self-attention mechanism

    CN115470828A

  • Heart beat classification model training method, classification method, equipment and storage medium

    CN116342963A

  • Electrocardiogram image processing method and device, medium, and electrocardiograph

    US20230293079A1