Multi-scale Transformer Encoder Network, Its Construction Method and Application

By building a multi-scale Transformer encoder network, using multi-scale feature extraction and feature interaction, the problems of key features loss and long training time in the electrical signal analysis of existing technology are solved, achieving higher analysis accuracy and shorter training time.

CN117114052BActive Publication Date: 2025-07-01TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311241751.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2025-07-01
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

The existing Transformer encoder network has problems of critical features loss and long training time in ECG signal analysis.

Method used

Using a multi-scale Transformer encoder network, three Transformer encoder units, Convolution module, liner module, Interaction output module and Pooling and Normalize module are built to realize multi-scale feature extraction and feature interaction, reducing calculation parameters and time.

Benefits of technology

It effectively avoids the loss of key features, significantly improves the accuracy of ECG signal analysis, and greatly shortens the training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117114052B_ABST
    Figure CN117114052B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of electrocardiogram signal processing and analysis, and particularly relates to a multi-scale Transformer encoder network, a method for constructing the same, and an application, which solves the technical problems in the background art. It includes data preprocessing; constructing a multi-scale Transformer encoder network model; training, optimizing, and testing the multi-scale Transformer encoder network model; and when it is determined that the accuracy rate meets the requirements, the optimal multi-scale Transformer encoder network model is obtained. The method has a simple process flow and greatly shortens the training time; the multi-scale Transformer encoder network constructed by the method of the present invention can make a more accurate analysis of electrocardiogram signals and is of great help to the diagnosis of heart failure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of processing and analysis of electrocardiogram signals, and particularly relates to a multi-scale Transformer encoder network, a construction method thereof, and an application. Background Art

[0002] There are approximately 26 million adults suffering from heart failure in the world, and the mortality rate caused by this disease is particularly high among all diseases. The characteristics of this disease are: rapid onset, extremely easy to be severe, and wide distribution area. Therefore, accurately identifying these diseases and treating them early is a very crucial link. Electrocardiogram is non-invasive and intuitive, and it is a commonly used auxiliary tool for doctors to diagnose most heart diseases. Due to the complexity of electrocardiogram signals (ECG signals), even experts in the industry are still unable to obtain sufficient information from electrocardiogram signals to ensure accurate diagnosis. In recent years, many researchers have started to apply deep learning to process electrocardiogram signal information, and networks such as CNN, ResNet, DenseNet, RNN, Fast-RNN, Transfomer, and BERT have emerged in an endless stream, achieving good results.

[0003] However, firstly, due to the influence of multiple factors such as the instability of electrocardiogram signals themselves and external noise interference in device detection, the use of single-modal electrocardiogram signal data for deep learning has severely limited classification, resulting in inaccurate calculation results. In order to comprehensively, stably, and accurately analyze electrocardiogram signals, deeply explore other characteristic attributes of electrocardiogram signals, and comprehensively analyze all attribute features, a dataset combining multi-modal data features based on electrocardiogram signals and their generated images is constructed, so as to form multi-modal data with feature fusion.

[0004] In 2017, Google proposed the Transformer model in the paper "Attention is All you need", which uses the Self-Attention structure to replace the RNN network structure commonly used in NLP tasks. The Transformer model has the advantages of parallelizability in processing sequence data and considering global dependencies. This model can still pay attention to the mutual parts for clinical problems, especially for the processing of electrocardiogram signal features. Therefore, in 2019, at the IEEE International Conference, a top international conference, it was proposed to use the Transformer encoder network for arrhythmia classification and recognition. The advantages of this network are: (1) Using multi-head attention can describe the complete relationships of multiple dependencies and pay attention to the representation information at different positions. (2) The residual connection adds the original input position embedding and the position weight specific information after multi-head attention to obtain more accurate information of the mutual dependence attention mechanism. (3) Finally, it can form full feature information. However, there are still some disadvantages: (1) Using the encoder unit in this network to finally perform full connection feature extraction can still cause the loss of most key features. (2) Using multiple Transformer encoder units increases the time and space complexity, resulting in a long training time. Summary of the Invention

[0005] To overcome the technical defects that when the encoder unit in the existing network finally performs full connection feature extraction, most key features are still lost; and using multiple Transformer encoder units increases the time and space complexity, resulting in a long training time, the present invention provides a multi-scale Transformer encoder network, its construction method, and applications.

[0006] The present invention provides a construction method for a multi-scale Transformer encoder network for analyzing electrocardiogram signals, including the following steps:

[0007] Step 1, data preprocessing: Respectively convert the normal electrocardiogram signals and the electrocardiogram signals of heart failure patients into three different image data using Gramian angular field, recurrence plot, and Markov transition field, and stack them into a three-channel image rectangle; perform denoising processing on the original electrocardiogram signals, and then segment the generated image data after denoising and add it to the three-channel image matrix to form a multi-modal fusion data set;

[0008] Step 2: Construct a multi-scale Transformer encoder network model, which includes three Transformer encoder units, three Convolution modules, three liner modules, three Interaction output modules, and a Pooling and Normalize module. The output result of each Transformer encoder unit is divided into two paths. One path is sequentially input into the next Transformer encoder unit, and the other path first undergoes multi-scale convolution operations through the Convolution module, and then extracts features through the liner module. The output results of the three liner modules are pairwise interacted through the three Interaction output modules. The outputs of the three Interaction output modules are subjected to fully connected linear operations and classification through the Pooling and Normalize module, and finally, data output is performed through the Final output module. Each Transformer encoder unit includes three cascaded Transformer encoder block modules, and each Convolution module includes three cascaded Dense Block modules.

[0009] Step 3: Divide the multi-modal fusion dataset formed in Step 1 into a training set and a test set according to a certain proportion. The multi-scale Transformer encoder network model is trained and optimized multiple times through the training set, and then the optimized multi-scale Transformer encoder network model is tested through the test set to obtain the judgment accuracy rate of the optimized multi-scale Transformer encoder network model. When the judgment accuracy rate meets the requirements, the optimal multi-scale Transformer encoder network model is obtained.

[0010] In the model construction of Step 2, the output results of the three liner modules are pairwise interacted through the three Interaction output modules to realize the interaction of the three outputs to extract their features. In this way, through feature extraction from multiple angles, the loss of key information in the image can be avoided to the greatest extent, and the pairwise interaction can ultimately enhance the comprehensive understanding of data features from different angles. Multi-scale feature extraction helps to avoid the loss of key features and form a more comprehensive and integrated feature result; constructing three Transformer encoder blocks can reduce the number of calculation parameters and reduce the calculation time; finally, pairwise interaction feature extraction can make the output result more robust, greatly shortening the training time of the model, and significantly improving the analysis accuracy of electrocardiogram signals.

[0011] The present invention also provides a multi-scale Transformer encoder network for analyzing electrocardiogram signals, which includes three Transformer encoder units, three Convolution modules, three liner modules, three Interaction output modules, and a Pooling and Normalize module. The output result of each Transformer encoder unit is divided into two paths. One path is sequentially input into the next Transformer encoder unit, and the other path first undergoes multi-scale convolution operations through the Convolution module, and then extracts features through the liner module. The output results of the three liner modules are pairwise interacted through the three Interaction output modules. The outputs of the three Interaction output modules undergo fully connected linear operations and classification through the Pooling and Normalize module, and finally, data output is performed through the Final output module. Each Transformer encoder unit includes three Transformer encoder block modules, and each Convolution module includes three Dense Block modules.

[0012] The present invention also provides an application of the multi-scale Transformer encoder network for analyzing electrocardiogram signals in the diagnosis of heart failure.

[0013] The technical solution provided by the present invention has the following advantages compared with the prior art: In the multi-scale Transformer encoder network of the present invention, multi-scale feature extraction helps to avoid the loss of key features and form a more comprehensively integrated feature result; constructing three Transformer encoder blocks can reduce the calculation parameters and calculation time; finally, pairwise interaction feature extraction can make the output result more robust, and the accuracy of electrocardiogram signal analysis is significantly improved; the method process is simple and the training time is greatly shortened; the multi-scale Transformer encoder network constructed by the method of the present invention can make a more accurate analysis of electrocardiogram signals and is of great help to the diagnosis of heart failure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 It is a schematic structural diagram of the Transformer encoder unit described in the present invention;

[0017] Figure 2 It is a schematic structural diagram of the Transformer encoder network described in the background art;

[0018] Figure 3 It is a schematic structural diagram of the multi-scale Transformer encoder network described in the present invention;

[0019] Figure 4 It is a flowchart of the construction method of the multi-scale Transformer encoder network for analyzing electrocardiogram signals described in the present invention. Detailed implementation manners

[0020] In order to be able to more clearly understand the above objects, features and advantages of the present invention, the following will further describe the solutions of the present invention. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0021] In the description, it should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. It should be noted that unless otherwise clearly defined and limited, the terms "installation", "connection" and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.

[0022] Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present invention, rather than all the embodiments.

[0023] The following will detail the specific embodiments of the present invention with reference to the drawings.

[0024] In one embodiment, as Figure 1As shown, a construction method of a multi-scale Transformer encoder network for analyzing electrocardiogram signals is disclosed, and the steps include:

[0025] Step 1. Data preprocessing: Obtain normal electrocardiogram signals from the MIT-BIH Normal Sinus Rhythm database, and obtain electrocardiogram signals of heart failure patients from the Beth Israel Deaconess Medical Centre heart failure database. Respectively convert the normal electrocardiogram signals and the electrocardiogram signals of heart failure patients into three different image data using Gramian angular field, recurrence plot, and Markov transition field, and stack them into a three-channel image rectangle; Denoise the original electrocardiogram signals, and the denoising process includes removing zero drift interference and electromyogram interference; Label the normal electrocardiogram signal dataset as 1, and label the electrocardiogram signal dataset of heart failure patients as 0; Then segment the image data generated after denoising and add it to the three-channel image matrix to form a multi-modal fusion dataset;

[0026] Step 2. Construct a multi-scale Transformer encoder network model: It includes three Transformer encoder units, three Convolution modules, three liner modules, three Interaction output modules, and one Pooling and Normalize module. The output result of each Transformer encoder unit is divided into two paths, one of which is sequentially input into the next Transformer encoder unit, and the other path first undergoes multi-scale convolution operation through the Convolution module, and then extracts features through the liner module. The output results of the three liner modules are pairwise interacted through the three Interaction output modules. The output of the three Interaction output modules undergoes full connection linear operation and classification through the Pooling and Normalize module, and finally data output is performed through the Final output module; Each Transformer encoder unit includes three cascaded Transformer encoder block modules, and each Convolution module includes three cascaded Dense Block modules;

[0027] Step 3: Divide the multi-modal fusion dataset formed in Step 1 into a training set and a test set according to a certain ratio. Use the training set to train and optimize the multi-scale Transformer encoder network model. The training and optimization process is as follows: Input the training set into the multi-scale Transformer encoder network model constructed in Step 2, adopt the Adam optimizer and the Relu loss function, and set the initial learning rate to 0.001. Stop training and save the model parameters until the loss of the multi-scale Transformer encoder network model does not decrease and the accuracy is stable. Test the optimized multi-scale Transformer encoder network model with the test set to obtain the judgment accuracy of the optimized multi-scale Transformer encoder network model. When the judgment accuracy meets the requirements, the optimal multi-scale Transformer encoder network model is obtained. Specifically, when inputting data, divide the dataset into a training set and a test set according to a ratio of 8:2. And the data will be input in a dimension of 132×132×3 per batch size, and train with batch_size = 64, that is, input 64 data with a size dimension of 132×132×3 each time.

[0028] In the model construction of Step 2, the output results of the three liner modules are interacted pairwise through three Interaction output modules to realize the interaction of the three outputs and extract their features. In this way, through feature extraction from multiple angles, the loss of key information in the image can be avoided to the greatest extent, and the pairwise interaction can ultimately enhance the comprehensive understanding of data features from different angles. Multi-scale feature extraction helps to avoid the loss of key features and form a more comprehensively integrated feature result; constructing three Transformer encoder blocks can reduce the calculation parameters and reduce the calculation time; finally, pairwise interaction feature extraction can make the output result more robust, greatly shortening the training time of the model, and significantly improving the analysis accuracy of electrocardiogram signals.

[0029] The following uses a confusion matrix to compare the analysis accuracies of electrocardiogram signals of two comparative examples and the multi-scale Transformer encoder network described in this application.

[0030] Among them, the first group of data comes from the first comparative document G. Zhang, Y. J. Si, W. Y. Yang, and D. Wang, A robust multilevel DWT densely network for cardiovascular disease classification, Sensors, vol. 20, no. 17, p. 4777, 2020. The second group of data comes from the second comparative document Li D, Shi C, Zhao J, et al. Intra-Patient and Inter-Patient Multi-Classification of Severe Cardiovascular Diseases Based on CResFormer[J]. Tsinghua Science and Technology, 2022, 28(2): 386-404.

[0031] The confusion matrix records the numbers of true positives, false positives, true negatives, and false negatives. OA is the proportion of correct classifications among all classifications. True positive (TP), false positive (FP), true negative (TN), and false negative (FN) are defined as follows: True positive (TP) is the category that is predicted as positive and is actually positive. False positive (FP) is the category that is predicted as positive but is actually negative. True negative (TN) is the category that is predicted as negative but is actually positive. False negative (FN) is the category that is predicted as negative and is actually negative. Accuracy (Acc) is the ratio of the correct classification of the current category to the current category. Specificity (Spe) is the proportion of samples judged to be negative among the actual negative samples. Sensitivity (Sen) is the proportion of samples judged to be positive among the samples that are actually in the positive category. Result analysis: The experimental results for heart failure classification are shown in the above table. It is found from the above comparison that the model proposed in the present invention is the highest among all performance evaluation indicators such as accuracy, sensitivity, precision, and specificity, which proves that the method described in the present invention can fully detect heart failure according to the actual electrocardiogram signals.

[0032] In one embodiment, a multi-scale Transformer encoder network for analyzing electrocardiogram signals is also disclosed, which includes three Transformer encoder units, three Convolution modules, three liner modules, three Interaction output modules, and a Pooling and Normalize module. The output result of each Transformer encoder unit is divided into two paths. One path is sequentially input into the next Transformer encoder unit, and the other path first undergoes multi-scale convolution operations through the Convolution module, and then extracts features through the liner module. The output results of the three liner modules are pairwise interacted through the three Interaction output modules. The outputs of the three Interaction output modules undergo fully connected linear operations and classification through the Pooling and Normalize module, and finally, data output is performed through the Finaloutput module; each Transformer encoder unit includes three Transformerencoder block modules, and each Convolution module includes three Dense Block modules.

[0033] In one embodiment, an application of the multi-scale Transformer encoder network for analyzing electrocardiogram signals in the diagnosis of heart failure is also provided.

[0034] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Although the foregoing embodiments have been described in detail, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments, and they should all be covered by the protection scope of the claims.

Claims

1. A construction method of a multi-scale Transformer encoder network for analyzing electrocardiogram signals, characterized in that, The steps included are as follows: Step 1, data preprocessing: The normal electrocardiogram (ECG) signals and the ECG signals of heart failure patients are respectively converted into three different feature images using Gramian Angular Field, recurrence plot, and Markov transition field. The three feature images are superimposed into a three-channel image rectangle. The original ECG signals are denoised, and then the image data generated after denoising is segmented and added to the three-channel image matrix to form a multi-modal fusion dataset. Step 2, constructing a multi-scale Transformer encoder network model: It includes three Transformer encoder units, three Convolution modules, three liner modules, three Interaction output modules, and one Pooling and Normalize module. The output result of each Transformer encoder unit is divided into two paths. One path is sequentially input into the next Transformer encoder unit, and the other path first undergoes multi-scale convolution operations through the Convolution module, and then extracts features through the liner module. The output results of the three liner modules are pairwise interacted through the three Interaction output modules. The outputs of the three Interaction output modules undergo fully connected linear operations and classification through the Pooling and Normalize module, and finally, data is output through the Final output module. Each Transformer encoder unit includes three cascaded Transformer encoder block modules, and each Convolution module includes three cascaded Dense Block modules. Step 3, dividing the multi-modal fusion dataset formed in Step 1 into a training set and a test set in proportion. The multi-scale Transformer encoder network model is trained and optimized multiple times through the training set, and then the optimized multi-scale Transformer encoder network model is tested through the test set to obtain the judgment accuracy rate of the optimized multi-scale Transformer encoder network model. When the judgment accuracy rate meets the requirements, the optimal multi-scale Transformer encoder network model is obtained.

2. The construction method of the multi-scale Transformer encoder network for analyzing electrocardiogram signals according to claim 1, characterized in that, In Step 1, the normal ECG signals are obtained from the MIT-BIH Normal Sinus Rhythm database, and the ECG signals of heart failure patients are obtained from the Beth Israel Deaconess Medical Centre heart failure database.

3. The method for constructing a multi-scale Transformer encoder network for analyzing electrocardiogram signals according to claim 1, wherein, The denoising process in Step 1 includes removing zero-drift interference and electromyogram interference.

4. The construction method of the multi-scale Transformer encoder network for analyzing electrocardiogram signals according to claim 1, wherein In Step 1, the normal ECG signal dataset is labeled as 1, and the ECG signal dataset of heart failure patients is labeled as 0.

5. The construction method of the multi-scale Transformer encoder network for analyzing electrocardiogram signals according to claim 1. In step three, the training and optimization process is as follows: Input the training set into the multi-scale Transformer encoder network model constructed in step two, adopt the Adam optimizer and the Relu loss function, set the initial learning rate to 0.001, and stop training and save the model parameters until the loss of the multi-scale Transformer encoder network model does not decrease and the accuracy is stable.