Training method, device, equipment and storage medium for electrocardiogram analysis model

By using the Gumbel-Softmax operator in the LSTM module of the ECG analysis model to binarize the gate weights, the problem of slow LSTM inference speed in the prior art is solved, and more efficient computing resource utilization and better model performance are achieved.

CN115886828BActive Publication Date: 2025-05-09GUANGZHOU SHIYUAN ELECTRONICS CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111138917.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-05-09
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

The inference speed of long and short-term memory networks (LSTMs) in electrocardiogram signal analysis cannot be effectively improved in the prior art, resulting in waste of computing resources and inefficient efficiency.

Method used

By introducing the Gumbel-Softmax operator into the long and short-term memory network module of the electrocardiogram analysis model, the gate weight tends to be binarized, thereby improving the inference speed of LSTM.

Benefits of technology

It realizes that the inference speed of LSTM is improved without reducing the inference accuracy, optimizes the utilization of computing resources, and improves the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115886828B_ABST
    Figure CN115886828B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a training method, device, equipment and storage medium for an electrocardiogram analysis model, which includes: obtaining an electrocardiogram training set and an electrocardiogram verification set; inputting multiple electrocardiogram signals in the electrocardiogram training set and multiple electrocardiogram signals in the electrocardiogram verification set into the electrocardiogram analysis model respectively, obtaining multiple training output results and multiple verification output results, and the electrocardiogram analysis model includes a long short-term memory network module; updating the model parameters of the electrocardiogram analysis model according to the multiple training output results, and the model parameters include the gate weights used by the long short-term memory network module; updating the parameters of the Gumbel‑Softmax operator in the long short-term memory network module according to the multiple verification output results, and the Gumbel‑Softmax operator is used to make the gate weights tend to be binary; continue to use the electrocardiogram training set and the electrocardiogram verification set to train the electrocardiogram analysis model until the electrocardiogram analysis model meets the training stop condition. The above method can solve the technical problem that the reasoning speed of LSTM cannot be improved in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of neural network technology, and in particular to a training method, apparatus, device and storage medium for an electrocardiogram analysis model. Background Art

[0002] ECG signals can reflect the electrophysiological process of cardiac activity and are often used to assist in the diagnosis of heart disease. Among them, common ECG waveforms in ECG signals include P wave, QRS wave and T wave. By analyzing ECG signals, the classification results of ECG signals (such as atrial fibrillation type, atrial flutter type, etc.) can be obtained, and then the diagnosis of heart disease can be assisted based on the classification results. However, due to the low amplitude, high complexity and nonlinearity of ECG signals, it is difficult to manually classify ECG signals accurately and quickly.

[0003] At present, with the development of computer technology, more and more neural network models based on deep learning are used to assist in the diagnosis of heart diseases. In order to obtain accurate ECG signal classification results, it is necessary to enable the neural network model to abstract feature information with strong expressive ability from the ECG signal. At this time, the convolutional neural network structure with spatial feature extraction and the recurrent neural network structure with temporal feature extraction have gained more and more attention. For example, after processing the ECG signal and the corresponding RR interval (the time limit between two adjacent R waves) using a neural network model composed of a convolutional neural network and a long short-term memory network (LSTM), it is possible to detect whether the ECG signal is of atrial fibrillation type. However, the LSTM reasoning process requires a lot of computing resources. Therefore, how to improve the reasoning speed of LSTM has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] The embodiments of the present application provide a training method, apparatus, device and storage medium for an electrocardiogram analysis model to solve the technical problem in the related art that the inference speed of LSTM cannot be improved.

[0005] In a first aspect, an embodiment of the present application provides a method for training an electrocardiogram analysis model, comprising:

[0006] Get ECG training set and ECG validation set;

[0007] Inputting the multiple ECG signals in the ECG training set and the multiple ECG signals in the ECG verification set into the ECG analysis model respectively to obtain multiple training output results and multiple verification output results, wherein the ECG analysis model includes a long short-term memory network module;

[0008] updating the model parameters of the electrocardiographic analysis model according to the plurality of training output results, wherein the model parameters include gate weights used by the long short-term memory network module;

[0009] updating the parameters of the Gumbel-Softmax operator in the long short-term memory network module according to the plurality of verification output results, wherein the Gumbel-Softmax operator is used to make the gate weights tend to be binary;

[0010] The operation of respectively inputting the plurality of ECG signals in the ECG training set and the plurality of ECG signals in the ECG verification set into the ECG analysis model is repeatedly performed until the ECG analysis model meets the training stop condition.

[0011] In a second aspect, an embodiment of the present application further provides a training device for an electrocardiogram analysis model, comprising:

[0012] A data set acquisition module is used to obtain an ECG training set and an ECG verification set;

[0013] A model reasoning module, used for inputting the multiple ECG signals in the ECG training set and the multiple ECG signals in the ECG verification set into the ECG analysis model respectively, to obtain multiple training output results and multiple verification output results, wherein the ECG analysis model includes a long short-term memory network module;

[0014] A parameter updating module, used for updating the model parameters of the electrocardiographic analysis model according to the plurality of training output results, wherein the model parameters include gate weights used by the long short-term memory network module;

[0015] An operator updating module, used for updating the parameters of the Gumbel-Softmax operator in the long short-term memory network module according to the plurality of verification output results, wherein the Gumbel-Softmax operator is used for making the gate weight tend to be binary;

[0016] The repeated training module is used to repeatedly execute the operation of inputting the multiple ECG signals in the ECG training set and the multiple ECG signals in the ECG verification set into the ECG analysis model respectively until the ECG analysis model meets the training stop condition.

[0017] In a third aspect, an embodiment of the present application further provides a training device for an electrocardiogram analysis model, including:

[0018] one or more processors;

[0019] A memory for storing one or more programs;

[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the electrocardiogram analysis model as described in the first aspect.

[0021] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the training method for the electrocardiogram analysis model as described in the first aspect.

[0022] In one embodiment of the present application, multiple ECG signals in an ECG training set and an ECG verification set are respectively input into an ECG analysis model, and then the gate weights of the long short-term memory network module in the ECG analysis model are updated according to the training output results corresponding to the ECG training set, and the model parameters of other layers are optionally updated, and the parameters of the Gumbel-Softmax operator used in the long short-term memory network module are updated according to the verification output results corresponding to the ECG verification set. The Gumbel-Softmax operator acts on the gate weights of the long short-term memory network module, so that the gate weights tend to be binarized after the update. The technical problem that the reasoning speed of LSTM cannot be improved in the related art is solved. By using the Gumbel-Softmax operator in the long short-term memory network module, and by updating the parameters and gate weights of the Gumbel-Softmax operator during the training process to achieve binarization of the gate weights, the reasoning speed of LSTM can be improved, and by updating the parameters of the Gumbel-Softmax operator through the verification output results, the reasoning accuracy can be guaranteed while the reasoning speed is improved, and a better optimization effect is achieved, thereby avoiding the decrease in reasoning accuracy. Furthermore, during the model training process, the gate weights of the LSTM network module are mainly optimized, which retains the numerical accuracy of the gate bias (i.e., the gate bias remains unchanged) and maintains the abstractness of feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A schematic diagram of a long short-term memory network provided for one embodiment of the present application;

[0024] Figure 2 A flowchart of a method for training an electrocardiogram analysis model provided in one embodiment of the present application;

[0025] Figure 3 A schematic diagram of a BiLSTM calculation process provided for an embodiment of the present application;

[0026] Figure 4 A schematic diagram of a long short-term memory network module provided for one embodiment of the present application;

[0027] Figure 5 A schematic diagram of an electrocardiogram analysis model structure provided for one embodiment of the present application;

[0028] Figure 6 A schematic diagram of the change of parameters of the Gumbel-Softmax operator provided in one embodiment of the present application;

[0029] Figure 7 A schematic diagram of changes in validation loss and training loss provided for one embodiment of the present application;

[0030] Figure 8 A schematic diagram of the structure of a training device for an electrocardiogram analysis model provided in an embodiment of the present application;

[0031] Fig. 9 A structural schematic diagram of a training device for an electrocardiogram analysis model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only the parts related to the present application, rather than all structures, are shown in the accompanying drawings.

[0033] Figure 1 A schematic diagram of a long short-term memory network provided for an embodiment of the present application, referring to Figure 1 , which includes a long short-term memory network processing of three time steps. t-1 、x t 、x t+1 are the input contents of the long short-term memory network at the previous time step, the current time step, and the next time step, respectively. t-1 、h t 、h t+1 are the outputs of the long short-term memory network at the previous time step, the current time step, and the next time step, respectively. t-1 、c t 、c t+1 are the unit states of the LSTM network in the previous time step, the current time step, and the next time step, which can also be understood as the cell state. The LSTM network includes a forget gate, an input gate, and an output gate. Taking the current time step as an example, the forget gate is the output h of the previous time step (i.e., the previous layer). t-1 and the input x of the current time step (i.e. this layer) t The output is obtained through an activation function sigmoid and is recorded as f t .f t The output value of is in the range [0,1], indicating the probability that the state of the previous layer of cells is forgotten, 1 is "completely retained", and 0 is "completely discarded". The input gate consists of two parts. The first part uses the sigmoid activation function, and the output is i t The second part uses the tanh activation function and the output is g t , then, through i t With g tThe product of indicates how much new information is retained in the current input feature signal. The output gate is the output h of the previous layer. t-1 And the input x of this layer t Through an activation function sigmoid, we get a value o in the interval [0, 1]. t , then the cell state c t After being processed by the tanh activation function and o t Multiply, that is, output h of this layer t , which is used to control how much of the cell state of this layer is filtered. At this time, the calculation formula of LSTM is:

[0034] i t =σ(W ii ·x t +b ii +W hi h t-1 +b hi )

[0035] f t =σ(W if ·x t +b if +W hf h t-1 +b hf )

[0036] o t =σ(W io ·x t +b io +W ho h t-1 +b ho )

[0037] g t =tanh(W ig ·x t +b ig +W hg h t-1 +b hg )

[0038] c t =f t ⊙c t-1 +i t ⊙g t

[0039] h t =o t ⊙tanh(c t )

[0040] Among them, x trepresents the time series feature graph of the current time step input LSTM, t∈(0,h), h is the length of the hidden layer state sequence output by LSTM (i.e., the output length of LSTM), i t is the output of the LSTM input gate, W ii and b ii are the input weight and bias of the input gate, W hi and b hi are the weight and bias of the hidden layer corresponding to the input gate, f t is the output of the forget gate of LSTM, W if and b if are the input weight and bias of the forget gate, W hf and b hf are the weight and bias of the hidden layer corresponding to the forget gate, o t is the output of the output gate of LSTM, W io and b io are the input weight and bias of the output gate, W ho and b ho are the weight and bias of the hidden layer corresponding to the output gate, g t is the output of the cell state of the input gate, W ig and b ig are the input weight and bias of the unit state, W hg and b hg are the weight and bias of the hidden layer corresponding to the unit state, c t and c t-1 are the cell states of the LSTM at the current time step and the previous time step, respectively, and h t and h t-1 are the elements in the hidden state sequence of the LSTM output at the current time step and the previous time step, σ is the activation function sigmoid, and ⊙ is the element product. The weights used by each gate in the above LSTM can be considered as the gate weight of LSTM, that is, W ii , W hi , W if , W hf , W io , W ho is the gate weight.

[0041] In order to improve the reasoning speed of LSTM, a cyclic and repeatable weight matrix can be used to replace the gate weight of LSTM to reduce the storage of weights, avoid the problem of irregular memory access during LSTM reasoning due to weight storage space, and speed up the reasoning speed of LSTM. Alternatively, the gate weight of LSTM can be pruned according to the threshold and the absolute value of the weight to make the weight sparse and improve the reasoning speed of LSTM. However, the above method requires the use of specific hardware, such as the use of Field Programmable Gate Array (FPGA), which has high hardware requirements and is not conducive to promotion.

[0042] Therefore, an embodiment of the present application provides a training method, apparatus, device and storage medium for an electrocardiogram analysis model to improve the reasoning speed of LSTM and balance the reasoning accuracy of LSTM. At the same time, there are no specific requirements for hardware, which is convenient for promotion and application.

[0043] A training method for an electrocardiogram analysis model provided in an embodiment of the present application can be executed by a training device for an electrocardiogram analysis model, and the training device for the electrocardiogram analysis model can be implemented by software and / or hardware. The training device for the electrocardiogram analysis model can be composed of two or more physical entities, or can be composed of one physical entity, and the embodiment does not limit this. In one embodiment, the training device for the electrocardiogram analysis model can be an electronic device with data processing and analysis capabilities, such as a desktop computer, a laptop computer, an interactive smart tablet, a server, an electrocardiograph, an electrocardiogram monitor, etc.

[0044] For example, Figure 2 A flowchart of a method for training an electrocardiogram analysis model provided in one embodiment of the present application. Figure 2 , the training method of the ECG analysis model specifically includes:

[0045] Step 110: Obtain an ECG training set and an ECG verification set.

[0046] Exemplarily, the ECG training set refers to the training set used for training the ECG analysis model, and the ECG verification set refers to the verification set used for training the ECG analysis model. The ECG analysis model is trained using the data of the ECG training set, and the parameters involved in the ECG analysis model are fine-tuned using the data of the ECG verification set. The ECG training set and the ECG verification set contain the same data type, for example, both the ECG training set and the ECG verification set contain ECG signals of a set length and corresponding ECG analysis results, for example, both the ECG training set and the ECG verification set contain ECG features of ECG signals of a set length and corresponding ECG analysis results. Optionally, the ECG training set and the ECG verification set can be constructed based on the same data set, for example, a data set is obtained, 80% of the data is selected as the ECG training set, and the remaining 20% ​​of the data is selected as the ECG verification set, for example, the ECG training set and the ECG verification set are constructed based on a data set, and the ECG training set and the ECG verification set contain overlapping data. It can be understood that the source of the data set is not currently limited.

[0047] In one embodiment, an example is described in which both the ECG training set and the ECG verification set contain ECG signals of a set length. The ECG signal of the set length refers to an ECG signal that has been preprocessed. Preprocessing refers to obtaining an ECG signal (analog signal) directly collected from the human body, first performing impedance matching, filtering, amplification, and other processing on the ECG signal to improve the quality of the ECG signal, and then using an analog-to-digital converter to convert the analog ECG signal into a digital ECG signal. Then, the noise interference in the ECG signal is filtered out, and the ECG signal is resampled to resample the ECG signal to a target frequency, wherein the value of the target frequency can be set according to actual conditions. Afterwards, the resampled ECG signal is cut into a preset length (such as 10s) to obtain at least one ECG signal segment, and then the ECG signal segment is normalized to obtain the preprocessed ECG signal.

[0048] Step 120: Input the multiple ECG signals in the ECG training set and the multiple ECG signals in the ECG verification set into the ECG analysis model respectively to obtain multiple training output results and multiple verification output results. The ECG analysis model includes a long short-term memory network module.

[0049] Exemplarily, the ECG analysis model can realize the analysis of ECG signals. In the ECG signal abnormality detection scenario, the ECG analysis model is used to identify the abnormal type of the ECG signal. For example, in the atrial fibrillation detection scenario, after the ECG analysis model analyzes the ECG signal, it can be determined whether the ECG signal belongs to the atrial fibrillation (also known as atrial fibrillation) type or the non-atrial fibrillation type. The structure of the ECG analysis model can be set according to actual needs. In one embodiment, the ECG analysis model includes at least a feature extraction module and a long short-term memory network module. The feature extraction module is used to extract the ECG features of the ECG signal, wherein the ECG features can reflect the relevant activities of the ECG signal, and its content is related to the ECG analysis result, that is, the ECG analysis result can be obtained according to the ECG features. The structure of the feature extraction module can be set according to the actual situation. For example, the feature extraction module can adopt a neural network of the type of residual convolution network (Resnet), VGG net (Visual Geometry Group Network) or AlexNet, and the ECG features extracted by the feature extraction module can be reflected in the form of a time series feature graph. The long short-term memory network module is used to enhance the features of the ECG features so as to obtain the accuracy of the ECG analysis results later. The long short-term memory network module includes at least one LSTM or at least one bidirectional long short-term memory network (BiLSTM), wherein the BiLSTM can be decomposed into two LSTMs with different directions, for example, Figure 3 A schematic diagram of a BiLSTM calculation process provided for an embodiment of the present application. Figure 3 , the ECG feature input to BiLSTM is recorded as feature graph Fin, where feature graph Fin enters one LSTM, and after being flipped, it enters another LSTM. After that, both LSTMs output hidden layer state sequences. The hidden layer state sequence corresponding to the unflipped ECG feature graph can be recorded as {h0, h1, h2, ..., h n}, the hidden layer state sequence corresponding to the flipped ECG characteristic graph can be recorded as Afterwards, the two hidden layer state sequences are concatenated to obtain the final hidden layer state sequence Fout. The hidden layer state sequence can be considered as an enhanced feature. In one embodiment, the long short-term memory network module also includes a Gumbel-Softmax operator. Figure 4 A schematic diagram of a long short-term memory network module provided by an embodiment of the present application. Figure 4, the long short-term memory network module includes: a long short-term memory network (LSTM) and a Gumbel-Softmax operator acting on the gate weight of the long short-term memory network, or, the long short-term memory network module includes: a bidirectional long short-term memory network (BiLSTM) and a Gumbel-Softmax operator acting on the gate weight of the bidirectional long short-term memory network, and the bidirectional long short-term memory network (BiLSTM) is composed of two long short-term memory networks (LSTM) in different directions. Among them, the Gumbel-Softmax operator is used to make the gate weight of the long short-term memory network module tend to be binary. When the long short-term memory network module includes BiLSTM, the processing method of the two LSTMs is the same, and the processing direction of the LSTM is the same as when the long short-term memory network module includes LSTM. Therefore, the application of the Gumbel-Softmax operator in one LSTM is described as an example.

[0050] In one embodiment, the Gumbel-Softmax operator can be expressed as:

[0051]

[0052] Among them, G i is the i-th Gumbel-Softmax operator, π i is the sample currently processed by the Gumbel-Softmax operator (a weight value in the gate weight in the embodiment), τ is the parameter of the Gumbel-Softmax operator, π j is the jth sample processed by the Gumbel-Softmax operator, g i =-log(-log(u i )),u i represents uniformly distributed sampling on (0, 1). In the above formula, the smaller τ is, the more the output of the Gumbel-Softmax operator tends to be discrete.

[0053] In one embodiment, the Gumbel-Softmax operator is applied in LSTM, specifically, the gate weights of LSTM are processed by the Gumbel-Softmax operator to discretize the gate weights and approach binarization. It can be understood that when the gate weights of LSTM tend to be binarized, the reasoning speed of LSTM can be accelerated when the gate weights are used to process the time series feature graph. Exemplarily, the Gumbel-Softmax operator is applied to the input gate, forget gate and output gate of LSTM to make the gate weights of the three tend to be binarized to realize the gradient back propagation in the binarized gate weights. In one embodiment, when the Gumbel-Softmax operator is used to make the gate weights tend to be binarized, the calculation formula of the long short-term memory network in the long short-term memory network module is:

[0054] i t =σ(G′(W ii )·x t +b ii +G′(W hi )h t-1 +b hi )

[0055] f t =σ(G′(W if )·x t +b if +G′(W hf )h t-1 +b hf )

[0056] o t =σ(G′(W io )·x t +b io +G′(W ho )h t-1 +b ho )

[0057] g t =tanh(W ig ·x t +b ig +W hg h t-1 +b hg )

[0058] c t =f t ⊙c t-1 +i t ⊙g t

[0059] h t =o t ⊙tanh(c t )

[0060] G′(W)=G(W k,; )k∈[1,size(W)]

[0061] Among them, x t represents the time series feature graph of the current time step input to the long short-term memory network, t∈(0,h), h is the length of the hidden layer state sequence output by the long short-term memory network, i t is the output of the input gate of the LSTM network, W ii and b ii are the input weight and bias of the input gate, W hi and b hi are the weight and bias of the hidden layer corresponding to the input gate, ft is the output of the forget gate of the long short-term memory network, W if and b if are the input weight and bias of the forget gate, W hf and b hf are the weight and bias of the hidden layer corresponding to the forget gate, o t is the output of the output gate of the LSTM network, W io and b io are the input weight and bias of the output gate, W ho and b ho are the weight and bias of the hidden layer corresponding to the output gate, g t is the output of the cell state of the input gate, W ig and b ig are the input weight and bias of the cell state of the input gate, W hg and b hg are the weight and bias of the hidden layer corresponding to the unit state of the input gate, c t and c t-1 are the cell states of the LSTM network at the current time step and the previous time step, respectively, h t and h t-1 are the elements in the hidden state sequence output by the long short-term memory network at the current time step and the previous time step, respectively. G′(W) means that the weights in the first dimension of W are currently used to sample the weights in the last dimension of W using the Gumbel-Softmax operator. G(W k,; ) indicates that the weights in the last dimension of W are sampled using the Gumbel-Softmax operator using the k-th dimension weight in the first dimension of W. W is the gate weight, W∈(W ii , W hi , W if , W hf , W io , W ho ), size(W) represents the size of W (which can also be understood as the dimension of the weights contained in each dimension of W), σ is the activation function sigmoid, ⊙ is the element product, and G is the Gumbel-Softmax operator mentioned above. From the above formula, we can see that the Gumbel-Softmax operator is applied to the gate weights (W) related to the input gate. ii and W hi ), the gate weights related to the forget gate (W if and W if ), gate weights related to the output gate (W io and W ho ).

[0062] In one embodiment, the gate weights in the long short-term memory network used by the current long short-term memory network module are two-dimensional matrices, wherein the first dimension represents the weights used by the Embedding in the long short-term memory network, and the second dimension represents the weights of the time window multiplied by the Embedding in the long short-term memory network, and the time window may include the weights of each sampling point in the time series feature graph in the current time step (including W ii , W if , W io , W ig The specific weight value corresponding to the current time step) can also include the weights of each hidden layer at the current time step (including W hi , W hf , W ho , W hg The specific weight value corresponding to the current time step). By multiplying the time window by Embedding, discrete variables (such as time series feature graphs) can be converted into continuous vectors to facilitate gradient back propagation. ii For example, at this time, W in the current time step ii When the weights of the first dimension in traverse the weights of the second dimension, the Gumbel-Softmax operator is used for sampling. At this time, W ii The weight of the second dimension (i.e. the last dimension) can be considered as the π of the Gumbel-Softmax operator i , when the weights of the first dimension of each input weight used by the input gate in all time steps traverse the weights of the second dimension, the weights of the second dimension can be considered as π of the Gumbel-Softmax operator j , in order to achieve sampling using Gumbel-Softmax. When updating the gate weights, by changing the parameters of the Gumbel-Softmax operator, the updated gate weights can be discretized, that is, the effect after using Gumbel-Softmax is close to that of Embedding as One-Hot encoding, in order to achieve binarization of gate weights, where One-Hot encoding is the representation of categorical variables as binary vectors.

[0063] In one embodiment, the ECG analysis model further includes a feature recognition module, wherein the feature recognition module can be considered to recognize the features enhanced by the long short-term memory network module to output the ECG analysis result (the result is a classification result). Optionally, the feature recognition module of the ECG analysis model includes a Flatten layer, a Dense layer, and a SoftMax function. In this case, the long short-term memory network module is followed by a Flatten layer, a Dense layer, and a SoftMax function, wherein the Flatten layer is used to convert multi-dimensional input into one dimension, the Dense layer is a fully connected layer, and the SoftMax function is used to obtain the ECG analysis result.

[0064] When the above-mentioned ECG analysis model is trained based on the ECG training set and the ECG verification set, a certain number of ECG signals are selected from the ECG training set, and the ECG signals are input into the ECG analysis model. The feature extraction module of the ECG analysis model extracts the ECG features of the ECG signals, and the long short-term memory network model of the ECG analysis model enhances the ECG features. After that, after the Flatten layer, the Dense layer and the SoftMax function, the ECG analysis results are obtained. At this time, each ECG feature corresponds to an ECG analysis result. In one embodiment, the ECG analysis result corresponding to the ECG training set is recorded as the training output result. Similarly, a certain number of ECG signals are selected from the ECG verification set to obtain the corresponding ECG analysis results. In one embodiment, the ECG analysis result corresponding to the ECG verification set is recorded as the verification output result. It can be understood that the number of ECG signals selected in the ECG training set and the ECG verification set can be the same or different, and is not currently limited. In addition, the ECG signals can be selected by random selection or sequential selection. In one embodiment, ECG prior features can also be used. The feature category of the ECG verification feature is set manually. For example, when the ECG analysis model detects atrial fibrillation, the feature category of the ECG prior feature is related to the clinical manifestations of the ECG signal during the attack of atrial fibrillation. At this time, the ECG features and the ECG prior features are fused and input into the long short-term memory network module to improve the feature expression capability, thereby improving the reasoning accuracy of the ECG analysis model.

[0065] Step 130: Update the model parameters of the ECG analysis model according to the multiple training output results, wherein the model parameters include gate weights used by the long short-term memory network module.

[0066] Exemplarily, when the ECG analysis model is trained using the training data set, the model parameters of the ECG analysis model are optimized by back propagation. The back propagation process includes a chain derivation process, the purpose of which is to use the error index as the basis for adjusting the parameters of the back propagation, and then adjust the model parameters (parameters such as weights and / or biases) of the relevant layers in the ECG analysis model, so that the ECG analysis results output by the ECG analysis model conform to the distribution of the ECG signal (that is, the ECG analysis results are as close as possible to the real ECG analysis results of the ECG characteristics). In the back propagation process, a loss function is constructed according to the training output results output by the ECG analysis model. In one embodiment, the loss function constructed according to the training output results is recorded as the training loss. The training loss can reflect the error index when the ECG analysis model uses the ECG training set, that is, it reflects the difference between the training output results and the real ECG analysis results of the corresponding ECG characteristics. After that, the model parameters of the ECG analysis model are updated by the training loss to improve the reasoning accuracy of the ECG analysis model. It can be understood that when training the ECG analysis model, a training loss is constructed in each training process, and the model parameters of the ECG analysis model are updated through the training loss, so as to achieve the purpose of improving the performance of the ECG analysis model.

[0067] Exemplarily, the type of training loss can be set according to the function of the ECG analysis model. In one embodiment, when the ECG analysis model is used to classify ECG signals, such as atrial fibrillation classification and non-atrial fibrillation classification detection of ECG signals, the loss function can be a cross-entropy loss function. At this time, the category to which each training output result belongs is substituted into the cross-entropy loss function, and the specific value of the training loss in this training process can be obtained. Afterwards, the model parameters of the ECG analysis model are updated according to the specific value of the training loss. Optionally, when updating the model parameters, the gate weights used in the long short-term memory network module are mainly updated to improve the reasoning speed of the long short-term memory network module. When updating the gate weights, the model parameters in other layers can also be selectively updated. For example, the model parameters in the feature extraction module can also be updated. In one embodiment, the model parameters that need to be updated can be selected in combination with the reasoning speed and reasoning accuracy of the ECG analysis model to optimize the performance of the ECG analysis model.

[0068] Step 140: Update the parameters of the Gumbel-Softmax operator in the long short-term memory network module according to the multiple verification output results. The Gumbel-Softmax operator is used to make the gate weights tend to be binary.

[0069] Exemplarily, when the ECG analysis model is fine-tuned using the validation data set, the parameters of the Gumbel-Softmax operator in the long short-term memory network module are optimized using the 0-1 discretization module, wherein the purpose of the 0-1 discretization module is to adjust the parameters of the Gumbel-Softmax operator so that the gate weights of the long short-term memory network module tend to be binarized. In one embodiment, binarization refers to making the gate weights tend to be distributed to 0 or 1. It can be understood that from the formula in step 120, the Gumbel-Softmax operator acts on the gate weights of the long short-term memory network module. By adjusting the parameters of the Gumbel-Softmax operator, the gate weights in step 130 can be gradually distributed to 0 or 1 when updated. During the operation of the 0-1 discretization module, the validation loss is calculated based on the validation output result output by the ECG analysis model. The validation loss can reflect the error index of the ECG analysis model when using the ECG validation set. Afterwards, the parameters of the Gumbel-Softmax operator are updated by the validation loss.

[0070] In one embodiment, when the ECG analysis model is used to classify ECG signals, such as when the ECG signals are classified into atrial fibrillation and non-atrial fibrillation, the verification loss corresponding to the ECG verification set is calculated according to the verification output results. The verification loss can be determined according to the average detection rate (a relevant indicator that can reflect that the ECG signal is correctly detected) and the average specificity (a relevant indicator that can reflect that the ECG signal is actually classified as non-atrial fibrillation but is identified as atrial fibrillation) of each verification output result. Afterwards, the parameters of the Gumbel-Softmax operator are adjusted according to the verification loss.

[0071] In one embodiment, updating the parameters of the Gumbel-Softmax operator in the long short-term memory network module according to multiple verification output results includes steps 141 to 143:

[0072] Step 141: Obtain the previous historical verification loss of the ECG analysis model and the parameters of the Gumbel-Softmax operator.

[0073] Exemplarily, the validation loss obtained in the previous training process of the electrocardiogram analysis model is recorded as the historical validation loss, and the historical validation loss obtained in the previous training process and the parameters of the Gumbel-Softmax operator updated in the previous training process are currently obtained.

[0074] Step 142: Obtain the current verification loss of the ECG analysis model according to the multiple verification output results.

[0075] Exemplarily, the average detection rate and the average specificity are calculated according to the verification output results output during this training process, and then the average specificity and the average detection rate are multiplied to obtain the verification loss.

[0076] Step 143: When it is determined based on the historical verification loss that the verification loss satisfies the first update condition and it is determined that the parameter satisfies the second update condition, the parameter is multiplied by the parameter step to update the parameter.

[0077] In one embodiment, the first update condition and the second update condition are the limiting conditions required to update the parameters of the Gumbel-Softmax operator. When the first update condition and the second update condition are met, the parameters of the Gumbel-Softmax operator are updated, otherwise, the parameters of the Gumbel-Softmax operator are not updated to avoid excessive updating of the parameters of the Gumbel-Softmax operator and the impact on the performance of the ECG analysis model. The first update condition and the second update condition can be set according to actual conditions. In one embodiment, the first update condition is that the difference between the verification loss and the historical verification loss is less than the first threshold, and the second update condition is that the number of iterations of the parameter reaches the first number threshold and the parameter is greater than the second threshold.

[0078] Exemplarily, the difference between the validation loss obtained in the current time and the historical validation loss obtained in the previous time is calculated. The smaller the difference is, the more stable the validation loss of the ECG analysis model is during the two training processes. Afterwards, the difference is compared with a pre-set first threshold, wherein the first threshold can be set according to actual conditions. For example, the first threshold is set to 0.01. When the difference is less than 0.01, it is determined that the first update condition is met, and the judgment of the second update condition is continued. When the difference is not less than 0.01, it is determined that the first update condition is not met, and the parameters of the Gumbel-Softmax operator are not updated during this training process, that is, the parameters of the current Gumbel-Softmax operator remain unchanged.

[0079] When judging the second update condition, determine whether the number of iterations of the parameters of the Gumbel-Softmax operator reaches the first number and whether the parameters of the Gumbel-Softmax operator are greater than the second threshold. In one embodiment, the number of iterations can be considered as the current number of training times. After obtaining the number of iterations, the number of iterations is compared with the preset first number threshold. The first number threshold can be considered as the minimum number of critical values ​​for updating the parameters of the Gumbel-Softmax operator. The first number threshold can be set according to actual conditions. For example, the first number threshold is 2, that is, when the ECG analysis model is trained more than twice, the parameters of the Gumbel-Softmax operator can be updated. It can be understood that when the ECG analysis model is just started to be trained, the Gumbel-Softmax operator is first used to process the gate weights. After that, as the number of iterations increases, the parameters of the Gumbel-Softmax operator are updated, which can make the Gumbel-Softmax operator act more effectively on the gate function, thereby improving the training effect of the ECG analysis model. In addition to comparing the number of iterations, it is also necessary to compare whether the parameters of the Gumbel-Softmax operator are greater than the second threshold value. The second threshold value can be considered as the minimum critical value of the parameters of the Gumbel-Softmax operator. The second threshold value can be set according to the actual situation. For example, the second threshold value is 1. When the parameters of the Gumbel-Softmax operator are greater than 1, it means that the parameters of the Gumbel-Softmax operator are large, and the output of the Gumbel-Softmax operator is not discretized enough. Therefore, the parameters of the Gumbel-Softmax operator need to be updated. When the parameters of the Gumbel-Softmax operator are less than or equal to 1, it means that the parameters of the Gumbel-Softmax operator are small, and the output of the Gumbel-Softmax operator can achieve the discretization goal. Therefore, there is no need to update the parameters of the Gumbel-Softmax operator. When it is determined that the second update condition is met, the parameters of the Gumbel-Softmax operator are updated. When the second update condition is not met, the parameters of the Gumbel-Softmax operator are not updated during this training process, that is, the parameters of the current Gumbel-Softmax operator remain unchanged.

[0080] Exemplarily, when updating the parameters of the Gumbel-Softmax operator, the parameters of the Gumbel-Softmax operator may be reduced to ensure the discretization of the output of the Gumbel-Softmax operator. In one embodiment, the parameters of the Gumbel-Softmax operator are updated by a pre-set parameter step. Currently, the parameter step can be understood as the magnitude of the change in the parameter each time the parameters of the Gumbel-Softmax operator are updated. The parameter step is a decimal in the range of (0,1), and its specific value can be set according to actual conditions. The result obtained by multiplying the parameters of the Gumbel-Softmax operator by the currently used parameter step is used as the updated parameters of the Gumbel-Softmax operator. For example, the parameter step is 0.8, and the parameters of the Gumbel-Softmax operator are multiplied by 0.8 as the updated parameters of the Gumbel-Softmax operator.

[0081] It should be noted that during each training, the first update condition and the second update condition need to be judged to determine whether the parameters of the Gumbel-Softmax operator need to be updated.

[0082] Step 150 : repeatedly inputting the multiple ECG signals in the ECG training set and the multiple ECG signals in the ECG verification set into the ECG analysis model, until the ECG analysis model meets the training stop condition.

[0083] It should be noted that after step 140 is completed, it can be considered that a training process is completed. In one embodiment, a training process can be recorded as an epoch. After the end of this training process, the number of training times corresponding to the epoch is increased by 1. Afterwards, it is determined whether the training stop condition is met. If the training stop condition is not met, the process returns to step 120, that is, multiple ECG signals are selected again in the ECG training set and multiple ECG signals are selected again in the ECG verification set, and are respectively input into the ECG analysis model to train the ECG analysis model according to the training output results and the verification output results until the training stop condition is met. If the training stop condition is met, the training of the ECG analysis model is stopped. Among them, the training stop condition is a critical condition for stopping the training of the ECG analysis model. The training stop condition can be set according to actual conditions. For example, the training stop condition is that the training loss obtained based on the training output results in the continuous training times is stable and the verification loss meets the expected numerical range. For another example, the training stop condition is that the number of training times of the ECG analysis model exceeds the second number threshold. In one embodiment, the training stop condition is described as the number of training times of the ECG analysis model exceeding the second number threshold. The second number threshold is the maximum critical value of the number of training times, and its specific value can be set according to the actual situation. After each training process is completed, the number of training times is increased by 1, and then the number of training times is compared with the second number threshold. If it exceeds the second number threshold, it is determined that the training stop condition is met, and the training of the ECG analysis model is stopped. If it does not exceed the second number threshold, it is determined that the training stop condition is not met, and a new training is started. For example, the second number threshold is 35. At this time, the ECG analysis model must be trained at least 35 times before the training of the ECG analysis model is stopped.

[0084] Optionally, after the training is completed, the ECG analysis model can be tested to determine the performance of the ECG analysis model, and then the ECG analysis model can be applied.

[0085] In the above, multiple ECG signals in the ECG training set and the ECG verification set are respectively input into the ECG analysis model, and then the gate weights of the long short-term memory network module in the ECG analysis model are updated according to the training output results corresponding to the ECG training set, and the model parameters of other layers are optionally updated, and the parameters of the Gumbel-Softmax operator used in the long short-term memory network module are updated according to the verification output results corresponding to the ECG verification set. The Gumbel-Softmax operator acts on the gate weights of the long short-term memory network module so that the gate weights tend to be binarized after the update, which solves the technical problem that the reasoning speed of LSTM cannot be improved in the related technology. By using the Gumbel-Softmax operator in the long short-term memory network module, and by updating the parameters and gate weights of the Gumbel-Softmax operator during the training process to realize the binarization of the gate weights, the reasoning speed of LSTM can be improved, and by updating the parameters of the Gumbel-Softmax operator through the verification output results, the reasoning accuracy can be guaranteed while the reasoning speed is improved, and a better optimization effect is achieved, and the decrease in reasoning accuracy is avoided. Furthermore, during the model training process, the gate weights of the long short-term memory network module are mainly optimized, the numerical accuracy of the gate bias is retained (i.e., the gate bias remains unchanged), and the abstractness of feature extraction is maintained. Furthermore, during the training process, by setting the first update condition and the second update condition as well as the training stop condition, and by introducing the verification loss and parameter stepping, an adaptive binary optimization strategy can be implemented. Under this strategy, the training process does not require manual intervention, and the ECG analysis model automatically iterates the training to obtain the optimal result.

[0086] In one embodiment of the present application, in the above process, when the gate weight is binarized by the Gumbel-Softmax operator, in order to further improve the inference speed of the ECG analysis model, the gate weight can be set to a binarized weight and the Gumbel-Softmax operator can be deleted. At this time, after step 150, steps 160 to 180 are also included:

[0087] Step 160: In the last dimension of the gate weight, retain the first value that is greater than or equal to the preset threshold.

[0088] Since the Gumbel-Softmax operator samples the last dimension of the gate weight, the last dimension of the gate weight after the Gumbel-Softmax operator tends to be binarized, and the larger the value of the gate weight (i.e., the specific weight value in the gate weight), the closer its specific value is to 1, and vice versa. If you want to make the gate weight binary, you need to determine which values ​​in the gate weight become 1 and which values ​​become 0. In one embodiment, a preset threshold is set, and among the weight values ​​of the last dimension of the gate weight, a value greater than or equal to the preset threshold is retained, and the value is recorded as the first value. It can be understood that when binarized, the first value corresponds to 1, and other values ​​in the gate weight except the first value (i.e., values ​​less than the preset threshold) correspond to 0. The preset threshold can be set according to actual conditions. For example, the preset threshold is currently set to 0.5. At this time, in the last dimension of the gate weight, each first value greater than or equal to 0.5 is retained, and values ​​less than 0.5 are eliminated. After this step, it can be considered that the specific values ​​in the gate weight are truncated using the preset threshold, so that the gate weight becomes discretized.

[0089] Step 170, establish a binary weight tensor of the same size as the gate weight, the first numerical position in the binary weight tensor is disposed as the second numerical value, and the remaining positions in the binary weight tensor are disposed as the third numerical value, and the second numerical value is greater than the third numerical value.

[0090] Exemplarily, a new binary weight tensor is created, and the size of the binary weight tensor is consistent with the size of the gate weight, that is, the dimension, number of row vectors and number of column vectors of the binary weight tensor are the same as the gate weight. At this time, the dimension of the binary weight tensor is also 2. In one embodiment, the binary weight tensor contains two values, which are respectively recorded as the second value and the third value, wherein the second value is greater than the third value, and the second value and the third value can be considered as two values ​​after binarization. In one embodiment, the second value is 1 and the third value is 0. At this time, the binary weight tensor is a 0-1 weight tensor, that is, each value contained in the binary weight tensor is 0 or 1. Exemplarily, since in step 160, only the first value greater than or equal to the preset threshold is retained in the gate weight, and the first value corresponds to 1 when binarized, therefore, after creating a new two-dimensional weight tensor, the specific position of each first value in the gate weight is determined, and 1 is filled in the same position of the two-dimensional weight tensor (that is, the position is set to 1), and then 0 is filled in other positions of the two-dimensional weight tensor (that is, other positions are set to 0). At this point, the binary weight tensor can be expressed as:

[0091]

[0092]

[0093] Among them, W' represents the binary weight tensor, is the value of the last dimension of the gate weight acted on by the Gumbel-Softmax operator in step 160, and W is the gate weight. Indicates that W' is a 0-1 binary weight tensor, The specific value of The last dimension of is determined by truncating the result by a preset threshold, where The value greater than or equal to α is set to 1, and the values ​​in other cases are all 0, and α is set to 0.5 here. It can be understood that when there are multiple long short-term memory networks in the long short-term memory network module, each long short-term memory network corresponds to a binary weight tensor.

[0094] Step 180: Load the binary weight tensor into the long short-term memory network module as a new gate weight, and delete the Gumbel-Softmax operator in the long short-term memory network module.

[0095] Exemplarily, the binary weight tensor is used as a new gate weight and loaded into the corresponding LSTM network in the LSTM network module to replace the gate weight in the LSTM network. It can be understood that since the new gate weight in the LSTM network is already a binary weight, there is no need to use the Gumbel-Softmax operator to process the gate weight. Therefore, when the new gate weight is loaded, the Gumbel-Softmax operator in the LSTM network module is deleted, that is, the gate weight is deleted. Figure 4 After that, the LSTM module only has the LSTM network. In the subsequent application process, the calculation formulas of the input gate, forget gate and output gate of the LSTM network are changed back to the following formulas:

[0096] i t =σ(W ii ·x t +b ii +W hi h t-1 +b hi )

[0097] f t =σ(W if ·x t +b if +W hf h t-1 +b hf )

[0098] o t =σ(W io ·x t +b io +W ho h t-1 +b ho )

[0099] It can be understood that the meaning of each symbol in the above formula is the same as the meaning of the symbols involved in the aforementioned LSTM calculation formula, and will not be elaborated on at present.

[0100] It can be understood that steps 160 to 180 can be considered as the inference processing of the ECG analysis model. After the inference is initialized, the ECG analysis model can be applied for inference.

[0101] As mentioned above, after using the Gumbel-Softmax operator to change the gate weights in the long short-term memory network into a binary weight tensor, the reasoning speed of the ECG analysis model is accelerated, and the phenomenon of serious decrease in the reasoning accuracy of the ECG analysis model after binarization is avoided. In addition, the positions of 0 and 1 in the binary weight tensor are determined by the Gumbel-Softmax operator, which ensures the reasoning accuracy of the ECG analysis model and better balances the reasoning speed and reasoning accuracy. In addition, only the long short-term memory network is binarized, which maintains the data structure of the input and output features of the long short-term memory network, avoiding changes in the structure of the ECG analysis model after optimization. At the same time, the ECG analysis model can continue to be optimized by pruning, quantization, etc., and the optimization process will not be affected by the current binarization.

[0102] The following is an exemplary description of the ECG analysis model. Figure 5 A schematic diagram of an electrocardiogram analysis model structure provided by an embodiment of the present application. Figure 5 The ECG analysis model adopts an end-to-end LSTM+CNN (convolutional neural network) network structure, where CNN is a feature extraction module. CNN includes modules composed of IDConvolution, Batch Normalization, ReLU function and MaxPool, and also includes 4 residual convolution networks ( Figure 5denoted as Residual block), where a residual convolution network includes: IDConvolution (ID convolution layer), Batch Normalization (batch normalization layer), ReLU function, IDConvolution (ID convolution layer), Batch Normalization (batch normalization layer) and Downsample (downsampling layer). The four residual convolution networks are followed by a long short-term memory network module, where the long short-term memory network module includes a BiLSTM layer, and the long short-term memory network module is followed by a Flatten layer, a Dense layer and a SoftMax function. CNN can extract ECG features. Afterwards, the BiLSTM layer can achieve feature enhancement, and then it is beneficial for the subsequent Flatten layer, the Dense layer and the SoftMax function to output the ECG analysis results. Optionally, the number of neurons in the BiLSTM layer is 32, and the number of neurons in the Dense layer is 64. In the residual convolution network, the convolution kernel size and channels of IDConvolution are set to 11 and 64 respectively. In the module composed of IDConvolution, Batch Normalization, ReLU function and MaxPool, the convolution kernel size and channels of IDConvolution are set to 11 and 128 respectively, and the size of MaxPool is 2. Dropout = 0.2, where Dropout is a simple method to prevent overfitting of the neural network (currently the ECG analysis model).

[0103] When the above training method is used to train the ECG analysis model, Figure 6 A schematic diagram of the change of parameters of the Gumbel-Softmax operator provided for one embodiment of the present application. Figure 6 The horizontal axis represents the number of training times, and the vertical axis represents the specific value of the parameter. Figure 7 A schematic diagram of the changes in validation loss and training loss provided for one embodiment of the present application. Figure 7 The horizontal axis represents the number of training times, and the vertical axis represents the specific values ​​of the validation loss and training loss. Figure 6 and Figure 7 It can be seen that as the number of training times increases, the parameters of the Gumbel-Softmax operator tend to 0. Moreover, as the parameters of the Gumbel-Softmax operator decrease, the training loss does not oscillate and the verification loss tends to converge, that is, the inference accuracy of the ECG analysis model is not affected.

[0104] Table 1 Figure 5The performance indicators of sensitivity and specificity of the ECG validation set and the ECG test set when adjusting the model parameters of different layers during the training process of the ECG analysis model shown. The ECG test set represents the data set used when testing the ECG distraction model after the ECG analysis model is trained. The ECG signals contained in the ECG test set are of the same type as those contained in the ECG training set and the ECG validation set. The ECG test set is used to test the accuracy of the ECG analysis model.

[0105]

[0106]

[0107] Table 1

[0108] Among them, τ is the parameter of the Gumbel-Softmax operator, SE is sensitivity, SP is specificity, and Rb2-Rb4 respectively represent the second residual convolution network to the fourth residual convolution network of the four residual convolution networks included in CNN. According to the parameters of SE and SP in the ECG test set, when the model parameters of Rb4+BiLSTM+Dense are adjusted, the model performance of the ECG analysis model is optimal. It can be understood that when adjusting the model parameters of Rb4+BiLSTM+Dense, it can be avoided that when only adjusting the gate weight of BiLSTM, it is not enough to compensate for the loss of reasoning accuracy caused by binarization of the gate weight. When adjusting the model parameters of too many layers, the distribution of the ECG analysis model will change, making the ECG analysis model unable to adapt to the distribution of the ECG test set.

[0109] Table 2 Figure 5 During the training process of the ECG analysis model shown in Figure 2, when the gate weights of different gates of BiLSTM are binarized, the inference time of the ECG analysis model is shown. Table 2 shows the inference time of the ECG analysis model when it is running on the Kirin 990 CPU of the embedded platform.

[0110] Binarized gate weights Inference time (ms) Baseline 108.024447±7.638560 Input gate, forget gate, output gate 92.737328±6.047918 Input gate, forget gate 95.824341±6.396740 Forget Gate 99.624357±11.957540

[0111] Table 2

[0112] Among them, Baseline is the inference time of the ECG analysis model when the gate weights are not binarized. It can be seen from Table 2 that the inference time of the ECG analysis model is the shortest when the gate weights of the input gate, forget gate and output gate are binarized. Compared with the Baseline, the inference speed can be relatively increased by 14.15% when the gate weights of the input gate, forget gate and output gate are binarized, which effectively improves the inference speed of the ECG analysis model.

[0113] Understandable, Figure 5After the training of the ECG analysis model shown is completed, that is, the gate weight of the long short-term memory network module in the ECG analysis model is 0-1 weight and the Gumbel-Softmax operator has been deleted, the ECG analysis model can be used to classify ECG signals. In one embodiment, the ECG analysis model is used to detect whether the ECG signal is an atrial fibrillation type ECG signal as an example for description. At this time, the types of ECG signals classified by the ECG analysis model include atrial fibrillation type and non-atrial fibrillation type. At this time, the application process of the ECG analysis model may include the following steps 210-220:

[0114] Step 210: Obtain the ECG signal to be classified and the ECG priori features.

[0115] Exemplarily, the ECG signal to be classified is the ECG signal to be processed currently, which has the same length as the ECG signals included in the ECG training set and the ECG verification set, and both are pre-processed ECG signals.

[0116] The feature category of the ECG verification feature is set manually. For example, when the ECG analysis model detects atrial fibrillation, the feature category of the ECG prior feature is related to the clinical manifestation of the ECG signal during the onset of atrial fibrillation. For example, for atrial fibrillation, the irregular RR interval of the ECG signal is a more significant clinical manifestation. Therefore, the standard deviation of the RR interval can be used as a feature category of the ECG prior feature to reflect whether the RR interval is regular through the standard deviation of the RR interval. In one embodiment, the ECG prior feature is constructed using the ECG signal in the public data set. For example, the ECG signal in the public data set is obtained, and each RR interval in the ECG signal is determined. After that, the standard deviation of the RR interval is calculated to obtain the ECG prior feature. It can be understood that in addition to the standard deviation of the RR interval, other ECG prior features can be constructed in combination with the clinical manifestations of atrial fibrillation to enrich the feature category of the ECG prior feature. For example, the ECG prior feature can also include the coefficient of variation of the RR interval, the standard deviation of the variability of the PR interval, and the standard deviation of the variability of the P wave. The coefficient of variation of the RR interval refers to the ratio of the standard deviation of the RR interval to the average value of the RR interval. The standard deviation of the variability of the PR interval can evaluate whether the PR interval is regular, wherein the PR interval refers to the time limit (time length) between the P wave in the electrocardiogram signal and the starting point of the adjacent QRS wave. The ratio of a PR interval to the average PR interval of the electrocardiogram signal can be used as the PR interval variability of the PR interval, and the standard deviation of the PR interval variability can be obtained through the variability of each PR interval. P wave variability refers to the ratio of the width of the current P wave to the average value of the width of each P wave in the electrocardiogram signal. The P wave width refers to the time length of the P wave in the electrocardiogram signal, and the standard deviation of the P wave variability can be obtained based on the P wave variability of each P wave in the electrocardiogram signal. Optionally, after the electrocardiogram prior feature is constructed, the electrocardiogram prior feature is first normalized for subsequent use, wherein the implementation method of normalization is not limited, for example, min-max normalization is used so that the value of the electrocardiogram prior feature is mapped to between [0-1] for subsequent use.

[0117] Step 220: Input the ECG signal to be classified and the ECG priori feature into the ECG analysis model, and the ECG analysis model outputs the classification result of the ECG signal to be classified.

[0118] In one embodiment, after the ECG signal to be classified and the ECG prior features are input into the ECG analysis model, the ECG analysis model can determine the classification result of the ECG signal to be classified in combination with the ECG prior features, that is, whether the ECG signal to be classified is of the atrial fibrillation type or the non-atrial fibrillation type. Figure 5Taking the ECG analysis model shown in the figure as an example, the ECG analysis model processing process is as follows: the ECG features of the ECG signal to be classified are extracted by the feature extraction module. After that, before entering the long short-term memory network module, the ECG features and the ECG prior features are fused to obtain a fusion feature with stronger discrimination ability. After that, the fusion features are enhanced by the long short-term memory network module, and then the category to which the ECG signal to be classified belongs is obtained through the Flatten layer, the Dense layer and the SoftMax function.

[0119] It can be understood that since the gate weights of the long short-term memory network module have been binarized, the reasoning speed of the ECG analysis model is improved and the reasoning accuracy is not reduced when classifying the ECG signals to be classified, which effectively balances the reasoning speed and reasoning accuracy.

[0120] Figure 8 A schematic diagram of a training device for an electrocardiogram analysis model provided in one embodiment of the present application, referring to Figure 8 The training device of the electrocardiogram analysis model includes a data set acquisition module 301, a model reasoning module 302, a parameter updating module 303, an operator updating module 304 and a repeated training module 305.

[0121] Among them, the data set acquisition module 301 is used to obtain the ECG training set and the ECG verification set; the model inference module 302 is used to input multiple ECG signals in the ECG training set and multiple ECG signals in the ECG verification set into the ECG analysis model respectively to obtain multiple training output results and multiple verification output results, and the ECG analysis model includes a long short-term memory network module; the parameter update module 303 is used to update the model parameters of the ECG analysis model according to the multiple training output results, and the model parameters include the gate weights used by the long short-term memory network module; the operator update module 304 is used to update the parameters of the Gumbel-Softmax operator in the long short-term memory network module according to the multiple verification output results, and the Gumbel-Softmax operator is used to make the gate weights tend to be binary; the repeated training module 305 is used to repeatedly execute the operation of inputting multiple ECG signals in the ECG training set and multiple ECG signals in the ECG verification set into the ECG analysis model respectively until the ECG analysis model meets the training stop condition.

[0122] In one embodiment of the present application, it also includes: a stage module, which is used to retain a first value greater than or equal to a preset threshold in the last dimension of the gate weight after the electrocardiogram analysis model meets the training stop condition; a tensor establishment module, which is used to establish a binary weight tensor with the same size as the gate weight, the first value position in the binary weight tensor is disposed as the second value, and the remaining positions in the binary weight tensor are disposed as the third value, and the second value is greater than the third value; a weight loading module, which is used to load the binary weight tensor as a new gate weight in the long short-term memory network module, and delete the Gumbel-Softmax operator in the long short-term memory network module.

[0123] In one embodiment of the present application, the long short-term memory network module includes: a long short-term memory network and a Gumbel-Softmax operator acting on the gate weights of the long short-term memory network, or the long short-term memory network module includes: a bidirectional long short-term memory network and a Gumbel-Softmax operator acting on the gate weights of the bidirectional long short-term memory network, and the bidirectional long short-term memory network is composed of two long short-term memory networks in different directions.

[0124] In one embodiment of the present application, when the Gumbel-Softmax operator is used to make the gate weights tend to be binary, the calculation formula of the long short-term memory network in the long short-term memory network module is:

[0125] i t =σ(G′(W ii )·x t +b ii +G′(W hi )h t-1 +b hi )

[0126] f t =σ(G′(W if )·x t +b if +G′(W hf )h t-1 +b hf )

[0127] o t =σ(G′(W io )·x t +b io +G′(W ho )h t-1 +b ho )

[0128] g t =tanh(W ig ·x t +b ig +W hgh t-1 +b hg )

[0129] c t =f t ⊙c t-1 +i t ⊙g t

[0130] h t =o t ⊙tanh(c t )

[0131] G'(W)=G(W k,; )k∈[1,size(W)]

[0132] Among them, x t Represents the time series feature graph of the current time step input to the long short-term memory network, t∈(0,h), h is the length of the hidden layer state sequence output by the long short-term memory network, i t is the output of the input gate of the LSTM network, W ii and b ii are the input weight and bias of the input gate, W hi and b hi are the weight and bias of the hidden layer corresponding to the input gate, f t is the output of the forget gate of the long short-term memory network, W if and b if are the input weight and bias of the forget gate, W hf and b hf are the weight and bias of the hidden layer corresponding to the forget gate, o t is the output of the output gate of the LSTM network, W io and b io are the input weight and bias of the output gate, W ho and b ho are the weight and bias of the hidden layer corresponding to the output gate, g t is the output of the cell state of the input gate, W ig and b ig are the input weight and bias of the cell state of the input gate, W hg and b hg are the weight and bias of the hidden layer corresponding to the unit state of the input gate, c t and c t-1 are the cell states of the LSTM network at the current time step and the previous time step, respectively, h t and h t-1are the elements in the hidden state sequence output by the long short-term memory network at the current time step and the previous time step, respectively. G'(W) means that the weights in the first dimension of W are currently used to sample the weights in the last dimension of W using the Gumbel-Softmax operator. G(W k,; ) indicates that the weight of the kth dimension in the first dimension of W is currently used to sample the weights in the last dimension of W using the Gumbel-Softmax operator, W is the gate weight, W∈(W ii ,W hi ,W if ,W hf ,W io ,W ho ), size(W) represents the size of W, σ is the activation function sigmoid, and ⊙ is the element-wise product.

[0133] In one embodiment of the present application, the operator update module 304 includes: a loss acquisition unit, used to obtain the previous historical verification loss of the ECG analysis model and the parameters of the Gumbel-Softmax operator, the Gumbel-Softmax operator is used to make the gate weights tend to be binary; a loss calculation unit, used to obtain the current verification loss of the ECG analysis model according to multiple verification output results; an operator parameter update unit, used to determine that the verification loss meets the first update condition based on the historical verification loss, and when it is determined that the parameter meets the second update condition, multiply the parameter by the parameter step to update the parameter.

[0134] In one embodiment of the present application, the first update condition is that the difference between the verification loss and the historical verification loss is less than a first threshold, and the second update condition is that the number of iterations of the parameter reaches a first number threshold and the parameter is greater than a second threshold.

[0135] In one embodiment of the present application, the training stop condition is that the number of training times of the electrocardiogram analysis model exceeds a second number threshold.

[0136] In one embodiment of the present application, an ECG analysis model is used to classify ECG signals. During the application of the ECG analysis model, it also includes: a prior acquisition module, used to obtain the ECG signals to be classified and ECG prior features; a classification module, used to input the ECG signals to be classified and the ECG prior features into the ECG analysis model, and the ECG analysis model outputs the classification results of the ECG signals to be classified.

[0137] The training device for the electrocardiogram analysis model provided above can be used to execute the training method for the electrocardiogram analysis model provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0138] It is worth noting that in the embodiment of the training device for the above-mentioned electrocardiogram analysis model, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application.

[0139] Fig. 9 A schematic diagram of a training device for an electrocardiogram analysis model provided in one embodiment of the present application. Fig. 9 As shown, the training device for the electrocardiogram analysis model includes a processor 40, a memory 41, an input device 42, and an output device 43; the number of the processors 40 in the training device for the electrocardiogram analysis model can be one or more. Fig. 9 A processor 40 is taken as an example. In the training device for the electrocardiogram analysis model, the processor 40, the memory 41, the input device 42, and the output device 43 can be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.

[0140] The memory 41, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the training method of the electrocardiogram analysis model in one embodiment of the present application (for example, a data set acquisition module, a model reasoning module, a parameter update module, an operator update module and a repeated training module in the training device of the electrocardiogram analysis model). The processor 40 executes various functional applications and data processing of the training device of the electrocardiogram analysis model by running the software programs, instructions and modules stored in the memory 41, that is, implements the above-mentioned training method of the electrocardiogram analysis model.

[0141] The memory 41 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the training device of the electrocardiographic analysis model, etc. In addition, the memory 41 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 41 may further include a memory remotely arranged relative to the processor 40, and these remote memories may be connected to the training device of the electrocardiographic analysis model via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0142] The input device 42 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the training device of the ECG analysis model, and can also include devices required by each lead when collecting ECG. The output device 43 can include display devices such as a display screen.

[0143] The above-mentioned training equipment for the ECG analysis model includes a training device for the ECG analysis model, which can be used to execute a training method for any ECG analysis model and has corresponding functions and beneficial effects.

[0144] In addition, an embodiment of the present application also provides a storage medium containing computer executable instructions, which, when executed by a computer processor, are used to perform relevant operations in the training method of the electrocardiogram analysis model provided in any embodiment of the present application, and have corresponding functions and beneficial effects.

[0145] Those skilled in the art should understand that the embodiments of the present application may be provided as methods, systems, or computer program products.

[0146] Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0147] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash memory (flashRAM). The memory is an example of a computer-readable medium.

[0148] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0149] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0150] Note that the above are only preferred embodiments of the present application and the technical principles used. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.

Claims

1. A training method for an electrocardiogram analysis model, characterized in that: include: Get ECG training set and ECG validation set; Inputting a plurality of ECG signals in an ECG training set and a plurality of ECG signals in an ECG verification set into an ECG analysis model respectively to obtain a plurality of training output results and a plurality of verification output results, wherein the ECG analysis model includes a long short-term memory network module; updating model parameters of the electrocardiogram analysis model according to a plurality of training output results, the model parameters including gate weights used by a long short-term memory network module; Updating the parameters of the Gumbel-Softmax operator in the long short-term memory network module according to the multiple verification output results, wherein the Gumbel-Softmax operator is used to make the gate weight tend to be binary; Repeating the operation of inputting a plurality of ECG signals in an ECG training set and a plurality of ECG signals in an ECG verification set into an ECG analysis model respectively until the ECG analysis model meets a training stop condition; When the Gumbel-Softmax operator is used to make the gate weights tend to be binary, the calculation formula of the long short-term memory network in the long short-term memory network module is: t =σ(G'(W ii )·x t +b ii +G'(W hi )h t-1 +b hi ), f t =σ(G'(W if )·x t +b if +G'(W hf )h t-1 +b hf ), o t =σ(G'(W io )·x t +b io +G'(W ho )h t-1 +b ho ), g t =tanh(W ig ·x t +b ig +W hg h t-1 +b hg ),c t =f t ⊙c t-1 +i t ⊙g t 、h t =o t ⊙tanh(c t )、G'(W)=G(W k,; )k∈[1,size(W)]; where x t is the time series feature graph of the input network at the current time step, t∈(0,h), h is the length of the hidden layer state sequence output by the network, W ii , b ii , W hi and b hi is the input weight and bias of the input gate, the corresponding hidden layer weight and hidden layer bias, W if , b if , W hf and b hf is the input weight and bias of the forget gate, the corresponding hidden layer weight and hidden layer bias, W io , b io , W ho and b ho is the input weight and bias of the output gate, the corresponding hidden layer weight and hidden layer bias, W ig 、b ig , W hg and b hg is the input weight and bias of the unit state of the input gate, the weight of the hidden layer corresponding to the unit state, and the bias of the hidden layer, c t and c t-1 is the unit state of the network at the current and previous time steps, h t and h t-1 is the element in the hidden state sequence output by the network at the current and previous time steps, G'(W) is the weight in the first dimension of W currently used to sample the weights in the last dimension of W using the Gumbel-Softmax operator, G(W k,; ) is the k-th dimension weight in the first dimension of W currently used. The weights in the last dimension of W are sampled using the Gumbel-Softmax operator. W is the gate weight, W∈(W ii ,W hi ,W if ,W hf ,W io ,W ho ), size(W) is the size of W, σ is the activation function sigmoid, and ⊙ is the element-wise product.

2. The training method according to claim 1, characterized in that: After the electrocardiogram analysis model meets the training stop condition, the method further includes: In the last dimension of the gate weight, retaining a first value greater than or equal to a preset threshold; Establishing a binary weight tensor of the same size as the gate weight, wherein the first value position in the binary weight tensor is disposed as a second value, and the remaining positions in the binary weight tensor are disposed as a third value, and the second value is greater than the third value; The binary weight tensor is loaded into the long short-term memory network module as a new gate weight, and the Gumbel-Softmax operator in the long short-term memory network module is deleted.

3. The training method according to claim 1, characterized in that: The long short-term memory network module includes: a long short-term memory network and a Gumbel-Softmax operator acting on the gate weights of the long short-term memory network, or the long short-term memory network module includes: a bidirectional long short-term memory network and a Gumbel-Softmax operator acting on the gate weights of the bidirectional long short-term memory network, and the bidirectional long short-term memory network is composed of two long short-term memory networks in different directions.

4. The training method according to claim 1, characterized in that The updating of the parameters of the Gumbel-Softmax operator in the long short-term memory network module according to the plurality of verification output results comprises: Obtain the previous historical verification loss of the electrocardiogram analysis model and the parameters of the Gumbel-Softmax operator; Obtaining the current verification loss of the electrocardiogram analysis model according to the plurality of verification output results; When it is determined according to the historical verification loss that the verification loss satisfies a first update condition and when it is determined that the parameter satisfies a second update condition, the parameter is multiplied by a parameter step to update the parameter.

5. The training method according to claim 4, characterized in that: The first update condition is that the difference between the verification loss and the historical verification loss is less than a first threshold, and the second update condition is that the number of iterations of the parameter reaches a first number threshold and the parameter is greater than a second threshold.

6. The training method according to claim 1, characterized in that: The training stop condition is that the number of training times of the electrocardiogram analysis model exceeds a second number threshold.

7. The training method according to claim 2, characterized in that: The electrocardiogram analysis model is used to classify electrocardiogram signals, and the training method further includes: Obtaining the ECG signal to be classified and ECG prior features; The ECG signal to be classified and the ECG priori feature are input into the ECG analysis model, and the ECG analysis model outputs the classification result of the ECG signal to be classified.

8. A training device for an electrocardiogram analysis model, characterized in that: include: A data set acquisition module is used to obtain an ECG training set and an ECG verification set; A model inference module, used for inputting a plurality of ECG signals in an ECG training set and a plurality of ECG signals in an ECG verification set into an ECG analysis model respectively, to obtain a plurality of training output results and a plurality of verification output results, wherein the ECG analysis model includes a long short-term memory network module; A parameter updating module, used for updating the model parameters of the ECG analysis model according to a plurality of training output results, the model parameters including gate weights used by the long short-term memory network module; An operator updating module, used for updating the parameters of the Gumbel-Softmax operator in the long short-term memory network module according to a plurality of verification output results, wherein the Gumbel-Softmax operator is used for making the gate weight tend to be binary; A repeated training module, used for repeatedly executing the operation of inputting a plurality of ECG signals in an ECG training set and a plurality of ECG signals in an ECG verification set into an ECG analysis model respectively, until the ECG analysis model satisfies a training stop condition; When the Gumbel-Softmax operator is used to make the gate weights tend to be binary, the calculation formula of the long short-term memory network in the long short-term memory network module is: t =σ(G'(W ii )·x t +b ii +G'(W hi )h t-1 +b hi ), f t =σ(G'(W if )·x t +b if +G'(W hf )h t-1 +b hf ), o t =σ(G'(W io )·x t +b io +G'(W ho )h t-1 +b ho ), g t =tanh(W ig ·x t +b ig +W hg h t-1 +b hg ),c t =f t ⊙c t-1 +i t ⊙g t 、h t =o t ⊙tanh(c t )、G'(W)=G(W k,; )k∈[1,size(W)]; where x t is the time series feature graph of the input network at the current time step, t∈(0,h), h is the length of the hidden layer state sequence output by the network, W ii , b ii , W hi and b hi is the input weight and bias of the input gate, the corresponding hidden layer weight and hidden layer bias, W if , b if , W hf and b hf is the input weight and bias of the forget gate, the corresponding hidden layer weight and hidden layer bias, W io , b io , W ho and b ho is the input weight and bias of the output gate, the corresponding hidden layer weight and hidden layer bias, W ig 、b ig , W hg and b hg is the input weight and bias of the unit state of the input gate, the weight of the hidden layer corresponding to the unit state, and the bias of the hidden layer, c t and c t-1 is the unit state of the network at the current and previous time steps, h t and h t-1 is the element in the hidden state sequence output by the network at the current and previous time steps, G'(W) is the weight in the first dimension of W currently used to sample the weights in the last dimension of W using the Gumbel-Softmax operator, G(W k,; ) is the k-th dimension weight in the first dimension of W currently used. The weights in the last dimension of W are sampled using the Gumbel-Softmax operator. W is the gate weight, W∈(W ii ,W hi ,W if ,W hf ,W io ,W ho ), size(W) is the size of W, σ is the activation function sigmoid, and ⊙ is the element-wise product.

9. A training device for an electrocardiogram analysis model, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the electrocardiogram analysis model as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the training method of the electrocardiogram analysis model as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Multi-task cascade neural network ECG signal arrhythmia classification model and method

    CN110638430A