An Electronic Medical Record Multi-Label Text Classification Method Based on Longformer
Through Longformer pre-trained model and multi-filter residual convolutional neural network, combined with the recalibration aggregation module and attention mechanism, the problem of long sentence feature extraction and noise in electronic medical record text classification is solved, and the classification performance is improved.
Patent Information
- Application Number
- CN202311073044.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-08-24
AI Technical Summary
The prior art is difficult to effectively extract text features of long sentences in electronic medical record text classification. Different sentence lengths lead to poor feature extraction results and a large amount of noise information, resulting in low classification performance.
The Longformer pre-trained model is used to combine the multi-filter residual convolutional neural network and the recalibration aggregation module to extract text features through convolution kernels of different sizes, add residual convolutional neural network to solve the problem of gradient disappearance and explosion, and use attention mechanism and feedforward neural network for classification.
Improves the accuracy and robustness of text classification of electronic medical records, can effectively process long texts and reduce noise impacts, adapting to large-scale and unbalanced datasets.
Smart Images

Figure CN117112787B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text classification, and particularly relates to a multi-label text classification method for electronic medical records based on Longformer. Background Art
[0002] With the popularization of computer technology and electronic medical records, people have obtained a huge amount of electronic medical record data. For the data analysis of electronic medical records, text classification of electronic medical records is an important part. The currently commonly used models have certain limitations in extracting text features of electronic medical records, which are mainly reflected in the following aspects: (1) The text length of electronic medical records is relatively long, and the model cannot accurately extract the text features of long sentences. (2) The lengths of sentences in electronic medical record texts vary, and the effect of text feature extraction is not good. (3) There are a lot of associated information and a large amount of noise information in electronic medical records, resulting in a low overall classification performance. Summary of the Invention
[0003] In order to solve the above existing problems, the present invention proposes: a multi-label text classification method for electronic medical records based on Longformer, including the following steps:
[0004] S1. The text of the electronic medical record is converted into word vectors through the Longformer pre-trained model, and the word vectors are used as the input of the multi-filter residual convolutional neural network, which is composed of a multi-filter convolutional neural network and a residual convolutional neural network;
[0005] S2. The multi-filter convolutional neural network is composed of convolutional kernels of different sizes, and the convolutional kernels of different sizes extract the feature information of texts of different lengths in the model;
[0006] S3. The residual convolutional neural network is a residual structure, which helps the model solve the problem of neural network degradation and solve the problems of gradient disappearance and gradient explosion in deep neural networks;
[0007] S4. After the word vector features are extracted by the multi-filter residual convolutional neural network with n convolutional kernels of different sizes respectively, n hidden feature matrices are output. A recalibration aggregation module is added to the model to perform noise reduction processing on the data set. The recalibration aggregation module accepts the output of the multi-filter residual convolutional neural network as input. This module recalibrates the extracted text features, aggregates the original text features and the recalibrated features, and finally combines the new representation with the original representation;
[0008] S5. An MLP layer is added to deepen the depth of the model, and the attention mechanism is used to further process the feature matrix;
[0009] S6. After inputting the results into the feedforward neural network, use the sigmoid function for classification to predict the probabilities of each label.
[0010] The beneficial effects of the present invention are as follows: After being processed by the Longformer pre-trained model, the multi-filter residual convolutional neural network, the re-aggregation calibration module, and the label attention mechanism, the text semantics in the dataset have been fully extracted. However, considering the large scale of the dataset and the imbalance of label distribution, in order to increase the overall depth of the model, expand the number of neuron layers in deep learning, enable the model to better learn the text features among some labels, and fit the information of the data, a feedforward neural network is added at the end of the model. Finally, the results are output through the sigmoid function. And after adding the feedforward neural network, the performance of the model has been slightly improved, which further proves the role of the feedforward neural network in this model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is the overall technical roadmap of the present invention;
[0012] Figure 2 is the overall model structure diagram of the present invention;
[0013] Figure 3 is the process diagram of the Longformer processing word vectors of the present invention;
[0014] Figure 4 is the segmentation mechanism diagram of the present invention;
[0015] Figure 5 is the multi-filter residual convolutional neural network diagram of the present invention;
[0016] Figure 6 is the multi-filter convolutional neural network diagram of the present invention;
[0017] Figure 7 is the residual block structure diagram of the present invention;
[0018] Figure 8 is the RAM calculation process diagram of the present invention;
[0019] Figure 9 is the upsampling calculation process diagram of the present invention;
[0020] Figure 10 is the label dependence example diagram of the present invention
[0021] Figure 11 is the label attention mechanism model diagram of the present invention;
[0022] Figure 12 Overall model input length comparison Figure 1 ;
[0023] Figure 13 Overall model input length comparison Figure 2 。 Specific implementation manners
[0024] The present invention provides a multi-label text classification method for electronic medical records based on Longformer, comprising the following steps:
[0025] S1. The text of the electronic medical record is converted into word vectors through the Longformer pre-trained model, and the word vectors are used as the input of the multi-filter residual convolutional neural network, which is composed of a multi-filter convolutional neural network and a residual convolutional neural network;
[0026] S2. The multi-filter convolutional neural network is composed of convolutional kernel groups of different sizes, and the convolutional kernels of different sizes extract the feature information of texts of different lengths in the model;
[0027] S3. The residual convolutional neural network is a residual structure, which helps the model solve the problem of neural network degradation and solve the problems of gradient vanishing and gradient explosion in deep neural networks;
[0028] S4. After the word vector features are extracted by the multi-filter residual convolutional neural network with n convolutional kernels of different sizes respectively, n hidden feature matrices are output. A recalibration aggregation module is added to the model to perform noise reduction processing on the data set. The recalibration aggregation module receives the output of the multi-filter residual convolutional neural network as the input. This module recalibrates the extracted text features, aggregates the original text features and the recalibrated features, and finally combines the new representation with the original representation to achieve the effect of noise reduction;
[0029] S5. An MLP layer is added to deepen the depth of the model, and the attention mechanism is used to further process the feature matrix;
[0030] S6. After the result is input into the feedforward neural network, the sigmoid function is used for classification to predict the probabilities of each label.
[0031] The overall architecture of the model is specifically as follows:
[0032] First, since the length of electronic medical record texts is generally very long, it is difficult for ordinary pre-trained models to accommodate the overall length of electronic medical record texts. Therefore, in this paper, Longformer is used as the pre-trained model of the model. The text of the electronic medical record will be transformed into word vectors through the Longformer pre-trained model, and the word vectors will become the input of the multi-filter residual convolutional neural network. The multi-filter residual convolutional neural network consists of a multi-filter convolutional neural network and a residual convolutional neural network. Among them, the multi-filter convolutional neural network is composed of convolutional kernel groups of different sizes, and convolutional kernels of different sizes help the model extract feature information of texts of different lengths. The residual convolutional neural network is a residual structure, which can help the model solve the problem of neural network degradation, and solve the problems of gradient disappearance and gradient explosion in deep neural networks. Since the MIMIC dataset used in this paper is relatively large in scale and the depth of the model network in this paper is relatively deep, the characteristics of the residual convolutional neural network are very suitable for the model in this paper.
[0033] After extracting the word vector features through the multi-filter residual convolutional neural network with n convolutional kernels of different sizes respectively, it will output n hidden feature matrices. Since there is a lot of noise information in the electronic medical record text, it is necessary to denoise the dataset. Therefore, a recalibration aggregation module is added to the model to denoise the text, thereby improving the performance of the overall model. The recalibration aggregation module will accept the output of the multi-filter residual convolutional neural network as input. This module recalibrates the extracted text features, aggregates the original text features and the recalibrated features, and finally combines the new representation with the original representation to achieve the denoising effect. Considering that the MIMIC dataset is a large dataset, when using the residual convolutional neural network to solve the problem of network degradation, the depth of the model should be deepened to ensure that the model can be fully trained. Therefore, an MLP layer is added to deepen the depth of the model. Since there are connections between the labels in the electronic medical record text, the attention mechanism is used to further process the feature matrix. After the result is input into the feedforward neural network, the sigmoid function is used for classification to predict the probabilities of each label. Figure 1 The technical route of this paper is given. The overall model structure proposed in this paper is as Figure 2 shown.
[0034] In step S1 described above, Longformer extracts text word vectors
[0035] Since the length of the text of the electronic medical record is very long, ordinary pre-trained models and pre-trained models based on the Transformer structure cannot load the input of the electronic medical record, which will cause the loss of text information data. Therefore, the attention mechanism proposed by Longformer is used to improve the Transformer in Bert.
[0036] Longformer proposed three attention mechanisms, namely the sliding window mechanism, the expansion sliding window mechanism and the sliding window mechanism that integrates global information. The first two attention mechanisms are suitable for text classification tasks, while the third attention mechanism is generally used for QA tasks. Therefore, we should choose from the two attention mechanisms to improve the Transformer. This paper uses the expansion sliding window mechanism to improve the Transformer, because experiments have shown that the expansion sliding window mechanism performs better than the ordinary sliding window mechanism due to the consideration of more comprehensive contextual information. Table 3.1 gives an example of a certain expansion sliding window mechanism being better than the ordinary sliding window mechanism. The red characters indicate the objects of the attention mechanism, w indicates the window size, and d indicates the interval size.
[0037] Under the sliding window mechanism, because the word "lost" is not associated with "casino" using the attention mechanism, "lost" is likely to be judged by the model as meaning "lost", but in fact the real meaning of "lost" should be "losing money". In the sliding window mechanism, because "lost" and "casino" have established an attention relationship, "lost" is likely to be judged by the model as meaning "losing money". Therefore, in this example, the expansion sliding window mechanism is obviously better than the sliding window mechanism. And due to its attention method, the expansion sliding window mechanism can have a larger receptive field than the sliding window mechanism.
[0038] Table 3.1 Comparison of examples of sliding window mechanism and expansion sliding window mechanism
[0039]
[0040] The maximum input length that the Longformer pre-trained model can accept is 4096, which can accommodate the input of long text sequences of electronic medical records. The length statistics of the MIMIC-Ⅲ dataset in this article are shown in Table 3.2.
[0041] Table 3.2 Dataset sentence length statistics
[0042]
[0043]
[0044] As can be seen from Table 3.2, the sentence lengths of the dataset are roughly divided into seven intervals. Among them, the length of the electronic medical record text in the dataset is less than 500, accounting for only 2.07%, while the length of the text above 500 is as high as 97.93%. The maximum input length of the Bert pre-trained model is 512. This indicates that when dealing with the electronic medical record text dataset, most of the texts need to be truncated. This approach will lose a large amount of text information, resulting in the model being unable to accurately extract text features. Therefore, using Bert as the pre-trained model to build a neural network for the task of electronic medical record text classification will undoubtedly weaken the performance of the model.
[0045] And the proportion of sentences with a length above 4,000 is only 2.14%. This shows that for the Longformer pre-trained model, its maximum accommodable length is 4,096. Therefore, the Longformer pre-trained model can accommodate all the information of most sentences. For other sentences with a length exceeding 4,096, Longformer can also extract most of the information. Moreover, the core information of the sentence is generally concentrated in the middle of the sentence. Therefore, for most of the discarded parts, the data importance is not strong. Therefore, it is very suitable to use the Longformer pre-trained model to extract long text semantic information. The process of the Longformer pre-trained model processing word vectors is as Figure 3 shown. After improving the attention mechanism in the Transformer, the text outputs the corresponding word vectors through the stacked encoder layers and decoder layers.
[0046] When exploring the processing of long text sequences before this article, a segmentation mechanism was once used. The structure diagram of the segmentation mechanism is as Figure 4 shown. Due to the excessive length of the long text sequence, the text information cannot be completely extracted by the pre-trained model. Therefore, the core of the segmentation mechanism is to split the long text input sequence into several subsequences, so that the maximum length of the subsequences is less than the maximum input length of the pre-trained model to ensure that the pre-trained model can effectively extract text information. Then these several sub-texts are respectively sent into the Embedding layer to output the corresponding sub-word vectors, and the splicing of these several word vectors will obtain a complete word vector feature matrix. Thus, the problem that the sentence length of the long text sequence is greater than the pre-trained model can be solved.
[0047] However, the use of the segmentation mechanism has certain limitations. Since splitting a long text sequence will disrupt the integrity of the original sentence. There is a strong connection between subsequences because these subsequences belong to the same long text. Feeding the subsequences into the Embedding layer separately can obtain the word vectors of each subsequence, but the strong correlation between each subsequence is lost. This also makes the performance of the segmentation mechanism far inferior to that of the Longformer pre-trained language model in extracting sentence text information. The method of Longformer to broaden its maximum input length retains the integrity of the input long text to the greatest extent. Therefore, Longformer is very suitable for extracting the semantic information of long text sequences. For this reason, this paper uses the Longformer pre-trained model as part of the overall experimental model, which also lays a good foundation for the overall experimental model to effectively learn the features of electronic medical record texts.
[0048] In step S2, the multi-filter residual convolutional neural network
[0049] In the task of electronic medical record text classification, different long text sequences are not only longer in length compared to ordinary texts, but also the lengths of each sentence in the long text are different. Therefore, for long text sequences of different lengths, traditional deep learning neural network models cannot extract the feature information of the text to the maximum extent. Therefore, in order to capture text sequences of different lengths, this paper uses a multi-filter convolutional neural network, where each filter has a different kernel size. The overall structure diagram of the multi-filter residual convolutional neural network is as Figure 5 shown.
[0050] The multi-filter residual convolutional neural network includes two modules: the multi-filter convolutional neural network and the residual convolutional neural network. The structure diagram of the multi-filter convolutional neural network is as Figure 6 shown.
[0051] When the multi-filter convolutional neural network receives word vectors as input, further feature extraction operations are performed on the word vectors through convolutional operations. The process of the convolution operation is as shown in (3-1).
[0052]
[0053] Among them, represents a left-to-right convolution operation, f1 and f n represent the corresponding filters, H1 and H n represent the feature matrices output by the multi-filter convolutional neural network. E represents the word vector matrix, and refer to the weight matrices of the corresponding filters. Represents a subset of the word vector matrix E, which starts from the j-th row and ends at the th row. Similarly, it starts from the j-th row and ends at the j + k m -1th row.
[0054] In step S3, vanishing gradients and exploding gradients are problems that are likely to be encountered during the training of deep networks. Due to the deepening of the network layers, the expansion or shrinkage effect of gradients accumulates continuously, and ultimately it is very easy to cause the model to not converge. To solve the problems of vanishing gradients and exploding gradients, a residual neural network is added on top of each filter in the multi-filter convolutional layer. A residual neural network consists of p residual blocks. For a residual block, it is internally composed of three convolutional filters, and the structure diagram of the residual block is as Figure 7 shown.
[0055] The residual block receives the input H of the multi-filter convolutional neural network and processes it through the filter r1. The processing process is as shown in (3-2).
[0056]
[0057] Among them, X1 represents the output of the filter r1, represents the weight matrix of the filter r1, is a subset of the feature matrix H, which starts from the j-th row of the H matrix and ends at the j + k m -1th column. After the filter r1 outputs X1, X1 will be used as the input of the next filter r2, and the processing process is as shown in (3-3).
[0058]
[0059] Among them, X2 represents the output of the filter r1, represents the weight matrix of the filter r2, is a subset of X1, which starts from the j-th row of the H matrix and ends at the j + k m -1th column.
[0060] The output H of the multi-filter convolutional neural network will become the input of the filter r3. After being convolved by the filter r3, the output X3 is obtained, and the processing process is as shown in (3-4).
[0061]
[0062] Among them, X3 represents the output of the filter r3, represents the weight matrix of the filter r3, H j:j is a subset of H, which refers to the feature vector of the j-th row in the feature matrix H.
[0063] The output X of the residual convolutional network can be obtained by adding the elements of X2 and X3. And the output X of each residual block ip , will be concatenated to form a new feature matrix Y. It will be output to the recalibrated aggregation module, and the output will be further denoised by the recalibrated aggregation module.
[0064] In step S4, the recalibrated aggregation module
[0065] Since the electronic medical record text contains a lot of noise information, these noise information will have a certain negative impact on the final performance of the model. Therefore, it is very important to add a noise removal module to the model. In this paper, a recalibrated aggregation module (RAM) is added, which can extract the text features learned by the multi-filter residual convolutional neural network, recalibrate the extracted text features, aggregate the original text features and the recalibrated features, and finally combine the new representation with the original representation. In this way, the RAM module can reduce the influence of noise in the electronic medical record text and improve the performance of the ICD coding task model. Specifically, the RAM module extracts and aggregates context information by using a nested convolutional structure, and adds the newly aggregated features to the original input features to achieve the result of removing noise. In addition, through convolution, RAM obtains a global receptive field in the feature extraction process. Therefore, RAM can improve the encoding of noisy and lengthy electronic medical record text. RAM consists of feature aggregation and recalibration. The calculation process of the RAM model is as Figure 8 shown.
[0066] First, the hidden representation Y extracted by the multi-filter residual convolutional neural network can obtain matrices A and A' through two downsampling operations. The process of transforming the hidden representation Y into A through downsampling can be expressed as:
[0067]
[0068] where represents dislocation addition, that is, the second matrix moves one unit to the right based on the position of the first matrix. Repeat this operation until the last matrix, and cut off the unit vectors connecting both sides of the matrix. The overlapping areas are summarized. and represent two groups of convolutional kernels on the output A matrix when the hidden matrix Y undergoes a downsampling operation. When the matrix A generates the matrix A' through a downsampling operation, the specific process of the operation is also similar to (3-5). However, the dimensions of the corresponding convolutional kernel groups change, from the original and to and Then, use another pair of convolutional kernel groups and to transform matrix A' into a horizontal feature matrix L. Then, use the upsampling operation to restore matrix L. At this time, the shape of matrix L is the same as that of matrix A, and the matrix addition operation can be performed. This will combine the newly extracted matrix information with the original matrix to achieve the purpose of recalibrating aggregation and eliminating the noise information in the electronic medical record text. The transposed convolutional kernel group used in this upsampling operation is and Through this operation, adding matrix A and the output after the upsampling operation of matrix L can obtain matrix B. The specific calculation process is shown in (3-6).
[0069]
[0070] Then, the process of generating the weight matrix O from the aggregated matrix B through the upsampling operation is as Figure 9 shown.
[0071] As Figure 9 shown, the aggregated matrix B undergoes a transposed convolution operation through the transposed convolutional kernel group to obtain matrix T'. Matrix T' will obtain the intermediate representation T after the dislocation addition operation. The specific calculation process is shown in (3-7).
[0072]
[0073] Performing a transposed convolution operation on the intermediate representation T to obtain O', where the transposed convolutional kernel group is Then, performing a dislocation addition operation on O' can obtain the weight matrix O. The specific process is shown in (3-8).
[0074]
[0075] Finally, performing a matrix dot product operation on the weight matrix O and the feature matrix Y output by the initial multi-filter residual convolutional neural network can obtain the recalibrated feature matrix Y'. The specific calculation process is shown in (3-9).
[0076] Y' = tanh (O ⊙ Y) (3-9)
[0077] The recalibration operation enhances the original features by injecting context information through the weight matrix O, which contains rich semantic information and is thus insensitive to the noise in the electronic medical record text. It enables the overall model to have better generalization ability and ultimately improves the performance of the electronic medical record text classification task. After the feature matrix Y' is output by the recalibration aggregation module, it serves as the input to the MLP. The overall model obtains more training opportunities due to the increased depth of the framework, which further deepens the training of the overall model. The output Q of the MLP layer serves as the input to the label attention mechanism.
[0078] In step S5, the label attention mechanism
[0079] In long documents, the occurrence of certain words may be crucial for predicting labels. For example, when words such as "Tuberculosis infection", "hepatitis", "cold", "lack strength" appear, it is very likely that "fever" will appear. However, this relationship may not hold in reverse. When "fever" appears, words such as "Tuberculosis infection", "hepatitis", "cold" may not appear, but it is very likely that "lack strength" will appear. An example diagram of label dependence is as Figure 10 shown.
[0080] Therefore, among various different labels, there is a certain label dependence relationship with the long documents in the dataset. If the information existing in the label dependence relationship can be combined with the features of the label itself, the model can obtain higher prediction performance and improve the accuracy of the model. Therefore, this paper uses the attention mechanism to assign different weights to different features, and those words that are more significant for predicting labels will be assigned higher weights, which helps to improve the performance of model prediction.
[0081] The overall model diagram of the attention mechanism is as Figure 11 shown.
[0082] The label attention mechanism takes the feature vector matrix Q output by the MLP layer as input, sets an adjustable initial matrix U of the attention mechanism, and sends both the Q matrix and the U matrix into the softmax layer for processing to obtain the weight matrix S of the label attention. The processing process is as shown in (3-10).
[0083] S = Softmax(QU) (3-10)
[0084] Perform a transpose operation on the attention weight matrix S to obtain the matrix S T , and then S TPerform a matrix multiplication operation with Q to give the original feature matrix the weights of the important words for each label, enabling the model to capture the label dependencies in the text. The processing is shown in (3-11).
[0085] Z = S T Q (3-11)
[0086] In the step S6,
[0087] After the label attention mechanism layer outputs the weight matrix Z with label attention, the final prediction result is output through a feed-forward neural network and sigmoid.
[0088] In this experiment, the Adam optimizer and the backpropagation algorithm are used to train the model. The loss function of this experiment is shown in (3-12).
[0089]
[0090] Where w represents the input text, y represents the true label, θ represents all the parameters, and l represents the size of the word vector dimension, represents the predicted value of the model for the label.
[0091] After being processed by the Longformer pre-trained model, the multi-filter residual convolutional neural network, the re-aggregation calibration module, and the label attention mechanism, the text semantics in the dataset have been fully extracted. However, considering the large scale of the dataset and the imbalance of label distribution, in order to increase the overall depth of the model, expand the number of neuron layers in deep learning, enable the model to better learn the text features in some labels, and fit the information of the data, a feed-forward neural network is added at the end of the model. Finally, the result is output through the sigmoid function. And after adding the feed-forward neural network, the performance of the model has been slightly improved, which further proves the role of the feed-forward neural network in this model.
[0092] The experimental results and analysis are as follows:
[0093] 4.1 Experimental dataset
[0094] The MIMIC-Ⅲ dataset is a free and open public resource intensive care unit research database. This database was jointly released in 2006 by the Computational Physiology Laboratory of the Massachusetts Institute of Technology, Beth Israel Deaconess Medical Center (BIDMC) and Philips Healthcare, attracting an increasing number of researchers in academia and industry to use this medical database for medical research. Due to the problem of label imbalance in the MIMIC dataset, the MIMIC-Ⅲ dataset is divided into MIMIC-Ⅲ (TOP-50) and MIMIC-Ⅲ (Full). Among them, the MIMIC-Ⅲ (TOP-50) dataset represents the dataset composed of the 50 most frequently occurring labels in MIMIC-Ⅲ, while MIMIC-Ⅲ (Full) represents the entire MIMIC-Ⅲ dataset.
[0095] This paper mainly uses five data tables in the MIMIC-Ⅲ dataset, namely D_ICD_DIAGNOSES.csv, D_ICD_PROCEDURES.csv, DIAGNOSES_ICD.csv, NOTEEVENTS.csv, and PROCEDURES_ICD.csv.
[0096] Tables 4.1, 4.2, 4.3, 4.4, and 4.6 respectively show the structures of the five data tables in the five csv files of D_ICD_DIAGNOSES.csv, D_ICD_PROCEDURES.csv, DIAGNOSES_ICD.csv, NOTEEVENTS.csv, and PROCEDURES_ICD.csv. Table 4.5 gives an example of the electronic medical record text. The detailed description of the dataset is shown in Table 4.7.
[0097] Table 4.1 D_ICD_DIAGNOSES structure table Table 4.1 D_ICD_DIAGNOSES structure table
[0098]
[0099] Table D_ICD_DIAGNOSES is the ICD disease diagnosis dictionary table, which mainly records the correspondence between ICD codes and the corresponding diagnoses. Among them, ROW_ID represents the row number, ICD9_CODE represents the corresponding ICD code, SHORT_TITLE represents the abbreviation of the corresponding diagnosis, and LONG_TITLE represents the full name of the corresponding diagnosis. For example, in the 19th row of the table, the row number of the data is 19, the corresponding ICD code is 1,743, the diagnosis abbreviation is "TB of ear-micro dx", and the full name of the diagnosis is "Tuberculosis of ear, tubercle bacilli found (in sputum) by microscopy", that is, ear tuberculosis, tubercle bacilli found by microscopy (in sputum). When the row number of the data is 36, the corresponding ICD code is 1,766, the diagnosis abbreviation is "TB of adrenal-oth test", and the full name of the diagnosis is "Tuberculosis of adrenal glands, tubercle bacilli not found by bacteriological or histological examination, but tuberculosis confirmed by other methods [inoculation of animals]", that is, "adrenal tuberculosis, tubercle bacilli not found by bacteriological or histological examination, but tuberculosis confirmed by other methods [animal inoculation]".
[0100] Table 4.2 D_ICD_PROCEDURES structure table
[0101]
[0102] Table D_ICD_PROCEDURES is the ICD treatment process dictionary table, which mainly records the correspondence between ICD codes and the corresponding treatment processes. Among them, ROW_ID represents the row number, ICD9_CODE represents the corresponding ICD code, SHORT_TITLE represents the abbreviation of the diagnosis corresponding treatment process, and LONG_TITLE represents the full name of the corresponding treatment process. For example, when the row number is 271, the corresponding ICD code is 869, the short title of the treatment process is "Lid reconstr w graft NEC", and the full title is "Other reconstruction of eyelid with flaps or grafts", that is, "Reconstruct the eyelid with flaps or grafts". When the row number is 314, the corresponding ICD code is 411, the short title of the treatment process is "Clos periph nervebiopsy", and the full title of the treatment process is "Closed[percutaneous][needle]biopsy of cranial orperipheral nerve or ganglion", that is, "Closed [percutaneous] [needle] biopsy of cranial or peripheral nerve or ganglion".
[0103] Table 4.3 DIAGNOSES_ICD Structure Table
[0104]
[0105] Table DIAGNOSES_ICD is the diagnosis information table, where ROW_ID represents the row number, SUBJECT_ID represents the patient number, HADM_ID represents the hospitalization number, SEQ_NUM represents the order of ICD diagnosis related to the patient, and ICD diagnoses are sorted according to priority, and this order will affect the treatment of the patient, while ICD9_CODE represents the ICD code. For example, when the row number is 1,297, SUBJECT_ID is 109, representing the patient numbered 109, HADM_ID is 172,335, and SEQ_NUM is 1. At this time, the ICD code is 40,301. In the table, this patient is associated with multiple ICD codes, such as ICD codes 486, 58,281, etc. are all ICD codes related to this patient. And SEQ_NUM indicates the order of the relevance of these ICD codes to this patient.
[0106] Table 4.4 NOTEEVENTS Structure Table
[0107]
[0108]
[0109] Table NOTEEVENTS is the text record event table, which mainly records the medical record text content of doctors and some related information. Among them, ROW_ID represents the row number, SUBJECT_ID represents the patient number, HADM_ID represents the hospitalization number, CHARTDATE represents the date of recording the text, CHARTTIME represents the time of recording the text, STORETIME represents the time when the text is saved to the system, CATEGORY represents the type of recorded text, DESCRIPTION represents the detailed classification of the type of recorded text, CGID represents the caregiver identification number, ISERROR represents whether the information has errors. If the text contains errors, the value of ISERROR will be marked as 1, and TEXT represents the content of the medical record text. The length of the electronic medical record text is generally long. Table 4.5 gives part of the medical record text content when ROW_ID = 223, SUBJECT_ID = 5,350, and HADM_ID = 169,684.
[0110] Table 4.5 Example of medical record text
[0111]
[0112]
[0113]
[0114] Table 4.6 PROCEDURES_ICD structure table
[0115]
[0116] Table PROCEDURES_ICD represents the ICD operation record table, where ROW_ID represents the row number, SUBJECT_ID represents the patient number, HADM_ID represents the hospitalization number, SEQ_NUM represents the order executed in the treatment process, such as surgery or physical therapy, and ICD9_CODE represents the ICD code. The data set statistical information is shown in Table 4.7.
[0117] Table 4.7 Data set statistics
[0118]
[0119] Two datasets are used in this paper, namely the MIMIC-Ⅲ (TOP-50) dataset and the MIMIC-Ⅲ (Full) dataset. The MIMIC-Ⅲ (TOP-50) dataset is a sub-dataset of the MIMIC-Ⅲ (Full) dataset, which contains the 50 most frequent tags of the MIMIC-Ⅲ (Full) dataset.
[0120] 4.2 Evaluation Metrics for Experiments
[0121] For the evaluation of the performance of the electronic medical record text classification model, this chapter uses multiple evaluation metrics to evaluate the overall performance of the model, mainly including: macro AUC value, micro AUC value, macro F1 value, micro F1 value, and precision at k (Precision@k). For the MIMIC-Ⅲ (TOP-50) dataset, each electronic medical record text contains an average of 5.8 ICD codes, so Precision@5 is selected as the corresponding evaluation metric. Similarly, for the MIMIC-Ⅲ (Full) dataset, Precision@8 is selected as the corresponding evaluation metric.
[0122] First, construct the corresponding confusion matrix as shown in Table 4.8.
[0123] Table 4.8 Confusion matrix
[0124] Table 4.8 Confusion matrix
[0125]
[0126] Among them: TP (True Positive) represents the true positive example, which is the number of samples where the instance is a positive class and the prediction result is also a positive class. FN (False Negative) represents the false negative example, which is the number of samples where the instance is a positive class but is predicted as a negative class. FP (False Positive) represents the false positive example, which is the number of samples where the instance is a negative class but is predicted as a positive class. TN (True Negative) represents the true negative example, which is the number of samples where the instance is a negative class and is also predicted as a negative class.
[0127] Therefore, the calculation formula for the true positive rate (TPRate) is shown in (4-1), and the calculation formula for the false positive rate (FPRate) is shown in (4-2).
[0128]
[0129]
[0130] Taking the true positive rate as the x-axis of the coordinate and the false positive rate as the y-axis of the coordinate, a receiver operating characteristic (ROC) curve is plotted, and the AUC (Area Under Curves) is defined as the area under the ROC curve.
[0131] The precision P represents the proportion of actual true positives among all the data samples predicted as positive by the model. It reflects the ability of the model to predict positive samples. The higher the precision, the better the model's ability to predict positive samples. Its specific calculation formula is:
[0132]
[0133] The recall rate R, also known as the true positive rate, represents the proportion of samples predicted as true positives among all the data samples that are actually positive in the model. The better the recall rate of the model, the better its classification performance. The calculation formula for the recall rate is shown in (4-1).
[0134] In order to evaluate the advantages and disadvantages of different algorithms, the concept of the F1 value is proposed, and its calculation formula is:
[0135]
[0136] When performing multi-class classification, macro-average and micro-average are often used to evaluate the performance of the classification model. The difference between macro-average and micro-average is that macro-average weighs each class, while micro-average weighs each sample. If the number of samples in each class is the same, then macro-average and micro-average can obtain similar results. If there are significant differences in the number of samples in each class, then the results obtained by macro-average and micro-average will also be quite different.
[0137] 4.3 Experimental Environment and Parameter Selection
[0138] The system used in the experiment is Windows 10. On this basis, the Pycharm programming software is used in terms of software. The experiment is carried out based on multiple Python libraries such as Pytorch. The specific hardware information is shown in Table 4.9. After multiple comparative experiments, the final parameter selection for the experiment is shown in Table 4.10.
[0139] Table 4.9 Overview of the experimental environment
[0140]
[0141] Table 4.10 Main parameter Settings in the experiment
[0142]
[0143] 4.4 Experimental Results and Analysis
[0144] To verify the effectiveness of the method, this paper conducted a comparative experiment. The main models for comparison are as follows:
[0145] (1) C-MemNN: Proposed a compressed memory neural network, which equips the neural network with an iterative compressed memory representation.
[0146] (2) C-LSTM-Att: Proposed an attention model based on feature-aware LSTM to assign ICD codes to clinical notes. They used an LSTM-based language model to generate representations of clinical annotations and ICD codes, and proposed an attention mechanism to address the mismatch between annotations and codes.
[0147] (3) Bi-GRU: Used stacked Bi-GRU modules to extract text features.
[0148] (4) CAML: It consists of a convolutional layer and an attention layer for generating label-aware features for multi-label classification.
[0149] (5) DR-CAML: It is an extension of CAML, which includes the text description of each code to regularize the model.
[0150] (6) MultiResCNN: Used a multi-filter residual convolutional neural network to encode text.
[0151] (7) TransICD: Used a deep learning method based on Transformer, adopted the Label Distribution Aware Margin (LDAM) loss to combat the imbalanced dataset, and used an attention mechanism for multi-label prediction.
[0152] (8) MultiResCNN+MADE: Proposed two sorters, MADE and Mask-SA. Taking MultiResCNN as the baseline and adding the sorter on this basis.
[0153] (9) MultiResCNN+Mask-SA: Used MultiResCNN and the re-sorter Mask-SA.
[0154] (10) LongBERT: Divide the long text into several sub-texts with a length of 512 and send them into BERT respectively, and use the attention mechanism for prediction.
[0155] In the framework of the whole model, first use the pre-trained language model to vectorize the words, then use the multi-filter convolutional neural network, the residual convolutional network and the attention model to extract and process the segment features of the sentence. Finally, at the output layer, let the hidden features be output through the feed-forward neural network, so as to achieve the purpose of text classification. In the experiment, the sentence lengths in the MIMIC-Ⅲ dataset were counted, and the ratio of the sentence length to the corresponding interval was counted. In the overall model, in order to keep the dimension of the feature matrix of the overall model consistent, a unified value needs to be set for the maximum length of the input sentence. The length of the sentence needs to be set reasonably. If the length of the input sentence is too small, it will affect the model's insufficient extraction of sentence meaning features, resulting in a decline in the overall performance of the model; if the length of the input sentence is too large, it will not only affect the training time of the overall model, but also for shorter sentences in the dataset, in order to conform to the input sentence length set in the experiment, padding will be performed on the insufficient part after the short sentence, thus affecting the true meaning of the sentence and causing an impact on the performance of the model. Therefore, in this experiment, only the Longformer pre-trained model plus the multi-filter residual convolutional neural network model, that is, the Longformer+MultiCNN model, was used to conduct a comparative experiment for different input lengths as shown in Table 4.11. At this time, the model for the comparative experiment is not the complete model of this article. The dataset used at this time is MIMIC-Ⅲ (Top50).
[0156] Table 4.11 Longformer+MultiCNN Model input length comparison test
[0157]
[0158] Table 4.11 Longformer+MultiCNN Model input length comparison test (continued)
[0159] Table 4.11 Longformer+MultiCNN Model input length comparison test (continued)
[0160]
[0161] As can be seen from Table 4.11, when the value of MaxLength, i.e., the maximum input length, is 3,000, the overall performance of the model is significantly better than that when the maximum input length is 2,500 or 3,500. Therefore, continue to change the value of the maximum input length and observe the overall situation of the model when the maximum input length is around 3,000. Experiments have found that when the values of the maximum input length are 2,750, 2,850, and 3,100, the overall performance of the model is inferior to that when the value is 3,000. Moreover, when the value of the maximum input length is 3,000, the loss value reaches the lowest, and many data are the best values in the table. Therefore, select the value of the maximum input length as 3,000 and continue the experiment. And, in order to fully study the impact of different maximum input lengths on the complete model, the bar charts of various indicators of the model in this paper under different maximum input lengths are as Figure 12 and Figure 13 shown.
[0162] F1
[0163] Using the complete experimental model in this paper, five relatively competitive values of the maximum input length are selected from Table 4.11 as parameters to continue the comparative experiment. Thus, explore the optimal value of the maximum input length of the model. The maximum input lengths for the comparative experiment are set to 2,750, 2,850, 3,000, 3,100, and 3,500 respectively. The reason for selecting these five values as the maximum input length is that, as shown in Table 4.11, when using Longformer+MultiResCNN as the model, when the values of the maximum input length are these five values, the performance of the model is relatively good. The dataset used in this experiment is MIMIC-Ⅲ(Top50). Observe Figure 12 and Figure 13 , when the value of the maximum input length is 3,500, the value of the macro AUC reaches the highest, but the difference is not very large compared with when the maximum input length is 3,000 and 3,100. And when the maximum input length is 3,000, the value of the micro AUC reaches the highest. And in Figure 12 , with MaxLength = 3,000 as the axis, the values of the micro AUC on both sides show a downward trend. And in Figure 13 , when the value of the maximum input length is set to 3,000, the macro F1 value, the micro F1 value, and Precision@8 all reach the highest among the five parameters. Therefore, this further proves that when the maximum input length of the model in this paper is set to 3,000, the effect of the model in this paper is the best.
[0164] In this experiment, the MIMIC-Ⅲ (Top50) dataset and the MIMIC-Ⅲ (Full) dataset were used for a comparative experiment. As shown in Table 4.12, the experimental results on the MIMIC-Ⅲ (Top50) dataset are given. As shown in Table 4.13, the experimental results on the MIMIC-Ⅲ (Full) dataset are given.
[0165] Table 4.12 MIMIC-Ⅲ (Top50) Dataset experimental results
[0166] Table 4.12 MIMIC-Ⅲ(Top50) Dataset experimental results
[0167]
[0168]
[0169] Table 4.13 MIMIC-Ⅲ(Full) Dataset experimental results
[0170]
[0171] Since there are many labels in the MIMIC-Ⅲ (Full) dataset, and due to the problem of unbalanced label set distribution, there are few training sets corresponding to many labels. However, the MIMIC-Ⅲ (Top50) dataset has a large amount of data for training for 50 labels. Therefore, the two datasets are distinguished and trained based on this difference. On these two datasets, the model of this experiment and the previous models were used for comparative experiments. Observing Table 4.12 and Table 4.13, it can be found that in the MIMIC-Ⅲ (Full) dataset, the macro value of the F1 value is much lower than the micro value. This is because there is an unbalanced label distribution in the MIMIC-Ⅲ (Full) dataset. The training sets of many labels may only have a few or dozens of data in the MIMIC-Ⅲ (Full) dataset, which results in insufficient training of the model for such labels. Too few training sets result in the model not fully learning the hidden features of these labels, resulting in the model showing poor performance when predicting these labels. In the MIMIC-Ⅲ (Top50) dataset, although the macro F1 value is still smaller than the micro F1 value, the gap between the two is greatly reduced. This is because in the MIMIC-Ⅲ (Top50) dataset, the 50 most frequent labels of the MIMIC-Ⅲ dataset are selected as the training set. These labels have a large amount of training data, so the model has been fully trained, and the gap between the macro F1 value and the micro F1 value has also been greatly reduced. Moreover, in Table 4.12 and Table 4.13, the model of this paper has achieved good performance on two different MIMIC-Ⅲ datasets. This is because the Longformer pre-training model retains the semantic feature information of long texts to the greatest extent, which also proves the effectiveness of the experimental model of this paper.
[0172] This paper also conducted ablation experiments on two datasets to prove the effectiveness of this paper based on Longformer on the MIMIC-Ⅲ (Top50) dataset and MIMIC-Ⅲ (Full) dataset. The ablation experiment results are shown in Tables 4.14 and 4.15.
[0173] Table 4.14 Ablation experiment results of MIMIC-Ⅲ (Top50) dataset
[0174] Table 4.14MIMIC-Ⅲ(Top50)Dataset experimental results of ablation experiment
[0175]
[0176] Table 4.15 Ablation experiment results of MIMIC-Ⅲ (Full) dataset
[0177] Table 4.15 Experimental results of ablation experiment on MIMIC-Ⅲ (Full) Dataset
[0178]
[0179]
[0180] Observing Table 4.14 and Table 4.15, through the ablation experiment, when the pre-trained model was replaced with Word2vec, the performance of the model reached the lowest, which also proved the effectiveness of Longformer in processing long texts. When the model lost attention, the various data of the model also decreased. This is because after losing attention, the model cannot effectively establish strong relationships such as labels and sensitive words, which makes the performance of the model decline. When the model lost RAM, the performance of the model also decreased. This also shows that in a dataset with a large amount of noise, the operation of the model for noise reduction is very important. The roles of MLP and FNN are mainly to increase the number of hidden layers and add more nodes in the hidden layer. Because for large-scale datasets or complex datasets, increasing the number of network layers and the number of nodes in the network is effective. In the ablation experiment, after the model lost MLP and FNN, the performance of the model also decreased slightly. Therefore, after being demonstrated by the ablation experiment, the effectiveness of the overall framework of this paper was proved.
[0181] From the experimental results, it can be seen that the model proposed in this paper has obtained good experimental results through comparison with the models proposed by other scholars and ablation experiments on the MIMIC dataset, thus proving that the model has good performance.
[0182] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An electronic medical record multi-label text classification method based on Longformer, characterized in that It includes the following steps: S1. The text of the electronic medical record is converted into word vectors through the Longformer pre-trained model. The word vectors are used as the input of the multi-filter residual convolutional neural network, which is composed of a multi-filter convolutional neural network and a residual convolutional neural network; S2. After the word vector features are extracted by the multi-filter residual convolutional neural network with n different-sized convolutional kernels respectively, n hidden feature matrices are output. A recalibration aggregation module is added to the model to denoise the data set. The recalibration aggregation module RAM takes the output of the multi-filter residual convolutional neural network as the input. This module recalibrates the extracted text features, aggregates the original text features and the recalibrated features, and finally combines the new representation with the original representation; In step S2, the calculation process of the RAM is as follows: First, the hidden representation Y extracted by the multi-filter residual convolutional neural network can obtain matrices A and A' through two downsampling operations. The process of converting the hidden representation Y into A through downsampling is expressed as: Among them, represents dislocation addition, that is, the second matrix moves one unit to the right based on the position of the first matrix, and this operation is repeated until the last matrix, and the unit vectors connecting the two sides of the matrix are cut off, and the overlapping area is summarized. and represents two convolutional kernel groups on the output A matrix during the downsampling operation by the hidden matrix Y. When the matrix A generates the matrix A' through the downsampling operation, the specific process of the operation is also similar to (3-5), but the dimensions of the corresponding convolutional kernel groups have changed, from the original and changed to and Then, another pair of convolutional kernel groups and are used to transform the matrix A' into a horizontal feature matrix L; then, the matrix L is restored using the upsampling operation. At this time, the shape of the matrix L is the same as the shape of the matrix A, and the matrix addition operation is performed to combine the newly extracted matrix information with the original matrix to achieve the purpose of re-calibrating aggregation and eliminating the noise information in the electronic medical record text. The deconvolution kernel groups used in this upsampling operation are and Through this operation, the matrix A and the output after the upsampling operation of the L matrix are added to obtain the B matrix, and the specific calculation process is shown in (3-6). Then, the aggregation matrix B generates the weight matrix O through the upsampling operation; The aggregation matrix B is subjected to a deconvolution operation via a set of deconvolution kernels to obtain a matrix T'. The matrix T' undergoes a dislocation addition operation to obtain an intermediate representation T. The calculation process is shown in (3-7). Performing a deconvolution operation on the intermediate representation T to obtain O′, where the deconvolution kernel group is Then, performing a dislocation addition operation on O′ can obtain the weight matrix O, and the process is as shown in (3-8). Finally, the matrix dot product operation is performed on the weight matrix O and the feature matrix Y output by the initial multi-filter residual convolutional neural network to obtain the recalibrated feature matrix Y′. The calculation process is shown in (3-9), Y′ = tanh(O⊙Y)(3-9); S3. An MLP layer is added to deepen the depth of the model, and the attention mechanism is used to further process the feature matrix; S4. After the result is input into the feedforward neural network, the sigmoid function is used for classification to predict the probabilities of each label.
2. The method for multi-label text classification of electronic medical records based on Longformer according to claim 1, wherein In step S1, Longformer extracts text word vectors, The attention mechanism proposed by Longformer is used to improve the Transformer in Bert. The sliding window mechanism of the dilated sliding window mechanism adopted by Longformer is applicable to text classification tasks, considering comprehensive context information, and the Transformer is improved using the dilated sliding window mechanism.
3. The method for multi-label text classification of electronic medical records based on Longformer according to claim 1, wherein, In step S1, when the multi-filter convolutional neural network receives word vectors as input, further feature extraction operations are performed on the word vectors through convolutional operations. The convolutional process is shown in (3-1), Among them, represents a left-to-right convolution operation, f1 and f n represents the corresponding filter, H1 and H n represents the feature matrix output by the multi-filter convolutional neural network, E represents the word vector matrix, and refers to the weight matrix of the corresponding filter, represents a subset of the word vector matrix E, which starts from the j-th row and ends at row, while similarly, it starts from the j-th row and ends at the (j + k) m - 1-th row.
4. The method for multi-label text classification of electronic medical records based on Longformer according to claim 1, characterized in that In step S1, a residual neural network is added to the top of each filter in the multi-filter convolutional layer to solve the problems of gradient disappearance and gradient explosion. A residual neural network is composed of p residual blocks. For a residual block, it consists of three convolutional filters inside. The residual block receives the input H of the multi-filter convolutional neural network and processes it through the filter r1. The processing process is shown in (3-2), Among them, X1 represents the output of filter r1, represents the weight matrix of filter r1, is a subset of the feature matrix H, which starts from the j-th row of the H matrix and ends at the (j + k) m - 1 column. After the filter r1 outputs X1, X1 will be used as the input of the next filter r2, and the processing process is shown in (3-3). Among them, X2 represents the output of filter r1, represents the weight matrix of filter r2, is a subset of X1, which starts from the j-th row of the H matrix and ends at the (j + k) m - 1 column; The output H of the multi-filter convolutional neural network will become the input of the filter r3, and after being convolved by the filter r3, the output X3 is obtained. The processing process is shown in (3-4), Among them, X3 represents the output of filter r3, represents the weight matrix of filter r3, H j:j is a subset of H, which refers to the eigenvector of the j-th row in the feature matrix H; The output X of the residual convolutional network can be obtained by adding the elements of X2 and X3, and the output X of each residual block ip , is concatenated to form a new feature matrix Y, which is output to the recalibration aggregation module, and the output is further denoised by the recalibration aggregation module.
5. The method for multi-label text classification of electronic medical records based on Longformer according to claim 1, wherein In the step S3, the label attention mechanism receives the feature vector matrix Q output by the MLP layer as input, sets an initial matrix U of the attention mechanism with adjustable parameters, and sends the Q matrix and the U matrix into the softmax layer for processing at the same time, and the weight matrix S regarding label attention is obtained. The processing process is shown in (3-10), S = Softmax(QU) (3-10) Transpose the attention weight matrix S to obtain the matrix S T , and then multiply S T by Q through matrix multiplication, so that the original feature matrix has the weights of important words for each label, enabling the model to capture the label dependencies in the text. The processing process is shown in (3-11). Z = S T Q(3 - 11).
6. The method for multi-label text classification of electronic medical records based on Longformer according to claim 1, wherein, In the step S4, after the label attention mechanism layer outputs the weight matrix Z with label attention, the final prediction result is output through a feedforward neural network and sigmoid, and the Adam optimizer and the backpropagation algorithm are used to train the model. The loss function is shown in (3-12), where w represents the input text, y represents the true label, θ represents all parameters, and l represents the size of the word vector dimension, indicating the predicted value of the label by the model.
Citation Information
Patent Citations
Multi-label text classification processing method and system and information data processing terminal
CN111428026A
Label perception-based gated recurrent acquisition method
WO2023124110A1