Multi-level Semantic Text Classification Method Based on Li's Artificial Liver Medical Records

Through multi-level semantic text representation network and feature extraction network, the problem of inefficient classification of Li's artificial liver record text is solved, and more accurate and efficient judgment of patient suitability is achieved.

CN116467440BActive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310330319.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2025-06-27
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the complaints, current medical history and previous history text data in the Li's artificial liver record, resulting in inefficient judgment on whether patients are suitable for Li's artificial liver treatment.

Method used

A text representation network based on multi-level semantics is adopted, and a word vector position coding fusion network, a word vector position coding fusion network and a word vector context information extraction network are extracted to generate richer semantic features, combining Bi-LSTM network and full connection layer to realize feature extraction and classification.

Benefits of technology

It improves the accuracy of the classification of Li's artificial liver medical record texts, makes full use of the semantic information of the main complaint, current medical history and previous historical texts, and improves the efficiency of judging the suitability of patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116467440B_ABST
    Figure CN116467440B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of medical natural language processing, specifically a multi-level semantic text classification method based on the medical records of Li's artificial liver. The process is as follows: collect the text data of the current medical history, chief complaint, and past medical history in the medical records; preprocess the text of the chief complaint and past medical history first and then input them into the text representation network based on multi-level semantics to obtain the chief complaint word vector matrix and the past medical history word vector matrix respectively. Segment the text of the current medical history into sentence sequences, preprocess them, then pass them through the text representation network based on multi-level semantics, and then obtain the current medical history word vector matrix through average pooling. The chief complaint word vector matrix, the past medical history word vector matrix, and the current medical history word vector matrix respectively pass through the feature extraction network and then output the classification result through the feature fusion prediction network. Compared with the existing method of directly splicing each paragraph of text for classification, the present invention makes full and effective use of the semantic information of the three texts of the current medical history, chief complaint, and past medical history, and improves the accuracy of prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical natural language processing, and specifically to a multi-level semantic text classification method based on Li's artificial liver medical records. Background Art

[0002] Liver failure refers to a group of clinical syndromes mainly manifested by coagulation mechanism disorders, jaundice, hepatic encephalopathy, ascites, etc. when the liver is severely damaged by various factors (such as viruses, alcohol, drugs, etc.), resulting in a large number of hepatocyte necrosis, leading to serious disorders or decompensation of liver function. At present, the main treatment methods for liver failure are comprehensive medical treatment, stem cell treatment, and artificial liver support treatment.

[0003] The first step for a patient to visit a doctor is for the doctor to ask questions. Through this process, the doctor initially understands the development of the patient's symptoms and records them in the form of medical record texts. Next, the patient will further undergo a large number of routine item examinations and medical imaging examinations. Finally, the doctor makes a diagnosis based on the complex medical data generated from a series of examinations, determines whether the patient is suitable for Li's artificial liver treatment, and formulates a specific treatment plan. The process of the doctor's evaluation of complex medical data (such as medical records, imaging examination results, test indexes, etc.) is time-consuming and laborious, and requires high professional experience of the doctor. However, the condition of liver failure often develops rapidly and critically. Therefore, it is necessary to design a method that can automatically assist the doctor in evaluating medical data (such as medical record texts) and quickly give an analysis to assist the doctor in making a judgment to improve the treatment efficiency at the Li's artificial liver site.

[0004] In the process of text classification based on the understanding of Li's artificial liver medical records, the effect of the text representation (also known as text vectorization) stage has a great impact on the final classification performance. The traditional text vector generation model - the bag-of-words model is mainly based on statistics, and its calculation is relatively simple, but the effect is also relatively weak, such as the One-hot Representation, TF-IDF model, and latent semantic analysis model. With the development of deep learning, models based on neural networks have begun to appear in the field of natural language processing, and have also greatly improved the quality of text vectorization, such as the word2vec and glove models, etc. These models contain some semantic information to a certain extent, but do not consider the position information of words (the position of words also contains semantic information. For example, swapping the positions of two words may change the meaning of the sentence), and there is still a lot of potential to be explored. In order to generate better word vectors, it is necessary to optimize the text representation stage and integrate more and richer semantic information. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a multi-level semantic text classification method based on the medical records of Li's artificial liver, which is used to perform feature splicing and classification of multi-level semantics on the text data of each field of the chief complaint, current medical history, and past medical history in the medical records.

[0006] To solve the above technical problem, the present invention provides a prediction of suitable treatment of Li's artificial liver based on multi-level semantic medical record understanding, including the following process:

[0007] S1. Collect the medical record text data of the patient, including the text data of the current medical history, chief complaint, and past medical history;

[0008] S2. Input the respective word segmentation sequences and character segmentation sequences obtained after preprocessing the text data of the chief complaint and past medical history into the feature splicing and classification network based on multi-level semantic medical record understanding. After splitting the text data of the current medical history into sentence sequences according to full stops and then preprocessing, the obtained word segmentation sequence and character segmentation sequence of the current medical history are input into the feature splicing and classification network based on multi-level semantic medical record understanding. The feature splicing and classification network based on multi-level semantic medical record understanding includes a text representation network based on multi-level semantics, a feature extraction network, and a feature fusion and prediction network;

[0009] S2.1. After passing through the text representation network based on multi-level semantics, the word segmentation sequence and character segmentation sequence of the chief complaint after preprocessing obtain the chief complaint word vector matrix;

[0010] S2.2. After passing through the text representation network based on multi-level semantics, the word segmentation sequence and character segmentation sequence of the past medical history after preprocessing obtain the past medical history word vector matrix;

[0011] S2.3. After passing through the text representation network based on multi-level semantics and then average pooling, the word segmentation sequence and character segmentation sequence of the current medical history after preprocessing obtain the current medical history word vector matrix;

[0012] S2.4. Respectively extract features from the chief complaint word vector matrix, past medical history word vector matrix, and current medical history word vector matrix through the feature extraction network to obtain the chief complaint feature, past medical history feature, and current medical history feature respectively;

[0013] S2.5. Output the classification result through the feature fusion and prediction network for the chief complaint feature, past medical history feature, and current medical history feature.

[0014] As an improvement to the multi-level semantic text classification method based on the medical records of Li's artificial liver of the present invention:

[0015] The feature extraction network is a Bi-LSTM network;

[0016] The feature fusion and prediction network refers to performing a feature splicing operation and then passing through a fully connected layer.

[0017] As a further improvement of the multi-level semantic text classification method based on Li's artificial liver medical records of the present invention:

[0018] The preprocessing is as follows: perform word segmentation, character segmentation, stop word removal, word frequency statistics and vocabulary construction, character frequency statistics and character table construction, and establish a bidirectional hash mapping on the text data, and output the word segmentation sequence and the character segmentation sequence.

[0019] As a further improvement of the multi-level semantic text classification method based on Li's artificial liver medical records of the present invention:

[0020] The process of mapping the input text to a word vector matrix by the text representation network based on multi-level semantics is as follows:

[0021] Perform word segmentation and character segmentation on the input text, and replace the words and characters with the corresponding sequence numbers in the vocabulary and character table to obtain the word segmentation sequence and the character segmentation sequence; the word segmentation sequence is mapped to a word vector matrix with complex form position encoding fused through a word vector position encoding fusion network, and the character segmentation sequence is mapped to a character vector matrix with complex form position encoding fused through a character vector position encoding fusion network; then each word vector matrix and the corresponding character vector matrix are concatenated and passed through a fully connected layer to obtain a new word vector matrix E ∈ R k×D , k represents the length of the input text, D represents the dimension of a word vector mapped from a word, and the new word vector matrix E ∈ R k×D After passing through the word vector context information extraction network, for each new word vector matrix E ∈ R k×D Calculate and fuse the context information within the specified context window to obtain a word vector matrix.

[0022] As a further improvement of the multi-level semantic text classification method based on Li's artificial liver medical records of the present invention:

[0023] The processing process of the word vector position encoding fusion network is as follows:

[0024] Pass the word segmentation sequence through two embedding layers respectively to generate the amplitude vector r w and the angular frequency ω w of the position encoding, and calculate:

[0025]

[0026] where g(·): N → (F) D is a function that maps the sequence number of a word to a D-dimensional vector, and pos refers to the position sequence number of a character in a word;

[0027] Then concatenate the real part and the imaginary part of the formula (1) function [r w cos(ω w pos); r wsin(r w pos)] as the output of the word vector position encoding fusion network.

[0028] As a further improvement of the multi-level semantic text classification method based on the Li's artificial liver medical record of the present invention:

[0029] The processing process of the word vector position encoding fusion network is as follows:

[0030] Pass the character splitting sequence through two embedding layers respectively to generate the amplitude vector r c of the word vector and the angular frequency ω c of the position encoding, and calculate:

[0031]

[0032] Then splice the real part and the imaginary part of the formula (2) function [r c cos(ω c pos); r c sin(ω c pos)] as the output of the word vector position encoding fusion network.

[0033] As a further improvement of the multi-level semantic text classification method based on the Li's artificial liver medical record of the present invention:

[0034] The calculation process of the word vector context information extraction network is as follows:

[0035] (1) Define the context window weight matrix as W ∈ R (2×c+1)×D , where c is the window radius;

[0036] Then calculate the projection vector of the (i + t)-th word vector within the window range centered on the i-th word after being weighted by the weight matrix

[0037]

[0038] where, V i represents the vector corresponding to the i-th word of the input text, and W t (-c ≤ t ≤ c) represents the (c + t)-th row of W;

[0039] (2) Perform reallocation weight weighted summation calculation on to obtain a new word vector V i ' as the output:

[0040]

[0041] where, the reallocation weight r t (-c ≤ t ≤ c) is:

[0042] [r -c , …, r t , r c = softmax([V i-c V i , …, V i+t V i , …, V i+c V i ) (5).

[0043] As a further improvement of the multi - level semantic text classification method based on the Li's artificial liver medical records of the present invention:

[0044] The training and testing process of the feature splicing classification network based on multi - level semantic medical record understanding is as follows:

[0045] Obtain the text data of medical records from the hospital medical record system and perform classification annotation. Construct a data set from the three fields of chief complaint, current medical history, and past medical history in each medical record. Perform the above - mentioned pre - processing on the text of the chief complaint, current medical history, and past medical history in the data set, and randomly divide the pre - processed data set into a training set and a testing set according to a ratio of 4:1; then input the training set into the feature splicing classification network based on multi - level semantic medical record understanding to optimize the model performance, and use the testing set for verification.

[0046] The beneficial effects of the present invention are mainly reflected in:

[0047] 1. The present invention jointly analyzes the three fields of chief complaint, current medical history, and past medical history of the Li's artificial liver medical record text, and proposes a feature splicing classification network based on multi - level semantics according to the characteristics of the text in each field. Compared with the existing method of directly splicing the text of each field for classification, it makes full and effective use of the semantic information of the text in the three fields and improves the accuracy of prediction;

[0048] 2. According to the existing work on plural - form position encoding, fuse the position encoding information in the word vector. Improvedly fuse the plural - form position encoding (the position of the character in the word) with the character vector, combine the word vector and the corresponding character vector, and then extract the context semantic information between words through the word vector context information extraction network. Together, they constitute a text representation network containing multi - level semantics, construct word vectors with richer semantic features, and improve the accuracy of the text classification task on the basis of the unchanged baseline classification network.

[0049] It should be emphasized that the present invention only processes and analyzes the text of characters and words in the Li's artificial liver medical record, belonging to the scope of natural language processing technology, and does not output the results of disease diagnosis and treatment, nor does it belong to the scope of disease treatment. Brief Description of the Drawings

[0050] The following further elaborates on the specific implementation manners of the present invention in conjunction with the accompanying drawings.

[0051] Figure 1 It is a schematic structural diagram of a feature splicing classification network based on multi-level semantic medical record understanding of the present invention;

[0052] Figure 2 It is a schematic structural diagram of a multi-level semantic text representation network of the present invention;

[0053] Figure 3 It is a schematic structural diagram of a word and character vector position encoding fusion network;

[0054] Figure 4 It is a schematic structural diagram of a word vector context information extraction network. Specific implementation manners

[0055] The following further describes the present invention in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:

[0056] Embodiment 1. A multi-level semantic text classification method based on the Li's artificial liver medical record, as Figures 1 - 4 shown, the method includes the following steps:

[0057] S1. Collect and construct a medical record text data set of liver failure patients;

[0058] The data set is sourced from the medical record system export of a cooperative hospital, and experts annotate and classify each patient's medical record text as to whether it is suitable for Li's artificial liver treatment. The chief complaint, current medical history, and past medical history fields in each medical record are taken to construct the data set. The specific data set information is as shown in Table 1 below:

[0059] Table 1. Information of the medical record text data set of liver failure patients

[0060] dataset Medical record text dataset of liver failure patients Number of categories 2 Category details Suitable for Li's artificial liver treatment, not suitable for Li's artificial liver treatment Number of texts 1087

[0061] The constructed data set contains a total of 1087 medical record texts, among which 499 medical record texts are marked as suitable for Li's artificial liver treatment, and 588 medical record texts are marked as not suitable for Li's artificial liver treatment.

[0062] S2. Preprocess the data set;

[0063] Preprocess all the case texts in the data set: Use PkuSeg for word segmentation, character segmentation, stop word removal, word frequency statistics and construction of a word table, character frequency statistics and construction of a character table, and establish a bidirectional hash mapping between words / characters and their indexes in the word table / character table; After preprocessing the input text, a word segmentation sequence and a character segmentation sequence are obtained;

[0064] Then, the preprocessed dataset is randomly divided into a training set and a test set according to a ratio of 4:1.

[0065] S3. Construct a feature splicing classification network based on multi-level semantic medical record understanding;

[0066] The feature splicing classification network based on multi-level semantic medical record understanding includes a text representation network based on multi-level semantics, a feature extraction network, and a feature fusion prediction network.

[0067] S3.1. Construct a text representation network based on multi-level semantics

[0068] The text representation network based on multi-level semantics includes a word vector position encoding fusion network built based on existing complex position encoding work, a character vector position encoding fusion network proposed by the present invention, a word vector context information extraction network, and the output splicing and fusion of the first two networks, as Figure 2 shown.

[0069] The segmented sequence obtained after preprocessing is mapped into a word vector matrix integrating complex form position encoding through the word vector position encoding fusion network, and the character sequence is mapped into a character vector matrix integrating complex form position encoding through the character vector position encoding fusion network; then each word vector matrix and its corresponding character vector matrix are spliced, and dimensionality reduction is performed through a fully connected layer to obtain a new word vector matrix. Finally, the new word vector matrix passes through the word vector context information extraction network structure, and the fusion context information is calculated within a specified context window for each new word vector matrix to obtain the final required word vector matrix.

[0070] The text representation network based on multi-level semantics of the present invention builds a word vector position encoding fusion network based on existing complex form position encoding methods, and proposes two structures: a character vector position encoding fusion network and a word vector context information extraction network. On the basis of generating word vectors by traditional methods, it includes adding three levels of semantic information: the level of words and sentences, the level of words and characters, and the level of words and words. Specifically;

[0071] (1). Word vector position encoding fusion network;

[0072] Build a word vector position encoding fusion network based on existing complex form position encoding methods (Wang B, Zhao D, Lioma C, et al. Encoding word order in complex embeddings [J]. arXiv preprint arXiv: 1912.12333, 2019.), and its objective function is a vector-valued function with position as a variable:

[0073]

[0074] where \(g(\cdot): N\rightarrow (F)\) D is a function that maps the serial number of a word to a \(D\)-dimensional vector, and pos refers to the position serial number of a character in a word.

[0075] Segment the input text into words, and replace each word with the corresponding serial number in the word list. Pass the segmented sequence through two embedding layers respectively to generate two vectors \(r\) w and \(\omega\) w , \(r\) w represents the amplitude related to the word vector, and \(\omega\) w represents the angular frequency related to the position encoding. Combine the two vectors into the vector function in Equation (1), and then concatenate the real part and the imaginary part of this function \([r\) w \cos(\omega\) w pos); \(r\) w \sin(\omega\) w pos)] as the output of this network.

[0076] (2) Character vector position encoding fusion network;

[0077] Apply the complex position encoding to the character vectors. Here, the position encoding refers to the position encoding of a character in a word, and introduce the semantic information of the basic unit that makes up a word - the character.

[0078] The improved application of this method based on the complex position encoding is to fuse the position encoding on the character vectors. Importantly, this position encoding corresponds to the position of a character in a word, rather than the position of a character in a sentence. Based on the traditional word segmentation method, each word is segmented into characters, each character is mapped to a character vector, and the complex position encoding reflecting the position relationship of characters in a word is introduced. Its objective function is the same as the original complex position encoding, that is, Equation (1):

[0079]

[0080] where.g(\cdot): N\rightarrow (F) D is a function that maps the serial number of a word to a \(D\)-dimensional vector, and pos refers to the position serial number of a character in a word.

[0081] Segment the input text into words, segment each word into characters, and replace each character with the corresponding serial number in the character list. Pass the segmented sequence through two embedding layers respectively to generate two vectors \(r\) c and \(\omega\) c , \(r\) c represents the amplitude related to the character vector, and \(\omega\) c represents the angular frequency related to the position encoding. Combine the two vectors into the vector function in Equation (2), and then concatenate the real part and the imaginary part of this function \([r\) c \cos(\omega\)c pos); r c sin(ω c pos)] as the output of this network.

[0082] (3) Output fusion of the word and character vector position encoding fusion network;

[0083] The output of the word vector position encoding fusion network (word vector) corresponds to the outputs of multiple character vector position encoding fusion networks (character vectors). The word vector is concatenated with multiple character vectors and then reduced in dimension through a fully connected layer to obtain a new word vector matrix E ∈ R k ×D , where k represents the length of the input text, usually set to the length of the longest text after word segmentation, and D represents the dimension of the word vector mapped from a word.

[0084] (4) Word vector context information extraction network;

[0085] The word vector context information extraction network introduces semantic information at the word-to-word level, that is, the context semantic information between adjacent words. To utilize the semantic information of adjacent word contexts, this method defines a context window weight matrix W ∈ R (2×c+1)×D , where c is the window radius. Let V i represent the vector corresponding to the i-th word in the input text, and W t (-c ≤ t ≤ c) represents the (c + t)-th row of W, represents the projection vector of the (i + t)-th word vector within the window range centered on the i-th word after being weighted by the weight matrix, which is calculated by element-wise multiplication and activating with the tanh function:

[0086]

[0087] The new word vector V i ′ ∈ R D is obtained through weighted summation calculation with reallocated weights as the output of this network:

[0088]

[0089] Among them, the reallocated weight r t (-c ≤ t ≤ c) is calculated by taking the inner product of the central word vector V i with the word vectors within the context window range in sequence and normalizing through the softmax function, as follows:

[0090] [r -c , …, r t , r c = softmax([V i-c Vi , …, V i+t V i , …, V i+c V i ) (5)

[0091] For the pre-output word vector matrix E ∈ R k×D , perform the above operations centered on each word vector to obtain a new word vector matrix V i ′ (with the same dimension as the original word vector matrix), which contains the context semantic information between adjacent words.

[0092] S32. Build a feature splicing classification network based on multi-level semantic medical record understanding

[0093] For the feature splicing classification network based on multi-level semantic medical record understanding, as Figure 1 shown, the respective word segmentation sequences and character segmentation sequences obtained after preprocessing the text data of the chief complaint and the past history are respectively sent into the multi-level semantic-based text representation network constructed in step S31 to map the input text into the chief complaint word vector matrix and the past history word vector matrix. Since the length of the current history text is quite different from that of the chief complaint and past history texts, the text data of the current history needs to be first segmented into sentence sequences according to full stops and then preprocessed. Then, the word segmentation sequence and character segmentation sequence of the preprocessed current history are passed through the multi-level semantic-based text representation network constructed in step S31 and average pooling to map the input text into the current history word vector matrix; the obtained chief complaint-current history, past history-current history, and current history word vector matrices are respectively passed through a feature extraction network (Bi-LSTM network) to extract features to obtain chief complaint features, past history features, and current history features. Finally, the chief complaint features, past history features, and current history features are output through a feature fusion prediction network (referring to the full connection layer after feature splicing) the multi-level semantic text classification result based on the Li's artificial liver medical record.

[0094] The input of the dataset used in this method is three texts: the chief complaint, the current history, and the past history, while the input of a single text classification network is a single text. There are two ways to solve this mismatch. One way is to splice the three texts into one text, which matches the input of a single text classification network. However, due to the characteristics of the dataset used in the present invention, the information of shorter texts may be diluted. Another way is to increase the number of text classification networks, use a three-way text classification network to receive the three text inputs (the text classification network at this time is generally called a feature extraction network), and then splice the outputs of the three networks for final classification. After experimental comparison, the present invention selects the second method to build a feature splicing classification network for multi-level semantic medical record understanding, which can improve the prediction accuracy compared with the first method.

[0095] S4. The overall network structure parameters of the present invention

[0096] As shown in Table 2, the overall network structure parameters of the feature splicing classification network based on multi-level semantic medical record understanding of the present invention are shown, including the output sizes and parameters of each layer (128 in the table is the batch size):

[0097] Table 2

[0098]

[0099]

[0100] S5. Training and testing

[0101] Experimental environment configuration; the network training and verification of the present invention are carried out on a Linux server, accelerated by GPU, and developed based on the PyTorch framework. The specific configuration is as shown in Table 3 below:

[0102] Table 3

[0103] Name Environment configuration Operating system CentOS 7.9.2009 Processor 32 * E5 - 2667v4 @ 3.20GHz Graphics card Tesla P4 8GB(440.33) Memory 251G Development environment Python3.7.9 PyTorch1.8.1

[0104] Input the training set in step 2 into the feature splicing classification network based on multi-level semantic medical record understanding established in step S3, and use the Adam optimizer to optimize the network. The weight decay is 0.0001. The initial learning rate (lr) of the optimizer is set to 0.001, and the learning rate is decayed in the form of cosine decay. The window radius c in step S34 is used as a hyperparameter (set to 1, 2, 3, and 4 respectively) to conduct experimental verification to select the best parameters to make the model performance optimal. The input text data volume of each batch is 128, and 400 rounds of iteration are carried out, and the test set is used for verification.

[0105] The training and testing of the present invention use the medical record text dataset of the actual Li's artificial liver suitable treatment in the cooperative hospital, and use the combination of the chief complaint, current medical history, and past medical history text fields for classification. In view of the large difference in the text length distribution of the three fields, the feature splicing method rather than the direct text splicing method is selected for classification. In view of the long text of the current medical history, segmentation and pooling operations are carried out before and after the text representation respectively. The trained feature splicing classification network based on multi-level semantic medical record understanding is suitable for classifying and predicting whether the Li's artificial liver is suitable for treatment.

[0106] S6. Experimental verification of the feature splicing classification network based on multi-level semantic medical record understanding proposed by the present invention;

[0107] Experiment 1: Ablation experiment on the chief complaint text representation method

[0108] In order to verify that the designed text representation network based on multi-level semantics (as shown in Figure 2 ) can generate higher-quality vectors, thereby improving the classification accuracy of the feature splicing classification network suitable for multi-level semantic medical record understanding, the commonly used Bi-LSTM and TextCNN networks are selected as the baseline feature extraction networks for comparison. The classification accuracy differences with and without the text representation network based on multi-level semantics of the present invention are experimentally compared on the basis of the baseline feature extraction network, and the performance improvements of the baseline feature extraction networks combined with the word vector position encoding fusion, character vector position encoding fusion, and word vector context information extraction network of the present invention are experimentally observed.

[0109] When the baseline feature extraction network uses the structure of Bi-LSTM, it is as shown in Figure 1 . If TextCNN is used, all three Bi-LSTM network parts in Figure 1 are replaced with TextCNN, and the rest of the network structure remains unchanged. The original word vectors in Table 4 refer to the input tokenized sequence mapped into word vectors through an embedding layer. The connection methods of some or all of the structures of the rest of the text representation network refer to Figure 2 .

[0110] Group experiments are carried out based on different levels of semantic information and different baseline networks. The experimental evaluation indicators are classification accuracy (acc) and F1 value, and the experimental results are shown in Table 4 below:

[0111] Table 4

[0112]

[0113] After experiments, the performance of the text representation network based on multi-level semantics reaches the best when the window radius is set to 4. It can be seen from the results in Table 4 that the text representation network based on multi-level semantics of the present invention has good performance on different baseline classification networks, and the three levels of semantic information mentioned in the present invention all contribute to the final classification accuracy, proving that the text representation network based on multi-level semantics in the method of the present invention can indeed make more full use of the semantic information of the text to improve the classification performance.

[0114] Experiment 2: Comparative experiment of the feature splicing classification network for multi-level semantic medical record understanding of the present invention

[0115] In order to verify the effectiveness of each operation in the network structure of the present invention, a comparative experiment is designed for verification. The comparative network and results of the experiment are shown in Table 5. "Text representation" in Table 5 represents the text representation network based on multi-level semantics of the present invention, and "Bi-LSTM classification" means directly sending the input text into the Bi-LSTM network for classification (refer to Figure 1, its structure is the structure after removing the text representation method from the feature extraction network for the chief complaint or past history in the method of the present invention. "Text representation + Bi-LSTM classification" means using the outputs of the feature extraction networks for the chief complaint text and the past history text in the feature splicing classification network based on multi-level semantic medical record understanding for classification (in the method of the present invention, the three texts are respectively fed into their respective feature extraction networks, and then the outputs are spliced for classification). "Sentence splitting + text representation + pooling + Bi-LSTM classification" means using the output of the feature extraction network for the present illness text in the feature splicing classification network based on multi-level semantic medical record understanding of the present invention for classification. "Text direct splicing + Bi-LSTM classification" means splicing the chief complaint, present illness, and past history texts before preprocessing, and then feeding the spliced text into the Bi-LSTM network for classification (the structure after splicing the text is the same as the structure for Bi-LSTM classification of a single text).

[0116] Table 5

[0117] Input text Method Accuracy rate F1 value Chief complaint Bi - LSTM classification 0.8028 0.8230 Chief complaint Text representation + Bi - LSTM classification 0.8624 0.8819 History of present illness Bi - LSTM classification 0.8394 0.8560 History of present illness Text representation + Bi - LSTM classification 0.8624 0.8870 History of present illness Sentence splitting + text representation + pooling + Bi - LSTM classification 0.8807 0.8943 Past history Bi - LSTM classification 0.7798 0.7983 Past history Text representation + Bi - LSTM classification 0.8073 0.8235 Chief complaint, history of present illness, past history Text direct splicing + Bi - LSTM classification 0.8440 0.8682 Chief complaint, history of present illness, past history The method of the present invention 0.8899 0.9016

[0118] It can be seen from Table 5 that the performance of the model for directly splicing and classifying the chief complaint, present illness, and past history texts is higher than that of the single-text classification model, verifying the necessity and effectiveness of jointly using the texts of the three fields in the present invention; at the same time, compared with the traditional basic model of directly splicing and classifying texts, the multi-level semantic feature splicing classification model of the present invention also has a significant improvement in the classification performance of the model; and the improvement points of the present invention (multi-level semantic text representation, separate segmentation and pooling of the present illness) are also verified to be effective through experiments.

[0119] Experiment 3: Literature comparison experiment

[0120] To prove the effectiveness of the method of the present invention, an experiment was designed to compare with similar methods in the same field. The dataset used was the medical record text dataset of liver failure patients of the present invention. In the model comparing the literature method, the text directly spliced from the chief complaint, present illness, and past history was used as the input. The results are shown in Table 6:

[0121] Table 6

[0122] Network model Accuracy rate F1 value Literature 1 0.8486 0.8736 Literature 2 0.8578 0.8745 Literature 3 0.8716 0.8814 The method of the present invention 0.8899 0.9016

[0123] Literature 1 has a Bi-LSTM + CNN structure, and Literature 2 has a relatively common CNN + Bi-LSTM + Attention structure. The method of the present invention has the advantage of fusing semantic information at the word vector level and the character vector level compared with Literature 1 and Literature 2, and introduces the information of position encoding compared with Literature 3, having richer semantic information. In addition, the feature splicing method of the present invention also has advantages compared with directly splicing the three parts of the text. In summary, the effectiveness of the method of the present invention is proved.

[0124] Reference 1: Abu Kwaik K, Saad M, Chatzikyriakidis S, et al. LSTM-CNN deeplearning model for sentiment analysis of dialectal Arabic[C] / / Arabic LanguageProcessing:FromTheory to Practice:7th International Conference,ICALP 2019,Nancy,France,October 16–17,2019,Proceedings 7.Springer InternationalPublishing,2019:108-121.

[0125] Reference 2: See Jang B, Kim M, Harerimana G, et al. Bi-LSTM model to increase accuracy in text classification: Combining Word2vec CNN and attentionmechanism[J]. Applied Sciences, 2020, 10(17):5841.

[0126] Reference 3: Zhang Mohan. CNN-LSTM short text classification based on word mixed vectors[J]. Information Technology and Informatization, 2019, No.226(01):77-80.

[0127] Finally, it should be noted that the above examples are only some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by a person skilled in the art should be considered as the protection scope of the present invention.

Claims

1. A multi-level semantic text classification method based on the medical records of Li's artificial liver, characterized in that The process is as follows: S1. Collect the medical record text data of the patient, including the text data of the current medical history, chief complaint, and past medical history; S2. Input the respective word segmentation sequences and character segmentation sequences obtained after preprocessing the text data of the chief complaint and past medical history into the feature splicing classification network based on multi-level semantic medical record understanding. After splitting the text data of the current medical history into sentence sequences according to full stops and then preprocessing, input the obtained word segmentation sequence and character segmentation sequence of the current medical history into the feature splicing classification network based on multi-level semantic medical record understanding. The feature splicing classification network based on multi-level semantic medical record understanding includes a text representation network based on multi-level semantics, a feature extraction network, and a feature fusion prediction network; S2.

1. After passing through the text representation network based on multi-level semantics, the word segmentation sequence and character segmentation sequence of the chief complaint after preprocessing obtain the chief complaint word vector matrix; S2.

2. After passing through the text representation network based on multi-level semantics, the word segmentation sequence and character segmentation sequence of the past medical history after preprocessing obtain the past medical history word vector matrix; S2.

3. After passing through the text representation network based on multi-level semantics and then through average pooling, the word segmentation sequence and character segmentation sequence of the current medical history after preprocessing obtain the current medical history word vector matrix; S2.

4. Respectively extract features from the chief complaint word vector matrix, past medical history word vector matrix, and current medical history word vector matrix through the feature extraction network to obtain the chief complaint feature, past medical history feature, and current medical history feature respectively; S2.

5. Output the classification result through the feature fusion prediction network with the chief complaint feature, past medical history feature, and current medical history feature; The process by which the text representation network based on multi-level semantics maps the input text into a word vector matrix is: Tokenize and character-segment the input text, and replace the words and characters with the corresponding numbers in the vocabulary and character list to obtain the tokenized sequence and character-segmented sequence; the tokenized sequence is mapped by the token vector position encoding fusion network into a token vector matrix with complex-form position encoding fused, and the character-segmented sequence is mapped by the character vector position encoding fusion network into a character vector matrix with complex-form position encoding fused; then each token vector matrix and the corresponding character vector matrix are concatenated and passed through a fully connected layer to obtain a new token vector matrix E ∈ R k×D , where k represents the length of the input text and D represents the dimension of the token vector mapped from a word, and the new token vector matrix E ∈ R k×D is passed through the token vector context information extraction network to calculate the fused context information within a specified context window for each new token vector matrix E ∈ R k×D to obtain the token vector matrix; The processing process of the word vector position encoding fusion network is: Pass the tokenized sequences through two embedding layers respectively to generate the amplitude vector r of the word vectors w and the angular frequency ω of the positional encoding w , and calculate: where \(g(\cdot): \mathbb{N} \to \mathcal{F}\) D is a function that maps the serial number of a word to a \(D\)-dimensional vector, and pos refers to the position serial number of a character in a word; The real and imaginary parts of the recombined (1) function [r w cos(ω w pos); r w sin(ω w pos)] are used as the output of the word vector position encoding fusion network.

2. The multi-level semantic text classification method based on the Li's artificial liver medical record according to claim 1, characterized in that: The feature extraction network is a Bi-LSTM network; The feature fusion prediction network refers to performing a feature splicing operation and then passing through a fully connected layer.

3. The multi-level semantic text classification method based on the Li's artificial liver medical record according to claim 2, characterized in that: The preprocessing is: performing word segmentation, character segmentation, removing stop words, counting word frequencies and constructing a word table, counting character frequencies and constructing a character table, and establishing a bidirectional hash mapping on the text data, and outputting a word segmentation sequence and a character segmentation sequence.

4. The multi-level semantic text classification method based on the Li's artificial liver medical record according to claim 3, characterized in that: The processing process of the character vector position encoding fusion network is: Pass the segmented character sequences through two embedding layers respectively to generate the amplitude vector r of the character vectors c and the angular frequency ω of the positional encoding c And calculate: The real and imaginary parts of the recombined (2) function [r c cos(ω c pos); r c sin(ω c pos)] are used as the output of the word vector position encoding fusion network.

5. The multi-level semantic text classification method based on the Li's artificial liver medical record according to claim 4, characterized in that: The calculation process of the word vector context information extraction network is: (1) Define the context window weight matrix as \(W\in\mathbb{R}\) (2×c+1)×D , where \(c\) is the window radius; Then calculate the projection vector of the word vector of the (i + t)-th word within the window range centered on the i-th word after being weighted by the weight matrix Among them, V i represents the vector corresponding to the i-th word of the input text, and W t (-c ≤ t ≤ c) represents the (c + t)-th row of W; (2) For perform weighted summation calculation with reallocated weights to obtain a new word vector V i ′ as the output: Among them, the reallocation weight r t (-c ≤ t ≤ c) is as follows: [r -c , …, r t , r c = softmax([V i-c V i , …, V i+t V i , …, V i+c V i ) (5).

6. The multi-level semantic text classification method based on the Li's artificial liver medical record according to claim 5, characterized in that: The training and testing process of the feature splicing classification network based on multi-level semantic medical record understanding is: Obtain the text data of medical records from the hospital medical record system and classify and label them. Construct a data set from the three fields of the chief complaint, current medical history, and past medical history in each medical record. Perform the above-mentioned preprocessing on the text of the chief complaint, current medical history, and past medical history in the data set, and randomly divide the preprocessed data set into a training set and a test set according to a ratio of 4:1; then input the training set into the feature splicing classification network based on multi-level semantic medical record understanding to optimize the model performance, and use the test set for verification.

Citation Information

Patent Citations

  • A hierarchical BiLSTM Chinese electronic medical record disease code labeling method capable of enhancing semantic representation

    CN109697285A

  • Significance-based prediction from unstructured text

    US20230061731A1