A Medical Text Classification Method Based on Knowledge Graph and Multi-Head Pooling Graph Convolution

Through the method based on knowledge graph and multi-pole pooling graph convolution, the word set of medical text is expanded and the self-attention mechanism is used to classify, which solves the problems of semantic sparseness and feature sparseness in medical text classification, significantly improving the classification accuracy.

CN116975282BActive Publication Date: 2025-07-22BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310787599.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-07-22
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

The existing medical text classification methods are difficult to effectively identify professional terms in medical texts, resulting in poor feature extraction results and the semantic sparseness of short medical texts, resulting in low classification accuracy.

Method used

The text extension method based on knowledge graph is adopted to expand the word set by constructing the smallest Steiner tree, and text classification is performed using a multi-pooled graph convolution network, combining the self-attention mechanism to improve feature extraction and classification accuracy.

Benefits of technology

Effectively capture the implicit semantic information of medical texts, improve the classification accuracy of short medical texts, solve the problems of semantic sparseness and feature sparseness, and improve the classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975282B_ABST
    Figure CN116975282B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical text classification method based on a knowledge graph and multi-head pooling graph convolution, belonging to the technical fields of text classification and natural language processing. First, text preprocessing is performed, that is, a word set is constructed from the text in the dataset. Secondly, text expansion is carried out, that is, the word set S i is expanded by using a text expansion method based on a knowledge graph. Then, a text-word heterogeneous graph is constructed according to the expanded word set. Finally, the nodes of the text-word heterogeneous graph are classified to obtain the text classification result. The feature of this text expansion method is to use the minimum Steiner tree modeling to expand the text and capture the implicit semantic information of the text, so as to solve the problem of semantic sparsity of medical short texts. By mining the relevance between words and texts through the text-word heterogeneous graph and based on the multi-head pooling graph convolution of the self-attention mechanism, the accuracy of text classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a medical text classification method based on a knowledge graph and multi-head pooling graph convolution, and belongs to the technical fields of text classification and natural language processing. Background Art

[0002] Text classification is an important research topic in the field of natural language processing. The text classification task refers to the process of mapping a text or document to one or more categories given classification categories or preset classification categories. Text classification methods include shallow learning methods and deep learning methods. The classification methods based on shallow learning mainly include five methods: classification based on probabilistic graphs, classification based on support vector machines (Support Vector Machine, abbreviated as SVM), classification based on the K-nearest neighbor algorithm (K-Nearest Neighbor, abbreviated as KNN), classification based on decision trees, and classification based on ensemble learning.

[0003] The classification method based on probabilistic graphs is the integration of probability theory and graph theory, including Bayesian networks, hidden Markov networks, maximum entropy, and conditional random fields, etc.

[0004] For example, Yi et al. proposed a medical text classification method based on a hidden Markov model in the literature "A Hidden Markov Model-based Text Classification of Medical Documents" (Journal of Information Science, 2009), using medical subject headings as prior knowledge and adopting a hidden Markov model for text classification. Joachims et al. proposed a text classification method based on support vector machines in the literature "Text Categorization with Support Vector Machines: Learning with Many Relevant Features" (ECML, 1998). First, a sample set is constructed through preprocessing, feature extraction, and normalization processing, and then a support vector machine is used to classify the text.

[0005] The text classification methods based on deep learning mainly include convolutional neural networks, recurrent neural networks, long short-term memory networks, etc. For example, Yoon Kim proposed a text classification method based on convolutional neural networks in the literature "Convolutional Neural Networks for Sentence Classification" (In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2014). Word vectors were introduced to vectorize the text, and then it was input into the convolutional neural network for classification. Sun Hong et al. proposed a Chinese text classification model that combines the pre-trained model BERT (Bidirectional Encoder Representation from Transformers) and the attention mechanism in the literature "Chinese Text Classification Combining BERT Word Embedding and Attention Mechanism" (Mini-Micro Systems, 2022). Yao et al. proposed a text classification method based on graph convolutional neural networks in the literature "Graph Convolutional Networks for Text Classification" (In Proceedings of the AAAI Conference on Artificial Intelligence, 2019). First, a heterogeneous corpus graph with documents and words as nodes was constructed, then the graph convolutional network was used to learn node representations, and finally text classification was achieved.

[0006] The main problems of existing medical text classification methods are as follows: First, it is difficult for existing medical text classification methods to effectively identify a large number of medical domain-specific terms in medical texts, which affects the effect of feature extraction; Second, there is a problem of semantic sparsity in short medical texts in the medical field, and it is difficult to capture the semantic information of the text, resulting in low accuracy of text classification. Summary of the Invention

[0007] The purpose of the present invention is to solve the problems of semantic sparsity of short medical texts and low accuracy of short medical text classification, and a medical text classification method based on knowledge graph and multi-head pooling graph convolution is proposed.

[0008] This method first adopts a text expansion method based on the knowledge graph to expand the semantics of the text and then extract features to mine the implicit semantic information of the text, and then uses a multi-head pooling graph convolutional network based on the self-attention mechanism to classify the text.

[0009] To achieve the above object, the present invention adopts the following technical solutions.

[0010] A medical text classification method based on a knowledge graph and a multi-head pooling graph convolutional network, comprising the following steps:

[0011] Step 1: Text preprocessing, that is, constructing a word set from the text in the dataset;

[0012] Step 1.1: Data cleaning, that is, screening out empty data and invalid data in the dataset;

[0013] Empty data refers to data where the text does not contain specific content. For example, the text "", which only contains spaces.

[0014] Invalid data refers to data where the text content or label content is missing. For example, the text "[text: 'The patient's age is between 18 and 35 years old', label: "]" is missing a label and belongs to invalid data.

[0015] Step 1.2: Segment the text;

[0016] Use a word segmentation tool to segment the medical text. Incorporate the drug name vocabulary corpus and the full International Classification of Diseases ICD corpus into the dictionary of the word segmentation tool. Let for the i-th medical text T i , perform word segmentation and construct a word set S i .

[0017] Step 1.3: Remove stop words from the word set;

[0018] According to the stop word list, remove the modal particles, conjunctions, auxiliary words, and punctuation marks in the word set S i ;

[0019] Step 2: Text expansion, that is, expanding the word set S i using a text expansion method based on a knowledge graph;

[0020] Step 2.1: Construct a keyword set;

[0021] For the words in the word set S i , if they appear in the entity set of the knowledge graph, add them to the keyword set keyword(S i ).

[0022] Step 2.2: Query the minimum Steiner tree of the keyword set;

[0023] Query the minimum Steiner tree SteinerTree(S i ) of the keyword set in the knowledge graph;

[0024] Step 2.3: Construct an extended word set;

[0025] For the node set in SteinerTree(S i ), which is the word set, perform a union operation with the word set S i to construct the extended word set Extend(S i );

[0026] Step 3: Construct a text-word heterogeneous graph based on the extended word set;

[0027] Step 3.1: Construct a text vocabulary library, which is composed of all the words in the extended word set Extend(S i );

[0028] For the text sequences T1, T2, …, T n , construct the extended word sets Extend(S1), Extend(S2), …, Extend(S n ), and the text vocabulary library is Extend(S1) ∪ Extend(S2) ∪ … ∪ Extend(S n ).

[0029] Step 3.2: Construct a text-word heterogeneous graph and calculate the distance values of word-text edges and word-word edges;

[0030] Construct a text node from a word set, construct a word node for each word in the text vocabulary library, further construct a "word-text edge", that is, an edge connecting a word node and a text node, and construct a "word-word edge", that is, an edge connecting a word node and a word node.

[0031] Construct the adjacency matrix A C of the text-word heterogeneous graph, as shown in formula (1):

[0032]

[0033] Step 3.2A: Calculate the distance value of the word-text edge;

[0034] For the edge between a word node and a text node, use the BM25 algorithm to calculate the distance value of the edge, as shown in formula (2):

[0035]

[0036] Among them, TF i is the word frequency of the word w i in the document d. L d represents the length of the document d, L avg represents the average length of all documents, N is the total number of documents in the corpus, N i is the total number of documents containing the word w i , and k and b are coefficients.

[0037] Step 3.2B: Calculate the distance value of the word-word edge;

[0038] For the edge between word nodes, the point mutual information (PMI) is used to calculate the distance value of the edge, as shown in formula (3):

[0039]

[0040] where p(w i , w j ) is the number of times that words w i and w j appear simultaneously divided by the square of the total number of words, representing the probability that words w i and w j appear simultaneously. p(w i ) is the number of times that word w i appears divided by the total number of words, representing the probability that word w i appears. p(w j ) is the number of times that word w j appears divided by the total number of words, representing the probability that word w j appears.

[0041] Step 3.3: Construct the adjacency matrix of the text-word heterogeneous graph;

[0042] The adjacency matrices A T and A P of the text-word heterogeneous graph are shown in formulas (4) and (5):

[0043]

[0044]

[0045] Step 4: Classify the nodes of the text-word heterogeneous graph to obtain the text classification result;

[0046] Use the multi-head pooling graph convolutional neural network based on the self-attention mechanism to classify the nodes of the text-word heterogeneous graph;

[0047] Step 4.1: Perform graph convolution operation on the adjacency matrix A C of the text-word heterogeneous graph;

[0048] Enter the first layer of graph convolution, and perform the graph convolution operation on the adjacency matrix A C as shown in formula (6);

[0049] Z (0) = A C XW (0), (6)

[0050] Among them, X is the feature matrix of the input node, and its initial value is obtained randomly or by using the pre-trained model BERT. W (0) is a trainable parameter matrix;

[0051] Step 4.2: Perform multi-head pooling operations on the adjacency matrices A T and A P respectively;

[0052] Perform multi-head pooling operations on the adjacency matrices A T and A P respectively, as shown in formulas (7) and (8);

[0053]

[0054]

[0055] Among them, RS is an n-dimensional random sampling operation, TopK is used to obtain the corresponding scores of the top K nodes with the highest scores and set the scores of the remaining nodes to zero, softmax is a normalization exponential function, and the activation function σ is the tanh function. W (0) is a trainable parameter matrix;

[0056] Step 4.3: Residual connection and weighted operation;

[0057] Perform residual connection and weighted operation on the results of Step 4.1 and Step 4.2, as shown in formulas (9), (10), and (11):

[0058]

[0059]

[0060]

[0061] Among them, mean is the operation of calculating the average value, max is the operation of calculating the maximum value, ⊙ is the element-wise product, ReLU is the activation function, softmax is the normalization function, and Z (0) is the output of Step 4.1.

[0062] Step 4.4: Output the text classification result;

[0063] Enter the second-layer graph convolution to obtain the final output of the multi-head pooling graph convolution network, that is, the text classification result matrix, as shown in formula (12):

[0064] Z (2) =softmax(A C Z (1)W (l) ), (12)

[0065] Among them, softmax is a normalization function, and Z (1) is the output of step 4.3, and W (l) is a trainable parameter matrix.

[0066] Beneficial effects

[0067] In view of the medical text classification problem, the present invention proposes a medical text classification method based on a knowledge graph and multi-head pooling graph convolution. Compared with the prior art, it has the following beneficial effects:

[0068] 1. The method proposes a text expansion method based on a knowledge graph for the problem of semantic sparsity of medical short texts. The core idea of this method includes: first, obtaining the word set of medical texts through part-of-speech filtering, and then, through knowledge graph query, generating a minimum Steiner tree as the expansion of the original text. This minimum Steiner tree expands the related entities of the entities in the medical text on the basis of the original text. The feature of this text expansion method is that it uses the minimum Steiner tree modeling to expand the text and capture the implicit semantic information of the text, thereby solving the problem of semantic sparsity of medical short texts.

[0069] 2. The method extracts features by using a method based on a text-word heterogeneous graph, calculates the word-word edge distance by using point mutual information, and calculates the word-text edge distance by using BM25, so as to be able to mine the relevance between words and texts, thereby improving the accuracy of text classification.

[0070] 3. The method uses a multi-head pooling graph convolution network model based on a self-attention mechanism for text classification. The multi-head pooling method based on the self-attention mechanism is introduced to highlight the word nodes that contribute greatly to the sentence semantics, and the multi-head pooling is used to avoid ignoring the implicit important word nodes, which helps to solve the problem of feature sparsity of medical short texts, thereby improving the accuracy of text classification. Description of the drawings

[0071] Figure 1 It is a schematic flowchart of a medical text classification method based on a knowledge graph and multi-head pooling graph convolution of the present invention. Specific embodiments

[0072] The medical text classification system based on the method of the present invention uses Pycharm as the development tool and Python as the development language. The preferred embodiments of a medical text classification method based on a knowledge graph and multi-head pooling graph convolution of the present invention will be described in detail below in conjunction with the embodiments.

[0073] Embodiment 1

[0074] This embodiment describes the process of a medical text classification method based on a knowledge graph and multi-head pooling graph convolution according to the present invention, as Figure 1 shown. It can be seen from Figure 1 that it specifically includes the following steps:

[0075] Step 1: Text preprocessing, that is, constructing a word set from the text in the dataset;

[0076] Step 1 specifically includes the following sub-processes:

[0077] Step 1.1: Data cleaning, that is, screening out empty data and invalid data in the dataset;

[0078] Empty data refers to data where the text does not contain specific content. For example, the text "", which only contains spaces.

[0079] Invalid data refers to data where the text content or label content is missing. For example, the text "[text: 'The patient's age is between 18 and 35 years old', label: ”]" is missing a label and belongs to invalid data.

[0080] Step 1.2: Tokenize the text;

[0081] Use a tokenization tool to tokenize the medical text. Add the drug name vocabulary corpus and the full International Classification of Diseases ICD corpus to the dictionary of the tokenization tool. Suppose for the i-th medical text T i , tokenize it and construct a word set S i .

[0082] In this embodiment, the jieba tokenizer is used to tokenize the medical text. Add the drug name vocabulary corpus and the full International Classification of Diseases ICD corpus to the dictionary of the jieba tokenizer.

[0083] In this embodiment, for the medical text "9. Complicated with other diseases such as tuberculosis, bronchiectasis, heart failure, etc. causing asthma;", the word set "['complicated with', 'disease', 'tuberculosis', 'bronchiectasis', 'heart failure', 'asthma']" is constructed.

[0084] Step 1.3: Remove stop words from the word set;

[0085] According to the stop word list, remove the modal particles, conjunctions, auxiliary words, and punctuation marks from the word set S i ;

[0086] Step 2: Text expansion, that is, expand the word set S i using a text expansion method based on a knowledge graph;

[0087] Step 2 includes the following sub-processes:

[0088] Step 2.1: Construct a keyword set;

[0089] For the words in the word set S i if they appear in the entity set of the knowledge graph, then add them to the keyword set keyword(S i ).

[0090] For the word set "['merge', 'disease', 'pulmonary tuberculosis', 'bronchiectasis', 'cardiac insufficiency', 'asthma']", the constructed keyword set is "['bronchiectasis', 'asthma']".

[0091] Step 2.2: Query the minimum Steiner tree of the keyword set;

[0092] Query the minimum Steiner tree SteinerTree(S i ) of the keyword set in the knowledge graph;

[0093] In this embodiment, put "['bronchiectasis', 'asthma']" into the knowledge graph for minimum Steiner tree query, and the obtained minimum Steiner tree node set is "['bronchiectasis', 'disease', 'asthma']".

[0094] Step 2.3: Construct an extended word set;

[0095] Perform a set union operation on the node set (i.e., the word set) in SteinerTree(S i ) and the word set S i to construct the extended word set Extend(S i );

[0096] In this embodiment, perform a union operation on "['bronchiectasis', 'disease', 'asthma']" and the word set "['merge', 'disease', 'pulmonary tuberculosis', 'bronchiectasis', 'cardiac insufficiency', 'asthma']" to construct the extended word set ['merge', 'disease', 'pulmonary tuberculosis', 'bronchiectasis', 'cardiac insufficiency', 'asthma', 'disease'].

[0097] Step 3: Construct a text-word heterogeneous graph according to the extended word set;

[0098] Step 3 includes the following sub-processes:

[0099] Step 3.1: Construct a text vocabulary library, which is composed of all the words in the extended word set Extend(S i );

[0100] For the text sequences T1, T2, …, T n , construct the extended word sets Extend(S1), Extend(S2), …, Extend(Sn ) and the text vocabulary library is Extend(S1) ∪ Extend(S2) ∪ … ∪ Extend(S n ).

[0101] Step 3.2: Construct a text-word heterogeneous graph and calculate the distance values of word-text edges and word-word edges;

[0102] Construct a text node from a word set, construct a word node for each word in the text vocabulary library, further construct a "word-text edge", that is, an edge connecting the word node and the text node, and construct a "word-word edge", that is, an edge connecting the word node and the word node.

[0103] Construct the adjacency matrix A of the text-word heterogeneous graph C , as shown in formula (1):

[0104]

[0105] Step 3.2A: Calculate the distance value of the word-text edge;

[0106] For the edge between the word node and the text node, use the BM25 algorithm to calculate the distance value of the edge, as shown in formula (2):

[0107]

[0108] Among them, TF i is the word frequency of word w i in document d. L d represents the length of document d, L avg represents the average length of all documents, N is the total number of documents in the corpus, N i is the total number of documents containing word w i , and k and b are coefficients.

[0109] Step 3.2B: Calculate the distance value of the word-word edge;

[0110] For the edge between the word node and the word node, use the point mutual information PMI (Point Mutual Information) to calculate the distance value of the edge, as shown in formula (3):

[0111]

[0112] Among them, p(w i , w j ) is the number of times that words w i and w j appear simultaneously divided by the square of the total number of words, indicating that words w i and w jThe probability of co-occurrence. p(w i ) is the number of occurrences of the word w i divided by the total number of words, representing the probability of the word w i ) is the number of occurrences of the word w j divided by the total number of words, representing the probability of the word w j ) is the number of occurrences of the word w j divided by the total number of words, representing the probability of the word w

[0113] Step 3.3: Construct the adjacency matrix of the text-word heterogeneous graph;

[0114] The adjacency matrix A of the text-word heterogeneous graph T and A P , as shown in Formulas (4) and (5):

[0115]

[0116]

[0117] Step 4: Classify the nodes of the text-word heterogeneous graph to obtain the text classification result;

[0118] Use the multi-head pooling graph convolutional neural network based on the self-attention mechanism to classify the nodes of the text-word heterogeneous graph;

[0119] Step 4 includes the following sub-processes:

[0120] Step 4.1: Perform graph convolution operation on the adjacency matrix A C of the text-word heterogeneous graph;

[0121] Enter the first layer of graph convolution and perform graph convolution operation on the adjacency matrix A C as shown in Formula (6);

[0122] Z (0) = A C XW (0) , (6)

[0123] where X is the feature matrix of the input nodes, and its initial value is obtained randomly or by using the pre-trained model BERT, and W (0) is the trainable parameter matrix;

[0124] Step 4.2: Perform multi-head pooling operations on the adjacency matrix A T and A P of the text-word heterogeneous graph respectively;

[0125] Perform multi-head pooling operations on the adjacency matrix A T and A P respectively, as shown in Formulas (7) and (8);

[0126]

[0127]

[0128] Among them, RS is an n-dimensional random sampling operation, TopK is used to obtain the corresponding scores of the top K nodes with the highest scores and set the scores of the remaining nodes to zero, softmax is a normalization exponential function, the activation function σ is the tanh function, and W (0) is a trainable parameter matrix;

[0129] Step 4.3: Residual connection and weighted operation;

[0130] Perform a residual connection and a weighted operation on the results of Step 4.1 and Step 4.2, as shown in Formulas (9), (10), and (11):

[0131]

[0132]

[0133]

[0134] Among them, mean is the operation of calculating the average value, max is the operation of calculating the maximum value, ⊙ is the element-wise product, ReLU is the activation function, softmax is the normalization function, and Z (0) is the output of Step 4.1.

[0135] Step 4.4: Output the text classification result;

[0136] Enter the second-layer graph convolution to obtain the final output of the multi-head pooling graph convolutional network, that is, the text classification result matrix, as shown in Formula (12):

[0137] Z (2) = softmax(A C Z (1) W (l) ), (12)

[0138] Among them, softmax is the normalization function, Z (1) is the output of Step 4.3, and W (l) is a trainable parameter matrix.

[0139] To illustrate the medical text classification effect of the present invention, this experiment was carried out under the same conditions, and four methods were compared using the same training set, validation set, and test set.

[0140] The first method is a text classification method based on a convolutional neural network. By tokenizing and constructing a text vector representation, a convolutional neural network is used for text classification. The second method is a text classification method based on a recurrent neural network. By tokenizing and constructing a text vector representation, a recurrent neural network is used for text classification. The third method is a method based on a graph convolutional neural network. The fourth method is the medical text classification method of the present invention.

[0141] The evaluation metrics used are: Accuracy, and the calculation method is shown in formula (13), where N represents the total number of samples, and N right represents the number of samples correctly classified.

[0142]

[0143] The medical text classification results are as follows: the accuracy of the first method is 28.48%, the accuracy of the second method is 32.35%, the accuracy of the third method is 58.77%, and the accuracy of the method of the present invention is 60.20%. Experiments have shown the effectiveness of the proposed medical text classification method based on a knowledge graph and multi-head pooling graph convolution.

[0144] The above are the preferred embodiments of the present invention, and the present invention should not be limited to the content disclosed in this embodiment and the drawings. Any equivalent or modified implementation completed without departing from the spirit disclosed by the present invention falls within the protection scope of the present invention.

Claims

1. A medical text classification method based on knowledge graph and multi-head pooling graph convolution, characterized in that: It includes the following steps: Step 1: Text preprocessing, that is, constructing a set of words from the text in the dataset; Step 2: Text expansion, that is, expanding the word set S by using a text expansion method based on a knowledge graph i for expansion; Step 2.1: Constructing a keyword set; For the word set S i For the words in it, if they appear in the entity set of the knowledge graph, then add them to the keyword set keyword(S i ); Step 2.2: Querying the minimum Steiner tree of the keyword set; Query the minimum Steiner tree of the keyword set in the knowledge graph SteinerTree(S i ) Step 2.3: Constructing an extended word set; For the node set in SteinerTree(S i ), that is, the word set, perform a union operation with the word set S i to construct the extended word set Extend(S i ); Step 3: Constructing a text-word heterogeneous graph according to the extended word set; Step 4: Classifying the nodes of the text-word heterogeneous graph to obtain the text classification result; Using a multi-head pooling graph convolutional neural network based on the self-attention mechanism to classify the nodes of the text-word heterogeneous graph; Step 4.1: Perform graph convolution operation on the adjacency matrix A of the text-word heterogeneous graph C ; Enter the first layer of graph convolution and perform operations on the adjacency matrix A C Perform graph convolution operations as shown in formula (6); Z (0) = A C XW (0) , (6) Among them, X is the feature matrix of the input node, and its initial value is obtained randomly or by using the pre-trained model BERT, and W (0) is a trainable parameter matrix; Step 4.2: Perform multi-head pooling operations on the adjacency matrices A T and A P of the text-word heterogeneous graph respectively; For the adjacency matrix A T and A P perform multi-head pooling operations respectively, as shown in formulas (7) and (8); Among them, RS is an n-dimensional random sampling operation, TopK is used to obtain the corresponding scores of the top K nodes with the highest scores and set the scores of the remaining nodes to zero, softmax is the normalized exponential function, the activation function σ is the tanh function, and W (0) is a trainable parameter matrix; Step 4.3: Residual connection and weighted operation; Performing residual connection and weighted operation on the results of Step 4.1 and Step 4.2, as shown in Formulas (9), (10) and (11): where mean is the average operation, max is the maximum operation, ⊙ is the element-wise product, ReLU is the activation function, softmax is the normalization function, and Z (0) is the output of step 4.1; Step 4.4: Outputting the text classification result; Entering the second-layer graph convolution to obtain the final output of the multi-head pooling graph convolutional network, that is, the text classification result matrix, as shown in Formula (12): Z (2) = softmax(A C Z (1) W (l) ), (12) Among them, softmax is the normalization function, and Z (1) is the output of step 4.3, and W (l) is the trainable parameter matrix.

2. The medical text classification method based on a knowledge graph and multi-head pooling graph convolution according to claim 1, characterized in that: Specifically, Step 1 is as follows: Step 1.1: Data cleaning, that is, screening out the empty data and invalid data in the dataset; Empty data refers to the data whose text does not contain specific content; Invalid data is the data with missing text content or label content; Step 1.2: Segmenting the text; Tokenize medical texts using a tokenization tool; add the drug name vocabulary corpus and the entire International Classification of Diseases (ICD) corpus to the dictionary of the tokenization tool; assume that for the i-th medical text T i , tokenize it and construct a set of words S i ; Step 1.3: Removing the stop words in the word set; Remove the modal particles, conjunctions, auxiliary words, and punctuation marks in the word set S according to the stop word list. i ​ 3. A medical text classification method based on a knowledge graph and multi-head pooling graph convolution according to claim 1, characterized in that: Step 3: Constructing a text-word heterogeneous graph according to the extended word set, specifically: Step 3.1: Construct a text vocabulary, which is composed of all the words in the extended word set Extend(S i ); For text sequences T1, T2, …, T n , construct extended word sets Extend(S1), Extend(S2), …, Extend(S n ), and the text vocabulary is Extend(S1) ∪ Extend(S2) ∪ … ∪ Extend(S n ); Step 3.2: Constructing a text-word heterogeneous graph and calculating the distance values of the word-text edges and word-word edges; Constructing a text node from a set of words, constructing each word in the text vocabulary as a word node, further constructing a "word-text edge", that is, the edge connecting the word node and the text node, and constructing a "word-word edge", that is, the edge connecting the word node and the word node; Construct the adjacency matrix A of the text-word heterogeneous graph C , as shown in formula (1): Step 3.2A: Calculating the distance value of the word-text edge; For the edge between the word node and the text node, using the BM25 algorithm to calculate the distance value of the edge, as shown in Formula (2); Among them, TF i is the word frequency of the word w i in the document d; L d represents the length of the document d, and L avg represents the average length of all documents. N is the total number of documents included in the corpus, and N i is the total number of documents containing the word w i , and k and b are coefficients; Step 3.2B: Calculating the distance value of the word-word edge; For the edge between the word node and the word node, using the point mutual information (PMI) to calculate the distance value of the edge, as shown in Formula (3); Among them, p(w i , w j ) is the number of times that words w i and w j appear simultaneously divided by the square of the total number of words, representing the probability that words w i and w j appear simultaneously; p(w i ) is the number of times that word w i appears divided by the total number of words, representing the probability that word w i appears; p(w j ) is the number of times that word w j appears divided by the total number of words, representing the probability that word w j appears; Step 3.3: Constructing the adjacency matrix of the text-word heterogeneous graph; Adjacency matrix A of the text-word heterogeneous graph T and A P , as shown in formulas (4) and (5):

Citation Information

Patent Citations

  • Chinese short text classification method based on graph attention network

    CN112434720A

  • Inducing rich interaction structures between words for document-level event argument extraction

    US20220318505A1