A method and system for generating medical text data
By using MetaMap and Siamese CNN-BIGRU models in medical document classification, and searching synonyms with WordNet dictionary, the problem of lack of training data sets is solved and the performance and efficiency of classifiers are improved.
Patent Information
- Application Number
- CN202211582461.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-12-09
AI Technical Summary
In medical document classification, the lack of training data sets leads to poor classifier performance, especially when the training data content is short, the performance problems of classifiers are more serious.
MetaMap is used to match professional medical names, use the Siamese CNN-BIGRU model to calculate text similarity, and find synonyms through WordNet dictionary, generate a new set of medical record texts, combine the conceptual semantic similarity model to calculate vocabulary similarity, and finally generate medical text data for training.
By adding training data samples, the classification accuracy and efficiency of neural network models in medical document classification are improved.
Smart Images

Figure CN115985435B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology and relates to a method and system for generating medical text data. Background Art
[0002] In medical document classification, detecting meaningful features from unstructured medical texts is a challenging task. Due to the particularity of medical documents in terms of text vocabulary content, specifically, each medical document includes a set of clinical events for accurately and comprehensively describing a patient's health history; at the same time, medical documents have terms and synonyms specific to the medical field, so the medical document classification task is different from general text classification. In addition, different orders of specific domain medical events in medical documents can illustrate a person's health condition in completely different ways. Obtaining meaningful information to examine medical records is crucial.
[0003] In the text classification task in the field of clinical medicine, there is a problem of lack of training data sets. The size of the training data set and the content of the training data provided to the learning model are a prominent point affecting the classification performance of the learning model. When the size of the training data set is not large enough, the trained classifier will not have enough samples for learning, and the performance of the trained classifier will not be convincing. When the training set contains notes with short content (such as abstract texts), the performance problem of the classifier is even more serious.
[0004] Therefore, there is an urgent need for an efficient and highly accurate medical text data enhancement method. Summary of the Invention
[0005] To solve the above problems, the present invention provides a method and system for generating medical text data.
[0006] In a first aspect, the present invention provides a method for generating medical text data, including the following steps:
[0007] S1. Obtain a set of medical record samples, use MetaMap to find important phrases in the set of medical record samples, and perform professional medical name matching in MetaMap;
[0008] S2. Each medical record sample is used to replace the important phrase with multiple obtained professional medical names to synthesize multiple new texts. The similarity between each new text and the medical record sample is calculated by a Siamese CNN - BIGRU model, and the new text with the maximum similarity is used as the first new medical record text. Finally, a set of first new medical record texts corresponding to the set of medical record samples is obtained;
[0009] S3. Preprocess the medical record sample set to obtain a processed sample set, and find multiple synonyms for each word in the processed sample set through the WordNet dictionary;
[0010] S4. Combine the concept semantic similarity model and the path similarity calculation model to construct an improved lexical similarity model. Calculate the similarity between each word and its corresponding multiple synonyms through the improved lexical similarity model, and replace the word with the synonym belonging to the maximum similarity to generate a second new medical record text set;
[0011] S5. Use the medical record sample set, the first new medical record text set, and the second new medical record text set as training samples, and use Word2vec to convert the training samples into word embedding vectors and input them into the medical text classification model for training to obtain a medical text classification model with excellent classification ability.
[0012] Furthermore, the process of obtaining the first new medical record text includes:
[0013] S11. Obtain the medical record sample set N represents the number of medical record samples. Divide each medical record sample into sentences to obtain medical documents. Denote the medical document S n corresponding to the medical record sample D n as S n = {s1, s2,..., s m ,..., s M}, where M represents the number of sentences in the medical document;
[0014] S12. Transmit the sentence s n in the medical document S m to UMLS through MetaMap, and find the important phrases and their phrase concepts in the sentence s m in UMLS;
[0015] S13. Match the corresponding multiple professional medical names for the phrase concept in UMLS, and use the multiple professional medical names to replace the important statements referred to by the phrase concept in the sentence s m respectively to form multiple new sentences. Calculate the similarity between each new sentence and the sentence s m through the Siamese CNN - BIGRU model, and replace the sentence s n in the medical record sample D m with the new sentence with the maximum similarity. After all the sentences in the medical document S n are processed, the first new medical record text D n ' is obtained.
[0016] Further, the Siamese CNN-BIGRU model includes a convolutional layer, a bidirectional GRU layer, a pooling layer, a fully connected layer, and a similarity calculation layer; the process of calculating similarity using the Siamese CNN-BIGRU model includes:
[0017] S131. Preprocess the original sentence to obtain a standard original sentence, and preprocess the new sentence corresponding to the original sentence to obtain a first preprocessed sentence;
[0018] S132. Use a word vector table to map the standard original sentence and the first preprocessed sentence respectively to obtain a standard high-dimensional dense vector and a first high-dimensional dense vector;
[0019] S133. Perform convolution operations on the standard high-dimensional dense vector and the first high-dimensional dense vector respectively in the convolutional layer to obtain a standard one-dimensional vector and a first one-dimensional vector;
[0020] S134. Use the bidirectional GRU layer to obtain the standard refined features and the first refined features of the standard one-dimensional vector and the first one-dimensional vector respectively;
[0021] S135. After reducing the dimensions of the standard refined features and the first refined features respectively through the pooling layer, input them into the fully connected layer to obtain standard semantic information and first semantic information;
[0022] S136. Send the standard semantic information and the first semantic information into the similarity calculation layer to calculate the similarity between the original sentence and the new sentence, expressed as:
[0023]
[0024] where, v represents the standard semantic information of the original sentence, v' represents the first semantic information of the new sentence, d(v, v') represents the similarity between the original sentence and the new sentence, v x represents the vector of the standard semantic information in the x-th dimension, and v' x represents the vector of the first semantic information in the x-th dimension.
[0025] Further, the process of preprocessing the medical record sample set includes:
[0026] S21. Use the jieba word segmentation tool to divide the medical record samples into a list of words;
[0027] S22. Delete the words with little or no semantic contribution in the list of words through a stop word list to obtain a new list of words;
[0028] S23. Use a Porter stemmer to perform stemming removal processing on all the words in the new list of words. If the word is a derivative, change it back to its base form to finally obtain the processed sample of the medical record sample.
[0029] Furthermore, an improved lexical similarity model is used to calculate the similarity between each word in the processed sample and its corresponding multiple synonyms, expressed as:
[0030]
[0031] where sim MICS () represents the conceptual semantic similarity model, w o represents the word in the processed sample, w q represents the q-th synonym of w o , represents the j-th semantic concept of w q , represents the i-th semantic concept of w o , represents the enhancement coefficient; ω represents the weight parameter, represents the set of semantic concepts of w o , represents the set of semantic concepts of w q , represents the path length between two sets of semantic concepts.
[0032] In a second aspect, based on the method of the first aspect, the present invention further provides a medical text data generation system, including a data acquisition module, a first new medical record text generation module, a data preprocessing module, and a second new medical record text generation module, where:
[0033] The data acquisition module is used to obtain a set of medical record samples required for training a medical text classification model;
[0034] The first new medical record text generation module is used to process the set of medical record samples according to the MetaMap and Siamese CNN - BIGRU models to obtain a set of first new medical record texts;
[0035] The data preprocessing module is used to preprocess the medical record samples to obtain processed samples;
[0036] The second new medical record text generation module is used to perform synonym replacement on the processed samples through the WordNet dictionary and the conceptual semantic similarity model to generate a set of second new medical record texts.
[0037] Furthermore, the first new medical record text generation module includes:
[0038] The matching unit is used to use MetaMap to find important phrases in the set of medical record samples and perform professional medical name matching in MetaMap;
[0039] A synthesis unit, configured to replace corresponding important phrases with multiple matched professional medical names in each medical record sample to synthesize multiple new texts;
[0040] A text similarity calculation unit, configured to calculate the similarity between each new text and its corresponding original medical record sample through a Siamese CNN - BIGRU model, and use the new text with the maximum similarity as the first new medical record text;
[0041] The second new medical record text generation module includes:
[0042] A synonym lookup unit, configured to look up multiple synonyms of each word in the processed sample through a WordNet dictionary;
[0043] A word similarity calculation unit, configured to calculate the similarity between each word and its corresponding multiple synonyms through an improved lexical similarity model, and replace the word with the synonym belonging to the maximum similarity.
[0044] Advantages of the present invention:
[0045] The present invention proposes a method based on the field of medical professional dictionaries and a combination method to increase clinical data to solve binary and multi - class medical document classification problems. The method aims to solve the data shortage problem in medical document classification by replacing meaningful expressions with medical names in UMLS and generating new documents using synonyms in the WordNet dictionary based on the original document. Through the present invention, more samples can be provided in the medical document classification training stage, improving the efficiency of models such as CNN, RNN, and HAN models, and enhancing the classification accuracy in neural network models. Description of the drawings
[0046] Figure 1 It is a flowchart for calculating similarity by the Siamese CNN - BIGRU model of the present invention;
[0047] Figure 2 It is a flowchart of the present invention;
[0048] Figure 3 It is a schematic diagram of semantic concept relationships in an embodiment of the present invention. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] The present invention provides a method for generating medical text data, as Figure 2 shown, which includes the following steps:
[0051] S1. Obtain a medical record sample set, use MetaMap to find important phrases in the medical record sample set, and perform professional medical name matching in MetaMap;
[0052] S2. For each medical record sample, replace the important phrases with the obtained professional medical names to synthesize multiple new texts. Calculate the similarity between each new text and the medical record sample through the Siamese CNN - BIGRU model, and use the new text with the maximum similarity as the first new medical record text. Finally, obtain a first new medical record text set corresponding to the medical record sample set;
[0053] S3. Preprocess the medical record sample set to obtain a processed sample set, and find multiple synonyms of each word in the processed sample set through the WordNet dictionary;
[0054] S4. Combine the concept semantic similarity model to propose an improved lexical similarity model to calculate the similarity between each word and its corresponding multiple synonyms, and replace the word with the synonym belonging to the maximum similarity to generate a second new medical record text set;
[0055] S5. Use the medical record sample set, the first new medical record text set, and the second new medical record text set as training samples, use Word2vec to convert the training samples into word embedding vectors and input them into the medical text classification model for training to obtain a medical text classification model with excellent classification ability.
[0056] Specifically, the process of obtaining the first new medical record text includes:
[0057] S11. Obtain the medical record sample set N represents the number of medical record samples. Divide each medical record sample into sentences to obtain medical documents. Represent the medical document S n corresponding to the medical record sample D n as S n ={s1, s2,..., s m ,..., s M}, and M represents the number of sentences in the medical document;
[0058] S12. Transmit the sentence s n in the medical document S m to UMLS through MetaMap, and find the important phrases and their phrase concepts in the sentence s m in UMLS;
[0059] S13. Match the phrase concepts of the important phrases in sentence s m in the Unified Medical Language System (UMLS) to obtain multiple professional medical names corresponding to the phrase concept. In sentence s m , replace the important sentences referred to by the phrase concept with multiple professional medical names respectively to form multiple new sentences. Calculate the similarity between each new sentence and sentence s m using the Siamese CNN-BIGRU model. In the medical record sample D n , replace sentence s m with the new sentence with the maximum similarity. After all the sentences in the medical document S n are processed, the first new medical record text D' is obtained.
[0060] n
[0061] Specifically, the Siamese CNN-BIGRU model includes a convolutional layer, a bidirectional GRU layer, a pooling layer, a fully connected layer, and a similarity calculation layer. As Figure 1 shown, the process of calculating similarity using the Siamese CNN-BIGRU model includes:
[0062] S131. Clean the text content of the original sentence and the new sentences synthesized from the original sentence (there is more than one new sentence synthesized from each original sentence). Remove stop words, abbreviations, special characters, etc. from the original sentence and the new sentences according to the stop word list. Then perform normalization processing on the cleaned original sentence and new sentences, that is, truncate the text lengths of the original sentence and the new sentences to a preset value of 30.
[0063] S132. In order to be able to put the original sentence and the new sentences into the Siamese BIGRU model, it is necessary to map the text content of the original sentence and the new sentences respectively through the word vector table numberbatch-en; map the text content of the original sentence and the new sentences into a standard high-dimensional dense vector and a first high-dimensional dense vector respectively through the word vector library.
[0064] S133. Perform convolutional calculations on the standard high-dimensional dense vector and the first dense vector in the convolutional layer to obtain a standard one-dimensional vector and a first one-dimensional vector containing important semantic information in the text respectively.
[0065] S134. Use the bidirectional GRU layer to obtain the standard refined features and the first refined features of the standard one-dimensional vector and the first one-dimensional vector respectively;
[0066] The bidirectional GRU layer is composed of two unidirectional GRUs with opposite directions. Therefore, during training, each input will be trained on two GRUs in opposite directions, so as to establish connections with the context content simultaneously. The tensor output by the convolutional layer is refined by half.
[0067] S135. The data output by the bidirectional GRU layer is dimensionally reduced through the pooling layer, and then the feature vectors after dimensional reduction are reassembled via the weight matrix through the fully connected layer. Finally, a 128-dimensional vector is output as the semantic information of the whole sentence, that is, the standard semantic information and the first semantic information are finally obtained.
[0068] S136. Through the similarity calculation layer of the Siamese CNN-BIGRU model, using the Manhattan distance calculation method, the absolute value is taken after subtracting each dimension between the two vectors (semantic information), and then the sum is calculated to represent the distance between the vectors in this dimension. The following is the Manhattan distance calculation formula:
[0069]
[0070] Among them, v represents the standard semantic information of the original sentence, v' represents the first semantic information of the new sentence, d(v, v') represents the similarity between the original sentence and the new sentence, v x represents the vector of the standard semantic information in the x-th dimension, and v' x represents the vector of the first semantic information in the x-th dimension.
[0071] Furthermore, the process of preprocessing the medical record sample set includes:
[0072] S21. Using the jieba word segmentation tool, the medical record samples are divided into a list of words;
[0073] S22. By using the stop word list, the words with little or no semantic contribution in the word list are deleted to obtain a new word list;
[0074] S23. Using the Porter stemmer to remove the inflection of all words in the new word list, and if the word is a derivative, it is changed back to its basic form, and finally the processed sample of the medical record sample is obtained.
[0075] In one embodiment, the concept semantic similarity model is used to calculate the similarity between each word in the processed sample and its corresponding multiple synonyms, which is expressed as:
[0076]
[0077] Combined with the method of calculating word similarity based on paths, an improved word similarity model is proposed:
[0078]
[0079] Among them, sim MICS () represents the conceptual semantic similarity model, and w o represents the word in the processing sample, and w q represents the q-th synonym of w o . represents the j-th semantic concept of w q . represents the i-th semantic concept of w o . represents the enhancement coefficient, and ω is the weight parameter that combines the MICS model and the Path model. represents the semantic concept set of w o . represents the semantic concept set of w q . represents the path length between two semantic concept sets. First, calculate the path length between two semantic concept sets, and use the reciprocal of the path length as the similarity between the two semantic concept sets. At the same time, use the conceptual semantic similarity model to calculate the pairwise similarity between the semantic concepts in the semantic concept set of w o and the semantic concepts in the semantic concept set of w q , and take the maximum value as the semantic similarity between w o and w q ; then, by assigning a weight of 1 - ω, the similarity between the two words is finally obtained.
[0080] Specifically, in the WordNet dictionary, based on various semantic relationships (hyponymy, meronymy, etc.), synsets are connected into a network-like graph structure (similar to a tree structure), as Figure 3 shown, c0, c1, etc. represent semantic concepts. Similar to flower, rose, chrysanthemum, etc. are all one semantic concept, and r, r' represent the connection relationship between c0 and c1. Similar to flower and rose belonging to the hyponymy relationship, flower includes rose. The conceptual semantic similarity model of this implementation considers the semantic relationships in the WordNet dictionary. The process of obtaining the conceptual semantic similarity model includes:
[0081] Based on the IC model, comprehensively consider the path factor in WordNet, use the conditional probability between a semantic concept and its adjacent semantic concept to weight their connection edges, and use the mutual information between two semantic concepts to characterize the conceptual semantic similarity. Mutual information is a way to express the similarity degree between information. Mutual information is the amount of information given by one event about another event, represented by I(,):
[0082] Then, taking the mutual information value (MI) between two semantic concepts as the conceptual semantic similarity value, we can obtain:
[0083] sim(c i , c j ) = I(c i , c j )
[0084]
[0085] Combined with the joint probability p(c0c1) in information theory, the concept semantic similarity is further obtained:
[0086]
[0087] In the calculation effect test of the model, the coefficient is increased through information comparison Finally, an improved concept semantic similarity model is obtained:
[0088]
[0089] In one embodiment, the present invention provides a medical text data generation system, including a data acquisition module, a first new medical record text generation module, a data preprocessing module, and a second new medical record text generation module, wherein:
[0090] The data acquisition module is used to acquire a medical record sample set required for training a medical text classification model;
[0091] The first new medical record text generation module is used to process the medical record sample set according to the MetaMap and Siamese CNN - BIGRU models to obtain a first new medical record text set;
[0092] The data preprocessing module is used to preprocess the medical record samples to obtain processed samples;
[0093] The second new medical record text generation module is used to perform synonym replacement on the processed samples through the WordNet dictionary and the concept semantic similarity model to generate a second new medical record text set.
[0094] Specifically, the first new medical record text generation module includes:
[0095] The matching unit is used to use MetaMap to find important phrases in the medical record sample set and perform professional medical name matching in MetaMap;
[0096] The synthesis unit is used to, in each medical record sample, replace the corresponding important phrases with the multiple professional medical names obtained by matching to synthesize multiple new texts;
[0097] A text similarity calculation unit is configured to calculate the similarity between each new text and its corresponding original medical record sample through a Siamese CNN - BIGRU model, and use the new text with the maximum similarity as the first new medical record text;
[0098] The second new medical record text generation module includes:
[0099] A synonym lookup unit is configured to look up multiple synonyms for each word in the processing sample through a WordNet dictionary;
[0100] A word similarity calculation unit is configured to calculate the similarity between each word and its corresponding multiple synonyms through an improved lexical similarity model, and replace the word with the synonym belonging to the maximum similarity.
[0101] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for generating medical text data, characterized in that, It includes the following steps: S1. Obtain a medical record sample set, use MetaMap to find important phrases in the medical record sample set, and perform professional medical name matching in MetaMap; S2. For each medical record sample, replace the important phrases with the obtained professional medical names to synthesize multiple new texts. Calculate the similarity between each new text and the medical record sample through the Siamese CNN-BIGRU model, and use the new text with the maximum similarity as the first new medical record text. Finally, obtain a first new medical record text set corresponding to the medical record sample set; S3. Preprocess the medical record sample set to obtain a processed sample set, and find multiple synonyms for each word in the processed sample set through the WordNet dictionary; S4. Combine the concept semantic similarity model and the path similarity calculation model to construct an improved lexical similarity model. Calculate the similarity between each word and its corresponding multiple synonyms through the improved lexical similarity model, and replace the word with the synonym belonging to the maximum similarity to generate a second new medical record text set; Calculate the similarity between each word in the processed sample and its corresponding multiple synonyms using the improved lexical similarity model, expressed as: Among them, sim MICS () represents the conceptual semantic similarity model, w o represents the word in the processed sample, w q represents the q-th synonym of w o . represents the j-th semantic concept of w q . represents the i-th semantic concept of w o . represents the enhancement coefficient; ω represents the weight parameter, represents the semantic concept set of w o . represents the semantic concept set of w q . represents the path length between two semantic concept sets; S5. Use the medical record sample set, the first new medical record text set, and the second new medical record text set as training samples, use Word2vec to convert the training samples into word embedding vectors and input them into the medical text classification model for training to obtain a medical text classification model with classification ability.
2. A method for generating medical text data according to claim 1, characterized in that, The process of obtaining the first new medical record text includes: S11. Obtain a medical record sample set D = {D1, D2,..., D n ,..., D N}, where N represents the number of medical record samples. Divide each medical record sample into sentences to obtain medical documents. Denote the medical document S n corresponding to the medical record sample D n as S n = {s1, s2,..., s m ,..., s M}, where M represents the number of sentences in the medical document; S12. Transfer the sentence s in the medical document S n to UMLS through MetaMap, and search for important phrases and their phrase concepts in the sentence s m in UMLS; m S13. Match its corresponding multiple professional medical names in UMLS through the phrase concept, and use the multiple professional medical names to replace the important sentences referred to by the phrase concept in the sentence s m respectively to form multiple new sentences. Calculate the similarity between each new sentence and the sentence s m through the Siamese CNN-BIGRU model. In the medical record sample D n , replace the sentence s m with the new sentence with the maximum similarity. After all the sentences in the medical document S n are processed, the first new medical record text D' n is obtained.
3. A medical text data generation method according to claim 2, characterized in that The Siamese CNN-BIGRU model includes a convolutional layer, a bidirectional GRU layer, a pooling layer, a fully connected layer, and a similarity calculation layer; The process of calculating similarity using the Siamese CNN-BIGRU model includes: S131. Preprocess the original sentence to obtain a standard original sentence, and preprocess the new sentence corresponding to the original sentence to obtain a first preprocessed sentence; S132. Use the word vector table to map the standard original sentence and the first preprocessed sentence respectively to obtain a standard high-dimensional dense vector and a first high-dimensional dense vector; S133. Perform convolutional operations on the standard high-dimensional dense vector and the first high-dimensional dense vector respectively in the convolutional layer to obtain a standard one-dimensional vector and a first one-dimensional vector; S134. Use the bidirectional GRU layer to obtain the standard refined features and the first refined features of the standard one-dimensional vector and the first one-dimensional vector respectively; S135. Perform dimensionality reduction processing on the standard refined features and the first refined features respectively through the pooling layer and input them into the fully connected layer to obtain standard semantic information and first semantic information; S136. Send the standard semantic information and the first semantic information into the similarity calculation layer to calculate the similarity between the original sentence and the new sentence, expressed as: Among them, v represents the standard semantic information of the original sentence, v′ represents the first semantic information of the new sentence, d(v, v′) represents the similarity between the original sentence and the new sentence, and v x represents the vector of the standard semantic information in the x-th dimension, and v′ x represents the vector of the first semantic information in the x-th dimension.
4. A method for generating medical text data according to claim 1, wherein The process of preprocessing the medical record sample set includes: S21. Use the jieba word segmentation tool to divide the medical record samples into a vocabulary list; S22. Delete the words with little or no semantic contribution in the vocabulary list through the stop word list to obtain a new vocabulary list; S23. Use a Porter stemmer to perform stemming on all the words in the new vocabulary list. If a word is a derivative, change it back to its base form to finally obtain a processed sample of the medical record sample.
5. A medical text data generation system using a medical text data generation method according to any one of claims 1-4, characterized in that, It includes a data acquisition module, a first new medical record text generation module, a data preprocessing module, and a second new medical record text generation module, where: The data acquisition module is used to obtain a medical record sample set required for training a medical text classification model; The first new medical record text generation module is used to process the medical record sample set according to the MetaMap and Siamese CNN-BIGRU models to obtain a first new medical record text set; The data preprocessing module is used to preprocess the medical record sample to obtain a processed sample; The second new medical record text generation module is used to perform synonym replacement on the processed sample through the WordNet dictionary and the concept semantic similarity model to generate a second new medical record text set.
6. A medical text data generation system according to claim 5, characterized in that, The first new medical record text generation module includes: A matching unit, which is used to use MetaMap to find important phrases in the medical record sample set and perform professional medical name matching in MetaMap; A synthesis unit, which is used to replace the corresponding important phrases with multiple professional medical names obtained by matching in each medical record sample to synthesize multiple new texts; A text similarity calculation unit, which is used to calculate the similarity between each new text and its corresponding original medical record sample through the Siamese CNN-BIGRU model, and use the new text with the maximum similarity as the first new medical record text; The second new medical record text generation module includes: A synonym search unit, which is used to search for multiple synonyms of each word in the processed sample through the WordNet dictionary; A word similarity calculation unit, which is used to calculate the similarity between each word and its corresponding multiple synonyms through an improved lexical similarity model, and use the synonym belonging to the maximum similarity to replace the word.