A word2vec medical similarity question retrieval method based on improved tf-idf weighting
Through the improved tf-idf weighted word2vec method, a medical Q&A database is built, which solves the problem that users find it difficult to retrieve similar questions in the medical Q&A system, and realizes accurate matching of user questions and return reliable answers, improving user experience.
Patent Information
- Application Number
- CN202111106824.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-09-22
AI Technical Summary
The existing medical Q&A system is difficult to accurately retrieve similar questions after user inquiries, especially in the recognition and modeling of proper nouns. At the same time, there are advertising and error messages in medical information, making it difficult for users to obtain reliable answers.
Using the improved tf-idf weighted word2vec method, a medical Q&A database is constructed, and through word segmentation, word embedding and cosine similarity calculation, accurate retrieval of user questions and return of similar Q&A pairs.
Improve user experience, and by strengthening the attention to and use of reliable data resources for proper nouns, it solves the difficulties in identifying and modeling of proper nouns and the mixed problems of medical information, and achieves accurate matching of user problems and return of reliable answers.
Smart Images

Figure CN113868387B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of online medical treatment, and in particular relates to a word2vec medical similarity question retrieval method based on improved tf-idf weighting. Background Art
[0002] With the widespread use of the Internet today, it has become the norm to obtain medical knowledge through retrieval. However, due to the unknown source of knowledge and the mixed quality of medical information, traditional search engines do not meet the medical needs of users. The question-answering system identifies the user's question intention and returns the answer the user wants, such as Alibaba's housekeeper "Xiao Mi", Baidu's Xiaodu, Huawei's Xiaoyi, etc. Automatic question-answering systems have been applied to a wide range of fields. Therefore, it has become an urgent need to automatically reply to users' questions based on existing medical question-answering data.
[0003] Wong et al. established a medical question-answer knowledge base through extraction methods and designed the enquireMe system. Ni et al. built a question-answer system that determines answers through multiple rounds of dialogue. Cui et al. created a knowledge base question-answer system by analyzing large-scale corpora and extracting rules to build templates. Domestic related research started late. Wang Lei et al. combined the knowledge base to build a question-answer system for auxiliary diagnosis. Yuan Kaiqi et al. established a knowledge graph in the medical field.
[0004] In the field of medical question answering, after receiving the user's inquiry, the results the user wants are returned based on the existing medical question answering data. Similar retrieval is both a key point and a difficult point. Similar medical question retrieval can be regarded as a rough match of the question. Although some results have been achieved, there are still some problems in the current search-based question answering that have not been properly solved, which are mainly reflected in the following two aspects.
[0005] One aspect is the difficulty in identifying and modeling proper nouns. Since some proper nouns in the medical field are relatively rare in daily life, such as drug names, examination item names, etc., it is difficult to model and identify such words in the medical question-and-answer database.
[0006] On the other hand, medical information is mixed. Currently, the answers to medical questions found using search engines are full of advertisements or erroneous information. Although there are many medical websites, users need to search one by one, and the websites themselves also have the problem of good and bad information. The automatic question-and-answer systems on the market are mainly aimed at daily life, and their ability to answer questions about medical information is insufficient, making it difficult to solve user problems. Summary of the invention
[0007] 1. Technical issues to be resolved
[0008] The technical problem to be solved by the present invention is how to provide a word2vec medical similar question retrieval method based on improved tf-idf weighting to solve the problem of returning the results desired by the user based on the existing medical Q&A data after receiving the user's query; the difficulties in the recognition and modeling of proper nouns and the problem of the mixed and disorderly medical information.
[0009] (II) Technical Solution
[0010] To solve the above technical problem, the present invention proposes a word2vec medical similar question retrieval method based on improved tf-idf weighting, and the method includes the following steps:
[0011] S1. Construct a medical Q&A pair database;
[0012] S2. Perform data preprocessing, and perform word segmentation on the question input by the user and the questions in the medical Q&A pair database;
[0013] S3. Calculate the improved tf-idf values of each term after word segmentation of the question input by the user and the questions in the medical Q&A pair database respectively;
[0014] S4. Perform word embedding processing on each term after word segmentation of the question input by the user and the questions in the medical Q&A pair database to generate word vectors;
[0015] S5. Multiply the word vector of each term by the tf-idf value to obtain the final text vector;
[0016] S6. Calculate the similarity of the text vectors through cosine similarity;
[0017] S7. Return the Q&A pairs with higher similarity in the medical Q&A pair database to the user.
[0018] Further, the step S1 specifically includes:
[0019] First, collect medical Q&A data, and crawl common medical Q&A pair data through public medical-related websites;
[0020] Then, perform screening, delete the unqualified Q&A pairs with invalid answers and repeated answers, and form a Q&A pair database.
[0021] Further, in the step S2, the Jieba word segmentation tool is used for word segmentation, and then stop words are removed, and words such as "is", "of", and modal particles, auxiliary words, and punctuation marks are removed.
[0022] Further, in the step S2, the Chinese medical word list THUOCL_medical is added to the word segmentation tool library to help distinguish medical exclusive vocabulary.
[0023] Furthermore, the step S3 specifically includes:
[0024] First, feature matching is performed, and the words that match the Chinese medical vocabulary THUOCL_medical are added to the calculation of feature weights;
[0025]
[0026] Among them, c i Expressing word a i The word frequency in the Chinese medical vocabulary N, n represents the number of all words in the Chinese medical vocabulary, w i It is the word a i Combine the weight of the keywords matched with the Chinese medical vocabulary;
[0027] Then, calculate the improved tf-idf value, the calculation formula is:
[0028] w ij =tf ij *idf i *w i (3);
[0029] Among them, tf ij Expressing word a i In the question and answer j The frequency of occurrence, idf i Expressing word a i Inverse document frequency in a set of question-answer pairs.
[0030] Furthermore, step S4 uses the word vector generation tool word2vec to perform word embedding processing.
[0031] Furthermore, step S4 uses the CBOW model of the word vector generation tool word2vec to perform word embedding processing.
[0032] Furthermore, the step S4 specifically includes: using the word vector generation tool Word2vec, setting the context window size to 2, the dimension of the word vector to 100, the minimum threshold of word occurrence to 1, and using the tool Word2vec to train each word into a set of word vectors using a three-layer neural network.
[0033] Furthermore, the step S6 specifically includes: calculating the cosine value of the angle between the text vector of each word in the user's question and the text vector of each word in the question in the medical question and answer database, assuming that the length of the user's question is M1, and assuming that the length of the question in the medical question and answer database is M2, obtaining the text similarity matrix of the two sentences, the size of which is M1M2, and taking the mean in the matrix as the similarity.
[0034] Furthermore, the step S7 specifically includes: setting a threshold, returning the question-answer pairs whose similarity exceeds the threshold to the user, and completing the retrieval.
[0035] (III) Beneficial effects
[0036] The present invention proposes a word2vec medical similar question retrieval method based on improved tf-idf weighting. The present invention first establishes a medical question-answer pair database. After the user asks a question, the present invention specifically searches for questions in the medical field, and finally returns the question-answer pair with a high similarity to the user's question as a search result for output. While searching, the attention to proper nouns is strengthened. Using TF-IDF, weights are added to all words to indicate the importance of the words and the impact on the similarity calculation of the entire sentence. For words such as proper nouns that need to be strengthened, the greater the TF-IDF value, the greater the impact on the similarity. The recognition of proper nouns is made clearer, solving the problem of difficulty in identifying and modeling proper nouns. When constructing a medical question-answer pair database, reliable data resources are selected to solve the problem of mixed medical information. The present invention can accurately return question-answer pairs with a high similarity to the user's question as search results, greatly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0038] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the drawings and examples.
[0039] In response to the above problems, this patent proposes a word2vec medical similar question retrieval method based on improved tf-idf weighting. First, a medical question-answer pair database is established. After the user asks a question, questions in the medical field are searched specifically. Finally, question-answer pairs with a high degree of similarity to the user's question are returned as retrieval results.
[0040] While searching, we pay more attention to proper nouns. We use TF-IDF to add weights to all words to indicate their importance and their impact on the similarity calculation of the entire sentence. For words that need to be emphasized, such as proper nouns, the larger the TF-IDF value, the greater the impact on similarity. This makes the recognition of proper nouns clearer and solves the problem of difficulty in recognizing and modeling proper nouns.
[0041] When building a medical question-and-answer database, select reliable data resources to solve the problem of mixed medical information.
[0042] This patent belongs to the field of natural language processing technology. Aiming at the retrieval of medical similar questions, a word2vec medical similar question retrieval method based on improved tf-idf weighting is proposed.
[0043] First, construct a medical Q&A pair database. In subsequent retrievals, retrieve and answer through the medical Q&A pairs in the database. Then, preprocess the data in the database and the user's question, and calculate the improved tf-idf values of each term in the user input and medical question pairs respectively. After that, perform word embedding, and then multiply the word2vec word vector of each word by the tf-idf value to obtain the final text vector. Thus, the parsing of the user's question and the database is completed, serving as the basis for subsequent retrievals. Finally, according to the parsed data, calculate the similarity between questions, find semantically related questions from the Q&A pair database, and return the questions with higher similarity.
[0044] The medical similar question retrieval method identifies the user's intention and returns results with higher similarity. The specific steps are as Figure 1 shown.
[0045] S1. Construct a medical Q&A pair database
[0046] First, collect medical Q&A data and crawl common medical Q&A pair data through public medical-related websites. Currently, the amount of Q&A data in the medical field is very large. The answers on medical health Q&A platforms such as "Good Doctor Online①" are all answered by doctors registered on the platform. Therefore, the amount of Q&A pair data is not only large but also the answers are reliable. By constructing the database in this way, new questions from patients can be answered based on previous medical Q&A pairs. The reason is that the general public has limited knowledge of medical-related knowledge, making the descriptions of the same question likely to be similar.
[0047] Then, screen and delete unqualified Q&A pairs such as invalid answers and duplicate answers to form a Q&A pair database.
[0048] S2. Data preprocessing, perform word segmentation on the questions input by the user and the questions in the medical Q&A pair database;
[0049] First, perform Chinese word segmentation. Since it is medical data, some professional medical terms need to be accurately identified. Use the Jieba word segmentation tool for word segmentation, and add the Chinese medical word list THUOCL_medical sorted out by Tsinghua University to the word segmentation tool's word library to help distinguish medical-specific vocabulary.
[0050] Then, remove stop words, removing words such as "is", "of", as well as modal particles, auxiliary words, punctuation marks, etc.
[0051] S3. Calculate the improved tf-idf value of each term after the user input question and the medical question and answer have been segmented for the question in the database.
[0052] TF-IDF (Term Frequency-inverse Document Frequency) is a statistical method that allows words that appear many times in a document to have higher weights in the representation of word vectors than words that appear less frequently, while weakening the role of words that always appear in all documents, so that rare or important words are given higher weights.
[0053] The calculation formula of tf-idf value is:
[0054] w ij =tf ij *idf i (1)
[0055] Among them, tf ij Expressing word a i In the question and answer j The frequency of occurrence, idf i Expressing word a i Inverse document frequency in a set of question-answer pairs.
[0056] The characteristics of TF-IDF can help improve the problems of medical data being difficult to obtain and having a low proportion in all data. However, the traditional TF-IDF simply uses "word frequency" as a measure of importance, which cannot take into account the fact that words in medical data may not appear many times. Therefore, we try to adjust the weight calculation of feature items by TF-IDF, and combine it with the Chinese medical vocabulary N to perform weight optimization calculation, so that the calculation of the improved tf-idf value increases the weight of words in the medical vocabulary.
[0057] First, feature matching is performed. For words that match the Chinese medical vocabulary THUOCL_medical, feature weight calculations are added, and for words that do not match, the original calculations are performed.
[0058]
[0059] Among them, c i Expressing word a i The word frequency in the Chinese medical vocabulary N, n represents the number of all words in the Chinese medical vocabulary, w i It is the word a i Combine the weight of keywords matched with the Chinese medical vocabulary.
[0060] The improved calculation formula is:
[0061] w ij =tf ij*idf i *w i (3)
[0062] S4, performing word embedding processing on each word item after word segmentation of the question input by the user and the medical question and answer in the database to generate a word vector;
[0063] The word vector generation tool word2vec is used for word embedding. Word2vec is a simple neural network that represents the words in the text with a dense low-dimensional real number vector, and then the spatial distance between two words can be calculated to represent the semantic relationship between the words. There are two main models: CBOW and Skip-Gram. The training method of the Skip-Gram model is to predict its context based on the given input word, while the CBOW model predicts the current word based on the given context. According to the attributes and quantity of medical data, the CBOW model is selected for word embedding.
[0064] We use Word2vec, an open source word vector generation tool from Google. After data preprocessing, the length of each question basically does not exceed 20 characters. Therefore, we set the context window size to 2, the dimension of the word vector to 100, and the minimum threshold for word occurrence to 1. Through the tool, we train each word into a set of word vectors using a three-layer neural network.
[0065] Word2vec solves the problem that the traditional bag-of-words model cannot represent the contextual semantic information of the text and the resulting dimensionality disaster, so that words with similar semantics have similar vector representations.
[0066] S5, multiply the word vector of each term by the tf-idf value to obtain the final text vector;
[0067] The text vector is obtained by multiplication, and the similarity of the text vector is used to determine the sentence similarity.
[0068] S6, calculating the similarity of the text vectors by using cosine similarity;
[0069] Since word2vec generates distributed word vectors, the cosine similarity of words with similar meanings will be higher, so cosine similarity is used as an indicator to measure text similarity. The cosine value of the angle between the text vector of each word in the user's question (assuming the length is M1) and the text vector of each word in the question in the medical question-answer database (assuming the length is M2) is calculated pairwise to obtain the text similarity matrix of the two sentences (size is M1M2), and the mean value in the matrix is used as the similarity.
[0070] S7. Return the question and answer pairs with higher similarity in the medical question and answer database to the user.
[0071] Set a threshold and return question-answer pairs whose similarity exceeds the threshold to the user to complete the retrieval.
[0072] The present invention first establishes a medical question-answer database, and after the user asks a question, searches specifically for questions in the medical field, and finally returns question-answer pairs with a high degree of similarity to the user's question as search results for output.
[0073] While searching, we pay more attention to proper nouns. We use TF-IDF to add weights to all words to indicate their importance and their impact on the similarity calculation of the entire sentence. For words that need to be emphasized, such as proper nouns, the larger the TF-IDF value, the greater the impact on similarity. This makes the recognition of proper nouns clearer and solves the problem of difficulty in recognizing and modeling proper nouns.
[0074] When building a medical question-and-answer database, select reliable data resources to solve the problem of mixed medical information.
[0075] The present invention can accurately return question-answer pairs with a high degree of similarity to the user's question as retrieval results, greatly improving the user experience.
[0076] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A word2vec medical similarity question retrieval method based on improved tf-idf weighting, characterized in that: The method includes the following steps: S1. Construct a medical Q&A pair database; S2. Data preprocessing, perform word segmentation on the question input by the user and the questions in the medical Q&A pair database; S3. Calculate the improved tf-idf values of each term after word segmentation of the question input by the user and the questions in the medical Q&A pair database respectively; S4. Perform word embedding processing on each term after word segmentation of the question input by the user and the questions in the medical Q&A pair database to generate word vectors; S5. Multiply the word vector of each term by the tf-idf value to obtain the final text vector; S6. Calculate the similarity of the text vectors through cosine similarity; S7. Return the Q&A pairs with higher similarity in the medical Q&A pair database to the user; Wherein, The specific steps of step S3 include: First, perform feature matching, and add the words that match the Chinese medical word list THUOCL_medical to the calculation of feature weights; Among them, c i Expressing word a i The word frequency in the Chinese medical vocabulary N, n represents the number of all words in the Chinese medical vocabulary, w i It is the word a i Combine the weight of the keywords matched with the Chinese medical vocabulary; Then, calculate the improved tf-idf value, and the calculation formula is: w ij =tf ij *idf i *w i (3); Among them, tf ij Expressing word a i In the question and answer j The frequency of occurrence, idf i Expressing word a i Inverse document frequency in a set of question-answer pairs.
2. The word2vec medical similar question retrieval method based on improved tf-idf weighting as claimed in claim 1, characterized in that: The specific steps of step S1 include: First, collect medical Q&A data, and crawl common medical Q&A pair data through public medical-related websites; Then, perform screening, delete the unqualified Q&A pairs with invalid answers and repeated answers to form a Q&A pair database.
3. The word2vec medical similar question retrieval method based on improved tf-idf weighting as claimed in claim 1, characterized in that: In step S2, the Jieba word segmentation tool is used for word segmentation, and then stop words are removed, and words such as "is", "of" and modal particles, auxiliary words, and punctuation marks are removed.
4. The word2vec medical similar question retrieval method based on improved tf-idf weighting as claimed in claim 2, characterized in that: In step S2, the Chinese medical word list THUOCL_medical is added to the word segmentation tool's thesaurus to help distinguish medical exclusive vocabulary.
5. The word2vec medical similar question retrieval method based on improved tf-idf weighting as claimed in claim 1, characterized in that: In step S4, the word vector generation tool word2vec is used for word embedding processing.
6. The word2vec medical similar question retrieval method based on improved tf-idf weighting as claimed in claim 5, characterized in that: In step S4, the CBOW model of the word vector generation tool word2vec is used for word embedding processing.
7. The word2vec medical similar question retrieval method based on improved tf-idf weighting as claimed in claim 1, characterized in that: The specific steps of step S4 include: adopting the word vector generation tool Word2vec, setting the context window size to 2, the dimension of the word vector to 100, and the minimum threshold of word appearance to 1, and training each word into a group of word vectors through the tool Word2vec with a three-layer neural network.
8. The word2vec medical similar question retrieval method based on improved tf-idf weighting according to any one of claims 5 to 7, characterized in that: The specific steps of step S6 include: calculating the cosine value of the angle between the text vectors of each word in the user's question and the text vectors of each word in the question in the medical Q&A pair database pairwise. Assume the length of the user's question is M1, and assume the length of the question in the medical Q&A pair database is M2, to obtain a text similarity matrix of the two sentences with a size of M1M2, and take the mean value in the matrix as the similarity.
9. The word2vec medical similar question retrieval method based on improved tf-idf weighting as claimed in claim 8, characterized in that: The specific steps of step S7 include: setting a threshold, and returning the Q&A pairs with similarity exceeding the threshold to the user to complete the retrieval.
Citation Information
Patent Citations
Determination method and device for similarities of text semantics
CN106776503A
Text matching method based on term frequency-inverse document and CRF
CN108255813A