A method for optimizing the retrieval accuracy of electronic medical records using a tagging and scoring mechanism

Optimized electronic medical record search through customized ElasticSearch analyzer and BiLSTM-CRF model marking system, solving false positive and false negative problems, and improving search accuracy and effectiveness.

CN114936268BActive Publication Date: 2025-07-08NANJING UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210528684.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-07-08
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

The existing electronic medical record search system is prone to false positive and false negative results when doctors query, and cannot effectively match synonyms, resulting in low search accuracy.

Method used

Query text processing is performed through a customized ElasticSearch analyzer, and negative word and synonym marking is performed in combination with BiLSTM-CRF deep learning model, query conditions are reconstructed, and the results are returned using the scoring system.

Benefits of technology

It significantly improves the accuracy of electronic medical record search, reduces false positive and false negative results, and improves the validity and accuracy of the query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114936268B_ABST
    Figure CN114936268B_ABST
Patent Text Reader

Abstract

The present invention provides a method for optimizing the retrieval accuracy of electronic medical records using a tagging and scoring mechanism, belonging to the technical field of information retrieval and analysis. By customizing the Elasticsearch analyzer, filtering negative descriptions in the user's original disease input and filling in synonyms to obtain the original word segmentation list, a tagging system for negative words and their coverage domains is trained and deployed using the BiLSTM-CRF deep learning model. Through this tagging system, the tagging sequence of the original input is obtained, and the word segmentation list with positive and negative features is obtained by combining the tagging sequence with the original word segmentation list; after reconstructing the query condition and inputting it into Elasticsearch, each returned result is re-input into the scoring system to recalculate the score according to the scoring algorithm of the invention, and the final result is returned to the user in descending order of the final score. The present invention greatly improves the false positive and false negative problems generated by medical text disease retrieval, and provides a more reliable query service for doctors and patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information retrieval and analysis, and particularly relates to a method for optimizing the retrieval accuracy of electronic medical records using a marking and scoring mechanism. Background Art

[0002] As a record text of a doctor's description of a patient's series of diseases, electronic medical case text has important significance in the doctor's auxiliary decision-making medical treatment. However, a problem arises here: how to find the record that best meets the doctor's needs from the vast electronic medical texts, which is a relatively realistic and severe problem. Electronic medical texts are different from general text corpora such as news and books. The core theme of their content is the description of the patient's diseases. As is well known, diseases have negative and positive aspects, and doctors often retrieve electronic cases according to the characteristics of the positive and negative aspects of the diseases for matching.

[0003] The existing solution technology is to build a case text search engine based on Elasticsearch (hereinafter referred to as ES for short). ES is a distributed and open-source search and analysis engine. The core underlying technology of the search engine generally uses a data structure called inverted index to store the mapping relationship between words and documents. By importing the electronic medical case text into ES, a series of processes such as word segmentation and index building are carried out to form an electronic medical case text search service for doctors to query and use. When the doctor inputs query conditions, etc., since ES will first perform basic word segmentation processing on the input query conditions, and then match the processed results with the inverted index, ES will return the case text results that contain these query conditions as much as possible in descending order according to the TF-IDF or BM25 scoring rules.

[0004] However, the search service implemented based on the above technology has the following defects: The positive diseases searched by doctors often match many false positive diseases, and the negative diseases searched often match many false negative diseases. For example, when a doctor searches for a cold, many texts without the fields of having a cold or not having a cold are included in the previously returned results; when a doctor inputs no fever, many texts related only to fever are also included in the returned results; moreover, due to the existence of many synonyms in medical terms, the current query service cannot provide query results related to synonyms. This brings great inconvenience to doctors. How to improve the current query accuracy and effectiveness and provide better services for doctors and patients is a problem that needs to be considered and solved. Summary of the Invention

[0005] Object of the Invention: To propose a method for optimizing the retrieval accuracy of electronic medical records using a marking and scoring mechanism to solve the false positive and false negative problems existing in medical text search.

[0006] Technical solution: A method for optimizing the retrieval accuracy of electronic medical records using a tagging and scoring mechanism, including the following steps:

[0007] Step 1, customize the ElasticSearch analyzer to process the input query text according to the strategy;

[0008] Step 2, based on the text electronic case data, mark the negative words and their negative scopes in the sample data. Based on the marked training text data, use the BiLSTM-CRF deep learning model to train a sequence annotation model for the negative words and their negative scopes in the text, and encapsulate and deploy the trained model to obtain a marking system;

[0009] Step 3, input the user's original disease query text into the ElasticSearch analyzer in Step 1 to obtain a disease word segmentation list rawTokens;

[0010] Step 4, input the user's original disease query text into the marking system in Step 2 to obtain the marking sequence of the original disease query text;

[0011] Step 5, based on the disease word segmentation list obtained in Step 3 and the marking sequence obtained in Step 4, if the word segmentation in the disease word segmentation list is within the negative scope in the marking sequence, assign a negative feature to the word segmentation, otherwise assign a positive feature to the word segmentation, and finally obtain a word segmentation list with symbol markings signTokens;

[0012] Step 6, based on the word segmentation list with positive and negative marking features obtained in Step 5, reconstruct the query condition in combination with the positive and negative word library;

[0013] Step 7, input the query condition obtained in Step 6 into Elasticsearch to obtain a return result list;

[0014] Step 8, re-score the return result in Step 7 using a scoring system, where the scoring system operates based on the following process:

[0015] 8.1, construct a regular expression based on the rawTokens obtained in Step 3, and extract the sentences in the return result that contain at least one item in the disease list;

[0016] 8.2, input each extracted sentence into the marking system constructed in Step 2 to obtain the marking sequence of the sentence;

[0017] 8.3, analyze and compare the positive and negative features of the disease word segmentation in the sentence with the word segmentation signTokens with positive and negative features obtained in Step 5. If they are consistent, record the word segmentation in the correct hit dictionary R, and if they are inconsistent, record the word segmentation in the wrong hit dictionary E;

[0018] 8.4. Calculate the final score of the current return result according to the following formula;

[0019]

[0020] where R is the set of correctly hit word segments and their frequency dictionaries, E is the set of incorrectly hit word segments and their frequency dictionaries, and m is the attenuation factor;

[0021] 8.5. Sort all the scoring results in descending order and return the final query result.

[0022] According to one aspect of the present invention, for the customized ES analyzer in step 1, the specific implementation steps are as follows:

[0023] Step 1.1. Configure a character filter, a character filter of the customized type mapping, where the path of the negation word mapping file is specified in the mappings_path field. The negation word character filter will delete the negation words included in the user's original input according to the mapping rules;

[0024] Step 1.2. Configure a tokenizer. Based on the text obtained by deleting the negation words from the original input in step 1.1, the tokenizer will tokenize the text to obtain a token sequence;

[0025] Step 1.3. Configure a token filter (optional), a token filter of the customized type synonym, where the path of the synonym mapping file is specified in the synonyms_path field. For the token sequence obtained in step 1.2, the token filter will fill in synonyms for the tokens;

[0026] Step 1.4. Create an analyzer, where the attribute field type is specified as custom, the char_filter field is filled with the character filter configured in step 1.1, the tokenizer field is configured with the tokenizer configured in step 1.2, the filter field is filled with the token filter configured in step 1.3, and this analyzer is set as the default analyzer for a certain index.

[0027] According to one aspect of the present invention, for the tagging system constructed in step 2, the specific implementation steps are as follows:

[0028] Step 2.1. Construct a training data set. Based on the text electronic medical record data, annotate the text according to the I / O tagging scheme. The label I indicates that the disease-related word in the sentence is within the negative focus, and the label O indicates that the disease-related word in the sentence is not within the negative focus;

[0029] Step 2.2: Based on the tags, obtain a good training dataset, train it using the BiLSTM-CRF deep learning model, and optimize and evaluate the model;

[0030] Step 2.3: Based on the trained model, by encapsulating the model in the form of a Web service, provide a text tagging service through an external RestFul API interface.

[0031] According to one aspect of the present invention, in step 6, constructing the query conditions, the specific implementation steps are as follows:

[0032] The disease word segmentation with positive tags obtained from step 5 has a specific format of {token-1:negative, token2:positive,..}, where negative indicates that token-1 belongs to a negative description, and positive indicates that token-2 represents a positive description; for the disease word segmentation with negative descriptions, construct query conditions by splicing each negative word in the negative word library with this word segmentation one by one, and combine these query conditions according to the ES boolean query logic OR; for the disease word segmentation with positive descriptions, first use the word segmentation itself as a query condition and combine it according to the logical OR, and combine these query conditions according to the boolean query logic OR, then construct query conditions by splicing each negative word in the negative word library with this word segmentation one by one, and combine them according to the logical NOT;

[0033] Beneficial effects: The present invention greatly improves the false positive and false negative problems generated by medical text disease retrieval, and the accuracy and effectiveness of the query have been recognized by doctors, providing a more reliable query service for doctors and patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is the overall flowchart of the present invention.

[0035] Figure 2 is the composition diagram of the customized ES analyzer.

[0036] Figure 3 is the overall architecture diagram of the tagging system.

[0037] Figure 4 is the network structure of the BiLSTM-CRF deep learning model.

[0038] Figure 5 is the overall flowchart of the scoring system. DETAILED DESCRIPTION OF THE INVENTION

[0039] The following further describes the specific implementation of the present invention with reference to the drawings.

[0040] A method for optimizing the retrieval accuracy of electronic medical records using a tagging and scoring mechanism, comprising the following steps:

[0041] Step 1, customize the ElasticSearch analyzer to process the input query text according to a certain strategy;

[0042] Step 2, based on the tagged training text data, use deep learning technology to train a sequence annotation model for negative words and their negative scopes in the text, and encapsulate and deploy the trained model to obtain a tagging system;

[0043] Step 3, input the user's original disease query text into the ElasticSearch analyzer in Step 1 to obtain a disease word segmentation list rawTokens;

[0044] Step 4, input the user's original disease query text into the tagging system in Step 2 to obtain the tagging sequence of the original disease query text;

[0045] Step 5, based on the disease word segmentation list obtained in Step 3 and the tagging sequence obtained in Step 4, if a word segmentation in the disease word segmentation list is within the negative scope in the tagging sequence, assign a negative feature to the word segmentation, otherwise assign a positive feature to the word segmentation, and finally obtain a word segmentation list with symbol tags signTokens;

[0046] Step 6, based on the word segmentation list with positive and negative tagging features obtained in Step 5, reconstruct the query condition in cooperation with the positive and negative word library;

[0047] Step 7, input the query condition obtained in Step 6 into Elasticsearch to obtain a return result list;

[0048] Step 8, re-score the return results in Step 7 using a scoring system, and return the query results in descending order of scores.

[0049] In a further embodiment, the specific composition of the customized ElasticSearch analyzer in Step 1 is as follows Figure 2As shown, this embodiment is completed on the Elasticsearch 7.8 version. It mainly consists of three components: character filter, tokenizer, and token filter. In order to filter negative descriptive words from the user's original input text, a character filter neg_filter of type mapping is customized. The path of the negative word mapping file is specified in the mappings_path field. As shown in the figure, the path is the negchar text file in the negcharfilter directory. The specific path of the negcharfilter directory needs to be stored in the config directory under the Elasticsearch installation directory. The text content format in negchar is "negative word =>", indicating that the negative word is mapped to an empty string, that is, the negative word is erased. After configuring the character filter, the tokenizer is then configured. The tokenizer tokenizes the text filtered by the character filter. The IK tokenizer is configured and the ik_smart format is selected for tokenization. After configuring the tokenizer, for the disease tokens obtained by the tokenizer, if the search for synonyms needs to be implemented, a synonym_filter token filter is configured. The defined type is synonym, which is used to add synonyms for the tokens. The path of the synonym mapping file is specified in the synonyms_path field. As shown in the figure, the path is the terminology text file in the synonyms directory. The specific path of the synonyms directory needs to be placed in the config directory under the Elasticsearch installation directory. The text content format in terminology is "disease, disease synonym", and different diseases are placed on different lines, indicating that the disease synonyms will also be searched. A custom analyzer is created. The property field type is specified as custom, the char_filter field is filled with the configured character filter neg_filter, the tokenizer field is configured with the configured tokenizer ik_smart, and the filter field is filled with the configured token filter synonym_filter. And this analyzer is set as the default analyzer for a certain index.

[0050] In a further embodiment, the tagging system constructed in step 2 has a specific architecture as Figure 3 shown. By constructing a training dataset, based on the text electronic medical record data, the text is annotated according to the I / O tagging scheme. The label I indicates that the disease-related word in the sentence is within the negative focus, and the label O indicates that the disease-related word in the sentence is not within the negative focus. Based on the well-tagged training dataset, a BiLSTM-CRF deep learning model is used for training. The specific network architecture of the BiLSTM-CRF model is as Figure 4As shown, it mainly consists of an Embedding layer, a BiLSTM layer, and a CRF layer. To accelerate the model convergence speed, the word embedding features are concatenated with the learning results of the BiLSTM and then input into the CRF. After optimizing and evaluating the model, based on the trained model, the model is encapsulated in the form of a Web service, and text marking services are provided through the RestFulAPI interface.

[0051] In a further embodiment, the construction of the query conditions in step 6 is specifically implemented as follows: The disease segmentation SignTokens with positive and negative markings obtained in step 5 have a specific format of {token-1:negative, token2:positive,...}, where negative indicates that token-1 belongs to a negative description, and positive indicates that token-2 belongs to a positive description. For the disease segmentation with negative descriptions, each negative word in the negative word library is concatenated with the segmentation one by one to construct query conditions, and these query conditions are combined according to the ES boolean query logic OR. For the disease segmentation with positive descriptions, first, the segmentation itself is used as a query condition and combined according to the logical OR, and these query conditions are combined according to the boolean query logic OR. Then, each negative word in the negative word library is concatenated with the segmentation one by one to construct query conditions, and they are combined according to the logical NOT.

[0052] In a further embodiment, the scoring system in step 8 has a specific implementation process as Figure 5 shown. For each returned result, first, a regular expression is constructed based on the segmentation in the rawTokens obtained in step 3 to extract the sentences in the returned result that contain at least one item in the disease list. For each extracted sentence, it is input into the marking system obtained in step 2 to obtain the marking sequence of the sentence. The positive and negative features of the disease segmentation in the sentence are analyzed and compared with the segmentation signTokens with positive and negative features obtained in step 5. If they are consistent, the segmentation is recorded in the dictionary R with the segmentation as the key and the correct number as the value, which indicates a correct hit. If they are inconsistent, the segmentation is recorded in the dictionary E with the segmentation as the key and the wrong number as the value, which indicates a wrong hit. When each extracted sentence completes the above steps, based on the obtained dictionaries R and E, the final score of the current returned result is calculated according to the following formula:

[0053]

[0054] Among them, by traversing each item in the correct dictionary R that is hit, if it is hit once, 1 point is added. To weaken the query bias caused by multiple hits, the same query that is hit multiple times is divided by the attenuation factor set by the user; traverse each item in the incorrect dictionary E that is hit. If it is hit incorrectly once, 1 point is deducted. Similarly, to attenuate the query bias of multiple hits, the same query that is hit incorrectly multiple times is also divided by the attenuation factor m; among them, to ensure that as many of the original input case features are hit as possible, a penalty term of len(R)-len(RawTokens) is added based on the greedy rule to ensure that the sentences that hit as many case features as possible have higher scores; finally, all the scores are sorted from largest to smallest, and the sorted results are returned to the user.

Claims

1. A method for optimizing the retrieval accuracy of electronic medical records using a tagging and scoring mechanism, comprising the following steps: Step 1: Customize the Elasticsearch analyzer to process the input query text according to the strategy; Step 2: Based on the text electronic case data, mark the negative words and their negative ranges in the sample data. Based on the marked training text data, use the BiLSTM-CRF deep learning model to train a sequence annotation model for the negative words and their negative ranges in the text, and encapsulate and deploy the trained model to obtain a tagging system; Step 3: Input the user's original disease query text into the Elasticsearch analyzer in Step 1 to obtain a disease token list rawTokens; Step 4: Input the user's original disease query text into the tagging system in Step 2 to obtain the tagging sequence of the original disease query text; Step 5: Based on the disease token list obtained in Step 3 and the tagging sequence obtained in Step 4, if a token in the disease token list is within the negative range in the tagging sequence, assign a negative feature to the token, otherwise assign a positive feature to the token, and finally obtain a token list with symbol tags signTokens; Step 6: Based on the token list with positive and negative tagging features obtained in Step 5, reconstruct the query condition in combination with the positive and negative word library; Step 7: Input the query condition obtained in Step 6 into the Elasticsearch analyzer to obtain a return result list; Step 8: Re-score each result returned in Step 7 using a scoring system, where the scoring system operates based on the following process: 8.1: Construct a regular expression based on the rawTokens obtained in Step 3, and extract the sentences in the return results that contain at least one item in the disease list; 8.2: Input each extracted sentence into the tagging system constructed in Step 2 to obtain the tagging sequence of the sentence; 8.3: Analyze and compare the positive and negative features of the disease tokens in the sentence with the signTokens of the tokens with positive and negative features obtained in Step 5. If they are consistent, record the token in the correct hit dictionary R, and if they are inconsistent, record the token in the wrong hit dictionary E; 8.4: Calculate the final score of the current return result according to the following formula; where R is the set of dictionaries of correctly hit tokens and their frequencies, E is the set of dictionaries of wrongly hit tokens and their frequencies, and m is the attenuation factor; 8.5: Sort all the scoring results in descending order and return the final query result.

2. According to the method for optimizing the retrieval accuracy of electronic medical records using a tagging and scoring mechanism as described in claim 1, the specific implementation steps of the customized Elasticsearch analyzer in Step 1 are: Step 1.1: Configure the character filter, customize the character filter of type mapping, and specify the path of the negative word mapping file in the mappings_path field. The negative word character filter will delete the negative words contained in the user's original input according to the mapping rules; Step 1.2: Configure the tokenizer. Based on the text obtained by deleting the negation words from the original input in Step 1.1, the tokenizer will tokenize the text to obtain a token sequence; Step 1.3: Configure the token filter. The customized type is the token filter in synonym. Specify the path of the synonym mapping file in the synonyms_path field. Based on the token sequence obtained in Step 1.2, the token filter will perform synonym filling on the tokens; Step 1.4: Create an analyzer. Specify the attribute field type as custom, fill in the character filter configured in Step 1.1 in the char_filter field, configure the tokenizer configured in Step 1.2 in the tokenizer field, fill in the token filter configured in Step 1.3 in the filter field, and set this analyzer as the default analyzer for one of the indexes.

3. A method for optimizing the retrieval accuracy of electronic medical records using a tagging and scoring mechanism according to claim 1, the tagging system constructed in the said Step 2, its specific implementation steps are: Step 2.1: Construct a training data set. Based on the text electronic medical record data, annotate the text according to the I / O tagging scheme. The label I indicates that the disease-related word in the sentence is within the negative focus, and the label O indicates that the disease-related word in the sentence is not within the negative focus; Step 2.2: Based on the well-tagged training data set, use the BiLSTM-CRF deep learning model for training and perform tuning and evaluation on the model; Step 2.3: Based on the trained model, by encapsulating the model in the form of a Web service, provide a text tagging service in the form of an external RestFul API interface.

4. A method for optimizing the retrieval accuracy of electronic medical records using a tagging and scoring mechanism according to claim 1, the construction of the query conditions in the said Step 6, its specific implementation steps are: The disease tokens with positive tags obtained from Step 5 are in the specific format of {token1:negative, token2:positive,..}, where negative indicates that token1 belongs to the negative description, and positive indicates that token2 represents the positive description; for the disease tokens with negative descriptions, the negative words in the negative word library are concatenated with the token one by one to construct query conditions, and these query conditions are combined according to the ES boolean query logic OR; for the disease tokens with positive descriptions, first use the token itself as the query condition and combine them according to the logical OR, and then combine these query conditions according to the boolean query logic OR, and then concatenate the negative words in the negative word library with the token one by one to construct query conditions and combine them according to the logical NOT.

Citation Information

Patent Citations

  • Method and system for detecting medical negative terms

    CN110019641A

  • Electronic medical record data quality evaluation method based on text classification

    CN113626591A