Method for extracting entity in text, electronic equipment and storage medium
By using entity score calculation and preset expert entity library matching methods in case text, the problem of time-consuming and easy to miss or incorrectly raising entities in case text is solved, and the rapid and accurate extraction of entities in case text is achieved, and the efficiency of medical management and cost accounting is improved.
Patent Information
- Application Number
- CN202510646887.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The information in the case text is dense and complex, which makes it time-consuming to extract entities, easy to omit or mistakenly withdraw, affecting medical management, expense accounting and medical insurance reimbursement.
By introducing the word frequency, word attributes, character length and position information of entity keywords in the case text, the entity score is calculated, and the preset expert entity library is combined for matching and reorganization, ensuring that the extracted entity words are more accurate and comprehensive.
It realizes rapid and accurate extraction of entities in case text, reduces information omissions and misreports, and improves the efficiency of medical management and cost accounting.
Smart Images

Figure CN120163151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and particularly to a method for entity extraction in text, an electronic device, and a storage medium. Background Art
[0002] Case text refers to the text materials that record in detail the patient's condition, diagnosis process, etc. during the medical treatment, and usually includes information such as the patient's basic information, diagnosis results, examination results, and treatment methods. However, due to the large amount of content in case text, it takes time to view case text and it is easy to have situations such as information omission or incorrect information extraction, which has an impact on aspects such as medical record file management, medical expense accounting, and medical insurance reimbursement scope.
[0003] In recent years, with the development of computer technology in the medical field, the method of using a computer system to extract entities from case text instead of humans has gradually been implemented. However, due to the complexity of case text and professional vocabulary, the accuracy of computer entity extraction still needs to be improved. Summary of the Invention
[0004] In view of the above technical problems, the present invention provides a method for entity extraction in text, an electronic device, and a storage medium, which can extract more comprehensive and accurate target entity words, and is beneficial to quickly and accurately obtain key information in case text.
[0005] According to a first aspect of the present invention, there is provided a method for entity extraction in text, including the following steps: S100, for any entity keyword extracted from a given case text, calculate an entity score corresponding to the entity keyword based on the word frequency, word attribute, character length, and position information of the entity keyword in the given case text.
[0006] S200, when the entity score corresponding to the entity keyword is not less than a preset entity score threshold, use the entity keyword as the extracted target entity word; otherwise, traverse the preset entity words in a preset expert entity library, and when there is a preset entity word identical to the entity keyword, use the entity keyword as the extracted target entity word.
[0007] S300, when there is no preset entity word identical to the entity keyword, find the target sentence where the entity keyword is located in the given case text, and perform character decomposition and character combination on the target sentence to generate a number of recombined entity words.
[0008] For S400, traverse the preset entity words in the preset expert entity library again. When there is no preset entity word identical to any of the recombined entity words, delete the entity keyword. Otherwise, based on the recombined entity words corresponding to the same preset entity words, determine the finally extracted target entity word.
[0009] According to a second aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the method for entity extraction in the above text.
[0010] According to a third aspect of the present invention, there is provided an electronic device including a processor and the above non-transitory computer-readable storage medium.
[0011] The present invention has at least the following beneficial effects: The present invention provides a method for entity extraction in text. For any entity keyword extracted from a given case text, four indicators including the word frequency, word attribute, character length, and position information of the entity keyword are introduced to comprehensively analyze the entity keyword, and an entity score is obtained by quantification, which is beneficial to the judgment of the extraction accuracy of the entity keyword. When the entity score is not less than the entity score threshold, it is considered that the extraction of the entity keyword is accurate. Otherwise, the accuracy of the entity keyword needs to be further judged. The entity keyword is matched with the preset entity words in the preset expert entity library, and those that match successfully are used as the target entity words. When the match fails, the target sentence where the entity keyword is located is found, and the target sentence is disassembled and combined character by character to generate several recombined entity words. Then, according to the matching situation between the recombined entity words and the preset entity words, the final target entity word is determined. Through the above process, the accuracy of the entity keyword and the recombined entity word is finally determined, making the selected entity words more accurate. At the same time, through the recombination of characters, the omission of entity words can also be prevented. Therefore, the present invention makes the finally extracted target entity words more comprehensive and accurate, which is beneficial to quickly and accurately obtaining the key information in the case text. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0013] Figure 1 It is a flowchart of the method for entity extraction in text provided by the embodiment of the present invention. DETAILED DESCRIPTION
[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0015] An embodiment of the present invention provides a method for entity extraction in text, as Figure 1 shown, the method includes the following steps: S100, for any entity keyword among a plurality of entity keywords extracted from a given case text, calculate the entity score corresponding to the entity keyword based on the word frequency, word attribute, character length, and position information of the entity keyword in the given case text.
[0016] Further, a plurality of entity keywords are extracted from the given case text through a named entity model. For example, the ner model.
[0017] Specifically, in step S100, the entity score corresponding to the entity keyword is calculated through the following steps: S101, according to a plurality of historical case texts, obtain the index-confidence fitting curves corresponding to the word frequency and character length respectively.
[0018] Further, both the word frequency and the character length are positively correlated with the entity score corresponding to the entity keyword.
[0019] Specifically, the index-confidence fitting curve corresponding to the word frequency is obtained through the following steps: S1011, based on the word frequency of each entity keyword in the corresponding historical case text, obtain the entity confidence corresponding to the entity keywords with different word frequencies. For example, for a plurality of entity keywords that appear 2 times in the same historical case text, match them with the known entity keywords. For the matched entity keywords, consider them as accurate keywords, and calculate the entity confidence corresponding to the entity keywords at each word frequency.
[0020] S1012, perform curve fitting on the word frequency and the entity confidence corresponding to the word frequency to obtain the index-confidence fitting curve.
[0021] Further, the method for obtaining the index-confidence fitting curve corresponding to the character length is the same as that for the word frequency, and will not be repeated here.
[0022] S102, according to the word frequency and character length of the entity keyword in the given case text, respectively find out the word frequency confidence and character length confidence corresponding to the entity keyword.
[0023] S103. Calculate the entity score corresponding to the entity keyword through weighted sum based on the word frequency confidence, character length confidence, pre-acquired word attribute confidence, and position confidence corresponding to the entity keyword, as well as the preset index weights corresponding to the word frequency confidence, character length confidence, word attribute confidence, and position confidence respectively.
[0024] Specifically, the word attribute confidence and position confidence are confidences set according to actual needs. For example, in actual operations, since most conclusive entities are nouns, a higher confidence should be set for nouns, the confidence of adjectives should be less than that of nouns, and the confidence of verbs should be less than that of adjectives; and since the end of the case text is generally the conclusion, a higher confidence should be set for entity keywords that are located later.
[0025] As described above, when extracting entities such as diagnoses, tests, surgeries, and treatment methods from case texts, it is necessary to judge the accuracy of the extracted entities. Therefore, the present application introduces the above four dimensions to comprehensively analyze the extracted entity keywords and quantifies them in numerical form, providing a clear and definite judgment basis for whether the extraction of entity keywords is accurate.
[0026] Furthermore, the entity score corresponding to the entity keyword is calculated through the following steps: S110. Based on the obtained several case text samples, divide the several case text samples into corresponding character quantity intervals according to several preset character quantity intervals. For example, the character quantity intervals can be set as 50 - 100, 100 - 150, etc.
[0027] S120. For the several case text samples corresponding to any character quantity interval, take each case text sample as the given case text and execute steps S100 - S500 to obtain several target entity words corresponding to each case text sample.
[0028] S130. According to the several target entity words corresponding to any case text sample and the corresponding several preset entity word samples, calculate the initial extraction accuracy rates corresponding to the target entity words with different character quantities under any character quantity interval. For example, for any case text sample within the 50 - 100 interval, obtain the target entity words with 2 characters and 3 characters respectively, and match them with the preset entity word samples to obtain the initial extraction accuracy rates corresponding to the target entity words with 2 characters and 3 characters respectively.
[0029] S140. Based on the several character quantity intervals, determine the average value of the several initial extraction accuracy rates corresponding to the target entity word with any character quantity as the final extraction accuracy rate corresponding to the target entity word with any character quantity.
[0030] S150. Calculate a new entity score corresponding to the entity keyword based on the entity score corresponding to the entity keyword and the final extraction accuracy rate corresponding to the character count of the entity keyword. It can be understood that the final extraction accuracy rate is used as the confidence factor of the entity score.
[0031] As described above, on the basis of obtaining the entity score, the case texts in different character intervals are further considered. Since the large number of characters in the case text is likely caused by the large number of characters in the entity keyword, and the number of entity keywords extracted in the actual extraction process may not be accurate. Therefore, the extraction accuracy rates of entity words with different character counts are obtained by analyzing a large number of case text samples in different character count intervals and used as reference factors to make the obtained entity score more reliable.
[0032] S200. When the entity score corresponding to the entity keyword is not less than the preset entity score threshold, take the entity keyword as the extracted target entity word. Otherwise, traverse the preset entity words in the preset expert entity library. When there is a preset entity word that is the same as the entity keyword, take the entity keyword as the extracted target entity word. It can be understood that the preset expert entity library refers to an entity library that stores a number of entity words in this field added by experts in this field. Those skilled in the art set the entity score threshold according to actual needs or experience, which will not be elaborated here.
[0033] As described above, when the entity score corresponding to the entity keyword is low, it indicates that the accuracy of the entity keyword is unreliable. To further determine its accuracy, compare it with the preset entity words in the expert entity library to achieve secondary confirmation of the entity keyword and make the selected entity keyword more accurate.
[0034] S300. When there is no preset entity word that is the same as the entity keyword, find the target sentence where the entity keyword is located in the given case text, and perform character decomposition and character combination on the target sentence to generate a number of recombined entity words.
[0035] Specifically, the performing character decomposition and character combination on the target sentence to generate a number of recombined entity words includes the following steps: S301. Obtain the total number of characters K of the target sentence.
[0036] S302. In the order of character count from 1 to K, respectively use the character-level N-gram model to perform character decomposition and character combination on the target sentence to generate a number of recombined entity words corresponding to each character count. It can be understood that N is sequentially taken as 1 to K.
[0037] As described above, when there is no preset entity word that is the same as the entity keyword, considering that the entity keyword extracted by the NER model may be incorrect, the sentence where the entity keyword is located is found, and several recombined entity words are generated through the N-gram model to judge the accuracy of the recombined entity words again, so as to prevent omission of the extraction of entity keywords.
[0038] S400. Traverse the preset entity words in the preset expert entity library again. When there is no preset entity word that is the same as any of the recombined entity words, delete the entity keyword. Otherwise, determine the finally extracted target entity word according to the recombined entity word with the same preset entity word.
[0039] Specifically, the step of determining the finally extracted target entity word according to the recombined entity word with the same preset entity word includes the following steps: S401. When the number of recombined entity words with the same preset entity word is 1, use the recombined entity word with the same preset entity word as the finally extracted target entity word, and delete the entity keyword.
[0040] S402. When the number of recombined entity words with the same preset entity word is greater than 1, obtain the position information of each recombined entity word with the same preset entity word in the target sentence; it can be understood that the position information refers to the character order corresponding to each character of the recombined entity word in the target sentence.
[0041] S403. For the position information of each recombined entity word obtained in the target sentence, judge whether there is character overlap between the recombined entity words. When there is no character overlap, use all the recombined entity words with the same preset entity word as the finally extracted target entity words, and delete the entity keyword.
[0042] S404. When there is character overlap between the recombined entity words, use the recombined entity word with more corresponding characters or the recombined entity word with a higher position as the finally extracted target entity word, and delete the entity keyword; it can be understood that when the number of characters is different, use the recombined entity word with more corresponding characters as the finally extracted target entity word. When the number of characters is the same, use the recombined entity word with a higher position as the finally extracted target entity word. For example, for XX resection, where "XX" represents a body part and is an important entity, the general meanings of XX resection and XX excision are the same, indicating that the entity word with a higher position has higher accuracy.
[0043] As described above, when there are overlapping characters in the recombined entity words, it indicates that only one of the overlapping recombined entity words is accurate. Therefore, this application introduces the character overlap situation of the recombined entity words to further judge the entity words, and filters out the target entity words considering the actual situation during the judgment process, making the finally extracted target entity words more comprehensive and accurate.
[0044] In summary, the present invention provides a method for entity extraction in text. For any entity keyword extracted from a given case text, four indicators, namely the word frequency, word attribute, character length, and position information of the entity keyword, are introduced to comprehensively analyze the entity keyword, and an entity score is obtained through quantification, which is beneficial to the judgment of the extraction accuracy of the entity keyword. When the entity score is not less than the entity score threshold, it is considered that the extraction of the entity keyword is accurate. Otherwise, the accuracy of the entity keyword needs to be further judged. The entity keyword is matched with the preset entity words in the preset expert entity library. Those that match successfully are used as the target entity words. When the match fails, the target sentence where the entity keyword is located is found, and the target sentence is disassembled and combined into several recombined entity words, and then the final target entity words are determined according to the matching situation between the recombined entity words and the preset entity words. Through the above process, the accuracy of the entity keyword and the recombined entity words is finally determined, making the selected entity words more accurate. At the same time, by using the N-gram model to recombine adjacent characters, the omission of entity words can also be prevented. Therefore, the present invention makes the finally extracted target entity words more comprehensive and accurate, which is beneficial to quickly and accurately obtaining the key information in the case text.
[0045] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which can be set in an electronic device to store at least one instruction or at least one program related to a method for implementing a method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method for entity extraction in text provided in the above embodiment.
[0046] An embodiment of the present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0047] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration purposes and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A method for extracting entities from text, characterized in that: The method comprises the following steps: S100, for any entity keyword among several entity keywords extracted from a given case text, calculating an entity score corresponding to the entity keyword based on word frequency, word attribute, character length and position information of the entity keyword in the given case text; S200, when the entity score corresponding to the entity keyword is not less than a preset entity score threshold, the entity keyword is used as the extracted target entity word; On the contrary, the preset entity words in the preset expert entity library are traversed, and when there is a preset entity word that is the same as the entity keyword, the entity keyword is used as the extracted target entity word; S300, when there is no preset entity word identical to the entity keyword, find the target sentence containing the entity keyword from the given case text, and perform character decomposition and character combination on the target sentence to generate a plurality of recombined entity words; S400, traverse the preset entity words in the preset expert entity library again, and when there is no preset entity word that is the same as any recombined entity word, delete the entity keyword. Otherwise, determine the final extracted target entity word based on the recombined entity word corresponding to the same preset entity word.
2. The method for extracting entities from text according to claim 1, characterized in that: In step S100, the entity score corresponding to the entity keyword is calculated through the following steps: S101, obtaining an index-confidence fitting curve corresponding to word frequency and character length respectively according to a number of historical case texts; S102, according to the word frequency and character length of the entity keyword in the given case text, respectively finding the word frequency confidence and character length confidence corresponding to the entity keyword; S103, based on the word frequency confidence, character length confidence, pre-acquired word attribute confidence and position confidence corresponding to the entity keyword, and the preset indicator weights corresponding to the word frequency confidence, character length confidence, word attribute confidence and position confidence, a weighted sum is calculated to obtain the entity score corresponding to the entity keyword.
3. The method for extracting entities from text according to claim 1, characterized in that: In step S100, the entity score corresponding to the entity keyword is calculated through the following steps: S110, based on the obtained case text samples, according to the preset character quantity intervals, the case text samples are divided into corresponding character quantity intervals; S120, for a number of case text samples corresponding to any character quantity interval, take each case text sample as a given case text and execute steps S100-S500 to obtain a number of target entity words corresponding to each case text sample; S130, calculating the initial extraction accuracy rates corresponding to target entity words with different numbers of characters in any character number interval according to a number of target entity words corresponding to any case text sample and a number of corresponding preset entity word samples; S140, based on a plurality of character number intervals, determining an average of a plurality of initial extraction accuracy rates corresponding to a target entity word of any character number as a final extraction accuracy rate corresponding to a target entity word of any character number; S150, calculating a new entity score corresponding to the entity keyword according to the entity score corresponding to the entity keyword and the final extraction accuracy corresponding to the number of characters of the entity keyword.
4. The method for extracting entities from text according to claim 1, characterized in that: The step of performing character decomposition and character combination on the target sentence to generate a plurality of reorganized entity words comprises the following steps: S301, obtaining the total number of characters K of the target sentence; S302, in the order of the number of characters from 1 to K, respectively use the character-level N-gram model to perform character decomposition and character combination on the target sentence to generate a number of recombined entity words corresponding to each number of characters.
5. The method for extracting entities from text according to claim 1, characterized in that: The step of determining the target entity word finally extracted according to the reorganized entity words corresponding to the same preset entity words comprises the following steps: S401, when the number of reorganized entity words corresponding to the same preset entity word is 1, the reorganized entity words corresponding to the same preset entity word are used as the target entity words finally extracted, and the entity keywords are deleted; S402, when the number of reorganized entity words corresponding to the same preset entity word is greater than 1, obtaining position information of each reorganized entity word corresponding to the same preset entity word in the target sentence; S403, for each acquired position information of the reorganized entity word in the target sentence, determine whether there is character overlap between the reorganized entity words, and when there is no character overlap, use the reorganized entity words corresponding to the same preset entity word as the target entity words finally extracted, and delete the entity keyword; S404, when there are overlapping characters between reorganized entity words, the reorganized entity words with more corresponding characters or the reorganized entity words at the front are taken as the target entity words finally extracted, and the entity keywords are deleted.
6. The method for extracting entities from text according to claim 1, characterized in that: Several entity keywords are extracted from the given case text through the named entity model.
7. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the method for extracting entities from text as described in any one of claims 1-6.
8. An electronic device, characterized in that: The invention comprises a processor and the non-transitory computer-readable storage medium as claimed in claim 7.
Citation Information
Patent Citations
Data mining-oriented text processing system and method
CN105243130A
Knowledge graph-based entity description extraction method and apparatus, and computing device
CN111984794A
Knowledge graph confidence evaluation method and device, electronic equipment and medium
CN115757837A
IDI full-project periodic risk digital assessment method and system
CN119294820A
Named entity extraction device, named entity extraction method, named entity extraction model and program
JP2024032206A