Medical record entity extraction method and device
Patent Information
- Application Number
- CN202311220212.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-20
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-09-20
AI Technical Summary
由于自然语言的复杂性,这种方式适用性不高,每出现一种不适配当前代码的实体描述,就需要通过修改代码来满足相应的实体匹配
[0016]本发明提供了一种病历实体提取方法及装置,该方法包括:提取病历中的病历段落,并对病历段落进行分割,得到不同的病历短句;根据词语词性对病历短句进行分割,并识别得到病历短句对应的词汇数组;其中,词汇数组中包括多个词汇以及各词汇对应的词性;将词汇数组与函数规则进行匹配得到多个病历实体,病历实体有多个词汇组成,函数规则中包括多个词性及各个词性的位置关系。利用规则函数的方法将自然语言转换为抽象语言,当需要产生新的病历实体出现的时候,直接调用相应的函数规则就可以完成对病历实体的提取,避免了需要预先进行复杂编码的情况,同时在使用的过程中提高了适配性。
Smart Images

Figure CN117273001B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information processing technology, and in particular to a method and apparatus for extracting medical record entities. Background Technology
[0002] Extracting entities from medical records is an essential process in clinical and research medicine. Traditional rule-based entity extraction schemes require extensive hard coding, with different rule codes needed for different entity content. Due to the complexity of natural language, this approach has limited applicability; whenever an entity description that doesn't fit the current code appears, the code needs to be modified to meet the corresponding entity matching requirements.
[0003] Therefore, existing solutions for extracting medical record entities often involve processing natural language, and the traditional encoding process is often quite complex and has limitations in adaptability during application. Summary of the Invention
[0004] The purpose of this invention is to provide a method and apparatus for extracting medical record entities, so as to improve the efficiency of medical record entity extraction.
[0005] In a first aspect, the present invention provides a method for extracting medical record entities. The method includes: extracting medical record paragraphs from the medical record and segmenting the medical record paragraphs to obtain different medical record short sentences; segmenting the medical record short sentences according to the parts of speech of words and identifying the vocabulary array corresponding to the medical record short sentences; wherein, the parts of speech of words are the classification features of each word in the medical record short sentences, and the correspondence between the parts of speech of words and each word is pre-stored in a database; the vocabulary array includes multiple words and the parts of speech of each word; matching the vocabulary array with a function rule to obtain multiple medical record entities, wherein the medical record entity is composed of multiple words, and the function rule contains multiple parts of speech and the positional relationship between each part of speech.
[0006] Furthermore, the steps of segmenting medical record phrases according to word parts of speech and identifying different word arrays include: segmenting medical record phrases using a preset medical lexicon to obtain different words; determining the part of speech of the words; associating the words with the part-of-speech codes corresponding to the part of speech to generate corresponding word arrays.
[0007] Furthermore, the step of matching the vocabulary array with function rules to obtain multiple medical record entities includes: extracting part-of-speech encoding strings from the vocabulary array; matching the part-of-speech encoding strings with multiple function rules pre-stored in the function database to obtain multiple medical record entities corresponding to each function rule; wherein, each function rule corresponds to multiple different functions, and each function corresponds to a matching rule and an output rule.
[0008] Furthermore, the above functions include at least one of the following: an empty function; a ver function; a hor function; an any function; an eq function; and a gt function. The empty function matches an entity, and its output rule is to add the matched words from the vocabulary array to the current entity. The ver function matches entities multiple times consecutively, and its output rule is to create a new entity for each matched word from the vocabulary array, with the new entity copying all attributes of the old entity. The hor function matches entities multiple times consecutively, and its output rule is to add the matched words from the vocabulary array to the current entity. The any function matches any entity, and its output rule is to add the matched words from the vocabulary array to the current entity; if no match is found, the next function is executed. The eq function matches an entity at a specified position, and its output rule is to add the matched words to the current entity. The gt function matches an entity after a specified position, and its output rule is to add the matched words to the current entity.
[0009] Further, the step of matching the part-of-speech (POS) encoding string with multiple function rules pre-stored in the function database to obtain multiple case entities corresponding to each function rule includes: determining the digit values of different POS codes in the POS encoding string; matching the POS encoding string according to the function rules and generating corresponding POS encoding groups; wherein, the POS encoding group includes the POS code and the digit code of the POS code in the POS encoding string; extracting the corresponding words in the vocabulary array based on the POS encoding group, and combining the words into a medical record entity.
[0010] Furthermore, the steps for extracting corresponding words from the vocabulary array based on part-of-speech coding groups include: determining the digit codes in the part-of-speech coding groups; determining the digits in the vocabulary array to be extracted based on the digit codes; and extracting the words corresponding to the digits in the vocabulary array.
[0011] Furthermore, each function rule is generated by concatenating multiple functions or by nesting multiple functions.
[0012] Furthermore, before matching the part-of-speech encoded string with multiple function rules pre-stored in the function database, the method also includes: sequentially calling preset function rules to match the vocabulary array; if a function rule matches the vocabulary array, then calling the function rule and stopping the calling of other function rules.
[0013] Furthermore, the method also includes: if none of the function rules match the vocabulary array, then stop calling the function rules.
[0014] Secondly, embodiments of the present invention also provide a medical record entity extraction device, wherein the device includes: a paragraph segmentation module, used to extract medical record paragraphs from medical records and segment the medical record paragraphs to obtain different medical record short sentences; a short sentence segmentation module, used to segment the medical record short sentences according to the parts of speech of words and identify the vocabulary array corresponding to the medical record short sentences; wherein the vocabulary array includes multiple words and the parts of speech of each word; and an entity extraction module, used to match the vocabulary array with function rules to obtain multiple medical record entities, wherein the medical record entities are composed of multiple words, and the function rules include multiple parts of speech and the positional relationship of each part of speech.
[0015] The embodiments of the present invention bring the following beneficial effects:
[0016] This invention provides a method and apparatus for extracting medical record entities. The method includes: extracting medical record paragraphs from the medical record and segmenting the paragraphs to obtain different medical record phrases; segmenting the medical record phrases according to word parts of speech and identifying the corresponding vocabulary array; wherein the vocabulary array includes multiple words and the part of speech corresponding to each word; matching the vocabulary array with function rules to obtain multiple medical record entities, each medical record entity consisting of multiple words, and the function rules including multiple parts of speech and the positional relationships of each part of speech. By using a rule-based function method to convert natural language into an abstract language, when new medical record entities need to be generated, the extraction of medical record entities can be completed by directly calling the corresponding function rules, avoiding the need for complex pre-coding and improving adaptability during use.
[0017] Other features and advantages disclosed in this embodiment will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.
[0018] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating a method for extracting medical record entities according to an embodiment of the present invention;
[0021] Figure 2A flowchart illustrating another method for extracting medical record entities provided in an embodiment of the present invention;
[0022] Figure 3 This is a schematic diagram of a medical record entity extraction device provided in an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0025] Traditional rule-based solutions for extracting medical record entities often directly process natural language, which requires extensive hard coding. Furthermore, different rule codes need to be written for different entity content.
[0026] Due to the complexity of natural language, this approach has limited applicability. Whenever an entity description becomes incompatible with the current code, the code needs to be modified to meet the corresponding entity matching requirements. Therefore, traditional coding processes are often complex and have limitations in adaptability during application.
[0027] Based on this, embodiments of the present invention provide a method and apparatus for extracting medical record entities. This technology can solve the aforementioned technical problems and improve the efficiency of medical record entity extraction. To facilitate understanding of the embodiments of the present invention, a detailed description of a method for extracting medical record entities disclosed in the embodiments of the present invention will be provided first.
[0028] Example 1
[0029] The method for extracting medical record entities provided in this embodiment of the invention. Figure 1 This is a flowchart illustrating a method for extracting medical record entities according to an embodiment of the present invention.
[0030] Depend on Figure 1 As can be seen, the above methods include:
[0031] Step S101: Extract medical record paragraphs from the medical records and segment the medical record paragraphs to obtain different medical record short sentences;
[0032] Entity extraction is the process of identifying medical entities such as diseases and symptoms from unstructured medical texts. Therefore, in the description of a patient's condition recorded in an electronic medical record, the description of an entity is often composed of words of different parts of speech including nature, degree, location, symptom, time and so on. The task of entity extraction is to extract the main entity such as symptoms and related attributes of the entity, wherein nature, degree, time and the like are all entity attributes.
[0033] Therefore, in practical applications, short sentences in medical record paragraphs can be segmented according to medical record ending words, and the paragraph is split into individual short sentences, so as to obtain different medical record short sentences, and the obtained medical record short sentences are processed to obtain the medical record entities required for subsequent processing.
[0034] In practical applications, ending words for medical records can be collected based on actual medical record data, and corresponding ending words can be extracted according to different medical record filling habits. The following description takes various medical record descriptions that may occur in actual situations as examples for illustration:[|END]]
[0035] For example, in the medical record description "The patient has had persistent abdominal pain, nausea and vomiting for 3 days.", based on previous data experience, the combination of the time description word "3 days", the modal particle "了" and the punctuation "." can be used as an ending word. The sentence where this ending word is located is regarded as one medical record short sentence, a segmentation boundary is set after this ending word, and the ending word of the next medical record short sentence is identified.
[0036] Meanwhile, in the above process of splitting a paragraph into individual short sentences, the solution provided by the present application is not limited to handwritten medical records by medical staff or electronic medical records input on digital media.
[0037] Specifically, when recognizing handwritten medical records, the combination of modal particles and punctuation marks can be used as the ending word; the blanks between sentences existing in handwritten medical records can also be identified as the aforementioned ending words to segment medical record paragraphs; when recognizing electronic medical records input on digital media, in addition to judging sentence ending words in the aforementioned manner, character markers typed in electronic medical records can also be read as the aforementioned ending words. For example, when sentence-changing symbols commonly used such as carriage return symbols and space symbols are typed in a medical record paragraph, these sentence-changing symbols can also be used as ending words to segment the medical record paragraph.
[0038] Step S102: segment medical record short sentences according to word parts of speech, and identify and obtain a word array corresponding to the medical record short sentence; wherein the word parts of speech are classification features of each word in the medical record short sentence, the corresponding relationship between the word parts of speech and each word is pre-stored in a database; and the word array includes a plurality of words and the corresponding part of speech of each word;
[0039] Specifically, the aforementioned word parts of speech can serve as a classification standard for vocabulary in medical record language. It can be pre-stored in a database according to different needs, and different word-of-speech definitions can be called for different systems.
[0040] Here, we illustrate one possible classification method: The part of speech for words like "continuous" and "occasionally" can be defined as a quality, denoted as x; the part of speech for words like "abdominal pain," "vomiting," and "nausea" can be defined as a symptom, denoted as z; and the part of speech for words like "3 days," "a week," and "half a month" can be defined as time, denoted as t. In practical applications, other classification methods can also be used to define parts of speech; there are no restrictions on the classification method used here.
[0041] After segmenting short sentences in a medical record paragraph, keywords can be identified and extracted from these sentences, thereby obtaining a vocabulary array corresponding to each sentence. Furthermore, the part-of-speech information from this vocabulary array can be retrieved from a pre-stored database. Those skilled in the art should understand that segmenting medical record sentences using part-of-speech is not intended to limit this invention; other methods of segmenting medical record sentences to obtain vocabulary arrays are also within the scope of this invention, such as segmenting medical record sentences using machine learning algorithms.
[0042] Taking a patient's chief complaint as "persistent abdominal pain, nausea, and vomiting for 3 days" as an example, the vocabulary array for this short sentence is [persistent / x, abdominal pain / z, nausea / z, vomiting / z, 3 days / t]. The English words after the slashes represent the parts of speech of these words. Specifically, different words in the database can be set to different parts of speech such as nature / x, degree / c, location / b, symptom / z, and time / t. The specific parts of speech settings are not restricted here.
[0043] Step S103: Match the vocabulary array with the function rules to obtain multiple medical record entities. Each medical record entity consists of multiple words, and the function rules include multiple parts of speech and the positional relationships of each part of speech.
[0044] Specifically, the above function rules consist of multiple functions. A function can be composed of a function description and corresponding parameters. The function description indicates the matching and output methods of the function, while the corresponding parameters describe the parts of speech of the entity words to be matched. Because it is necessary to match entity words with various parts of speech, the parameter values can change dynamically. Each function has matching rules and output rules corresponding to its function.
[0045] In practical applications, different functions can be selected and combined according to different needs. The parts of speech in the vocabulary array are used as the corresponding parameter input values in the function description to extract the words corresponding to different parts of speech. Thus, different combinations of functions can yield different positional relationships of parts of speech, resulting in word arrangements with different positional relationships, thereby enabling the extraction of medical record entities from the required vocabulary array.
[0046] This invention provides a method for extracting medical record entities. This method extracts medical record paragraphs and segments them to obtain different short sentences. The short sentences are then segmented according to word parts of speech, and a corresponding vocabulary array is identified. This vocabulary array includes multiple words and their corresponding parts of speech. The vocabulary array is then matched with a function rule to obtain multiple medical record entities. Each medical record entity consists of multiple words, and the function rule includes multiple parts of speech and their positional relationships. By using a rule-based function to convert natural language into an abstract language, when new medical record entities need to be generated, the extraction can be completed simply by inputting the corresponding function rule, avoiding the need for complex pre-coding and improving adaptability during use.
[0047] Example 2
[0048] exist Figure 1 Based on the method shown, this invention also provides another method for extracting medical record entities. Figure 2 This is a flowchart illustrating another method for extracting medical record entities provided in an embodiment of the present invention.
[0049] Depend on Figure 2 As seen, the method includes the following steps:
[0050] Step S201: Extract medical record paragraphs from the medical records and segment the medical record paragraphs to obtain different medical record short sentences;
[0051] In practical applications, short sentences within a medical record paragraph can be segmented based on the closing words. Specifically, the entire text can be broken down into short sentences according to punctuation marks, resulting in different medical record sentences. These sentences can then be processed to obtain the required medical record entities. Furthermore, the specific algorithms used to segment the medical record paragraph and obtain different medical record sentences are not specifically limited here.
[0052] Step S202: Use a preset medical lexicon to segment the medical record sentences into words to obtain different vocabulary;
[0053] Specifically, a pre-trained medical lexicon can be used to segment short sentences in medical records. After segmentation, useless redundant single characters or phrases may appear in the segmented short sentences, and these single characters or phrases can be eliminated during the segmentation process of the short sentences.
[0054] Taking the short sentence "Abdominal pain, nausea and vomiting have persisted for 3 days." as an example, the short sentence can be segmented into: "持续了腹痛恶心呕吐3天." According to the preset medical lexicon, the above segmentation result contains non-medical single character "了" in addition to medical terms, so this single character "了" can be eliminated during the segmentation process to improve the accuracy of word segmentation.
[0055] In practical applications, after segmenting short sentences, redundant punctuation marks often remain. Therefore, interfering punctuation marks also need to be eliminated to improve the accuracy of word segmentation.
[0056] Taking the segmentation result of the above "持续了腹痛恶心呕吐3天." as an example, after eliminating the single character "了", the obtained segmentation result still contains the punctuation mark ".", so this punctuation mark also needs to be removed, thereby obtaining the required segmentation result: "持续腹痛恶心呕吐3天".
[0057] Step S203: determining the part of speech of a word, associating the word with a part of speech code corresponding to the part of speech, and generating a corresponding word array;
[0058] In practical application, after parts of speech are defined for nouns in advance, they are stored in a preset database table, that is, there is a medical lexicon with relevant parts of speech marked. Still taking the above segmentation result "持续腹痛恶心呕吐3天" as an example, the segmentation result can be determined as a segmentation array [持续 / x, 腹痛 / z, 恶心 / z, 呕吐 / z, 3天 / t].
[0059] Step S204: extracting a part of speech code string from the word array;
[0060] Specifically, before determining function rules, part-of-speech analysis needs to be performed on multiple medical records, and the parts of speech of medical record features can be counted according to requirements, so as to generate a part-of-speech arrangement order corresponding to the requirements, and then different functions are selected and arranged according to the arrangement order and the arrangement scheme of medical record features corresponding to the arrangement order, so as to generate corresponding function rules;
[0061] Since function rules are a tool for processing abstract language, it is necessary to extract part of speech codes from the word array and generate corresponding part of speech code strings.
[0062] Taking the vocabulary array [persistent / x, abdominal pain / z, nausea / z, vomiting / z, 3 days / t] as an example, extracting the part-of-speech encoding string will yield the part-of-speech encoding string "xzzzt".
[0063] Step S205: Match the part-of-speech encoding string with multiple function rules pre-stored in the function database to obtain multiple medical record entities corresponding to each function rule; wherein, each function rule corresponds to multiple different functions, and each function corresponds to a matching rule and an output rule.
[0064] Specifically, each function rule corresponds to multiple different functions, and each function corresponds to a matching rule and an output rule.
[0065] In practical applications, functions include at least one of the following: empty function; ver function; hor function; any function; eq function; gt function;
[0066] The function description for an empty function can be “()”; the corresponding matching rule is to match an entity, and the output rule for an empty function is to put the matched words into the current entity.
[0067] The function description for the ver function can be "ver()"; the corresponding matching rule is to match entities multiple times consecutively, and the output rule for the ver function is that each matched word will create a new entity storage, and the new entity will copy all the attributes of the old entity.
[0068] The function description for the hor function can be "hor()"; the matching rule is to match entities multiple times consecutively; the output rule for the hor function is to put the matched words into the current entity.
[0069] The function description for the any function can be "any()"; the matching rule is to match any entity; the output rule for the any function is to add the matched word to the current entity, and execute the post-function if no match is found.
[0070] The function description for the eq function can be "eq()"; the corresponding matching rule is to match an entity at a specified position, and the output rule for the eq function is to put the matched word into the current entity;
[0071] The function description for the gt function can be "gt()"; the corresponding matching rule is to match an entity following a specified position, and the output rule for the gt function is to add the matched word to the current entity.
[0072] Taking a practical example, if we want to match any entity of type "symptom", where the function for matching any entity is defined as `any()`, and the part-of-speech encoding of "symptom" is `z`, then the function can be written as `any(z)`. This achieves the goal of extracting the entity to be matched.
[0073] In practical applications, special function rules can be created to match certain irregular entities. Different functions can also be nested and combined. The specific function rule is ultimately determined by the output entity content, so the defined rules can be stored in a rule table in the database for continuous database improvement. This supports more complex semantic matching, such as ver((x)ver(z)), where the function parameter can be one or more functions.
[0074] Specifically, the process of matching the part-of-speech encoded string with multiple function rules pre-stored in the function database to obtain multiple case entities corresponding to each function rule can be achieved by the following steps A1-A3:
[0075] Step A1: Determine the digit values of the different part-of-speech codes in the part-of-speech coding string;
[0076] Step A2: Match the part-of-speech encoding string according to the function rules and generate the corresponding part-of-speech encoding group;
[0077] Step A3: Extract the corresponding words from the vocabulary array based on part-of-speech coding groups, and combine the words into a medical record entity.
[0078] The aforementioned part-of-speech coding group includes part-of-speech codes and their digital encoding within the part-of-speech coding string;
[0079] Therefore, the process of extracting the corresponding words from the vocabulary array based on part-of-speech coding groups can be achieved by the following steps B1-B3:
[0080] Step B1: Determine the digit codes in the part-of-speech coding group;
[0081] In practical applications, the positional relationship of each part of speech in a part-of-speech coding group can be represented by a digital code.
[0082] Taking the part-of-speech encoding string "xzzzt" as an example again, the numbers in [x / 0z / 1t / 2z / 3t / 4] represent the digits of the part-of-speech encoding before the forward slash " / ". In practical applications, as in the example above, computers often use 0 to represent the first element when sorting; therefore, the first part-of-speech encoding in the string starts with 0.
[0083] Step B2: Determine the digits in the vocabulary array to be extracted based on the digit encoding;
[0084] Step B3: Extract the words at the corresponding positions from the word array.
[0085] Specifically, taking a function rule of any(x)ver(z)any(t) as an example, entity extraction is performed on the above vocabulary array [continuous / x, abdominal pain / z, nausea / z, vomiting / z, 3 days / t].
[0086] First, the function in the rule needs to be parsed, and entity extraction is performed according to the matching and output rules of the function. The specific matching process is as follows: first, match the word with part of speech x, then match the word with part of speech z multiple times, and finally match any word with part of speech t. Since the output rule of the ver(z) function is that a new entity needs to be generated every time an entity is matched, the final output result is three part-of-speech encoding groups [x / 0z / 1t / 4, x / 0z / 2t / 4, x / 0z / 3t / 4].
[0087] Based on the positional markers of characters in the part-of-speech tag relative to the vocabulary array, entities in the vocabulary array are extracted. For example, [x / 0z / 1t / 4] corresponds to the 1st, 2nd, and 5th characters in the vocabulary array. Since computers often use 0 to represent the first element when sorting, the digit encoding of the first part-of-speech tag in the tag string starts with 0, where 0 represents the 1st character, 2 represents the 3rd character, and 4 represents the 5th character. Therefore, the entities extracted from [x / 0z / 1t / 4] are: [persistent] (part of speech x at the 1st character), [abdominal pain] (part of speech z at the 2nd character), and [3 days] (part of speech t at the 5th character). Thus, the final medical record entities are [persistent, abdominal pain, 3 days].
[0088] In practical applications, before matching the part-of-speech encoded string with multiple function rules pre-stored in the function database, the method also includes: sequentially calling preset function rules to match the vocabulary array; if a function rule matches the vocabulary array, then calling the function rule and stopping the call request.
[0089] If none of the function rules match the vocabulary array, then stop calling the function rules.
[0090] Example 3
[0091] This invention also provides a medical record entity extraction device, such as... Figure 3 The diagram shown is a structural schematic of a medical record entity retrieval device provided in an embodiment of the present invention. The device includes:
[0092] The paragraph segmentation module 301 is used to extract medical record paragraphs from medical records and segment the medical record paragraphs to obtain different medical record short sentences;
[0093] The short sentence segmentation module 302 is used to segment medical record short sentences according to word parts of speech and identify the corresponding vocabulary array of the medical record short sentences; wherein, word parts of speech are the classification features of each word in the medical record short sentence, and the correspondence between word parts of speech and each word is pre-stored in the database; the vocabulary array includes multiple words and the corresponding parts of speech of each word;
[0094] The entity extraction module 303 is used to match the vocabulary array with the function rules to obtain multiple medical record entities. The medical record entities are composed of multiple words, and the function rules contain multiple parts of speech and the positional relationships between each part of speech.
[0095] The aforementioned short sentence segmentation module 302 is also used to segment medical record short sentences using a preset medical lexicon to obtain different words;
[0096] Determine the part of speech of a word, associate the word with the part-of-speech code corresponding to the part of speech, and generate a corresponding word array.
[0097] The entity extraction module 303 described above is also used to extract part-of-speech encoding strings from the vocabulary array; match the part-of-speech encoding strings with multiple function rules pre-stored in the function database to obtain multiple case entities corresponding to each function rule; wherein, each function rule corresponds to multiple different functions, and each function corresponds to a matching rule and an output rule.
[0098] Specifically, the functions include at least one of the following: an empty function; a ver function; a hor function; an any function; an eq function; and a gt function. The empty function matches an entity, and its output rule is to add the matched words from the vocabulary array to the current entity. The ver function matches entities multiple times consecutively, and its output rule is to create a new entity for each matched word from the vocabulary array, with the new entity copying all attributes of the old entity. The hor function matches entities multiple times consecutively, and its output rule is to add the matched words from the vocabulary array to the current entity. The any function matches any entity, and its output rule is to add the matched words from the vocabulary array to the current entity; if no match is found, the next function is executed. The eq function matches an entity at a specified position, and its output rule is to add the matched words to the current entity. The gt function matches an entity after a specified position, and its output rule is to add the matched words to the current entity.
[0099] The entity extraction module 303 is also used to determine the digit values of different part-of-speech codes in the part-of-speech encoding string; to match the part-of-speech encoding string according to the function rules and generate the corresponding part-of-speech encoding group; wherein, the part-of-speech encoding group includes the part-of-speech code and the digit code of the part-of-speech code in the part-of-speech encoding string; to extract the corresponding words in the vocabulary array based on the part-of-speech encoding group and to combine the words into a medical record entity.
[0100] Specifically, the entity extraction module 303 is also used to determine the digit code in the part-of-speech coding group; determine the digit in the word array to be extracted based on the digit code; and extract the word corresponding to the digit in the word array.
[0101] Each of the above function rules is generated by concatenating multiple functions or by nesting multiple functions.
[0102] The entity extraction module 303 is also used to sequentially call function rules to match the vocabulary array; if a function rule matches the vocabulary array, the function rule is called and the matching of other function rules is stopped.
[0103] If none of the function rules match the vocabulary array, then stop calling the function rules.
[0104] Example 4
[0105] This embodiment provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the medical record entity extraction method.
[0106] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a medical record entity extraction method.
[0107] See Figure 4 The diagram shows the structure of an electronic device, which includes a memory 41 and a processor 42. The memory 41 stores a computer program that can run on the processor 42. When the processor executes the computer program, it implements the steps provided by the above-mentioned medical record entity extraction method.
[0108] like Figure 4 As shown, the device also includes a bus 43 and a communication interface 44, with the processor 42, the communication interface 44 and the memory 41 connected via the bus 43; the processor 42 is used to execute executable modules, such as computer programs, stored in the memory 41.
[0109] The memory 41 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 44 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0110] Bus 43 can be an ISA bus, PCI bus, or EISA bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 4 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0111] The memory 41 stores the program, and the processor 42 executes the program after receiving the execution instruction. The method executed by the medical record entity retrieval device disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 42, or implemented by the processor 42. The processor 42 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 42 or by instructions in the form of software. The processor 42 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this invention can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 41, and processor 42 reads information from memory 41 and, in conjunction with its hardware, completes the steps of the above method.
[0112] Furthermore, this embodiment of the invention also provides a machine-readable storage medium storing machine-executable instructions. When the machine-executable instructions are invoked and executed by the processor 42, the machine-executable instructions cause the processor 42 to implement the above-described medical record entity extraction method.
[0113] The electronic devices and computer-readable storage media provided in the embodiments of the present invention have the same technical features, so they can solve the same technical problems and achieve the same technical effects.
[0114] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0115] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
Claims
1. A method for extracting medical record entities, characterized in that, The method includes: Extract medical record segments from the medical records and segment the medical record segments to obtain different medical record short sentences; The medical record sentences are segmented based on word parts of speech, and a vocabulary array corresponding to each sentence is identified. The word parts of speech are the classification features of each word in the medical record sentence, and the correspondence between word parts of speech and each word is pre-stored in a database. The vocabulary array includes multiple words and their corresponding parts of speech. Extract part-of-speech tag strings from the vocabulary array; match the part-of-speech tag strings with multiple function rules pre-stored in a function database to obtain multiple medical record entities corresponding to each function rule; wherein, each function rule corresponds to multiple different functions, and each function corresponds to a matching rule and an output rule; each medical record entity is composed of multiple words, and each function rule contains multiple parts of speech and the positional relationship between each part of speech; each function rule is generated by concatenating multiple functions or by compounding and nesting multiple functions; The function includes at least one of the following: an empty function; a ver function; a hor function; an any function; an eq function; a gt function; Wherein, the matching rule corresponding to the empty function is to match an entity, and the output rule corresponding to the empty function is to put the matched words in the word array into the current entity; The matching rule corresponding to the ver function is to match entities multiple times consecutively, and the output rule corresponding to the ver function is that each time a word in the matched word array is used, a new entity is created and stored, and the new entity copies all the attributes of the old entity. The matching rule corresponding to the hor function is to match entities multiple times consecutively, and the output rule corresponding to the hor function is to put the matched words in the word array into the current entity. The matching rule corresponding to the any function is to match any entity, and the output rule corresponding to the any function is to add the matched words in the word array to the current entity, and to execute the next function if no match is found. The matching rule corresponding to the eq function is to match an entity at a specified position, and the output rule corresponding to the eq function is to put the matched word into the current entity; The matching rule corresponding to the gt function is to match an entity following a specified position, and the output rule corresponding to the gt function is to place the matched word into the current entity.
2. The method for extracting medical record entities according to claim 1, characterized in that, The steps of segmenting the medical record sentences according to the parts of speech and identifying different word arrays include: The medical record sentences are segmented using a pre-defined medical thesaurus to obtain different words; Determine the part of speech of the word, associate the word with the part-of-speech code corresponding to the part of speech, and generate a corresponding word array.
3. The method for extracting medical record entities according to claim 1, characterized in that, The step of matching the part-of-speech encoded string with multiple function rules pre-stored in a function database to obtain multiple case entities corresponding to each function rule includes: Determine the digit values of the different part-of-speech codes in the part-of-speech coding string; The part-of-speech encoding string is matched according to the function rules to generate a corresponding part-of-speech encoding group; wherein, the part-of-speech encoding group includes the part-of-speech encoding and the digital encoding of the part-of-speech encoding in the part-of-speech encoding string; Based on the part-of-speech coding group, the corresponding words in the vocabulary array are extracted, and the words are combined into the medical record entity.
4. The method for extracting medical record entities according to claim 3, characterized in that, The step of extracting the corresponding words from the vocabulary array based on the part-of-speech coding group includes: Determine the digit codes in the part-of-speech coding group; The digits in the vocabulary array to be extracted are determined based on the digit encoding. Extract the words at the corresponding positions from the vocabulary array.
5. The method for extracting medical record entities according to claim 1, characterized in that, Before matching the part-of-speech encoded string with multiple function rules pre-stored in the function database, the method further includes: The preset function rules are called sequentially to match the vocabulary array; If the function rule matches the vocabulary array, then the function rule is invoked, and invocation of other function rules is stopped.
6. The method for extracting medical record entities according to claim 5, characterized in that, The method further includes: If none of the function rules match the vocabulary array, then the function rules will no longer be invoked.
7. A medical record entity retrieval device, characterized in that, The apparatus is used to perform the method of any one of claims 1-6, and the apparatus comprises: The paragraph segmentation module is used to extract medical record paragraphs from the medical records and segment the medical record paragraphs to obtain different medical record short sentences; The short sentence segmentation module is used to segment the medical record short sentences according to the parts of speech of words and identify the vocabulary array corresponding to the medical record short sentences; wherein, the parts of speech of words are the classification features of each word in the medical record short sentences, and the correspondence between the parts of speech of words and each word is pre-stored in a database; the vocabulary array includes multiple words and the parts of speech of each word. The entity extraction module is used to match the vocabulary array with the function rules to obtain multiple medical record entities. The medical record entities are composed of multiple words, and the function rules contain multiple parts of speech and the positional relationships between each part of speech.
Citation Information
Patent Citations
Medical knowledge graph construction method and device based on electronic medical record
CN110427491A
Multi-task question and answer driven medical entity relationship extraction method
CN113609868A