Medical record text processing method and device, electronic equipment and computer readable storage medium
By retrieving candidate codes for non-standard name entities from the coding database and then searching the standard name database, the standard name of the non-standard name entity is determined using statistical information of aligned candidate pairs. This solves the problem of automated standardization processing of medical record texts in existing technologies, and improves efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA INNOVATION PRIVATE LIMITED
- Filing Date
- 2021-05-12
- Publication Date
- 2026-05-22
AI Technical Summary
Existing technologies cannot automatically discover the matching relationship between natural language medical record descriptions and standard names, resulting in low efficiency and high cost of manual annotation.
By retrieving candidate codes for non-standard name entities from the coding database and then searching the standard name database, the standard name of the non-standard name entity is determined using statistical information from the aligned candidate pairs.
It has achieved automated and standardized processing of medical record texts, reducing the workload of manual annotation and improving the efficiency of medical record text content recognition and management.
Smart Images

Figure CN115344664B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text processing technology, and in particular to a method and apparatus for processing medical case texts, an electronic device, and a computer-readable storage medium. Background Technology Background Technology
[0003] With the development of big data technology, more and more industries are adopting electronic databases to manage the data generated in their operations. However, while a large amount of operational data generated in daily operations has been digitized using computer technology, actual operational staff typically use natural language to write various records. Big data databases, however, only allow records using specific terminology for easier classification and retrieval. For example, in the hospital field, doctors often use natural language to describe patients' conditions and symptoms when writing medical records. However, medical terminology databases typically assign only one specific term for the same condition or symptom. Therefore, it's necessary to match the doctor's natural language medical record with the corresponding terminology in the database to manage all patient medical data using big data technology. This matching process is highly dependent on the matching database. Therefore, it's necessary to continuously add new standard terms and their corresponding natural language fields to the matching database to keep up with the development of new technologies.
[0004] Therefore, a technical solution is needed that can automatically mine the correspondence between natural language medical case descriptions and standard names from a large corpus of natural language. Summary of the Invention
[0005] This application provides a method and apparatus for processing medical record text, an electronic device, and a computer-readable storage medium to address the shortcomings of existing technologies that cannot automatically mine the matching relationship between natural language medical record descriptions and standard names.
[0006] To achieve the above objectives, embodiments of this application provide a method for processing medical record text, including:
[0007] Obtain the medical record text to be processed, wherein the medical record text to be processed includes multiple entities, including standard name entities and non-standard name entities;
[0008] A search is performed in the encoding database to obtain candidate codes corresponding to the non-standard name entity;
[0009] The candidate codes are used to search the standard name database to obtain multiple search results;
[0010] The multiple search results are respectively paired with the non-standard name entity to form multiple first alignment candidate pairs;
[0011] The standard name of the non-standard name entity is determined based on the statistical information of the plurality of first alignment candidate pairs.
[0012] This application also provides a medical record text processing device, including:
[0013] The first acquisition module is used to acquire the medical record text to be processed, wherein the medical record text to be processed includes multiple entities, including standard name entities and non-standard name entities;
[0014] The first retrieval module is used to perform a retrieval in the encoding database to obtain candidate codes corresponding to the non-standard name entity;
[0015] The second retrieval module is used to search the standard name database using the candidate codes to obtain multiple retrieval results;
[0016] The alignment candidate pair generation module is used to form multiple first alignment candidate pairs by combining multiple search results with the non-standard name entities;
[0017] The first determining module is used to determine the standard name of the non-standard name entity based on the statistical information of the plurality of first alignment candidate pairs.
[0018] This application also provides an electronic device, including:
[0019] Memory, used to store programs;
[0020] A processor is configured to run the program stored in the memory, wherein the program executes the medical record text processing method provided in the embodiments of this application.
[0021] This application also provides a computer-readable storage medium storing a computer program executable by a processor, wherein the program, when executed by the processor, implements the medical record text processing method provided in this application.
[0022] The medical record text processing method, apparatus, electronic device, and computer-readable storage medium provided in this application retrieve the possible candidate codes corresponding to non-standard name entities in the medical record text to be processed from the coding database, and search for these candidate codes in the standard name database to find and determine whether the non-standard name entity is used in the standard name database. Based on this, the statistical information of such a pair of aligned candidate pairs can be used to determine whether the non-standard name entity in the aligned candidate pair is frequently used with the standard name, thereby determining a qualified standard name for the non-standard name entity.
[0023] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0024] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0025] Figure 1 A schematic diagram illustrating an application scenario of the medical record text processing method provided in this application embodiment;
[0026] Figure 2 A flowchart of one embodiment of the medical record text processing method provided in this application;
[0027] Figure 3 A flowchart of another embodiment of the medical record text processing method provided in this application;
[0028] Figure 4 A schematic diagram of the structure of an embodiment of the medical record text processing device provided in this application;
[0029] Figure 5 A schematic diagram of the structure of an embodiment of the electronic device provided in this application. Detailed Implementation
[0030] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0031] Example 1
[0032] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0033] With the development of big data technology, more and more industries are adopting electronic databases to manage the data generated in their operations. However, while a large amount of operational data generated in daily operations has been digitized using computer technology, actual operational staff typically use natural language to write various records. Big data databases, on the other hand, only allow records using specific terminology to facilitate classification and retrieval. For example, in the medical field, when doctors write patient medical records, they often use natural language familiar to them to describe the patient's condition and corresponding symptoms. However, the corresponding medical terminology database usually assigns only one standard name to the same condition or symptom. Therefore, in managing hospital operations, it is necessary to match the doctor's natural language records to the standard names in the corresponding database to enable the use of big data technology to manage all patient medical data.
[0034] For example, in medical management within hospitals or healthcare institutions, big data technology is needed for various medical statistics. A crucial foundation for such statistics is the standardization of medical terminology, which involves converting doctors' descriptions of diagnoses or conditions written in natural language during treatment into standardized medical terms. However, in actual clinical practice, different doctors may describe the same condition, surgery, medication, examination, or even symptom differently based on their own writing styles. In such cases, healthcare institutions collect a large amount of medical records written in natural language and need to perform various statistical analyses based on these records.
[0035] In existing technologies, these medical record texts need to be matched with specific standard names in a database to facilitate the classification and statistical analysis of these medical records. For example, when classifying patients, Diagnosis Related Groups (DRGs) are typically used to categorize patients into 500-600 groups based on attributes such as length of hospital stay, clinical diagnosis, symptoms, surgery, disease severity, comorbidities, and complications. In other words, in the medical records doctors routinely write for patients, these attributes are expressed in the doctor's natural language. Therefore, when classifying patients based on DRGs, it is necessary to match these natural language forms or medical record texts that include natural language to predetermined standard classification terms (i.e., standardized codes).
[0036] However, such matching processing is highly dependent on a standard name database. For example, a standard name database typically stores the correspondence between medical record texts in various natural languages and standard names. For instance, the standard name "hypertension" can correspond to various natural language expressions such as "high blood pressure" and "high blood pressure reading." Therefore, if the database only records the natural language correspondence for the standard term "hypertension" as "high blood pressure," the accuracy of matching other natural languages will be low. To address this, in existing technologies, manual annotation is typically used to generate aligned corpora from a large amount of collected natural language data (e.g., medical records) based on the standard terminology list used in the standard name database. This manual annotation obviously requires significant human resources, especially demanding high professional skills from the personnel performing the annotation and alignment. For example, those performing standard alignment need not only to understand the meaning and common symptoms of the diseases corresponding to the standard terms, but also to be able to quickly read and understand the natural language text describing the patient's condition in various natural language forms, and especially to quickly and accurately identify the corresponding fields. Therefore, manual annotation is not only inefficient but also labor-intensive.
[0037] therefore, Figure 1 This is a schematic diagram illustrating an application scenario of the medical record text processing method provided in this application embodiment. For example... Figure 1As shown, medical record texts written in various natural language forms can be obtained from various text sources. In this embodiment, the text to be processed obtained in this way includes natural language text, that is, medical record text written in natural language. For example, in the above-mentioned medical management field, such natural language medical record text can be obtained from data sources such as medical record databases storing medical records, medical paper databases, hospital management platforms, or even doctors' forums, and the natural language text can be parsed into various entities including standard name entities and non-standard name entities by performing, for example, structured parsing on such input medical record text. In this embodiment, these entities can include standard name entities and non-standard name entities. Of course, the entities can be further preprocessed afterward, for example, numbers and serial numbers, such as 1, or (1), can be identified and removed, or punctuation marks contained in the entities can be removed, such as question marks (?), commas (,), etc., or the parsed entities can be further parsed and split. For example, after the structured parsing of the medical case text, the obtained entities include concatenated entities such as "hypertension + diabetes". Therefore, such entities can be further split. For example, they can be split into two entities: "hypertension" and "diabetes". Thus, entities preprocessed in this way are suitable for retrieval and matching in the embodiments of this application.
[0038] For example, in this embodiment, the parsed or preprocessed entities can be directly input into the standard name database for retrieval, thus excluding natural language entities that can directly match the standard names. For instance, some doctors are accustomed to using standard names to describe illnesses when writing medical records, so such natural language entities can directly match the standard names in the standard name database, eliminating the need for alignment annotation. Therefore, in this embodiment, these natural language entities that already match the standard names can be removed from the acquired medical record text.
[0039] Subsequently, in this embodiment, the obtained non-standard name entities can be input into the coding database for retrieval to obtain candidate codes corresponding to the non-standard name entities. These candidate codes are then used to search the standard name database to find all retrieval results related to the candidate codes of such non-standard name entities. These retrieval results can be sorted, for example, according to the matching degree between the retrieval results and the corresponding entities, and the top-ranked retrieval results are selected as the corresponding retrieval results. That is, in this embodiment, since natural language entities that are completely identical to standard terms have been excluded, the remaining entities will encounter situations where they cannot be directly matched or will result in matching errors during actual matching. Therefore, in this embodiment, these entities can be retrieved from the coding database to obtain corresponding candidate codes, and then the candidate codes can be used to search the standard name database to obtain relevant standard name retrieval results. In this embodiment, these standard name retrieval results can all be related to the medical case text to be processed, and can be sorted according to relevance, with a predetermined number of top-ranked standard terms selected as candidates.
[0040] Specifically, in this embodiment, the number of search results can be determined by acquiring various attribute information related to the medical record text to be processed. For example, in the medical field, different hospitals have varying levels of informatization, and even different doctors within the same hospital have significantly different abilities and habits. Therefore, considering these differences, when searching for relevant standard names for a non-standard entity, the number of search results can be adjusted based on such difference information. For example, in this embodiment, for hospitals with lower levels of informatization, the number of search results can be set to a larger number, such as 100 or 200 or even more. Alternatively, the difference between the non-standard entity and each standard name can be calculated. When the difference exceeds a predetermined threshold or the difference with a certain number of standard names exceeds a certain threshold, it can be determined that the doctor who wrote the non-standard entity does not frequently use standard names. Therefore, the number of search results for the non-standard entity can be set to a larger number, such as 100 or 200 or even more. Conversely, when the hospital has a high level of informatization, the number can be set to a smaller value, such as 10 or 20 or even lower. When the difference between the non-standard entity and each standard name is less than the predetermined threshold, it can be determined that the doctor to whom the non-standard entity belongs has a good habit of using standardized terminology. Therefore, the number of search results for the non-standard entity can be set to a smaller value, such as 10 or 20 or even lower.
[0041] After obtaining the search results, according to the embodiments of this application, these search results can be standard term candidates that are consistent with the standard terms used in the standard name database. It can be determined that these search results are both standard terms used in the standard name database and have a high relevance to the field to be processed. Therefore, the search results can form a set of aligned data candidates with the matching data.
[0042] Subsequently, the statistical information of the set of alignment data candidates can be calculated to determine whether the set of alignment data candidates can be used as qualified alignment data for standard name databases or machine learning. For example, in this embodiment, the number of times a non-standard entity appears in a predetermined text can be calculated first, and then the number of times the set of alignment data candidates appears in the predetermined text can be calculated. That is, the number of times the determined standard term candidate and the corresponding non-standard entity appear together is counted. For example, if the number of times they appear together is large, it can be said that the non-standard entity is likely to be frequently used to describe the standard term candidate. For example, in the solution of this embodiment, the standard term candidate "hypertension" is determined for the non-standard entity "high blood pressure value", that is, "high blood pressure value" and "hypertension" are combined into a set of alignment data candidates, and if there are descriptions such as "the patient's blood pressure value is high, and he is likely a hypertensive patient" or "the patient's blood pressure value is high, and hypertension treatment can be performed first" in the predetermined text, then it can be determined that the set of alignment data candidates "high blood pressure value" and "hypertension" is a qualified set of alignment data.
[0043] The medical record text processing scheme provided in this application retrieves the possible candidate codes corresponding to non-standard name entities in the medical record text to be processed from the coding database, and searches for these candidate codes in the standard name database to find and determine whether the non-standard name entity is used in the standard name database. Based on this, the statistical information of such a pair of alignment candidate pairs can be used to determine whether the non-standard name entity in the alignment candidate pair is frequently used with the standard name, thereby determining a qualified standard name for the non-standard name entity.
[0044] Therefore, the medical record text processing scheme of this application embodiment can automatically determine matching candidate codes in the coding database based on non-standard name entities in the input medical record text, and then search for corresponding retrieval results in the standard name database based on the candidate codes. This allows for the determination of the standard name of the non-standard name entity based on the statistical information of the aligned candidate pairs formed by the retrieval results and the non-standard name entities, greatly saving the workload of manual annotation. This enables efficient standardization of medical record texts written in natural language by doctors in their daily work. In particular, public medical institutions, private medical institutions, or national medical management agencies, in the process of automating the management of electronic medical records, can use the medical record text processing scheme provided in this application embodiment through technology licensing or service purchase to efficiently establish the correspondence between a large number of non-standard entities and standard names in the medical record text, and based on this, further process the identification and management of the medical record content written by doctors. In particular, as described above, the medical record text processing scheme of this application can match identified non-standard name entities from a large number of medical record texts, or even the original medical record texts, to the corresponding standard names in the standard name database without manual intervention. This adds matching relationships between non-standard name entities that have not yet been labeled and potential standard names. The established matching relationships between non-standard name entities in a large number of medical record texts and standard names in the standard name databases used by these medical institutions or medical management agencies form a very important foundation for the identification and management of medical record text content. This enables efficient and highly accurate identification of medical record text content, greatly improving the management efficiency of medical managers over the medical work of doctors and other primary healthcare personnel.
[0045] The above embodiments illustrate the technical principles and exemplary application framework of the embodiments of this application. The specific technical solutions of the embodiments of this application will be further described in detail below through multiple embodiments.
[0046] Example 2
[0047] Figure 2 This is a flowchart of one embodiment of the medical record text processing method provided in this application. The subject executing this method can be various terminals or server devices with text processing capabilities, or it can be a device or chip integrated into these devices. Figure 2 As shown, the processing method for this medical record text includes the following steps:
[0048] S201, Obtain the medical record text to be processed.
[0049] In step S201, medical record texts written in various natural language formats can be obtained from various medical record text sources. In this embodiment, the medical record text to be processed obtained in this way includes medical record texts written in natural language, that is, medical record texts written in natural language. For example, in the aforementioned medical management field, such natural language medical record texts can be obtained from data sources such as medical record databases storing medical records, medical paper databases, hospital management platforms, or even doctors' forums. The input medical record text to be processed is then parsed using, for example, structured parsing, to parse the natural language text into various entities containing standard name entities and non-standard name entities. In this embodiment, the medical record text can contain standard name entities and non-standard name entities, that is, entities written with standard names in a standard name database and entities not written with such standard names. In this embodiment, the parsed and split entities or pre-processed entities can be directly input into a standard name database for retrieval to exclude natural language entities that can directly match standard terms. For example, some doctors are already accustomed to using standard terminology to describe medical conditions when writing medical records. Therefore, such natural language entities can be directly matched with standard name entities in the standard name database during actual use, thus eliminating the need for alignment annotation of such entities. Therefore, in this embodiment of the application, these natural language entities that are already consistent with standard terminology can be removed from the acquired medical record text first, thereby saving the computational load of subsequent matching and retrieval processing.
[0050] S202, Search the coding database to obtain candidate codes corresponding to non-standard name entities.
[0051] In this embodiment, the non-standard name entities obtained in step S201 can be input into the encoding library for searching, and candidate codes related to the non-standard name entities can be used as the basis for retrieval in the standard name database. For the convenience of subsequent processing, only a portion of the candidate codes can be selected for retrieval. For example, all candidate codes can be sorted, such as by the matching degree between the candidate codes and the corresponding non-standard name entities, and the top few candidate codes can be selected as the corresponding candidate codes.
[0052] In this embodiment, since standard name entities that are completely identical to standard terminology have been excluded, the remaining non-standard name entities will encounter situations where they cannot be directly matched or will result in matching errors during actual matching. Therefore, in this embodiment, these non-standard name entities can be searched in the coding database to obtain relevant candidate codes. In this embodiment, these codes can all be related to the medical record text, and can be sorted according to relevance, with a predetermined number of candidate codes ranked at the top selected as candidates.
[0053] In particular, in this embodiment of the application, the coding database may be a coding database that stores standard terms for a specific industry.
[0054] S203, using candidate codes to search in the standard name database, obtains multiple search results.
[0055] After obtaining candidate codes in step S202, according to embodiments of this application, these candidate codes can be used to further determine which standard names in the standard name database may correspond to the non-standard name entity. In embodiments of this application, standard terms are standard names that can be used in the standard name database. For example, in embodiments of this application, the standard name database may be a database that stores the correspondence between standard terms and natural language fields, so that an industry platform, such as a hospital management system that uses standard terms for data management, can use the standard name database to automatically match text fields written in natural language into the database.
[0056] S204, combine multiple search results with non-standard name entities to form multiple first alignment candidate pairs.
[0057] Therefore, in this embodiment of the application, in step S204, the search results obtained in step S203, namely the standard terms related to non-standard name entities, can be combined with the field to be processed to form a set of alignment candidate pairs.
[0058] S205, determine the standard name of the non-standard name entity based on the statistical information of multiple first alignment candidate pairs.
[0059] In this embodiment, in step S204, the various search results related to non-standard name entities obtained in step S203 are combined with the non-standard name entities to form alignment candidate pairs. Therefore, in step S205, the statistical information of the alignment candidate pair can be calculated to determine whether the alignment candidate pair can be used as a qualified standard name for standard name databases or machine learning.
[0060] The medical record text processing method provided in this application retrieves the possible candidate codes corresponding to non-standard name entities in the medical record text to be processed from the coding database, and searches for these candidate codes in the standard name database to find and determine whether the non-standard name entity is used in the standard name database. Based on this, the statistical information of such a pair of alignment candidate pairs can be used to determine whether the non-standard name entity in the alignment candidate pair is frequently used with the standard name, thereby determining a qualified standard name for the non-standard name entity.
[0061] Example 3
[0062] Figure 3 A flowchart of another embodiment of the medical record text processing method provided in this application. (See flowchart for example.) Figure 3 As shown, the medical record text processing method provided in this embodiment may include the following steps:
[0063] S301, retrieve multiple original medical record texts.
[0064] In this embodiment, the medical record text to be processed comes from multiple original medical record texts. Various original medical record texts can be obtained from various data sources. These original medical record texts can be texts from a specific domain, or original medical record texts obtained from data sources based on given keywords. In this embodiment, the original medical record texts obtained in this way can include natural language text, that is, medical record texts written in natural language. For example, in the aforementioned medical management field, such natural language text can be obtained from data sources such as medical record databases storing medical records, medical paper databases, hospital management platforms, or even doctors' forums.
[0065] S302, Obtain the medical record text to be processed from the original medical record text.
[0066] After obtaining various original medical record texts in step S301, step S302 allows the acquisition of medical record texts written in various natural language formats from these sources. Furthermore, by performing, for example, structured parsing on these input medical record texts, the natural language texts are parsed into various entities containing standard and non-standard name entities.
[0067] In this embodiment, the parsed or preprocessed entities can be directly input into the standard name database for retrieval, thus excluding standard name entities that directly match standard terms. For example, some doctors are accustomed to using standard terms to describe illnesses in their actual writing; therefore, such natural language entities can directly match the standard names in the standard name database during actual use, eliminating the need for alignment annotation. Therefore, in this embodiment, standard name entities that already match standard terms can be removed from the entities in the acquired medical record text, thereby saving computational resources in subsequent matching and retrieval processing.
[0068] S303: Obtain the medical record text to be processed from the original medical record text and search the coding database to obtain candidate codes corresponding to non-standard name entities.
[0069] In this embodiment, the non-standard name entities obtained in step S302 can be input into the encoding library for searching, to find all candidate codes related to the non-standard name entities. For the convenience of subsequent processing, only a portion of the candidate codes can be selected as the codes used for subsequent standard name retrieval. For example, all candidate codes can be sorted, such as by the matching degree between the candidate codes and the corresponding non-standard name entities, and the top few candidate codes can be selected as the candidate codes used for standard name retrieval.
[0070] In this embodiment, since standard name entities that are completely identical to standard terminology have been excluded, the remaining non-standard name entities will encounter situations where they cannot be directly matched or will result in matching errors during actual matching. Therefore, in this embodiment, these non-standard name entities can be searched in the coding database to obtain relevant candidate codes. In this embodiment, these codes can all be related to the medical record text, and can be sorted according to relevance, with a predetermined number of codes ranked at the top selected as candidates, thus obtaining candidate codes.
[0071] Specifically, in this embodiment, the number of candidate codes can be determined by acquiring various attribute information related to the field to be processed. For example, in the medical field, different hospitals have varying levels of informatization, and even different doctors within the same hospital may have significantly different abilities and habits. Therefore, considering these differences, when searching for related codes for a non-standard name entity, the number of candidate codes can be adjusted based on such difference information. For example, step S303 in this embodiment may further include: calculating the degree of difference between the non-standard name entity and the standard name entity; and determining the number of search results based on the degree of difference.
[0072] Therefore, by calculating the difference between non-standard name entities and standard name entities after retrieving all candidate codes, or by directly using the relevance of the standard names corresponding to each code calculated during retrieval, a more suitable number of candidate codes can be selected.
[0073] For example, for hospitals with a low level of informatization, the number of candidate codes can be set to a larger number, such as 100 or 200 or even more. Alternatively, the difference between non-standard name entities and standard name entities can be calculated. When the difference exceeds a predetermined threshold, or when the difference with a certain number of standard entities exceeds a certain threshold, it can be determined that the doctor who wrote the entity does not frequently use standard terminology. Therefore, the number of candidate codes for that field can be set to a larger number, such as 100 or 200 or even more. Conversely, when the hospital has a high level of informatization, the number can be set to a smaller number, such as 10 or 20 or even lower. And when the difference between non-standard name entities and standard name entities is less than a predetermined threshold, it can be determined that the doctor who wrote the entity has a good habit of using standardized terminology. Therefore, the number of candidate codes for that entity can be set to a smaller number, such as 10 or 20 or even less.
[0074] S304 uses candidate codes to search the standard name database and obtains multiple search results.
[0075] After obtaining candidate codes in step S303, according to embodiments of this application, these candidate codes can be used to search in a standard name database to further determine which search result matches the non-standard name entity. For example, standard terms are standard names that can be used in the standard name database. For instance, in embodiments of this application, the standard name database can be a database that stores the correspondence between standard terms and natural language fields. Thus, for example, an industry platform that uses standard terms for data management in a hospital management system can use the standard name database to automatically match text fields written in natural language into the database.
[0076] S305, combine multiple search results with non-standard name entities to form multiple first alignment candidate pairs.
[0077] Therefore, in this embodiment of the application, in step S305, the various search results in step S304, namely the search results of standard names related to non-standard name entities, can be combined into a set of alignment candidate pairs.
[0078] S306, Obtain source information for multiple original medical record texts.
[0079] S307, Determine the first preset threshold based on the source information.
[0080] S308, determine the standard name of the non-standard name entity based on the statistical information of multiple first alignment candidate pairs in multiple original medical record texts.
[0081] After obtaining the alignment candidate pairs that may be used in the standard name database in step S305, the source information of the original medical record text obtained in step S301 can be obtained in step S306, and a first preset threshold can be determined based on the source information in step S307. This preset threshold can be used to filter the alignment candidate pairs. For example, if the original medical record texts all come from one hospital and the hospital has a low level of digitization, there are fewer original medical record texts available, i.e., fewer medical records. Therefore, in this case, the threshold for that hospital can also be set low to avoid not being able to obtain a standard name.
[0082] Furthermore, for step S308, the method for processing the medical record text in this application may further include:
[0083] S309, Calculate the number of times a non-standard name entity appears first in the medical record database composed of all medical record texts;
[0084] S310, Calculate the second occurrence count of the first alignment candidate pair in the medical record database;
[0085] S311, when the ratio of the second occurrence count to the first occurrence count is greater than the first preset threshold, the first alignment candidate pair is determined as the standard name of the non-standard name entity.
[0086] For example, in this embodiment of the application, the statistical information of the alignment candidate pair obtained in step S305 in the medical record database composed of all medical record texts can be used to determine whether the alignment candidate pair can be used as a qualified standard name for the standard name database or machine learning.
[0087] For example, in this embodiment, the number of times a non-standard name entity appears in the medical record database can be calculated first, and then the number of times the non-standard name entity and the standard search result that forms an alignment candidate pair with it appear together in the medical record database can be calculated. That is, the number of times the standard search result that can form an alignment candidate pair with the non-standard name entity, as determined in step S304, appears together with the corresponding non-standard name entity is counted.
[0088] For example, in the embodiments of this application, the medical record database may be a database consisting of all the data that can be applied to the medical record text processing scheme of this application, or it may be a database consisting of all the original medical record texts obtained in step S301, or it may be a database consisting of predetermined medical record texts specified by the operator or operator.
[0089] Therefore, if the non-standard name entity and the search results that form an alignment candidate pair with it appear frequently in step S310, it can be concluded that the non-standard name entity is likely frequently used to describe the search results. For example, in the solution of this application embodiment, the standard name "hypertension" is determined for the non-standard name entity "high blood pressure value". That is, "high blood pressure value" and "hypertension" form a set of alignment data candidates. If there are descriptions such as "the patient's blood pressure value is high, and he is likely a hypertensive patient" or "the patient's blood pressure value is high, and hypertension treatment can be performed first" in the predetermined medical record text, then in step S310, it can be determined that the set of alignment data candidates "high blood pressure value" and "hypertension" is a qualified set of alignment data.
[0090] The medical record text processing method provided in this application retrieves the possible candidate codes corresponding to non-standard name entities in the medical record text to be processed from the coding database, and searches for these candidate codes in the standard name database to find and determine whether the non-standard name entity is used in the standard name database. Based on this, the statistical information of such a pair of alignment candidate pairs can be used to determine whether the non-standard name entity in the alignment candidate pair is frequently used with the standard name, thereby determining a qualified standard name for the non-standard name entity.
[0091] Example 4
[0092] Figure 4 This is a schematic diagram of the structure of an embodiment of the medical record text processing apparatus provided in this application, which can be used to perform tasks such as... Figure 2 and Figure 3 The method steps are shown. (As shown) Figure 4 As shown, the medical record text processing device may include: a first acquisition module 41, a first retrieval module 42, a second retrieval module 43, an alignment candidate pair generation module 44, and a first determination module 45.
[0093] The first acquisition module 41 can be used to acquire the medical record text to be processed.
[0094] In this embodiment, the medical record text processing device can acquire text written in various natural language forms from various text sources. For example, in this embodiment, the medical record text processing device may include a second acquisition module 46 to acquire multiple original medical record texts from various data sources or platforms. In particular, the medical record text to be processed acquired by the second acquisition module 46 includes medical record text written in natural language. For example, in the aforementioned medical management field, such natural language medical record text can be acquired from data sources such as medical record databases storing medical records, medical paper databases, hospital management platforms, or even doctors' forums. Thus, the first acquisition module 41 can perform, for example, structured parsing on the input medical record text to be processed to parse the natural language text into various entities containing standard name entities and non-standard name entities. In this embodiment, the parsed and split entities or the pre-processed entities can be directly input into a standard name database for retrieval to exclude natural language entities that can directly match standard terms. For example, some doctors are already accustomed to using standard terminology to describe medical conditions when writing medical records. Therefore, such natural language entities can be directly matched with standard name entities in the standard name database during actual use, thus eliminating the need for alignment annotation of such entities. Therefore, in this embodiment of the application, these natural language entities that are already consistent with standard terminology can be removed from the acquired medical record text first, thereby saving the computational load of subsequent matching and retrieval processing.
[0095] The first retrieval module 42 can be used to perform retrieval in the coding database to obtain candidate codes corresponding to non-standard name entities.
[0096] In this embodiment, the first retrieval module 42 can input non-standard fields obtained by the first acquisition module 41 that differ from standard terms into the coding database for searching, to find all candidate codes related to non-standard name entities. For ease of subsequent processing, only a portion of the retrieved codes can be selected as candidate codes. For example, all retrieved codes can be sorted, such as by their matching degree with non-standard name entities, and the top few codes can be selected as corresponding candidate codes.
[0097] For example, the medical record text processing apparatus of this application embodiment may further include a first calculation module 47 and a second determination module 48. The first calculation module 47 may be used to calculate the degree of difference between non-standard name entities and standard name entities. And the second determination module 48 may be used to determine the number of candidate codes based on the degree of difference.
[0098] Therefore, after retrieving candidate codes, the first calculation module 47 calculates the difference between non-standard name entities and standard names, and the second determination module 48 determines the preset number based on the difference or directly uses the relevance to each standard term calculated during retrieval to determine the preset number. The first retrieval module 42 can select a more suitable number of code candidates.
[0099] For example, for hospitals with a low level of information technology integration, the number of candidate codes can be set to a larger number, such as 100 or 200 or even more. Alternatively, the difference between the non-standard name entity and each standard name can be calculated. If the difference exceeds a predetermined threshold, or if the difference with a certain number of standard names exceeds a certain threshold, it can be determined that the doctor who wrote the non-standard name entity does not frequently use standard terminology. Therefore, the number of candidate codes for this field can be set to a larger number, such as 100 or 200 or even more. Conversely, when the hospital has a high level of information technology integration, the number can be set to a smaller number, such as 10 or 20 or even lower. And if the difference between the non-standard name entity and each standard term is less than a predetermined threshold, it can be determined that the doctor to whom the entity belongs has a good habit of using standardized terminology. Therefore, the number of candidate codes for the non-standard name entity can be set to a smaller number, such as 10 or 20 or even less.
[0100] Specifically, in this embodiment, the number of candidate codes can be determined by obtaining various attribute information related to the non-standard name entity. For example, in the medical field, different hospitals have varying levels of informatization, and even different doctors within the same hospital may have significantly different abilities and habits. Therefore, considering these differences, when searching for relevant codes for a non-standard name entity, the number of candidate codes can be adjusted based on such difference information.
[0101] In this embodiment, since standard name entities that are completely identical to standard terminology have been excluded, the remaining non-standard name entities will encounter situations where they cannot be directly matched or will result in matching errors during actual matching. Therefore, in this embodiment, these non-standard name entities can be searched in an encoding database to obtain relevant codes. In this embodiment, these codes can all be related to non-standard name entities, and can be sorted according to relevance, with a predetermined number of codes ranked at the top selected as candidates, i.e., candidate codes are obtained.
[0102] In particular, in this embodiment of the application, the encoding database may be a standard field database that stores standard terms for a specific industry.
[0103] The second retrieval module 43 can be used to search in the standard name database using candidate codes to obtain multiple search results.
[0104] After the first retrieval module 42 obtains candidate codes, according to embodiments of this application, these candidate codes can be used to search in a standard name database to further determine which search result matches the non-standard name entity. For example, standard terms are standard names that can be used in the standard name database. For example, in embodiments of this application, the standard name database can be a database that stores the correspondence between standard terms and natural language fields. Thus, for example, an industry platform that uses standard terms for data management in a hospital management system can use the standard name database to automatically match text fields written in natural language into the database.
[0105] The alignment candidate pair generation module 44 can be used to form multiple first alignment candidate pairs by combining multiple search results with non-standard name entities.
[0106] Therefore, in this embodiment of the application, the alignment candidate pair generation module 44 can form a set of alignment candidate pairs with each retrieval result obtained by the second retrieval module 43 and the non-standard name entity.
[0107] The first determining module 45 can be used to determine the standard name of a non-standard name entity based on statistical information of multiple first alignment candidate pairs.
[0108] In this embodiment, the alignment candidate pair generation module 44 obtains alignment candidate pairs that may be used in the standard name database. Therefore, the first determining module 45 can determine whether the alignment candidate pair can be used as a qualified standard name for the standard name database or machine learning by calculating the statistical information of the alignment candidate pair.
[0109] The medical record text processing device of this application may further include: a third acquisition module 49 and a third determination module 491.
[0110] The third acquisition module 49 can be used to acquire source information of the original medical record text.
[0111] The third determining module 491 can be used to determine the first preset threshold based on the source information.
[0112] Therefore, after the alignment candidate pair generation module 44 obtains alignment candidate pairs that may be used in the standard name database, the third acquisition module 49 can further obtain the source information of the original medical record text obtained by the second acquisition module 46, and the third determination module 491 can determine a first preset threshold based on the source information. This preset threshold can be used by the first determination module 45 to select alignment candidate pairs. For example, if the original medical record texts all come from one hospital and the hospital has a low level of digitization, there are fewer original medical record texts available, i.e., fewer medical records. Therefore, in this case, the threshold for that hospital can also be set lower to avoid failing to obtain a standard name.
[0113] In addition, the first determining module 45 may include: a first calculation unit 451, a second calculation unit 452, and a first determining unit 453.
[0114] The first calculation unit 451 can be used to calculate the first occurrence count of a non-standard named entity in the full dataset;
[0115] The second calculation unit 452 can be used to calculate the second occurrence count of the first alignment candidate pair in the full data;
[0116] The first determining unit 453 can be used to determine the first alignment candidate pair as the standard name of a non-standard name entity when the ratio of the second occurrence count to the first occurrence count is greater than a first preset threshold.
[0117] For example, in this embodiment of the application, the first calculation unit 451 can calculate the statistical information of the alignment candidate pairs obtained by the alignment candidate pair generation module 44 in the medical record database composed of all medical record texts to determine whether the set of alignment candidate pairs can be used as qualified standard names for use by the standard name database or machine learning.
[0118] For example, in this embodiment, the number of times a non-standard name entity appears in the medical record database can be calculated first, and then the number of times the non-standard name entity and the standard search result that forms an alignment candidate pair with it appear together in the medical record database can be calculated. That is, the number of times the standard search result that can form an alignment candidate pair with the non-standard name entity, as determined by the alignment candidate pair generation module 44, appears together with the corresponding non-standard name entity is counted.
[0119] For example, in this embodiment of the application, the medical record database may be a database consisting of all the data that can be applied to the medical record text processing scheme of this application, or it may be a database consisting of all the original medical record texts acquired by the second acquisition module 46, or it may be a database consisting of predetermined medical record texts specified by the operator or operator.
[0120] Therefore, if the first determining unit 453 determines that the non-standard name entity and the search results that form an alignment candidate pair with it appear frequently, it can be concluded that the non-standard name entity is likely to be frequently used to describe the search results. For example, in the solution of this application embodiment, the standard name "hypertension" is determined for the non-standard name entity "high blood pressure value", that is, "high blood pressure value" and "hypertension" form a set of alignment data candidates. If there are descriptions such as "the patient's blood pressure value is high, and he is likely a hypertensive patient" or "the patient's blood pressure value is high, and hypertension treatment can be performed first" in the predetermined medical record text, then the first determining unit 453 can determine that the set of alignment data candidates "high blood pressure value" and "hypertension" is a qualified set of alignment data.
[0121] The medical record text processing apparatus provided in this application retrieves the possible candidate codes corresponding to non-standard name entities in the medical record text to be processed from the coding database, and searches for these candidate codes in the standard name database to determine whether the non-standard name entity is used in the standard name database. Based on this, the statistical information of such a pair of aligned candidate pairs can be used to determine whether the non-standard name entity in the aligned candidate pair is frequently used with the standard name, thereby determining a qualified standard name for the non-standard name entity.
[0122] Example 5
[0123] The above describes the internal functions and structure of the medical record text processing device, which can be implemented as an electronic device. Figure 5 A schematic diagram illustrating the structure of an embodiment of the electronic device provided in this application. (See attached diagram.) Figure 5 As shown, the electronic device includes a memory 51 and a processor 52.
[0124] Memory 51 is used to store programs. In addition to the programs described above, memory 51 can also be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.
[0125] The memory 51 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0126] Processor 51 is not limited to a central processing unit (CPU), but may also be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an embedded neural network processor (NPU), or an artificial intelligence (AI) chip. Processor 52 is coupled to memory 51 and executes the program stored in memory 51. When the program runs, it executes the medical record text processing method of embodiment two or three described above.
[0127] Furthermore, such as Figure 5 As shown, the electronic device may also include other components such as a communication component 53, a power supply component 54, an audio component 55, and a display 56. Figure 5 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 5 The components shown.
[0128] Communication component 53 is configured to facilitate wired or wireless communication between electronic devices and other devices. The electronic devices can access wireless networks based on communication standards, such as WiFi, 3G, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 53 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 53 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0129] Power supply component 54 provides power to various components of the electronic device. Power supply component 54 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.
[0130] Audio component 55 is configured to output and / or input audio signals. For example, audio component 55 includes a microphone (MIC) configured to receive external audio signals when the electronic device is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 51 or transmitted via communication component 53. In some embodiments, audio component 55 also includes a speaker for outputting audio signals.
[0131] Display 56 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation.
[0132] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing medical record texts, comprising: Obtain the medical record text to be processed, wherein the medical record text to be processed includes multiple entities, including standard name entities and non-standard name entities; A search is performed in the encoding database to obtain candidate codes corresponding to the non-standard name entity; The candidate codes are used to search the standard name database to obtain multiple search results; wherein the number of search results is determined according to the degree of difference between the non-standard name entity and the standard name entity. The multiple search results are respectively paired with the non-standard name entity to form multiple first alignment candidate pairs; The standard name of the non-standard name entity is determined based on the statistical information of the plurality of first alignment candidate pairs; Determining the standard name of the non-standard name entity based on the statistical information of the plurality of first alignment candidate pairs includes: Calculate the number of times the non-standard name entity appears for the first time in the medical record database composed of all medical record texts; Calculate the second occurrence count of the first alignment candidate pair in the medical record database; When the ratio of the second occurrence count to the first occurrence count is greater than the first preset threshold, the retrieval result in the first alignment candidate pair is determined as the standard name of the non-standard name entity; The method further includes: Obtain source information for multiple original medical record texts; The first preset threshold is determined based on the source information.
2. The method for processing medical record text according to claim 1, wherein, The method further includes: acquiring multiple original medical record texts, wherein the medical record text to be processed comes from the multiple original medical record texts; The step of determining the standard name of the non-standard name entity based on the statistical information of the plurality of first alignment candidate pairs includes: The standard name of the non-standard name entity is determined based on the statistical information of the plurality of first alignment candidate pairs in the plurality of original medical record texts.
3. A medical record text processing device, comprising: The first acquisition module is used to acquire the medical record text to be processed, wherein the medical record text to be processed includes multiple entities, including standard name entities and non-standard name entities; The first retrieval module is used to perform a retrieval in the encoding database to obtain candidate codes corresponding to the non-standard name entity; The second retrieval module is used to retrieve multiple retrieval results in the standard name database using the candidate codes; the medical record text processing device further includes: a first calculation module, used to calculate the difference degree between the non-standard name entity and the standard name entity; and a second determination module, used to determine the number of retrieval results based on the difference degree. The alignment candidate pair generation module is used to form multiple first alignment candidate pairs by combining multiple search results with the non-standard name entities; The first determining module is used to determine the standard name of the non-standard name entity based on the statistical information of the plurality of first alignment candidate pairs; The first determining module includes: The first calculation unit is used to calculate the first occurrence count of the non-standard name entity in the medical record database composed of all medical record texts; The second calculation unit is used to calculate the second occurrence number of the first alignment candidate pair in the medical record database; The first determining unit is used to determine the retrieval result in the first alignment candidate pair as the standard name of the non-standard name entity when the ratio of the second occurrence count to the first occurrence count is greater than a first preset threshold. The medical record text processing device further includes: The third acquisition module is used to acquire the source information of the original medical record text; The third determining module is used to determine the first preset threshold based on the source information.
4. An electronic device, comprising: Memory, used to store programs; A processor is configured to run the program stored in the memory, wherein the program, when running, executes the medical record text processing method as described in any one of claims 1 to 2.
5. A computer-readable storage medium having a computer program stored thereon that can be executed by a processor, wherein, When executed by a processor, the program implements the medical record text processing method as described in any one of claims 1 to 2.