Automated Information Extraction and Improvement within Pathology Reports Using Natural Language Processing

An automated workflow using OCR and NLP transforms unstructured pathology reports into structured medical data, addressing the inefficiencies of manual extraction and enhancing healthcare delivery through accurate and timely data analysis.

JP7714617B2Active Publication Date: 2025-07-29F HOFFMANN LA ROCHE & CO AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023197189
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-06
Filing Date
2023-11-21
Publication Date
2025-07-29
Estimated Expiration
2040-09-08

AI Technical Summary

Technical Problem

Clinical data, particularly pathology reports, is often unstructured and difficult to access and analyze, leading to time-consuming, laborious, and error-prone manual extraction processes that burden healthcare workers and hinder efficient healthcare delivery.

Method used

An automated workflow using optical character recognition (OCR) and natural language processing (NLP) to extract and improve information from pathology reports, converting unstructured data into structured medical data aligned with standards like SNOMED, enabling efficient retrieval and analysis.

Benefits of technology

Substantially expedites the extraction process, reduces errors, and facilitates large-scale analysis of pathology reports, providing timely and accurate insights for clinical decision-making and improving healthcare quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007714617000003
    Figure 0007714617000003
  • Figure 0007714617000004
    Figure 0007714617000004
  • Figure 0007714617000005
    Figure 0007714617000005
Patent Text Reader

Abstract

To provide a technique for automated information extraction and enrichment in a pathology report.SOLUTION: In one example, a method being performed by a computer system comprises: receiving an image file containing a pathology report; performing an image recognition operation on the image file to extract input text strings; detecting, using a natural language processing (NLP) model, entities from the input text strings, each entity including a label and a value; extracting, using the NLP model, the values of the entities from the input text strings; converting, based on a mapping table that maps entities and values to pre-determined terminologies, the values of at least some of the entities to the corresponding predetermined terminologies; and generating a post-processed pathology report including the entities detected from the input text strings and the corresponding predetermined terminologies.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 897,252, filed on September 6, 2019, the entire content of which is incorporated herein by reference for all purposes.

Background Art

[0002]

[0002] Every day, hospitals around the world generate a vast amount of clinical data. Healthcare workers, such as clinicians and clinical staff, need to analyze the clinical data in order to care for patients. The analysis of this data is also important in providing detailed insights into healthcare delivery and the quality of care, and in providing a basis for improving healthcare.

[0003]

[0003] Unfortunately, most clinical data is difficult to access and analyze because most of the data is either in paper form or in the form of scanned images. The data may include, for example, pathology reports or any other data that is not associated with a structured data model and is not organized in a predefined manner to define the context and / or meaning of the data. Due to the physical form of the data and the fact that the data is unstructured, clinicians and clinical staff typically have to spend a great deal of time reading through a patient's pathology report to obtain important clinical data such as diagnostic history, treatment history, etc., and the time accumulates when reading the pathology reports of a large number of patients. Moreover, manual extraction is also prone to being broken, slow, costly, and error - prone. Manually processing and extracting clinical data from pathology reports places a heavy burden on healthcare workers and may affect the ability of healthcare workers to care for patients. The large - scale manual processing of pathology reports to provide detailed insights into healthcare delivery and the quality of care is also not feasible due to cost and time constraints.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Provide techniques for the automated extraction and improvement of information within a pathology report.

Means for Solving the Problem

[0005]

[0004] Techniques for the automated extraction and improvement of information within a pathology report are disclosed herein. The pathology report can include electronic reports from various primary information sources (e.g., in one or more medical facilities), such as an EMR (Electronic Medical Record) database, a PACS (Picture Archiving and Communication System), a Digital Pathology (DP) system, a LIS (Laboratory Information System) including genomic data, a RIS (Radiology Information System), a patient-reported outcome database, wearable and / or digital technologies, and social media. The pathology report can also be in paper form and can be derived from clinicians / clinical staff. The pathology report can be in the form of an image file (e.g., Portable Document Format (pdf), Bitmap Image File (BMP file)) obtained by scanning a paper-form pathology report.

[0006]

[0005] In some examples, a workflow is provided to extract pathological entities from images of a pathology report. The workflow can begin by extracting text strings from an image file of the pathology report. The extraction of text strings from the image file can be based on an image recognition process that recognizes characters and / or text strings from the image, such as optical character recognition (OCR), optical word recognition, etc. The workflow can further include using a natural language processor (NLP) to recognize entities from the input text strings, where each entity includes a label and a value, and identifying the value of the entity from the text strings. Entities can generally refer to predefined medical categories and classifications, such as medical diagnoses, medical treatments, medications, specific locations / organs in a patient's body, etc. Each entity can have a label indicating the category / classification and a value corresponding to the categorized / classified data. In some examples, the workflow further includes mapping the values of at least some of the entities to standard terms, such as clinical terms and codes defined based on the International Classification of Diseases for Oncology (ICD) standard. The workflow can then generate structured medical data that associates the label of the entity with at least one of the value of the entity or the standardized term based on the mapping.

[0007] It should be noted that there is an inaccuracy in the original text where it mentions "International Classification of Diseases for Oncology (ICD)" which should likely be "Systematized Nomenclature of Medicine (SNOMED)". The translation has been corrected accordingly.

[0006] Structured medical data can be provided for various applications. For example, structured medical data can be stored in a searchable database, from which entities and their values (whether standardized or not) can be retrieved based on a search query. The searchable database as well as the structured medical data can also be made available for various applications such as clinical decision support applications, analysis applications, etc. for processing. For example, a clinical decision support application can retrieve entities related to clinical decisions (e.g., diagnosis history, treatment history, medication history) and their values from a database to support clinical decisions, and process the entities to generate an output. An analysis application can, for example, retrieve entities related to treatment history and diagnosis from the pathology reports of a number of patients and perform an analysis to gain insights into the quality of medical care and nursing. In other examples, a clinical portal application can be provided to display structured medical data and / or an image of a pathology report in which the extracted entity information is overlaid on an image.

[0008]

[0007] An NLP model can be trained to identify sequences of text strings containing entities and values and extract the entities and values based on the identification. NLP can be trained in a two-step process. As a first step, the NLP model can be trained based on documents containing common medical terms to build a baseline NLP submodel. As a second step, the baseline NLP submodel can then be trained using text strings from pathology reports to extend the model to include specific pathology terms. The second step of the training operation can be performed using a CoNLL (Conference on Natural Language Learning) file.

[0009]

[0008] In addition, various techniques can determine various parameters of an image recognition operation to improve the extraction accuracy of NLP. In some examples, a parameter sweep operation can be performed to obtain different combinations of parameter values. Then, the image recognition operation can be repeatedly executed, with each repetition being executed based on a combination of parameter values. Then, the text recognition accuracy for each repetition can be measured, and a specific combination of parameter values leading to the highest text recognition accuracy can be used to configure the image recognition operation for the workflow. As another example, the determination of the parameters of the image recognition operation can be based on the output of NLP. Specifically, the image recognition operation can be preconfigured based on a first set of parameter values. The preconfigured image recognition operation can be executed on the image of the pathology report to extract a text string, which can be input to NLP to extract pathology entities. Then, the parameters of the image recognition operation can be adjusted based on the extraction accuracy by NLP.

[0010]

[0009] The above and other embodiments of the present invention are described in detail below. For example, other embodiments relate to systems, devices, and computer-readable media associated with the methods described herein.

[0011]

[0010] A better understanding of the essence and advantages of the embodiments of the present invention may be obtained by referring to the following detailed description of the invention and the accompanying drawings.

[0011] The detailed description of the invention is described with reference to the accompanying drawings.

Brief Description of the Drawings

[0012]

Figure 1

[0012] FIG. shows an example of a conventional pathology report.

Figure 2A

[0013] FIG. shows an example of post-processing of a conventional pathology report that can be implemented by an example of the present disclosure.

Figure 2B

Figure 3

[0014] The figure shows an example of a system that executes automated information extraction and improvement of a pathological report.

Figure 4A

[0015] The figure shows exemplary internal components of the system of FIG. 3 and their operations.

Figure 4B

Figure 4C

Figure 4D

Figure 4E

Figure 5A

[0016] The figure shows an example of a training operation of a natural language processing model of the system of FIG. 3.

Figure 5B

Figure 5C

Figure 5D

Figure 5E

Figure 6

[0017] The figure shows an exemplary operation of determining parameters of an image recognition operation within the system of FIG. 3.

Figure 7

[0018] The figure shows an exemplary application supported by the output of the system of FIG. 3.

Figure 8

[0019] The figure shows a method of executing automated information extraction and improvement of a pathological report.

Figure 9

[0020] A diagram showing an exemplary computer system that can be used to implement the techniques disclosed in this specification.

Mode for Carrying Out the Invention

[0013]

[0021] Techniques for the extraction and improvement of automated information within a pathology report are disclosed herein. The pathology report can be derived from electronic reports from various primary information sources (e.g., in one or more medical facilities), including, for example, an EMR (Electronic Medical Record) database, a PACS (Picture Archiving and Communication System), a Digital Pathology (DP) system, a LIS (Laboratory Information System) including genomic data, a RIS (Radiation Information System), a patient-reported outcome database, wearable and / or digital technology, and social media. The pathology report can also be in paper form and can be derived from a clinician / clinical staff. The pathology report can be in the form of an image file (e.g., Portable Document Format (pdf), Bitmap Image File (BMP file)) obtained by scanning a paper-form pathology report.

[0014]

[0022] In some embodiments, a workflow is provided for extracting pathological entities from images of pathological reports. The workflow can begin by extracting text strings from the image files of the pathological reports. The extraction of text strings from the image files can be based on an image recognition process for recognizing characters and / or text strings from images, such as optical character recognition (OCR), optical word recognition, etc. The workflow is to use a natural language processor (NLP) to recognize entities from the text strings, including recognizing each entity having a label and a value, and further identifying the value of the entity from the text strings. Entities generally refer to predefined medical categories and classifications, such as medical diagnoses, medical treatments, medications, specific locations / organs within a patient's body, etc. Each entity has a label indicating the category / classification and a value indicating the categorized / classified data. In some examples, the workflow includes mapping the values of at least some of the entities to standard terms. The mapping can be part of an improvement process, in which the values of at least some of the entities, which can be non-standardized representations of the categorized / classified data, are converted to standardized data, such as clinical terms and codes defined based on the International Classification of Diseases (ICD) standard. The workflow can then generate structured medical data associating the labels of the entities with at least one of the values of the entities or the standardized terms.

[0015]

[0023] Structured medical data can be provided to various applications. For example, the structured medical data can be stored in a searchable database, from which entities (standardized or not) and their values can be retrieved based on a search query. The searchable database as well as the structured medical data can also be made available to various applications such as clinical decision support applications, analysis applications, etc. for processing. For example, a clinical decision support application can retrieve entities related to clinical decisions (e.g., diagnosis history, treatment history, medication history) and their values from a database to support clinical decisions, and process the entities to generate an output. An analysis application can, for example, retrieve entities related to treatment history and diagnosis from the pathology reports of a number of patients, and perform an analysis to obtain insights into the quality of medical care and nursing.

[0016]

[0024] As another example, a clinical portal application may be provided that performs an end-to-end improved workflow operation. The clinical portal application can receive an image of a pathology report from a patient database and perform an optical character recognition (OCR) operation on the image to generate first data including the extracted text strings and their image positions within the image. The clinical portal application can then use NLP to extract pathology entities (including labels and values) from the extracted text strings. The clinical portal application then aggregates the entities into structured medical data and stores the structured medical data back in the patient database. The clinical portal application can also display the structured medical data. In some examples, the clinical portal application can display the structured medical data in a structured format (e.g., in the form of a table, input form) to enable the user of the portal (e.g., clinician, clinical staff) to efficiently identify the medical information they are looking for. In some examples, the clinical portal application can include a display interface for displaying the image and selectable highlighting markings superimposed on the text strings that the NLP has determined represent the pathology entities. The display interface can also detect a selection of the highlighting markings on a set of text strings and display a pop-up window including the entity labels and values, as well as other refined information for the selected text strings (e.g., standardized data based on SNOMED).

[0017]

[0025] An NLP model can be trained to identify a sequence of text strings that include entities and values and extract the entities and values based on the identification. NLP can be trained in a two-step process. As a first step, the NLP model can be trained based on documents that include common medical terms to build a baseline NLP sub-model. The baseline NLP sub-model can be used to provide a primary context for identifying a sequence of text strings that include common medical terms, which may or may not include pathological entities. The baseline NLP sub-model can be trained / built based on biomedical articles from various major information sources, such as, for example, PubMed Central®, a free full-text archive of biomedical and life science journal literature at the U.S. National Library of Medicine of the U.S. National Institutes of Health. As a second step, the baseline NLP sub-model is then trained using text strings from pathology reports to extend the sub-model to include pathological entities. The second step of the training operation can be performed using CoNLL (Conference on Natural Language Learning) files. The CoNLL files may include text strings extracted from other pathology reports, and each text is tagged with either an entity label or a label indicating that it is a non-entity. NLP can be trained based on CoNLL files from multiple pathology reports. In some examples, the training can be specific to a hospital, clinical group, or individual clinician, and as a result, NLP can be trained to learn the word preferences of the hospital / clinical group / clinician that can maximize the extraction accuracy of entities and their values. In some embodiments, statistical data on the extraction accuracy of entities can be retained. If the statistical data indicates that the NLP has a low extraction accuracy when extracting entities from the input text string, the input text string can be tagged to generate a new CoNLL file, and the NLP can be retrained using the new CoNLL file to improve the extraction accuracy.

[0018]

[0026] In addition, various techniques are proposed to determine various parameters of the image recognition operation so as to improve the extraction accuracy of NLP. The parameters may include, for example, an erosion value, a page iterator level, a page segmentation mode, or a magnification factor. The erosion value can indicate whether a smoothing operation of a blurred line has been performed. The page iterator level can indicate how the image recognition operation has been performed by treating the entire page as a block or by treating sections (such as paragraphs, lines, words, characters, etc.) within the page as blocks in order to increase the granularity of the image recognition operation. The page segmentation mode can detect the inclined orientation of the page being processed and adjust the image recognition operation to correct the inclined orientation. The magnification factor can set the zoom level to zoom in or out the image to be processed.

[0019]

[0027] In some examples, a parameter sweep operation can be performed to obtain different combinations of parameter values. Then, the image recognition operation can be repeatedly performed on a set of pathology reports, and each repetition is performed based on a combination of parameter values. Then, the text recognition accuracy for each repetition can be measured, and a specific combination of parameter values leading to the highest text recognition accuracy can be used to configure the image recognition operation for the workflow.

[0020]

[0028] As another example, the determination of the parameters of the image recognition operation can be based on the output of NLP. Specifically, the image recognition operation can be preconfigured based on a first set of parameter values. The preconfigured image recognition operation can be performed on the image of the pathology report to extract a text string, and the text string can be input into NLP to extract pathology entities. Then, the parameters of the image recognition operation can be adjusted based on the extraction accuracy by NLP. The text string can be input into NLP to extract pathology entities. Then, the parameters of the image recognition operation can be adjusted based on the extraction accuracy by NLP.

[0021]

[0029] Adjusting the parameters of the image recognition operation based on the output of NLP can be advantageous, particularly when the image file includes annotations by a particular physician that may include non-standard codes and phrases. When the output of OCR is compared to a standardized phrase to determine text recognition accuracy, the comparison can lead to incorrect conclusions regarding the text recognition accuracy for a particular set of OCR parameters when the text string includes non-standard codes and phrases. On the other hand, since the NLP model is trained to recognize non-standard codes and phrases, as well as standardized terms, using the output of NLP to determine text recognition accuracy can ensure that the text recognition accuracy measurements are not overly affected by the presence of non-standard codes and phrases within the output of OCR.

[0022]

[0030] The disclosed technique can enable an automated workflow that begins by processing an image of a pathology report to extract a text string, uses NLP to extract entities and their values from the text string, improves them by mapping the extracted entities and values to standard terms, and subsequently generates structured medical data that includes at least one of the extracted entities and the extracted values or standard terms. Compared to the case where clinicians and clinical staff need to manually read through pathology reports to extract relevant information, the disclosed technique can substantially expedite the extraction process and reduce the time / resources required for clinicians and clinical staff to obtain the necessary information from pathology reports, thereby enabling clinicians and clinical staff to allocate more time / resources to finding the correct treatment and treating patients. Moreover, by making the structured medical data accessible by other applications such as clinical support applications, analysis applications, etc., large-scale analysis of pathology reports of large patient populations can be performed to provide insights into the quality of medical care and nursing, to provide relevant data to support clinical judgments made by clinicians, etc. Improvements in the overall speed of the data flow and in the accuracy and completeness of medical data extraction can provide broader and faster access to high-quality patient data for clinical and research purposes, which can facilitate developments in treatment and medical technology and improvements in the quality of care provided to patients.

[0023] I. Examples of Information Extraction and Improvement from Pathology Reports

[0031] FIG. 1 shows an example of a conventional pathology report 100. A pathology report is a medical document written by a pathologist and can provide a histological diagnosis based on the pathologist's examination of a tissue sample taken from a patient's tumor. From the tumor tissue, the pathologist can find, for example, whether the tissue is cancerous or non-cancerous and other specific details regarding the characteristics of the tumor. All of this information can be part of the pathology report. Based on this information, a treatment can be formulated.

[0024]

[0032] Referring to FIG. 1, the pathology report 100 may include multiple sections of diagnostic information. For example, the pathology report 100 may include, among others, a section 102 indicating the location of the tumor (e.g., right lung / middle lobe), a section 104 indicating the number of lesions (e.g., squamous cell carcinoma of the lung), a section 106 indicating the size of the tumor (e.g., 5.3×4.0×3.0 cm), a section 108 indicating the histological diagnosis (e.g., well-differentiated or moderately-differentiated keratinizing squamous cell carcinoma), a section 110 indicating the lymph node status (e.g., N2(8 / 28)), and a section 112 indicating the TNM (tumor lymph node metastasis) stage (e.g., pT3 (pericardial cavity invasion) N2(8 / 28) G2 R0). The pathology report 100 can be in paper form or stored as an image file (e.g., a pdf file, a BMP file) generated by scanning the pages containing the pathology report 100.

[0025]

[0033] A clinician and / or clinical staff member can read through the pathology report 100 and manually extract the medical information being sought. However, such an arrangement is laborious, slow, costly, and error-prone. Specifically, pathology reports may not be organized in a uniform format and structure, especially reports created from different hospitals and groups. As a result, the reader may need to read through the entire pathology report 100 to search for a particular piece of medical information, which can be very time-consuming and laborious, especially when the reader has to read through a large number of pathology reports for a large patient population.

[0026]

[0034] Manual extraction processes can also be error - prone. Readers may have very limited time to read through a pathology report to find the information they need, and they may make mistakes when reading and / or transcribing the information obtained from the pathology report. So, one cause of errors can be the cumbersome extraction process. Another cause of errors can be the fact that different clinicians may have different ways of documenting diagnostic results, which can lead to confusion and incorrect interpretations. For example, in section 110, readers may have difficulty understanding the meaning of "lymph node status" and the associated value "N2 8 / 28". As a result, readers may have an incorrect interpretation of section 110. Another cause of errors can be the mapping of important entities to standard terms. Default standard terms may have a lot of redundancy, and just looking them up may not help in converting the extracted entities into normalized terms. For example, the word "lung" may be associated with more than 20 normalized concepts. Identifying the concept to which the word "lung" maps can be difficult to do manually.

[0027]

[0035] Figures 2A and 2B show exemplary results of post - processing a pathology report 100 that can be implemented by the techniques of the present disclosure. As shown in Figure 2A, the diagnostic information within sections 102 - 112 of the pathology report 100 can be mapped to various medical entities. Medical entities can refer to pre - defined medical categories and classifications. Medical entities can include, for example, medical diagnoses, medical treatments, medications, and specific locations / organs within a patient's body. Medical entities can be defined based on a universal standard such as SNOMED, so that any clinician and medical provider can attach the same meaning to those medical entities. A list of typical medical entities in a pathology report and their meanings can be as follows.

[0028]

Table 1 - 1

[0029]

Table 1-2

[0030]

[0036] Referring to FIG. 2A, the diagnostic information within sections 102-112 of the pathology report 100 can be mapped to various medical entities in Table 1 in order to generate a digital pathology report 200 that includes structured data organized based on medical entities. For example, the information in section 102 can be split and mapped to both the entity "specimen laterality" (having the value "right") and the entity "tumor location" (having the value "middle lobe"). The information in section 104 can be mapped to the entity "tissue structure" having the value "squamous cell carcinoma". The information in section 106 can be mapped to the entity "tumor size" having the value "5.3×4.0×3.0 cm". The information in section 108 can be mapped to the entity "histological malignancy" having the value "well-differentiated or moderately differentiated keratinizing squamous cell carcinoma". The information in section 110 can be mapped to the entity "regional lymph nodes / category (pN)" having the value N2, and the information in section 112 can be split and mapped to the entity "primary tumor (pT)" (having the value pT3) and the entity "overall malignancy" (having the value G2). Since each medical entity in the digital pathology report 200 is defined based on a worldwide common standard and has a clearly defined meaning, the risk that a reader misinterprets the meaning of a medical entity and its associated value can be reduced.

[0031]

[0037] In some examples, the digital pathology report 200 can be a plain text file in which entities and associated values are stored in the form of text strings and can be easily parsed / queried by other applications. Moreover, the placement of entities and their associated values within the digital pathology report 200 can follow a structured and standardized order, such that each entity has its own predetermined location within the digital pathology report 200. In such an arrangement, instead of searching through the entire pathology report to find an entity, an application (or a human reader familiar with the standardized order) can search for a particular entity and its value within the pathology report 200 based on the entity's predetermined location, which can substantially accelerate the extraction of medical information from the digital pathology report 200.

[0032]

[0038] As part of the improvement process, the combinations of entities and values in the digital pathology report 200 can be mapped to predetermined medical terms defined based on a common global standard such as SNOMED. Such an arrangement enables the diagnostic results represented by the combinations of entities and values to conform to a common global standard, thereby further reducing the risk of misinterpretation and ambiguity. For example, returning to FIG. 2A, section 210 indicates that the histological tumor site has the value "mid-lobe", but the organ is not specified, which can introduce uncertainty and potential confusion regarding the exact location of the tumor site. However, if section 210 is converted into a standardized and globally accepted format, the uncertainty / confusion regarding the exact location of the tumor site can be avoided.

[0033]

[0039] Figure 2B shows a mapping table 250 that illustrates an example of the mapping between entity-value pairs and SNOMED concepts, which can remove the risk of misinterpretation and ambiguity. For example, the entity "tissue structure" with the value "squamous cell carcinoma" can be mapped to the SNOMED concept "intraepithelial carcinoma" with concept ID 59529006. Additionally, the entity "tumor site" with the value "lower lobe" can be mapped to the SNOMED concept "structure of the lower lobe of the lung" with concept ID 90572001. Such mappings can be based on the pairing between the entity "tumor site" and the value "lower lobe", as well as the information included in section 102 that is extracted as context information such as the text "lung" which is not part of the entity. Similarly, the entity "specimen laterality" with the value "left" can be mapped to the SNOMED concept "left lung structure" with concept ID 44029006 and can also be based on the entity-value pairing and context information. In all these cases, the SNOMED concept can clarify the exact location of the tumor site to remove potential confusion / ambiguity.

[0034]

[0040] As part of the improvement process, each entity-value pair of the digital pathology report 200 that maps (matches) to a SNOMED concept can be replaced with the SNOMED concept. For example, the entity-value pair (tumor site - lower lobe) within section 210 can be replaced with the SNOMED concept "structure of the lower lobe of the lung" and / or SNOMED concept ID 90572001. On the other hand, entity-value pairs within the digital pathology report 200 that do not have a corresponding SNOMED concept are not replaced. If there is no match, the report can include the entity-value pair. NLP can be trained to provide the corresponding SNOMED concept when applicable.

[0035]

[0041] The replacement of entity-value pairs with their SNOMED concepts can improve digital pathology report 200 by including standard terms in the report, which can reduce the risk of misinterpretation and ambiguity associated with non-standard values of entities for human readers. In some examples, the entity-value pairs of digital pathology report 200 can also be replaced with SNOMED concept IDs to reduce the data size of digital pathology report 200. Such arrangements can also facilitate the processing of digital pathology report 200 by applications. Specifically, since entity-value pairs may have multiple alternative versions of values representing the same concept, an application that extracts and interprets entity-value pairs needs to have a built-in function to recognize multiple alternative versions of values and recognize the associated concepts. On the other hand, an application can syntactically interpret SNOMED concept IDs and uniquely link concepts to concept IDs, which can reduce the complexity of the application.

[0036] II. Pathological Entity Extraction and Improvement System

[0042] As described above, conventional pathology reports such as pathology report 100 are difficult to access and analyze data in either paper form or the form of scanned images. Due to the physical form of the data and the fact that the data is not structured, clinicians and clinical staff usually have to spend a great deal of time reading through the pathology report to obtain important clinical data, which is time-consuming, slow, costly, and error-prone. Moreover, since the clinical data in the report may include non-standardized terms, potential ambiguity and confusion may occur when clinicians interpret the non-standardized terms in the report, which may lead to errors in the extraction of clinical data from the pathology report.

[0037] A. System Architecture

[0043] Figure 3 shows a system 300 that can perform automated information extraction and refinement of a pathology report to address at least some of the problems described above. System 300 can be part of a clinical portal application that implements an end-to-end refinement workflow operation. Referring to Figure 3, system 300 can receive, as input, a pathology report image file 302 (e.g., of pathology report 100) from a patient database 301. System 300 can generate, as output, post-processed pathology report data 304 (e.g., of pathology report 200). As described below, the post-processed pathology report data 304 can include information extracted from the pathology report image file 302, including pathology entities such as those described in Figure 2A and Table 1 above, and associated values identified from the pathology report image file 302. Additionally, the post-processed pathology report data 304 may also include refinement information such as standardized pathology entity values (e.g., SNOMED concepts). The post-processed pathology report data 304 can be written back to the patient database 301 (or another clinical database) as structured medical data for the patient. In some examples, system 300 also includes a display interface 305 for displaying the post-processed pathology report data 304 in a structured format (e.g., in the form of a table, input form). In some examples, the display interface 305 can also display the pathology report image file 302 overlaid with text and graphical information based on the post-processed pathology report data 304.

[0038]

[0044] System 300 may include an optical processing module 306, an entity extraction module 308, and an improvement module 310 to perform information extraction and improvement. Each module can include software instructions that can be executed on a computer system (e.g., within a server or a cloud computing environment including multiple servers). In some examples, system 300 can be part of a clinical software platform (not shown in FIG. 3). Each module of system 300 can include an application programming interface (API) to communicate with the software platform and access different databases such as patient database 301.

[0039]

[0045] Referring to FIG. 3, the optical processing module 306 can receive an image file 302. The image file 302 can be received from various primary information sources (e.g., in one or more medical facilities) including, for example, an EMR (electronic medical record) database, a PACS (picture archiving and communication system), a digital pathology (DP) system, a LIS (laboratory information system) including genomic data, a RIS (radiology information system), a patient-reported outcome database, wearable and / or digital technologies, and social media. The image file can be in various formats such as, for example, a portable document file (pdf) or a bitmap image file (BMP file). In some examples, the image file can be obtained by scanning a paper-based pathology report.

[0040]

[0046] After receiving the image file 302, the optical processing module 306 can perform an image recognition operation to identify a text image from the image file 302, generate text data from the text image, and generate an intermediate text file 312 containing the text data. The image recognition operation may include, for example, optical character recognition (OCR) or optical word recognition. In both operations, the optical processing module 306 can extract a pixel pattern of a character (by, for example, identifying a pattern of pixels having a dark color), compare each pixel pattern with a predefined pixel pattern of a character, and determine based on the comparison which character (or which word / phrase) each pixel pattern represents. Next, the optical processing module 306 can store the character / word / phrase in the text file 312. The optical processing module 306 can scan through the image file 312 according to a predetermined pattern (e.g., raster scan) to extract and process the pixel patterns of the rows from left to right, and can repeat the scan for each row. Based on the scan pattern, the optical processing module 306 can generate a sequence of text strings (e.g., characters, words, phrases) and store the sequence of text strings in the text file 312. In some examples, a metadata file 314 indicating the pixel positions of each sequence of text strings can also be generated by the optical processing module 306. The metadata file 314 can be used by other applications as described below. An example of the metadata file 314 is shown in FIG. 4D.

[0041]

[0047] The entity extraction module 308 can process the text file 312, recognize entities (e.g., the entities listed in Table 1) from the text file 312, and extract the values associated with the entities. The entity extraction module 308 can generate entity-value pairs 320, where each pair includes the extracted entity and the corresponding value. The entity extraction module 308 may include a natural language processing (NLP) model 328 to perform entity recognition and value extraction. The NLP model 328 processes a sequence of text from the text file 312 and, based on recognizing a specific sequence of text strings, determines that a subset of the sequence of text is the value of a specific entity and can identify the entity-value pair for the subset.

[0042] B. Natural Language Processor Model

[0048] FIG. 4A shows an example of the NLP model 328. As shown in FIG. 4A, the NLP model 328 includes a graph having nodes such as nodes 402, 404a, 404b, 406a, 406b, 406c, and 408. Each node can correspond to a text string. In the graph, the nodes are connected by arcs, and the direction of the arcs defines the sequence of text strings to be detected by the NLP model 328. For example, nodes 402 and 404a are connected by arc 410, and nodes 404a and 406b are Node 406b is connected by arc 412 and nodes 406b and 408 are connected by arc 414. These nodes and arcs can define the text sequence "right lung middle lobe". The nodes can also be organized hierarchically, and detection outputs that can be entity-value pairs, contexts, etc. can be generated from each hierarchy. In the example of FIG. 4A, node 402 can be in the first hierarchy that detects the entity "specimen laterality", nodes 404a and 404b can be in the second hierarchy that detects the context, and nodes 406a to 406c and 408 can be in the third hierarchy that detects the entity "tumor site". The detection can be based on a parameterized formula that calculates a score based on, for example, the similarity between the input sequence of text strings and the text strings represented by the nodes, and predetermined entity-pair and / or context information can be output based on the score.

[0043]

[0049] The NLP model 328 can process sequences of text strings such as sequence 420 from text file 312. The NLP model 328 can search for a sequence of nodes from the graph that matches sequence 420 (either exactly or up to a proximity threshold), while skipping text strings (such as words, punctuation, symbols) not found in the graph. In some examples, the text strings of the nodes can be represented by vectors, and the proximity can be defined by a threshold of the total Euclidean distance between the text strings in the sequence of nodes and the text strings in sequence 420. In some examples, the proximity can also be defined by a threshold number of matching words between the sequence of nodes and sequence 420. In the example of FIG. 4A, the NLP model 328 can process sequence 420 "Location: Right Lung / Middle Lobe" by searching for the sequence of nodes from the graph that is closest to sequence 420, and can identify the sequence of nodes 402, 404a, 406b, and 408 that are closest to sequence 420 while ignoring the word "Location" and the punctuation ": " and " / ". From the identified sequence, the NLP model 328 can output the entity-value pair 422 (specimen laterality, right) from node 402 and the context 424 (lung) from node 404a. Moreover, based on the context 424 indicating that the entity is related to the lung, the NLP model 328 can further output the entity-value pairs 426 (tumor location, middle lobe of the lung) from nodes 406b and 408 from sequence 420. In some examples, the NLP model 328 can output the entity-value pair 426 based on detecting that the sequence of text strings "Right", "Lung", and "Middle", and that such a sequence leads to the entity-value pair 426, even if the text string "Lobe" is not found in sequence 420. The extracted entities and their values can be collected into structured medical data and stored back in the patient database 301.

[0044]

[0050] In some examples, the NLP model 328 can include a hierarchy of submodels, such as a baseline NLP submodel and a pathology NLP submodel specific to pathology entities. The baseline NLP submodel can be used to provide a primary context for identifying a sequence of text strings that include common medical terms, which may (or may not) include pathology entities. The primary context can be used to infer the identification of a sequence of text strings that include pathology entities.

[0045]

[0051] FIG. 4B shows another example of the NLP model 328. As shown in FIG. 4B, the NLP model 328 can include a baseline NLP submodel 430 and a pathology NLP submodel 440. The baseline NLP submodel 430 can include, for example, nodes 430a, 430b, and 430c. Nodes 430a and 430b can be associated with general medical terms related to tissue structures such as lesions, tissues, etc., and node 430c can be associated with general medical terms not related to tissue structures such as surgeries. In addition, the pathology NLP submodel 440 can include nodes 440a, 440b, 440c, 440d, 440e, and 440f. Nodes 440a, 440b, 440c, and 440d can be linked by edges 442, 444, and 446 to form the sequence "squamous cell carcinoma of the lung". On the other hand, nodes 440e and 440f are associated with different organs that undergo surgeries such as the heart and breast.

[0046]

[0052] The baseline NLP sub-model 430 can provide context / advice on which part of the pathology NLP sub-model 440 to select to process a sequence of text strings such as sequence 450 shown in FIG. 4B. Specifically, from the text string "number of lesions" within the text string sequence 450, the baseline NLP sub-model 430 can select nodes 440a - 440d of the pathology sub-model 440 to process the remainder of the text string sequence 450. Next, the pathology sub-model 440 can compare the sequences ("squamous cell carcinoma of the lung") associated with nodes 440a - 440d with the remainder of the text string sequence 450. Based on finding that the sequences match, the NLP sub-model 430 can output the entity-value pair 452 (tissue structure, squamous cell carcinoma of the lung).

[0047]

[0053] Note that the NLP model topologies of FIGS. 4A and 4B are provided as examples for illustration. The NLP model 328 can take other forms such as a CRF (conditional random field) classifier as a linear chain sequence model, a CNN Bi-LSTM (convolutional neural network bidirectional long short-term memory).

[0048] C. Improved operation

[0054] Returning to FIG. 3, the improvement module 310 can perform improved operations to improve the quality of medical information extracted from the pathology report image file 302. One exemplary improved operation may include converting entity values within the pathology report to standardized values such as SNOMED concepts, as shown in FIG. 2B. The system 300 may further include a term mapping database 370 to assist with the improved operations by the improvement module 310.

[0049]

[0055] FIG. 4C shows an exemplary refinement operation performed by refinement module 310 using a term mapping database 370 that can include mappings between entity-value pairs and standard terms such as SNOMED concepts and concept IDs. In FIG. 4C, the mapping can be in the form of a mapping table that includes an entity column 454, a value column 456, and a SNOMED concept column 458. For each entity-value pair, refinement module 310 can perform a search for the entity and value within entity column 454 and value column 456, respectively, and the associated SNOMED concept and concept ID within SNOMED concept column 458. In the example of FIG. 4C, for entity-value pair 452 of "tumor site, lower lobe", refinement module 310 can identify the "tumor site" within entity column 454, the "lower lobe" within value column 456, and the SNOMED concept of "Structure of the lower lobe of the lung" and the concept ID of 90572001 within SNOMED concept column 458 370.

[0050]

[0056] In some examples, as part of the refinement process, the refinement module 310 can replace each entity-value pair extracted by the entity extraction module 308 having a mapping to a SNOMED concept with an entity-SNOMED concept pair and store the entity-SNOMED concept pairs in the post-processed pathology report data 304. The replacement of the entity-value pair with its SNOMED concept can improve the post-processed pathology report data 304 by including standard terms in the report, which can reduce the risk of misinterpretation and ambiguity associated with non-standard values of the entity for a human reader. In some examples, the entity-value pair can also be replaced with a SNOMED concept ID to reduce the data size of the post-processed pathology report data 304. Such an arrangement can also facilitate the processing of the post-processed pathology report data 304 by an application. Specifically, since an entity-value pair may have multiple alternative versions of a value representing the same concept, an application that extracts and interprets the entity-value pair needs to have a built-in function to recognize the multiple alternative versions of the value and recognize the associated concept. On the other hand, an application can syntactically interpret the SNOMED concept ID and uniquely link the concept to the concept ID, which can reduce the complexity of the application.

[0051] D. Display Interface to Assist Refinement Operations

[0057] Returning to FIG. 3, the system 300 may include a display interface 305 for displaying the post - processed pathology report data 304. In some examples, the display interface 305 can display the structured medical data of the post - processed pathology report data 304 in a structured format (e.g., in the form of a table, an input form), enabling the portal users (e.g., clinicians, clinical staff) to efficiently identify the medical information they are seeking. In some examples, the display interface 305 can display the pathology report image file 302, as well as highlighting markup (text) overlaid on the text strings that NLP 328 has determined to display pathology entities. The highlighting markup is selectable. The display interface 305 can also detect the selection of highlighting over a set of text strings and display a pop - up window that includes the labels and values of the entities, as well as other refined information of the selected text strings (e.g., standardized data based on SNOMED).

[0052]

[0058] The operation of the display interface 305 can be based on a metadata file 314 that indicates that the pixel positions of each sequence of text strings can also be generated by the optical processing module 306. FIG. 4D shows an example of the metadata file 314. As shown in FIG. 4D, from the pathology report 100, the metadata 462, 464, and 466 can be generated based on entity - value pairs extracted from sections 108, 110, and 112, respectively. Each metadata set can indicate the start and end pixel positions ("start_offset" and "end_offset") of the text string from which the entity - value pair was extracted, the label of the entity, and the value of the entity ("mention"). In some examples, the start and end pixel positions can be presented by pixel numbers that start from the upper - left corner of the image and are counted in a rasterized manner. In some examples, the start and end pixel positions can also be represented by two - dimensional pixel coordinates within the image.

[0053]

[0059] FIG. 4E shows an example of the display interface 305. As shown in FIG. 4E, the display interface 305 can display an image 470 of a pathology report, as well as highlighting markup such as highlighting marks 472, 474, 476, and 480. Each highlighting is overlaid on the image 470 at the start and end pixel positions indicated in the metadata of the text string from which the entity-pair was extracted. In addition, each highlighting is selectable to display the underlying metadata (e.g., by moving the mouse cursor over the highlighting). For example, in FIG. 4E, the display interface 305 can detect that the mouse cursor has moved over the highlighting mark 476 for the text string "excisional biopsy". Based on the pixel position of the mouse cursor, the display interface 305 can identify metadata having a range of pixel positions (represented by start_offset and end_offset) from all the metadata generated for the image 470. The display interface 305 can then extract SNOMED information, the text string, the label of the entity, and the confidence level (score) of the extraction from the identified metadata and display the extracted information in the pop-up window 482.

[0054] E. Training of the natural language processor

[0060] Returning to FIG. 3, the NLP model 328 can be a machine learning model to be trained. As shown in FIG. 3, the system 300 may include a training module 340 that can train the NLP model 328. The training module 340 can train the NLP model 328 based on the labeled general medical documents 348 and the labeled pathology reports 350. The general medical documents 348 can include various categories of biomedical literature, reports, and the like. The training creates nodes representing words of medical terms and edges representing the order relationships between words, such as the edges of the NLP model 328 in FIG. 4A. As part of the training operation, a sequence of text strings with specific labels (e.g., labeled entities, labeled entity values, labeled contexts) can be input into the NLP model 328 to determine whether the NLP outputs the correct entity-value pairs and / or context information. If the training module 340 determines that the NLP model 328 does not output the correct entity-value pairs and / or context information (based on comparing the labeled entities / entity values of the text string sequence with the entity-value pairs output by the NLP model for the text string sequence), the training module 340 can modify the NLP model 328, such as by creating new nodes representing new words and adding edges between existing nodes. The decision mechanism (e.g., a parameterized formula) for outputting entity-value pairs can also be updated (e.g., by updating parameters) to increase the likelihood of outputting the correct entity-pairs and / or context information.

[0055]

[0061] Figures 5A, 5B, 5C, 5D, and 5E illustrate examples of the training operations of the NLP model 328. As shown in Figure 5A, the training operation 500 of the NLP model 328 can be executed in a two-step process. In step 502, a baseline NLP submodel, such as the baseline NLP submodel 430, can be constructed based on labeled general medical documents. As described above, the baseline NLP submodel 430 can be used to provide a primary context for identifying a sequence of text strings that includes common medical terms (which may or may not include pathology report terms). The baseline NLP submodel 430 can be trained based on training data derived from biomedical articles from various major information sources, such as PubMed Central®, a free full-text archive of biomedical and life science journal literature in the National Library of Medicine of the National Institutes of Health in the United States. The training data can include a sequence of text strings with specific labels extracted from biomedical articles (e.g., labeled entities, labeled entity values, labeled contexts).

[0056]

[0062] In step 504, the baseline NLP submodel can be trained using a sequence of text strings from a pathology report, thereby extending the baseline NLP submodel to include a pathology NLP submodel (e.g., the pathology submodel 440) that can detect a sequence of pathology terms. Step 504 can be executed using a CoNLL (Conference on Natural Language Learning) file. The CoNLL file may include text extracted from other pathology reports, and each text can be tagged with a label indicating whether it is an entity label or a non-entity. NLP can be trained based on CoNLL files from multiple pathology reports. In some examples, the training can be specific to a hospital, clinical group, individual clinician, etc., such that NLP can be trained to learn the word preferences of the hospital / clinical group / clinician that maximize the extraction accuracy of entities and their values.

[0057]

[0063] FIG. 5B shows an example of a labeled pathology report 350 that can be in CoNLL format. The labeled pathology report 350 includes a text string to be input into the NLP model 328, as well as labels indicating entities of the text string, which can be used by the training module 340 to induce the output of the NLP model 328 to perform training. The labels can represent reference entities to be output by the NLP model 328 for a sequence of text strings. The training module 340 can then update the parameters of the NLP model 328 based on the difference between the reference entities and the entities actually output by the NLP model 328 for a sequence of text strings. The labeled pathology report 350 can be generated by a human (e.g., a clinician, clinical staff) who can identify the information contained in the pathology report and associate the information with labels. The identification of information and association with labels can be based on a worldwide standard (e.g., SNOMED), and can also be specific to the customs / practices of a particular clinician, medical group, healthcare provider, etc. For example, a clinician may have a particular way of reporting the location of a tumor site, and a pathology report from the clinician can be labeled as such to train the NLP model 328.

[0058]

[0064] As shown in FIG. 5B, each line of the labeled pathology report 350 may include text characters / text strings / text phrases such as text strings 510a, 512a, 514a, 516a, 518a, etc. Each text string is linked with a label, and the label can indicate the context, entity, skipped word, and their locations in the sequence. For example, the label 512b for the word "lung" is "I - localization", which indicates that the word "lung" belongs to the context "localization", and "I" indicates that the word "lung" was found at the beginning of the sequence where the context "localization" should be identified. As another example, the label 514b is "I - laterality", which indicates that the word "right" belongs to the entity "laterality", and "I" indicates that the word "right" was found at the beginning of the sequence where the entity "laterality" should be identified. Further, the labels 516b and 518b are "I - tumor site" and "B - tumor site", respectively. Those labels can indicate that the words "middle" and "lobe" belong to the entity "tumor site", the word "middle" should be found at the beginning of the sequence for the entity, and "B" indicates that the word "lobe" should be found in the middle of the sequence for the entity. Further, the label 510b indicates that the word "4" is skipped text that is not processed by the NLP model 328.

[0059]

[0065] FIG. 5C shows how a sequence of labeled text strings can be processed by the NLP model 328. For each text in the sequence, the training module 340 can determine whether the text is within a node of the NLP model 328, and if the text string is not found, the model can add nodes and / or edges. Moreover, the training module 340 can compare a label (e.g., the entity "laterality") with the output of the NLP model 328, and if the output does not match, the decision mechanism can be updated.

[0060]

[0066] FIG. 5D shows an exemplary distribution 520 of different entities within a labeled sequence of text strings used to train NLP328, and FIG. 5E shows various metrics when measuring the accuracy of entity extraction by NLP328. As shown in FIG. 5D, a relatively large portion of the text string sequence is labeled with "B-Malignancy", "B-Laterality", "B-Size", "B-Type", and "B-Location" (6% - 11%) because these text strings are more commonly found in the middle of the sequence. Moreover, a relatively small portion of the text string sequence is labeled with "B-Outcome", "I-Vessel", "I-Bronchus", and "I-Margin" (0.003% - 0.275%) because these text strings are rarer. The distribution 520 can be based on a corpus of documents from PubMe d Central (registered trademark) and can contain approximately two million words.

[0061]

[0067] Figure 5E shows a table 530 of extraction accuracy metrics for entities output by the NLP model 328 after the model has been trained based on a corpus of documents from PubMed Central® having an entity distribution 520. The extraction accuracy metrics include, for each entity, a true positive (tp) count, a false positive (fp) count, a false negative (fn) count, precision (prec), recall (rec), and an F1 score (f1). The true positive count counts the number of text string sequences that NLP 328 correctly detected as containing a particular entity. The false positive count counts the number of text string sequences that do not contain a particular entity but that NLP 328 incorrectly detected as containing that entity. The false negative count counts the number of text string sequences that contain a particular entity but that NLP 328 incorrectly detected as not containing that entity. Precision, also known as positive predictive value, refers to the proportion of correct positive detections (flagged as sequences containing the entity) out of all positive detections (correct and incorrect). Recall, also known as sensitivity, refers to the proportion of correct positive detections out of all detection results (true positive detections and false negative detections). Precision and recall can be compared based on the following equations. Precision = tp / (tp + fp) (Equation 1) Recall = tp / (tp + fn) (Equation 2)

[0062]

[0068] The F1 score is calculated to provide a measure of the confidence of the detections. A good F1 score is a holistic reflection of both good precision and good recall. Since the NLP model is used in the medical domain, high precision is preferred over high recall. F1 = (Precision × Recall) / (Precision + Recall) (Equation 3)

[0063]

[0069] As shown in FIG. 5E, the average F1 score is approximately 0.85, and the F1 scores of most entities exceed approximately 0.9. Entities with low F1 scores, such as I-margin (0.4), are generally entities that are not well represented in FIG. 5D, thereby making it difficult for the NLP model to accurately detect those entities.

Note 64

[0070] The training of the NLP model 328 can be performed offline or while processing the pathology report image file to dynamically update the NLP model 328. For example, the training of the NLP model 328 can be performed as part of a maintenance operation before the NLP model 328 is used to process the pathology report image file. As another example, the system 300 may include an analysis module 360 that can analyze the correctness of the output (e.g., entity-value pairs, context) of the NLP model 328 from processing the pathology report image file. If the output is incorrect (or the number of incorrect outputs exceeds a threshold), the analysis module 360 can trigger the training module 340 to retrain the NLP model 328. As part of the retraining, text sequences in the pathology report image file that produced incorrect outputs and were labeled with the correct labels can be added to the labeled pathology reports 350 for retraining the NLP model 328.

Note 65

[0071] In addition, various techniques can determine various parameters of the image recognition operation to improve the extraction accuracy of NLP. Parameters for optical character recognition (OCR) operations may include an erosion value, a page iterator level, a page segmentation mode, or a magnification factor. The erosion value can indicate whether a smoothing operation for blurred lines has been performed. The page iterator level can indicate how the image recognition operation was performed by treating the entire page as a block or treating sections (paragraphs, lines, words, characters, etc.) within the page as blocks to increase the granularity of the image recognition operation. The page segmentation mode can detect the tilted orientation of the page being processed and adjust the image recognition operation to correct the tilted orientation. The magnification factor can set the zoom level to zoom in or out on the image to be processed.

[0066]

[0072] In some examples, the adjustment of these OCR parameters can be based on the output of NLP328. Specifically, the image recognition operation can be preconfigured based on a first set of parameter values. The preconfigured OCR operation can be performed on the image of the pathology report to extract a text string, and the text string can be input into NLP to extract pathology entities. Then, the OCR parameters can be adjusted based on the extraction accuracy by NLP.

[0067]

[0073] FIG. 6 shows an example of an adjustment operation 600 for adjusting OCR parameters based on the output of NLP328.

[0074] In step 602, a set of OCR parameters such as erosion value, page iterator level, page segmentation mode, magnification, etc. can be determined. Those parameters can be set to default values or values determined from a parameter sweep operation. The parameter sweep operation can be performed for the image recognition operation on the same set of images of the pathological report, in which the image recognition operation can be repeatedly executed, and each repetition is executed based on a different combination of parameter values. Then, the text recognition accuracy for each repetition can be measured, and the combination of parameter values leading to the highest recognition accuracy can be used to configure the image recognition operation for the workflow.

[0068]

[0075] In step 604, by applying an OCR model with OCR parameters to the image of the pathological report, pathological report text data 312 can be generated.

[0076] In step 606, the pathological report text data can be processed using NLP to extract entity-value pairs.

[0069]

[0077] In step 608, the extraction accuracy of the entity-value pairs by NLP is specified. The accuracy can be specified, for example, based on determining the F1 score according to the above formulas 1 to 3.

[0070]

[0078] In step 610, it is determined whether the extraction accuracy exceeds a threshold. For example, it is determined whether the F1 score exceeds 0.75.

[0079] When the extraction accuracy exceeds the threshold, the OCR parameter adjustment operation can be stored in step 612. However, when the extraction accuracy is below the threshold, the OCR parameters are adjusted in step 614, and then step 604 is repeated. The parameters to be adjusted can be selected based on identifying the entity-value pair with the lowest accuracy. As an example for illustration, it may be determined that some words in a pathology report belonging to an entity-value pair with low accuracy have a very small image size. In such an example, the magnification of the OCR operation can be increased.

[0071]

[0080] In addition to providing an accurate measurement of the extraction of entity-value pairs to precisely indicate the specific OCR parameters to be adjusted, adjusting OCR parameters based on NLP output can be advantageous in other scenarios. For example, in a case where an image file contains notes by a specific doctor that may include non-standard codes and phrases, when the OCR output is compared with a standardized phrase to determine text recognition accuracy, the comparison may lead to incorrect conclusions regarding text recognition accuracy. For example, a text string containing non-standard codes and phrases may be wrongly flagged as an error when in fact the OCR operation has correctly extracted the text string. On the other hand, since the NLP model is trained to recognize non-standard codes and phrases as well as standardized terms, using the output of NLP to determine text recognition accuracy can ensure that the text recognition accuracy measurement is not overly affected by the presence of non-standard codes and phrases in the OCR output.

[0072] IV. Exemplary Applications of the Post-Processed Pathology Report Data

[0081] FIG. 7 shows an exemplary application of the post - processed pathology report data 304 and the metadata file 314. As shown in FIG. 7, the post - processed pathology report data 304 can be provided to a clinician portal 702 that can include the display interface 305 of FIG. 4E. In some examples, the clinician portal 702 can display entity - value pairs (and / or SNOMED concepts) to the user in a predetermined structured format (e.g., in the form of a table, an input form) to enable the user of the portal (e.g., a clinician, clinical staff) to efficiently identify the medical information they are looking for. As another example, the clinician portal 702 can also display the image of the original pathology report, and some or all of the text strings can be replaced with entity - value pairs and / or SNOMED concepts, or the text strings are highlighted and tagged with entity - value pair / SNOMED concepts. The clinician portal 702 can perform highlighting of text strings in the image based on the pixel positions of the text strings shown in the metadata file 314 as described in FIG. 4E.

[0073]

[0082] As another example, the post - processed pathology report data 304 can be provided to a searchable database 704, and entities and their values (standardized or not) can be retrieved therefrom based on a search query. The searchable database as well as the structured medical data can also be made available to various applications such as a clinical decision support application 706, an analysis application 708, etc. for processing. For example, the clinical decision support application can retrieve entities (e.g., diagnosis history, treatment history, medication history) related to a clinical decision and their values from the database to process the entities and generate an output to support the clinical decision. The analysis application can also, for example, retrieve entities related to treatment history and diagnosis from the pathology reports of a number of patients and perform an analysis to obtain insights into the quality of medical care and nursing.

[0074] V. Method

[0083] FIG. 8 shows a method 800 for automated information extraction and refinement. The method 800 can be executed, for example, by the system 300 of FIG. 3.

[0075]

[0084] In step 802, the optical processing module 306 receives an image file (e.g., image file 302) that includes a pathology report. The image file can be received from various primary information sources (e.g., in one or more medical facilities), including, for example, an EMR (Electronic Medical Record) database, a PACS (Picture Archiving and Communication System), a digital pathology (DP) system, a LIS (Laboratory Information System) that includes genomic data, a RIS (Radiology Information System), a patient-reported outcome database, wearable and / or digital technology, and social media. The image file can be in various formats, such as, for example, a portable document format (pdf), or a bitmap image file (BMP file), and can be obtained by scanning a paper-based pathology report.

[0076]

[0085] In step 804, after receiving the image file, the optical processing module 306 can perform an image recognition operation to extract an input text string from the image file. The extraction may include identifying a text image from the image file, generating text data represented by the text image, and generating an intermediate text file (e.g., text file 312) containing the text data. The image recognition operation may include, for example, optical character recognition (OCR) or optical word recognition. In both operations, the optical processing module 306 can extract a pixel pattern of a character (by, for example, identifying a pattern of pixels having a dark color), compare each pixel pattern with a predefined pixel pattern of a character, and determine which character (or which word / phrase) each pixel pattern represents based on the comparison. The optical processing module 306 can then store the character / word / phrase in the text file 312. The optical processing module 306 can scan through the image file 312 according to a predetermined pattern (e.g., raster scan) to extract and process the pixel patterns of the rows from left to right, and can repeat the scan for each row. Based on the scan pattern, the optical processing module 306 can generate a sequence of text strings (e.g., characters, words, phrases) and store the sequence of text strings in the text file 312.

[0077]

[0086] In step 806, the entity extraction module 308 can use a natural language processing (NLP) model (e.g., NLP model 328) to identify entities from the input text string, and each entity includes a label and a value.

[0078]

[0087] In step 808, the entity extraction module 308 can also extract entity values from the input text string using an NLP model. Specifically, the NLP model 328 processes a sequence of text from the text file 312 and, based on recognizing a specific sequence of the text string, determines that a subset of the sequence text is an entity value and can identify an entity-value pair for the subset. As described above, the NLP model 328 includes a graph with nodes. Each node may correspond to a text string and can be connected to another node via an arc. The nodes and arcs can define a sequence of text. The nodes are also organized hierarchically, and detection outputs such as entity-value pairs, contexts, etc. can be generated from each layer. The detection can be based on, for example, a parameterized formula that calculates a score based on the similarity between the input sequence of the text string and the text string represented by the nodes, and predetermined entity-pair and / or context information can be output based on the score. The NLP model 328 can process a sequence of text strings by searching for a sequence of nodes from the graph that matches the sequence (exactly or up to a predetermined proximity). From the identified sequence, the NLP model 328 can output entity-value pairs. In some examples, the NLP model 328 may include a baseline NLP submodel 430 and a pathology NLP submodel 440, and the NLP model 328 can be trained in a two-step process, first with a sequence of text strings from general medical documents and then with a sequence of text strings from pathology reports, as described in FIGS. 5A-5D.

[0079]

[0088] In some examples, the parameters of the image recognition operation can also be adjusted based on the accuracy of the output of the NLP model 328. Specifically, as described in FIG. 6, the image recognition operation in the image processing module 306 can be preconfigured based on a first set of parameter values. The preconfigured image recognition operation can be performed on the image of the pathology report to extract a text string, and the text string can be input to the NLP to extract pathology entities. Then, the parameters of the image recognition operation can be adjusted based on the extraction accuracy by the NLP.

[0080]

[0089] In step 810, the refinement module 310 can convert the values of at least some entities into corresponding predetermined terms using a mapping table that maps entities and values to predetermined terms. The predetermined terms can include standard terms defined based on a universal standard such as SNOMED. The mapping table can be based on data stored in a term mapping database that can include mapping between entity-value pairs and standard terms such as SNOMED concepts and concept IDs. For each entity-value pair and associated context, the refinement module 310 can perform a search for the associated SNOMED concepts and concept IDs within the term mapping database 370.

[0081]

[0081]

[0090] In step 812, the improvement module 310 can generate a post-processed pathology report that includes entities detected from the input text string and corresponding predetermined terms. Specifically, the improvement module 310 can replace each entity-value pair from the NLP model 328 having a mapping to SNOMED concepts with SNOMED concepts and store the SNOMED concepts in the post-processed pathology report text file. In some examples, the entity-value pairs can also be replaced with SNOMED concept IDs to reduce the data size of the post-processed pathology report text file. The post-processed pathology report can then be provided to support various applications, such as for display on a clinician portal, for storage in a searchable database, for processing by a clinical decision support application, an analysis application, etc.

[0082] VI. COMPUTER SYSTEM

[0091] Any of the computer systems referred to herein can utilize any suitable number of subsystems. Examples of such subsystems are shown in FIG. 9 in computer system 10. In some embodiments, the computer system includes a single computer device, and the subsystems can be components of the computer device. In other embodiments, the computer system can include multiple computer devices, each of which is a subsystem and has internal components. The computer system can include desktop and laptop computers, tablets, mobile phones, and other mobile devices. In some embodiments, a cloud infrastructure (e.g., Amazon Web Services), a graphics processing unit (GPU), etc. can be used to implement the disclosed techniques.

[0083]

[0092] The subsystems shown in FIG. 9 are interconnected via system bus 75. Further subsystems are illustrated, such as printer 74, keyboard 78, storage device 79, and monitor 76 coupled to display adapter 82. Peripheral devices and input / output (I / O) devices coupled to I / O controller 71 can be coupled to the computer system by any number of means known in the art, such as input / output (I / O) port 77 (e.g., USB, FireWire (registered trademark)). For example, I / O port 77 or external interface 81 (e.g., Ethernet, Wi-Fi) can be used to connect computer system 10 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus 75 enables central processor 73 to communicate with each subsystem and to control the execution of a plurality of instructions from system memory 72 or storage device 79 (e.g., a fixed disk such as a hard drive, or an optical disk), as well as the exchange of information between subsystems. System memory 72 and / or storage device 79 can embody computer-readable media. Another subsystem is data collection device 85, such as a camera, microphone, accelerometer, etc. Any of the data mentioned herein can be output from one component to another and can be output to the user.

[0084]

[0093] A computer system can include a plurality of the same components or subsystems connected together, for example, by an external interface 81 or an internal interface. In some embodiments, a computer system, subsystem, or device can communicate via a network. In such cases, one computer can be considered a client and another computer can be considered a server, and each can be part of the same computer system. A client and a server can each include a plurality of systems, subsystems, or components.

[0085]

[0094] Aspects of the embodiments may be implemented in the form of control logic using hardware (e.g., application specific integrated circuits or field programmable gate arrays) and / or using computer software having a processor that is generally programmable in a modular or integrated manner. Processors used herein include single-core processors, multi-core processors on the same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, those skilled in the art will know and understand other ways and / or methods of implementing the embodiments of the present invention using hardware and combinations of hardware and software.

[0086]

[0095] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language, for example, using conventional techniques or object-oriented techniques, such as Java, C, C++, C#, Objective-C, Swift, etc., or script languages such as Perl or Python. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media can include random access memory (RAM), read-only memory (ROM), magnetic media such as hard drives or floppy disks, or optical media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, etc. The computer-readable media may be any combination of such storage devices or transmission devices.

[0087]

[0096] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks that comply with various protocols including the Internet. Accordingly, a computer-readable medium may be created using a data signal encoded with such a program. A computer-readable medium encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable medium may exist on or within a single computer product (e.g., a hard drive, CD, or an entire computer system), or on or within different computer products within a system or network. A computer system may include a monitor, printer, or other appropriate display to provide a user with any of the results mentioned herein.

[0088]

[0097] Any of the methods described herein may be executed, in whole or in part, by a computer system including one or more processors configured to execute the steps. Accordingly, embodiments may be directed to a computer system configured to execute any of the steps of the methods described herein, and potentially, different components may execute respective steps or respective groups of steps. Although presented as numbered steps, the steps of the methods herein may be executed simultaneously or in a different order. Further, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of the steps may be optional. Further, any of the steps of any of the methods may be executed by modules, units, circuits, or other means for performing these steps.

[0089]

[0098] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the present invention. However, other embodiments of the present invention may be directed to specific embodiments related to individual aspects, or specific combinations of these individual aspects.

[0090]

[0099] The above description of the exemplary embodiments of the present invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise forms described, and many improvements or modifications are possible in light of the above teachings.

[0091]

[0100] The description of "a", "an", or "the" means "one or more" unless otherwise specified. The use of "or" means "inclusive or" rather than "exclusive or" unless otherwise specified. References to a "first" component do not necessarily require that a second component be provided. Moreover, references to a "first" or "second" component do not limit the referenced component to a particular location unless explicitly stated. The term "based on" means "at least partially based on".

[0092]

[0101] All patents, patent applications, publications, and specifications mentioned herein are hereby incorporated by reference in their entirety for all purposes. There is no admission that any of the foregoing is prior art.

Claims

1. A method executed by a computer system, comprising: receiving an image file including a report containing clinical data; performing an image recognition operation on the image file to extract an input text string; detecting entities from the input text string using a natural language processing (NLP) model, each entity including a label and a value, the NLP model including a baseline NLP sub-model and a pathology NLP sub-model, the baseline NLP sub-model being trained based on a first training text string from general medical documents, and the pathology NLP sub-model being trained based on a second training text string from pathology reports; extracting the value and the label of the entity from the input text string using the NLP model; converting the values of at least some of the entity and value pairs to the corresponding predetermined terms based on a mapping table that maps the entity and value pairs to the predetermined terms; generating a post-processed report including the entities detected from the input text string and the corresponding predetermined terms; and a method.

2. The method according to claim 1, wherein the image recognition operation includes at least one of an optical character recognition (OCR) process or an optical word recognition process.

3. The method according to claim 1, wherein the image file is in a portable document format (pdf) format.

4. The NLP model includes a graph having nodes and edges, each node corresponding to a text string, an edge between two nodes indicating an order relationship between two text strings represented by the two nodes, and the step of detecting the entity includes comparing a sequence of text strings of the input text string with a sequence of text strings represented in the graph. The method according to claim 1.

5. The method according to claim 4, further comprising updating the graph based on a training text string tagged with the name of the entity.

6. specifying the accuracy of recognizing the entity from the input text string by the NLP model; Updating the training text string based on the input text string based on the accuracy; Updating the graph based on the updated training text string The method according to claim 5, further comprising.

7. The method according to claim 1, wherein a plurality of entities are recognized from a set of adjacent text strings of the input text string.

8. The input text string is a first input text string, The parameters of the image recognition operation are determined based on the accuracy of recognizing entities from a second input text string by the NLP model, and the second input text string is generated by the image recognition operation using the parameters. The method according to claim 1.

9. The method according to claim 1, wherein the predetermined term is based on the International Classification of Diseases (ICD) standard, and the predetermined term includes at least one of an ICD concept or an ICD concept identifier (ID).

10. The method according to claim 9, wherein the mapping is based on a plurality of entities.

11. The method according to claim 1, further comprising providing structured medical data to at least one of a clinical decision support tool, a healthcare provider portal, or a searchable medical database.

12. The image recognition operation outputs an image position of the input text string in the image file, The method is, Displaying the image file in a display interface; Displaying a highlight markup on a subset of the input text strings where entities are detected based on the image position; Detecting a selection of at least one of the highlight markups; Displaying a pop-up window on the selected highlight markup in response to detecting the selection, the pop-up window including the predetermined term of the entity detected from the input text string of the selected highlight markup. [[ID= ​ The method according to claim 1, wherein the image file is received from one or more information sources including at least one of an EMR (Electronic Medical Record) database, a PACS (Picture Archiving and Communication System), a digital pathology (DP) system, an LIS (Laboratory Information System), an RIS (Radiation Information System), a patient-reported outcome database, a wearable device, or a social media website.

14. A computer-readable medium storing a plurality of instructions for controlling a computer system to execute the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Generation of natural language processing models for information domains

    JP2015505082A

  • Method and system for visualization of patient history

    JP2017513590A

  • Automatic identification and extraction of medical condition and fact from electronic medical treatment record

    JP2019049964A

  • Mobile supplementation, extraction, and analysis of health records

    US10395772B1