Automatically extracting structured labels from medical text using deep convolutional networks and using them to train computer vision models

By generating structured tags associated with medical text and images, the difficulties of unstructured data in medical information processing are solved, and the efficiency of automated information extraction and machine learning model training is achieved.

CN111727478BActive Publication Date: 2025-05-06GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201880089613.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-02-16
Publication Date
2025-05-06
Estimated Expiration
2038-02-16

AI Technical Summary

Technical Problem

The unstructured forms of medical texts and images lead to difficulties in information retrieval, automation tool development, quality monitoring and bill processing, especially in machine learning applications that require structured data.

Method used

This process is automated using natural language processors and computer vision models in the form of a one-dimensional deep convolutional neural network by generating structured tags associated with free text medical reports and associating these tags with medical images.

Benefits of technology

It realizes automatic extraction of structured information from unstructured medical texts and images, improves the efficiency of information retrieval and machine learning model training, and reduces the need and cost of manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111727478B_ABST
    Figure CN111727478B_ABST
Patent Text Reader

Abstract

A method for processing medical text and associated medical images is provided. A natural language processor configured as a deep convolutional neural network is trained on a first corpus of selected free-text medical reports, each of which has one or more structured tags assigned by a medical expert. The network is trained to learn to read additional free-text medical reports and generate predicted structured tags. The natural language processor is applied to a second corpus of free-text medical reports associated with medical images. The natural language processor generates structured tags for the associated medical images. A computer vision model is trained using the medical images and the generated structured tags. Thereafter, the computer vision model can assign structured tags to other input medical images. In one example, the medical image is a chest X-ray.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to systems and methods for processing medical text and associated medical images. Background Art

[0002] Medical opinions and observations are often communicated in the form of natural language, free-text reports prepared by physicians. This unstructured text prevalence in medical records can create significant difficulties for downstream consumers of this information. Healthcare providers maintain archives of large volumes of medical reports, but due to the low level of standardization of medical notes, it is difficult to fully utilize the potential of such reports.

[0003] Unstructured medical notes can present challenges in many situations:

[0004] 1. Retrieve cases of interest for review and study (e.g., identify all patients with a rare disease).

[0005] 2. Quickly learn about pre-existing conditions that plague patients before providing care. Reading through many lengthy reports can be time consuming and prone to errors.

[0006] 3. Develop automated tools to support medical decision making. Most AI techniques require structured data.

[0007] 4. Quality monitoring, such as determining how often patients with condition X receive treatment Y.

[0008] 5. Billing. In most healthcare systems, medical encounters must be converted into a standard set of codes to qualify for reimbursement.

[0009] In use case (1), a clinician or researcher is trying to identify records of patients suffering from a specific condition (e.g., or receiving a specific form of treatment) from a large hospital archive. This might be for the purpose of recruiting subjects to a clinical trial based on eligibility criteria, or for educational / training purposes. Without a database of structured clinical variables to query, the task would be daunting. Simple string matching of text notes would likely have very low specificity. And later, medical staff would have to comb through each record to determine whether the specific condition was confirmed.

[0010] In use case (2), a healthcare provider is trying to grasp a patient's medical history during an initial encounter. The patient may have been in the healthcare system for some time; in order to effectively treat the patient, they must understand previous encounters based on scattered notes written by numerous providers (residents, radiologists, and other specialists). This can be time-consuming and error-prone when there is a lot of documentation to review. An automated system that surfaces specific clinical variables based on previous documentation and draws the reader's attention to relevant sections of a longer note can save valuable time in the clinic.

[0011] Research in medical informatics often attempts to replicate the judgment of human experts through automated systems. Such efforts often rely on retrospective data that includes many instances of patient information (e.g., laboratory tests, medical images) paired with diagnoses, treatment decisions, or survival data. In the paradigm of supervised learning, patient data are the "inputs" and diagnoses, outcomes, or treatment recommendations are the "outputs." The modeling task is then to learn to predict this information given access to patient records at an earlier point in time. For example, an application of computer vision might attempt to automatically determine the cause of a patient's chest pain by examining an X-ray image of the patient's chest. Such a system can be trained by leveraging previously captured images and their corresponding opinions by radiologists.

[0012] In order to train and evaluate these models, it is important to have highly structured outputs on which clinically important metrics can be defined. An example of a highly structured output would be a code for a subsequently prescribed medication. In many areas of medicine (e.g. radiology, pathology), the "input" to the diagnostic system is always archived as digital images and is therefore easily available for modern machine learning methods. Unfortunately, the historical diagnosis (our "output") is almost always buried in natural language provided by the physician. This text is often unstructured and poorly standardized, making it difficult to use in a machine learning context. Recasting this information using standardized schemas or specifications requires a lot of effort from medical staff and is therefore both time-consuming and costly. Summary of the invention

[0013] The present disclosure relates to a method of generating structured labels for free text medical reports (e.g., doctor's notes) and associating such labels with medical images (such as chest X-rays) attached to the medical reports. The present disclosure also relates to a computer vision system that generates structured labels or classifies structured labels from input image pixels only. The computer vision model can help diagnose or evaluate medical images (e.g., chest X-rays).

[0014] In this document, the term "structured label" refers to data according to a predetermined pattern, or in other words, data in a standardized format. The data is used to convey diagnosis or medical condition information for a specific sample (e.g., a free text report, an associated image, or a separate image). An example of a structured label is a set of one or more binary values ​​(e.g., positive (1) or negative (0)) for a medical condition or diagnosis (such as pneumothorax (collapsed lung), pulmonary embolism, gastric feeding tube malposition, etc. in the context of a chest x-ray or associated medical report). Alternatively, a structured label can be a pattern or format in the form of a set of one or more binary values ​​for assigning diagnostic billing codes, such as "1" indicating that a specific diagnostic billing code for pulmonary embolism should be assigned, and "0" indicating that this code should not be assigned. The specific pattern will depend on the application, but in the example of a chest x-ray and associated medical report, it may consist of a set of binary values ​​to indicate the presence or absence of one or more of the following conditions: airspace opacity (including emphysema and consolidation), pulmonary edema, pleural effusion, pneumothorax, cardiomegaly, nodules or masses, nasogastric tube malposition, endotracheal tube malposition, central venous catheter malposition, and the presence of a chest tube. In addition, each of these conditions may have a binary modifier such as laterality (e.g., left is 1, right is 0) or severity. The structured label may also include a severity term, which in one possible example may be a binary modifier that can be 1, 2, or 3 bits of information in the structured label to encode different severity values. Or, as another example, the structured label may encode some specification or pattern of severity, such as some integer value (e.g., absent = 0, mild = 1, moderate = 2, severe = 3, or severity levels from 1 to 10).

[0015] For example, a structured label for a given medical report or for a given medical image may take the form of [100101], where each bit in the label is associated with a positive or negative finding, billing code, or diagnosis for a particular medical condition, and in this example, there are six different possible findings or billing codes for this example. Structured labels can be either categorical or binary. For example, a model may produce a label within the set {absent, mild, moderate, severe} or some numerical equivalent thereof.

[0016] The method uses a Natural Language Processor (NLP) in the form of a one-dimensional deep convolutional neural network trained on a corpus of curated medical reports with structured labels assigned by medical experts. The NLP learns to read free-text medical reports and produces predicted structured labels for such reports. The NLP is validated on a set of reports and associated structured labels to which the model has not been previously exposed.

[0017] NLP is then applied to a corpus of medical reports (free text) that have associated medical images (such as chest X-rays) but no clinical variables of interest or structured labels are available. The output of NLP is a structured label associated with each medical image.

[0018] The medical images with structured labels assigned by NLP are then used to train a computer vision model (e.g., a deep convolutional neural network pattern recognizer) to assign or copy the structured labels to the medical images based solely on the image pixels. Thus, the computer vision model essentially acts as a radiologist (or more generally, acts as an intelligent assistant to the radiologist) that produces structured outputs or labels for medical images (such as chest X-rays) rather than natural language or free text reports.

[0019] In one possible embodiment, the method includes a technique or algorithm known as integrated gradients, implemented as a software module that assigns attributions to words or phrases in a medical report that contribute to the assignment of a structured label to an associated image. The method allows such attributions to be presented to a user, for example, by displaying an excerpt of a medical report with the relevant terms highlighted. This has great applicability in healthcare scenarios where healthcare providers often sift through long patient records to find information of interest. For example, if NLP identifies that a medical image shows signs of a cancerous lesion, the relevant text in the associated report can be highlighted.

[0020] In one configuration of the present disclosure, a system for processing medical text and associated medical images is described. The system includes: (a) a computer memory storing a first corpus of curated free-text medical reports, each of which has one or more structured tags assigned by a medical expert; (b) a natural language processor (NLP) configured as a deep convolutional neural network, trained on the first corpus of curated free-text medical reports to learn to read additional free-text medical reports and generate predicted structured tags for the additional free-text medical reports; and (c) a computer memory storing a second corpus of free-text medical reports associated with medical images (and typically without structured tags). The NLP is applied to the second corpus of free-text medical reports and responsively generates structured tags for the associated medical images. The system also includes (d) a computer vision model trained on the medical images and the structured tags generated by the NLP. The computer vision model operates (i.e., performs inference) to assign structured tags to other input medical images (e.g., chest X-rays without associated free-text medical reports). In this way, the computer vision model acts as an expert system to read medical images and generate structured labels, like a radiologist, or as an assistant to a radiologist. NLP and computer vision models are typically implemented in a dedicated computer configured with hardware and software for implementing deep learning models as is customary in the art.

[0021] In one embodiment, the computer vision model and NLP are trained in an integrated manner, as will be apparent from the following description.

[0022] In one possible configuration, the system includes a module implementing an integrated gradient algorithm that assigns attributions to structured tags generated by NLP to input words in a free-text medical report. The system may also include a workstation having a display for displaying the medical image, the free-text report, and the attributions of elements in the report calculated by the integrated gradient algorithm.

[0023] As will be explained below, in one embodiment, the medical image is a chest X-ray. A computer vision model is trained to generate structured labels for the chest X-ray based solely on the image pixels, without the need for an associated free-text medical report. As described above, the structured labels can be, for example, a series of binary values ​​indicating the positivity or negativity of a particular finding, or a diagnostic billing code, optionally including laterality or severity.

[0024] In another aspect of the present disclosure, a method for processing medical text and associated medical images is provided, the method comprising:

[0025] training a natural language processor on a first corpus of free-text medical reports, each of the medical reports having one or more structured tags assigned by a medical expert, training the natural language processor to learn to read additional free-text medical reports and to generate predicted structured tags for the additional free-text medical reports;

[0026] applying a natural language processor to a second corpus of free text medical reports without structured tags associated with medical images, and wherein the natural language processor generates structured tags for the associated medical images; and

[0027] Medical images and structured labels generated by a natural language processor are used to train a computer vision model to assign structured labels to other input medical images.

[0028] In another aspect of the present invention, there is provided a machine learning system, the combination of which comprises:

[0029] a computer memory storing a training dataset in the form of a plurality of training examples, each of the training examples comprising a free-text medical report and one or more associated medical images, wherein a subset of the training examples contain ground truth structured labels assigned by a medical expert; and

[0030] A computer system configured to operate on each of a plurality of training examples, and wherein the computer system is configured to

[0031] a) a feature extractor receiving one or more medical images as input and generating a vector of extracted image features from the one or more medical images;

[0032] b) a diagnosis extraction network that receives as input a free text medical report and the vector of extracted image features and generates a structured label;

[0033] c) an image classifier trained on vectors of extracted features and corresponding structured labels generated by a feature extractor from one or more medical images of a large number of training examples; wherein the image classifier is also configured to generate structured labels for other input medical images based on the feature vectors generated by the feature extractor from the other input medical images.

[0034] The term "curated" is used to indicate that the corpus of data includes at least some elements generated by human intelligence, such as structured labels assigned by medical experts during training. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1is a schematic diagram of the machine learning process used to generate structured labels for a collection of medical images and associated free-text medical reports.

[0036] Figure 2 is a schematic diagram of the process of developing a natural language processor (NLP) in the form of a deep convolutional neural network by training it on a corpus of curated free-text medical reports; the neural network learns to read additional free-text medical reports and produces predicted structured tags for the additional free-text medical reports.

[0037] Figure 3 is used Figure 2 NLP is the process of generating structured labels and attributions for free-text medical reports associated with one or more medical images.

[0038] Figure 4 is a schematic diagram of a method called integrated gradients that generates attributions for input tokens; although the schematic diagram shows attribution of tokens to pixels in an input image, the method can be performed similarly for attribution of tokens where the input is words in a free-text medical report.

[0039] Figure 5 and Figure 6 is a diagram of a monitor on a workstation with highlighted text in a free-text medical report. Figure 3 The generated structured tags shown have highly attributed text elements (words).

[0040] Figure 7 is a schematic diagram of a method for training a computer vision model (e.g., a deep convolutional neural network pattern recognizer) from a collection of medical images and associated structured labels. The medical images and associated labels can be, for example, from Figure 3 Medical images and labels in .

[0041] Figure 8 is used Figure 7 A computer vision model is used to generate structured labels for an input medical image (e.g., a chest X-ray).

[0042] Fig. 9 is a schematic diagram of an integrated approach to training computer vision models and NLP in the form of Fig. 9 The network is identified as the “diagnostic extraction network”. DETAILED DESCRIPTION

[0043] In one aspect, the work described in this document relates to methods for developing computer vision models, typically embodied in a computer system including memory and processor capabilities and consisting of certain algorithms, that are capable of emulating human expert interpretation of medical images. To build such models, we follow the highly successful supervised learning paradigm: deep learning computer vision models are trained to replicate the expert diagnostic process using many examples of radiological images (inputs) paired with a set of corresponding findings generated by radiologists (outputs). The goal is that if we train on enough data from enough doctors, our models may be able to surpass human performance in certain areas.

[0044] Figure 1 is a schematic diagram of a machine learning process for generating structured labels for a collection of medical images and associated free-text medical reports, which will explain the motivation for the present disclosure. In this scenario, the goal is to generate a machine learning model 14 that can use one or more medical images 10 and free-text reports 12 associated with the medical images as input (X) and produce output (Y) in the form of structured labels 16. For each input X1...X N , model 14 generates outputs Y1…Y N The model 14 is shown to have two parts, natural language processing (NLP) applied to the text portion, and inference (pattern recognizer) applied to the medical images.

[0045] Unfortunately, in most medical clinics, radiologists convey their opinions back to referring physicians in the form of unstructured natural language text reports of hundreds of words. Thus, while the input to our computer vision models — radiological images — are always archived in digital format and require little preprocessing, the output — the medical diagnosis, or findings, or equivalent structured labels that the computer vision model should replicate — is buried in unstructured, poorly standardized natural language.

[0046] Although frameworks exist for directly converting images to text strings, our application requires a more structured output, i.e. Figure 1 of labels 16 on which clinically important indicators can be defined. One way to address this problem could be to recruit radiologists to re-review historical scans and annotate them using a structured form. However, given that all of these scans have already been interpreted, trying to meet the data requirements of state-of-the-art computer vision models in this way would be slow, expensive, and wasteful. Therefore, the present disclosure provides a more efficient and cost-effective method to generate structured labels and subsequently train computer vision models.

[0047] As described in this document, structured data can be extracted from free-text radiology reports ( Figure 1 An automated system for labeling of images16) could unlock vast amounts of training data for our computer vision models and associated algorithms. Investing in such a system would have extremely high leverage by collecting data that has been manually labeled explicitly for this purpose, rather than collecting new image-based labels for training. The annotation task would be several times faster per case and could be performed by less specialized humans. Additionally, once a natural language processing model has been trained, it can be applied to any number of cases, providing a virtually endless supply of training data for computer vision models.

[0048] Thus, one aspect of the present disclosure relates to a method for generating structured tags for free text medical reports (e.g., physician's notes) and associating such tags with medical images, such as chest X-rays, attached to the medical reports. Figure 2 , the method uses a natural language processor (NLP) 100 in the form of a one-dimensional deep convolutional neural network trained on a curated corpus 102 of medical reports consisting of a large number of individual free text reports 104, each of which is associated with a corresponding structured label 106 assigned by a medical expert (i.e., a human expert). Such input by the expert is shown at 108, and can cause, for example, one or more medical experts to review the reports 104 using a workstation and, using the workstation's interface, assign one or more binary values ​​to various diagnoses or medical conditions that, in their judgment, occur in a patient associated with the report in a table-type document, e.g., as previously described, a label such as [110110], where each bit is associated with a positive or negative for a particular medical condition, a diagnostic billing code, etc. Additionally, as previously described, the label can contain bit or alphanumeric data to reflect laterality or severity.

[0049] The NLP 100 learns to read free-text medical reports from a training corpus 102 and to generate predicted structured labels for such reports. For example, after the network has been trained, it can be applied to a collection 110 of free-text reports and for each report 112A, 112B, and 112C, it generates structured labels 114A, 114B, and 114C, respectively. Additionally, the NLP 100 is trained on a collection of reports and associated structured labels ( Figure 2, 120). Background information on machine learning methods for generating labels from text reports is described in Chen et al., Deep Learning to Classify Radiology Free-Text Reports, Radiology, November 2017, the contents of which are incorporated herein by reference. The scientific literature includes various descriptions of convolutional neural networks for text classification that may be suitable for use in the present context, e.g., Yoon Kim, Convolutional Neural Networks for Sentence Classification, https: / / arxiv.org / pdf / 1408.5882.pdf, and Ye Zhang et al., A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification, https: / / arxiv.org / pdf / 1510.03820.pdf, and Alexis Conneau et al., Very Deep Convolutional Networks for Text Classification, arXiv:1606.01781 [cs.CL] (January 2017), so a more detailed description of NLP 100 is omitted for the sake of brevity.

[0050] Reference now Figure 3 , and then according to Figure 2 The trained NLP 100 is applied to a corpus 300 of medical reports (free text) 302, each of which is associated with one or more medical images 304 (such as, for example, chest X-rays), but for which no structured labels encoding the clinical variables of interest are available. The output of the NLP is the structured labels 304 associated with each image.

[0051] For example, the NLP 100 may have as input a free text report 302 indicating the presence of a misplaced gastric feeding tube and an associated chest X-ray showing this condition. Figure 2 The network is trained to generate a label for a misplaced feeding tube (e.g., a label such as [100000], where 1 indicates a misplaced feeding tube and 0 indicates no other condition), and the label [100000] is assigned to the associated chest X-ray.

[0052] Generating structured labels from free text reports and assigning them to the associated medical images allows for the automatic generation of large numbers of medical images with assigned structured labels. We previously pointed out that in theory, radiologists could be recruited to re-review historical scans and annotate them using a structured form. However, given that all of these scans have already been interpreted, it would be slow, expensive, and wasteful to try to meet the data requirements of state-of-the-art computer vision models in this way. Instead, using Figure 2 The training process and Figure 3 We automatically generate structured labels for medical images without requiring a lot of time from trained radiologists.

[0053] In one embodiment, our method includes a technique called integrated gradient, which is a technique for Figure 3 In the method of assigning attributions to words or phrases in a medical report that contributed to the generation of a label for an associated image, we use a software module 310 that implements an integrated gradient algorithm that generates weights or attributions for specific words or strings of words in a medical report that contributed significantly to the generation of a label. This approach is conceptually similar to an "attention mechanism" in machine learning. The output of software module 310 is displayed as attributions 308. Our method allows such attributions to be presented to a user, for example, by displaying an excerpt of a medical report with the relevant terms highlighted, see discussion below. Figure 5 and Figure 6 This has great applicability in medical healthcare scenarios where healthcare providers often sift through long patient records to find information of interest. For example, if the model identifies that a medical image shows signs of a cancerous lesion, relevant text in the associated report can be highlighted. Integrated gradients have been described as having applications in object recognition, diabetic retinopathy prediction, question classification, and machine translation, but its application in this scenario (attribution of textual elements for binary classification of medical images) is new.

[0054] The integrated gradient algorithm is described in M. Sundararajan et al., Axiomatic Attribution for Deep Networks, arXiv:1703.01365 [cs.LG] (June 2017), which is incorporated herein by reference in its entirety. Figure 4 The method is conceptually described in the context of attributing the attributes of individual pixels in an image to the classification of the entire image, and is similarly applicable to individual words in a free-text medical report. Basically, Figure 4As shown, the integrated gradient score IG for each pixel i in the image is calculated on a uniform scaling (α) of the input image information content (in this example, the brightness spectrum) from the baseline (zero information, each pixel is black, α = 0) to the full information in the input image (α = 1) i (or attribution weight or value), where IG i The score of each pixel is given by equation (1)

[0055] (1)

[0056] Where F is the prediction function of the label;

[0057] image i is the RGB intensity of the i-th pixel;

[0058] IG i (image) is the integrated gradient with respect to the ith pixel, i.e., the attribution of the ith pixel; and

[0059] ▽It’s about image i The gradient operator.

[0060] In the context of free-text medical reports, the report is a one-dimensional string of words of length L, with each word represented as, for example, a 128-dimensional vector x in a semantic space. The dimensions of the semantic space encode semantic information, co-occurrence statistics of the occurrence or frequency of a word with other words, and other information content. α = 0 means that each word in the report has no semantic content and no meaning (or can be represented as a zero vector), and as α approaches 1, each word approaches its full semantic meaning.

[0061] A more general expression for the integrated gradient is

[0062] (2)

[0063] Here, the integrated gradient of the input x and the baseline x' along the i-th dimension is defined according to equation (2). Here, is the gradient of F(x) along the i-th dimension. The algorithm is further explained in Section 3 of the paper by Sundararajan et al., which is incorporated herein by reference. The gradient vector with respect to each word is itself a 128-dimensional word vector. Note that the number 128 is somewhat arbitrary, but not an uncommon choice. To obtain the final integrated gradient value for each word, the components are summed. Their signs are preserved, so a net positive score implies an attribution toward the positive class (a value of 1), while a net negative score implies an attribution toward the negative or "null" class (a value of 0).

[0064] As the integrated gradient algorithm is Figure 4 For each pixel in the image example, the attributed value IG is calculated i Similarly, it calculates for each word in the free text report 302 an attribute value associated with the label assigned to the associated image 304. In one possible configuration, the free text report may be associated with a thumbnail of the associated image. Figure 1 Displayed on the workstation, such as Figure 5 and Figure 6 The free text report is color-coded or highlighted in any suitable manner to draw the user's attention to key words or phrases with the highest attribution value scores, thereby helping the user understand why the labels were assigned to the associated images. For example, all words in the report with attribution scores above a certain threshold (which can be specified by the user) are shown in bold, red, larger font, underlined, or other manner to indicate that these words are most important in generating structured labels.

[0065] Then, Figure 3 The medical image 304 with structured labels 306 generated by NLP 100 is used to train a computer vision model (pattern recognizer) to assign or copy structured labels to the medical image based solely on the image pixels. Figure 7 The training of the computer vision model 700 is described in . In this example, the computer vision model 700 is in the form of a deep convolutional neural network. The computer vision model can be implemented in several different configurations. Generally, in the field of pattern recognition and machine vision, Figure 7The type of deep convolutional neural network pattern recognizer used in is known, so for the sake of brevity, its detailed description is omitted. One implementation is the Inception-v3 deep convolutional neural network architecture, which is described in the scientific literature. See the following references, the contents of which are incorporated herein by reference: C. Szegedy et al., Going Deeper with Convolutions, arXiv: 1409.4842 [cs.CV] (September 2014); C. Szegedy et al., Rethinking the Inception Architecture for Computer Vision, arXiv: 1512.00567 [cs.CV] (December 2015); see also U.S. patent application serial number 14 / 839,452 filed by C. Szegedy et al. on August 28, 2015, "Processing Images Using Deep Neural Networks". The fourth generation, known as Inception-v4, is considered an alternative architecture for computer vision models. See C. Szegedy et al., Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning, arXiv:1602.0761 [cs.CV] (February 2016). See also U.S. patent application Ser. No. 15 / 395,530 filed Dec. 30, 2016 by C. Vanhoucke, “Image Classification Neural Networks”. The descriptions of convolutional neural networks in these papers and patent applications are incorporated herein by reference.

[0066] Another alternative approach to computer vision models is described in the paper ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classificationand Localization of Common Thorax Diseases by X. Wang et al., arXiv:1705.02315v5[cs.CV] (December 2017), the contents of which are incorporated herein by reference.

[0067] For the training of Computer Vision Model 700, we used Figure 3 The image 304 and Figure 3The labels generated by the convolutional neural network 100 are explained. In essence, by Figure 7 As shown, trained on this body of medical images 304 and associated structured labels 306, the computer vision model 700 is able to accept input images (pixel data only) and generate structured labels. In other words, the computer vision model does not require an associated free text report to generate structured labels. Instead, it can generate the labels on its own. Therefore, now refer to Figure 8 , Figure 7 Once trained and validated, the computer vision model 700 is able to accept a given input image 800, such as a chest X-ray that is not associated with a medical report, and generate structured labels 802, for example, as a digital assistant for a radiologist.

[0068] One example of how computer vision models can be used is in a radiology scenario in a hospital or clinic setting. In one configuration, the hospital or clinic will have a computer system that is configured with the necessary hardware to implement Figure 7 and Figure 8 Computer vision model 700 (the details of which are not particularly important). In another configuration, the hospital or clinic can also simply request predictions from a computer vision model hosted by a third-party service provider in the cloud. In this way, the hospital or clinic will not need to run model inference on-site, but can configure its software using an application programming interface (API) to call the service provider's computer system in the cloud, which hosts the computer vision model and performs inference on the input image (e.g., chest X-ray) and returns structured labels via the API.

[0069] The patient undergoes a chest X-ray, and the digital X-ray image is provided as input to the computer vision model. Structured labels 802 are generated and then interpreted by the radiologist. For example, the label [110101] is interpreted as positive for pneumothorax, left sided, moderate in severity, and negative for nasogastric tube malposition, and negative for cancerous lesions or masses. The radiologist reviews the X-ray with the help of this additional information from the computer vision model, and it confirms her own findings and diagnosis. As another example, the radiologist reviews the X-ray and has initial findings of pulmonary edema and multiple effusions, but after considering the structured label [010110] indicating cardiac hypertrophy and negative for pulmonary edema, she reconsiders her assessment of the X-ray, confirms that the computer vision model correctly assessed the X-ray, and then makes the correct entry and correct diagnosis in the patient's chart, thereby avoiding a potentially serious medical error.

[0070] One of the advantages of the system and method described in this document is that it allows the computer vision model 700 to be trained from a large number of medical images and associated structured labels, but the labels are automatically generated without requiring extensive input from a human operator. Figure 3 Convolutional neural networks (NLP) do utilize a corpus of medical reports with labels assigned by medical experts in initial training, but in comparison, the amount of human involvement is relatively small and the method of the present disclosure allows the generation of very large numbers of training images with structured labels, on the order of thousands or even tens of thousands, and this large number of training images helps computer vision models both generalize and avoid overfitting to the training data.

[0071] Reference now Fig. 9 , is a schematic diagram of an integrated method for training both a computer vision model (700) and a diagnosis extraction network 900, which is similar to Figure 2 A natural language processor (NLP) or 1D deep convolutional neural network 100 is used. Each training example 902 (only one of which is shown) consists of a medical image 906 (e.g., a chest x-ray) and an associated free text report 904 provided by a radiologist. These reports are made of unstructured natural language and are prone to subjectivity and errors. These are lower quality "noisy" labels. Some portion of the training examples will have associated images that have been independently reviewed by a panel of experts and evaluated according to a structured label pattern or specification (e.g., a series of binary variables representing the presence or absence of various findings as described above). These labels are considered "ground truth" labels.

[0072] The CNN feature extractor 910 is a convolutional neural network or pattern recognizer model (see example above) that processes medical image pixels to develop vector representations of image features that correlate to the classifications produced by the feature extractor. The diagnosis extraction network 900 uses the vector representations of image features as additional input in addition to the free text report 904 to learn to extract clinical variables of interest from the report 904. This allows the network 900 to not only convert natural language into a useful format, but also to correct confusion, bias, or errors in the original report 904. (Note that the convolutional neural network NLP model 100 described previously can be used ( Figure 2) to initialize the diagnosis extraction network, but the difference is that the diagnosis extraction network also receives the output of the CNN feature extractor 910. ) The diagnosis extraction network 900 generates structured findings 910, which can be used to supervise a multi-label classifier 930 (e.g., a deep convolutional neural network pattern recognizer), which operates only on image features when ground truth labels are not available. In this way, we rely on a smaller number of more expensive ground truth labels to train medical image diagnosis models using the abundant but noisy signal contained in the reports.

[0073] exist Fig. 9 The method may optionally include an integrated gradient module 940 that queries structured findings and text in the medical report 904 and generates attribution data 950 that can be presented to a user or stored in a computer memory for use in evaluating the performance of the diagnosis extraction network 900.

[0074] In use, after training Fig. 9 After the computer vision model 700, it can be configured in the computer system and the auxiliary memory and used Figure 8 The illustrated approach is used to generate labels for an input image (e.g., a chest X-ray). In this case, the multi-label classifier 930 generates labels for the input image 800 even if the image is not associated with a free text report.

[0075] Therefore, reference Fig. 9 In one aspect of the present disclosure, we have described a machine learning system that includes a computer memory (not shown) storing a training data set 902 in the form of a plurality of training examples, each of the training examples comprising a free text medical report 904 and one or more associated medical images 906. A subset of the training examples contains ground truth structured labels assigned by a medical expert. The system also includes a computer system configured to operate on each of the plurality of training examples 902. The computer system ( Fig. 9 ) is configured as follows: it includes: a) a feature extractor 910 (which may take the form of a deep convolutional neural network) receiving one or more medical images 906 as input and generating a vector 911 of extracted image features; b) a diagnosis extraction network 900 (substantially similar to Figure 2The NLP 1-D convolutional neural network 100 of FIG. 1 ) receives a free text medical report 904 and a vector 911 of extracted image features as input and generates a structured label 920; and c) an image classifier 930 (e.g., a deep convolutional neural network pattern recognizer) trained on the structured label 920 and the extracted feature vector 911. The image classifier 930 is also configured to generate structured labels for other input medical images (e.g., Figure 8 ).

[0076] like Fig. 9 As shown, the system may optionally include an integrated gradient module 940 that generates attribution data for words in the free-text medical report that contribute to the structured labels generated by the diagnosis extraction network 900. In one embodiment, the medical image is a chest X-ray. The structured label can take the form of a binary label for the presence or absence of a medical condition or diagnosis and / or a label for assigning a specific diagnostic billing code. For example, the medical condition may be one or more of the following conditions: air cavity density (including emphysema and consolidation), pulmonary edema, pleural effusion, pneumothorax, cardiac hypertrophy, nodules or masses, nasogastric tube malposition, endotracheal tube malposition, and central venous catheter malposition.

[0077] The data used for model training complies with all disclosure and use requirements required by HIPAA and provides appropriate delisting of identifiers. Patient data is not linked to any Google user data. In addition, for records used to create models, our system includes a sandbox infrastructure that keeps each record separate from each other based on rules, data licenses, and / or data use agreements. Data in each sandbox is encrypted; all data access is controlled, logged, and audited at the individual level.

Claims

1. A machine learning system for classifying a medical image based on the presence or absence of a medical condition, the system comprising: a computer memory storing a first corpus of free-text medical reports, each of the free-text medical reports having an associated medical image and one or more ground truth structured labels assigned by a medical expert; a computer vision model trained to assign another structured label to an input medical image, the computer vision model comprising a feature extractor configured to generate a vector of extracted image features of the medical image and an image classifier configured to generate another structured label using the vector of extracted image features; a natural language processor configured to receive as input a free-text medical report and a vector of extracted image features generated by a feature extractor of the computer vision model, and to generate a predicted structured label, wherein the natural language processor is trained using a first corpus of free-text medical reports and associated ground truth structured labels to generate predicted structured labels using, in addition to the free-text medical reports, vectors of extracted image features generated by a feature extractor for medical images associated with the free-text medical reports, The system further comprises: computer memory storing a second corpus of free-text medical reports associated with the medical images, and wherein a natural language processor is applied to the second corpus of free-text medical reports and responsively generates a predicted structured tag for the associated medical image, wherein the computer vision model is trained on medical images and predicted structured labels generated by a natural language processor for free-text medical reports of a second corpus to generate another structured label for other input medical images using a feature extractor and an image classifier, wherein the ground truth structured label, the another structured label, and the predicted structured label include labels for the presence or absence of a medical condition.

2. The system of claim 1, further comprising a module implementing an integrated gradient algorithm that assigns to input words in the free-text medical report an attribution for the predicted structured tags generated by the natural language processor.

3. The system of claim 1 or 2, further comprising a workstation having a display for displaying both the medical image and the attributions assigned by the integrated gradient algorithm.

4. The system according to claim 1 or 2, wherein: The medical image includes a chest X-ray.

5. The system according to claim 1 or 2, wherein: The ground truth structured label, the other structured label, and the predicted structured label include a label for assigning a specific diagnosis billing code.

6. The system according to claim 1, wherein: The medical condition includes at least one of the following: airspace density, pulmonary edema, pleural effusion, pneumothorax, cardiomegaly, nodule or mass, malplaced nasogastric tube, malplaced endotracheal tube, malplaced central venous catheter.

7. The system according to claim 1 or 2, wherein: The other input medical images are not associated with free-text medical reports.

8. A method for classifying a medical image based on the presence or absence of a medical condition, the method comprising: A natural language processor is trained on a first corpus of free-text medical reports, each of which has an associated medical image and one or more ground truth structured labels assigned by a medical expert, wherein: The computer vision model is configured to assign another structured label to an input medical image, the computer vision model comprising a feature extractor configured to generate a vector of extracted image features of the medical image and an image classifier configured to generate another structured label using the vector of extracted image features; and the natural language processor being configured to receive as input a free-text medical report and a vector of extracted image features generated by a feature extractor of the computer vision model, the natural language processor being trained during the training to generate predicted structured labels using the vector of extracted image features generated by the feature extractor in addition to the free-text medical report; The method further comprises: applying a natural language processor to a second corpus of free-text medical reports associated with medical images, and wherein the natural language processor generates another structured label for the associated medical images; and A computer vision model is trained using the medical image and the predicted structured labels generated by the natural language processor to assign another structured label to another input medical image using a feature extractor and an image classifier, wherein the ground truth structured label, the predicted structured label, and the another structured label include a label for the presence or absence of a medical condition.

9. The method according to claim 8, further comprising the steps of: An integrated gradient algorithm is applied to assign, to input words in the free-text medical report, attributions to predicted structured labels generated by the natural language processor.

10. The method of claim 8 or 9, further comprising displaying both the medical image and the attributions assigned by the integrated gradient algorithm on a workstation.

11. The method according to claim 8 or 9, wherein: The medical image includes a chest X-ray.

12. The method according to claim 8 or 9, wherein: The ground truth structured label, the other structured label, and the predicted structured label include a label for assigning a specific diagnosis billing code.

13. The method according to claim 8, wherein: The medical condition includes at least one of the following: airspace density, pulmonary edema, pleural effusion, pneumothorax, cardiomegaly, nodule or mass, malplaced nasogastric tube, malplaced endotracheal tube, malplaced central venous catheter.

14. The method according to claim 8 or 9, wherein: The other input medical images are not associated with free-text medical reports.

15. The method according to claim 8 or 9, wherein: The natural language processor is trained in an integrated manner with the computer vision model.

Citation Information

Patent Citations

  • Image classification neural networks

    US10460211B2

  • Processing images using deep neural networks

    US20160063359A1

  • Method for extracting concepts in Chinese electronic medical record based on deep learning

    CN106484674A

  • Recurrent neural feedback model for automated image annotation

    WO2017151757A1