Method and device for processing medical text data and medical image data

The apparatus and method efficiently match medical image abnormalities with text entities using an integrated processing unit, addressing the challenge of costly manual annotation in CAD training by leveraging radiology reports for quicker and more effective dataset preparation.

JP2025141829APending Publication Date: 2025-09-29CANON MEDICAL SYST CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025029773
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-13
Filing Date
2025-02-27
Publication Date
2025-09-29

AI Technical Summary

Technical Problem

Annotating medical image datasets for training computer-aided detection (CAD) algorithms is time-consuming and expensive, requiring expert tools and medical expertise, while existing data platforms provide large datasets but lack efficient methods for segmentation annotation.

Method used

An apparatus and method utilizing an image data processing unit, text data processing unit, and data matching unit to identify abnormalities in medical images and entities in text data, performing a matching process to obtain correspondence and generate training data efficiently, leveraging available radiology reports for quicker annotation.

Benefits of technology

Enables efficient construction of CAD systems by automating or semi-automating the extraction of training segmentations, reducing the time and cost associated with manual annotation, and facilitating the use of clinically recognized features for anomaly and entity detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025141829000001_ABST
    Figure 2025141829000001_ABST
Patent Text Reader

Abstract

To specify a correspondence between information included in medical image data and information included in medical text data easily.SOLUTION: A device includes an image data processing unit, a text data processing unit, and a data collation unit. The image data processing unit acquires image data indicating a medical image, and specifies an abnormal part by applying the acquired image data to a first model. The text data processing unit acquires medical text data corresponding to the medical image and specifies an entity and an associated attribute associated with the entity by applying the acquired medical text data to a second model. The data collation unit performs collation processing between the specified abnormal part and the specified entity on the basis of the property of the specified abnormal part and the associated attribute associated with the specified entity, and acquires coincidence data based on the correspondence between the specified abnormal part and the specified entity.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION The embodiments described herein generally relate to methods and apparatus for processing medical text and medical image data. [Background technology]

[0002] Machine learning algorithms, such as Computer Aided Detection (CAD) algorithms, can require large amounts of annotated or labeled training datasets. Currently, available data platforms provide access to large datasets. However, annotating such datasets for use as training data can be time-consuming and expensive, requiring expert tools and medical expertise. Platforms that provide access to large datasets are available. However, obtaining the segmentation annotations (at the pixel or voxel level) necessary to train CAD algorithms can be time-consuming and expensive, requiring expert tools and medical expertise. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] US Patent Application Publication No. 2020357117 [Patent Document 2] U.S. Patent No. 10,755,413 [Patent Document 3] International Publication No. 2021122267 Summary of the Invention [Problem to be solved by the invention]

[0004] One of the problems to be solved by the embodiments disclosed in this specification and the drawings is to easily identify the correspondence between information contained in medical image data and information contained in medical text data. However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problem. Problems corresponding to the effects of each configuration shown in the embodiments described below can also be positioned as other problems. [Means for solving the problem]

[0005] An apparatus according to an embodiment includes an image data processing unit, a text data processing unit, and a data matching unit. The image data processing unit acquires image data representing a medical image and identifies an abnormal portion by applying the acquired image data to a first model. The text data processing unit acquires medical text data corresponding to the medical image and identifies an entity and related attributes associated with the entity by applying the acquired medical text data to a second model. The data matching unit performs a matching process between the identified abnormal portion and the identified entity based on the properties of the identified abnormal portion and the related attributes associated with the identified entity, and obtains matching data based on the correspondence between the identified abnormal portion and the identified entity. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 is a schematic diagram of an apparatus according to an embodiment. [Figure 2] FIG. 2 is a flow chart outlining an image and text processing process according to an embodiment. [Figure 3] FIG. 3 shows an overview of a method according to an embodiment. [Figure 4] FIG. 4 shows the abnormal part detection process. [Figure 5] Figure 5 shows entities, specifically findings and impressions. [Figure 6] FIG. 6 is an exemplary user interface for evaluating the output of a method according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0007] One embodiment provides an apparatus including an image data processing unit, a text data processing unit, and a data matching unit. The image data processing unit acquires image data representing a medical image and identifies an abnormal portion by applying the acquired image data to a first model. The text data processing unit acquires medical text data corresponding to the medical image and identifies entities and related attributes associated with the entities by applying the acquired medical text data to a second model. The data matching unit performs a matching process between the identified abnormal portion and the identified entity based on properties of the identified abnormal portion and related attributes associated with the identified entity, and obtains matching data based on the correspondence between the identified abnormal portion and the identified entity.

[0008] An embodiment provides a method for processing medical text data and image data, the method including the steps of acquiring image data representing a medical image and identifying an abnormality by applying the acquired image data to a first model, acquiring medical text data corresponding to the medical image and identifying entities and associated attributes associated with the entities by applying the acquired medical text data to a second model, and performing a matching process between the identified abnormality and the identified entity based on properties of the identified abnormality and the associated attributes associated with the identified entity to obtain matched data based on the correspondence between the identified abnormality and the identified entity.

[0009] The following embodiments relate to methods for obtaining or generating matching data for the purpose of further training machine learning models, such as computer-aided diagnosis models. In some embodiments, the methods include automatically or at least semi-automatically extracting training segmentations using a combination of heatmap-based detection methods and text data processing. In specific embodiments, the methods leverage available expert text descriptions in radiology reports. Compared to images, radiology reports may be quicker to annotate and train predictive models. Reviewing images may be slower due to the time required to load them into a viewer and the viewing time required to understand the content, requiring expert analysis by a radiologist. Additionally, creating pixel-level annotations may be particularly slow.

[0010] The following embodiments relate to separately extracting a set of abnormalities and entities from medical image data, and performing a matching process between the extracted image portions and the entities using a clinically recognized feature set to obtain training data. This may enable efficient construction of a CAD system from a large imaging dataset. In some embodiments, given that a large dataset of medical images and associated text reports contains examples of pathologies to be detected, the method includes performing anomaly and / or abnormality detection on the medical image data and entity detection on the associated text reports to obtain representations of the anomalies and entities, and performing further matching processes to match the identified anomalies / abnormalities with the identified entities.

[0011] Identifying abnormalities in medical images is performed using one or more models. These models may include artificial neural networks or other machine learning-based models. The identification process may include, for example, detecting the presence of abnormal medical image data and / or determining the presence and / or location of one or more anomalies in the medical image data, e.g., indicative of or associated with pathology. Such abnormalities may include anomalous samples that are outliers compared to a normal distribution. For example, attention is focused on local anomalies corresponding to anomalous regions associated with or indicative of pathology. These include, but are not limited to, tumors, hemorrhages, and cysts. Depending on the model used and the training of the model, other anomalies in the medical image data may be detected. These may include, for example, artifacts such as streak artifacts or motion artifacts, and / or implants such as artificial eyes, metal pins, and / or abnormal tissue.

[0012] Additionally, the method includes identifying entities within text data, such as electronic medical records, using one or more models. These models include artificial neural networks or other machine learning-based models. Generally, the text data may correspond to or be derived from any suitable free text or unstructured or at least semi-structured documents. In the embodiments described below, the text data corresponds to medical text data obtained from radiology reports corresponding to one or more medical images. The text data may be the output of a medical professional or other human review of the medical images and may include the medical professional's observations of the images. In this context, entities may correspond to the medical professional's findings and / or impressions and / or any observations of real-world objects or anomalies within the images. A finding may be understood as an observation in the medical image, and an impression may be understood as a diagnosis based on the finding. In some embodiments, the identified entities may include pairings of findings and impressions from the findings. Such entities may correspond to diseases, symptoms, medications, and / or other observables or observed properties of the images and may be associated with pathology. The entities may be findings and impressions reported by a radiologist.

[0013] In a medical context, the text to be analyzed may be a clinician's text note. Clinical text notes may be stored in Electronic Medical Records. Clinical text notes may be free-text radiology reports. The text may be analyzed to obtain information about, for example, a medical condition or type of treatment.

[0014] A data processing device 10, also referred to as device 10 according to an embodiment, is shown schematically in Figure 1. In this embodiment, the data processing device 10 is configured to process medical imaging data and to process medical text data. In other embodiments, the data processing device 10 may be configured to process any other suitable data, for example any suitable non-medical image or text data.

[0015] Data processing apparatus 10 includes a computing device 12, which in this example is a personal computer (PC) or workstation. Computing device 12 includes a processing unit 14. Computing device 12 is connected to a display screen 16, or other display device, and one or more input devices 18, such as a computer keyboard and mouse. Display screen 16 is an example of a display.

[0016] Computing device 12 is configured to acquire an image dataset from a first data store, also referred to as medical image data store 20. The image dataset may be generated by processing data acquired by a scanner 22 and stored in data store 20.

[0017] The scanner 22 is configured to generate medical imaging data, which may comprise two-dimensional, three-dimensional, or four-dimensional data from any imaging modality. For example, the scanner 22 may comprise a magnetic resonance (MR or magnetic resonance imaging (MRI)) scanner, a computed tomography (CT) scanner, a cone-beam CT scanner, an x-ray scanner, an ultrasound scanner, a positron emission tomography (PET) scanner, or a single photon emission computed tomography (SPECT) scanner. The medical imaging data may comprise or be associated with additional condition data, which may include, for example, non-imaging data. The image data may be 1D, 2D, 3D, or 4D data, and / or the medical image data includes at least one of medical imaging data acquired using CT, MRI, fluoroscopy, ultrasound data or other modalities, ECG data or other medical measurement data, volumetric or slice data, and / or time series data.

[0018] Computing device 12 may receive medical image data or other data from one or more additional data stores (not shown) instead of or in addition to data store 20. For example, computing device 12 may receive medical image data from one or more remote data stores (not shown) that may form part of a Picture Archiving and Communication System (PACS) or other information system.

[0019] The computing device 12 receives the medical text data from a second data store, also referred to as a medical data text store 24. In alternative embodiments, the computing device 12 receives the medical text data from one or more additional data stores (not shown) instead of or in addition to the data store 24. For example, the computing device 12 may receive the medical text data from one or more remote data stores (not shown) that may form part of an Electronic Medical Records (EMR) or Electronic Health Records (EHR) system or a Picture Archiving and Communication System (PACS). In some embodiments, the computing device 12 receives the medical image data and the medical text data from a single data store.

[0020] The computing device 12 provides processing resources 14 or processing circuitry for automatically or semi-automatically processing medical image data and medical text data, for example, to obtain training data.

[0021] The processing device 14 comprises an image data processing circuit 102 for obtaining identified anomalous regions in one or more images, a text data processing circuit 104 for obtaining one or more entities from text data, a data matching circuit 106 for performing a matching process using the identified anomalous regions and the identified entities to synthesize matched data, a training circuit 108 for training at least one further machine learning model using the matched data, and a further image processing circuit 110 for applying the trained model(s) to unseen image data. The image data processing circuit 102 is an example of an image data processing unit. The text data processing circuit 104 is an example of a text data processing unit. The data matching circuit 106 is an example of a data matching unit. The training circuit 108 is an example of a training unit. The image processing circuit 110 is an example of an image processing unit.

[0022] In this embodiment, the circuits 102, 104, 106, 108, and 110 are each implemented on the computing device 12 by a computer program having computer-readable instructions executable to perform the method of the embodiment. However, in other embodiments, the various circuits may be implemented as one or more Application Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs). Specifically, matching, training, and further application of the trained model may be performed by different processing resources. Generally, the embodiments described below describe the use of a model, and it is understood that using a model may include applying the model to input data to generate output data. Using a model may include providing input data in a predetermined data structure to a rule set or algorithm and outputting the output data. The model may be predetermined by a separate training process. The predetermined model may be a partially trained or pre-trained model.

[0023] Computing device 12 also has a hard drive and other components of a PC, including RAM, ROM, a data bus, an operating system including various device drivers, and hardware devices including a graphics card, although such components are not shown in FIG.

[0024] The data processing apparatus 10 of FIG. 1 performs the methods shown and / or described below, such as the methods described with reference to FIGS.

[0025] The device 10 of Figure 1 performs the method described below. In other embodiments, different devices may be used to perform different processes or portions of the processes described below. For example, a first device may be used to perform image analysis and anomaly detection, a second device may be used to perform entity identification, a third device may be used to perform matching and generate additional training data, a fourth device may be used to train at least one additional model on the generated training data, and a fifth device may be used to apply the at least one additional model to new image data. In this manner, any suitable combination of devices may be used.

[0026] 2 illustrates a method 200 for obtaining consensus data. The method may be a method for obtaining data for training, refining, and / or tuning a computer-aided diagnostic algorithm. The method may include steps of allocating data and selecting the allocated data. The output consensus data relates to annotated segmentation data that can be used to train further models.

[0027] In step 202, image data and associated text data are obtained. In the embodiment of Figure 2, the image data includes medical image data in the form of CT scan data. In the embodiment of Figure 2, the text data corresponds to medical text data obtained from a radiology report corresponding to one or more medical images, although as noted above, the text data may be any suitable medical text data. Medical text data typically includes descriptive information, e.g., medically and / or clinically relevant information, and the entity detection process includes extracting the descriptive information and representing the information in entity representations. The text data represents or is derived from a radiology report.

[0028] In step 204, an anomaly detection process is performed on the image data using the first trained model, for example by the image data processing circuit 102. The anomaly detection process is a pixel-wise anomaly detection process. The anomaly detection method uses n pixels extracted from I images. ,a,I As described below with reference to FIG. 3, anomaly candidates and / or anomalous regions may be extracted from the anomaly heatmap at different intensity thresholds.

[0029] In step 206, an entity detection process is performed on the text data using the second trained model, for example by the text data processing circuit 104. The text data corresponds to text reports. The entity detection process is performed on n entities extracted from the corresponding i text reports. e,i This results in a set of entity candidates. It is understood that for each image-report pairing, the number of extracted anomalous regions does not correspond to the number of extracted entities. As will be explained with reference to Figure 4, the entity detection process yields entities and a set of attributes corresponding to each entity.

[0030] In step 208, a matching process is performed between the first and second representations, for example by the data matching circuitry 106. In this embodiment, the matching process is modeled as a combinatorial optimization assignment problem to find the best fitted match between entities and anomalies. In this embodiment, the best fit corresponds to the lowest cost of an entity-anomaly cost function. Therefore, the matching process includes minimizing such a cost function.

[0031] The entity-anomaly cost function can be constructed using quantities, such as characteristics of the candidate anomaly ROI (intensity, texture, shape), location of the image ROI (e.g., registering images to an atlas to have a common coordinate system & mapping of anatomical regions), absolute anomaly score of the image ROI, attributes of the entity (anatomical location, severity, chronicity, size, certainty), radiologist who wrote the report (e.g., different radiologists may provide different levels of detail and content in their reports). In some embodiments, any suitable assignment algorithm may be used to calculate the best match, for example, the Jonker-Volegenant algorithm may provide benefits. The matching process is described in more detail below.

[0032] In some embodiments, the cost function is based on further information determined from the medical text data. For example, author information or other quality measures of the report may be used. Other metadata associated with the text may be used. The author itself may be considered a category, or the author's seniority may be used as relevant information. For example, junior radiologist, consultant radiologist, specialist consultant radiologist, e.g., neuroradiologist, may be used. As a non-limiting example, these categories may be represented as one-hot vectors. In some embodiments, the cost function has a term that rewards text expressions derived from the same radiologist or author.

[0033] A matching process is described below, although it is understood that alternative matching processes may be performed. Generally, the matching process involves obtaining a mapping between captured abnormal image portions and entities based on one or more properties derived from the image and attributes extracted from the text. The matching process involves pairwise (or, in some cases, groupwise) evaluating the match between each detected abnormal portion and entity and assigning a score representing the match between the anomaly representation and the entity representation. The score may be referred to as a matching score. Thus, each pairing (or grouping) has a resulting matching score. The process includes selecting one or more pairs (or groups) based on the scores. In some embodiments, the selection is based on the highest matching score. In some embodiments, the matching is a combinatorial optimization assignment problem, involving a group or pairwise assignment process that finds the best match between the identified anomalies and entities. The properties and attributes used in the matching process are clinically informed.

[0034] In this embodiment, performing the matching process includes creating at least a first representation, e.g., a feature vector, for each identified anomaly and creating at least a second representation, e.g., a further feature vector, for each identified entity. The matching process includes evaluating the two representations and obtaining a score representing their degree of agreement. Matching of the identified anomalies to the entities is based on the obtained score. In some embodiments, the representations may be combined when calculating the cost function, e.g., the feature vectors may be combined into a single representation, such as a matrix or other multidimensional array.

[0035] Feature vectors are created based on the properties of the identified parts and the corresponding attributes of the anomalous regions; for example, the feature vector of a first anomalous region incorporates the properties of the anomalous region. The feature vector of an entity incorporates the attributes of the entity. By constructing feature vectors in the form of image properties that correspond to, or at least relate to, the attributes of the entity, matches can be calculated using other mathematical operations, such as cosine similarity or dot-product-based functions. In some embodiments, the properties selected for the image part feature vector have a one-to-one correspondence with the properties selected for the entity feature vector. Thus, the matching process includes obtaining pairwise (or group-wise) matches between each identified region and each identified entity based on the corresponding properties and attributes, respectively, and selecting the closest matching pair.

[0036] In step 210, the match data including the paired or assigned matches are stored together as training data. The match data provides training data in which image portions assigned to entities provide segmentation ground truth for further model training. The training data provides anomalous regions / anomalies as segmentation ground truth for the class of matching entities.

[0037] In step 212, an evaluation step is performed on the matches. The evaluation step may include human input. Poor / failed matches may be discarded or submitted for review / annotation by a human expert. The evaluation step may be a separate step to generate training data and / or verify results, or may form part of the training in step 214. An exemplary evaluation step is described with reference to FIG. 6.

[0038] In step 214, one or more further models are trained using at least the consensus data, e.g., by training circuitry 108 or further model refining circuitry. In this embodiment, computer-aided diagnostic models are trained. In some embodiments, the consensus data is provided as input to a training data generation procedure for generating synthetic training data.

[0039] In step 216, an optional step of retraining and / or refining the anomaly detection algorithm is performed, for example, by training circuitry 108. In such a step, the agreement data or other output data may be used to refine the parameters of the model used in step 204 as part of the semi-supervised learning or training process.

[0040] FIG. 3 illustrates in diagram form the data flow and data processing steps corresponding to method 200. As described with reference to step 202, image data 302 representing an input medical image and text data 304 in the form of a report corresponding to the input medical image are obtained. The image data 302 represents the input medical image of a region of interest, and the text data corresponds to corresponding medical text data regarding the medical image of the region of interest. The image data 302 is provided as input to an anomaly detection model 306, which in this embodiment is an artificial neural network, which performs entity detection processing as described with reference to step 204. The output of the anomaly detection processing corresponds to identified abnormal portions of the medical image that correspond to potential anomalies. In this example, the abnormal portions correspond to a first abnormal portion 310a related to the posterior outer edge of the brain scan region, a second abnormal portion 310b on the left hemisphere of the brain, and a third abnormal portion 310c on the left side of the brain region.

[0041] The text data 304 is provided as input to a separate entity detection model 308, which performs entity detection processing as described with reference to step 206. In this embodiment, the entity detection model 308 is an artificial neural network. The output of the entity detection model 308 is represented by a first identified entity 314a corresponding to "hypotenchyma," a second identified entity 314b corresponding to "infarction," and a third identified entity 316c corresponding to "sulcal loss." It will be appreciated that the hypoattenuation and infarction may be identified as a single entity, which in some embodiments may correspond to a finding-impression pairing.

[0042] The anomaly detection model 306 and the entity detection model 308 are independent and predetermined by training on training data. Models 306 and 308 represent multi-layer artificial neural networks that are trained by optimizing the parameters of the neural network by minimizing the difference between the predicted and measured quantities using a suitable technique, such as minimizing an error or cost function. It will be appreciated that other neural network or machine learning based optimization methods may be used.

[0043] A matching process is then performed on the outputs of the two models according to an embodiment. Specifically, the matching process is performed between one or more identified anomalous portions identified by the anomaly detection process and one or more identified entities identified by the entity identification process. As generally indicated by reference numeral 316, a first anomalous region matches a second identified entity, and a second anomalous region matches the first and / or second identified entities. A third anomalous region does not match any of the identified regions. The matching process is performed as described with reference to step 208. Generally, the matching process is performed based on properties of the identified anomalous portions and attributes of the identified entities.

[0044] The result of the matching process is match data represented at 318. The match data represents labeled or annotated images suitable for use in training one or more further models, such as computer-aided diagnostic models, and thus represents a grouping of one or more identified anomalies with one or more identified entities.

[0045] In some embodiments, the method includes training further model(s). The output may be stored as described with reference to step 210, or may be further evaluated as described with reference to step 212, or may be used to train one or more models as described with reference to step 214.

[0046] An optional model refinement step (corresponding to step 216) is indicated by reference numeral 320, where the anomaly detection algorithm is retrained and / or refined using supervision from the matched anomaly pseudo-labels.

[0047] In the method of Figures 2 and 3, an abnormality or anomaly detection image processing method is described in which a medical image is provided as input to a predetermined neural network or other suitable model, which outputs a number of identified abnormalities.

[0048] A non-limiting example of a model for abnormality detection or anomaly detection is a denoising autoencoder, such as the one described by O'Neil et al., "Denoising Autoencoders for Unsupervised Anomaly Detection in Brain MRI," Proceedings of the 5th International Conference on Medical Imaging with Deep Learning, the contents of which are incorporated herein by reference. Such an example has a U-Net structure (as described in Section 4.3 and Appendix A) and is trained on a dataset of MRI scans (four sequences: T1, T1Gd, T2, and FLAIR) of healthy brains.

[0049] The encoder-decoder architecture has three downsampling / upsampling stages. Each encoder stage consists of two weight-standardized convolutions with a convolution kernel size of 3. Following Swish activation and group normalization, the three stages have 64, 128, and 256 output channels, respectively. Average 2x2 pooling is used for downsampling. The decoder architecture is the inverse of the encoder and uses transposed convolution layers for upsampling.

[0050] In the above-described methods, a deep learning model, such as a denoising autoencoder and encoder model, is applied to output an anomaly score for each trained pixel. Some models compare the image to a normal distribution. Each pixel / voxel in the heatmap has an anomaly score whose value represents or indicates the probability that the pixel or voxel is anomalous.

[0051] The method is a heatmap-based approach that involves generating a pixel-level heatmap based on the difference between an image and a normal distribution.

[0052] In this way, by applying the model, it is possible to identify abnormal areas based solely on their spatial distribution, without classification or labeling.

[0053] In some embodiments, applying the model includes comparing the image data to a mean or normal spatial distribution of the region of interest learned from training data of healthy subjects, and identifying abnormal regions based on the comparison, hi some embodiments, each pixel / voxel has a value that represents its degree of difference from the learned normal distribution.

[0054] In some embodiments, many different threshold levels are applied to identify anomalous regions. It will be appreciated that a high threshold will capture only those regions most likely to be anomalies. For example, if a high threshold results in two anomalous regions, but three entities are extracted from the corresponding radiology report, the threshold for identifying anomalous regions may be lowered to identify weaker anomaly candidates.

[0055] Identifying anomalous medical image data may include detecting the presence of anomalous medical / image data and / or determining the presence and / or location of one or more anomalies within the medical image data, e.g., indicative of or associated with pathology. As such, any suitable anomaly detection algorithm may be used. In particular, unsupervised heatmap-based algorithms, such as those described in "Anomaly Detection via Context and Local Feature Matching," 2022 IEEE 19th International Symposium on Biomedical Imaging (ISBI), IEEE, are suitable, the contents of which are incorporated herein by reference. Such methods may be extended to use human annotations used to perform semi-supervised learning.

[0056] Figure 4 shows an input medical image 402 corresponding to a CT scan of the brain. Figure 4 also shows an intermediate heatmap of the image 404. Figure 4 also shows the output of the abnormal region identification model, in which five abnormal regions have been identified. These include a first region 414 at a first high threshold, three regions 408, 410, and 412 at a second threshold, and an additional region 406 at a lower threshold. In practice, a high threshold is applied first, followed by subsequent application of lower thresholds to identify additional regions.

[0057] Thresholding can be understood as selecting all pixels / voxels with an anomaly score above a selected threshold score as positive detections. In addition, morphology or connected components may be applied to the results. Morphology can be understood as performing spatial operations such as dilation, erosion, opening, and closing to, for example, remove random single-pixel detections or single-pixel holes in the detected region, smooth "rough" edges, and make the detected region more spatially coherent. Additionally, connected components may be identified by identifying regions of connected (i.e., adjacent) pixels / voxels and discarding small regions where the number of connected components (pixels / voxels) is less than a selected minimum.

[0058] In this embodiment, entities generally refer to real-world objects such as findings and / or impressions. The entities may correspond to pathologies and may relate to the effects of a disease, a medical condition, or the effects of a treatment. The text extraction described above is performed using any suitable text extraction algorithm, such as the text extraction algorithm described in Schrempf et al., "Templated text synthesis for expert-guided multi-label extraction from radiology reports," 2021 Machine Learning and Knowledge Extraction, 3(2), pp. 299-317. The method uses the neural network transformation model PubMedBERT as a base model. Figure 5 is a schematic diagram of 33 label sets of entities related to neurological abnormalities from head CT reports of stroke patients. The model was trained on a dataset containing 28,687 radiology reports provided by the West of Scotland Safe Haven within NHS Greater Glasgow and Clyde (GGC). Figure 5 shows 13 radiological findings (labeled 502), 16 clinical impressions (labeled 504), and four crossover labels indicated by an asterisk. The links from findings to impressions are shown diagrammatically in Figure 5. Many labels fit both the finding and impression categories and are labeled with an asterisk. Also, example pathologies associated with the findings and citations are labeled 506.

[0059] In addition to extracting entities from the text data, the method also includes extracting attributes of the entities from the text. Specifically, in addition to the deep learning model that identifies entities, a further deep learning model is applied to obtain entity attributes associated with the identified entities. Such attributes may relate to the location or appearance of the entities, for example. In some embodiments, obtaining the entities and their attributes is a two-step process, where the text data is applied as input to a first trained model, which extracts the entities. The entities are then applied as input to a second model along with the text, and the output from the second model is a set of attribute values.

[0060] As a non-limiting example, one embodiment chains two BERT (Bidirectional Encoder Representations from Transformers) models. The BERT models are described by Devlin et al. in "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding."

[0061] The model allows the following attributes to be identified and annotated in the text data for each pathology: The attributes may relate to polarity and / or certainty, e.g., one of {positive, negative, uncertain}, laterality, e.g., one of {left, right, bidirectional}, chronicity, e.g., one of {acute, subacute, chronic, acute-on-chronic}, severity, e.g., one of {mild, moderate, severe}, or change in severity, e.g., one of {unchanged, alleviated, worsening}, anatomical location, e.g., tissue (one of {white matter, gray matter, region, bone, vessel, dura}) or region (lesion) ({extracranial, intracranial [extraaxial[...], intraaxial[...]]}) or artery (vascular lesion) (one of {protocerebrum, middle cerebrum [M1, M2, M3, M4], ...}), anatomical distribution (e.g., one of {regional, global}). The above list of attributes is given as a non-limiting example. At annotation time, the label is represented in text as a human-readable name.

[0062] In one example, an entity identification model is trained to obtain predictions for a given set of labels for each sentence in a radiology report and / or for the entire radiology report. Each label relates to a corresponding entity, in this embodiment, a finding and / or an impression. For example, such labels may include hemorrhage and tumor. A deep learning model may be trained to classify each sentence or the entire report to indicate whether hemorrhage and tumor are present, respectively. For example, for each sentence, each label is classified into one of multiple certainty classes. The certainty classes include positive, negative, and uncertain. If the model determines from the sentence that the finding or impression represented by the label is present, the classification is in the positive certainty class. If the model determines from the sentence that the finding or impression represented by the label is not present, the classification is in the negative certainty class. If the model determines from the sentence that the model is uncertain about the presence of the finding or impression represented by the label, the classification is in the uncertain certainty class. For example, a sentence may suggest that an observation or impression exists but is not strong enough to be classified as positive.

[0063] In the method 200 described above, a matching process has been described. A further example of the matching process is given below. In further detail, a possible injective matching function m:a that matches an anomaly 'a' to an entity 'e' mentioned in the report in each image 'i' is given. i →e i For each, we can calculate a cost function that we define. The cost function may include many terms. Below, we consider terms that relate to or represent the goodness or closeness of the entity-anomaly match and the intra-anomaly consistency of the anomaly group. In some embodiments, the cost function represents or is proportional to the degree of match between the properties of the anomaly and the attributes of the entity. The cost function may also be referred to as a matching function.

[0064] A first example of a suitable cost function term is an entity-anomaly matching cost function that uses cosine similarity to evaluate whether matched entities and anomalies have similar properties. Thus, the matching process involves calculating a cost function value for each pairing of an anomaly with an identified entity, and determining the pair with the lowest cost function value and / or one or more pairs with a value below a predetermined threshold. The cost function is defined as follows:

[0065]

number

[0066] The cost function considers potential matches across all pairs. During the matching process, we are interested in identifying the pair combination with the lowest total cost. We calculate a score for each pairing of an entity and anomaly and sum them up. The cost function then outputs a value that represents the similarity between the region and the entity.

[0067] feature vector f a and f mcorresponds to a first feature vector extracted from the image and a second feature vector extracted from the associated text. For example, the image feature vector represents properties of the anomaly, such as properties derived from one or more image processing steps for the abnormality. In this example, the anomaly feature vector includes entries corresponding to measurements of the anomaly region of interest (ROI). The entity feature vector corresponds to attributes detected in the text description, which may be mapped to numerical values.

[0068] As a non-limiting example, the feature vector f a The (anomaly feature vector) has the following entries obtained by measuring the attributes of the anomaly region of interest (ROI): (mean intensity of the anomaly ROI, X coordinate of the center of the anomaly ROI, size of the anomaly in voxels). The associated feature vector f for each entity is obtained by detecting attributes in the text description and mapping them to numerical values. m(a) has the following entries: {"low absorption"=0, "low absorption"=1}, {"left"=0, "bilateral"=0.5, "right"=1}, {"small"=0, "medium"=0.5, "large"=1}. In some embodiments, the feature vector is constructed such that there is a one-to-one correspondence between at least one property derived from the anomaly and at least one attribute of the identified entity used in the matching process. In some embodiments, the relationship between the entries of the anomaly feature vector and the entity feature vector is predetermined. In some embodiments, the relationship is learned through further machine learning processes.

[0069] The cost function may include additional terms to penalize high variance within a group of entity candidates during the training process (low variance means better consistency) and ensure that anomalies matched to the same entity type have consistent properties. For example, one may consider radiomics-style features extracted from intensity ROIs, such as intensity mean, intensity variance, size, Gray-level co-occurrence matrix (GLCM), and fractal features. These features may be the same as or different from the matching features depicted in the previous slide. -1 denotes the inverse of the matching function depicted in the previous slide (i.e., m -1 :a i →e i ), n f If we have a set of features f, we may add a difference term. More specifically, in such an example the cost function is:

[0070]

number

[0071] The feature f for which dissimilarity is penalized may be different for each entity e. In some embodiments, a feature reduction process is performed, e.g., using principal component analysis, to obtain a reduced / compressed feature set, and one or more portions of the cost function are applied to the reduced / compressed feature set. For example, the dissimilarity measure may be applied to the reduced features rather than the original features. The cost function may be used when the total number of features is n f and the total number of entities is n eThe second term iterates over all features and all entities and sums over all pairwise differences. Differences violate independence and require considering sets of pairs. Thus, the second term does not loop over entities & anomalies, but over features & entities (and computes them as part of matching for all anomalies assigned to a given entity label). Therefore, the above cost function has a first term that outputs a value representing the similarity between a region and an entity, and a second term that penalizes changes in anomalous image parts assigned to the same class of entity.

[0072] In embodiments where some ground truth is available, for example in situations where a known set of anomaly-entity matches is available, rather than measuring simple difference, one may use the distance from the feature vector of the known true anomaly to the feature value of the assigned anomaly, as described above.

[0073] The method may apply an optimization process to find the minimum (and therefore best) cost mapping from entities to anomalies. The optimization process may be based, for example, on the Jonker-Volegenant algorithm applied to solve as a multiple linear assignment problem.

[0074] FIG. 6 shows an exemplary evaluation step where user input is solicited. FIG. 6 is presented as an illustrative example and shows two proposed matches based on the data described with reference to FIG. 3. A graphical representation 602 of the two proposed matches is presented to the user, e.g., on the display screen 16, along with an interactive element for confirming or rejecting the determined matches. In this example, the interactive element is a tick box. The exemplary user interface allows the user to quickly review the match data, e.g., via the input device 18. In some embodiments, the displayed match data includes a grouping of the identified anomaly(s) and entities and receives further user input representing the user's evaluation. The user input may represent a confirmation of the match or a rejection of the match. The evaluation step may form part of a training method, which provides the radiologist with an opportunity to view and edit the matches between image regions and text passages (as well as the matching features). This allows domain expert controllability at the voxel and / or region level without requiring full-scale voxel-wise annotation.

[0075] Although the above embodiments have been described with respect to medical data, in other embodiments, the above methods may be used to process any text and / or image data.

[0076] In the embodiment described above, the entity detection model is an unsupervised anomaly detection model, with the idea being to learn something that visually approximates a normal distribution and then detect "outliers" from that distribution at test time. However, in further embodiments, the method may be implemented using an anomaly detection method trained on a dataset containing both normal and abnormal images, and / or potentially trained on an unlabeled dataset of all images, where the method learns to find outliers in the training data distribution. Because there are likely not enough labeled anomalies to train a fully supervised algorithm to detect specific anomalies (e.g., labeled as hemorrhage, tumor, etc.), we want to leverage information from the associated radiology reports to assign (match) these labels.

[0077] One embodiment provides a method for detecting radiological findings. The method may include a semi-supervised anomaly detection method for medical imaging data that may be trained using a dataset of normal and abnormal images to generate a heat map indicating the degree of anomaly at each pixel. The method may include extracting regions of interest from the anomaly heat map. The method may include an entity recognition method for a corresponding textual radiology report. The method may include a method for extracting attributes (e.g., laterality, size, severity) of each entity in the textual radiology report. The method may include creating a representation of the imaging anomaly. The method may include creating a representation of the textual entity. The method may include a matching function, or "cost function," that quantifies the quality of the match between each anomaly / entity representation pair. The cost function may also be referred to as a matching function. The method may include an optimizer that finds the minimum cost of (best) mapping from entities to anomalies. The optimizer may be based, for example, on the Jonker-Volegenant algorithm applied to solve multiple linear assignment problems. Laterality, size, and severity are examples of mathematical representations of attributes of entities. A cost function is an example of a similarity function in an embodiment.

[0078] The function for may be the cosine similarity between the image and text vector representations, which may be constructed such that equal elements exhibit equal properties.

[0079] A cost term in the matching function may penalize high variations in the image representation of all anomalies assigned to the same entity class.

[0080] A cost term in the matching function may penalize high variations in the image representation of all anomalies that have similar corresponding text representations.

[0081] A cost term in the matching function may reward low variance of image representations of all anomalies with similar corresponding text representations, text representations derived from the same radiologist but not from different radiologists.

[0082] Deep learning methods may be used to learn image representations. Deep learning models may be used to obtain anomalous portions of images and / or entities. Deep learning methods may be used to learn text representations. Other suitable image analysis and natural language processing methods may be used to obtain anomalous portions of images and text.

[0083] The method may include retraining the anomaly detection method using supervision from the matched anomaly as well. The method may include training on both healthy data and matched pathology data.

[0084] The method may include an evaluation step, for example, where a human user reviews the automatic matches and filters / selects correct matches before using the human-confirmed matches as additional supervision.

[0085] The method may include retraining the anomaly detection method using supervision from the matched anomaly as well. The method may include training on both healthy data and matched pathology data.

[0086] The method may include an evaluation step, for example, in which a human user reviews the automatic matches and filters / selects correct matches before using the human-confirmed matches for additional supervision, which may include displaying a user interface to the user via a display showing the matched pairs or other output.

[0087] Although particular circuits are described herein, in alternative embodiments, the functionality of one or more of these circuits may be provided by a single processing resource or other component, or the functionality provided by a single circuit may be provided by a combination of two or more processing resources or other components. A reference to a single circuit encompasses multiple components that provide the functionality of that circuit, whether or not such components are separate from one another. A reference to multiple circuits encompasses a single component that provides the functionality of those circuits.

[0088] While certain embodiments have been described, these embodiments are presented for illustrative purposes only and are not intended to limit the scope of the invention. Indeed, the novel methods and systems described herein may be embodied in a variety of other forms. Furthermore, various omissions, substitutions, and changes in the form of the methods and systems described herein may be made without departing from the spirit of the invention. The appended claims and their equivalents are intended to cover such forms and modifications as fall within the scope of the invention. [Explanation of symbols]

[0089] 10 Data processing device (device) 12 Computing Devices 14 Processing Unit (Processing Resources) 16 display screens 18 Input Devices 20 Data storage unit (medical image data storage unit) 22 Scanner 24 Data storage unit (medical data text storage unit) 102 Image data processing circuit 104 Text data processing circuit 106 Data matching circuit 108 Training Circuit 110 Image processing circuit 200 ways 302 Image data 304 Text Data 306 Anomaly Detection Model 308 Entity Detection Model 310a 1st abnormal part 310b 2nd abnormal part 310c 3rd abnormal part 314a First Specific Entity 314b Second specific entity 316c Third Specified Entity 402 Input Medical Images 404 images 406,408,410,412 parts 414 Part 1 602 Graphical Representation

Claims

1. an image data processing unit that acquires image data representing a medical image and identifies an abnormal portion by applying the acquired image data to a first model; a text data processing unit that acquires medical text data corresponding to the medical image and applies the acquired medical text data to a second model to identify entities and related attributes associated with the entities; a data matching unit that performs a matching process between the identified abnormal part and the identified entity based on the property of the identified abnormal part and the related attribute associated with the identified entity, and obtains matching data based on a correspondence relationship between the identified abnormal part and the identified entity; A device comprising:

2. The matching process includes: creating a first representation representing at least one of the identified abnormal portions and a second representation representing at least one of the identified entities; obtaining a score representing a degree of match between the first representation and the second representation; and matching the abnormal portion to at least one of the identified entities based on the obtained score.

10. The apparatus of claim 1.

3. generating training data for at least one model using the matched data.

10. The apparatus of claim 1.

4. the image data processing unit identifies the abnormal portion based on a difference between the spatial distribution of the medical image and a predetermined normal distribution; 10. The apparatus of claim 1.

5. the image data processor identifies the abnormalities using thresholding, morphology, or a connected component-based pixel or voxel level approach; 10. The apparatus of claim 1.

6. the matching process includes determining a match based on at least one of a similarity and a consistency between a property of the anomaly and an attribute of the entity; 10. The apparatus of claim 1.

7. The matching process includes determining based on a similarity function or distance between a mathematical expression of the property of the anomaly and an attribute of the identified entity.

10. The apparatus of claim 1.

8. The matching function used in the matching process includes a term that represents the similarity between the identified abnormal portions and a term that penalizes differences between abnormal portions assigned to the same class of the entity.

10. The apparatus of claim 1.

9. The matching process includes performing an optimization process of a Jonker-Volegenant algorithm.

10. The apparatus of claim 1.

10. and further retraining or refining at least one of the first model and the second model using the obtained match data.

10. The apparatus of claim 1.

11. and further displaying the matched data on a display unit.

10. The apparatus of claim 1.

12. At least one of the first model and the second model comprises a deep learning or other artificial neural network based model.

10. The apparatus of claim 1.

13. the image data includes at least one of CT data, MRI data, fluoroscopy data, ultrasound data, ECG data, medical measurement data, volumetric data, slice data, and time series data; 10. The apparatus of claim 1.

14. the properties of the identified anomaly include at least one of anomaly intensity, texture, shape, location, and measurement of the anomaly; 10. The apparatus of claim 1.

15. the entities include at least one of observations, impressions, and observables; 10. The apparatus of claim 1.

16. acquiring image data representing a medical image and applying the acquired image data to a first model to identify an abnormality; obtaining medical text data corresponding to the medical image, and applying the obtained medical text data to a second model to identify entities and associated attributes associated with the entities; performing a matching process between the identified abnormal part and the identified entity based on the properties of the identified abnormal part and the related attributes associated with the identified entity, to obtain matching data based on a correspondence relationship between the identified abnormal part and the identified entity; a method for processing the medical text data and the image data, the method comprising:

Citation Information

Patent Citations

  • Method and system for medical imaging evaluation

    US10755413B1

  • Heat map generating system and methods for use therewith

    US20200357117A1

  • Temporal disease state comparison using multimodal data

    WO2021122267A1