Medical Image Processing
The method uses machine learning to generate and compare image embedding vectors for accurate and efficient medical image segmentation, addressing inefficiencies in existing methods by leveraging medical reports for training, thereby improving the detection of indistinct features.
Patent Information
- Application Number
- JP2024134713
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-08-10
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-08-10
AI Technical Summary
Existing medical image segmentation methods lack efficiency and accuracy in identifying and delineating features, particularly those with gradual transitions, and require labor-intensive manual labeling.
A computer-implemented method using machine learning systems to generate and compare image embedding vectors and feature vectors, enabling automated segmentation and classification of medical image features, including those with indistinct boundaries, by leveraging medical reports for training data.
Enhances the accuracy and efficiency of medical image segmentation by automatically identifying and delineating medical features, reducing manual labor, and improving the detection of abnormalities with blurred boundaries.
Smart Images

Figure 0007725674000001 
Figure 0007725674000002 
Figure 0007725674000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method, an apparatus and a computer program for segmenting medical image data. [Background technology]
[0002] Medical images can be segmented to identify regions of interest, such as organs, medical abnormalities, etc. Segmentation may, among other things, enable quantitative data to be obtained from medical images (e.g., the size of certain medical features), enable accurate planning of radiation treatments, and enable identification of features, such as medical abnormalities, that may not be readily apparent to medical professionals.
[0003] It would be desirable to provide an automated method for processing medical images to identify such features and their locations within the medical image, such as for segmentation purposes. Summary of the Invention
[0004] According to a first aspect of the present invention, there is provided a computer-implemented method for processing medical image data, the method comprising: receiving, in a first machine learning based system, medical image data representing a medical image; generating, based on the received medical image data, a plurality of image embedding vectors corresponding to each of a plurality of medical image features, each of the plurality of image embedding vectors comprising medical image feature data associated with a respective separate medical image feature and indicative of the presence or absence of each medical image feature at each of a plurality of locations within the medical image; receiving, in a second machine learning based system, an indication of a first medical image feature contained in the medical image; generating, by the second machine learning based system, a feature vector based on the indication; performing a comparison between the feature vector and the plurality of image embedding vectors; and identifying the first image embedding vector from the plurality of image embedding vectors based on the comparison.
[0005] Optionally, a first image embedding vector is identified from the plurality of image embedding vectors as having the highest similarity to the feature vector.
[0006] Optionally, the method includes determining position data representing a position within the medical image of the indicated first medical image feature within the medical image data based on the identified first image embedding vector.
[0007] Optionally, the medical image feature data represents a segmentation map for each medical image feature, and the determined location data comprises the segmentation map represented by the identified first image embedding vector.
[0008] Optionally, the segmentation map indicates the probability that each of the medical image features is present at each location within the medical image.
[0009] Optionally, the method includes determining whether the medical image represents a medical abnormality, and in response to determining that the medical image represents a medical abnormality, displaying the segmentation map on a display device.
[0010] Optionally, the indication received at the second machine learning based system includes region data representing a region of the first medical image feature within the medical image, and the comparison of the feature vector with the plurality of medical image embedding vectors is performed based at least in part on the region data.
[0011] Optionally, the indication received at the second machine learning based system includes data representing a text prompt indicative of the first medical image feature.
[0012] Optionally, the method includes performing a training process for training a first machine learning based system and a second machine learning based system, the training process including inputting first training data including a plurality of sets of medical image data into the first machine learning based system to generate a plurality of trial image embedding vectors for each set of medical image data; and inputting second training data into the second machine learning based system to generate a trial feature vector for each set of medical image data, the second training data including data representing a plurality of medical reports, each medical report including data indicating the presence of one of a plurality of medical image features in a corresponding one of the sets of medical image data, each medical report including data representing a region of the feature indicated to be present in a medical image represented by the corresponding set of medical image data; and the training process including training the first machine learning based system and the second machine learning based system together to minimize a loss function between the trial image embedding vectors and the corresponding trial feature vectors.
[0013] Optionally, the method includes inputting each of the medical reports into a natural language processing system to generate a set of data representing findings of the medical reports, wherein the data representing the plurality of medical reports is data representing findings of the medical reports.
[0014] Optionally, the method includes inputting at least one of the plurality of image embedding vectors into a third machine learning based system to generate natural language text describing findings related to the medical image data.
[0015] Optionally, the method includes performing a training method for training a third machine learning based system, the training method including inputting image embedding vectors generated by the first machine learning based system based on a set of input medical image data to the third machine learning based system to generate, for each set of input medical image data, trial natural language text describing findings related to the set of input medical image data, and training the third machine learning based system to minimize a loss function between the trial natural language text and data representing a medical report corresponding to the set of input medical image data.
[0016] Optionally, at least one of the plurality of medical image features comprises a medical abnormality.
[0017] According to a second aspect of the present invention there is provided an apparatus configured to carry out the method according to the first aspect.
[0018] According to a third aspect of the present invention, there is provided a computer program which, when executed by a computer, causes the computer to carry out the method according to the first aspect. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a flow chart that illustrates generally a method for processing medical image data. [Figure 2] 2 is a flow chart illustrating a schematic example of the method illustrated in FIG. 1; [Figure 3] 2 is a flow diagram that generally illustrates an example of a method used by the image encoder 204 to generate multiple image embedding vectors. [Figure 4] 2 is a flow diagram that schematically illustrates a method for training a machine learning-based system to perform the method illustrated in FIG. 1 . [Figure 5] 1 is a diagram illustrating a schematic diagram of an apparatus according to an example. DETAILED DESCRIPTION OF THE INVENTION
[0020] The word encompasses individuals of male, female, and other gender identities, regardless of grammatical usage.
[0021] FIG. 1 shows a flow diagram of a computer-implemented method for processing medical image data 202 .
[0022] In summary, the method comprises: In step 102, at a first machine learning based system, receiving medical image data 202 representing a medical image 202; In step 104, the first machine learning based system generates, based on the received medical image data 202, a plurality of image embedding vectors 206 corresponding to a plurality of medical image features, each of the plurality of image embedding vectors 206 being respectively associated with a different medical image feature and including medical image feature data indicative of the presence or absence of each medical image feature at each of a plurality of locations within the medical image 202; In step 106, receiving, at a second machine learning based system, an indication of a first medical image feature included in the medical image 202, and generating, by the second machine learning based system, a feature vector 212 based on the indication; Step 108 includes performing a comparison between the feature vector 212 and the plurality of image embedding vectors 206, and identifying a first image embedding vector from among the plurality of image embedding vectors 206 based on the comparison.
[0023] In this manner, the method enables the generation and identification of medical image feature data indicative of the presence or absence of each medical image feature at each of a plurality of locations within the medical image. Comparison of the feature vector with a plurality of image embedding vectors can generate an image embedding vector with content related to the medical image feature, as indicated in step 106.
[0024] The data provided by the identified image embedding vectors, indicating the presence or absence of each medical image feature at each location, can be used in a variety of ways. As described below, a segmentation map can be generated that indicates the probability that the indicated medical feature is present at each pixel in the medical image, which can also be used by medical professionals for the purposes described in the Background section. The identified image embedding vectors can be used to determine whether a medical image represents a medical abnormality. Furthermore, the identified image embedding vectors can be used to generate text that describes the findings in the medical image. Alternatively, image embedding vectors can be generated for multiple patients, each with a medical abnormality, and each medical abnormality can be classified by observing cluster patterns in the image embedding vectors.
[0025] Next, an example method 100a will be described in detail with reference to Figure 2. The method 100a of Figure 2 is a specific example of the method 100 of Figure 1 and includes each step of the method 100 of Figure 1.
[0026] As previously described, method 100 includes receiving medical image data 202 at a first machine learning based system at step 102. In method 100a, the first machine learning based system is an image encoder 204. Step 102 may include retrieving the medical image data 202 from storage, such as a memory (see, e.g., memory 504 in FIG. 5, described below).
[0027] The medical image data 202 may include a 3D array of elements, each having a value, that collectively represent a 3D medical image. The elements are voxels, and each voxel has at least one value. This at least one value may represent the output signal of the medical imaging technique used to generate the medical image data 202. For example, in magnetic resonance imaging, the value of an element (e.g., a voxel) represents the rate at which excited nuclei return to equilibrium in the region corresponding to the element. As another example, in a SPECT image, the value of an element represents the blood flow rate of the capillary represented by a given voxel. As another example, in a CT image, the value of an element corresponds to or represents X-ray attenuation. In some examples, each element has only one value, while in other examples, each element has multiple values or is otherwise associated with multiple values. For example, the multiple values of a given element represent respective values of multiple signal channels. For example, each signal channel may represent a separate medical image signal or characteristic of the imaged subject. In some examples, the at least one value includes an element (e.g., voxel) intensity value. For example, output signals from medical imaging are mapped to voxel intensity values, e.g., values within a defined intensity value range. For example, in the case of grayscale images, intensity values correspond to values ranging from 0 to 255, where, for example, 0 represents a "black" voxel and 255 represents a "white" voxel. As another example, in the case of USHORT medical image data, intensity values correspond to values ranging from 0 to 65536. As another example, in the case of color images (e.g., where different characteristics of the imaged object are represented by different colors), each pixel has three intensity values, e.g., one for each of the red, green, and blue channels. Of course, other values can be used. While the methods described herein primarily use 3D medical images, they are, of course, also applicable to 2D medical images.
[0028] As previously mentioned, at step 104, the method 100 includes an image encoder 204 that generates, based on the received medical image data 202, a plurality of image embedding vectors 206 corresponding to each of a plurality of medical image features.
[0029] Method 100a includes inputting medical image data 202 to an image encoder 204 to generate image embedding vectors 206. In method 100a, each of the plurality of medical image embedding vectors 206 is associated with a different medical image feature. In method 100a, each medical image feature is a different medical abnormality. The medical abnormalities in method 100a include ground-glass opacities and lesions. However, in some examples, the plurality of medical image embedding vectors 206 includes two image embedding vectors, one associated with normal tissue and the other associated with abnormal tissue. Furthermore, in some examples, the medical image feature is associated with multiple body parts that are not medical abnormalities. For example, an image embedding vector may be associated with a kidney, pancreas, or rib. The medical image features associated with the image embedding vectors 206 depend on the content of the training data used to train the image encoder 204, as described below with reference to FIG. 4. An example of a method used by the image encoder to generate image embedding vectors is described below with reference to FIG. 3.
[0030] Each of the plurality of medical image embedding vectors 206 includes medical image feature data indicating the presence or absence of a respective medical image feature at each of a plurality of locations within the medical image 202. In method 100a, the feature data represents a segmentation map 218 of the medical image feature with which the image embedding vector is associated, the segmentation map 218 indicating the probability that each medical image feature is present at each location within the medical image 202. In method 100a, each location is a portion 302 of the medical image 202. The first image embedding vector, for example, represents a segmentation (probability) map for a lesion. If the medical image 202 does not contain a lesion, the segmentation map 218 is likely to indicate a zero probability for all locations within the image. However, because the object depicted in the medical image 202 in this example has ground-glass opacity, the segmentation map 218 of the image embedding vector for ground-glass opacity indicates a non-zero probability for locations with ground-glass opacity. In method 100a, these probabilities are the entries of a predetermined medical image embedding vector, with each component of the predetermined medical image embedding vector corresponding to a different one of the image portions 302. However, in some examples, the medical image feature data is binary data indicating the presence or absence of the medical feature in each of the image portions, with a "1" indicating the presence of the feature and a "0" indicating the absence of the feature.
[0031] As previously described, in step 106, the method 100 includes receiving, at a second machine learning based system, an indication of a first medical image feature contained in the medical image 202, and generating, by the second machine learning based system, a feature vector 212 based on the indication.
[0032] In method 100a, the indication of the first medical image feature is a text prompt 208 that describes the first medical image feature. For example, a user, such as a radiologist, may input text, such as "ground glass opacity," into the device 500 with the second machine learning-based system via input interface 506. Alternatively, the user may select the first medical image feature from a drop-down box that lists multiple medical image features.
[0033] In method 100a, the indication includes region data representing a region of a first medical image feature in medical image 202. For example, a radiologist can input text such as "of the upper lobe of the right lung" via input interface 506. A single text prompt 208 can include both the region and the first medical image feature, e.g., the text prompt 208 is "ground-glass opacity of the upper lobe of the right lung." Alternatively, the radiologist can input the region data by clicking on a region of a template image of a human body (or a portion thereof) stored in apparatus 500 and displayed on display device 510; for example, apparatus 500 stores a mapping between regions in the template image and names of each region (e.g., "upper lobe of the right lung"), and retrieves the name of the clicked region to use in text prompt 208.
[0034] In method 100a, the second machine learning-based system is a text encoder 210. The text encoder 210 generates a feature vector 212, which in this example represents the text "ground-glass opacity in the upper lobe of the right lung." The text encoder 210 includes a recurrent (recurrent) neural network that generates the feature vector 212. The feature vector 212 may have the same number of components as each of the image embedding vectors 206. Training of the text encoder 210 is described below.
[0035] As previously described, in step 108, the method 100 includes performing a comparison between the feature vector 212 and a plurality of image embedding vectors 206 and identifying a first image embedding vector from among the plurality of image embedding vectors 206 based on the comparison.
[0036] In the method 100a, the apparatus 500 incorporating the text encoder 210 and the image encoder 204 generates, for each image embedding vector, a similarity measure indicating the degree of similarity between the image embedding vector and the feature vector 212. The identified (first) image embedding vector is selected as the image embedding vector identified from all image embedding vectors 206 as having the highest similarity to the feature vector 212.
[0037] In method 100a, image embedding vector 206 and feature vector 212 are represented in a common embedding space, and the similarity measure is a distance measure in the common embedding space between image embedding vector and feature vector 212. Any distance measure described herein may be, for example, Euclidean distance or cosine distance. By representing image embedding vector 206 and feature vector 212 in a common embedding space, the distance measure can be used to identify the image embedding vector whose content most closely matches feature vector 212. This allows for the identification of an image embedding vector that represents the medical image feature described by text prompt 208.
[0038] Since the feature vector 212 encodes information about the region of the medical image feature pointed out in step 106, the comparison is based on region data representing this region. Pointing out the region in addition to the medical image feature provides more information on which to base the comparison, which means that the image embedding vectors identified through the comparison are more likely to be related to the medical image feature that the medical expert wishes to segment.
[0039] The method 100 a includes determining location data representing locations in the medical image 202 of medical image features within the medical image data 202 based on the identified image embedding vector 214 .
[0040] The location data may include a segmentation map 218 represented by the identified image embedding vectors 214. The segmentation map 218 indicates the probability that a medical image feature is present at each location in the medical image 202. In this example, each location is the location of each pixel in the medical image 202.
[0041] In this example, the medical image feature data referred to in step 104 includes a probability that the medical image feature is present at each of a plurality of locations within the medical image 202. As discussed above, each of these locations is a portion of the medical image 202. Each portion of the image includes a plurality of voxels.
[0042] In this example, the medical image feature data is interpolated by the image upscaler 216 to provide, for each voxel of the medical image 202, the probability that each feature is present in the voxel. This interpolation may comprise, for example, trilinear interpolation. For example, the probability that each feature is present in a given voxel may be calculated by trilinear interpolation of the probabilities that each feature is present in each of the eight portions of the image nearest to the voxel. In this manner, the low-dimensional representation of the probability map may be upscaled to provide a sufficient probability map of the medical image 202.
[0043] Showing a probability of presence, rather than just a binary indication of whether a feature is present or not, means that the segmentation map 218 can show the blurred boundaries of medical abnormalities.
[0044] The method 100a includes displaying the segmentation map 218 on a display device 510. The segmentation map 218 may be displayed over the medical image 202 to indicate areas of the medical image 202 in which abnormalities are present. The display device 510 may be a computer monitor or other display screen of a computer and is connected to the processor 502 of the apparatus 500. A medical professional reviewing the segmentation map 218 displayed on the display device 510 can, for example, use the segmentation map 218 to make a qualitative assessment of the extent of the abnormality and determine a course of treatment based on this assessment.
[0045] The location data may additionally or alternatively include natural language text describing a finding 222 associated with the medical image data 202. Method 100a includes inputting at least one, and optionally all, of the plurality of image embedding vectors 206 to a text decoder 220, also referred to as a fourth machine learning-based system, to generate this natural language text. The natural language text may include a finding 222 such as "ground-glass opacity in the upper lobe." Text decoder 220 may include, for example, a recurrent neural network. Training of text decoder 220 is described below with reference to FIG. 4.
[0046] A medical professional may use the natural language text, for example, to diagnose a medical condition or recommend a treatment. In either case, generating such natural language text allows a radiologist or other medical professional to understand the patient's condition, as represented by the image embedding vector 206, by reading the natural language text.
[0047] The method 100a includes performing a clustering process on a plurality of selected image embedding vectors determined from separate sets of medical image data. For example, k-means clustering can be used to determine to which of a plurality of clusters a particular image embedding vector belongs. In this manner, it can be determined whether two image embedding vectors representing different sets of medical image data represent the same medical abnormality.
[0048] The method 100a includes determining whether the medical image 202 represents a medical abnormality using a classifier 224. In one example, the classifier 224 is a large-scale language model, such as ChatGPT, to which the generated natural language text is input along with a text prompt such as, "Does this represent a medical abnormality? Please answer yes or no." In another example, the generated natural language text is input to a natural language processing system that compares each word in the natural language text describing the finding to a list of words related to medical abnormalities, such as lesions, and outputs an indication (e.g., a binary output) that the medical image 202 represents an abnormality if and only if the natural language text contains a word from the list.
[0049] In another example, the image embedding vectors 206 are all input to a fourth machine learning-based system to determine whether the medical image 202 represents a medical abnormality. The fourth machine learning-based system is trained using supervised learning, for example, by inputting the image embedding vectors generated by the trained image encoder 204 into the fourth machine learning-based system to generate a trial classification of whether the medical image 202 represents a medical abnormality. The supervised learning includes minimizing a loss function between the trial classification and ground truth data representing whether each medical image used to generate the image embedding vector represents a medical abnormality.
[0050] The output of the classifier 224 can be used to determine whether further processing of the medical image 202 is required. If the output of the classifier 224 indicates that the medical image 202 does not represent an abnormality, the medical image 202 can be removed from the radiology workflow. For example, the processor may display the segmentation map 218 on the display device 510 or perform other processing of the medical image 202 in response to a determination that the medical image 202 represents a medical abnormality. In this example, if it is determined that the medical image 202 does not represent a medical abnormality, the processor 502 does not display the segmentation map 218 on the display device 510. Alternatively, the processor 502 may indicate on the display device 510 that the medical image 202 does not represent a medical abnormality using text such as, for example, "No abnormalities detected in the image."
[0051] Determining whether a medical image 202 represents a medical abnormality and performing further processing in response to a determination that the medical image 202 represents an abnormality can improve the efficiency of analysis of the medical image 202. Removing medical images that do not represent a medical abnormality from the radiology workflow reduces redundant processing of medical images.
[0052] 3 shows a flow diagram of an example process used by the image encoder 204 of method 100a to generate image embedding vectors 206. The image encoder 204 is trained to generate multiple image embedding vectors 206 based on medical image data 202. The training process is described in detail below. The example image encoder 204 described below is a group vision transformer similar to that described in Section 3 of Xu et al., "GroupViT: Semantic Segmentation Emerges from Text Supervision," arXivorg, 2022. However, other transformer-based attention mechanisms, such as the Swin Transformer, may also be used.
[0053] The image encoder 204 first divides the image into equal-sized portions 302, for example, equal-sized rectangular portions 302. The image encoder 204 assigns a label to each portion 302. Initially, the labels are all different. In the illustrative 2D example shown in FIG. 3, there are a total of 6×6=36 portions 302, labeled 1 through 36. In practice, the number of portions 302 is typically much greater. If the medical image data 202 is 3D, each portion 302 may consist of a cube of voxels within the medical image 202. For example, each portion 302 includes a cube of voxels with a side length of three voxels.
[0054] For each label, each portion 302 is assigned a probability that the portion 302 belongs to that label. The initial probability assignment simply indicates that the portion 302 belongs to only one label and no other labels. For example, the portion in the upper left corner of FIG. 3 has a 100% probability of belonging to label 1 and a 0% probability of belonging to any other label. The portion immediately to the right of this portion has a 100% probability of belonging to label 2 and a 0% probability of belonging to any other label.
[0055] The medical image data 202 for each portion 302 (e.g., voxel) is input to a transformer layer 304. The transformer layer 304 determines a similarity metric for each pair of portions 302. The similarity metric represents the visual similarity between the portions 302.
[0056] The similarity metric is input to a grouping layer 306, which groups each portion 302 according to the similarity metric and determines the number of labels that are less than the number of labels that existed before the grouping. In the example shown in Figure 3, 36 initial labels have been reduced to 14 intermediate labels after the first grouping layer 306. The number of labels before and after grouping is a fixed parameter of the image encoder 204.
[0057] The intermediate labels represent specific image features. For example, intermediate label 1 in the intermediate image of Figure 3 represents the empty black space. Intermediate label 3 represents the tissue surrounding the lungs of the image object. Other intermediate labels represent other image features.
[0058] Through the above transformation and grouping, each image portion 302 is assigned a probability of belonging to each intermediate label. The intermediate label shown in FIG. 3 for each portion 302 of the medical image 202 represents the intermediate label with the highest probability for that portion. The black space at the top of the image is likely to have a nearly 100% probability of belonging to intermediate label 1 and a nearly 0% probability of belonging to any other intermediate label. On the other hand, for example, a portion 302 with a higher probability of belonging to intermediate label 10 than any other intermediate label may have a 40% probability of belonging to intermediate label 10, a 30% probability of belonging to intermediate label 9, and a 10% probability of belonging to intermediate label 8. The probabilities for a given image portion do not need to total 100%.
[0059] The transforming and grouping process may be repeated a certain number of times until a desired number of labels emerge. In the case shown in FIG. 3, two final labels emerge after the final transforming and grouping stage. Depending on the content of the training data, these two final labels may represent specific medical image features, which constitute the multiple medical image features mentioned above. For example, final label 1 in FIG. 3 represents normal tissue, while final label 2 represents abnormal tissue. Again, for simplicity, only the label associated with the highest probability is shown for each portion 302 of the medical image 202. Each portion 302 has a probability of belonging to final label 1 (i.e., representing normal tissue) and a probability of belonging to final label 2 (i.e., representing abnormal tissue).
[0060] Each image embedding vector generated by the image encoder 204 is associated with one of the final labels or medical image features. In this example, each component of the image embedding vector represents the probability that each image portion 302 belongs to the label or medical image feature associated with the image embedding vector.
[0061] 4 shows a flow diagram of an example training process 400 for training the image encoder 204, text encoder 210, and text decoder 220 used in method 100a. In the example shown in FIG. 4, a large-scale language model (e.g., ChatGPT) is trained in a separate process and used in its already-trained form in training process 400.
[0062] The training process 400 includes inputting multiple sets of medical image data 202 into the image encoder 204 to generate multiple trial image embedding vectors 228 for each set of medical image data 202. The trial image embedding vectors 228 referred to here are image embedding vectors generated during the training process 400, as opposed to image embedding vectors generated in the inference method 100 by the trained image encoder 204.
[0063] The sets of medical image data 202 can be selected from a bank of medical images. The selection process can include selecting medical images that each include one of a plurality of medical image features. For example, if the plurality of medical image features are a plurality of medical abnormalities, images that include one of the medical abnormalities are selected, and images that include the other abnormalities are excluded. If the plurality of medical image features include normal tissue and abnormal tissue, multiple sets of medical image data 202 representing normal tissue are selected, and approximately the same number of sets of medical image data 202 representing abnormal tissue are selected.
[0064] The inventors recognized that because medical reports typically contain indications of medical image features along with associated images, they can be used as training data, with the indications in the reports serving as ground truth data. This avoids the difficulty of labeling medical image training data, which generally disqualifies the human practitioner of the training process from accurately assessing medical images. Accordingly, each set of medical image data 202 is associated with a medical report 224 about the medical image 202. The medical report 224 includes data indicating the presence of one of multiple medical image features. For example, the medical report 224 includes text such as "Hypoattenuating Lesion Present." Each report also includes data describing the region of the feature indicated to be present. For example, the medical report 224 includes text such as "Right Lobe of the Liver." Such reports may be written by a medical professional or other person. The report may also include data indicating the presence of the feature in the same sentence as the data describing the region of the feature. For example, the medical report 224 may include a sentence such as "Hypoattenuating Lesion Present in the Right Lobe of the Liver."
[0065] Through a first loss function, described below, the region data can guide the image encoder 204 toward portions of the image that contain medical image features. For example, if the pre-trained image encoder 204 is only told that there are hypoattenuating lesions somewhere in the medical images 202 used for training, it will likely incorrectly identify portions of the image that do not contain hypoattenuating lesions as containing hypoattenuating lesions. Specifically, it may not output any image embedding vectors that indicate a high probability for the image portion 302 that contains hypoattenuating lesions and a low probability for the image portion 302 that does not contain hypoattenuating lesions. On the other hand, if the pre-trained image encoder 204 is told that there are hypoattenuating lesions in the right lobe of the liver, it will be more likely to focus attention on the correct portion of the image. It will be more likely to output at least one image embedding vector that accurately indicates a high probability value for the image portion 302 that contains hypoattenuating lesions and low probability values for other portions. The inclusion of region data allows the image encoder 204 to train using fewer training examples because it can more easily identify the portions of the image to which it should direct its attention. The accuracy of the image embedding vectors generated by the image encoder 204 after training is also improved, which can result in improved accuracy of the segmentation map 218 described above.
[0066] In this example, the report is input into a natural language processing system that generates a set of data representing findings 222 in a medical report 224. The natural language processing system is a large-scale language model (LLM) 226, such as ChatGPT, as shown in FIG. 4. Prior to input to the LLM 226, the medical report 224 may contain information that is not useful for training the text encoder 210 and the image encoder 204. For example, the medical report 224 may include the patient's name and measurements about the patient, such as the patient's height and weight. The medical report 224 may be input into the LLM 226 along with a prompt such as "Summarize the findings of this medical report 224." Data representing the output of the LLM 226 is then input into the text encoder 210. This helps eliminate redundant information that may mislead the text encoder 210.
[0067] The text encoder 210 is configured to generate a trial feature vector 230 based on the input summary of findings 222. The trial feature vector 230 efficiently encodes information about the medical image features (e.g., abnormalities) described in the medical report 224 and the regions in which the image features reside.
[0068] The training process 400 involves training the image encoder 204 and the text encoder 210 together to minimize a first loss function between the trial image embedding vectors 228 and the corresponding trial feature vectors 230. In this example, the trial image embedding vectors 228 generated for a given medical image are averaged to generate an averaged embedding vector 232. Each component of the averaged embedding vector 232 represents, for each image portion, the average, across all medical image features, of the probability that the image portion 302 contains the given medical image feature. Note that the probabilities for a given image portion 302 do not need to sum to one, so the average probability is not necessarily constant and is likely to vary across images. It is also worth noting that in this example, image embedding vectors are not manually assigned to each medical image feature prior to training—instead, this is performed by training, and the specific assignment of medical image features to image embedding vectors is not arbitrarily set by the user.
[0069] The first loss function evaluates the similarity between the generated trial feature vector 230 and the averaged image embedding vector 232. The similarity may be evaluated by, for example, Euclidean distance or cosine distance. The value of the first loss function increases as the level of similarity between the averaged image embedding vector 232 and the feature vector of a given medical image decreases.
[0070] In some examples, a contrastive loss function may be used as the first loss function. For example, medical images used for training may be grouped into pairs, with each pair including a normal medical image representing only normal tissue and an abnormal medical image representing a medical abnormality. The contrastive loss function evaluates the similarity between the averaged image embedding vectors 232 generated for the normal medical images and the averaged image embedding vectors 232 generated for the abnormal medical images. The value of the contrastive loss function increases with the level of similarity between the averaged image embedding vectors 232 of the normal medical images and the averaged image embedding vectors 232 of the abnormal medical images.
[0071] In either case, the first loss function can be used to update parameters of the image encoder 204 and the text encoder 210. Weights of both the transformer layer 304 and the grouping layer 306 of the image encoder 204 can be updated, while the number of labels before and after each grouping layer 306 is typically a fixed parameter of the image encoder 204. Furthermore, in this example, parameters of the text decoder 220 are not updated using the first loss function.
[0072] Through training, the image encoder 204 can be trained to generate, based on the medical image 202, an image embedding vector that represents information that would be contained in a summary of findings 222 in a medical report 224 about the medical image 202.
[0073] The assignment of image embedding vectors to each medical image feature (e.g., separate medical anomalies) is automatic and performed by training. Medical reports 224 for abnormal medical images typically include findings describing the anomaly. Therefore, to minimize the first loss function, it is beneficial for the image encoder 204 to generate at least one image embedding vector for each medical image that represents a high probability value for the image portion 302 containing the anomaly. Because it is likely that the same anomaly has been mentioned in several of the medical reports 224 for each medical image, generating such a vector reduces the total value of the first loss function for these medical images. Therefore, to minimize the first loss function, the image encoder will output image embedding vectors that adequately describe anomalies commonly mentioned in the medical reports. The actual medical image features represented by the image embeddings therefore depend primarily on the content of the medical report and secondarily on the number of image embeddings generated for each medical image.
[0074] Consequently, the image encoder 204 can be trained to generate multiple image embedding vectors corresponding to multiple respective medical image features based on the received medical image data 202. In particular, the image encoder 204 can generate an image embedding vector, or segmentation map, for each of multiple different types of abnormalities.
[0075] Furthermore, using a first loss function, the text encoder can be trained to generate a feature vector based on the medical image feature indications. The first loss function is minimized when the averaged image embedding vector 232 matches the trial feature vector 230 generated by the text encoder. Because the medical report describes features (e.g., abnormalities) present in the medical image 202, the text encoder will generate a feature vector that represents the information encoded in the summary of findings generated by the LLM 226 to minimize the first loss function. If this were not done and instead the feature vectors were generated randomly, the image embedding vectors generated by the image encoder 204 would not match the feature vectors generated by the text encoder 210.
[0076] Some automated segmentation methods can be used to identify clearly defined medical image features. However, some features, including ground-glass opacities and bone in x-rays, do not present clearly defined boundaries when represented in a medical image. In other words, the transition between normal tissue and tissue affected by a medical image feature is gradual. Therefore, image segmentation methods that involve determining the boundary between two segments of an image are inadequate for segmenting such features.
[0077] To address this issue, machine learning-based systems can be trained using medical images where doctors have manually designated the boundaries of abnormalities, but this is very labor-intensive.
[0078] The above-described training process 400 and resulting runtime method 100a address this problem by using readily available medical reports as training data.
[0079] As described above, method 100a includes inputting at least one of the image embedding vectors into a text decoder 220 to generate natural language text describing findings 222 associated with the medical image data 202. An example training method for training the text decoder 220 will now be described. This training method may form part of the training process 400.
[0080] The training method includes inputting at least one image embedding vector generated by an image encoder 204 based on a set of input medical image data 202 to a text decoder 220. Based on the input image embedding vector, the text decoder 220 generates trial natural language text describing findings 222 associated with the set of input medical image data 202.
[0081] The sets of training data used in training with the first loss function are also used to train the text decoder 220. As mentioned above, these sets of training data include sets of medical image data 202 and corresponding medical reports 224. The text decoder 220 is trained after the image encoder 204 and text encoder 210 are trained. Alternatively, in alternating training sessions, the image encoder 204 / text encoder 210 are trained using the first loss function and the text decoder 220 is trained using the second loss function. In either case, in some examples, the parameters of the image encoder 204 and text encoder 210 are not updated in training with the second loss function.
[0082] In one example, the text decoder 220 generates trial natural language text based on some or all of the image embedding vectors generated by the image encoder 204 for a given set of input medical image data 202 .
[0083] The training method includes training the text decoder 220 to minimize a second loss function between trial natural language text generated for each set of input medical image data 202 and data representing the corresponding medical report 224. In one example, the data representing the medical report 224 includes a summary of findings generated by the LLM 226 for training using the first loss function. For example, a vector representing the natural language text generated by the LLM 226 can be compared with a vector representing the trial natural language text generated by the text encoder 210. Here, the natural language text can be represented as a vector, for example, by using a mapping between words in the natural language text and vectors such as word2vec and averaging the vectors to generate a vector representing the entire natural language text. The comparison can include, for example, calculating the Euclidean distance or cosine distance between the vectors. Based on the comparison, a value of the second loss function for a given set of input medical image data 202 can be calculated. The value of the second loss function increases as the distance between the vectors increases, i.e., the trial natural language text matches the summary of findings generated by the LLM 226 less closely.
[0084] The second loss function can be used to update the parameters (e.g., weights) of the text decoder 220. In this example, it is not used to update the parameters of the image encoder 204.
[0085] Through training using the second loss function, the text decoder 220 is trained to generate natural language text representing information that would be included in a summary of findings in a medical report 224 for the medical image 202 based on at least one image embedding vector output by the image encoder 204 based on the medical image 202.
[0086] FIG. 5 illustrates an example of an apparatus 500. The apparatus 500 may be implemented as a processing system and / or a computer. The apparatus 500 includes an input interface 506, an output interface 508, a processor 502, a memory 504, and a display device 510. The processor 502 and the memory 504 are configured to perform the method 100, the method 100a, and / or the training process 400 according to any of the examples described above with reference to FIGS. 1-4. The memory stores instructions that, when executed by the processor 502, cause the processor 502 to perform the method 100, the method 100a, according to any of the examples described above with reference to FIGS. 1-3, and / or the training method 400 according to any of the examples described above with reference to FIG. 4. The instructions may be stored on a computer-readable medium, such as a non-transitory computer-readable medium.
[0087] For example, the input interface 506 receives the medical image data 202, and the processor 502 performs the method 100a described above with reference to Figures 1-3 to identify image embedding vectors. The apparatus 500 also displays the segmentation map 218 of the identified image embedding vectors 214 on the display device 510. In some examples, the natural language text generated by the text decoder 220 is sent to structured storage (not shown), where the natural language text is stored.
[0088] The foregoing examples should be understood as illustrative of the present invention. It should also be understood that a feature described in connection with one example can be used alone or in combination with other features described, or can be used with one or more features of other examples or in combination with other examples. Furthermore, equivalents and modifications not described herein may be employed without departing from the scope of the present invention, as defined in the claims.
Claims
1. 1. A computer-implemented method for processing medical image data, comprising: receiving medical image data representing a medical image at a first machine learning based system; the first machine learning based system generates, based on the received medical image data, a plurality of image embedding vectors corresponding to a plurality of medical image features, each of the plurality of image embedding vectors being associated with a respective separate medical image feature and including medical image feature data indicative of the presence or absence of each of the medical image features at each of a plurality of locations within the medical image; receiving, at a second machine learning based system, an indication of a first medical image feature included in the medical image, the indication including region data representing a region of the first medical image feature within the medical image; and generating, by the second machine learning based system, a feature vector based on the indication; performing a comparison of the feature vector with the plurality of image embedding vectors based at least in part on the region data, and identifying a first image embedding vector from among the plurality of image embedding vectors based on the comparison; The method of claim 1, wherein the first image embedding vector is identified from the plurality of image embedding vectors as having the highest similarity to the feature vector.
2. The method of claim 1 , further comprising determining position data representing a position within the medical image of the indicated first medical image feature within the medical image data based on the identified first image embedding vector.
3. 3. The method of claim 2, wherein the medical image feature data represents a segmentation map of each of the medical image features, and the determined location data includes the segmentation map represented by the identified first image embedding vector.
4. The method of claim 3 , wherein the segmentation map indicates the probability that each of the medical image features is present at each location within the medical image.
5. determining whether the medical image is indicative of a medical abnormality; The method of claim 3 , further comprising displaying the segmentation map on a display device in response to determining that the medical image is indicative of a medical abnormality.
6. The method of claim 1 , wherein the indication received at the second machine learning based system includes data representing a text prompt indicative of the first medical image feature.
7. performing a training process to train the first machine learning based system and the second machine learning based system; This training process is inputting first training data, the first training data including a plurality of sets of medical image data, into the first machine learning based system to generate a plurality of trial image embedding vectors for each of the sets of medical image data; inputting second training data into the second machine learning based system to generate a trial feature vector for each set of medical image data; the second training data includes data representing a plurality of medical reports, each of the medical reports including data indicating the presence of one of the plurality of medical image features in a corresponding one of the sets of medical image data, each of the medical reports including data representing a region of the feature indicated to be present in the medical image represented by the corresponding set of medical image data; 2. The method of claim 1 , wherein the training process comprises jointly training the first machine learning based system and the second machine learning based system to minimize a loss function between the trial image embedding vectors and the corresponding trial feature vectors.
8. 8. The method of claim 7, comprising inputting each of the medical reports into a natural language processing system to generate a set of data representing findings of the medical reports, wherein the data representing the plurality of medical reports is data representing findings of the medical reports.
9. 10. The method of claim 1, further comprising inputting at least one of the plurality of image embedding vectors into a third machine learning based system to generate natural language text describing findings related to the medical image data.
10. performing a training method for training the third machine learning based system; This training method is inputting the image embedding vectors generated by the first machine learning based system based on sets of input medical image data into the third machine learning based system to generate, for each set of input medical image data, trial natural language text describing findings related to the set of input medical image data; 10. The method of claim 9, comprising training the third machine learning based system to minimize a loss function between the trial natural language text and data representing a medical report corresponding to the set of input medical image data.
11. The method of claim 1 , wherein at least one of the plurality of medical image features comprises a medical abnormality.
12. An apparatus configured to carry out a method according to any one of claims 1 to 11.
13. A computer program which, when executed by a computer, causes the computer to carry out the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Treatment order determining method, computer program, and computing device
JP2020149682A
Automated medical scanning triage system and methods for its use
JP2023537619A
Automatic generation of medical imaging reports based on fine grained finding labels
US11244755B1
Medical visual question answering
US20220130499A1
Method and apparatus for annotating a portion of medical imaging data with one or more words
US20230033783A1