Method and system for image segmentation and identification
The system addresses the inefficiencies of manual and semi-manual medical image segmentation by integrating annotation and segmentation, enabling continuous model refinement and reducing human dependency, thus enhancing the accuracy and efficiency of medical image analysis.
Patent Information
- Application Number
- JP2020101099
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2020-06-10
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2040-06-10
AI Technical Summary
Existing medical image segmentation and identification methods are labor-intensive, time-consuming, and dependent on expert expertise, with challenges in training and validating machine learning models due to the need for human-annotated ground truth data and limited training images, making it difficult to efficiently improve the models.
A system and method for integrating annotation and segmentation, using a training subsystem to refine machine learning models by comparing segmentations with existing annotations, allowing continuous retraining and deployment based on performance thresholds, and incorporating non-machine learning image processing methods to generate candidate annotations.
Facilitates efficient and accurate medical image segmentation and identification by continuously refining machine learning models, reducing reliance on human input and improving model performance over time.
Smart Images

Figure 0007709821000001 
Figure 0007709821000002 
Figure 0007709821000003
Abstract
Description
Technical Field
[0001] The present invention relates to methods and systems for image segmentation and identification, and more particularly to segmentation and identification of medical images (e.g., bones, other anatomical structures, masses, tissues, landmarks, lesions, pathological conditions, etc.), and also relates to medical imaging modalities such as computed tomography (CT), magnetic resonance (MR), ultrasound, lesion scanner imaging, etc., but is not limited to these applications. The present invention also relates to methods and systems for automating the annotation (annotation) of medical image data for training a machine learning model for segmentation and identification purposes, and for evaluating and improving a machine learning model.
[0002] Related Application This application claims priority based on and claims the benefit of U.S. Patent Application No. 16 / 448,252, filed on June 21, 2019, the entire content of which is incorporated herein by reference.
Background Art
[0003] Accurate segmentation and identification of medical images are required for quantitative analysis and disease diagnosis. Segmentation is the drawing of an object (e.g., an anatomical structure or tissue) within a medical image from its background. Identification is the process of identifying an object and correctly labeling it. Traditionally, segmentation and identification are performed manually or semi-manually. Manual techniques require an expert with sufficient knowledge in the field to be able to draw the contours of the target object and label the extracted object.
[0004] There also exist computer-aided systems that provide semi-manual segmentation and identification. For example, such systems can detect an approximate contour of an object of interest based on selected parameters including signal strength, edges, 2D / 3D curvature, shape, or other 2D / 3D geometric features. Then, an expert manually refines the segmentation or identification. Alternatively, an expert can provide input data such as the approximate position of a target object to such a system, and the computer-aided system performs the segmentation and identification.
[0005] Whether manual or semi-manual, both are labor-intensive and time-consuming methods. Also, the quality of the results strongly depends on the expertise of the expert. Depending on the operator / expert, considerable differences can result in the segmented and identified objects.
[0006] In the past few years, machine learning, especially deep learning (e.g., deep neural networks and deep convolutional neural networks), has come to outperform humans in many visual recognition tasks including medical imaging.
[0007] Patent Document 1 discloses a method and system for medical image segmentation based on artificial intelligence. The method includes the steps of receiving a medical image of a patient, automatically determining a current segmentation context based on the medical image, and automatically selecting at least one segmentation algorithm from a plurality of segmentation algorithms based on the current segmentation context. Using the selected at least one segmentation algorithm, a target anatomical structure is segmented within the medical image.
[0008] Patent Document 2 discloses a method for segmenting an image of a target patient, the method comprising the following steps: providing a target 2D slice and a neighboring 2D slice for a 3D anatomical site image; calculating a segmentation region by a trained multi-slice fully convolutional neural network (multi-slice FCN), the region including defined intra-body anatomical features that spatially extend across the target 2D slice and the neighboring 2D slices, and each of the 2D slices and each neighboring 2D slice being processed by a corresponding downsampling component of the sequential downsampling components of the multi-slice FCN, the processing being in accordance with the order of the target 2D slice and the neighboring 2D slices, which is based on a sequence of 2D slices extracted from the 3D anatomical image, and the output of the sequential downsampling components being combined and processed by a single upsampling component that outputs a segmentation mask for the target 2D slice.
[0009] Patent Document 3 discloses a system and method for applying a deep convolutional neural network to medical images to generate a diagnosis or a recommended diagnosis plan in real time or near real time. The method includes: a step of performing image segmentation on a plurality of medical images, the image segmentation step including separating a region of interest from each image; a step of applying a cascaded deep convolutional neural network detection structure to the segmented images, the detection structure including: i) a first stage of screening all possible positions within each two-dimensional slice of the segmented medical images by means of a sliding window method using a first convolutional neural network to identify one or more candidate positions; and ii) a second stage of screening a three-dimensional solid constructed from the candidate positions by using a second convolutional neural network, the screening being performed by selecting at least one random position within each solid with a random scale and a random viewing angle, thereby identifying one or more refined positions and classifying the refined positions; and a step of automatically generating a report including a diagnosis or a recommended diagnosis plan.
[0010] However, there are problems with applying such a method to the segmentation and identification of medical images. That is: To train and validate a machine learning algorithm, ground truth data is required, but this data is provided by human experts who annotate the data, which involves a great deal of time and cost; improving a machine learning model ideally requires additional misrecognition results (for example, when a trained model fails in the segmentation or identification of a medical image, etc.), but it is highly difficult to efficiently manage the training and retraining processes; for some biomedical applications, it is difficult or impossible to obtain a large number of training images, and it becomes difficult to efficiently train a segmentation and identification model if the training data is limited.
Prior Art Documents
Patent Documents
[0011]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
[0012] An object of the present invention includes providing a segmentation system capable of integrating annotations.
[0013] According to a first aspect, the present invention is an image segmentation system comprising: A training subsystem configured to train a segmentation machine learning model using annotated training data (annotated training data) including images (such as medical images) associated with each segmentation annotation to generate a trained segmentation machine learning model, A model evaluator, A segmentation subsystem configured to perform segmentation on structures or materials (including, for example, bone, muscle, fat, or other biological tissues) within an image using a trained segmentation machine learning model. The model evaluator evaluates the segmentation machine learning model by (i) Controlling the segmentation subsystem to segment at least one evaluation image associated with existing segmentation annotations using the segmentation machine learning model, thereby generating a segmentation for the annotated evaluation image; (ii) Forming a comparison between the segmentation of the annotated evaluation image and the existing segmentation annotations; And if the comparison indicates that the segmentation machine learning model is qualified, deploying or releasing the trained segmentation machine learning model for use.
[0014] Thus, the model evaluator evaluates the segmentation machine learning model using the operational segmentation subsystem used for segmentation, providing an integrated training and segmentation system.
[0015] In an embodiment, the system is configured to deploy or release the model for use if the segmentation of the annotated evaluation image and the existing annotation match within a predetermined threshold.
[0016] In an embodiment, the system is configured to continue training the model when the segmentation of the annotated evaluation image and the existing annotation do not match within a predetermined threshold. For example, the system is configured to continue training the model by modifying the model algorithm and / or adding additional annotated training data. Thereby, the predetermined threshold can be tuned according to the desired application, and refinement can be made if the initial tuning is not satisfactory.
[0017] In an embodiment, the training subsystem (i) receives an annotation for an image (some of which may optionally be generated by the segmentation subsystem using a segmentation machine learning model) and (ii) a score associated with the annotation, the score indicating the degree of success or failure of the segmentation machine learning model in segmenting the image, where a higher score indicates failure and a lower weighting indicates success; and is configured to retrain or refine the segmentation machine learning model using the image and the annotation, including weighting the annotation of the image according to the score.
[0018] Therefore, the segmentation machine learning model can be continuously retrained or refined before and during deployment. Note that the annotation for an image can include multiple items of annotation information, and optionally, the annotation can be partially generated by the segmentation subsystem using a segmentation machine learning model.
[0019] In an embodiment, the training subsystem generates a segmentation machine learning model by refining or modifying an existing segmentation machine learning model.
[0020] In an embodiment, the system includes an annotation subsystem, which includes steps to provide at least one annotated training image (i.e., each image with a segmentation annotation) from an unannotated image or a partially annotated image, which includes generating one or more candidate image annotations for the unannotated image or the partially annotated image, receiving an input to identify one or more parts of the one or more candidate image annotations (where the part can be composed of the whole of the candidate image annotation), and providing an annotated training image with at least one or more parts. For example, the annotation subsystem is configured to generate at least one of the candidate image annotations using a) a non-machine learning image processing method, or b) a segmentation machine learning model.
[0021] Therefore, the annotation subsystem can be used to show the validity of the annotation (which may be generated by a non-machine learning image processing method), and the valid part of the annotation can be retained and used.
[0022] In an embodiment, the system further includes an identification subsystem, the annotated training data further includes identification annotations, and the segmentation machine learning model is a segmentation and identification machine learning model.
[0023] Combining the annotation subsystem and the segmentation subsystem promotes the continuous improvement of the segmentation system, which can be beneficial for applying artificial intelligence to medical image analysis.
[0024] In an embodiment, the system further comprises an identification subsystem, the annotated training data further provides identification annotation, and the training subsystem is further configured to train an identification machine learning model using (i) the annotated training data after each image has been segmented by a segmentation machine learning model and (ii) the identification annotation.
[0025] The training subsystem may include a model trainer, which is configured to utilize machine learning to train a segmentation machine learning model to determine the classification of each pixel / voxel of an image. The model trainer can employ, for example, a support vector machine, a random forest tree, a deep neural network, etc.
[0026] The structure or material may include bone, muscle, fat, or other biological tissues. For example, the act of segmentation can be to separate bone from non-bone materials (such as surrounding muscle and fat), or to separate one bone from another bone.
[0027] According to a second aspect, the present invention provides a computer-implemented image segmentation method, the method comprising: Training a segmentation machine learning model using annotated training data comprising each image (e.g., a medical image) and segmentation annotations to generate a trained segmentation machine learning model; Evaluating the segmentation machine learning model, (i) Segmenting at least one evaluation image associated with existing segmentation annotations using the segmentation machine learning model to thereby generate a segmentation for the annotated evaluation image; (ii) Forming a comparison between the segmentation of the annotated evaluation image and the existing annotations to perform this evaluation step. If the comparison indicates that the segmentation machine learning model is qualified, the method includes the step of deploying or releasing the trained segmentation machine learning model for use.
[0028] In an embodiment, the method involves deploying or releasing the model for use if the segmentation of the annotated evaluation image and the existing annotation match within a predetermined threshold.
[0029] In an embodiment, if the segmentation of the annotated evaluation image and the existing annotation do not match within a predetermined threshold, the method involves continuing to train the model. For example, the method involves continuing to train the model by modifying the model algorithm and / or adding additional annotated training data.
[0030] In an embodiment, the training includes the steps of receiving (i) annotated images for use in retraining or refining the segmentation machine learning model and (ii) a score associated with the annotated images, the score indicating the degree of success or failure of the segmentation machine learning model in segmenting the annotated images, where a higher score indicates failure and a lower weighting indicates success; and weighting the annotated images according to the score during retraining or refining the segmentation machine learning model.
[0031] In an embodiment, the method includes the step of generating a segmentation machine learning model by refining or modifying an existing segmentation machine learning model.
[0032] In an embodiment, the method includes the step of providing at least one annotated training image from an unannotated image or a partially annotated image, which includes the steps of generating one or more candidate image annotations for the unannotated image or the partially annotated image, receiving an input for identifying one or more portions of the one or more candidate image annotations, and providing an annotated training image with at least the one or more portions. By way of example, the method includes the step of generating at least one of the candidate image annotations using a) a non-machine learning based image processing method, or b) a segmentation machine learning model.
[0033] In an embodiment, the annotated training data further comprises identification data, and the method includes the step of training a segmentation machine learning model as a segmentation and identification machine learning model.
[0034] In an embodiment, the annotated training data further provides identification annotations, and the method includes the step of training an identification machine learning model using (i) the annotated training data after each image has been segmented by a segmentation machine learning model and (ii) the identification annotations.
[0035] The method may include the step of training a segmentation machine learning model using machine learning (including, for example, support vector machines, random forest trees, deep neural networks, etc.) to determine the classification of each pixel / voxel of the image.
[0036] According to a third aspect, the present invention provides an image annotation system, the system comprising: a) an input for receiving at least one unannotated or partially annotated (such as a medical image) image, b) an annotator configured to yield respective annotated training images from images that are not annotated or are partially annotated, and to do so by generating one or more candidate image annotations for an image that is not annotated or is partially annotated; receiving an input identifying one or more portions of the one or more candidate image annotations, where the portions can consist of the entire candidate image annotation; and yielding an annotated training image from at least the one or more portions,
[0037] In an embodiment, the annotator is configured to generate at least one of the candidate image annotations using a) a non-machine learning based image processing method, or b) a segmentation machine learning model (such as of the type described above).
[0038] According to a fourth aspect, the present invention provides an image annotation method, the method comprising: a) receiving or accessing an image that is not annotated or is partially annotated; b) yielding respective annotated training images from an image that is not annotated or is partially annotated, and to do so by generating one or more candidate image annotations for an image that is not annotated; receiving an input identifying one or more portions of the one or more candidate image annotations; and yielding an annotated training image from at least the one or more portions,
[0039] In an embodiment, the method includes generating at least one of candidate image annotations using a) a non-machine learning-based image processing method, or b) a segmentation machine learning model.
[0040] According to a fifth aspect, the present invention provides computer program code configured to implement either the second aspect or the fourth aspect when executed by one or more processors. This aspect may also provide a computer-readable medium (which may be non-volatile) comprising the computer program code as described above.
[0041] Thus, certain aspects of the present invention assist in the continuous training and evaluation of a machine learning (e.g., deep learning) model made with an integrated annotation and segmentation system.
[0042] The annotation system not only simply uses images annotated / identified by humans, but also combines segmented / identified results obtained by the following methods: image processing algorithms, annotation deep learning models, correction models, and segmentation models. Human input is only used as a final step for verifying or correcting annotations, if needed. As more images are annotated, the deep learning model is improved and requires less human intervention. In many cases, it has been found that accurate annotations can be obtained by combining segmented / identified results from different models (such as annotation, correction, and segmentation / identification models).
[0043] Regarding any of the various individual features of each of the above aspects of the present invention, and regarding any of the various individual features of the embodiments described herein including the claims, they can be combined in appropriate and desired manners.
Brief Description of the Drawings
[0044] To make the present invention more clearly defined, exemplary embodiments will be described below with reference to the accompanying drawings.
[0045]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7A
Figure 7B
Figure 7C
Figure 8A
Figure 8B
Figure 8C
DETAILED DESCRIPTION OF THE INVENTION
[0046] FIG. 1A is a schematic diagram of the architecture of a segmentation and identification system 10 with built-in annotation and training functions. The system 10 includes an annotation and training subsystem 12, a segmentation and identification subsystem 14, a trained segmentation and identification model 16, and a user interface 18 including a graphical user interface (GUI) 20.
[0047] Generally, the annotation and training subsystem 12 is configured to annotate medical images selected as training data with "ground truth (correct answer)" (regarding the targeted anatomical structures and tissues) under the management of an operator with expertise. The annotated training data is input into a machine learning algorithm (such as a deep learning algorithm, etc.) to train a segmentation and identification model (stored in the trained segmentation and identification model 16). The trained segmentation and identification model is used by the segmentation and identification subsystem 14 (26) to segment and identify objects of interest in new medical images, including pixel / voxel classification and labeling of the images.
[0048] The GUI 20 can be implemented in several different ways. For example, it can be a GUI installed as software on a computing device (such as a personal computer, laptop, tablet computer, or mobile phone, etc.) used by an annotator (i.e., an operator with appropriate skills to recognize the features of the image). In another example, the GUI 20 can be provided to annotators as a web page and accessed by a web browser.
[0049] In addition to training (22) the segmentation and identification model using the annotated images, the annotation and training subsystem 12 evaluates (24) the performance of the trained segmentation and identification model 16 so as to be able to determine whether the model is ready to be deployed, and also to be able to retrain or improve the model.
[0050] However, if the segmentation and identification model fails for one or more specific images, the images for which the model has failed are added to the annotation and training subsystem 12 as a new training data set, and the associated segmentation and identification model is retrained (28).
[0051] Figure 1B is a schematic diagram of the system 10. The system 10 includes a segmentation and identification controller 32 and the aforementioned user interface 18. The segmentation and identification controller 32 includes at least one processor 34 (or, in some embodiments, a plurality of processors) and a memory 36. The system 10 can be implemented, for example, as a combination of software and hardware on a computer (e.g., as a personal computer or a mobile computing device), or as a dedicated image segmentation system. Optionally, the system 10 can be described as being distributed; for example, all or some components of the memory 36 can be located remotely from the processor 34; the user interface 18 can be located remotely from the memory 36 and / or the processor 34, and in fact it can comprise a web browser and a mobile device application.
[0052] The memory 36 is capable of communicating data with the processor 34, and typically comprises both volatile and non-volatile memory (and can include one or more of each memory type), including RAM (random access memory), ROM, and one or more mass storage devices.
[0053] The processor 34 includes an annotation and training subsystem 12 and a segmentation and identification subsystem 14. As will be described in more detail hereinafter, the annotation and training subsystem 12 includes an initial segmenter and identifier 38 (which processes the original image according to a non-machine learning image processing method), an image annotator 40, a model trainer 42, and a model evaluator 44. The segmentation and identification subsystem 14 includes a pre-processor 46, a segmenter 48, a structure identifier 50, and a result evaluator 52. The processor 34 also includes an I / O interface 54 and a result output 56.
[0054] The memory 36 includes program code 58, image data 60, training data 62, evaluation images 64, ground truth images 66, a trained segmentation and identification model 16, a trained annotation model 68, and a trained correction model 69.
[0055] The segmentation and identification controller 32 is at least partially implemented by the processor 34 that executes the program code 58 from the memory 36.
[0056] Generally speaking, the I / O interface 54 is configured to read or receive image data (e.g., in DICOM format) regarding the object under test or the patient into the image data 60 or the training data 62, which are each intended to be used for analysis and / or training (or both, as described hereinafter). Once the system 10 has segmented and identified the various structures from the image data 60, the I / O interface 54 outputs the results of the analysis (optionally in the form of a report) via, for example, the result output 56 and / or the GUI 20.
[0057] Figure 2 is a flowchart 20 of the general workflow of system 10. At S72, a first set of images is imported (typically from a database of such images stored remotely or locally) and used as training images / data. These images are input into system 10 via interface 54 and stored in training data 62. The training data 62 should represent images that are expected to be encountered in a clinical environment and should include both basic and edge cases (i.e., each of the "normal" and "extreme" cases). The structures and tissues in the basic cases should be more easily segmented or identified than those in the edge cases. For example, for a wrist CT scan, it is necessary to segment the bone from the surrounding muscle and fat. This task is easier in scans of young and healthy subjects. Because for these, the porosity of the bone boundary is lower and clearer than that of older and frail patients, for whom the bone is very porous and the boundary clarity is lower. However, it is desirable to collect examples for both basic and edge cases as training data.
[0058] In this step, a second set of images for evaluating the performance of the trained model is also imported. These evaluation images for use in the evaluation should represent clinical images and are stored in evaluation images 64.
[0059] In S74, an operator with appropriate expertise annotates both the training data 62 and the evaluation image 64 using the image annotator 40. Different applications require different annotations. For example, if the model is to be trained to segment bone material from non-bone material, the annotation needs to distinguish and identify bone pixels / voxels from those of non-bone, and thus segmentation data should be included. On the other hand, if the model is more advanced and is required to identify each bone mass, the annotation needs to distinguish and label the pixels / voxels of each bone mass, and thus additional identification data should be included.
[0060] In S76, the model trainer 42 trains a classifier model using the annotated training data 62, which in this embodiment is a segmentation and identification model, which is a classifier that determines the classification of each pixel / voxel on the image. In training the model, the aim is to determine a decision pattern to reach from the input (training image) to the correct answer (annotation). The model can be trained by utilizing machine learning algorithms such as support vector machines or random forest trees.
[0061] In this embodiment, a deep neural network is used. As will be described later (see Fig. 6), this deep neural network consists of an input layer, an output layer, and various layers in between. Each layer consists of artificial neurons. An artificial neuron is a mathematical function that receives one or more inputs, sums them up, and produces an output. Usually, each input is individually weighted, and the sum is passed through a non-linear function. As the neural network learns, the weight values of the model are adjusted, and the adjustment is made according to the resulting error (the difference between the network output and the annotation), and the adjustment is made until the error can no longer be reduced.
[0062] In S78, the model evaluator 44 evaluates the trained model by controlling the segmentation and identification subsystem 14, and performs segmentation and identification on the evaluation image using the model. In S80, the model evaluator 44 checks whether the trained model has reached a passing grade, and does this by comparing the result of this process with the annotation used in S74. In this embodiment, this is done by determining whether the difference between the result generated by the trained model and the annotation data associated with the annotated image (correct image) 66 used in S74 is less than a predetermined threshold. For example, the difference can be calculated as the overlap between the segmentation generated by the model and the segmentation suggested by the annotation data, and the ratio to the segmentation suggested by the annotation data. If these match exactly, the overlap is clearly 100%, and implicitly, the result is a pass. In some applications, this threshold is set to 90%.
[0063] If in S80 the model evaluator 44 determines that the model has failed, the process proceeds to S82, where one or more new images are imported (to supplement the original training data) in the same way as the original training data was imported in S72, and / or the learning algorithm is adjusted / changed (for example, this can be done by tuning the parameters of the neural network, making changes to the layers of the neural network, or making changes to the activation function of the neurons in the neural network, etc.). Then, the process returns to S74.
[0064] When the model evaluator 44 determines at S80 that the difference between the result generated by the trained model and the annotation used at S74 is less than a predetermined threshold, that is, when it is determined that the trained model is qualified, the process proceeds to S84, and the trained model is stored and deployed (as one of the trained segmentation and identification models 16). The training phase is (at least at this stage) completed. Therefore, at S85, the trained model is used by the segmentation and identification subsystem 14 to segment and identify one or more structures or materials / tissues in the new medical image that is input and stored in the image data 60. At S86, the result evaluator 52 verifies the result of this segmentation and identification by comparing the result with one or more predefined conditions or parameters known to characterize the targeted structure or material. If it is shown in this verification that the performance of the segmentation and identification model is unqualified, the process proceeds to S88, and the new image is added to the training set as new training data, and the process proceeds to S74 (where these images are annotated or used for retraining the model, etc.).
[0065] If the result evaluator 52 does not determine at S86 that the result is unqualified, the process proceeds to S90, and the segment / identification result from the new medical image is output, such as being displayed by the user interface 18 of the system 10. (Optionally at S86, if the result evaluator 52 does not determine that the result is unqualified, the segmentation and identification results can be displayed to the user so that the user can perform complementary manual verification. And if the user flags the result as unqualified (for example, by selecting the "unqualified" or "reject" button on the GUI 20), the result is regarded as unqualified and the process proceeds to S88. If the user does not dispute the decision of the result evaluator 52 by selecting the "qualified" or "accept" button on the GUI 20, the process proceeds to S90.)
[0066] After S90, the process ends. For example, the results can be used for diagnostic purposes or as input for further qualitative analysis.
[0067] FIG. 3 is a pseudo flow diagram 70 of the operation of system 10. As described above, the annotation and training subsystem 12 is configured to annotate the segmentation and identification made on medical images. The medical images and their annotations are used to train the segmentation and identification models. The annotator annotates the medical image set via the GUI 20, which may also be referred to herein as the annotation interface.
[0068] The annotator combines information from different resources to complete the annotation. In this embodiment, the annotator combines information from one or more candidate image annotations and, if necessary, uses the annotation tools of the GUI 20 to complete the annotation. In this embodiment, the candidate image annotations can be in the following format: preliminary segmentation and identification results generated using existing non-machine learning image processing methods; or results generated by a partially trained segmentation and identification model. The candidate image annotations are displayed together (i.e., adjacent to each other) so that the annotator can easily compare which parts of each are the best.
[0069] Also, since the preliminary segmentation results are generated using existing image processing methods, you will understand that the preliminary results typically require improvement. Examples of existing image processing methods include contour detection, blob detection, and threshold-based object segmentation. These methods can quickly segment the pixels corresponding to the bones in a CT scan from the surrounding pixels, but in many cases, they result in a general or rough preliminary segmentation.
[0070] The GUI 20 includes a preliminary result window 102 and an annotated image window 104, and the annotated image window 104 includes a plurality of annotation tools 106. The annotation tools 106 include both manual and semi-manual annotation tools, which can be displayed to and operated by an annotator. These tools are schematically shown in FIG. 4, and these tools include manual tools, which include a pen tool 130 and an eraser tool 132. The pen tool 130 can be controlled by an annotator to label pixels of different structures or materials with different values or colors, and the eraser tool 132 can be controlled by an annotator to remove excess segments or over-identified pixels from the target structure or material. The semi-manual annotation tool includes a fill tool 134, and the annotator can control this to draw a contour surrounding the target structure or material, and can cause the annotated image window 104 to annotate all the pixels / voxels within the contour. Further, the semi-manual annotation tool includes a region expansion tool 136, and after an annotation is applied to a very small part of the target structure or material, the annotator can control this to control the annotated image window 104 to expand the annotation to include the entire structure or material.
[0071] System 10 also includes machine learning-based tools, which assist an annotator in quickly performing accurate segmentation and identification. For example, annotation tool 106 includes a machine learning-based annotation model control 138, which calls a trained annotation model stored within annotation model 68. Annotation model 68 is controlled by the annotator to select a plurality of points for use by annotation model 68 when segmenting an object. In this embodiment, annotation model 68 prompts the annotator to identify four extreme points of the object, i.e., the four extreme points are the leftmost, rightmost, uppermost, and lowermost pixels. Then, the annotator sequentially selects each point, for example, by touching the image with a stylus (if displayed on a touch screen) or by using a mouse. Then, annotation model 68 segments the object using these extreme points. In this way, the selected points (e.g., extreme points) constitute annotation data 118 used to train annotation 68.
[0072] Annotation tool 106 includes an annotation motion capture tool 140, which is another mechanism for recording annotation data 118 and can be activated by the annotator. Movements are recorded by system 10 to train a correction model (stored within trained correction model 69). For example, the annotator can move an over-segmented contour inward or an under-segmented contour outward. The inward or outward movement and the location where they are made are recorded by annotation motion capture tool 140 as inputs for training a correction model (stored within trained correction model 69). Also, annotation motion capture tool 140 records the amount of effort, such as the number of mouse movements or the amount of mouse scrolling, and is used to give greater weight to more cumbersome cases during training.
[0073] To return to FIG. 3, during use, the preliminary segmentation and the identification result 110 will be presented to the annotator via the preliminary result window 102 of the GUI 20. If the annotator is satisfied with the accuracy of the preliminary segmentation and the identification result 110, the annotator moves the preliminary result 110 to the annotated image window 104 of the GUI 20, for example, by clicking the mouse on the preliminary result window 102. If the annotator is satisfied with only a part (but not all) of the preliminary result 110, the annotator moves only a sufficient part of the preliminary result 110 into the annotated image window 104, and corrects / finishes the segmentation and the identification using the annotation tool 106 of the annotated image window 104.
[0074] If the annotator is not satisfied with any part of the preliminary segmentation and the identification result, the annotator annotates the image from scratch using the annotation tool 106.
[0075] If the annotator either completes the annotation of a partially sufficient preliminary result 110 or annotates the image himself / herself, the annotated original image set 114 is within the annotated image window 104 and is stored as the correct image 66.
[0076] After one or more ground truth images 66 have been collected in such a manner, a segmentation and identification model is trained using the annotated original images, i.e., the ground truth images 66. Once the trained segmentation and identification model becomes available (within the trained segmentation and identification model 68), the segmentation and identification generated using the trained model are also provided to the annotator for reference. If the annotator is satisfied with the results generated by the trained segmentation and identification model or a portion of such results, the annotator moves a sufficient portion to the annotated image window 104. The annotator can combine a sufficient portion of the preliminary results with a sufficient portion from the trained model within the annotated image window 104. If there are still images or image portions with insufficient annotation, the annotator can correct or supplement the annotation using the annotation tool 106.
[0077] The collected ground truth images and original images 60 within the ground truth image 66 are used to train the segmentation and identification model. To evaluate the trained model, a different set of ground truth images and original images is used. In some embodiments, the same images are used for both training and evaluation. When the performance of the trained model is satisfied during image evaluation, the model is sent to the trained segmentation and identification model 68 and is used by the segmentation and identification subsystem 14 to process any new images. If the performance is insufficient, more ground truth images are collected. The criteria used by the system 10 when evaluating the model (i.e., when determining whether the segmentation and identification model is qualified) are adjustable. Such adjustments are usually made by the developer of the system 10 or the person who first trained the model and are adjusted according to the requirements of the application. For example, when determining vBMD (volumetric bone mineral density), the criteria are set to verify whether the entire bone is accurately segmented from the surrounding material; when calculating cortical porosity, the criteria are set to verify the segmentation accuracy of the entire bone and cortical bone.
[0078] Figure 5 is a flowchart 150 for the annotation and training workflow. Referring to Figure 5, at S152, the original image 60 is selected or input, and at S154, the original image is processed by the initial segmenter and identifier 38, thereby generating preliminary segmentation and identification results. The process proceeds to S156, where the segmenter 48 determines whether the trained segmentation and identification model is available within the trained segmentation and identification model 16. If so, the process proceeds to S158, and the original image 60 is also processed using the trained segmentation and identification model 16. Then, the process proceeds to S160. If at S156 the system 10 determines that the trained segmentation and identification model is not available, the process proceeds to S160.
[0079] At S160, the annotator verifies, using the system 10, whether any part of either the preliminary model result or the trained model generation result is qualified (usually by inspecting these results on the display of the user interface 18). If so, the process proceeds to S162, where the annotator moves the qualified part to the annotated image window 104 using the annotation tool 106, and the process proceeds to S164. If at S160 the annotator determines that none of the results obtained are qualified, the process proceeds to S164.
[0080] At S164, if none of the results obtained are qualified, the annotator is prompted to annotate the image from scratch or to complete, correct, or supplement the annotation in some way. The annotator does this by controlling the manual and semi-manual tools 106 until the annotation is satisfactory. Optionally, the annotator can also call the correction model 69 using the correction model control 142 (if the correction model 69 is available) to improve the annotation.
[0081] At S166, the segmentation and identification models are trained using the ground truth image and the annotation data. In parallel, at S168, the annotation model 68 and the correction model 69 are trained using the ground truth image and the annotation data. (Note that if there is no correction model available at S164, for the case after the correction model is generated at S168, the correction model becomes available within the correction model 69 in a later pass up to S164.)
[0082] The process proceeds to S170, and the trained segmentation and identification model is evaluated by the model evaluator 44 of the annotation and training subsystem 12. At S172, the result of the evaluation is verified. If it fails, the process proceeds to S174, and the trained segmentation and identification model is saved to the trained segmentation and identification model 68 and deployed for use in processing the new image 60. In the remaining cases, the process returns to S152, one or more additional ground truth images 66 are collected or selected, and the process is repeated.
[0083] The new image is processed by the segmentation and identification subsystem 14, which segments and identifies the target structure or material (e.g., tissue). As described above, the segmentation and identification subsystem 14 comprises four modules: a pre-processor 46, a segmenter 48, a structure identifier 50, and a result evaluator 52. The pre-processor 46 verifies the validity of the input image (a medical image in this embodiment), including whether the information of the medical image is complete and whether the image is damaged. Also, the pre-processor 46 even reads the medical image into the image data 60 within the system 10. The images can be of different formats, and the pre-processor 46 is configured to be able to read images in a format having major relevance, which includes DICOM. Also, since there may be a plurality of trained segmentation and identification models 16 available, the pre-processor 46 can be configured to determine which of the trained segmentation and identification models 16 should be used for performing segmentation and identification based on the information extracted by the pre-processor 46 from the image. For example, in one scenario, two classifier models can be trained, the first classifier model being for segmenting and identifying the radius from a wrist CT scan, and the second classifier model being for segmenting and identifying the tibia from a leg CT scan. When a new CT scan in DICOM format is to be processed, the pre-processor 46 extracts the scanned body part information from the DICOM header and decides, for example, to use the radius classifier model.
[0084] Segmenter 48 and structure identifier 50 each segment and identify a target structure or material in the image using a model selected from the trained segmentation and identification model 68. Then, the results of the segmentation and identification are automatically evaluated by result evaluator 52. In this embodiment, the evaluation performed by result evaluator 52 involves the following: verifying the results in relation to one or more predefined conditions or parameters of the target structure or material (e.g., the acceptable range of one or more structures, or the dimensions of the material, or the volume of the structure or material, etc.) to determine whether the segmentation and identification results are clearly insufficient. Result evaluator 50 outputs a result indicating whether the segmentation and identification results are actually sufficient.
[0085] The automatic evaluation performed by result evaluator 52 can be enhanced manually. For example, the results of the segmentation and identification and / or the evaluation of result evaluator 52 can be displayed to a user (e.g., a doctor) for further evaluation, and the automatic evaluation can be refined.
[0086] Images for which the segmentation and identification results are rejected are used as additional training images for retraining the segmentation and identification model. If result evaluator 52 approves the segmentation and identification results, the segmentation and identification results are output to result output 56 and / or user interface 18 via I / O interface 54 (e.g., on the display of a computer or mobile device). In some embodiments, the segmentation and identification results are used as input for further quantitative analysis. For example, if the radius is segmented and identified by system 10 from a wrist CT scan, attributes such as the volume and density of the extracted radius can be passed to another application (executed locally or remotely) for other further analysis.
[0087] A convolutional neural network for segmentation and identification according to an embodiment of the present invention is generally shown at reference numeral 180 in FIG. 6 and is in the form of a deep learning segmentation and identification model. The deep learning segmentation and identification model 180 includes a contracting convolutional network (CNN) 182 and an expanding convolutional network (CNN) 184. The contracting or downsampling CNN 182 processes the input image into feature maps that reduce their resolution through various layers. The expanding or upsampling CNN 184 processes these features through layers that increase the resolution and eventually generates a mask for segmentation and identification.
[0088] In the example of FIG. 6, the contracting CNN 182 has four layers 186, 188, 190, 192, and the expanding CNN has four layers 194, 196, 198, 192, but it should be noted that other numerical values can be used for the number of each layer. The bottom layer (i.e., layer 192) is shared by the contracting and expanding networks 182, 184. Each layer 186,..., 198 has an input and a feature map (illustrated as a hollow box in each layer). The input of each layer 186,..., 198 is processed by a convolution process and converted into a feature map (illustrated by a hollow arrow in the figure). The convolution process or activation function can be, for example, a rectified linear unit (ReLU), a sigmoid unit, or a Tanh unit.
[0089] In the contracting CNN 182, the input of the first layer 186 is the input image 200. In the contracting CNN 182, the feature map on each layer is downsampled into another feature map and is used as the input 202, 204, 206 of each of its next lower layers. The downsampling process 208 can be, for example, a max pooling operation or a stride operation.
[0090] In the extended CNN 184, the feature maps 208, 201, 212 on each layer 192, 198, 196 are each upsampled to separate feature maps 214, 216, 218 and are used within the next layers 198, 196, 194. The upsampling process can be, for example, an upsample convolution operation, a transposed convolution, or a bilinear upsampling.
[0091] To capture the localization pattern, the high-resolution feature maps 220, 222, 224 in each corresponding layer 190, 188, 186 within the contraction CNN 182 are each copied and concatenated to the upsampled feature maps 214, 216, 218 (illustrated as dashed arrows in the figure), and the resulting high-resolution feature map / feature map pairs 220 / 214, 222 / 216, 224 / 218 are used as the input for each layer 198, 196, 194. In the final layer 194 of the extended CNN 184, the feature map 226 resulting from the convolution process is transformed and output as the final segmentation and identification map 228, and this transformation is illustrated as a solid arrow in the figure. The transformation can be implemented, for example, as a convolution operation or sampling.
[0092] In this way, the deep learning segmentation and identification model 180 of FIG. 6 combines the location information from the contraction path with the contextual information in the expansion path, ultimately obtaining general information that combines localisation and context.
[0093] Exemplary implementations related to segmentation and identification within system 10 according to embodiments of the present invention are shown in FIGS. 7A through 7C. Referring to FIG. 7A, an implementation example 240 related to the first segmentation and identification is characterized as an example of a "one-step" implementation. The first segmentation and identification 240 includes an annotation and training phase 242 and a segmentation and identification phase 244. In the former, the original image 246 is subjected to annotation 248 for segmentation and identification, and after training is performed, a segmentation and identification model 250 is generated. In the segmentation and identification phase 244, a new image 252 is subjected to segmentation and identification act 254 according to the segmentation and identification model 250, and then a segmentation and identification result 256 is output.
[0094] Accordingly, in the embodiment of FIG. 7A, the model 250 is trained to perform segmentation and identification simultaneously. For example, when training a segmentation and identification model for the radius and fibula of a wrist CT scan, the training images are annotated, and different values are assigned to the voxels of the radius and fibula at that time. Once the model is ready, any new wrist CT scan 252 is processed by the segmentation and identification model 250, and the result 256 is provided as a map of the radius and fibula in a segmented and identified form.
[0095] A second segmentation and identification implementation example 260 is schematically shown in FIG. 7B. In implementation example 260, segmentation and identification are implemented in two steps. Referring to FIG. 7B, implementation example 260 also includes an annotation and training phase 262 and a segmentation and identification phase 264, where the annotation and training phase 262 includes training of a segmentation model and training of an identification model as separate processes. Thus, annotation 268 is performed on the original image 266 to train the segmentation model 270. Annotation 274 is performed on the segmented image 272 to train the identification model 276. After the two models 270, 276 are trained, any new image 264 is first segmented 280 by the segmentation model 270, and the resulting segmented image 282 is then subjected to identification by the identification model 276, and an identified result 286 is output. For example, by these two steps and in these two steps, all bones in a wrist CT scan can first be segmented from the surrounding material, and different types of bones can be differentiated and identified.
[0096] A third segmentation and identification implementation example 290 is schematically shown in FIG. 7C. In this implementation example, the segmentation model achieves very accurate results, and the identification can be performed using a heuristic-based algorithm instead of a machine learning based model. Referring to FIG. 7C, the implementation example 290 includes an annotation and training phase 292 and a segmentation and identification phase 294. The original image 266 is annotated 268 for segmentation and used to train the segmentation model 270 in the annotation and training phase 292. In the segmentation and identification phase 264, a new image 278 is segmented 280 using the segmentation model 270 to yield a segmented image 282, and identification 284 is performed using a heuristic-based algorithm, and the identified result 286 is output. For example, once accurately segmented 280 using the segmentation model 270, a wrist CT scan can identify different bones, for example, by volume calculation and differentiation.
[0097] For example, in a wrist HRpQCT scan, after bone segmentation, the radius can be easily distinguished and identified from the ulna by calculating and comparing bone volumes. This is because the radius is larger than the ulna of the same subject.
[0098] Three exemplary deployments of system 10 are schematically shown in FIGS. 8A-8C, each as follows: on-premises deployment; cloud deployment; and hybrid deployment. As shown in FIG. 8A, in the first deployment example 300, system 10 is deployed locally (although it may be in a distributed manner). Any data processing (such as model training and image segmentation and identification) is done by one or more processors 302 of system 10. The data is encrypted and stored in one or more data storages 304. The user interacts with system 10 (e.g., annotating an image or inspecting the results) via user interface 18.
[0099] As shown in FIG. 8B, in the second deployment example 310, system 10 is deployed within encrypted cloud 312, except for user interface 18. User interface 18 is located locally. The communication 314 between user interface 18 and encrypted cloud 312 is encrypted.
[0100] As shown in FIG. 8C, in the third deployment example 320, system 10 is partially deployed within encrypted cloud service 322, while local part 324 with data storage 326, some processing power holders (e.g., processors) 328, and user interface 18 is located locally. In such a deployment example, usually, most of system 10 (except for user interface 18) is deployed within cloud service 322, but data storage 326 and processing power holders 328 are sufficient to support user interface 18 and communication 330 between local part 324 and cloud service 322. The communication 330 between local part 324 and cloud service 322 is encrypted.
[0101] Those skilled in the art should realize that many changes can be made without departing from the scope of the present invention, and in particular, note that specific features of the embodiments of the present invention can be used to bring about further embodiments.
[0102] Note that even if a reference to the prior art is present in this specification, such reference does not constitute an admission that the prior art forms part of the common general knowledge in any country.
[0103] In the appended claims and the foregoing detailed description of the invention, unless the context requires otherwise by express or necessary implication, the terms "comprising", "comprises" (third person singular present tense), "comprising" and the like are used in an inclusive sense (i.e., in the sense of specifying the presence of the stated features), but do not preclude the presence or addition of further features in various embodiments of the invention.
Description of Reference Numerals
[0104] 10 System 12 Annotation and Training Subsystem 14 Segmentation and Identification Subsystem 16 Trained Segmentation and Identification Model 18 User Interface 20 GUI 32 Segmentation and Identification System 34 Processor 36 Memory 66 Ground Truth Image 106 Annotation Tool 118 Annotation Data 182 Reduction Network 184 Expansion Network 242 Annotation and Training 244 Segmentation and Identification 262 Annotation and Training 264 Segmentation and Identification 292 Annotation and Training 294 Segmentation and Identification 302 Processor 304 Data Storage Unit 312 Encrypted Cloud Service 322 Encrypted Cloud Service 324 Local Portion 326 Data Storage Unit 328 Processor
Claims
1. An image segmentation system, the system comprising: A training subsystem configured to train a segmentation machine learning model using annotated training data comprising images associated with respective segmentation annotations to generate a trained segmentation machine learning model; A model evaluator; A segmentation subsystem configured to perform segmentation on structures or materials within an image using the trained segmentation machine learning model; The model evaluator is configured to evaluate a segmentation machine learning model by: (i) Controlling the segmentation subsystem to segment at least one evaluation image associated with existing segmentation annotations using the segmentation machine learning model, thereby generating a segmentation for the annotated evaluation image; (ii) Forming a comparison between the segmentation of the annotated evaluation image and the existing segmentation annotations; If the comparison indicates that the segmentation machine learning model is qualified, deploying or releasing the trained segmentation machine learning model for use; And is configured to do so by: The system includes an annotation subsystem configured to form at least one of the annotated training images from unannotated images or partially annotated images, the forming being done by generating one or more candidate image annotations for the unannotated image or partially annotated image, receiving an input identifying one or more portions of the one or more candidate image annotations, and forming the annotated training image from at least the one or more portions. An image segmentation system.
2. The system according to claim 1, wherein the system is configured to deploy or release the model for use if the segmentation of the annotated evaluation image and the existing annotation match within a predetermined threshold.
3. In the system according to claim 1 or 2, when the segmentation of the annotated evaluation image and the existing annotation do not match within a predetermined threshold, the system is configured to continue training the model.
4. In the system according to any one of claims 1 to 3, the training subsystem includes a step of receiving (i) an image and an annotation for the image, and (ii) a score associated with the annotation, the score indicating the degree of success or failure of the segmentation machine learning model for segmenting the image, where a higher score indicates failure and a lower weighting indicates success, the system including a step of receiving the score and a step of retraining or refining the segmentation machine learning model using the image and the annotation, including weighting the annotation of the image according to the score.
5. In the system according to any one of claims 1 to 4, the training subsystem generates the segmentation machine learning model by refining or modifying an existing segmentation machine learning model.
6. In the system according to any one of claims 1 to 5, the system further includes an identification subsystem, and the annotated training data further includes identification annotations. a) Whether the segmentation machine learning model is a segmentation and identification machine learning model; or b) The training subsystem is further configured to train an identification machine learning model using (i) the annotated training data after each image has been segmented by the segmentation machine learning model and (ii) the identification annotations.
7. In the system according to any one of claims 1 to 6, the training subsystem includes a model trainer, and the model trainer is configured to train the segmentation machine learning model using machine learning to determine the classification of each pixel / voxel of the image.
8. A computer-implemented image segmentation method, the method comprising: Training a segmentation machine learning model using annotated training data each comprising an image and a segmentation annotation, to generate a trained segmentation machine learning model; Evaluating the segmentation machine learning model by: (i) Segmenting at least one evaluation image associated with an existing segmentation annotation using the segmentation machine learning model, thereby generating a segmentation for the annotated evaluation image; (ii) Forming a comparison between the segmentation of the annotated evaluation image and the existing annotation; And if the comparison indicates that the segmentation machine learning model is qualified, deploying or releasing the trained segmentation machine learning model for use; The method includes forming at least one of the annotated training images from an unannotated image or a partially annotated image, the forming being done by: generating one or more candidate image annotations for the unannotated image or the partially annotated image; receiving an input identifying one or more portions of the one or more candidate image annotations; and forming the annotated training image from at least the one or more portions.
9. The method according to claim 8, wherein: (a) If the segmentation of the annotated evaluation image and the existing annotation match within a predetermined threshold, deploying or releasing the model for use and / or (b) If the segmentation of the annotated evaluation image and the existing annotation do not match within a predetermined threshold, continuing to train the model.
10. The method according to claim 8 or 9, wherein the training comprises: Receiving: (i) An annotated image for use in retraining or refining the segmentation machine learning model; (ii) A score associated with the annotation, the score indicating the degree of success or failure of the segmentation machine learning model for segmenting the image, where a higher score indicates failure and a lower weighting indicates success, and receiving the score; weighting the annotated image according to the score when retraining or refining the segmentation machine learning model. A method comprising the step of **Claim 11** The method according to any one of claims 8 to 10, comprising the step of generating the segmentation machine learning model by refining or modifying an existing segmentation machine learning model. **Claim 12** In the method according to any one of claims 8 to 11, (a) the annotated training data further comprises identification data, and the method comprises training the segmentation machine learning model as a segmentation and identification machine learning model; or (b) the annotated training data further comprises identification annotations, and the method comprises training an identification machine learning model using (i) the annotated training data after each image has been segmented by the segmentation machine learning model and (ii) the identification annotations. A method **Claim 13** The method according to any one of claims 8 to 12, comprising the step of training the segmentation machine learning model using machine learning to determine the classification of each pixel / voxel of the image. **Claim 14** The method according to any one of claims 8 to 13, wherein the trained segmentation machine learning model is configured to perform segmentation on a structure or material within an image, the structure or material including bone, muscle, fat, or other biological tissue. **Claim 15** An image annotation system, a) an input unit for receiving at least one unannotated or partially annotated image; b) an annotator configured to form each annotated training image from an image that has not been annotated or has been partially annotated, the forming being generating, for the image that has not been annotated or has been partially annotated, one or more candidate image annotations; receiving an input for identifying one or more portions of the one or more candidate image annotations; forming the annotated training image from at least the one or more portions, the annotator; An image annotation system comprising. **Claim 16** An image annotation method, comprising: a) receiving or accessing an image that has not been annotated or has been partially annotated; b) forming each annotated training image from the image that has not been annotated or has been partially annotated, the forming being generating, for the image that has not been annotated, one or more candidate image annotations; receiving an input for identifying one or more portions of the one or more candidate image annotations; forming the annotated training image from at least the one or more portions, the step; A method comprising. **Claim 17** Computer program code configured to implement the method according to any one of claims 8 to 14 and 16 when executed by one or more processors. **Claim 18** A computer-readable medium comprising the computer program code according to claim 17.
Citation Information
Patent Citations
Method and device for classifying object of image, and corresponding computer program product and computer-readable medium
JP2017062778A
Image processing apparatus, imaging device, image processing method and program
JP2018007078A
Method and system for simultaneous scene analysis and model fusion for endoscopic and laparoscopic navigation
JP2018522622A
Convolutional neural network for segmentation of medical anatomical images
US20180240235A1
Nonvolatile semiconductor memory device and method for manufacturing same
US9589974B2