Image segmentation

The use of tissue-specific loss functions and a 3D U-Net CNN with attribute-based tools enhances medical image segmentation efficiency and accuracy, facilitating rapid and collaborative surgical planning.

GB2639595BActive Publication Date: 2026-04-14HOLOCARE AS
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
HOLOCARE AS
Filing Date
2024-03-18
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Current methods for medical image segmentation, such as those used in CT and MRI images, are time-consuming and prone to errors, requiring manual intervention by experts like radiologists, and lack efficient collaboration tools for multiple users.

Method used

A method involving a neural network trained with tissue-specific loss functions and a convolutional neural network (CNN) for medical image segmentation, utilizing a 3D U-Net architecture, combined with attribute-based smart paintbrush tools for semi-automatic segmentation.

Benefits of technology

Reduces segmentation time from hours to minutes, improves accuracy, and enables collaborative surgical planning across geographically dispersed teams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000002_0001
    Figure 00000002_0001
Patent Text Reader

Abstract

Training a neural network for medical image segmentation by: selecting a loss function based on a tissue type represented in a set of medical images; training a convolutional neural network using the
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND The present invention relates to segmentation of medical images, for example computed tomography (CT) images or magnetic resonance imaging (MRI) images. A patient (whether human or animal) may be scanned by a computed tomography (CT) scanner. The scanner may use X-rays or positron emission tomography (PET) to obtain a series of two-dimensional (2D) images of the patient. Each 2D image represents a slice through the patient. Each pixel of each 2D image represents tissue density in a slice. A stack (that is, sequence or series) of two-dimensional CT images may be converted by a computer into a three-dimensional (3D) image comprising voxels. Each voxel of the 3D image represents a small volume. The 3D image may be used for planning a surgical operation; for example, a patient may have a damaged or diseased liver or a damaged or diseased heart. To facilitate use of a 3D image (comprising voxels), the 3D image may be segmented by a computer. Segmentation is a process in which a computer recognises various portions of a 3D image. For example, for liver tissue, the segmentation process may recognise distinct portions of a 3D image such as: (i) liver tissue, (ii) blood vessels in the liver, and (iii) holes within the parenchyma of the liver. The segmented 3D image may be used for planning a surgical operation, such as removing a tumour. A surgical procedure may be planned so that the procedure, for example, avoids large blood vessels. Conventionally, it may take an expert (such as a radiologist) about 4 to 6 hours to manually segment an image. It would be desirable to reduce the time taken for 2 segmentation to, for example, 1 hour or, even better, to a few minutes, or less. It would also be desirable to improve the quality of a segmentation, for example by reducing the proportion of a 3D image that is incorrectly segmented. It would also be desirable to allow several users (such as clinicians or radiologists) to collaborate in the planning of a surgical procedure. For example, two or more spatially or geographically separated users may wish to collaborate in planning a surgical procedure. The inventors have realised that one problem with applying machine-learning techniques to segmentation is that different tissue types have different characteristics. The inventors have realised that using different loss functions, when training a neural network, can improve image segmentation. US 8,346,695 relates to a method of performing oilfield operations for an oilfield, the oilfield having a subterranean formation. The method includes collecting a first volume data set of seismic data and a second volume data set of seismic data, co-rendering a visually-melded scene directly from the first volume data set and the second volume data set, displaying the visually-melded scene comprising a visualized geobody, where the visualized geobody corresponds to a portion of the first volume data set and the second volume dataset, and selectively adjusting the oilfield operations based on the visualized geobody. BRIEF DESCRIPTION OF INVENTION According to an aspect of the present invention, there is provided a method of training a neural network for medical image segmentation, the method comprising the steps of: receiving data indicative of one or more medical images; receiving data indicative of a tissue type of the one or more medical images; selecting a loss function based on the tissue type; and training a convolutional neural network, based on the selected loss function. In another aspect of the present invention, there is provided a method of tissue segmentation, the method comprising the steps of: receiving data indicative of one or more medical images; receiving data indicative of a tissue type of the one or more medical images; selecting a trained convolutional neural network, based on the tissue type; and using the trained convolutional neural network that has been selected, to perform segmentation. In another aspect of the present invention, there is provided a data-processing apparatus. An aspect of the present invention relates to a segmentation editor that may be implemented as a web-based segmentation application, thereby allowing clinicians to undertake a segmentation process. The segmentation editor may semi-automatically create 3D segmentation from stacks of CT images. Another aspect of the present invention is an attribute-based smart "paintbrush"; the paintbrush is an editing tool that helps clinicians to perform segmentation. FIGURES Figure 1 shows a system for training a convolutional neural network, according to an embodiment. Figure 2 shows a system for performing segmentation based on medical images, according to an embodiment. Figure 3 shows a system that provides an attribute-based smart paintbrush, according to an embodiment. Figure 4 shows a chain of attribute processing, according to an embodiment. Figure 5 shows a branched chain of operations, according to an embodiment. Figure 6 shows a denoised brush system that uses a curvature flow filter for segmentation, according to an embodiment. Figure 7 shows an example (left) of a CT image, and (right) an attribute volume with a denoised image of the CT image, according to an embodiment. Figure 8 shows an example (left) of a threshold segmentation performed on a CT image, and (right) threshold segmentation performed on an attribute volume of the CT image, according to an embodiment. Figure 9 shows a vessel segmentation method, according to an embodiment. Figure 10 shows (left) a CT image, and (right) vessel segmentation performed on a vessel likelihood attribute volume, according to an embodiment. Figure 11 shows a bone brush tool for determining regions of an image that contain bone structures, according to an embodiment. Figure 12 shows an output of a watershed filter, according to an embodiment. Figure 13 shows voxel classes after discriminating based on average class voxel values, according to an embodiment. Figure 14 shows a result of a paintbrush segmentation performed on a bone likelihood attribute, according to an embodiment. Figure 15 shows a method for processing a discrete attribute volume consisting of classes, according to an embodiment. Figure 16 shows an example of a discrete attribute volume with classes, according to an embodiment. Figure 17 shows an example of segmentation of liver, gall bladder, pancreas, kidneys and spleen, according to an embodiment. DESCRIPTION When processing medical images, segmentation is the process of partitioning medical images into multiple regions (sets of pixels) that represent anatomical structures. A goal of segmentation is to simplify and / or change the representation of an image (or images). The image or images may be labelled and therefore easier to analyse. By undertaking segmentation on a set of medical images, the collection of 2D images, can be converted to 3D volumes and graphically rendered for study or treatment planning. Embodiments of the present invention may use an inference algorithm to produce a "likely" (that is, probable) organ segmentation, based on 2D CT images. Embodiments may use machine learning techniques to produce a labelled file comprising data points within a 3D voxel space (x, y, z) that are representative of, for example, the liver. A convolutional neural network that may be used by the present invention is the 3D U-Net. Il-Net networks are fully convolutional networks that are useful for medical image segmentation. The 3D ll-Net consists of multiple convolutional layers that look at progressive levels of detail. Each convolutional layer moves a filter across the image to learn patterns and structures which result in a valid segmentation output, given enough training. Figure 1 shows a system 100 fortraining a neural network, and shows: 110 a machine learning model undergoing training, 115 an optional pre-processing step, 120 a loss function, as an input to the machine learning model, 130a - 130d labelled slices (4 are shown) from a CT scanner, 140a a labelled region of slice 130a, 150 a parameter set, indicative of the trained neural network, and 160 a tissue-type input to the machine learning model. Figure 1 shows a machine learning model 110 that is undergoing training. A series of labelled 2D slices 130a - 130d of image data are presented to the machine learning model 110 that is undergoing training. The machine learning model 110 that is undergoing training may be a convolutional neural network (CNN). The CNN may be a 3D ll-Net neural network. An optional pre-processing step 115 is described below in connection with Figure 2 (which shows a pre-processing step 215 that is similar to the pre-processing step 115). The loss function 120 acts (as those skilled in the art will appreciate) as a guide to the machine learning model 110 that is undergoing training. In the case of a 3D ll-Net neural network, the loss function 120 measures how far off the neural network’s predictions are from a target. (The target is a correct segmentation of the medical image.) The machine learning model 110 that is undergoing training makes predictions on what the segmented image should like. The loss function 120 calculates how different these predictions are from a real segmented image. Then, using this difference information, internal parameters of the machine learning model 110 that is undergoing training are adjusted to make better predictions and thus minimise the difference between its predictions and the actual segmented image. For the liver, binary weighted cross entropy may be used as the loss function 120. This may be calculated by assessing the amount of entropy (that is, uncertainty) that exists between the known probability distribution of the ground truth, and the probability distribution of the inferred model. For the heart, a Sorensen-Dice coefficient (also known as a Dice score) may be used as the loss function 120. The Dice score counts how many points that are present I missing within an inferred segmentation are also present / missing in a ground truth segmentation. Figure 1 shows 4 slices 130a - 130d but more, or fewer, slices may be presented to the machine learning model 110 that is undergoing training. Slice 130a is shown as being labelled with data 140a. (The other slices, such as slices 130b - 130d, may also be labelled with respective label data.) The label data 140 may represent a segmentation that has previously been performed by a clinician, for example a radiologist. In other words, the label data 140 represents a 'ground truth' that is used to teach the machine learning model 110 that is undergoing training. In some cases, the machine learning model 110 that is undergoing training may receive data (for example, from an MRI machine) that is already in a 3D format, such as voxels. In such cases, the voxels may be labelled with label data 140. The label data 140 teaches the machine learning model 110 a segmentation that has previously been performed by a clinician. In some embodiments, the label data 140 may be generated by a machine learning model. The tissue-type input 160 may be used to inform the machine learning model 110 that is undergoing training what type of tissue is represented by the slices 130. For example, the slices 130a - 130d may represent liver tissue. Other slices (not shown) may represent heart tissue. Each slice may be associated with respective data that indicates a tissue type of the slice. The inputs 160 may be numerical data (for example, heart tissue may have a value of "57" whereas liver tissue may have value of "86") or text data, for example in the ASCII format. In some embodiments, some or all of the slices 130 may include a tissue-type input 160. For example, as those skilled in the art will appreciate, an example of a format that is used for medical images is the digital imaging and communications in medicine (DICOM) format. The DICOM format may include metadata, such as the tissue type of the image(s). Examples of other medical image formats are: Analyze, neuroimaging informatics technology initiative (NlfTI) and MINC. When the machine learning model 110 has been trained, the trained machine learning model 110 may be outputted as parameter set 150. The parameter set 150 may be a set of numerical weights for various parameters of the trained machine learning model 110. Figure 1 shows a single loss function 120. In some embodiments, two or more loss functions 120 may be used. For example, a first loss function 120 may be used for a first tissue type (for example, heart tissue), resulting in a first parameter set 150. A second loss function 120 may be used for a second tissue type (for example, liver tissue), resulting in a second parameter set 150. The training data (such as the slices 130 or MRI voxel data) may be split into two or more categories. For example, only heart training data may be provided to the first loss function 120, resulting in the first parameter set 150. For example, only liver training data may be provided to the second loss function 120, resulting in the second parameter set 150. After a neural network has been trained, the trained neural network may be used to perform inference. As those skilled in the art will appreciate, various hyper-parameters may be set to control the model structure and training procedure. For example, the following hyper-parameters may be used: (1) Learning rate details the network's rate of learning. This parameter may be set based on common standard values found in research and literature reviews. (2) Batch size sets the amount of data in each batch processed by the neural network. This parameter may be set based on computational limits and efficiency. (3) A decay parameter reduces learning rate over time to help the neural network learn finer details. This parameter may be set based on common standard values found in research and literature reviews. Figure 2 shows an embodiment 200 that may use a trained neural network to perform inference and thereby perform image segmentation, and shows: 150 a parameter set, indicative of a trained neural network, 160 an optional tissue type input, 210 medical image data that is to be segmented, 215 a pre-processing step, that may be performed by the computer server 220, 217 a post-processing step, 220 a computer server, 230 a processor of the computer 220, 240 segmented voxel data, generated by the processor 230, 250 a remote computer, that is remote from the server 220, 260 commands from the remote computer 250 to the server 220, 270 a rendered image, representing the segmented voxel data 240. Figure 2 shows an embodiment 200 that may be used by a user (such as clinician, not shown) to segment medical images 210. Optional pre-processing 215 of the medical images will be discussed in more detail below. Optional post-processing 217 that may be performed on the segmented images generated by the computer server 220 will be discussed in more detail below. Figure 2 shows an optional tissue type input 160 to the computer server 220. Based on the tissue type input 160, the server 220 may automatically select an appropriate parameter set 150. The tissue type input 160 may be metadata in a medical image 210. Alternatively, a clinician may arrange for the server 220 to be loaded with a parameter set 150 that is suitable for liver tissue, or with a different parameter set 150 that is suitable for heart tissue. A computer server 220 is used to perform inference (using a trained neural network, based on a parameter set 150) and thereby produce voxel data 240 indicative of a segmented medical image. For example, for a liver, each voxel (the voxels are not shown) may indicate whether a voxel contains: (i) non-liver tissue, (ii) liver parenchyma, or (iii) a blood vessel. Medical image data 210 is presented to the computer server 220. The computer server 220 comprises a processor 230 that processes the medical image data 210. As will be explained below in more detail, one function performed by the processor 230 is to convert slices of medical image data 210 into voxel data 240. The processor 230 then uses a trained neural network (which is represented by the parameter set 150) to infer and classify the voxels. The processor 230 may comprise, for example, one or more central processing units (CPUs), and / or one or more graphics processing units (GPUs), and / or one or more field programmable gate arrays (FPGAs). For clarity, Figure 2 does not show random access memory (RAM) or storage (such as a hard disk drive, HDD, or a solid-state disk, SSD). The parameter set 150 may be stored in RAM of the server 220. In some embodiments, two or more parameter sets 150 may be stored in the server 220. For example, a first parameter set 150 may be stored, for heart tissue. A second parameter set 150 may be stored, for liver tissue. The user may send commands 260 from the remote computer 250 to the server 220. As one example of a command 260, the user may inform the server 220 whether to use a first parameter set 150 or a second parameter set 150. As another example of a command 260, and as will be explained below in more detail, the user may use a digital "paintbrush" to tell the server 220 that a region of the voxel data 240 is a particular tissue type, for example blood vessel or parenchyma. The processor 230 may then use a trained neural network (as represented by the parameter set 150) to automatically segment a region (or regions) of the voxel data 240 as belonging to the same tissue type. The processor 230 may use the parameter set 150 to automatically produce a "likely" (that is, plausible) segmentation of the voxel data 240. Each voxel may be labelled with data (not shown) to indicate a segmentation parameter for that voxel. For example, each voxel may be labelled with data that indicates whether the voxel is a blood vessel or is parenchyma. The user may then review and, if necessary, amend, the segmented voxel data 240. As those skilled in the art will appreciate, processing voxel data can be challenging for a computer. For example, voxel data 240 that has a three-dimensional (3D) array of 1024 x 1024 x 1024 voxels has a total of 1,073,741,824 voxels. Thus, the computer server 220 may be a relatively high-performance computer device. The remote computer 250 may be a relatively low-performance computer device that communicates with the server 220 over a communications network, for example the internet. The server 220 may render the voxel data 240 to form a two-dimensional (2D) image. The server 220 may then send the rendered image 270 to the remote computer 250 for display to the user. A benefit of rendering the 2D image on the server 220 is that it reduces the quantity of data 270 that is transferred to the remote computer 250. For example, the rendered image 270 may have 1920 x 1080 = 2,073,600 pixels which is dramatically fewer than the 1,073,741,824 voxels. Figure 2 shows the computer server 220 and the remote computer 250 as separate devices; the server 220 may be a cloud-based device and the computer 250 may be a web browser running on a laptop. In an alternative embodiment (not shown), the user may directly communicate with the server 220. For example, the user may use a keyboard (not shown) and a mouse (not shown) of the server 220. Figure 2 shows a single remote computer 250. In alternative embodiments, two or more remote computers 250 may receive the rendered image 270. This can allow users who are spatially or geographically separated to participate in the planning of a surgical procedure. For example, two or more clinicians may use respective laptop computers 250 to view the rendered image 270. As will be appreciated, the server 220 performs relatively intense mathematical processing, allowing relatively low performance laptops 250 (or desktop computers that are less capable than the server 220) to display the rendered image 270. The pre-processing 215 (and the pre-processing 115) will now be explained in more detail. For liver, prior to the use either within an inference procedure (Figure 2) or for training a model (Figure 1), the images may be resampled and reoriented so that the voxel size and axis orientation are uniform. The images may then be normalized, with different methods for each organ model, and the segmentations may be binarized where applicable. Some specific pre-processing steps for the liver are: (1) Resampling: as different images can have different voxel sizes, the images may be resampled to give a uniform voxel size. The liver may be resampled to a voxel size of 1 in each dimension, although it is possible to resample to other voxel sizes. (2) Reorientation: To cope with images potentially having different orientations, images may be re-oriented to specific axes codes before being fed into the neural network. Images typically contain information about their orientation in their image meta-data, this enables re-orienting images to a standard orientation. Liver images may be reorientated to axes codes R, A, S. This means that the first voxel axis is left to right, the second voxel axis is posterior to anterior and the third voxel axis is inferior to superior. (3) Hounsfield thresholding and normalization: for liver, thresholding based on the intensity values in a CT image may be performed to clip intensity values that are far away from the values usually associated with liver intensity. For example, a CT image may be clipped so that, after clipping, the image has a minimum threshold of-50 units on the Hounsfield scale and a maximum threshold of 600 units on the Hounsfield scale. Intensity values under and above these thresholds are therefore clipped up or down to then range of -50 to 600. Values in the range of -50 to 600 may then be normalized to a new range of 0 to 1, respectively. (4) Label binarization: if the ground truth segmentations (for the training procedure of Figure 1) are divided into more than 1 class (excluding background), a label binarization procedure may be performed. This means that all the segmented classes are combined into a single class, allowing the neural network (of Figure 1) to train on segmenting only the background and a single segmentation class. (5) Data augmentation: the CT images may optionally be augmented with additional data. (6) Patch sampling: if the CT images are too large to fit in the memory of a computer (whether for Figure 1 or Figure 2), the images may need to be split up into smaller patches which are then sampled from, and fed into, the neural network. For inference, each patch outputted from the neural network is assembled to make up a resulting segmentation having the same resolution as the original input CT image. (6) Masking: if there is a previous segmentation data available, this can be used crop the input volume. This can reduce the processing time and increase accuracy of the model. (7) Padding: padding may be applied to keep volumes at a consistent size. The volumes may be padded with the value -1024, which is the Hounsfield value of air (background). (8) Adaptive Histogram Equalization: histogram equalization may be performed to modify the contrast in an image. An adaptive histogram equalization image filter is a superset of two or more contrast-enhancing filters. This may be done for some images to highlight specific features, for example vessels (such as blood vessels). (9) Curvature Flow Filter: Curvature flow filtering is an anisotropic diffusion method that may be used for smoothing images while preserving edges. A curvature flow filter may be applied in situations where maintaining clear edges is necessary, in addition to the benefit of smoothing. For example, a curvature flow filter may be used for blood vessel models. Some specific pre-processing steps for the heart are: (1) Reorientation: to cope with images potentially having different orientations, images may be re-oriented to specific axes codes before being fed into the neural network. Heart images may be reorientated to axis codes Radial, Axial, Sagittal (RAS). (2) Brightness normalization: an image-level normalization for heart images may be performed based on the lowest voxel value in the image. For example, values below 0 may be clamped to 0. Values above 2048 may be clamped to 2048. The range of 0 to 2048 may then be normalised to 0 to 1. (3) Binary Label: ground truth segmentations (that may be used for the training of Figure 1) are only one class excluding background. This means the blood volume is labelled as one segmented class, making the neural network train on segmenting only the background and a single segmentation class. (4) Patch sampling (binary balanced): if the images are too large to fit in memory (whether the memory of the computer of Figure 1 or Figure 2), the images may need to be split up into smaller patches which are then sampled from, and fed into, the neural network. For inference, each of the patches outputted from the neural network is assembled to make up a resulting segmentation having the same resolution as the original input file. This process can also ensure a balanced amount of ground truth versus background. (5) Spatial window size: a pre-set spatial window size may be applied to ensure the 17 images are of a fixed size. A spatial window may be used for accurate sampling and / or to ensure that the computer (of Figure 1 or Figure 2) does not use more memory than is available. Post-processing 217 may be performed on voxel images that have been segmented by the server 220. Some examples of post-processing steps are: (1) Connectivity: unconnected masses may be removed using a connectivity algorithm. The connectivity algorithm may prioritise a largest collection of voxels. All unconnected islands may be assigned as noise. The algorithm used for this removal process may be the Connected Components 3D (CC3D) algorithm. As those skilled in the art will appreciate, CC3D achieves this by checking if sets of voxels (islands) are connected to the largest collection of voxels and removing the islands if they are not connected to the largest collection of voxels. To mitigate excessive voxel removal, the CC3D algorithm may be supplemented with a failsafe mode: if the number of voxels removed is higher than a pre-set threshold then the CC3D post-processing process may be stopped, returning to the original segmentation (before the use of the CC3D algorithm was attempted). A resampling (a new connectivity algorithm is run) is then performed to counteract the resampling done during pre-processing. This may be necessary to produce an output matching the original input image dimension to validate the process. (2) Hole Filling Algorithm: in some cases, unidentifiable voxel volume (holes) may appear within the largest collection of voxels (such as liver parenchyma). These holes in the parenchyma are often due to lesions. Identification and marking of lesions may be performed manually and may be the responsibility of a medical specialist. There is a risk that the segmentation algorithm may incorrectly classify holes as background. To compensate for this risk, holes within the parenchyma itself are filled using a filling algorithm and classified as parenchyma. (3) Small Island Algorithm: where appropriate, a small island removal algorithm is applied. This may be for cases where the CC3D algorithm is not used, such as in a vessel model. This is due to vessels sometimes being disconnected due to large slice thickness or incorrect classification in other parts of a vessel tree. To compensate for this, the small island removal algorithm is applied to remove islands below a pre-set threshold. As a prelude to the discussion of Figures 3-17, some terms will be defined: (1) Brush: a 2D or 3D shape (typically a circle or sphere) by which the user can interactively fill or erase voxels in a segmented volume. (2) Medical volume: a medical 3D voxel image, produced for example by CT or MRI. (3) Segmented volume: a discrete voxel volume with labels. The labels may specify which voxels define organ structure or pathology of interest. (4) Attribute volume: a 3D voxel volume which is derived from a medical image. This includes, but is not limited to, denoising, thresholded values, gradient magnitude, likelihood volumes and other metrics. Embodiments of the present invention can aid a user in more easily segmenting parts of a 3D model by using an attribute "brush" tool (that is, "paintbrush"). Tools according to the present invention can also reduce the number of errors in a segmentation by auto-generating the bulk of the segmentation for the user, thereby allowing a user to focus on key parts which may then be validated by a clinician. Figure 3 shows a system that provides an attribute-based smart paintbrush, according to an embodiment. As those skilled in the art will appreciate, the "paintbrush" is a computer-implemented tool that may be used by a user to modify an image that has been generated by, for example, the computer server 220. The paintbrush may have 3 modes: (1) Fill I erase sphere. This paints in 2D or 3D inside a sphere, regardless of the underlying CT I MRI data. (2) Threshold-based paint I erase. This will paint or erase segmented voxels based on whether the underlying medical voxel intensity is inside a user-defined interval. (3) Threshold-based connected body paint. This behaves like (2), but may perform a flood fill inside a sphere starting at a paintbrush location, only growing into regions that are spatially connected with the centre of the brush. Attribute volumes may be calculated automatically on high-performance hardware (for example, the server 220) in the cloud and then downloaded to a client device 250 for interactive segmentation. The specific attribute volumes depend on the type of the organ structure. For example, a bone segmentation will require other attributes than a vessel segmentation. The user may view the original CT image so that their understanding of the anatomy is not biased by the derived attribute volume. A benefit of using an attribute volume as a basis for the paintbrush is that the attribute volume is better suited for semi-automatic flood fill or threshold-based fill than the original CT volume as it better describes the boundaries of the structure of interest. Figure 4 shows a chain of attribute processing, according to an embodiment. Figure 4 shows two operations and two attribute volumes but there may be more than two. Figure 5 shows a branched chain of operations, according to an embodiment. As shown, operation 3 receives multiple (in this case, two) inputs. Other embodiments may receive 3 or more inputs. Figure 6 shows a denoised brush system that uses a curvature flow filter for segmentation, according to an embodiment. A denoised brush may allow a user to do segmentation on an image which has been denoised. A benefit is that random image noise will not cause issues for threshold-based segmentation. An optimal denoising parameter may be chosen for a specific organ type. A curvature flow filter may be used to reduce noise in the image. A curvature flow filter may apply the partial differential equation with the contour curvature * V to every image channel f. This filter reduces random noise in the image which is typically present in CT scans with thin slices. When segmenting e.g. liver vessels where there is a low signal I noise ratio, a plain threshold-based segmentation is problematic because the vessels will be equally noisy. When instead doing the threshold-based segmentation based on the denoised attribute volume, the vessels are smooth. Figure 7 shows an example (left) of a CT image and (right) an attribute volume with a denoised image of the CT image, according to an embodiment. Figure 8 shows an example (left) of a threshold segmentation performed on a CT 21 image, and (right) threshold segmentation performed on an attribute volume, according to an embodiment. Figure 9 shows a vessel segmentation method, according to an embodiment. For vessel segmentation, a threshold flood fill paintbrush may be used on a vessel likelihood attribute. The vessel likelihood is an attribute where each voxel indicates the likelihood of that voxel representing a vessel. The paintbrush's threshold may be set so that the paintbrush paints only voxels that have a high likelihood of being a vessel. The first attribute volume may be a denoised image as described above. The local contrast attribute volume may be calculated, for each voxel, based on calculating a biased mean in the neighbourhood. The local contrast is then calculated as the difference between the voxel value at each voxel and the biased mean in the neighbourhood. Figure 10 shows (left) a CT image and (right) vessel segmentation performed on a vessel likelihood attribute volume, according to an embodiment. Figure 11 shows a bone brush for determining regions of an image that contain bone structures, according to an embodiment. This attribute volume has two branches (see also Figure 5) which are eventually joined into one operation resulting in the bone likelihood attribute. One branch produces a denoised image; the other branch calculates a gradient magnitude attribute. The denoised image and the gradient magnitude attribute volumes are used as an input for a bone likelihood attribute calculation. Based on the gradient magnitude volume, a morphological watershed image filter may be performed, yielding a discrete voxel volume that is segmented into multiple classes based on the gradient magnitude. Figure 12 shows an output of a watershed filter, according to an embodiment, with a unique colour assigned to each class. Figure 13 shows voxel classes after discriminating based on average class voxel values, according to an embodiment. A discrimination of the classes, where classes that are not likely to represent bone structures are discarded, may be performed using the following steps: (1) For each class from the water-shedding volume, go through each voxel in that class and calculate the average voxel value from the same voxel index in the denoised CT image. (2) Create a new voxel volume with the same resolution. (3) For each class where the average voxel value is above a given threshold (e.g. where the Hounsfield value (radio opacity) is high, generate a new class identifier and assign this identifier to all voxels from the same class. For each class where the average voxel value is below this threshold, assign the value 0 to that voxel. In some areas parts of the bone structure may be missing because of a low gradient magnitude. Therefore, an additional “snap to bone edge” may be performed, using the following steps: (1) For each voxel that has a non-zero class from a previous step, do a connected body flood fill. (2) If a neighbouring voxel has a voxel value from the denoised CT attribute which corresponds to a reasonable bone Hounsfield value and also has not been assigned to a class, assign the class from (1) to that voxel and continue the search to the next neighbouring voxels. (3) Continue (2) until all connected neighbours are visited. Figure 14 shows a result of a paintbrush segmentation performed on the bone likelihood attribute, according to an embodiment. Figure 15 shows a method for processing a discrete attribute volume consisting of classes, according to an embodiment. Organs in the abdomen often have a uniform, but similar radio opacity. Doing a traditional intensity-based flood fill segmentation will typically result in over-segmentation. In a generic abdominal organ brush, an attribute volume is generated which delineates organ structures based on gradients. This provides a paintbrush which limits itself to areas that are likely to belong to the organ of interest. As shown by Figure 15, a discrete attribute volume consisting of classes may be generated as follows: (1) A curvature flow filter is used to generate a denoised image. (2) A gradient magnitude volume is generated. (3) The watershed volume is generated. Figure 16 shows an example of a discrete attribute volume with classes, according to an 5 embodiment. Note that each organ is assigned a low number of classes, making it efficient and easy to do semi-automatic segmentation with a paintbrush. Figure 17 shows an example of segmentation of liver, gall bladder, pancreas, kidneys and spleen, according to an embodiment. A user may interactively paint on the attribute 10 volume. With a flood fill based on the attribute classes, a segmentation of multiple organs can be achieved in minutes. The reference numerals in the claims are for the convenience of the relevant searching authority and are not to be regarded as limiting the scope of the claims. 15 The abstract as filed is hereby included by reference. 12 08 24

Claims

1. A method of training a convolutional neural network (110) for medical image segmentation, the method comprising the steps of:receiving data of one or more medical images (130a - 130d);selecting one of a plurality of loss functions (120) based on a tissue type associated with the one or more medical images (130a - 130d); andtraining the convolutional neural network using the selected loss function (120) to perform medical image segmentation.

2. A method according to claim 1, further comprising the step of: outputting a parameter set (150) indicative of the trained convolutional neural network.

3. A method according to claim 1 or 2, further comprising the step of: receiving data indicative of a tissue type (160) of the one or more medicalimages.

4. A method according to any of claims 1 to 3, wherein the loss function (120) comprises at least one of: binary weighted cross entropy, and Dice score.

5. A method according to any one of claims 1 to 4, further comprising the step of pre-processing (115) the data of the one or more medical images (130a - 130d).

6. A method according to claim 5, wherein the pre-processing (115) comprises at12 08 24least one of: resampling, reorientation, Hounsfield thresholding and normalization, label binarization, data augmentation, patch sampling, masking, padding, adaptive histogram equalization, curvature flow filtering, reorientation, brightness normalization, patch sampling, and using a spatial window size.

7. A method according to any of claims 1 to 6, wherein the neural network (110) implements a 3D ll-Net neural network.

8. A method of tissue segmentation comprising the steps of:receiving a tissue type input (160) indicating a tissue type associated with one or more medical images;selecting, based on the tissue type input, one of a plurality of stored parameter sets (150) indicative of a trained neural network (110) trained using a selected loss function associated with the tissue type,receiving data indicative of the one or more medical images (210); and using the trained neural network (110) in accordance with the selected parameter set to perform segmentation of the medical image data (210).

9. A method according to claim 8, further comprising the step of:pre-processing (215) the data of the one or more medical images (130a -130d).

10. A method according to claim 9, wherein the pre-processing (215) comprises at least one of: resampling, reorientation, Hounsfield thresholding and normalization, label binarization, data augmentation, patch sampling, masking, padding, adaptive12 08 24histogram equalization, curvature flow filtering, reorientation, brightness normalization, patch sampling, and using a spatial window size.

11. A method according to any one of claims 8 to 10, further comprising the step of:post-processing (217) the segmented medical image data (240).

12. A method according to claim 11, wherein the post-processing (217) comprises at least one of: using a connectivity algorithm to remove unconnected masses, using a hole-filling algorithm, and using a small island removal algorithm.

13. A method according to claim 12, wherein the post-processing (217) comprises using the Connected Components 3D (CC3D) algorithm, and wherein if a number of voxels removed is higher than a pre-set threshold then the CC3D post-processing process is stopped.

14. A method according to any of claims 8 to 13, wherein the segmentation is performed on a computer server (220), and wherein the computer server (220) is operable to transfer a rendered image (270) to a remote computer (250).

15. A method according to any of claims 8 to 14, wherein the segmentation is operable to use two or more chained operations (Fig 4).

16. A method according to any of claims 8 to 15, wherein the segmentation is operable to use two or more branched operations (Fig 5).

17. A method according to any of claims 8 to 16, wherein the segmentation is operable to use at least one of: a curvature flow filter, and a morphological watershed.

18. A computer system (100, 200, 220), wherein the computer system comprises data stored on a non-transitory medium, wherein the computer system (100, 200, 220) is programmed to implement the method of any one of claims 1 to 17.

Citation Information

Patent Citations

  • CNN-based image processing

    EP4053752A1