Method and system of segmentation of medical images
The method and system co-register PET and CT scans with a bias factor and machine learning for accurate lung lesion segmentation, addressing misregistration and variability, enhancing tuberculosis research and treatment monitoring.
Patent Information
- Application Number
- PCT/ZA2025/050039
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-05
- Filing Date
- 2025-08-05
- Publication Date
- 2026-02-12
AI Technical Summary
Misregistration and low resolution in PET-CT scans complicate lung lesion segmentation, particularly in tuberculosis research, leading to labor-intensive manual delineation and variability in lung field segmentation, which is exacerbated by organ movement and breathing artifacts.
A method and system for co-registering PET and CT scans, followed by combining them with a bias factor, and using a machine learning model for accurate organ segmentation, addressing misregistration and leveraging both modalities' information.
Enables automated, accurate, and reproducible lung lesion segmentation with reduced computational requirements, suitable for diseases like tuberculosis, sarcoidosis, and malignancies, providing robust quantification and treatment monitoring.
Smart Images

Figure ZA2025050039_12022026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND SYSTEM OF SEGMENTATION OF MEDICAL IMAGES
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority from United Kingdom patent application number 2411477.9 filed on 5 August 2024, which is incorporated by reference herein.
[0004] FIELD
[0005] This disclosure relates to a method of segmentation of medical images. In particular, this invention relates to segmentation using both Positron Emission Tomography (PET) scan images and Computed Tomography (CT) scan images.
[0006] BACKGROUND
[0007] Biomedical imaging includes Positron Emission Tomography (PET) images and Computed T omography (CT) images. These images, or scans, have various medical applications. CT scans provide images of a patient’s body organs, tissues, and bones. PET scans provide a functional imaging technique that uses radioactive substances known as radiotracers to visualize and measure changes in physiological activities in the body.
[0008] Positron Emission Tomography - Computed Tomography (PET-CT) is a dual-modality imaging technique that combines two sequential scans of the same subject, providing both functional (radiotracer uptake in target cells) and structural (anatomical details) information. PET-CT is widely used for detecting and monitoring cancer and infectious lesions.
[0009] A problem when combining the information from CT and PET scans is misregistration when the anatomical features on PET and CT scans do not properly align. There are three main reasons for misregistration. Firstly, the patient may move during the scan. Secondly, there is organ movement that is captured in a PET scan. Lastly, PET scans have inherent low resolution.
[0010] Medical image segmentation can be used for quantifying areas of abnormality. As an example, areas of abnormality may be lesions in a patient’s lungs. The most time-consuming and least reproducible aspect of lung-lesion segmentation is delineating the lung fields (i.e., creating a lungmask), which is required before quantification prior to analysis. This is usually done manually by readers on CT scans and is prone to inter- and intra- reader variability, as well as variability between patients, and time points. The lungs are closely related and situated to soft tissue structures in the mediastinum and abdomen. Also, the lungs are not static. The shape and structure of the lungs change during breathing cycles and the heart border moves with each beat.
[0011] Lung segmentation of tuberculosis (TB) lesions have additional challenges: they are heterogenic, diffuse, often of extreme high or low density, may extend into the chest wall, and often cause extensive damage to the lungs and distort the anatomy. Combining PET and CT scans help clinicians identify and segment the lungs and lung lesions in pursuit of T uberculosis (TB) research, specifically tracking extent and nature of lesions, and the treatment response of TB. These scans assist in tracking the effectiveness of the treatment administered to a patient over a period by analysing the lung area, and identifying active TB lesions before, during, and after treatment.
[0012] The process of identifying the lungs on these images or scans can be extremely labour intensive, requiring many hours of clinicians’ time to manually draw, or correct, these delineated segmentations.
[0013] The preceding discussion of the background is intended only to facilitate an understanding of the present disclosure. It should be appreciated that the discussion is not an acknowledgment or admission that any of the material referred to was part of the common general knowledge in the art as at the priority date of the application.
[0014] SUMMARY
[0015] In accordance with an aspect of the invention there is provided a computer implemented method for combining a three-dimensional positron emission tomography (PET) scan and a three- dimensional computed tomography (CT) scan of a body into a combined three-dimensional image for segmentation, comprising: co-registering each of the PET scan and the CT scan as an input scan with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image, the co-registering including: extracting metadata from the input scan and the reference scan relating to positional and orientation data; matching an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan; matching dimensions, orientation, and axes of the input scan and the reference scan; and outputting a co-registered PET scan and a co-registered CT scan; and combining the co-registered PET scan and the co-registered CT scan into a single 3D image with a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan.
[0016] The PET scan and CT scan of the method may encompass an organ for organ segmentation.
[0017] The method may include using the combined PET and CT images to train a machine learning model, and inputting the combined PET and CT scan images as a single input channel into the machine learning model for image segmentation.
[0018] Combining the co-registered PET scan and the co-registered CT scan may include; normalizing the co-registered PET scan and multiplying the normalized PET scan by the PET bias factor weight; normalizing the co-registered CT scan and multiplying the normalized CT scan by the CT bias factor weight; and, adding the normalized and biased scans to output the combined PET and CT scan.
[0019] The co-registration may include reorienting one or both of the input scan and reference scan to be parallel or perpendicular to a world axis and to each other. Reorienting one or both of the input scan and reference scan to be parallel or perpendicular to a world axis and to each other may include; when the input scan and the reference scan are parallel or perpendicular to a world axis, reorientating one of the input scan or reference scan to fit the direction of the other by transpositioning of a scan or reversing of a scan axis. Reorienting one or both of the input scan and reference scan to be parallel or perpendicular to a world axis and to each other may include; when the input scan and the reference scan are not parallel or perpendicular to a world axis, resampling one of the input scan or the reference scan to fix a grid orientation.
[0020] Matching the origin points of the input scan and the reference scan may include for each of planes x, y and z; cropping by removing voxels at the start of one scan in a plane; or padding by adding voxels at the start of one scan in a plane, whilst taking spacing and direction into account.
[0021] The method may include matching the input scan and the reference scan in real world dimensions including padding or cropping voxels on a corner of a scan opposing the origin point and is repeated for each direction x, y and z.
[0022] The method may include cropping the co-registered PET scan and a co-registered CT scan before combining the co-registered PET scan and the co-registered CT scan into a single 3D image. The method may include resizing the co-registered CT scan and / or the co-registered PET scan before combining with the other scan. The reference scan of the method may be a cropped portion of the CT scan. The reference scan of the method may be a CT scan of an area of interest.
[0023] The method may include isolating an area of interest of a CT scan by: applying a user-defined threshold and identifying continuous slices of the scan at a start and an end of the scan below the user-defined threshold; excluding slices that form a boundary from the identified continuous slices; and removing identified slices from the scan leaving the area of interest.
[0024] The method may include co-registering at least one of the PET scan and the CT scan with the reference scan in the form of an isolated area of interest (520, 620) of a CT scan.
[0025] The method may include co-registering the combined PET and CT scan and an isolated area of interest of a CT scan to output a cropped combined PET and CT scan.
[0026] The method may include: co-registering an input scan of a PET scan and a reference scan of an isolated area of interest (620) to obtain a co-registered PET scan in the form of cropped PET scan; and co-registering an input scan of a CT scan and a reference scan of the cropped PET scan to obtain a co-registered CT scan in the form of a cropped CT scan; and wherein combining the co-registered PET scan and the co-registered CT scan combine the cropped PET scan and the cropped CT scan.
[0027] The method may include: co-registering an input scan of a CT scan and a reference scan of an isolated area of interest to obtain a co-registered CT scan in the form of cropped CT scan; and co-registering an input scan of a PET scan and a reference scan of the cropped CT scan to obtain a co-registered PET scan in the form of a cropped PET scan; and wherein combining the coregistered PET scan and the co-registered CT scan combine the cropped PET scan and the cropped CT scan.
[0028] The method may include training a machine learning model on a combined positron emission tomography (PET) scan and computed tomography (CT) scan of a portion of a body for segmentation, comprising: receiving manually segmented combined PET and CT scans; and iteratively training a 3D segmentation machine learning model on combined images to receive as input a combined PET and CT scan to output a segmentation image that accurately delineates the organ or anatomical structure of interest.
[0029] The method may include: initially training the machine learning model on the received manually segmented combined images; applying the machine learning model (822) to create masks for further training images; and manually correcting the further training images for additional training.
[0030] In accordance with an aspect of the invention there is provided a system for combining a three- dimensional positron emission tomography (PET) scan and a three-dimensional computed tomography (CT) scan of a body into a combined three-dimensional image for segmentation, comprising: a processor including a graphical processing unit accelerator; and a memory, said memory containing instructions executable by said processor, to execute functions of components including: co-registering each of the PET scan and the CT scan as an input scan with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image, the co-registering including: extracting metadata from the input scan and the reference scan relating to positional and orientation data; matching an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan; matching dimensions, orientation, and axes of the input scan and the reference scan; outputting a co-registered PET scan and a co-registered CT scan; and combining the co-registered PET scan and the co-registered CT scan with a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan.
[0031] The system may include a system for training the 3D segmentation machine learning model using combined PET and CT scans. The system may include a combined-image trained 3D segmentation machine learning model having as input the combined PET and CT scan as a single channel.
[0032] In accordance with an aspect of the invention there is provided a computer program product for combining a three-dimensional positron emission tomography (PET) scan and a three- dimensional computed tomography (CT) scan of a body into a combined three-dimensional image for segmentation, comprising a computer-readable medium having stored computer-readable program code for performing the steps of: co-registering each of the PET scan and the CT scan as an input scan with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image, the co-registering including: extracting metadata from the input scan and the reference scan relating to positional and orientation data; matching an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan; matching dimensions, orientation, and axes of the input scan and the reference scan; outputting a co-registered PET scan and a co-registered CT scan; and combining the co-registered PET scan and the co-registered CT scan with a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan.
[0033] In accordance with an aspect of the invention there is provided a method of training a machine learning model for combined positron emission tomography (PET) scan and computed tomography (CT) scans of a portion of a body, comprising: receiving manually segmented combined PET and CT scans; and iteratively training a co-registration trained 3D segmentation machine learning model to receive as input a combined PET and CT scan to output a segmented scan.
[0034] The method of training a machine learning model may include; initially training the machine learning model on the received manually segmented images; applying the machine learning model to create masks for further training images; and manually correcting the further training images for additional training.
[0035] The computer-readable medium may be a non-transitory computer-readable medium and for the computer-readable program code to be executable by a processing circuit.
[0036] Embodiments of the technology will now be described, by way of example only, with reference to the accompanying drawings.
[0037] BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In the drawings:
[0039] Figure 1 is a flow diagram illustrating an example embodiment of the described method;
[0040] Figures 2A to 2D are flow diagrams illustrating example embodiments of aspects of the described method;
[0041] Figure 3 is a schematic diagram of a voxel arrangement for the purposes of definitions;
[0042] Figure 4 is a flow diagram illustrating an example embodiment of a first implementation of the described method;
[0043] Figure 5 is a flow diagram illustrating an example embodiment of a second implementation of the described method;
[0044] Figure 6A to 6C are flow diagrams illustrating example embodiments of a third, fourth and fifth implementation of the described method;
[0045] Figure 7 is a flow diagram illustrating an example embodiment of an aspect used to isolate an area of interest;
[0046] Figure 8A is a flow diagram illustrating an example embodiment of machine learning processes and models that may be used in embodiments of the described method and system;
[0047] Figure 8B is a flow diagram illustrating an example embodiment of a method of training machine learning models of the described method and system;
[0048] Figure 9 is a block diagram illustrating an example embodiment of the described system; and
[0049] Figure 10 illustrates an example of a computing device in which various aspects of the disclosure may be implemented.
[0050] DETAILED DESCRIPTION WITH REFERENCE TO THE DRAWINGS
[0051] A method and system of segmentation of medical images is provided that may receive as inputs a three-dimensional CT scan and a three-dimensional PET scan of a whole or part of a human or animal body. The three-dimensional CT and PET scans may each include stacked slices of two- dimensional images, forming a three-dimensional scan. The described method and system address the problem of misregistration between the CT and PET images to generate a combined PET and CT scan that is used to isolate regions of interest (segmentations), while using information from both modalities.
[0052] Direct transfer of segmented volumes from CT to PET is problematic due to motion artifacts (caused by heartbeat and breathing) and differences in imaging acquisition, leading to spatial misalignment that prevents transfer from regions of interest (ROIs) obtained from CT to PET. Misalignment between PET and CT scans is addressed by the described methods and systems.
[0053] In addition, an advanced deep learning model is provided that accurately performs 3D anatomical segmentation of PET images. This technology addresses the challenges of automated anatomical segmentation of PET scans by integrating both imaging modalities and leveraging a trained convolutional neural network (CNN). The described methods and system may be applied to medical imaging involving solid organs affected by motion artifacts. These organs may include, as examples, lungs, liver, spleen, mediastinal lymph nodes, central nervous system, and bone marrow. Furthermore, the methods and systems may be applicable to imaging of diseases with diffuse or widespread lesions. The imaging may be applied to determine extent and characteristics of lesions or treatment monitoring. Example applications include imaging in diseases such as tuberculosis, disseminated fungal infections, granulomatous disease (such as sarcoidosis and granulomatosis with polyangiitis), and malignancies (such as infiltrative metastases and lymphomas).
[0054] Referring to Figure 1 , a flow diagram (100) shows an example embodiment of the described computer implemented method for combining a three-dimensional PET scan and a three- dimensional CT scan of a body into a combined three-dimensional image for segmentation. The method refers to a scan of a body and this may be a part or whole of a human or animal.
[0055] The method inputs (111) a PET scan and a reference scan into a co-registration method (112). In a parallel branch, the method inputs (121) a CT scan and a reference scan into a co-registration method (122). The described method co-registers each of the PET scan and the CT scan as an input with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image to output a co-registered PET scan and a co-registered CT scan. The method then combines (131) the output co-registered PET scan (117) and the output coregistered CT scan (127) with a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan.
[0056] The combined PET and CT scan may be input (132) as a single input channel into a machine learning model for image segmentation. The single input channel has more relevant information than each of the single PET scan or CT scan. Having only one input channel to a machine learning model means that there are fewer weights required in the model and more accurate training can be provided.
[0057] The combined PET and CT scan may include an organ of interest and is used for training a convolutional network model to perform organ segmentation. As an example application, a training process may include Fluorodeoxyglucose PET-CT scans predominantly used from patients with pulmonary tuberculosis. As such, these included lesions with a vast spectrum of extent, and morphological and inflammatory characteristics, improving the robustness of the model. The purpose may be to quantify the burden and nature of lung lesions disease on PET images for research or clinical applications. This can be performed, at a single or multiple timepoints to measure changes over time. The technology is especially suited for diseases with widespread or diffuse lesions such as tuberculosis, sarcoidosis, cystic fibrosis, pneumoconiosis, and interstitial lung disease.
[0058] With the training of the model completed, the method of combining the PET and CT scan may be included in a software pipeline in which the trained model performs segmentation of organs on PET-CT images. The pipeline produces as outputs accurate regions of interests for respectively the PET and the CT scan. This allows reproducible and informative downstream analysis through sub-segmentation, quantification, texture analysis, and radiomics. The method and system are highly adaptable and can be applied to different organ systems with minimal modifications, expanding its potential clinical and research applications.
[0059] The method of automated segmentation offers several benefits to other automated segmentation techniques in that it can segment organs on PET with high accuracy, and it overcomes the challenge of misregistration with reference CT scans. This is attained with relatively low computational system requirements.
[0060] A method of co-registering (112, 122) two scans is described and is a process used in various aspects and implementations of the described method. The co-registering method has two scans as input referred to as an input scan and a reference scan. The co-registering method is carried out for each of the PET scan and the CT scan as an input scan with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image.
[0061] In one aspect, a PET scan and a CT scan are used as the input scan and the reference scan so that they are co-registered to each other. In this aspect, only a single instance of the coregistration method may be required, as shown in the embodiment of Figure 4 and Figure 5. Two outputs are obtained in the form of each of the co-registered PET scan and the co-registered CT scan.
[0062] In another aspect as shown in the embodiment of Figure 6A, the input scans of the PET scan and the CT scan may both be co-registered against a reference scan in the form of a third image in the form of a CT mask of an area of interest. In this aspect, two instances of the co-registration method are carried out, one for each of the PET scan and the CT scan.
[0063] In an even further embodiment as shown in Figure 6B and 6C, the input scans may be one of the CT scan and cropped PET scan, or the inputs may be a PET scan and a cropped CT scan.
[0064] In a further aspect, the input scan may be the combined PET and CT scan and the reference scan may be a CT mask of an area of interest, as shown in the embodiment of Figure 5. In this aspect, a single instance of the co-registration method is carried out.
[0065] The PET scan or CT scan may be an image in the form of a scan file including multiple stacked slices of two-dimensional images forming a single three-dimensional image. Each two- dimensional image may contain rows and columns of values. The stacked two-dimensional images may be arranged into an array structure for computations applied to the scans.
[0066] The co-registering method (112, 122) may be carried out by scaling and fitting the PET image and the CT image using matrix calculation and manipulation of the values of the arrays of the PET and CT scans.
[0067] The co-registering method (112, 122) as used in the various embodiments, includes extracting (113, 123) metadata from the input scan and the reference scan relating to positional and orientation data. The co-registering method (112, 122) may include reorienting (114, 124) one or both of the input scan and reference scan to be parallel or perpendicular to a world axis and to each other. This step may not be needed if the scans are taken in a given orientation to each other.
[0068] The co-registering method (112, 122) includes matching (115, 125) an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan. The matching of the origin point may include matching reference points that are offset from the origin point by a known amount of pixels or dimensions.
[0069] The co-registering method (112, 122) includes matching (116, 126) dimensions, orientation, and axes of the input scan and the reference scan. This may be referred to as matching in real world dimensions.
[0070] After the co-registration method, the co-registered PET scan and the co-registered CT scan should both take up the same physical / reference space; however, the scans may still be different sizes and rescaling of one of the scans is required to essentially bring the scan from a projected real space to a “matching” representation in the form of an array, which can be directly used in the described implementations.
[0071] The method outputs (117, 127) a co-registered PET scan and a co-registered CT scan and combines (131) the co-registered PET scan and the co-registered CT scan with a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan.
[0072] After scaling and cropping, the co-registered PET scan and co-registered CT scan may be combined (131) using a percentile addition of matrices. The combined scan is biased towards the PET scan. The bias ratio may be based on a ratio determined through experimentation, where a range of PET:CT ratios between 90% and 50% were tested, with the most effective ratio found for the normalisation being approximately 75%.
[0073] The method may create a combined PET and CT image by scaling and fitting the PET image and the CT image using matrix calculation and manipulation. These matrix calculations and manipulations may include the use of axes permutation, matrix multiplication, affine transformations, reversing array direction, scaling, normalization, padding and cropping of the PET / CT pixel array.
[0074] Referring to Figure 2A, a flow diagram (200) shows an example method of the aspect of combining (131) the co-registered PET scan and the co-registered CT scan. The method may receive as inputs the co-registered PET scan (201) and the co-registered CT scan (202) as output from the co- registration method (112, 122). The method may also have a PET bias factor (203) as an input. The method may normalize (210) the co-registered PET scan and may multiply (211) the normalized PET scan by the PET bias factor weight (203). The method may normalize (212) the co-registered CT scan and multiply (213) the normalized CT scan by the CT bias factor weight in the form of (1-the PET bias factor (203)). The method may add (214) the normalized and biased scans to output (215) the combined PET and CT scan.
[0075] The weighting is biased towards PET scans. The bias is affected by the normalization of the two scans and therefore may vary depending on the normalization. In one embodiment, the PET and CT scans are linearly normalized to between 0 and 1 . As an example, for the CT scan, the lower bound may be -1024HU and an upper bound may be 1024HLI. For the PET scan, a lower bound may be 0 standard uptake value (SUV) units and an upper bound may be 5 SUV units. These bounds may be changed depending on requirements.
[0076] Various PET biases have been tested between 90% and 50% PET bias. All the PET biases worked better than with no co-registration and it is likely that it would still work beyond those bounds. However, the most effective bias tested was 75% towards the PET scan.
[0077] An alternative method of biasing may be required if normalizing between different values other than 0 and 1 , (e.g-1 to 1) or if different normalization is used for PET and CT.
[0078] Further details are now provided of an example embodiment of the steps of the co-registration method (112, 122) with reference to Figures 2B to 2D.
[0079] The co-registration method (112, 122) includes extracting (113, 123) metadata from the scan image folders of the PET scan and the CT scan. The metadata may be used both in preprocessing, for SUV conversion of the PET scan, and to determine each scan's orientation, spacing, and origin. The step of extracting (113, 123) metadata may include extracting required parameters for a scan image and assigning them to a variable in memory. The parameters required to determine the pixel arrays' physical position size and orientation are stored differently depending on the file type.
[0080] For Digital Imaging and Communications in Medicine (DICOM) files, the headers Image Position(patient), Pixel Spacing, Spacing Between Slices, and Image Orientation(patient) may be extracted. For Neuroimaging Informatics Technology Initiative (NIFTI) files pixdim[*](*=0,1 ,2,3), qoffset_*(*=x,y,z), quatren_* (*=b,c,d); srow_* (*=x,y,z) may be extracted.
[0081] A programming interface such as SimplelTK™ may be used to read an image for use with programming languages. SimplelTK™ stores the image spatial descriptors as origin, spacing, size, and direction. The spatial descriptors describe the position, spacing, size, and orientation of the pixel array with respect to a real-world axis and coordinates. The real-world may be either set to a known anatomical position or a fixed point within the scanner.
[0082] The extracted spatial descriptors are used to match the input scan and reference scan ensuring they are oriented the same and take up the same physical region of space. This may include reorientating the pixel array as described in Figure 2B, which may require transposing or reversing one or more axes if the PET and CT scan directions do not match. With this reorienting, the spacing and origin of the scan may change to describe the new pixel array. If one or both of the scans are not orthogonal and axis-aligned to the world axis, then a resampling of the scan is required. One method of resampling is through the use of the library SimplelTK™, using either the PET or CT image as a fixed / reference image and the other as a moving image, an Affine transformation scales rotates, sheers, and translates the moving image onto the geometric grid space of the reference image using interpolation a new 3D pixel array is created which matches the orientation, offset, spacing, and size of the fixed scan. Alternatively, one may resample both images to a fixed, or partially fixed geometric space, using a combination of any one of a specified direction, spacing, and size. The co-registering method (112, 122) may include reorienting (114, 124) one or both of the input scan and the reference scan to be parallel or perpendicular to a world axis and to each other. Reorienting may include, when the input scan and the reference scan are parallel or perpendicular to a world axis, reorientating one of the input scan or reference scan to fit the direction of the other by trans-positioning of a scan or reversing of a scan axis. Reorienting may also include, when the input scan and the reference scan are not parallel or perpendicular to a world axis, resampling one of the input scan or the reference scan to fix a grid orientation.
[0083] An example embodiment of a reorienting method (114, 124) is shown in Figure 2B. The coregistering method (112, 122) receives an input scan and a reference scan as input, referred to as scan A and B arrays (221). The metadata from scans A and B (222) may be obtained from the extraction (113, 123) steps. The arrays (221) and metadata (222) are input into a check (223) to determine if DA and DB are equal, where DA and DB may each be a direction cosine matrix of the two input scans (221), respectively. If the direction cosine matrices are equal, then the reorienting method (114, 124) may output (224) the array B without any change. If the direction cosine matrices of scans A and B are different, then a second check (225) may be completed to determine if they are parallel or perpendicular to a world axis. If the second check (225) is true, then array B may be reorientated (228) to fit array A, forming a new array B. The new array B may be a newly generated array, or the previous array B with a new structure after the reorientation (228) of B. The new array B, along with a new origin B, a new spacing B, and a new direction B may be output (229). If the second check is false, array B may be resampled (226) to fix the grid orientation of A (226). The new array B, along with a new origin B, a new spacing B, and a new direction B may be output (227).
[0084] Once both pixel arrays are orientated in the same direction, it is essential to ensure they take up the same physical space by removing or padding any areas that do not overlap. This can be done by first ensuring that the origin of the scans match. The co-registering method (112, 122) may then match (115, 125) an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan.
[0085] This is done by determining the difference in the origin between each scan. Then a direction matrix may be used to determine which array axis (i, j, k) corresponds to each world axis (x,y,z) and the direction of the next pixel regarding the world axis. This can be used to determine if the second scan needs to be padded or cropped to fit the first scan. Alternatively, the larger of the two scans may be cropped to fit the smaller. The difference of the origin and the spacing between pixels is used to determine how many voxels need to be either removed or added in each direction for the two origins to match (note this is unlikely to be an exact match due to the pixel size difference and pixel size being indivisible, if the difference is too large scaling and resampling may be implemented).
[0086] Origin coordinates may be defined as the distance of the array's first element to the reference / world origin in terms of the world co-ordinate system. This first element of the scan is the first element of the array and not necessarily the closest point to the real-world origin. Padding may add elements (voxels) to the array and cropping may remove elements (voxels) from the array. Direction is a direction cosine matrix which describes the orientation of each axis relative to the world axes. Spacing is the distance between the centre of two voxels / pixels within a scan.
[0087] The matching (115, 125) of the origin involves moving the origin with respect to the direction of the scan, so that both origins match. Any difference in origin can be used to calculate the physical space between the two origins, using the spacing between voxels as well as the direction of the scan. The number of voxels that need to be removed or added in each direction are calculated and then removed or padded.
[0088] Matching the origin points of the input scan and the reference scan includes for each of planes x, y and z: cropping by removing voxels at the start of one scan in a plane or padding by adding voxels at the start of one scan in a plane; whilst taking spacing and direction into account.
[0089] An example embodiment of a method of matching (115, 125) an origin point of the input scan and an origin point of the reference scan is shown in Figure 2C. The matching (115, 125) of origins of scans receives the same input A array as the reorientation (114, 124) method, and one of the outputs (227, 229) of the reorientation (114, 124) of the scans as array B. This is referred to as scan A and B arrays (231). The match (115, 125) may also receive the associated metadata (232) of scans A and B. The inputs may be checked (233) to see if the product of the origin and direction of A (O_Ai*D_Ai) is less than the product of the origin and direction of B (O_Bi*D_Bi). The check (233) may be performed for each axis (I, j, k). If true, array B may be padded (235) by the difference between the origins. If false, array B may be cropped (234) by removing elements from the start of the array B. The new array B and the new origin B is output (236).
[0090] Once the origins of the scans have been aligned, both the CT scan and the PET scan may be cropped by determining the number of voxels that need to be removed for the CT or PET scan. This may be done by determining the real-life distance between the mask origin and the scan origin. Similar methods that may align the origins of the scans may be used in the above step. The co-registering method (112, 122) may include matching (116, 126) the input and the reference scan in real world dimensions.
[0091] Once a common origin is created the two pixel arrays will have a matching corner in physical space to ensure the whole scan is matching, opposing side of the array to the origin must be matched. One can match these arrays by determining the difference in physical space on each axis, this is done by comparing the array size multiplied by their corresponding spacing in each direction and then determining if the second array needs to be padded or cropped.
[0092] Matching the input scan and the reference scan in real world dimensions may include padding or cropping voxels on a corner or an edge of a scan opposing the origin point and is repeated for each direction x, y and z.
[0093] An example embodiment of a matching of dimensions, orientations, and axes (116, 126) is shown in Figure 2D. The matching (116, 126) may match to real world dimension. The matching (116, 126) may receive the same input A array of the reorientation (114, 124) method, and the output (236) of the match origins (115, 125) as array B. This is referred to as scan / mask A & B arrays
[0094] (241). The matching (116, 126) to real world dimensions may receive the associated metadata
[0095] (242) of scan / mask A & B. A check (243) may be performed to determine if the product of: A_SIZEi x A_SPACINGi is larger than the product of B_SIZEi x B_SPACINGi. The check (243) may be performed for each axis (I, j, k). If true, array B is padded (245) by appending the end of array B. If false, array B is cropped (244) by removing elements at the end of array B. The output (246) may be a new array B.
[0096] The scan world origin coordinates may be determined from the metadata. If the scans are in the same orientation for a scan to match in real space, the origin of each scan should match. Hence, voxels are either removed or added / padded from the plane of each axis until a close match is found. The difference in the physical dimensions of the two scans are calculated, and used to determine the number of voxels which need to be added or removed for the two scans to have a matching real-world size.
[0097] The various described implementations may require resizing the co-registered CT scan and / or the co-registered PET scan before combining with the other scan. The physical space taken up by each scan should match for a seamless combination of the two scans, hence the difference of the physical size in each direction is matched by either removing or padding of the scans until a close match is found. Once the two scans match in physical space, the array sizes must match. Scaling of the images may be carried out using an open-source image processing library for image interpretation (for example, a sickie-image or skimage library). An alternative method may be to scale the images before cropping.
[0098] Referring to Figure 3, a visual explanation (300) of the arrangement of voxels (301) of a scan image is illustrated. A world origin reference (302) is shown with three dimensional axes. A scan image origin (303) is shown with three dimensional axes showing the directions (304) of the scan image. Dashed lines (305) show a spacing of the voxels (301).
[0099] Various implementations of the described method are now described with reference to Figures 4, 5 and 6.
[0100] Figure 4 shows a flow diagram (400) of an example embodiment of a first implementation of the described method. In this implementation, no CT segmentation is required, and the implementation combines a PET scan (401) and a CT scan (402). The described co-registration method (112) of Figure 1 is used to co-register the PET scan (401) with the CT scan (402) to result in a cropped co-registered PET scan (403). The described co-registration method (122) of Figure 1 is used to co-register the CT scan (402) with the PET scan (401) to result in a cropped CT scan (404). In this implementation, the same co-registration method (112 / 122) may be carried out for the co-registration of each of the PET scan and the CT scan as they are referenced against each other.
[0101] The cropped co-registered CT scan (404) may be resized (405) to match the cropped coregistered PET scan (403). The method then combines the cropped co-registered PET scan (403) and the resized CT scan (405) using the scan combining method (131) described in relation to Figure 1 including applying a PET biasing factor. The output of the combining method (131) may be resized (406) and the combined PET and CT scan output (407).
[0102] Figure 5 shows a flow diagram (500) of an example embodiment of a second implementation of the described method. In the second implementation, the same process may be carried out as in the first implementation of using described co-registration method (112) of Figure 1 to coregister the PET scan (501) with the CT scan (502) to result in a cropped co-registered PET scan (503). The described co-registration method (122) of Figure 1 is used to co-register the CT scan (502) with the PET scan (501) to result in a cropped CT scan (504).
[0103] The cropped CT scan is (504) may be resized to form a resized CT scan (505). The resized CT scan (505) may be combined with the cropped PET scan (503) in the scan combining method (131) to form a first combined scan (506).
[0104] An additional process may be used in this implementation to focus on an area of interest such as a lung mask when carrying out a lung segmentation. Although the examples use a lung mask, a mask of any area or object of interest may be used. In Figure 5, a CT lung mask (510) is used and a method for isolating (520) the area of interest is applied. An example embodiment of a method of isolating (520) the area of interest is described below in relation to Figure 7.
[0105] An isolated lung area and the first combined scan (506) are input into a third co-registration method (512), where the two inputs are processed through the steps of; extracting (513) metadata, reorientation (514), origin matching (515), and matching of real-life dimensions (516).
[0106] The output of the co-registration (512) step is a cropped combined scan (530). The cropped combined scan (530) may be resized (531) to form a resized / cropped combined scan (532).
[0107] Figure 6A shows a flow diagram (600) of an example embodiment of a third implementation of the described method. In the third implementation, a CT lung mask (610) is used as the reference scan for the co-registration methods (112, 122) for each of the PET scan (601) and the CT scan (604). The CT lung mask (610) is input to a method of isolation (620) an example embodiment of which is described below in relation to Figure 7.
[0108] The PET scan (601) is co-registered (112) with the isolated area of interest (620) of the CT lung mask (610) to output a co-registered cropped PET scan (607). The CT scan (604) is co-registered (122) with the isolated area of interest of the CT lung mask (610) to output a co-registered cropped CT scan (608). One or both of the co-registered cropped PET scan (607) and the co-registered cropped CT scan (608) may be resized (611 , 612).
[0109] The method then combines the cropped co-registered PET scan (607) and the resized CT scan (608) using the scan combining method (131) described in relation to Figure 1 including applying a PET biasing factor. The output of the combining method (131) may be a combined PET and CT scan output (613) that may be input into a co-registered image model (614).
[0110] Further embodiments of the described method of a fourth implementation and a fifth implementation are shown in the flow diagrams (640) and (660) in Figure 6B and Figure 6C, respectively. Instead of the reference scans in the co-registration methods (112, 122) being the other of the PET scan (501) and the CT scan (502), the cropped CT scan (504) as the output of the coregistration method (122) may be used as the reference scan for the PET scan (501) coregistration method (112). Similarly, the cropped PET scan (503) as the output of the coregistration method (112) may be used as the reference scan for the CT scan (502) co-registration method (122).
[0111] Figure 6B shows an implementation where the cropped PET scan (607) is input into the coregistration method (122) along with the CT scan (604). In this implementation, the co-registration (112) of the PET scan (601) and the isolated (620) area of the CT lung mask (610) must be performed first to obtain the cropped PET scan (607). The steps of resizing (611 , 612) the cropped PET scan (607) and cropped CT scan (612), combining the scans (131), obtaining a combined scan (613) and inputting the combined scan into a co-registered image model (614) remains the same as Figure 6A.
[0112] Similarly for Figure 6C, the cropped CT scan (608) is input into the co-registration method (112) along with the PET scan (601). In this implementation, the co-registration (122) of the CT scan (604) and the isolated (620) area of the CT lung mask (610) must be performed first to obtain the cropped CT scan (608). The steps of resizing (611 , 612) the cropped PET scan (607) and cropped CT scan (612), combining the scans (131), obtaining a combined scan (613) and inputting the combined scan into a co-registered image model (614) remains the same as Figure 6A and Figure 6B.
[0113] In all the implementations, some additional pre-processing of the PET and CT scans may be carried out before the co-registration methods (112, 122). For the CT scan, the pre-processing may convert a DICOM file to a NIFTI file and resizing and normalizing the scan. For the PET scan, the pre-processing may convert the DICOM file to a NIFTI file and may calculate a standard uptake value (SUV) correction factor for the PET slices and may transform the PET into SUV units and may normalise the PET scan.
[0114] In the second, third, fourth and fifth implementations described with reference to Figures 5 6A, 6B, and 6C, isolation (520, 620) of an area of interest is used. Figure 7 shows an example embodiment of a method of isolation (520, 620).
[0115] Isolating an area of interest of a CT scan may be carried out by: applying a user-defined threshold and isolating continuous columns, rows, and slices at a start and end of the scan below the threshold; removing columns, rows, and slices from isolated columns, rows, and slices to create an additional boundary for the area of interest; and removing remaining columns, rows, and slices from the scan, leaving the area of interest remaining. The additional boundary may be required if the edges of the lungs on the PET scan may be outside the area of interest created by the CT lung mask due to misregistration.
[0116] Figure 7 shows a flow diagram of an example embodiment of isolating an area of interest (520, 620) of a CT scan. A scan or mask array (701) and a user defined threshold (702) value are input into a first step. The first step may create (704) three lists of; a column index, a row index, and a slice index.
[0117] The column index list may be populated by all column indices from the scan array whereby all values in a column fall below the user defined threshold (702) value. Similarly for the row index list and slice index list, if all values in the corresponding row or slice are below the user defined threshold (702) value, then the corresponding index value is added to the row index list or slice index list, respectively.
[0118] Once the slice index list is populated with all slice indices where each value in the slice falls below the user defined threshold (702) value, all continuous slice index numbers starting from a start of the array and all continuous slice index numbers starting from an end of the array, are retained in the slice index and the other non-continuous slice index values are removed from the slice index list. The start and end of the array may be at opposing ends of the array. A similar procedure may be performed to maintain or remove indices for the row index and column index lists.
[0119] The three index lists may be used to remove columns, rows and slices from the scan or mask array (701) that correspond to the indices in the lists. Alternatively, a boundary to the isolated area can be maintained, by first removing indices from the index lists, with the addition of the metadata (708) of the scan or mask array (701), as well as a user defined boundary size (710).
[0120] The size of the boundary of the area of interest is proportional to the number of elements removed (712) from the list. This boundary size (710) could be based on the real-world size using the metadata (708). Any remaining indices stored in the lists after removing (712) the boundary indices based on the real-world size, are removed (714) from the scan or mask array (701). A new origin may be calculated (716). An output of an array of the area of interest and a new origin may be output (718). An alternative method to using lists, is to perform the same process by determining a cropping position and use array indexing to crop the array.
[0121] The method may isolate an area of interest of the co-registered, combined image that may be input into a trained 3D segmentation machine learning model (140) to output (142) a PET lung mask which fits the input co-registered pixel array. The output of the model is padded and scaled to match the original PET image. This lung mask can be used to isolate the lungs, creating an anatomical segmentation of the lungs on the PET image.
[0122] The co-registering method has two main uses: it is used to match the PET and CT scan (by cropping and rearranging) so that co-registration can occur; and it is also used to crop the scans to the isolated lung region.
[0123] If the isolation step has been used on the mask before the co-registering method, then it has a similar effect of the isolation step, as it crops the image to fall into the same bounds as the isolated mask. This can also be said if the mask is isolated, then a CT scan is matched to it, and then the PET scan is matched to the CT scan region; the PET scan will now only consist of the region isolation of the mask.
[0124] The described method and system consist of a combination of machine learning models for semantic segmentation of lungs, including severely damaged lungs. The models use a unique combination of 3D ll-Net models to provide lung segmentation for PET and CT scans. The models may be trained on scans from research participants which creates a vast variety of lesions and anatomical distortion. In one example application, the models are trained on research participants with pulmonary tuberculosis. The result is that the models are robust enough to be more accurate when used on CT scans of lungs severely damaged by disease than existing options.
[0125] Using a co-registration algorithm to combine the PET and CT images before applying a trained machine learning model, the method and system create anatomical segmentation of the lungs on PET, overcoming misregistration and low resolution. Fully automated anatomical segmentation facilitates faster, more reproducible downstream analysis, including densitometric, radiomic and texture analysis, as well as further sub-segmentation of specific lesion types.
[0126] Architecture of 3-D U-Net
[0127] The machine learning models may be convolutional neural networks (CNNs). CNNs are a subset of artificial neural networks (ANNs) that may consist of interconnected units, commonly referred to as neurons, as they are inspired by and resemble neurons of the brain. The units may be made up of nodes and edges forming a connected network. The edges may connect nodes together. ANNs may be configured in the form of a layered structure with an input at the first layer and an output provided by the final layer. The layers between the first layer and final layer are hidden layers.
[0128] The input layer may include one or more nodes. An edge may extend from each node. Each edge may be connected to a node in a subsequent hidden or output layer. Each node may include more than one edge that connects the node to a plurality of other nodes in other layers. In some examples, an edge may feed back into a previous node in a preceding layer (a node not in subsequent layers but in a further layer), or to a different node in the same layer.
[0129] The output of a node may be computed by an activation function, which may be a linear or a nonlinear function of the sum of the inputs into each node in each layer. The output value of each node in the preceding layer is multiplied by a weighting value, which determines the strength of each nodes’ output value. Finally, the value that is determined at the node(s) of the final layer is the output of the ANN. For regression type AN Ns, the output may contain only a single node with a value, or many nodes. For classification type ANNs, the output may include multiple nodes, where each node is an output of the probability of a classification type.
[0130] In addition to the weights and activation functions of a regular ANN, a CNN applies a filter (or a kernel) onto a two-dimensional data structure, which may reduce the number of edges between the hidden layers in the neural network. This may in turn reduce the number of weights within the neural network. A convolutional neural network may find application in image-based tasks, where image data may be structured as a two-dimensional data structure. A convolutional neural network may be extended into further dimensions by increasing the dimensions of the filter / kernel to match the number of dimensions of the input data.
[0131] In one embodiment, a 3D ll-Net convolutional neural network is used. A 3D-based architecture has the benefit of having access to more contextual features than a 2D-based architecture. The additional contextual and spatial features provided by a 3D architecture become even more valuable when considering the significant structural variance damaged lungs have, on both PET and CT scans. Additionally, the unclear border regions of PET scans make spatial context crucial.
[0132] Each layer of the CNN may consist of two convolution blocks, each consisting of a convolution with a kernel size of [3x3x3], a batch normalisation and a rectified linear unit activation function. This is followed by max pooling with a kernel size of [2x2x2], Each up convolution consists of a convolution with a kernel size of [3x3x3] and up-sampling with a kernel size of [2x2x2], The final convolution uses a kernel size of [1x1x1] and a sigmoid activation function.
[0133] Software pipeline to combine the models A fully automatic pipeline may be implemented. The pipeline allows the user to input a combined PET and CT DICOM scan. The combined scan may be obtained from combining (131) the coregistered PET and co-registered CT scans. The pipeline may output the CT lung mask, PET lung mask, and cavity mask with minimal or no intervention required.
[0134] Lung segmentation on CT scan
[0135] Pre-processing of the CT image may be carried out. This may involve:
[0136] • removing the less relevant data outside the body using a thresholding method;
[0137] • resizing and down-sampling of the matrix which allows for a uniform input shape of [128,128,128] and yielding a low-resolution version of the CT scan;
[0138] • a linear normalisation with a cut-off maximum of 1024 Hounsfield; and,
[0139] • a linear normalisation with a cut-off minimum of -1024 Hounsfield.
[0140] The uniform input shape may be of sizes other than [128,128,128], Once pre-processing is complete, the low-resolution scan may be run through a 3D ll-net to roughly isolate the lung and surrounding region of the original high-resolution scan. This isolated lung region can then go through similar pre-processing described above. This lung region may then be run through a second model to perform further segmentation of the lungs on CT (also referred to as a lung mask). In the case where segmentation is to be performed on any organ, the CT scan may encompass the entirety of the organ.
[0141] Lung segmentation on PET scan
[0142] Quantitative analysis of medical images for assisted expert decision making has several advantages over qualitative expert reading alone. These include improved reproducibility, more time efficient, more complete assessment, and similar or improved correlation to pathology and function. A critical step in PET-CT analysis is segmentation, which involves delineating specific regions of interest (ROIs) by isolating relevant structures from surrounding tissue. This segmentation is essential for quantitative image analysis, which is increasingly used in research and clinical decision-making. Automated segmentation offers advantages over manual methods, including improved accuracy, efficiency, and reproducibility.
[0143] The application of the method to lung segmentation of a PET scan is used as an example. PET scans generally have low resolution and a significant variance in their raw values, complicating automated lung segmentation. In addition to the low resolution, there are often artefacts caused by movement, which is prevalent around the heart, diaphragm, and lung wall. Misregistration due to movement and the low resolution of structural borders means that a segment created using the CT scan does not accurately fit the PET scan. The method created was used to address the issue of misregistration. In the case where segmentation is to be performed on any organ, the PET scan may encompass the entirety of the organ.
[0144] For a PET scan, the significant variance in the intensity of the pixel data is caused by many factors, including the radiotracer dose varying between scans, the radioactive decay of the scan and the body weight of the scan subject. During pre-processing, to decrease the variation of the scans, a SUV conversion may be done. This assists in correcting the scan for the radiotracer dose given to the patient, the decay of said dose, and the patient's body weight. A linear normalisation may be performed on the scan with a cut-off maximum of 5 SUV. The pipeline then co-registers the PET and CT images to create a combined PET-CT image.
[0145] The combined images are biased towards the PET scan. This provides the model with a clearer distinction between the lung's boundary while preserving and favouring the PET features. The broader lung area of the combined image may then be isolated using the CT lung mask before being accurately segmented using a trained 3D U-Net model.
[0146] There are many outputs that may be produced by the pipeline. Forms of outputs that may be produced include volumes of interest and statistical data. Examples of the volumes of interest are: CT lung mask, PET lung mask, SUV threshold mask, CT hard density tissue mask, CT medium density tissue mask, and CT low density tissue mask.
[0147] The statistical data can include many different first-order parameters, including patient weight, scan size, and time between scans, lung volume on CT, lung volume on PET, cavity volume, SUV mean, SUV max, SUV volume, hard tissue volume, medium tissue volume, soft tissue volume, and total glycolytic activity. The format of the masks produced by the pipeline allows for a vast range of statistical formulas or models to be implemented to allow calculations for densitometry, radiomics and texture analysis.
[0148] In one embodiment, the models are trained on scans of tuberculosis patients and healthy controls. The range of disease states in the training set could, however, allow accuracy in most lung diseases. Further training could allow for sub-segmentation of complex lesions, as was done for cavities, and for other organs. Automating the process of segmenting CT and PET scans can save a considerable amount of time compared to manual segmentation. Alternative methods of automating segmentation of lungs on CT, such as threshold and seed growth, can often miss complex lung lesions, especially when the density ranges and outlines overlap with those of the lungs or bordering anatomical structures. The scan dataset used for training the current models includes scans with a wide variety of lung lesions, structural damage, and parenchymal changes due to the nature of pulmonary tuberculosis disease which results in a model that is robust.
[0149] Currently, the use of machine learning for the segmentation of lungs for PET scans are not widely known, if at all. Segmentations created to fit a linked CT scan generally do not fit the PET component due to misregistration. The current method and system aim to address, at least to some extent, the challenge faced with misregistration low resolution of PET images.
[0150] The segmentation step is needed for downstream quantification of the PET-CT images and analysis. There is a growing need for the quantification of radiological images because it has been shown to aid research findings and clinical decisions by reducing variability, subjectivity, and time required for image interpretation. The software was designed to address lungs damaged by TB but may be used for other lung diseases like cystic fibrosis and interstitial lung disease to determine severity or measure changes over time. It may also be applied to other organs systems.
[0151] The described system and models are intended to be run with minimal graphical memory, meaning most laptops or PCs with about 6GB of video memory should be able to run the model. Further, it may be slotted into other workflows that require anatomical segmentation (for example if facial features need to be removed as part of image anonymisation), 3D printing, or subsegmentation. Further, the pipeline could be re-trained and applied on other organ systems for which PET segmentation may also be a challenge.
[0152] An example embodiment of the method is described below. PET DICOM and CT DICOM files may be used for inputs. The PET DICOM files show metabolic activity and generally have a high contrast. Uncorrected PET DICOM files may be used, and these files may be of a low resolution. The CT DICOM files show radiodensity and may have a high resolution and consequently have a large file size. During pre-processing, several adaptations may be made to the input data. File conversion may be needed to convert the DICOMs to NIFTIs and to extract relevant metadata.
[0153] Maintaining spatial features is important to ensure accurate and meaningful interpretation of the data. Maintaining spatial features in medical imaging involves using a 3D format to represent volumetric data accurately, defining a clear scan origin for spatial registration, ensuring appropriate voxel spacing to prevent distortion, and establishing the correct orientation for accurate anatomical reference. These concepts collectively contribute to the fidelity of spatial information in medical images, supporting precise diagnosis and treatment planning.
[0154] SUV (Standardized Uptake Value) transformation may be included in pre-processing (320). SUV transformation in PET imaging is associated with standardizing images by correcting for variations in tracer dose and accounting for radioactive decay. This standardization enhances the reliability and comparability of PET images, allowing for more consistent and meaningful quantitative analysis in clinical and research settings.
[0155] Normalization is another pre-processing step that may contribute to making scans more uniform for the machine learning models. It may also reduce the processing requirements and complexity of the models, leading to more efficient training and inference. Additionally, the customization of normalization methods per model helps emphasize different features.
[0156] Trimming and resizing are further pre-processing steps that involve removing irrelevant portions and adjusting image dimensions. Proportional considerations are important to preserving spatial relationships and avoiding distortions. Ensuring a uniform size for model input through these processes is needed for optimizing the model's processing power and facilitating effective training and inference.
[0157] A PET / CT co-registration model may be designed to synergistically utilize information from both CT and PET modalities to segment lung structures. The model benefits from the functional insights of PET and the anatomical clarity provided by CT for more accurate and detailed analysis.
[0158] The architecture of models like a 3D U-net leverages spatial information to enhance robustness. The architecture may include mechanisms for adaptive spatial processing, allowing the model to adjust its focus based on the spatial characteristics of the input data.
[0159] The step of training the model may involve employing data augmentation to increase dataset diversity, iterative training for continuous improvement, transfer learning for leveraging pre-trained knowledge, and careful consideration of dataset characteristics to ensure representative and high-quality training data. These strategies collectively contribute to the development of a robust model.
[0160] Data augmentation may be used in training the models. It may involve enlarging the training dataset by applying various transformations to the existing images. The goal is to improve the variety of the dataset, leading to a more generalized and robust model. Augmentation methods may include rotation, cropping, flipping, and noise-generation.
[0161] Iterative training may be used in training the model. It may involve an approach where the training process is conducted in multiple cycles, beginning with hand-segmented scans for initial training. Subsequently, the initial trained model is then used to increase the speed of segmentation for generating further training data by creating (initially imperfect) segments that may be manually corrected faster than creating new segments from scratch.
[0162] Transferred learning may be used to train the model. It focuses on enhancing the initial feature mapping of the model, leading to potentially improved generalization. Knowledge may be transferred from the pre-trained CT model to enhance the performance of the PET co-registration task. This approach may enable the model to use prior knowledge for improved learning and adaptation.
[0163] Training the model on TB-affected lungs with PET and CT data may include a diversity of TB damage levels, a wide range of lung conditions, both PET and CT modalities, and variability in imaging characteristics. The dataset should include diverse examples of TB-affected lungs, covering a spectrum of damage severity. Also, accurate annotations or labels indicating the presence and severity of TB damage in the lung images are needed for supervised training.
[0164] The output step may involve volumes of interest including the identification or segmentation of specific structures. The format of the output (CSV, NIFTI) may be chosen based on the intended use of the results. The volumes of interest may be determined by applying specific criteria to identify regions with distinct characteristics. The criteria may be based on SUV thresholds on PET, density thresholds on CT, detection of cavities, or focusing on lung structures such as cavities. Statistical analysis (360) may involve exploring correlations and associations between first-order radiomics features and clinical parameters, such as patient outcomes or molecular markers.
[0165] Referring to Figure 8A and Figure 8B, the machine learning models used in the described method and system require training. Keras™ and Tensorflow™ are exemplary systems and software libraries that can be used to create and train the model. The co-registered model takes an input of a combined PET and CT scan at a set size. The input may be isolated to just the lungs area, the model outputs a mask which is padded and scaled to fit the original PET scan. The model is trained with 50 PET and corresponding CT scans. The training pool size may be artificially increased by including augmentation methods to distort or change the orientation of the scans. The augmentation may lead to improved accuracy of the 3D ll-NET inference step. An Adam optimiser may be used. A dice- coefficient loss function or a Tversky loss function are two exemplary loss functions that may be used. Variables for feedback for learning include instance of union, the used loss function, and binary accuracy. Additional model feature implementations include the use of dropout to prevent overtraining.
[0166] Figure 8A illustrates a general overview of a training and user of a machine learning model. The training involves a data preparation process (811) to prepare training data (812). A training process (813) iterates training of the machine learning model (814). At runtime of the model (814), an input (821) is received into the runtime process (822) that uses the model (814) to obtain an output (823) that may be used in a downstream process (824).
[0167] Figure 8B shows a flow diagram of a training process for training of the co-registration 3D segmentation machine learning model using the co-registered images to provide outputs including a lung mask to fit the PET image.
[0168] The training may receive manually segmented high-resolution Computed Tomography (CT) images from a CT scan including a region of a lung and may receive manually segmented Positron Emission Tomography (PET) image from a PET scan including the region of the lung and may iteratively train a co-registration trained 3D segmentation machine learning model to receive as input a co-registered image to output a PET lung mask to create an anatomical segmentation of the lungs on a PET image.
[0169] The flow may start (830) with the iteration for model X=0 (831) where X is the number of training iterations. The training flow may acquire (832) the CT and PET DICOMS and may determine (833) if X=0. If X=0, the training flow may hand segment (834) the CT and the PET images, augment (835) the CT / PET scans with their segmentations, and output training data (836) in the form of the CT / PET scans and the segmentations. The training data is used to train (842) model version X+1 , so in this first iteration case Model Version 1. X is incremented (838) to X=X+1 and the training method loops to acquire (832) the CT / PET DICOMS for further training iteration.
[0170] In the further iterations where X does not equal 0 (833), the training method may augment (841) the CT / PET scans, run (842) the trained Model Version X, output segmentation (843), and correct (844) the segmentation. The corrected segmentation (844) is added to the training data (836) to train (837) model version X+1. The method iterates through model versions, until the training is stopped. In this way, the method may initially train the machine learning model on the received manually segmented images; may apply the machine learning model to create masks that may be manually corrected for further training.
[0171] The training method may apply data augmentation on the scan-mask combinations including one or more of the augmentations of: rotation, mirroring, Gaussian noise, stretching, and cropping, to increase the training dataset to make the machine learning model more generalised. The method may also apply dropout to reduce overfitting.
[0172] Training an effective model is a resource-intensive process that requires two essential elements: appropriate source images and segments created to accurately fit these images. The training requires many scans which have already been segmented.
[0173] Scans used for training may be sourced from existing study. The scans may have various levels of lung TB damage, from minor to severe, as well as other abnormal pathology, such as emphysema, as well as some healthy lung scans. The variation in structural damage that TB introduces to the lungs is crucial in making a more robust segmentation model. The scans may be obtained from a number of different machines to increase the robustness of the model.
[0174] Accurate manual segmentation of the hundreds of slices of a single lung scan displaying extensive damage is very time-consuming for a trained, skilled individual. To improve efficiency, an iterative training process may be used to develop the training data needed and data augmentation may be used to expand the scan-mask combinations and maximize the returns.
[0175] A trained medical professional may manually segment an initial 10 scans, providing a PET lung mask, a CT lung mask, and a cavity mask for each scan. Once the CNN models are trained on these initial scans, the ML models are used to create masks for new scans. These new scan masks were then corrected by the skilled person and used to train the model further. After each iteration, the process of correcting the scans became quicker as the models became increasingly accurate.
[0176] Due to the relatively small number of scans, data augmentation may be performed on the scanmask combinations. This included rotation, mirroring, Gaussian noise, and stretching / cropping, to increase the training dataset, which results in making the model more generalised. In addition to data augmentation, dropout may be used to reduce overfitting.
[0177] Referring to Figure 9, a block diagram 900 shows one or more computing system(s) (920) used to implement the described medical images segmentation system (930). Each of the computing system(s) (920) may include a processor (921) for executing the functions of components described below, which may be provided by hardware or by software units executing on the computing system(s). The software units may be stored in a memory component (922) and instructions (923) may be provided to the processor (921) to carry out the functionality of the described components. A graphical processing unit (GPU) (924) may be used to carry out the functionality of the processor (921) when a GPU (912) is available. In some cases, for example in a cloud computing implementation, software units arranged to manage and / or process data on behalf of the computing system(s) may be provided remotely.
[0178] The medical images segmentation system (930) may receive inputs obtained from a CT imaging machine (901) and from a PET imaging machine (902). The medical images segmentation system (930) may include a CT processing system (940) including a CT image component (941) for obtaining a high-resolution 3D CT image from a CT scan. The CT processing system (940) may include a CT pre-processing component (942). The medical images segmentation system (930) may include a PET processing system (950) including a PET image component (951) for obtaining a 3D PET image 9. The PET processing system (950) may include a PET preprocessing system (952).
[0179] The medical images segmentation system (930) includes a co-registration system (960) for coregistering each of a PET scan and a CT scan as an input with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image. The coregistration system (960) may include a metadata extracting component (961) for extracting metadata from the input scan and the reference scan relating to positional and orientation data. The co-registration system (960) may include a reorienting component (964) for reorienting one or both of the input can and reference scan to be parallel or perpendicular to a world axis and to each other. The co-registration system (960) may include a common origin component (962) for matching an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan. The co-registration system (960) may include a matching component (963) for matching the dimensions, orientation, and axes of the input scan and the reference scan.
[0180] The medical images segmentation system (930) may include a combining component (966) for combining the co-registered PET scan and the co-registered CT scan with a PET bias factor component (967) for setting a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan. The medical images segmentation system (930) may include a resizing component (965) for resizing the co-registered CT scan and / or the co-registered PET scan before combining with the other scan.
[0181] The medical images segmentation system (930) may include an isolating component (981) for isolating an area of interest of a CT scan.
[0182] The medical images segmentation system (930) may include an output component (982) for outputting a combined PET and CT scan. The combined PET and CT scan may be used as an input to a co-registration image machine learning model (970). The co-registration image machine learning model (970) may have an associated co-registration model training system (971) and a co-registration model applying system (972).
[0183] Figure 10 illustrates an example of a computing device (1000) in which various aspects of the disclosure may be implemented. The computing device (1000) may be embodied as any form of data processing device including a personal computing device (e.g. laptop or desktop computer), a server computer (which may be self-contained, physically distributed over a number of locations), a client computer, or a communication device, such as a mobile phone (e.g. cellular telephone), satellite phone, tablet computer, personal digital assistant or the like. Different embodiments of the computing device may dictate the inclusion or exclusion of various components or subsystems described below.
[0184] The computing device (1000) may be suitable for storing and executing computer program code. The various participants and elements in the previously described system diagrams may use any suitable number of subsystems or components of the computing device (1000) to facilitate the functions described herein. The computing device (1000) may include subsystems or components interconnected via a communication infrastructure (1005) (for example, a communications bus, a network, etc.). The computing device (1000) may include one or more processors (1010) and at least one memory component in the form of computer-readable media. The one or more processors (1010) may include one or more of: CPUs, graphical processing units (GPUs) (1012), microprocessors, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs) and the like. In some configurations, a number of processors may be provided and may be arranged to carry out calculations simultaneously. In some implementations various subsystems or components of the computing device (1000) may be distributed over a number of physical locations (e.g. in a distributed, cluster or cloud-based computing configuration) and appropriate software units may be arranged to manage and / or process data on behalf of remote devices.
[0185] The memory components may include system memory (1015), which may include read only memory (ROM) and random access memory (RAM). A basic input / output system (BIOS) may be stored in ROM. System software may be stored in the system memory (1015) including operating system software. The memory components may also include secondary memory (1020). The secondary memory (1020) may include a fixed disk (1021), such as a hard disk drive, and, optionally, one or more storage interfaces (1022) for interfacing with storage components (1023), such as removable storage components (e.g. magnetic tape, optical disk, flash memory drive, external hard drive, removable memory chip, etc.), network attached storage components (e.g. NAS drives), remote storage components (e.g. cloud-based storage) or the like.
[0186] The computing device (1000) may include an external communications interface (1030) for operation of the computing device (1000) in a networked environment enabling transfer of data between multiple computing devices (1000) and / or the Internet. Data transferred via the external communications interface (1030) may be in the form of signals, which may be electronic, electromagnetic, optical, radio, or other types of signal. The external communications interface (1030) may enable communication of data between the computing device (1000) and other computing devices including servers and external storage facilities. Web services may be accessible by and / or from the computing device (1000) via the communications interface (1030).
[0187] The external communications interface (1030) may be configured for connection to wireless communication channels (e.g., a cellular telephone network, wireless local area network (e.g. using Wi-Fi™), satellite-phone network, Satellite Internet Network, etc.) and may include an associated wireless transfer element, such as an antenna and associated circuitry.
[0188] The computer-readable media in the form of the various memory components may provide storage of computer-executable instructions, data structures, program modules, software units and other data. A computer program product may be provided by a computer-readable medium having stored computer-readable program code executable by the central processor (1010). A computer program product may be provided by a non-transient or non-transitory computer- readable medium, or may be provided via a signal or other transient or transitory means via the communications interface (1030).
[0189] Interconnection via the communication infrastructure (1005) allows the one or more processors (1010) to communicate with each subsystem or component and to control the execution of instructions from the memory components, as well as the exchange of information between subsystems or components. Peripherals (such as printers, scanners, cameras, or the like) and input / output (I / O) devices (such as a mouse, touchpad, keyboard, microphone, touch-sensitive display, input buttons, speakers and the like) may couple to or be integrally formed with the computing device (1000) either directly or via an I / O controller (1035). One or more displays (1045) (which may be touch-sensitive displays) may be coupled to or integrally formed with the computing device (1000) via a display or video adapter (640).
[0190] The foregoing description has been presented for the purpose of illustration; it is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Persons skilled in the relevant art can appreciate that many modifications and variations are possible in light of the above disclosure.
[0191] Any of the steps, operations, components or processes described herein may be performed or implemented with one or more hardware or software units, alone or in combination with other devices. Components or devices configured or arranged to perform described functions or operations may be so arranged or configured through computer-implemented instructions which implement or carry out the described functions, algorithms, or methods. The computer- implemented instructions may be provided by hardware or software units. In one embodiment, a software unit is implemented with a computer program product comprising a non-transient or non- transitory computer-readable medium containing computer program code, which can be executed by a processor for performing any or all of the steps, operations, or processes described. Software units or functions described in this application may be implemented as computer program code using any suitable computer language such as, for example, Java™, C++, or Perl™ using, for example, conventional or object-oriented techniques. The computer program code may be stored as a series of instructions, or commands on a non-transitory computer-readable medium, such as a random access memory (RAM), a read-only memory (ROM), a magnetic medium such as a hard-drive, or an optical medium such as a CD-ROM. Any such computer-readable medium may also reside on or within a single computational apparatus, and may be present on or within different computational apparatuses within a system or network.
[0192] Flowchart illustrations and block diagrams of methods, systems, and computer program products according to embodiments are used herein. Each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, may provide functions which may be implemented by computer readable program instructions. In some alternative implementations, the functions identified by the blocks may take place in a different order to that shown in the flowchart illustrations. Some portions of this description describe the embodiments of the invention in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations, such as accompanying flow diagrams, are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. The described operations may be embodied in software, firmware, hardware, or any combinations thereof.
[0193] The language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention set forth in any accompanying claims.
[0194] Finally, throughout the specification and any accompanying claims, unless the context requires otherwise, the word ‘comprise’ or variations such as ‘comprises’ or ‘comprising’ will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.
Claims
CLAIMS:
1. A computer-implemented method for combining a three-dimensional positron emission tomography (PET) scan and a three-dimensional computed tomography (CT) scan of a body into a combined three-dimensional image for segmentation, comprising: co-registering (112, 122) each of the PET scan and the CT scan as an input scan with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image, the co-registering including: extracting (113, 123) metadata from the input scan and the reference scan relating to positional and orientation data; matching (115, 125) an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan; and matching (116, 126) dimensions, orientation, and axes of the input scan and the reference scan; and outputting (117, 127) a co-registered PET scan and a co-registered CT scan; and combining (131) the co-registered PET scan and the co-registered CT scan into a single3D image with a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan.
2. The method of claim 1 , including using the combined PET and CT images to train a machine learning model, and inputting (132) the combined PET and CT scan images as a single input channel into the machine learning model for image segmentation.
3. The method of claim 1 or claim 2, wherein the co-registering (112, 122) includes: reorienting (114, 124) one or both of the input scan and reference scan to be parallel or perpendicular to a world axis and to each other, including when the input scan and the reference scan are not parallel or perpendicular to a world axis, resampling (226, 228) one of the input scan or the reference scan to fix a grid orientation. (Figure 2B)4. The method of any of claims 1 to 3, wherein combining (131) the co-registered PET scan and the co-registered CT scan includes: normalizing (210) the co-registered PET scan and multiplying the normalized PET scan by the PET bias factor weight; normalizing (212) the co-registered CT scan and multiplying the normalized CT scan by the CT bias factor weight; and adding (214) the normalized and biased scans to output the combined PET and CT scan.
5. The method of any of the preceding claims, wherein matching (115, 125) the origin points of the input scan and the reference scan includes for each of planes x, y and z: cropping (234) by removing voxels at the start of one scan in a plane; or padding (235) by adding voxels at the start of one scan in a plane; whilst taking spacing and direction into account.
6. The method of any of the preceding claims, including matching (116, 126) the input scan and the reference scan in real world dimensions including padding (245) or cropping (244) voxels on a corner of a scan opposing the origin point and is repeated for each direction x, y and z.
7. The method of any of the preceding claims, wherein the PET scan and CT scan encompass an organ for organ segmentation.
8. The method of any of the preceding claims, including cropping (403, 404, 503, 504) the co-registered PET scan and a co-registered CT scan before combining the co-registered PET scan and the co-registered CT scan into a single 3D image.
9. The method of any of the preceding claims, including isolating (520, 620) an area of interest of a CT scan by: applying a user-defined threshold and identifying continuous slices of the scan at a start and an end of the scan below the user-defined threshold; excluding slices that form a boundary from the identified continuous slices; and removing identified slices from the scan leaving the area of interest.
10. The method of claim 9, including co- registering at least one of the PET scan and the CT scan with the reference scan in the form of an isolated area of interest (520, 620) of a CT scan.
11. The method of claim 9 or claim 10, including co-registering (512) a combined PET and CT scan and an isolated area of interest of a CT scan to output a cropped combined PET and CT scan.
12. The method of claim 9 or claim 10, including: co-registering (112) an input scan of a PET scan (601) and a reference scan of an isolated area of interest (620) to obtain a co-registered PET scan in the form of cropped PET scan; and co-registering (122) an input scan of a CT scan (604) and a reference scan of the cropped PET scan (607) to obtain a co-registered CT scan in the form of a cropped CT scan; andwherein combining the co-registered PET scan and the co-registered CT scan combine the cropped PET scan and the cropped CT scan.
13. The method of claim 9 or claim 10, including: co-registering (122) an input scan of a CT scan (604) and a reference scan of an isolated area of interest (620) to obtain a co-registered CT scan in the form of cropped CT scan; and co-registering (112) an input scan of a PET scan (601) and a reference scan of the cropped CT scan (608) to obtain a co-registered PET scan in the form of a cropped PET scan; and wherein combining the co-registered PET scan and the co-registered CT scan combine the cropped PET scan and the cropped CT scan.
14. The method of any of the preceding claims, including training (813) a machine learning model on a combined positron emission tomography (PET) scan and computed tomography (CT) scan of a portion of a body for segmentation, comprising: receiving manually segmented combined PET and CT scans; and iteratively training a 3D segmentation machine learning model on combined images to receive as input a combined PET and CT scan to output a segmentation image that accurately delineates the organ or anatomical structure of interest.
15. The method of claim 14, including: initially training (813) the machine learning model on the received manually segmented combined images; applying the machine learning model (822) to create masks for further training images; and manually correcting the further training images for additional training.
16. A system for combining a three-dimensional positron emission tomography (PET) scan and a three-dimensional computed tomography (CT) scan of a body into a combined three- dimensional image for segmentation, comprising: a processor including a graphical processing unit accelerator; and a memory, said memory containing instructions executable by said processor, to execute functions of components including: co-registering (112, 122) each of the PET scan and the CT scan as an input scan with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image, the co-registering including: extracting (113, 123) metadata from the input scan and the reference scan relatingto positional and orientation data; matching (115, 125) an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan; and matching (116, 126) dimensions, orientation, and axes of the input scan and the reference scan; and outputting (117, 127) a co-registered PET scan and a co-registered CT scan; and combining (131) the co-registered PET scan and the co-registered CT scan with a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan.
17. The system as claimed in claim 16, including: a system for training the 3D segmentation machine learning model using combined PET and CT scans.
18. The system as claimed in claim 16 or claim 17, including: a combined-image trained 3D segmentation machine learning model having as input the combined PET and CT scan as a single channel.
19. The system as claimed in any of claims 16 to 18, wherein the body is a portion of a human body for organ segmentation.
20. A computer program product for combining a three-dimensional positron emission tomography (PET) scan and a three-dimensional computed tomography (CT) scan of a body into a combined three-dimensional image for segmentation, comprising a computer-readable medium having stored computer-readable program code for performing the steps of: co-registering (112, 122) each of the PET scan and the CT scan as an input scan with reference to a reference scan in the form of the other of the PET scan or CT scan, or in the form of a third image, the co-registering including: extracting (113, 123) metadata from the input scan and the reference scan relating to positional and orientation data; matching (115, 125) an origin point of the input scan and an origin point of the reference scan including padding or cropping one of the input scan and reference scan to match the origin point of the other input scan and reference scan; matching (116, 126) dimensions, orientation, and axes of the input scan and the reference scan; and outputting (117, 127) a co-registered PET scan and a co-registered CT scan; andcombining (131) the co-registered PET scan and the co-registered CT scan with a PET bias factor weight and a CT bias factor weight with a selected bias factor favouring the PET scan to produce a combined PET and CT scan.
Citation Information
Patent Citations
Quantification of medical image data
US20120033865A1