Systems and methods for classifying lesions

JP2025500348A5Pending Publication Date: 2025-10-22KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024537384
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-01-10
Filing Date
2022-12-16
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing lesion classification methods, particularly for prostate cancer, rely heavily on invasive biopsies and lack the accuracy needed for optimal treatment planning, as image-based classification systems have not adequately differentiated anatomical and functional imaging modalities.

Method used

A system utilizing dual-modal CT and PET images, registered and normalized to account for physiological variations, extracts radiomics features to classify lesions based on severity, employing machine learning architectures for precise lesion identification and treatment planning.

Benefits of technology

The system provides accurate, non-invasive lesion classification, enabling optimal biopsy targeting and treatment planning by differentiating between malignant and benign prostate tumor regions, reducing the need for invasive procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system 100 is provided comprising an image providing unit 101 for providing a dual modal image of an imaging region comprising a CT image 10, 10' and a registered PET image 20', the imaging region comprising a tissue region of interest including a lesion and a reference tissue region. The system further comprises an identification unit 102 for identifying respective lesion image segments in the CT image and the PET image and for identifying a reference image segment in the PET image. The system also comprises a normalization unit 103 for normalizing the lesion image segment in the PET image with respect to the reference image segment, an image feature extraction unit 104 for extracting image feature values ​​from both lesion image segments and a classification unit 105 for classifying the lesion based on the extracted values.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a system, method and computer program for classifying lesions, and to a method for selecting a set of image features and machine learning architectures used by a system or in a method for classifying lesions. [Background technology]

[0002] Classifying lesions according to known lesion classification schemes can support optimal treatment in clinical practice. Due to the ever-increasing quality of medical imaging, image-based diagnostic tools have become standard for many applications. However, lesions are still usually classified based on lesion samples taken by biopsy. This is especially true for prostate cancer, which is the second most commonly occurring cancer and the fifth leading cause of death in men worldwide. In fact, biopsy-based diagnosis can still be considered the gold standard for most lesion types, including inflammatory diseases, e.g. in the kidney or liver, and various types of infections, if not easily accessible via biopsy. Early diagnosis and treatment planning are generally of high medical importance. In particular, they are crucial in terms of reducing the mortality rate due to prostate cancer. Since both biopsies necessary for diagnosis and biopsies for the treatment of lesions, e.g. by radiation, usually rely on image guidance, accurate identification of lesion areas, especially with regard to their location and size, in guided images by image-based lesion classification is desirable. Indeed, it would be desirable to eventually facilitate a purely image-based diagnosis of lesions in order to avoid any unnecessary invasive steps. Thus, there is a need in the art for improved image-based classification of lesions, particularly of prostate tumor regions.

[0003] In the prior art, machine learning has been used in attempts to meet this need, where different machine learning architectures have been proposed that operate with different sets of image features to classify lesions, however, the known attempts indicate that there is still room to improve the performance of image-based lesion classification.

[0004] US Patent Application Publication No. 2020 / 0245960A1 describes an automated analysis of three-dimensional medical images of a subject to automatically identify specific three-dimensional volumes within the three-dimensional images that correspond to specific anatomical regions. Identification of one or more such volumes can be used to automatically determine quantitative metrics representative of accumulation of a radiopharmaceutical in specific organs and / or tissue regions. These accumulation metrics can be used to assess a disease state in the subject, determine a prognosis for the subject, and / or determine a course of treatment. Summary of the Invention [Problem to be solved by the invention]

[0005] It is an object of the present invention to provide a system, method and computer program for classifying lesions with improved accuracy, and to provide a method for selecting a set of image features and a machine learning architecture to be used by a system or in a method for classifying lesions such that lesion classification accuracy is improved. [Means for solving the problem]

[0006] In a first aspect, the invention relates to a system for classifying lesions, the system comprising an image providing unit for providing a dual modal image of an imaging region, the dual modal image including a mutually registered computed tomography (CT) image and a positron emission tomography (PET) image, the imaging region including a) a tissue region of interest including a lesion to be classified and b) a reference tissue region including healthy tissue. The system further comprises an identification unit for identifying in each of the CT image and the PET image a respective lesion image segment corresponding to the lesion in the tissue region of interest and for identifying in the PET image a reference image segment corresponding to the reference tissue region. The system also comprises a normalization unit for normalizing the lesion image segment in the PET image with respect to the reference image segment in the PET image, an image feature extraction unit for extracting values ​​for a set of image features associated with the lesion from the lesion image segment in the CT image and the normalized lesion image segment in the PET image, and a classification unit for classifying the lesion according to a severity based on the extracted image feature values.

[0007] Dual modal images including CT and PET images are used and lesions are classified based on the values ​​of image features extracted from both CT and PET images, thus combining anatomical and functional imaging to enable accurate lesion classification. Physiological variations and the PET imaging modality and the parameters selected therefor are accounted for by normalizing the lesion image segments in the PET images to the reference image segments, which would otherwise affect the PET images and thus the image feature values ​​extracted from the lesion image segments in the PET images, thus further improving the lesion classification accuracy.

[0008] The image providing unit is configured to provide a dual modal image of the imaging region, the dual modal image including, in particular consisting of, a CT image and a PET image, which are mutually registered. The CT and PET images can be acquired using an integrated imaging system including a CT imaging modality and a PET imaging modality. They are preferably volumetric, i.e. three-dimensional images, but can also be two-dimensional images. In the case of volumetric images, they can be decomposed as two-dimensional image slices along a given axis, for example the scanning axis. The image elements of the CT images are associated with image values ​​given in Hounsfield Units (HU) and the image elements of the PET images are associated with image values ​​given in Standard Uptake Values ​​(SUV).

[0009] Registration between the CT and PET images can be achieved according to any known registration technique. With registration between the CT and PET images, which together show the imaging region, the dual-modal image can be thought of as associating two image elements to each point in the imaging region: one image element in the CT image and one image element in the PET image.

[0010] The imaging region includes a tissue region of interest that includes the lesion to be classified, and further includes a reference tissue region that includes healthy tissue. The tissue region of interest may be, for example, a part of an organ that includes the lesion to be classified, or the whole organ. The reference tissue region may refer, for example, to a part of an organ that includes tissue that is at least presumably and / or substantially healthy in comparison to or in contrast to the tissue region of interest, or the whole organ. Thus, the tissue included in the reference tissue region being "healthy" may refer to the tissue being healthy, for example, in comparison to the tissue region of interest. The reference tissue region is preferably selected to behave similarly or even similarly to the tissue region of interest in response to physiological variations and / or variations in the PET imaging modality and the parameters selected therefor. The imaging region includes or consists of, for example, a part of the body, such as the body or upper body, of the patient.

[0011] The identification unit is configured to identify in each of the CT image and the PET image a respective lesion image segment corresponding to a lesion in the tissue region of interest, and to identify in the PET image a reference image segment corresponding to the reference tissue region. Thus, the identification unit is configured to identify lesion image segments in both the CT image and the PET, and to identify possibly a single reference image segment in the PET image. Due to the registration between the CT image and the PET image, the identification unit may also be seen as identifying a joint lesion image segment, i.e. a segment in the dual modal image corresponding to a lesion in the tissue region of interest, where a joint lesion image segment may be understood as a combination of the lesion image segments in the CT image and the PET image, or an overlay image segment corresponding thereto. Indeed, due to the registration between the CT image and the PET image, the identified reference image segment in the PET image also corresponds to a reference image segment in the CT image corresponding to the reference tissue region.

[0012] In an embodiment, the identification unit can be configured to identify in each of the CT image and the PET image a respective reference image segment corresponding to the reference tissue region, the reference image segment in the PET image being identified as the image segment aligned to the reference image segment identified in the CT image, thereby allowing a particularly accurate identification of the reference image segment in the PET image.

[0013] The identification unit may be configured to identify lesion image segments in the CT image and the PET image, respectively, as well as reference image segments in the PET image, and possibly also reference image segments in the CT image, based on a user input or based on an automatic segmentation process. Preferably, at least the reference image segments in the PET image and possibly also the reference image segments in the CT image are identified based on an automatic segmentation process.

[0014] The normalization unit is configured to normalize the lesion image segment in the PET image to the reference image segment in the PET image. Normalizing the lesion image segment refers to normalizing the image values ​​in the lesion image segment. For example, the normalization unit may be configured to normalize the SUV in the lesion image segment of the PET image to the median SUV in the reference image segment of the PET image. In fact, the normalization unit may be configured to normalize the entire PET image to the reference image segment in the PET image, for example by normalizing the SUV of the entire PET image to the median SUV in the reference image segment of the PET image.

[0015] For CT imaging, a standardized scale for Hounsfield units exists, but for PET imaging, such a standardized scale of image values ​​is not readily available. The detected brightness from which image values ​​are derived in PET imaging depends on the amount of nuclear tracer injected, the scintillation material used in the detector, whether time of flight differences are accounted for, and the reconstruction parameters used in the PET imaging modality. Usually, therefore, great attention to detail is required in PET scanning to avoid errors. Normalizing the lesion image segments in the PET image to a reference image segment in the PET image allows lesion classification that is less susceptible to such manual and system errors.

[0016] The image feature extraction unit is configured to extract values ​​of a set of image features associated with the lesion from the lesion image segment in the CT image and from the normalized lesion image segment in the PET image. Thus, the set of image features comprises CT image features and PET image features. Preferably, the set of image features comprises radiomics features associated with the lesion, such as semantic features, in particular size, shape and / or location of the lesion, independent features, such as texture-based first order statistics associated with the lesion, and / or features resulting from higher order algorithms, such as Gray Level Co-occurrence Matrix (GLCM) and / or Gray Level Run Length Matrix (GLRLM). The first order features are based on statistical information of the image and may comprise maximum intensity, minimum intensity, mean intensity, standard deviation of intensity, variance of intensity, energy, entropy, sharpness, skewness, kurtosis and / or grayscale. These features may provide statistical information based on the frequency distribution of different gray levels in the image. Lesion classification based on extracted radiomics features can be viewed as involving radiomics feature analysis, which may be generally defined as the extraction of quantitative image features for characterization of disease patterns.

[0017] In one example, the lesion image segments in the CT and PET images are cubic shaped image regions that depict the lesion and its immediate surroundings in the tissue region of interest, and each cubic shaped image region or its margin may also be referred to as a bounding box of the region of interest. Radiomics features are then extracted from the respective image values ​​or intensities in these rectangular image regions.

[0018] The classification unit is configured to classify the lesion according to its severity based on the extracted image feature values. For example, the classification of the lesion may be achieved based on a possibly predefined relationship between the values ​​of the extracted image features and known severity grades. The predefined relationship is stored, for example in the form of a table, by or for access by the classification unit. In particular, in the case of a large number of image features, this table can be seen as a database relating the values ​​of the image features associated with the lesions and the assigned severity classification, where the assignment is based on an assessment of the lesion severity by a trained pathologist. Preferably, however, the classification unit is configured to classify the lesion based on a predefined procedure, possibly stored in terms of a computer program by or for access by the classification unit. This predefined procedure may have the form of an artificial intelligence, in particular a machine learning model undergoing supervised learning.

[0019] The lesion is classified according to its severity, so that the optimal biopsy target and / or optimal treatment for the lesion can be determined depending on the classification. For example, if the lesion is a tumor, the classification indicates whether it is malignant or benign, where a malignant tumor requires intensive treatment, whereas a benign tumor may require less treatment or no treatment at all. Also, accurate location of the malignant tumor allows, for example, optimal biopsy targeting and / or optimal planning of subsequent removal treatment. Using image features of both CT and PET images allows for higher accuracy in the classification of the lesion, and therefore improved biopsy targeting for diagnosis and / or improved treatment targeting.

[0020] In a preferred embodiment, the tissue region of interest refers to the prostate of a patient, the lesion is a prostate tumor region, and the PET image is generated using an imaging agent that binds to the prostate specific membrane antigen (PSMA). In particular, the lesion to be classified may be a single habitat or a set of adjacent habitats of a prostate tumor, and the imaging agent, also referred to as a PET tracer, is in particular gallium 68 ( 68 Ga), in which case the imaging technique used is 68 This is referred to as Ga-PSMA-PET / CT imaging. More specifically, the imaging material is 68 Ga-PSMA-11. Over 90% of prostate cancers overexpress PSMA, so these tumor cells can be precisely tagged by using a PET imaging agent that binds to PSMA. 68 A molecular imaging technique called Ga-PSMA-PET / CT imaging appears to be clinically replacing CT imaging alone and appears to be superior to magnetic resonance (MR) imaging for the detection of metastatic disease. 68 Ga-PSMA-PET / CT imaging has the ability to reliably differentiate the stage of prostate cancer relative to presentation and may help inform optimal treatment approaches.

[0021] In a preferred embodiment, the lesions are classified according to a binary classification of severity based on the Gleason score, which distinguishes between normal / inactive and malignant prostate tumor areas. The Gleason grade is the prostate cancer prognostic tool most commonly used by physicians. It has been used for a long time to determine the malignancy of prostate cancer in order to plan treatment options. However, this process usually requires well-trained pathologists to view multiple biopsy samples under a microscope and assign a Gleason grade to the cancer based on its severity as assessed by the pathologists. The system disclosed herein allows a purely image-based, and therefore non-invasive and objective, determination of the Gleason grade of prostate cancer. In prostate tumors, there are different habitats with distinct volumes, each with a specific combination of flow, cell density, necrosis and edema. The habitat distribution in patients with prostate cancer is particularly 68 Based on Ga-PSMA-PET / CT images, it can be used to distinguish between rapidly progressing and more inactive cancer regions in order to classify cancer regions with high accuracy according to the Gleason grade system. A deeper understanding of tumor heterogeneity can be obtained by regional analysis, in which image elements, e.g. voxels corresponding to tumors, are assigned specific colors according to their combination of brightness, and clusters of voxels with specific colors give rise to regions reflecting different habitats that can be referred to as physiological microenvironments, linking regional analysis with Gleason score and hormonal status, as well as with neuroendocrine components within the tumor.

[0022] In a preferred embodiment, the reference tissue region is the liver of the patient, in particular the liver of the patient whose tissue region of interest refers to its prostate. Thus, the identification unit may be configured to identify in the PET image a reference image segment corresponding to the liver of the patient. The liver contains the largest amount of blood in the body at any given time. Thus, the contrast in the PET image values, such as SUV, between the liver and the tissue region of interest containing the lesion to be classified is directly correlated with the patient's circulatory system. Using the liver as the reference tissue region thus allows a normalization of the lesion image segment that is particularly well removed from the variability of the patient's physiology. However, in principle, other organs, preferably normal or healthy organs, may also be selected as the reference tissue region. Selecting the liver as the reference tissue region is particularly suitable when the lesion to be classified is a prostate tumor region.

[0023] In a preferred embodiment, the set of image features includes total energy, maximum two-dimensional diameter in the column direction, sphericity, surface area, large area enhancement and dependent entropy as features whose image feature values ​​are extracted from lesion image segments in CT images, and 90th percentile, minimum, flatness and sphericity as features whose image feature values ​​are extracted from normalized lesion image segments in PET images. Classifying lesions based on the value of this combination of image features allows particularly accurate results. Optionally, the set of image features can include or consist of any combination of one or more of the above-mentioned features (total energy, maximum two-dimensional diameter in the column direction, sphericity, surface area, large area enhancement and dependent entropy) whose values ​​are extracted from lesion image segments in CT images, and one or more of the above-mentioned features (90th percentile, minimum, flatness and sphericity) whose values ​​are extracted from normalized lesion image segments in PET images. This also allows very accurate classification results.

[0024] In a preferred embodiment, the identification unit comprises a machine learning architecture configured and preferably trained to receive as input a CT image of the imaging region including the reference tissue region and determine as output a reference image segment in the CT image, so as to identify a reference image segment in the CT image, and then a reference image segment in the PET image is identified as an image segment in the PET image that corresponds to the determined reference image segment in the CT image via a registration between the CT image and the PET image.

[0025] In particular, if the CT image is a volumetric image, the machine learning architecture may be configured to receive the CT image as slices, with each input being formed by three adjacent image slices. For example, if the CT image includes 100 slices, three adjacent slices at a time may be provided as inputs to the machine learning architecture. The machine learning architecture is trained using binary cross entropy and a loss function.

[0026] The identification unit may be configured to identify the reference image segment in the CT image by first modifying the CT image and then identifying the reference image segment in the modified CT image. Rescaling refers to clipping the CT image values ​​to a range or window that includes the image values ​​corresponding to the reference tissue region and the image values ​​corresponding to the tissue region of interest, and then mapping the range or window to the entire range of possible CT image values, i.e., for example, brightness. For example, the CT image data may be clipped to a range between -150 and 250 HU, which is then mapped to a range from 0 to 255 HU. The soft tissue of the liver and the lesions typically lie in the range between -150 and 250 HU.

[0027] In a preferred embodiment, the machine learning architecture included in the identification unit includes a convolutional neural network having convolution, activation and pooling layers as a base network, and respective side sets of convolutional layers in multiple convolution stages of the base network connected to respective ends of the convolutional layers of the base network at each convolution stage, and the reference image segments in the CT image are determined based on outputs of the multiple side sets of convolutional layers. In particular, the base network may correspond to a VGG-16 architecture with the last fully connected layer removed, as described in the conference contributions by M. Bellver et al., "Detection-aided liver lesion segmentation using deep learning", Conference of Neural Information Processing, pages 1-4 (2017), arXiv:1711.11069, and K.-K. Maninis et al., "Deep Retinal Image Understanding", Medical Image Computing and Computer-Assisted Intervention, vol. 9901, pages 140-148 (2016), arXiv:1609.01103. However, the base network may also include, for example, a U-Net or Unet++ architecture.

[0028] The machine learning architecture included in the identification unit can be trained, for example, on ImageNet, which allows for fast and robust model training. More specifically, the machine learning architecture included in the region unit can be trained on a set of training CT and / or PET images showing reference tissue regions and / or lesions in the tissue region of interest of the type to be regioned, where the training images are or can be a subset of the images on ImageNet. In each training image, the lesion image segments, the reference image segments corresponding to the reference tissue regions and / or lesions in the tissue region of interest can be identified, for example, by a pathologist, during preparation for training, and this identification during preparation for training can be considered as ground truth for training, which can refer, for example, to annotation, indication and / or segmentation. If the reference tissue region is the liver of a patient, the segmentation process automated by the trained machine learning architecture of the identification unit can give a result corresponding to a Dice score of, for example, 0.96. Via training, the selected machine learning architecture, which initially implements a general machine learning model, is tuned to be suitable for identifying reference and / or lesion image segments, which may include tuning hidden parameters of the architecture in particular. According to a more general training scheme, other parameters such as hyperparameters of the architecture may also be tuned. Hyperparameters can be seen as parameters of a given type of machine learning architecture that control its learning process.

[0029] In a preferred embodiment, the classification unit comprises a possibly already trained machine learning architecture suitable for receiving image feature values ​​as input and determining an associated lesion severity classification as output in order to classify the lesion. Thus, the system may comprise two separate machine learning architectures: a first machine learning architecture that is a discrimination unit for identifying reference tissue regions in the CT and / or PET images, said first machine learning architecture being referred to as the machine learning architecture for discrimination, and a second machine learning architecture that is a classification unit for classifying lesions based on image feature values ​​extracted from lesion image segments in the CT and PET images, said second machine learning architecture being referred to as the machine learning architecture for classification.

[0030] The machine learning architecture included in the classification unit can be trained using the same training data as that for the machine learning architecture included in the identification unit, i.e., for example, on ImageNet, which allows for fast and robust model training. More specifically, the machine learning architecture included in the classification unit can be trained on a set of CT and / or PET images for training showing lesions in the tissue region of interest of the type to be classified, which is or can be a subset of the images on ImageNet, or on lesion image segments identified therein. From each training image or lesion image segment used for training, a respective image feature value associated with each lesion therein can be extracted, and a classification such as a lesion severity classification can be assigned to each training image, lesion image segment, or extracted image feature value in preparation for training, for example by a pathologist, and this assignment in preparation for training is considered as the ground truth for training. Through training, the selected machine learning architecture, which initially implements a general machine learning model, is tuned to be suitable for classifying lesions, which includes tuning hidden parameters of the architecture. According to a more general training scheme, other parameters such as hyperparameters of the architecture can also be tuned.

[0031] For example, the machine learning architecture included in the classification unit, specifically trained, may be a random forest with entropy as splitting criterion, but with a maximum tree depth of 5, a maximum number of features considered for each split of 2, and a maximum number of features of 500. However, different hyperparameters for the random forest may be suitable depending on the training dataset used to train the machine learning architecture.

[0032] In some embodiments, the set of image features and the machine learning architecture for classification are predetermined by training a candidate machine learning architecture for classifying the lesion using values ​​of a candidate set of image features extracted from training images of tissue regions containing the lesion with the respective assigned lesion severity classifications, determining the performance of each trained candidate machine learning architecture for the values ​​of the respective candidate set of image features extracted from test images of tissue regions containing the lesion with the respective assigned lesion severity classifications, and selecting the candidate set of image features and the candidate machine learning architecture to be used in the system based on the determined performance. In this way, the machine learning architecture of the classification unit used for classification and the set of image features are tuned, i.e., determined, in a joint training process. This enables efficient and accurate classification of the lesion by the system, because the machine learning architecture of the classification unit used for classification and the set of image features are mutually adapted.

[0033] In a further aspect of the invention, a method for classifying lesions is presented, comprising the steps of providing a dual modal image of an imaging region, the dual modal image including a mutually registered CT image and a PET image, the imaging region including a) a tissue region of interest including a lesion to be classified and b) a reference tissue region including healthy tissue, identifying in each of the CT image and the PET image a respective lesion image segment corresponding to the lesion in the tissue region of interest and identifying in the PET image a reference image segment corresponding to the reference tissue region. The method comprises the steps of normalizing the lesion image segment in the PET image with respect to the reference image segment in the PET image, extracting values ​​of a set of image features associated with the lesion from the lesion image segment in the CT image and the normalized lesion image segment in the PET image, and classifying the lesion according to a severity based on the extracted image feature values. The method can be performed, for example, using the system described above, in particular the machine learning architecture for discrimination and / or the machine learning architecture for classification described above.

[0034] In another aspect, a method is presented for selecting a set of image features and a machine learning architecture, in particular a machine learning architecture for classification, for use by a system or method for classifying lesions, which may also be referred to as a tuning and training method, comprising the steps of: i) providing a set of dual-modal training images, each training image including a mutually registered CT image and a PET image, in each of which a respective lesion image segment corresponding to a respective lesion in a tissue region of interest has been identified, and an assigned severity classification has been provided for each of the lesions in the training images; ii) selecting a candidate set of image features from a predefined collection of possible image features; iii) generating a training data set by extracting corresponding values ​​of each of the candidate set of image features from each of the identified lesion image segments in the training images and assigning them to the respective severity classifications assigned to the lesions in the respective training images; iv) selecting a candidate machine learning architecture suitable for receiving the image feature values ​​as input and determining an associated lesion severity classification as output; and v) training the candidate machine learning architecture on the training data set. The method also includes the steps of: vi) providing a set of dual-modal test images, each including a CT image and a PET image, in each of which a respective lesion image segment corresponding to a respective lesion in the tissue region of interest is identified, and an assigned severity classification is provided for each of the lesions in the test images; and vii) generating a test data set by extracting corresponding values ​​of each of a candidate set of image features from each of the identified lesion image segments in the test images and assigning them to the respective severity classifications assigned to the lesions in the respective test images.The method further comprises the steps of: viii) determining the performance of the candidate machine learning architecture on the test dataset based on a relationship between the assigned severity classifications provided for the lesions in the test images and the corresponding lesion severity classifications determined by the candidate machine learning architecture; ix) repeating the above steps, i.e. steps i) to xiii), using a larger number of image features in the candidate set of image features until a stopping criterion is met, where for each iteration, the candidate set of image features is expanded with further image features from a predefined collection of possible image features, the further image features being those that result in the highest increase in performance; and x) repeating the above steps, i.e. steps i) to ix), using at least one, but preferably multiple, candidate machine learning architectures, where the set of image features and the machine learning architectures used to classify the lesions are selected from each candidate according to the determined performance.

[0035] The relationship between the assigned severity classifications provided for the lesions in the test images and the corresponding lesion severity classifications determined by the candidate machine learning architecture, on which performance is determined, may refer to, for example, a one-to-one comparison. More generally, performance may be a measure of how often the severity classifications determined by the candidate machine learning model match the severity classifications assigned to the respective lesion image segments in the test images, perhaps at least to some margin of error. Thus, for example, performance may refer to classification accuracy.

[0036] In some embodiments, the values ​​of image features extracted from identified lesion image segments in the training and test images are Gaussian standardized, i.e. scaled based on their mean and their standard deviation to form a standard normal distribution, to avoid attributing erroneous significance to features with larger values. The mean and standard deviation used to scale the values ​​of a particular feature extracted from a lesion image segment in a given image may, for example, correspond to the mean and standard deviation of values ​​of a particular feature previously extracted from lesion image segments in other images, in particular training images.

[0037] In a preferred embodiment, the candidate machine learning architectures correspond to the same type of machine learning architecture, but in particular to random forests, with different hyper-parameters. The hyper-parameters are preferably selected successively according to a grid search. The test images are cross-validated by including different subsets of the training images. However, the test images can also be distinct images that are not included in the entire set of training images.

[0038] In another aspect, a system for performing a method for selecting a set of image features and a machine learning architecture, in particular a classifying machine learning architecture, for use by a system for classifying lesions or in a method for classifying lesions is presented, also referred to as a tuning and training system, which includes a unit for performing, for each of steps i) to x), a respective one of steps i) to x), i.e., the system comprises: i) a providing unit for providing a set of dual-modal training images, each training image including a mutually registered CT image and a PET image, in each of which a respective lesion image segment corresponding to a respective lesion in a tissue region of interest has been identified and an assigned severity classification has been provided for each of the lesions in the training images; and ii) a selecting unit for selecting a candidate set of image features from a predefined collection of possible image features. The system further comprises iii) a generation unit for generating a training dataset by extracting corresponding values ​​of each of a candidate set of image features from each of the lesion image segments identified in the training images and assigning them to respective severity classifications assigned to the lesions in the respective training images; iv) a selection unit for selecting a suitable candidate machine learning architecture for receiving the image feature values ​​as input and determining an associated lesion severity classification as output; and v) a training unit for training the candidate machine learning architecture on the training dataset.The system further comprises: vi) a providing unit for providing a set of dual-modal test images, each test image including a CT image and a PET image, in each of which a respective lesion image segment corresponding to a respective lesion in the tissue region of interest is identified and an assigned severity classification is provided for each of the lesions in the test images; and vii) a generating unit for generating a test dataset by extracting corresponding values ​​of each of a candidate set of image features from each of the identified lesion image segments in the test images and assigning them to the respective severity classifications assigned to the lesions in the respective test images. The system further comprises a determining unit for determining the performance of the candidate machine learning architecture on the test dataset based on a relationship between the assigned severity classifications provided for the lesions in the test images and the corresponding lesion severity classifications determined by the candidate machine learning architecture, the units of the system being configured to: ix) repeating each of steps i) to xiii) using a larger number of image features in the candidate set of image features until a stopping criterion is met, where for each iteration the candidate set of image features is expanded by further image features from a predefined collection of possible image features, the further image features resulting in the highest increase in performance; and x) repeating each of steps i) to ix) with at least one, but preferably a plurality of, candidate machine learning architectures, where the set of image features and the machine learning architectures used to classify the lesions are selected from the respective candidates according to the determined performance.

[0039] In yet a further aspect, a computer program for classifying lesions is presented, the program including program code means for causing a system for classifying lesions to perform the method for classifying lesions when the program is run on a computer controlling the system. Also presented is a computer program for selecting a set of image features and a machine learning architecture, in particular a machine learning architecture for classification, for use by a tuning and training system or in a tuning and training method, the program including program code means for causing the tuning and training system to perform the tuning and training method when the program is run on a computer controlling the tuning and training system. In this aspect, each system may in fact be a computer itself, and each computer program may be executed by one or more processing units of that computer.

[0040] It is to be understood that the system of claim 1, the method of claim 10, the method of claim 11 and the computer program of claim 14 have similar and / or identical preferred embodiments, in particular as defined in the dependent claims.

[0041] It shall be understood that a preferred embodiment of the invention can also be any combination of the dependent claims or the respective independent claims of the above embodiments.

[0042] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. [Brief description of the drawings]

[0043] [Figure 1] FIG. 1 illustrates a schematic and exemplary system for classifying lesions. [Diagram 2]FIG. 2 shows a schematic and exemplary illustration of a base network underlying the machine learning architecture included in the identification unit. [Diagram 3] FIG. 1 shows a schematic and exemplary method for classifying lesions. [Figure 4] 1 is a schematic, exemplary illustration of steps performed in an embodiment to normalize lesion image segments in a PET image. [Diagram 5] 1 is a schematic, exemplary illustration of steps performed in one embodiment of a method for selecting a set of image features and a machine learning architecture for use by a system for classifying lesions or in a method for classifying lesions; [Figure 6] 1 is a schematic and exemplary illustration of an embodiment according to some of the above-mentioned aspects; [Figure 7A] 1 is an illustration of the classification accuracy of a system for cross-validation in an embodiment with a confusion matrix. [Figure 7B] 1 illustrates the classification accuracy of a system for test data in an embodiment with a confusion matrix. [Figure 8] 1 is an illustration of the classification accuracy of a system in one embodiment using Receiver Operating Characteristics (ROC). [Figure 9] 1 is an illustration of the dependence of the classification accuracy of the system on extracted image features in one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0044] 1 shows, in a schematic and exemplary manner, a system 100 for classifying lesions. The system 100 comprises an image providing unit 101 for providing a dual-modal image of an imaging region, the dual-modal image comprising mutually registered CT and PET images, the imaging region comprising a) a tissue region of interest comprising a lesion to be classified and b) a reference tissue region comprising healthy tissue.

[0045] The CT and PET images may be acquired using an integrated imaging system with a CT imaging modality, such as a CT scanner, and a PET imaging modality, such as a PET scanner. However, the CT and PET images may also be acquired separately by separate imaging modalities. In the case of an integrated imaging system, the registration between the CT and PET images may be provided by the imaging system. However, the registration between the CT and PET images may also generally be provided separately, especially after imaging.

[0046] The image providing unit may be, for example, a storage unit that stores the dual modal images for access by other units of the system 100, or may be an interface through which the system 100 receives the dual modal images. The image providing unit 101 may also include an interface through which the system 100 receives the separate CT and PET images, and further a processing unit that provides registration of the CT and PET images, and the image providing unit 101 may then be configured to provide the dual modal image based on the registration of the separate CT and PET images. The dual modal image may be viewed as a single image corresponding to an overlay of the CT and PET images, the overlay being constructed based on the registration between the two images.

[0047] The system 100 further includes an identification unit 102 for identifying respective lesion image segments corresponding to lesions in the tissue region of interest in each of the CT image and the PET image, and for identifying reference image segments corresponding to reference tissue regions in the PET image. In case the system 100 is configured to classify prostate tumor regions, such as different habitats inside the prostate tumor, the lesion image segments refer to image segments including the respective prostate tumor regions. The lesion image segments have a standard form, which is a cubic shape, in which case the lesion image segment or its boundary can be referred to as the bounding box of the region of interest. The lesion image segments then show the prostate and / or the tissue surrounding the prostate tumor apart from the specific prostate tumor region of interest.

[0048] The lesion image segments are generated by applying a mask to the respective images, which may be generated in response to a user input and / or by the identification unit alone, such as by using a machine learning architecture for identification. In the embodiment described herein, the habitat of the prostate tumor is marked by a trained user of the system, but the lesion image segments may also be identified using a deep neural network-based segmentation architecture such as U-Net, or using a dynamic threshold on a PET scan. Identifying reference image segments in a CT image similarly refers to applying a mask to the CT image.

[0049] In case the lesion to be classified is a prostate tumor habitat and the PET image is generated using an imaging agent that binds to PSMA, it is preferable to identify the lesion image segment, i.e., for example, generate a bounding box around each habitat based on the PET image and transfer the bounding box to the CT image via registration of the PET image and the CT image.The identification unit is then preferably configured to identify a respective reference image segment in each of the CT image and the PET image, which corresponds to a reference tissue region, in which case the reference image segment in the PET image is identified as the image segment registered to the reference image segment identified in the CT image.Thus, in case of classifying the prostate tumor habitat, it is preferable to identify the lesion image segment in the PET image and the reference image segment in the CT image, in which case the lesion image segment in the CT image and the reference image segment in the PET image are identified via registration.

[0050] In order to extend the masks in the CT and PET images to each other through registration, it is necessary that the resolutions of the CT and PET images are equalized. For example, a CT volume usually has a spatial resolution of 512X512, whereas a PET image typically has a spatial resolution of 192X192. However, for example, a mask of a CT reference image segment can be transferred to a PET image by increasing the spatial resolution of the PET image to 512X512 by trilinear interpolation.

[0051] The system 100 further includes a normalization unit 103 for normalizing the lesion image segment in the PET image to a reference image segment in the PET image. For example, the reference tissue region may be the liver of the patient, in which case the identification unit may be configured to identify the image segment corresponding to the liver of the patient as the reference image segment, in which case the normalization unit may then be configured to normalize the lesion image segment in the PET image, i.e., for example, the image segment of the PET image corresponding to the classified prostate tumor habitat, to the image segment in the PET image corresponding to the liver of the patient. Normalization refers to normalizing the image value in the lesion image segment of the PET image to a characteristic value of the image value in the PET image corresponding to the liver of the patient, in which case the characteristic value may be, for example, the median SUV.

[0052] The system 100 further includes an image feature extraction unit 104 for extracting values ​​of a set of image features associated with the lesion from the lesion image segments in the CT images and from the normalized lesion image segments in the PET images. In particular, the set of image features extracted by the image feature extraction unit 104 are radiomics features, where the radiomics features are extracted from bounding boxes around the lesions to be classified, which bounding boxes are identified in the CT images and in the PET images, respectively. The set of image features extracted by the image feature extraction unit is a predefined set from a grand set, i.e., a base collection of possible image features, where the predefined is a result of training the system 100 or a preconfiguration of the system 100 based on user input.

[0053] The system 100 further comprises a classification unit 105 for classifying the lesion according to its severity based on the extracted image feature values. The lesion is classified based on the values ​​of the image features extracted from both the CT image and the PET image, so that anatomical and functional information enters into the classification. Also, since the lesion image segment in the PET image from which the image feature values ​​are extracted has been previously normalized with respect to a reference image segment corresponding to a reference tissue region, the functional information can be normalized, thus enabling an objective and accurate classification of the lesion.

[0054] Where the tissue region of interest refers to the patient's prostate and the lesion is a prostate tumor, and PET images can be generated using an imaging agent that binds to PSMA, the lesion is preferably classified according to a binary classification of severity based on Gleason score, where the binary classification distinguishes between inactive and malignant prostate tumor regions.

[0055] In a preferred embodiment, the set of image features extracted by image feature extraction unit 104 includes, and in particular consists of, the features listed in Table 1 below, where the left column lists the features whose values ​​are extracted from lesion image segments in CT images, and the right column lists the features whose values ​​are extracted from normalized image segments in PET images.

[0056] [Table 1]

[0057] The features listed in Table 1 above have been found to be the 10 most important features for classifying lesions, particularly prostate tumor habitats. In fact, any subset of these features has been found to result in lower accuracy for classifying prostate tumor habitats into inactive and malignant. In an embodiment, therefore, any combination of one or more of the CT images listed in Table 1 and one or more of the PET images can be used. On the other hand, using more features, especially more than the 10 features listed in Table 1, may result in a larger amount of calculations without a significant increase in classification accuracy.

[0058] The importance of using features from both CT and PET images can be seen from the results of three independent experiments that were performed. In the first of these experiments, only features from the CT images were used. In the second of these experiments, only features from the PET images were used. And in the third of these experiments, concatenated (fused) features from both CT and PET images were used. The best results were achieved using concatenated features, as can be seen from Table 2 below.

[0059] [Table 2]

[0060] The identification unit 102 preferably comprises a machine learning architecture configured and trained to receive as input a CT image of an imaging region including a reference tissue region and to determine as output a segment in the CT image corresponding to the reference tissue region in order to identify the reference image segment in the CT image. Thus, the machine learning architecture included in the identification unit 102 is configured and trained to segment a reference tissue region in a CT image, the input of which consists of a CT image and the output of which consists of a CT image in which the reference tissue region has been segmented or in a corresponding masked CT image.

[0061] As shown in FIG. 2, the machine learning architecture included in the identification unit 102 includes a convolutional neural network having convolution, activation and pooling layers as a base network 115, and in multiple convolution stages of the base network 115, each side set of convolution layers is connected to each end of the convolution layer of the base network 115 in each convolution stage, and the reference image segment in the CT image is determined based on the output of the multiple side sets of convolution layers. As seen in FIG. 2, the base network 115 corresponds to setting B of the general VGG-16 architecture originally disclosed in the contribution to the conference by K. Simonyan and A. Zisserman, "Very Deep Convolutional Networks for Large-Scale Image Recognition", 3rd International Conference on Learning Representations (2015). The filter used for the convolution layer has a relatively small size of 3X3, and in this case, the convolution layer has a depth ranging from 64 to 512. The pooling layer, in this embodiment, is a max pooling layer, where max pooling is performed on a window of size 2X2, as indicated by the shorthand " / 2" in FIG.

[0062] Although FIG. 2 shows only the base network 115, details regarding the side set of convolution layers connected to the ends of the convolution layers of the base network 115 in each convolution stage and how the reference image segments can be determined in the CT image based on the output of the side set of convolution layers can be inferred from the conference contributions by M. Bellver et al., “Detection-aided liver lesion segmentation using deep learning”, Conference of Neural Information Processing, pages 1-4 (2017), arXiv:1711.11069, and by K.-K. Maninis et al., “Deep Retinal Image Understanding”, Medical Image Computing and Computer-Assisted Intervention, vol. 9901, pages 140-148 (2016), arXiv:1609.01103. The activation layer used in the machine learning architecture of the identification unit can in particular be a normalized unit (ReLU) activation layer.

[0063] The base network 115 may be pre-trained on ImageNet. The deeper the network, the coarser the information it carries, and the more semantically relevant the learned features are. On the other hand, in shallower feature maps that work at higher resolutions, the filters capture more local information. In order to utilize the information learned in feature maps that work at different resolutions, several side outputs are used in the machine learning architecture included in the identification unit 102. In particular, supervision is provided to a side set of convolution layers connected to the ends of the convolution layers of the base network 115 at each convolution stage, and a reference image segment in the CT image is determined based on the output of the multiple side sets of the convolution layers to which supervision is provided. Each of the side outputs, i.e., the set of convolution layers connected to each end of the multiple convolution stages, differentiates in different types of features depending on the resolution at each connection point. In an embodiment, the outputs of the multiple side sets of the convolution layers may be resized and linearly combined to identify the reference image segment in the CT image, i.e., to generate, for example, a segmented liver mask.

[0064] 3 shows a schematic and exemplary method 200 for classifying a lesion, comprising a step 201 of providing a dual-modal image of an imaging region, the dual-modal image including a mutually registered CT image and a PET image, the imaging region including a) a tissue region of interest including a lesion to be classified and b) a reference tissue region including healthy tissue, and a step 202 of identifying in each of the CT image and the PET image a respective lesion image segment corresponding to the lesion in the tissue region of interest and identifying in the PET image a reference image segment corresponding to the reference tissue region. The method 200 further comprises a step 203 of normalizing the lesion image segment in the PET image with respect to the reference image segment in the PET image, a step 204 of extracting values ​​of a set of image features associated with the lesion from the lesion image segment in the CT image and the normalized lesion image segment in the PET image, and a step 205 of classifying the lesion according to its severity based on the extracted image feature values.

[0065] Fig. 4 shows a schematic and exemplary PET SUV normalization with CT liver segmentation. From a CT image 10 (see the first image in Fig. 4 when viewed from left to right) showing an imaging region including the upper body of a patient, the liver is segmented in step 202a, for example by application of a machine learning architecture, which can also be considered as a deep learning model, as described with respect to Fig. 2 above. The CT image 10 in which the liver is segmented is shown in the second image of Fig. 4 when viewed from left to right. The CT liver segmentation is then extended in step 202 to the PET image, which is aligned to the CT image 10 via registration, as shown in the third image of Fig. 4 when viewed from left to right. Thus, in the embodiment shown in Fig. 4, in each of the CT image 10 and the PET image, respective reference image segments corresponding to reference tissue regions, in this case the liver, are identified, and a reference image segment in the PET image is identified as an image segment aligned to the reference image segment identified in the CT image, by application of a machine learning architecture for identification. Finally, the PET image is normalized with respect to its SUV relative to the median SUV in the reference image segment identified in the PET image corresponding to the liver, resulting in a normalized brightness of the PET image as shown in the last image of Figure 4 when viewed from left to right. In this case, the entire PET image is normalized to specifically normalize the lesion image segment, here corresponding to the prostate tumor region, which corresponds to step 203.

[0066] Preferably, to classify the lesion, a machine learning architecture is used that is suitable for receiving the image feature values ​​as input and determining an associated lesion severity classification as output. The classification unit 105 preferably includes this machine learning architecture, which may be a random forest as described above. The machine learning architecture used to classify the lesion and the set of image features provided to that machine learning architecture as input for classifying the lesion are preferably determined in a joint decision process.

[0067] FIG. 5 illustrates, in a schematic and exemplary manner, parts of such a joint decision process as one embodiment of a method 300 for selecting a set of image features and a machine learning architecture to be used by the system 100 or in the method 200 to classify a lesion. The method 300 comprises a step 301 of providing a set of dual-modal training images, each training image including a mutually registered CT image and a PET image, in each of which a respective lesion image segment corresponding to a respective lesion in a tissue region of interest is identified, and an assigned severity classification is provided for each of the regions in the training images. In a further step 302 of the method 300, a candidate set of image features is selected from a predefined collection of possible image features. The selected candidate set of image features is thus a subset generated from the set of all possible image features, as shown in FIG. 5. The method 300 further comprises a step 303 of generating a training data set by extracting corresponding values ​​of each of the candidate set of image features from each of the lesion image segments identified in the training images and assigning them to the respective severity classifications assigned to the lesions in the respective training images. In a further step 304 of the method 300, a candidate machine learning architecture is selected that is suitable to receive the image feature values ​​as input and determine the associated lesion severity classification as output. Thus, for example, a random forest with certain hyper-parameters is selected as a candidate machine learning architecture to be included in the classification unit of the system 100. In a further step 305 of the method 300, the candidate machine learning architecture is trained on a training data set. The training of the candidate machine learning architecture can also be viewed as a learning algorithm, as shown in FIG. 5.In a further step 306 of the method 300, a set of dual-modal test images is provided, where each test image includes a CT image and a PET image, in each of which a respective lesion image segment corresponding to a respective lesion in the tissue region of interest is identified, and an assigned severity classification is provided for each of the lesions in the test images. The severity classifications provided for the lesions in the test images may be understood as ground truth for testing purposes. In a further step 307 of the method 300, a test data set is generated by extracting corresponding values ​​of each of the candidate set of image features from each of the lesion image segments identified in the test images and assigning them to the respective severity classifications assigned to the lesions in the respective test images. In a further step 308 of the method 300, the performance of the candidate machine learning architecture on the test data set is determined based on a relationship between the assigned severity classifications provided for the lesions in the test images and the corresponding lesion severity classifications determined by the candidate machine learning architecture, as shown in FIG. 5. Performance may be measured, for example, in terms of the accuracy, sensitivity, specificity of the candidate machine learning architecture, and / or the speed with which the candidate machine learning architecture arrives at a classification based on a candidate set of image features.

[0068] The above-mentioned steps 301 to 308 of the method 300 are repeated in step 309 with a larger number of image features in the candidate set of image features, where for each iteration the candidate set of image features is expanded by an additional image feature from a predefined collection of possible image features, which results in the highest increase in performance, until a stopping criterion is met. Thus, for a fixed candidate machine learning architecture, an increasing set of extracted image features is used until a stopping criterion is met, which is illustrated by a cycle in FIG. 5. The stopping criterion may be that the number of extracted candidate image features increases above a predefined threshold, i.e., the performance increase falls below a predefined threshold. Explicitly illustrated in FIG. 5 is that all previous steps, i.e., steps 301 to 309, may be repeated in a further step x) of the method 300 with multiple candidate machine learning architectures, i.e., for example, with random forests with different hyperparameters. A set of image features and machine learning architectures used to classify the lesions are then selected from each candidate according to the performance determined for each combination, or pair.

[0069] Thus, after concatenation and standardization of image features related to CT and PET images, not all image features are used for model training. To reduce overfitting, complexity, training time, and to arrive at a stable and generalized model, feature selection is performed before model training. To select the most important features, sequential feature selection (SFS) can be applied, for example, at the beginning of each of the cycles shown in FIG. 5. In the SFS method, we start with an empty set of image features, and then add, one by one, the features that perform best in terms of improving the classification results with the already selected features. At each iteration, the image feature that gives the largest improvement is included in the set of selected features. The pseudocode that follows for the SFS method is expressed as follows:

[0070] The first step is to create an empty set.

number

[0071] The second step is to select the best remaining features.

number

[0072] And in the third step, J(Y k ∪{x +})>J(Y k ) (3) a.Y k+1 =Y k ∪{x + Update} (4a) b. Increase k to k+1 (4b) c. Go back to step 2 (4c) Here, Y refers to the set of all possible, i.e., predefined, collections of image features, and J is a measure of the performance of each machine learning architecture with a subset of image features in its arguments.

[0073] Candidate machine learning architectures considered for classification may include logistic regression, support vector machine (SVM), ensemble methods such as random forest, XGBoost and balanced random forest, especially for classification of inactive and aggressive prostate cancer. The performance of the candidate machine learning architectures is evaluated using, for example, five-fold cross-validation. Thus, the test images may include different subsets of the training images, thereby performing cross-validation. In five-fold cross-validation, five different subsets of the training images are selected to form the test images.

[0074] In a preferred embodiment, the multiple candidate machine learning architectures correspond to random forests with different hyperparameters, and the hyperparameters are successively selected according to a grid search. This allows for relatively low complexity while showing very good performance in cross-validation. By grid search, the hyperparameters of the random forests can be tuned, for example, using the sklearn library. The possible hyperparameter values ​​resulting from the grid search are shown in Table 3 below.

[0075] [Table 3]

[0076] FIG. 6 illustrates, in a schematic and exemplary manner, an embodiment according to some of the described aspects. A CT image 10' and a PET image 20' of a patient's upper body are provided, which can be viewed as a single dual image. In this dual image, a lesion image segment corresponding to the patient's prostate is identified, thereby generating a mask 30' for the dual image. Both the CT image 10' and the PET image 20' are pre-processed, where pre-processing 230 of the CT image 10', which can also be viewed as a correction, refers to clipping of the CT image values ​​to a range or window that includes image values ​​corresponding to a reference tissue region and image values ​​corresponding to a tissue region of interest, and a subsequent mapping of the range or window to the full range of possible CT image values, i.e., for example, intensities. The reference tissue region corresponds in particular to the patient's liver, and the tissue region of interest refers to the patient's prostate. Pre-processing 203 of the PET image 20' refers to normalization of the PET image values ​​to a reference image segment identified in the PET image 20' and corresponding to a reference tissue region, i.e., in particular to the patient's liver.

[0077] 6, CT image features are extracted from the preprocessed CT image 10' (block 204a), PET image features are extracted from the preprocessed PET image 20' (block 204b), and the extracted CT image features and the extracted PET image features are then fused into a joint set of dual CT / PET image features (204c). These steps together culminate in the extraction of a joint set of dual CT / PET image features (204).

[0078] Block 320 shown in Figure 6 indicates that not all previously extracted image features are used for classification, but only a selected subset thereof. Thus, while with reference to Figures 4 and 5 an embodiment has been described in which the selection of the image features to be used for classification is performed before their values ​​are extracted from the image, in other embodiments the values ​​of all possible, or at least a predefined, possibly relatively large collection of image features are extracted from the image, in which case only the values ​​of a selected subset of these image features are sent as input to the classification unit, in particular to its machine learning architecture.

[0079] The machine learning architecture included in the classification unit, which may also be considered as a machine learning model, is indicated in Fig. 6 by block 125. In the embodiment shown, it classifies on its output the lesions shown in the CT and PET images and masked by a mask 30' with a binary classification into inactive and malignant lesions, where in particular the lesions may be prostate tumor regions.

[0080] The classification performance of the system is measured using several performance measures for the embodiment in which prostate tumors are classified into inactive tumors and malignant tumors, and the liver is selected as the reference tissue region, both for 5-fold cross-validation with training data and for 5-fold cross-validation with separate test data.It should be noted that training data and test data can refer to training images and test images, respectively, or can refer to the image features extracted from training images and test images.The performance achieved in this particular embodiment is shown in Table 4 below.

[0081] [Table 4]

[0082] Figures 7A and 7B show confusion matrices corresponding to the performance of the system in the above embodiment, where the confusion matrix in Figure 7A is for validation on five-fold cross-validation, and the confusion matrix in Figure 7B is for validation on separate test data. As can be seen from Table 4 above and Figures 7A and 7B, the risk of confusion, i.e., the risk that the system will produce false negative or false positive classifications, is relatively low. In particular, the risk of misclassifying actually malignant prostate tumor regions as inactive tumor regions, i.e., the false negative rate, which is particularly problematic clinically, is found to be low for both validation schemes.

[0083] FIG. 8 shows a receiver operating characteristic (ROC) curve for an embodiment also related to Table 4 and FIGS. 7A and 7B when using test data for validation, and the area under the curve (ROC-AUC) is determined to be 0.90.

[0084] In Fig. 9, the dependence of classification accuracy on an increasing number of image features selected to be used in classification is shown for the cross-validation case. It can be seen from Fig. 9 that the performance measured in terms of classification accuracy increases relatively quickly initially, i.e., the first few most important image features are added. At the same time, the increase in performance becomes slower as more features of less importance are added. Also, as can be seen in Fig. 9, the accuracy may even decrease as more image features are added. However, using such more image features may be acceptable or even preferable, since it may lead to generalization of the model.

[0085] The embodiments disclosed herein relate to assessing the Gleason grade for prostate cancer. Gleason grading is the most commonly used tool to determine the prognosis of prostate cancer. However, Gleason grading is a manual process in which a pathologist manually evaluates multiple biopsy samples under a microscope and assigns a grade to the cancer based on its severity. Thus, in one embodiment, the Gleason grade group (inactive vs. aggressive) is used to determine the prognosis of prostate cancer. 68A machine learning based model is proposed to automatically classify tumor habitats in Ga-PSMA-PET / CT images. The present disclosure thus relates in particular to an artificial intelligence model for mapping Gleason grade scores to PSMA-PET / CT scans. Gleason grading capabilities are added to PSMA-PET / CT scans for tumor habitats with the intent of classifying them as malignant or inactive. Tagging habitats with Gleason grades to categorize them as malignant or inactive types aids in biopsy planning to extract correct tissue samples, and tagging aids in targeting malignant tumors during radiation therapy. The model can be trained by learning from ground truth of cancer regions (ROIs) classified with Gleason grade groups of PET and CT imaging data. The ground truth Gleason grades are classified by a physician after reviewing biopsy samples. This allows for extraction of correct biopsy samples and also allows for targeting of malignant tumors during treatment.

[0086] The proposed system and method address the following issues and opportunities in existing systems and methods: 1) SUV values ​​vary between vendors based on detector type, physiological conditions, acquisition and reconstruction parameters, 2) tagging habitats with Gleason grade to categorize them as aggressive or inactive type helps target the correct tissue for biopsy, 3) prostate biopsy entails the risk of false negatives based on sample location, which may affect cancer grading, and 4) PSMA-PET / CT imaging potentially offers superior detection of prostate cancer lesions with sensitivity exceeding that of multi-parametric MRI.

[0087] It is proposed to normalize the SUV values ​​of PSMA for liver tissue (e.g., using an adapted VGG-16 neural network), especially for the corresponding median SUV to support multiple vendors and machine variants based on the different luminescent crystal materials (LSO, BGO, LYSO) used in the machine. PET normalization also helps to normalize the SUV values ​​that vary with the patient's physiological conditions and the uptake time of the nuclear tracer. A multimodal fusion of handcrafted features from CT and PET images, such as radiomics features extracted from ROI bounding boxes, has been proposed along with a machine learning model to map Gleason scores and classify the habitat tumors as malignant or inactive. Concatenated features from PET and CT gave the best results. Various machine learning models have been tried to classify inactive and malignant prostate cancers, especially random forests, with good performance. To add Gleason grading capabilities to PSMA-PET / CT scans for tumor habitats, it is possible to follow the process shown in Figure 6 above, for example. In particular, this process may include pre-processing, mask generation of habitat ROIs (either by user or by automatic methods), multi-modal feature extraction, feature selection and fusion, and machine learning models to classify tumor habitats into inactive and malignant types.

[0088] In the embodiments described above, the dual-modal images each include a CT image and a PET image, but more generally, for example, the dual-modal images may each include a CT image and a functional image, which may be different from the PET image. Also, in many of the embodiments described above, the tissue region of interest is or is within the patient's prostate and the reference tissue region is or is within the patient's liver, but generally any other combination of tissue regions of interest and reference tissue regions is possible.

[0089] Those skilled in the art will understand and effect other variations to the disclosed embodiments, from a study of the drawings, the disclosure, and the appended claims, and in practicing the claimed invention.

[0090] In the claims, the words "comprises" and "comprises" do not exclude other elements or steps and the words "a" or "an" do not exclude a plurality.

[0091] A single unit or device may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0092] The procedures by one or more units or devices, such as providing dual images, identifying lesion image segments from reference image segments, normalizing lesion image segments in PET images, extracting image feature values, and classifying lesions, can be performed by any other number of units or devices, and these procedures can be implemented as program code means of a computer program and / or as dedicated hardware.

[0093] The computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless communication systems.

[0094] Any reference signs in the claims shall not be construed as limiting the scope.

[0095] A system for classifying lesions is provided, comprising an image providing unit for providing a dual-modal image of an imaging region, the dual-modal image including a CT image and a registered PET image, the imaging region including a tissue region of interest including a lesion and a reference tissue region, the system further comprising an identification unit for identifying respective lesion image segments in the CT image and the PET image corresponding to the lesion, and for identifying a reference image segment in the PET image corresponding to the reference tissue region, the system also comprising a normalization unit for normalizing the lesion image segment in the PET image relative to the reference image segment, an image feature extraction unit for extracting values ​​for a set of image features from both lesion image segments, and a classification unit for classifying the lesion based on the extracted values.

Claims

1. an image providing unit for providing a dual-modal image of an imaging region, the dual-modal image comprising a mutually registered computed tomography image and a positron emission tomography image, the imaging region comprising a) a tissue region of interest comprising a lesion to be classified, and b) a reference tissue region comprising healthy tissue; an identification unit for identifying, in each of the computed tomography image and the positron emission tomography image, a respective lesion image segment corresponding to the lesion in the tissue region of interest, and for identifying, in the positron emission tomography image, a reference image segment corresponding to the reference tissue region; a normalization unit for normalizing the lesion image segment in the positron emission tomography image to the reference image segment in the positron emission tomography image; an image feature extraction unit for extracting values ​​of a set of image features associated with the lesion from the lesion image segment in the computed tomography image and the normalized lesion image segment in the positron emission tomography image; a classification unit for classifying the lesions according to their severity based on the extracted image feature values; 1. A system for classifying lesions, comprising:

2. 2. The system of claim 1, wherein the identification unit identifies respective reference image segments in each of the computed tomography image and the positron emission tomography image corresponding to the reference tissue region, and the reference image segment in the positron emission tomography image is identified as the image segment aligned to the reference image segment identified in the computed tomography image.

3. 3. The system of claim 1, wherein the tissue region of interest refers to the patient's prostate, the lesion is a prostate tumor region, and the positron emission tomography image is generated using an imaging agent that binds to a prostate-specific membrane antigen.

4. The system of claim 3 , wherein the lesions are classified according to a binary classification of severity based on a Gleason score, the binary classification distinguishing between inactive and aggressive prostate tumor regions.

5. The system of claim 1 , wherein the reference tissue region is the patient's liver.

6. The set of image features may be: The image feature values ​​extracted from the lesion image segment in the computed tomography image include total energy, largest two-dimensional diameter in the column direction, sphericity, surface area, large area enhancement, and dependent entropy; The image feature values ​​are extracted from the normalized lesion image segments in the positron emission tomography images, and the image feature values ​​include 90th percentile, minimum, flatness, and sphericity. The system of claim 1 .

7. 3. The system of claim 2, wherein the classification unit comprises a machine learning architecture that receives as input a computed tomography image of the imaging region including the reference tissue region to identify the reference image segment in the computed tomography image, and determines as output the reference image segment in the computed tomography image.

8. 8. The system of claim 7, wherein the machine learning architecture included in the identification unit includes a convolutional neural network having a convolutional layer, an activation layer, and a pooling layer as a base network, and respective side sets of convolutional layers in a plurality of convolutional stages of the base network, the side sets being connected to respective ends of the convolutional layer of the base network at the respective convolutional stages, and the reference image segment in the computed tomography image is determined based on outputs of the plurality of side sets of convolutional layers.

9. The system of claim 1 , wherein the classification unit comprises a machine learning architecture suitable for receiving image feature values ​​as input and determining an associated lesion severity classification as output, in order to classify the lesion.

10. providing a dual-modal image of an imaging region, the dual-modal image comprising a mutually registered computed tomography image and a positron emission tomography image, the imaging region comprising: a) a tissue region of interest comprising a lesion to be classified; and b) a reference tissue region comprising healthy tissue; identifying respective lesion image segments in each of the computed tomography image and the positron emission tomography image corresponding to the lesion in the tissue region of interest, and identifying a reference image segment in the positron emission tomography image corresponding to the reference tissue region; normalizing the lesion image segment in the positron emission tomography image to the reference image segment in the positron emission tomography image; extracting values ​​of a set of image features associated with the lesion from the lesion image segment in the computed tomography image and the normalized lesion image segment in the positron emission tomography image; classifying the lesions according to severity based on the extracted image feature values; 1. A computer-implemented method for classifying lesions, comprising:

11. 11. A computer-implemented method for selecting a set of image features and a machine learning architecture for use by the system of claim 1 or for use in the method of claim 10 to classify a lesion, the method comprising: providing a set of dual-modal training images, each training image including a co-registered computed tomography image and a positron emission tomography image, each of which has identified a respective lesion image segment corresponding to a respective lesion in a tissue region of interest, and has provided an assigned severity classification for each of the lesions in the training images; selecting a candidate set of image features from a predefined collection of possible image features; generating a training data set by extracting corresponding values ​​of each of the candidate set of image features from each of the lesion image segments identified in the training images and assigning these values ​​to each of the severity classifications assigned to the lesions in each of the training images; selecting a candidate machine learning architecture suitable for receiving the image feature values ​​as input and determining an associated lesion severity classification as output; training the candidate machine learning architecture on the training dataset; providing a set of dual-modal test images, each test image including a computed tomography image and a positron emission tomography image, each of which has identified a respective lesion image segment corresponding to a respective lesion in a tissue region of interest, and has provided an assigned severity classification for each of the lesions in the test images; generating a test data set by extracting corresponding values ​​of each of the candidate set of image features from each of the lesion image segments identified in the test images and assigning these values ​​to each of the severity classifications assigned to the lesions in each of the test images; determining the performance of the candidate machine learning architecture on the test dataset based on a relationship between the assigned severity classifications provided for the lesions in the test images and the corresponding lesion severity classifications determined by the candidate machine learning architecture; repeating all the above steps with a larger number of image features in the candidate set of image features until a stopping criterion is met, where for each iteration, the candidate set of image features is expanded with a further image feature from a predefined collection of possible image features, the further image feature being the one that results in the highest increase in performance; repeating all the above steps with at least one candidate machine learning architecture, wherein the set of image features and the machine learning architecture used to classify lesions are selected from the respective candidates according to the determined performance; A method comprising:

12. The method of claim 11 , wherein the plurality of candidate machine learning architectures correspond to random forests with different hyperparameters.

13. The method of claim 12 , wherein the hyperparameters are selected successively according to a grid search.

14. The method of claim 11 , wherein the test images comprise different subsets of training images to perform cross-validation.

15. 11. A computer program for classifying lesions, the computer program comprising program code means for causing a system according to claim 1 to carry out the method according to claim 10 when the program is run on a computer controlling the system.