Method for detecting at least one lesion of pancreas of patient in at least one medical image
By segmenting the pancreas in medical images and measuring its features, combined with deep learning and logistic regression models, the challenge of early pancreatic cancer detection has been solved, improving detection accuracy and efficiency, and providing a risk assessment tool for early lesions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies are insufficient for the early and effective detection of pancreatic cancer, leading to diagnostic delays and limited treatment options. Furthermore, deep learning methods have failed to effectively identify the imaging features of pancreatic cancer, affecting the detection rate of early lesions.
A method implemented using computer devices predicts the probability map of the pancreas, main pancreatic duct, and lesions in medical images, segments the pancreas and identifies its head, body, and tail, calculates a score for the patient's pancreatic lesions by combining feature measurements, and assesses the risk of lesions using convolutional neural networks and logistic regression models.
It improves the accuracy and efficiency of early pancreatic cancer detection, reduces the workload of radiologists, provides objective tools to help clinicians make management decisions, and increases the detection rate of early lesions.
Smart Images

Figure CN121753065A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a method implemented by a computer device for detecting at least one lesion of a pancreas of a patient in at least one medical image. BACKGROUND
[0002] Pancreatic cancer is currently the 11th most common cancer worldwide and the 7th leading cause of cancer-related deaths. It is projected to increase by 78% between 2018 and 2040, making it a major and growing health problem. Due to the rising incidence of pancreatic cancer, coupled with its very low 5-year survival rate (only 9%), the disease is expected to become the third leading cause of cancer-related deaths by 2025.
[0003] Most patients with early-stage pancreatic cancer present with non-specific symptoms and, if examined, usually receive a routine portal phase computed tomography (CT). Image interpretation can be difficult because, at an early stage, pancreatic lesions are usually small (<2 cm) and isoattenuating, with reported sensitivities ranging from 58% to 77%. In addition, the heavy workload of radiologists and their level of expertise and experience can further affect the interpretation of CT scans. Due to the rapid progression of the disease, pancreatic cancer is mostly discovered at an advanced stage, when treatment options are limited, explaining the low 5-year survival rate. To date, only 10% of patients have undergone pancreatic resection, which is the only curative treatment.
[0004] However, in the last few years, there has been an increase in the proportion of patients diagnosed with stage 1A pancreatic cancer. In general, the description of stage 1A pancreatic cancer includes: the cancer is confined to the pancreas, with a maximum diameter of no more than 2 cm (0.8 inches), and has not spread to adjacent lymph nodes or distant sites. At this stage 1A, patients are more often eligible for pancreatic resection and adjuvant chemotherapy, and thus the 5-year survival rate for these patients is now more than 80%, highlighting the importance of early detection of pancreatic cancer. To identify imaging findings that might alert radiologists to the potential presence of pancreatic cancer, a retrospective analysis of CT scans of patients prior to histopathological diagnosis was performed. It is agreed in the scientific community that subtle secondary signs, such as main pancreatic duct (MPD) dilation, are often visible 1 year before cancer diagnosis. This is because pancreatic cancer is mainly ductal adenocarcinoma, with the malignant tumor causing MPD stenosis, leading to upstream dilation. Dilation is generally defined as a duct diameter greater than 3 mm in the head, and greater than 2 mm in the body and tail. In addition, dilation upstream of the stenosis, i.e. focal disappearance of the MPD lumen, is also a pathological pattern.
[0005] In this context, deep learning (DL) methods can play an important role in the daily practice of radiologists by issuing alerts that a patient is at risk of developing pancreatic cancer. For some diseases, such as breast cancer, promising results have been obtained: DL models significantly reduced the false positive and false negative rates on two large datasets, while substantially reducing the workload of radiologists. Related efforts have also been directed at pancreatic cancer, with many studies attempting to use DL for lesion detection. These studies proposed DL models to detect pancreatic tumors and validated them on independent patient databases. Although the results showed good prospects, these methods purely rely on DL and do not address the need to identify imaging findings that are predictive of pancreatic cancer, which can improve the detection rate of early lesions.
[0006] In this document, a method to predict that a patient is at risk of developing pancreatic cancer is proposed and evaluated. SUMMARY
[0007] To this end, a method for detecting at least one lesion of the pancreas of a patient in at least one medical image, implemented by a computer device, is proposed herein, the method comprising an inference phase, the inference phase comprising the steps of: (a) predicting at least one first probability map representing the probability that each pixel or voxel of the image belongs to a part of a pancreatic lesion, (b) predicting at least one second probability map representing the probability that each pixel or voxel of the image belongs to a part of the pancreas, (c) predicting at least one third probability map representing the probability that each pixel or voxel of the image belongs to a part of the main pancreatic duct, (d) segmenting the pancreas, the main pancreatic duct and the at least one lesion in the image based on the probability maps, (e) determining the head, body and tail of the segmented pancreas, (f) determining a first feature representing the risk of lesion based on the first probability map, (g) determining a second feature representing the maximum size of the lesion, (h) determining a third feature representing the maximum size of the main pancreatic duct in the head of the pancreas, (i) determining a fourth feature representing the maximum size of the main pancreatic duct in the body of the pancreas, (j) determining a fifth feature representing the maximum size of the main pancreatic duct in the tail of the pancreas, (k) determining a first score representing the probability that the patient has a pancreatic lesion based on at least the first feature, the second feature, and based on at least one of the third, fourth and fifth features.
[0008] A medical image can be defined as a visual representation of internal body structures or functions obtained using various imaging modalities. These images are captured and processed by hardware and software systems to generate two-dimensional (2D) or three-dimensional (3D) images that aid in the diagnosis, treatment planning, and monitoring of various medical conditions. Medical images often contain information about tissue density, composition, and function, which are interpreted by radiologists, physicians, or other medical professionals trained in image analysis.
[0009] The image can be a 3D image.
[0010] Medical images can take various forms depending on the imaging modality used and the type of medical condition being investigated. However, in accordance with the present disclosure, the image can be a CT scan image, such as a portal phase CT scan image.
[0011] CT scan images, also known as computed tomography images, are created by combining a series of X-ray images taken from different angles of the body. CT scans can provide detailed information about the internal structures and composition of organs, bones, and tissues.
[0012] The main difference between a portal phase CT scan image and a CT scan image lies in the focus of the imaging. A CT scan image is a medical imaging modality that uses a computed tomography (CT) scanner to provide detailed imaging of any part of the body. CT scans can be used to evaluate many different organs and tissues, such as the brain, chest, abdomen, and pelvis, and can be performed with or without the use of contrast agents.
[0013] On the other hand, a portal phase CT scan image specifically focuses on imaging the portal venous system, which includes veins that drain blood from the gastrointestinal tract, spleen, and pancreas to the liver. A portal phase CT scan is a type of CT scan specifically designed to image the abdominal portal venous system using a dedicated protocol that optimizes the imaging of this system.
[0014] A portal phase CT scan can help diagnose and monitor a range of diseases that affect the liver and pancreas, including tumors, infections, and inflammation. It can also help evaluate treatment response and guide further management of the disease. The portal venous system is an important component of the blood supply to the liver and other abdominal organs, and a portal phase CT scan can provide valuable information about its function and anatomy.
[0015] In the context of machine learning, the inference phase refers to the stage where a trained model is used to make predictions on new or unseen data. In this phase, the model receives input data and produces an output based on the patterns learned from the training data.
[0016] In other words, the inference phase is the process of applying a trained model to real-world data to make predictions or decisions.
[0017] A probability map (also called probability density map or probability distribution map) can refer to a representation that indicates the likelihood or probability of a certain event occurring at different locations or regions within an input space.
[0018] More specifically, in the context of image processing, a probability map can be a 2D or 3D grid or heat map that assigns a probability value to each pixel or voxel in an image, indicating the likelihood of the pixel or voxel belonging to a certain class or category. For example, in object detection, a probability map can be used to identify the location and extent of objects in an image by assigning higher probabilities to pixels or regions that are more likely to belong to the object of interest.
[0019] Probability maps are commonly used in machine learning to represent uncertainty or variability associated with different outcomes, and can be used to make predictions, classify inputs, or guide decision-making processes.
[0020] Segmentation can refer to the process of dividing an input image or data into multiple segments or regions (collections of pixels or voxels) based on certain criteria such as color, texture, intensity, shape, or other features. More precisely, image segmentation is the process of assigning a label, class, or category to each pixel or voxel in an image, such that pixels with the same label share certain features. The goal of segmentation can be to identify and isolate different objects or regions of interest within an image or data so that they can be analyzed, processed, or classified individually.
[0021] Algorithms and techniques that can be used for segmentation include thresholding, clustering, edge detection, region growing, and deep learning-based methods.
[0022] In digital image processing and computer vision, a feature is a piece of information about the content of an image, typically about whether a certain region of the image has certain properties. A feature can be a specific structure in the image, such as a point, edge, or object. A feature can also be the result of a general neighborhood operation applied to the image or a feature detection. Other examples of features are related to motion in a sequence of images or shapes related to curves or boundaries defined between different image regions.
[0023] More broadly, a feature is any piece of information relevant to solving a computational task related to a certain application. This is the same as the meaning of features in machine learning and pattern recognition, although image processing has a very complex set of features.
[0024] The pancreas is a glandular organ located in the abdomen that plays an important role in the digestive system and in regulating blood sugar levels. The head of the pancreas is the widest and most rightward portion of the organ, located next to the duodenum (the first part of the small intestine) and the bile duct. The tail of the pancreas is the narrower left end of the organ, which extends to the spleen. The body of the pancreas is the middle portion of the organ, which connects the head and tail and is located behind the stomach.
[0025] These three parts of the pancreas can be important for diagnostic and treatment purposes, as different diseases and conditions can affect specific parts of the organ.
[0026] Pancreatic lesions can refer to any abnormal growth, mass, or area of tissue in the pancreas that differs in appearance or texture from the surrounding tissue. This can include cysts, tumors, or other abnormal tissue growths, and can be benign or malignant. Pancreatic lesions can be detected through medical imaging examinations such as ultrasound, CT scans, or MRI, and may require further evaluation through biopsy or other diagnostic procedures to determine their nature and potential impact on an individual's health. The management and treatment of pancreatic lesions will depend on factors such as the size, location, and type of the lesion, as well as the individual's overall health condition and medical history.
[0027] The first score quantifies the likelihood that an area detected in the pancreas is a lesion (such as a tumor, cyst, or other abnormality), whether benign or malignant. The goal of this first score is to provide clinicians with an objective tool to help guide their management decisions, such as whether to recommend further diagnostic testing or treatment.
[0028] The first score can be displayed as is on the interface, or explained to the user to indicate whether a lesion exists based on a threshold. For example, if the first score is greater than the threshold, the interface can display a message informing the user of the presence of pancreatic lesions or a risk of lesions. In other words, a classification (e.g., lesion / no lesion) can be derived based on the first score.
[0029] The first score may be determined based solely on the first and second features, and on at least one of the third, fourth, and fifth features.
[0030] The first score can be determined based at least on the first feature, the second feature, and the maximum value of the third, fourth, and fifth features. It should be noted that in this case, the maximum value of the third, fourth, and fifth features is a value derived from the third, fourth, and fifth features, because the third, fourth, and fifth features need to be calculated to determine their maximum value.
[0031] Several types of pancreatic lesions can occur, some of which are benign (non-cancerous) and some are malignant (cancerous). Here are some examples of possible pancreatic lesions: - Pancreatic cysts: These are fluid-filled sacs that can develop within the pancreas. Most pancreatic cysts are benign, but some may be precancerous or cancerous.
[0032] - Pancreatic pseudocyst: This is a collection of fluid and debris that can form after an acute attack of pancreatitis. Pseudocysts are usually benign, but they can become infected or rupture.
[0033] - Serous cystadenoma: is a benign pancreatic tumor filled with clear, watery fluid.
[0034] - Mucinous cystic neoplasm: is a pre-cancerous pancreatic tumor filled with thick, sticky mucus.
[0035] - Intraductal papillary mucinous neoplasm: is a pre-cancerous pancreatic tumor that grows within the pancreatic duct and produces mucus.
[0036] - Solid-pseudopapillary neoplasm: is a rare pancreatic tumor that is usually benign but can become malignant.
[0037] - Pancreatic neuroendocrine tumor: is a rare tumor that originates from hormone-producing cells in the pancreas. Most pancreatic neuroendocrine tumors are non-cancerous, but some can be malignant.
[0038] - Pancreatic adenocarcinoma: is the most common type of pancreatic cancer, usually originating from cells lining the pancreatic ducts.
[0039] - Pancreatic lymphoma: is a rare type of pancreatic cancer that originates from the lymphatic tissue of the pancreas.
[0040] - Metastatic pancreatic tumor: is a tumor that originates from another part of the body and spreads to the pancreas, such as from the lung, breast, or colon.
[0041] The first feature representing the risk of a lesion based on the first probability map can be a number between 0 and 1.
[0042] To compute this first feature, all candidate pixels or voxels located outside the segmented pancreas can be removed (e.g., all pixels or voxels associated with a probability higher than 0). Then, for each remaining connected component, the risk of a lesion between 0 and 1 can be computed by averaging the probabilities of all its pixels or voxels.
[0043] Connected components with a risk of a lesion lower than 0.05 can be automatically removed.
[0044] In other words, wherein step (f) can comprise the following sub-steps: - removing all pixels or voxels not connected to the segmented pancreas, - for each remaining connected component, computing the risk of a lesion by averaging the probabilities of all its pixels or voxels, and removing connected components with a risk of a lesion lower than a threshold value.
[0045] The threshold value can be, for example, lower than 0.05.
[0046] A connected component is a group of pixels connected to each other by an adjacency relation. Connected components can be groups of pixels in an image that form a coherent object or region.
[0047] A common way to identify connected components in an image is by using a connected component labeling algorithm. These algorithms analyze the relationships between pixels in the image and assign a unique label to each connected component. Once the connected components are labeled, various operations can be performed on them, such as calculating their size, shape, or orientation, or applying filters to enhance or remove them.
[0048] The second feature representing the size of the lesion can be the largest diameter of the segmented pancreatic lesion. To determine the largest diameter of the lesion, once the lesion is segmented, in the case of a 3D image, its diameter can be measured along the axial view layer by layer, for example using the skimage library. The 2D Feret diameter of each layer can be calculated. If the lesion is not segmented, the diameter can be automatically set to 0.
[0049] The second feature representing the size of the lesion can also be determined by: - Extracting the lesion contour using edge detection algorithms or morphological operations. The contour is the boundary of the lesion and can be used to measure its size - Measuring the largest diameter of the lesion by finding the longest distance between two points on the lesion contour. This can be done through distance transform, skeletonization, or other geometric measurements.
[0050] The first, second, and third probability maps can be obtained using at least one convolutional neural network or model, for example a fully convolutional neural network.
[0051] The convolutional neural network can be nnUNet.
[0052] nnUNet automatically designs a segmentation pipeline based on the UNet architecture by relying on heuristics applied to the data to estimate key parameters. Dataset properties are estimated to automatically perform preprocessing steps. The design choices of the model (number of layers, kernel size, convolution blocks, etc.) are then automatically defined. The training pipeline (data augmentation, scheduled learning rate, etc.) is also implemented.
[0053] The first, second, and third probability maps can be obtained using multiple or a set of models or convolutional neural networks, each model being able to predict the first, second, and third probability maps.
[0054] The set of models (for example a set of k models) can be trained using a k-fold cross-validation setup. The outputs of the models (each probability map as well as the segmentation of the pancreas, pancreatic lesions, and main pancreatic duct) are then averaged pixel-wise or voxel-wise to generate a single final probability map. The final segmentation is obtained by assigning to each pixel or voxel its most likely class (i.e. the class with the highest probability). The lesion probability map is obtained by looking at the probability of each pixel or voxel to be a lesion.
[0055] More generally, k-fold cross-validation is a known technique in machine learning for evaluating model performance and preventing overfitting. In k-fold cross-validation, the original dataset is split into k subsets or folds of approximately equal size. The model is then trained on k-1 folds and validated on the remaining fold.
[0056] The k-fold cross-validation process is repeated k times, with each fold being used as the validation set once. The performance of the model is then averaged over all k iterations to obtain a more reliable estimate of the performance.
[0057] During training, image preprocessing (i.e. resampling and intensity normalization) can be determined and performed entirely by nnUNet.
[0058] The identification of the head, body and tail of the segmented pancreas can be obtained by performing the following steps.
[0059] - Along the axial view, extract the morphological skeleton of the pancreas segmentation layer by layer, for example using the skeletonize function in the skimage (or scikit-image) library, - Obtain the 3D skeleton of the pancreas and convert the obtained 3D pancreas skeleton into a graph, for example using the network library, - Given the image point located at the most right lower front part of the abdomen, identify the point in the graph that is farthest from it. This point can thus be considered as the end point of the tail of the pancreas, - Identify the head of the pancreas by looking at the point in the graph that is farthest from the tail, - Once the head and tail are identified, calculate the shortest path between the head and the tail (for example using the Dijkstra algorithm) to finally obtain the centerline that passes through the pancreas and connects its two end points, - Divide the centerline into three parts, namely the head, body and tail. The first 25% can be considered as the tail, the next 50% can be the body and the last 25% can be the head.
[0060] - For each voxel of the pancreas segmentation, find its nearest point on the centerline and assign it to the corresponding head, body or tail location or part.
[0061] Once the pancreas is subdivided, the main pancreatic duct can also be subdivided. For each voxel of the segmented main pancreatic duct, its nearest point on the centerline can be calculated. Then, depending on the location of its nearest point on the centerline, the voxel is assigned a head, body or tail location of the segmented pancreas.
[0062] Once this is done, the main pancreatic duct diameter for each part of the pancreas (head, body, tail) can be measured. The main pancreatic duct diameter is calculated layer by layer along the axial view, for example using the IMEA library. In addition, it can be checked whether the diameter is calculated in the head, body or tail.
[0063] Based on this, a number of measurements can be extracted: minimum, maximum, mean, median, percentile of the main pancreatic duct diameter in the head, body and tail. The measurements of the three parts can be aggregated to compute the minimum, maximum, mean, median, percentile of the main pancreatic duct throughout the pancreas.
[0064] Another advantage of subdividing the pancreas into a tail, body and head is that the lesion location can also be extracted by computing the distance between the lesion voxels and the centerline. Once done, the part (head, body, tail) where most of the lesion voxels are located can be determined and assigned as the lesion location.
[0065] Overall, the pancreas subdivision can be used to identify the location of other anatomical structures in the pancreas and to extract features regionally.
[0066] The computations can be performed using the IMEA library.
[0067] The maximum diameter of the main pancreatic duct can also be computed by taking the maximum diameter between the head, body and tail. For each part, the main pancreatic duct diameter is set to 0 if there is no segmentation.
[0068] The determination of the first score can be performed by a first logistic regression model.
[0069] The method described herein can further comprise a step of determining a second score based on at least the third, fourth and fifth features, the second score being representative of the probability of main pancreatic duct dilation.
[0070] The second score can be displayed in the interface as is or interpreted to inform the user of the dilated or non-dilated state of the main pancreatic duct based on a threshold. For example, if the second score is greater than the threshold, the interface can display a message informing the user that the main pancreatic duct is in a dilated state. In other words, a classification (e.g. dilated / non-dilated) can be derived based on the second score.
[0071] The second score can be determined based on the third, fourth and fifth features only.
[0072] The determination of the second score can be performed by a second logistic regression model.
[0073] The first and / or second logistic regression model can be trained using the Scikit-Learn library with default hyperparameters.
[0074] In addition to the features described above, the first score and / or the second score can be determined based on at least one other feature from the following list: - first order statistics of the pancreas diameter, lesion diameter, main pancreatic duct diameter and / or common bile duct diameter. These features can be regionalized to each part (head, body, tail) using the pancreas subdivision, creating new features.
[0075] - the location of the lesion, for example the lesion is located in the head, body and / or tail, - at least one radiomic feature of the pancreas, the lesion, the main pancreatic duct and / or the common bile duct. These features can be regionalized to the different parts of the pancreas (head, body, tail) using pancreas segmentation, creating new features.
[0076] - 3D shape radiomic features of the pancreas, the lesion, the main pancreatic duct and / or the common bile duct. These features can be regionalized to the different parts of the pancreas (head, body, tail) using pancreas segmentation.
[0077] - 2D shape radiomic features of the pancreas, the main pancreatic duct, the lesion and / or the common bile duct and / or first order statistics of these features. These features can be regionalized to the different parts of the pancreas (head, body, tail) using pancreas segmentation.
[0078] In the medical field, radiomics is a method that uses data characterization algorithms to extract a large number of features from medical images. These features, called radiomic features, have the potential to reveal image patterns and characteristics that are not visible to the naked eye.
[0079] First order statistics, also known as descriptive statistics, are summary measures that describe the basic features of a data set without making any assumptions about its underlying distribution. These statistics include measures of central tendency such as mean, median, and mode, which provide information about the typical or average value of the data. They also include measures of dispersion or variability such as range, variance, and standard deviation, which provide information about the extent of the spread of the data. Other first order statistics include measures of skewness and kurtosis, which describe the shape of the distribution. These statistics are widely used in data analysis and can help provide insights into the characteristics of a data set.
[0080] Radiomic features can be grouped into the following groups: - First order features - 3D shape features - 2D shape features - Gray Level Co-occurrence Matrix (GLCM) features - Gray Level Size Zone Matrix (GLSZM) features - Gray Level Run Length Matrix (GLRLM) features - Neighboring Gray-Tone Difference Matrix (NGTDM) features - Gray Level Dependence Matrix (GLDM) features More specifically: - First order features describe the distribution of voxel intensities within the image region defined by the mask by common and basic measures. The first order features can include at least one of the following features: Energy, Total Energy, Entropy, Minimum, 10th percentile, 90th percentile, Maximum, Mean, Median, Interquartile Range, Range, Mean Absolute Deviation (MAD), Robust Mean Absolute Deviation (rMAD), Root Mean Square (RMS), Standard Deviation, Skewness, Kurtosis, Variance, Homogeneity.
[0081] - 3D shape features are descriptors of the 3D size and shape of the ROI. These features are independent of the distribution of grey intensity within the ROI and are therefore computed based on the non-derived image and the mask only. The 3D shape features can include at least one of the following features: Mesh volume, Voxel volume, Surface area, Surface area to volume ratio, Sphericity, Compactness 1, Compactness 2, Spherical asymmetry, Maximum 3D diameter, Maximum 2D diameter (slice, row or column), Length of the long axis, Length of the short axis, Length of the minimum axis, Elongation, Flatness, - 2D shape features are descriptors of the 2D size and shape of the ROI. These features are independent of the distribution of grey intensity within the ROI and are therefore computed based on the non-derived image and the mask only. The 2D shape features can include at least one of the following features: Mesh surface, Pixel surface, Perimeter, Perimeter to surface ratio, Sphericity, Spherical asymmetry, Maximum 2D diameter, Length of the long axis, Length of the short axis, Elongation.
[0082] These radiomics features are all defined at the following website: https: / / pyradiomics.readthedocs.io / en / latest / features.html# and can be implemented by using the pyradiomics library.
[0083] The mathematical definitions of these features are independent of the imaging modality and can also be found in the literature: Galloway, Mary M (1975). "Texture analysis using gray level runlengths". Computer Graphics and Image Processing. 4 (2): 172--179. doi:10.1016 / S0146-664X(75)80008-6. Pentland AP (June 1984). "Fractal-based description of natural scenes". IEEE Transactions on Pattern Analysis and Machine Intelligence. 6(6): 661-74. doi:10.1109 / TPAMI.1984.4767591. PMID 22499648. S2CID 17415943. Amadasun M, King R (1989). "Textural features corresponding to textural properties". IEEE Transactions on Systems, Man, and Cybernetics. 19(5): 1264-1274. doi:10.1109 / 21.44046. Thibault G, Angulo J, Meyer F (March 2014). "Advanced statistical matrices for texture characterization: application to cell classification". IEEE Transactions on Bio-Medical Engineering. 61(3): 630-7. doi:10.1109 / TBME.2013.2284600. PMID 24108747. S2CID 11319154. The common bile duct can be determined in a similar manner as the main pancreatic duct, by using probability maps and / or segmentation.
[0084] It is also proposed herein a computer software comprising instructions for implementing at least one of the methods according to at least one part of the present document, when said software is executed by a processor.
[0085] It is also proposed herein a computer device comprising: - an input interface for receiving a medical image, - a memory for storing instructions of a computer program according to at least the preceding claim, - a processor for accessing said memory to read the above instructions and to execute a method according to the present document, - an output interface for providing an indication based on said first score and / or said second score.
[0086] It is also proposed a computer readable non-transitory recording medium having recorded thereon computer software for implementing the method according to the present description when said computer software is executed by a processor. BRIEF DESCRIPTION OF DRAWINGS
[0087] Other features, details and advantages will appear in the following detailed description and in the appended drawings, in which: - Figure 1 An example of a computer device according to the present description is schematically illustrated, - Figure 2 Data used to build the training and test sets are illustrated, - Figure 3 A flow of the training phase of the method according to the present description is illustrated, - Figure 4 A flow of the inference phase of the method according to the present description is illustrated, - Figure 5 Subdivisions of three different segmented pancreases are illustrated, - Figures 6 to 9 Examples of model segmentation are illustrated, where the left image illustrates an input image and the right image illustrates the segmentation elements on said image (pancreas in red, pancreatic lesions in green, main pancreatic duct in blue).
[0088] - Figure 10 and Figure 11 The performance of the model on the test set is illustrated.
[0089] The accompanying drawings contain meaningful colors. Although the present application will be published in black and white form, color versions of the drawings have been previously filed with the patent office. DETAILED DESCRIPTION
[0090] Figure 1 An example of a computer device 1 according to the present description is schematically illustrated. The computer device 1 comprises: - an input interface 2, - a memory 3 for storing at least instructions of a computer program, - a processor 4 accessing the memory 3 to read the above instructions and to execute the method described herein, - an output interface 5.
[0091] Dataset Data were collected from 9 medical centers in Europe, the United States and Brazil. Inclusion criteria were as follows: (i) presence of a portal phase CT scan; (ii) maximum slice thickness of 3 mm; (iii) For patients diagnosed with pancreatic tumors, the study is obtained before treatment or surgery.
[0092] This resulted in a total of 2,890 cases, which were further divided into a training set and an independent test set, with 2,134 and 756 participants, respectively (see [link]). Figure 2 ).
[0093] The training set consisted of portal venous phase CT scans from 2,134 patients from 5 institutions, of whom 1,692 had pancreatic tumors. Diagnosis was obtained from biopsy reports in 78% of the subjects, or from the International Classification of Diseases (ICD-10) C25 code for the remaining subjects.
[0094] The training set also included 442 control subjects who had no evidence of pancreatic lesions according to radiological reports.
[0095] Table 1 below describes the patient characteristics.
[0096]
[0097] Table 1: Demographic and clinical information for different datasets. Unless otherwise specified, data represent the percentage of patients in parentheses. * Data represent the median interquartile range in square brackets. CT: Computed tomography. GE: General Electric. PDAC: Pancreatic ductal adenocarcinoma. NET: Neuroendocrine tumor. IPMN: Intraductal papillary mucinous tumor. MCN: Mucinous cystic tumor. SCN: Serous cystadenoma. NA: Not applicable.
[0098] The 2,134 participants included 1,174 (55%) women and 960 (45%) men, with a median age of 64 years (range [56, 73] years). 1,184 patients had pancreatic ductal adenocarcinoma (PDAC), 134 had neuroendocrine tumors (NETs), and 158 had unclassified solid lesions. Additionally, 81 participants had intraductal papillary mucinous tumors (IPMNs), 34 had mucinous cystic tumors (MCNs), 42 had serous cystadenomas (SCAs), and 59 had unclassified cystic lesions. Finally, 43% (907) of the participants had dilated MPD.
[0099] The independent test set includes public and private data from 756 participants across 4 institutions (see [link to test set]). Figure 2These institutions differ from those used in the training set. The public dataset contains 361 portal venous phase CT scans, of which 281 showed pancreatic lesions and 80 were healthy cases. Cancer cases were obtained from the Medical Decathlon Challenge. Control subjects were drawn from the Cancer Imaging Archive pancreatic CT dataset. The private dataset consists of 212 portal venous phase CT scans from patients with histopathologically confirmed pancreatic tumors, and 183 portal venous phase CT scans from a second medical institution without radiographic evidence of pancreatic lesions. No demographic information was available for the test set subjects. This database consisted of patients with PDAC (n=360), NET (n=48), and unclassified solid lesions (n=12), and patients with IPMN (n=18) and unclassified cystic lesions (n=53) (see Table 1). Finally, 276 subjects had dilated MPD. Annotation protocol Each portal venous phase CT scan was reviewed and annotated by one of a team of nine radiologists. The annotator segmented the pancreas on each image. Segmentation was performed systematically as the radiologist identified lesions. The radiologist characterized the tumor type based on the CT scan. Possible types were: PDAC, NET, IPMN, MCN, and SCA. Cases where the lesion type could not be determined were marked as unclassified. Finally, the radiologist segmented the MPD when it was visible and visually assessed its expansion.
[0100] Detection pipeline Segmentation A 3D nnUNet was trained in a 5-fold cross-validation setting to segment the pancreas, lesions, and MPD. Image preprocessing (i.e., resampling and intensity normalization) was entirely determined and performed by the nnUNet. When applied to new images, the trained nnUNet generated segments of the pancreas, pancreatic lesions, and MPD, as well as a lesion probability map assigning a lesion probability to each voxel (see [link to documentation]). Figure 3 ).
[0101] During the inference phase, five nnUNets are used to predict probability maps for each voxel, assigning a probability of pancreas, lesion, or MPD. These five lesion probability maps are then averaged voxel-by-voxel to generate a single final probability map. The final segmentation is obtained by assigning each voxel its most likely class (i.e., the class with the highest probability). The lesion probability map is obtained by examining the probability of each voxel being a lesion.
[0102] Feature extraction Using the output of nnUNet, the feature extraction is as follows: (i) Calculate the lesion risk between 0 and 1 based on the 3D probabilistic map provided by nnUNet. To do this, all candidate lesions outside the pancreas are eliminated using a segmentation map. Then, for each connected component, the lesion risk between 0 and 1 is calculated by averaging the probabilities of all its voxels. Connected components with a lesion risk below 0.05 are automatically removed; (ii) Maximum diameter of lesions segmented by nnUNet. Once a lesion is segmented, its diameter can be measured layer by layer along the axial view using the skimage library. The 2D Feret diameter of each layer can be measured, and many values can be calculated: minimum, maximum, mean, median, percentile. If the lesion is not segmented, the diameter is automatically set to 0. The common bile duct can also be segmented. Once the segmentation network is trained, the diameter of the common bile duct can be calculated using the exact same technique.
[0103] (iii) Maximum MPD diameter of the pancreatic head, body, and tail. To regionally measure the MPD diameter, we first need to identify the head, body, and tail of the pancreas. This is accomplished by performing the following steps: 1) Segment the pancreas (using the nnUNet described above).
[0104] 2) Along the axial view, use the skeletonize function in skimage to extract the morphological skeleton of the pancreas segmentation layer by layer.
[0105] 3) Obtain the 3D skeleton of the pancreas by converting it into a graph using the network library.
[0106] 4) Given an image point located at the rightmost lower front part of the abdomen, identify the point in the image that is furthest from it. Therefore, this point can be regarded as the endpoint of the pancreatic tail.
[0107] 5) Identify the head of the pancreas by looking at the point furthest from the tail in the image.
[0108] 6) Once the head and tail are identified, use Dijkstra's algorithm to calculate the shortest path between the head and tail to obtain the centerline that passes through the pancreas and connects its two endpoints.
[0109] 7) Divide the centerline into three parts. The first 25% is considered the tail, the next 50% is the body, and the last 25% is the head.
[0110] 8) For each voxel of the pancreas segment, find the point closest to it on the center line and assign it its corresponding position.
[0111] Figure 5 This illustrates how this technique allows the pancreas to be identified as three sub-parts. The figure shows three distinct segments of the pancreas (red in the top image) and the subdivision results (bottom image), where the head (green), body (blue), and tail (red) are subdivided.
[0112] Once the pancreas is subdivided, the MPD can also be subdivided. For each voxel in the MPD subdivision, its nearest point on the center line is calculated. Then, based on the location of its nearest point on the center line, the voxel is assigned a position (head, body, tail).
[0113] Once completed, the MPD diameter for each section can be measured. The MPD diameter is calculated layer by layer along the axial view using the IMEA library. Additionally, we check whether the diameter is calculated for the head, body, or tail. Based on this, many measurements can be extracted: the minimum, maximum, average, median, and percentile of the MPD diameter for the head, body, and tail.
[0114] The measurements from the three parts can be aggregated to calculate the minimum, maximum, average, median, and percentile of the entire MPD.
[0115] Another advantage of pancreatic subdivision is that lesion locations can be extracted by calculating the distance between the lesion voxel and the centerline. Once completed, the locations of most lesion voxels (head, body, and tail) can be identified, and the results can be assigned as lesion locations.
[0116] Overall, pancreatic subdivision can be used to identify the location of other anatomical structures within the pancreas and to extract features regionally.
[0117] The calculations were performed using the IMEA 20 library. The maximum pancreatic MPD diameter was also calculated by taking the largest diameter between the head, body, and tail. For each region, if no segmentation was performed, the MPD diameter was set to 0.
[0118] Logistic regression It is recommended to use two logistic regression models to predict lesion presence and MPD expansion based on previously defined features. To this end, nnUNet is applied to the validation set at each fold, generating a total of 2134 3D probability maps and segments from which features are extracted. The two logistic regression models are trained using the Scikit-Learn library with default hyperparameters.
[0119] The first method predicts lesion presence based on three features: 1) Risk of lesions; 2) Diameter of the lesion; 3) Maximum MPD diameter of the pancreas.
[0120] The second method uses the MPD diameter of the pancreatic head, body, and tail to predict MPD dilation.
[0121] Once the two logistic regression models are trained, the evaluation of test subjects involves three steps: 1) Use nnUNet to calculate the lesion probability map and segment the pancreas, lesions, and MPD; 2) Feature extraction; 3) Apply the two previously trained logistic regression models to predict lesion presence and MPD expansion (see...). Figure 4 ).
[0122] In this example, lesion risk, maximum lesion diameter, and maximum MPD diameter are used to predict lesion presence. Regarding MPD, the MPD diameters of the head, body, and tail are considered.
[0123] In addition to the features mentioned above, each logistic regression model may also use at least one other feature from the following list as input features: - First-order statistics of pancreatic diameter, lesion diameter, main pancreatic duct diameter, and / or common bile duct diameter. These features can be regionalized into different parts (head, body, tail) using pancreatic subdivision, thereby creating new features.
[0124] - Location of the lesion, for example, the lesion is located in the head, body, and / or tail. - At least one radiomics feature of the pancreas, lesion, main pancreatic duct, and / or common bile duct. These features can be regionalized into different parts (head, body, tail) using pancreatic subdivision, thereby creating new features.
[0125] - 3D shape imaging features of the pancreas, lesions, main pancreatic duct, and / or common bile duct. These features can be regionalized to individual parts (head, body, tail) using pancreatic subdivision.
[0126] - 2D shape radiomics features of the pancreas, main pancreatic duct, lesions, and / or common bile duct, and / or first-order statistics of these features. These features can be regionalized to different parts (head, body, tail) using pancreatic subdivision.
[0127] Statistical analysis To evaluate performance, receiver operating characteristic (ROC) curves were constructed by plotting sensitivity against false positive rate at different thresholds. The area under the curve (AUC) and sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were measured at the operating point that maximizes the balance between accuracy and performance. The performance of logistic regression in estimating lesion presence and predicting MPD expansion was evaluated at the case level. For example, in the case of lesion detection, sensitivity was defined as the ratio of the number of patients whose lesions were correctly detected by the model to the total number of patients with pancreatic lesions. Other metrics were calculated and defined accordingly.
[0128] Use bootstrapping to provide median and 95% confidence intervals (CI) for AUC, sensitivity, specificity, PPV, and NPV.
[0129] Finally, segmentation performance was evaluated by calculating the Dice and Normalized Surface Dice (NSD) scores between the reference and predicted segments for the test patients. NSD is well-suited for evaluating small structures such as MPD because it allows for a tolerance error between the reference and predicted segments21. In this study, the tolerance error was set at 2 mm along each spatial dimension.
[0130] Results Detection of pancreatic tumor patients The model's performance on the test set is reported as follows: Figure 10 and 11 And in Table 2 below.
[0131] Figure 10 The ROC (Receiver Operating Characteristic) curve (AUC: Area Under the Curve) of the logistic regression model predicting the presence of lesions is shown. The center line represents the mean curve, and the shaded area represents the 95% confidence interval. Figure 11 This is the model confusion matrix obtained at the operation point that maximizes the balance accuracy.
[0132]
[0133] Table 2: Lesion detection and assessment metrics obtained by the model on the test set and in specific patient subgroups. For each group within parentheses, the number of pancreatic tumor patients is compared to the total number of patients. Data are median values of the 95% confidence interval within square brackets. AUC: Area under the curve. PPV: Positive predictive value. NPV: Negative predictive value. D: Diameter. PDAC: Pancreatic ductal adenocarcinoma. NET: Neuroendocrine tumor. IPMN: Intraductal papillary mucinous tumor.
[0134] The model achieved an AUC of 0.98 (95% CI: 0.97, 0.99), a sensitivity of 0.94 (469 out of 493 cases, 95% CI: 0.92, 0.97), and a specificity of 0.95 (246 out of 262 cases, 95% CI: 0.92, 0.98).
[0135] To compare with other methods, Table 2 also reports the model's performance on public data ( Figure 2 The performance on the test set is as defined in [the standard test set]. The evaluation metrics are similar to those obtained on the entire test set.
[0136] The model also assessed subjects with specific lesion characteristics and types. Further investigation was conducted in five subgroups of the cohort: (i) Patients with lesions less than 2 cm in diameter; (ii) Patients with isodense lesions; (iii) PDAC patients; (iv) NET patients; (v) IPMN patients.
[0137] The performance of these subgroups is reported in Table 2. AUC, sensitivity, and specificity remained consistent across the subgroups and were comparable to those obtained on the entire test set. The model performed best on the NET subgroup with a sensitivity of 1.0 (95% CI: 0.98, 1.0). Specificity was slightly weaker on small lesions compared to other subgroups.
[0138] Feature importance on lesion detection sensitivity The logistic regression model used to predict the presence of lesions in patients is based on three features: lesion risk, lesion diameter, and MPD diameter. To evaluate the impact of combining these features on performance, an ablation study was conducted. Two additional logistic regression models were trained: the first using two features, lesion risk and lesion diameter; the second using only lesion risk. Table 3 below reports the sensitivity of these three models on the test set and the aforementioned subgroups.
[0139]
[0140] Table 3: The sensitivity of logistic regression depends on the features used to train it. Median, with the 95% confidence interval in square brackets. D: Diameter. PDAC: Pancreatic ductal adenocarcinoma. NET: Neuroendocrine tumor. IPMN: Intraductal papillary mucinous tumor.
[0141] In all groups, the use of MPD diameter and lesion diameter systematically improved the sensitivity of lesion detection. Across the entire test set, the addition of MPD diameter and lesion diameter improved sensitivity by 4% compared to the baseline model using only lesion risk. The impact of these two features was particularly significant for isodense lesions, with a 10% improvement in sensitivity compared to the model using only one feature.
[0142] MPD dilation detection performance Regarding the performance of MPD expansion, the logistic regression model was also evaluated on the test set. The results are reported in Table 4 below.
[0143]
[0144] Table 4: Evaluation metrics obtained by the model for MPD dilation detection. Data are the medians of the 95% confidence intervals within square brackets. MPD: Main pancreatic duct. AUC: Area under the curve. PPV: Positive predictive value. NPV: Negative predictive value.
[0145] The results showed an AUC of 0.97 (95% CI: 0.96, 0.98), a sensitivity of 0.94 (259 out of 276 cases, 95% CI: 0.89, 0.97), and a specificity of 0.90 (432 out of 480 cases, 95% CI: 0.86, 0.94).
[0146] Segmentation performance The segmentation predicted by nnUNet was evaluated quantitatively and qualitatively. The Dice scores between the reference segmentation of the patient's pancreas, lesion, and MPD and the nnUNet segmentation are reported in Table 5 below.
[0147]
[0148] Table 5: Dice and NSD measured across the entire test set and public data. Dice was calculated for the pancreas, lesions, and main pancreatic duct. NSD was measured only for the MPD. Data are mean plus / minus standard deviation. NSD: Normalized surface Dice.
[0149] NSD scores were calculated only for MPD. To compare the segmentation model with other studies, Dice and NSD scores were also measured only on publicly available data. Figures 6 to 9 Qualitative examples of model segmentation are provided. In these figures, the left image shows the input image (corresponding to a patient's axial portal venous phase CT slice). Figures 6 to 8 The white arrow indicates the location of the lesion. Figure 9 The right image shows the segmentation elements on the image (pancreas in red, pancreatic lesions in green, main pancreatic duct in blue - PDAC: pancreatic ductal adenocarcinoma, IPMN: intraductal papillary mucinous tumor, NET: neuroendocrine tumor, MPD: main pancreatic duct).
[0150] The mean Dice scores for pancreas, lesions, and MPD across the entire test set were 0.91 (±0.06), 0.69 (±0.34), and 0.58 (±0.37), respectively. The NSD score for MPD was 0.71 (±0.39).
[0151] Discussion This paper presents a method for automatically detecting pancreatic lesions (such as tumors) and identifying cases of MPD expansion. The proposed method was validated on an independent cohort of 756 subjects. It demonstrates how utilizing MPD expansion information can improve the sensitivity of lesion detection compared to basic methods that rely solely on segmentation network output. Finally, the model's ability to correctly locate lesions was evaluated by assessing its segmentation performance.
[0152] Patient subgroups were assessed based on lesion characteristics and type. Similar AUC, sensitivity, and specificity were observed across different subgroups, highlighting the robustness of the proposed method. Greater variability was observed in PPV and NPV cases, primarily due to imbalances between healthy and diseased subjects in some subgroups, such as NET (49 out of 312) and IPMN (18 out of 290).
[0153] Results on publicly available data can be compared with other deep learning models that use this database to test their method. Our method shows significantly better performance than existing techniques, with an AUC of 0.99 (95% CI: 0.98, 0.99).
[0154] The performance improvement can be primarily attributed to the use of multiple specific features to predict lesion presence. While conventional methods often rely solely on segmentation generated by convolutional neural networks to predict lesions on images, the proposed method combines lesion risk with MPD diameter and lesion diameter to predict lesion presence in subjects via logistic regression. To assess the impact of combining these features, the sensitivity of logistic regression based on the features used to train it was measured. Compared to models using two or three features, the logistic regression model, which is closest to existing methods and uses only lesion risk, had lower sensitivity across all subgroups. In particular, using MPD diameter systematically improved sensitivity, especially for isodense lesions (an increase of 5%). An improvement in PDAC detection sensitivity (an increase of 5%) was also observed when using lesion and MPD diameter. A milder effect (a 2% increase in sensitivity) was observed for pancreatic NETs. However, since NETs are not associated with MPD expansion, the effect on sensitivity in NET patients was not expected. Finally, although the addition of MPD diameter and lesion diameter significantly improved IPMN sensitivity (an increase of 6%), the small case count (18 out of 290 cases) prevented any definitive conclusions.
[0155] We also utilized a segmentation network to design a logistic regression model to predict MPD expansion based on the MPD diameter of the pancreatic head, body, and tail. The reported AUC was 0.97 (95% CI: 0.96, 0.98). We emphasize that while other deep learning methods allow for MPD segmentation, none have utilized MPD segmentation to provide an early warning of its potential expansion, a crucial finding for radiologists assessing the pancreas.
[0156] The algorithm's segmentation performance on pancreas, lesions, and MPD was also evaluated. For pancreas, no competing deep learning methods evaluated Dice scores on the same dataset as this study. However, three deep learning methods were found to report an average Dice score of 0.87 on their test set. The segmentation network proposed in this paper achieves similar results on public data, but achieves a higher average Dice score on the overall test set. For lesion segmentation, an average Dice score of 0.63 was obtained on public data. This is a 9% improvement over the existing Dice score of 0.54 reported using nnUNet. Finally, we were unable to find other deep learning methods that achieved MPD Dice scores. However, given the small size of MPD, the model appears to exhibit satisfactory performance compared to the Dice scores achieved for lesions, further confirmed by an NSD score of 0.71 on the test set.
Claims
1. A method implemented by a computer device for detecting at least one lesion of a patient's pancreas in at least one medical image, the method comprising an inference phase, the inference phase comprising the following steps: (a) Predict at least one first probability map, the first probability map representing the probability that each pixel or voxel of the image belongs to part of a pancreatic lesion, (b) Predict at least one second probability map, the second probability map representing the probability that each pixel or voxel of the image belongs to part of the pancreas. (c) Predict at least one third probability map, the third probability map representing the probability that each pixel or voxel of the image belongs to a part of the main pancreatic duct. (d) Based on the probability map, segment the pancreas, main pancreatic duct, and at least one lesion in the image. (e) Identify the head, body, and tail of the segmented pancreas. (f) Based on the first probability map, determine a first feature representing the risk of lesion. (g) Determine the second feature representing the maximum size of the lesion. (h) Determine the third feature representing the maximum size of the main pancreatic duct in the head of the pancreas. (i) Determine the fourth feature representing the maximum size of the main pancreatic duct in the body of the pancreas. (j) Determine the fifth feature representing the maximum size of the main pancreatic duct in the tail of the pancreas. A first score is determined based on at least the first feature, the second feature, and at least one of the third, fourth, and fifth features, wherein the first score represents the probability that the patient has pancreatic lesions.
2. The method according to the preceding claims, characterized in that, The medical image is a CT scan image.
3. The method according to the preceding claims, characterized in that, The image is a CT scan image from the portal venous phase.
4. The method according to any one of the preceding claims, characterized in that, Step (f) includes the following sub-steps: - Eliminate all pixels or voxels located outside the segmented pancreas. - For each remaining connected component, the lesion risk is calculated by averaging the probabilities of all its pixels or voxels, and connected components with a lesion risk below a threshold are eliminated.
5. The method according to any one of the preceding claims, characterized in that, The second feature is the maximum diameter of the corresponding segmented pancreatic lesion.
6. The method according to the preceding claim, characterized in that, The image is a 3D image, and the maximum diameter of the corresponding segmented pancreatic lesion is determined by measuring the 2D Ferrette diameter of the segmented lesion layer by layer along the axial view.
7. The method according to any one of the preceding claims, characterized in that, The first, second, and third probability maps are obtained through at least one convolutional neural network or model, such as nnUnet.
8. The method according to any one of the preceding claims, characterized in that, The first, second, and third probability maps are obtained through multiple or a set of convolutional neural networks or models, each model being able to predict the first, second, and third probability maps, and the output of the models being used to generate each probability map by pixel-wise or voxel-wise averaging.
9. The method according to the preceding claim, characterized in that, The model set was trained using a k-fold cross-validation process.
10. The method according to any one of the preceding claims, characterized in that, The image is a 3D image, and step (e) includes the following sub-steps: - The morphological framework of the pancreas is extracted layer by layer along the axial view. - Obtain the 3D skeleton of the pancreas and convert the skeleton into a graph. - Given an image point located at the far right anterior lower part of the abdomen, identify the point in the image furthest from it, which is considered the endpoint of the pancreatic tail. - Identify the point furthest from the tail in the image as the head of the pancreas. - Calculate the shortest path between the head and tail to obtain a centerline that passes through the pancreas and connects its head and tail endpoints. - Divide the centerline into head, body, and tail. - For each voxel in the pancreas segment, find its nearest point on the center line and assign it to the corresponding head, body or tail.
11. The method according to the preceding claim, characterized in that, For each voxel of the segmented main pancreatic duct, calculate its nearest point on the center line, and assign the segmented pancreas head, body, or tail to the voxel based on the position of its nearest point on the center line. The diameter of the main pancreatic duct is measured separately in the head, body, and tail.
12. The method according to any one of the preceding claims, characterized in that, The first score was determined using a first logistic regression model.
13. The method according to any one of the preceding claims, comprising the step of determining a second score based on at least the third, fourth and fifth features, the second score representing the probability of dilation of the main pancreatic duct.
14. The method according to the preceding claim, characterized in that, The second score was determined using a second logistic regression model.
15. The method according to any one of the preceding claims, characterized in that, The first or second score is determined based on at least one other feature from the following list: - First-order statistics of pancreatic diameter, lesion diameter, main pancreatic duct diameter, and / or common bile duct diameter. - Location of the lesion, for example, the lesion is located in the head, body, and / or tail. - At least one radiomics feature of the pancreas, lesions, main pancreatic duct, and / or common bile duct. - 3D shape features of the pancreas, lesions, main pancreatic duct, and / or common bile duct. - 2D shape characteristics of the pancreas, main pancreatic duct, lesions and / or common bile duct and / or first-order statistics of these characteristics.