Computer-implemented methods for localizing and correcting segmentation failures in automatically obtained predicted segmentations of medical images

The method addresses segmentation failures in medical imaging by calculating voxel-wise correctness and regional influence scores to prioritize corrections on critical areas, improving accuracy and efficiency in medical image analysis.

WO2026158986A1PCT designated stage Publication Date: 2026-07-30UNIVERSITY OF BERN
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
UNIVERSITY OF BERN
Filing Date
2026-01-16
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing automatic segmentation methods in medical imaging, such as MRI and CT scans, suffer from segmentation failures due to variations in image data and anatomy, requiring time-consuming manual corrections by clinicians without considering the impact on clinical metrics.

Method used

A computer-implemented method that localizes and corrects segmentation failures by calculating voxel-wise correctness estimates, generating feature-derived regions, and assigning risk scores based on regional influence values to prioritize corrections on critical areas affecting clinical metrics.

Benefits of technology

This method efficiently targets and corrects segmentation errors with the highest impact on clinical metrics, improving accuracy and reducing manual effort by focusing on critical regions, thus enhancing the efficiency and effectiveness of the correction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2026051050_30072026_PF_FP_ABST
    Figure EP2026051050_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The invention pertains to computer-implemented methods for localizing and correcting segmentation failures in automatically obtained predicted segmentations of medical images, such as MRI or CT scans.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Computer-implemented methods for localizing and correcting segmentation failures in automatically obtained predicted segmentations of medical images

[0002] Technical Field

[0003] The present invention relates to the field of medical image analysis and processing. In particular, the present application describes computer-implemented methods and system for localizing and correcting segmentation failures in medical images, in particular three-dimensional medical images, such as those obtained from magnetic resonance imaging (MRI) or computed tomography (CT) scans.

[0004] Background of the invention

[0005] Medical imaging is a crucial tool in modern healthcare, allowing for non-invasive examination of the body to detect, diagnose, and monitor diseases.

[0006] Techniques such as Magnetic Resonance Imaging (MRI) and Computed Tomography (CT) scans generate three-dimensional images of the body's internal structures. These images are often complex and detailed, requiring significant expertise to interpret accurately.

[0007] Segmentation, the process of identifying and delineating specific structures within these images, is a critical step in many medical imaging applications. It allows for the isolation of regions of interest, enabling precise measurements, visualization, and further analysis. Manual segmentation, however, is a time-consuming and labour-intensive process, subject to inter and intra-observer variability.

[0008] Automatic segmentation of medical images using convolutional neural networks (CNN) has recently become more prevalent. In this approach, CNNs trained with ground truth manual segmentation data are subsequently used to infer the segmentation of images not used in the training of the model, eliminating the need for time-consuming manual segmentation of data. Using network architectures designed for medical image segmentation, automatic multi-structure segmentation with average accuracies equivalent to manual segmentation has been enabled for a range ofanatomical regions. Once a model has been trained, segmentation can be performed in a matter of seconds, facilitating the incorporation of more quantitative three-dimensional (3D) analysis into the clinical routine.

[0009] The robustness of these deep learning-based segmentation algorithms, however, remains inferior compared to expert manual segmentation in practice, and segmentation networks might produce inadequate segmentations when, for instance, the input image data is dissimilar from the data used to train the model. Moreover, the use of scanners from various manufacturers and different imaging protocols employed by different clinics can result in variations in image intensities and resolution that might diverge from the input data used for training. The inherent heterogeneity of the imaged anatomy can also result in segmentation failure. Due to this subpar robustness, there are concerns relating to the integration of these techniques into clinical routine, as clinicians are currently required to review each segmentation, slice by slice, to ensure adequate segmentation quality. If segmentation failures are identified, clinicians have to manually segment the affected regions. This time-consuming process, which requires specialised skills and software, detracts from the benefits gained from these automatic segmentation networks.

[0010] While there is ongoing research dedicated to improving the robustness of these automatic segmentation models, another research trend focuses on the detection of segmentation failures by analysing network-generated uncertainties. Recently, methods have been proposed to detect the segmentation failures of a given structure. These methods, which have shown promising results, exploit network-generated uncertainties to predict the segmentation quality of each target structure. This structure-wise failure detection is useful for an expert to flag affected segmentations. However, these methods do not localise the error within the segmentation and are therefore unable to assess the clinical significance of the error. Segmentation failure can be regional and might not be pertinent to the quantification of relevant metrics. Furthermore, a flagged segmentation still requires the time-consuming inspection and manual correction of each slice containing the structure.

[0011] A semi-automatic correction technique in which a user can indicate segmentation failures with mouse clicks, and scribbles has also been proposed. Fromthis input, geodesic distance maps are generated and used as input for a CNN together with the original image and segmentation. This technique can achieve three-dimensional corrections from user interaction on two-dimensional MRI slices. It was demonstrated that this method could expedite the correction process and achieve better results than unaided segmentation models. Others used network-generated uncertainties instead of user interactions to obtain an error map. They then applied an auxiliary model for the re-segmentation of regions likely containing a segmentation error. They showed that this approach could increase the segmentation accuracy. However, network-generated uncertainties can only predict which regions are likely to contain a segmentation error. Some user interaction would, therefore, likely be beneficial in confirming which regions are actually erroneously segmented.

[0012] Segmentation correction techniques previously reported have been designed for use on the entire structure and do not consider the impact of the correction on the final application. In clinical applications, segmentations are often used for anatomy analyses that are limited to certain regions of the anatomical structure or for treatment planning that depends on accurate anatomy segmentations in specific anatomical regions only. Thus, the importance of segmentation accuracy is not always uniform over the whole structure. Undertaking corrections of the entire structure may thus be an unnecessary effort that has potential to retract from attention on more critical regions.

[0013] Therefore, there is a need for methods that allow for localizing and correcting segmentation failures in automatically obtained predicted segmentations of medical images, in particular three-dimensional medical images, such as a magnetic resonance imaging or computed tomography image, in a more efficient and precise manner.

[0014] Summary of the invention

[0015] The object of the present invention is in particular to propose novel computer-implemented methods, with which the above-described drawbacks of the known methods are completely overcome or at least greatly diminished.The present invention addresses the above-mentioned disadvantages of the known techniques by providing, in a first aspect, a computer-implemented method for localizing segmentation failures in an automatically obtained predicted segmentation of a medical image, in particular a three-dimensional medical image, in view of verification and / or correction, that incorporates information pertaining to the impact of the regional segmentation accuracy on the final metric for which the segmentation is used, to optimise and guide correction.

[0016] The method involves obtaining voxel-wise correctness estimates for the predicted segmentation, which are measures of the likelihood that each voxel belongs to its assigned predicted label. This is advantageously obtained through the entropy calculation of the calibrated confidence of the neural network used for the automatic segmentation or using a machine learning model trained to estimate the correctness of the predicted labels.

[0017] Feature-derived regions are then generated from the medical image and the voxel-wise predicted segmentation. These regions are advantageously formed by clustering the voxels based on their similarity in image contrast and proximity and optionally on their predicted labels. Optionally, the feature-derived regions are generated from multiple images of the same anatomical structure.

[0018] The voxel-wise correctness estimates are then aggregated within each feature-derived region to obtain region correctness estimates. This involves computing a measure of the overall correctness of the predicted segmentation within each feature-derived region. Advantageously, the correctness estimates associated with the voxels contained in each respective feature-derived region were averaged to obtain the region correctness estimates.

[0019] A metric, advantageously a clinical metric, value is derived from the predicted segmentation, such as the volume, surface area or thickness of an anatomic element. Alternatively, the metric value is derived from multiple predicted segmentations, advantageously of the same anatomic element or from different anatomic elements. Other relevant metrics are for instance principal axis, angle between anatomical structures and closest distance between structures. The methodthen determines a regional influence value for each feature-derived region on the metric value, or a voxel-wise influence value for each voxel in the predicted segmentation on the metric value. The influence value quantifies the influence of the presence of the respective voxels in the predicted segmentation on the metric value.

[0020] Finally, the method combines the region correctness estimates and regional influence values to obtain a risk score for each feature-derived region, or the region correctness estimates and voxel-wise influence values to obtain a risk score for each voxel. This risk score represents the likelihood of a segmentation failure occurring within the respective feature-derived region or at the individual voxel, weighted by its influence value on the metric value.

[0021] Important to understand is that some errors are inherent to the process of automatic segmentation and are typically within the range of precision of manual segmentation and would thus typically not be corrected by a user. These errors are often attributed to noise at the surface of the segmented structure, which refers to random variations in image intensity that can interfere with the segmentation process.

[0022] However, there are certain errors that are considered more significant and are referred throughout this application to as segmentation failures. These failures represent instances where the automatic segmentation process has deviated significantly from the expected or desired result. These failures can have a substantial impact on the derived I metrics, such as the volume, surface area, or thickness of an anatomical structure.

[0023] The method described here is designed to localize these segmentation failures, thereby enabling targeted verification and correction efforts. This is done by assigning a risk score to each feature-derived region or voxel, which represents the likelihood of a segmentation failure occurring within the respective feature-derived region or at the individual voxel, weighted by its influence value on the metric value.

[0024] The key advantage of the present method is its ability to allow focusing correction efforts on regions that have a significant impact on the final I metric, thereby improving the efficiency and effectiveness of the correction process. This approach isparticularly beneficial in clinical applications where the importance of segmentation accuracy varies across the anatomical structure. By prioritizing regions based on their influence on the metric, the method ensures that critical regions receive the necessary attention, while less critical regions are not unnecessarily corrected.

[0025] In one preferred embodiment of the first aspect of the present invention, the feature-derived regions are super-voxels. Super-voxels are clusters of voxels that share similar characteristics, such as intensity or texture, and are spatially contiguous in the medical image. The use of super-voxels as feature-derived regions can provide a more efficient and robust representation of the medical image, as compared to individual voxels. This is because super-voxels can capture more complex and meaningful structures in the image, while also reducing the dimensionality of the data. The use of super-voxels can therefore improve the accuracy and efficiency of the subsequent steps in the method, such as the aggregation of voxel-wise correctness estimates and the determination of influence values.

[0026] To calculate super voxels, the simple linear iterative clustering (SLIC) algorithm [1], [2] can advantageously be used. The algorithm first partitions the input image volume into a grid of cuboids of a specified dimension. The centres of these clusters are then iteratively shifted by applying -means until they optimally partition the volume into voxel groups of similar contrast and spatial proximity.

[0027] In another preferred embodiment of the first aspect of the present invention, the borders of the feature-derived regions are aligned with the voxels of the predicted segmentation. This alignment allows for a semi-automatic correction approach in which the segmentation label of a whole feature-derived region is corrected by a user to an uniform value. To achieve this, the segmentation image can advantageously be used as an additional input channel for the SLIC algorithm. To ensure the alignment, the segmentation channel is scaled so that the maximum scalar value is the same as in the medical image.

[0028] In another preferred embodiment of the first aspect of the present invention, a visual representation of the risk score for each feature-derived region, respectively voxel, is generated. This visual representation can be a graphical display, such as acolor-coded map, where feature-derived regions, respectively voxels, with different risk scores are highlighted in a visually distinct manner. This provides a clear and intuitive way for users to identify the areas of the predicted segmentation that are most likely to contain segmentation failures and that have the greatest influence on the derived metrics.

[0029] Advantageously, the method can include a step of displaying the visual representation to a user for review and correction. This can involve presenting the visual representation on a computer screen or other display device, and providing user interface controls that allow the user to interact with the visual representation. For example, the user may be able to select a region, respectively voxel, on the visual representation, view its risk score and other associated data, and manually correct, advantageously with computer mouse clicks or scribbles, the predicted segmentation label if necessary. This interactive review and correction process can help to improve the accuracy of the predicted segmentation, by allowing the user to focus their efforts on the areas of the segmentation that are most likely to contain errors and that have the greatest impact on the metrics.

[0030] In a further preferred embodiment of the first aspect of the present invention, the risk scores for each feature-derived region, respectively voxel, are ranked in descending order. This ranking can be based on the magnitude of the risk scores, with regions, respectively voxels, having higher risk scores being ranked higher. This ranking provides a prioritized list of regions, respectively voxels, for review and correction, allowing the user to focus their efforts on the areas of the predicted segmentation that are most likely to contain segmentation failures and that have the greatest influence on the derived metrics.

[0031] The visual representation is advantageously displayed to the user in a sequential manner reflecting this ranking. For example, the regions, respectively voxels, can be highlighted on the visual representation in order of their ranking, starting with the region, respectively voxel, with the highest risk score. This provides a clear and intuitive way for the user to navigate through the predicted segmentation, and can help to streamline the review and correction process. By presenting the regions, respectively voxels, in order of their risk scores, the method ensures that the user'sattention is directed to the most critical areas of the segmentation first, thereby improving the efficiency and effectiveness of the review and correction process.

[0032] In another preferred embodiment of the first aspect of the present invention, steps e. to g. are performed for each segmentation label and each metric. This means that the determination of influence values, the calculation of risk scores, and the combination of these values are performed separately for each predicted segmentation label and each derived metric. By performing the steps e. to g. for each segmentation label and each metric, the method provides a more accurate and meaningful assessment of the quality of the predicted segmentation, and more effectively guides the user in the review and correction process. In this case, the ranking and / or the visual representation of the risk scores can be made for one or more particular labels or all labels together, advantageously in descending order.

[0033] In yet another preferred embodiment of the first aspect of the present invention, steps e. to g. are performed only for the feature-derived regions, respectively voxels, in contact with the segmented structure in the predicted segmentation. This means that the determination of influence values, the calculation of risk scores, and the combination of these values are performed only for those regions, respectively voxels, that are adjacent to or within the boundaries of the segmented anatomical structures in the predicted segmentation. This helps to focus the analysis on the areas of the segmentation that are most relevant to the derived metrics, and reduces the computational complexity of the method by excluding regions or voxels that are unlikely to contain segmentation failures or to have a significant influence on the metrics. By performing these steps only for the feature-derived regions, respectively voxels, in contact with the segmented structure, the method provides a more efficient and targeted analysis of the predicted segmentation, and more effectively guides the user in the review and correction process.

[0034] In another preferred embodiment of the first aspect of the present invention, the process begins with the acquisition of a medical image This image can be obtained using any suitable medical imaging modality, such as magnetic resonance imaging (MRI), computed tomography (CT), positron emission tomography (PET), or

[0035] ultrasound.Following the acquisition of the medical image, the method can include a step of automatic segmentation of the medical image. This segmentation can be performed using a trained machine-learning model, especially a trained neuronal network model. The machine-learning model can be trained on a large dataset of medical images, with each image in the dataset being associated with a ground truth or reference segmentation. The model learns to predict the segmentation of a new image based on the patterns and features it has learned from the training dataset.

[0036] The automatic segmentation step results in a predicted segmentation of the medical image, where each voxel in the image is assigned a predicted segmentation label (I) indicating the anatomical structure to which it is predicted to belong. This predicted segmentation serves as the basis for the subsequent steps in the method, including the generation of feature-derived regions, the computation of correctness estimates, and the determination of influence values and risk scores.

[0037] The use of a machine-learning model for automatic segmentation provides a number of advantages. For example, it significantly reduces the time and effort required to segment a medical image, as compared to manual segmentation methods. It can also provide a more consistent and objective segmentation, as it is not subject to the variability and bias that can occur with human observers. Furthermore, the accuracy of the automatic segmentation can be continually improved as the machinelearning model is trained on more and more data. The model can for instance be a 3D ll-Net [3] trained and evaluated using the self-configuring nnll-Net framework from Isensee et al. [4],

[0038] In another preferred embodiment of the first aspect of the present invention, the voxel-wise correctness estimates are derived using an ensemble method. This ensemble method involves generating a set of / V automatic predicted segmentations of the medical image using different segmentation models or configurations. Each of these predicted segmentations provides a different interpretation of the medical image, based on the specific characteristics and assumptions of the segmentation model or configuration used to generate it. The ensemble method then involves determining, for each voxel, the probability that it belongs to its assigned predicted label based on its mean predicted label value among the / V segmentations. This involves calculating theaverage of the predicted labels assigned to each voxel across the / V segmentations, and using this average value as an estimate of the probability that the voxel belongs to its assigned predicted label.

[0039] The use of an ensemble method for deriving voxel-wise correctness estimates provides a number of advantages. For example, it can improve the robustness and accuracy of the correctness estimates, as it takes into account the variability and uncertainty in the predicted segmentations. It can also provide a more comprehensive and nuanced assessment of the predicted segmentation, as it considers multiple different interpretations of the medical image. Furthermore, the ensemble method can help to mitigate the limitations and biases of individual segmentation models or configurations, by combining their predictions into a single, consensus estimate.

[0040] In another preferred embodiment of the first aspect of the present invention, step g., which involves combining the region correctness estimates and regional influence values to obtain a risk score for each feature-derived region, respectively voxel, is performed using a second trained model. This model is advantageously a logistic regression model. The parameters of the logistic regression model are obtained through training on a dataset comprising pairs of automatic predicted risk scores maps and corresponding manually obtained ground truth segmentation failures maps, i.e. the difference between the manual ground truth segmentation and the predicted segmentation, wherein this difference is advantageously filtered to only take into account errors significantly deviating from the expected or desired result (i.e. segmentation failures). The ground truth segmentation failures serve as a reference standard for the correct segmentation failure map of the medical image, and are typically obtained by comparing an automatically obtained predicted segmentation with a corresponding ground truth segmentation performed by a human expert. The training process involves minimizing a loss function value, advantageously the binary crossentropy loss value, which quantifies the difference between the predicted risk scores and the ground truth segmentation failures, with a lower loss value indicating a better match between the predictions and the ground truth. By training the logistic regression model on this data, the model learns to predict the likelihood of a segmentation failure occurring within each feature-derived region, respectively voxel, based on the regioncorrectness estimates and regional influence values. The output of the logistic regression model is then used as the risk score for each feature-derived region, respectively voxel.

[0041] Advantageously, the logistic model p(s,l), where s is the feature-derived region, respectively voxel, and / the segmentation label, takes the form:

[0042]

[0043] where

[0044]

[0045] is the bias, linear factor

[0046]

[0047] defines if the feature-derived region, respectively voxel, is part of a different segmented structure,

[0048]

[0049] is the linear factor for the region correctness estimates, p3to pkthe linear factors for each metric, m^s, I) to mk(s, I) are the metric influence values based on feature-derived regions, voxels respectively.

[0050] The use of a logistic regression model for computing the risk scores provides a number of advantages. For example, it can provide a more accurate and robust estimation of the risk scores, as it takes into account the complex relationships and interactions between the region correctness estimates and regional influence values. It can also provide a more interpretable and transparent prediction, as the logistic regression model provides an explicit formula for the risk score based on the input features. Furthermore, the logistic regression model can be easily trained and updated as new data becomes available, allowing the method to continually improve its performance over time.

[0051] The invention also encompasses, in a second aspect, a computer-implemented method for correcting segmentation failures in an automatically obtained predicted segmentation of a medical image, in particular a three-dimensional medical image, such as a magnetic resonance imaging (MRI) or computed tomography (CT) image.

[0052] The method begins by performing the previously described steps to obtain a risk score for each feature-derived region, respectively voxel. The method then involves iteratively correcting the labels of the voxels of the feature-derived regions, respectively of the voxels, according to their risk score in descending order. At eachiteration step, the label of the voxels of a target feature-derived region, respectively of the voxels, is manually corrected by a user. This correction process can involve changing the predicted segmentation label of the voxels to match the correct or desired label. After each adjustment, the relative difference between at least one metric value, advantageously a clinical metric value, derived from the updated predicted segmentation and the same metric value derived from the predicted segmentation of the previous iteration is advantageously computed. This relative difference provides a measure of the impact of the correction on the derived metric, and can be used to assess the effectiveness of the correction. The iterative adjustment process continues until the relative difference in the at least one metric value converges below a predefined threshold, reaches a maximum number of iterations, or upon user intervention. This ensures that the correction process continues until a satisfactory result is achieved, as determined by the convergence of the metric, the completion of a specified number of iterations, or the judgement of the user.

[0053] If steps e. to g. of the method for localizing segmentation failures, according to the first aspect of the present invention, are performed only for the feature-derived regions, respectively voxels, in contact with the segmented structure in the predicted segmentation, it is important to note that during the correction process and based on the changing segmentation results the region(s) in contact with the segmented structure can be reevaluated.

[0054] This method for correcting segmentation failures provides a systematic and efficient approach for improving the accuracy of the predicted segmentation. By prioritizing the correction efforts based on the risk scores, the method allows the user to focus on the areas of the segmentation that are most likely to contain errors and that have the greatest impact on the derived metrics. Furthermore, by iteratively adjusting the segmentation and monitoring the impact on the metrics, the method provides a dynamic and responsive correction process that can adapt to the specific characteristics and requirements of each individual case.

[0055] In one preferred embodiment of the second aspect of the preset invention, at each iteration, the current predicted segmentation with the visually highlighted voxels of the feature-derived region, respectively voxels, to be corrected are displayed to theuser. This visual representation can be a graphical display, such as a color-coded map, where the voxels to be corrected are highlighted in a visually distinct manner. This can provide a clear and intuitive way for users to identify the areas of the predicted segmentation that need correction. The visual representation can be displayed on a computer screen or other display device, and can be updated at each iteration, or at a different rate, to reflect the current state of the predicted segmentation and the voxels to be corrected. This interactive display can help to guide the user through the correction process, by clearly indicating the areas of the segmentation that need attention and showing the impact of the corrections on the predicted segmentation. By displaying the current predicted segmentation with the visually highlighted voxels to be corrected, the method can provide a more user-friendly and efficient correction process. This helps to reduce the time and effort required for the correction process, and can improve the accuracy and consistency of the corrected segmentation.

[0056] In another preferred embodiment of the second aspect of the present invention, during the correction step, the label of each voxel within the target feature-derived region is corrected separately by the user. This means that the user has the ability to individually adjust the predicted segmentation label of each voxel within the target region. This can provide a high level of control and precision in the correction process, allowing the user to make detailed adjustments to the segmentation based on their expert knowledge and judgement. This individual voxel correction can be facilitated by user interface controls that allow the user to select individual voxels and change their predicted segmentation labels. These controls can provide a simple and intuitive way for the user to interact with the predicted segmentation and make the necessary corrections. Advantageously, the user is able to correct the label by mouse clicks or scribbles, enabling larger areas to be corrected quickly.

[0057] In another preferred embodiment of the second aspect of the present invention, during the correction step, the label of the voxels within the target feature-derived region is corrected to a uniform value defined by the user. This means that all voxels within the target region are assigned the same corrected segmentation label, as specified by the user. This can provide a quick and efficient way to correct large regions of the predicted segmentation, particularly in cases where the entire region is known to belong to a single anatomical structure. This uniform correction can befacilitated by user interface controls that allow the user to select a target region and specify a corrected segmentation label. These controls can provide a simple and intuitive way for the user to make large-scale corrections to the predicted segmentation. Advantageously, the user is able to correct the label by mouse clicks or scribbles, enabling larger areas to be corrected quickly.

[0058] Important to note is that in case steps e. to g. of the localizing method according to the fist aspect of the present invention are performed for different segmentation labels and different metrics, the correction can be made for one or more particular labels or all labels together.

[0059] The invention also encompasses, in a further aspect, a data processing system configured to carry out the steps of the described method for localizing segmentation failures in an automatically obtained predicted segmentation of a medical image, and / or the steps of the described method for correcting segmentation failures in an automatically obtained predicted segmentation of a medical image.

[0060] The data processing system can include a processor, memory, and a user interface. The processor is configured to execute instructions stored in the memory to perform the steps of the method, in particular obtaining a voxel-wise predicted segmentation of the medical image, obtaining voxel-wise correctness estimates for the predicted segmentation, generating feature-derived regions, aggregating the voxelwise correctness estimates within each feature-derived region to obtain region correctness estimates, deriving at least one metric value from the predicted segmentation, determining a regional influence value for each feature-derived region, respectively voxel, on the at least one metric value, and combining the region correctness estimates and regional influence values to obtain a risk score for each feature-derived region, respectively voxel.

[0061] The user interface can provide a means for a user to interact with the data processing system, for example, to manually correct the labels of the voxels of the feature-derived regions according to their risk score in descending order, and to review and correct the predicted segmentation based on the visual representation of the riskscores. Advantageously, user interface allows the user to correct the label by mouse clicks or scribbles, enabling larger areas to be corrected quickly.

[0062] The data processing system can be a standalone system, or it can be part of a larger system, such as a medical imaging system or a hospital information system. The data processing system can be implemented using any suitable hardware and software technologies, and can be configured to communicate with other systems and devices, for example, to receive medical images for processing, and to send the corrected segmentation and derived metrics for further analysis or reporting.

[0063] The invention also includes, in a further aspect, a computer program product comprising instructions which, when executed by a computer, cause the computer to carry out the steps of the described method for localizing segmentation failures in an automatically obtained predicted segmentation of a medical image, and / or the steps of the described method for correcting segmentation failures in an automatically obtained predicted segmentation of a medical image. The computer program product can be stored on a computer-readable medium, such as a hard drive, solid-state drive, optical disc, flash memory, or any other non-transitory storage medium. The computer program product can also be distributed over a network, for example, for download over the internet.

[0064] The instructions in the computer program product can be written in any suitable programming language, and can be compiled or interpreted for execution by the computer. The instructions include in particular performing the steps of the method, such as obtaining a voxel-wise predicted segmentation of the medical image, obtaining voxel-wise correctness estimates for the predicted segmentation, generating feature-derived regions, aggregating the voxel-wise correctness estimates within each feature-derived region to obtain region correctness estimates, deriving at least one metric value from the predicted segmentation, determining a regional influence value for each feature-derived region on the at least one metric value, and combining the region correctness estimates and regional influence values to obtain a risk score for each feature-derived region.The computer program product can also include instructions for interfacing with a user, for example, to display a visual representation of the risk scores, and to receive user input for the manual correction of the predicted segmentation. The computer program product can be used in a variety of settings, such as in a hospital, clinic, or research institution, to improve the accuracy and efficiency of medical image segmentation and the derived metrics.

[0065] The invention also includes, in a further aspect, a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the steps of the described method for localizing segmentation failures in an automatically obtained predicted segmentation of medical image, and / or the steps of the described method for correcting segmentation failures in an automatically obtained predicted segmentation of a medical image. The computer-readable storage medium can be any type of non-transitory medium that can store data that can be read by a computer system. Examples include, but are not limited to, hard drives, solid-state drives, optical discs, flash memory devices, and memory cards.

[0066] The instructions stored on the computer-readable storage medium can be written in any suitable programming language, and can be compiled or interpreted for execution by the computer. The instructions include performing the steps of the method, such as obtaining a voxel-wise predicted segmentation of the medical image, obtaining voxel-wise correctness estimates for the predicted segmentation, generating feature-derived regions, aggregating the voxel-wise correctness estimates within each feature-derived region to obtain region correctness estimates, deriving at least one metric value from the predicted segmentation, determining a regional influence value for each feature-derived region on the at least one I metric value, and combining the region correctness estimates and regional influence values to obtain a risk score for each feature-derived region.

[0067] The computer-readable storage medium can also include instructions for interfacing with a user, for example, to display a visual representation of the risk scores, and to receive user input for the manual correction of the predicted segmentation. The computer-readable storage medium can be used in a variety ofsettings, such as in a hospital, clinic, or research institution, to improve the accuracy and efficiency of medical image segmentation and the derived metrics.

[0068] Brief description of the drawings

[0069] - Figure 1 displays a flowchart for a preferred embodiment of a method for localizing segmentation failures;

[0070] - Figure 2 displays a flowchart for a preferred embodiment of a method for correcting segmentation failures;

[0071] - Figure 3 illustrates a schema of a preferred embodiment of a data processing system;

[0072] - Figure 4 illustrates in more detail the segmentation process of a medical image and the outputs of this process;

[0073] - Figure 5 illustrates in more detail the combining process of a predicted segmentation with a medical image for generating super-voxels;

[0074] - Figure 6 illustrates combining the super-voxels and the voxel-wise uncertainties for generating regions uncertainties; and

[0075] - Figure 7 illustrates the process of deriving metric values from the predicted segmentation.

[0076] Detailed description of preferred embodiments

[0077] Figure 1 illustrates a flowchart depicting the workflow of a computer-implemented method 100 for localizing segmentation failures in an automatically obtained predicted segmentation of a medical image, in particular a three-dimensional (3D) medical image, such as a magnetic resonance imaging (MRI) or computed tomography (CT) image, in view of verification and / or correction, according to a preferred embodiment of the first aspect of the present invention.

[0078] The workflow begins with the obtaining 101 of a voxel-wise predicted segmentation of the input medical image. In this step, each voxel within the image is assigned a predicted segmentation label, indicating its presumed anatomical structure or region.Subsequently, voxel-wise correctness estimates 102, such as voxel-wise uncertainty values, are obtained for the predicted segmentation. These estimates represent the likelihood that each voxel belongs to its assigned predicted label, quantifying the confidence in the initial prediction.

[0079] The method then proceeds with the generation 103 of feature-derived regions from the medical image and optionally from the voxel-wise predicted segmentation. Feature-derived regions are formed by clustering voxels based on their similarity in image contrast, proximity, and optionally predicted labels.

[0080] Next, region correctness estimates 104 are aggregated within each feature-derived region by combining the voxel-wise correctness estimates obtained earlier.

[0081] In parallel or subsequently, metric(s) is derived 105 from the predicted segmentation. These relevant metrics can be for instance the volume, surface area, or thickness of anatomic elements.

[0082] For each feature-derived region, a regional influence value 106 is determined on the at least one metric value. This influence value quantifies how significantly the presence of voxels within their respective feature-derived regions impacts the computed metrics.

[0083] In a final step, the method culminates with the combining 107 of region correctness estimates and regional influence values to obtain a risk score for each feature-derived region. These risk scores represent the likelihood of segmentation failures occurring within respective regions, weighted by their influence on metric values. This risk-based prioritization enables focused validation and correction efforts, ensuring that areas with higher failures probabilities and greater relevance are addressed first.

[0084] It is important to note that alternatively to regional influence values, voxelwise influence values can be determined for each voxel in the predicted segmentation on the at least one metric value. Voxel-wise influence values provide a localizedassessment of the impact that individual voxels' segmentation failures may have on relevant measurements. If voxel-wise influence values are computed, the step 107 comprises combining region correctness estimates and voxel-wise influence values to obtain a risk score for each voxel.

[0085] While the workflow presented in Figure 1 follows a specific sequential order for clarity and understanding, it is essential to note that the steps described can be performed in alternative orders.

[0086] Figure 2 illustrates a flowchart outlining the computer-implemented method for correcting segmentation failures in an automatically obtained predicted segmentation of a medical image according to a preferred embodiment of the present invention. This workflow builds upon the risk scores generated by the localization method described in Figure 1 and facilitates targeted, iterative correction of segmentation errors based on their relevance.

[0087] The process begins with the obtaining 201 of risk scores for each feature-derived region and / or voxel within the predicted segmentation. These risk scores are generated using the method of the present invention for quantifying the likelihood of segmentation failures within respective regions or voxels while considering their influence on relevant metrics.

[0088] The correction workflow is, as illustrated by the dashed arrow in this Figure, iterative, with each cycle focusing on a subset of voxels with the highest risk scores. At the start of each iteration, the method selects 202 the target feature-derived region(s), respectively voxel(s), with the highest risk score(s), prioritizing areas where segmentation failures are most likely to have significant consequences.

[0089] Once the target regions / voxels are identified, a manual correction 203 process is initiated. In this step, a user, such as an experienced medical professional, adjusts the predicted labels of the voxels within the selected feature-derived region(s) or individual voxel(s). This hands-on approach allows for expert knowledge and visual context to be integrated into the segmentation refinement process.After each label adjustment, the method advantageously computes 204 the relative difference between at least one I metric value derived from the updated predicted segmentation and the same metric obtained from the previous iteration's segmentation. This step quantifies the impact of the user's corrections on relevant metric, such as volume, surface area, or thickness.

[0090] The iterative adjustment process continues 205 until a convergence condition is met. The method checks if:

[0091] The relative difference in the at least one I metric value has converged below a predefined threshold, indicating that further adjustments are unlikely to significantly improve the segmentation's relevance;

[0092] - A maximum number of iterations has been reached, ensuring that the correction process does not become excessively time-consuming or exhaustive; or

[0093] User intervention occurs, allowing experts to halt the automatic refinement if necessary.

[0094] Once the convergence condition is satisfied, the method proceeds to the next iteration 206, selecting a new set of feature-derived regions, respectively voxels, with the highest remaining risk scores. This prioritized, iterative approach ensures that segmentation errors with the most significant implications are addressed first, optimizing user effort and improving overall segmentation quality.

[0095] Figure 3 presents a schematic diagram of a data processing system 300, configured to carry out the methods for localizing and / or correcting segmentation failures in an automatically obtained predicted segmentation of medical images. The system 300 comprises a processor 301, responsible for executing instructions stored in memory 302 to perform the steps of the described methods. The processor can be a central processing unit (CPU), a more specialized graphics processing unit (GPU), a neural processing unit (NPU), or a quantum processing unit (QPU) depending on the specific implementation and performance requirements. The memory 302 serves as a storage unit, housing executable instructions for method execution, intermediate data results, and other relevant information. It can consist of various types of volatile andnon-volatile memory technologies, such as random-access memory (RAM), flash memory, or hard disk drives. The system features a user-friendly user interface 303 that facilitates interaction between users and the data processing system 300. The user interface comprises a display 304, typically a monitor or touchscreen device, for visualizing medical images, predicted segmentations, risk scores, and other pertinent information throughout the processing workflow. Input devices 305, such as mouse, keyboard, or touchpad, enabling users to manually correct segmentations based on risk scores and visual representations, by means of mouse clicks or scribbles, as well as interact with other system functionalities.

[0096] To facilitate data exchange between the data processing system 300 and external systems or devices, such as medical imaging systems or hospital information systems, a communication module 306 is integrated. This interface enables the receipt of medical images for processing and the transmission of corrected segmentations and derived metrics for further analysis or reporting.

[0097] Moreover, the system includes a dedicated medical image database 307, storing medical images obtained from various sources, such as magnetic resonance imaging (MRI) or computed tomography (CT) scans. This database serves as a repository for input data processed by the system.

[0098] The data processing system 300 can be implemented using any suitable hardware and software technologies and can function either as an independent standalone system or integrated into larger systems, depending on specific use cases and deployment requirements.

[0099] Figure 4 illustrates in more detail the steps 101 and 102 of the method 100. A medical image A (here as example a hip MRI image) is acquired and in B automatically segmented by means of a trained CNN. The CNN outputs a predicted segmentation C and a voxel-wise uncertainty map D. As displayed in Figure 5, the medical image A and optionally the predicted segmentation C are used, in step 103 of the method 100 (see above), to generate super-voxels as shown in E. Following the generation of the super-voxels, the latter and the uncertainty map D are combined to generate super-voxel uncertainties F (see Figure 6).In order to compute for each super-voxel an influence value on a I metric and a corresponding risk score, metric values are derived from the predicted segmentation as illustrated in Figure 7. In the particular example of the hip, the chondrolabral complex (cartilage and labrum) is typically the first structure in the hip joint to sustain damage when hip biomechanics are compromised. The specific damage patterns on the cartilage and labrum, as well as 3D cartilage and labrum volume, offer critical insights for the surgeon, helping to determine whether the underlying issue is hip instability or a bony conflict between the femur and acetabulum. Based on the predicted segmentation, the metrics values such as the cartilage volume, surface area, and thickness are thus derived. Additionally, the labrum volume and surface area (joint-facing area which has contact with the femoral head) are calculated. The average cross-section and average length of the labrum are advantageously additionally calculated by dividing the volume, and the surface area by the angle (a) that covers the labrum on the clockface times the radius (r) of the femoral head estimated by a fitted sphere (Figure 7 right).

[0100] References

[0101]

[0001] R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. Susstrunk, “SLIC Superpixels Compared to State-of-the-Art Superpixel Methods,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 34, no. 11, pp. 2274-2282, Nov. 2012, doi:

[0102] 10.1109 / TPAMI.2012.120.

[0103] [2] D. Stutz, A. Hermans, and B. Leibe, “Superpixels: An evaluation of the state-of-the-art,” Comput. Vis. Image Underst., vol. 166, pp. 1-27, Jan. 2018, doi:

[0104] 10.1016 / j.cviu.2017.03.007.

[0105] [3] O. Ronneberger, P. Fischer, and T. Brox, “ll-Net: Convolutional Networks for Biomedical Image Segmentation,” in Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015, vol. 9351, N. Navab, J. Hornegger, W. M. Wells, and A. F. Frangi, Eds., in Lecture Notes in Computer Science, vol. 9351. , Cham:

[0106] Springer International Publishing, 2015, pp. 234-241. doi: 10.1007 / 978-3-319-24574-4_28.

[0107] [4] F. Isensee, P. F. Jaeger, S. A. A. Kohl, J. Petersen, and K. H. Maier-Hein, “nnll-Net: a self-configuring method for deep learning-based biomedical imagesegmentation,” Nat. Methods, vol. 18, no. 2, Art. no. 2, Feb. 2021, doi: 10.1038 / S41592-020-01008-z.

Claims

24Claims1. A computer-implemented method for localizing segmentation failures in an automatically obtained predicted segmentation of a) medical image, in particular a three-dimensional (3D) medical image, such as a magnetic resonance imaging (MRI) or computed tomography (CT) image, in view of verification and / or correction, the method comprising:a. obtaining a voxel-wise predicted segmentation of the medical image, wherein each voxel is assigned a predicted segmentation label;b. obtaining voxel-wise correctness estimates for the predicted segmentation, wherein said correctness estimate corresponds to the likelihood that the voxel belongs to its assigned predicted label;c. generating feature-derived regions from the medical image and optionally from the voxel-wise predicted segmentation, wherein the feature-derived regions are advantageously formed by clustering the voxels based on their similarity in image contrast and proximity and optionally on their predicted labels;d. aggregating the voxel-wise correctness estimates within each feature- derived region to obtain region correctness estimates;e. deriving at least one metric value, advantageously a clinical metric value, from the predicted segmentation, wherein the metric is for instance the volume, surface area or thickness of an anatomic element;f. determining at least one of:i. a regional influence value for each feature-derived region on the at least one metric value; orii. a voxel-wise influence value for each voxel in the predicted segmentation on the at least one metric value, wherein said influence value quantifies the influence of the presence of the respective voxels in the predicted segmentation on the metric value; and g. combining at least one of:i. the region correctness estimates and regional influence values to obtain for each feature-derived region a risk score; orii. the region correctness estimates and voxel-wise influence values to obtain for each voxel a risk score,wherein the risk score represents the likelihood of a segmentation failure occurring within the respective feature-derived region respectively at the individual voxel, weighted by its influence value on the metric value, allowing for prioritized validation and correction efforts.

2. The computer-implemented method according to claim 1 , wherein the feature- derived regions are super-voxels.

3. The computer-implemented method according to any one of the claims 1 or 2, wherein the borders of the feature-derived regions are aligned with the voxels of the predicted segmentation.

4. The computer-implemented method according to any one of the preceding claims, further comprising a step of generating a visual representation of the risk score for each feature-derived region, respectively voxel, wherein regions, respectively voxels, with different risk scores are highlighted in a visually distinct manner and a step of displaying the visual representation to a user for review and correction.

5. The computer-implemented method according to claim 4, wherein the risk scores for each feature-derived regions, respectively voxels, are ranked in descending order and the visual representation is displayed to the user in a sequential manner reflecting the ranking.

6. The computer-implemented method according to any one of the preceding claims, wherein steps e. to g. are performed for each segmentation label and each metric.

7. The computer-implemented method according to any one of the preceding claims, wherein steps e. to g. are performed only for the feature-derived regions, respectively voxels, in contact with the segmented structure in the predicted segmentation.

8. The computer-implemented method according to any one of the preceding claims, comprising a step of acquiring a medical image and a step of automatic segmentation of the medical image, advantageously by means of a first trained machine-learning model, especially a trained neuronal network model.

9. The computer-implemented method according to any one of the precedent claims, wherein the voxel-wise correctness estimates provided in step b. are derived using an ensemble method comprising:a. generating a set of / V automatic predicted segmentations of the medical image using different segmentation models or configurations; and b. determining for each voxel, the probability that it belongs to its assigned predicted label based on its mean predicted label value among the / V segmentations.

10. The computer-implemented method according to any one of the preceding claims, wherein step g. is performed by using a second trained model, advantageously a logistic regression model, to compute a risk score for each feature-derived region, respectively voxel, wherein:a. the model’s parameters are obtained through training on a dataset comprising pairs of automatic predicted risk scores and corresponding manually obtained ground truth segmentation failures; andb. the training process involves minimizing a loss function value, advantageously the binary cross-entropy loss, between the predicted risk scores and the ground truth segmentation failures.

11. A computer-implemented method for correcting segmentation failures in an automatically obtained predicted segmentation of a medical image, in particular a three-dimensional (3D) medical image, such as a magnetic resonance imaging (MRI) or computed tomography (CT) image, the method comprising:27a. Performing the computer-implemented method according to any one of claims 1 to 10 to obtain for each feature-derived region, respectively voxel, a risk score;b. Iteratively correcting the labels of the voxels of the feature-derived regions, respectively the voxels, according to their risk score in descending order, wherein:i. at each iteration step, the label of the voxels of a target feature- derived region, respectively the voxels, is manually corrected by a user; andii. after each adjustment, the relative difference between at least one metric value, advantageously a clinical metric value, derived from the updated predicted segmentation and the same metric value derived from the predicted segmentation of the previous iteration is advantageously computed; andiii. the iterative adjustment process continues until the relative difference in the at least one metric value converges below a predefined threshold, reaches a maximum number of iterations or upon user intervention.

12. The method of claim 11, wherein at each iteration the current predicted segmentation with the visually highlighted voxels of the feature-derived region, respectively voxels, to be corrected are displayed to the user.

13. The method of any one of the claims 11 or 12, wherein in step a. the label of each voxel within the target feature-derived region is corrected separately by the user.

14. The method of any one of the claims 11 or 12, wherein in step a. the label of the voxels within the target feature-derived region is corrected to a uniform value defined by the user.

15. A data processing system comprising means for carrying out the steps of the method of claim 1 to 10 and / or the steps of the method of claim 11 to 14.2816. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the steps of the method of claim 1 to 10 and / or the steps of the method of claim 11 to 14.

17. A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the steps of the method of claim 1 to 10 and / or the steps of the method of claim 11 to 14.