Identify the boundaries of lesions within image data

By using uncertainty data in medical image analysis to identify and correct lesion boundaries, the problem of inaccurate lesion boundary identification in the prior art is solved, achieving higher identification accuracy and better diagnostic support.

CN113661518BActive Publication Date: 2025-05-20KONINKLIJKE PHILIPS NV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080026839.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-04
Filing Date
2020-04-03
Publication Date
2025-05-20
Estimated Expiration
2040-04-03

AI Technical Summary

Technical Problem

When analyzing medical images using machine learning algorithms, the prior art faces the problem of inaccurate training data caused by inaccurate lesion annotation, especially at the lesion boundary, which leads to inaccurate lesion boundary identification.

Method used

By generating uncertain data, these data are used to identify and correct lesion boundaries in medical images, thereby improving the identification accuracy of boundaries.

Benefits of technology

Improves the identification accuracy of lesion boundaries and provides more accurate information to help clinicians diagnose and evaluate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113661518B_ABST
    Figure CN113661518B_ABST
Patent Text Reader

Abstract

The present invention provides a method, computer program and processing system for identifying the boundary of a lesion in image data. The image data is processed using a machine learning algorithm to generate probability data and uncertainty data. The probability data provides a probability data point indicating the probability that the image data point is part of a lesion for each image data point of the image data. The uncertainty data provides an uncertainty data point indicating the uncertainty of the probability data point for each probability data point. The uncertainty data is processed to identify or correct the boundary of the lesion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automated processing of medical images, and in particular, to the identification of the boundaries of lesions within the image data of medical images. Background Art

[0002] There is increasing interest in using machine learning algorithms to analyze medical images. Specifically, machine learning algorithms have the potential to perform highly accurate medical image segmentation tasks, such as organ segmentation, nodule segmentation, etc. A particularly interesting area is the use of machine learning algorithms to identify lesions within medical images, such as CT (Computed Tomography) scans, ultrasound images, or X-ray images.

[0003] A challenge in implementing machine learning algorithms is that the training data used to train the machine learning algorithms may be not entirely accurate, for example, due to incorrect or incomplete annotations by trained clinicians. Inaccurate annotation of lesions is particularly common at the boundaries of lesions, where clinicians may be unable or unwilling to accurately identify the exact boundaries of the lesions. This problem is exacerbated by the need to resample these often inaccurate / ambiguous lesion annotations to a uniform resolution to allow training of the machine learning algorithms. Upsampling of lower-resolution images and blurry / inaccurate annotations is required, resulting in less accurate training data for the machine learning algorithms.

[0004] In addition, different boundary identification machine learning algorithms processing the same image may identify different boundaries of the lesion, for example, due to differences in the training data or loss functions used to train the machine learning annotation algorithms. Thus, there may be a disagreement among different annotation algorithms on the exact location where the boundary of the lesion lies.

[0005] The present inventors have recognized that there is a need, therefore, to improve the identification of the boundaries of lesions within medical images. Improved identification of the boundaries of lesions provides more accurate information to clinicians, enabling them to more accurately diagnose or evaluate a subject. Summary of the Invention

[0006] The present invention is defined by the claims.

[0007] According to an example aspect of the present invention, there is provided a method for identifying one or more boundaries of a lesion within N-dimensional medical image data of a region of a subject.

[0008] The method includes: receiving N-dimensional medical image data, which includes image data points; using a machine learning algorithm to process the N-dimensional medical image data to generate: N-dimensional probability data, including corresponding probability data points indicating the probability that an image data point is part of a lesion for each image data point; and N-dimensional uncertainty data, including corresponding uncertainty data points indicating the uncertainty of the indicated probability for each probability data point; and identifying one or more boundaries of a lesion in the medical image data using at least the uncertainty data.

[0009] The present invention proposes to identify the boundaries of lesions within (medical) image data based on uncertainty data. The uncertainty data includes uncertainty data points. Each uncertainty data point indicates the uncertainty of a predicted probability associated with a particular image data point. The predicted probability is a prediction as to whether the particular image data point is part of a lesion (e.g., within a range between 0 and 1), i.e., the "lesion probability".

[0010] The probability data can be used to predict the location, size, and shape of a lesion. For example, if clusters of image data points are each associated with the probability data, then the cluster of image data points can be regarded as a predicted lesion, where the probability data indicates that each image data point in the cluster has a certain probability (higher than a predetermined value) that it is part of a lesion. This method is well known to those skilled in the art.

[0011] The present invention relies on the understanding that the probability data may not be completely accurate when predicting whether a given image data point is part of a lesion. In fact, the present invention identifies that the accuracy level decreases near the boundary or edge of the lesion.

[0012] The inventors have recognized that the uncertainty data can be used to more accurately identify the location of the boundary of a lesion relative to the image data. Specifically, the uncertainty of the lesion probability can be used to distinguish or determine the location where the boundary or edge of the lesion is located. Thus, the problem of an inaccurately trained machine learning algorithm can be overcome by using the uncertainty data to modify, identify, or correct the boundary of the lesion. These inaccuracies are particularly prevalent at the boundary of the lesion.

[0013] The proposed invention is thus able to more accurately identify the boundaries of lesions. This directly helps the user perform the task of diagnosing or assessing the condition of a subject, as the number, location, and size of the lesions within the subject will be more accurately identified. Specifically, knowing the extent and location of the lesions is essential for correctly analyzing the condition of the subject.

[0014] The step of identifying one or more boundaries of a lesion may include: identifying one or more potential lesions in medical image data based on probability data; and processing each potential lesion using at least uncertainty data to identify one or more boundaries of the lesion in the medical image data.

[0015] Thus, the boundaries of a lesion can be identified by identifying potential lesions in medical image data and at least using uncertainty data to correct or modify the boundaries of the identified potential lesions.

[0016] The step of identifying one or more potential lesions in medical image data may include identifying groups of image data points associated with probability data points that indicate a probability that the image data points are part of a lesion that exceeds a predetermined probability, each group of image data points thus forming a potential lesion. Thus, a cluster of image data points forms a potential lesion if each image data point in the cluster is associated with a high enough (above a predetermined threshold) probability of being part of a lesion.

[0017] Embodiments may include the step of selectively excluding (i.e., excluding from further consideration) potential lesions based on uncertainty data. For example, if the image data points forming a potential lesion are associated with an average uncertainty above a predetermined value, then the potential lesion can be excluded (i.e., removed from one or more potential lesions). As another example, if more than a certain percentage of the image data points forming a potential lesion are associated with uncertainty data points having an uncertainty above a predetermined value, then the potential lesion can be excluded.

[0018] Thus, uncertain potential lesions can be excluded. In this way, potential lesions can be excluded based on uncertainty data. This improves the accuracy of correctly identifying potential lesions, specifically the true positive rate, by excluding those lesions with uncertain predictions.

[0019] In some embodiments, the step of identifying one or more boundaries of a lesion includes using a region growing algorithm to process the image data, probability data, and uncertainty data to identify one or more boundaries of the lesion.

[0020] Region growing methods can be used to appropriately expand or shrink a predicted lesion by using a set of rules to decide whether neighboring image data points should be added to the predicted lesion.

[0021] The region growing rules can be image features based on the image data, such as the magnitude of the image data points (e.g., Hounsfield units or HU), where the decision rule is a simple threshold of the magnitude of the image data points.

[0022] Specifically, the inclusion threshold can be based on the uncertainty level for the same image data point. Specifically, the inclusion threshold can be generated by a machine learning algorithm (e.g., based on the calculated uncertainty for the image data point). As another example, the uncertainty level can weight the inclusion value.

[0023] Similarly, these rules can be used to reduce the size of the lesion, e.g., by examining the voxels included in the perimeter of the lesion and applying a decision rule that combines the image data point amplitude and the uncertainty threshold.

[0024] Thus, there are embodiments where identifying one or more boundaries of a lesion includes applying a region growing algorithm and / or a region shrinking algorithm to each potential lesion.

[0025] The region growing algorithm can include iteratively performing the following steps: identifying perimeter image data points that form the perimeter of the potential lesion; identifying neighboring image data points that are image data points outside the potential lesion and adjacent to any of the perimeter image data points; and for each neighboring image data point, adding the neighboring image data point to the potential lesion in response to the amplitude of the neighboring image data point being greater than a first amplitude threshold for the neighboring image data point, where the first amplitude threshold is based on the uncertainty data point associated with the neighboring image data point, and where the first region growing algorithm ends in response to no new neighboring image data points being identified.

[0026] The region shrinking algorithm can include iteratively performing the following steps: identifying perimeter image data points that form the perimeter of the potential lesion; for each perimeter image data point, removing the perimeter image data point from the potential lesion in response to the amplitude of the perimeter image data point being less than a second amplitude threshold for the perimeter image data point, where the second amplitude threshold is based on the uncertainty data point associated with the perimeter image data point, and where the second region growing algorithm ends in response to no new perimeter image data points being identified.

[0027] The step of identifying one or more boundaries of a lesion can include: using probability data to identify one or more predicted boundary portions of each potential lesion, each predicted boundary portion being a predicted location of a part of the boundary of the potential lesion; using uncertainty data to identify the uncertainty of each predicted boundary portion and / or the uncertainty of each potential lesion; selecting one or more of the predicted boundary portions based on the identified uncertainty of each boundary portion and / or the identified uncertainty of each potential lesion; presenting the selected predicted boundary portions to the user; after presenting the selected predicted boundary portions to the user, receiving user input indicating one or more boundaries of the lesion; and identifying the boundary portions based on the received user input.

[0028] Optionally, the step of selecting one or more predicted boundary portions includes selecting those boundary portions associated with an uncertainty below a first predetermined uncertainty value and / or those boundary portions associated with a lesion having an uncertainty below a second predetermined uncertainty value.

[0029] The method may further include generating corresponding one or more graphical annotations, each graphical annotation indicating the location of each of the one or more identified boundaries.

[0030] Preferably, the machine learning algorithm is a Bayesian deep learning segmentation algorithm. The Bayesian deep learning segmentation algorithm can be easily adapted to generate uncertainty data and thus provides a suitable machine learning algorithm for processing image data. In a preferred embodiment, the image data is computer tomography (CT) image data. In some embodiments, N equals 3. In such an embodiment, the term "N-dimensional" may be replaced by the term "3-dimensional" (3D).

[0031] There is also provided a computer program including code components for implementing any of the described methods when the program is run on a processing system.

[0032] According to an example of an aspect of the present invention, there is also provided a processing system for identifying one or more boundaries of a lesion in N-dimensional medical image data of a region of a subject. The processing system is adapted to: receive N-dimensional medical image data including image data points; use a machine learning algorithm to process the N-dimensional medical image data to generate: N-dimensional probability data including corresponding probability data points indicating the probability that an image data point is part of a lesion for each image data point; and N-dimensional uncertainty data including corresponding uncertainty data points indicating the uncertainty of the indicated probability for each probability data point; and identify one or more boundaries of the lesion in the medical image data using at least the uncertainty data.

[0033] The processing system may be adapted to: identify one or more boundaries of the lesion by identifying one or more potential lesions in the medical image data based on the probability data; and process each potential lesion using at least the uncertainty data to identify one or more boundaries of the lesion in the medical image data.

[0034] In some embodiments, the processing system may be adapted to process each potential lesion by using a region growing algorithm to process the image data, the probability data, and the uncertainty data to identify one or more boundaries of the lesion.

[0035] In other embodiments, the processing system may be adapted to process each potential lesion by: using probability data to identify one or more predicted boundary portions, each predicted boundary portion being a predicted location of a part of the boundary of the lesion; using uncertainty data to identify the uncertainty of each predicted boundary portion; selecting one or more of the predicted boundary portions based on the identified uncertainty of each boundary portion; presenting the selected predicted boundary portions to the user; after presenting the selected predicted boundary portions to the user, receiving user input indicating one or more boundaries of the lesion; and identifying the boundary portions based on the received user input.

[0036] The processing system may be adapted to select one or more of the predicted boundary portions by selecting those boundary portions associated with an uncertainty below a predetermined uncertainty value.

[0037] In some embodiments, the processing system is further adapted to generate corresponding one or more graphical annotations, each graphical annotation indicating the location of each one or more identified boundaries.

[0038] The machine learning algorithm may be a Bayesian deep learning segmentation algorithm.

[0039] These and other aspects of the invention will become apparent from the embodiments described hereinafter, and will be elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] For a better understanding of the present invention, and to more clearly show how the present invention may be implemented, reference will now be made, by way of example only, to the accompanying drawings,

[0041] Figure 1 showing a 2D slice of a computed tomography of a blood vessel with a lesion;

[0042] Figure 2 Conceptually showing an embodiment of the present invention;

[0043] Figure 3 being a flowchart showing a method according to an embodiment;

[0044] Figure 4 showing a region growing algorithm used in an embodiment;

[0045] Figure 5 showing a region shrinking algorithm used in an embodiment; and

[0046] Figure 6 showing a system according to an embodiment. DETAILED DESCRIPTION

[0047] The present invention will be described with reference to the drawings.

[0048] It should be understood that the detailed description and specific examples are merely intended to illustrate exemplary embodiments of the indicating device, system, and method, and are not intended to limit the scope of the present invention. These and other features, aspects, and advantages of the device, system, and method of the present invention will become better understood from the following description, the appended claims, and the accompanying drawings. It should be understood that these drawings are merely schematic and not drawn to scale. It should also be understood that throughout the drawings, the same reference numerals are used to indicate the same or similar parts.

[0049] The present invention provides a method, computer program, and processing system for identifying the boundary of a lesion within image data. The image data is processed using a machine learning algorithm to generate probability data and uncertainty data. The probability data provides a probability data point indicating the probability that the image data point is part of a lesion for each image data point of the image data. The uncertainty data provides an uncertainty data point indicating the uncertainty of the probability data point for each probability data point. The uncertainty data is processed to identify or correct the boundary of the lesion.

[0050] The inventive concept is based on the recognition that the boundary of a lesion can be more accurately identified by processing or further analyzing the uncertainty data, which indicates the uncertainty of the prediction that a particular image data point is or is not part of a lesion.

[0051] Embodiments can be used to identify lesions within medical images of a subject. By way of example, embodiments can be used to identify calcified lesions within the cardiac vasculature or tumors within one or more lungs.

[0052] As previously described, embodiments relate to the concept of identifying one or more boundaries of a lesion within N-dimensional medical image data.

[0053] Figure 1 An example of a medical image 1 for understanding the context of the present invention is illustrated. Here, the medical image is a 2D slice of a CT scan of a blood vessel 5 (i.e., the aorta with a lesion 10). The lesion 10 has a boundary (or boundaries) that can be identified by employing the inventive concept described herein.

[0054] The boundary separates the part of the image data that is considered to be part of the lesion ("lesion part") from the part of the image data that is not considered to be part of the lesion ("non-lesion part"). Specifically, the boundary can surround or define the perimeter of the lesion part.

[0055] For example, the boundary can be presented in the form of a graphical annotation that can, for example, overlay the medical image 1 on a display, thereby drawing attention to the location of the boundary with respect to the image data. In another example, the boundary can be presented in the form of coordinate information that identifies the location of the boundary of the lesion portion.

[0056] Although only 2D slices of a CT scan are illustrated here, embodiments of the present invention can be extended to any N-dimensional medical image data, such as data for 3D CT scans, ultrasound images, X-ray images, etc. Medical image data is any data suitable for (re)constructing a visual representation of a subject (internal or external) for clinical analysis.

[0057] Figure 2 The basic concept of the present invention is illustrated for the purpose of improving the understanding of the inventive concept.

[0058] Embodiments relate to identifying one or more boundaries of a lesion within N-dimensional medical image data 21. The medical image data 21 is formed by image data points (such as pixels or voxels) that form an overall N-dimensional image. As illustrated, the medical image data can correspond to a two-dimensional (2D) image or a three-dimensional (3D) image.

[0059] The medical image data 21 is processed by a machine learning algorithm 22 to generate probability data 23 and uncertainty data 24. The probability data 23 and the uncertainty data 24 have the same number of dimensions N as the medical image data 21.

[0060] The probability data 23 includes probability data points. Each probability data point corresponds to a corresponding image data point and indicates the probability that the corresponding image data point forms part of a lesion, such as a tumor or a calcium deposit. Thus, there are the same number of probability data points as there are image data points.

[0061] By way of example, each probability data point can be a continuous value between 0 and 1 that directly indicates the probability that the corresponding image data point is part of the lesion, or can be a binary value (such as 0 or 1) that indicates whether the corresponding image data point is predicted to form part of the lesion (e.g., associated with a probability above a threshold probability value). Other suitable formats or values for the probability data points will be apparent to the person skilled in the art.

[0062] The uncertainty data 24 includes uncertainty data points. Each uncertainty data point corresponds to a corresponding probability data point (and thus to a corresponding image data point). Each uncertainty data point indicates the level of uncertainty of the indicated probability, such as a measure or value representing the (in)certainty that the probability data point is correct. Thus, there are the same number of uncertainty data points as there are probability data points and thus the same number as there are image data points.

[0063] By way of example, each uncertainty data point can be a continuous value between 0 and 1, which continuous value indicates a measure of the relative uncertainty that the probability data point is correct. In another example, each uncertainty data point can indicate a margin of error or standard deviation of the probability data point (e.g., in the case where the probability data point is a continuous value between 0 and 1). Other suitable formats or values for the uncertainty data points will be apparent to the person skilled in the art.

[0064] Then, the uncertainty data is used to identify or correct the boundaries of lesions within the medical image. The identified boundaries can be provided as boundary data 25, e.g., identifying the location of the boundaries within the image data 21. By way of example, the boundary data can identify the coordinates of one or more boundaries within the image data, or can indicate which image data points are part of the boundary.

[0065] It will generally be understood that the probability data itself can be used to define the predicted location of the boundary (or boundaries) of one or more lesions. This is because the probability data can be used to predict whether a given image data point is part of a lesion. Thus, the perimeter of the lesion can be identified using the probability data.

[0066] Merely by way of example, if the probability data point is a continuous value between 0 and 1, if a given image data point is associated with a probability data point greater than a predetermined value, then that image data point can be considered part of the lesion, and otherwise not. By another example, if the probability data point is a binary (e.g., 0 or 1) prediction, if a given image data point is associated with the probability data point 1, then that image data point can be considered part of the lesion, and otherwise not.

[0067] In this way, the probability data can be used to identify potential lesions within the image data.

[0068] The present invention relies on the understanding that the probability data may not be completely accurate when predicting whether a given image data point is part of a lesion. In fact, the present invention recognizes that the level of accuracy decreases near the boundary or edge of the lesion.

[0069] The inventors have recognized that the uncertainty data can be used to more accurately identify the location of the boundaries of the lesions relative to the image data.

[0070] In a particular embodiment, the uncertainty data can be used to perform further processing to more accurately identify the location of the boundary than the probability data alone. Identifying the location of the boundary can include correcting the location of the predicted boundary derived from the probability data.

[0071] Figure 3FIG. 30 illustrates a method 30 according to an embodiment of the present invention. The method identifies one or more boundaries of a lesion within N-dimensional medical image data of a region of a subject.

[0072] Method 30 includes step 31 of receiving N-dimensional medical image data including image data points.

[0073] Method 30 further includes step 32 of processing the N-dimensional medical image data using a machine learning algorithm to generate N-dimensional probability data and N-dimensional uncertainty data. The N-dimensional probability data includes corresponding probability data points indicating the probability that an image data point is part of a lesion for each image data point. The N-dimensional uncertainty data includes corresponding uncertainty data points indicating the uncertainty of the indicated probability for each probability data point.

[0074] Method 30 further includes step 33 of identifying one or more boundaries of a lesion in the medical image data using at least the uncertainty data.

[0075] Step 33 can be implemented in a number of different possible ways.

[0076] Specifically, step 33 can include: identifying one or more potential lesions in the medical image data based on the probability data; and processing each potential lesion using at least the uncertainty data to identify one or more boundaries of a lesion in the medical image data.

[0077] Thus, an initially predicted lesion can be identified based on the probability data. The boundaries of the lesion can be modified based at least on the uncertainty data.

[0078] The boundaries of the lesion can be modified in an automated manner using a region growing algorithm and / or a region shrinking algorithm.

[0079] The region growing method can be used to expand the size of an initial lesion by using a set of rules to determine whether neighboring image data points (i.e., image data points neighboring the predicted lesion) should be added to the lesion.

[0080] The region growing method can be based on the value or magnitude (e.g., Hounsfield units) of neighboring image data points, where the decision rule (i.e., including or not including an image data point in the lesion) is a simple inclusion threshold magnitude for the image data point. Thus, if the magnitude of a neighboring image data point is higher than the inclusion threshold, then the neighboring image data point is included in the potential lesion.

[0081] One way to implement this can be to vary the inclusion threshold based on the uncertainty level for the same image data point or the average uncertainty of the associated predicted lesion. The threshold can be a parameter tuned during hyperparameter optimization of the segmentation model.

[0082] When it is predicted that the lesion cannot grow further, i.e., when there are no new pixels adjacent to the (updated) predicted lesion, the region growing algorithm can be stopped.

[0083] Figure 4 An example method 33A of step 33 in which the region growing algorithm 42 is executed is illustrated.

[0084] Method 33A includes a step 41 of identifying one or more potential lesions in the medical image data based on probability data.

[0085] Step 41 may include identifying a set of image data points associated with a probability data point that indicates a probability (that the image data point forms part of a lesion) higher than a predetermined probability value. Those skilled in the art will know other methods of using probability data to identify potential lesions.

[0086] Step 41 may include selectively excluding (i.e., excluding from further consideration) a subset of potential lesions (not shown) based on uncertainty data. For example, if the image data points forming a potential lesion are associated with an average uncertainty higher than a predetermined value, then the potential lesion can be excluded (i.e., removed from one or more potential lesions). As another example, if more than a certain percentage of the image data points forming a potential lesion are associated with uncertainty data points having an uncertainty higher than a predetermined value, then the potential lesion can be excluded. This improves the accuracy of correctly identifying potential lesions, specifically the true positive rate, by excluding those lesions with uncertain predictions.

[0087] Then, the region growing algorithm is applied to each identified potential lesion. This can be performed in parallel or sequentially. In the illustrated embodiment, each region growing algorithm is sequentially applied to each identified potential lesion. Thus, there is a step 49A of determining whether there are any unprocessed potential lesions (i.e., lesions to which the region growing algorithm has not yet been applied). If each potential lesion has been processed, then method 33A ends at step 49C. If not every potential lesion has been processed, then a step 49D of selecting an unprocessed lesion is performed.

[0088] An embodiment of the region growing algorithm 42 is provided below.

[0089] Region growing algorithm 42 includes a step 42A of identifying perimeter image data points that form the perimeter of a potential lesion. Perimeter image data points are the outermost image data points of a potential lesion, i.e., image data points that border, adjoin, are adjacent to, or are in close proximity to image data points that do not form part of the potential lesion.

[0090] The region growing algorithm 42 further includes a step 42B of identifying neighboring image data points, which are image data points outside the potential lesion and in close proximity to any perimeter image data points. Thus, each neighboring image data point in the image data points is adjacent to the image data points of the potential lesion.

[0091] In response to no new neighboring image data points being identified in step 42B, the region growing algorithm 42 ends. Thus, there may be a step 42C of determining whether any new neighboring image data points have been identified.

[0092] If, for example, in step 42C, new neighboring image data points have been identified, then the region growing algorithm 42 moves to step 42D of adding any neighboring image data points to the potential lesion having an amplitude greater than a first amplitude threshold for that neighboring image data point. This effectively increases the size of the potential lesion.

[0093] The first amplitude threshold is based on the uncertainty data points associated with the neighboring image data points. Thus, the first amplitude threshold may be different for each neighboring image data point.

[0094] By way of example, the value of the uncertainty data point may act as a variable for an equation for calculating the first amplitude threshold. In one example, the uncertainty data point is used to modify a predetermined value for the amplitude threshold to generate the first amplitude threshold. As another example, the first amplitude threshold for each neighboring image data point may be a parameter generated during the generation of the uncertainty data (i.e., by a machine learning algorithm). Thus, each image data point may be associated with a corresponding first amplitude threshold (generated by a machine learning algorithm) for use by the region growing algorithm.

[0095] The region growing algorithm is iteratively repeated until no new neighboring image data points are identified, i.e., until the lesion stops growing.

[0096] Return reference Figure 3 , step 33 may include performing a region shrinking or region reducing algorithm. This may be performed instead of or in addition to the region growing algorithm.

[0097] The region shrinking algorithm is similar to the region growing algorithm, except that it is used to reduce the size of the lesion. For example, this may be performed by examining the image data points forming the perimeter of the lesion and applying a decision rule based on the uncertainty data to combine the image data points and the amplitude including a threshold. When it is predicted that the lesion cannot be further shrunk, for example, when all the image data points forming the perimeter of the lesion have an amplitude higher than their corresponding thresholds, the region shrinking algorithm may be stopped.

[0098] Figure 5Illustrated is an example method 33B of step 33 in which the region shrinking algorithm is executed.

[0099] Method 33B includes step 51 of identifying one or more potential lesions in the medical image data based on probability data in substantially the same manner as step 41 of method 33A. Step 51 may be executed simultaneously with step 41 (if executed).

[0100] The region shrinking algorithm 52 is applied to each identified potential lesion. This can be executed in parallel or sequentially. In the illustrated embodiment, each region shrinking algorithm is sequentially applied to each identified potential lesion. Thus, there is step 59A of determining whether there are any unprocessed potential lesions (i.e., lesions to which the region shrinking algorithm has not yet been applied). If each potential lesion has been processed, then method 33B ends at step 59C. If not every potential lesion has been processed, then step 59D of selecting an unprocessed lesion is executed.

[0101] The region shrinking algorithm 52 includes step 52A of identifying perimeter image data points that form the perimeter of the potential lesion. As previously described, the perimeter image data points are the outermost image data points of the potential lesion, i.e., the image data points that abut, adjoin, are adjacent to, or are contiguous with image data points that do not form part of the potential lesion.

[0102] In response to no new perimeter image data points being identified in step 52A, the region growing algorithm 52 ends. Thus, there may be step 52B of determining whether any new perimeter image data points have been identified. If no new perimeter image data points are identified, then the region shrinking algorithm ends.

[0103] If new perimeter image data points have been identified, then step 52C is executed. Step 52C includes removing from the potential lesion any perimeter image data points having an amplitude less than a second amplitude threshold for the perimeter image data points. The second amplitude threshold is based on uncertainty data points associated with the perimeter image data points.

[0104] The second amplitude threshold is based on uncertainty data points associated with neighboring image data points. Thus, the second amplitude threshold may be different for each neighboring image data point.

[0105] By way of example, the value of an uncertainty data point can serve as a variable in an equation for calculating a second amplitude threshold. In one example, the uncertainty data point is used to modify a predetermined value for the amplitude threshold to generate the second amplitude threshold. As another example, the second amplitude threshold for each neighboring image data point can be a parameter generated during the generation of the uncertainty data (i.e., by a machine learning algorithm). Thus, each image data point can be associated with a corresponding second amplitude threshold (generated by a machine learning algorithm) for use by a region growing algorithm. The first amplitude threshold and the second amplitude threshold can be the same.

[0106] Return reference Figure 3 In another example, step 33 includes presenting a predicted boundary portion to a user and enabling the user to indicate a correction or modification to the predicted boundary portion. Thus, identifying the boundary can include receiving user input indicating one or more boundaries of the lesion (the input being received after presenting the predicted boundary portion to the user).

[0107] In some embodiments, the user may be able to exclude the presented boundary, thereby excluding the predicted lesion.

[0108] Specifically, step 33 can include using probability data to identify one or more predicted boundary portions of each potential lesion, each predicted boundary portion being a predicted location of a part of the boundary of the potential lesion; using uncertainty data to identify the uncertainty of each predicted boundary portion and / or the uncertainty of each potential lesion; selecting one or more predicted boundary portions from the predicted boundary portions based on the identified uncertainty of each boundary portion and / or the uncertainty of each potential lesion; presenting the selected predicted boundary portions to the user; after presenting the selected predicted boundary portions to the user, receiving user input indicating one or more boundaries of the lesion; and identifying the boundary portion based on the received user input.

[0109] In this way, the method can utilize the understanding and experience of a clinician to more accurately identify the boundary without burdening the clinician with identifying where the automatically predicted boundary needs to be corrected. Specifically, by drawing attention to the most (or least) uncertain predicted boundaries, the clinician does not need to correct boundaries that are considered to be identified with sufficient accuracy. This increases the ease and speed of identifying boundaries within the image data.

[0110] Two levels of uncertainty can be considered, one being the uncertainty of the entire potential lesion ("lesion uncertainty") and the other being the uncertainty of the boundary portion ("boundary uncertainty"). Each boundary portion can be associated with both lesion uncertainty and boundary uncertainty. For example, the lesion uncertainty can be the average uncertainty associated with each image data point of the potential lesion or the relative number (e.g., percentage) of image data points of the potential lesion having an uncertainty higher than a specific threshold. The selected boundary to be presented to the user may depend on one or both of these uncertainty measures.

[0111] For example, the boundaries selected for presentation can be those associated with a boundary uncertainty higher than a first predetermined uncertainty value and a lesion uncertainty lower than a second predetermined uncertainty value. The first predetermined uncertainty value and the second predetermined uncertainty value can be the same. The first predetermined uncertainty value and the second predetermined uncertainty value are preferably in the range of 40 to 70% of the maximum possible uncertainty value.

[0112] In another example, the boundaries selected for presentation can be those associated with a lesion uncertainty higher than a third predetermined uncertainty value. Thus, the boundaries of the uncertain lesions can be presented to the user for modification and / or exclusion.

[0113] Thus, in general, probability and uncertainty data can be used to extend the detected lesions in an automated manner by combining the uncertainty data with a region growing algorithm, or in a semi-automated manner where the least uncertain boundaries of the lesions are presented to a clinician or user to allow effective correction of any misidentification of the boundaries.

[0114] Return reference Figure 3 Method 30 may further include step 34 of generating a corresponding one or more graphical annotations, each graphical annotation indicating the location of each one or more identified boundaries.

[0115] Thus, for each identified boundary, a graphical annotation identifying the location and extent of the identified boundary can be generated. The graphical annotation can be designed to overlay the image data (e.g., on a display) to indicate the location and / or presence of the boundary of the lesion. The graphical annotation can be included in the boundary data.

[0116] In some embodiments, the method may further include displaying, e.g., on a display, the N-dimensional image data and the one or more graphical annotations. This display enables the boundaries of the lesions to be brought to the attention of the subject.

[0117] The proposed invention can more accurately identify the boundaries of lesions. This directly helps the user perform the task of diagnosing or evaluating the condition of a subject, as the number, location, and size of the lesions within the subject will be more accurately identified.

[0118] Therefore, the proposed invention provides reliable assistance to the user in performing technical tasks such as diagnosing or otherwise evaluating the condition of a subject / patient. Specifically, the boundaries of the lesions can be more accurately identified, thereby improving the information available to the clinician.

[0119] The proposed invention also recognizes that a direct link can be established between the actual location of the lesion boundary and the uncertainty of the predicted location of the lesion boundary.

[0120] It will be understood that a machine learning algorithm is any self-training algorithm that processes input data to produce output data. In the context of the present invention, the input data includes N-dimensional image data and the output data includes N-dimensional probability data and N-dimensional uncertainty data.

[0121] Suitable machine learning algorithms for the present invention will be apparent to those skilled in the art. Examples of suitable machine learning algorithms include decision tree algorithms, artificial neural networks, logistic regression, or support vector machines.

[0122] Preferably, the machine learning algorithm is a Bayesian deep learning segmentation algorithm such as a Bayesian neural network. The Bayesian-based machine learning model provides a relatively simple method for calculating the uncertainty of probability data points, as the uncertainty calculation is built into the Bayesian method.

[0123] The structure of an artificial neural network (or simply neural network) is inspired by the human brain. A neural network consists of layers, each layer including a plurality of neurons. Each neuron includes mathematical operations. Specifically, each neuron can include different weighted combinations of a single type of transformation (e.g., the same type of transformation, sigmoid, etc., but with different weights). During the process of processing the input data, the mathematical operations of each neuron are performed on the input data to produce a numerical output, and the output of each layer in the neural network is sequentially fed into the next layer. The last layer provides the output.

[0124] Methods for training machine learning algorithms are well known. Generally, such methods include obtaining a training data set that includes training input data entries and corresponding training output data entries. In the context of the present invention, the training input data entries correspond to example N-dimensional image data. The training output data entries correspond to the boundaries of the lesions within the N-dimensional image data (which effectively identifies the probability that a given image data point of the image data forms part of a lesion).

[0125] An initialized machine learning algorithm is applied to each input data entry to generate a predicted output data entry. The error or loss function between the predicted output data entry and the corresponding training output data entry is used to modify the machine learning algorithm. This process can be repeated until the error converges and the predicted output data entry is similar enough (e.g., ±1%) to the training output data entry. This scenario is commonly referred to as a supervised learning technique.

[0126] In a Bayesian neural network, each weight is associated with a probability distribution. During the training of the Bayesian neural network, the probability distribution of the weights can be updated to reflect the latest teachings of the training data. This effectively provides a model with its own probability distribution that can output prediction data with a probability distribution for each data point.

[0127] In such an example, the probabilistic data point can be the mean or central probability of the probability distribution, while the uncertainty data point can be the standard deviation of the probability distribution. This provides a simple, effective, and low-cost approach for generating probabilistic and uncertainty data.

[0128] In the context of the present invention, the machine learning algorithm can be trained using a loss function that takes into account the fact that predicted lesions should form connected components and that the uncertainty estimate should be concentrated on the boundaries of the lesions.

[0129] An example of a suitable loss function is shown in Equation (1) below.

[0130] Loss(P, G) = Dice(P, G) + γCCscore(P) (1)

[0131] In Equation 1, P represents the probabilistic data point prediction for the image data, G represents the ground truth, Dice is the dice similarity coefficient between the model prediction and the ground truth, and CCscore is a measure of whether the predicted lesion is compact or has missing pixels. CCscore can be measured by creating a fully compact lesion from the prediction (i.e., by adding the missing pixels from the predicted lesion) and then measuring the difference (e.g., the number of pixels) between this compact lesion and the predicted lesion. γ is a real number that provides a weight for this part of the loss function and is optimized during the training of the model.

[0132] In summary, it is clear that the machine learning algorithm is trained to directly compute probabilistic data from image data.

[0133] Uncertainty data can be calculated using the Monte Carlo dropout method. Specifically, a machine learning model can be trained with dropout layers in the model architecture. During the inference of probabilistic data, we turn dropout "on", which means that if the inference is run multiple times, then we will get different probabilistic data (i.e., different predictions). Then, the uncertainty can be measured as the variance of these predictions. Those skilled in the art will know other methods for calculating uncertainty data.

[0134] More information about Bayesian deep learning models and uncertainty information can be found in the paper “What uncertainties do we need in Bayesian deep learning for computer vision?” by Kendall, Alex, and Yarin Gal, Advances in Neural Information Processing Systems 2017. Those skilled in the art will consider referring to this document to identify more information about Bayesian deep learning models.

[0135] Those skilled in the art will be able to easily develop a processing system for implementing any of the methods described herein. Thus, each step of the flowchart can represent different actions performed by the processing system and can be executed by the corresponding modules of the processing system.

[0136] Figure 6 System 60 is illustrated, in which a processing system 61 according to an embodiment is implemented. The system includes a processing system 61, an image data generator 62, and a display unit 63.

[0137] The processing system 61 is adapted to receive N - dimensional medical image data, including image data points; use a machine learning algorithm to process the N - dimensional medical image data, thereby generating: N - dimensional probabilistic data, including corresponding probability data points indicating the probability that an image data point is part of a lesion for each image data point; and N - dimensional uncertainty data, including corresponding uncertainty data points indicating the uncertainty of the indicated probability for each probability data point; and identify one or more boundaries of the lesion in the medical image data using at least the uncertainty data.

[0138] The image data generator 62 is adapted to generate or provide N - dimensional medical image data. The image data generator can include, for example, a memory system storing medical image data (such as a stack of medical images) or an image data generating element, such as a CT scanner. The image data generator generates medical image data for analysis by the processing system 61.

[0139] The display unit 63 is adapted to receive information regarding the boundaries identified by the processing system 61 (e.g., boundary information) and display such information on, for example, the display 63A. The display unit may be adapted to display medical image data associated with the boundaries (e.g., beneath the illustrated boundaries). The medical image data may be received directly from the image data generator 62.

[0140] The display unit 63 may also include a user interface 63B, which, as is known to those skilled in the art, may allow a user to change or manipulate the view of the medical image (and thus the view of the visible boundaries).

[0141] Embodiments thus utilize a processing system. The processing system may be implemented in software and / or hardware in a variety of ways to perform the various required functions. A processor is an example of a processing system that employs one or more microprocessors, which may be programmed using software (e.g., pseudocode) to perform the required functions. However, the processing system may be implemented with or without a processor and may also be implemented as a combination of dedicated hardware for performing some functions and a processor (e.g., one or more programmed microprocessors and associated circuitry) for performing other functions.

[0142] Examples of processing system components that may be used in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs).

[0143] In various embodiments, the processor or processing system may be associated with one or more storage media, such as volatile and non-volatile computer memories, such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when executed on one or more processors and / or processing systems, perform the required functions. The various storage media may be fixed within the processor or processing system or may be transportable such that the one or more programs stored thereon may be loaded into the processor or processing system.

[0144] It should be understood that the disclosed methods are preferably computer-implemented methods. As such, the concept of a computer program including code components for implementing any of the described methods when the program is run on a processing system (such as a computer) is also proposed. Thus, different parts, lines, or code blocks of a computer program according to an embodiment may be executed by the processing system or computer to perform any of the methods described herein. In some alternative embodiments, the functions labeled in the blocks may not occur in the order labeled in the figures. For example, two consecutive blocks shown may in fact be executed substantially in parallel, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved.

[0145] By studying the drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments when practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously. If a computer program is discussed above, the computer program may be stored / distributed on a suitable medium (such as an optical storage medium or a solid-state medium supplied with or as part of other hardware), but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. If the term "adapted to" is used in the claims or the specification, it should be noted that the term "adapted to" is intended to be equivalent to the term "configured to". Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. A method (30) for identifying one or more boundaries of a lesion (10) in N-dimensional medical image data (1) of a region of a subject, the method comprising: Receiving (31) the N-dimensional medical image data (21), including image data points; The N-dimensional medical image data is processed (22) using a machine learning algorithm to generate: N-dimensional probability data (23), including a corresponding probability data point for each image data point, the corresponding probability data point indicating the probability that the image data point is part of a lesion; and N-dimensional uncertainty data (24) comprising, for each probability data point and thus for each image data point, a corresponding uncertainty data point indicating the uncertainty of the probability indicated by the probability data point; as well as One or more boundaries of a lesion in the medical image data are identified (33, 33A, 33B) using at least the probability data and the uncertainty data.

2. The method of claim 1, wherein the step of identifying one or more boundaries of the lesion comprises: identifying (41, 51) one or more potential lesions in the medical image data based on the probability data; as well as Each potential lesion is processed using at least the uncertainty data to identify one or more boundaries of the lesion in the medical image data.

3. The method of claim 2, wherein the step of identifying one or more boundaries of the lesion comprises: A region growing algorithm (42) is applied to each potential lesion, the region growing algorithm using the image data and the uncertainty data to define the boundaries of each potential lesion.

4. The method of claim 3, wherein applying the region growing algorithm (42) to each potential lesion comprises iteratively performing the following steps for each potential lesion: identifying (42A) perimeter image data points forming a perimeter of the potential lesion; identifying (42B) adjacent image data points, the adjacent image data points being image data points outside the potential lesion and immediately adjacent to any perimeter image data points among the perimeter image data points; as well as for each neighboring image data point, in response to the magnitude of the neighboring image data point being greater than a first magnitude threshold for the neighboring image data point, adding (42D) the neighboring image data point to the potential lesions, wherein the first magnitude threshold is based on the uncertainty data point associated with the neighboring image data point, The region growing algorithm terminates in response to (42C) no new neighboring image data points being identified.

5. The method according to any one of claims 2 to 4, wherein the step of identifying one or more boundaries of the lesion comprises applying a region reduction algorithm (52) to each potential lesion, the region reduction algorithm comprising iteratively performing the following steps: identifying (52A) perimeter image data points forming a perimeter of the potential lesion; for each perimeter image data point, in response to the amplitude of the perimeter image data point being less than a second amplitude threshold for the perimeter image data point, removing (52C) ​​the perimeter image data point from the potential lesions, wherein the second amplitude threshold is based on the uncertainty data point associated with the perimeter image data point, The region reduction algorithm terminates in response to (52B) no new perimeter image data points being identified.

6. The method of claim 2, wherein the step of identifying one or more boundaries of the lesion comprises: identifying one or more predicted boundary portions of each potential lesion using the probability data, each predicted boundary portion being a predicted location of a portion of a boundary of the potential lesion; using the uncertainty data to identify an uncertainty in each predicted boundary portion and / or an uncertainty in each potential lesion; selecting one or more of the predicted boundary portions based on the identified uncertainty of each boundary portion and / or the uncertainty of each potential lesion; presenting the selected predicted boundary portion to a user; receiving user input indicating one or more boundaries of a lesion after presenting the selected predicted boundary portion to the user; as well as The boundary portion is identified based on the received user input.

7. The method of claim 6, wherein the step of selecting one or more predicted boundary portions comprises: Those boundary portions that are associated with an uncertainty above a first predetermined uncertainty value and / or those boundary portions that are associated with lesions having an uncertainty below a second predetermined uncertainty value are selected.

8. The method according to any one of claims 2 to 4 and 6 to 7, wherein the step of identifying one or more potential lesions in the medical image data comprises: Groups of image data points associated with probability data points indicating that the probability that the image data point is part of a lesion exceeds a predetermined probability are identified, each group of image data points thereby forming a potential lesion.

9. The method according to any one of claims 1 to 4 and 6 to 7, further comprising: One or more corresponding graphical annotations are generated (34), each graphical annotation indicating the location of each of the one or more identified boundaries.

10. A computer program product comprising code means for performing the method according to any one of claims 1 to 9 when said program is run on a processing system.

11. A processing system for identifying one or more boundaries of a lesion within N-dimensional medical image data of a region of a subject, the processing system being adapted to: Receiving the N-dimensional medical image data, including image data points; The N-dimensional medical image data is processed using a machine learning algorithm to generate: N-dimensional probability data, including a corresponding probability data point for each image data point, the corresponding probability data point indicating a probability that the image data point is part of a lesion; and N-dimensional uncertainty data, including, for each probability data point, a corresponding uncertainty data point indicating the uncertainty of the indicated probability; and One or more boundaries of a lesion in the medical image data are identified using at least the probability data and the uncertainty data.

12. The processing system of claim 11, wherein the processing system is adapted to: identifying one or more potential lesions in the medical image data based on the probability data; and Each potential lesion is processed using at least the uncertainty data to identify one or more boundaries of the lesion in the medical image data.

13. The processing system of claim 12, wherein the processing system is adapted to process each potential lesion by applying a region growing algorithm to each potential lesion, the region growing algorithm using the image data and the uncertainty data to define the boundaries of each potential lesion.

14. The treatment system of claim 12, wherein the treatment system is adapted to treat each potential lesion by: identifying one or more predicted boundary portions of each potential lesion using the probability data, each predicted boundary portion being a predicted location of a portion of a boundary of the potential lesion; using the uncertainty data to identify an uncertainty in each predicted boundary portion and / or an uncertainty in each potential lesion; selecting one or more of the predicted boundary portions based on the identified uncertainty of each boundary portion and / or the uncertainty of each potential lesion; presenting the selected predicted boundary portion to a user; receiving user input indicating one or more boundaries of a lesion after presenting the selected predicted boundary portion to the user; as well as The boundary portion is identified based on the received user input.

15. The processing system according to any one of claims 11 to 14, further adapted to generate a corresponding one or more graphical annotations, each graphical annotation indicating the position of each one or more identified boundaries.

Citation Information

Patent Citations

  • System and method for whole body landmark detection, segmentation and change quantification in digital images

    US20070081712A1

  • Image processing apparatus

    US20170098301A1