Interactive image segmentation of abdominal structures
By introducing input data preprocessor and subset specifyer in medical image segmentation, and optimizing size designation with trained machine learning models, the problem of frequent user interaction in the prior art is solved, and segmentation efficiency and user experience are improved.
Patent Information
- Application Number
- CN202380072804.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-07
- Filing Date
- 2023-10-11
- Publication Date
- 2025-06-06
AI Technical Summary
The existing medical image segmentation technology is inefficient in clinical practice and frequent user interactions, which leads to user fatigue, especially in a clinical environment facing a large number of patients.
Design an input data preprocessor, based on the trained machine learning model through a subset specifier, optimizes the size specification during image segmentation, reduces user interaction, and improves segmentation efficiency.
By optimizing the size specification, the annotation work is significantly reduced, the performance of the segmentation algorithm is improved, the number of user interactions is reduced, and the fatigue of clinical users is reduced.
Smart Images

Figure CN120112941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an input data preprocessor for facilitating image segmentation, a related method, a system using such a preprocessor, in particular for segmentation, a medical imaging apparatus, a training system for training a machine learning model to implement such a preprocessor, a related training method, a computer program unit and a computer-readable medium. Background Art
[0002] Medical imaging has been a front-line staple of medicine at least since Wilhelm Roentgen discovered the X-ray in the 1800s. Being able to see "inside" a patient non-invasively is extremely valuable for diagnosis, treatment, scans, and other medical tasks.
[0003] Segmentation of medical images is another useful tool in image-based medical resources. Segmentation (a type of image processing) allows medical meaning to be given to images. For example, it helps users (such as radiologists or other medical users) quickly and safely find structures of interest in images during image viewing sessions. Images may not always be easy to interpret: for example, images may be high-dimensional, have high-resolution pixel density, and may not always be easy to quickly locate and distinguish interfaces of interest therein, especially under pressure caused by workload, etc. Being able to quickly navigate and distinguish image information is also a consideration in therapeutic image-guided interventions, where real-time pressure is increased, such as in catheter laboratories during cardiovascular applications. In such scenarios, projection images from various spatial directions can be acquired, which increases the complexity of locating structures of interest and medical tools (e.g., catheters, etc.). Therefore, medical imaging combined with segmentation adds more value. In addition, some organs can exist as highly complex spatial structures. Examples include tumor tissue in abdominal anatomical structures, such as tumor tissue in and around the pancreas. This makes it difficult for clinical users to sometimes tell "which is which". This is a particularly unfortunate combination of circumstances because pancreatic cancer happens to be one of the most deadly types of cancer.
[0004] Recently, segmentation algorithms based on medical machine learning ("ML") have been used, which can produce impressive results, but sometimes fall short of expectations in terms of producing outputs that do not meet clinical standards. Certain interactive medical segmentation algorithms have been proposed, in which the user can provide additional input data to better guide the machine learning model in the segmentation task. However, such user interaction models lead to user fatigue, as multiple user interventions are sometimes required to achieve acceptable results. This can be prohibitive in a clinical setting facing a large number of patients, due to an aging population, but also due to backlogs caused by, for example, epidemics, or due to workforce turnover or any other reason. Summary of the invention
[0005] Therefore, improvements in the medical field may be needed, especially to make image-based tasks more effective in clinical practice.
[0006] The objects of the invention are achieved by the subject-matter of the independent claims, wherein further embodiments are incorporated in the dependent claims.
[0007] It should be noted that the aspects of the present invention described below also apply to related methods, systems using such preprocessors in particular for segmentation, medical imaging devices, training systems for training machine learning models to implement such preprocessors, related training methods, computer program elements and computer-readable media.
[0008] According to one aspect, there is provided an input data preprocessor for facilitating (medical) image segmentation, which, when used, comprises:
[0009] an input port for receiving an input image to be segmented by the interactive machine learning based segmenter;
[0010] a subset designator configured to determine a size designation for a subset within an image based at least on the input image (b); and
[0011] An output interface for passing the size designation to a user interface or segmenter for interacting with the segmenter.
[0012] Wherein, the subset size specifier (SS) is based on a trained machine learning model (M', M).
[0013] Wherein, the trained machine learning model (M', M) is trained on a set of training images to determine an optimal size specification, and when the training images are segmented by the segmenter using the size specification as input, the optimal size specification maximizes the segmentation quality indicator.
[0014] The preprocessor can be used with any interactive segmentation model. The preprocessor ensures that the size specification (i.e., user input (e.g., circle)) has the "correct" size to improve the performance of interactive segmentation operations with fewer user interactions (such as "clicks"). For example, the performance can be measured by the Dice index of the overlap (in the set theoretic sense of intersection) of the segmentation output with the ground truth for a given number of interactions.
[0015] The trained machine learning model (M', M) that optimizes size specification is extended to using machine learning to optimize the user interface for annotation to train the actual segmentation machine learning algorithm. This is based on the insight that one of the expensive and time-consuming steps in segmentation machine learning is the annotation for generating ground truth. In particular, in the case of volumetric images, annotation has to be done slice by slice. By optimizing the user interface to suggest the best size for segmentation, the annotation effort is significantly reduced.
[0016] In an embodiment, the subset size specifier is based on a trained machine learning model. The machine learning model of the subset size specifier may be different from the model used for the segmenter. However, in some embodiments, an integrated model for both is also contemplated.
[0017] Subset designation is usually different from segmentation in that the prior may not follow edges or contours, but may also include background image matter, whereas segmentation is usually designed to follow such contours. The specified subset is usually more liberally cast, and is used to indicate areas on which subsequent segmentations of the segmenter will focus primarily, or areas to which more weight can be attached in subsequent segmentations.
[0018] In an embodiment, the size specification comprises geometric parameters of the subset within the image.
[0019] In an embodiment, the geometric parameter is at least any one or more of: i) a radius or diameter of the subset within the image, ii) a diagonal of the subset within the image, iii) a length of the subset within the image, iv) a volume of the subset within the image.
[0020] In an embodiment, the subset specifier is configured to determine the size specification further based on a previous size specification provided by a user via a user interface.
[0021] On the other hand, a user interface is provided herein, which is configured to allow a user to define a subset designation associated with an image structure in an input image by setting a direction vector next to or across the structure. The direction vector can then be extended with a geometric shape representing the subset to be specified. The setting of the direction vector can be supported as dragging, strokes, gestures, etc. along or across operations. The user interface can be used with any interactive segmenter. The user interface can be used in conjunction with an input data preprocessor, but can also be used without it. The user interface can provide a graphical user interface ("GUI").
[0022] In an embodiment, the subset specifier is further used to specify the shape and / or position of the subset within the image.
[0023] In an embodiment, a subset designator is used to determine a size designation based on the medical purpose of the segmentation.
[0024] In an embodiment, the subset is an n-sphere / n-1 sphere, where n=2, 3, or any one or more of polygons / polyhedra, such as a bounding box, free-line curve (closed or open), or any other.
[0025] In an embodiment, the user input interface comprises any one or more of the following: i) a pointer tool, ii) a touch screen, iii) a speech recognizer / NL component, iii) a keyboard. The speech recognizer / NL component may translate the user's verbal description into such a subset specification, such as a size.
[0026] In an embodiment, the segmentation is for an abdominal region of interest and / or a cancerous lesion.
[0027] In an embodiment, the abdominal region of interest includes at least a portion of any one or more of the liver, pancreas, pancreatic duct, pancreatic cancer tissue.
[0028] In another aspect, a system for medical image segmentation is provided, comprising: an input data preprocessor of any one of the above embodiments, and one or more of the following: an interactive segmenter and a user interface based on machine learning.
[0029] In an embodiment, the system comprises a subset adapter configured to adapt a size of a previous subset based on the size specification.
[0030] In an embodiment, the machine learning model of the segmenter and / or the machine learning model of the subset specifier is configured such that the resolution of the data propagable through the respective model is maintained.
[0031] In an embodiment, the input data pre-processor is part of the segmenter.
[0032] In an embodiment, the segmenter is configured to calculate a segmentation result based on i) the input image and based on a subset of the input image receivable via the user interface, and the subset has the specified size according to the input data preprocessor.
[0033] According to another aspect, a medical apparatus is provided, comprising a system of any one of the above aspects or embodiments, and any one or more of the following: i) a medical imaging device providing an input image, ii) a memory from which the input image can be retrieved, and iii) a resectability index calculation unit that can be configured to calculate at least one resectability index based on the segmentation result.
[0034] According to yet another aspect, a machine learning training system is provided, which is configured to train a machine learning model of a subset specifier (SS) based on training data according to any one of the above aspects or embodiments.
[0035] In some embodiments, the training system is driven by a cost function having weights that depend on the in-image distance to structures that have been segmented in previous training cycles.
[0036] According to yet another aspect, a training data generator system is provided, which is configured to generate training data for training a training system.
[0037] In an embodiment, the training data generator system is capable of operating the training of a first machine learning model for interactive image segmentation based on a first training data set, wherein the first training data set includes subset size specifications for subsets in the images of the first training data set, wherein the training data generator system is used to: change the size specifications to produce different training results in the training of the first machine learning model, and evaluate such results; and, provide second training data as training data based on the evaluated results.
[0038] According to yet another aspect, there is provided a training data generator system configured to operate training of a first machine learning model for interactive image segmentation based on a first training data set, the first training data set including subset size specifications for subsets in images of the first training data set, wherein the training data generator system is used to: vary the size specifications to produce different training results in the training of the first machine learning model, and evaluate such results; and, based on the results of the evaluation, provide second training data for training a second machine learning model, so that the second machine learning model can determine the subset size specifications for interactive segmentation of such images based on the input images in deployment. The second training data may include instances of varying subset size specifications associated with corresponding training input images used in the training of the first machine learning model. The evaluation may be based on the output of the objective function used in the training of the first machine learning model.
[0039] In another aspect, an apparatus is provided, comprising the training system of any one of the above aspects or embodiments, and the training data generator system of any one of the above aspects or embodiments.
[0040] On the other hand, the following uses are provided: ii) use of a training data generator system in a system for training a first machine learning model for interactive image segmentation, and ii) use of training data generated by a training data generator system (TDGS) of any of the above aspects of the embodiments in a system for training a first machine learning model for interactive image segmentation.
[0041] According to yet another aspect, there is provided a computer-implemented method for facilitating medical image segmentation, comprising:
[0042] receiving an input image to be segmented by the interactive machine learning based segmenter;
[0043] determining a size designation for a subset within an image based at least on the input image; and
[0044] The size specification is provided to a user interface for interacting with the segmenter.
[0045] According to yet another aspect, a method is provided for training a machine learning model of a subset designator based on training data according to any of the above aspects or embodiments.
[0046] According to yet another aspect, a method of generating at least part of training data for a training method is provided.
[0047] According to a further aspect, at least one computer program element is provided which, when run by at least one processing unit, is adapted to cause at least one processing unit to perform the method.
[0048] According to yet another aspect, at least one computer-readable medium is provided, on which a program element is stored, or on which a machine learning model of a subset designator is stored.
[0049] Thus, the subset specifier may be operable to dynamically adapt at least the size of a user input specifying a subset within an image for use with an interactive segmentation algorithm that processes such combined input of the image and the subset specification together, possibly adapted by the subset specifier to the determined size to compute a segmentation of the image. The segmenter comprises a computer-based implementation of a segmentation algorithm / arrangement of an interactive type, which is configured to accept such additional input (subset specification) alongside the actual image. Depending on the implementation, the subset size specification may be integrated in the image, or may be provided separately from the input image as a different data item. However, in either case, the subset size specification is additional information in addition to and on top of the information in the original input image, and is a result of an image processing operation performed by a user of the subset specifier. In the following, the subset size specification may be simply referred to as a "subset specification", it being understood that such a specification may include more information than a size. That is, the subset specification (the subset so specified) may involve an image structure (such as its size, shape or position, etc.) according to the input image. The image structure may be an initial image structure in the image, or an image structure previously segmented by a segmenter in a previous cycle, since a segmenter supported by an input data (pre)processor may be used one or more times in an iterative interactive manner. During such an interaction / segmenter cycle, such a subset designation that is changed / adapted is provided by a user via a user interface, possibly adapted (replaced) by a subset designation determined by a subset designator.
[0050] The operation of the subset designator may be based not only on image information (image structure, etc.), but may also be additionally based on context data, such as may be found in the imaging protocol, metadata, or as data specifically provided by the user via the interface UI (or via another (user) interface). The context data may define (such as via a medical code) the purpose of the imaging (which produces the input image), and / or may define the intended segmentation purpose. Thus, the determined size designation may correspond to the size, shape, position, etc. of the image structure relevant to the task, but the said purpose of the imaging or segmentation task at hand may also be taken into account. Thus, the determined size designation may require a size that is larger or smaller than a certain image structure size (representation of an organ / part thereof, cancerous / non-cancerous tissue / lesion, etc.) in the input image, which size is medically suitable and recommended for the medical purpose / task at hand. The image structure size may be understood as an in-image representation of an organ or anatomical structure or a part thereof, or an in-image representation of a tissue such as a cancerous / non-cancerous tissue / lesion, etc. The determined subset designation may represent an in-image subset (a subset of voxels or pixels). As desired, the in-image subset defined by the subset designation may cover a certain image structure of interest, or may only overlap with the image structure or may cover the image structure. For example, the subset designation may focus on the edge portion of the image structure. The subset designation may or may not represent a topologically connected in-image set.
[0051] The proposed system input data preprocessor allows to improve the performance of segmentation algorithms, especially those implemented by machine learning, preferably segmentation algorithms of the interactive type. The user input to the segmenter in terms of subset designation can be pre-adapted by the subset designator in this article to correspond to the image structure in a given input image in terms of size and / or position, so that the user interaction required to achieve the set performance level is completely reduced. Therefore, the proposed input data preprocessor can be preferably used in an interactive segmentation setting to reduce the number of interactive iteration loops. More specifically, one goal of this article is to reduce the number of such interactive segmentation loops, that is, the number of times the input is passed to the segmenter given a certain segmentation performance level. Given a set segmentation level, the performance level can be objectively quantified by associating the segmentation performance with the required number of interactions. Multiple scores or metrics can be used to measure the segmentation performance, such as based on Dice scores or similar, using verification / test data including real-world truth values. With the user input adapted according to the proposed system, the execution of interactive segmentation can be better than non-interactive fully automatic segmenters, and also better than other interactive segmentation settings, especially in applications where it is necessary to segment abdominal structures and / or cancer tissues associated with such abdominal structures.
[0052] The operation of the input data preprocessor can be based on the analysis of the input image, which is an initial / original image or a previously segmented image from an earlier segmentation cycle. The image input image can be received at the input port of the designator of the input data preprocessor. Optionally, the subset designator based on machine learning then processes the input image to output the size designation of the subset through its output interface, which will be considered by the segmentor in the next cycle. Therefore, in the next cycle, the input image can be processed by the segmentor together with the subset size designation provided, suggested or otherwise communicated by the input data processor. Therefore, the segmentor combines both the information item, the subset size information and the input image, and processes the combined data together to produce a segmented image with segmentation, which is the final image, or can be used as an intermediate segmentation result of (one or more) subsequent cycles, etc. Therefore, the system can operate in user-controlled iterations within one or more segmentation cycles. In the first or subsequent iterations, the segmentor uses the current subset designation to jointly process the input image or (one or more) segmentations of the previous cycle. Therefore, the subset designation can be updated based on the current segmentation result until the final segmentation is defined.
[0053] The operation of the subset designator may be corrective or advisory, and / or may be prospective or retrospective, as desired.
[0054] In some embodiments, the user interface includes a pointer tool. In some such or other embodiments, the user interface is configured to allow the user to define the length along or even across the image structure, thereby marking the structure. It has been found that some medical practitioners better receive this marking interaction by one or more drag / stroke operations compared to the click operation interaction of setting a predefined geometric shape area (such as a circle, etc.) into the input image. However, this click operation type interaction using a pointer tool is not excluded here.
[0055] The study showed that for certain image structures defined by image regions, using certain click radii results in a reduction in the need for clicks to achieve the same segmentation quality metric. Since during the training of a machine learning model for segmenting an image, an image is being segmented using a set of one or more clicks, the radius of the clicks can be easily adjusted and the effect of the adjustment can be evaluated. In the case where the set of one or more clicks is a simulated set of clicks, the radius can be varied between the sets of clicks and the resulting corresponding segmentations can be compared, and the best radius can be selected, associated with the segmented image in the machine learning model, and later, when the learned machine learning model is used to interactively segment other images by accepting clicks from the user, an appropriate click radius for use with a click input tool is suggested to the user, so that not only does the user not have to select the radius on his own, but also the number of clicks required to achieve the desired segmentation quality metric is lower.
[0056] During the training of the machine learning model for segmenting the image, instead of or in addition to the simulation, a set of one or more clicks may be derived from the user. In this case, the radius of the click input by the user is varied while, for example, keeping the center of the click at the same location, and then comparing the segmentation quality metrics for the various click radii used, thereby optimizing the machine learning model for segmenting the image.
[0057] In an embodiment of a machine learning model for segmenting an image, a first click radius and a second click radius are selected based on properties of an image region to be segmented.
[0058] Studies have shown that when a relatively small click radius is used, smaller structures tend to require fewer clicks, while when a relatively large click radius is used, larger structures tend to require fewer clicks. However, when using machine learning models for medical images to segment images, there is also an inverse correspondence between structure size and click radius. However, by varying the click radius during the training phase while training the image segmentation, the appropriate association between image structure size and anatomical structure, respectively, will be learned.
[0059] Although some of the above and following descriptions may involve user input in the form of clicks, strokes, gestures, scribbles, etc., any other form / type of specifying or indicating a subset (particularly its size) in a given image (such as a region in an image) may be equally used herein as is practically contemplated.
[0060] Although the above description has been mainly focused on applications in the medical field, non-medical applications are not excluded herein, and segmentation of structures other than the abdominal structure is not excluded.
[0061] Although, as described above, the input data preprocessor is preferably implemented as a machine learning model, either on its own / standalone model or in combination with a segmenter model, such an ML implementation is not required in all embodiments. Instead, non-machine learning models can be envisioned. For example, the input data preprocessor can be represented as a lookup table (LUT) in which appropriate subset sizes are tabulated for structural sizes or purposes (imaging purposes, segmentation purposes, clinical goals), etc. Thus, the subset specifier can look up the most appropriate size specification and suggest it to the user, or automatically apply the size specification. As mentioned, the subset size used for user interaction with the segmenter can also be derived not only from the image data itself, but also from certain metadata that can be found in the header data file of the image to be segmented. Such data can indicate the purpose / clinical goal of the intended segmentation.
[0062] In addition, the input data preprocessor can use other methods, such as non-machine learning based methods, such as classic analytical segmentation (region growing, snake, contour modeling, etc.). A first segmentation run can be applied to the input image to segment some structures and calculate, for example, a suggested size for the input subset size adaptation or a suggested subset. Sometimes, analytical segmentation methods may be computationally cheaper than using machine learning. However, such analytical pre-segmentation may produce a large number of structures, while ML-driven embodiments of the subset designator are able to better estimate task-relevant image structures only from the overall appearance of the image or based on background data.
[0063] The following relates to other embodiments in which subset size designations are expressed in terms of radius, but this is not limiting as any other subset designations are also contemplated and therefore references hereinafter to "first", "second", "radius", "position", etc. may be replaced by "first subset (size) designation" or "second subset (size) designation", respectively.
[0064] In one aspect, there is provided a computer-implemented method of training a machine learning model for segmenting an image, the image comprising an image region to be segmented, the method comprising:
[0065] - receiving an input training dataset comprising a set of images;
[0066] - receiving a set of one or more clicks having a click position and a click radius for a first image in the set of images, the click position and the click radius indicating a region to be segmented in the first image,
[0067] - segmenting the image using the set of one or more clicks, and
[0068] - evaluating the result of the step of segmenting the image based on a set of segmentation quality indicators (such as a cost function, a utility function driving the segmentation, etc.),
[0069] wherein the step of segmenting the image is performed using a first click radius that produces a first set of segmentation quality indicators and a second click radius that produces a second set of segmentation quality indicators;
[0070] If the first set of segmentation quality indicators is higher than the second set of segmentation quality indicators, the first click radius is associated with the first image, and if in the machine learning model, the first set of segmentation quality indicators is higher than the second set of segmentation quality indicators, the second click radius is associated with the first image, respectively.
[0071] In an embodiment, the set of one or more clicks is a set of one or more simulated clicks.
[0072] In an embodiment, a set of one or more clicks is derived from a user.
[0073] In an embodiment, the first click radius and the second click radius are selected based on properties of the image region to be segmented.
[0074] In some aspects, a computer-implemented method of segmenting an image is provided, the image comprising an image region to be segmented, the method comprising:
[0075] - receiving images;
[0076] - obtaining a click radius from the first trained machine learning model obtained by the method;
[0077] - Providing the image to the user
[0078] - providing the user with a click input tool having the click radius;
[0079] - receiving one or more click locations from the click input tool;
[0080] -Using the received image, the received one or more click locations, and the click radius, generating an output dataset comprising a segmented image from a second trained machine learning model.
[0081] In an embodiment, the image region to be segmented includes providing the segmented image to a user.
[0082] In an embodiment, the set of one or more clicks is derived from a user during performance of a computer-implemented method of segmenting an image comprising an image region to be segmented.
[0083] In another aspect, there is provided an ML training system comprising:
[0084] - a receiver arranged for receiving an input training data set comprising a set of images;
[0085] a receiver for receiving a set of one or more clicks having a click position and a click radius for a first image in the set of images, the click position and the click radius indicating a region to be segmented in the first image,
[0086] a processing unit arranged to segment the image using the set of one or more clicks, and
[0087] - an evaluator arranged for evaluating the result of the step of segmenting the image based on a set of segmentation quality indicators,
[0088] in,
[0089] The processing unit is further arranged to segment the image using a first click radius producing a first set of segmentation quality indicators and using a second click radius producing a second set of segmentation quality indicators;
[0090] And it is also arranged to associate the first click radius with the first image if the first set of segmentation quality indicators is higher than the second set of segmentation quality indicators, and to associate the second click radius with the first image if the second set of segmentation quality indicators in the machine learning model is higher than the first set of segmentation quality indicators.
[0091] In an embodiment, the set of one or more clicks is a set of one or more simulated clicks.
[0092] In an embodiment, the ML training system comprises a click input tool arranged to receive a set of one or more clicks from a user.
[0093] In an embodiment, the ML training system comprises a selector arranged to select the first click radius and the second click radius based on a property of the image region to be segmented.
[0094] In another aspect, there is provided an ML system for segmenting an image, the image comprising an image region to be segmented, the ML system comprising:
[0095] - a receiver for receiving an image;
[0096] - a processing unit arranged for obtaining a click radius from a first trained machine learning model obtained by the method;
[0097] - a display arranged to provide said image and a click input tool having said click radius to a user;
[0098] - the point input tool being arranged to receive one or more point locations from the user;
[0099] -The processing unit is also arranged to generate an output dataset comprising a segmented image according to a second trained machine learning model using the received image, the received one or more click locations and the click radius.
[0100] In general, the term "machine learning" includes a computerized device (or module) that implements a machine learning ("ML") algorithm. Some such ML algorithms operate to adjust the parameters of a machine learning model that is configured to perform ("learn") a task. Other MLs operate directly on training data, not necessarily using such a model. Such adjustment or updating of parameters based on a training data set is called "training". In general, the performance of a task by an ML module can be measurably improved by training experience. The training experience can include suitable training data and the exposure of the model to such training data. Task performance can be improved, the better the data represents the task to be learned. If the training data well represents the distribution of examples that measure the performance of the final system, the training experience helps to improve performance. Performance can be measured by objective testing based on the output generated by the module in response to feeding the module with test or verification data. Performance can be defined according to a specific error rate to be achieved for a given test data. See, for example, TM Mitchell, "Machine Learning", page 2, section 1.1, page 6, section 1.2.1, McGraw-Hill, 1997.
[0101] "Subset (specification)" is any data that defines a subset (sub-region, region portion) of a related image (initial input image, intermediate image, etc.). The subset (specification) can be provided by a user in interaction, or provided to be adapted / computed by a subset specifier. The specification can describe any one or more of the main envisioned size, position, shape. The specification can be explicit, such as "span of k pixels / voxel" or "length of k pixels / voxel", and / or can be a geometric equation of a trajectory, etc. Implicit definitions can include surface definitions through examples or instances, such as image masks (binary or probabilistic), which represent subsets and implicitly represent their in-image size, position (within the image), shape, etc. Implicit definitions can also include markings in the related image, such as by superposition of masks and images or otherwise. The subset specification can actually be encoded in the related image, such as by such markings or metadata, or in any other way. As described herein, the various methods of subset / subset specification / definition, etc. described above can be applied to any of training, training data generation, or deployment / inference. For the present purposes, sometimes, a subset designation (whether provided by the user or determined by a subset designator herein) may be identified with the subset itself. Thus, as desired and depending on the implementation, a segmenter or subset designator may operate on a designation and / or subsets (of image values comprised by a subset). Thus, shorthand language may be used herein, such as when referring to "a subset designation" covering (or not covering) "a certain image structure", which is explicit that the subset defined by the subset designation covers (or does not cover) the image structure. This shorthand may relate to other statements herein regarding subset designations.
[0102] Designation of drawings
[0103] Exemplary embodiments of the present invention will now be described with reference to the following drawings, which are not drawn to scale, in which:
[0104] Figure 1 A block diagram of a medical imaging apparatus preferably used for interactive image segmentation is shown;
[0105] Figure 2 A machine learning model that can be used for or facilitate machine learning based segmentation is shown;
[0106] Figure 3 A block diagram of a training system for training a machine learning model is shown;
[0107] Figure 4 A flow chart illustrating a computer-implemented method for facilitating medical segmentation of images;
[0108] Figure 5 A flowchart illustrating at least a portion of a computer-implemented method of generating training data for training a machine learning model; and
[0109] Figure 6 A flow chart of a computer-implemented method for training a machine learning model is shown. DETAILED DESCRIPTION
[0110] Reference now Figure 1 , which shows a block diagram of a medical imaging apparatus MIA. The medical imaging apparatus MIA may comprise a medical imaging apparatus IA of any suitable modality operable to acquire images of a patient PAT or a portion (region of interest, “ROI”) of a patient PAT in an imaging session. The medical imaging apparatus MIA may also comprise an imaging processing system IPS which may be arranged on one or more computing platforms PU. The imaging session may take place in an examination room of a medical institution such as a hospital, a clinic, a GP practice, etc.
[0111] In a broad sense, images acquired by the imaging device IA may be processed by the image processing system IPS to produce a segmented image m'. Specifically, the image processing system IPS may process a received input image m to produce a segmented image m'. When the input image m is acquired in an examination room, the input image m may be received directly from, for example, the imaging device IA via a wired or wireless connection in an online environment. Alternatively, the input image m is first stored in a preferably non-volatile memory IM (a database, storage device or other image repository, such as a PACS, etc.) after acquisition and then retrieved therefrom at a later stage after the image session when and if required, for example when a radiologist performs an image review, etc. A suitable interface / retrieval system may be run by a computing system (such as a workstation) for requesting the image m and its segmentation.
[0112] The computer-based image processing system IPS itself can run on one or more computing platforms PU, processing units, etc., which can be arranged in software or hardware or partially in both. All or part of the image processing system IPS can be implemented using a single such platform or computing system PU (such as a workstation or other in practice). Alternatively or additionally, there can be multiple such platforms PU, and all or part of the functions of the image processing system IPS are effectively distributed between those platforms, such as in a distributed computer environment, in a cloud setting, etc.
[0113] The segmented image m' produced by the image processing system IPS may be stored in another image storage device DB or in the same data image storage device IM from which the input image m is received. Alternatively or additionally, the segmented image m' may be visualized by a visualizer VIZ on a display device DD for visual inspection by a medical user to inform treatment or diagnosis. Additionally or conversely, the segmented image m' may be processed in other ways, for example it may be used for planning or controlling medical instruments, etc., whatever the medical task at hand, which may be usefully supported by the well-defined segmented images available herein. In one embodiment, the segmented image m' may be used in a tumor environment. For example, the resectability index calculation unit RCU may use the segmented image m' to calculate a resectability index r in order to assess the feasibility of removing cancer tissue from a patient PAT. The resectability index r provides an indication of the manner and extent to which cancer and non-cancerous tissues mesh with each other, and whether there is a well-defined tissue interface along which the removal of cancer tissue has an expectation of success. In order to distinguish tissue types, the accuracy of the calculation of the resectability index r relies heavily on detailed and accurate segmentation. Due to the good quality of segmentation results achievable herein, the accuracy of the calculation of the resectability index r can be improved.
[0114] Although the present application is primarily envisioned for medical imaging, in particular for tumor support, other medical applications or indeed non-medical applications are also envisioned herein. However, in the following, we will exclusively refer to medical applications, but again, this does not exclude non-medical uses or other medical applications for the proposed image processing system IPS as envisioned herein.
[0115] The image processing system IPS provides effective and in some cases excellent segmentation results, particularly for abdominal structures such as the pancreatic duct (PD), pancreas, liver or a combination thereof. The segmentation results are superior in the sense that they provide good results with lower user involvement, both of which are considerations in the medical field where stress, user fatigue, etc. are often encountered.
[0116] The image processing system IPS comprises a segmentation component SEG which is preferably driven by machine learning. In other words, the segmentation SEG comprises or is implemented by a trained machine learning model M' previously trained on suitable training data by a computer-implemented training system TS. Figure 2The training aspect is explained in more detail ahead, while the follow-up to this point will focus primarily on deployment, the post-training phase. In short, in machine learning ("ML"), training and deployment are different stages of use of a segmenter SEG or more generally an image processing system IPS. Training can be a one-time operation or can be repeated as needed when new training data becomes available. In some embodiments, training can also occur during deployment.
[0117] The segmentation setting SEG envisaged in this article is particularly of the interactive machine learning ("IML") type, also referred to as "interactive artificial intelligence" ("IAI") in some places. Interactive machine learning typically allows the user to provide input on top of the actual input image m to be segmented. In particular, the segmentation result m' is calculated not only based on the actual input image m itself, but also based on the user input (data). In addition to the input image, the segmenter SEG also considers the user input to calculate the segmentation result m'. The user input can represent domain knowledge expressed based on the input image. The user input can form useful background data related to the imaging goals and purposes, which the segmenter SEG can take into account when calculating the segmented image m' based on the input image m.
[0118] The functionality of the IML-based image processor IPS is optionally iterative in that the user may repeatedly provide user input in one or more loops j, respectively, in various intermediate segmentation results mj that may be produced by the segmenter SEG until a clinically acceptable result m' may be achieved. This interactive ML-based segmentation system SEG is in contrast to non-interactive systems, which are fully automatic and where the segmentation is computed without any user input. However, it has been observed that in certain such IML segmentation scenarios, such as those of structures involved at various scales that may be frequently encountered in cancer tissue including in abdominal structures, user inputs may usefully aim and guide the capabilities of the machine learning model M' to focus on certain regions of interest and thus provide excellent results in some cases even over non-interactive ML. Thus, the proposed system IPS is a good example where the capabilities of the user operate in useful combination with the capabilities of the machine learning, as opposed to "projecting man to machine".
[0119] Segmentation is generally a computer-implemented operation that assigns meaning to an image. It can be conceptualized as a classification task. An image is a spatial data structure of 2D, 3D, or 4D that is acquired. Image values constitute an image and are spatially arranged as image elements (pixels or voxels) in a 2D matrix or 3D matrix / tensor or even in a 4D structure such as a time series of a 3D volume. Any imaging device or modality is contemplated herein, such as X-ray, spectroscopy (e.g., dual-energy imaging), CT, MRI, PET / SPECT, US, or any other. Image structure is a subset of such image values in an image. The variation across image values (its spatial distribution) imparts image contrast, and thus generates (image) structure. Some of the (one or more) image structures represent background, while other image structures represent spatial extent, shape, distribution, etc. of different tissue types, organs, combinations of organs / tissues, solids (foreign matter, internal medical devices, etc.) or fluids (such as contrast agents that can be used in some imaging protocols, etc.). In some X-ray-based embodiments, spectral imaging is particularly contemplated herein. In the segmentation at the image element level (or patches thereof), a classification label is assigned to each image element by the segmenter SEG, which represents a certain tissue, organ or substance of interest. By segmenting essentially all image elements, their classification labels can be used by the visualizer VIZ to appropriately visualize the segmented image by color or gray value encoding based on the labels. This allows for better semantic visualization and differentiation, where the definition of different image regions represents different tissues, organs, substances of interest, respectively. Thus, medical users can interpret the images to derive added value for diagnosis and / or treatment or other. Image structures can have spatial interfaces (boundaries, edges, etc.) of high geometric complexity, possibly with intricately intertwined transitions from one tissue type to another, making such tissue interfaces difficult to discern. In this case, the good segmentation capabilities provided by the proposed interactive setting combined with the high image resolution provide a competitive differentiation between different tissue / organ types or substances.
[0120] Before explaining the image processing system IPS for segmentation purposes in more detail, reference will now first be made to the imaging apparatus IA in more detail to provide more background and to better assist in the later, more in-depth explanation of the operation of the imaging processing system IPS.
[0121] Typically, the imaging apparatus IA may comprise a signal source SiS and a detector device DeD. The signal source SS generates a signal, such as an interrogation signal, which interacts with the patient to generate a response signal, which is then measured by the detector device DeD and converted into measurement data, such as the medical image. An example of an imaging apparatus or imaging device is an X-ray based imaging apparatus, such as a radiographic apparatus, which is configured to generate a protective image. Volumetric tomographic (cross-sectional) imaging, such as via a C-arm imager or a CT (computed tomography) scanner or others is not excluded herein.
[0122] During the imaging session, the patient may reside on a patient couch, bed, etc., but this is not necessary, as the patient may also stand, squat or sit or assume any other body posture in the examination area during the imaging session. The examination area is formed by the portion of the space between the signal source SiS and the detector device DeD.
[0123] For example, in a CT setup, during an imaging session, an X-ray source SiS is rotated around an examination region with a patient therein to acquire projection images from different directions. The projection images are detected by a detector device DeD, in this case an X-ray sensitive detector. The detector device DeD can be rotated together with the X-ray source SS relative to the examination region, although such a joint rotation is not necessarily required, such as in CT scanners of the 4th generation or higher. A signal source SS, such as an X-ray source (X-ray tube), is activated so that an X-ray beam XB is emitted from a focal point in the tube during the rotation. The beam XB passes through the examination region and the patient tissue therein and interacts with the examination region and the patient tissue therein so that modified radiation is generated. The modified radiation is detected as intensity by the detector device DeD. The detector DeD device is coupled to an acquisition circuit, such as a DAQ, to capture the projection image in digital form as a digital image. The same principle applies to (planar) radiography, except that there is no rotation of the source SS during imaging. In such a radiographic setup, it is this projection image that can then be examined by a radiologist. In a tomography / rotational setting, the multi-directional projection images are first processed by a reconstruction algorithm, which transforms the projection images from the projection domain into cross-sectional images in the image domain. The image domain is located in the examination area. Projection images or reconstructed images will no longer be distinguished in this article, but simply referred to as (input) images / images m. It is such input images m that can be processed by the system IPS.
[0124] However, the input image m may not necessarily be generated by X-ray imaging. In contrast to the previously mentioned transmission imaging modalities, other imaging modalities are also contemplated herein, such as emission imaging, such as SPECT or PET, etc. Furthermore, in some embodiments, magnetic resonance imaging (MRI) is also contemplated herein.
[0125] In an MRI embodiment, the signal source SS is formed by a radio frequency coil, which may also be used as a detector device(s) DD, which is configured to receive, in a receive mode, a radio frequency response signal emitted by a patient residing in a magnetic field. Such a response signal is generated in response to a previous RF signal emitted by the coil in a transmit mode. However, in some embodiments, there may be dedicated transmit and receive coils, rather than the same coils used in the different modes.
[0126] In emission imaging, the source SS is in the patient's body in the form of a previously administered radiotracer that emits radioactive radiation that interacts with the patient's tissue. This interaction produces a gamma signal that is detected by a detection device DD (in this case, a gamma camera), which is preferably arranged in an annulus around the examination region where the patient is located during imaging.
[0127] Instead of or in addition to the above mentioned modalities, ultrasound (US) is also envisaged, wherein the signal source and detector device DeD is a suitable acoustic US transducer.
[0128] We now turn to the image processing system IPS in more detail, and continue to refer to Figure 1 , which preferably includes not only the segmenter SEG, but also an input data processor IDP that can operate in conjunction with the segmenter SEG. In a distributed computing setting, the input data processor IDP can be arranged on one computing platform, while the segmenter SEG can be arranged on another platform, possibly remote from each other, and the two can be interconnected in a wireless or wired or hybrid network setting. For example, the input data processor IDP can be run on a computer platform that can be operated by a user (such as a laptop, tablet, smart phone, workstation, etc.), while the segmenter SEG runs on a server that may be located remote from the input data processor IDP computing platform. For example, the computing platform running the segmenter SEG can be arranged outside the medical facility, while the input data processor IDP computer platform resides on site as needed.
[0129] It has been observed that in an interactive machine learning arrangement, the nature of the user input provided and processed together with the input image M is related to the overall performance of the segmenter SEG. In particular, it has been observed that the user input in some interactive segmentation settings includes the specification of subsets of the input image M to be segmented. These subset specifications may relate to clinical goals that the user wishes to achieve, and typically include at least part of the structure that the user wishes to segment.
[0130] Such subset designation can typically be defined / specified using a pointer tool, which is an example of many embodiments of the user interface UI contemplated herein. Such and other embodiments of the user interface UI are contemplated herein, which allow the user to specify such a subset. The designation may include an intra-image size (such as the number of pixels / voxels across) and a position within the image plane / volume of the input image m. For example, a user may use a stylus or a computer mouse to specify a subset of the input image to be input into the segmenter SEG along with the image m itself. The user may operate the user interface UI to specify, for example, a circular graphic, such as a circle, a sphere, an ellipse / ellipsoid, or a bounding box ("bbox"), such as a rectangle / square / prism, a triangle or other polygon, or indeed any other geometric figure. The dimension of the subset is preferably a function of the dimension of the input image m. For example, a 3D shape subset may be specified for a 3D image m, while a 2D shape subset may be used for a 2D image. The designation of a Lebesgue zero-metric subset is also contemplated, wherein the dimension of the specified subset is lower than the dimension of the surrounding image volume. A UI with subset propagation through dimensions can be used, where the user initially specifies in a lower dimension, but the UI's through-dimension functionality then propagates or expands the initial specification to higher spatial dimensions, adapting to local structure along the propagation depth. For example, in a tomographic image, for a volume m, a user can specify subsets in one or more slices. The subsets are then propagated through multiple adjacent slices, essentially expanding the initially specified (one or more) 2D subsets to s subvolumes of the input volume m, thereby producing 3D subsets. In the following, without ambiguity, we will use the concept "b" to refer to a subset specification or to a subset specified in terms of one defining another.
[0131] Regardless of the shape, the boundaries of the geometric figure / shape b so specified mark the image region that the user considers to include the image structure of interest for segmentation. That is, the area / volume covered by the geometric figure specifies a subset that includes at least some part, preferably all, of the structure of interest ms in its area or volume. The specification of the subset b (size, location and optionally the type of shape) is driven by the user's domain knowledge, such as clinical knowledge primarily contemplated herein, while some of the specified parameters (e.g., shape) may be preset.
[0132] The subset specification b describes the input image associated with the structure of interest ms. It can specify a combination or subset of coordinates. Alternatively, it can be defined by an analytical formula of locations that constitute a geometric figure, such as the center and radius / diameter of a circle, or any other geometry as required.
[0133] One way of specifying such a subset as contemplated herein is by "clicking" on the image m using a pointer tool UI while visualizing the image on the display device DD. Thus, a pointer tool such as a stylus or a computer mouse or other can be used. Typically, the pointer tool is operable to convert the physical movement caused by the user into the movement of a pointer symbol on the screen of the display device. They also allow the selection of image points or regions via the location accessed by the pointer symbol. The event handler detects user-system interaction events, such as the click event, and associates the current on-screen position of the pointer symbol with a subset that can be associated with the pointer position. For example, when a user clicks on an area / point in the image at or around the structure of interest ms, a graphical subset depictor (such as the circle or ball, sphere, etc.) can be presented together with the input image on the display DD for optional visualization. The subset depictor is a graphical representation of the subset so specified. This single-click interaction is preferred, which enables expansion into the described subset. However, multi-point interaction is not excluded, in which multiple points are set around the structure of interest ms, and when the user issues a further confirmation event or after a preset number or points have been set in this way, a surface is formed, or a curve is run to connect the points, or at least surround the set points, thereby setting the subset b. Therefore, as contemplated herein, the subset designation b can be specified via such a pointer tool (mouse, stylus) or indeed by the action of touching the screen. As an example of a single-point interaction, the center point of a circle is set, which extends in a circle at a certain radius. As mentioned above, a circle is just an example of such a subset definition. Any other shape can be used, and a circle or other shape for the subset b can be specified in addition to a mouse click action. It should also be understood herein that the visualization of the subset as a geometric figure / shape is not strictly required, because the most important thing in this article is the spatial definition of the subset (for example, in coordinates or other aspects), and its processing by the splitter, rather than the display of the specified subset b, although such a display is preferred in this article. Once a subset is specified, it should be understood that the subset b describes or at least relates to the clinical goal or property of the image, in particular to the location / region of the structure of interest ms, which in turn relates to the clinical goal sought by the user.
[0134] The pointer tool described is an example of a user input interface UI for interacting with the segmenter SEG. For the clinical goals to be achieved, this article can envision any other input interface UI type (not necessarily a single-point type) for the user to describe various aspects of the image structure based on the input image. However, in the following, we sometimes refer to user interaction as "clicking". But this example or represents any other means provided by the user input, because using a computer mouse with a familiar click operation is a popular user interface. This article specifically includes other ways of UI interaction. An example of another such UI will be described in more detail below. In addition, one purpose of this article is to reduce the interactive segmentation cycle, that is, the number of times the user input is passed to the segmenter SEG. Any given such user input may require one ("single point / click") or more than one ("multiple points / clicks") input operation to define any given subset designation in a given cycle.
[0135] For example, in another embodiment, the user may simply use scribbles (free line curves) to roughly surround the structure of interest ms. In another embodiment, the user simply defines a direction vector, such as by using a pointer tool. The drag interaction can be dragged along the structure of interest ms, wherein the length of the drag interaction roughly corresponds to the projection π of the length dimension of the image structure ms. In some embodiments, such drag-along type interactions are actually contemplated. In addition to the more familiar click-type operations described above, such drag-along interaction functionality may be preferred for medical practitioners, wherein a predefined geometric shape (such as a circle) is set into the image. Figure 1A This type of "drag-and-drop" interaction is shown in the illustration of ("drag vector") indicates a drag-along operation. A drag-along interaction may be accomplished by a click and drag interaction using a computer mouse or similar device such as a stylus, or may be accomplished by a touch screen action where the user drags along the structure ms on the screen with their finger. Drag vector The length dimension of the marking structure ms, for example in a projection. The dragging operation can actually be visualized on the screen of the display device DD (e.g. Figure 1A or similar), or no such display exists. When the drag operation ends, the drag vector The appropriate area above or below or to the left or right is captured as a subset b. For example, the vector length can be expanded to a rectangular subset that has the vector length as one of its sides, or at least one side has the vector length Therefore, the drag vector can be directed toward the set as Figure 1Aas shown, but this may not be the case as the drag operation can be run on either side of the expected area. Whether the area to the left or right of the drag vector is captured as subset b can be configurable or user controlled. The ML component of the UI can be used to check which side has the structure that better conforms to the expected segmentation, and the appropriate decision is made by the ML component of the UI. The range of the subset away from the drag vector can again be predefined, or again decided by ML based on the image structure, or can be temporarily controlled by the user. For example, two non-parallel drag vectors can be run along both sides of the expected structure, and the intersection of the corresponding areas away from and perpendicular to the corresponding drag vectors can be used to define subset b.
[0136] Despite the above Figure 1A The explanation of the user interaction in is constructed according to the drag operation, but this is only for clarity in this article. Although it is indeed envisioned in this embodiment, other embodiments that implement the basic principles of interaction are also so. Therefore, any other interaction methods that can indicate direction and length are also envisioned herein, and this article is centered around Figure 1A The interpretation is not limited to the drag-along type interaction. As a refinement of the above-mentioned drag-based operation, the drag vector can be extended to a 2D shape. The shape can be curved and is therefore not limited to a linear boundary shape b. In addition, the drag operation is not limited to a linear vector. For example, a drag operation along a structure can allow tracking a certain arc length of a non-linear curve. Therefore, the curve can at least partially (if not completely) provide an indication of the shape of the enclosing structure, and the curve can be extended to such an enclosing shape. Multi-segment vectors or curves can be defined by a user interface.
[0137] In addition, it should be understood that according to Figure 1A The proposed drag-along embodiment of the (G)UI is not necessarily associated with an image processing system having a subset specifier SS as described above. Figure 1A GUI and the above Figure 1A The variants discussed here may be standalone components that can be used with any interactive segmentation system, whether ML-based or not, even without the subset specifier SS. Figure 1A The GUI can still have certain performance benefits in combination with the subset specifier SS proposed in this paper.
[0138] Interactions supported by computer graphics widgets are not necessarily required in all embodiments herein. For example, subset designations may also be issued by a user as verbal instructions, which are interpreted by a natural language ("NL") processing pipeline (such as an NL machine learning model) as subset designations of structures in the image m. Such natural language-based subset designations may be used in conjunction with, or in place of, the geometry-driven user input methods described previously. Additionally or alternatively, a user interface UI for capturing and inputting subset designations may be based on eye / gaze tracking technology. The user interface UI may include a virtual reality (VR) headset, or the like.
[0139] The input data preprocessor IDP facilitates the interaction between the user and the segmenter SEG via the user interface UI. Specifically, the user can participate in one or more interaction loops j with the segmenter SEG via the user interface UI to initiate one or more segmentation calculations performed by the segmenter SEG, and the interaction is mediated and facilitated by the input data processor IDP. Broadly speaking, the IDP is operable to ensure that the subset designation b corresponds in size within the image to the in-image size of the image structure ms expected for the clinical goal at hand. It has been found that if the subset designation b corresponds in size to the size of the image structure, fewer interaction loops j with the segmenter SEG are required in order to achieve a (e.g., clinically) acceptable segmentation m'. Otherwise, for example, if a subset designation that is too large or too small in area is input for consideration by the segmenter SEG, more interaction loops j may be required to achieve the same segmentation result, as opposed to when the subset corresponds in size to the structure ms. Therefore, the proposed input data processor IDP facilitates finding a clinically acceptable segmentation while keeping the number of user interaction loops low. Therefore, the proposed input data pre-processor IDP (also sometimes referred to herein as "input data processor") helps to combat user fatigue, which is an important consideration especially in clinical settings.
[0140] In operation, the image processing system IPS receives an input image to be segmented. The input data processor IDP receives the input image m at its input IN and outputs an estimated appropriate subset size designation b to the user input device UI using its subset designator SS. The user input device UI uses the proposed subset designation to help the user define the input image with respect to the image structure ms. The input image is forwarded to the segmenter SEG together with the subset designation b for segmentation. The subset designation b may be forwarded as a separate data item from the input image m to which it belongs, or the subset designation b is forwarded as part of the input image, as an overlay, metadata, header data, or in any other form as required. Forwarding as separate data items and applying both data items, the image and its subset designation b may have some benefits when processed by the segmenter SEG. In either case, the segmenter SEG processes the image and its subset designation b together to calculate a first version m of the segmented image comprising a first segmentation j . It is the first segmented image m j This may then be presented to the user via a visualization on the display DD for cross-checking. If it is not standard, the user may again operate the user interface UI to specify a new subset designation bj, but now in the segmented image with the first segmentation mj, thus entering a second loop, wherein it is now the new subset designation bj (optionally again adapted by the input data processor IDP) together with the current version of the segmentation mj, which are then processed together by the segmenter SEG to produce the next version m(j+1) of the segmented image, and so on through user-controlled iterative loops until a clinically acceptable segmentation result m' is achieved.
[0141] During the iteration, the input data processor IDP may or may not intervene in the subset adaptation, depending on whether the current subset provided by the user via the UI has a volumetric correspondence. The actual size calculation by the subset designator SS may only need to be performed once in or before the first loop. In subsequent loops, the designator can only monitor the attachment to the calculated subset size, and adapt the current user-provided subset designation when necessary before passing it in the next loop for processing by the segmenter SEG. However, if necessary, re-adaptation / new calculation of the size can still be performed, because the size of the segmented structure can change during iterations during the loop. Whether to recalculate can be based on monitoring the size segmented by the designator SS, or such recalculation can be performed when requested by the user. Depending on the needs, the monitoring itself can be done by the ML, or can be based on rules, analysis, etc. The user interface can include a suitably configured user function that allows a request such as a recalculation of the size from the user interface UI to be issued for input into the data processor IDP. Therefore, in this embodiment, the user's initial or current specification b is intercepted by the event handler of the input data processor IDP and then resized based on the analysis of the subset specifier SS' on the medical input image m or the current segmented image mj produced by the segmenter SEG in the current cycle.
[0142] Thus, in the correction operating mode of the designator SS', the designation provided by the user can be resized, in particular can be enlarged or reduced as required by the subset adapter / sizer SA. Thus, the resizer SA can be resized in this way based on the subset designation b calculated by the designator SS. The resizer SA can be part of the subset designator SS, or can be arranged elsewhere in the processing pipeline, for example in the image input data processor IDP, or indeed at the input of the segmenter SEG, etc. In addition to this change in size, a repositioning in the image plane / volume can also be done by the input data processing system IDP. Sometimes it can also (or alternatively) be a geometric shape, which can be adapted as required, for example from a circle to a polygon or other. In some cases, the subset designator can only change the size according to the current designation b, without changing the position and / or shape as required.
[0143] Different from the above described corrective operation mode of the input data processor, in other embodiments the input data processor IDP may be configured to only suggest to the user a subset designation b of the corresponding size when calculated by the designator SS, for example by means of a visualization on the screen DD, preferably superimposed on the input image m or on its current segmented version mj in the current cycle j. The user may then accept this potentially resized subset compared to the subset initially designated by the user using the user interface UI.
[0144] The input data processor IDP may operate before receiving any user input and, based on its analysis of the current image, pre-suggest a subset designation of suitable size. The user may then accept that designation, or may actually choose to override the IDP's suggestion if desired.
[0145] Thus, according to the above, the input data (pre)processor IDP can have various modes of operation: correction or suggestion. The overturning function is not only envisaged in this suggestion mode of operation, but can also be used in any other mode of operation (such as correction mode). In addition, other modes involve the timing of the operation of the designator SS to calculate the corresponding subset size for segmentation. This timing can be said to be retrospective or prospective. If the designator SS is operable after a user issues a request to specify a subset in the input image via the user interface UI, the timing is retrospective. If the designator SS is operable before (any) such user request comes in, the timing can be prospective. In fact, once the image has been acquired by the imager IA, or once the acquired image m has been stored, before any inspection session, the subset designator SS can calculate the subset size. The subset size b calculated in this way can then be associated with some or each such future input image m. For example, the calculated subset size for use with the interactive segmenter SEG can be written out as metadata, can be written into the header of the image file, or can be stored in any other way, which can later be associated with image m when the inspection of image m is completed and applied in the manner described above in either correction or suggestion mode.
[0146] In some (but not all) embodiments, the subset designation b computed by the designator SS may be the smallest (area / volume size) geometry of a given type that still includes (i.e., covers) the structure of interest ms (either in the initial image or according to the current segmentation in a given segmentation cycle).
[0147] The subset that includes / covers the structure of interest ms is not necessarily defined herein by the subset designation determined by the subset designator SS. Therefore, there may or may not be an intersection area between the subset according to the determined subset designation and the image structure ms. In some embodiments, in addition to the image data m, the subset designator SS may also take into account background data indicating the clinical purpose of the imaging or segmentation. This other background data may be taken from patient files, protocol data, metadata, etc.
[0148] It is envisaged herein that the subset specifier SS is itself driven by a suitably trained machine learning model M', which may be different from the machine learning model M" driving the segmenter SEG. However, in other embodiments, there is a unified model M that includes both functionalities, namely, adaptation according to user input of the specifier SS, and segmentation via the segmenter SEG. Thus, in an embodiment, there is a single such machine learning model M that is trained to achieve both user input adaptation and segmentation.
[0149] The performance of the segmenter SEG in collaboration with the input data preprocessor IDP may be described by some metric q related to the segmentation performance, such as the number of clicks, or more generally the “number of interactions (“NOC”)”. For example, the number of interactions may be related to the correspondence of the segmentation produced by SEG to a test / validation set of ground-truth segmentations. The correspondence may be measured by the metric or score q. Metrics using concepts from discrete set theory, such as the Dice coefficient or others, may be used to measure the set-theoretic intra-image overlap between the produced segmentation results and certain ground-truth segmentations included in the test data. For example, the metric q may be defined based on the NOC of the score level q (“NOC@x”). Using such or similar performance metrics, it has been found that the proposed interactive machine learning system including user input size adaptation in combination with the segmenter SEG can achieve better scores than others (particularly other interactive types of image segmentation systems) within a certain number of NOCs. Non-interactive systems may also sometimes outperform, sometimes with NOC=1, particularly for the segmentation tasks of abdominal structures, such as the pancreas, PD, around the liver area, particularly still with associated cancer lesions. Specifically, and surprisingly, the proposed method achieves satisfactory scores for PD or pancreatic cancer vs. normal tissue segmentation with only a single interaction. In most cases around pancreas-related segmentation, the NOC is less than 3 or 4 for a given q-score level. Sometimes, based on the synergy of the input data processor IDP and the segmenter SEG, the proposed method outperforms other interactive ML segmentation systems in segmentation performance by up to 30%.
[0150] Instead of the subset specifier SS first providing the adapted / computed subset size specification b to the input user interface, it may alternatively or additionally be forwarded directly to the segmenter SEG as required, without a detour via the UI.
[0151] Reference now Figure 2 , which shows a block diagram of a machine learning model M" that can be used by the segmenter SEG. A similar or different ML model M' can be used to implement a subset specifier SS of the input data processor IDP, or indeed such a model M can be used to implement an integrated system of both the segmenter SEG and the input data processor IDP.
[0152] Broadly speaking, turning first to the ML implementation of the segmenter SEG, a machine learning model M' for the segmenter (or for the ensemble model M) jointly processes a current subset designation b adapted by an input data processor IDP according to user input and a current input image m, mj (which may be an initial image or a currently segmented image provided by the segmenter from a previous cycle or initial input image) to produce a segmented image m', mj+1 with its segmentations s, sj.
[0153] Due to the spatial correlation structure of the images to be processed, it has been found that the neural network type NN model has particular benefits in this article. More specifically, it has been found that the convolutional type neural network CNN model produces good results. However, other ML models, such as ML models using attention mechanisms, can also be used beneficially in this article.
[0154] In some of these NN model types, which are mainly contemplated herein, processing is arranged in individual layers Lj, where input is received at the input layer IL. The input layer responds with an intermediate data output, which is passed through the intermediate layer Lj to produce additional intermediate hidden layer intermediate output data (referred to as feature maps FM), which are then appropriately combined by the output layer OL into a structure having the same size as the input image. Feature maps are spatial data that can be represented as matrices or tensors like the input image, so each feature map can be said to have size and resolution. Deep architectures may be preferred in this article, meaning that there is usually at least one hidden layer, although their number is sometimes in the tens or hundreds or even more. Feedforward models can be used, but recursive types of NNs are not included, in which case the "depth" of the model is measured by the number of hidden layers multiplied by the number of passes. In other words, if there is a sufficient number of passes for repeated processing, in which feedback data is passed back from the downstream layer to the upstream layer, a recursive model with fewer layers than another such model can still be called a deeper architecture.
[0155] The representation, storage and processing of the feature maps FM and input / output data as matrices / tensors at the input / output layers provide accelerated processing. Many operations within the model can be implemented as matrix or dot product multiplications. Therefore, the use of matrix / tensor data representation allows a certain degree of parallel processing according to the single instruction multiple data ("SIMD") processing paradigm. Computing platforms with multi-core CPU designs (such as GPUs) or other computing devices adapted to parallel computing can be used to achieve responsive throughput in training and deployment. In particular, the input data processor IDP and / or the segmenter SEG and / or in fact the training system TS (see below) Figure 3) can be beneficially run on a computing platform PU with such a multi-core CPU, GPU, etc. SIMD processing can help achieve efficiency gains, especially on the training system TS during training.
[0156] Preferably, some or all of the layers IL, Lj, OL are convolutional layers, i.e., include one or more convolutional filters CV that process the input feature map FM from an earlier layer into an intermediate output, sometimes called a logit. For example, an optional bias term can be applied by addition. The activation operator of a given layer processes the logit into the next generation feature map in a nonlinear way, which is then output and passed as input to the next layer, and so on. The activation operator can be implemented as a rectified linear unit ("RELU"), or as a soft-max function, a sigmoid function, a hyperbolic tangent (tanh) function, or any other suitable nonlinear function. Optionally, there may be other functional layers, such as a pooling layer P or a drop-out layer, to facilitate more robust learning. The pooling layer P reduces the dimensionality of the output, while the drop-out layer cuts off the connections between nodes from different layers.
[0157] A specific type of machine learning model of the type contemplated herein is one in which a representation of an input image m, mj maintains its resolution as it is propagated through the various layers as feature maps and ultimately combined at the output layer into a final segmented image m or intermediate output mj. This feature of the model (which may be referred to herein as “resolution preservation”) is Figure 2 The dashed arrows represent the processing chain with the resolution-preserving feature map FM, terminating in the output layer OL, where the full-resolution data is combined with other feature maps to output a segmented image m'.
[0158] The resolution preservation as envisioned herein for some models M" for segmenter SEG is different from certain other models with bottleneck structures, such as encoder-decoder NN type. In such bottleneck structure models, the resolution of the data is changed by using a sequence of convolution operators followed by, for example, a sequence of deconvolution operators. Convolution operators with stride > 1 are used to first down-transform the input images m, mj into a lower blurred representation / resolution, called a latent representation or "code". The resolution of the code is then up-transformed / increased by sequential operations of a sequence of deconvolution operators to output data with a resolution matching the resolution of the input image at the output layer. However, it has been found that such down-transformation and up-transformation of resolution produces poor results in the interactive segmentation network proposed herein, especially for abdominal structure / cancer and non-cancer differentiation. In contrast, the model M" for segmenter SEG is configured to provide at least one chain of feature map representations propagated through the model (as shown by the dashed arrows), which preserves the resolution of the input images m, mj throughout the process until it is fed into the output layer OL. This preferred feature of resolution preservation does not exclude the model from having other processing chains through convolution / deconvolution operators (or by any other means), where such up-conversion and down-conversion are in fact envisaged. However, preferably, and in some embodiments, there should be at least one such chain that is resolution-preserving. Resolution preservation can be accomplished by using a series of convolution operators, each with a stride of 1. However, in principle, the ML setup and model selection / configuration of the segmenter SEG model M" can be of any type and kind described in the literature, as long as its layers are configured for multi-channel processing to accept input data in the input layer IL of the form (b[or bj], m[or mj]), as the (possibly adapted) subset specifies b to be processed jointly with the image data m, mj. See, for example, Jingding Wang et al., "Deep High-Resolution Representation Learning for Visual Recognition", in IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, MARCH 2020, available online at arXiv:1908.07919v2.
[0159] The subset designations b, bj processed by the model M" of the segmenter SEG together with the images m, mj (optionally adapted by the subset designator SS) can be understood as background data. The subset designations b, bj can be processed like images with the same size as the input image, wherein the image elements falling within the subset are marked accordingly. Therefore, the subset designations b, bj can be implemented as a map or mask. Such a map / mask (also referred to as a (subset) designation map / mask in this document) can be merged with the image data during propagation through the layers. For example, an image m having an image associated therewith Subset-specified masks b of the same size (e.g., a×b pixels / voxels) may be combined into a (2×a×b) tensor. As a form or co-compression, the merging may be achieved by applying a convolutional filter CV to the tensor, preferably with a depth of at least 2 to enable cross-channel convolutions. Using subset-specified b, bj as background data can be conceptualized as data-driven regularization in deployment. Regularization can guide the deployment / inference process to stabilize the true estimate for segmentation. The parameter landscape on which the model M" is trained is typically of very high dimensionality and may have a complex structure. Such complex parameter landscapes may cause the trained model M' to be ill-conditioned, with a large condition number and therefore sensitive to even small fluctuations in the input. Using user inputs b, bj as background data during training allows for more balanced learning (see below). Figure 3 ), and providing user-specified inputs b, bj and input images m, mj during deployment after training allows for more robust operation of the segmenter SEG.
[0160] Turning now to the ML model M' driving the operation of the input data processor IDP, much of what has been explained above also concerns the application of this model M'. Thus, M' may have a similar architecture to the model M", and in particular, may also be of convolutional NN type. M' may also be configured to be resolution preserving as described above. Its layers must be configured to accept input images m, mj as input, and the layers operate together with the output layer OL to combine the feature maps into representations of appropriately sized subset designations b, bj as output. Thus, the model M' processes the medical image data m, mj as input, and outputs a recommended size designation, such as the diameter of a circle or sphere, or indeed a mask / map representing a possibly resized subset designation, at its output layer OL as required. In addition, an indication of the location of the center of the circle / sphere may be provided, such as is done in the mask / map rendering. Thus, in a preferred embodiment, the representation of the output of the model M' may be in the form of such a mask or map, preferably having the same size as the input image. The mask may be a binary mask or a fuzzy mask having probability values indicating membership of each image element subset. A Gaussian or other appropriately parameterizable probability density function may alternatively be used to indicate the probability of membership of an image element (voxel / pixel) in a possibly resized subset.
[0161] Therefore, the two models M', M" have different scopes and objectives. This can be understood as follows: the segmentation SEG model M" is trained to learn the underlying relationship between the image m and its segmentation s. This can be referred to as the "image-to-segmentation relationship" (II). On the other hand, the model M' of the subset specifier SS is trained to learn the underlying relationship between the image m for subsequent segmentation and its appropriate subset designation b. This can be referred to in this article as the "image-to-subset size relationship" (I). It can be expected that both of these relationships are highly complex and not easy (if ever) to model explicitly using analytical methods that use evaluable closed / known function expressions. Machine learning does not rely on such explicit analytical modeling, but instead learns patterns in existing or temporarily generated training data. Instead of being used for analytical modeling, a general machine learning model M, M', M" architecture can be used instead. Based on training data, such as can be found in a medical database, the corresponding relationships (I)(II) are implicitly modeled by adjusting certain parameters of the machine learning models M', M", M based on the training set in the training phase. Once sufficiently trained (as can be determined by testing), the trained models M', M" can then be deployed in medical practice as described above to estimate subsets and segmentations, respectively. However, analytical modeling is not excluded herein and may be well used in some embodiments, particularly non-ML type embodiments. Furthermore, for the implementation of the subset specifier SS, LUT-based embodiments as described above or any other embodiments are not excluded herein.
[0162] Reference now Figure 3 ,pass Figure 3 A block diagram of a training system TS as may be envisaged herein in an embodiment for training any of the models M′, M″ or a common model M as described above is shown. Figure 3 , (x, y) represents training data, where x is the training data input and y is the associated ground truth or target. The training data (x, y) differs depending on whether M' or M" is trained, as will be explained in more detail below.
[0163] The data item x may be historical medical image data retrieved from a medical database, and y may be based on annotations provided by human experts. In some embodiments, at least some of the training data may be simulated by a training data generator system TDGS. The training system TS and the training data generator TDGS may be implemented on the same or different computing platforms. Preferably, the platform has a multi-core chipset to obtain better throughput, particularly when it comes to the implementation of the training system TS.
[0164] Turning now to the training aspect in more detail, in some cases, when training any of the models M′, M″, M, the operation of the training system TS may be controlled / driven by a cost function F, such as Figure 3 shown.
[0165] In general, one can therefore usually define two processing phases with respect to a machine learning model: a training phase and a later deployment (or inference) phase.
[0166] In the training phase, before the deployment phase, the model is trained by adapting its parameters based on the training data. Once trained, the model can then be used in the deployment phase to compute alternative discoveries (if any). Training can be a one-time operation, or can be repeated as new training data becomes available.
[0167] During the training phase, the architecture of the machine learning models M', M", M (such as Figure 2 The NN type shown in ) can be pre-populated with an initial set of weights. The weights θ of the models M', M', M represent the parameterization M' θ (which is the same for M' and M), and the purpose of the training system TS is to optimize and thus adapt the parameters θ based on the training data (x,y) pairs. In other words, learning can be formulated mathematically as an optimization scheme where a cost function F is minimized, although a dual formulation that maximizes a utility function may be used instead.
[0168] Training is the process of adapting the parameters of a model based on training data. An explicit model is not necessarily required, as in some examples the training data itself constitutes a pattern, such as in clustering techniques or k-nearest neighbors. In explicit modeling, such as in NN-based methods and many other methods, the model may include a system of model functions / computational nodes whose inputs and / or outputs are at least partially interconnected. The model function or node is associated with parameters θ that are adapted in training. The model function may include weights of convolution operators and / or nonlinear units (such as RELU), which are combined with NN-type models in Figure 3 In the case of NN, the parameters θ may include the weights of the convolution kernels of the operator CV and / or nonlinear units. The parameterized model can be formally written as M θ(M” and M are the same). Parameter adaptation can be achieved through a numerical optimization process. The optimization can be iterative. The objective function F can be used to guide or control the optimization process. The parameters are adapted or updated so that the objective function is improved. The input training data x is applied to the model. The model responds to produce the training data output. The objective function maps from the parameter space to a set of numbers. The objective function F measures the combined deviation between the training data output and the corresponding target. The parameters are iteratively adjusted so that the combined deviation decreases until a stopping condition preset by the user or designer is met. The objective function can quantify the deviation using a distance metric D[…].
[0169] In some embodiments, but not all, the combined deviation may be implemented as the sum of some or all residuals based on the training data instances / pairs, and the optimization problem in terms of the objective function may be formulated as:
[0170] argmin θ F=∑ k D[M θ (x k ),y k ] (1)
[0171] In setting (1), the optimization is formulated as the minimization of a cost function F, but this is not limiting in this context as the dual formulation maximizing the utility function may be used instead. The sum is the sum over the training data instances i.
[0172] The specifier model M' can be formulated as a classifier model or a regression model.
[0173] The cost function F can be pixel / voxel based, such as L1 or smoothed L1 norm, L2 norm, Hubert or soft margin cost function, etc. For example, in least squares or similar methods, the (squared) Euclidean distance type cost function in (1) can be used for regression tasks. When the model is configured as a classifier, the sum in (1) is alternatively formulated as one of the cross entropy or negative log-likelihood (NLL) divergence or the like. This article does not exclude formulations as generative models, nor does it exclude other formulations, including hybrids.
[0174] The updater UP of the system TS updates the model parameters based on the function F. The exact operation or functional composition of the updater UP depends on the optimization procedure implemented. For example, an optimization scheme (such as back / forward propagation or other gradient-based methods) can then be used to adapt the parameters θ of the model M so as to reduce the combined residual of all or a subset of the training pairs from the complete training dataset. Such subsets are sometimes called batches, and the optimization can be performed in batches until all the training dataset is exhausted, or until a predefined number of training data instances have been processed.
[0175] Segmenter Model M ” Training
[0176] When training a segmenter model M", x may likewise denote some input images as training input, and its target y is the ground-truth segmentation. The training input x may be obtained from a stock of historical images, such as may be found in a medical database, PACS, etc. However, when the model M" is trained for interactive segmentation, the training data input is contextualized and thus has the form x = (m, b), comprising the image m to be segmented (but also associated with the image), a relevant subset designation b. For training, the ground-truth segmentation y and the background data b, such as may be accessed in a medical database, may be provided by a human expert reviewing historical imagery. Alternatively, and preferably, background subset data b for the training input data x = (m, b) may be generated, for example by simulation d, and fed into the training system TS.
[0177] For example, one such pipeline TS for training an interactive segmentation model may proceed as follows, wherein background data such as a user-provided subset designation b is permitted for co-processing. The subset designation for training may be simulated at least for a first cycle to start the pipeline TS. The simulation preferably includes "positive" and "negative" for more robust training with good generalization. For example, a positive type of subset designation is one in which the designated subset at least partially includes (overlapping) foreground / structure. For example, the subset may overlap at least a portion of the structure, such as the center of gravity of the structure. A negative type of subset designation includes background or only background. The simulation may include random sampling to initialize a subset designation processing channel / chain in the training system TS. The training input image (such as a 2DCT image slice), the current subset designation (which may be empty initially) and the (initial) blank segmentation mask are combined channel by channel and provided as input to the segmentor model M" for co-processing, such as by cross-channel convolution or otherwise. The segmentor model M" then predicts a segmentation mask based on the provided input. Various statistical subset-specific sampling strategies can be used to avoid bias, such as those described in S. Mahadevan et al., “Iterically trained interactive segmentation”, arXiv preprint arXiv:1805.04398, 2018. In the first iteration loop, the initial b0 can be blank.
[0178] For some or each batch of training data during training, a maximum number of iterations may be randomly selected for the inner loop used for model parameter adaptation. Parameter adaptation may be driven by an objective function F. In some or each iteration, the predicted segmentation mask may be preferably used for sampling specified for another subset for the next cycle, and so on. This feedback sampling using the predicted segmentation mask from the current cycle is advantageous because it allows for improved training performance. Such as by combining the predicted mask with the input image by cascading along one spatial dimension, and simulating a new subset designation a, and feeding the data so combined into the same model, and so on. For each anatomical structure or medical target purpose, a separate model M" may be trained, although this is not required. Therefore, in some embodiments, the model M" may actually be a set of models M"=M" a , one or more for each anatomical structure / tissue, etc., "a" or medical purpose.
[0179] It is proposed herein, in some embodiments, that local training of a segmenter model M" may be a more feasible approach and provide better results than seeking optimization over the entire image. Proposed herein is a cost function F having weights that vary with proximity to the current image edge when presented in the segmentation of a given training input image during optimization seeking to improve the cost function. Specifically, F is configured to attach more weight to image elements that are spatially closer to an intra-image gradient edge (which may represent an interface between tissue types) than to image elements that are further away from such a (current) edge. In a given During iterations of the input image, the progression and / or clarity of the edge may change, and thus the weights of the function F will need to be dynamically and accordingly reallocated during iterations. For example, F may include a modulator or weight term d: F = d(v)*F, where d(vp)>d(vd), where a proximal voxel vp is closer to the edge than a distal voxel vd, and thus an edge-proximal voxel vp will incur a higher cost during training than a distal voxel vd. Depending on the implementation, the weight term d may assign a smaller weight to a voxel the closer it is to the interface / structure in the training output image M'(x).
[0180] This edge-related positioning of the cost function F allows focusing the learning on tissue interface surfaces, as these are of particular importance for accurate and relevant segmentation. In particular such tissue interface edges are considered for resectability assessment in oncology applications, as mentioned previously with respect to the index r. The exact location and progression of the actual interface are of course unknown, but the edges representing such interfaces appear as structures in the intermediate results mj calculated by the system TS (or indeed obtained elsewhere by some other process / model) during the learning of the model M” parameters. The model parameters are adapted based on the deviation of M” (x) from its target y. According to the example, as more weight is given to voxels or pixels that are closer to the currently defined image edge (interface), the learning system TS configures the model M” to focus more on such edge structures.
[0181] The positioning of the cost function F described above is optional. In addition, in addition to considering the contextualized form of the training data input data x = (m, b), any existing segmenter machine learning model and its training can be used in this paper.
[0182] Training of the specifier model M'
[0183] When training a specifier SS model M', x refers to some input image, and y is an appropriate subset designation (which preferably at least indicates an appropriate subset size) to be used when wishing to interactively segment the input image with the segmenter model M". As mentioned before, in this case y can be represented as a binary or fuzzy mask to represent / designate a subset b of an appropriate size.
[0184] As contemplated herein, the training system TS is based on training data provided at least in part by a training data generator system, although any suitable training system TS with or without a training data generator TDGS is contemplated herein, including those training systems that rely on human expert annotations. However, in some preferred embodiments, and as Figure 3AAs shown, the training is based on a computer-implemented training data generator TDGS, as will now be explained in more detail. Specifically, and as envisaged in some embodiments, the generation of training data for the subset designator model M' (learning the m vs b relationship) can be "attached" to the training for the segmenter model M" (learning the relationship m vs s). Broadly speaking, this can be done by changing the simulated background data b (subset designation data) in the training scheme of the model M". In this way, certain subset designations can be associated with related training input images m, thereby building a reserve of image vs subset designation pairs (x', y'). This training data (x', y') can then be fed into a new independent training, now dedicated to training the model M' to learn the relationship m vs b based on the pair (x', y'), which relationship is of primary interest in this context. This approach is advantageous because there is an unexpected effect: not only is training data provided for training the designator model M', but the segmenter model M" is also trained at the same time.
[0185] In particular, the training of the designator model M' can be achieved by first generating (e.g., by simulation) subset designations and then varying these subset designations over the dimensions bk (k indicating the dimension). The steps of training the segmenter model M" described above can be used for this, although in principle the simulations and / or the size variations can also be done manually by a human participant. For each bk and a given training input image m, the segmenter model M" is trained for the target y=s. The cost function F is evaluated by the evaluator EV. Any given training input image m is then associated by the selector SL with the corresponding bk* that produces the lowest cost among the bk considered for the given training input image x=m. In this way, an association or mapping with the training pairs (x'y') can be generated (see step S510, as follows Figure 5). Thus, each m is associated with its "best" bk that serves as a target, to obtain the pair j (x'=mj, y'=bk*j). When training M", the same operation is repeated for some or each training data input image m to build up a reserve of training data pairs (x', y') in the training of the model M" to train the specifier model M'. For example, an initial specification may first be associated with m generated by simulation / sampling in the training of the model M" as described above. The initial specification so simulated may then be deterministically altered in size, for example by scaling down or scaling up. Alternatively, the simulation is redone for each bk considered for a given training input image m. Once a sufficiently large reserve of training data has been accumulated, a new model M' may be used and then trained in a new separate training run on the set of (x'=mj, y'=bk*, j) pairs (now independent of the training of the model M"). Therefore, once the training of model M" is completed, the designator model M' is trained to predict, for a given image, its associated subset designation bk, which has the optimal size for subsequent segmentation of model M". The same model used for segmenter model M" can be used as model M' for designator SS, for example for transfer training. Alternatively, a new model M", is established with newly initialized parameters and then trained on the generated training data pairs (x', y'). The latter approach may be more beneficial because the image vs subset relationship m vs b that model M' will be trained on may have different characteristics from the segmentation relationship m vs s that segmenter model M", is trained to estimate. As mentioned earlier, the subset designations used in this paper can be simulated / randomized. The above-mentioned training data generation is coordinated in the training of segmenter model M", by the training data generator system TDGS.
[0186] As mentioned in conjunction with the training of the segmenter model M", preferably, in the training of the subset specifier SS model M', it is ensured that the simulation / randomization is configured such that there are subset designations of positive and negative types provided for one or more image structures of interest, for example, with respect to the current medical task for which the model M' is trained. This allows for improved performance. As described above, it is also preferred in this context to use the training output segmentation as feedback data on a training loop to sample the current subset designation for the next loop.
[0187] When training the specifier model M', the above-mentioned edge-dependent positioning of the cost function F relative to the model M" may also be considered.
[0188] i) The two processes of generating training data for training the specifier model M' by the generator TDGS and training the segmenter model M" by the system TS can be performed sequentially, one after another. However, preferably, the two are performed in combination, such as by interleaving or in parallel. However, for better processing efficiency, it is preferred in this article to combine the two processes i) and ii) by interleaving or other means.
[0189] It should be understood that Figure 3 , 3A The training setting in is just an example in this article. Any other ML training setting for training the subset specifier model M' may be used instead in this article.
[0190] Reference now Figure 4-6 , which shows a corresponding flow chart explaining the operation of the above-mentioned system IPS and training system / training data generation. However, it should be understood that the following flow chart is not necessarily associated with the system described above, but can also be understood separately as its own corresponding teaching.
[0191] Broadly speaking, Figure 4 A flow chart is explained of a computer implemented method of user interactive segmentation operation using user input adaptation as described above in conjunction with the operation of the specifier SS and the segmenter SEG after training. Figure 5 involves the generation of training data for training the specifier SS model M', while Figure 6 A flowchart of a method for training a specific model M' or an ensemble model M given training data is shown.
[0192] Now first refer to Figure 4 , which assumes that the system has been pre-trained, if indeed implemented in ML. Likewise, although ML is preferred in this paper, other analysis settings are not excluded in this paper.
[0193] In step S410, an input image to be segmented is received.
[0194] In step S420, a subset designation for a subset in the input image is determined, preferably using a trained subset designator machine learning model M'. The subset designation may be based on additional input from the user, including an earlier subset designation from an earlier segmentation cycle or a user-provided subset designation. Thus, determining the subset designation at step S420 may be based on image information in or around the subset in the image to which the earlier subset designation belongs, and / or may include consideration of the entire image and image information in and around the designated subset. The earlier subset designation may be adapted in step S420 based on the newly determined subset designation. If intended for use in iterative segmentation, as is in fact contemplated herein, the initial subset designation may be determined based solely on image information (possibly with background data), without such an earlier subset designation. The determined subset designation, and in fact the earlier subset designation, may preferably relate to the in-image size of the subset, but may also relate to the subset position and / or subset shape.
[0195] At step S430, the thus determined subset designation is output and used by a segmentation algorithm, preferably based on machine learning, to calculate a segmented image m' based on the subset designation b and the input image m, which includes a segmentation result s, which can then be output at step S440 for further processing.
[0196] The further processing step S450 may include any one of storing, displaying / visualizing or using it for control operation of the medical device, etc.
[0197] In an embodiment, step S450 comprises calculating the above-mentioned resectability index based on the segmentation s of the segmented image m'.
[0198] Further processing may include feedback to step S420, where new subset designations may be determined based on the segmentation s now presented in the segmented image m', using new user input (new user-provided subset designations bj) received in one or more segmentation cycles regarding the current segmentation m', mj, etc. Thus, the segmentation step S430 driven by the subset user input adaptation at step S420 may operate in one or more iterations, such as Figure 4 , as shown by the feedback arrow in . That is, the segmentation output at step S440 is not necessarily the final result, but may only be an intermediate result sj. For example, the intermediate result sj may still be displayed to allow the user to verify. If found to be below standard, the current segmentation is sent back to step S420, where a new user input bj about the new (current) segmentation sj is obtained, and then optionally adapted as described above, and then the new segmentation cycle is passed forward again, etc.
[0199] Now turn to Figure 5 , 6, these flowcharts specifically relate to the training aspects of the specifier model M'. The training of the segmenter model M" is not considered in detail herein, as any existing training method / system can be used as described above, as long as the segmenter training scheme allows processing of contextualized training inputs x = (mb), as described above. Although the two training settings of the models M', M" are different, they can be as described above in Figure 3 , 3A , where training data generation (particularly for training the subset specifier model M') is integrated into the training of the segmenter model M".
[0200] Reference now Figure 5 , which shows computer-aided training data generation for training in conjunction with interactive segmentation of imagery, in particular for a machine learning model M' for input adaptation by a specifier SS, as Figure 4 Such user input may include a click or other user interaction, wherein a subset in the input image is specified, which is to be segmented.
[0201] Although it is possible to train the model M' by using human participants / experts, this would be cumbersome. Therefore, in a preferred embodiment shown in step S510, the various subset designations b during training the segmenter model M' are generated based on simulations, for example by random sampling as described above. The training data generator step is preferably used in the context of training the segmenter model M", which is therefore a separate process. However, surprisingly, it has been found that associating the two processes produces good results for training the model M' for subset designation. Preferably, the generation of the training data set for the model M' is associated with the training of the model M', and the designator model M' will collaborate with the model M' after training.
[0202] The subset designation bk* associated with the image m thus generated in step S510 can be provided as training data for training the designator model M'. The underlying image m can be obtained as a historical image from an image database. The training image x'=m used to train the designator model M' can be different from the image used when training the segmenter model M".
[0203] Specifically, in an embodiment, the step S510 of generating a training data set for the specifier model M′ may be integrated in the training context of the segmenter model M″. Therefore, the training data step S510 may include:
[0204] In step S510_1, an input training image x'=(m, b) is received, including a subset designation b as background data. The subset designation b may specifically relate to an in-image size of a subset in the training input image m.
[0205] In step S510_2, a segmenter model M'' is trained based on the input training image, wherein the training includes changing a specification b of a subset in the training input image at multiple specifications of the subset to obtain different training results.
[0206] In step S510_3, the training result is evaluated based on a segmentation quality indicator (such as a quality indicator). An objective function F (such as a cost or utility function or other quality indicator) may be used.
[0207] Based on the evaluation, at step S510_4, an input image associated with one bk* among the multiple subset designations bk is provided as a training data item for training the / another machine learning model M' for subset designation adaptation. The learning model M' may be different from the one M' used for training segmentation. The input image associated with one bk* among the multiple subset designations may be an input image that produces the best (or good enough according to a threshold) quality indicator value or a quality indicator value that is at least higher than a preset quality threshold.
[0208] Steps S510_1 to S510_4 may be repeated on multiple training input images m to thereby obtain a set of training data items (x', y'), as described above in Figure 3 .3A.
[0209] Then you can refer to Figure 6 The training method described in the flowchart is used to process Figure 5 The training data (x', y') obtained in this way or the training data (x', y') obtained by other means (such as with the assistance (annotation) of human experts) is used to train the model M'.
[0210] In step S610, training data is received in the form of pairs (x', y'). Each pair includes a training input x'=m and an associated target subset specification y'=bk*, as described above. Figure 3 Location defined.
[0211] In step S620, the training input x' is applied to the initialized machine learning model M' to produce a training output M'(x').
[0212] The deviation or residual of the training output from the associated target y' is quantified at S630 by a cost function F. In one or more iterations in the inner loop, one or more parameters of the model are adapted at step S640 to improve the cost function. For example, the model parameters are adapted to reduce the residual measured by the cost function. In the case of using a convolutional model M' as actually envisaged in the embodiments, the parameters particularly include the weights of the convolution operator.
[0213] The training method then returns to step S610 in an outer loop, where the next pair of training data is fed. In step S620, the parameters of the model are adapted so that the aggregate residual of all pairs considered is reduced, in particular minimized. A cost function quantifies the aggregate residual. Forward-backward propagation or similar gradient-based techniques can be used for the inner loop. Instead of looping over individual data training items, such a loop can be done over a collection ("batch") of training data items, which is described more below.
[0214] Examples of gradient-based optimization may include gradient descent, stochastic gradient, conjugate gradient, maximum likelihood method, EM maximization, Gauss-Newton, etc. Methods other than gradient-based methods are also contemplated, such as simplex (Nelder-Mead), Bayesian optimization, simulated annealing, genetic algorithms, Monte Carlo methods, etc.
[0215] More generally, the parameters of the model M' are adjusted to improve the objective function F, which is a cost function or utility function. In an embodiment, the cost function is configured to measure the aggregated residual. In an embodiment, the aggregation of the residuals is achieved by summing all or some of the residuals of all pairs considered. In particular, in a preferred embodiment, the outer summation is performed in batches (subsets of training instances), and when adjusting the parameters in the inner loop, all of their summed residuals are considered at once. The outer loop then proceeds to the next batch, and by now, the necessary number of training data instances have been processed. Unlike pair-by-pair processing, the outer loop accesses multiple pairs of training data items at a time and loops batch by batch. Therefore, the sum of the index "k" in equation (1) above can be extended batch by batch over the entire corresponding batch. Figure 6 The cost function used in Figure 5 The cost function used for quality assessment in may be the same or different. As mentioned above, Figure 5 , 6 The cost function used in can be positioned relative to the edge.
[0216] Although reference has been made above mainly to NN models, the principles disclosed herein are not limited to NNs. For example, instead of using NNs for the preprocessor PP, another approach such as a Hidden Markov Model (HMM) or a Sequential Dynamic System (SDS), in particular a Cellular Automaton (CA), may be used.
[0217] Other models contemplated herein include support vector machines (SVM), or boosted decision trees may be used. Other models include generative models such as GANs or transformer type models.
[0218] As previously Figure 3A Mentioned in Figure 5 , 6The two processes may be interleaved to obtain better efficiency, such that both the generation of training data and the training of the specifier model M' may be completed quasi-simultaneously / in parallel as needed to obtain better efficiency.
[0219] Although the above has been explained mainly with reference to supervised learning, this is not necessary in this context, as unsupervised learning methods are not excluded herein. The model M' may usefully be implemented as a generative ML type model. Any such model type and appropriately configured training system may be envisaged herein. The training data generator system TDGS may be part of (integrated into) the training system TS.
[0220] As previously described, the model M' can be trained to learn not only to determine the subset size, but also to learn any one or more of the (in-image) position and / or shape of the subset as part of the subset designation. In any of these cases, the training data generator system TDGS can be configured to change not only the subset size, but also the position and / or shape. The shape change can be taken from a set of predefined shape types (e.g., bounding boxes, circles, etc.), or the subset shape can be changed without being limited to predefined shape types. A generative ML model can be used in the training data generator system TDGS to suggest new shapes during learning. Such a generative type ML model can be used in the training data generator system TDGS to discover new shapes that can improve the cost function. When an image mask / map is used as a representation of a subset designation with the same size / resolution (e.g., rows and columns) as the input image, the designator SS will learn to output its determined subset designation in the mask / map format, thereby automatically determining not only the size, but also the position and shape within the image.
[0221] In the above, the same concepts m, b, s, etc. (and related symbols such as mj, bj, etc.) are used for training data and data encountered in deployment / inference. But this is purely for simplicity and does not hinder the symbols. The training data and the data encountered during deployment / inference are different. Moreover, the concepts m, b, s, etc. used in this article refer to concepts rather than individual data items. Therefore, the use of the same symbols for training data and deployment data in this article does not mean that the data are equal, because they are not equal.
[0222] The components of the system IPS may be implemented as one or more software modules, running on one or more general purpose processing units PU, such as a workstation associated with an imager IA, or on a server computer associated with a group of imagers.
[0223] Alternatively, some or all components of the system IPS may be arranged in hardware, such as an appropriately programmed microcontroller or microprocessor, such as an FPGA (field programmable gate array), or as a hardwired IC chip, an application specific integrated circuit (ASIC), integrated into the imaging system IA. In yet another embodiment, the system IPS may be implemented both partially in software and partially in hardware.
[0224] The different components of the system IPS may be implemented on a single data processing unit PU. Alternatively, some or more components are implemented on different processing units PU, possibly arranged remotely in a distributed architecture and connectable in a suitable communication network, such as in a cloud setup or a client-server setup, etc.
[0225] One or more features described herein may be configured or implemented as or with circuits encoded in a computer readable medium and / or a combination thereof. The circuits may include discrete and / or integrated circuits, systems on a chip (SOCs) and combinations thereof, machines, computer systems, processors and memories, computer programs.
[0226] In a further exemplary embodiment of the present invention, a computer program or a computer program element is provided, characterized in that it is adapted to execute the method steps of the method according to one of the preceding embodiments on a suitable system.
[0227] Therefore, the computer program element can be stored on a computer unit, which can also be part of an embodiment of the present invention. The computing unit can be adapted to perform or cause the steps of the above method to be performed. In addition, it can be adapted to operate the components of the above device. The computing unit can be adapted to automatically operate and / or execute the user's commands. The computer program can be loaded into a working memory of a data processor. Therefore, a data processor can be equipped to perform the method of the present invention.
[0228] This exemplary embodiment of the invention covers both a computer program which right from the beginning uses the invention and a computer program which with the aid of an up-to-date program turns an existing program into a program which uses the invention.
[0229] Furthermore, the computer program element may be able to provide all necessary steps to implement the procedures of an exemplary embodiment of the method as described above.
[0230] According to a further exemplary embodiment of the present invention, a computer readable medium, such as a CD-ROM, is proposed, wherein the computer readable medium has a computer program element stored thereon, the computer program element being described by the preceding section.
[0231] The computer program may be stored and / or distributed on a suitable medium, in particular but not necessarily a non-transitory medium, such as an optical storage medium or a solid-state medium provided together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.
[0232] However, the computer program may also be presented over a network like the World Wide Web and may be downloaded from such a network into a working memory of a data processor. According to a further exemplary embodiment of the present invention, a medium for making a computer program element available for downloading is provided, said computer program element being arranged to perform a method according to one of the aforementioned embodiments of the present invention.
[0233] It must be noted that embodiments of the present invention are described with reference to different subject matters. Specifically, some embodiments are described with reference to method type claims, while other embodiments are described with reference to apparatus type claims. However, a person skilled in the art will gather from the above and following descriptions that, unless otherwise specified, any combination of features relating to different subject matters, in addition to any combination of features belonging to one type of subject matter, is also considered to be disclosed by the present application. However, all features may be combined, thereby providing a synergistic effect that is more than the simple sum of the features.
[0234] Although the present invention has been specified and described in detail in the drawings and the foregoing description, such specification and description are to be considered as specific or exemplary rather than restrictive. The present invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments may be understood and effected by those skilled in the art in practicing the claimed invention by studying the drawings, the disclosure, and the dependent claims.
[0235] In the claims, the word "comprising" does not exclude other elements or steps and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope. Such reference signs may consist of numbers, letters or any alphanumeric combination.
Claims
1. An input data preprocessor (IDP) for facilitating image segmentation of medical images, when used, include: an input port (IN) for receiving an input image to be segmented by an interactive machine learning based segmenter (SEG, M", M); a subset size specifier (SS) capable of determining a size specification (b) for a subset within an image based at least on said input image; and an output interface (OUT) for providing said size specification to a user to allow said user to interact with said segmenter (SEG) to segment said image, Wherein, the subset size specifier (SS) is based on the trained machine learning model (M', M) Wherein, the trained machine learning model (M', M) is trained on a set of training images to determine an optimal size specification, and when the training images are segmented by the segmenter using the size specification as input, the optimal size specification maximizes the segmentation quality indicator.
2. The input data preprocessor (IDP) according to claim 1, in, The in-image subset is a predefined shape or volume to be used as input to a segmentation algorithm.
3. The input data preprocessor (IDP) according to claim 2, in, The size specification includes geometric parameters of the subset within the image, wherein the geometric parameters are preferably any at least one or more of the following: i) a radius or diameter of the subset within the image, ii) a diagonal of the subset within the image, iii) a length of the subset within the image, iv) a volume of the subset within the image.
4. An input data preprocessor (IDP) according to any one of the preceding claims, in, The subset specifier (SS) is configured to determine the size specification (b) further based on a previous size specification provided by a user via a user interface (UI).
5. The input data preprocessor (IDP) according to claim 4, in, The user input interface (UI) includes any one or more of the following: i) a pointer tool, ii) a touch screen, iii) a speech recognizer component, iii) a keyboard.
6. The input data preprocessor (IDP) according to claim 5, in, The segmenter (SEG) is capable of segmenting an abdominal region of interest and / or a cancer lesion, wherein the abdominal region of interest includes at least a portion of any one or more of the following: liver, pancreas, pancreatic duct, pancreatic cancer tissue.
7. A medical image segmentation system (IPS), include: An input data pre-processor (IDP) as claimed in any one of the preceding claims; And, one or more of: the interactive machine learning based segmenter (SEG) and the user interface (UI).
8. The system according to claim 7, in, The segmenter (SEG) is used to calculate a segmentation result based on i) the input image and based on a subset of the input image that can be received via the user interface (UI) and the subset having a size specified according to the input data preprocessor (IDP).
9. A medical device (MA), include: The system according to claim 7 or 8; And, any one or more of the following: i) a medical imaging device (IA), which provides the input image, ii) a memory (IM), which can retrieve the input image from the memory (IM), iii) a resectability index calculation unit (RCU), which can be configured to calculate at least one resectability index based on the segmentation result.
10. A machine learning training system (TS) capable of training the machine learning model (M') of the subset specifier (SS) according to any one of claims 2-9 based on training data.
11. A training data generator system (TDGS) capable of generating the training data for a training system according to claim 10.
12. A computer-implemented method for facilitating image segmentation of medical images, include: receiving ( S410 ) an input image to be segmented by an interactive machine learning based segmenter; determining (S420) a size specification (b) for a subset within the image based at least on the input image; and The size specification is provided (S430) to interact with the segmenter. wherein the subset specifier (SS) is based on a trained machine learning model (M', M), Wherein, the trained machine learning model (M', M) is trained on a set of training images to determine an optimal size specification, and when the training images are segmented by the segmenter using the size specification, the optimal size specification maximizes a segmentation quality indicator.
13. At least one computer program element, which, when run by at least one processing unit (PU), is adapted to cause the at least one processing unit to perform the method according to claim 12.
14. At least one computer-readable medium having stored thereon a program unit according to claim 12, or having stored thereon a machine learning model according to any one of claims 2-10.