Computer-implemented method and image processing device for marking an object on a 3D image
By using 3D geometry tools and trained algorithms to automatically correct user input in 3D images, the problems of time-consuming manual labeling and reliance on user skills are solved, enabling faster and more reliable object labeling.
Patent Information
- Application Number
- CN202480035251.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-30
- Filing Date
- 2024-05-28
- Publication Date
- 2025-12-23
AI Technical Summary
Manually labeling objects in 3D images is time-consuming and depends on user skill and diligence, resulting in inconsistent labeling quality.
By providing 3D geometry tools to select locations on 2D slices, combining trained algorithms to automatically define 3D regions of interest, and applying neural networks or edge detection filters to correct user input, 3D volume markers are generated.
It significantly reduces tagging time, improves tagging reliability and accuracy, reduces reliance on user skills, and achieves a more intuitive and objective tagging process.
Smart Images

Figure CN121195282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computer-implemented method for marking at least one object on a 3D image, as well as a corresponding image processing apparatus and a method for training a neural network. Background Technology
[0002] Many medical applications, such as measurement or treatment planning, require segmenting structures, particularly depicting them, within volumetric (i.e., three-dimensional (3D)) images. For example, labels and annotations can be generated by manually drawing the boundaries of objects. However, manually generating labels for objects in an image can be a very time-consuming task. Furthermore, it requires the user to draw very accurately to correctly depict the object boundaries. Therefore, the quality of the labels is highly dependent on the user's skill and diligence.
[0003] Purpose of the invention Therefore, the object of the present invention is to provide a method to improve upon the aforementioned problems, particularly a more time-saving way to label at least one object on a 3D image. Furthermore, it is desirable to provide a more reliable method for labeling objects, one that relies less on user skill and diligence. Summary of the Invention
[0004] To achieve these objectives, the method according to claim 1, the image processing apparatus according to claim 12, and the method according to claim 13 are provided. Advantageous embodiments are set forth in the dependent claims. Any features, advantages, or alternative embodiments described herein with respect to the claimed methods are also applicable to other categories of claims, and vice versa.
[0005] According to a first aspect, a computer-implemented method is provided for labeling at least one object on 3D (three-dimensional) images, particularly medical 3D image data. The method includes the following steps: (a) Provide a 3D image including at least one object to be labeled, wherein the 3D image is capable of being represented by a plurality of 2D (two-dimensional) slices; (b) Displaying at least one 2D slice of the 3D image to the user and locating a 3D geometry within the 3D image, wherein a cross-section of the 3D geometry is displayed on the 2D slice; (c) Provide the user with a unit to select the position of the 3D geometry on the 2D slice, such that the 3D geometry at least partially overlaps with the object to be labeled, and receive position information of the 3D geometry selected by the user; (d) Automatically define a 3D region of interest within the 3D image, including the 3D geometry at a selected location, and crop the image data from the 3D image from the 3D region of interest; (e) The trained algorithm is applied to the image data from the 3D region of interest, wherein the output of the trained algorithm is the 3D volume within the 3D region of interest predicted to cover the object, wherein steps (c) to (e) are repeated at least once, such that an additional 3D volume (9) is generated depending on additional user input at each repetition, and wherein each additional 3D volume (9) is combined with the currently existing 3D volume (9) to produce a larger existing 3D volume (9). (f) Optionally, together with the 2D slice, output the labeled 3D volume.
[0006] This method can be performed by an image processing device. The image processing device can be, or may include, for example, a computer (such as a personal computer, cloud computer, tablet, or server). Additionally and / or alternatively, the image processing device can be part of an imaging system (particularly a medical imaging system). 3D images can include 3D image data. 3D images can be represented by volumetric rendering or by providing 2D slices cut through a 3D volume. This method is particularly advantageous when applied to medical 3D image data because providing labeling of objects within a 3D image is useful for many medical applications, including measurement and examination or treatment planning, and training machine algorithms, such as automatically identifying objects such as specific organs. However, this method can also be applied to other 3D images to label objects within those images. Objects can typically be 3D objects. In the case of medical 3D image data, an object can be, for example, an organ (such as a kidney) or a portion of an organ. In the context of this invention, labeling an object can mean that the image volume in the 3D image representing the object or a portion of an object is labeled. An object can be labeled by providing image coordinates located within the object. The output of the trained algorithm can be a "mask" of a 3D region of interest, where all pixels / voxels predicted to be outside the object have a value of 0, and all pixels / voxels predicted to be inside the object have a value ≠ 0, preferably the same value, such as 1. Voxels predicted to be inside the object can typically be contiguous, forming a single labeled 3D volume; however, this condition cannot be programmed into the trained algorithm and is therefore not an absolute requirement. The output of the trained algorithm can be equivalent to object segmentation, but not obtained through classical segmentation algorithms, but rather through the trained algorithm. In an alternative embodiment, the output of the trained algorithm can be voxels forming the boundaries between the object and other structures in the 3D image. In the labeled 3D volume output in step (f) (optionally along with 2D slicing), pixels or voxels belonging to the object can be identified, for example, by highlighting these pixels / voxels or otherwise marking or labeling them or their locations within the 3D image.
[0007] The term "labeled object" may sometimes also be referred to as "annotation object." The terms "pixel" and "voxel" are used interchangeably in this article, both referring to the values of the image matrix.
[0008] To provide a 3D image to a user, 2D slices of the 3D image are displayed. The user can identify the objects to be labeled within the slice and note the location of the object's boundaries within the slice. Therefore, the user can relatively clearly understand which parts of the 2D slice belong to objects and which do not. At this point, objects within the 2D slice can be manually labeled, and then switching to other 2D slices can be done to label objects in all these individual slices. However, this is a very time-consuming task. Furthermore, the reliability of this labeling will depend on the user's skill and diligence, typically related to the time spent on the process. Advantageously, 3D geometry is provided, and the cross-section of this 3D geometry is displayed to the user, specifically the cross-section located in the plane of the 2D slice. For example, the cross-section can be displayed by showing the boundary of the cross-section on top of the 2D image slice. The boundary can be displayed in a color with strong contrast to the image data of the 2D slice, making the cross-section easily identifiable by the user. For example, the image data on the 2D slice can be represented by grayscale values, and the 2D slice can be represented by color, such as a bright color like bright green. Alternatively, the area within the cross-section can also be colored, especially with transparent colors, so that the image data remains visible.
[0009] The user can then position the 3D geometry, specifically by moving the displayed cross-section. For example, the user can select the location using an input device, such as by moving a computer mouse and left-clicking or by touching a touchscreen. The 3D geometry can be a 3D mouse cursor. Typically, the user selects a location such that the cross-section overlaps with the object to be marked in the currently displayed slice. The user can select a location such that the boundary of the cross-section lies within the boundary of the object to be marked, so that the cross-section is completely within the object. However, the user can also at least slightly overlap or “overdraw” the boundary, so that the 3D geometry extends slightly beyond the object, i.e., beyond the object's boundary. Because the method of the present invention will still correctly identify and mark the object, the user can work more easily, thus saving significant time when marking objects. The user's input can be processed by an image processing device.
[0010] Therefore, due to the 3D shape of the 3D geometry, not only areas on 2D slices can be input and processed, but also the entire 3D region defined by the shape of the 3D geometry can be input and processed. Thus, the 3D geometry can be a tool that allows a user to simultaneously input positional information on multiple slices in a single action. This information can then advantageously be received, for example, by an image processing device. Advantageously, such a tool can reduce the effort required to depict an object by automatically drawing the 3D geometry around the click location, for example. For example, the tool could be a spherical tool corresponding to the 3D geometry of a sphere. Therefore, instead of generating only 2D circle markers, 3D sphere markers are created around the click location. Thus, multiple slices of a volumetric image can be edited at once, thereby significantly reducing the time required to label a complete 3D object. However, in general, since the shape of the 3D geometry will differ from the shape of the object, the shape of the 3D geometry at the selected location may not perfectly correspond to the boundaries of the object in all 2D slices of the 3D image. Therefore, unexpected variations may also occur in adjacent slices, as the shape of the object to be labeled is often not a perfect sphere itself or is not typically formed according to the 3D geometry. In particular, if the user does not select the correct radius, the 3D geometry may lie outside the object in some slices or even in the current slice. Therefore, if the user only uses this information as input, additional editing is required to correct the problem, or the markings may be highly inaccurate. This problem can also be addressed by reducing the radius of the sphere to decrease inaccuracy, but this will come at the cost of more editing time. In practice, choosing the correct radius can therefore be a trade-off between more 3D editing required and less correction for unintended changes.
[0011] Advantageously, a trained algorithm is provided that can automatically correct user input that extends beyond object boundaries. Specifically, the trained algorithm can be trained to segment objects within a 3D region of interest (ROI). For this purpose, the 3D ROI is automatically defined such that it includes 3D geometry. Preferably, the size and location of the 3D ROI are designed such that the boundary of the 3D ROI has a predefined distance to the outer shape of the 3D geometry located within that region. Preferably, the 3D ROI can be defined such that its center point corresponds to the center point of the 3D geometry. The corresponding center point can provide particularly reliable results. Image data from the 3D ROI is cropped from the 3D image, and the trained algorithm is applied to this image data. This specifically means extracting the image data from the 3D ROI and providing it to the trained algorithm, while not providing image data outside the 3D ROI to the trained algorithm. The trained algorithm can be, for example, a neural network. Alternatively, the trained algorithm can be, for example, an edge detection filter. The region from the center point of the ROI to the boundary edge can be selected. Edge detection filters can be graph-cut-based, specifically for dividing image data of a 3D region of interest (ROI) into inner and outer parts. The trained algorithm can have at least one input channel that takes the image data from the 3D ROI as input, and at least one output channel that provides the output of the trained algorithm. The image data from the 3D ROI can be or includes a 3D tensor containing image information at the 3D ROI. Specifically, the trained algorithm is trained or configured to automatically identify differences between objects in a 3D image and surrounding areas and output corresponding information. Advantageously, the trained algorithm provides information about which part of the ROI overlaps with the object to be labeled and outputs a 3D volume or information corresponding to the 3D volume within the 3D ROI predicted to cover the object.
[0012] According to an embodiment, a trained algorithm can be trained to output a 3D volume entirely within a 3D geometry. In other words, a 3D volume can be selected so that it lies within the 3D geometry. Therefore, the method of the present invention can provide a means of correcting overdrawing by the user on the boundaries of an object while maintaining proximity to user input. Advantageously, the user thus has more control over the drawing result. Additionally, a more intuitive method can be provided. Therefore, the 3D geometry can indicate the maximum volume that can be labeled.
[0013] According to an embodiment, the trained algorithm can be trained to output a 3D volume that lies within and additionally overlaps the 3D geometry by a predetermined amount, provided that the amount is within the object. In other words, the 3D volume can be selected such that it does not exceed a predefined maximum amount of the 3D geometry. This can be advantageous because it can be automatically corrected if the user has a habit of selecting areas that are not at the object's boundary but slightly away from it (e.g., by systematically misjudging the object's boundary). Advantageously, the labeling results can therefore be more objective and less dependent on individual users. This embodiment may be more advantageous because lower user accuracy may be required, so the user can potentially label the entire object faster. The labeled 3D volume can be binary information that distinguishes between in-object and out-of-object volumes. For example, a binary annotation of the 3D volume can be output. The labeled 3D volume can be output along with 2D slices. The 3D volume can be transferred to the image domain and drawn as a label on the displayed 2D slice. For example, the 2D slice can be displayed to the user along with labels showing the portions of the 2D slice that lie within the labeled 3D volume.
[0014] The 3D geometry can have a circular external shape. According to an embodiment, the first 3D geometry is an ellipsoid, preferably a sphere. Therefore, the 3D geometry can be spherical. For example, the 3D geometry can be a spherical mouse cursor. Correspondingly, the cross-section displayed in a 2D slice can have an elliptical outline, particularly a circular outline. Advantageously, a sphere can fit particularly well to the outline of an object. The size of the 3D geometry can be adjustable and can be selected before starting the marking method, particularly by the user. In useful embodiments, the size of the 3D geometry is small, particularly having a volume less than 1 / 5, preferably less than 1 / 10, of the object to be marked. Therefore, steps (c) to (e) or (f) are repeated several times to mark the complete object. This can advantageously create a user experience of "colored" objects, where the result is far less accurate than the user's actual drawing work.
[0015] According to an embodiment, in addition to image data, information corresponding to the 3D geometry is also input into the trained algorithm. This information specifically includes information about the size of the 3D geometry and, optionally, its relative position within the region of interest. For example, the information corresponding to the 3D geometry could be a click map corresponding to a user's mouse click, indicating the position and size of the 3D geometry relative to the 3D region of interest.
[0016] According to an embodiment, the 3D region of interest is a cuboid shape, particularly a cuboid shape, such as a cube. Correspondingly, within a 2D slice, a rectangular region around the cross-section can be considered within the 3D region of interest. Rectangular cubes have proven particularly advantageous when used as trained algorithms with trained neural networks, as the rectangular shape works well with typical neural network architectures. The 3D region of interest can be represented by a 3D image matrix or a 3D tensor.
[0017] According to an embodiment, the 3D region of interest is defined as being a predetermined amount larger than the 3D geometry. This feature allows for more specific training of the neural network to identify objects. In particular, speed and reliability can be optimized such that the size of the 3D region is chosen to be large enough to provide reliable results, but on the other hand, as small as possible to allow for optimization of computational speed, thereby optimizing the time required for the algorithm to provide output. The actual size to be applied can depend on the specific algorithm and its architecture used. The 3D region of interest can be 2 to 4 times larger than the 3D geometry, preferably 2.5 to 3 times larger. When the 3D region of interest is cuboid in shape, the side length of the 3D region of interest can be a predetermined amount larger than the diameter of the 3D geometry. It has been found that being 2 to 4 times larger is particularly suitable for determining the 3D volume sufficiently reliably. More preferably, being 2.5 to 3 times larger may be particularly suitable for combining the good robustness of the method with relatively fast computation time, especially when combined with the neural network explained herein, which may allow for near real-time processing of the label.
[0018] According to an embodiment, to determine the 3D volume within the 3D region of interest (ROI), a 3D Gaussian kernel, whose maximum value is at the center point of the cross-section, is fed together with the image data from the ROI into a trained algorithm, particularly a neural network, wherein the trained algorithm further determines the 3D volume based on the 3D Gaussian kernel. The image data from the ROI may be a 3D tensor containing image information at the ROI, accompanied by a second 3D tensor of the same size containing the 3D Gaussian kernel. The 3D Gaussian kernel can be passed as a second channel through the trained algorithm, particularly a trained neural network. The 3D tensor containing image information at the ROI and the second 3D tensor containing the 3D Gaussian kernel can together be a 4D tensor, particularly a 4D tensor with dimensions of 2×Z×Y×X, where X, Y, and Z are 3D image coordinates. Advantageously, the trained algorithm can thus correlate each intensity value with a relative distance to the center point of the cross-section and / or the center point of the 3D ROI. Therefore, a trained algorithm can, for example, emphasize intensity values closer to the center for greater reliability, i.e., reduce weights farther from the center. The 3D Gaussian kernel can be a general (three-dimensional) Gaussian form. According to an embodiment, the user is provided with the unit to change the size of the 3D geometry. For example, the user can be provided with the ability to change the size by scrolling a computer mouse wheel. The trained algorithm may include at least an input channel indicating the size of the 3D geometry. The size of the 3D geometry can be indicated relative to the size of the 3D region of interest. An interface corresponding to existing marker or annotation interfaces can be provided, particularly a classic spherical tool interface. Advantageously, the user can thus select the most convenient and suitable size of the 3D geometry. For example, a smaller 3D geometry may be advantageous for smaller objects or objects with more irregular shapes, while a larger size may be advantageous for larger and more regularly shaped objects.
[0019] According to embodiments, for each size of the 3D geometry, different trained algorithms, particularly different trained neural networks, are applied for that specific size. Training algorithms specifically for the specific size of the 3D geometry may be particularly advantageous, as it has been found that the algorithm can therefore be more reliable. For example, neural networks with output sizes of 8×8×8, 12×12×12, or 16×16×16 can be trained, where the output size corresponds to the size of the labeled 3D volume. The trained algorithm may include additional input channels indicating the size of the 3D geometry relative to the region of interest, and optionally its location. For example, when using a neural network, the additional input channels may be cascaded with an image-based input. The size of the region of interest relative to the output size can be determined based on the applied trained algorithm, particularly based on the network architecture of the trained neural network. For example, in the case of a convolutional neural network, using real convolutions instead of padding convolutions may generally mean that the input is larger than the output size. Therefore, as an example, for an input size of 18×18×18, the output size could be 6×6×6. For another example, with an input size of 40×40×40, the output size could be 28×28×28. However, these relative sizes can vary, especially due to the application architecture of the network.
[0020] According to an embodiment, the trained algorithm is a trained neural network. The trained neural network can preferably be a trained convolutional neural network. The trained neural network can be, in particular, a semantic network, preferably a U-Net-based network or an F-Net-based network. The trained neural network can preferably be trained to predict the expected segmentation output based on user selection. Preferably, the neural network has a shallow network architecture and is configured to provide output in real time. For example, the shallow network architecture can be a neural network with a U-Net structure having two levels, each with three convolutional blocks, or anything of that order of magnitude. Providing output in real time can specifically mean creating an output in less than 3 seconds, preferably less than 1 second. Surprisingly, it has been found that neural networks can deliver reliable results in a very short time, which even allows for almost immediate updating of labels after a user selection, resulting in a user experience with virtually no delay compared to users manually labeling objects entirely without neural network support. Furthermore, it has been found that neural networks are more reliable than other algorithms such as edge detection filters. For example, edge detection filters may fail to find edges on images with heavy texture or high noise, while trained neural networks have proven quite reliable in providing smooth edges that are closer to the true edges of objects. Semantic networks can be networks trained to output information about whether a pixel or voxel is part of an object for each pixel or voxel. According to a preferred embodiment, the U-trained convolutional neural network may include 1 to 3 levels, preferably 2 levels, each level including 2 to 4 convolutional blocks, preferably 3 convolutional blocks. This architecture has been found to be remarkably fast and reliable. In a particularly preferred embodiment, the trained neural network is a U-Net-based neural network. The general concept of the U-Net convolutional network is described in Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation” (MICCAI'15, LNCS, Vol. 9351, pp. 234-241, 2015). According to a preferred embodiment, U-Net may include 1 to 3 levels, preferably 2 levels, each level including 2 to 4 convolutional blocks, preferably 3 convolutional blocks. In an alternative preferred embodiment, the trained neural network is an F-Net-based neural network. F-Net-based neural networks can be particularly suitable for different organs, i.e., different kinds of objects. Furthermore, F-Net is optimized to reduce the amount of GPU memory required to process large images. That is, advantageously, F-Net may be particularly advantageous for processing large 3D images.The concept of F-Net is described in Brosch and Saalbach's "Foveal fully convolutional nets for multi-organ segmentation" (Proc. Spie, 105740u, 2018). Convolutional networks typically reduce the number of rows compared to the input. This can be achieved by padding the rows of data with zeros before feeding image data into the convolutional neural network, for example. This can be referred to as padding. Padding can help ensure that the output size is not too small compared to the input. According to alternative embodiments, the trained neural network may include real convolutional layers without padding convolutional layers. It has been found that avoiding padding can help avoid problems associated with padding. Furthermore, when using real convolutional layers instead of padding convolutional layers, the receptive field associated with the output region can be increased.
[0021] According to the invention, steps (c) to (e), particularly providing the user with units to select the position of the 3D geometry on the 2D slice, automatically defining the 3D region of interest, and applying the trained algorithm to image data from the 3D region of interest, are repeated at least once, such that an additional 3D volume, depending on additional user input, is generated at each repetition, and wherein each additional 3D volume is combined with the currently existing 3D volume to produce a larger existing 3D volume. Optionally, the currently existing 3D volume is displayed to the user and updated after each repetition. Advantageously, it has been found that the determination of the 3D volume can be so fast that updating the display can be essentially real-time, especially when using a trained neural network. Therefore, user selection can be processed quickly, and the 3D region of interest can be repeatedly input into the trained algorithm based on the current user selection. Thus, by adding additional markers corresponding to the additional user selections, the currently displayed markers can be updated based on the additional user selections. Thus, the user can “draw” markers by moving the 3D geometry, wherein the markers are automatically corrected to be within the object to be marked, particularly also for 2D slices not currently displayed to the user. Therefore, the process of tagging objects can be significantly accelerated. Furthermore, reliability can be improved, and it becomes less dependent on individual users.
[0022] According to an embodiment, the user is given the opportunity to switch the displayed view to another 2D slice of the 3D image and select a location within the currently displayed 2D slice, wherein method steps following the user's selection are performed based on the user's selection on the currently displayed slice. Advantageously, the user can thus switch between different slices, for example, add markers to slices where objects are wider than others, and generally control the marking of objects across the entire 3D image.
[0023] According to an embodiment, the user can select the intensity window of the displayed 2D slice. Therefore, the user can change the intensity window to achieve optimal contrast between the object and the background. This can advantageously improve the performance of the trained algorithm. For example, to mark the liver in a CT (computed tomography) scan (liver ~50-100 HU, surrounding tissue ~-100 HU; "HU" is the Henley scale), the user can select a window / level setting of 300 / 0 HU.
[0024] According to another aspect, an image processing apparatus is provided. This image processing apparatus is configured to perform the following steps: (a) Receive a 3D image including at least one object to be labeled, wherein the 3D image can be represented by a plurality of 2D slices; (b) Display at least one slice of the 3D image to the user; (c) Provide the user with a unit to select the position of the 3D geometry on the 2D slice such that the 3D geometry at least partially overlaps with the object to be labeled; (d) In response to the user selection, automatically define a 3D region of interest within the 3D image that includes the 3D geometry at the selected location, and crop image data from the 3D region of interest from the 3D image; (e) Applying a trained algorithm to the image data from the 3D region of interest, wherein the output of the trained algorithm is the 3D volume within the 3D region of interest predicted to cover the object; (f) Output the 2D slice together with the labeled 3D volume.
[0025] Specifically, the image processing device can be configured to perform method steps as described herein with respect to a computer-implemented method for marking at least one object on a 3D image. The image processing device may include a computer-readable storage medium on which specific instructions for performing the methods described herein may be stored. The image processing device may be or may include, for example, a computer, such as a personal computer, cloud computer, tablet computer, or server. Additionally and / or alternatively, the image processing device may be part of an imaging system (particularly a medical imaging system). The image processing device may include a user interface that allows information (such as cross-sections of 3D images and 3D geometries) to be displayed to a user. For this purpose, the image processing device may include or may be connected to a display device, such as a computer screen. Furthermore, the interface may include units for receiving user input (such as selection of 3D geometries on 2D slices). For this purpose, the image processing device may include input devices, such as a computer mouse and / or keyboard or touchpad. The user interface may be similar to a conventional user interface for marking at least one object on a 3D image. For example, the user interface may be configured to allow a user to select the size of the 3D geometry via a mouse wheel. For example, the user interface can be configured to allow users to select the location of 3D geometry by moving the computer mouse and clicking buttons such as the left mouse button.
[0026] According to another aspect, a method is provided for training a neural network, particularly for training a neural network corresponding to a trained neural network as described herein. The method includes the following steps: (a) Provide a set of 3D regions of interest from a 3D image as input training data, wherein each 3D region of interest includes at least a portion of an object on the 3D image; (b) Provide a 3D dataset as output training data, the 3D dataset having the size of the 3D region of interest and containing information for each voxel, the voxel being labeled as a part of the object or not corresponding to the object for each provided 3D region of interest; (c) The neural network is trained using the input training data and the output training data.
[0027] Therefore, a neural network can be trained to predict the segmentation output expected by the user based on user selection. This is simulated by providing a 3D region of interest (ROI) corresponding to a 3D region of interest, which is automatically created based on the user's selection of the location of a 3D geometry in a computer-implemented method for object labeling as described herein. The “expected” segmentation is simulated by the output training data, as voxels are labeled as whether they are part of an object. Preferably, the input training data comprises a set of different images with different kinds of objects to be labeled, such as different organs. Advantageously, a different set of images allows for the generation of a more flexible trained neural network capable of recognizing different kinds of objects, i.e., more independent of the specific structure to be labeled. Thus, better generality can be advantageously achieved. The images used as training data can be randomly selected from several 3D images. Furthermore, objects to be labeled within these 3D images can be randomly selected. The selected structures can have similar contrast, such as all dark contrasts. This can allow for the creation of more specialized trained neural networks that may be more robust in their specialized domain or may require less training data. On the other hand, the selected structures can cover different kinds of contrast, such as objects that appear bright and objects that appear dark. This can generate neural networks that can be applied more generally to a wider range of objects. Regions of interest (ROIs) can be systematically selected, particularly to simulate typical user behavior. On the other hand, randomly selecting ROIs can also be an option. This can, for example, allow for faster preparation of training data. Furthermore, random selection can allow the neural network to be trained to be more robust to unpredictable user behavior. For example, the neural network can also be trained to work when the ROI is entirely within the object, making the boundaries invisible on the ROI. Therefore, the selected ROI does not need to be perfectly reasonable relative to the object's anatomy or perfectly match the appearance of a given imaging modality, as smaller deviations can increase the robustness of the trained neural network.
[0028] According to an embodiment, each provided 3D region of interest is selected such that the distance between its center point and the boundary of the object has a specified value, wherein the specified value is a fixed value or a fixed value with added random deviation within a predefined maximum deviation amount. The center point can be specifically determined in the plane of the 3D image (i.e., within a 2D slice of the 3D image). Thus, the center point can optionally be only the center point of a cross-section of the 3D region of interest. Advantageously, the specified value can correspond to the size of the 3D geometry or the cross-section of the 3D geometry, for example, corresponding to the radius of a sphere or circle. Thus, by means of this embodiment, the user's selection of the position of the 3D geometry can be simulated, particularly such that the boundary of the cross-section is placed according to the boundary of the object. Therefore, the specified value can be a fixed value corresponding to the size (especially the radius) of the cross-section. Adding random deviation to the fixed value may be particularly advantageous. This can simulate different user behaviors, for example, different habits in selecting the position of the 3D geometry and / or different application precision during different user selections. Advantageously, the trained neural network can thus become more robust to user-related deviations. Random deviation can also be regarded as noise in the position of the center point of the 3D region of interest. In one embodiment, random bias can be constrained such that the distance between the object's boundary and center point decreases only randomly relative to a fixed value. This can advantageously simulate the selection of locations that slightly overlap with the boundary, i.e., the user "drawing" on the boundary. This can be particularly advantageous because the automatic correction of the method of the present invention can automatically correct this "overdrawing" by applying a trained neural network, thus overdrawing can be beneficial in ensuring that striped portions of the object at the boundary are not missed. Therefore, training the network specifically for this user behavior can be particularly advantageous.
[0029] According to an embodiment, the maximum deviation is predefined relative to the size of the 3D geometry to be used after training, such that a larger size corresponds to a larger maximum deviation. Advantageously, the predefined maximum deviation can be selected relative to the size of the 3D region of interest, and therefore particularly relative to a fixed value. This relationship with the size of the 3D region of interest can specifically correspond to the size of the 3D geometry. Thus, a larger size of the 3D region of interest can correspond to a larger size of the predefined maximum deviation. This can advantageously take into account user behavior, which is often drawn or selected more accurately when using smaller 3D geometry. Therefore, the neural network can be trained to be specifically prepared for this user behavior, and thus more robust when the object labeling method of the present invention is applied.
[0030] According to an embodiment, multiple neural networks are trained, each trained for a different specified size and a specific output size of the 3D geometry. Therefore, advantageously, each neural network can be dedicated to a specific size of the 3D geometry. This allows for the application of object labeling methods, enabling users to select different sizes of the 3D geometry, specifically corresponding to different trained specified sizes. Dedicated neural networks trained specifically for the corresponding sizes can be applied. For example, neural networks can be trained for output sizes of 6×6×6, 8×8×8, 12×12×12, 16×16×16, and / or 28×28×28, where the output size can specifically correspond to the output size of the neural network, i.e., the labeled 3D volume. For example, a neural network with an input size of 18×18×18 and an output size of 6×6×6 has been successfully tested. As another example, an input size of 40×40×40 has been trained together with an output size of 28×28×28. However, these relative sizes may vary, particularly due to the application architecture of the neural networks.
[0031] According to an embodiment, the intensity window of the 3D region of interest is randomly determined based on at least one PERT distribution, wherein the at least one PERT distribution is specifically based on the minimum / maximum (i.e., minimum / maximum) data values of the corresponding 3D image, the average data value inside the object, the average data value outside the object, and the center between the average data values. The at least one PERT distribution may include two partial PERT distributions, a first partial PERT distribution and a second partial PERT distribution, which share a common center value, wherein the center value is specifically the maximum value of the first partial PERT distribution and the minimum value of the second partial PERT distribution. The first partial PERT distribution may be defined by the minimum data value of the corresponding 3D image, the average data value inside the object, and the center between the average data values. Specifically, the average data value inside the object may be the most probable value of the first partial PERT distribution, the minimum data value of the corresponding 3D image may be the minimum value of the first partial PERT distribution, and the center between the average data values may be the maximum value of the first partial PERT distribution. The second partial PERT distribution may be defined by the maximum data value of the corresponding 3D image, the average data value outside the object, and the center between the average data values. In this model, the average data value outside the object can be the most probable value of the second part of the PERT distribution, the maximum data value of the corresponding 3D image can be the maximum value of the second part of the PERT distribution, and the center between the average data values can be the minimum value of the second part of the PERT distribution. 3D images are typically based on a range of data values, such as grayscale values ranging from dark to light. The range of measured data values can depend on the image modality and the precision applied. For example, the range of data values can be from 0 to 255, where 0 represents light and 255 represents dark. Therefore, the minimum value can be the lowest data value in the range, and the maximum value can be the highest data value in the range. Furthermore, to allow for differentiation between the object and its surrounding area (background), an image modality is applied to create the 3D image, resulting in a clear contrast between the object and its surrounding area; for example, the object has only or at least most of high data values, while the area surrounding the object has only or at least most of low data values, and vice versa. For example, considering the range of grayscale values, in a 3D image, the object may appear bright, while the surrounding area may appear dark. However, adjusting window settings and leveling (i.e., intensity window) can be beneficial for optimizing the contrast between the object and the background. Window settings and leveling define a range of data values resolved using grayscale values, where all data values outside this range are cropped, either completely darkened or completely brightened. By cropping some data values, the contrast between the object and the background can be made sharper. Therefore, it can be expected that users will set window settings and leveling accordingly. However, the exact settings can vary based on user preferences.Advantageously, applying the Pert distribution to prepare training data with randomly distributed intensity windows may be an effective way to simulate various user-relevant intensity windows. The average data values inside and outside the object preferably correspond to the most probable values of the Pert distribution. Therefore, this simulates how users typically set intensity windows such that the average data values of the object contrast with the average data values of the background to the greatest extent possible, but also takes into account deviations from this approach. Thus, the intensity window will be selected between a lower and an upper limit, where both ends are randomized according to the Pert distribution (specifically, a two-part Pert distribution). Preferably, the lower limit of the window setting can be drawn from the first part of the Pert distribution, and the upper limit of the window setting can be drawn from the second part of the Pert distribution.
[0032] According to another aspect, a computer program is provided that includes instructions, which, when run by a computer, cause the computer to perform a method as described herein for marking at least one object on a 3D image. The computer program may have the features and advantages described herein.
[0033] Features and advantages of different embodiments can be combined. Features and advantages of one aspect of the invention (e.g., the method) can be applied to other aspects (e.g., computer programs and data processing systems), and vice versa. Attached Figure Description
[0034] The invention will now be described with reference to embodiments in the accompanying drawings, in which: Figure 1 A flowchart illustrating a computer-implemented method for marking objects on a 3D image according to the present invention is shown; Figure 2 An example of two 2D slices of a 3D image with a cross-section of a 3D geometric object is shown; Figure 3 A schematic diagram of a method for marking at least one object according to the present invention is shown; Figure 4 A method for training a neural network according to the present invention is shown; Figure 5 The PERT distribution, which can be used to generate training data for the neural network according to the present invention, is shown.
[0035] Figure Labels 4. Trained Neural Network 5 objects 6. Cross-section of 3D geometry 7 3D Region of Interest 8. Information corresponding to the 3D geometric shape 9 Marked 3D Volume 10 2D slices 11 First 2D slice 12 Second 2D slice 21 Minimum value 22 Maximum value 23. Central value 24. Average value within the object 25. Average value outside the object 101-106 Method steps for marking objects 201-203 Methods and Steps for Training Neural Networks Detailed Implementation
[0036] In all the accompanying drawings, the same or corresponding features / elements of various embodiments are indicated by the same reference numerals.
[0037] Figure 1A flowchart illustrating a computer-implemented method for marking an object 5 on a 3D image according to the present invention is shown. In a first step 101, a 3D image is provided. The 3D image may preferably be a 3D medical image, such as a computed tomography image of an organ. Notably, the 3D image includes the object 5 to be marked. The 3D image may, for example, be taken from a database of 3D images. Alternatively, in this first step 101, the 3D image may be measured using image modalities. Thus, the object 5 can be marked directly after the 3D image measurement (in the following steps). In a further step 102, a 2D slice 10 of the 3D image and a cross-section 6 of a 3D geometry located within the 3D image are displayed to the user, for example, via a computer screen. The 3D geometry may preferably be a sphere (i.e., a ball), and correspondingly, the cross-section 6 may be circular. In a further step 103, a unit is provided to the user to select the position of the 3D geometry on the 2D slice 10 such that the 3D geometry at least partially overlaps with the object 5. Therefore, the user can attempt to mark the position of the object by this selection, specifically by moving the displayed cross-section 6 to overlap with the object 5, for example, by moving the cross-section 6 relative to the 2D slice 10 with a computer mouse. The user can then input a command, for example, by left-clicking the computer mouse, indicating that the current position will be the selected position. When the user selects the position by moving the cross-section 6 relative to the 2D slice 10 with the computer mouse and left-clicking when the desired position is reached, the position information is received by an image processing device (such as a computer). Optionally, in this step 103, particularly before position selection, the user can be provided with a unit to change the size of the 3D geometry, for example, by scrolling using the computer mouse wheel. Furthermore, the user can be given the opportunity to switch the displayed view to another 2D slice 10 of the 3D image and select a position in the other displayed 2D slice 10. In a further step 104, a 3D region of interest 7 is automatically defined such that the 3D region of interest 7 includes the 3D geometry at its current position, and image data from the 3D region of interest is cropped, i.e., image data is acquired for further processing. The 3D region of interest can preferably be a cuboid shape, particularly a cube shape. In a further step 105, the trained algorithm is applied to the cropped image data. The trained algorithm can be, in particular, a trained neural network 4, preferably a U-Net-based network or an F-Net-based network. If the user can choose to adjust the size of the 3D geometry, one of several trained algorithms can be applied. These several different trained algorithms can include different trained algorithms, each trained for another selectable size of the 3D geometry. The algorithm is trained such that its output is a 3D volume 9 within the 3D region of interest 7 predicted to cover the object 5. According to an embodiment, the 3D volume can be selected to lie within the 3D geometry.Alternatively, a 3D volume can be selected such that it does not exceed a predefined maximum amount beyond the 3D geometry. That is, there may be a specific overlap that ensures the location selected by the user system (i.e., a small portion of the boundary of object 5) is automatically corrected. To determine the 3D volume 9 within the 3D region of interest 7, a 3D Gaussian kernel whose maximum value is at the center point of cross-section 6 can be fed into a trained algorithm along with image data from the 3D region of interest 7. In this case, the trained algorithm can also determine the 3D volume based on the 3D Gaussian kernel. In a further step 106, the labeled 3D volume 9 is output. For example, the 3D volume can be displayed along with a 3D image or a 2D slice 10 of a 3D image. In particular, the displayed 2D slice 10 can be upgraded so that it now labels object 5 or a portion of that object corresponding to the 3D volume 9. For example, the 3D volume 9 can be labeled on the 2D slice 10 via a circle displayed to the user on the screen. Steps 103 to 105 and optionally 106 may be repeated at least once, such that at each repetition an additional 3D volume 9 is generated depending on additional user input during step 103, and wherein each additional 3D volume 9 is joined with the currently existing 3D volume 9 to produce a larger existing 3D volume 9. Optionally, the currently existing 3D volume 9 may be displayed to the user and updated after each repetition, particularly during step 106, in which case step 106 may also be repeated a corresponding number of times. For example, the user may switch to different 2D slices of the 3D image for each repetition.
[0038] Figure 2 An example of two 2D slices of a 3D image is shown, with a first 2D slice 11 on the left and a second 2D slice 12 on the right. The 3D image includes object 5. Furthermore, in both 2D slices 11 and 12, a corresponding cross-section 6 of the circular shape of a 3D geometry (here, a sphere or ball) is shown. The left 2D slice 11 is the slice currently displayed to the user when the user selects the position of the 3D geometry. It can be seen that the user has positioned the 3D geometry such that the cross-section 6 is placed directly at the lower right boundary of object 5. However, due to the rigid shape of the 3D geometry, the boundary of the 3D geometry does not match the boundary of object 5 in the second 2D slice 12 on the right. Instead, the cross-section 6 of the 3D geometry extends beyond object 6 on the second 2D slice 12, causing part of the surrounding background to be covered by the cross-section 6. If a marker directly corresponding to the selected position of the 3D geometry is applied, the result will therefore be quite inaccurate. Therefore, manual correction is required, which takes additional time from the user. This problem is addressed below. Figure 3 The method of the present invention, detailed in the text, solves this problem.
[0039] Figure 3A schematic diagram of a method for labeling at least one object 5 according to the present invention is shown. A 2D slice 10 of a 3D image including the object 5 (in this example, a human kidney) is displayed to the user along with a cross-section 6 of a 3D geometry. The user can then move the cross-section 6 such that it at least partially overlaps with the object, and, for example, forward his / her selection at the current location by clicking a computer mouse. Upon receiving the user's selection, a 3D region of interest 7, shown here by dashed lines, is automatically generated. The 3D region of interest 7 is defined such that its center point corresponds to the center point of the 3D geometry. In this example, it can be seen that the center point of the cross-section 6 corresponds to the center point of the rectangular cross-section of the 3D region of interest 7. The 3D region of interest is larger than the 3D geometry, preferably by a predetermined amount. In this example, the side length of the 3D region of interest 7 is approximately 2.6 times the diameter of the cross-section 6 of the 3D geometry. The image data of the 3D region of interest 7 is then cropped and fed to a trained algorithm, in this case, a relatively simple trained neural network 4 with a shallow network architecture. Along with the image data of the 3D region of interest (ROI), information corresponding to the 3D geometry 8 (in this example, a click map indicating the position and size of the 3D geometric object relative to the 3D ROI 7) is fed into a trained neural network 4. The trained neural network 4 may preferably consist only of real convolutional layers, excluding filled convolutional layers. The neural network 4 is trained to segment objects within the ROI and output a predicted region as a 3D volume 9. The predicted 3D volume 9 is then transmitted back to the image domain and displayed on a 2D slice 10 as a labeled cross-section of the 3D volume 9. It can be seen that although the shape of the 3D geometry selected by the user does indeed extend beyond the boundary of the object 5, the 3D volume 9 is automatically corrected by the trained neural network 4 and is cut off at the boundary of the object 5; that is, only the image data belonging to the object is within the 3D volume 9. Therefore, in this embodiment, the cross-section 6 indicates the largest area to be labeled via the 3D volume 9, but this area is confined within the object. Surprisingly, this automatic marker correction process can be performed with good robustness in near real-time, meaning that user input can be updated with markers almost instantaneously, effectively allowing users to "draw" automatically corrected markers, for example, by repeatedly clicking the mouse or even holding down the mouse button and moving the mouse.
[0040] Figure 4A method for training a neural network according to the present invention is shown to provide a trained neural network 4 for use in a method for labeling at least one object 5. The training method includes a first step 201, namely, providing a set of 3D regions of interest 7 from a 3D image as input training data. Each 3D region of interest 7 includes at least a portion of the object 5 on the 3D image. Preferably, each provided 3D region of interest 7 is selected such that the distance between its center point and the boundary of the object 5 has a specified value. This specified value can be a fixed value, or it can be a fixed value with added random deviation within a predefined maximum deviation amount. In other words, a specified amount of noise can be added to the location of the 3D region of interest 7. Preferably, the maximum deviation amount is predefined relative to the size of the 3D geometry to be used after training during the method for labeling at least one object. To simulate more accurate user behavior when using larger 3D geometries, it is anticipated that a larger size of the 3D geometry can correspond to a larger maximum deviation amount. In a further step 202, a 3D dataset is provided as output training data. This 3D dataset has the size of a 3D region of interest 7 and contains information for each voxel, which is either labeled as a portion of an object or does not correspond to the object for each provided 3D region of interest I. Multiple corresponding input and output training datasets can be provided for expected 3D geometries of different sizes, corresponding to different specified values of the distance between their center point and the boundary of the object 5. Multiple training datasets can be used to train multiple neural networks, such that each neural network is trained for different specified sizes and specific output sizes of the 3D geometry. In a further step 203, the neural networks are trained using the input training data and the output training data. Optionally, multiple neural networks can be trained, such that each neural network is trained for different specified sizes and specific output sizes of the 3D geometry.
[0041] Figure 5The PERT distribution according to the invention can be used to generate training data for a neural network. Specifically, the PERT distribution can be used to determine the intensity / window setting of an extracted 3D region of interest 7, which serves as input training data to simulate user behavior when setting a window / level. Specifically, the lower / upper intensity values of the window can be extracted from the PERT distribution. The PERT distribution is defined by a minimum value 21, a maximum value 22, an average value 24 inside the object, an average value 25 outside the object, and a center value 23 between the averages 24 and 25. The averages 24 and 25 correspond to the most probable values of the PERT distribution. Specifically, the PERT distribution can be viewed as two partial PERT distributions, namely a left-side partial PERT distribution and a right-side partial PERT distribution, which share a common center value 23, wherein the center value 23 is specifically the maximum value of the left-side partial PERT distribution and the minimum value of the right-side partial PERT distribution. For example, the lower limit of the window can be derived from the PERT distribution given by the minimum intensity value 21 (minimum) in the image, the internal mean 24 (i.e., the most likely value) of the structure to be segmented, and the center value 23 (therefore, the maximum value of the partial PERT distribution on the left) between the internal mean intensity value 24 and the external mean intensity value 25. The upper limit of the window can be derived from the PERT distribution given by the maximum intensity value 22 (maximum) in the image, the external mean (i.e., the most likely value) outside the object, and the center value 23 (therefore, the minimum value of the partial PERT distribution on the right) between the internal mean intensity value 24 and the external mean intensity value 25.
[0042] The above discussion is intended to illustrate the system only and should not be construed as limiting the claims to any particular embodiment or group of embodiments. Therefore, the specification and drawings should be considered illustrative and not intended to limit the scope of the claims. By studying the drawings, disclosure, and claims, those skilled in the art can understand and implement other variations of the disclosed embodiments in practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the quantifiers "a" or "an" do not exclude multiple. A single processor or other unit can perform the functions of several items recited in the claims. The fact that certain measures are referenced in mutually different dependent claims does not mean that a combination of these measures cannot be used advantageously. Computer programs can be stored / distributed on suitable media, such as optical storage media or solid-state media provided with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. Any reference numerals in the claims should not be construed as limiting the scope of the claims.
Claims
1. A computer-implemented method for marking at least one object (5) on 3D images, particularly medical 3D image data, comprising the following steps: (a) Provide a 3D image including at least one object (5) to be labeled, wherein the 3D image can be represented by a plurality of 2D slices (10); (b) Displaying at least one 2D slice (10) of the 3D image to the user and locating a 3D geometry within the 3D image, wherein a cross section (6) of the 3D geometry is displayed on the 2D slice (10); (c) Provide the user with a unit to select the position of the 3D geometry on the 2D slice (10) such that the 3D geometry at least partially overlaps with the object (5) to be marked, and receive position information of the 3D geometry selected by the user; (d) Automatically define a 3D region of interest (7) within the 3D image that includes the 3D geometry at the selected location, and crop image data from the 3D image from the 3D region of interest (7); (e) The trained algorithm is applied to the image data from the 3D region of interest (7), wherein the output of the trained algorithm is a 3D volume (9) within the 3D region of interest (7) predicted to cover the object (5), wherein steps (c) to (e) are repeated at least once, such that an additional 3D volume (9) depending on additional user input is generated at each repetition, and wherein each additional 3D volume (9) is joined together with the currently existing 3D volume (9) to produce a larger existing 3D volume (9). (f) Output the labeled 3D volume (9).
2. The computer-implemented method according to claim 1, in, The 3D geometric shape is an ellipsoid, preferably a sphere.
3. The computer-implemented method according to any one of the preceding claims, in, The 3D region of interest (7) is defined such that its center point corresponds to the center point of the 3D geometry.
4. The computer-implemented method according to any one of the preceding claims, in, The 3D region of interest (7) is defined as being a predetermined amount larger than the 3D geometry.
5. The computer-implemented method according to any one of the preceding claims, in, In order to determine the 3D volume (9) within the 3D region of interest (7), a 3D Gaussian kernel whose maximum value is at the center point of the cross section (6) is fed together with the image data from the 3D region of interest (7) into the trained algorithm, in particular into the neural network, and wherein the trained algorithm also determines the 3D volume (9) based on the 3D Gaussian kernel.
6. The computer-implemented method according to any one of the preceding claims, in, The 3D region of interest (7) is a cuboid shape, especially a cube shape.
7. The computer-implemented method according to any one of the preceding claims, in, The user is provided with the unit to change the size of the 3D geometry.
8. The computer-implemented method according to claim 7, in, For each dimension of the 3D geometry, different trained algorithms trained for that specific dimension are applied, in particular different trained neural networks (4).
9. The computer-implemented method according to any one of the preceding claims, in, The trained algorithm is a trained neural network (4), particularly a semantic network, preferably a U-Net-based network or an F-Net-based network.
10. The computer-implemented method according to any one of the preceding claims, in, The current existing 3D volume (9) is displayed to the user and updated after each repetition.
11. The computer-implemented method according to any one of the preceding claims, in, The user is given the opportunity to switch the displayed view to another 2D slice (10) of the 3D image and select a location within the currently displayed 2D slice (10). The method steps following the user's selection are performed based on the user's selection on the currently displayed slice.
12. An image processing apparatus configured to perform the following steps, particularly the steps of the method according to any one of the preceding claims: (a) Receive a 3D image including at least one object (5) to be labeled, wherein, The 3D image can be represented by multiple 2D slices (10); (b) Display at least one slice of the 3D image to the user; (c) Provide the user with a unit to select the position of the 3D geometry on the 2D slice (10) such that the 3D geometry at least partially overlaps with the object (5) to be labeled; (d) In response to the user selection, automatically define a 3D region of interest (7) within the 3D image that includes the 3D geometry at the selected location, and crop image data from the 3D image from the 3D region of interest (7); (e) The trained algorithm is applied to the image data from the 3D region of interest (7), wherein the output of the trained algorithm is a 3D volume (9) within the 3D region of interest (7) predicted to cover the object (5), wherein steps (c) to (e) are repeated at least once, such that an additional 3D volume (9) depending on additional user input is generated at each repetition, and wherein each additional 3D volume (9) is joined together with the currently existing 3D volume (9) to produce a larger existing 3D volume (9). (f) Output the 2D slice (10) together with the labeled 3D volume (9).
13. A method for training a neural network, particularly a neural network according to any one of claims 1-11, comprising the following steps: (a) Provide a set of 3D regions of interest from a 3D image as input training data, wherein each 3D region of interest (7) includes at least a portion of an object (5) on the 3D image; (b) Provide a 3D dataset as output training data, the 3D dataset having the size of the 3D region of interest (7) and containing information for each voxel, the voxel being labeled as a part of the object (5) or not corresponding to the object (5) for each provided 3D region of interest (7). (c) The neural network is trained using the input training data and the output training data.
14. The method according to claim 13, in, Each 3D region of interest (7) is selected such that the distance between its center point and the boundary of the object (5) has a specified value. The specified value is either a fixed value or a fixed value with an added random deviation within a predefined maximum deviation.
15. The method according to any one of claims 13 to 15, in, The intensity window of the 3D region of interest (7) is randomly determined based on at least one PERT distribution. The at least one PERT distribution is specifically based on the minimum data value (21) / maximum data value (22) of the corresponding 3D image, the average data value (24) inside the object (5) and the average data value (25) outside the object (5), and the center (23) between the average data values.