Processing 3d images
By applying object detection technology and machine learning methods on 2D slices of 3D medical images, predicting the desired plane and the location of the anatomical marks of the target anatomy structure, the problems of information overload and analysis inconvenience in 3D medical images are solved, and the images of expected views are automatically generated are improved, improving the accuracy and efficiency of analysis.
Patent Information
- Application Number
- CN202380075361.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-26
- Filing Date
- 2023-10-23
- Publication Date
- 2025-06-06
AI Technical Summary
The large amount of information and details contained in 3D medical images may affect the ease of analysis for clinicians, and standardized testing often requires the use of specific 2D image views.
By obtaining 2D slices of 3D images, the boundary area of the target anatomy is identified using object detection techniques and a set of initial positions is defined within that area. The 3D image is then processed using a machine learning method to predict the first predicted offset and the second predicted offset of each initial position for determining the location of the desired plane and anatomical marker.
An image of an automated supply of anatomical structures without the need for movement or additional input from a clinician or image capture system operators, improving the accuracy and efficiency of 3D image analysis.
Smart Images

Figure CN120112948A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D images, and in particular to the processing of 3D images. Background Art
[0002] In the medical field, there is increasing interest in using 3D images (i.e., 3D medical images) to evaluate and / or analyze target anatomical structures. Specifically, 3D images have been shown to provide additional contextual information to assist in analyzing target anatomical structures.
[0003] However, the large amount of information and details contained in the 3D image may affect the ease of analysis by the clinician. In addition, many standardized tests for analyzing target anatomical structures require the use of 2D images containing specific views of the target anatomical structure and / or specific anatomical landmarks within the anatomical structure.
[0004] For example, in fetal screening, abdominal circumference (AC) is a standard measurement for estimating fetal size and growth. Existing screening guidelines define criteria that 2D images must meet to achieve valid and comparable biometric measurements. For example, the stomach and umbilical veins must be visible, while the heart and / or kidneys should not be visible, the scanning plane of the 2D image should be orthogonal to the head-to-toe axis, and the shape of the abdomen should be as round as possible. Summary of the invention
[0005] The invention is defined by the independent claim. The dependent claims define advantageous embodiments.
[0006] According to an example of one aspect of the invention, a computer-implemented method of processing a 3D image of a target anatomical structure is provided.
[0007] The computer-implemented method includes: obtaining a 2D slice of a 3D image; processing the 2D slice using an object detection technique to identify a bounding region containing a representation of the target anatomical structure; defining a set of initial positions within the bounding region; and processing the 3D image using a machine learning method to predict, for each initial position: a set of one or more first prediction offsets, each of which is a predicted spatial offset, relative to the 3D image, between the initial position and a predetermined position of the 3D image within the desired plane; and a set of one or more second prediction offsets, each of which is a predicted spatial offset, relative to the 3D image, between the initial position and a position of an anatomical landmark of the target anatomical structure outside the desired plane.
[0008] The proposed method provides a technique for identifying or defining a spatial offset between a 2D slice and a desired 2D plane or anatomical landmark within a 3D image. For example, the defined spatial offset can be used to generate a 2D image containing a representation of the desired plane or landmark and / or provide guidance on how to move / change an image capture system to capture an image containing the desired plane or landmark.
[0009] The techniques presented herein utilize an initial position defined within a bounding region containing a representation of the target anatomy. This provides a consistent baseline for defining any spatial offsets for easy image generation and / or modification of the image capture system.
[0010] The bounding region may have a predetermined shape, such as a rectangle, a circle or a square.The bounding region defines a predicted region containing the representation of the target anatomical structure.
[0011] Since a set of one or more first prediction offsets is generated for each initial position, respective sets (ie, a plurality of sets) of one or more first prediction offsets are produced.
[0012] The method may also include the steps of processing the set of initial positions, the set of one or more first predicted offsets, and the 3D image to produce at least one image that is smaller than the 3D image, the at least one image containing a representation of one or more of the one or more desired positions. The present disclosure recognizes that the 3D image may contain sufficient information to produce an image (e.g., a 2D image) containing the desired position or view plane. Thus, an automated method is provided to provide an image of a desired view of an anatomical structure without requiring movement or additional input from a clinician or operator of an image capture system.
[0013] In some examples, each initial position is associated with a different corresponding desired position within a desired plane of the 3D image for viewing the target anatomical structure; and for each initial position, the set of one or more first predicted offsets includes a single predicted offset that is a predicted spatial offset between the initial position and its corresponding desired position.
[0014] Optionally, the method further comprises: estimating the position of a desired plane within the 3D image by processing a set of initial positions and a set of first prediction offsets generated for each initial position; and processing the 3D image using the estimated position of the desired plane to generate a 2D image representing the desired plane.
[0015] In this way, a 2D image is generated containing a desired view representation of anatomical landmarks. This is particularly advantageous when one or more 2D views are essential or mandatory for performing a standardized analysis or monitoring technique.
[0016] Estimating the position of the desired plane within the 3D image may be performed by processing the set of initial positions and the set of first prediction offsets generated for each initial position using a linear regression technique.
[0017] Each second predicted offset may be a predicted spatial offset between the initial position and the position of the same first anatomical landmark of the target anatomical structure relative to the 3D image.
[0018] Preferably each set of one or more second prediction offsets comprises a plurality of second prediction offsets.
[0019] In some examples, the step of defining a set of initial positions within the boundary region includes: defining a grid of cells within the boundary region; and identifying a single position in each cell, thereby defining a plurality of initial positions. Preferably, the grid is a regular grid (i.e., cells of equal or nearly equal size). Each cell may be rectangular and / or square. The single position in each cell may be a center point of the cell, or an edge or corner point of the cell.
[0020] The 3D image may be defined by an initial 2D slice stack, and the 2D slice may be closer to a center-most slice in the initial 2D slice stack than a first or last 2D slice in the initial 2D slice stack. For example, the 2D slice may be a center-most slice in the 2D slice stack.
[0021] One recognition of this embodiment is that during the capture of a 3D image by an operator of the image capture system, it can be assumed that the operator will control the image capture system so that the desired plane and / or (one or more) anatomical landmarks are approximately centered with respect to the 3D image. This means that more central slices of the 3D image are more likely to be only partially offset with respect to the desired plane and / or landmark. Therefore, the proposed technique is more accurate when operating under this valid assumption.
[0022] The step of processing the 2D slices using a machine learning method may further include generating a confidence value for each predicted offset, the confidence value representing the confidence or certainty of the predicted offset.
[0023] In some examples, the method further includes estimating the position of the desired plane within the 3D image by processing the set of initial positions and the set of first predicted offsets generated for each initial position using a linear regression technique, wherein the values of one or more weighting parameters of the linear regression technique are responsive to the confidence values of the one or more predicted offsets.
[0024] A computer-implemented method is also presented for training a machine learning method to predict the location of a desired plane in a desired 3D image.
[0025] The computer-implemented method includes: acquiring a training data set including a plurality of training items, each training item including: an example 3D image of a target anatomical structure; first information identifying a location of a desired plane in the example 3D image, and second information identifying locations of one or more anatomical landmarks outside the desired plane in the example 3D image; and using the training data set to train a machine learning method to predict information identifying the location of the desired plane in the desired 3D image and information identifying the locations of one or more 3D landmarks outside the desired plane in the desired 3D image.
[0026] Also proposed is a computer program product comprising computer program code means which, when executed on a computing device having a processing system, causes the processing system to perform all the steps of any method described herein.
[0027] A processing system configured to perform any of the methods described herein is also presented.
[0028] Therefore, a processing system for processing a 3D image of a target anatomical structure is proposed, the processing system being configured to: acquire 2D slices of the 3D image; process the 2D slices using object detection techniques to identify a boundary region containing a representation of the target anatomical structure; define a set of initial positions within the boundary region; and process the 3D image using a machine learning method to predict, for each initial position: a set of one or more first prediction offsets, each first prediction offset being a predicted spatial offset, relative to the 3D image, between the initial position and a predetermined position of the 3D image within a desired plane; and a set of one or more second prediction offsets, each second prediction offset being a predicted spatial offset, relative to the 3D image, between the initial position and a position of an anatomical landmark of the target anatomical structure outside the desired plane.
[0029] In some examples, the processing system is further configured to process the set of initial positions, at least one set of one or more predicted offsets, and the 3D image to produce at least one image smaller than the 3D image, the at least one image containing representations of one or more of the one or more desired positions.
[0030] A processing system for training a machine learning method is proposed for predicting the position of a desired plane in a desired 3D image, the processing system being configured to: acquire a training data set comprising a plurality of training entries, each training entry comprising: an example 3D image of a target anatomical structure; first information identifying the position of the desired plane in the example 3D image, and second information identifying the positions of one or more anatomical landmarks outside the desired plane in the example 3D image; and use the training data set to train the machine learning method to predict information identifying the position of the desired plane in the desired 3D image and information identifying the positions of one or more 3D landmarks outside the desired plane in the desired 3D image.
[0031] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] For a better understanding of the invention, and to show more clearly how it may be put into practice, reference will now be made, by way of example only, to the accompanying drawings, in which:
[0033] Figure 1 is a flow chart illustrating a method according to an embodiment;
[0034] Figure 2 illustrates the relationship between different locations within a 3D image;
[0035] Figure 3 illustrates the architecture of a machine learning algorithm for an embodiment;
[0036] Figure 4 is a flowchart illustrating a method according to another embodiment; and
[0037] Figure 5 is a flow chart illustrating a method according to another embodiment;
[0038] Figure 6 is a schematic diagram of a processor according to an embodiment. DETAILED DESCRIPTION
[0039] The present invention will be described with reference to the accompanying drawings.
[0040] It should be understood that the detailed description and specific examples, although indicating exemplary embodiments of the present invention, are for illustrative purposes only and are not intended to limit the scope of the present invention. These and other features, aspects and advantages of the present invention will be better understood from the following description, the appended claims and the accompanying drawings. It should be understood that the drawings are schematic only and are not drawn to scale. It should also be understood that throughout the drawings, the same reference numerals are used to represent the same or similar parts.
[0041] The present invention provides a mechanism for determining a desired plane position within a 3D image. A machine learning algorithm is used to determine a first set of one or more offsets between an initial position within a 2D slice of the 3D image and a predetermined position within the desired plane, and a second set of one or more offsets between the same initial position and the position of an anatomical landmark. The offset is a spatial offset between an initial position and another position.
[0042] The present disclosure recognizes that a machine learning algorithm that has been trained to simultaneously identify an offset to a desired plane and an offset to at least one anatomical landmark can more accurately identify the offset to the desired plane. This identification can be used to improve the accuracy of identification of the desired plane in a 3D image, such as generating a 2D image in the desired plane.
[0043] Embodiments may be employed in any scenario where identification of 2D planes of a 3D medical image is required. One example environment is fetal monitoring and / or analysis.
[0044] Figure 1 is a flow chart illustrating the overall method of 3D image processing adopted by the present invention. Figure 1 Thus, a computer-implemented method 100 of processing a 3D image of a target anatomical structure is illustrated.
[0045] The method 100 includes a step 110 of obtaining a 2D slice of a 3D image. A 2D slice is a planar selection of a portion of a 3D image. Specifically, a 2D slice may be a slice perpendicular to an axis of the 3D image. Thus, if the 3D image has dimensions X×Y×Z, then the 2D slice has dimensions X×Y.
[0046] The obtained 2D slice is preferably a slice that is closer to the most central slice of the 3D image than the end slices. Thus, if the 3D image is defined by an initial 2D slice stack, the obtained 2D slice is closer to the most central slice of the initial 2D slice stack than the first or last 2D slice in the initial 2D slice stack. More preferably, the obtained 2D slice is the most central slice of the 3D image.
[0047] The method 100 further comprises a step 120 of processing the 2D slice using an object detection technique to identify a boundary region containing a representation of the target anatomical structure (ie, a 2D image containing the target anatomical structure).
[0048] The bounding region may have a certain shape known in the art (eg, rectangular or circular). When having a rectangular shape, the bounding region is usually marked as a bounding box.
[0049] 2D object detection is a well-analyzed problem in computer vision, especially for generating boundary regions of objects or structures that are expected to be identified. Therefore, there are multiple methods to perform step 120. A suitable example is the "You Only Look Once" (YOLO) algorithm, as illustrated by Joseph Redmon, Ali Farhadi: YOLO9000: Better, Faster, Stronger, CVPR 2018. Other suitable methods are proposed by Wang, Guotai et al.: "Interactive medical image segmentation using deep learning with image-specific fine tuning." IEEE transactions on medical imaging 37.7 (2018): 1562-1573.
[0050] The size of the boundary region is smaller than the size of the 2D slice. Specifically, the boundary region is sized and positioned to attempt to define the boundaries of the target anatomical structure representation in the 2D slice.
[0051] The method 100 further comprises a step 130 of defining a set of initial positions within the boundary area.
[0052] For example, step 130 may be performed to define a grid of cells within the boundary region and identify a single position in each cell, thereby defining a plurality of initial positions. The grid may be a regular grid (i.e., cells having equal or nearly equal sizes). Each cell may be a rectangle and / or a square, for example, depending on the size and / or shape of the boundary region. The initial position may be defined as being located at a predetermined or defined position relative to each cell, for example, at a center point of the cell or at a predetermined or predefined corner of the cell.
[0053] If used, the grid of cells may be of size N x M. Preferably, N=M.
[0054] The method 100 further performs step 140 of processing the 3D image using a machine learning method to predict, for each initial position, a set of one or more first prediction offsets and a second one of the one or more second prediction offsets.
[0055] The prediction offset is the predicted spatial offset between an initial position and another position.
[0056] The spatial offset defines the distance and direction between two points in 3D space, i.e., relative to the 3D image. The spatial offset may, for example, be in the form of a vector identifying the direction and distance between two relative locations (e.g., between an initial location and a predetermined location, or between an initial location and the location of an anatomical landmark). For example, the vector may be formatted in Cartesian form or in polar coordinate form. Other suitable data formats for defining the spatial offset between two points in a 3D image will be apparent to those skilled in the art.
[0057] Each first predicted offset is a predicted spatial offset relative to the 3D image between the initial position and a predetermined position of the 3D image within a desired plane.
[0058] The desired plane may represent a plane in the 3D image that contains a representation of a desired view and / or cross section of the target anatomical structure. For example, if the target anatomical structure is a fetus, the desired plane may be a plane that depicts the stomach and umbilical cord orthogonal to the head-foot axis.
[0059] Each second predicted offset is a predicted spatial offset, relative to the 3D image, between the initial position and a position of an anatomical landmark of the target anatomical structure that is outside of the desired plane.
[0060] In some embodiments, each set of one or more second prediction offsets includes multiple second prediction offsets, wherein each second prediction offset in the same set is a predicted spatial offset between an initial position relative to the 3D image and a position of the same anatomical landmark outside a desired plane of the target anatomical structure.
[0061] Thus, any given set of second predicted offsets may include a plurality of second predicted offsets between the same initial position and the position of the same anatomical landmark.Different sets are associated with different initial positions.
[0062] Examples of anatomical landmarks may depend on the type of target anatomical structure. For example, if the target anatomical structure is a fetus, then the anatomical landmarks may be the eyes of the fetus, the fingers or toes of the fetus, and / or the base of the fetal spine. As another example, if the target anatomical structure is a heart, then the anatomical landmarks may be the entrances to specific vessels or chambers of the heart. Other suitable examples will be apparent to those skilled in the art.
[0063] Thus, the machine learning method is a machine learning method or technique that is trained to perform simultaneous identification of the relative position of the desired plane and the relative position of (one or more) anatomical landmarks. It has been recognized that a machine learning method that has been trained to perform both of these features simultaneously has improved performance with respect to the desired plane information. This means that the performance of identifying the relative position of the desired plane (i.e., determining the set of first offsets) is improved.
[0064] In a preferred example, each initial position is associated with a different corresponding desired position within a desired plane of the 3D image for viewing the target anatomical structure. For each initial position, the set of one or more first prediction offsets includes a single prediction offset that is a predicted spatial offset between the initial position and its corresponding desired position.
[0065] Each desired position may effectively represent a particular portion or region of a desired plane of the 3D image.
[0066] To improve the performance of step 140, the input to the machine learning method for performing step 140 may be a portion or part of the 3D image defined by the geometry of the boundary region. The geometry of the boundary region defines the size, shape and position of the boundary region. Therefore, the physical position of the voxel value input to the machine learning method for performing step 140 may depend on the boundary region.
[0067] Figure 2 A visual representation of the process performed through steps 110 - 140 is provided.
[0068] Figure 2 Illustrated is a 2D slice 210 for which a boundary region 220 has been identified by processing the 2D slice using a suitable object detection technique. The size of the boundary region 220 is smaller than the size of the 2D slice.
[0069] The boundary region 220 is used to define a plurality of initial positions 221, 222. In the illustrated example, each initial position 221, 222 is located at a center point of a cell within a regular cell grid.
[0070] The relative positions of desired plane 230 and anatomical landmarks 240 are also conceptually illustrated.
[0071] The purpose of step 140 is to predict or determine one or more first offsets 251, 252 between each initial position 221, 222 and a predetermined position (for that initial position) within the desired plane 230, and one or more second offsets 261, 262 between each initial position 221, 222 and (one or more) anatomical landmarks.
[0072] like Figure 2 As illustrated in , each initial position may be associated with only a single first offset 251, which defines an offset between the initial position 221 and a corresponding predetermined position 231 within the desired plane 230. Therefore, each initial position 221 maps or corresponds to a predetermined position 231 within the desired plane 230.
[0073] Obviously, for clarity of explanation, only a subset of all possible initial positions, first offsets and second offsets are illustrated and / or labeled. A person skilled in the art will readily appreciate that in practice there may be more positions and / or offsets.
[0074] The proposed embodiment utilizes a machine learning algorithm to process the 3D image to generate a set of first and second offsets. A machine learning algorithm is any self-training algorithm that processes input data to generate or predict output data. Here, the input data includes the 3D image (or a portion thereof), and the output data includes a set of first and second offsets.
[0075] Suitable machine learning algorithms for the present invention will be apparent to the skilled person. Examples of suitable machine learning algorithms include decision tree algorithms and artificial neural networks. Other machine learning algorithms, such as logistic regression, support vector machines, or naive Bayes models are suitable alternatives.
[0076] However, one particularly advantageous form of machine learning algorithm for generating the first and second offsets (i.e., for use in step 140) is a neural network. Based on current understanding, neural networks appear to be by far the best choice for performing difficult learning tasks while being able to meet nearly real-time performance.
[0077] One particularly advantageous approach is to use multi-resolution neural networks as machine learning algorithms. Multi-resolution neural networks have been shown to achieve state-of-the-art accuracy in several segmentation tasks in medical imaging.
[0078] Standard multi-resolution neural networks use the concept of combining image patches of different resolutions for pixel-by-pixel classification. At each resolution level, standard convolutional layers are used to extract image features. The feature maps of the coarse level are then successively upsampled and combined with the next finer level. In this way, a segmentation at the original image resolution is obtained while having a large receptive field to consider a wide range of image context.
[0079] For the purposes of this disclosure, this general idea is used to generate or predict offsets. Thus, if used, the multi-resolution neural network has been trained and / or configured such that an output layer of size N×M×P holds the offsets, where P is the number of offsets generated for each initial position (and N×M equals the number of initial positions).
[0080] It should be appreciated that, by design, the output layer containing the offsets has a much smaller resolution than the input layer of the multi-resolution neural network. Therefore, the standard multi-resolution neural network (described above) should be modified to utilize a downsampling strategy, rather than an upsampling strategy. This means that instead of upsampling a coarse level and combining it with a finer level as in the architecture described above, the finer levels are sequentially downsampled and combined with a coarser level.
[0081] Methods for training machine learning algorithms are well known. Typically, such methods include obtaining a training data set, which includes training input data entries and corresponding training output data entries. This is often referred to as "ground truth data." An initialized machine learning algorithm is applied to each input data entry to generate a predicted output data entry. The error (e.g., mean square error) between the predicted output data entry and the corresponding training output data entry is used to modify the machine learning algorithm. The process can be repeated until the error converges and the predicted output data entry is sufficiently similar to the training output data entry (e.g., ±1%). This is often referred to as a supervised learning technique.
[0082] Figure 3 The architecture of a multi-resolution neural network 300 that may be employed in embodiments of the present invention is illustrated.
[0083] Neural network 300 generates output data 350 that includes a set of one or more first offsets for each initial position and a set of one or more second offsets for each initial position.
[0084] Where appropriate, blocks labeled "C" represent convolution processes performed on the data (e.g., a process that includes convolution, batch normalization, and ReLU (Rectified Linear Unit), also known as CBR (Convolution-Batch Normalization-ReLU)). Similarly, blocks labeled "P" represent pooling processes performed on the data. Blocks labeled D represent downsampling processes, which themselves can be pooling processes, such as maximum or average pooling processes.
[0085] Image data 311 is processed to produce output data. Image data 311 includes 3D image data and / or portions thereof. The image data is successively downsampled to produce downsampled image data 312, 313, 314 of smaller resolution. At each resolution level, image features are extracted using one or more standard convolutional layers (in a convolution process C). The feature maps at finer levels are then successively downsampled (e.g., using a pooling process P) and combined (e.g., summed) with feature maps at the next coarser level.
[0086] Figure 4 A method 400 according to an embodiment is shown which utilizes the generated offset.
[0087] The method 400 includes performing the previously described method 100 to generate an offset.
[0088] The method 400 further includes step 410 of estimating a position of the desired plane within the 3D image by processing the set of initial positions and the set of one or more first prediction offsets generated for each initial position.
[0089] For example, step 410 may be performed by processing the set of initial positions and the set of first prediction offsets generated for each initial position using a linear regression technique to identify the position / location of the desired plane within the 3D image.
[0090] For example, suppose each initial position in a 2D slice (in the xy plane) can be represented by the value g xy Indicates that each initial position can be offset from the first xy Associated. First offset op xy represents the determined offset between the initial position and the corresponding predetermined position in the desired plane p. By using linear regression to xy +op xy (Among them, op xy :=(0,0,op xy ) T ) to interpolate and locate the desired plane in the 3D image.
[0091] It is worth noting that for each initial position, the corresponding predetermined position is the desired plane and the direction perpendicular to the 2D slice plane passing through the initial position g xy Therefore, a single offset op xy to estimate the offset from the initial position to the desired plane.
[0092] Schmidt-Richberg, Alexander et al. "Offset regression networks for viewplane estimation in 3D fetal ultrasound." Medical Imaging 2019: ImageProcessing. Vol. 10949. SPIE, 2019, discloses a method for determining the position of a desired plane by linear regression. The same document also discloses a method for determining (one or more) offsets.
[0093] In this case, each initial position is associated with a (single) different corresponding desired position within a desired plane of the 3D image for viewing the target anatomical structure. Thus, for each initial position, the set of one or more first prediction offsets contains (only) a single prediction offset, i.e. the predicted spatial offset between the initial position and its respective desired position.
[0094] Then, method 400 may proceed to step 420 of processing the 3D image using the estimated position of the desired plane to generate a 2D image representing the desired plane. This may effectively include extracting a slice of the 3D image to form a 2D plane.
[0095] It should be appreciated that it is not necessary to perform step 420. Instead, the determined position of the desired plane in the 3D image may be utilized for other purposes.
[0096] For example, the determined position may be used to guide a personal mobile image capture system to more accurately capture a desired plane.
[0097] In another use case scenario, the determined positions may be used to initialize a deformable segmentation model to perform segmentation of a 3D image. Thus, the method 400 may include step 430 of processing the 3D image using an image segmentation technique, initializing the segmentation using the determined positions of the desired planes.
[0098] It will be appreciated that a set of one or more second offsets may be used to a similar effect to identify the location of anatomical landmark(s). For example, this may be used to generate one or more images containing representations of the landmark(s).
[0099] Figure 5 This approach is conceptually illustrated, showing a method according to an embodiment.
[0100] The method 500 also includes performing the previously described method 100 to generate the offsets. In the method, each set of one or more second offsets generated for each initial position contains the offsets between the corresponding initial position and the same first anatomical landmark. Thus, each set of one or more second offsets represents an offset to the same anatomical landmark.
[0101] The method 500 may then include step 510 of processing each set of one or more second offsets to estimate or identify the location of the anatomical landmark 1.
[0102] For example, suppose each initial position in a 2D slice (in the xy plane) can be represented by the value g xy Each initial position can be offset from the second lxy (representing the spatial offset from the first anatomical landmark l). For example, g can be calculated across all initial positions.xy +o lxy The position of the first anatomical landmark in the 3D image may be determined by using an average or median value of the values of the first anatomical landmark. Other interpolation processes (rather than determining the mode or median) may be used to predict the position of the first anatomical landmark.
[0103] The method 500 may then perform step 520 of processing the 3D image using the determined location of the first anatomical landmark to produce an output image (smaller than the 3D image) containing a representation of the first anatomical landmark. For example, this may include selecting a slice of the 3D image containing the determined location of the first anatomical landmark. Alternatively, this may include extracting a smaller 3D image from the 3D image containing the determined location of the first anatomical landmark.
[0104] Similarly, it should be understood that the performance of step 520 is not required. Rather, the determined locations of anatomical landmarks may be used for other purposes.
[0105] For example, the determined position may be used to guide a personal mobile image capture system to more accurately capture anatomical landmarks.
[0106] In another use case scenario, the determined positions may be used to initialize a deformable segmentation model to perform segmentation of the 3D image. Thus, the method 500 may include step 530 of processing the 3D image using an image segmentation technique, initializing the segmentation using the determined positions of the anatomical landmarks.
[0107] As previously described, in some embodiments, each set of one or more second prediction offsets includes multiple second prediction offsets, wherein each second prediction offset in the same set is a predicted spatial offset between an initial position relative to the 3D image and a position of the same anatomical landmark outside a desired plane of the target anatomical structure.
[0108] The same averaging procedure may be employed to determine the location of the first anatomical landmark in the 3D image.
[0109] It should be appreciated that the proposed techniques may be employed to generate images for each of a plurality of anatomical landmarks.
[0110] Specifically, the machine learning method can be configured to generate a set of one or more predicted offsets for each initial position and each anatomical landmark, each offset in the same set representing an offset between the initial position and the additional anatomical landmark. For each anatomical landmark, the set of predicted offsets associated therewith can then be processed to determine or predict the position of the anatomical landmark within the 3D image, for example, using the previously described method.
[0111] The determined positions may be used to similar effect as determining a single landmark, for example, to initialize an image segmentation process or to guide an operator in generating new image data that more completely captures the landmark(s).
[0112] Of course, in some embodiments, method 400 and method 500 may be performed simultaneously.
[0113] At least for reference Figure 1 In the method described previously, the machine learning algorithm outputs a plurality of offsets, including at least a set of one or more first offsets and a set of one or more second offsets for each initial position.
[0114] In some examples, the step 140 of processing the 2D slices using a machine learning method further includes generating a confidence value for each predicted offset, the confidence value representing the confidence or certainty of the predicted offset.
[0115] Methods for determining confidence values for predicted data using machine learning methods are well known in the art. Example methods are presented by: Kendall, A., Gal, Y., 2017 in Advances in Neural Information Processing Systems: What uncertainties do we need in bayesian deep learning for computer vision? , and van der Waa, Jasper, et al. "ICM: an intuitive model independent and accurate certainty measure for machine learning." ICAART (2). 2018.
[0116] If the predicted offset(s) are used to determine the position of the desired plane and / or the position of the desired anatomical feature in the 3D image, the confidence value(s) may be used in the determination process. Specifically, the confidence value may be used to weight the value of each offset used to determine the position(s).
[0117] Therefore, for each offset, a confidence value can be estimated to express the confidence of the machine learning algorithm in its estimate. These confidence values can then be used to perform a weighted interpolation of the plane / landmark.
[0118] For example, if we use linear regression to xy +o pxy (Among them, pxy :=(0,0,o pxy )T ) to locate the position of the expected plane, the confidence value can be used to perform a weighted linear regression, such as a weighted least squares algorithm. The value of the weight or weighting factor can be equal to (one or more) confidence values.
[0119] As another example, if we calculate g across all initial positions xy +o lxy If the position of the first anatomical landmark in the 3D image is determined by the average or median of the anatomical landmarks, a weighted average or median may be determined using an appropriate formula (e.g., a weighted arithmetic mean). The weight used when determining the average / median using this technique may be a confidence value.
[0120] The previously described embodiments present methods of using machine learning algorithms to determine or predict a deviation from an initial position to a position within a desired plane and / or position of an anatomical landmark.
[0121] Specifically, embodiments utilize the recognition that a machine learning algorithm trained to determine or predict information defining the location of a desired plane within a 3D image shows improved performance if it is able to simultaneously predict the location of one or more anatomical landmarks outside of the desired plane. This is because, early in the training process, the location of the one or more anatomical landmarks provides additional ground truth data for training the machine learning algorithm and reduces the likelihood of overfitting.
[0122] This same underlying creative insight can be exploited to perform improved training of machine learning algorithms.
[0123] Therefore, a computer-implemented method for training a machine learning method to predict the position of a desired plane in a desired 3D image is also proposed. The computer-implemented method utilizes the same inventive insights as the previously described method for using a machine learning method.
[0124] The computer-implemented method includes obtaining a training data set including a plurality of training items. Each training item includes: an example 3D image of a target anatomical structure; first information identifying a location of a desired plane in the example 3D image; and second information identifying a location of one or more anatomical landmarks outside of the desired plane in the example 3D image.
[0125] For example, the training data set may be obtained from a database or other storage system. The training data set may be initially generated by a suitably trained or experienced clinician who labels different instances or examples of the 3D image data to identify the locations of desired planes and / or one or more anatomical landmarks.
[0126] The computer-implemented method also includes training a machine learning method using the training data set to predict information identifying the location of a desired plane in the desired 3D image and information identifying the locations of one or more anatomical landmarks outside of the desired plane in the desired 3D image.
[0127] For example, the information identifying the desired plane position in the desired 3D image may be a set of one or more first offsets. Each first predicted offset is a predicted spatial offset between an initial position relative to the 3D image and a predetermined position within the desired plane of the 3D image. For example, each predetermined position may be located within a cell within a cell grid covering the 2D plane. The positions of the predetermined positions and their corresponding cells may be predetermined.
[0128] Each initial position may be defined by processing the 2D slice using an object detection technique to identify a boundary region containing a representation of the target anatomical structure; and defining a set of initial positions within the boundary region.
[0129] Information identifying the position of one or more anatomical landmarks outside the desired plane may include a set of one or more second predicted offsets, each second predicted offset being a predicted spatial offset relative to the 3D image between the initial position and the position of an anatomical landmark of the target anatomical structure outside the desired plane.
[0130] Suitable examples of machine learning algorithms have been described above, but the machine learning algorithm is preferably a neural network, more preferably a multidimensional neural network.
[0131] A skilled person will be able to easily develop a processing system for executing any of the methods described herein. Therefore, each step of the flowchart may represent a different action performed by a processing system and may be performed by a corresponding module of the processing system.
[0132] Thus, embodiments may utilize a processing system. A processing system may be implemented in a variety of ways using software and / or hardware to perform the various functions desired. A processor is an example of a processing system that employs one or more microprocessors that may be programmed using software (e.g., microcode) to perform the desired functions. However, a processing system may be implemented with or without a processor, and may also be implemented as a combination of dedicated hardware for performing some functions and a processor (e.g., one or more programmed microprocessors and associated circuits) for performing other functions.
[0133] Examples of processing system components that may be used in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs).
[0134] In various implementations, a processor or processing system may be associated with one or more storage media, such as volatile and non-volatile computer memory, such as random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). The storage media may be encoded with one or more programs that perform desired functions when run on one or more processors and / or processing systems. The various storage media may be fixed within a processor or processing system, or may be transportable so that one or more programs stored thereon may be loaded into a processor or processing system.
[0135] Figure 6 A schematic diagram of a processor circuit 600 according to an embodiment is shown. As shown, the processor circuit 600 may include a processor 606, a memory 603, and a communication module 608. These elements may communicate with each other directly or indirectly, for example, via one or more buses.
[0136] The processor 606 contemplated by the present disclosure may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a controller, a field programmable gate array (FPGA) device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein. The processor 606 may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessor cores in combination with a DSP, or any other such configuration. The processor 606 may also implement various deep learning networks, which may include hardware or software implementations. The processor 606 may also include a preprocessor implemented in hardware or software.
[0137] The memory 603 contemplated by the present disclosure may be any suitable storage device, such as a cache memory (e.g., a cache memory of the processor 606), a random access memory (RAM), a magnetoresistive RAM (MRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a solid-state memory device, a hard drive, other forms of volatile and non-volatile memory, or a combination of different types of memory. The memory may be distributed among multiple memory devices and / or remotely located relative to the processor circuitry. In one embodiment, the memory 603 may store instructions 605. The instructions 605 may include instructions that, when executed by the processor 606, cause the processor 606 to perform the operations described herein.
[0138] Instructions 605 may also be referred to as code. The terms "instructions" and "code" should be broadly interpreted as including any type of (one or more) computer-readable statements. For example, the terms "instructions" and "code" may refer to one or more programs, routines, subroutines, functions, processes, etc. "Instructions" and "code" may include a single computer-readable statement or multiple computer-readable statements. Instructions 605 may be in the form of an executable computer program or script. For example, routines, subroutines and / or functions may be defined in a programming language, including but not limited to C, C++, C#, Pascal, BASIC, API calls, HTML, XHTML, XML, ASP scripts, JavaScript, FORTRAN, COBOL, Perl, Java, ADA, .NET, etc. Instructions may also include instructions for training anatomical structures, deep learning and / or machine learning modules.
[0139] The communication module 608 may include any electronic circuit and / or logic circuit to facilitate direct or indirect data communication between the processor circuit 600 and, for example, an external display (not shown) and / or an imaging device or system (such as an X-ray imaging system). In this regard, the communication module 608 may be an input / output (I / O) device. Communication may be performed via any suitable means. For example, the communication means may be a wired link, such as a universal serial bus (USB) link or an Ethernet link. Alternatively, the communication means may be a wireless link, such as an ultra-wideband (UWB) link, an Institute of Electrical and Electronics Engineers (IEEE) 802.11 WiFi link, or a Bluetooth link.
[0140] It should be understood that the disclosed method is preferably a computer-implemented method. Thus, the concept of a computer program is also proposed, which includes a code for implementing any described method when the program is run on a processing system (e.g., a computer). Therefore, different parts, lines or code blocks of a computer program according to an embodiment can be executed by a processing system or a computer to perform any method described herein.
[0141] A non-transitory storage medium storing a computer program is also proposed.
[0142] In some alternative embodiments, the functions recorded in the block diagram(s) or flow chart(s) may not occur in the order shown in the drawings. For example, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functions involved.
[0143] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.
[0144] In the claims, the word "comprising" does not exclude other elements or steps, and the word "a" or "an" does not exclude a plurality. A single processor or other unit may perform the functions of several items recited in the claims. Even though certain measures are recited in mutually different dependent claims, this does not indicate that a combination of these measures cannot be used to advantage.
[0145] If a computer program is described above, it may be stored / distributed on a suitable medium such as an optical storage medium or a solid-state medium provided together with other hardware or as part of other hardware, but may also be distributed in other forms such as via the Internet or other wired or wireless telecommunications systems.
[0146] If the term "suitable for" is used in the claims or the description, it should be noted that the term "suitable for" is intended to be equivalent to the term "configured to". If the term "arranged" is used in the claims or the description, it should be noted that the word "arranged" is intended to be equivalent to the term "system" and vice versa. Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A computer-implemented method (100) for processing a 3D image of a target anatomical structure, the computer-implemented method include: obtaining (110) a 2D slice (210) of the 3D image; processing (120) the 2D slice using an object detection technique to identify a boundary region (220) containing a representation of the target anatomical structure; defining (130) a set of initial positions (221, 222) within the boundary region; and The 3D image is processed (140) using a machine learning method to predict for each initial position: a set of one or more first predicted offsets (251, 252), each first predicted offset being a predicted spatial offset relative to the 3D image between the initial position and a predetermined position of the 3D image within a desired plane (230); as well as A set of one or more second predicted offsets (261, 262), each second predicted offset being a predicted spatial offset relative to the 3D image between the initial position and a position of an anatomical landmark (240) of the target anatomical structure outside of the desired plane.
2. The computer-implemented method of claim 1 , further comprising: The following steps are involved: The set of initial positions, the set of one or more first prediction offsets, and the 3D image are processed to produce at least one image that is smaller than the 3D image, the at least one image containing a representation of one or more of the one or more desired positions.
3. The computer-implemented method according to claim 1 or 2, in: Each initial position is associated with a different respective desired position within the desired plane of the 3D image for viewing the target anatomical structure; and The set of one or more first prediction offsets comprises, for each initial position, a single prediction offset, the single prediction offset being a predicted spatial offset between the initial position and its corresponding desired position.
4. The computer-implemented method of claim 3, further comprising: include: estimating (410) the position of the desired plane within the 3D image by processing the set of initial positions and the set of first prediction offsets generated for each initial position; and The 3D image is processed (420) using the estimated position of the desired plane to produce a 2D image representative of the desired plane.
5. The computer-implemented method of claim 4, in, The position of the desired plane within the 3D image is estimated by processing the set of initial positions and the set of first prediction offsets generated for each initial position using a linear regression technique.
6. A computer-implemented method according to any one of claims 1 to 5, in, Each second predicted offset is a predicted spatial offset between an initial position and a position of the same first anatomical landmark of the target anatomical structure relative to the 3D image.
7. A computer-implemented method according to any one of claims 1 to 6, in, Each set of one or more second prediction offsets comprises a plurality of second prediction offsets.
8. A computer-implemented method according to any one of claims 1 to 7, in, The step of defining a set of initial positions within the boundary area comprises: defining a grid of cells within the boundary region; and A single position in each cell is identified, thereby defining a plurality of said initial positions.
9. A computer-implemented method according to any one of claims 1 to 8, in, The 3D image is defined by an initial 2D slice stack, and the 2D slice is closer to a center-most slice of the initial 2D slice stack than a first or a last 2D slice in the initial 2D slice stack.
10. A computer-implemented method according to any one of claims 1 to 9, in, The step of processing the 2D slices using a machine learning method also includes generating a confidence value for each predicted offset, the confidence value representing the confidence or certainty of the predicted offset.
11. The computer-implemented method of claim 10, further comprising: include: estimating (410) the position of the desired plane within the 3D image by processing the set of initial positions and the set of first prediction offsets generated for each initial position using a linear regression technique, Wherein the values of one or more weighting parameters of the linear regression technique are responsive to the confidence values of one or more prediction offsets.
12. A computer-implemented method for training a machine learning method to predict the location of a desired plane in a desired 3D image, the computer-implemented method include: obtaining a training data set including a plurality of training items, each training item including: an example 3D image of a target anatomical structure; first information identifying a location of a desired plane within the example 3D image, and second information identifying a location of one or more anatomical landmarks in the example 3D image outside of the desired plane; and The machine learning method is trained using the training data set to predict information identifying the location of the desired plane in the desired 3D image and information identifying the locations of one or more anatomical landmarks in the desired 3D image outside of the desired plane.
13. A computer program product comprising computer program code which, when executed on a computing device having a processing system, causes the processing system to perform all the steps of the method according to any one of claims 1 to 12.
14. A processing system for processing a 3D image of a target anatomical structure, the processing system being configured to perform the method according to any one of claims 1 to 11.
15. The processing system according to claim 14, in, The processing system is also configured to process the set of initial positions, at least one set of the one or more prediction offsets, and the 3D image to produce at least one image smaller than the 3D image, the at least one image containing a representation of one or more of the one or more desired positions.