Medical image processor and medical image processing method
The medical image processing apparatus addresses the challenge of accurately detecting and labeling vertebrae by using a spline-based method that normalizes and adjusts landmark positions, resulting in improved accuracy and diagnostic confidence.
Patent Information
- Application Number
- JP2024201763
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-30
AI Technical Summary
Existing medical image processing technologies face challenges in accurately detecting and labeling anatomical landmarks, particularly vertebrae in the spine, due to physiological curvatures and similarities between adjacent vertebrae, leading to incorrect labeling and reduced diagnostic confidence.
A medical image processing apparatus and method that includes an acquisition unit for obtaining anatomical landmark positions, a generation unit for creating a spline based on these positions, a normalization processing unit for converting the spline into one-dimensional or two-dimensional data, an adjustment unit for correcting landmark positions, and a remapping unit for repositioning the landmarks in three-dimensional space.
This approach enhances the accuracy of landmark detection and labeling, reduces errors in vertebrae identification, and improves diagnostic confidence by correctly numbering vertebrae and maintaining anatomical accuracy.
Smart Images

Figure 2025083326000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments disclosed in this specification generally relate to a medical image processing apparatus and a medical image processing method for identifying the positions of anatomical landmarks, for example, for identifying the positions of human spinal vertebrae.
Background Art
[0002] It is known to process medical images to identify anatomical landmarks. An anatomical landmark is generally a clearly defined location within a biological structure, such as a human anatomical structure. Anatomical landmarks are defined in relation to anatomical structures such as bones, blood vessels, organs, etc.
[0003] Landmark Automatic Detection (ALD) is used as bone diagnosis support. ALD is used for detecting bone metastases in the chest and anatomical landmarks of bones around the lesion site. Labels are added to each vertebra of the spine. By using ALD, the position of bone metastases can be efficiently and accurately identified by the labels of the vertebrae, and diagnosis support can be enabled by making it easier to read scans and report results.
[0004] Since the spine has physiological curvatures, it may be difficult to accurately detect anatomical landmarks. Vertebrae are usually used as anatomical landmarks when processing human spinal images, but there are multiple vertebrae in the human body that have few external differences. The appearance of adjacent vertebrae is similar, and it may be difficult to uniquely and accurately detect each individual vertebra. It is difficult to evaluate the accuracy of anatomical landmarks, which makes it difficult to automatically number the vertebrae in post-processing. For example, in order for a neural network used in image processing to clearly identify landmarks, consistent features are required. Detecting repetitive structures is difficult.
[0005] If the detection network fails to properly detect adjacent vertebrae, it may not be able to detect one or more vertebrae along the spine. Also, by labeling two or more vertebrae within the space of a single vertebra, an incorrect label may be added to that vertebra. Further, for example, labels that are all offset by a common offset amount, such as offset by 1 in all cases, may be incorrectly attached to multiple vertebrae from the ground truth. Such offset processing may cause a chain reaction of errors.
[0006] Accurate automatic labeling can potentially save, for example, several minutes per examination, i.e., more than one hour per day, of the time of radiologic technologists. Accurate automatic labeling makes it possible to improve the workflow of radiologic technologists. By correctly identifying vertebrae and labeling them in the expected anatomical order, the confidence of clinicians can be increased and accurate diagnosis can be supported.
[0007] Figure 1 shows an example of offset processing with incorrect labeling of vertebrae in a medical image 2. The labels shown in Figure 1 were added by the detection network. The first vertebra 10 is correctly identified as T6. The detection network missed the second vertebra 12 that should be identified as T7 and did not label it. For the third vertebra 14 that should be identified as T8, the detection network attached the label of T7. Similarly, the vertebrae 16, 18, 20, 22, 24, 26 that should be identified as T9, T10, T11, T12, L1, L2 respectively have been labeled as T8, T9, T10, T11, T12, L1 respectively by the detection network. The detection network has added two labels L2 and L3 to the vertebra 28 that should be identified as L3.
[0008] Incorrect identification may reduce the usefulness of the labeling and may also cause clinicians to lose confidence in the labels.
Prior Art Documents
Patent Documents
[0009]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0010] One of the problems to be solved by the embodiments disclosed in this specification and the drawings is to perform landmark detection of each vertebra in volume data with high accuracy. However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problems. The problems corresponding to the respective effects of each configuration shown in the embodiments described later can also be regarded as other problems.
Means for Solving the Problems
[0011] The medical image processing apparatus according to the embodiment includes an acquisition unit, a generation unit, a normalization processing unit, an adjustment unit, and a remapping processing unit. The acquisition unit acquires the position of each of a plurality of anatomical landmarks in a three-dimensional space. The generation unit generates a spline based on the positions of the plurality of anatomical landmarks. The normalization processing unit normalizes the spline into one-dimensional data or two-dimensional data. The adjustment unit adjusts at least one of the plurality of anatomical landmarks based on the one-dimensional data or the two-dimensional data. The remapping processing unit remaps the adjusted at least one anatomical landmark into a three-dimensional space.
[0012] Next, the embodiments shown in the following drawings will be described as non-limiting examples.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
[0014] A medical image processing apparatus according to one embodiment includes an acquisition unit, a generation unit, a normalization processing unit, an adjustment unit, and a remapping processing unit. The acquisition unit acquires the position of each of a plurality of anatomical landmarks in a three-dimensional space. The generation unit generates a spline based on the positions of the plurality of anatomical landmarks. The normalization processing unit normalizes the spline into one-dimensional data or two-dimensional data. The adjustment unit adjusts at least one of the plurality of anatomical landmarks based on the one-dimensional data or the two-dimensional data. The remapping processing unit remaps the adjusted at least one anatomical landmark into the three-dimensional space.
[0015] A medical image processing method according to one embodiment includes acquiring the position of each of a plurality of anatomical landmarks in a three-dimensional space, generating a spline based on the positions of the plurality of anatomical landmarks, normalizing the spline into one-dimensional data or two-dimensional data, adjusting at least one of the plurality of anatomical landmarks based on the one-dimensional data or the two-dimensional data, and remapping the adjusted at least one anatomical landmark into the three-dimensional space.
[0016] FIG. 2 shows an overview of a medical image processing apparatus 20 according to an embodiment. The medical image processing apparatus 20 includes an arithmetic unit 22 that is directly or indirectly connected to a scanner 24 and a data storage unit 30. In this example, the arithmetic unit 22 is a personal computer (PC) or a workstation.
[0017] The medical image processing apparatus 20 also includes one or more display screens 26 and one or a plurality of input devices 28 such as a computer keyboard, a mouse, and a trackball.
[0018] In the present embodiment, the scanner 24 is a computed tomography (CT) scanner. The scanner 24 generates image data representing at least one anatomical region of a patient or other subject. The image data includes a plurality of voxels each having a corresponding data value. In the present embodiment, the data value represents the luminance of a Hounsfield unit.
[0019] In another embodiment, the scanner 24 acquires two-dimensional, three-dimensional, or four-dimensional image data by any imaging modality. For example, the scanner 24 may include a magnetic resonance (MR) scanner, a computed tomography (CT) scanner, a cone beam CT scanner, a positron emission tomography (PET) scanner, an X-ray scanner, or an ultrasonic scanner.
[0020] In the present embodiment, the image data set acquired by the scanner 24 is stored in the data storage unit 30 and provided to the arithmetic unit 22. In another embodiment, the image data set is supplied from a data storage unit (not shown) located at a remote location. The data storage unit 30 or the data storage unit at a remote location includes any form of memory. In one embodiment, the medical image processing apparatus 20 is not connected to the scanner.
[0021] The computing device 22 includes a processing device 32 for data processing. The processing device 32 includes a Central Processing Unit (CPU) and a Graphics Processing Unit (GPU). Additionally, it may further include a Tensor Processing Unit (TPU). In another embodiment, any other processing device may be utilized. The processing device 32 provides processing resources for automatically or semi-automatically processing a medical image dataset. In another embodiment, the data to be processed may include any image data other than medical image data.
[0022] The processing device 32 includes a landmark detection circuit 34 for detecting landmarks, a normalization circuit 36 for normalizing a three-dimensional landmark position to one dimension, and an adjustment circuit 38 for adjusting (e.g., correcting) the one-dimensional landmark position and remapping the adjusted (e.g., corrected) position to three dimensions.
[0023] In one embodiment, the adjustment circuit 38 may be referred to as a correction circuit.
[0024] In this embodiment, the circuits 34, 36, 38 are each realized by a computer program including computer-readable instructions for causing the method of the embodiment to be executed within the CPU and / or GPU and / or TPU. In another embodiment, these circuits may be realized as one or more application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs).
[0025] Further, the arithmetic unit 22 includes a hard drive of the PC and other components, namely, a RAM, a ROM, a data bus, an operating system including various device drivers, and hardware devices including a graphic card, etc. For the sake of clarity, FIG. 2 does not show these components.
[0026] The device 20 in FIG. 2 executes a method according to FIG. 3. FIG. 3 is a flowchart showing an outline of the landmark detection method according to the embodiment.
[0027] In step 40, the landmark detection circuit 34 receives a volume image data set obtained by a CT scan of the scanner 24. The volume image data set may also be referred to as an input CT image. The volume image data set includes luminance values of each of a plurality of voxels in the coordinate space of the data set. The coordinate space of the volume image data set may also be referred to as the patient space. The luminance value is also called an image data value. In other embodiments, the volume image data set may be acquired from one or more arbitrary scanners. In yet another embodiment, the volume image data set may include data of one or more arbitrary modalities.
[0028] The CT scan for acquiring the volume image data set is a CT scan of the spine of the subject. The CT scan of the spine in the embodiment of FIG. 3 includes all vertebrae of the spinal column. In other embodiments, the CT scan may include a subset of the spinal vertebrae. For example, the CT scan may be a scan of only the upper spine or only the lower spine.
[0029] The spine of the subject has a physiological curvature as represented by the curve 42 in FIG. 3. Although the curvature is shown two-dimensionally in FIG. 3, the actual curvature of the spine may also be three-dimensional. The actual spine may curve in the left-right direction and the front-back direction.
[0030] In step 50, the landmark detection circuit 34 performs an initial localization step and finds the vertebrae of the spine by processing the volume image dataset.
[0031] In the embodiment of FIG. 3, the initial localization step includes inputting the volume image dataset into a trained model that has been trained to identify a plurality of anatomical landmarks. The plurality of anatomical landmarks includes one anatomical landmark per vertebra, located at the central position of the vertebra. In other embodiments, different anatomical landmarks may be used.
[0032] The trained model comprises, for example, a trained convolutional neural network. This convolutional neural network is trained to output an output dataset that includes a set of points (e.g., C1…C7, T1…T12, and L1…L3) each with a vertebra label added. The output data includes the three-dimensional position of each point within the coordinate space of the volume image dataset. The three-dimensional position is also referred to as the initial landmark position, which in the case of FIG. 3 is the initial landmark 52. For example, the initial landmark 52 includes the center of the vertebra within the coordinate space of the volume image dataset. Also, the initial landmark 52 includes a set of three-dimensional positions that span the whole or part of the volume of the vertebra. Also, the initial landmark 52 may include the three-dimensional positions of the uppermost vertebra 74 and the lowermost vertebra 76, which are also referred to as anchor landmarks.
[0033] Here, in steps 40 and 50 of FIG. 3, it can be said that the landmark detection circuit 34 obtains the volume data including a plurality of vertebrae and obtains the position of each of the plurality of anatomical landmarks by detecting the plurality of anatomical landmarks within the volume data. Therefore, the landmark detection circuit 34 in steps 40 and 50 of FIG. 3 is an example of an acquisition unit.
[0034] As described above, it is difficult for a learned model to correctly detect all vertebrae. For example, in the first localization step, the learned model may fail to detect one or more vertebrae along the spine, add labels to multiple vertebrae in a single vertebra space, add incorrect labels to one or more vertebrae, or cause a chain of errors such as omitting one from the ground truth and adding a label.
[0035] Subsequent steps in FIG. 3 are aimed at correcting inaccurate data obtained in the first localization step, such as correcting inaccurate labels and / or inaccurate three-dimensional positions.
[0036] In one embodiment, for example, adjustments such as corrections are made as appropriate. For example, adjustment of anatomical landmarks may include addition of one or more landmarks, deletion of one or more landmarks, reordering of landmarks, or change of at least one position or label within the landmarks.
[0037] In step 60, the normalization circuit 36 fits a smooth three-dimensional parameter curve to the three-dimensional points output in step 50. In the embodiment of FIG. 3, the smooth three-dimensional parameter curve is a three-dimensional spline 62. Any spline-fitting method can be used. For example, a low-order spline is fitted to the centerline consisting of the vertebra centers. The spline can be fitted to the three-dimensional positions of the detected initial landmarks without the need for labeled landmarks.
[0038] For spline fitting, hyperparameters representing smoothness and a measure of how well the spline passes through the three-dimensional points are used. That is, a balance is taken between the smoothness of the spline and the proximity of the spline to the initial landmarks output in step 50.
[0039] In step 60, the normalization circuit 36 may automatically evaluate whether the added label makes sense from an anatomical point of view. The intervertebral displacement vector is a useful means for evaluating whether the added label matches the anatomical structure of the spinal column. In one embodiment, the intervertebral displacement vector connects between the vertebral centers within the coordinate space of the volume image dataset. For example, the intervertebral displacement vector between vertebra T8 and vertebra T9 extends from the center of the T8 vertebra to the center of the T9 vertebra and towards the upper end of the spinal column from T8 to T9. In other embodiments, the vector connects between other landmarks detected within the volume of each vertebra, or is in the direction of or towards another landmark or reference position.
[0040] The normalization circuit 36 may evaluate the accuracy of the detected landmarks by comparing each intervertebral displacement vector (e.g., including magnitude and direction respectively) with a spline and determining the combination of the spatial position, relative position, and orientation of the intervertebral displacement vector. In one embodiment, this accuracy evaluation may include an evaluation of the relative position and orientation of consecutive landmarks such as consecutive vertebrae. In one embodiment, the midpoint of the displacement vector is projected onto the tangent vector of the spline recorded at the intersection with the spline. The angle between the intervertebral displacement vector and the tangent vector of the spline may be determined. In other embodiments, other criteria may be defined to evaluate the distance between the landmarks detected in the coordinate space of the volume image dataset and each relative position.
[0041] The normalization circuit 36 can automatically evaluate the accuracy of labeling. In one embodiment shown in FIG. 3, the landmark detection accuracy is evaluated based on the accuracy of the relative orientation and the mutual distance of the intervertebral displacement vectors 44, 46, 48, 54, 56, 58 indicating the spatial arrangement of the initial landmarks 52, and the anatomical accuracy of the distance from the detected landmarks. In one embodiment, the ground truth in the orientation compared with the initial landmark 52 is the average value of the patient's records, for example, based on the reference distance between vertebrae and the deviation from the reference distance, or other criteria. Also, for example, based on the deviation from the reference direction of the reference direction of the vector in the space. In one embodiment, the correct direction is determined based on the relative position of the consecutive vertebrae on the vertical axis parallel to the spine, that is, whether a certain vertebra is above or below the neighboring vertebrae. In other embodiments, other means and techniques are used to compare the labeled volume image dataset with the anatomically correct reference.
[0042] FIG. 4 shows a two-dimensional display 92 of the intervertebral displacement vectors 44, 46, 48, 54, 56, 58 connecting the three-dimensional spline 62 and the vertebrae 64, 66, 68, 78, 84, 86, 88 from T1 to T7 in the volume image data. In one embodiment, the normalization circuit 36 determines the intervertebral displacement vectors 58, 56 that connect the vertebra T7 88 to the vertebra T6 86 and the vertebra T6 86 to the vertebra T5 84 in the correct size and correct orientation. Further, the normalization circuit 36 may record that the landmark detection accuracy regarding T7 and / or T6 and / or T5 is high. For example, if the intervertebral displacement vector 54 connecting the vertebra T5 84 to the vertebra T4 78 is considered to be smaller than the acceptable size, it is recorded that the landmark detection accuracy of the vertebra T5 and / or T4 is low. Also, for example, if the intervertebral displacement vector 48 connecting the vertebra T4 78 to the vertebra T3 68 is considered to be larger than the acceptable size, it is recorded that the landmark detection accuracy of the vertebra T4 and / or T3 is low. Until the accuracy evaluation of all the intervertebral displacement vectors is completed, the normalization circuit 36 traverses the two-dimensional display 92 of the three-dimensional spline 62 in one direction.
[0043] The output of step 60 is the adapted 3D spline 62.
[0044] Here, in step 60 of FIG. 3, it can be said that the normalization circuit 36 generates a spline based on anatomical landmarks. Therefore, the normalization circuit 36 in step 60 of FIG. 3 is an example of a generation unit.
[0045] In step 70, the normalization circuit 36 converts the landmark positions from the coordinate space of the volume image dataset to the reference coordinate space based on the adapted 3D spline 62. The reference coordinate space, also referred to as the reference frame, is a coordinate space in which the spine appears straight and the length of the spine is normalized.
[0046] If there are changes in the spinal curvature and orientation of the subject to be imaged, it means that useful statistical information may be lost as a result of post-processing of the spline data and the identified landmarks 52 in the patient space. In embodiments where the use of a neural network is included in the post-processing, the training dataset of the neural network includes changes in spinal curvature and orientation. By converting the adapted 3D spline 62 into a low-dimensional coordinate space such as 1D, training of a neural network using labeled 1D data, or 2D data in some embodiments, can be facilitated. In this embodiment, the 3D spline 62 is converted into a 1D coordinate space, but in other embodiments, the conversion is performed so as not to result in a 1D space. For example, a normal spine exhibits most of its physiological curvature in the anterior-posterior direction (due to the narrowness of the back and shoulders). In one embodiment, only this dimensional change is removed to obtain a 2D space.
[0047] Certain types of normalization are beneficial for removing at least some, and optionally as much as possible, of the physiological and anatomical variations. This makes it possible to effectively limit the possible post-processing options to those that maintain anatomical variations.
[0048] The conversion of 3D space to 1D or 2D space is part of dimensional normalization. In one embodiment, lengths and scales, which are another form of variation, are eliminated or reduced. Component decomposition can be performed, and each component can be solved or determined individually and appropriately. That is, the remaining problems can be simplified. For example, in the case of a convolutional neural network (CNN) or at least some other type of network, during training, the CNN is made to learn to detect certain features. Due to the nature of how the features are defined, the features learned by the CNN may not be scale or rotation invariant (although the downsampling layer provides a certain degree of scale invariance). That is, if it is desirable for the CNN to learn feature detection at multiple scales or rotation angles, the CNN needs to learn to detect features multiple times, for example, at multiple scales or rotation angles. Thus, a part of the CNN that has the potential to detect vertebrae 10 pixels wide may not be able to detect vertebrae 15 pixels wide. By removing these variations through normalization, the CNN and other networks can be focused on other relevant features.
[0049] In one embodiment, axial and oblique multi-planar reformatting (MPR) images are generated from CT data or other medical imaging modalities to observe human anatomical structures and other subjects. Axial and oblique MPR images include axial and oblique slices of a 3D image dataset, such as slices of a patient's spine. Using the curved planar reformatting (CPR) images of each vertebra, a composite image is generated from the set of CPR images of the spine. It is effective to represent each vertebra linearly within the image. The adapted 3D spline 62 is converted into a linear MPR image. In other embodiments, other images may be used as appropriate. In one embodiment, the normalization circuit 36 normalizes the volume data by creating curved multi-planar reformatting images based on the 3D spline.
[0050] By such coordinate transformation, the 3D spline 62 is made into a straight line, and a 1D line 72 shown in FIG. 5 is obtained. FIG. 5 shows the z-dimension in the vertical direction. Therefore, the coordinates along the spline 62 are shifted from the 3D space to the 1D space. Next, normalization is performed to reduce and expand the coordinates in the 1D space to a standard length, which can also be referred to as a fixed size. FIG. 5(a) shows a 2D display 92 of the 3D spline 62 and the vertebrae 64, 66, 68, 78, 84, 86, 88 between T1-T7. FIG. 5(b) shows the 1D line 72 after coordinate transformation with labels added to the vertebrae T1-T7. The 1D line 72 is an example of 1D data. In another embodiment, a shift is made from the 3D space to the 2D space, and normalization is performed within the 2D space.
[0051] In the embodiment of FIG. 3, the normalization circuit 36 determines the length of the spine based on two 3D positions, namely the anchor landmarks determined for the uppermost vertebra 74 and the lowermost vertebra 76 at step 50. The uppermost vertebra is the vertebra at the highest position within the volume image data set. The lowermost vertebra is the vertebra at the lowest position within the volume image data set. For example, when the entire spine is included in the volume image data set, in this example, the uppermost vertebra is C1 and the lowermost vertebra is L5 or L6.
[0052] The normalization circuit 36 determines a magnification factor representing the scale of the subject's spine based on the spine length determined from the positions of the uppermost and lowermost vertebrae. The magnification factor depends on the spine length. The normalization circuit 36 obtains the normalized 1D line 72 by expanding and shrinking the 1D space to the standard length with this magnification factor.
[0053] Depending on the situation, the uppermost and lowermost vertebrae are often more characteristic than other vertebrae and are thus often correctly located in the first localization step.
[0054] In other embodiments, different normalization methods may be used.
[0055] In one embodiment, the scale is estimated by dividing one or more of the vertebrae within the volumetric image dataset. The normalization circuit 36 determines the size or volume of each vertebra segment and determines a magnification by comparing the determined size or volume to a standard size or volume. Normalization includes using the magnification to transform coordinates to a one-dimensional line of standard length.
[0056] In this case, it can be said that the normalization circuit 36 divides one or more anatomical structures within the volume data. Therefore, the normalization circuit 36 in this case is an example of a dividing unit. Also, in this case, it can be said that the normalization circuit 36 determines the size or volume of the one or more divided anatomical structures. Therefore, the normalization circuit 36 in this case is also an example of a determining unit.
[0057] In one embodiment, measurements along the spine within the volumetric image dataset are used as the scale. The measured length may be scaled to a standard length.
[0058] In the embodiment of FIG. 3, the normalization circuit 36 resamples the volumetric image dataset by taking a plurality of cross-sectional images perpendicular to the spline. In one embodiment, it may be a CPR image of the spine. Each cross-sectional image becomes a slice of the resampled volume. The spine within the resampled volume is substantially straight.
[0059] For example, a fixed number of slices such as 50 or 100 can be used. The resampled volume has a fixed size. One of the dimensions of the fixed size is the standard length.
[0060] Here, in step 70 of FIG. 3, it can be said that the normalization circuit 36 normalizes the spline to one-dimensional data or two-dimensional data. Therefore, the normalization circuit 36 in step 70 of FIG. 3 is an example of a normalization processing unit.
[0061] Convert each initial landmark position to the normalized coordinate space of the resampling volume. Most of the landmark positions after conversion are expected to be on or near a one-dimensional line. Suppressing the spatial change in landmark position caused by this conversion is beneficial for the post-processing in step 80. Also, suppressing the scale and rotation changes in the image due to normalization is beneficial when using a neural network for the post-processing in step 80.
[0062] In step 80, the adjustment circuit 38 applies post-processing to obtain the final landmark position within the normalized coordinate space of the resampling volume. The post-processing may include correction of one or more landmark positions and other adjustments. The correction includes addition, deletion, and adjustment of one or more landmark positions or labels.
[0063] The adjustment circuit 38 may train a neural network, execute post-processing using the trained neural network, and obtain the corrected landmark position in the normalized coordinate space. This process includes using training data for obtaining the average position in the normalized space. For example, in an embodiment of a normalized space less than three dimensions such as a vertical one-dimensional space, based on the training data, an average z value on the z-axis is determined as shown in FIG. 5. Then, using any method, outliers are detected by comparing the landmark position converted to the reference space with the average z position. Examples of such methods include Gaussian likelihood and comparison with a threshold based on absolute coordinates (as an example, coordinates that are more than 10 mm away from the average position are outliers). Or, a percentage threshold may be used. Nearest neighbor comparison with the training data such as a one-class or other support vector machine (SVM) may also be used. In other embodiments, other outlier detection methods are available as appropriate. Also, it is possible to correct labeling errors such as skipping one by comparing the landmark position converted to the reference space with the average z position of neighboring landmarks. Landmarks are added, deleted, or repositioned. Hereinafter, such a method is referred to as "absolute one-dimensional average post-processing".
[0064] Any learning data can be used as appropriate. For example, the learning data includes scans of the spine with symbols added to landmarks. Normalization is performed using the landmark with the symbol, and the intermediate position is calculated in the normalized space.
[0065] In one embodiment, structural errors that repeat can be eliminated by absolute one-dimensional mean post-processing.
[0066] The adjustment circuit 38 may perform spatial optimization of the landmark position to determine the final landmark position. Based on the learning data and other data representing the related anatomical structure, a coordinate point with a high predicted value that matches a specific landmark in the transformed space is determined. Further, based on the data, the change in distance between anatomical landmarks may be identified, and the average distance between all landmarks may be determined. The algorithm for spatially optimizing the landmark position may include steps that encourage transforming the coordinate points in the transformed space into coordinate points with high predicted values. Also, the algorithm may include steps that prevent the landmark positions from approaching each other too closely. Hereinafter, this method is referred to as "global optimization post-processing".
[0067] In an embodiment where global optimization post-processing is applied to the normalized one-dimensional space, all landmarks are moved in the upward or downward direction of the spine, or in the ±Z^ direction of FIG. 5(b).
[0068] Alternatively, points are projected onto the spine of the vertebra by a method in the form of a heat map. To expand on this point, the CNN structure is not immediately suitable for landmark detection. A landmark only occupies a single voxel. That is, the CNN will attempt to learn a very imbalanced problem (many negative voxels versus very few positive voxels). Therefore, when it is preferable to have the CNN learn landmark identification, it is known to replace the symbol for each voxel with the symbol of the heat map. Instead of a single positive voxel, a Gaussian distribution of positive values, also called a heat map, is used. The Gaussian distribution reaches a peak at the landmark position and thins out as it moves away from the landmark. In one embodiment, a CPR resampling volume is created based on the normalization information and provided to the CNN. In another embodiment, an unnormalized volume may be used as an input to perform an estimation in the form of a heat map, and later the predicted value may be projected onto the nearest point of the spine.
[0069] Adjustment circuit 38 performs absolute one-dimensional or two-dimensional mean post-processing and global optimization post-processing using a neural network.
[0070] FIG. 6 schematically shows a neural network 94 as an example. The neural network 94 includes a node input layer, a node hidden layer, and a node output layer. In reality, there are often many more nodes and hidden layers than shown. Each node Ni in the input layer receives one value of the input data and generates an activation value or node value at its output. The activation value or node value is generated by performing an activation function (e.g., sigmoid) on the input value. Each node N in the input layer i is connected to each node N in the hidden layer h . The vector of node values from the input layer is scaled by the vector of respective weights at the input of each node in the hidden layer. Each weight defines the connectivity between a particular node and the node connected to the hidden layer. In FIG. 7, the weight applied to one of the inputs of node N h is w 0 ...w3 is shown as follows. The input value of each node N in the hidden layer h is given by the dot product of the weight vector connecting the node to the input layer and the output value of the input node. Next, by applying the activation function to the input value of node N h , the output value of node N h is obtained. The output vector of the hidden layer is supplied to each node in the next layer of the network (i.e., in this case, the output node N 0 of the output layer N 0 ) and is similarly used to generate the output value of the next layer.
[0071] For example, in the case of one embodiment according to the Unet architecture, the network has three pooling layers before the bottleneck and three non-pooling layers after the bottleneck to restore the spatial resolution to the same as the input resolution. At the beginning of the network, 32 feature maps are prepared, and after each pooling layer, the number of feature maps doubles, for example, to 32, 64, 128, 256. Also, the number of feature maps after the bottleneck becomes, after each non-pooling layer, for example, 256, 128, 64, 32 in the reverse pattern. When using 3D CNN, each kernel can have a shape of 3×3×3. Optionally, any model structure can be utilized.
[0072] The network 94 can be trained by various methods such as supervised or unsupervised learning. In one embodiment, through supervised learning, the network 94 identifies at least one set of output values, compares the output values with known labels representing ground truth values, and calculates an error or loss associated with the network 94 (e.g., based on the difference between the output values and the ground truth values), thereby training the network. Then, the loss is backpropagated through the network 94 to update the weights so that the network 94 can more appropriately approximate the labels from the input values. By this update, the weights can be optimized according to the objective function (e.g., adjusting the weights to reduce the error of the output values). In the next cycle, the updated weights and other training data are used to further update the weights. In this way, the network is trained to perform the desired operation.
[0073] A convolutional neural network can be utilized for image classification. A convolutional neural network (CNN) is a neural network that utilizes a convolutional operation in at least one of its layers. Since a convolutional neural network is shift-invariant, it is particularly suitable for image analysis and processing. However, a CNN filter is not rotation or scale-invariant. Therefore, to perform vertebral detection, it is necessary to learn the appearance of the vertebrae multiple times. For feature recognition in a series of medical images, a 2D convolutional neural network (CNN) is applied to identify the features of individual frames of a series of medical images. Alternatively, a 3D convolutional neural network (CNN) may be applied to feature recognition including the identification of temporal relationships between frames.
[0074] Refer to FIGS. 7 and 8 showing an operation example of a convolutional neural network. A convolutional neural network is utilized to identify specific features within an image and perform its classification. In the illustrated example, the input image 96 shows a CT scan of a patient's spine. The convolutional neural network can be utilized to identify the locations where each vertebra is located.
[0075] Apply the kernel 98 to determine the convolution of the input image 96 by the kernel 98. Perform an activation function on the output of this convolution to add non-linearity. The activation function used in FIG. 7 is a rectified linear unit (RELU) that outputs the input if it is positive and zero if it is not positive. In this way, by performing the convolution between the input image and different kernels representing different types of features such as vertical lines and horizontal lines, many feature maps 100 are generated from the input image.
[0076] Next, perform pooling processing on each feature map generated by convolution and the activation function. The pooling processing performed for reducing the spatial size of the features after convolution involves moving the kernel across the entire feature map, sampling pixel groups, and returning the maximum or average value from each sampled pixel within the feature map. Using the kernel, further convolution processing (along with the application of the RELU function) is performed on each of the resulting pooled feature maps to generate another set of feature maps on which pooling is again performed.
[0077] Using the pooled feature maps obtained from multiple stages of convolution and pooling, generate a one-dimensional array to be provided as input to a feed-forward neural network. The resulting output values represent the positions of each vertebra in the CT scan image. The convolutional neural network estimates whether the position of the vertebra is correct by processing the output values.
[0078] The convolutional neural network may be trained by comparing the output values for different images with the labels of those images and adjusting the weights of the feed-forward part of the convolutional neural network.
[0079] The normalization performed in step 70 can remove variations in scale and rotation, which is beneficial for the operation of the CNN. For other types of analysis, landmark positions with reduced inter-patient discrepancies are used. In such analysis, post-processing using probability models of landmark positions such as principal component analysis and a mixture Gaussian model of different vertebrae is performed. Also, by detecting mis-placed landmarks, certain predictable lesions such as age-related degenerative diseases can be identified. For example, lesions such as degenerative disc disease and scoliosis can be appropriately identified. In this case, it can be said that the normalization circuit 36 in step 70 identifies at least one lesion based on one-dimensional data. Therefore, the normalization circuit 36 in step 70 in this case is an example of an identification unit.
[0080] Here, in step 80 of FIG. 3, it can be said that the adjustment circuit 38 adjusts at least one of a plurality of anatomical landmarks based on the one-dimensional data obtained by normalizing the spline. Therefore, the adjustment circuit 38 in step 80 of FIG. 3 is an example of an adjustment unit.
[0081] In step 90, the adjustment circuit 38 maps the final landmark position from the normalized coordinate space to the coordinate space of the volume image dataset. This output is a series of coordinate points representing anatomical landmarks within the volume of the scanned patient.
[0082] The method described in this specification can be appropriately applied to any subject, such as a human or animal subject.
[0083] In various embodiments, for example, any landmark such as a rib or a tooth can be utilized.
[0084] Although the use of neural networks such as CNNs has been described, in other embodiments, other pre-trained models may be used for, for example, alignment, post-processing, adjustment, normalization, and other processing steps. Depending on the embodiment, for example, graph convolutional neural networks or Transformer architectures may be used.
[0085] In one embodiment, adjustments such as correction can be appropriately made for landmarks. For example, the adjustment of anatomical landmarks includes the addition of one or more landmarks, the deletion of one or more landmarks, the rearrangement of the order of landmarks, and the change of at least one position or label within the landmarks, by implementing a suitable pre-trained model and other processing procedures.
[0086] In one embodiment, the normalized volume may be used for other processing such as scaling in three dimensions. For example, one vertebra is segmented and normalized in all three dimensions.
[0087] There is a feature that other processing is more effective in the normalized space.
[0088] In one embodiment, it is possible to eliminate repetitive structural errors by absolute 1D mean post-processing. Based on the training data, the average z position in the normalized space is obtained. Outliers at inference time can be replaced with the average z value of that landmark. Outliers are detected by comparing the landmark predicted position with the average z position (e.g., based on Gaussian likelihood). The comparison between the landmark predicted position and neighboring landmarks can be used for corrections such as the detection and correction of a one-time chain reaction.
[0089] In one embodiment, the final position may be calibrated by global optimization post-processing. For example, one term makes it easier to place points at positions with high predicted values, and another repulsive term prevents the relative positions of the landmarks from getting too close. This enables all the landmarks in the 1D space to be surely pushed up / pulled down from the spinal spline.
[0090] Depending on the embodiment, one or more CNNs can be utilized. Since CNN filters are not rotation or scale invariant, it is generally necessary to learn the appearance of the vertebrae multiple times. By attempting to remove rotation and scale changes through normalization, the CNN becomes more effective.
[0091] A medical image processing apparatus according to one embodiment includes a processing circuit. The processing circuit receives volume data including a plurality of vertebrae, detects a plurality of anatomical landmarks corresponding to the plurality of vertebrae, generates a 3D spline based on the plurality of anatomical landmarks, normalizes the 3D spline into 1D data, corrects at least one of the anatomical landmarks based on the 1D data, and remaps the corrected anatomical landmarks into 3D space.
[0092] The processing circuit may further correct at least one of the anatomical landmarks based on teacher data which is data in which a plurality of anatomical landmarks and position information are associated with each other.
[0093] The processing circuit may further correct at least one of the anatomical landmarks on the 1D data normalized based on the anchor landmark.
[0094] A medical image processing apparatus according to one embodiment includes a processing circuit. The processing circuit performs the following operations. · Receive an image of the spine, · First detection of vertebral landmarks, ·(1) Estimate the spinal curvature by fitting the detected spline curve to the vertebrae, and (2) detect the parameters of spinal changes in the image by estimating the scale of the spine. · Normalize the image and landmark data using the detected parameters. · Apply post-processing enhanced by normalization to the vertebral landmarks by the following method. (1) Absolute 1D mean post-processing to obtain the average z position in the normalized space. This leads to an improvement in the accuracy of landmark positions based on the normalized representation of the image and / or landmarks. (2) Global optimization post-processing that promotes position prediction and calibrates the final position using terms that avoid the vicinity of adjacent landmarks. (3) CNN heatmap detection method that projects the landmark positions onto the spine spline.
[0095] The scale parameter may be estimated by measuring the distance between the top and bottom vertebrae.
[0096] The scale parameter may be estimated by dividing a vertebra and measuring the size or volume characteristics of the segmentation.
[0097] The image data may be normalized by creating a surface multi-segment reconstruction image using the fitted spline curve.
[0098] The landmarks may be normalized by converting the 3D coordinates to 1D coordinates that determine the position along the length of the spline curve.
[0099] Although specific circuits have been described in this specification, in other embodiments, one or more of the functions of these circuits can be implemented by a single processing resource or other components. Alternatively, the functions implemented by a single circuit can be realized by combining two or more processing resources or other components. A single circuit includes the meaning of a plurality of components that implement the functions of the circuit, regardless of whether they are separated from each other. A plurality of circuits includes the meaning of a single component that implements the functions of those circuits.
[0100] Although several embodiments have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. The novel methods and systems can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and its equivalent scope.
Description of Reference Numerals
[0101] 20 Device 22 Arithmetic Unit 26 Display Screen 28 Input Device 30 Data Storage Unit 32 Processing Unit 34 Landmark Detection Circuit 36 Normalization Circuit 38 Adjustment Circuit
Claims
1. an acquisition unit that acquires the positions of each of a plurality of anatomical landmarks in a three-dimensional space; a generator for generating a spline based on the positions of the plurality of anatomical landmarks; a normalization unit that normalizes the spline to one-dimensional data or two-dimensional data; an adjustment unit that adjusts at least one of the plurality of anatomical landmarks based on the one-dimensional data or the two-dimensional data; a remapping processor for remapping the adjusted at least one anatomical landmark into the three-dimensional space; A medical image processing device comprising:
2. the plurality of anatomical landmarks corresponds to a plurality of vertebrae; The medical image processing device according to claim 1 .
3. the acquiring unit further acquires volume data including the plurality of vertebrae, and detects the plurality of anatomical landmarks within the volume data to acquire positions of the anatomical landmarks. The medical image processing device according to claim 2 .
4. the normalization processing unit further normalizes the volume data by creating a curved multilevel reconstruction image using the spline. The medical image processing device according to claim 3 .
5. The detection of the plurality of anatomical landmarks includes automatic landmark detection using a trained model. The medical image processing device according to claim 3 .
6. The at least one anatomical landmark is adjusted based on supervised data, which is data in which the plurality of anatomical landmarks and position information are mutually associated. The medical image processing device according to claim 1 .
7. adjusting the at least one anatomical landmark includes comparing the one-dimensional or two-dimensional data to a set of one-dimensional or two-dimensional reference positions and adjusting the one-dimensional or two-dimensional data based on the one-dimensional or two-dimensional reference positions. The medical image processing device according to claim 1 .
8. the reference position of the one-dimensional data or the two-dimensional data is an average position of the plurality of anatomical landmarks in a reference data set; The medical image processing device according to claim 7 .
9. adjusting the at least one anatomical landmark includes facilitating position prediction through calibration and avoiding proximity of adjacent landmarks; The medical image processing device according to claim 1 .
10. Acquiring the positions of the plurality of anatomical landmarks by a registration process in the acquisition unit; Normalization in the normalization processing unit; Adjusting the at least one anatomical landmark with the adjustment unit; Remapping in the remapping processing unit; using at least one neural network or other trained model; The medical image processing device according to claim 4 .
11. and normalizing the position includes estimating a scale and normalizing according to the estimated scale. The medical image processing device according to claim 1 .
12. The scale is estimated based on the positions of two or more anchor landmarks. The medical image processing device according to claim 11.
13. The first anchor landmark corresponds to the top vertebra of the spine and the second anchor landmark corresponds to the bottom vertebra of the spine. The medical image processing device according to claim 12.
14. a segmentation unit for segmenting one or more anatomical structures within the volumetric data; and a determiner for determining a size or volume of the segmented one or more anatomical structures, The normalization unit estimates a scale of the determined size or volume, and normalizes the position according to the estimated scale. The medical image processing device according to claim 3 .
15. The spline comprises a cubic spline. The medical image processing device according to claim 1 .
16. adjusting at least one of the anatomical landmarks includes correcting at least one of the anatomical landmarks. The medical image processing device according to claim 1 .
17. adjusting at least one of the anatomical landmarks includes adding at least one of the landmarks, removing at least one of the landmarks, rearranging the order of the landmarks, or changing a position or label of at least one of the landmarks. The medical image processing device according to claim 1 .
18. and further comprising an identification unit for identifying at least one lesion based on the one-dimensional data or the two-dimensional data. The medical image processing device according to claim 1 .
19. the plurality of anatomical landmarks correspond to a plurality of vertebrae, ribs, or teeth; The medical image processing device according to claim 1 .
20. obtaining a position of each of a plurality of anatomical landmarks in three-dimensional space; generating a spline based on the locations of the plurality of anatomical landmarks; normalizing said spline to one-dimensional or two-dimensional data; adjusting at least one of the plurality of anatomical landmarks based on the one-dimensional data or the two-dimensional data; remapping the adjusted at least one anatomical landmark into the three-dimensional space; A medical image processing method comprising:
Citation Information
Patent Citations
Three-dimensional imaging and modeling of ultrasound image data
US20210045715A1