Image processing system and image processing method
The image processing system employs 3D convolutional neural networks to accurately estimate coronary artery stenosis from multiple images, overcoming the limitations of 2D convolution methods and achieving high precision in stenosis classification.
Patent Information
- Application Number
- JP2023189649
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-19
AI Technical Summary
Existing methods using 2D convolution algorithms for estimating coronary artery stenosis from coronary angiography images are limited by processing only one image at a time, resulting in low estimation accuracy.
An image processing system utilizing a 3D convolutional neural network for feature extraction and a fully connected neural network for identifying the stenosis state of the coronary artery, allowing for the processing of multiple coronary arteries simultaneously and estimating the stenosis state into healthy, mild, and severe classes.
The system achieves high precision in estimating the stenosis state of coronary arteries by performing machine learning on multiple non-contrast images at once, reducing patient burden by not requiring a contrast agent.
Smart Images

Figure 2025077457000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing system using three-dimensional convolution and an image processing method thereof. In particular, the present invention relates to an image processing system that takes a non-contrast CT image as an input and outputs an estimation of the stenosis state of the coronary artery.
Background Art
[0002] Prior to the present invention, a technique for quantifying coronary artery stenosis from a plurality of Computed Tomography images (hereinafter referred to as "CT images") using three-dimensional convolution calculation processing has been disclosed. For example, Patent Document 1 discloses a calculation algorithm for evaluating coronary artery disease. Quantification of blood flow is performed from angiography and blood flow data, and a calculation algorithm for segmentation of angiographic images and blood flow evaluation, and a hydrodynamic simulation for quantification of blood flow are disclosed.
[0003] Prior art using deep learning includes Non-Patent Documents 1 and 2. In order to determine stenosis from coronary angiography images, a Convolutional Neural Network (hereinafter referred to as "CNN") is used to realize segmentation of coronary artery branches.
[0004] Also, in Non-Patent Document 3, a CNN is used to estimate a QCA (Quantitative Coronary Angiography) score representing the stenosis state of the coronary artery from an MPI (Myocardial Perfusion Image). In Non-Patent Documents 4 and 5, an image recognition model lnception-V3 is used to estimate the stenosis state from a plurality of coronary angiography images.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Non-Patent Literature
[0006]
Non-Patent Literature 1
Non-Patent Literature 2
Non-Patent Literature 3
[0007] In the above prior art, a 2D convolution algorithm is used to estimate the stenosis state from coronary angiography images, and there is a problem that the angiography images that can be processed at one time are limited to only one, resulting in low estimation accuracy.
[0008] The present invention is made to assist in the diagnosis of the stenosis state of the coronary artery, particularly the degree (severity) of the disease, by a specialist doctor or medical technician. In view of the problems of the above prior art, an object of the present invention is to provide an image processing system and an image processing method for accurately estimating the stenosis state of the coronary artery.
Means for Solving the Problems
[0009] The first image processing system according to the present invention is an image processing system for estimating the stenosis state of the coronary artery from a medical image in which the coronary artery is imaged by a processing device, The processing device A feature quantity extractor that extracts feature quantities of the image by machine learning using a 3D convolutional neural network, An identifier that identifies the stenosis state of the coronary artery by computational processing using a fully connected neural network from the feature quantities, And is characterized by estimating the stenosis state of the coronary artery for each desired class Thereof.
[0010] The second image processing system according to the present invention is the first image processing system according to the present invention, wherein the feature quantities are extracted by one-time machine learning from medical images in which two or more coronary arteries are imaged Thereof.
[0011] The third image processing system according to the present invention is the first or second image processing system according to the present invention, and is characterized by identifying the stenosis state of the coronary artery into healthy, mild, and severe classes Thereof.
[0012] The fourth image processing system according to the present invention is the third image processing system according to the present invention, The processing device includes preprocessing means for feature extraction. The preprocessing means extracts data from a file in DICOM (Digital Imaging and Communications in Medicine) format, performs conversion processing from luminance values to image density values, window processing, and conversion processing for converting the number of input channels, and performs processing such as resizing, cropping, and normalization on the image after the conversion processing. This is the feature.
[0013] The first image processing method according to the present invention is an image processing method for estimating the stenosis state of a coronary artery from a medical image in which the coronary artery is photographed, and a feature amount extraction step of extracting a feature amount of the image by machine learning using a three-dimensional convolutional neural network, and an identification step of identifying the stenosis state of the coronary artery by calculation processing using a fully connected neural network from the feature amount, and is provided to estimate the stenosis state of the coronary artery for each desired class. This is the feature.
[0014] The second image processing method according to the present invention is the first image processing method according to the present invention, wherein the feature amount is extracted by one machine learning from medical images in which two or more coronary arteries are photographed. This is the feature.
[0015] The third image processing method according to the present invention is the first or second image processing method according to the present invention, and identifies the stenosis state of the coronary artery into classes of healthy, mild, and severe. This is the feature.
Advantages of the Invention
[0016] With the image processing system and image processing method according to the present invention, it is possible to perform machine learning on a plurality of non-contrast images of coronary arteries at once and accurately estimate the stenosis state of the coronary arteries into a desired class with high precision. In addition, since no contrast agent is used, it is expected to reduce the burden on the patient.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0018] The image processing system according to the present invention integrates main hardware devices such as an arithmetic processing unit (CPU, etc.), a random access memory (RAM), and a read-only memory (ROM), and stores a program for controlling them. It is a storage medium, an external storage device (such as a non-volatile memory storage device) that stores programs and databases, a high-speed image processing device (GPU, etc.) that repeatedly executes element operations constituting a neural network, an image processing device for display, input / output devices such as a display and CT images, and a user interface operation device. It is a system that closely cooperates with the electrically connected hardware and sequentially executes program processing according to a predetermined rule.
[0019] Hereinafter, embodiments for carrying out the invention will be described with reference to diagrams and the like. The image processing system according to the present invention pre-learns feature amounts of coronary arteries and their stenosis states from CT images, and extracts feature amounts from new CT images other than during pre-learning based on the results. It is configured to include a feature amount extractor, and a discriminator that pre-learns the feature amounts of CT images during pre-learning and meta-information pre-annotated by a specialist, and based on the results, discriminates the stenosis state of the coronary artery (hereinafter sometimes referred to as "degree of stenosis") from new CT images other than during pre-learning.
[0020] A three-dimensional convolutional neural network (hereinafter sometimes referred to as a "three-dimensional convolutional model") is used for the feature amount extractor. A neural network consisting of fully connected layers is used for the discriminator. A block diagram of the configuration of the image processing system according to the present invention is shown in FIG. 1. The image processing system according to the present invention takes a non-contrast CT image as input, and extracts feature amounts of a plurality of CT images using a three-dimensional convolutional model. Next, the degree of stenosis of the coronary artery is estimated (discrimination result) from the feature amounts using a neural network consisting of fully connected layers.
[0021] In addition, the image processing method according to the present invention is divided into two process flows: training and estimation of a feature extractor and a discriminator. The flow of the training process is shown in FIGS. 2, 3, and 4. Also, the flow of the estimation process is shown in FIG. 5.
[0022] (1) Pretraining processing (S1, S2, S3, S5) FIG. 2 shows the flow of the pretraining process. In pretraining, S3 and S4 are processed a plurality of times (for the number of epochs) in the loop of S1, S2, and S5. S1 initializes a variable representing the number of loop iterations. S2 determines whether the maximum number of loop iterations has been reached, and if not, continues the learning of the neural network model. S3 is the learning process of the neural network model, and the flow of the learning process is shown in FIG. 3. S4 is the process of evaluating the neural network model after learning, and the flow of the evaluation process is shown in FIG. 4. S5 is the process of incrementing the related number (number of epochs) by 1 each time the learning process of the neural network model ends.
[0023] (2) Preprocessing of images used for pretraining The acquisition process (S6) of the image group acquires the image data to be used in one process from all the image data for learning. The image data for one patient consists of a plurality of non-contrast CT images, and furthermore, the learning image data includes the image data for a plurality of patients. Therefore, batch processing is performed to acquire data for N persons (N is a natural number) at a time from the learning image data group.
[0024] The input data of the image processing system according to the present invention is in the DICOM (Digital Imaging and Communications in Medicine) format, which is an international standard. The conversion process (S7) for the image extracts data from a DICOM format file and performs conversion from luminance values to CT values (image density values; in the medical field, water is set to 0 and air is set to -1000), window processing, and conversion processing of the number of input channels.
[0025] The preprocessing of the converted image (S8) converts the grayscale image or binary image into a 3-channel (Red, Green, Blue) image when input. Furthermore, conversion processing to an image size suitable for the feature extractor, augmentation processing, normalization of the image, etc. are performed. In the case of medical images, when there are few cases, data cannot be obtained sufficiently, and overfitting may occur when learning using a 3D convolutional model. Therefore, overfitting is prevented by augmentation processing.
[0026] The feature extraction process (S9) inputs the image data group (RGB image) that has undergone the preprocessing of the image into a feature extractor using a 3D convolutional neural network model, and obtains a 1D feature vector as the output.
[0027] The identification process (S10) of the stenosis state of the coronary artery outputs the degree of stenosis with the L-dimensional feature amount obtained from the feature extractor as the input. For the degree of stenosis of the coronary artery, for example, a QCA score, FFR, or severity class based on the QCA score can be used.
[0028] The loss calculation (S11) performs derivation of the estimation result by the neural network model and loss calculation of the correct answer. When estimating with the severity class, the Cross Entropy function is used.
[0029] The optimization process (S12) updates the parameters of the neural network model using the loss. In the embodiment according to the present invention, ADAM is used as the optimization function.
[0030] The loop process of the training data (S13) performs a loop process until all the image data of all patients are learned for batch processing using the image data group for N people per learning. When all patients are processed, one epoch of learning is completed, and the process proceeds to S4.
[0031] (3) Evaluation process of the model The acquisition process of the image group (S14) is performed on the evaluation image data in the same manner as S6. This evaluation is performed by extracting the image data of each patient one by one.
[0032] The conversion process to an image (S15) performs the conversion from the DICOM format to image data in the same manner as S7.
[0033] The preprocessing of the image (S16) is performed on the evaluation image data. Here, when inputting a grayscale image or a binary image, it is converted into a 3-channel (Red, Green, Blue) image. Also, conversion processing to an image size suitable for the feature extractor and normalization of the image are performed.
[0034] The feature extraction process (S17) extracts features using a 3D convolutional neural network model with the image data that has undergone the preprocessing of the image as the input.
[0035] The identification process of the stenosis state (S18) estimates the degree of stenosis with the features of S17 as the input.
[0036] The loss calculation (S19) uses a neural network model in the same manner as S11 to derive the estimation result and calculate the loss of the correct answer.
[0037] The saving of the loss and weights (S20) averages the losses of all the evaluation data, and further compares the minimum loss up to the previous epoch and the current loss. If the current loss is lower than the previous value, the weights of the current neural network model are saved. Finally, the weights with the lowest loss in the evaluation data are saved and used for the next estimation.
[0038] The loop process for the evaluation data (S21) performs a loop process until the data of all patients are learned because one patient's data is used in one evaluation. When all patients are evaluated, it ends as one epoch of evaluation and proceeds to S5.
[0039] The flow of estimation when unknown image data is input to the image processing system according to the present invention will be described. The acquisition process of the image group (S22) reads the image data groups of each patient one by one, similar to S14. Therefore, when there are multiple patients, batch processing is performed for each patient one by one.
[0040] In the image conversion process (S23), since the image data for which the presence or absence of coronary artery stenosis is unknown, M images (M is a natural number) before and after the target image for estimation are extracted.
[0041] In the preprocessing of the image (S24), when inputting a grayscale image or a binary image, conversion is performed to a 3-channel (Red, Green, Blue) image. Also, conversion processing to an image size suitable for the feature extractor and normalization of the image are performed. In particular, the same parameters as in S16 are used.
[0042] In the feature extraction process (S25), using the image data obtained by performing the preprocessing of the S24 image as input, features are extracted using a 3D convolutional neural network model.
[0043] In the stenosis state identification process (S26), the degree of coronary artery stenosis is estimated using the features of S25 as input.
[0044] In the output of the estimation result (S27), the degree of coronary artery stenosis is output from the derivation of the result of S26. For example, the QCA score, FFR, and classes of the degree of stenosis are used.
Example
[0045] In the example according to the present invention, operation confirmation was performed on the following computer. CPU: Intel (registered trademark) Xeon (registered trademark) Silver 4114 CPU @ 2.20GHz GPU: Quadro (registered trademark) RTX6000 / 8000 CUDA: 11.7 Pytorch: 1.13.1
[0046] A dataset was used with non-contrast CT images taken at a medical hospital. Labels regarding coronary artery stenosis were created by expert technicians, and for the coronary arteries, the presence or absence of stenosis, the image numbers of the stenosis, and the QCA score as the degree of stenosis were assigned.
[0047] The above labels were divided into three classes according to the following criteria. The QCA score is calculated for vascular sites selected from angiographic images by radiation specialist technicians, and takes values from 0.0 to 100.0, that is, 0.0 means completely normal and 100.0 means complete stenosis of the coronary artery. In the embodiments according to the present invention, the QCA score is learned and the range of the values of the QCA score is estimated.
[0048] Class 0 is the healthy class, with no stenosis in the coronary artery or a QCA score less than 50. Class 1 is the class of patients with moderate stenosis, with stenosis in the coronary artery and a QCA score of 50 or more and less than 75. Class 2 is the class of patients with severe stenosis, with stenosis in the coronary artery and a QCA score of 75 or more.
[0049] In the embodiments according to the present invention, the entire image data group of all patients was divided into three: for training, for evaluation (validation), and for testing. Table 1 shows the number of image data. From the image data group of 646 patients, 201 (locations) of training image data, 234 (locations) of evaluation image data group, and 299 (locations) of test image data group were obtained. The determination of the number of data was performed as follows. First, the entire image data group was divided into 7 to 3, and 3 / 10 of the image data group was used for testing. Next, the remaining 7 / 10 of the image data group was divided into 7 to 3 for training and evaluation. The training image data group was equalized to reduce the bias for each class. In the training data, when comparing Class 0 (no stenosis) and Class 2 (severe), there is a bias in the number of data. Therefore, if learning is performed as it is, the 3D convolutional model may be biased towards Class 0 for output. Thus, the bias of the learning data was reduced by equalizing the data of Class 0 and Class 1 to match the number of data of Class 2.
[0050]
Table 1
[0051] In the embodiments of the present invention, accuracy, precision, recall, and F1 score are used as evaluation metrics. The calculation formulas for each metric are shown below. Accuracy represents the proportion of data that can be correctly predicted out of all the data. Precision represents the proportion of data that can be correctly predicted out of the data predicted as positive. Recall represents the proportion of data that can be correctly predicted out of the data that is actually positive. The F1 score is the harmonic mean of precision and recall.
[0052]
Equation
[0053]
Equation
[0054]
Equation
[0055]
Equation
[0056] In the image conversion, the following processes were performed sequentially. 1) Conversion from luminance value to CT value Since the data included in the DICOM format may be expressed in luminance values, conversion to CT values representing the X-ray transmittance is performed. Organs, calcified sites, etc. can be separated by value thresholds. 2) Window processing Since the CT values have a wide range in terms of value magnitude, conversion processing (window processing) to the range of 0 to 255 is performed to clearly capture the necessary objects. 3) Conversion of the number of input channels The CT images in the embodiments according to the present invention use grayscale. The number of channels of the images is increased to input them into the neural network model. Any one channel is added and converted into three channels. In the case of color images, they can be used as they are.
[0057] For extracting neighboring images of the image to be estimated, M images before and after are extracted from the image to be estimated and used as input. M images before and after the target image for estimating the stenosis of the coronary artery are extracted from the image group of one patient. If there is stenosis in the coronary artery, N images before and after it (N is a natural number) are extracted, and if there is no stenosis in the coronary artery, N images before and after are extracted from images at random positions.
[0058] In the preprocessing of the training images, the following processing was performed. 4) Resizing of images The original image size of 512×512 pixels is converted to 455×256 pixels. 5) Cropping of images A random position on the target image is cropped at 244×244 pixels. 6) Data augmentation Horizontal flipping is performed on any image in the image group with a probability of 50%. 7) Normalization Normalization is performed for each channel (r, g, b) of the image. The normalization is performed with the mean μr = 0.485, μb = 0.456, μg = 0.406 and the variance σr = 0.229, σb = 0.224, σg = 0.225.
[0059] The preprocessing of the test images is to perform the following processing. 8) Resizing of images The original image size of 512×512 pixels is converted to 455×256 pixels. In the embodiments according to the present invention, the estimation performance was most improved when the image size was 455×255. 9) Cropping of images The center position of the image is cropped at 244×244 pixels. 10) Normalization Normalize for each channel (r, g, b) of the image. Normalize with the mean μr = 0.485, μb = 0.456, μg = 0.406 and the variance σr = 0.229, σb = 0.224, σg = 0.225.
[0060] FIG. 6 illustrates a schematic diagram showing the configuration of an image processing system according to an embodiment of the present invention. The feature extractor of the image processing system according to the present invention uses I3D as a three-dimensional neural network model. I3D is a three-dimensional convolutional model based on Inception-V1 and was proposed for performing video recognition with OpticalFlow and an RGB image group as inputs. In this embodiment, an RGB image is used as an input, and OpticalFlow is not used. Further, the three-dimensional neural network model is trained using the video recognition dataset Kineticcs-400 as transfer learning, and in addition, non-contrast CT images are added and trained, enabling the extraction of features even with a small amount of input data. Furthermore, after pre-training, except for the layers subsequent to the Average-Pooling layer, a feature amount of 1024 dimensions is output.
[0061] The discriminator of the image processing system according to the present invention uses a neural network model composed of two layers of FC (Fully-Connected) layers. The first layer takes a feature amount of 1024 dimensions as an input and outputs a feature amount of 256 dimensions. The ReLU function is used as the activation function. The second layer takes a feature amount of 256 dimensions as an input and outputs probability values for three classes. The Softmax function is used as the activation function.
[0062] In the embodiment according to the present invention, as a comparison target, the feature extractor according to the present invention is replaced with a two-dimensional convolutional model (DenseNet-201, EfficientNet-BO, Inception-V3, ResNet-152). After extracting the feature amount of each image, the feature amounts are averaged and used. On the other hand, the same discriminator is used.
[0063] The parameters in the embodiments according to the present invention will be described. First, the number of input sheets was set to 13 sheets (6 sheets before and after). The ADAM optimization function was used. The learning rate was set to 0.0001 for the feature extractor and 0.001 for the discriminator. Also, the Cross Entropy loss function was used. Learning was performed with a batch size of 16 and the number of epochs set to 50.
[0064] The performance comparison between the image processing system according to the present invention and the 2D convolutional neural network model was conducted. In the estimation of the degree of coronary artery stenosis using non-contrast CT images as input, Table 2 shows the results obtained using the 3D convolutional neural network model and the 2D convolutional neural network model. In terms of accuracy, precision, recall, and F1 value, the image processing system according to the present invention using 3D convolution showed the highest performance. In the medical field, it is important not to miss diseases. In terms of recall, the image processing system according to the present invention showed high performance, indicating that the rate of missing diseases is low.
[0065]
Table 2
[0066] Next, a comparison is made with the image processing system using EfficientNet-B0, which has high accuracy, in the image processing system according to the present invention. FIGS. 7 and 8 show the ROC curves, and Table 3 shows the AUC. Comparing the ROC curves in both figures, it can be seen that there is a significant difference in performance between Class 1 and Class 2. For example, when TPR = 0.8 is used as a reference, it can be seen that the image processing system according to the present invention showed a low value of FPR. Furthermore, even when comparing the AUC in Table 3, the difference in values between Class 1 and Class 2 is large, and the image processing system according to the present invention shows high performance results.
[0067]
Table 3
Example
[0068] The performance of estimating the presence or absence of coronary artery stenosis was compared. The presence or absence of coronary artery stenosis was defined as no stenosis when the QCA score was less than 50% and stenosis when it was 50% or more, based on the QCA score. The results of the evaluation indicators are shown in Table 4. In terms of accuracy, recall, and F1 value, the image processing system according to the present invention showed the highest performance results. Similar to the three-class estimation results, in the estimation of the presence or absence of coronary artery stenosis, the image processing system according to the present invention also showed the highest performance.
[0069]
Table 4
Explanation of Signs
[0070] S1 Initialization of the number of loop iterations S2 Judgment of the number of loop iterations S3 Learning process S4 Evaluation of the neural network model after learning S5 Process of increasing the number of epochs by 1 for each learning process S6 Acquisition of the image group S7 Image format conversion S8 Image preprocessing S9 Feature extraction S10 Identification of the stenosis state of the coronary artery S11 Loss calculation S12 Optimization S13 Learning loop process
Claims
1. An image processing system for estimating a stenosis state of a coronary artery from a medical image of the coronary artery by a processing device, comprising: The processing device includes: A feature extractor that extracts features of the image by machine learning using a three-dimensional convolutional neural network; a classifier that classifies a state of stenosis of a coronary artery from the feature quantity by computational processing using a fully connected neural network; and estimating the stenosis state of the coronary artery by desired class.
1. An image processing system comprising:
2. The feature amount is extracted by one machine learning process from medical images of two or more coronary arteries.
2. The image processing system according to claim 1.
3. Classify the stenosis state of the coronary artery into normal, mild, and severe classes.
3. The image processing system according to claim 1, wherein the first and second inputs are input to the image processing apparatus.
4. The processing device includes a preprocessing means for extracting features, The preprocessing means extracts data from a file in Digital Imaging and Communications in Medicine (DICOM) format, Executes conversion processing from luminance values to image density values, window processing, and conversion processing for converting the number of input channels, After conversion, the image is resized, cropped, and normalized.
4. The image processing system according to claim 3.
5. 1. An image processing method for estimating a state of stenosis of a coronary artery from a medical image of the coronary artery, comprising: A feature extraction step of extracting features of the image by machine learning using a three-dimensional convolutional neural network; a classification step of classifying a state of stenosis of a coronary artery from the feature quantity by a computational process using a fully connected neural network; and estimating the stenosis state of the coronary artery by desired class.
13. An image processing method comprising:
6. The feature amount is extracted by one machine learning process from medical images of two or more coronary arteries.
6. The image processing method according to claim 5.
7. Classify the stenosis state of the coronary artery into normal, mild, and severe classes.
7. The image processing method according to claim 5 or 6.
Citation Information
Patent Citations
Anatomical and functional assessment of coronary artery disease using machine learning
JP2023502429A