Generation method, training method, generation device, control program, and recording medium
The method addresses the challenge of image quality variations by generating high-quality learning images through image processing, enhancing the accuracy of machine learning models in medical image analysis.
Patent Information
- Application Number
- PCT/JP2024/038382
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-08
AI Technical Summary
Existing methods for generating medical images for machine learning models face challenges due to variations in image quality, which can lead to reduced accuracy in detecting medical abnormalities.
A method and device for generating learning images by acquiring medical images and performing image processing, including correction and conversion, to enhance image quality and improve estimation accuracy for machine learning models.
The proposed method effectively reduces the risk of reduced accuracy in machine learning models by generating high-quality learning images, thereby improving the estimation of bone density and other medical parameters.
Smart Images

Figure JP2024038382_08052025_PF_FP_ABST
Abstract
Description
Generation method, learning method, generation device, control program, and recording medium
[0001] The present disclosure relates to a generation method, a learning method, a generation device, a control program, and a recording medium.
[0002] A technique is known in which a medical image is loaded into a trained machine learning model to determine whether or not there is a medical abnormality in the tissue in the image. To train the machine learning model, it is known that annotated medical images in which an abnormality has actually been confirmed and annotated medical images in which no abnormality has been confirmed are used as training data.
[0003] Medical images used as training data are typically images taken in actual medical settings. Examples of medical images include X-ray images, MRI images, and CT scan images. It is known that image quality of such medical images varies depending on the imaging device, imaging method, imaging conditions, and other factors. For example, X-ray images of bones can result in images with excessive contrast, resulting in overexposed bones. Using such images for training a machine learning model may result in a machine learning model with poor accuracy in determining (or estimating) the presence or absence of bone tissue abnormalities. Therefore, in order to prepare training images for a machine learning model, images taken in a medical setting (original images) may be preprocessed by correcting the brightness, hue, and other aspects of the original images. The corrected original images are also called augmented images.
[0004] For example, Non-Patent Document 1 reports the effect of image preprocessing on classification when classifying chest X-ray images into three classes: normal (healthy), novel coronavirus disease (COVID-19), and pneumonia using a machine learning model.
[0005] Furthermore, Non-Patent Document 2 reports the results of comparing the performance of a deep learning algorithm for screening for novel coronavirus disease (COVID-19) using chest X-ray images with and without geometric augmentation of the training data. Non-Patent Document 2 indicates that the learning effect may not necessarily be higher when augmented images are used.
[0006] Furthermore, Non-Patent Document 3 reports a deep learning model that predicts bone mineral density (BMD) and T-score from chest X-ray images. In the processing described in Non-Patent Document 3, image deformation / brightness / saturation / hue image expansion is performed.
[0007] Morteza Heidari, et. al., “Improving the performance of CNN to predict the likelihood of COVID-19 using chest X-ray images with preprocessing algorithms”, International Journal of Medical Informatics, 144, 104284, 2020. Mohamed Elgendi, et al., “The Effectiveness of Image Augmentation in Deep Learning Networks for Detecting COVID-19: A Geometric Transformation Perspective”, Front. Med., Translational Medicine, 8-202, 2021. Yoichi Sato, et al., “Deep Learning for Bone Mineral Density and T-Score Prediction from Chest X-rays: A Multicenter Study”, Biomedicines., 10(9), 2323, 2022.
[0008] In order to solve the above problem, a generation method according to one aspect of the present disclosure includes an acquisition step of acquiring a first image that includes at least a portion of a subject, and a generation step of performing image processing on the first image to generate a training image to be used for training a machine learning model.
[0009] In addition, a generation device according to one aspect of the present disclosure is a generation device that generates training images to be used for training a machine learning model that makes a predetermined estimation from a target image, and includes an acquisition unit that acquires a first image that contains at least a portion of a subject, and a generation unit that performs image processing on the first image to generate the training image.
[0010] In addition, a learning system according to one aspect of the present disclosure is a learning system comprising a generation device that generates learning images used to train a machine learning model that makes a predetermined inference from a target image, and a learning device that trains the machine learning model, and comprises an acquisition unit that acquires a first image that contains at least a portion of a subject, a generation unit that performs image processing on the first image to generate the learning image, an output unit that outputs the generated learning image to the learning device, a memory unit that stores a parameter set that defines the machine learning model, and a learning unit that trains the machine learning model.
[0011] The generation device and learning device according to each aspect of the present disclosure may be realized by a computer. In this case, the present disclosure also includes a control program for the generation device that causes the computer to operate as each unit (software element) of the generation device or the learning device, thereby realizing the generation device or the learning device on the computer, and a computer-readable recording medium on which the control program is recorded. The computer may also include an edge computer, a cloud server, etc. Furthermore, the functions of the generation device and the functions of the learning device may each be realized by multiple devices. For example, the present disclosure also includes an aspect in which a first generation device performs the image correction described below and a second generation device performs the image conversion described below.
[0012] 1 is an example of a block diagram showing the configuration of a learning system according to embodiment 1. FIG. 2 is an example of a flowchart showing the flow of a method for generating medical images for machine learning. FIG. 3 is a schematic diagram showing an example of the flow of clustering processing performed by a selection unit. FIG. 4 is a diagram showing an example of a luminance histogram before and after image correction processing. FIG. 5 is a diagram showing an example of image conversion processing by an image conversion unit. FIG. 6 is a diagram showing an example of image conversion processing by the image conversion unit. FIG. 7 is an example of a block diagram showing the configuration of a learning system according to embodiment 2. FIG. 8 is a diagram showing an example of image conversion processing by the image conversion unit and a second image conversion unit. FIG. 9 is a diagram showing an example of an image in which the second conversion unit has performed processing to erase a partial region included in an area captured in an image that has undergone image conversion processing by the image conversion unit. FIG. 10 is an example of a block diagram showing the configuration of a computer that executes instructions of a program that realizes the control block of the generation device. FIG. 11 is an example of a chest X-ray image of a human. FIG. 12 is an example of a flowchart for X-ray images included in a dataset. FIG. 13 is an example of a schematic diagram of a workflow for lumbar BMD estimation. FIG. 14 is an example of a BMD estimation result by gender. FIG. 15 is an example of a BMD estimation result by BMI category. FIG. 16 is an example of a typical attention heat map overlaid on a lumbar X-ray image using an ANN. 1 is an example of a flowchart for lumbar spine X-ray images included in the training and test datasets; 2 is an example of a schematic diagram of an osteoporosis classification flowchart; 3 is an example of a result of lumbar spine BMD estimation and osteoporosis classification by gender and BMI category; 4 is an example of a result of femur BMD estimation and osteoporosis classification by gender and BMI category; 5 is an example of a typical attention heat map based on estimation results using an ANN; 6 is an example of a flowchart for chest X-ray images included in the training and test datasets; 7 is an example of a schematic diagram of an osteoporosis classification flowchart; 8 is an example of a result of lumbar spine BMD estimation and osteoporosis classification by gender and BMI category; 9 is an example of a result of femur BMD estimation and osteoporosis classification by gender and BMI category; 10 is an example of a typical attention heat map based on estimation results using an ANN;
[0013] [Embodiment 1] A medical image generating device 1 (hereinafter also simply referred to as "generating device 1") according to an embodiment of the present disclosure will be described in detail below with reference to the drawings. First, what the generating device 1 does will be described.
[0014] Methods for pre-processing images for training a machine learning model to improve the learning effect include the above-mentioned image preprocessing method for correcting images, the image enhancement method, and a method for selecting training images. However, when images that capture at least a portion of a subject are used as training images, there is a risk of reducing the accuracy of the machine learning model that performs a predetermined estimation. For example, when the subject is a bone, there is a risk of reducing the accuracy of the machine learning model that estimates bone density if an image in which the bone portion is blown out or an image in which the bone is poorly captured is used.
[0015] A medical image generating apparatus 1 according to an embodiment of the present disclosure generates training images that can reduce the risk of a decrease in the accuracy of a machine learning model when images that include at least a portion of a subject are used as training images. When images that include at least a portion of a subject are used as training images, the medical image generating apparatus 1 can generate training images that improve the estimation accuracy of the machine learning model. For example, when using a machine learning model to estimate bone density of a bone that is a subject, the medical image generating apparatus 1 can generate training images that improve the estimation accuracy of the machine learning model.
[0016] The three images shown in Figure 11 are examples of chest X-ray images of a human subject. In image 111, in addition to human bones, the upper right region 1111 shows the letter "L," which indicates the left side (left shoulder), and other characters indicating other information. In image 112, the upper left region 1121 shows an arrow indicating the direction toward the head and other characters indicating other information. In addition, the lower region 1122 shows soft tissue below the diaphragm. In image 113, the upper right region 1131 shows a symbol indicating specific information, and the region 1132 below shows an artificial object.
[0017] Thus, in actual medical images, in addition to the object of the subject, intentionally inserted information, human tissue other than the part where it is desired to determine the presence or absence of an abnormality, artificial objects such as pacemakers or necklaces, or lesions occurring in parts other than the target part may be captured. In other words, the above-mentioned objects and intentionally inserted information are examples of subjects contained in each image such as a medical image.
[0018] The target object may be, for example, a target tissue and / or a target organ. The target tissue may be, for example, at least one of epithelial tissue, connective tissue, muscle tissue, and nervous tissue. The target organ may be, for example, at least one of the digestive system, the cardiovascular system, the endocrine system, and the musculoskeletal system. The target object may be, for example, a joint and / or a bone. As mentioned above, the image quality of medical images varies depending on the imaging device, imaging method, imaging conditions, and the like. When training a machine learning model, the entire image, including unnecessary parts of the actual medical image (original medical image), is used for training, and therefore training is performed including various elements unrelated to the target training. Therefore, using the original image as is for training may result in a decrease in training accuracy.
[0019] The medical image may be, for example, an image of the subject captured with an endoscope. More specifically, the medical image may include an endoscopic image of at least one of the subject's nasal cavity, esophagus, stomach, duodenum, rectum, large intestine, small intestine, anus, and colon. Medical images of these areas may output analysis results that clearly indicate, using a learning model, areas of interest that include at least one of inflammation, polyps, and cancer. In such cases, the learning model may be, for example, a model trained based on first learning images including images of the areas of interest and first training data indicating the presence of the areas of interest, and second learning images including images without the areas of interest and second training data indicating the absence of the areas of interest. The first training data may include information indicating the degree of inflammation (degree of inflammation) or malignancy (degree of malignancy) of the areas of interest. The analysis result may, for example, be displayed by surrounding the areas of interest, pointing to the areas of interest, or superimposing a color on the areas of interest. Derivation basis information indicating the basis for deriving the analysis result may be displayed together with the analysis result.
[0020] Alternatively, the medical images may be, for example, images of the subject's eyes, skin, etc. captured with a digital camera. Medical images of these areas may output analysis results that clearly indicate signs of interest using a learning model. For example, if the signs of interest are of the eye, they may include signs indicating diseases such as at least one of glaucoma, cataracts, age-related macular degeneration, conjunctivitis, hordeolum, retinopathy, and blepharitis. Alternatively, if the signs of interest are of the skin, they may include signs such as skin cancer, hives, atopic dermatitis, and herpes. The analysis results may be displayed as a display that surrounds these signs of interest, a display that indicates the signs of interest, a display that superimposes a color on the signs of interest, or a display that indicates the name of the disease. For example, the learning model may be a model trained based on first training images including images of these areas of interest and first training data indicating the presence of signs of interest, and second training images including images without signs of interest and second training data indicating the absence of signs of interest. Along with such analysis results, derivation basis information indicating the basis on which the analysis results were derived may be displayed.
[0021] The generating device 1 generates training images used for training a machine learning model that performs predetermined estimations from target images. In one aspect, the generating device 1 generates medical images for machine learning to train a machine learning model that estimates the bone density of a bone from a target image that shows at least a portion of the bone. Examples of major bone densities include the bone densities of the lumbar vertebrae, femur, calcaneus, and radius. The machine learning model reads the training medical images and trains so that its output matches the specific bone densities associated and annotated with the medical images.
[0022] Here, the target image is an image in which at least a portion of a predetermined object to be estimated is captured as a subject. The target image may be an X-ray image including a plain X-ray image, an MRI image, a CT scan image, or the like. Furthermore, the subject captured in the target image may be the same as the subject. The captured regions of the target image and the learning images may include, for example, at least one of the head, neck, chest, lower back, temporomandibular joint, spinal intervertebral joint, hip joint, sacroiliac joint, knee joint, ankle joint, foot, toes, shoulder joint, acromioclavicular joint, elbow joint, wrist joint, hand, fingers, and temporomandibular joint. Furthermore, the target image may be a frontal image or a lateral image.
[0023] The X-ray image data may be simple X-ray images such as lumbar X-ray images or chest X-ray images, or may be X-ray images taken with a DXA (Dual Energy X-ray Absorptiometry) device. When measuring bone density in the lumbar vertebrae with a DXA device, X-rays are irradiated from the front of the subject's lumbar vertebrae. When measuring bone density in the proximal femur with a DXA device, X-rays are irradiated from the front of the subject's proximal femur. Here, "front of the lumbar vertebrae" and "front of the proximal femur" refer to the direction that correctly faces the imaging site, such as the lumbar vertebrae or the proximal femur, and may be the ventral side or the back side of the subject's body. In the MD (micro densitometry) method, X-rays are irradiated to the hand. The image data does not have to be X-ray images; it can be any image containing bone information. For example, it can be estimated from magnetic resonance imaging (MRI) images, computed tomography (CT) images, positron emission tomography (PET) images, ultrasound images, and the like.
[0024] The machine learning model may learn bone density, etc., of bones not shown in training images that show at least a portion of the bone, or may estimate bone density, etc., of bones not shown in target images that show at least a portion of the bone, from the target images. Additionally, the machine learning model may be configured to receive input images of a predetermined portion of the subject's body and learn and estimate the bone density of the subject. The machine learning model may also be configured to receive input information about the subject, such as age and gender, as well as each image, and learn and estimate. The machine learning model may also be configured to determine whether a subject has osteoporosis. The machine learning model may determine osteoporosis based on, for example, proprietary standards or already known guidelines.
[0025] For example, the machine learning model may learn bone density of the femur and / or lumbar bones from training images showing chest bones, or may estimate bone density of the femur and / or lumbar bones from target images showing chest bones.
[0026] For example, the machine learning model may learn the bone density of the femur and / or lumbar bones from a training image showing lumbar bones, or may estimate the bone density of the femur and / or lumbar bones from a target image showing lumbar bones.
[0027] Additionally, the machine learning model may be configured to receive input of images of a predetermined part of the subject's body, and to learn and estimate the subject's bone density. Also, the machine learning model may be configured to receive input of not only each image but also information about the subject, such as age and gender, and to learn and estimate.
[0028] The generation device 1 is a device that generates learning images that serve as training data for a machine learning model.
[0029] Bone density is expressed as bone mineral density per unit area (g / cm 2 ), bone mineral density per unit volume (g / cm 3The bone mineral density may be expressed by at least one of the following: YAM (%), T-score, and Z-score. YAM (%) is an abbreviation for "Young Adult Mean" and may be referred to as the young adult mean percentage. For example, bone mineral density may be expressed by bone mineral density per unit area (g / cm 2 The bone mineral density may be an index determined by the guidelines or an original index.
[0030] The bone density may be determined using a DXA device, an X-ray device, an ultrasonic bone density measuring device, etc. Alternatively, the bone density may be a value estimated using a device that estimates bone density from a plain X-ray image.
[0031] Furthermore, the description of the generating device 1 in the present disclosure is not limited to a machine learning model that estimates bone density, but can also be applied to a machine learning model that estimates bone abnormalities. The bone abnormality may be, for example, a bone abnormality related to a bone disease. The bone disease may be, for example, osteoporosis, scoliosis, fracture, spinal stenosis, intervertebral disc degeneration, ankylosing spondylitis, spinal cord injury, osteomyelitis, osteophyte, spinal muscular atrophy, etc.
[0032] Furthermore, the predetermined estimation is not limited to estimation of bone density, but may also be estimation of the component ratio of bone or other substances. Furthermore, the machine learning model may predict a time point different from the current estimation result of the subject estimated from the image data. The different time point may be a future and / or past time point relative to the time point at which the image data was captured. In the following explanation, the generating device 1 will be described as performing each process on a medical image, but the present invention is not limited to this, and may also be applied to images related to biology, engineering, or the like.
[0033] Furthermore, the description of the machine learning model in this disclosure is not limited to a configuration that makes a predetermined estimation from medical images, etc., but can also be applied to a machine learning model that determines the identity of an object appearing in an input image.
[0034] FIG. 1 is an example block diagram showing the configuration of a learning system 100 including a generation device 1 according to the present embodiment. As shown in FIG. 1, the learning system 100 includes a generation device 1, a display device 50, and a learning device 70. The generation device 1 includes a control unit 10, a memory unit 20, and an input / output interface (I / O IF) 30. The control unit 10 includes one or more processors (not shown) and controls the generation device 1. The memory unit 20 is a storage device that stores various data, and may store original medical image data (original data) 21 and converted image data (learning image data) 22, which is the original medical image subjected to image conversion processing (described below). Here, the original medical image data refers to the raw image data captured in a medical setting. The memory unit 20 may also store various control programs that are read and executed by the processor. The I / O IF 30 is an interface for communicating with the outside world. The I / O IF 30 may be a wired communication interface, such as a USB port, or a wireless communication interface, such as Bluetooth (registered trademark) or Wi-Fi (registered trademark). Furthermore, a user may be able to input data to the generation device 1 via a mouse, keyboard, or the like connected to the input / output IF 30. Here, the user refers to a medical professional, technician, or the like who uses the generation device 1, and the same applies hereinafter. Furthermore, the generation device 1 may be connected to the Internet via the input / output IF 30 to communicate information.
[0035] The present disclosure also includes a configuration in which the generation device 1 does not have an input / output IF 30 and is integrated with a display device 50 capable of displaying medical images, etc., or a learning device 70.
[0036] The components of the control unit 10 are described below. The control unit 10 includes an acquisition unit 11, a selection unit 12, an image correction unit 13, an image conversion unit 14, and an output unit 16. The acquisition unit 11 acquires an original medical image. The original medical image may be, for example, an X-ray image. The captured region of the original medical image may include at least one of the head, neck, chest, lower back, hip joint, knee joint, ankle joint, foot, toe, shoulder joint, elbow joint, wrist joint, hand, fingers, and temporomandibular joint. The original medical image may be, for example, a plain X-ray image including at least a portion of a bone such as the sternum, lumbar vertebrae, or hip joint. The acquisition unit 11 may acquire the original medical image from the storage unit 20 or from an external device or storage medium connected to the generation device 1 via the input / output IF 30. The original medical image and the training image generated from the medical image are annotated with bone density information indicating the bone density of the bones depicted in the medical image.
[0037] The selection unit 12 selects original medical images containing bones acquired by the acquisition unit 11 using a method for evaluating the quality or commonality of the images. The selection unit 12 may select images that satisfy predetermined criteria. In one aspect, the selection unit 12 basically selects images whose quality is evaluated to be better than a standard or whose commonality is evaluated to be higher than a standard. The reason for selecting images with high commonality is that images that are clustered into groups with low commonality with a large number of images are different from typical images and are therefore considered unsuitable for learning a machine learning model.
[0038] As an example, the selector 12 evaluates the commonality of images by combining t-Distributed Stochastic Neighbor Embedding (t-SNE) and Density-based spatial clustering of applications with noise (DBSCAN). In this case, t-SNE may be performed before or after DBSCAN. However, the method by which the selector 12 evaluates the commonality of images is not limited to a specific method, and another method may be performed between t-SNE and DBSCAN. FIG. 3 is a schematic diagram showing an example of the flow of the clustering process performed by the selector 12. The selector 12 performs a fast Fourier transform (FFT) on an original medical image 301 to obtain a spectral image 302. Next, the selector 12 dimensionally compresses the spectral image 302 into an image 303 using t-SNE. Here, dimensionality compression is a process of reducing multidimensional data to a lower number of dimensions so as to minimize the loss of features of the original information. Next, the selection unit 12 clusters the image 303 using DBSCAN. As a result, the clustering result shown in image 304 is obtained. In image 304, it can be seen that there is the largest cluster 3041, small clusters 3042 and 3043, and data that does not form clusters. A cluster to which many pieces of data belong is determined to have high commonality. Therefore, the selection unit 12 selects only the data included in the largest cluster 3041 and does not select other data.
[0039] Furthermore, images not selected by the selection unit 12 do not need to be used in the subsequent process of generating learning images. From another perspective, the acquisition unit 11 may also acquire images that are not actually used to generate learning images. Furthermore, an image that contains at least a portion of the object to be the subject and is actually used to generate learning images is an example of a first image in the present disclosure.
[0040] In other words, the acquisition unit 11 acquires a plurality of specimen images, including a first image that shows the bones of the subject and is actually used to generate a learning image, and the selection unit 12 selects a first image from the plurality of specimen images.
[0041] The control unit 10 may also have a function of recognizing a predetermined vertebral body in an image, and the acquisition unit 11 or another component included in the control unit 10 may trim the image at any timing so that at least the predetermined vertebral body is included in the image. The trimming range specified by the control unit 10 may be a preset range, or may be a range specified each time by the user via the input / output IF 30.
[0042] In the above-described clustering process, the selection unit 12 may classify the plurality of specimen images into a plurality of groups based on the distribution of pixel values of each of the plurality of specimen images acquired by the acquisition unit 11, and select a part of the groups that includes at least the first image. The plurality of groups may be classified based on, for example, a predetermined criterion.
[0043] Furthermore, in the clustering process, the selection unit 12 may classify the plurality of specimen images into a plurality of groups based on the result of dimensional compression of the result of analyzing the pixel value distribution of each of the plurality of specimen images.
[0044] The image quality evaluation method may use at least one of the luminance mean, luminance variance, and BRISQUE (Blind / Referenceless Image Spatial Quality Evaluator). For example, in a method using the luminance mean and luminance variance, the selector 12 determines that an image whose values fall within a range defined by a predetermined threshold is of high quality. In other words, the selector 12 does not select an image in which any of the values falls outside the range and is an outlier. BRISQUE is a method for evaluating the quality of an image using features that can classify the magnitude or type of image distortion. Since X-ray images have a certain range of luminance values, X-ray images with a luminance value range that is greater than a certain value are determined to have poor image quality. In a method using BRISQUE, images with a predetermined score or less are determined to be of high quality. The image selection method used by the selector 12 may be preset, or may be selectable by the user via the input / output IF 30 each time a learning image is generated.
[0045] The image correction unit 13 performs a predetermined image correction process on at least a part of the original medical image selected by the selection unit 12. Here, the image correction unit 13 may perform the image correction process on the entire medical image, or may perform the image correction process on only a part of the medical image, such as the central, upper, or lower part.
[0046] The image correction unit 13 may perform image correction using a trained image generation model. Here, when a first image is input, the image correction unit 13 may generate an image that approximates the first image to a predetermined image using the trained image generation model.
[0047] Furthermore, the predetermined image correction process may be at least one of pixel value normalization, black and white inversion, CLAHE, white balance correction using Retinex with adjust processing, unsharp masking, Detail Enhance, histogram equalization, edge enhancement, noise removal, sharpening processing, Registration, and image size adjustment.
[0048] Here, pixel value normalization is a process of reducing the range of possible pixel values to a fixed range. Furthermore, CLAHE is a process of equalizing the histogram for each region of an image to adjust the contrast. Retinex (with adjust) processing is a process of correcting the colors of objects contained in an image to make them as clear as they are seen by the human eye. Unsharp masking is a filter process that emphasizes unclear edges in an image. Detail Enhance is a process that emphasizes the details of an image. Histogram equalization is a process of correcting the histogram of pixel values to make it flatter. Registration is a process of correcting the positional deviation of multiple images.
[0049] These image correction processes reduce the visual variability of medical images, making bone features clearer and easier to see.
[0050] 4, image 401 illustrates an example of a luminance histogram of an original medical image before image correction, and image 402 illustrates an example of a luminance histogram of the same medical image after white balance correction and luminance normalization have been performed as image correction processes. As shown by the vertical and horizontal axes of images 401 and 402, the bias and variation in luminance of the medical image before and after this image correction process have been reduced, i.e., the difference in the number of pixels for each luminance value has been reduced.
[0051] Furthermore, the normalization process of pixel values may include normalization process of brightness of each of the plurality of sample images and Retinex process.
[0052] The image correction described above can remove noise and adjust the balance of brightness, hue, etc. while preserving the characteristics of the original image. Images corrected in this way become images more suitable for training a machine learning model. The type of image correction performed by the image correction unit 13 may be preset, or may be selectable by the user via the input / output IF 30 each time a training image is generated.
[0053] The image conversion unit 14 generates a new medical image to be used as a learning image by performing a predetermined image conversion process on at least a part of the first image that has been subjected to the predetermined image correction by the image correction unit 13. The predetermined image conversion process may be, for example, at least one of brightness change, contrast change, gamma conversion, optical distortion, CLAHE, Channel Shuffle, blurring (blurring), cropping, rotation, inversion, and scaling.
[0054] Here, gamma conversion is a process of applying a function, which takes a gamma value indicating the relationship between pixel values and brightness as an argument, to an image. Channel shuffle is a process of rearranging channels such as RGB.
[0055] Furthermore, the predetermined image conversion process may be configured to randomly execute at least one of the processes described above. The brightness or contrast of the image may also be changed before and after the image conversion process.
[0056] Furthermore, the image conversion unit 14 may convert, for example, at least a part of the image or the entire image of the first image that has been subjected to predetermined image correction by the image correction unit 13. Furthermore, the image conversion unit 14 may convert the images by uniformly applying a strength and / or type of image conversion, or may randomly change the strength and / or type of image conversion for each image.
[0057] The image conversion unit 14 may perform image conversion using a trained image generation model. Here, when a first image is input, the image conversion unit 14 may generate an image that approximates the first image to a predetermined image using the trained image generation model.
[0058] 5 and 6 are diagrams showing an example of image conversion processing by the image conversion unit 14. Image 501 in Fig. 5 is an example of an X-ray image of a human lumbar spine. Images 502 and 503 are examples of images obtained by performing a predetermined image conversion process on image 501. In images 502 and 503, processing for randomly changing the brightness and / or contrast of the images and blurring processing have been performed at positions and to different degrees between the images.
[0059] 6 is an example of a chest X-ray image of a person. Images 602 and 603 are examples of images obtained by performing a predetermined image conversion process on image 601 by image conversion unit 14. Specifically, image 602 is an image obtained by performing a rotation process on image 601, and image 603 is an image obtained by performing an enlargement process on image 601.
[0060] The image conversion process (Augmentation) described above generates a large number of new images under various different conditions. That is, the image conversion unit 14 generates one or more medical images by randomly converting a single medical image within a predetermined parameter range.
[0061] By training the machine learning model using a large number of new training images, the accuracy of the machine learning model can be improved. Furthermore, as described above, by associating the same bone density with multiple randomly converted images and using them for training the machine learning model, the influence of brightness differences, contrast differences, noise differences, etc., in the target image on the estimation of bone density can be reduced. The training images may be images generated by the image correction unit 13 or may be images generated by the image conversion unit 14.
[0062] The type of image conversion performed by the image conversion unit 14 may be set in advance, or may be selectable by the user via the input / output IF 30 each time a learning image is generated.
[0063] The image correction process by the image correction unit 13 and the image conversion process by the image conversion unit 14 and a second image conversion unit 15 described later are each an example of image processing in the present disclosure.
[0064] The output unit 16 outputs the new medical image generated by the image conversion unit 14 to the storage unit 20 or to an external device. The external output destination may be the display device 50, the learning device 70, a storage device such as a database, or another image processing device. Data may also be transmitted to an external device via the Internet.
[0065] The learning device 70 is a device for realizing a machine learning model and learning thereof. The learning device 70 includes a storage unit 701 that stores a parameter set that defines the machine learning model 702, and a learning unit 703 that trains the machine learning model 702. The learning unit 703 trains the machine learning model 702 by updating the values of the parameter set stored in the storage unit 701.
[0066] The machine learning model 702 may be generated in advance by the learning unit 703 using the training images generated by the generating device 1 as explanatory variables and bone density information indicating the bone density of the subject's bones corresponding to the training images as a target variable. The machine learning model 702 estimates the bone density of bones appearing in a target image that shows at least a portion of a bone and / or the bone density of bones not appearing in the target image from the target image.
[0067] Next, the flow of the method S1 for generating medical images for machine learning according to this embodiment will be described. Fig. 2 is an example of a flowchart showing the flow of the method S1 for generating medical images for machine learning, which is executed by the generating device 1.
[0068] In S11, the acquisition unit 11 acquires a plurality of original medical images, which are specimen images, from the storage unit 20 or the like (acquisition step).
[0069] In S12, the selection unit 12 evaluates the quality of each image acquired by the acquisition unit 11 and selects a medical image to be used for training the machine learning model as the first image (selection step).
[0070] In S13, the image correction unit 13 performs a predetermined image correction process on at least a part of the medical images selected by the selection unit 12. Furthermore, the image correction unit 13 may perform the normalization process described above as the predetermined image correction (normalization step).
[0071] In S14, the image conversion unit 14 generates new medical images to be used as learning images (generation step) by performing a predetermined image conversion process on at least a portion of the medical images that have been subjected to the predetermined image correction by the image correction unit 13. However, the above explanation does not mean that medical images that have not been subjected to the image conversion process by the image conversion unit 14 cannot be used as learning images.
[0072] In addition, the learning image generated by the image conversion unit 14 is output by the output unit 16 to the memory unit 20, or to a device that stores a parameter set that defines the machine learning model 702 that is the learning target, together with bone density information indicating the bone density that was annotated on the original medical image (output step).
[0073] In addition, in training the machine learning model 702, training images obtained by performing image conversion processing on a first image showing the bones of the subject may be used as explanatory variables, and bone density information indicating the bone density of the bones of the subject corresponding to the training images may be used as a response variable. In training the machine learning model 702, training images obtained by performing image conversion processing on a first image showing the bones of the subject may be used as explanatory variables, and values indicating the hormone concentration, vitamin D concentration, etc. of the subject corresponding to the training images may be used as a response variable.
[0074] The first image may be an image taken from the front of the subject, or may be an image taken from the side.
[0075] 2, the method S1 for generating medical images for machine learning uses medical images that have been subjected to image conversion processing such as blurring for machine learning, thereby reducing the influence of brightness differences and the like in the target image on estimation by the machine learning model. In other words, the method S1 makes it possible to generate learning images that can reduce the risk of a decrease in the accuracy of a machine learning model when estimating the bone density of bones shown in the target image using the machine learning model.
[0076] In addition, estimation using a machine learning model may be performed, for example, by the control unit 10 of the medical image generating device 1 via the input / output IF 30, or by a device not shown in the figure that is provided in the learning system 100, or by the learning device 70.
[0077] [Embodiment 2] A second embodiment of the present disclosure will be described below. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and redundant description will not be repeated.
[0078] 7 is an example of a block diagram showing the configuration of a learning system 100A including a medical image generating device 1A (hereinafter simply referred to as "generating device 1A") according to this embodiment. As shown in FIG. 7, the generating device 1A includes a control unit 10A in addition to the generating device 1 shown in FIG. 1, and further includes a second image conversion unit 15.
[0079] 2, the second image conversion unit 15 performs its own image conversion process following the image conversion process by the image conversion unit 14. The second image conversion unit 15 performs image conversion process including a process of erasing a partial area included in an area captured in the image or a process of changing the pixel values of the partial area to predetermined pixel values, thereby generating a learning image. Here, the second image conversion unit 15 may generate multiple learning images by performing a process such as erasing partial areas at different random positions on a single image that has been subjected to image conversion process by the image conversion unit 14.
[0080] Fig. 8 is a diagram showing an example of image conversion processing by the image conversion unit 14 and the second image conversion unit 15. Image 801 in Fig. 8 is an example of a chest X-ray image of a person. Images 802 and 803 are examples of images obtained by performing a predetermined image conversion processing on image 801 by the image conversion unit 14. Specifically, image 802 is an image obtained by performing a random translation processing on image 801, and image 803 is an image obtained by performing a random horizontal flip processing on image 801. The translation processing is a processing for moving the position of the object serving as the subject in the image to an arbitrary position while maintaining the angle of the object. The horizontal flip processing is a processing for flipping the image around a vertical axis at an arbitrary position as the rotation axis.
[0081] In contrast, image 804 is an image obtained by performing a process on image 801 by the second image conversion unit 15 to erase a partial area 8041 included in the area depicted in the image. Here, the erasure process means, for example, setting the pixel values of pixels included in the target area to zero or an undefined value.
[0082] 9 is a diagram showing an example of an image in which the second image conversion unit 15 has further processed an image that has been subjected to image conversion processing by the image conversion unit 14, and erased a partial area included in the area depicted in the image. That is, the image 901 shown in FIG. 9 is an image generated by combining the image conversion processing by the image conversion unit 14 and the image conversion processing by the second image conversion unit 15.
[0083] The image conversion process by the second image conversion unit 15 may include at least one of horizontal movement processing, vertical movement processing, rotation processing, inversion processing, enlargement processing, reduction processing, and affine transformation of the image after the image conversion process by the image conversion unit 14, or may be a combination of two or more of these processes.
[0084] Parallel translation processing includes, for example, horizontal translation processing and vertical translation processing. The parallel translation processing may, for example, move an object so that it extends beyond the outer edge of the image before the translation processing. In this case, the pixel values of the pixels in the portion of the image that has been removed by the translation processing may be set to zero or an undefined value, for example. Alternatively, the parallel translation processing may, for example, trim a portion of the image before the translation processing and move it within the outer edge of the original image. In this case, the pixel values of the pixels in the portion of the image other than the trimmed portion may be set to zero or an undefined value, for example. The horizontal translation processing is processing that moves the horizontal position (horizontal or left-right position) of an object in an image to a desired position while maintaining the angle of the object. The vertical translation processing is processing that moves the vertical position (vertical or up-down position) of an object in an image to a desired position while maintaining the angle of the object.
[0085] Rotation processing is a process of changing the angle of an object in an image to an arbitrary angle. Inversion processing is a process of inverting an image around an arbitrary axis as the rotation axis. Enlargement processing is a process of increasing the apparent size of an object in an image. Reduction processing is a process of decreasing the apparent size of an object in an image. Affine transformation is a process of transforming coordinates in an image by multiplying the coordinates by a determinant, thereby enlarging or reducing, rotating, or translating the image. The determinant is, for example, a 3x3 square matrix. Through affine transformation, the coordinates (x, y) of each point in the image are moved to a unique point (x', y'). The image transformation unit 14 and the second image transformation unit 15 may perform image transformation processing on the entire image or on a portion of the image. The portion of the image on which image transformation processing is performed may or may not include at least a portion of the object, such as a bone.
[0086] The second image conversion unit 15 may fill partial regions of any size with black, white, or other colors, or with random noise or grid noise. Random noise is noise in which each pixel value has random brightness. Grid noise is noise that mainly consists of geometric patterns such as a grid. The filling includes masking a figure such as black in front of the partial region. The second image conversion unit 15 may erase or fill partial regions of any shape, such as a rectangle, circle, or triangle. The number of partial regions erased or filled by the second image conversion unit 15 may be one or more. The partial region may overlap at least a portion of an object such as a bone, or may not overlap. The sizes of the partial regions may be the same or different from each other.
[0087] Conventionally, unlike lumbar spine X-ray images, chest X-ray images use the entire X-ray image for learning, without being limited to just the vertebrae. This means that medical images often contain characters indicating the shooting direction and soft tissue other than bone, which can lead to a decrease in accuracy.
[0088] According to the configuration of this embodiment, it is possible to randomly generate pseudo-learning images containing error data, and to construct a learning model that is resistant to missing bone images due to information other than bones appearing in the target image.
[0089] Furthermore, the control unit 10A may determine whether the specimen image is a lumbar spine X-ray image or a chest X-ray image by referring to annotations associated with the specimen image or by using machine learning. Subsequently, the generating device 1A may be configured to perform the processing of embodiment 1 if the specimen image is a lumbar spine X-ray image, and to perform the processing of this embodiment including image conversion processing by the second image converter 15 if the specimen image is a lumbar spine X-ray image.
[0090] In the present disclosure, the subject is described as a human, but the subject is not limited to a human. The subject may be, for example, a non-human mammal such as an equine, feline, canine, bovine, or porcine animal, or may be an animal other than a mammal (e.g., a bird, reptile, amphibian, or fish). In other words, the target image in the present disclosure may be an image showing an animal's bones.
[0091] [Example of Implementation by Software] The control blocks of the medical image generating device 1 (1A) (particularly the acquisition unit 11, selection unit 12, image correction unit 13, image conversion unit 14, second image conversion unit 15, and output unit 16) may be realized by a logic circuit (hardware) formed on an integrated circuit (IC chip) or the like, or may be realized by software. In the latter case, each function of the generating device 1 is realized, for example, by a computer that executes instructions of a program P, which is software.
[0092] An example of such a computer (hereinafter referred to as computer C) is shown in Figure 10. Computer C includes at least one processor C1 and at least one memory C2. Memory C2 stores a program P for causing computer C to operate as generation device 1. In computer C, processor C1 reads and executes program P from memory C2, thereby realizing each function of generation device 1.
[0093] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PU), a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0094] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input devices such as a keyboard and a mouse, and / or output devices such as a display and a printer.
[0095] The program P can be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communications network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0096] Some or all of the functions of the control blocks can be realized by logic circuits. For example, the scope of the present disclosure also includes integrated circuits in which logic circuits that function as the control blocks are formed. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.
[0097] The processes described in the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may run on the control device or on another device (for example, an edge computer or a cloud server).
[0098] The invention according to the present disclosure has been described above based on the drawings and examples. However, the invention according to the present disclosure is not limited to the above-described embodiments. In other words, the invention according to the present disclosure can be modified in various ways within the scope of the present disclosure, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the invention according to the present disclosure. In other words, it should be noted that a person skilled in the art can easily make various modifications or corrections based on the present disclosure. It should be noted that these modifications or corrections are included in the scope of the present disclosure.
[0099] [Summary] The generation method according to aspect 1 of the present disclosure includes an acquisition step of acquiring a first image that includes at least a portion of a subject, and a generation step of performing image processing on the first image to generate a training image to be used in training a machine learning model.
[0100] A generation method according to aspect 2 of the present disclosure may be a method in which, in the above-mentioned aspect 1, the acquisition step acquires a plurality of specimen images including the first image, and before the generation step, further includes a selection step of classifying the plurality of specimen images into a plurality of groups based on the distribution of pixel values of each of the plurality of specimen images, and selecting at least a portion of the groups including the first image.
[0101] A generation method according to aspect 3 of the present disclosure may be a method in the above-described aspect 2, in which in the selection step, the plurality of specimen images are classified into a plurality of groups based on the results of dimensional compression of the results of analyzing the pixel value distribution of each of the plurality of specimen images.
[0102] The generation method according to aspect 4 of the present disclosure may be a method in aspect 2 or 3 above, further including, between the acquisition step and the generation step, a normalization step in which a pixel value normalization process is performed on each of the plurality of sample images.
[0103] A generation method according to a fifth aspect of the present disclosure may be a method according to the fourth aspect, wherein the normalization process includes normalization of the brightness of each of the plurality of specimen images and / or Retinex processing.
[0104] A generation method according to aspect 6 of the present disclosure may be a method in any of aspects 1 to 5 above, in which the image processing includes at least one of a process of changing the brightness of the first image, a process of changing the contrast, and a blur process.
[0105] A generation method according to aspect 7 of the present disclosure may be a method according to any one of aspects 2 to 5 above, wherein the specimen image is a plain X-ray image.
[0106] A generation method according to aspect 8 of the present disclosure may be a method in which, in any of aspects 1 to 7 above, the subject is a bone, and the machine learning model estimates the bone density of a bone that appears in a target image in which at least a portion of the bone appears, and / or the bone density of a bone that does not appear in the target image, from the target image in which at least a portion of the bone appears.
[0107] A generation method according to aspect 9 of the present disclosure may be a method in which, in the above-mentioned aspect 8, the machine learning model estimates bone density of the femur and / or lumbar bones from a target image showing chest bones.
[0108] A generation method according to aspect 10 of the present disclosure may be a method in which, in aspect 8 above, the machine learning model estimates bone density of the femur and / or lumbar bones from a target image showing lumbar bones.
[0109] A generation method according to aspect 11 of the present disclosure may be a method according to any of aspects 8 to 10 above, wherein the bone mineral density of the bone is at least one of bone mineral density per unit area, bone mineral density per unit volume, percent of young adult mean, T-score, and Z-score.
[0110] A generation method according to aspect 12 of the present disclosure may be a method in which, in any of aspects 1 to 11 above, in the generation step, the image processing includes a process of erasing a partial area included in an area that appears in the image after the image processing, or a process of changing the pixel value of the partial area to a predetermined pixel value, and the learning image is generated.
[0111] A generation method according to aspect 13 of the present disclosure may be a method in aspect 12 above, in which the image processing includes at least one of horizontal movement processing, rotation processing, inversion processing, enlargement processing, and reduction processing.
[0112] A learning method according to aspect 14 of the present disclosure is a learning method for a machine learning model that estimates the bone density of a bone from a target image that shows the bone, and uses a training image generated by the generation method described in any of aspects 8 to 11 above as an explanatory variable, and uses bone density information indicating the bone density of the bone of a subject corresponding to the training image as a dependent variable.
[0113] A generation method according to aspect 15 of the present disclosure is a method for generating a machine learning model that estimates the bone density of a bone from a target image in which the bone is captured, the method using training images generated by the generation method described in any of aspects 8 to 11 above as explanatory variables and bone density information indicating the bone density of the bone of a subject corresponding to the training images as a target variable to generate a machine learning model that estimates the bone density of a bone captured in a target image in which at least a portion of the bone is captured and / or the bone density of a bone not captured in the target image.
[0114] A generation device according to aspect 16 of the present disclosure is a generation device that generates training images to be used in training a machine learning model that makes predetermined inferences from a target image, and is configured to include an acquisition unit that acquires a first image that contains at least a portion of a subject, and a generation unit that performs image processing on the first image to generate the training image.
[0115] A learning system according to aspect 17 of the present disclosure is a learning system comprising a generation device that generates learning images to be used in training a machine learning model that makes a predetermined inference from a target image, and a learning device that trains the machine learning model, and is configured to include an acquisition unit that acquires a first image that contains at least a portion of a subject, a generation unit that performs image processing on the first image to generate the learning image, an output unit that outputs the generated learning image to the learning device, a memory unit that stores a parameter set that defines the machine learning model, and a learning unit that trains the machine learning model.
[0116] A control program according to aspect 18 of the present disclosure is a control program for causing a computer to function as the generation device described in aspect 16 above, and may be configured to cause a computer to function as the acquisition unit and the generation unit.
[0117] The recording medium according to aspect 19 of the present disclosure may be a computer-readable non-transitory recording medium on which the control program according to aspect 18 above is recorded.
[0118] The present disclosure will be further explained below with reference to Examples 1 to 3, but the present disclosure is not limited to these. In each of the following Examples, in addition to matters corresponding to the above-mentioned Embodiments 1 and 2, matters derived from each of the embodiments are included as part of the present disclosure. In each Example, learning system 100 or 100A, or a system equivalent thereto, is used.
[0119] In this first embodiment, an osteoporosis diagnostic system using a frontal X-ray image of the lumbar spine will be described.
[0120] (1. Background) Osteoporosis is a bone disease characterized by reduced bone mass and bone fragility, resulting in an increased risk of fracture. The number of osteoporosis patients is increasing annually in many developed countries, which are experiencing an ultra-aging society. Patients with osteoporosis are more likely to suffer fragility fractures (e.g., of the vertebral body or proximal femur) even from minor external forces. Such fractures, along with fractures at the same site, can increase the risk of secondary fractures at other sites. Fractures in the elderly population are not only associated with reduced activities of daily living (ADL) and quality of life, but can also have a serious impact on prognosis. It has previously been reported that the risk of death after a fracture is several times higher than the risk of death from all causes, regardless of gender. Therefore, early diagnosis and treatment of osteoporosis and prediction of fragility fractures are urgent issues not only for medical care but also for society.
[0121] In recent years, significant progress has been made in the prediction, diagnosis, and treatment of osteoporosis, including the development of consensus statements and Japanese guidelines, and the clinical application of drugs based on basic research into the pathogenesis. However, in actual clinical practice, it is not uncommon for patients to visit a hospital after a fracture, receive a diagnosis of osteoporosis, and begin treatment. The medical system for early diagnosis of osteoporosis is inadequate. First, osteoporosis is usually asymptomatic, making it difficult to motivate patients to seek consultation. Second, DXA equipment is only available in 6% of general clinics in Japan and only 10.6 units per million in Europe, yet DXA is recommended for the diagnosis of osteoporosis, measuring BMD (bone mineral density) at the lumbar spine and proximal femur. Therefore, screening all elderly people is medically and socially difficult. Meanwhile, radiography (X-ray) equipment is available in 46.7% of general clinics in Japan and 212 units per million in the United States. Furthermore, X-rays are the initial examination for the majority of orthopedic patients (estimated to be approximately 619 million), so if it were possible to estimate BMD from X-ray images of patients with some disease, it would be possible to guide patients with potential osteoporosis to early treatment.
[0122] Several imaging diagnostic systems using artificial intelligence (AI) are being investigated to detect various diseases such as colonoscopy, pulmonary nodules, cerebral aneurysms, diabetic retinopathy, or osteoporosis.
[0123] The inventors of this application used a bone AI system (hereinafter sometimes simply referred to as the bone AI system) that outputs estimated BMD from frontal X-ray images of the lumbar spine to evaluate the performance of the AI system in terms of BMD estimation and osteoporosis classification accuracy. They hypothesized that the AI system can estimate BMD from frontal images of the lumbar spine (plain anterior-posterior X-ray images) and classify patients into those with osteopenia and those with osteoporosis.
[0124] (2. Premise and Structure) (2.1. Patients) The subjects for the performance evaluation of the bone AI system were patients who underwent both lumbar spine DXA frontal radiography and bone mineral density measurement within 180 days between April 2007 and January 2021 (Figure 12). Images meeting exclusion criteria included: (1) images of patients under the age of 20; (2) images of patients who had undergone lumbar spine surgery such as spinal instrumentation, vertebroplasty, or kyphoplasty (BKP; Balloon Kyphoplasty); and (3) images showing foreign objects such as staplers, sutures, or tubes overlapping the lumbar spine. All clinical data were retrospectively collected at the University of Tokyo Hospital. This study was approved by the hospital's Institutional Review Board and an external expert organization. All patients provided informed consent (opt-out method).
[0125] 2.2. Radiography and BMD Measurements. Frontal lumbar spine radiographs were acquired in Digital Imaging and Communications in Medicine (DICOM) format using a computer or digital radiography machine. BMD values were measured using a DXA whole-body densitometer (Lunar iDXA, GE Healthcare, Madison, WI). Each frontal lumbar spine radiograph was correlated with the measured lumbar BMD, which was the average of the specific lumbar spine regions L1-L4, resulting in a dataset. Each lumbar spine region from L1 to L4 was annotated by an orthopedic surgeon and trained reviewers.
[0126] (2.3. Segmentation Network) The artificial neural network (ANN) in this example was composed of two stages: a first stage was a segmentation network for detecting the spinal region, and a second stage was a deep neural network (DNN) for estimating the BMD value of the lumbar vertebrae on the X-ray image ( FIG. 13 ).
[0127] To prevent overfitting of the segmentation network, first image preprocessing, including cropping, rotation, and horizontal flipping using a cut-out region within the image, was randomly performed on the X-ray images in the training dataset. The segmentation network was based on an encoder-decoder architecture to reduce computational cost. A segmentation model in the segmentation network was constructed by training the first preprocessed images using annotation data for the L1-L4 lumbar vertebrae region as the ground truth for the X-ray image training data. Ground truth refers to the actual data used to train and test the output of AI models such as ANNs. Adaptive moment estimation was used as the optimization algorithm.
[0128] (2.4. BMD Estimation DNN) To improve the accuracy of BMD estimation, a second set of image preprocessing, including size adjustment, brightness normalization, white balance correction, and cropping around the segmented region of the lumbar vertebrae, was performed for both the training and estimation processes. To obtain a sufficient number of images to train the ANN for BMD estimation, data augmentation was performed to increase the number of X-ray images. The data augmentation corresponds to image generation by the generator 1. Data augmentation simulated imaging variations that may occur during actual X-ray imaging, including rotation, scaling, brightness, and contrast changes. The data augmentation process was performed randomly for each X-ray image, random types of changes, and a random number of images.
[0129] A DNN was used to estimate BMD. A transformer-based network was selected as an appropriate network for this example. The mean squared error was used as the loss function to calculate the difference between the training data (i.e., BMD measurements obtained by DXA) and the BMD estimate during training. A fully connected layer was provided in front of the loss function calculation unit, which corresponds to part of the training device 70, and a single scalar value was output as the estimated BMD. The number of DNN parameters and the learning speed of the optimizer (as a hyperparameter) were manually adjusted appropriately to control the behavior of the neural network. Here, an optimizer refers to software or a function used for optimization, which adjusts the target settings, structure, etc., and recombines them into a more favorable state. Stochastic gradient descent was used as the optimization algorithm. A fully connected layer was used as the output layer. The DNN was constructed by training BMD measurements by DXA as the ground truth for training data, and the second preprocessed X-ray image was cropped around the lumbar spine region by the segmentation network.
[0130] (2.5. Performance Evaluation of ANN) The 2,283 datasets were randomly divided into five groups, and no two patients were included in both the training and test datasets. Five-fold cross-validation was performed to evaluate the accuracy of BMD estimated using the ANN. Four groups of training datasets of X-ray images correlated with BMD measurements by DXA were input into the ANN, and training parameters were constructed using the ANN. After training, the only X-ray image in one untrained group of the dataset without annotation data was input into the trained ANN as a test dataset for estimation, and BMD estimates were output. These training and testing processes were repeated five times for each group of the dataset, obtaining BMD estimates for all patients with X-ray images. The estimation performance of the ANN was evaluated by comparing BMD measurements by DXA with BMD estimates by AI. The mean absolute error (MAE) and correlation coefficient between DXA-measured BMD and AI-estimated BMD were calculated using Microsoft Excel for Microsoft 365 (registered trademark, Microsoft Corporation, Redmond, WA, USA). DXA-measured and AI-estimated T-scores, based on the Japanese Osteoporosis Diagnostic Criteria, were calculated from the DXA-measured and AI-estimated BMD, respectively. The calculated T-scores were classified into three osteoporosis categories according to the World Health Organization (WHO) criteria, and the classification performance of the ANN was evaluated. The classification performance was evaluated as normal (T-score ≥ -1.0), osteopenia (-2.5 < T-score < -1.0), and osteoporosis (T-score ≤ -2.5). To evaluate the classification performance of the ANN, we calculated the sensitivity, specificity, accuracy, positive predictive value (PPV), and area under the curve (AUC).To evaluate the attention region of the DNN, a color image as an attention map was output from the information of the attention layer in the transformer-based network.All testing and evaluation was performed using PyTorch 1.13.1 (https: / / PyTorch.org / ) with a Windows 10® (Microsoft Corporation) device equipped with a central processing unit and graphics processing unit (RTX 3090, NVIDIA Corporation, Santa Clara, CA).
[0131] 3. Results A total of 2,283 X-ray images from 1,394 patients underwent BMD estimation using the ANN (Table 1), which shows the characteristics of the lumbar spine and femur datasets by total, gender, and BMI.
[0132] The number of X-ray images of females was 1,913 and that of males was 370. The mean age, BMI, and BMD of the subjects were 71.5, 21.8, and 0.930 g / cm, respectively. 2 In the female group, the values were 71.7, 21.7, and 0.911 g / cm 2 , and 70.5, 22.4, and 1.030 g / cm in the male group. 2 According to the World Health Organization report criteria, the numbers of underweight (BMI < 18.5), normal (18.5 ≤ BMI < 25.0), and obese (25.0 ≤ BMI) X-ray images were 387, 1,519, and 377, respectively (Table 2). Table 2 shows the accuracy of lumbar BMD estimation and osteoporosis classification using the ANN.
[0133] The mean BMD measurements by DXA for the underweight, normal, and obese groups were 0.806, 0.932, and 1.051 g / cm, respectively. 2 The number of radiographs classified by DXA BMD measurements was 703 (30.8%), 924 (40.5%), and 656 (28.7%) as normal, osteopenic, and osteoporotic, respectively.
[0134] Regarding BMD estimation performance, when comparing all lumbar DXA BMD measurements with all AI estimates using the ANN, the MAE was 0.070 g / cm 2The correlation coefficient between DXA-measured and AI-estimated BMD values was 0.895, suggesting a strong correlation (Figure 14). In Figure 14, graphs A, C, and E are examples of correlation scatter plots for estimated BMD for all patients, women, and men, respectively, and graphs B, D, and F are examples of ROC curves for estimated BMD for all patients, women, and men, respectively. No clear difference in estimation performance between women and men was observed. The MAE for women and men was 0.069 and 0.075 g / cm, respectively. 2 The correlation coefficients were 0.886 and 0.905, respectively. The MAEs were nearly identical, with values of 0.071, 0.070, and 0.070 for underweight, normal, and obese patients, respectively. The correlation coefficients for DXA-measured and AI-estimated BMD for patients in each BMI category were 0.867, 0.881, and 0.900, respectively, demonstrating that AI-estimated values were strongly correlated with DXA measurements regardless of patient BMI category ( Figure 15 ). In Figure 15 , graphs A, C, and E are examples of correlation scatter plots for estimated BMD for underweight, normal, and obese patients, respectively, and graphs B, D, and F are examples of ROC curves for estimated BMD for underweight, normal, and obese patients, respectively. The MAE between BMD measurements by DXA and BMD estimates by AI was 0.086, 0.061, and 0.066 g / cm for normal, osteopenic, and osteoporotic patients, respectively. 2 It was.
[0135] The classification performance of osteopenia and osteoporosis patients using the ANN was 91.8% sensitivity and 73.0% specificity, 82.1% specificity and 91.8% accuracy, 88.8% specificity and 86.4% accuracy, 92.0% specificity and 78.1% specificity, and AUCs of 0.951 and 0.935, respectively. The classification performance of female and male osteopenia patients was similar. The accuracy was 89.0% specificity and 87.8% specificity, and the AUCs were 0.948 and 0.953, respectively. The sensitivity of osteoporosis classification was lower (72.8% specificity and 75.0% specificity) than that of osteopenia classification (92.3% specificity and 87.4% specificity, respectively). For underweight osteopenic patients, both accuracy (92.2%) and AUC (0.972) were the best classification performance. On the other hand, for obese osteoporotic patients, both accuracy (95.0%) and AUC (0.970) were the best performance.
[0136] FIG. 16 shows typical attention heat maps overlaid on lumbar spine X-ray images with high and low estimation accuracy using an ANN. The areas surrounded by dotted lines in FIG. 16 represent the areas where the attention heat maps are displayed. In FIG. 16, A shows an attention heat map with high estimation accuracy, and B shows an attention heat map with low estimation accuracy. BMD estimation primarily used areas with high brightness, which the ANN focused on as the target. The attention heat map is a visualization of the area the ANN focused on as the target for easy understanding. For example, the area the ANN focused on as the target may be displayed with high brightness in the attention heat map. In the example shown in FIG. 16, the majority of the lumbar spine area surrounded by vertices L1-L4 had a more uniform brightness distribution, which tended to result in higher estimation accuracy. As shown in FIG. 16A, the BMD measured by DXA and the BMD estimated by AI were both 0.904 g / cm 2 On the other hand, when a part of the lumbar vertebrae was not colored as shown in FIG. 16B, the BMD measurement value by DXA (0.932 g / cm 2 ) and the AI-estimated BMD (0.795 g / cm 2 Compression fractures of L1 and L2 were observed in the patient in Figure 16B on a lateral radiograph (not shown).
[0137] In this second embodiment, a bone AI system that estimates the BMD of the lumbar vertebrae and femur from a frontal X-ray image (anterior-posterior X-ray image) of the lumbar region will be described. Furthermore, the same description of the matters already described in the previous embodiments will not be repeated. The same applies to the third embodiment described below.
[0138] (1. Background) The inventors of this application used a bone AI system trained on lumbar spine X-ray images obtained in a community cohort study to evaluate the accuracy of lumbar and femoral BMD estimation and osteoporosis classification when outputting not only lumbar spine BMD estimates but also femoral BMD estimates from frontal lumbar X-ray images. The hypotheses of this example are: first, that AI can estimate not only lumbar spine BMD but also femoral BMD from a simple frontal image of a single lumbar vertebra, and can classify osteopenia patients from osteoporosis patients. Second, the accuracy of femoral BMD estimation (not visible in X-ray images) is sufficient, but not higher than that of lumbar spine BMD estimation (visible in X-ray images). Third, AI can estimate BMD in the general community, including healthy individuals, and classify osteopenia patients from osteoporosis patients.
[0139] (2. Premise and Structure) (2.1. Participants) The "ROAD (Research on Osteoarthritis / Osteoporosis Against Disability) Study," which began in 2005, is a large-scale, population-based, prospective cohort study of osteoarthritis and osteoporosis in several communities in urban (Itabashi Ward, Tokyo, Japan), mountainous (Hidaka River, Wakayama, Japan), and coastal (Taiji, Wakayama, Japan) regions. In this example, data from the third survey of the mountainous and coastal regions of the ROAD study (2012-2013) was used ( FIG. 17 ). Participants in the third survey study included participants from the previous second ROAD study. In addition to the previous participants, residents aged 40 years or older who were willing to participate in the ROAD study were included. A total of 1,721 patients (769 in the mountainous region and 952 in the coastal region) participated in the third study. In this study, the following data were excluded: (1) data without lumbar frontal radiographs, (2) data without lumbar and / or femoral BMD data, (3) data in which the patient had lumbar instrumentation, (4) data from patients with large lumbar deformities, and (5) data in which the images were not clear enough for analysis.
[0140] 2.2. Radiography and BMD Measurement. Plain frontal radiographs of the lumbar spine were obtained in Digital Imaging and Communications in Medicine (DICOM) format using two devices manufactured by two companies (FUJIFILM Corporation, Tokyo, Japan, and Konica Minolta Inc., Tokyo, Japan). BMD values were measured at the lumbar spine (L2-L4) and proximal femur (neck) using DXA (Hologic Discovery; Hologic, Waltham, MA). The same DXA device was used for all participants. Osteopenia and osteoporosis were defined according to the World Health Organization (WHO) criteria as normal (T-score ≥ -1.0), osteopenia (-2.5 < T-score < -1.0), and osteoporosis (T-score ≤ -2.5). The cutoff values were based on the Japanese guidelines for BMD values of both the lumbar spine and femur. Osteopenia was defined as a BMD of 0.7135 to 0.8920 g / cm at the lumbar spine, regardless of gender. 2and 0.5650 to 0.7000 g / cm at the femoral neck. 2 Osteoporosis was defined as a bone mineral density between 0.7135 g / cm at the lumbar spine and 0.7135 g / cm at the lumbar spine, regardless of sex. 2 Below, 0.5650 g / cm at the femoral neck 2 It was defined by the following BMD:
[0141] (2.3. Segmentation Network for Lumbar Spine BMD Estimation) The artificial neural network (ANN) for lumbar spine BMD estimation in this example consisted of two stages: the first stage was a segmentation network for detecting spinal regions, and the second stage was a deep neural network (DNN) for estimating the BMD values of the lumbar spine on X-ray images ( FIG. 18 ).
[0142] Image preprocessing for the segmentation network, including cropping, rotation, and horizontal flipping, was randomly performed on the X-ray images in the training dataset to prevent overfitting of the network. The segmentation network was based on an encoder-decoder architecture to reduce computational cost. A segmentation model in the segmentation network was constructed by training the first preprocessed image using annotation data annotating the L1-L4 lumbar vertebrae region as the ground truth for the X-ray image training data. Adaptive moment estimation was used as the optimization algorithm. The femoral BMD estimation in this example consisted only of a DNN for estimating the BMD value of the femoral neck. This segmentation network was not used because the femoral neck did not appear in the lumbar spine X-ray images.
[0143] (2.4. Image Preprocessing and Data Augmentation) To improve the accuracy of the DNN for lumbar and femoral BMD estimation, image preprocessing, such as size adjustment, brightness normalization, white balance correction, and histogram equalization, was performed before DNN training or BMD estimation. Data augmentation was performed to increase the number of X-ray images to obtain enough images for training the DNN. Data augmentation simulated imaging variations that may occur during actual X-ray imaging, including rotation, scaling, brightness, and contrast changes. Data augmentation was performed randomly for each X-ray image, random types of changes, and a random number of images.
[0144] (2.5. DNN) The ANN in this example was constructed using a DNN to estimate BMD values of the lumbar spine or femoral neck from frontal X-ray images of the lumbar spine. A transformer-based network was selected as an appropriate network. The mean squared error was used as the loss function to calculate the difference between the training data (i.e., BMD measurements obtained by DXA) and the estimated BMD values during training. A fully connected layer was provided before the loss function calculation section, and a single scalar value was output as the estimated BMD. The number of DNN parameters and the optimizer learning speed as hyperparameters were manually adjusted appropriately to control the behavior of the neural network. Stochastic gradient descent was used as the optimization algorithm. The fully connected layer was used as the output layer. The DNN was constructed by training BMD measurements obtained by DXA as training data and ground truth data from preprocessed X-ray images.
[0145] (2.6. Performance Evaluation of ANN) The lumbar spine and femur datasets were randomly divided into five groups, with no identical participants included in both the training and test datasets. Five-fold cross-validation was performed to evaluate the accuracy of BMD estimation using the ANN. Four groups of training datasets of X-ray images correlated with DXA BMD measurements were input to each DNN for the lumbar spine and femur, and trained parameters were constructed for each DNN. After training, the only X-ray image in one untrained dataset group was input to each trained DNN as a test dataset for estimation, and estimated lumbar spine and femur BMD was output, respectively. These training and testing processes were repeated five times for each group of datasets. The ANN's estimation performance for lumbar spine and femoral neck BMD was evaluated by the correlation coefficient and mean absolute error (MAE) between DXA-measured BMD and AI-estimated BMD, calculated using Microsoft Excel for Microsoft 365. DXA-measured and AI-estimated T-scores for the lumbar spine and femur, based on the Japanese Osteoporosis Diagnostic Criteria, were calculated from the DXA-measured and AI-estimated BMD values, respectively. The calculated lumbar spine and femoral neck T-scores were classified into three osteoporosis categories according to the WHO criteria, and the classification performance of the ANN was evaluated. To evaluate the classification performance of the ANN, sensitivity, specificity, accuracy, positive predictive value (PPV), and area under the curve (AUC) were also calculated. Participants were classified into three categories according to BMI: underweight (BMI < 18.5), normal (18.5 ≤ BMI < 25.0), and obese (25.0 ≤ BMI). To evaluate the attention region of the ANN, a color image was output from the attention layer of the transformer-based network as an attention heatmap.All testing and evaluation was performed using PyTorch 1.13.1 (https: / / PyTorch.org / ) as the framework for building neural networks on a Windows (Windows 10, Microsoft Corporation) device equipped with a central processing unit and graphics processing unit (RTX 3090, NVIDIA Corporation, Santa Clara, CA).
[0146] 3. Results The total number of radiographs categorized by sex, BMI category (underweight, normal, and obese), and DXA BMD measurements (normal, osteopenic, and osteoporotic patients) are shown in Table 3. Table 3 shows the characteristics of the lumbar spine and femur datasets by total, sex, and BMI.
[0147] A total of 1,537 x-ray images from 1,537 participants were analyzed for lumbar spine and femur BMD estimation using each DNN. The mean age and BMI of all participants were 65.6 and 23.0, respectively. The mean DXA BMD measurements of the lumbar spine and femur were 0.957 and 0.663 g / cm, respectively. 2 It was.
[0148] The MAE, mean relative error (MRE), and correlation coefficient for BMD estimation performance using the ANN for all subjects were 0.073 g / cm for the lumbar spine. 2 , 7.9% and 0.887 for the femur, and 0.081 g / cm 2 , 12.7% and 0.665 (Table 4). Table 4 shows the accuracy of lumbar and femoral neck BMD estimation and osteoporosis classification using ANN.
[0149] Lumbar spine BMD measurements by DXA and lumbar spine BMD estimates by AI were strongly correlated regardless of gender and BMI (Fig. 19A) (Fig. 19C and E). In Fig. 19, graphs A, C, and E are examples of correlation scatter plots for estimated lumbar spine BMD for all patients, gender, and BMI category, respectively, and graphs B, D, and F are examples of ROC curves for estimated lumbar spine BMD for all patients, gender, and BMI category, respectively. Correlations between DXA BMD measurements and AI BMD estimates were also observed for femur BMD estimates regardless of gender and BMI (Fig. 20A) (Fig. 20C and E). In Figure 20, graphs A, C, and E are examples of correlation scatter plots for estimated femoral BMD by all patients, sex, and BMI category, respectively, and graphs B, D, and F are examples of ROC curves for estimated femoral BMD by all patients, sex, and BMI category, respectively. 2 The MAEs for lumbar and femoral BMD estimates were 0.069 and 0.077 g / cm for female participants (0.069 and 0.077 g / cm). 2 ) were larger than those for lumbar spine BMD estimation. Conversely, the MREs for male participants (7.7 and 12.0%) were similar to or smaller than those for female participants (8.0 and 13.0%). For lumbar spine and femur BMD estimation, each MAE was similar regardless of BMI category. MRE tended to decrease with BMI. Conversely, MRE tended to increase with osteoporosis category (i.e., decreased BMD / T-score). MRE for femur BMD estimation was larger compared with that for lumbar spine BMD estimation, regardless of sex or osteoporosis category. Conversely, the correlation coefficient for femur BMD estimation was lower compared with that for lumbar spine BMD estimation, regardless of sex or osteoporosis category.
[0150] For all participants, the lumbar spine classification performance for osteopenia and osteoporosis was 80.7% and 64.5%, respectively. The specificity was 89.0% and 95.9%, the accuracy was 85.6% and 92.7%, the PPV was 83.3% and 63.7%, and the AUC was 0.940 and 0.948, respectively. For all participants, the femur classification performance for osteopenia and osteoporosis was 82.2% and 45.0%, the specificity was 71.6% and 88.3%, the accuracy was 78.5% and 78.4%, the PPV was 84.3% and 53.5%, and the AUC was 0.830 and 0.810, respectively. The AUC for lumbar spine BMD estimation was higher than that for femoral BMD estimation. The sensitivity for male osteoporotic participants (25.0% and 18.9%) was significantly lower than that for other participants for both lumbar spine and femur BMD estimates. The PPV for osteoporotic participants was relatively lower than that for other participants for both lumbar spine and femur BMD estimates.
[0151] Figure 21 shows typical attention heat maps overlaid on lumbar spine X-ray images for high and low BMD estimation accuracy using each DNN. The dotted areas in Figure 21 represent the attention heat maps. In Figure 21, A and C show the attention heat maps for high BMD estimation accuracy, while B and D show the attention heat maps for low BMD estimation accuracy. For BMD estimation, the ANN primarily used high-intensity regions. As shown in Figure 21A (MRE value ±0.0%) and 21B (MRE value −13.4%), for lumbar spine BMD estimation with both high and low accuracy, the L1–L4 regions were highly bright. In Figure 21B, significant deformations and osteophytes were observed at L2 and L3. As shown in Figure 21C (MRE value ±0.0%), for high-accuracy femoral BMD estimation, the thoracic, lumbar, and pelvic regions were highly bright. As shown in FIG. 21D (MRE value -29.1%), in the femur BMD estimation with low accuracy, the thoracic, lumbar vertebrae, and rib regions were distributed with high brightness.
[0152] In this third embodiment, a bone AI system using frontal X-ray images of the chest (anterior-posterior X-ray images) will be described.
[0153] (1. Background) The inventors of this application have developed a bone AI system that outputs BMD estimates from frontal X-ray images of the lumbar spine, and have devised a system for screening for osteoporosis from frontal chest X-ray images taken during initial surgery for lung cancer, heart disease, or breast cancer, and / or during general medical surgery. This provides numerous opportunities for diagnosing patients for diseases other than osteoporosis. BMD estimation from chest X-ray images of patients with any organ disease, such as lung (e.g., lung cancer or viral pneumonia), heart (e.g., heart disease), and cancer (e.g., breast cancer), which affects 300,000 patients in Japan and 730 million patients worldwide, is expected to lead to early treatment for potential osteoporosis patients. Furthermore, approximately 12.5 million people annually undergo health checkups, including chest X-rays, primarily in Japan.
[0154] The purpose of this example was to evaluate the performance of a bone AI system using chest X-ray images to estimate lumbar and femoral BMD and to evaluate the accuracy of osteoporosis classification. The hypothesis was that AI could be used to effectively estimate lumbar and femoral BMD from plain frontal chest X-ray images and to classify patients with osteopenia and osteoporosis.
[0155] (2. Premise and Structure) (2.1. Patients) In this example, patients who underwent a frontal chest radiograph and DXA bone mineral density measurement of the lumbar spine or femur within 365 days between December 2006 and May 2022 were included ( Fig. 22 ). Images of patients under the age of 20 years, (2) patients with previous implantation procedures such as cervical braces, total shoulder arthroplasty, or cardiac pacemakers, and (3) images with foreign objects such as staples, sutures, or tubes overlapping the bone area met exclusion criteria. Images of patients with only the lumbar spine (4) previous spinal surgery such as spinal fusion, vertebroplasty, or kyphoplasty met additional exclusion criteria. BMI was classified into three categories according to the World Health Organization (WHO) criteria: underweight (BMI < 18.5), normal (18.5 ≤ BMI < 25.0), and obese (25.0 ≤ BMI). All clinical data were retrospectively collected from the University of Tokyo Hospital. The Institutional Review Boards of the hospital and external specialist institutions approved this study. All patients provided informed consent via an opt-out procedure.
[0156] 2.2. Radiography and BMD Measurement Frontal chest radiographs were acquired and formatted in Digital Imaging and Communications in Medicine (DICOM). BMD values were measured using a DXA whole-body bone densitometer (Lunar iDXA, GE Healthcare Technologies Inc., Madison, WI). Frontal chest radiographs were paired with corresponding measured lumbar spine (L2-L4) and femur (total hip) BMD to construct comprehensive lumbar spine and femur data sets, respectively.
[0157] (2.3. Image Preprocessing and Data Augmentation) To improve the BMD estimation accuracy of the artificial neural network (ANN), image preprocessing, such as size adjustment, brightness normalization, white balance correction, and histogram equalization, was performed for both ANN training and BMD estimation. Data augmentation was performed to increase the number of X-ray images to obtain a sufficient number of images for training the ANN. Data augmentation simulates imaging variations that may occur during X-ray imaging, including rotation, scaling, brightness, and contrast changes. Data augmentation was performed randomly for each X-ray image, type of change, and number of images. Enhancement processing (i.e., sharp kernel reconstruction) was performed on bone regions after image preprocessing and / or data augmentation. Image preprocessing and bone enhancement processing were performed for the training (learning) and estimation processes.
[0158] (2.4. ANN) The ANN in this example included a deep neural network (DNN) for estimating lumbar or femoral BMD values on frontal chest X-ray images (Figure 23). A transformer-based network was selected as an appropriate network for this example. The mean squared error was used as the loss function to calculate the difference between the training data (i.e., BMD measurements obtained by DXA) and the estimated BMD values during training. A fully connected layer was added before the loss function calculation section, outputting a single scalar value as the estimated BMD. Hyperparameters, including the number of parameters in the DNN and the learning speed of the optimizer, were manually adjusted to adjust the behavior of the neural network. Stochastic gradient descent was used for optimization. The fully connected layer was used as the output layer. The DNN was developed by training DXA BMD measurements as training data and preprocessed X-ray images as ground truth.
[0159] (2.5. Performance Evaluation of ANN) The lumbar spine and femoral datasets were randomly divided into five groups, with no identical patients included in both the training and test datasets. A five-fold cross-validation was performed to evaluate the accuracy of BMD estimation using the ANN. Four training datasets consisting of X-ray images correlated with DXA BMD measurements were input into the ANN to establish training parameters. After training, only the X-ray images from the untrained dataset group were input into the trained ANN as the test dataset for estimation, and estimated BMD values were output. These training and testing processes were repeated five times for each group of datasets. To evaluate the attention region of the DNN, a color image as an attention heatmap was output from the attention layer information in the transformer-based network. All testing and evaluation was performed using PyTorch 1.13.1 (https: / / PyTorch.org / ) as a framework for building neural networks on a Windows 10 (Microsoft Corporation, Redmond, WA) machine equipped with a central processing unit and a graphics processing unit (RTX 3090, NVIDIA Corporation, Santa Clara, CA).
[0160] (2.6. Statistical Analysis) The estimation performance of the ANN was evaluated by the mean absolute error (MAE) and the correlation coefficient between DXA-derived and AI-estimated BMD. DXA-measured T-scores and AI-estimated T-scores, based on the Japanese Osteoporosis Diagnostic Criteria, were calculated from DXA-measured BMD and AI-estimated BMD, respectively. T-scores were classified into three osteoporosis categories according to the World Health Organization criteria: normal (T-score ≥ −1.0), osteopenia (−2.5 < T-score < −1.0), and osteoporosis (T-score ≤ −2.5). A confusion matrix was presented as a 2 × 2 contingency table displaying the number of true positives, false positives, false negatives, and true negatives. To evaluate the ANN classification performance, sensitivity, specificity, accuracy, positive predictive value (PPV), and area under the curve (AUC) were calculated. All calculations were performed using Microsoft Excel (Microsoft Corporation).
[0161] 3. Results The total number of radiographs categorized by sex, BMI category (underweight, normal weight, and obese), or DXA BMD measurement (normal, osteopenic, and osteoporotic) is shown in Table 5. Table 5 shows the characteristics of the lumbar spine and femur datasets by total, sex, and BMI.
[0162] A total of 4,217 x-ray images from 1,249 patients and 5,047 x-ray images from 1,503 patients were analyzed for BMD estimation of the lumbar spine and femur, respectively, using the ANN. The mean age, BMI, and BMD of the subjects were 72.5, 21.2, and 0.957 g / cm for the lumbar spine dataset. 2 , 72.7, 21.3, and 0.712 g / cm for the femur dataset. 2 It was.
[0163] Regarding the BMD estimation performance using ANN, the MAE for the BMD estimates of all the target cases was 0.126 g / cm between the DXA-derived lumbar BMD values and the AI-estimated lumbar BMD values. 2 (Table 6) and 0.081 between their femoral BMD values (Table 7). Table 6 shows the lumbar spine BMD estimates and osteoporosis classification accuracy using the ANN for total, sex, BMI, and disease category. Table 7 shows the femoral BMD estimates and osteoporosis classification accuracy using the ANN for total, sex, BMI, and disease category.
[0164] MAE is male (0.138 g / cm 2 ), obesity BMI (0.136 g / cm 2 ), normal disease category (0.160 g / cm 2 ) and the estimated lumbar BMD values of patients were relatively large, and those of men (0.101 g / cm 2 ) and the normal disease category (0.117 g / cm 2The correlation coefficients between DXA measurements and AI estimates were 0.703 and 0.698 for lumbar and femoral BMD estimates, respectively, demonstrating a strong correlation ( Figures 24 and 25 ). In Figure 24 , graphs A, C, and E are examples of correlation scatter plots for estimated lumbar spine BMD for all patients, sex, and BMI category, respectively, and graphs B, D, and F are examples of ROC curves for estimated lumbar spine BMD for all patients, sex, and BMI category, respectively. Also, in Figure 25 , graphs A, C, and E are examples of correlation scatter plots for estimated femoral BMD for all patients, sex, and BMI category, respectively, and graphs B, D, and F are examples of ROC curves for estimated femoral BMD for all patients, sex, and BMI category, respectively.
[0165] The classification performance of the lumbar spine and femur for osteopenia patients was 86.2% and 92.9%, respectively, the specificity was 64.0% and 56.8%, the accuracy was 79.5% and 86.5%, the PPV was 84.6% and 91.0%, and the AUC was 0.845 and 0.880 (Tables 6 and 7). The classification performance of the lumbar spine and femur for osteoporosis patients was 59.9% and 63.9%, the specificity was 85.9% and 82.6%, the accuracy was 78.3% and 76.7%, the PPV was 63.3% and 63.0%, and the AUC was 0.839 and 0.830, respectively. The AUCs for the lumbar spine classification for male patients with osteopenia and osteoporosis (0.870 and 0.870, respectively) were better than those for female patients (0.827 and 0.822, respectively) (Table 6 and Figure 24D). Similarly, the AUCs for the femur classification for male patients with osteopenia and osteoporosis (0.883 and 0.862, respectively) were better than those for female patients (0.858 and 0.816, respectively) (Table 7 and Figure 25D). The higher the patient's BMI, the higher the AUCs for both the lumbar spine and femur classifications (Tables 6 and 7, and Figures 24F and 25F).
[0166] Figure 26 shows typical attention heat maps overlaid on chest X-ray images with high and low accuracy using an ANN. The dotted areas in Figure 26 represent the attention heat maps. In Figure 26, A and C show the attention heat maps for high-accuracy bone mineral density estimation, while B and D show the attention heat maps for low-accuracy bone mineral density estimation. For BMD estimation, we primarily used high-intensity regions focused by the ANN. In the high-accuracy lumbar spine case, as shown in Figure 26A (mean relative error [MRE] value +0.1%), the neck and rib regions tended to be more uniformly distributed with high intensity. In contrast, in the low-accuracy lumbar spine case, as shown in Figure 26B (MRE value -16.3%), there was less distribution of high intensity, and the MRE values for lumbar spine BMD estimation differed significantly. Similarly, in the high-accuracy femur case, as shown in Figure 26C (MRE value +0.2%), even high intensity regions were distributed. Furthermore, the region had less distribution of high intensity in the low accuracy femur case (MRE value was −20.4%).
[0167] 1, 1A Medical image generating device (generating device) 10, 10A Control unit 11 Acquisition unit 12 Selection unit 13 Image correction unit 14 Image conversion unit 15 Second image conversion unit 16 Output unit 20 Storage unit 30 Input / output interface (input / output IF) 50 Display device 70 Machine learning model 100, 100A Learning system C Computer C1 Processor C2 Memory
Claims
1. A generation method including: an acquisition step of acquiring a first image that contains at least a portion of a subject; and a generation step of performing image processing on the first image to generate a learning image to be used in training a machine learning model.
2. The generating method described in claim 1, wherein the acquisition step acquires a plurality of specimen images including the first image, and further includes a selection step, prior to the generation step, of classifying the plurality of specimen images into a plurality of groups based on the distribution of pixel values of each of the plurality of specimen images, and selecting at least a portion of the groups including the first image.
3. The generation method according to claim 2, wherein in the selection step, the plurality of sample images are classified into a plurality of groups based on a result of dimensional compression of a result of analyzing the pixel value distribution of each of the plurality of sample images.
4. The generating method according to claim 2 or 3, further comprising a normalization step between the acquisition step and the generating step, for performing a pixel value normalization process on each of the plurality of sample images.
5. The method according to claim 4, wherein the normalization process includes a brightness normalization process and / or a Retinex process for each of the plurality of sample images.
6. The generation method according to any one of claims 1 to 5, wherein the image processing includes at least one of a process of changing the brightness of the first image, a process of changing the contrast, and a blurring process.
7. The method according to any one of claims 2 to 5, wherein the specimen image is a plain X-ray image.
8. A method for generating a bone according to any one of claims 1 to 7, wherein the subject is a bone, and the machine learning model estimates the bone density of a bone that appears in a target image in which at least a portion of the bone appears, and / or the bone density of a bone that does not appear in the target image.
9. The method of claim 8, wherein the machine learning model estimates bone density of the femur and / or lumbar bones from a target image showing chest bones.
10. The method of claim 8, wherein the machine learning model estimates bone density of the femur and / or lumbar bones from a target image showing lumbar bones.
11. The method of any one of claims 8 to 10, wherein the bone mineral density of the bone is at least one of bone mineral density per unit area, bone mineral density per unit volume, percent of young adult mean, T-score, and Z-score.
12. A generating method according to any one of claims 1 to 11, wherein in the generating step, the image processing is performed including a process of erasing a partial area included in an area appearing in the image after the image processing, or a process of changing the pixel value of the partial area to a predetermined pixel value, thereby generating the learning image.
13. The generating method according to claim 12, wherein the image processing includes at least one of a horizontal movement process, a rotation process, a flip process, an enlargement process, and a reduction process.
14. A method for training a machine learning model that estimates bone density of a bone from a target image in which the bone is captured, the method using training images generated by a generation method according to any one of claims 8 to 11 as explanatory variables, and using bone density information indicating the bone density of the subject's bone corresponding to the training images as a target variable.
15. A method for generating a machine learning model that estimates bone density of a bone from a target image in which the bone is shown, comprising the steps of: using training images generated by the generation method described in any one of claims 8 to 11 as explanatory variables; and using bone density information indicating the bone density of the bone of a subject corresponding to the training images as a target variable; and generating a machine learning model that estimates the bone density of a bone shown in a target image in which at least a portion of a bone is shown and / or the bone density of a bone not shown in the target image.
16. A generation device that generates learning images to be used in training a machine learning model that performs a predetermined inference from a target image, comprising: an acquisition unit that acquires a first image that contains at least a portion of a subject; and a generation unit that performs image processing on the first image to generate the learning image.
17. A learning system comprising a generation device that generates learning images used to train a machine learning model that makes a predetermined inference from a target image, and a learning device that trains the machine learning model, the learning system comprising: an acquisition unit that acquires a first image that contains at least a portion of a subject; a generation unit that performs image processing on the first image to generate the learning image; an output unit that outputs the generated learning image to the learning device; a memory unit that stores a parameter set that defines the machine learning model; and a learning unit that trains the machine learning model.
18. A control program for causing a computer to function as the generating device according to claim 16, the control program causing a computer to function as the acquisition unit and the generating unit.
19. A non-transitory computer-readable recording medium having the control program according to claim 18 recorded thereon.
Citation Information
Patent Citations
Image sorting device
JP1996293025A
Estimation device
JP2020171785A
Image processing apparatus, medical image pick-up device, image processing method, and program
JP2022111704A
Image processing method and device, electronic device, storage medium, and computer program
JP2022530413A
Estimating bone mineral density from plain radiograph by assessing bone texture with deep learning
US20210212647A1