Generation method, training method, generation device, control program, and recording medium
By combining and sharpening images to enhance tissue features, the method addresses suboptimal preprocessing issues, leading to improved learning and estimation accuracy in machine learning models for medical image analysis.
Patent Information
- Application Number
- PCT/JP2025/002591
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2025-01-28
- Publication Date
- 2025-08-07
AI Technical Summary
Existing techniques for preprocessing images for neural networks do not optimize the learning and estimation accuracy of tissue states, leading to suboptimal performance in machine learning models.
A method involving the combination and sharpening of multiple images to generate training images, which are used to train a machine learning model, enhancing the model's estimation accuracy by improving the sharpness and contrast of tissue features.
The proposed method improves the learning and estimation accuracy of neural networks by generating training images that better capture tissue characteristics, resulting in enhanced performance of machine learning models in medical image analysis.
Smart Images

Figure JP2025002591_07082025_PF_FP_ABST
Abstract
Description
Generation method, learning method, generation device, control program, and recording medium
[0001] The present disclosure relates to a method for generating training images used in training a machine learning model, a training method, a generation device, a control program, and a recording medium.
[0002] 2. Description of the Related Art There is known a technique for inputting an image into a neural network (NN) to learn and / or estimate the image.
[0003] Non-Patent Document 1 describes a configuration in which preprocessed chest X-ray images are used in machine learning of a neural network (NN) to estimate pneumonia associated with the novel coronavirus infection (COVID-19). In the preprocessing described in Non-Patent Document 1, a low-pass filter, a bilateral filter, histogram equalization, etc. are applied to the chest X-ray image, and then a pseudo color image connected in the three channel (RGB) directions is generated.
[0004] Morteza Heidari, et al., “Improving the performance of CNN to predict the likelihood of COVID-19 using chest X-ray images with preprocessing algorithms”, International Journal of Medical Informatics 144, 104284, 2020
[0005] <1> A generation method according to one aspect of the present disclosure includes a combining step of generating a first combined image by combining multiple first images, each of which depicts at least a portion of a first subject, and a sharpening step of performing a sharpening process on the first images, and the first combined image, which includes at least one of the first images after the sharpening process, is a training image used for training a machine learning model.
[0006] <2> A learning method according to one aspect of the present disclosure is a learning method for a machine learning model that uses training images generated by the generation method described in <1> above as explanatory variables and characteristic information indicating the characteristics of at least a portion of the first subject corresponding to the training images as a target variable, and estimates at least a portion of the characteristics of a second subject from a second image that shows at least a portion of the second subject.
[0007] <3> A machine learning model according to one aspect of the present disclosure is a trained machine learning model generated by the learning method described in <2> above.
[0008] <4> A generating device according to one aspect of the present disclosure includes a combining unit that generates a first combined image by combining multiple first images that show at least a portion of a first subject, and a sharpening unit that performs a sharpening process on the first images, and the first combined image, which includes at least one of the first images after the sharpening process, is a training image used for training a machine learning model.
[0009] The generating device according to each aspect of the present disclosure may be realized by a computer. In this case, the present disclosure also includes a control program for the generating device that causes the computer to operate as each unit (software element) of the generating device to realize the generating device, and a computer-readable non-transitory recording medium on which the control program is recorded. The computer may also include an edge computer, a cloud server, etc. The functions of the generating device and the functions of the learning device may each be realized by multiple devices.
[0010] FIG. 1 is a block diagram showing an example of the configuration of a learning system. FIG. 2 is a flowchart showing an example of the flow of a method for generating learning images for machine learning. FIG. 3 is a schematic diagram showing an example of the flow of processing executed by a sharpening unit. FIG. 4 is a schematic diagram showing another example of the flow of processing executed by a sharpening unit. FIG. 5 is a block diagram showing an example of the configuration of a computer that executes instructions of a program that realizes a control block of an image generating device.
[0011] It is known that preprocessing for generating images that can improve the learning accuracy and / or estimation accuracy of a neural network differs depending on the type of image input to the neural network and the type of neural network.
[0012] For example, although a technique is known in which an image is input into a neural network to learn and / or estimate the state of the tissue of the subject shown in the image, the optimal preprocessing for generating an image that improves the learning accuracy and / or estimation accuracy of the neural network has not been determined.
[0013] One aspect of the present disclosure has been made in consideration of the above-mentioned problems, and provides preprocessing for generating images that can improve the learning accuracy and / or estimation accuracy of a neural network that estimates the state of tissue in a subject depicted in an image.
[0014] [Embodiment 1] (Learning System 100) A learning system 100 including an image generating device 1 (generating device) according to an embodiment of the present disclosure will be described in detail below with reference to the drawings. FIG. 1 is a block diagram showing an example of the configuration of the learning system 100. As shown in FIG. 1, the learning system 100 includes the image generating device 1 (generating device), a display device 50, and a learning device 70 that performs machine learning on a machine learning model 702. The learning device 70 may be a cloud-based device or an on-premise device installed in a medical facility or a company that provides analysis services.
[0015] (Overview of Display Device 50 and Learning Device 70) The display device 50 is a device capable of displaying on a display screen various images and information output from the image generating device 1 and the learning device 70. The learning device 70 includes a storage unit 701 that stores a parameter set that defines a machine learning model 702, and a learning unit 703 that trains the machine learning model 702. The learning unit 703 trains the machine learning model 702 by updating the values of the parameter set stored in the storage unit 701.
[0016] 1 shows, as an example, a learning system 100 in which an image generating device 1 is communicatively connected to a display device 50 and a learning device 70. The learning system 100 is not limited to the configuration shown in FIG. 1. For example, in the learning system 100, the image generating device 1 may have the configuration of the display device 50 and / or the configuration of the learning device 70.
[0017] (Overview of Image Generating Device 1) The image generating device 1 generates training images from a plurality of first images that capture at least a portion of a first subject. The appearance and contrast of the tissue to be estimated in the original (i.e., unsharpened) first images are not necessarily optimal for machine learning to generate a machine learning model 702. Therefore, the image generating device 1 generates training images by combining a plurality of first images, including at least one first image that has been sharpened, in the channel direction. By using such training images, first images whose sharpness has been changed by the sharpening process can also be used in the machine learning of the machine learning model 702. This makes it possible to generate a machine learning model 702 with improved estimation accuracy compared to a machine learning model 702 generated by machine learning using only the original first images.
[0018] That is, the image generating device 1 is a device that generates training images that can improve the training accuracy and / or estimation accuracy of a machine learning model 702, which is a neural network (NN) that estimates the state of tissue of a subject captured in an image. Here, the training images are images used as training data for training the machine learning model 702. The training images may be medical images (first images) captured of a subject at a medical institution such as a hospital. The medical images include at least a portion of the subject.
[0019] The part shown in the medical image may be, for example, a tissue and / or an organ of the subject. The tissue of the subject may be, for example, a joint and / or a bone. Alternatively, the tissue may be at least one of epithelial tissue, connective tissue, muscle tissue, and nervous tissue. Alternatively, the organ may be, for example, at least one of the digestive system, the cardiovascular system, the endocrine system, and the musculoskeletal system.
[0020] The machine learning model 702 is capable of estimating, from a medical image (second image) in which at least a portion of the second subject is visible, the characteristics of at least a portion of the second subject that is visible in the second image, and / or the characteristics of a portion of the second subject that is not visible in the second image.
[0021] The medical image may be, for example, an image of a subject captured using an endoscope. More specifically, the medical image may be an image of a region of the subject including at least one of the nasal cavity, esophagus, stomach, duodenum, rectum, large intestine, small intestine, anus, and colon captured using an endoscope. Such a medical image can be used to obtain an analysis result that indicates, using the trained machine learning model 702, a region of interest that includes at least one of inflammation, polyps, and tumors (e.g., malignant tumors) in the imaged region.
[0022] In training such a machine learning model 702, for example, a first training image showing a region of interest and first training data indicating the presence of the region of interest, and a second training image not showing the region of interest and second training data indicating the absence of the region of interest may be used. The first training data may include information indicating the inflammation level (degree of inflammation) or malignancy level (degree of malignancy) of the region of interest. The trained machine learning model 702 may be configured to output, as a result of analysis, an image in which a line is drawn indicating a position corresponding to the region of interest. Alternatively, the trained machine learning model 702 may be configured to output, as a result of analysis, information indicating a position corresponding to the region of interest, or an image superimposed with a color indicating that the position corresponds to the region of interest. The trained machine learning model 702 may also be configured to output, as a result of analysis, evidence information indicating the basis for estimating the position corresponding to the region of interest.
[0023] The medical images may be, for example, images of the subject's eyes, skin, etc., captured using a digital camera, etc. Medical images of these areas can be used to obtain analysis results that clearly indicate, for example, signs of interest in the imaged area using the trained machine learning model 702.
[0024] Such a machine learning model 702 may be trained using, for example, first training images depicting a region where a sign of interest is observed and first training data indicating that the sign of interest is observed, and second training images depicting a region where a sign of interest is not observed (i.e., a healthy region) and second training data indicating that the sign of interest is not observed. In the case of the eye, the sign of interest may include signs indicating diseases including at least one of glaucoma, cataracts, age-related macular degeneration, conjunctivitis, hordeolum, retinopathy, and blepharitis. In the case of the skin, the sign of interest may include signs including, for example, skin cancer, hives, atopic dermatitis, and herpes. The trained machine learning model 702 may be configured to output, as a result of analysis, an image in which lines indicating the locations where these signs of interest are observed are drawn. Alternatively, the trained machine learning model 702 may be configured to output, as a result of analysis, information indicating the locations where the signs of interest are observed, or an image in which a color indicating the location corresponding to the sign of interest is superimposed. The trained machine learning model 702 may be configured to output, as a result of the analysis, for example, identification information (e.g., disease name) corresponding to the symptom of interest. The trained machine learning model 702 may further display, as a result of the analysis, evidence information indicating the basis for deriving the position corresponding to the symptom of interest.
[0025] The medical image may be, for example, an X-ray image including a plain X-ray image of at least a portion of a subject, a magnetic resonance imaging (MRI) image, or a computed tomography (CT) image. Alternatively, the medical image may be, for example, a positron emission tomography (PET) image or an ultrasound image. The medical image may also be, for example, a dental image. The medical image may be, for example, an inspection device image acquired from an inspection device (e.g., an X-ray inspection device), or an inspection device image with noise reduction. Here, the at least a portion of the subject may include, for example, at least one of the head, neck, chest, lower back, temporomandibular joint, spinal facet joint, hip joint, sacroiliac joint, knee joint, ankle joint, foot, toe, shoulder joint, acromioclavicular joint, elbow joint, wrist joint, hand, finger, and temporomandibular joint. The medical image may be a frontal image or a lateral image.
[0026] The X-ray image may be a plain X-ray image such as a chest X-ray image or a lumbar X-ray image. The X-ray image may also be captured by dual-energy X-ray absorptiometry (DXA), microdensitometry (MD), or other methods. In a DXA device that measures bone density using the DXA method, when measuring the bone density of the lumbar vertebrae, X-rays are irradiated from the front of the lumbar vertebrae of a subject. In a DXA device that measures bone density of the proximal femur, X-rays are irradiated from the front of the proximal femur of a subject. Here, the terms "front of the lumbar vertebrae" and "front of the proximal femur" refer to the direction that correctly faces the imaging site, such as the lumbar vertebrae or the proximal femur, and may be the ventral side or the back side of the subject's body. The proximal femur includes, for example, at least one of the neck, trochanter, and shaft, and the entire proximal femur (neck, trochanter, and shaft). In the MD method, for example, the hand is irradiated with X-rays.
[0027] A medical image containing information about a subject's bones can be used, for example, to obtain an analysis result that clearly indicates the bone density of the subject's bones shown in the medical image using a trained machine learning model 702. Examples of major bone densities include the bone densities of the lumbar vertebrae, femur, calcaneus, and radius. For example, when generating a machine learning model 702 that estimates bone density, machine learning is performed so that the output when a training image is input matches the specific bone density annotated and associated with the training image.
[0028] Such a machine learning model 702 can be trained by using training images generated using a plurality of medical images showing the chest bones of a first subject as explanatory variables and information indicating the bone density of the chest bones, femur bones, and / or lumbar bones of the first subject corresponding to the training images as a target variable. Such a trained machine learning model 702 can estimate the bone density of the chest bones, femur bones, and / or lumbar bones of a second subject from a second image, which is a medical image showing the chest bones of a second subject.
[0029] The machine learning model 702 may also be trained by using training images generated using a plurality of medical images showing the lumbar bones of the first subject as explanatory variables and information indicating the bone density of the chest bones, femoral bones, and / or lumbar bones of the first subject corresponding to the training images as a target variable. Such a trained machine learning model 702 can output estimation results of the bone density of the femoral bones and / or lumbar bones from training images showing the lumbar bones. The trained machine learning model 702 may be configured to output, as a result of analysis, evidence information indicating the basis for estimating the bone density of the bones of the second subject, together with the estimation results.
[0030] In other words, the machine learning model 702 may learn the bone density, etc. of bones not shown in a training image that shows at least a portion of the bones of the first subject, or may estimate the bone density, etc. of bones not shown in a target image that shows at least a portion of the bones.
[0031] The machine learning model 702 may be configured to learn and estimate the bone mineral density of the bones of the second subject from a second image showing a predetermined part of the body of the second subject. The machine learning model 702 may be configured to learn and estimate the bone mineral density of the bones of the second subject from the second image and attribute information of the second subject, such as the age and gender. The machine learning model 702 may be configured to output a determination result as to whether the second subject has osteoporosis.
[0032] Bone mineral density may be a value related to the density of bone. Bone mineral density is defined as bone mineral density per unit area (g / cm 2 ), bone mineral density per unit volume (g / cm 3 The bone mineral density may be expressed by at least one of the following: YAM (%), T-score, and Z-score. YAM (%) is an abbreviation for "Young Adult Mean" and may be referred to as the young adult mean percentage. For example, bone mineral density may be expressed by bone mineral density per unit area (g / cm 2 ) and YAM (%). The bone mineral density may be an index defined by a guideline or a unique index. For example, the bone mineral density may be a value used in osteoporosis guidelines (such as, but not limited to, the 2015 edition of the Prevention and Treatment Guidelines of the Japan Osteoporosis Society).
[0033] The bone density may be determined using a DXA device, an X-ray device, an ultrasonic bone density measuring device, etc. The bone density may be a value estimated using a device that estimates bone density from a plain X-ray image.
[0034] The image generating device 1 according to the present disclosure can generate training images that can also be used in machine learning to generate a machine learning model 702 that estimates bone abnormalities. Here, the bone abnormality may be, for example, a bone abnormality associated with a bone disease. Examples of bone diseases include osteoporosis, scoliosis, fractures, spinal stenosis, degenerative disc disease, ankylosing spondylitis, spinal cord injury, osteomyelitis, osteophytes, and spinal muscular atrophy.
[0035] The image generating device 1 according to the present disclosure can generate training images that can also be used in machine learning to generate a machine learning model 702 that estimates a predetermined characteristic of a subject. Here, the predetermined characteristic may be a component ratio of the subject's bones or an object attached to the subject's body, or a disease that the subject is suffering from and its symptoms.
[0036] The machine learning model 702 may also use a medical image showing the subject's bones as an explanatory variable and a measurement value of the subject's characteristic at a time different from the time when the medical image was captured as a target variable to predict a predetermined characteristic of the subject at a time different from the current time. The predetermined characteristic of the subject may be, for example, the bone mineral density of the subject's bones. The different time may be a time in the future and / or past the time when the image data was captured.
[0037] (Configuration of Image Generating Device 1) The configuration of the image generating device 1 will be described below with reference to Fig. 1. In the following, a training image that can be used to generate a machine learning model 702 that estimates the bone density of the bones of a second subject is generated from a medical image (e.g., a chest X-ray image) captured of the second subject.
[0038] The image generating device 1 includes a control unit 10, a storage unit 20, and an input / output IF 30. The control unit 10 includes one or more processors (not shown) and controls the image generating device 1. The storage unit 20 is a storage device that stores various data, and may store original medical image data (original data) 21 and converted image data (learning image data) 22 in which the original medical image has been subjected to a sharpening process (described below). Hereinafter, the original medical image will also be referred to as an "original image." The storage unit 20 may also store various control programs that are read and executed by the processor. Here, original image data refers to medical image data that has not been subjected to a sharpening process.
[0039] The input / output IF 30 is an interface for communicating information with external devices such as the learning device 70 and the display device 50. The input / output IF 30 may be a wired communication interface such as a USB port, or a wireless communication interface such as Bluetooth (registered trademark) or Wi-Fi (registered trademark). A user may be able to input data to the image generation device 1 via a mouse, keyboard, or the like connected to the input / output IF 30. Here, the term "user" refers to a medical professional, technician, or the like who uses the image generation device 1, and the same applies hereinafter. The image generation device 1 may also be connected to the Internet via the input / output IF 30 to communicate information.
[0040] The following describes each unit of the control unit 10. The control unit 10 may include an acquisition unit 11, a selection unit 12, and a sharpening unit 13.
[0041] The acquisition unit 11 acquires an original image (first image). The original image may be, for example, an X-ray image. The original image may depict at least one of the following regions of the first subject: the head, neck, chest, lower back, hip joints, knee joints, ankle joints, feet, toes, shoulder joints, elbow joints, wrist joints, hands, fingers, and temporomandibular joints. The original image may be, for example, a simple X-ray image including at least a portion of a bone, such as the sternum, lumbar vertebrae, or hip neck. The acquisition unit 11 may acquire the original image from the storage unit 20, or from an external device (e.g., an imaging device) or storage medium connected to the image generation device 1 via the input / output IF 30. The original image and the training image generated from the medical image are annotated with characteristic information indicating the bone density of the bones depicted in the medical image.
[0042] The selection unit 12 selects original images containing bones acquired by the acquisition unit 11 using a method for evaluating the quality or commonality of images. The selection unit 12 may select images that satisfy predetermined criteria. In one aspect, the selection unit 12 basically selects images whose quality is evaluated to be better than a standard or whose commonality is evaluated to be higher than a standard. The reason for selecting images with high commonality is that images that are clustered into groups with low commonality with a large number of images are different from typical images and are therefore considered unsuitable for learning by the machine learning model 702.
[0043] As an example, the selection unit 12 evaluates the commonality of images by combining t-Distributed Stochastic Neighbor Embedding (t-SNE) and Density-based spatial clustering of applications with noise (DBSCAN). In this case, t-SNE may be performed before or after DBSCAN. However, the method by which the selection unit 12 evaluates the commonality of images is not limited to a specific method, and another method may be performed between t-SNE and DBSCAN.
[0044] Sharpening processing does not need to be performed on images not selected by the selection unit 12. From another perspective, the acquisition unit 11 may also acquire images that are not actually used to generate learning images. For example, the acquisition unit 11 may acquire multiple medical images that show the first subject, and the selection unit 12 may select multiple medical images that show at least a portion of the first subject from the acquired multiple medical images. The selected multiple medical images are used to generate learning images. In other words, the acquisition unit 11 and the selection unit 12 can function as an acquisition unit that acquires multiple medical images that show at least a portion of the first subject.
[0045] The sharpening unit 13 performs a sharpening process on at least one of the acquired original images, and combines multiple medical images, including at least one medical image after the sharpening process, in the channel direction to generate a training image to be used for training the machine learning model 702. The sharpening process may be performed once or multiple times. Alternatively, the sharpening process may be performed both before and after the combining process. Specifically, after combining multiple sharpened medical images, at least one additional sharpening process may be performed to generate a training image. In this case, at least one of the strength and processing method of the sharpening process may be changed before and after the combining process. For example, the strength of the sharpening process before the combining process may be increased while the strength of the sharpening process after the combining process may be decreased. Alternatively, an unsharp process may be used as the sharpening process before the combining process, while a sharpening filter may be used as the sharpening process after the combining process.
[0046] Each pixel in a medical image (color image) expresses one color using multiple planes corresponding to color information, such as an R channel, a G channel, and a B channel. When the planes corresponding to different color channels are arranged at different positions in the depth direction and illustrated, the channel direction is the direction corresponding to this depth.
[0047] The sharpening unit 13 may combine multiple medical images, including at least one medical image after sharpening processing, in the vertical or horizontal direction on the image plane to generate a training image to be used for training the machine learning model 702. That is, the sharpening unit 13 combines multiple medical images, including at least one medical image after sharpening processing, in at least one of the channel direction, vertical direction, and / or horizontal direction to generate a training image to be used for training the machine learning model 702. Here, the sharpening unit 13 may combine multiple medical images, including at least one medical image after sharpening processing, in the channel direction to create a combined image (first combined image) by superimposing the images. The combined image may be an image formed by the average pixel value of each superimposed image, an image formed by the highest pixel value of each superimposed image, or a combined image formed by the lowest pixel value of each superimposed image.
[0048] The sharpening unit 13 may perform sharpening processing of the medical image using multiple sharpening filters. The sharpening processing is an enhancement process for making the image appear clearer, and is a process of performing a conversion that increases the change in pixel values (change in shading) of the original image. The sharpening unit 13 may adjust the contrast of the medical image by performing gamma correction on the medical image using the sharpening filters. In the sharpening processing of the medical image, the sharpening unit 13 may use an edge filter to enhance the contours of tissues (e.g., bones) of the first subject that appear in the medical image. Here, the edge filter may be a filter process that enhances portions (edges) where the brightness of the image changes suddenly.
[0049] The sharpening unit 13 may also use the following filters to perform sharpening processing on medical images: Differential filter (a filter that outputs the difference between the value of a pixel and its adjacent pixel). Prueout filter (a filter that combines the properties of a differential filter and an averaging (smoothing) filter, and takes a simple average when smoothing). Sobel filter (a filter that combines the properties of a differential filter and an averaging (smoothing) filter, and takes a weighted average when smoothing). Second-order differential filter (an edge extraction filter). Laplacian filter (a filter that combines a horizontal second-order differential filter and a vertical second-order differential filter). Log (Laplacian of Gaussian) filter (a filter that performs weighted averaging with a Gaussian filter and detects edges with a Laplacian filter).
[0050] The sharpening unit 13 may use any neural network (not shown) capable of emphasizing regions corresponding to tissues (e.g., bones) in the medical image sharpening process. In this case, the neural network generates a sharpened image from a medical image showing the tissues of the first subject, in which the regions corresponding to the tissues of the first subject are emphasized. For example, the sharpening unit 13 may perform a sharpening process and / or a detail enhancement process to emphasize bone regions (regions corresponding to the tissues of the first subject).
[0051] When generating a training image from multiple medical images, the sharpening unit 13 may input each of the multiple medical images into a different neural network to generate a sharpened image, or may combine multiple medical images into one and then input the combined medical images into a neural network to generate a sharpened image.
[0052] In the sharpening process of the medical image, the sharpening unit 13 may apply a sharpening filter with a predetermined sharpening strength to the medical image, or may perform unsharp masking on the medical image. The unsharp masking is a process of sharpening the original image by taking the difference between an image smoothed by applying a low-pass filter to the original image and the original image after applying a weight to the original image.
[0053] Regardless of the type of device used to capture a medical image, all medical images are formatted in the Digital Imaging and Communications in Medicine (DICOM) format, an international standard for medical data communication. DICOM data includes medical image data and information about the device used to capture the medical image. Therefore, the sharpening process for a medical image may be configured to change depending on the device used to capture the medical image.
[0054] (Processing Performed by Image Generating Device 1) Next, the processing performed by the image generating device 1 will be described with reference to FIGS. 2 and 3. FIG. 2 is a flowchart showing an example of the flow of a method for generating training images for machine learning. FIG. 3 is a schematic diagram showing an example of the flow of processing performed by the sharpening unit 13. FIG. 3 shows an example in which the training images have three channels (e.g., R, G, and B channels). However, the number of channels of the training images is not limited to three, and may be two or less, or four or more.
[0055] Although the present disclosure will be described with reference to an example in which the subject is a human, the subject is not limited to humans. The subject may be a non-human mammal, such as an equine, feline, canine, bovine, or porcine animal, or may be a non-mammalian animal (e.g., a bird, reptile, amphibian, or fish).
[0056] First, the acquisition unit 11 acquires multiple original images (Step S11: Acquisition Step). Next, the selection unit 12 selects medical images to be used in training the machine learning model 702 as first images using a method for evaluating the quality or commonality of each medical image (Step S12: Acquisition Step). Next, the sharpening unit 13 combines the selected multiple first images, for example, in the channel direction, to generate a first combined image (Step S13: Combination Step). Then, the sharpening unit 13 performs a sharpening process on the first images corresponding to at least one channel of the first combined image to generate training images to be used in training the machine learning model 702 (Step S14: Sharpening Step).
[0057] FIG. 3 shows an overview of the process of generating training images using original images A to C from among the multiple original images selected by the selection unit 12. The sharpening unit 13 first generates a first combined image P by combining original images A to C in the channel direction. The sharpening unit 13 then applies a sharpening filter Fa with a first sharpening strength to original image A, a sharpening filter Fb with a second sharpening strength to original image B, and a sharpening filter Fc with a third sharpening strength to original image C. In this way, the sharpening unit 13 generates training images including images Af to Cf after the sharpening process (sharpened images). That is, the training images may be formed by combining multiple original images that have been subjected to sharpening processes with different sharpening strengths in the channel direction. In this case, the sharpening strength may be determined by a predetermined parameter. Alternatively, the sharpening strength may be determined by a combination of multiple parameters. The sharpening strength may be changed depending on the pixel values of the original images. The learning image is not limited to the example shown in FIG. 3, and may be a combination of a plurality of original images that have been subjected to sharpening processing with the same sharpening strength in the channel direction.
[0058] While FIG. 3 illustrates an example in which sharpening processing is performed on all original images A to C combined in the channel direction, the sharpening unit 13 is not limited to this configuration. For example, the sharpening unit 13 may perform sharpening processing on an original image corresponding to at least one channel among the original images A to C combined in the channel direction. In other words, the sharpening unit 13 performs sharpening processing on at least one of the multiple original images combined in the channel direction. The training images generated in this manner include images (sharpened images) after sharpening processing that have a different sharpness from the original images. When a machine learning model (NN) is trained using such training images, the machine learning model can select and train images with an appropriate level of sharpening for each part of the subject depicted in the original images. This can improve the training accuracy of the machine learning model and the estimation accuracy of the trained machine learning model.
[0059] The sharpening unit 13 may output the generated new learning image to the storage unit 20 or to an external device. The output destination may be the display device 50, the learning device 70, a storage device such as a database, or another image processing device. Data may be transmitted to an external device via the Internet.
[0060] [Embodiment 2] A second embodiment of the present disclosure will be described below. For ease of explanation, components having the same functions as those described in the above embodiment will be denoted by the same reference numerals, and redundant description will not be repeated. While an example in which the sharpening unit 13 combines multiple original images in the channel direction and then performs the sharpening process has been described in Figures 2 and 3, the present disclosure is not limited to this configuration. The sharpening unit 13 may perform the sharpening process on multiple original images and then combine multiple images including the sharpened image in the channel direction.
[0061] Fig. 4 is a schematic diagram showing another example of the flow of processing executed by the sharpening unit 13. Like Fig. 3, Fig. 4 shows an overview of processing for generating learning images using original images A to C from among the multiple original images selected by the selection unit 12. The sharpening unit 13 first generates sharpened images Af to Cf by applying sharpening filters Fa to Fc to each of the original images A to C. Next, the sharpening unit 13 combines the sharpened images Af to Cf in the channel direction to generate a learning image.
[0062] While FIG. 4 illustrates an example in which sharpening processing is performed on all of original images A to C, the sharpening unit 13 is not limited to this configuration. For example, the sharpening unit 13 may perform sharpening processing on an original image corresponding to at least one channel of original images A to C. In other words, the sharpening unit 13 generates training images including at least one sharpened image with a different sharpness from that of the original image. When training a machine learning model (NN) using such training images, the machine learning model can select and train images with an appropriate level of sharpening for each part of the subject appearing in the original image. This can improve the training accuracy of the machine learning model and the estimation accuracy of the trained machine learning model.
[0063] 2 to 4, the configurations in which multiple images are combined in the channel direction before or after performing sharpening processing on multiple original images have been described. However, the generation method according to the present disclosure is not limited to these configurations. For example, a configuration in which processing to combine multiple images in the channel direction both before and after performing sharpening processing on the original images may be used.
[0064] [Example of Implementation by Software] The control blocks of the image generating device 1 (particularly the acquisition unit 11, the selection unit 12, and the sharpening unit 13) may be implemented by a logic circuit (hardware) formed on an integrated circuit (IC chip) or the like, or may be implemented by software. In the latter case, each function of the image generating device 1 is implemented by, for example, a computer that executes instructions of a program P, which is software.
[0065] An example of such a computer (hereinafter referred to as computer C) is shown in Figure 5. Computer C includes at least one processor C1 and at least one memory C2. Memory C2 stores a program P for operating computer C as image generation device 1. In computer C, processor C1 reads and executes program P from memory C2, thereby realizing each function of image generation device 1.
[0066] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PU), a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0067] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may further include a communication interface for transmitting and receiving data to and from other devices. The computer C may further include an input / output interface for connecting input devices such as a keyboard and a mouse, and / or output devices such as a display and a printer.
[0068] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communications network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0069] In addition, some or all of the functions of each of the control blocks can be realized by logic circuits. For example, integrated circuits in which logic circuits that function as each of the control blocks are formed are also included in the scope of the present disclosure. In addition, the functions of each of the control blocks can also be realized by, for example, a quantum computer.
[0070] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may run on the control device or on another device (for example, an edge computer or a cloud server).
[0071] The invention according to the present disclosure has been described above based on the drawings and examples. However, the invention according to the present disclosure is not limited to the above-described embodiments. In other words, the invention according to the present disclosure can be modified in various ways within the scope of the present disclosure, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the invention according to the present disclosure. In other words, it should be noted that a person skilled in the art can easily make various modifications or corrections based on the present disclosure. It should also be noted that these modifications or corrections are included in the scope of the present disclosure.
[0072] [Summary] A generation method according to aspect 1 of the present disclosure is a method for generating training images used in training a machine learning model, and includes a combining step of generating a first combined image by combining a plurality of first images that show at least a portion of a first subject, and a sharpening step of performing a sharpening process on the first images, wherein the training image is the first combined image that includes at least one of the first images after the sharpening process.
[0073] In the generation method according to aspect 2 of the present disclosure, in the above aspect 1, the combining step may be performed before the sharpening step.
[0074] In the generation method according to aspect 3 of the present disclosure, in the above aspect 1, the combining step may be performed after the sharpening step.
[0075] In a generation method according to aspect 4 of the present disclosure, in the above aspect 1, the combining step may be performed both before and after the sharpening step.
[0076] A generation method according to aspect 5 of the present disclosure, in any one of aspects 1 to 4 above, may be such that the first combined image is an image obtained by combining multiple first images in the channel direction.
[0077] A generation method according to aspect 6 of the present disclosure, in any of aspects 1 to 5 above, may be an image in which the training image is an image in which multiple first images after the sharpening process, each having a different sharpening strength, are combined in the channel direction.
[0078] A generation method according to aspect 7 of the present disclosure may be a system for estimating, from a second image showing at least a portion of a second subject, characteristics of at least a portion of a second subject that is shown in the second image, and / or characteristics of a portion of the second subject that is not shown in the second image, in any of aspects 1 to 6 above.
[0079] A generation method according to aspect 8 of the present disclosure is, in the above-described aspect 7, wherein the second image is a medical image showing the chest bones of the second subject, and the machine learning model may estimate bone density of the chest bones, femur bones, and / or lumbar bones of the second subject from the second image.
[0080] In the generation method according to aspect 9 of the present disclosure, in the above aspect 8, the bone density of the bone may be at least one of bone mineral density per unit area, bone mineral density per unit volume, percent of young adult mean, T-score, and Z-score.
[0081] A learning method according to aspect 10 of the present disclosure is a method for learning a machine learning model that estimates characteristics of at least a portion of a second subject from a second image that shows at least a portion of the second subject, and uses a learning image generated by the generation method described in any one of aspects 1 to 9 above as an explanatory variable, and uses characteristic information indicating the characteristics of at least a portion of the first subject corresponding to the learning image as a target variable.
[0082] The trained machine learning model according to aspect 11 of the present disclosure is a machine learning model generated by the learning method described in aspect 10 above.
[0083] A generating device according to aspect 12 of the present disclosure is a generating device that generates training images used in training a machine learning model, and includes a combining unit that generates a first combined image by combining multiple first images that depict at least a portion of a first subject, and a sharpening unit that performs a sharpening process on the first images, and the training image is the first combined image that includes at least one of the first images after the sharpening process.
[0084] A control program according to aspect 13 of the present disclosure is a control program for causing a computer to function as the generating device described in aspect 12 above, and is a control program for causing a computer to function as the combining unit and the sharpening unit.
[0085] A recording medium according to aspect 14 of the present disclosure is a computer-readable non-transitory recording medium on which the control program according to aspect 13 above is recorded.
[0086] 1 Image generating device (generating device) 11 Acquisition unit 13 Sharpening unit 702 Machine learning model S13 Combining step S14 Sharpening step
Claims
1. A generation method comprising: a combining step of generating a first combined image by combining a plurality of first images, each of which shows at least a portion of a first subject; and a sharpening step of performing a sharpening process on the first images, wherein the first combined image, which includes at least one of the first images after the sharpening process, is a training image used for training a machine learning model.
2. The method of claim 1, wherein the combining step is performed before the sharpening step.
3. The method of claim 1, wherein the combining step is performed after the sharpening step.
4. The method of claim 1, wherein the combining step is performed both before and after the sharpening step.
5. A generation method according to any one of claims 1 to 4, wherein the first combined image is an image obtained by combining multiple first images in the channel direction.
6. A generation method according to any one of claims 1 to 5, wherein the learning image is an image in which multiple first images after the sharpening process, each having a different sharpening strength, are combined in the channel direction.
7. The method of any one of claims 1 to 6, wherein the machine learning model estimates, from a second image in which at least a portion of the second object is visible, characteristics of at least a portion of the second object visible in the second image and / or characteristics of a portion of the second object not visible in the second image.
8. The generation method of claim 7, wherein the second image is a medical image showing the chest bones of the second subject, and the machine learning model estimates bone density of the chest bones, femur bones, and / or lumbar bones of the second subject from the second image.
9. The method of claim 8, wherein the bone mineral density of the bone is at least one of bone mineral density per unit area, bone mineral density per unit volume, percent of young adult mean, T-score, and Z-score.
10. A method for learning a machine learning model that estimates at least a portion of a characteristic of a second subject from a second image that shows at least a portion of the second subject, using training images generated by the generation method described in claim 1 as explanatory variables and characteristic information indicating the characteristic of at least a portion of the first subject corresponding to the training images as a target variable.
11. A trained machine learning model generated by the learning method described in claim 10.
12. A generating device comprising: a combining unit that generates a first combined image by combining a plurality of first images that show at least a portion of a first subject; and a sharpening unit that performs a sharpening process on the first images, wherein the first combined image, which includes at least one of the first images after the sharpening process, is a training image used for training a machine learning model.
13. A control program for causing a computer to function as the generating device according to claim 12, the control program causing the computer to function as the combining unit and the sharpening unit.
14. A computer-readable non-transitory recording medium on which the control program according to claim 13 is recorded.
Citation Information
Patent Citations
In-vivo dwelling object detection device and in-vivo dwelling object detection method
JP2021137191A
Image processing device, method for operating image processing device, and program
JP2023154994A