Learning method, learning device, learning system, control program, and recording medium

The learning method improves region designation accuracy in medical images by generating and modifying attention maps to focus on relevant areas, enhancing training and estimation of subject information like bone density.

WO2025164632A1PCT designated stage Publication Date: 2025-08-07KYOCERA CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/002671
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2025-01-29
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing image analysis methods, particularly in medical imaging, suffer from low accuracy in designating regions of interest due to reliance on image information alone, leading to decreased training accuracy when unnecessary parts of the image are included.

Method used

A learning method that generates a first feature map from a medical image, calculates a first attention map indicating a region of interest, modifies it with a second attention map, and uses these maps to generate a second feature map for improved accuracy in training a learning model.

Benefits of technology

Enhances the accuracy of region designation in medical images by focusing on relevant areas, thereby improving the training and estimation of information related to the subject, such as bone density or abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025002671_07082025_PF_FP_ABST
    Figure JP2025002671_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a learning method capable of improving accuracy of region designation by means of an attention map. This learning method includes: a step for generating, from a medical image, a first feature map indicating a feature amount of the medical image; a step for calculating, from the medical image, a first attention map indicating a region of interest in the medical image; a step for calculating, from the medical image, a second attention map for correcting the first attention map; a step for generating a second feature map, in which the first feature map is weighted, by using the first attention map and the second attention map; and a step for training, by using the second feature map, a learning model for estimating information relating to a portion of a subject appearing in the medical image.
Need to check novelty before this filing date? Find Prior Art

Description

Learning method, learning device, learning system, control program, and recording medium

[0001] The present disclosure relates to a learning method, a learning device, a learning system, a control program, and a recording medium.

[0002] A technique is known in which, when extracting and classifying information from an image, a neural network automatically designates an image region that should be focused on (attention map) and classifies the image.

[0003] Non-Patent Document 1 describes a Residual Attention Network that mechanically calculates an attention map at each of the three stages of the Attention Module, i.e., the beginning, middle, and end, and multiplies the attention map with a feature map of the image.

[0004] Fei Wang, et al., “Residual Attention Network for Image Classification”, Proc. in Computer Vision and Pattern Recognition Conference(CVPR), 2017.

[0005] A learning method according to one aspect of the present disclosure includes the steps of: generating a first feature map from a medical image, the first feature map indicating a feature of the medical image; calculating a first attention map from the medical image, the first attention map indicating a region of interest in the medical image; calculating a second attention map from the medical image to modify the first attention map; using the first attention map and the second attention map to generate a second feature map by weighting the first feature map; and using the second feature map to train a learning model that estimates information regarding a part of a subject appearing in the medical image.

[0006] A learning model according to one aspect of the present disclosure is a trained learning model generated by the above-described learning method.

[0007] A learning device according to one aspect of the present disclosure includes a first feature map generation unit that generates a first feature map indicating features of a medical image from the medical image; a first attention map calculation unit that calculates a first attention map indicating a region of interest in the medical image from the medical image; a second attention map calculation unit that calculates a second attention map from the medical image for correcting the first attention map; a second feature map generation unit that uses the first attention map and the second attention map to generate a second feature map by weighting the first feature map; and a learning unit that uses the second feature map to train a learning model that estimates information regarding parts of a subject appearing in the medical image.

[0008] A learning system according to one aspect of the present disclosure includes an input unit that inputs a medical image, a first feature map generation unit that generates a first feature map indicating features of the medical image from the medical image, a first attention map calculation unit that calculates a first attention map indicating a region of interest in the medical image from the medical image, a second attention map calculation unit that calculates a second attention map from the medical image to modify the first attention map, a second feature map generation unit that uses the first attention map and the second attention map to generate a second feature map by weighting the first feature map, a learning unit that uses the second feature map to train a learning model that estimates information related to a part of a subject appearing in the medical image, and a memory unit that stores the trained learning model.

[0009] The learning device according to each aspect of the present disclosure may be realized by a computer. In this case, the scope of the present disclosure also includes a control program for the learning device that causes the computer to operate as each unit (software element) of the learning device, and a computer-readable recording medium on which the program is recorded. The computer may also include an edge computer, a cloud server, etc. Furthermore, the functions of the learning device may be realized by multiple devices.

[0010] 1 is a block diagram showing an example of the configuration of a learning system according to embodiment 1. FIG. 2 is an example of a diagram schematically showing the processing of the learning system according to embodiment 1. FIG. 3 is a diagram showing an example of a segmentation information creation unit. FIG. 4 is an example of a flowchart for explaining the processing procedure of the learning system according to embodiment 1. FIG. 4 is a block diagram showing an example of the configuration of a learning system according to embodiment 2. FIG. 5 is an example of a diagram schematically showing the processing of the learning system according to embodiment 2. FIG. 6 is a block diagram showing an example of the configuration of a part of a learning device according to embodiment 2. FIG. 7 is an example of a diagram (part 1) schematically showing the processing of the learning device according to embodiment 2. FIG. 8 is an example of a diagram (part 2) schematically showing the processing of the learning device according to embodiment 2. FIG. 9 is a block diagram showing an example of the configuration of a learning system according to embodiment 3. FIG. 10 is an example of a diagram schematically showing the processing of the learning device according to embodiment 3. FIG. 11 is a block diagram showing an example of the configuration of a learning system according to embodiment 4. FIG. 12 is an example of a diagram schematically showing the processing of the learning device according to embodiment 4. FIG. 13 is an example of a block diagram showing the configuration of a computer that executes instructions of a program that realizes a control block of the learning device.

[0011] If the designation of a region of interest in an image relies solely on image information, the accuracy of the region designation may be low.

[0012] One aspect of the present disclosure has been made in consideration of the above-mentioned problems, and provides a learning method that can improve the accuracy of area designation using an attention map.

[0013] First Embodiment A learning system 100 according to an embodiment of the present disclosure will be described in detail below with reference to the drawings.

[0014] The learning system 100 is a system for generating a learning model that estimates information about a region of a subject that appears in a medical image. Medical images typically use images that are actually captured in a medical setting. In addition to the subject's target object, actual medical images may also contain human tissue other than the region where the presence or absence of an abnormality is desired to be determined, or lesions that have occurred in regions other than the target region. In other words, the above-mentioned target objects are examples of subjects contained in each image, such as a medical image.

[0015] Although the present disclosure will be described with reference to an example in which the subject is a human, the subject is not limited to humans. The subject may be a non-human mammal, such as an equine, feline, canine, bovine, or porcine animal, or may be a non-mammalian animal (e.g., a bird, reptile, amphibian, or fish).

[0016] Furthermore, the part of the subject may be, for example, a target tissue and / or a target organ. The target tissue may be, for example, at least one of epithelial tissue, connective tissue, muscle tissue, and nervous tissue. The target organ may be, for example, at least one of the digestive system, the cardiovascular system, the endocrine system, and the musculoskeletal system. The part of the subject may be, for example, a joint and / or a bone.

[0017] Types of medical images include X-ray images, MRI (Magnetic Resonance Imaging) images, and CT (Computed Tomography) images. The image quality of such medical images varies depending on the imaging device, imaging method, imaging conditions, and the like. When training a machine learning model, the entire image, including unnecessary parts of the actual medical image (original medical image), is used for training, and therefore training is performed including various elements unrelated to the intended training. Therefore, if the original image is used for training as is, there is a risk of a decrease in training accuracy.

[0018] The medical image may be, for example, an image of the subject captured with an endoscope. More specifically, the medical image may include an endoscopic image of at least one of the subject's nasal cavity, esophagus, stomach, duodenum, rectum, large intestine, small intestine, anus, and colon. Medical images of these areas may output analysis results that clearly indicate, using a learning model, areas of interest that include at least one of inflammation, polyps, and cancer. In such cases, the learning model may be, for example, a model trained based on first learning images including images of the areas of interest and first training data indicating the presence of the areas of interest, and second learning images including images without the areas of interest and second training data indicating the absence of the areas of interest. The first training data may include information indicating the degree of inflammation (degree of inflammation) or malignancy (degree of malignancy) of the areas of interest. The analysis result may, for example, be displayed by surrounding the areas of interest, pointing to the areas of interest, or superimposing a color on the areas of interest. Derivation basis information indicating the basis for deriving the analysis result may be displayed together with the analysis result.

[0019] Alternatively, the medical images may be, for example, images of the subject's eyes, skin, etc. captured with a digital camera. Medical images of these areas may output analysis results that clearly indicate signs of interest using a learning model. For example, if the signs of interest are of the eye, they may include signs indicating diseases such as at least one of glaucoma, cataracts, age-related macular degeneration, conjunctivitis, hordeolum, retinopathy, and blepharitis. Alternatively, if the signs of interest are of the skin, they may include signs indicating skin cancer, hives, atopic dermatitis, herpes, etc. The analysis results may be displayed as a display that surrounds these signs of interest, a display that indicates the signs of interest, a display that superimposes a color on the signs of interest, or a display that indicates the name of the disease. For example, the learning model may be a model trained based on first training images including images of the areas of the signs of interest and first training data indicating the presence of the signs of interest, and second training images including images without the signs of interest and second training data indicating the absence of the signs of interest. Along with such analysis results, derivation basis information indicating the basis on which the analysis results were derived may be displayed.

[0020] The learning device 1 is a device that trains a machine learning model that performs a predetermined estimation from a second image. An image used to train the machine learning model is a first image, and an image used to estimate bone density using the machine learning model is a second image, both of which are medical images. In one aspect, the learning device 1 trains a machine learning model that estimates the bone density of a bone from a second image that depicts at least a portion of the bone. Examples of major bone densities include the bone densities of the lumbar vertebrae, femur, calcaneus, and radius. The machine learning model reads the first image for training and trains so that its output matches a specific bone density associated and annotated with the medical image.

[0021] Here, the target image is an image in which at least a portion of a predetermined object to be estimated is captured as a subject. The target image may be an X-ray image including a plain X-ray image, an MRI image, a CT image, or the like. Furthermore, the subject captured in the target image may be the same as the subject. The captured regions of the target image and the learning images may include, for example, at least one of the head, neck, chest, lower back, temporomandibular joint, spinal intervertebral joint, hip joint, sacroiliac joint, knee joint, ankle joint, foot, toes, shoulder joint, acromioclavicular joint, elbow joint, wrist joint, hand, fingers, and temporomandibular joint. Furthermore, the target image may be a frontal image or a lateral image.

[0022] The X-ray image data may be a simple X-ray image, such as a lumbar X-ray or a chest X-ray, or may be an X-ray image taken using a dual energy X-ray absorptiometry (DXA) or microdensitometry (MD) method. When measuring bone density in the lumbar spine with a DXA device, X-rays are irradiated from the front of the subject's lumbar spine. When measuring bone density in the proximal femur with a DXA device, X-rays are irradiated from the front of the subject's proximal femur. Here, the terms "front of the lumbar spine" and "front of the proximal femur" refer to the correct orientation of the imaging site, such as the lumbar spine or the proximal femur, and may be directed toward the ventral side or the back of the subject's body. The proximal femur includes, for example, at least one of the neck, trochanter, shaft, and the entire proximal femur (neck, trochanter, and shaft). In the MD method, X-rays are irradiated onto the hand. The image data does not have to be an X-ray image, but may be any image containing bone information. For example, the image may be estimated from an MRI image, a CT image, a PET (Positron Emission Tomography) image, an ultrasound image, or the like. The medical image may also be, for example, a dental image. The medical image may be, for example, an inspection device image acquired from an inspection device (e.g., an X-ray inspection device), or an image in which noise has been reduced from the inspection device image.

[0023] The machine learning model may also learn bone density, etc., of a portion of a bone not shown in a first image of a first subject, based on the first image showing at least a portion of the bone. The machine learning model may also estimate bone density, etc., of a portion of a bone not shown in a second image showing at least a portion of the bone. Additionally, the machine learning model may be configured to receive a third image of a predetermined portion of the body of a second subject, and learn and estimate bone density, which is a characteristic of the second subject. The machine learning model may also be configured to receive not only each image but also information about the second subject, such as age and gender, and perform learning and estimation. The machine learning model may also be configured to determine whether the second subject has osteoporosis. The machine learning model may determine osteoporosis based on, for example, proprietary standards or already known guidelines.

[0024] For example, the machine learning model may learn bone density of the femur and / or lumbar bones from training images showing chest bones, or may estimate bone density of the femur and / or lumbar bones from target images showing chest bones.

[0025] Furthermore, for example, the machine learning model may learn the bone density of the femur and / or lumbar bones from a training image showing lumbar bones, or may estimate the bone density of the femur and / or lumbar bones from a target image showing lumbar bones.

[0026] Additionally, the machine learning model may be configured to receive input of images of a predetermined part of the subject's body, and to learn and estimate the subject's bone density. Also, the machine learning model may be configured to receive input of not only each image but also information about the subject, such as age and gender, and to learn and estimate.

[0027] Bone mineral density may be a value related to the density of bone. Bone mineral density is defined as bone mineral density per unit area (g / cm 2 ), bone mineral density per unit volume (g / cm 3The bone mineral density may be expressed by at least one of the following: YAM (%), T-score, and Z-score. YAM (%) is an abbreviation for "Young Adult Mean" and may be referred to as the young adult mean percentage. For example, bone mineral density may be expressed by bone mineral density per unit area (g / cm 2 ) and YAM (%). The bone mineral density may be an index defined by a guideline or a unique index. For example, the bone mineral density may be a value used in osteoporosis guidelines (such as, but not limited to, the 2015 edition of the Prevention and Treatment Guidelines of the Japan Osteoporosis Society).

[0028] The bone density may be determined using a DXA device, an X-ray device, an ultrasonic bone density measuring device, etc. The bone density may be a value estimated using a device that estimates bone density from a plain X-ray image.

[0029] Furthermore, the description of the learning system 100 in the present disclosure is not limited to a machine learning model that estimates bone density, but can also be applied to a machine learning model that estimates bone abnormalities. The bone abnormality may be, for example, a bone abnormality related to a bone disease. The bone disease may be, for example, osteoporosis, scoliosis, fracture, spinal stenosis, intervertebral disc degeneration, ankylosing spondylitis, spinal cord injury, osteomyelitis, osteophyte, spinal muscular atrophy, etc.

[0030] Furthermore, the predetermined estimation is not limited to estimation of bone density, but may also be estimation of characteristics of a subject, such as the component ratio of bone or other substances. Furthermore, the machine learning model may use a medical image of the subject as an explanatory variable and characteristics of the subject at a time different from the time the medical image was captured as a target variable, to predict a time point different from the current estimation result of the subject's characteristics estimated from the image data. The different time point may be a time point in the future and / or in the past from the time the image data was captured. In the following explanation, the learning device 1 will be described as performing various processes on medical images, but this is not limited thereto, and images related to biology, engineering, etc. may also be used as the target.

[0031] Furthermore, the description of the machine learning model in this disclosure is not limited to a configuration that makes a predetermined estimation from medical images, etc., but can also be applied to a machine learning model that determines the identity of an object appearing in an input image.

[0032] Fig. 1 is a block diagram showing an example configuration of a learning system 100 including a learning device 1 according to this embodiment. As shown in Fig. 1, the learning system 100 includes a learning device 1, an image input unit 2, an input / output interface (IF) 3, an input device 4, and a display device 5. The learning device 1 may be a cloud-based device, or an on-premise device installed in a medical facility or a company that provides analysis services.

[0033] The image input unit 2 inputs, as a medical image, for example, an X-ray image of a subject from an X-ray device (not shown). The learning device 1 trains a learning model that estimates information about a part of the subject that appears in the X-ray image from the X-ray image input by the image input unit 2. Details of the learning device 1 will be described later.

[0034] The input / output IF3 is an interface for communicating with the outside world. The input / output IF3 may be a wired communication interface such as a USB port, or a wireless communication interface such as Bluetooth (registered trademark) or Wi-Fi (registered trademark). Furthermore, a user may be able to input data to the learning device 1 via an input device 4 such as a mouse or keyboard connected to the input / output IF3. Here, the term "user" refers to a medical professional or technician who uses the learning device 1, and this applies hereinafter as well. Furthermore, the learning device 1 may be connected to the Internet via the input / output IF3 to communicate information.

[0035] The present disclosure also includes a configuration in which the learning system 100 does not have an input / output IF 3 and the learning device 1 is integrated with a display device 5 capable of displaying medical images, etc.

[0036] The learning device 1 includes a segmentation information creation unit 11, a preprocessing unit 12, a first feature map generation unit 13, a first attention map calculation unit 14, a second attention map calculation unit 15, a second feature map generation unit 16, a learning unit 17, and a memory unit 18.

[0037] FIG. 2 is an example diagram schematically illustrating the processing of the learning device 1 according to the first embodiment. The components of the learning device 1 shown in FIG. 1 will be described with reference to FIG. 2 as needed. As shown in FIG. 2, the segmentation information creation unit 11 creates segmentation information for identifying regions appearing in an input image (X-ray image) 201 input by the image input unit 2. The attention region correction information estimation result 202 is information generated from the input image 201 with reference to the segmentation information. For example, if the thoracic vertebrae are designated as the attention region of the input image 201, the attention region correction information estimation result 202 is generated so that the region corresponding to the thoracic vertebrae in the input image 201 becomes the attention region, as shown in FIG. 2. For example, the segmentation information creation unit 11 may be configured using a U-Net. The region designated as the attention region may be predetermined by the user, as described above.

[0038] Figure 3 is a diagram for explaining the overview of U-Net. U-Net is classified as a semantic segmentation model within deep learning. Semantic segmentation is the division of an object into multiple regions at the pixel level. U-Net is a type of FCN (Fully Convolution Network) without a fully connected layer, and is a convolutional neural network with a U-shaped structure that performs fast and accurate image segmentation. Unlike ordinary CNNs, U-Net does not have a fully connected layer like FCNs and SegNets, and is composed of convolutional layers. U-Net is a model composed of an encoder and a decoder.

[0039] As shown in Figure 3, the downsampling encoder is composed of four 3x3 convolutional layers 111-114, and the number of convolutional channels is increased after each max-pooling. The feature map after each max-pooling is combined with a feature map of the same spatial size in the decoder.

[0040] The decoder for upsampling is composed of four steps of 3x3 convolutional layers 116 to 119. The decoder combines features from the encoder input via skip connections, reduces the number of channels for each convolutional layer 116 to 119, and creates a probability map (segmentation information) of the same size as the input image.

[0041] 2, the pre-processing unit 12 performs image processing to remove unnecessary areas from an input image 201, thereby generating image correction 203. Also, the pre-processing unit 12 performs image processing to remove unnecessary areas from an attention area correction information estimation result 202, thereby generating attention area correction information 204.

[0042] The first feature map generation unit 13 generates a first feature map 301 indicating the feature quantities of the original image (corrected image) 203. For example, when using the above-mentioned Residual Attention Network, the first attention map calculation unit 14 calculates a first attention map 302 indicating an attention region of the original image (corrected image) 203 by connecting an encoder-decoder structure similar to that of the segmentation information creation unit 11 to the feature map generation unit (ResNet). The attention region may be one region or multiple regions. The attention region may include not only a portion but also a certain extent of the region. The first feature map 301 and the first attention map 302 can be generated using, for example, the Residual Attention Network disclosed in the above-mentioned Non-Patent Document 1. The attention region is the portion to be noted indicated by the first attention map 302.

[0043] The second attention map calculation unit 15 calculates a second attention map 303 for correcting the first attention map 302 from the attention region correction information 204. Specifically, the second attention map calculation unit 15 calculates the second attention map by performing a convolution operation on the segmentation information and applying an activation function. For example, if the thoracic vertebrae are specified as the attention region, the second attention map calculation unit 15 calculates the second attention map 303 by referring to the segmentation information, as shown in FIG. 2, so that the region corresponding to the thoracic vertebrae in the input image becomes the attention region. The method is not limited to using convolution, a linear layer, and an activation function; other methods may be used, such as a method that uses only convolution and a linear layer without applying an activation function. The attention region correction information 204 of the second attention map 303 may include one or more attention regions.

[0044] The second feature map generation unit 16 generates a second feature map 304 by weighting the first feature map 301 using the first attention map 302 and the second attention map 303. As shown in FIG. 2 , the second feature map generation unit 16 multiplies the first attention map 302 by weight 1 and the second attention map 303 by weight 2. Then, the second feature map generation unit 16 multiplies the weight-multiplied first attention map 302 and the weight-multiplied second attention map 303 by the multiplier 101, and weights the first feature map 301 using the multiplication result. In this manner, the second feature map generation unit 16 generates the second feature map 304. Weight 1 and weight 2 may be fixed values ​​determined by the user as described above. The second feature map generation unit 16 calculates the second feature map such that a larger weighting value indicates a higher level of attention.

[0045] For example, three attention modules are described in the residual attention network disclosed in the aforementioned Non-Patent Document 1. The second feature map generation unit 16 may generate the second feature map 304 by multiplying one of the outputs of the three attention modules by the second attention map 303.

[0046] The learning unit 17 uses the second feature map 304 to train a learning model to estimate information (bone density estimation result 205) about the part of the subject appearing in the input image 201. The input image 201 and information (bone density) about the part of the subject appearing in the input image 201 are used as training data. The learning unit 17 then stores the trained learning model in the storage unit 18.

[0047] 4 is an example of a flowchart for explaining the processing procedure S1 of the learning system 100 according to embodiment 1. First, the image input unit 2 inputs an input image (X-ray image) 201 and outputs it to the segmentation information creation unit 11 and the preprocessing unit 12 (S11).

[0048] Next, the segmentation information creating unit 11 creates segmentation information for identifying regions appearing in the input image 201 (X-ray image) from the input image 201 input by the image input unit 2 (S12).

[0049] Next, the pre-processing unit 12 performs image processing to remove unnecessary areas from the input image 201, thereby generating image correction 203. The pre-processing unit 12 also performs image processing to remove unnecessary areas from the attention area correction information estimation result 202, thereby generating attention area correction information 204 (S13).

[0050] Next, the first feature map generation unit 13 generates a first feature map 301 indicating the feature amounts of the original image (corrected image) 203 (S14). Then, when using the above-mentioned Residual Attention Network, for example, the first attention map calculation unit 14 calculates a first attention map 302 indicating the attention region of the original image (corrected image) 203 by connecting an encoder-decoder structure similar to that of the segmentation information creation unit 11 to the feature map generation unit (S15).

[0051] Next, the second attention map calculation unit 15 calculates a second attention map 303 for correcting the first attention map 302 from the attention area correction information 204 (S16).

[0052] Next, the second feature map generation unit 16 generates a second feature map 304 by weighting the first feature map 301 using the first attention map 302 and the second attention map 303 (S17). This improves the accuracy of area designation using the attention map. That is, by multiplying the automatically created first attention map 302 by the second attention map 303 indicating specific attention area correction information 204, the accuracy of area designation improves.

[0053] The learning unit 17 uses the second feature map 304 to train a learning model to estimate information (bone density estimation result 205) about the part of the subject appearing in the input image 201 (S18). Then, the learning unit 17 stores the trained learning model in the storage unit 18 (S19).

[0054] [Embodiment 2] A second embodiment of the present disclosure will be described below. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and redundant description will not be repeated.

[0055] 5 is an example of a block diagram showing the configuration of a learning system 100A including a learning device 1A according to this embodiment. As shown in FIG. 5, the learning system 100A includes a learning device 1A, an image input unit 2, an input / output IF 3, an input device 4, and a display device 5.

[0056] The learning device 1A includes a segmentation information creation unit 11, a preprocessing unit 12, a first feature map generation unit 13, a first attention map calculation unit 14, a second attention map calculation unit 15A, a second feature map generation unit 16A, a learning unit 17A, and a memory unit 18.

[0057] Fig. 6 is an example diagram schematically illustrating the processing of the learning device 1A according to embodiment 2. The components of the learning device 1A shown in Fig. 5 will be described with reference to Fig. 6 as needed. As shown in Fig. 6, the segmentation information creation unit 11 creates segmentation information for identifying regions appearing in an input image 201 (X-ray image) input by the image input unit 2. An attention area correction information estimation result 202 is information generated from the input image 201 with reference to the segmentation information.

[0058] 6, the pre-processing unit 12 performs image processing to remove unnecessary areas from the input image 201, thereby generating image correction 203. Also, the pre-processing unit 12 performs image processing to remove unnecessary areas from the attention area correction information estimation result 202, thereby generating attention area correction information 204.

[0059] The first feature map generation unit 13 generates a first feature map 301 indicating the feature amount of the original image (corrected image) 203. For example, when the above-mentioned Residual Attention Network is used, the first attention map calculation unit 14 calculates a first attention map 302 indicating the attention area of ​​the original image (corrected image) 203 by connecting an encoder-decoder structure similar to that of the segmentation information creation unit 11 to the feature map generation unit.

[0060] The second attention map calculation unit 15A calculates a second attention map 303 for correcting the first attention map 302 from the attention area correction information 204. As will be described later, the second attention map calculation unit 15A generates an attention area map indicating an attention area of ​​the input image 201 and a background area map indicating a background area, based on the segmentation information. Then, the second attention map calculation unit 15A calculates the second attention map 303 by performing a convolution operation on the attention area map and the background area map.

[0061] The second feature map generation unit 16A generates a second feature map 304 by weighting the first feature map 301 using the first attention map 302 and the second attention map 303. As shown in FIG. 6 , the second feature map generation unit 16A multiplies the first attention map 302 by weight 1 and the second attention map 303 by weight 2. Then, the second feature map generation unit 16A multiplies the weight-multiplied first attention map 302 and the weight-multiplied second attention map 303 by the multiplier 101, and weights the first feature map 301 using the multiplication result. As described below, weight 1 and weight 2 are optimized by machine learning.

[0062] The learning unit 17A uses the second feature map 304 to train a learning model to estimate information (bone density estimation result 205) related to the part of the subject appearing in the input image 201. The learning unit 17A then stores the trained learning model in the storage unit 18. As will be described later, the learning unit 17A optimizes weight 1 and weight 2 through machine learning.

[0063] Fig. 7 is a block diagram showing an example of the configuration of a portion of a learning device 1A according to embodiment 2. Fig. 7 shows only the first feature map generation unit 13, the first attention map calculation unit 14, the second attention map calculation unit 15A, and the second feature map generation unit 16A of the learning device 1A.

[0064] The first feature map generation unit 13 includes Residual Modules 131-1 to 131-n. By connecting the Residual Modules 131-1 to 131-n in series, it is possible to train a highly accurate deep CNN. A Linear Module, a Transformer Block, an SE Block, or the like can also be used to generate a feature map. A feature map may also be generated by a feature extraction unit that does not use a Residual Module.

[0065] The first attention map calculation unit 14 includes an Encoder-Decoder 141, 1×1 convolution layers 142 and 143, and an activation function (Sigmoid function) 144. The Encoder-Decoder 141 downsamples the first feature map output from the Residual Module 131-2. The Encoder-Decoder 141 then upsamples the downsampled information while combining features from the encoder input via skip connections.

[0066] The 1×1 convolutional layers 142 and 143 perform convolutional operations on the information output from the encoder-decoder 141. The activation function 144 applies the activation function to the information output from the 1×1 convolutional layer 143 to calculate a first attention map.

[0067] The second attention map calculation unit 15A calculates a second attention map 303 for correcting the first attention map 302 from the attention area correction information 204. Based on the segmentation information, the second attention map calculation unit 15A generates an attention area map 151 that indicates the attention area of ​​the input image 201, and a background area map 152 that defines the area other than the attention area of ​​the input image 201 as the background area. The attention area map 151 and the background area map 152 have the same number of pixels (w × h) as the input image 201.

[0068] The 1x1 convolutional layer 153 performs a convolution operation on the attention region map 151 and the background region map 152 to calculate a second attention map 303 having the same number of channels and number of pixels (wxh) as the first attention map.

[0069] The second feature map generation unit 16A includes multipliers 101 and 102 and an adder 103. The multiplier 101 multiplies the result of multiplying the first attention map 302 by weight 1 by the result of multiplying the second attention map 303 by weight 2. The multiplier 102 multiplies the first feature map 301 by the multiplication result of the multiplier 101, and the adder 103 adds the first feature map 301 to the multiplied result, thereby generating the second feature map 304.

[0070] 8 and 9 are examples of diagrams schematically illustrating the processing of the learning device according to embodiment 2. First, the learning unit 17A performs multiple rounds of learning while updating parameters of the first feature map generation unit 13 and the first attention map calculation unit 14 that constitute the neural network.

[0071] Thereafter, as shown in FIG. 8 , the learning unit 17A performs 100 epoch learnings while fixing the parameters of the first feature map generation unit 13 and the first attention map calculation unit 14 and updating the parameters of the second attention map calculation unit 15A. The number of epoch learnings is an example and is not limited to this. Epoch learning refers to performing learning multiple times using the same training data. Since deep learning requires a large number of layers, epoch learning is used to facilitate updating of weights between neurons.

[0072] 9 , the learning unit 17A performs 300 epoch learnings while fixing the parameters of the first feature map generation unit 13 and updating the parameters of the first attention map calculation unit 14 and the second attention map calculation unit 15A. By performing learning in this manner, weight 1 for the first attention map 302 and weight 2 for the second attention map 303 are updated, and weight 1 and weight 2 are optimized. The number of epoch learnings is an example and is not limited to this.

[0073] [Embodiment 3] A third embodiment of the present disclosure will be described below. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiments, and redundant description will not be repeated.

[0074] 10 is an example of a block diagram showing the configuration of a learning system 100B including a learning device 1B according to this embodiment. As shown in FIG. 10, the learning system 100B includes a learning device 1B, an image input unit 2, an input / output IF 3, an input device 4, and a display device 5.

[0075] The learning device 1B includes a segmentation information creation unit 11, a preprocessing unit 12, a first feature map generation unit 13, a first attention map calculation unit 14, a second attention map calculation unit 15B, a second feature map generation unit 16B, a learning unit 17B, and a memory unit 18.

[0076] Fig. 11 is an example diagram schematically illustrating the processing of the learning device 1B according to embodiment 3. The components of the learning device 1B illustrated in Fig. 10 will be described with reference to Fig. 11 as needed. The first feature map generation unit 13 generates a first feature map 301 indicating the feature amounts of the original image (corrected image) 203. The first attention map calculation unit 14 performs a convolution operation on the first feature map 301 and applies an activation function to calculate a first attention map 302 indicating an attention region of the original image (corrected image) 203.

[0077] The second attention map calculation unit 15A calculates a second attention map 303B for correcting the first attention map 302 from the attention area correction information 204. As shown in Fig. 11 , the second attention map calculation unit 15A calculates the second attention map 303B using weights 206 for each body part determined based on a plurality of anatomical structures.

[0078] 11, the weighting of the clavicle is "2.0", the weighting of the ribs is "1.0", the weighting of the sternum is "0.5", and the weighting of the thoracic vertebrae is "1.5". The second attention map calculation unit 15B calculates the second attention map 303B so that areas with higher weighting values ​​attract more attention. This makes it possible to set different values ​​of attention area correction information 204 for multiple body parts.

[0079] The second feature map generating unit 16B generates a second feature map 304B by weighting the first feature map 301 using the first attention map 302 and the second attention map 303B.

[0080] The learning unit 17B uses the second feature map 304B to train a learning model to estimate information (bone density estimation result 205) about the part of the subject appearing in the input image 201. Then, the learning unit 17B stores the trained learning model in the storage unit 18.

[0081] [Embodiment 4] A fourth embodiment of the present disclosure will be described below. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiments, and redundant description will not be repeated.

[0082] 12 is an example of a block diagram showing the configuration of a learning system 100C including a learning device 1C according to this embodiment. As shown in FIG. 12, the learning system 100C includes a learning device 1C, an image input unit 2, an input / output IF 3, an input device 4, and a display device 5.

[0083] The learning device 1C includes a segmentation information creation unit 11, a preprocessing unit 12C, a first feature map generation unit 13, a first attention map calculation unit 14, a second feature map generation unit 16C, a learning unit 17, and a memory unit 18.

[0084] Fig. 13 is an example of a diagram schematically illustrating the processing of the learning device 1C according to embodiment 4. Each component of the learning device 1C shown in Fig. 12 will be described with appropriate reference to Fig. 13. As shown in Fig. 13, the segmentation information creation unit 11 creates segmentation information for identifying regions appearing in an input image 201 (X-ray image) from an input image 201 input by the image input unit 2.

[0085] The pre-processing unit 12C refers to the segmentation information and creates an attention area correction information estimation result 202 for weighting the attention area of ​​the input image 201. Then, the pre-processing unit 12C weights the input image 201 using the attention area correction information estimation result 202, deletes unnecessary areas from the attention area correction information estimation result 202 by image processing, and generates an image correction 203.

[0086] The first feature map generation unit 13 generates a first feature map 301 indicating the feature amounts of the original image (corrected image) 203. The first attention map calculation unit 14 calculates an attention map 302 indicating the attention area of ​​the original image (corrected image) 203 by performing a convolution operation on the first feature map 301 and applying an activation function.

[0087] The attention map 302 may be corrected after it is calculated. The attention map may be corrected by creating and correcting the second attention map 303, as described in the first to third embodiments, by having a medical professional or technician directly correct the attention map, or by other methods. The second attention map 303 may be created after the attention map 302 is calculated. The second attention map 303 may be created before the attention map 302 is calculated. The second attention map 303 may be created simultaneously with the calculation of the attention map 302. The second feature map generation unit 16C uses the corrected attention map 302 to generate the second feature map 304 by weighting the first feature map 301.

[0088] [Example of Implementation by Software] The control blocks of the learning device 1 (1A to 1C) (particularly the segmentation information creation unit 11, preprocessing unit 12, first feature map generation unit 13, first attention map calculation unit 14, second attention map calculation unit 15, second feature map generation unit 16, and learning unit 17) may be implemented by logic circuits (hardware) formed on an integrated circuit (IC chip) or the like, or by software. In the latter case, each function of the learning device 1 is implemented, for example, by a computer that executes instructions of a program P, which is software.

[0089] An example of such a computer (hereinafter referred to as computer C) is shown in Figure 14. Computer C has at least one processor C1 and at least one memory C2. Memory C2 stores a program P for operating computer C as learning device 1. In computer C, processor C1 reads and executes program P from memory C2, thereby realizing each function of learning device 1.

[0090] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PU), a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0091] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may further include a communication interface for transmitting and receiving data to and from other devices. The computer C may further include an input / output interface for connecting input devices such as a keyboard and a mouse, and / or output devices such as a display and a printer.

[0092] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communications network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0093] In addition, some or all of the functions of each of the control blocks can be realized by logic circuits. For example, integrated circuits in which logic circuits that function as each of the control blocks are formed are also included in the scope of the present disclosure. In addition, the functions of each of the control blocks can also be realized by, for example, a quantum computer.

[0094] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI ​​may run on the control device or on another device (for example, an edge computer or a cloud server).

[0095] The invention according to the present disclosure has been described above based on the drawings and examples. However, the invention according to the present disclosure is not limited to the above-described embodiments. In other words, the invention according to the present disclosure can be modified in various ways within the scope of the present disclosure, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the invention according to the present disclosure. In other words, it should be noted that a person skilled in the art can easily make various modifications or corrections based on the present disclosure. It should also be noted that these modifications or corrections are included in the scope of the present disclosure.

[0096] [Summary] The learning method according to aspect 1 of the present disclosure includes the steps of generating a first feature map from a medical image, the first feature map indicating the feature quantities of the medical image; calculating a first attention map from the medical image, the first attention map indicating a region of interest in the medical image; calculating a second attention map from the medical image to modify the first attention map; using the first attention map and the second attention map to generate a second feature map by weighting the first feature map; and using the second feature map to train a learning model that estimates information about parts of a subject appearing in the medical image.

[0097] A learning method according to aspect 2 of the present disclosure may be a method in which, in the above-described aspect 1, the medical image is an image showing at least a portion of the bones of the subject, and the information relating to the part of the subject is the bone density of a specific part of the bone of the subject.

[0098] A learning method according to aspect 3 of the present disclosure may be a method in which, in the step of calculating the first attention map in aspect 1 or 2 above, the first attention map is calculated by performing a convolution operation on the first feature map and applying an activation function.

[0099] A learning method according to aspect 4 of the present disclosure may be a method in which, in the step of training the learning model in aspect 1 or 2 above, the learning model is trained using the medical image and information relating to the part of the subject shown in the medical image as training data.

[0100] A learning method according to aspect 5 of the present disclosure may be a method in which, in the step of calculating the second attention map in aspect 1 or 2 above, segmentation information is created from the medical image to identify the area appearing in the medical image, and the second attention map is calculated based on the medical image and the segmentation information.

[0101] A learning method according to aspect 6 of the present disclosure may be a method in which, in the step of calculating the second attention map in the above-mentioned aspect 5, an encoder consisting of multiple convolutional layers downsamples the medical image, and a decoder consisting of multiple convolutional layers creates the segmentation information by upsampling the downsampled information while combining features from the encoder input via skip connections.

[0102] A learning method according to aspect 7 of the present disclosure may be a method in which, in the step of training the learning model in the above-mentioned aspect 5, the first attention map and the second attention map are weighted and multiplied together, and the result of the multiplication is used to generate the second feature map.

[0103] A learning method according to aspect 8 of the present disclosure may be a method in which, in the step of calculating the second attention map in aspect 7 above, an attention area map indicating an attention area of ​​the medical image and a background area map indicating a background area are generated based on the medical image and the segmentation information, and the second attention map is calculated by performing a convolution operation on the attention area map and the background area map.

[0104] A learning method according to aspect 9 of the present disclosure may be a method in which, in the step of calculating the second attention map in aspect 8 above, the segmentation information is referenced, the attention area of ​​the medical image is extracted, and the attention area map is generated, and the segmentation information is referenced, and the background area map is generated with the area other than the attention area of ​​the medical image as the background area.

[0105] A learning method according to aspect 10 of the present disclosure may be a method in which, in the step of calculating the second attention map in aspect 5 above, multiple parts appearing in the medical image are identified based on the segmentation information, and the second attention map is calculated from the medical image based on weightings set for each of the multiple parts.

[0106] The trained learning model according to aspect 11 of the present disclosure is a learning model generated by the learning method described in aspect 1 or 2 above.

[0107] A learning device according to aspect 12 of the present disclosure includes a first feature map generation unit that generates a first feature map indicating the feature amounts of a medical image from the medical image; a first attention map calculation unit that calculates a first attention map indicating a region of interest in the medical image from the medical image; a second attention map calculation unit that calculates a second attention map from the medical image for correcting the first attention map; a second feature map generation unit that uses the first attention map and the second attention map to generate a second feature map by weighting the first feature map; and a learning unit that uses the second feature map to train a learning model that estimates information regarding a part of a subject appearing in the medical image.

[0108] A learning device according to aspect 13 of the present disclosure may be a device in which, in aspect 12 above, multiple rounds of learning are performed while updating parameters of the first feature map generation unit and the first attention map calculation unit that constitute a neural network, and then multiple rounds of epoch learning are performed while fixing the parameters of the first feature map generation unit and the first attention map calculation unit and updating the parameters of the second attention map calculation unit.

[0109] A control program according to aspect 14 of the present disclosure is a control program for causing a computer to function as the learning device described in aspect 12 or 13 above, and may be configured to cause the computer to function as the first feature map generation unit, the first attention map calculation unit, the second attention map calculation unit, the second feature map generation unit, and the learning unit.

[0110] A recording medium according to aspect 15 of the present disclosure may be a computer-readable non-transitory recording medium on which the control program according to aspect 14 above is recorded.

[0111] A learning system according to aspect 16 of the present disclosure includes an input unit for inputting a medical image, a first feature map generation unit for generating a first feature map indicating the feature amounts of the medical image from the medical image, a first attention map calculation unit for calculating a first attention map indicating a region of interest in the medical image from the medical image, a second attention map calculation unit for calculating a second attention map for correcting the first attention map from the medical image, a second feature map generation unit for generating a second feature map by weighting the first feature map using the first attention map and the second attention map, a learning unit for training a learning model that uses the second feature map to estimate information regarding a part of a subject appearing in the medical image, and a memory unit for storing the trained learning model.

[0112] REFERENCE SIGNS LIST 1, 1A to 1C Learning device 2 Image input unit 3 Input / output IF 4 Input device 5 Display device 11 Segmentation information creation unit 12, 12C Preprocessing unit 13 First feature map generation unit 14 First attention map calculation unit 15, 15A, 15B Second attention map calculation unit 16, 16A, 16B, 16C Second feature map generation unit 17, 17A, 17B Learning unit 18 Storage unit 100, 100A to 100C Learning system C Computer C1 Processor C2 Memory

Claims

1. A learning method comprising the steps of: generating a first feature map from a medical image, the first feature map indicating the feature quantities of the medical image; calculating a first attention map from the medical image, the first attention map indicating a region of interest in the medical image; calculating a second attention map from the medical image to modify the first attention map; using the first attention map and the second attention map to generate a second feature map by weighting the first feature map; and using the second feature map to train a learning model that estimates information about parts of a subject appearing in the medical image.

2. The learning method described in claim 1, wherein the medical image is an image showing at least a portion of the bones of the subject, and the information about the part of the subject is the bone density of a specific part of the bones of the subject.

3. A learning method as described in claim 1 or 2, wherein in the step of calculating the first attention map, the first attention map is calculated by performing a convolution operation on the first feature map and applying an activation function.

4. A learning method according to any one of claims 1 to 3, wherein in the step of training the learning model, the learning model is trained using the medical image and information relating to the part of the subject that appears in the medical image as training data.

5. A learning method described in any one of claims 1 to 4, wherein in the step of calculating the second attention map, segmentation information is created from the medical image to identify each of the parts appearing in the medical image, and the second attention map is calculated based on the medical image and the segmentation information.

6. The learning method of claim 5, wherein in the step of calculating the second attention map, an encoder consisting of multiple convolutional layers downsamples the medical image, and a decoder consisting of multiple convolutional layers creates the segmentation information by upsampling the downsampled information while combining features from the encoder input via skip connections.

7. The learning method described in claim 5, wherein in the step of training the learning model, the attention level of the first attention map and the attention level of the second attention map are weighted and multiplied together, and the second feature map is generated using the result of the multiplication.

8. The learning method described in claim 7, wherein in the step of calculating the second attention map, an attention region map indicating an attention region of the medical image and a background region map indicating a background region are generated based on the medical image and the segmentation information, and the second attention map is calculated by performing a convolution operation on the attention region map and the background region map.

9. The learning method described in claim 8, wherein in the step of calculating the second attention map, the segmentation information is referenced, the attention area of the medical image is extracted, and the attention area map is generated; and the segmentation information is referenced, and the background area map is generated, with areas of the medical image other than the attention area as background areas.

10. The learning method described in claim 5, wherein in the step of calculating the second attention map, multiple parts appearing in the medical image are identified based on the segmentation information, and the second attention map is calculated from the medical image based on weighting set for each of the multiple parts.

11. A trained learning model generated by the learning method of claim 1 or 2.

12. A learning device comprising: a first feature map generation unit that generates a first feature map indicating the feature amounts of a medical image from the medical image; a first attention map calculation unit that calculates a first attention map indicating a region of interest in the medical image from the medical image; a second attention map calculation unit that calculates a second attention map from the medical image for correcting the first attention map; a second feature map generation unit that uses the first attention map and the second attention map to generate a second feature map by weighting the first feature map; and a learning unit that uses the second feature map to train a learning model that estimates information about parts of a subject appearing in the medical image.

13. The learning device described in claim 12, wherein multiple rounds of learning are performed while updating the parameters of the first feature map generation unit and the first attention map calculation unit that constitute a neural network, and then multiple rounds of epoch learning are performed while fixing the parameters of the first feature map generation unit and the first attention map calculation unit and updating the parameters of the second attention map calculation unit.

14. A control program for causing a computer to function as a learning device as described in claim 12 or 13, the control program causing the computer to function as the first feature map generation unit, the first attention map calculation unit, the second attention map calculation unit, the second feature map generation unit and the learning unit.

15. A computer-readable non-transitory recording medium on which the control program according to claim 14 is recorded.

16. A learning system comprising: an input unit for inputting a medical image; a first feature map generation unit for generating, from the medical image, a first feature map indicating the feature quantities of the medical image; a first attention map calculation unit for calculating, from the medical image, a first attention map indicating a region of interest in the medical image; a second attention map calculation unit for calculating, from the medical image, a second attention map for correcting the first attention map; a second feature map generation unit for generating, using the first attention map and the second attention map, a second feature map by weighting the first feature map; a learning unit for training a learning model using the second feature map to estimate information regarding parts of a subject appearing in the medical image; and a memory unit for storing the trained learning model.

Citation Information

Patent Citations

  • Multi-target semantic segmentation method for echocardiography image

    CN115578360A

  • Estimation device

    JP2020171785A

  • Sensor-specific image recognition device and method

    JP2021093144A

  • Image processing device, image processing method and image processing program

    JP2023051399A

  • Information processing device, information processing method, and information processing program

    JP7395767B2