Machine learning program, machine learning method, and information processing device
By incorporating imaging device setting information into the training data and using it as input to the neural network, the machine learning program addresses the challenge of inconsistent imaging conditions, resulting in improved accuracy and reduced false positives in image analysis.
Patent Information
- Application Number
- PCT/JP2023/042923
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-05
AI Technical Summary
Existing machine learning models for image analysis, such as lesion detection in medical images, face performance degradation when trained with images having inconsistent imaging conditions, leading to poor inference performance and increased false positives.
A machine learning program that acquires training data incorporating both pixel information and setting information from imaging devices, allowing the model to be trained with diverse imaging conditions. This approach includes using device setting information as input to the neural network and employing post-processing models to reduce false positives.
The proposed solution improves the performance of machine learning models by accounting for varying imaging conditions, leading to enhanced accuracy and reduced false positives in image analysis tasks, particularly in medical image diagnosis.
Smart Images

Figure JP2023042923_05062025_PF_FP_ABST
Abstract
Description
Machine learning program, machine learning method, and information processing device
[0001] The present invention relates to a machine learning program, a machine learning method, and an information processing device.
[0002] There is an image analysis technology that utilizes AI (Artificial Intelligence). For example, medical institutions can use AI to diagnose the presence or absence of pathology based on CT (Computed Tomography) images. Image analysis using AI requires the collection of a large number of training images to generate an image analysis model (e.g., a neural network). The more consistent the imaging conditions of the collected images, the better the learning performance will be, but it can sometimes be difficult to standardize the imaging conditions of the images used for learning.
[0003] If the number of training images is insufficient, it becomes difficult to improve the performance of a machine learning model. For example, if the inference performance of an image-based lesion detection model is poor, the detected lesions are likely to contain false positives. A false positive is when something that is not a lesion is judged to be a lesion.
[0004] As an image analysis technology, a technology has been proposed that makes it possible to detect the emergence of desired changes in a more suitable manner, even in situations where various changes may become apparent in the target image.
[0005] Japanese Patent Application Laid-Open No. 2022-156604
[0006] Devices that capture images have various setting items that affect the captured image. If the values of the setting items differ, the feature quantities of the captured image will also differ. If the setting information indicating the values of each setting item at the time of capture of images used for training a machine learning model is common, a high-performance model can be generated based on those images. However, if the only images available for training a machine learning model are images with inconsistent values for each setting item at the time of capture, it will be difficult to improve the performance of the model.
[0007] In one aspect, the present invention aims to improve the performance of machine learning models.
[0008] One proposal provides a machine learning program that causes a computer to perform the following process: The computer acquires a plurality of first training data, each of which is training data that associates input feature amounts, including pixel information of a captured image and setting information of an imaging device that captured the captured image at the time of imaging, with correct answers of a determination process for the captured image; and the computer trains the machine learning model based on the results of the determination process for the captured image to be determined when the input feature amounts indicated in the plurality of first training data are input to the machine learning model, and on the correct answers indicated in the plurality of first training data.
[0009] According to one aspect, the performance of a machine learning model can be improved. These and other objects, features, and advantages of the present invention will become apparent from the following description taken in conjunction with the accompanying drawings illustrating preferred embodiments of the present invention.
[0010] 1 is a diagram illustrating an example of a machine learning method according to a first embodiment; FIG. 2 is a diagram illustrating an example of a system configuration according to a second embodiment; FIG. 3 is a diagram illustrating an example of server hardware; FIG. 4 is a diagram illustrating an example of image diagnosis support using AI; FIG. 5 is a diagram illustrating an example of a lesion location to be detected; FIG. 6 is a diagram illustrating an example of image diagnosis support including post-processing; FIG. 7 is a diagram illustrating an example of image diagnosis support for each of images from a plurality of CT devices; FIG. 8 is a diagram illustrating an example of image diagnosis support performing post-processing for each examination device; FIG. 9 is a diagram illustrating an example of image diagnosis support using device setting information; FIG. 10 is a block diagram illustrating the functions of each device for image diagnosis support; FIG. 11 is a diagram illustrating an example of noise removal effect information; FIG. 12 is a diagram illustrating an example of image diagnosis support processing using device setting information as input to a neural network;
[0011] The present embodiment will be described below with reference to the drawings. Note that each embodiment can be implemented by combining multiple embodiments within a consistent range. [First Embodiment] The first embodiment is a machine learning method that improves the judgment performance (accuracy rate, precision rate, etc.) of a machine learning model for performing a predetermined judgment process on a captured image, even when only images with inconsistent values for setting items of the device at the time of capture are available.
[0012] Fig. 1 is a diagram illustrating an example of a machine learning method according to a first embodiment. Fig. 1 shows an information processing device 10 used to implement the machine learning method. The information processing device 10 can implement the machine learning method by, for example, executing a machine learning program.
[0013] To implement the machine learning method, the information processing device 10 includes a memory unit 11 and a processing unit 12. The memory unit 11 is, for example, a memory or storage device included in the information processing device 10. The processing unit 12 is, for example, a processor or an arithmetic circuit included in the information processing device 10.
[0014] The storage unit 11 stores a training data set 2. The training data set 2 includes a plurality of training data 2a, 2b, .... The training data 2a is data in which an input feature 3 is associated with a correct answer 4 of a judgment process for a captured image 3a. The input feature 3 includes pixel information of the captured image 3a and setting information 3b of the imaging devices 1a, 1b, ... that captured the captured image 3a at the time of imaging. Each of the other training data 2b, ... also includes the same type of information as the training data 2a.
[0015] The captured image 3a is, for example, an image of the outside or inside of a human body. The inside of a human body can be imaged using X-rays or the like. An example of the imaging devices 1a, 1b, ... that use X-rays is a CT device. When the imaging devices 1a, 1b, ... are CT devices, the setting information 3b includes, for example, the X-ray dose and the type of filtering applied to the captured image 3a.
[0016] The processing unit 12 trains the machine learning model 5 using the training dataset 2. For example, the processing unit 12 acquires first training data corresponding to the captured image to be determined from the training dataset 2. Then, the processing unit 12 trains the machine learning model 5 based on the result of the determination process on the captured image to be determined when the input feature amounts indicated in the acquired first training data are input to the machine learning model 5, and on the correct answer indicated in the first training data.
[0017] In this way, the setting information of the imaging devices 1 a, 1 b, ... is used to train the machine learning model 5. This makes it possible to improve the performance of the machine learning model 5 even when only images with inconsistent values for each setting item at the time of shooting can be prepared as images for training the machine learning model.
[0018] For example, the processing unit 12 performs inference using a trained machine learning model 5. In this case, the processing unit 12 inputs, to the trained machine learning model 5, an inference target feature 6 including pixel information of a captured image 6a to be inferred and setting information 6b of the imaging device 1 that captured the captured image 6a at the time of imaging. Then, the processing unit 12 performs a judgment process on the captured image 6a to be inferred in accordance with the machine learning model 5. In this way, a high-performance judgment is made.
[0019] If the machine learning model 5 is a model that identifies the location of a lesion from an image of a human body captured by a CT device, improving the performance of the machine learning model 5 enables accurate determination of the location of the lesion. For example, if the imaging device 1 is a CT device, a CT image of the patient is obtained as the captured image 6a. In this case, the setting information 6b at the time of imaging includes, for example, the X-ray dose, the type of filter processing, etc. The processing unit 12 inputs the inference target feature 6 to the machine learning model 5 and obtains an inference result 7. The inference result 7 indicates the lesion location 7a. By performing inference using the setting information 6b, the likelihood that a lesion actually exists at the location indicated by the lesion location 7a increases. This reduces the doctor's time required for image interpretation and prevents the doctor from overlooking a lesion.
[0020] The processing unit 12 can use a model combining two machine learning models as the machine learning model 5 that performs a predetermined determination process. In this case, the processing unit 12 trains the first machine learning model 5a and the second machine learning model 5b. The first machine learning model 5a is a model that receives, for example, pixel information of a captured image to be determined and setting information of the imaging device that captured the captured image to be determined, and determines a location in a shadow image where a predetermined object appears. In this case, the second machine learning model 5b is a model that receives, as input, a location determined to contain an object, and determines whether the location determined to contain an object is true or false.
[0021] In this way, by using the setting information of the imaging devices 1 a, 1 b, ... in training the first machine learning model 5 a for determining locations where an object is captured, the performance of determining locations where an object is captured in inference using the first machine learning model 5 a is improved. Furthermore, the processing unit 12 can improve the precision by determining whether or not the locations determined by the second machine learning model 5 b to contain an object are true.
[0022] The first machine learning model 5a may be a model that receives pixel information of the captured image to be judged as an input and judges a location in the captured image where a predetermined object appears. In this case, the second machine learning model 5b receives, for example, the location determined to include the object and setting information of the imaging device that captured the captured image to be judged as an input and judges whether the location determined to include the object is true or false.
[0023] In this way, the setting information of the imaging devices 1 a, 1 b, ... is used to train the second machine learning model 5 b for determining the authenticity of a portion determined to contain an object, thereby improving the accuracy of determining the authenticity of a portion determined to contain an object in inference using the second machine learning model 5 b.
[0024] It should be noted that, while a large amount of training data for images captured by some of the multiple imaging devices 1a, 1b, ... can be prepared, there are cases where only a small amount of training data for images captured by the other imaging devices can be prepared. In such cases, there is a possibility that the machine learning model 5 will overfit (overlearn) to the setting information of some of the imaging devices. For example, if the machine learning model 5, the first machine learning model 5a, or the second machine learning model 5b is a neural network, overfitting is likely to occur due to bias in the content of the setting information included in the training data 2a, 2b, ....
[0025] Therefore, the processing unit 12 may sample the first training data acquired for use in training so as to prevent bias in the content of the setting information included in the training data 2a, 2b, .... For example, the processing unit 12 classifies each of the multiple training data 2a, 2b, ... into multiple groups based on the similarity of the setting information. Then, the processing unit 12 acquires the first training data evenly from each group. This prevents the machine learning model 5 from overfitting to setting information with specific content.
[0026] Second Embodiment In recent years, image diagnosis support technology using AI has been attracting attention. Image diagnosis support is expected to prevent oversight of lesions and shorten the time required for doctors to interpret images. Therefore, as a second embodiment, a technology for improving the diagnostic performance in image diagnosis support using AI will be described.
[0027] 2 is a diagram illustrating an example of a system configuration according to the second embodiment. For example, systems in a plurality of facilities 30, 31, ... are connected to a server 100 via a network 20. The server 100 is a computer that performs learning using AI.
[0028] A system including an AI workstation 200, inspection devices 30a and 30b, and a monitor 30c is constructed in a facility 30. The AI workstation 200 is connected to a network 20. The AI workstation 200 is a computer that uses a trained model to provide diagnostic support based on medical images.
[0029] The inspection devices 30a and 30b are connected to the AI workstation 200. The inspection devices 30a and 30b are devices used to examine patients and output images as the examination results. The inspection devices 30a and 30b are, for example, CT devices, MRI (Magnetic Resonance Imaging), ultrasound diagnostic devices, X-ray devices, etc. The inspection devices 30a and 30b are examples of the imaging devices 1a, 1b, ... shown in the first embodiment.
[0030] The monitor 30c is connected to the AI workstation 200. The monitor 30c is a device that acquires and displays images captured by the inspection devices 30a and 30b, images showing analysis results, and the like from the AI workstation 200.
[0031] A system similar to that of facility 30 is also constructed in each of facilities 31, etc. other than facility 30. The system including the server 100 and the AI workstations in each of facilities 31, etc. is an example of the information processing device 10 shown in the first embodiment.
[0032] The server 100 collects annotated images from the facilities 30, 31, ..., performs learning using the collected images, and generates a trained model. The generated model is transmitted from the server 100 to each of the facilities 30, 31, ..., and is used by the systems within the facilities 30, 31, .... For example, the AI workstation 200 in the facility 30 analyzes images acquired from the inspection devices 30a, 30b using the model acquired from the server 100, and determines the presence or absence of a lesion, etc.
[0033] 3 is a diagram illustrating an example of server hardware. The entire server 100 is controlled by a processor 101. A memory 102 and multiple peripheral devices are connected to the processor 101 via a bus 109. The processor 101 may be a multiprocessor. The processor 101 is, for example, a central processing unit (CPU), a micro processing unit (MPU), or a digital signal processor (DSP). At least some of the functions realized by the processor 101 executing a program may be realized by an electronic circuit such as an application specific integrated circuit (ASIC) or a programmable logic device (PLD).
[0034] The memory 102 is used as a main storage device of the server 100. The memory 102 temporarily stores at least a portion of the OS (Operating System) programs and application programs to be executed by the processor 101. The memory 102 also stores various data used in processing by the processor 101. The memory 102 may be a volatile semiconductor storage device such as a RAM (Random Access Memory).
[0035] The peripheral devices connected to the bus 109 include a storage device 103 , a GPU (Graphics Processing Unit) 104 , an input interface 105 , an optical drive device 106 , a device connection interface 107 , and a network interface 108 .
[0036] The storage device 103 electrically or magnetically writes and reads data to and from a built-in recording medium. The storage device 103 is used as an auxiliary storage device for the server 100. The storage device 103 stores an OS program, application programs, and various data. Note that the storage device 103 may be, for example, a hard disk drive (HDD) or a solid state drive (SSD).
[0037] The GPU 104 is an arithmetic unit that performs image processing. The GPU 104 is an example of a graphics controller. A monitor 21 is connected to the GPU 104. The GPU 104 displays an image on the screen of the monitor 21 in accordance with an instruction from the processor 101. The monitor 21 may be a display device using organic electroluminescence (EL) or a liquid crystal display device.
[0038] The input interface 105 is connected to a keyboard 22 and a mouse 23. The input interface 105 transmits signals sent from the keyboard 22 and the mouse 23 to the processor 101. The mouse 23 is an example of a pointing device, and other pointing devices can also be used. Examples of other pointing devices include a touch panel, a tablet, a touch pad, and a trackball.
[0039] The optical drive device 106 uses a laser beam or the like to read data recorded on an optical disc 24 or write data to the optical disc 24. The optical disc 24 is a portable recording medium on which data is recorded so that it can be read by reflected light. Examples of the optical disc 24 include a DVD (Digital Versatile Disc), a DVD-RAM, a CD-ROM (Compact Disc Read Only Memory), and a CD-R (Recordable) / RW (Rewritable).
[0040] The device connection interface 107 is a communication interface for connecting peripheral devices to the server 100. For example, a memory device 25 or a memory reader / writer 26 can be connected to the device connection interface 107. The memory device 25 is a recording medium equipped with a function for communicating with the device connection interface 107. The memory reader / writer 26 is a device for writing data to the memory card 27 or reading data from the memory card 27. The memory card 27 is a card-type recording medium.
[0041] The network interface 108 is connected to the network 20. The network interface 108 transmits and receives data to and from other computers or communication devices via the network 20. The network interface 108 is a wired communication interface that is connected by a cable to a wired communication device such as a switch or a router. The network interface 108 may also be a wireless communication interface that is connected by radio waves to a wireless communication device such as a base station or an access point.
[0042] The server 100 can realize the processing functions of the second embodiment with the above-described hardware. The AI workstation 200 can also be realized with the same hardware as the server 100. Moreover, the information processing device 10 shown in the first embodiment can also be realized with the same hardware as the server 100.
[0043] The server 100 realizes the processing functions of the second embodiment by executing a program recorded on, for example, a computer-readable recording medium. The program describing the processing to be executed by the server 100 can be recorded on various recording media. For example, the program to be executed by the server 100 can be stored in the storage device 103. The processor 101 loads at least a portion of the program from the storage device 103 into the memory 102 and executes the program. The program to be executed by the server 100 can also be recorded on a portable recording medium such as the optical disk 24, the memory device 25, or the memory card 27. The program stored on the portable recording medium becomes executable after being installed on the storage device 103 under the control of, for example, the processor 101. The processor 101 can also read and execute the program directly from the portable recording medium.
[0044] In order to implement AI-based image diagnosis support in society, it is important to ensure that the performance of judgments does not deteriorate, regardless of the type of examination equipment used to capture the images. Here, we will explain the difficulty of preventing performance deterioration.
[0045] FIG. 4 is a diagram showing an example of image diagnosis support using AI. For example, the AI workstation 200 acquires an image 41 of a patient from the inspection device 30a. The image 41 includes information such as the brightness of each of multiple pixels. For example, the image 41 shows a cross-section of the patient's body (such as the head). The AI workstation 200 executes image diagnosis support using AI and identifies a suspected lesion from the image 41. The AI workstation 200 displays an inference result screen 42 on the monitor, which shows the inference result of the lesion location using AI. In the example of FIG. 4, a suspected lesion of a cerebral aneurysm is displayed on the inference result screen 42.
[0046] The doctor 40 determines whether or not there is a lesion based on the cross-sectional image of the patient's head displayed on the inference result screen 42. In doing so, the doctor 40 particularly carefully observes areas that have been indicated as suspected of having a lesion by the image diagnosis support, and determines whether or not there is a lesion in the relevant areas.
[0047] In this way, image diagnosis support using AI can prevent doctors 40 from overlooking lesions. It also reduces the time it takes doctors 40 to interpret images. The accuracy of the judgment of areas identified as lesions by image diagnosis support using AI depends on the performance of the AI model. For AI image diagnosis support using CT images, a trained model is generated using the patient's CT images. However, there are various difficulties in learning using CT images.
[0048] For example, the characteristics of CT images change depending on the dose used when capturing the image or the filtering used during image reconstruction. Dose refers to the amount of X-rays irradiated to the patient. The higher the dose, the clearer the image, but at the same time, the greater the patient's exposure. For this reason, the dose is set for each CT device to strike an appropriate balance between image clarity and exposure. Filtering is a noise removal technique used when reconstructing images from X-ray map information. Each CT device manufacturer uses different filtering methods, and each method results in different images. Even if a model is trained based on images captured with different device settings, it is difficult to improve the model's performance.
[0049] On the other hand, when detecting lesions using AI, it is important not to overlook lesions. Therefore, in image diagnosis support, it is desirable to detect lesions as lesions even if there is even the slightest suspicion of a lesion. In this case, there is a possibility that areas that are not actually lesions may be detected as lesions.
[0050] Neural networks can be used as AI for lesion detection. In image diagnosis support using neural networks, it is common to improve the neural network's discrimination performance to reduce the number of false positives. However, medical images also contain personal patient information, making it difficult to obtain a large amount of training data. Therefore, improving the performance of neural networks by increasing the amount of training data is not easy. As a result, it is inevitable that areas that are not actually lesions will be detected as lesions.
[0051] FIG. 5 is a diagram showing an example of detected lesion locations. Three lesion locations 43a to 43c are displayed on image 43-1 on inference result screen 43. Of these, lesion location 43a is a location where the AI correctly detected the lesion. A lesion correctly detected as lesion location 43a is called a true positive. The other two lesion locations 43b and 43c are locations where the AI incorrectly detected the lesions. Lesions incorrectly detected like lesion locations 43b and 43c are called false positives.
[0052] As shown in Figure 5, a single image may contain multiple lesions that result in false positives. The number of false positives per image is called the number of false positives. The greater the number of false positives, the greater the burden on doctors to interpret the images, resulting in increased interpretation time. Therefore, in image diagnosis support, it is important to reduce the number of false positives.
[0053] Therefore, it is conceivable to remove false-positive lesion locations included in the inference results using AI through post-processing. Fig. 6 is a diagram showing an example of image diagnosis support including post-processing. When an AI workstation 200 acquires an image 41, it infers the lesion location using, for example, a neural network. The neural network receives, for example, pixel values of the image 41 as input. Statistical information (such as the average) of the pixel values of the image 41 may also be used as input to the neural network.
[0054] As a result of the lesion location inference, multiple lesion locations 43a to 43c are detected, as shown in image 43-1. These lesion locations 43a to 43c may include false positive lesion locations 43b and 43c in addition to the true positive lesion location 43a.
[0055] Therefore, in post-processing, the AI workstation 200 determines whether each of the lesion locations 43a to 43c is a false positive or a true positive, and excludes the false positive lesion locations 43b and 43c from the inference result image 44. Only the true positive lesion location 44a is output as the inference result.
[0056] The AI workstation 200 can perform post-processing using, for example, a machine learning model. For example, the AI workstation 200 determines whether a result is a false positive or a true positive by inputting a value (such as a confidence level) indicated by the inference result of the neural network and a value of some attribute related to the image 43 into the model. As the attribute value of the image 43, for example, statistical information, pixel values, size, etc. of the image 43 can be used.
[0057] In this way, post-processing can reduce the number of false positives. Furthermore, in order to implement AI-based image diagnosis support in society, it is important that performance degradation does not occur regardless of which of the inspection devices installed in multiple facilities 30, 31, etc. is used to judge the images captured.
[0058] 7 is a diagram showing an example of image diagnosis support for each of images from a plurality of CT scanners. Image diagnosis support information is provided from the server 100 to the AI workstations 200, 200a, etc. at each of a plurality of facilities 30, 31, etc. The image diagnosis support information includes, for example, a neural network model for detecting lesion locations. The image diagnosis support information also includes, for example, a machine learning model for post-processing or parameters for determining whether an image is false positive or true positive.
[0059] The AI workstation 200 in the facility 30 performs image diagnosis support using multiple images (image groups 45 and 46) taken by the inspection devices 30a and 30b in the facility 30 as inputs to a trained model. Similarly, the AI workstation 200a in the facility 31 performs image diagnosis support using multiple images (image groups 47 and 48) taken by the inspection devices 31a and 31b in the facility 31 as inputs to a trained model.
[0060] Here, in order to enable accurate identification of lesion locations in images taken by any of the multiple inspection devices 30a, 30b, 31a, and 31b (reducing the number of false positives), it is possible to consider performing individual post-processing on each of the inspection devices 30a, 30b, 31a, and 31b, for example.
[0061] That is, the characteristics of a CT image change depending on the dose at the time of image capture and the filtering process used during image reconstruction. Dose is the amount of X-rays emitted to the patient. The higher the dose, the clearer the image, but at the same time, the greater the patient's exposure, which is a problem. Filtering is a noise removal technique used when constructing an image from X-ray map information. Examination equipment manufacturers have proposed several filtering methods, each of which results in a different image.
[0062] The radiation dose and type of filtering used during imaging in the examination device are set in the examination device as device setting information before imaging. Since the image appears differently depending on the device setting information, the optimal post-processing also differs for each examination device.
[0063] 8 is a diagram showing an example of image diagnosis support that performs post-processing for each inspection device. For example, the inspection devices 30a and 30b have different radiation doses during imaging and also perform different filter processing. The AI workstation 200 inputs images in image groups 45 and 46 acquired from the two inspection devices 30a and 30b into a common neural network and performs inference of the lesion location for each image. The AI workstation 200 performs different post-processing for each inspection device 30a and 30b and outputs inference results 45a and 46a.
[0064] In order to perform post-processing for each image according to the inspection device 30a, 30b used to capture the image, it is necessary to prepare parameters or machine learning models for post-processing for each device setting of the inspection device 30a, 30b. In other words, differences in the device settings of the inspection devices 30a, 30b result in differences in the distribution of pixel values. If the distribution of pixel values of the captured image differs, the post-processing parameters or models will also need to be changed. When the number of devices to be supported increases, it becomes practically difficult to prepare individual parameters or models for post-processing for each inspection device.
[0065] Therefore, the server 100 provides information that can support image diagnosis by taking into account the device setting information as information for supporting image diagnosis. This allows the AI workstation 200 to perform inference processing to detect lesion locations and post-processing to reduce the number of false positives, taking into account the device setting information.
[0066] 9 is a diagram showing an example of image diagnosis support using device setting information. The AI workstation 200 acquires image groups 45 and 46 and device setting information 51 and 52 from the inspection devices 30a and 30b. The device setting information 51 acquired from the inspection device 30a includes information on the X-ray dose and the type of filtering applied during CT imaging by the inspection device 30a. The device setting information 52 acquired from the inspection device 30b includes information on the X-ray dose and the type of filtering applied during CT imaging by the inspection device 30b.
[0067] For images in image group 45 acquired from inspection device 30a, AI workstation 200 uses device setting information 51 to perform neural network-based inference processing of lesion locations and post-processing to reduce false positives, and outputs inference results 45b. For images in image group 46 acquired from inspection device 30b, AI workstation 200 uses device setting information 52 to perform neural network-based inference processing of lesion locations and post-processing to reduce false positives, and outputs inference results 46b.
[0068] There are two possible ways to use device setting information: 1. Use device setting information in addition to the image as input to the neural network. 2. Use a machine learning model for post-processing, and include device setting information in the input to determine false positives.
[0069] When a machine learning model is used for post-processing, each of the AI workstations 200, 200a, ... can include, in addition to device setting information, pixel value statistical information of the image and the inference results of the neural network as input. The pixel value statistical information is, for example, a representative value (average, variance, maximum, etc.) of the pixel values of the image. The inference results of the neural network are, for example, the confidence level, position, size, etc. of the lesion location.
[0070] In the second embodiment, each of the AI workstations 200, 200a, ... uses device setting information in addition to images as input to the neural network, thereby improving the performance of image diagnosis support.
[0071] 10 is a block diagram showing the functions of each device for image diagnosis support. The server 100 includes a storage unit 110, a training image collection unit 120, a neural network training unit 130, a post-processing model training unit 140, and a trained model distribution unit 150.
[0072] The storage unit 110 stores information used for machine learning. For example, the storage unit 110 stores training image groups 111a, 111b, etc. acquired from the AI workstations 200, 200a, etc. installed in the multiple facilities 30, 31, etc. The training image groups 111a, 111b, etc. are images to which teacher data has been added by annotation. For example, the training image groups 111a, 111b, etc. are added with lesion locations identified by doctors as teacher data. The images included in the training image groups 111a, 111b, etc. include device setting information for the inspection device that captured the images. The storage unit 110 also stores noise removal effect information 112. The noise removal effect information 112 sets the noise removal effect of each filter process for each type of filter process.
[0073] The storage unit 110 further stores a neural network 113 and a post-processing model 114 as trained models. The neural network 113 is a trained neural network model generated based on training image groups 111a, 111b, .... The post-processing model 114 is a trained machine learning model for post-processing. The post-processing model 114 is, for example, a machine learning model such as a neural network, a decision tree, or a random forest.
[0074] The neural network 113 is an example of the first machine learning model 5a described in the first embodiment, and the post-processing model 114 is an example of the second machine learning model 5b described in the first embodiment.
[0075] The training image collection unit 120 acquires training image groups 111a, 111b, ... from the AI workstations 200, 200a, ... for the facilities 30, 31, .... The training image collection unit 120 stores the acquired training image groups 111a, 111b, ... in the memory unit 110.
[0076] The neural network training unit 130 trains the neural network 113 for detecting lesion locations using the training image group 111a, 111b, .... For example, the neural network training unit 130 trains the neural network 113 that receives images included in the training image group 111a, 111b, ... as input and outputs lesion locations. For example, the neural network training unit 130 adjusts the weight parameters of the neural network 113 so that the correct lesion location is output.
[0077] The neural network learning unit 130 can use the device setting information in learning the neural network 113. In doing so, the neural network learning unit 130 refers to, for example, the noise removal effect information 112, and includes the noise removal effect according to the type of filter processing in the input to the neural network 113 during learning. The neural network learning unit 130 stores the learned neural network 113 in the storage unit 110.
[0078] The post-processing model learning unit 140 generates a post-processing model 114 that reduces false-positive lesion locations through machine learning, based on the inference results when images in the training image group 111a, 111b, ... are input to the trained neural network 113. The post-processing model learning unit 140 can use, for example, device setting information when learning the post-processing model 114. In doing so, the post-processing model learning unit 140 refers to, for example, noise removal effect information 112, and includes the noise removal effect according to the filter processing type in the input to the post-processing model 114 during learning. The post-processing model learning unit 140 stores the trained post-processing model 114 in the storage unit 110.
[0079] The trained model distribution unit 150 distributes the trained model for image diagnosis support to each of the AI workstations 200, 200a, ... in each of the facilities 30, 31, .... For example, the trained model distribution unit 150 transmits a common neural network 113 and a post-processing model 114 to each of the AI workstations 200, 200a, ....
[0080] The AI workstation 200 has a memory unit 210, an image acquisition unit 220, an annotation unit 230, an image transmission unit 240, a trained model acquisition unit 250, and an inference unit 260.
[0081] The storage unit 210 stores noise removal effect information 211, an unannotated image group 212, and an annotated training image group 111a, in addition to the neural network 113 and post-processing model 114 generated by the server 100. The noise removal effect information 211 is information with the same content as the noise removal effect information 112 stored in the server 100.
[0082] The image group 212 is a plurality of images taken by the inspection device 30a or the inspection device 30b. Each image included in the image group 212 is assigned with the device setting information of the inspection device that took the image.
[0083] The training image group 111a is obtained by adding annotation results to images extracted from the image group 212. The annotation results are, for example, training data indicating the correct lesion location.
[0084] The image acquisition unit 220 acquires images from the inspection devices 30 a and 30 b. For example, the image acquisition unit 220 acquires captured images and imaging conditions from the inspection devices 30 a and 30 b, and stores the images with the imaging conditions assigned as an image group 212 in the storage unit 210.
[0085] The annotation unit 230 supports annotation of at least a portion of the image group 212. For example, the annotation unit 230 displays an image selected by a doctor on the monitor 30c. The doctor interprets the displayed image and makes a diagnosis, such as the presence or absence of a lesion. When the doctor inputs the diagnosis result, the annotation unit 230 stores the image with the diagnosis result annotated in the storage unit 210. For example, when annotating an image extracted from the image group 212, the annotation unit 230 stores the image with the diagnosis result annotated in the storage unit 210 as the training image group 111a.
[0086] The image sending unit 240 sends the training image group 111a to the server 100 in response to, for example, an image sending instruction from a user. The image sending unit 240 may also periodically read out unsent training images from the training image group 111a and send the read out training images to the server 100.
[0087] The trained model acquisition unit 250 acquires the neural network 113 and the post-processing model 114 from the server 100. The trained model acquisition unit 250 stores the acquired neural network 113 and post-processing model 114 in the storage unit 210.
[0088] The inference unit 260 executes the inference phase of AI based on the neural network 113, the post-processing model 114, and the noise removal effect information 211, and estimates the location of a lesion shown in an image in the image group 212, the type of the detected lesion, etc. The inference unit 260 displays the estimation result on a monitor, for example.
[0089] 10 shows the functions of the AI workstation 200, but the other AI workstations 200a, ... also have the same functions as the AI workstation 200. Furthermore, the functions of each element shown in FIG. 10 can be realized, for example, by having the processor 101 execute a program module corresponding to that element.
[0090] Before starting operation of the image diagnosis support system, the system administrator investigates the noise removal effects of various types of filter processing. The system administrator then creates noise removal effect information 112, 211 and stores it in the server 100 or the AI workstations 200, 200a, ....
[0091] 11 is a diagram showing an example of noise removal effect information. The noise removal effect information 112 indicates the degree of noise removal effect by the filter processing, in correspondence with the filter processing type, using a real number ranging from 0 to 1. The higher the number, the higher the noise removal effect (the more noise can be removed). The same noise removal effect as the noise removal effect information 112 is set in the noise removal effect information 211 stored in the AI workstations 200, 200a, ...
[0092] By quantifying the noise removal effect of the filter processing using the noise removal effect information 112, 211, it becomes easier to reflect the influence of the filter processing in image diagnosis support using AI.
[0093] Next, processing when device setting information is used as input to the neural network will be described. FIG. 12 is a diagram showing an example of image diagnosis support processing using device setting information as input to the neural network. For example, when the AI workstation 200 performs image diagnosis support for an image included in the image group 45 acquired from the inspection device 30a, the AI workstation 200 inputs the image and device setting information 51 acquired from the inspection device 30a to the neural network 113. When the AI workstation 200 performs image diagnosis support for an image included in the image group 46 acquired from the inspection device 30b, the AI workstation 200 inputs the image and device setting information 52 acquired from the inspection device 30b to the neural network 113.
[0094] AI workstation 200 performs inference using neural network 113 for each image to detect suspected lesion locations. AI workstation 200 then generates lesion location information 61, 62, ... for each detected lesion location. If multiple lesion locations are detected in one image, multiple pieces of lesion location information corresponding to that image are output.
[0095] The lesion location information 61, 62, ... includes pixel value statistical information 61a, 62a and AI output information 61b, 62b, respectively. The pixel value statistical information 61a, 62a is, for example, information on representative values (average, variance, etc.) of pixel values of pixels in an image. The AI output information 61b, 62b includes, for example, information on the degree of certainty that a lesion exists in the detected lesion location, the location of the lesion location, etc.
[0096] The AI workstation 200 performs post-processing to reduce false positives for each piece of lesion location information 61, 62, .... That is, the AI workstation 200 determines whether each piece of lesion location information 61, 62, ... is a false positive, and if it determines that it is a false positive, it excludes the corresponding lesion location from the inference result.
[0097] The AI workstation 200 outputs inference results 45d and 45e after removing false-positive lesion locations through post-processing. To realize such image diagnosis support, the server 100 generates a trained neural network 113 that includes the device setting information 51 and 52 as input. The AI workstations 200, 200a, ... then use the neural network 113 generated by the server 100 to perform inference using the device setting information 51 and 52.
[0098] Fig. 13 is a sequence diagram showing an example of the procedure for image diagnosis support processing in which device setting information is included in the input to the neural network. Fig. 13 shows the processing between the server 100 and the AI workstation 200, but similar processing is also performed between the server 100 and other AI workstations 200a, etc.
[0099] The training image collection unit 120 of the server 100 requests training images including device setting information from the AI workstation 200 (step S101). The image transmission unit 240 of the AI workstation 200 transmits a training image group 111a, which is a collection of annotated images, to the server 100 (step S102). Each image included in the training image group 111a includes device setting information. Upon receiving the training image group 111a, the training image collection unit 120 of the server 100 stores the training image group 111a in the storage unit 110.
[0100] The neural network training unit 130 of the server 100 trains the neural network 113 for detecting lesion locations using the training image groups 111a, 111b, ... acquired from the AI workstations 200, 200a, ... (step S103).
[0101] During learning, the neural network learning unit 130 includes, as input to the neural network 113, the dose value indicated in the device setting information included in each image. The neural network learning unit 130 also includes, as input to the neural network 113, the noise removal effect value corresponding to the filter processing type indicated in the device setting information. The neural network learning unit 130 stores the trained neural network 113 in the storage unit 110.
[0102] Next, the neural network training unit 130 uses the trained neural network 113 to create sample data for false positives and true positives (step S104). For example, for each image in the training image group 111a, 111b, ..., the neural network training unit 130 inputs device setting information corresponding to that image into the neural network 113 to obtain lesion location information indicating an estimated lesion. If the lesion indicated in the lesion location information of a certain image matches the lesion indicated as training data for that image, the neural network training unit 130 designates the lesion location information as a true positive sample. On the other hand, if the lesion indicated in the lesion location information of a certain image does not match the lesion indicated as training data for that image, the neural network training unit 130 designates the lesion location information as a false positive sample.
[0103] The post-processing model learning unit 140 acquires the created sample data from the neural network learning unit 130. Then, the post-processing model learning unit 140 uses the sample data to learn the post-processing model 114 (step S105). The post-processing model learning unit 140 stores the learned post-processing model 114 in the storage unit 110.
[0104] When the training of the neural network 113 and the post-processing model 114 is completed, the trained model distribution unit 150 transmits these trained models (neural network 113 and post-processing model 114) to the AI workstation 200 (step S106).
[0105] The trained model acquisition unit 250 of the AI workstation 200 acquires the trained model and stores it in the storage unit 210 (step S107). Then, the image acquisition unit 220 acquires an image of the patient, including the device setting information, from either the examination device 30a or 30b (step S108). The image acquisition unit 220 transmits the acquired image to the inference unit 260.
[0106] The inference unit 260 first uses the neural network 113, which inputs the device setting information, to infer lesion locations in the acquired image (step S109). If the inference reveals a suspected lesion in the image, the location is detected as a lesion. For each detected lesion, the inference unit 260 determines whether it is a false positive using the post-processing model 114, and removes false positive lesion locations from the inference results (step S110).
[0107] The inference unit 260 then displays the inference result of the image after eliminating false positives on the screen (step S111). When a lesion location is indicated in the inference result displayed on the screen, the doctor interprets the image, keeping in mind that there is a high possibility that a lesion is present at that lesion location. At this time, the doctor's interpretation becomes more efficient because the number of false positive lesion locations has been reduced.
[0108] [Third Embodiment] The third embodiment is a system that improves the performance of image diagnosis support by including device setting information in the input to post-processing using machine learning and determining false positives in the post-processing.
[0109] 14 is a diagram showing an example of image diagnosis support processing that performs post-processing using device setting information. For example, when the AI workstation 200 performs image diagnosis support for an image included in the image group 45 acquired from the inspection device 30a, the AI workstation 200 inputs the image to the neural network 113. When the AI workstation 200 performs image diagnosis support for an image included in the image group 46 acquired from the inspection device 30b, the AI workstation 200 inputs the image to the neural network 113.
[0110] AI workstation 200 performs inference using neural network 113 for each image to detect suspected lesion locations. AI workstation 200 then generates lesion location information 61, 62, ... for each detected lesion location. If multiple lesion locations are detected in one image, multiple pieces of lesion location information corresponding to that image are output.
[0111] AI workstation 200 performs post-processing to reduce false positives for each piece of lesion location information 61, 62, ..., using the device setting information of the inspection device used to capture the image. That is, AI workstation 200 determines whether each piece of lesion location information 61, 62, ... is a false positive, and if it determines a false positive, excludes the corresponding lesion location from the inference result.
[0112] The AI workstation 200 outputs inference results 45f, 45g after removing false-positive lesion locations through post-processing. To realize such image diagnosis support, the server 100 generates a trained post-processing model 114 that includes the device setting information 51, 52 as input. The AI workstations 200, 200a, ... then use the post-processing model 114 generated by the server 100 to reduce the number of false positives using the device setting information 51, 52.
[0113] In reality, it is difficult to collect enough images from all inspection devices as training images. Furthermore, if the post-processing model is a neural network, there is a possibility that the post-processing model 114 may overfit to a specific inspection device if there is a bias in the number of training images that can be acquired from each inspection device. Therefore, to prevent overfitting, the server 100 groups the device setting information according to the dose and the filtering type, and performs sampling so that the number of sample data corresponding to the device setting information in each group is equal.
[0114] For example, the server 100 divides the range of dose values that can be set into multiple divided ranges, and divides the range of noise removal effect values (0 to 1) for each filter processing type into multiple divided ranges. The server 100 generates a group for each combination of the divided range of dose and the divided range of noise removal effect. The server 100 then assigns each piece of device setting information to a group corresponding to a combination of the divided range of dose to which the dose in that device setting information belongs and the divided range of noise removal effect to which the noise removal effect of the filter processing type in that device setting information belongs. The server 100 may also cluster multiple pieces of device setting information by the dose and the noise removal effect to which the noise removal effect belongs, and assign device setting information that belongs to the same cluster to the same group.
[0115] Fig. 15 is a sequence diagram showing an example of the procedure for image diagnosis support processing in which device setting information is included in the input to the neural network. Fig. 15 shows the processing between the server 100 and the AI workstation 200, but similar processing is also performed between the server 100 and other AI workstations 200a, ...
[0116] The training image collection unit 120 of the server 100 requests training images including device setting information from the AI workstation 200 (step S201). The image transmission unit 240 of the AI workstation 200 transmits a training image group 111a, which is a collection of annotated images, to the server 100 (step S202). Each image included in the training image group 111a includes device setting information. Upon receiving the training image group 111a, the training image collection unit 120 of the server 100 stores the training image group 111a in the storage unit 110.
[0117] The neural network training unit 130 of the server 100 trains the neural network 113 for lesion detection using the training image groups 111a, 111b, ... acquired from the AI workstations 200, 200a, ... (step S203). The neural network training unit 130 stores the trained neural network 113 in the storage unit 110.
[0118] Next, the neural network training unit 130 generates sample data for false positives and true positives using the trained neural network 113 (step S204). The neural network training unit 130 then performs training data sampling processing (step S205). Details of the training data sampling processing will be described later (see FIG. 16).
[0119] The post-processing model training unit 140 acquires the sample data sampled in the training data sampling process from the neural network training unit 130. Then, the post-processing model training unit 140 trains the post-processing model 114 using the sample data (step S206). The post-processing model training unit 140 stores the trained post-processing model 114 in the storage unit 110.
[0120] When the training of the neural network 113 and the post-processing model 114 is completed, the trained model distribution unit 150 transmits these trained models (neural network 113 and post-processing model 114) to the AI workstation 200 (step S207).
[0121] The trained model acquisition unit 250 of the AI workstation 200 acquires the trained model and stores it in the storage unit 210 (step S208). Then, the image acquisition unit 220 acquires an image of the patient, including the device setting information, from either the examination device 30a or 30b (step S209). The image acquisition unit 220 transmits the acquired image to the inference unit 260.
[0122] The inference unit 260 first uses the neural network 113 to infer lesion locations in the acquired image (step S210). If the inference reveals a suspected lesion in the image, the location is detected as a lesion. For each detected lesion, the inference unit 260 determines whether it is a false positive using the post-processing model 114, which includes device setting information as input, and removes false positive lesion locations from the inference results (step S211).
[0123] The inference unit 260 then displays the inference result of the image after eliminating false positives on the screen (step S212). When a lesion location is indicated in the inference result displayed on the screen, the doctor interprets the image, keeping in mind that there is a high possibility that a lesion is present at that lesion location. At this time, the doctor's interpretation becomes more efficient because the number of false positive lesion locations has been reduced.
[0124] Furthermore, by using device setting information in the post-processing model rather than in the neural network for lesion location detection, the influence of the device setting information can be appropriately reflected in the trained model. That is, medical images include 3D medical images. The data size of 3D medical images is much larger than that of device setting information. Therefore, even if device setting information is used as input to a neural network along with the image, the influence of the image may be dominant in the learning of the neural network, and differences in the device setting information may hardly be reflected in the learning. In contrast, the lesion location information used as input to the post-processing model only includes pixel value statistical information and AI output information, and has a small data size. Therefore, by using device setting information as input to the post-processing model, a post-processing model that fully reflects the influence of the device setting information can be generated.
[0125] Furthermore, if the post-processing model is a machine learning model other than a neural network, there is little risk of the post-processing model overfitting to the device setting information of a specific inspection device. However, if the post-processing model is a neural network and there is a bias in the number of training images for each device setting information group, there is a possibility that the post-processing model will overfit to the device setting information of a specific inspection device. Therefore, training data sampling is used to equalize the number of training images for each device setting information.
[0126] FIG. 16 is a flowchart showing an example of a procedure for sampling training data. The process shown in FIG. 16 will be described below in order of step number. [Step S221] The neural network training unit 130 determines whether the post-processing model is a neural network. If it is a neural network, the neural network training unit 130 proceeds to step S222. If it is not a neural network, the neural network training unit 130 ends the training data sampling process.
[0127] [Step S222] The neural network training unit 130 counts the number of training samples for each group with similar doses and filter processing. [Step S223] The neural network training unit 130 equalizes the number of training samples among groups with similar doses and filter processing. For example, the neural network training unit 130 identifies the group with the fewest number of training samples. The neural network training unit 130 samples the same number of training samples for each group as the number of training samples for the identified group.
[0128] In this way, when the post-processing model is a neural network, the number of training samples for each group according to the dose and filtering process is equalized. By equalizing the number of training samples, even if only one training image for a specific device setting information is significantly more than others, the post-processing model does not overfit to the dose and filtering process indicated by that device setting information.
[0129] Other Embodiments In the second and third embodiments, the learning phase of machine learning is executed by the server 100, and the inference phase is executed by the AI workstations 200, 200a, .... However, the learning phase and the inference phase may be executed by a single device. For example, the server 100 may execute inference of the lesion location using the neural network 113 and the post-processing model 114. In this case, images of patients are transmitted to the server 100 from the AI workstations 200, 200a, ... of each facility 30, 31, ... such as a medical institution. The server 100 infers the lesion location and transmits the inference results to the AI workstation that transmitted the images.
[0130] The foregoing merely illustrates the principles of the present invention. Further, since numerous modifications and changes will be apparent to those skilled in the art, the present invention is not limited to the exact construction and application shown and described above, and all corresponding modifications and equivalents are deemed to be within the scope of the present invention as defined by the appended claims and their equivalents.
[0131] REFERENCE SIGNS 1, 1a, 1b, ... Imaging device 2 Training data set 2a, 2b, ... Training data 3 Input feature amount 3a, 6a Captured image 3b, 6b Setting information 4 Correct answer 5 Machine learning model 5a First machine learning model 5b Second machine learning model 6 Inference target feature amount 7 Inference result 7a Lesion location 10 Information processing device 11 Storage unit 12 Processing unit
Claims
1. Obtain a plurality of first training data, which are training data associating an input feature amount including pixel information of a captured image and setting information at the time of capturing the captured image by an imaging device that captured the captured image with the correct answer of a determination process for the captured image. Train the machine learning model based on the result of the determination process for the captured image to be determined when the input feature amount shown in the plurality of first training data is input to the machine learning model and the correct answer shown in the plurality of first training data. A machine learning program that causes a computer to execute the process.
2. In the process of training the machine learning model, using the pixel information of the captured image to be determined and the setting information of the imaging device that captured the captured image to be determined as inputs, train a first machine learning model to determine a location in the captured image where a predetermined object appears, and using the location determined to have the object appearing as an input, train a second machine learning model to determine the truth or falsehood of the location determined to have the object appearing. The machine learning program according to claim 1.
3. In the process of training the machine learning model, using the pixel information of the captured image to be determined as an input, train a first machine learning model to determine a location in the captured image to be determined where a predetermined object appears, and using the location determined to have the object appearing and the setting information of the imaging device that captured the captured image to be determined as inputs, train a second machine learning model to determine the truth or falsehood of the location determined to have the object appearing. The machine learning program according to claim 1.
4. Classify each of the plurality of the training data included in the training data set into a plurality of groups according to the similarity of the setting information, and evenly obtain the plurality of first training data from each of the groups. The machine learning program according to claim 1, which further causes a computer to execute the process.
5. Input an inference target feature amount including pixel information of a captured image to be inferred and setting information at the time of capturing the captured image to be inferred by an imaging device into the trained machine learning model, and perform the determination process for the captured image to be inferred according to the machine learning model. The machine learning program according to claim 1, which further causes the computer to execute the process.
6. In the process of obtaining the plurality of first training data, the plurality of first training data are obtained from the training data set including the training data in which the input feature amount including the pixel information of the captured image of the human body and the setting information at the time of imaging of the imaging device that captured the human body is associated with the lesion location shown in the captured image, and the machine learning model is trained based on the determination result of the lesion location for the captured image to be determined when the input feature amount shown in the plurality of first training data is input to the machine learning model and the lesion location shown in the plurality of first training data. The machine learning program according to claim 1.
7. The setting information includes the X-ray dose of the imaging device that captured the human body using X-rays. The machine learning program according to claim 1.
8. The setting information includes the type of filter processing performed by the imaging device on the captured image. The machine learning program according to claim 1.
9. A machine learning method in which a computer executes a process of obtaining a plurality of first training data, each of which is training data in which an input feature amount including pixel information of a captured image and setting information at the time of imaging of the imaging device that captured the captured image is associated with the correct answer of the determination process for the captured image, and training the machine learning model based on the result of the determination process for the captured image to be determined when the input feature amount shown in the plurality of first training data is input to the machine learning model and the correct answer shown in the plurality of first training data.
10. An information processing apparatus having a processing unit that obtains a plurality of first training data, each of which is training data in which an input feature amount including pixel information of a captured image and setting information at the time of imaging of the imaging device that captured the captured image is associated with the correct answer of the determination process for the captured image, and trains the machine learning model based on the result of the determination process for the captured image to be determined when the input feature amount shown in the plurality of first training data is input to the machine learning model and the correct answer shown in the plurality of first training data.
Citation Information
Patent Citations
Medical information processing apparatus, medical information processing system, and medical information processing program
JP2020010823A
Medical imaging apparatus, medical image processing apparatus and image processing program
JP2021133142A
Method and system for identifying skin texture and skin lesion using artificial intelligence cloud-based platform
US20200372639A1
Machine learning system and method, integration server, information processing device, program, and inference model creation method
WO2021079792A1
Information processing device, information processing method, program, and recording medium
WO2023112497A1