Formula-driven supervised learning method and system for ai model, and method and system for creating transfer-learned ai model
The mathematical-driven supervised learning method generates synthetic image data sets using multiple formulas to pre-train AI models, addressing the limitations of existing AI models in medical imaging by achieving higher accuracy and practicality with limited data.
Patent Information
- Application Number
- PCT/JP2024/038350
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-08
AI Technical Summary
Existing AI models based on convolutional neural networks (CNNs) for image classification, particularly in medical imaging, face limitations due to the need for large amounts of labeled teacher data, which can be costly, time-consuming, and subject to data bias and copyright issues.
A mathematical-driven supervised learning method and system that generates synthetic image data sets using multiple mathematical formulas, creating a mixed database for pre-training AI models, which can then be applied to target tasks with higher accuracy than models pre-trained using single formula data sets.
This approach enables AI models to achieve higher classification accuracy and practical applicability in medical imaging tasks, even with limited actual image data, while avoiding issues related to data bias, copyright, and labeling errors.
Smart Images

Figure JP2024038350_08052025_PF_FP_ABST
Abstract
Description
Method and system for formula-driven supervised learning of AI models, and method and system for creating transfer-trained AI models
[0001] The present invention relates to a method and system for formula-driven supervised learning of an AI model based on automatically generated images, and a method and system for creating a transfer-trained AI model.
[0002] Artificial intelligence (AI) models based on convolutional neural networks (CNNs) used in image classification generally require a massive amount of training data (images and their correct labels) for training. For example, AI-based medical image diagnostic support technologies have been developed and are being used in clinical practice for diagnostic imaging, particularly in departments with a large number of radiological, MRI, and gastrointestinal endoscopy images, which require a relatively large number of doctors. However, not all medical image diagnostic tasks can collect sufficient training data for CNN training. As a result, there are limitations to AI research and development for image diagnosis, where sufficient training data cannot be collected. Therefore, the introduction of methods to train diagnostic AI models with limited training data has been anticipated. In such cases, techniques such as transfer learning and data augmentation have been used as a solution in research and development. Transfer learning often involves training using large, general-purpose datasets, such as real images, in the pre-training stage before learning the dataset for the target task. In this process, when using large-scale real-image datasets such as ImageNet for AI training, problems arise, such as copyright and data bias, labeling errors and human costs due to manual annotation, and the inability to access the images themselves depending on the database. In particular, in the medical field, transparency of the data used for AI training is extremely important.
[0003] To address these issues, formula-driven supervised learning (FDSL) has been proposed, as disclosed in Non-Patent Document 1. This involves creating a dataset of fractal and contour images from mathematical formulas and using them for pre-training a CNN model for image recognition AI. Formula-driven supervised learning is a technique in which synthetic images are generated using mathematical formulas as a pre-training dataset, automatically labeled with the parameters used for generation, and then pre-training is performed using the automatically generated images and labels. Figure 9 shows a diagram including the formulas and parameters described in Non-Patent Document 1. The developers of this formula-driven supervised learning technique discovered that the shapes and patterns of the lines defined by the formulas are important, and that training a CNN model with generated fractal images allows the model to acquire feature representations similar to those achieved when training with real images. Formula-driven supervised learning allows training using only synthetic images generated from mathematical formulas, without using real images, thereby eliminating issues that arise when training using real images, such as copyright, privacy, and labeling errors during commercial use.
[0004] In the technology described in Non-Patent Document 1, a data set called FractalDB is constructed using fractal geometry.
[0005] Furthermore, technologies that are an extension of the formula-driven supervised learning disclosed in Non-Patent Document 1 are disclosed in Non-Patent Documents 2 and 3. The technology disclosed in Non-Patent Document 2 uses a dataset (RadialContourDB) with refined contour shapes and ExFractalDB (Extended FractalDB), an improved version of FractalDB, for pre-training. The technology disclosed in Non-Patent Document 3 uses a dataset (VisualAtomDB) with further refined contour shapes for pre-training.
[0006] Furthermore, Non-Patent Document 4 shows the results of verifying the effectiveness of ExFractalDB in mathematics-driven supervised learning for medical image tasks.
[0007] H. Kataoka, K. Okayasu, A. Matsumoto, E. Yamagata, R. Yamada, N. Inoue, A. Nakamura and Y. Satoh Pre-training without Natural Images, Proc. IEEE / CVF Asian Conference on Computer Vision (ACCV), pp.583-600, December 2020H. Kataoka, R. Hayamizu, R. Yamada, K. Nakashima, S. Takashima, X. Zhang, EJ Martinez-Noriega, N. Inoue and R. Yokota, Replacing Labeled Real-image Datasets with Auto-generated Contours, Proc. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.21200-21209, 2022. Takashima, R. Hayamizu, N. Inoue, H. Kataoka, and R. Yokota, Visual atoms: Pre-training vision transformers with sinusoidal waves, Proc. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.18579-18588, 2023. Takato Endo, Hideya Takahashi, and Eisaku Maeda presented a paper titled "On the Effectiveness of Mathematically Driven Supervised Learning in Medical Imaging Tasks" [IEICE Technical Report PRMU2022-71, IBISML2022-7(2023-03)].
[0008] In formula-driven supervised learning proposed in Non-Patent Document 1, actual images captured in the real world are not used. Instead, a large number of images are generated from randomly determined formulas, and an AI model that can be transferred to the target task is pre-trained (supervised learning) using these images. To create fractals, contour images, etc., hyperparameters are required when creating the data. One combination is randomly selected from the range of hyperparameters, and this is used as one class (hyperparameter combination). Approximately 1,000 images are created from this class, preparing approximately 1,000 classes. The task of classifying these 1,000 images is then performed using supervised learning. In formula-driven supervised learning, the greater the complexity of the mathematically generated images, the higher the accuracy. The complexity of the mathematically generated images can be increased by adjusting the image generation parameters for formula-driven supervised learning. In particular, Non-Patent Document 2 confirms that training with a dataset (RCDB) specialized for drawing object contours can achieve performance equivalent to FractalDB.
[0009] However, as shown in Non-Patent Document 1, Non-Patent Document 2, or Non-Patent Document 3, there is a limit to improving the performance of an AI model by pre-training an AI model that can be transferred to a target task using images generated by formula-driven supervised learning using one type of formula. In particular, when the target task becomes complex, although the performance is higher than other methods, the classification accuracy itself is low, making it difficult to use in a practical manner.
[0010] An object of the present invention is to provide a method and system for formula-driven supervised learning of an AI model that pre-trains an AI model that can be transferred to a target task with more practical accuracy than conventional methods.
[0011] Another object of the present invention is to provide a method and system for creating a transfer-trained AI model, which performs transfer learning using an AI model pre-trained by a formula-driven supervised learning method and real teacher images of a target task to create a transfer-trained AI model.
[0012] The formula-driven supervised learning method for an AI model of the present invention uses formula-driven supervised learning to create two or more image datasets consisting of two or more types of images each automatically generated using two or more formulas, mix these two or more image datasets to create a mixed database, and pre-train an AI model using the mixed database. This AI model is a general-purpose AI model that serves as a basis for creating an AI model for a target task. According to the present invention, it is possible to provide a pre-trained AI model that can be applied to a target task with higher accuracy than an AI model pre-trained using a single type of image dataset consisting of multiple images automatically generated using a single formula as described in Non-Patent Documents 1 to 3.
[0013] More specifically, in one aspect of the present invention, a first type of image dataset is created using a first mathematical formula, the first type consisting of a plurality of automatically generated images, and a second type of image dataset is created using a second mathematical formula. The first and second mathematical formulas may be the same or different. When the same mathematical formula is used, the formulas may be substantially different by changing the parameters used. A mixed database is then created by mixing the first type of image dataset and the second type of image dataset, and an AI model is pre-trained using this mixed database. Here, the mixed database refers to a database in which a plurality of automatically generated images from the first type of image dataset and a plurality of automatically generated images from the second type of image dataset are mixed randomly or according to a predetermined rule. Note that the predetermined rule and the mixing ratio of the plurality of automatically generated images from the first type of image dataset and the plurality of automatically generated images from the second type of image dataset used to create the mixed database may be determined through prior testing so as to enable application to a target task with higher accuracy. By using the mixture ratio determined based on this prior testing, it is possible to provide an AI model that can be applied to a target task with higher accuracy through pre-training than an AI model created using the techniques described in Non-Patent Documents 1 to 3, which is pre-trained using an image dataset created using a single mathematical formula. In other words, the pre-trained AI model created according to the present invention can be applied to a target task with higher accuracy than an AI model that is pre-trained using one type of image dataset consisting of multiple images automatically generated using one mathematical formula described in Non-Patent Documents 1 to 3.
[0014] Preferably, the first formula is a formula for creating an image dataset of shape and pattern images, and the second formula is a formula for creating an image dataset of contour images. Formulas used to automatically generate multiple images of a target task for formula-driven supervised learning have already been disclosed for standalone use in Non-Patent Documents 1 to 3, and therefore can certainly be put to practical use. However, these documents do not disclose or suggest creating a mixed database.
[0015] An example of an AI model that can use a pre-trained AI model is an endoscopic diagnosis support AI model. The endoscopic images targeted by the endoscopic diagnosis support AI model have all of the above-mentioned problems that arise when collecting real images. Therefore, the present invention is suitable for application. Experiments have confirmed that the present invention is suitable for application to a cystoscopic diagnosis support AI model, among endoscopic diagnosis support AI models. The mixing ratio of the first type of image dataset and the second type of image dataset is determined according to the target task. Incidentally, when creating a cystoscopic diagnosis support AI model, the mixing ratio of the first type of image dataset and the second type of image dataset during pre-training is preferably 1:3 to 3:1.
[0016] In the formula-driven supervised learning method for an AI model of the present invention, in order to use a pre-trained AI model created using a mixed database as an AI model to actually use for a target task, it is preferable to use the pre-trained AI model to perform transfer learning (two-stage learning) for each target task to create a transfer-trained AI model.When building an endoscopic diagnosis support AI model that classifies whether or not an image contains a lesion, it is preferable to change the fully connected layer to a binary classification based on whether or not the image contains an object (lesion), and perform transfer learning on all layers of the pre-trained AI model.Furthermore, when building an endoscopic diagnosis support AI model that detects the location of a lesion in an image, it is also possible to change the fully connected layer to a pixel-by-pixel classification that determines whether or not each pixel in the image contains an object (lesion), and perform transfer learning on all layers of the pre-trained AI model.
[0017] The present invention can be understood as a formula-driven supervised learning system for an AI model executed by a computer. A more specific formula-driven supervised learning system for an AI model of the present invention includes a first image dataset creation unit that creates a first type of image dataset consisting of a plurality of images automatically generated using a first formula, a second image dataset creation unit that creates a second type of image dataset consisting of a plurality of images automatically generated using a second formula, a mixed database creation unit that creates a mixed database by mixing the first type of image dataset and the second type of image dataset, and an AI model training unit that pre-trains an AI model using the mixed database. Of course, the mixed database may also be created by mixing two or more image datasets consisting of two or more types of images each automatically generated using two or more formulas.
[0018] The present invention can also be understood as a system for creating a transfer-trained AI model that includes a transfer learning unit that performs transfer learning using an AI model that has been pre-trained in a formula-driven supervised learning system for an AI model to create a transfer-trained AI model.
[0019] 1 is a diagram used to explain an overview of an embodiment of a method and system for formula-driven supervised learning of an AI model based on automatically generated images and a method and system for creating a transfer-trained AI model according to the present invention.
[0023] FIG. 1 is a block diagram showing the schematic configuration of a system for creating a transfer-trained AI model equipped with a formula-driven supervised learning system for an AI model that pre-trains an AI model for a target task using a plurality of automatically generated images of the target task in order to realize the concept of FIG. 1.
[0024] FIG. 1 is a diagram showing a sample of a bladder endoscopy image.
[0025] FIG. 1 is a diagram showing the classification performance of an AI model for each pre-training training database.
[0026] FIG. 1 is a diagram showing a comparison of the classification performance of an AI model with the classification performance of a doctor.
[0027] FIG. 1 is a diagram showing the accuracy rate for the proportion of bladder endoscopy images used for the target task in the classification performance of an AI model for each pre-training training database.
[0028] FIG. 1 is a diagram showing the Youden index for the proportion of bladder endoscopy images used for the target task in the classification performance of an AI model for each pre-training training database.
[0029] FIG. 2 is a diagram showing the results of pixel-by-pixel lesion classification performed by an AI model pre-trained with automatically generated images overlaid on an input bladder endoscopy image.
[0029] FIG. 2 is a diagram including formulas and parameters shown in Non-Patent Document 1.
[0020] An embodiment of the present invention will be described in detail below with reference to the drawings. FIG. 1 is a diagram used to explain an overview of an embodiment of a method and system for creating a transfer-learned AI model of the present invention, which creates a transfer-learned AI model using an AI model pre-trained by a formula-driven supervised learning method and system for an AI model. FIG. 2 is a block diagram showing the schematic configuration of a system 2 for creating a transfer-learned AI model, which includes a formula-driven supervised learning system 1 for an AI model that pre-trains an AI model using multiple types of automatically generated images, in order to realize the concept of FIG. 1. Note that the means constituting each block of the system shown in FIG. 2 are implemented within a computer by installing software on the computer. Of course, the system shown in FIG. 2 can also be configured by configuring the first image data creation unit 11, the second image data creation unit 12, the mixed database creation unit 13, the AI model learning unit 14, and the transfer learning unit 15 of the system shown in FIG. 2 using hardware such as a storage device, an arithmetic unit, and a control device, respectively.
[0021] In the embodiment of Figure 2, examples of AI models that have undergone transfer learning and are created by transfer learning a pre-trained AI model include an endoscopic diagnosis support AI model and a cystoendoscopic diagnosis support AI model. The endoscopic images targeted by the endoscopic diagnosis support AI model have all of the aforementioned problems that arise when collecting real images. Therefore, they are suitable for application of the present invention. In particular, the suitability of the present invention for application to a cystoendoscopic diagnosis support AI model, among endoscopic diagnosis support AI models, is supported by experiments described below.
[0022] In this embodiment, the formula-driven supervised learning system 1 for an AI model includes a first image dataset creation unit 11, a second image dataset creation unit 12, a mixed database creation unit 13 incorporating a mixed database 13A, and an AI model learning unit 14. The transfer-trained AI model creation system 2 includes a transfer learning unit 15 that performs transfer learning using an AI model pre-trained in the formula-driven supervised learning system 1 for an AI model to create a transfer-trained AI model. The first image dataset creation unit 11 creates a first type of image dataset DS1 consisting of a plurality of images automatically generated using a first formula. The second image dataset creation unit 12 creates a second type of image dataset DS2 consisting of a plurality of images automatically generated using a second formula. For example, the first formula is a formula for creating an image dataset DS1 of shape and pattern images, and the second formula is a formula for creating an image dataset DS2 of contour images. Formulas that can be used to automatically generate a plurality of images in formula-driven supervised learning have already been disclosed for use alone in Non-Patent Documents 1 to 3. In this embodiment, the known formula shown in Non-Patent Document 1 and shown in FIG. 8 is used. The first formula and the second formula may be the same formula or different formulas. When the same formula is used, it is sufficient to change the parameters appropriately depending on the characteristics of the target task.
[0023] The mixed database creation unit 13 creates a mixed database 13A by mixing the first type of image dataset DS1 and the second type of image dataset DS2. The AI model training unit 14 then pre-trains an AI model using the mixed, automatically generated images in the mixed database 13A. This pre-trained AI model is a general-purpose AI model that can serve as a basis for creating AI models for many target tasks.
[0024] Here, the mixed database 13A is a database in which a plurality of automatically generated images from a first type of image dataset DS1 and a plurality of automatically generated images from a second type of image dataset DS2 are mixed randomly or according to a predetermined rule. The predetermined rule and the mixing ratio of the plurality of automatically generated images from the first type of image dataset DS1 and the plurality of automatically generated images from the second type of image dataset DS2 used to create the mixed database 13A are determined through pre-testing depending on the target task. Incidentally, in the experiment described below, the mixing ratio of the first type of image dataset DS1 and the second type of image dataset DS2 is 1:1. Ideally, the pre-test should determine a mixing ratio that can obtain an AI model that can demonstrate higher classification performance in all evaluation indices than when pre-training is performed using an image dataset created using a single mathematical formula, as in the conventional techniques described in Non-Patent Documents 1 to 3.
[0025] Rather than using an AI model that has been pre-trained using one type of image dataset consisting of multiple images automatically generated by formula-driven supervised learning using one formula described in Non-Patent Document 1, a pre-trained AI model that has been pre-trained according to this embodiment and then further transferred to perform transfer learning such as the two-stage learning described below to suit the target task will generally show high classification performance in all evaluation indicators.
[0026] The transfer learning unit 15 of the transfer-trained AI model creation system 2 of this embodiment performs pre-training of an AI model using the AI model's formula-driven supervised learning system 1, and then performs transfer learning using this pre-trained AI model. When the target task for transfer learning the pre-trained AI model is endoscopic diagnosis, which classifies whether or not an image contains a lesion, the transfer learning unit 15 converts the fully connected layer to binary classification based on whether or not the image contains a lesion, and transfers all layers of the pre-trained AI model to construct an endoscopic diagnosis support AI model as a transfer-trained AI model. Furthermore, when constructing an endoscopic diagnosis support AI model that detects the location of a lesion in an image, the fully connected layer may be converted to pixel-by-pixel classification that determines whether each pixel in the image contains a target (lesion), and all layers of the pre-trained AI model may be transfer-trained. The performance of this transfer-trained AI model has been confirmed by experiments described below.
[0027] (Experimental Example of Cystoendoscopic Diagnosis Support AI Model) When creating a transfer-learned cystoendoscopic diagnosis support AI model using this embodiment, an AI model that has been pre-trained using formula-driven supervised learning system 1 is subjected to transfer learning using real images, i.e., cystoscopic images. That is, in the first stage of learning (pre-learning), pre-learning is performed using formula-driven supervised learning system 1 for the AI model. In the second stage of learning, transfer learning unit 15 is used to change the final layer (fully connected layer) of the pre-trained AI model (ResNet50) to a binary classification based on whether or not the image contains a lesion, and all layers are transferred to a cystoendoscopic diagnosis support AI model that diagnoses cystoscopic images.
[0028] Specifically, we assumed that both contour and texture features of the target cystoendoscopic images are strongly related to distinguishing between diseased and normal mucosa. As image datasets to be created using the formulas in the formula-driven supervised learning method, a first type of image dataset DS1 was created using a first formula, and second types of image datasets DS1 and DS2 were created using a second formula. The two automatically created image datasets were FractalDB (first image dataset DS1) and VisualAtomDB (second image dataset DS2), and these were mixed to create a mixed database. In the following description, the mixed database created by this process is referred to as the Mix database (MixDB).
[0029] (Summary of Experiments) In order to verify the effectiveness of the method of the present invention, a basic AI model was created that was pre-trained using the formula-driven supervised learning system 1 for the AI model, and the pre-trained AI model was subjected to transfer learning using cystoscope images, and the classification performance of the AI model was evaluated through the following three experiments.
[0030] In Experiment 1, we compared the classification performance of an AI model that had been pre-trained using automatically generated images with that of an AI model that had been trained using only actual cystoscopy images.
[0031] In Experiment 2, we used an AI model that had been pre-trained using automatically generated images and compared the classification performance of a urologist as shown in the paper in Non-Patent Document 5 below.
[0032] (Non-Patent Document 5) A. Ikeda, H. Nosato, Y. Kochi, T. Kojima, K. Kawai, H. Sakanashi, M. Murakawa and H. Nishiyama, Support System of Cystoscopic Diagnosis for Bladder Cancer Based on Artificial Intelligence, Journal of Endourology, vol. 34, issue. 3, pp. 352-358, March 2020. In Experiment 3, an AI model was created that had been pre-trained using automatically generated images, and then the classification performance of the AI model that had undergone transfer learning using a small number (25%) of cystoscopic images compared to Experiment 1 was compared with the classification performance of the AI model that had undergone transfer learning using only a small number (25%) of cystoscopic images, with the results of Experiment 1 in which transfer learning was performed using 100% of cystoscopic images.
[0033] In this experiment, 4 / 5 of the training data of bladder endoscopy images, the target task, was used for training, and the remaining 1 / 5 was used for validation, using a 5-fold cross-validation method, and the average value of the five models was used to evaluate the classification performance.
[0034] The cystoscopy images used in this experiment were taken during bladder cancer examinations at the University of Tsukuba Hospital. A sample cystoscopy image is shown in Figure 3. Cystoscopy images were labeled as "abnormal" if they contained a lesion, and "normal" if they did not. A urologist determined whether the image contained a lesion. Of the image data, 8,812 images were used as training data (1,259 abnormal, 7,553 normal), and 422 images were used as test data (87 abnormal, 335 normal).
[0035] Data augmentation was used during transfer learning of the AI model using cystoscope images. Data augmentation involved rotating and flipping the training images. The images were rotated in 30° increments between -60° and 60°, and images at -60°, -30°, 0°, 30°, and 60° were used for training. This process increased the number of images used for training to 44,060 (6,295 abnormal, 37,765 normal).
[0036] The experiment in this embodiment was conducted with the approval of the Ethics Committee of the University of Tsukuba Hospital (No. R01-051) and the National Institute of Advanced Industrial Science and Technology (No. I2019-0204). When obtaining consent from patients to participate in the study, an opt-out system was used.
[0037] The learning conditions for transfer learning were as follows: RAdam was used as the optimization algorithm, the learning rate was 0.001 (0.1 times at 30 and 50 epochs), and the loss function was cross-entropy error weighted by the inverse ratio of the number of images, and learning was performed up to 60 epochs.
[0038] (Evaluation method) To evaluate the classification performance of the AI model that underwent transfer learning, the accuracy rate, precision rate, F-score, Youden index (sensitivity - (1 - specificity)), and area under the curve (AUC) based on the receiver operating characteristic curve (ROC curve) were calculated, and the classification performance for each pre-training database was compared. The cutoff value at the threshold that maximized the Youden index is shown in Figure 4.
[0039] In addition, to compare with the classification performance of actual urologists, we cited the results in the paper in Non-Patent Document 5 (see Figure 5). Note that the 422 cystoscope images used as test data in this experiment are the same images used in Non-Patent Document 5 when investigating the classification performance of physicians.
[0040] (Experimental Results) Experiment 1: Figure 4 shows the classification results of an AI model pre-trained using automatically generated images and an AI model trained only with cystoscopy images. In the table in Figure 4, FDB indicates the results when the pre-training database (FractalDB) of Non-Patent Document 1 was used for pre-training, RCDB indicates the results when the pre-training database (RadialContourDB) of Non-Patent Document 2 was used for pre-training, and VADB indicates the results when the pre-training database (VisualAtomDB) of Non-Patent Document 3 was used for pre-training. Note that in Figure 4, the highest score is bolded and underlined, and the second highest score is bolded. As a result of this experiment, the AI model pre-trained using automatically generated images [the AI model pre-trained in this embodiment will be referred to as the MixDB pre-training model hereinafter] showed higher classification results in all evaluation indices than the AI model trained only with cystoscopy images. In particular, the MixDB pre-trained model showed a sensitivity of 0.933, a specificity of 0.975, a Youden index of 0.908, an accuracy rate of 0.966, a precision rate of 0.908, an F-measure of 0.920, and an AUC of 0.978. In all of the comprehensive evaluation indicators of AI models, including the Youden index, accuracy rate, precision rate, F-measure, and AUC, the model performed better than AI models pre-trained with other single types of automatically generated images.
[0041] Experiment 2: The classification results of the MixDB pre-trained model, which had the best classification results in Experiment 1, were compared with the classification results of a urologist. The results are shown in Figure 5. In Figure 5, the top score is shown in bold and underlined, and the second-place score is shown in bold.
[0042] As can be seen from Figure 5, the MixDB pre-training model of this embodiment showed classification performance between urology residents and specialists in four comprehensive evaluation indices: the Ouden index, accuracy rate, and F-measure. Experiment 3: After pre-training an AI model using the same automatically generated images as in Experiment 1 for pre-training, the classification performance of the AI model when transfer learning was performed using a small number (25%) of cystoscopy images was compared with the classification performance of the AI model when training using only this small number (25%) of cystoscopy images, including the results of Experiment 1 in which transfer learning was performed using 100% of cystoscopy images.
[0043] Figure 6 shows the accuracy rate of the classification performance of the AI model for each pre-training training database relative to the proportion of cystoscopy images used during transfer learning. Furthermore, Figure 7 shows the Youden index for the proportion of cystoscopy images used during transfer learning relative to the classification performance of the AI model for each pre-training training database. In Figures 6 and 7, MixDB indicates the results when the mixed database of this embodiment is used as the training database for pre-training. FDB indicates the results when the pre-training database (FractalDB) of Non-Patent Document 1 is used for pre-training. VADB indicates the results when the pre-training database (VisualAtomDB) of Non-Patent Document 3 is used for pre-training. RCDB indicates the results when the pre-training database (RadialContourDB) of Non-Patent Document 2 is used for pre-training. SC indicates the results when an AI model trained only on cystoscopy images without pre-training. 6 and 7 show that MixDB of this embodiment had the smallest rate of decline in classification performance even when the number of training images of bladder endoscopy, which is the target task during transfer learning, was small (25%), indicating that the pre-training method of MixDB of this embodiment functions effectively in creating an AI model when the target task is small.
[0044] (Conclusion) In an experiment (Experiment 1), automatically generated images created using a formula in an AI model's formula-driven supervised learning system were used in the first stage of pre-training, and then transfer learning was performed to cystoscopy images in the second stage of pre-training.The results showed that the AI model (MixDB pre-trained model), which used MixDB for pre-training and performed transfer learning to cystoscopy images, showed a Youden index of 0.908, a precision of 0.966, and an F-measure of 0.920 in a five-fold cross-validation.These three indices represent the classification performance between urology residents and specialists in a comparative experiment with doctors (Experiment 2).
[0045] Furthermore, in an experiment (Experiment 3) in which transfer learning was performed using a small number of cystoscope images, transfer learning was performed using 25% of the total cystoscope images in the second stage of transfer learning. Even with a small number of cystoscope images in the target images, the MixDB pre-trained model was shown to minimize the decline in classification performance and achieve higher classification results than any other method. This suggests that this embodiment may also be an effective method for medical image tasks in which there is less training data available for learning than the cystoscope images used in this experiment.
[0046] In the above embodiment, the mixing ratio of the plurality of automatically generated images in the first type of image dataset DS1 to the plurality of automatically generated images in the second type of image dataset DS2 used to create the mixed database 13A was set to 1:1. However, this mixing ratio can be determined by conducting a preliminary test depending on the target task. Research by the inventors has shown that, for various target tasks, preliminary testing can select a mixing ratio that can show higher classification performance for more evaluation indices than pre-training using an image dataset created using a single mathematical formula, as in the prior art disclosed in Non-Patent Documents 1 to 3. Incidentally, currently known mixing ratios of the first type of image dataset DS1 to the second type of image dataset DS2 are between 3:1 and 1:3. Even if classification performance is poor for some evaluation indices, the mixed database can still be used depending on the nature and purpose of the target task.
[0047] In this embodiment, the first formula is a formula for generating the image data set DS1 for images of shapes and patterns, and the second formula is a formula for generating the image data set DS2 for contour images, but it goes without saying that other formulas can be used depending on the target task. For other formulas, preferable formulas can also be selected by conducting experiments using various formulas for each target task.
[0048] Furthermore, in the above embodiment, the AI model that was pre-trained using the formula-driven supervised learning system 1 for the AI model was trained in two stages by transfer learning using bladder endoscopy images (real images), but the present invention only requires the use of an AI model that was pre-trained using the formula-driven supervised learning system 1, and the type of transfer learning to be performed as the second stage of transfer learning is arbitrary.
[0049] Furthermore, the transfer learning in the case of two-stage learning is not limited to the specific transfer learning in the above embodiment.
[0050] In addition, the images used for pre-learning may be acquired using a center crop method that uses the image portion of the available central area, or a random crop method that uses the image portion appropriately shifted from the available area.
[0051] Furthermore, in the above embodiment, two types of mathematical formulas are used, but the present invention can also be applied to cases where two or more types of image data sets are created using two or more types of mathematical formulas, and the two or more types of image data sets are mixed to create a mixed database.
[0052] According to the present invention, in a formula-driven supervised learning method and system for an AI model that uses a formula to create an image dataset, two or more image datasets are mixed to create a mixed database, and an AI model is pre-trained using the mixed database, thereby providing a pre-trained AI model that can be applied to a target task with higher accuracy than conventional methods.
[0053] REFERENCE SIGNS LIST 1 Formula-driven supervised learning system for AI model 2 Transfer-learned AI model creation system 11 First image dataset creation unit 12 Second image dataset creation unit 13 Mixed database creation unit 13A Mixed database 14 AI model learning unit 15 Transfer learning unit
Claims
1. A formula-driven, supervised learning method for an AI model, comprising: creating two or more image datasets each consisting of two or more types of images that are automatically generated using two or more formulas; mixing the two or more image datasets to create a mixed database; and pre-training an AI model using the mixed database.
2. A formula-driven supervised learning method for an AI model, comprising: creating a first type of image dataset consisting of a plurality of images automatically generated using a first formula; creating a second type of image dataset consisting of a plurality of images automatically generated using a second formula; creating a mixed database by mixing the first type of image dataset and the second type of image dataset; and pre-training an AI model using the mixed database.
3. The formula-driven supervised learning method for an AI model according to claim 2, wherein the first formula is a formula for generating an image dataset of shape or pattern images, and the second formula is a formula for generating an image dataset of contour images.
4. The formula-driven supervised learning method for an AI model according to claim 3, wherein a mixing ratio of the first type of image dataset and the second type of image dataset is determined according to a target task.
5. The formula-driven supervised learning method for an AI model according to claim 4, wherein a target task is endoscopic diagnosis support, and a mixing ratio of the first type of image dataset and the second type of image dataset is a mixing ratio determined by a prior test according to the target task.
6. A method for creating a transfer-learned AI model, comprising performing transfer learning using the AI model created in the formula-driven supervised learning method for an AI model described in claim 5 and a real teacher image of a target task, to create an AI model that has been transfer-learned.
7. A method for creating a transfer learned AI model as described in claim 6, wherein the target task includes endoscopic diagnosis assistance, a fully connected layer of the AI model is changed to a binary classification based on whether an image contains a tumor, and transfer learning is performed for all layers of the AI model to create the transfer learned AI model.
8. A formula-driven supervised learning system for an AI model comprising: a first image dataset creation unit that creates a first type of image dataset consisting of a plurality of images automatically generated using a first formula; a second image dataset creation unit that creates a second type of image dataset consisting of a plurality of images automatically generated using a second formula; a mixed database creation unit that creates a mixed database by mixing the first type of image dataset and the second type of image dataset; and a learning unit that pre-learns an AI model using the mixed database.
9. The formula-driven supervised learning system for an AI model according to claim 8, wherein the first formula is a formula for generating an image dataset of shape or pattern images, and the second formula is a formula for generating an image dataset of contour images.
10. A system for creating a transfer-trained AI model, comprising a transfer learning unit that performs transfer learning using the AI model created by the formula-driven supervised learning system for the AI model described in claim 9 to create a transfer-trained AI model.
11. The system for creating a transfer learned AI model as described in claim 10, wherein the transfer learning unit changes a fully connected layer of the AI model that has undergone the pre-training to a binary classification based on whether an object is included in an image, and performs transfer learning on all layers of the AI model to create the transfer trained AI model.
Citation Information
Patent Citations
Machine learning system
JP2022178892A