Training image screening apparatus, training image screening method, and program
Patent Information
- Application Number
- JP2024571559
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-03
AI Technical Summary
Existing learning data selection methods for machine learning models are inadequate, as they fail to appropriately select unbiased and evenly covered data, particularly when user-determined categories are incorrect or depend on human skill, leading to biased learning data.
A learning image sorting device and method that uses contrastive learning to generate feature amounts of images, with a first machine learning model and a second machine learning model, to determine the similarity of parameters and identify inappropriate training images, ensuring suitable selection of learning images.
The approach effectively sorts learning images by determining the inclusion of inappropriate images based on parameter similarity, ensuring unbiased and appropriate data for machine learning models, thereby improving the quality of learning data.
Smart Images

Figure 2024154316000001
Abstract
Description
Learning image selection device, learning image selection method, and program
[0001] The present invention relates to a training image selection device, a training image selection method, and a program for selecting training images used in training a machine learning model.
[0002] A technique for selecting learning images to be used in training a machine learning model is disclosed.
[0003] Patent Document 1 discloses a learning device that includes a first learning means that executes a first learning process to learn a first model that determines the category to which given data falls through machine learning using training data.
[0004] In addition, the learning device described in Patent Document 1 selects the top learning data as the first learning data and the bottom learning data as the second learning data from the learning data sorted in ascending order based on the difference between the determination result by the first learning means and the correct category determined by the user.
[0005] The learning device described in Patent Document 1 is equipped with a second learning means that executes a second learning process to learn a second model that evaluates the learning data through machine learning using the first learning data and the second learning data.
[0006] International Publication No. 2019 / 187594
[0007] However, the learning device described in Patent Document 1 has a problem in that it cannot properly select learning data if the correct category is incorrect. Because the correct category is determined by the user, the user may set an incorrect correct category. Furthermore, in cases where the correct category depends on the skill of the person setting it, such as pathological cells, the correct category may not always be set correctly.
[0008] Furthermore, in machine learning, it is preferable that the training data is unbiased and comprehensive. However, the learning device described in Patent Literature 1 cannot select training data that is unbalanced as inappropriate training data.
[0009] One aspect of the present invention has been made in consideration of the above-mentioned problems, and one example of its purpose is to provide a technology for suitably selecting training images to be used in training a machine learning model.
[0010] A training image selection device according to one aspect of the present invention comprises: a first learning means for training a first machine learning model, the first machine learning model having a first layer group that receives an image as input and generates features of the image, by contrastive learning using a training image set that is a plurality of training images; the first layer group; and a second layer group connected to the first layer group and that receives features of the image as input and classifies the image; a second learning means for training a second machine learning model, the second machine learning model having the first machine learning model as a pre-training model, using the training image set; a first calculation means for calculating a first similarity, which is the similarity between parameters in the first layer group before training by the second learning means and parameters in the first layer group after training by the second learning means; and a first determination means for determining whether or not an inappropriate training image is included in the training image set, based on the first similarity.
[0011] A method for selecting training images according to one aspect of the present invention includes: a training image selection device training a first machine learning model, the first machine learning model having a first layer group that receives an image as input and generates features of the image, by contrastive learning using a training image set that is a plurality of training images; training a second machine learning model, the second machine learning model having the first layer group and a second layer group connected to the first layer group and that receives features of the image as input and classifies the image, the second machine learning model having the first machine learning model as a pre-training model, using the training image set; calculating a first similarity that is the similarity between parameters in the first layer group before training by training the second machine learning model trained in the contrastive learning and parameters in the first layer group after training by training the second machine learning model; and determining whether or not an inappropriate training image is included in the training image set based on the first similarity.
[0012] A program according to one aspect of the present invention is a program that causes a computer to function as a training image selection device, and causes the computer to function as: first learning means that trains a first machine learning model by contrastive learning using a training image set that is a plurality of training images, the first machine learning model having a first layer group that receives an image as input and generates features of the image; second learning means that trains a second machine learning model using the training image set, the second machine learning model having the first layer group and a second layer group that is connected to the first layer group and receives features of the image as input and classifies the image, the second machine learning model having the first machine learning model as a pre-training model; first calculation means that calculates a first similarity that is the similarity between parameters in the first layer group before training by the second learning means and parameters in the first layer group after training by the second learning means; and first determination means that determine whether or not an inappropriate training image is included in the training image set based on the first similarity.
[0013] According to one aspect of the present invention, it is possible to suitably select learning images to be used in training a machine learning model.
[0014] FIG. 1 is a block diagram showing the configuration of a training image selection device according to exemplary embodiment 1 of the present invention. FIG. 2 is a flow diagram showing the flow of a training image selection method according to exemplary embodiment 1 of the present invention. FIG. 3 is a block diagram showing the configuration of a training image selection device according to exemplary embodiment 2 of the present invention. FIG. 4 is a diagram showing an example of a first machine learning model in exemplary embodiment 2 of the present invention. FIG. 5 is a diagram showing an example of a second machine learning model in exemplary embodiment 2 of the present invention. FIG. 6 is a flow diagram showing the flow of a training image selection method according to exemplary embodiment 2 of the present invention. FIG. 7 is a block diagram showing the configuration of a training image selection device according to exemplary embodiment 3 of the present invention. FIG. 8 is a diagram showing the attributes of each of a plurality of training images in exemplary embodiment 3 of the present invention. FIG. 9 is a flow diagram showing the flow of a training image selection method according to exemplary embodiment 3 of the present invention. FIG. 10 is a flow diagram showing the flow of a training image selection method according to a modification of exemplary embodiment 3 of the present invention. FIG. 11 is a block diagram showing an example of the hardware configuration of a training image selection device according to each exemplary embodiment of the present invention.
[0015] [First Exemplary Embodiment] A first exemplary embodiment of the present invention will be described in detail with reference to the drawings. This exemplary embodiment is a basic form of the exemplary embodiments described below.
[0016] (Configuration of Training Image Selection Device 1) The training image selection device 1 according to this exemplary embodiment is a device that selects training images to be used when training a machine learning model. For example, the training image selection device 1 selects training images by determining whether a training image set, which is a plurality of training images, includes inappropriate training images. An example of an inappropriate training image is a training image with an incorrect teacher label. Another example of a case in which an inappropriate training image is included in a training image set is when there is a bias among the multiple training images included in the training image set.
[0017] The configuration of a learning image selection device 1 according to this exemplary embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of a learning image selection device 1 according to this exemplary embodiment.
[0018] 1, the learning image selection device 1 includes a first learning unit 11, a second learning unit 12, a first calculation unit 13, and a first determination unit 14. In this exemplary embodiment, the first learning unit 11, the second learning unit 12, the first calculation unit 13, and the first determination unit 14 are configured to realize a first learning means, a second learning means, a first calculation means, and a first determination means, respectively.
[0019] The first learning unit 11 uses an image as input and trains a first machine learning model having a first group of layers that generates features of the image by contrastive learning using a training image set that is a plurality of training images.
[0020] Contrastive learning is a method of selecting one target image (anchor) from multiple training images, and training a machine learning model so that the dot product of the feature vectors of the target image and positive examples (training images classified in the same category as the target image, and images obtained by applying optional image extensions to the target image) is large, and the dot product of the feature vectors of the target image and negative examples (training images classified in a different category from the target image) is small.
[0021] The first machine learning model includes an Encoder (feature analysis model), which is a first layer group that receives an input image and generates features of the input image. The first machine learning model is used as a pre-training model for the second machine learning model described below.
[0022] The second learning unit 12 includes a first group of layers and a second group of layers connected to the first group of layers that classifies images using image features as input, and uses a learning image set to train a second machine learning model that uses the first machine learning model trained by the first learning unit 11 as a pre-training model.
[0023] The second machine learning model is a model in which a second layer group (Classifier) is connected to the first layer group (Encoder) of the first machine learning model. The second learning unit 12 mainly trains the Classifier part, but also trains the Encoder part so as to fine-tune it.
[0024] A known method may be used as the method by which the second learning unit 12 trains the first machine learning model and the second machine learning model. One example of the method by which the second learning unit 12 fine-tunes the first machine learning model and trains the second machine learning model is a method of learning using cross entropy loss as a loss function to minimize the error between the output from the machine learning model and ground truth data.
[0025] The first calculation unit 13 calculates a first similarity, which is the similarity between the parameters in the first layer group (encoder, feature analysis model) learned by the first learning unit 11 before learning by the second learning unit 12, and the parameters in the first layer group (encoder, feature analysis model) after learning by the second learning unit 12.
[0026] Hereinafter, the parameters of the first layer group (encoder, feature analysis model) after the first learning unit 11 has trained the first machine learning model but before the second learning unit 12 has trained it will also be referred to as "first parameters." Furthermore, the parameters in the first layer group (encoder, feature analysis model) of the second machine learning model after the second learning unit 12 has trained it will also be referred to as "second parameters."
[0027] That is, the first calculation unit 13 calculates a first similarity, which is the similarity between the first parameter and the second parameter, and supplies the calculated first similarity to the first determination unit 14.
[0028] The first determination unit 14 determines whether or not an inappropriate training image is included in the training image set, based on the first similarity calculated by the first calculation unit 13 .
[0029] For example, if the first parameter and the second parameter are similar, the first determination unit 14 determines that the training image set does not include an inappropriate training image. In this case, if the first similarity is equal to or greater than a threshold, the first determination unit 14 determines that the training image set does not include an inappropriate training image. Furthermore, if the first similarity is less than the threshold, the first determination unit 14 determines that the training image set includes an inappropriate training image.
[0030] As described above, the training image selection device 1 according to this exemplary embodiment employs a configuration including: a first training unit 11 that trains a first machine learning model, the first machine learning model having a first layer group that receives an image as input and generates features of the image, by contrastive training using a training image set that is a plurality of training images; a second training unit 12 that trains a second machine learning model using the training image set, the second machine learning model having the first machine learning model as a pre-training model, the second layer group being connected to the first layer group and receiving features of the image as input and having the first machine learning model as a pre-training model; a first calculation unit 13 that calculates a first similarity, which is the similarity between the parameters in the first layer group trained by the first training unit 11 before training by the second learning unit 12 and the parameters in the first layer group after training by the second learning unit 12; and a first determination unit 14 that determines whether or not an inappropriate training image is included in the training image set, based on the first similarity calculated by the first calculation unit 13.
[0031] In this configuration, if the first machine learning model is trained to extract highly invariant features, the first similarity will be high. On the other hand, if the first machine learning model is not trained to extract highly invariant features, the first similarity will be low. For example, if the training image set includes training images with inappropriate teacher labels or if there is a bias among the multiple training images included in the training image set, the first machine learning will not be trained to extract highly invariant features, and the first similarity will be low.
[0032] The training image selection device 1 according to this exemplary embodiment determines whether or not an inappropriate training image is included in the training image set based on the first similarity. Therefore, when the first similarity is high, the training image selection device 1 according to this exemplary embodiment can determine that the first machine learning model has been trained to extract highly invariant features, and can determine that an inappropriate training image is not included in the training image set.
[0033] On the other hand, if the first similarity is low, the training image selection device 1 according to this exemplary embodiment can determine that the first machine learning model has not been trained to extract highly invariant features, and can determine that inappropriate training images are included in the training image set.
[0034] Therefore, the learning image selection device 1 according to this exemplary embodiment has the effect of being able to suitably select learning images to be used in training a machine learning model.
[0035] (Flow of Learning Image Selection Method S1) The flow of the learning image selection method S1 according to this exemplary embodiment will be described with reference to Fig. 2. Fig. 2 is a flow chart showing the flow of the learning image selection method S1 according to this exemplary embodiment.
[0036] (Step S11) In step S11, the first learning unit 11 trains a first machine learning model having a first layer group that receives an image as input and generates features of the image by contrastive learning using a training image set that is a plurality of training images.
[0037] (Step S12) In step S12, the second learning unit 12 includes a first layer group and a second layer group connected to the first layer group and configured to classify images using image features as input, and uses the learning image set to train a second machine learning model that uses the first machine learning model trained by the first learning unit 11 as a pre-training model.
[0038] (Step S13) In step S13, the first calculation unit 13 calculates a first similarity, which is the similarity between the parameters in the first layer group (encoder, feature analysis model) trained by the first learning unit 11 before training by the second learning unit 12, and the parameters in the first layer group (encoder, feature analysis model) after training by the second learning unit 12. In other words, in step S13, the first calculation unit 13 calculates the first similarity, which is the similarity between the first parameter and the second parameter. The first calculation unit 13 supplies the calculated first similarity to the first determination unit 14.
[0039] (Step S14) In step S14, the first determination unit 14 determines whether or not an inappropriate learning image is included in the learning image set, based on the first similarity calculated by the first calculation unit 13.
[0040] For example, in step S14, if the first parameter and the second parameter are similar, the first determination unit 14 determines that the training image set does not include an inappropriate training image. In this case, if the first similarity is equal to or greater than a threshold, the first determination unit 14 determines that the training image set does not include an inappropriate training image. Furthermore, if the first similarity is less than the threshold, the first determination unit 14 determines that the training image set includes an inappropriate training image.
[0041] As described above, in the training image selection method S1 according to this exemplary embodiment, the first training unit 11 trains a first machine learning model, which has a first layer group that receives an image as input and generates features of the image, by contrastive training using a training image set that is a plurality of training images. The second training unit 12 trains a second machine learning model, which has the first layer group and a second layer group that is connected to the first layer group and classifies images using features of the image as input, using the training image set. The second machine learning model uses the first machine learning model trained by the first training unit 11 as a pre-training model. The configuration includes step S12 in which the first calculation unit 13 calculates a first similarity, which is the similarity between the parameters in the first layer group (encoder, feature analysis model) trained by the first learning unit 11 before training by the second learning unit 12 and the parameters in the first layer group (encoder, feature analysis model) trained by the second learning unit 12, and step S14 in which the first determination unit 14 determines whether or not an inappropriate training image is included in the training image set, based on the first similarity calculated by the first calculation unit 13. Therefore, the training image selection method S1 according to this exemplary embodiment can achieve the same effects as those of the above-described training image selection device 1.
[0042]
[0033] A second exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first exemplary embodiment are denoted by the same reference numerals, and their description will be omitted as appropriate.
[0043] (Overview of Training Image Selection Device 2) The training image selection device 2 according to this exemplary embodiment is a device that selects a portion of a plurality of images as a training image set, which is a plurality of training images for training a machine learning model, and outputs the training image set if the training image set is appropriate for machine learning. For example, the training image selection device 2 selects training images by determining whether the training image set, which is a plurality of training images, includes any inappropriate training images, and outputs the training image set if no inappropriate training images are included.
[0044] Furthermore, if the training image selection device 2 determines that the training image set includes an inappropriate training image, the training image selection device 2 selects a training image set different from the selected training image set. For example, the training image selection device 2 selects a training image set different from the selected training image set by changing at least one training image from among the training images included in the selected training image set to an unselected training image.
[0045] The learning image selection device 2 selects learning images by determining whether or not the newly selected learning image set contains any inappropriate learning images, and outputs the learning image set if it does not contain any inappropriate learning images.
[0046] Examples of inappropriate training images include training images with incorrect teacher labels, and training image sets containing biased training images.
[0047] (Configuration of the learning image selection device 2) The configuration of the learning image selection device 2 according to this exemplary embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of the learning image selection device 2 according to this exemplary embodiment. As shown in Fig. 3, the learning image selection device 2 includes a control unit 21, a storage unit 25, a communication unit 26, an input unit 27, and an output unit 28.
[0048] The storage unit 25 stores data referenced by the control unit 21. An example of the data stored in the storage unit 25 is training images and teacher labels corresponding to the training images.
[0049] The communication unit 26 is a communication module that communicates with other devices connected via a network. For example, the communication unit 26 receives training images and outputs a training image set that is determined not to include inappropriate training images.
[0050] The input unit 27 is an interface that acquires data from other connected devices. As an example, the input unit 27 acquires learning images.
[0051] The output unit 28 is an interface that outputs data to another device connected thereto. As an example, the output unit 28 outputs a set of training images that have been determined not to include inappropriate training images.
[0052] (Functions of the control unit 21) The control unit 21 controls each component included in the learning image selection device 2. As shown in Fig. 3, the control unit 21 also includes a first learning unit 11, a second learning unit 12, a first calculation unit 13, a first determination unit 14, and a selection unit 22. In this exemplary embodiment, the first learning unit 11, the second learning unit 12, the first calculation unit 13, the first determination unit 14, and the selection unit 22 are configured to realize a first learning means, a second learning means, a first calculation means, a first determination means, and a selection means, respectively.
[0053] The first learning unit 11 trains the machine learning model by contrastive learning. As an example, the first learning unit 11 trains a first machine learning model including a first layer group that receives an image as input and generates features of the image by contrastive learning using a training image set that is a plurality of training images selected by a selection unit 22 described later.
[0054] An example of the first machine learning model is shown in Fig. 4. Fig. 4 is a diagram showing an example of the first machine learning model in this exemplary embodiment. As shown in Fig. 4, the first machine learning model includes an Encoder (feature analysis model) which is a first layer group that receives an image as input and outputs a feature vector as a feature amount of the image.
[0055] The second learning unit 12 learns a machine learning model using a known method. As an example, the second learning unit 12 includes a first group of layers and a second group of layers connected to the first group of layers and configured to classify images using image features as input, and learns a second machine learning model using a training image set, the second machine learning model using the first machine learning model as a pre-training model.
[0056] An example of the second machine learning model is shown in Fig. 5. Fig. 5 is a diagram showing an example of the second machine learning model in this exemplary embodiment. As shown in Fig. 5, the second machine learning model includes a first group of layers (encoder, feature analysis model) included in the first machine learning model trained by the first learning unit 11, and a second group of layers (classifier) that classifies an input image and outputs a classification result. In other words, the second machine learning model is a combination of the first group of layers (encoder, feature analysis model) and the second group of layers (classifier).
[0057] As an example, a pathology image containing specimen cells as a subject is input to a first machine learning model, and a second machine learning model outputs a classification result classifying the specimen cells as benign or malignant.
[0058] Hereinafter, the first layer group (encoder, feature analysis model) of the first machine learning model trained by the first learning unit 11 and before training by the second learning unit 12 will also be referred to as the "feature analysis model M1." The first layer group (encoder, feature analysis model) trained by the second learning unit 12 will also be referred to as the "feature analysis model M2." When there is no particular need to distinguish between them, they will simply be referred to as the "feature analysis model."
[0059] The first calculation unit 13 calculates a first similarity, which is the similarity between a parameter (weight, first parameter) in the feature analysis model M1 and a parameter (second parameter) in the feature analysis model M2. As an example, the first calculation unit 13 calculates a second similarity, which is the similarity for each layer of a first layer group (encoder, feature analysis model) included in the first machine learning model. In this case, the first calculation unit 13 calculates the first similarity based on the calculated second similarity. An example of the process in which the first calculation unit 13 calculates the first similarity and the second similarity will be described later.
[0060] The first determination unit 14 determines whether or not an inappropriate training image is included in the training image set. As an example, the first determination unit 14 determines whether or not an inappropriate training image is included in the training image set based on the first similarity calculated by the first calculation unit 13.
[0061] For example, if the first similarity is equal to or greater than a threshold, the first determination unit 14 determines that the training image set does not include an inappropriate training image, and if the first similarity is less than the threshold, the first determination unit 14 determines that the training image set includes an inappropriate training image.
[0062] The selection unit 22 selects a portion of the multiple images as the training image set. As an example, the selection unit 22 selects a portion of the training images stored in the storage unit 25 as the training image set. The number of training images selected by the selection unit 22 is not particularly limited. As an example, the selection unit 22 may randomly select a predetermined number (e.g., 9,500 or more) from all training images (e.g., 10,000). The selection unit 22 supplies the selected training image set to the first training unit 11 and the second training unit 12.
[0063] Furthermore, when repeatedly selecting some of the multiple images as the training image set, the selection unit 22 selects a training image set different from the already selected training image set. As an example, when the first determination unit 14 determines that an inappropriate training image is included in the training image set, the selection unit 22 selects a training image set different from the already selected training image set. With this configuration, the selection unit 22 can cause the first determination unit 14 to determine whether an inappropriate training image is included in a training image set different from the training image set that has already been determined to include an inappropriate training image.
[0064] (Flow of Learning Image Selection Method S2) The flow of the learning image selection method S2 according to this exemplary embodiment will be described with reference to Fig. 6. Fig. 6 is a flow chart showing an example of the flow of the learning image selection method S2 according to this exemplary embodiment.
[0065] (Step S21) In step S21, the selection unit 22 selects, as a set of learning images, a portion of the learning images stored in the storage unit 25. The selection unit 22 supplies the selected set of learning images to the first learning unit 11 and the second learning unit 12.
[0066] (Step S22) In step S22, the first learning unit 11 trains a first machine learning model including a first layer group (encoder, feature analysis model) by contrastive learning using the training image set supplied by the selection unit 22. The first layer group (encoder, feature analysis model) of the first machine learning model after training by the first learning unit 11 in step S22 is the feature analysis model M1.
[0067] (Step S23) In step S23, the second learning unit 12 trains a second machine learning model that includes a first layer group and a second layer group and that uses the feature analysis model M1 as a pre-training model, using the training image set supplied by the selection unit 22. The first layer group (Encoder, feature analysis model) trained by the second learning unit 12 in step S23 is feature analysis model M2.
[0068] (Step S24) In step S24, the first calculation unit 13 calculates a second similarity, which is the similarity for each layer of the first layer group included in the first machine learning model. In other words, in step S24, the first calculation unit 13 calculates the similarity between the first parameter in each layer of the feature analysis model M1 and the second parameter in each layer of the feature analysis model M2. The first calculation unit 13 stores the calculated second similarity in the storage unit 25.
[0069] As an example, the first calculation unit 13 calculates the second similarity “similarity k (x, y)" is calculated using the following formula (1).
[0070] where x is the first parameter (weight vector) in the k-th layer of the feature analysis model M1, and x = (x 1 , x 2 , x 3 ,...xn ) where y is the second parameter (weight vector) in the k-th layer of the feature analysis model M2, and y = (y 1 , y 2 , y 3 ,...y n )
[0071] (Step S25) In step S25, the first calculation unit 13 calculates a first similarity based on the second similarity stored in the storage unit 25. The first calculation unit 13 stores the calculated first similarity in the storage unit 25.
[0072] As an example, the first calculation unit 13 calculates the first similarity by dividing the sum of the second similarities by the number of layers in the first layer group included in the first machine learning model. Specifically, the first calculation unit 13 calculates the second similarity “similarity k Using (x, y)', the first similarity, "similarity", is calculated using the following formula (2).
[0073] Here, m is the number of layers in the first machine learning model.
[0074] As another example, the first calculation unit 13 calculates the first similarity by dividing the weighted sum obtained by adding weights to the second similarities by the total sum of the weight values. Specifically, the first calculation unit 13 calculates the second similarity “similarity k Using (x, y)', the first similarity, "similarity", is calculated using the following formula (3).
[0075] Here, W k is the weight given to the k-th second similarity.
[0076] Furthermore, the first calculation unit 13 may increase the weight value of the second similarity of a layer (deeper layer) in the first layer group that is closer to the output of one machine learning model. With this configuration, the first calculation unit 13 can increase the influence of the second similarity on the first similarity of a layer that is closer to the output and focuses on global features.
[0077] (Step S26) In step S26, the first determination unit 14 determines whether the first similarity stored in the storage unit 25 is equal to or greater than a threshold value.
[0078] (Step S27) If it is determined in step S27 that the first similarity is equal to or greater than the threshold value (step S26: YES), the first determination unit 14 outputs the training image set in step S27. In other words, if the first determination unit 14 determines that the training image set does not include any inappropriate training images, it outputs the training image set.
[0079] On the other hand, if it is determined in step S26 that the first similarity is less than the threshold value (step S26: NO), the learning image selection device 2 returns to the process of step S21. In other words, if the first determination unit 14 determines that an inappropriate learning image is included in the learning image set, the learning image selection device 2 returns to the process of step S21.
[0080] In step S21, the selection unit 22 selects a training image set different from the previously selected training image set. Then, in the processing from step S22 onward, it is determined whether or not the newly selected training image set includes any inappropriate training images.
[0081] Effect of Exemplary Embodiment 2 As described above, in the training image selection device 2 according to this exemplary embodiment, when it is determined that the training image set includes an inappropriate training image, the selection unit 22 selects a training image set that is different from the previously selected training image set. The first determination unit 14 then determines whether the training image set newly selected by the selection unit 22 includes an inappropriate training image. With this configuration, the training image selection device 2 according to this exemplary embodiment does not output a training image set until it is determined that the training image set does not include an inappropriate training image, making it possible to output an appropriate training image set.
[0082] (Variation 1) The training image selection device 2A according to a variation of this exemplary embodiment executes, until a predetermined time has elapsed, processes from selecting a training image set to determining whether or not an inappropriate training image is included in the training image set. Note that, instead of (or in addition to) the predetermined time, the training image selection device 2A may be configured to execute, a predetermined number of times, processes from selecting a training image set to determining whether or not an inappropriate training image is included in the training image set.
[0083] The configuration of the learning image selection device 2A is the same as that of the learning image selection device 2, and therefore a description thereof will be omitted.
[0084] (Flow of Learning Image Selection Method S2A) The flow of the learning image selection method S2A according to the modified example of this exemplary embodiment will be described with reference to Fig. 7. Fig. 7 is a flow chart showing the flow of the learning image selection method S2A according to the modified example of this exemplary embodiment.
[0085] (Steps S21 to S25) The processes of steps S21 to S25, from when the selection unit 22 selects a learning image set to when the first calculation unit 13 calculates the first similarity, are the same as the processes described above, and therefore will not be described again.
[0086] (Step S26a) In step S26a, the first determination unit 14 determines whether a predetermined time has elapsed.
[0087] If it is determined in step S26a that the predetermined time has not elapsed (step S26a: NO), the learning image selection device 2A returns to the processing of step S21. Then, in step S21, the selection unit 22 selects a learning image set different from the previously selected learning image set, and the processing from step S22 onward is executed using the selected learning image set.
[0088] In addition, if the learning image selection device 2A is configured to perform the process from selecting a learning image set to determining whether or not an inappropriate learning image is included in the learning image set a predetermined number of times, instead of (or in addition to) a predetermined period of time, in step S26a, the first judgment unit 14 may be configured to determine whether or not the process of step S26a has been performed a predetermined number of times, instead of (or in addition to) determining whether or not a predetermined period of time has elapsed.
[0089] In this configuration, if it is determined in step S26a that the process of step S26a has been executed a predetermined number of times (step S26a: YES), the learning image selection device 2A proceeds to the process of step S27a. On the other hand, if it is determined in step S26a that the process of step S26a has not been executed a predetermined number of times (step S26a: NO), the learning image selection device 2A returns to the process of step S21. Here, in step S25, the first calculation unit 13 stores the calculated first similarity in the storage unit 25 each time. In other words, if steps S21 to S25 are repeated N times, the first calculation unit 13 stores the first similarity for N times in the storage unit 25.
[0090] (Step S27a) If it is determined in step S26a that a predetermined time has elapsed (step S26a: YES), in step S27a, the first judgment unit 14 determines whether or not any of the multiple first similarities stored in the memory unit 25 is greater than or equal to a threshold value.
[0091] (Step S28a) If it is determined in step S27a that the first similarity is equal to or greater than the threshold (step S27a: YES), in step S28a, the first determination unit 14 outputs the training image set corresponding to the highest first similarity among the first similarities equal to or greater than the threshold. In other words, the first determination unit 14 outputs the training image set determined to be most appropriate for training from among the multiple training image sets.
[0092] (Step S29a) If it is determined in step S27a that there is no first similarity greater than or equal to the threshold (step S27a: YES), in step S29a, the first judgment unit 14 outputs a message indicating that a learning image set suitable for learning could not be selected.
[0093] (Effects of Modification 1) The training image selection device 2A according to the modification of this exemplary embodiment executes processes from selecting a training image set to determining whether or not the training image set contains inappropriate training images until a predetermined time has elapsed (or until the processes have been executed a predetermined number of times). Therefore, in addition to the effects achieved by the training image selection device 2 according to exemplary embodiment 2, the training image selection device 2A according to the modification of this exemplary embodiment can output a training image set determined to be most appropriate for training from among the selected training image sets.
[0094]
[0033] A third exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the above exemplary embodiment are denoted by the same reference numerals, and their description will not be repeated.
[0095] (Configuration of Training Image Selection Device 3) In addition to the functions of the training image selection device 2 (and training image selection device 2A) described above, the training image selection device 3 according to this exemplary embodiment determines whether there is a bias in the attributes of each of the multiple training images included in the training image set. If the training image selection device 3 determines that there is a bias in the attributes of each of the multiple training images included in the training image set, it determines that an inappropriate training image is included in the training image set. The attributes of the training images will be described later.
[0096] The configuration of the learning image selection device 3 according to this exemplary embodiment will be described with reference to Fig. 8. Fig. 8 is a block diagram showing the configuration of the learning image selection device 3 according to this exemplary embodiment. As shown in Fig. 8, the learning image selection device 2 includes a control unit 31, a storage unit 25, a communication unit 26, an input unit 27, and an output unit 28.
[0097] The storage unit 25, communication unit 26, input unit 27, and output unit 28 are the same as those described in the first exemplary embodiment, and therefore a description thereof will be omitted.
[0098] (Functions of the control unit 31) The control unit 31 controls each component included in the learning image selection device 3. As shown in Fig. 8 , the control unit 31 also includes a first learning unit 11, a second learning unit 12, a first calculation unit 13, a first determination unit 14, a selection unit 22, a second calculation unit 32, and a second determination unit 33. In this exemplary embodiment, the first learning unit 11, the second learning unit 12, the first calculation unit 13, the first determination unit 14, the selection unit 22, the second calculation unit 32, and the second determination unit 33 respectively implement a first learning means, a second learning means, a first calculation means, a first determination means, a selection means, a second calculation means, and a second determination means.
[0099] The first learning unit 11, the second learning unit 12, the first calculation unit 13, the first judgment unit 14, and the selection unit 22 are as described in the above exemplary embodiment, so their description will be omitted.
[0100] The second calculation unit 32 calculates an index indicating the bias in the attributes of each of the multiple images. As an example, the second calculation unit 32 calculates an index indicating the bias in the attributes of each of the multiple training images included in the training image set selected by the selection unit 22. In the following, variance will be used as an example of the index, but the index is not limited to this. The attributes of the training images will be described with reference to FIG. 9. FIG. 9 is a diagram showing the attributes of each of the multiple training images in this exemplary embodiment.
[0101] 9 , if the training images were captured at any one of Hospital A to Hospital Z, the second calculation unit 32 sets the facility where the images were captured as an attribute. In this case, the second calculation unit 32 calculates the number of data items for each hospital where each of the multiple training images included in the training image set was captured. The second calculation unit 32 then calculates a variance as an index indicating the bias of the facilities where each of the multiple training images included in the training image set was captured.
[0102] 9, if the training images were images captured using any of the models of scanners A to Z, the second calculation unit 32 sets the model of the imaging device as the attribute. In this case as well, the second calculation unit 32 calculates the variance as an index indicating the bias in the models of the imaging devices that captured each of the multiple training images included in the training image set.
[0103] 9, when the training images are images containing cells such as normal epithelial cells, small cell carcinoma, adenocarcinoma, and squamous cell carcinoma as subjects, the second calculation unit 32 sets the type of cells contained as subjects as an attribute. In this case as well, the second calculation unit 32 calculates the variance as an index indicating the bias in the types of cells contained as subjects in each of the multiple training images included in the training image set.
[0104] The second determination unit 33 determines whether the index calculated by the second calculation unit 32 is equal to or greater than a threshold. In other words, the second determination unit 33 determines whether there is bias in the attributes of each of the multiple training images included in the training image set. As an example, if the second calculation unit 32 calculates variance as the index, the second determination unit 33 determines whether the value of the variance is equal to or greater than a threshold (whether there is bias) or whether the value of the variance is less than the threshold (whether there is no bias).
[0105] (Flow of Learning Image Selection Method S3) The flow of the learning image selection method S3 according to this exemplary embodiment will be described with reference to Fig. 10. Fig. 10 is a flow chart showing the flow of the learning image selection method S3 according to this exemplary embodiment.
[0106] (Steps S21 to S26) The processes of steps S21 to S26, from when the selection unit 22 selects a learning image set until when the first judgment unit 14 judges whether the first similarity is equal to or greater than a threshold, are the same as those described above, and therefore will not be described here.
[0107] (Step S31) If it is determined in step S26 that the first similarity is greater than or equal to the threshold (step S26: YES), in step S31, the second calculation unit 32 calculates an index indicating the bias in the attributes of each of the multiple training images included in the training image set selected by the selection unit 22.
[0108] (Step S32) In step S32, the second determination unit 33 determines whether the index calculated by the second calculation unit 32 is less than the threshold value.
[0109] If it is determined in step S32 that the index calculated by the second calculation unit 32 is not less than the threshold value (step S32: NO), the training image selection device 3 returns to the processing of step S21. Then, in step S21, the selection unit 22 selects a training image set different from the previously selected training image set. In other words, if there is a bias in the attributes of the multiple training images included in the training image set, the selection unit 22 selects a training image set different from the previously selected training image set.
[0110] As described above, step S32 is a process that is executed when the index is variance. However, even when the index is other than variance, if the second judgment unit 33 judges in step S32 based on the index that there is bias in the attributes of each of the multiple training images included in the training image set, the training image selection device 3 returns to the process of step S21.
[0111] (Step S27) If it is determined in step S32 that the index calculated by the second calculation unit 32 is less than the threshold value (step S32: YES), the second determination unit 33 outputs the training image set in step S27. In other words, if there is no bias in the attributes of the multiple training images included in the training image set, the second determination unit 33 outputs the training image set.
[0112] (Effects of Exemplary Embodiment 3) As described above, the training image selection device 3 according to this exemplary embodiment employs a configuration including the second calculation unit 32 that calculates an index indicating bias in the attributes of each of the multiple training images included in the training image set selected by the selection unit 22, and the second determination unit 33 that determines whether the index calculated by the second calculation unit 32 is less than a threshold value. With this configuration, the training image selection device 3 according to this exemplary embodiment can provide a training image set including unbiased training images, in addition to the effects achieved by the training image selection device 1 according to exemplary embodiment 1.
[0113] (Variant 2) In a training image selection device 3A according to a variant of this exemplary embodiment, before training the first machine learning model, it is determined whether there is any bias in the attributes of each of the multiple training images included in the training image set.
[0114] The configuration of the learning image selection device 3A is the same as that of the learning image selection device 3, and therefore a description thereof will be omitted.
[0115] (Flow of Learning Image Selection Method S3A) The flow of the learning image selection method S3A according to the modified example of this exemplary embodiment will be described with reference to Fig. 11. Fig. 11 is a flow chart showing the flow of the learning image selection method S3A according to the modified example of this exemplary embodiment.
[0116] (Step S21) In step S21, the selection unit 22 selects, as a set of learning images, some of the learning images stored in the storage unit 25. The selection unit 22 supplies the selected set of learning images to the second calculation unit 32.
[0117] (Step S31) In step S31, the second calculation unit 32 calculates an index indicating the bias in the attributes of each of the plurality of learning images included in the learning image set selected by the selection unit 22.
[0118] (Step S32) In step S32, the second determination unit 33 determines whether the index calculated by the second calculation unit 32 is less than the threshold value.
[0119] In step S32, if it is determined that the index calculated by the second calculation unit 32 is not less than the threshold value (step S32: NO), the training image selection device 3A returns to the processing of step S21. In other words, if there is a bias in the attributes of the multiple training images included in the training image set, the selection unit 22 selects a training image set different from the previously selected training image set in step S21.
[0120] On the other hand, if it is determined in step S32 that the index calculated by the second calculation unit 32 is less than the threshold value (step S32: YES), the selection unit 22 supplies the selected training image set to the first training unit 11 and the second training unit 12. In other words, if the second determination unit 33 determines that there is no bias in the attributes of the multiple training images included in the training image set, the selection unit 22 supplies the selected training image set to the first training unit 11 and the second training unit 12.
[0121] (Steps S22 to S27) The processing from steps S22 to S27, in which the first learning unit 11 trains the first machine learning model by contrastive learning and the first judgment unit 14 outputs a learning image set when it determines that the first similarity is equal to or greater than a threshold, is the same as the processing described above, and therefore will not be described again.
[0122] (Effects of Modification 2) As described above, in the training image selection device 3A according to the modification of this exemplary embodiment, before training the first machine learning model, it is determined whether there is any bias in the attributes of the plurality of training images included in the training image set. With this configuration, in addition to the effects achieved by the training image selection device 3 according to exemplary embodiment 3, the training image selection device 3A according to the modification of this exemplary embodiment can reduce the processing load by not training the machine learning model if there is any bias in the attributes of the plurality of training images included in the training image set.
[0123] [Example of Software Implementation] Some or all of the functions of the learning image selection devices 1, 2, 2A, 3, and 3A may be implemented by hardware such as an integrated circuit (IC chip), or may be implemented by software.
[0124] In the latter case, the learning image selection devices 1, 2, 2A, 3, and 3A are realized, for example, by a computer that executes program instructions, which are software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in FIG. 12 . Computer C includes at least one processor C1 and at least one memory C2. Memory C2 stores a program P for operating computer C as learning image selection devices 1, 2, 2A, 3, and 3A. In computer C, processor C1 reads and executes program P from memory C2, thereby realizing each function of learning image selection devices 1, 2, 2A, 3, and 3A.
[0125] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0126] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.
[0127] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0128] [Additional Note 1] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the above-described embodiments are also included in the technical scope of the present invention.
[0129] [Additional Note 2] Part or all of the above-described embodiment can also be described as follows: However, the present invention is not limited to the following described aspects.
[0130] (Supplementary Note 1) A training image selection device comprising: first learning means for training a first machine learning model by contrastive learning using a training image set that is a plurality of training images, the first machine learning model including a first layer group that receives an image as input and generates features of the image; second learning means for training a second machine learning model using the training image set, the second machine learning model having the first machine learning model as a pre-training model, the second layer group being connected to the first layer group and receiving features of the image as input and classifying the image; first calculation means for calculating a first similarity that is a similarity between parameters in the first layer group before training by the second learning means and parameters in the first layer group after training by the second learning means, the parameters having been trained by the first learning means; and first determination means for determining whether or not an inappropriate training image is included in the training image set, based on the first similarity.
[0131] (Supplementary Note 2) The training image selection device according to Supplementary Note 1, further comprising a selection means for selecting a portion of a plurality of images as the training image set, wherein if the first determination means determines that the inappropriate training image is included in the training image set, the selection means selects a training image set different from the already selected training image set.
[0132] (Supplementary Note 3) The training image selection device according to Supplementary Note 1 or 2, wherein the first calculation means calculates a second similarity, which is a similarity for each layer of the first layer group provided in the first machine learning model.
[0133] (Supplementary Note 4) The training image selection device according to Supplementary Note 3, wherein the first calculation means calculates the first similarity by dividing the sum of the second similarities by the number of layers in the first layer group included in the first machine learning model, and the first determination means determines that the inappropriate training image is included in the training image set if the first similarity is less than a threshold.
[0134] (Supplementary Note 5) The training image selection device described in Supplementary Note 3, wherein the first calculation means calculates the first similarity by dividing the weighted sum obtained by weighting each of the second similarities and adding the weighted sum by the total sum of the weight values, and the first determination means determines that the inappropriate training image is included in the training image set if the first similarity is less than a threshold value.
[0135] (Supplementary Note 6) The training data selection device according to Supplementary Note 5, wherein the first calculation means increases the weight value of the second similarity of a layer of the first layer group that is closer to the output of the first machine learning model.
[0136] (Supplementary Note 7) The training image selection device described in Supplementary Note 2 further includes a second calculation means for calculating an index indicating the bias in the attributes of each of the multiple training images included in the training image set selected by the selection means, and a second determination means for determining whether the index calculated by the second calculation means is less than a threshold value.
[0137] (Supplementary Note 8) The training image selection device according to Supplementary Note 7, wherein, when the second determination means determines that the index is not less than a threshold value, the selection means selects a training image set different from the already selected training image set.
[0138] (Supplementary Note 9) A training image selection method, comprising: a training image selection device training a first machine learning model, the first machine learning model having a first layer group that receives an image as input and generates features of the image, by contrastive learning using a training image set that is a plurality of training images; training a second machine learning model, the second machine learning model having the first layer group and a second layer group connected to the first layer group and that receives features of the image as input and classifies the image, the second machine learning model having the first machine learning model as a pre-training model, using the training image set; calculating a first similarity that is the similarity between parameters in the first layer group before learning by training the second machine learning model and parameters in the first layer group after learning by training the second machine learning model, the parameters being learned by the contrastive learning; and determining whether or not an inappropriate training image is included in the training image set based on the first similarity.
[0139] (Supplementary Note 10) A program that causes a computer to function as a training image selection device, the program causing the computer to function as: first learning means that trains a first machine learning model by contrastive learning using a training image set that is a plurality of training images, the first machine learning model having a first layer group that receives an image as input and generates features of the image; second learning means that trains a second machine learning model using the training image set, the second machine learning model having the first layer group and a second layer group that is connected to the first layer group and receives features of the image as input and classifies the image, the second machine learning model having the first machine learning model as a pre-training model; first calculation means that calculates a first similarity that is a similarity between parameters in the first layer group before training by the second learning means and parameters in the first layer group after training by the second learning means; and first determination means that determine whether or not an inappropriate training image is included in the training image set, based on the first similarity.
[0140] [Additional Note 3] Part or all of the above-described embodiment can also be expressed as follows.
[0141] a first learning process for training a first machine learning model, the first machine learning model having a first layer group that receives an image as input and generates features of the image, by contrastive learning using a learning image set that is a plurality of learning images; a second learning process for training a second machine learning model, the second machine learning model having the first machine learning model as a pre-learning model, using the learning image set, the second machine learning model having the first layer group and a second layer group connected to the first layer group and that receives features of the image as input and classifies the image; a first calculation process for calculating a first similarity, the first similarity being the similarity between parameters in the first layer group before learning by the second learning process and parameters in the first layer group after learning by the second learning process; and a first determination process for determining whether an inappropriate training image is included in the training image set, based on the first similarity.
[0142] The learning image selection device may further include a memory that stores a program for causing the processor to execute the first learning process, the second learning process, the first calculation process, and the first determination process. The program may also be recorded on a computer-readable, non-transitory, tangible recording medium.
[0143] 1, 2, 2A, 3, 3A Learning image selection device 11 First learning unit 12 Second learning unit 13 First calculation unit 14 First judgment unit 22 Selection unit 32 Second calculation unit 33 Second judgment unit M1, M2 Encoder (feature analysis model)
Claims
1. A training image selection device comprising: a first learning means for training a first machine learning model, the first machine learning model having a first layer group that receives an image as input and generates features of the image, by contrastive learning using a training image set that is a plurality of training images; the first layer group; and a second layer group connected to the first layer group and that receives features of the image as input and classifies the image. A second learning means for training a second machine learning model, the second machine learning model having the first machine learning model as a pre-training model, using the training image set; a first calculation means for calculating a first similarity which is the similarity between parameters in the first layer group before training by the second learning means and parameters in the first layer group after training by the second learning means, which are trained by the first learning means; and a first determination means for determining whether or not an inappropriate training image is included in the training image set based on the first similarity.
2. The training image selection device of claim 1, further comprising a selection means for selecting a portion of a plurality of images as the training image set, wherein when the first determination means determines that the inappropriate training image is included in the training image set, the selection means selects a training image set different from the training image set that has already been selected.
3. The learning image selection device according to claim 1 or 2, wherein the first calculation means calculates a second similarity which is a similarity for each layer of the first group of layers provided in the first machine learning model.
4. The training image selection device of claim 3, wherein the first calculation means calculates the first similarity by dividing the sum of the second similarities by the number of layers in the first layer group provided in the first machine learning model, and the first determination means determines that the inappropriate training image is included in the training image set if the first similarity is less than a threshold.
5. The training image selection device of claim 3, wherein the first calculation means calculates the first similarity by dividing the weighted sum obtained by weighting each of the second similarities and adding the weighted sum by the total sum of the weight values, and the first determination means determines that the inappropriate training image is included in the training image set if the first similarity is less than a threshold value.
6. The learning image selection device according to claim 5, wherein the first calculation means increases the weight value of the second similarity of a layer of the first group of layers that is closer to the output of the first machine learning model.
7. The training image selection device of claim 2, further comprising: a second calculation means for calculating an index indicating a bias in attributes of each of a plurality of training images included in the training image set selected by the selection means; and a second judgment means for judging whether the index calculated by the second calculation means is less than a threshold value.
8. The training image selection device according to claim 7, wherein, when the second determination means determines that the index is not less than a threshold value, the selection means selects a training image set different from the training image set already selected.
9. A method for selecting training images, comprising: a training image selection device training a first machine learning model, the first machine learning model having a first layer group that receives an image as input and generates features of the image, by contrastive learning using a training image set that is a plurality of training images; training a second machine learning model, the second machine learning model having the first machine learning model as a pre-training model, using the training image set, the second machine learning model comprising: the first layer group; and a second layer group connected to the first layer group and that receives features of the image as input and classifies the image, the second machine learning model being trained using the training image set; calculating a first similarity that is the similarity between parameters in the first layer group before learning by training the second machine learning model and parameters in the first layer group after learning by training the second machine learning model, the parameters being trained in the learning by contrastive learning; and determining whether or not an inappropriate training image is included in the training image set based on the first similarity.
10. A program causing a computer to function as a training image selection device, the program causing the computer to function as: a first learning means for training a first machine learning model, the first machine learning model having a first layer group that receives an image as input and generates features of the image, by contrast learning using a training image set that is a plurality of training images; a second learning means for training a second machine learning model, the second machine learning model having the first machine learning model as a pre-training model, using the training image set, the second machine learning model comprising the first layer group and a second layer group connected to the first layer group and that receives features of the image as input and classifies the image; a first calculation means for calculating a first similarity which is the similarity between parameters in the first layer group before training by the second learning means and parameters in the first layer group after training by the second learning means, which are trained by the first learning means; and a first determination means for determining whether or not an inappropriate training image is included in the training image set, based on the first similarity.