Pre-learning device, method, and program

The pre-training device and method enhance the accuracy of neural network pre-training by maintaining local image structure and deforming global structure, effectively generating usable unlabeled data even with limited labeled data, addressing the challenge of task-type variability.

JP2025091191APending Publication Date: 2025-06-18KK TOSHIBA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023206304
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Existing pre-training methods for neural networks face challenges in improving accuracy regardless of the task type, especially when a large amount of data cannot be collected due to measurement costs or when unlabeled data cannot be used.

Method used

A pre-training device and method that includes a conversion unit to maintain the local structure and deform the global structure of input images, a data augmentation unit to generate augmented images, and processing units to calculate feature amounts and update feature extractor parameters, enabling effective representation learning.

Benefits of technology

This approach allows for high-precision pre-training even with a small number of labeled data, generating a large amount of unlabeled data that can be used for pre-training, thereby improving the accuracy of pre-training regardless of the task type.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025091191000001_ABST
    Figure 2025091191000001_ABST
Patent Text Reader

Abstract

To provide a pre-learning device, a method, and a program that improve accuracy of pre-learning regardless of a type of task.SOLUTION: A pre-learning device 100 includes a conversion unit, a data extension unit, a first processing unit, a second processing unit, and an update unit. The conversion unit converts an input image to generate a converted image. The data extension unit generates, on the basis of a method different from a method for generating the converted image, a first extended image and a second extended image from the converted image. The first processing unit inputs the first extended image to a first feature extractor to calculate a first feature quantity. The second processing unit inputs the second extended image to a second feature extractor to calculate a second feature quantity. The update unit updates, on the basis of the first feature quantity and the second feature quantity, a parameter of at least one of the first feature extractor and the second feature extractor.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a pre-training device, method, and program.

Background Art

[0002] The effectiveness of neural networks for tasks targeting images such as image recognition has been confirmed. In the initial stage of introducing such neural networks or in the initial consideration phase, only a small number of labeled data may be available as learning data from the perspective of teaching costs. As an effective method in this case, a technique is known in which pre-training is performed by unsupervised learning using a large amount of unlabeled data, and then re-learning (fine-tuning) is performed using a small number of labeled data. However, depending on the type of task, there may be cases where a large amount of data cannot be collected due to reasons such as measurement costs, or where unlabeled data cannot be used for pre-training.

[0003] In addition, a method has been proposed in which a general-purpose model is learned using artificial images generated based on random numbers, taking advantage of the property of learning by focusing on the local structure of images. However, since different artificial images are suitable depending on the type of task, a general-purpose pre-training method that can be used regardless of the type of task is required.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Non-Patent Documents

[0005]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0006] The problem to be solved by the present invention is to provide a pre-training device, method, and program that can improve the accuracy of pre-training regardless of the type of task.

Means for Solving the Problems

[0007] To solve such problems, the pre-training device according to the embodiment includes a conversion unit, a data augmentation unit, a first processing unit, a second processing unit, and an update unit. The conversion unit converts an input image to generate a converted image. The data augmentation unit generates a first augmented image and a second augmented image from the converted image based on a method different from the method for generating the converted image. The first processing unit inputs the first augmented image to a first feature extractor to calculate a first feature amount. The second processing unit inputs the second augmented image to a second feature extractor to calculate a second feature amount. The update unit updates at least one parameter of the first feature extractor and the second feature extractor based on the first feature amount and the second feature amount.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0009] Hereinafter, embodiments of a pre-training device, a pre-training method, and a pre-training program will be described in detail with reference to the drawings. In the following description, components having substantially the same functions and configurations are denoted by the same reference numerals, and duplicate descriptions are made only when necessary.

[0010] (First Embodiment) FIG. 1 is a diagram showing an example of the configuration of a pre-training device 100 according to the first embodiment. The pre-training device 100 is a learning device that learns a neural network based on labeled training data. The neural network is a machine learning model that receives an input of image data and outputs a feature vector of the image data. The pre-training device 100 is connected to a database or the like that records training data via a network or the like. As the training data, for example, multi-dimensional image data is used. Hereinafter, the image data is simply referred to as an image. Also, an image input to the neural network is referred to as an input image.

[0011] The pre-training device 100 learns a neural network including a feature extractor by optimizing the feature extractor through representation learning using the feature extractor. As the representation learning, for example, any contrastive learning method can be used.

[0012] The learned neural network is used for various tasks targeting images, such as image recognition, object detection in images, image segmentation, anomaly detection, and the like.

[0013] The pre-training device 100 is a computer having a processing circuit 11, a storage device 12, an input device 13, a communication device 14, and a display device 15. Data communication among the processing circuit 11, the storage device 12, the input device 13, the communication device 14, and the display device 15 is performed via a bus. The input device 13 and the display device 15 may not be provided.

[0014] The processing circuit 11 includes a processor such as a CPU (Central Processing Unit) and a memory such as a RAM (Random Access Memory). The processing circuit 11 includes a conversion unit 111, a data expansion unit 112, a first processing unit 113, a second processing unit 114, and an update unit 115. The processing circuit 11 realizes the conversion function, the data expansion function, the first processing function, the second processing function, and the update function by the respective units by executing a pre-training program.

[0015] The storage device 12 is composed of a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), an integrated circuit storage device, or the like. The storage device 12 stores a pre-training program and the like.

[0016] The pre-training program is stored in a non-transitory computer-readable recording medium such as the storage device 12. The pre-training program may be implemented as a single program describing all the functions of the respective units, or may be implemented as a plurality of modules divided into several functional units. Further, the respective units may be implemented by an integrated circuit such as an application specific integrated circuit (ASIC). In this case, it may be implemented by a single integrated circuit or may be individually implemented by a plurality of integrated circuits.

[0017] The input device 13 inputs various commands from the operator. As the input device 13, a keyboard, a mouse, various switches, a touch pad, a touch panel display, etc. can be used. The output signal from the input device 13 is supplied to the processing circuit 11.

[0018] The communication device 14 is an interface for performing data communication between the pre-learning device 100 and an external device connected via a network. For example, the communication device 14 performs data communication with a database that stores input images used as labeled learning data.

[0019] The display device 15 displays various information. As the display device 15, a CRT (Cathode-Ray Tube) display, a liquid crystal display, an organic EL (Electro Luminescence) display, an LED (Light-Emitting Diode) display, a plasma display, or any other display known in the art can be appropriately used. Further, the display device 15 may be a projector.

[0020] Next, the functions executed by each part of the processing circuit 11 will be described in detail. The conversion unit 111 converts the input image to generate a converted image. At this time, the conversion unit 111 generates a converted image in which the local structure of the input image is maintained and the global structure of the input image is deformed. The local structure is an element related to the characteristics of the subject of interest, such as, for example, pixel values, colors, and the shapes of elements included in the image. Images with a changed local structure may not be appropriate as training data depending on the type of task. On the other hand, images with an unchanged local structure can be used as training data regardless of the type of task. The global structure is an element related to the position information and arrangement of pixels. Even if the global structure changes, it can be used as training data regardless of the type of task. As the conversion method, for example, a method of applying perturbation to a partial image and replacing it, a method using an affine transformation, a method of rearranging partial images, etc. can be used. Any other conversion method may be used as long as it is a conversion method that deforms the global structure while maintaining the local structure of the input image. The conversion unit 111 may be called a duplication unit that duplicates the input image.

[0021] As a method of applying perturbation to a partial image and replacing it, for example, a method using GPNN or SinGAN as described in Non-Patent Document 1 and Non-Patent Document 2 can be used. In this case, the conversion unit 111 generates a plurality of adjusted images with the resolution of one input image converted, generates partial images by cutting out the same part for each adjusted image, and generates perturbed partial images with perturbation added to each partial image. Then, the perturbed partial images are synthesized by weighted addition of a plurality of perturbed partial images with different resolutions. And by replacing the part corresponding to the perturbed partial image in the original input image with the synthesized perturbed partial image, a converted image is generated. Thereby, a converted image in which the local structure of the original input image is maintained and the global structure of the original input image is deformed can be generated.

[0022] The affine transformation includes, for example, geometric transformation and rotation. In the method using an affine transformation, a plurality of converted images are generated by applying an affine transformation to the input image. A plurality of converted images may be generated from one input image by performing a plurality of affine transformations using random numbers.

[0023] In the method of rearranging partial images, one input image is divided into a plurality of partial images, and a converted image is generated by rearranging and combining the plurality of divided images. Then, by changing the order of the partial images, a plurality of converted images can be generated from one input image. Further, the number of converted images to be generated may be increased by performing left-right flipping, up-down flipping, or rotation processing on one or more partial images and then rearranging them.

[0024] Alternatively, a converted image may be generated from the input image using a pre-trained model based on the target image or an image similar to the target image. For example, in the case of a GAN that performs image conversion based on one or more learned parameters as described in Non-Patent Document 3, after calculating an input vector for generating the input image by the generator, a converted image can also be generated by inputting the input vector with noise superimposed into the generator.

[0025] Note that by changing the parameters in one of the above methods, a plurality of converted images may be generated from one input image, or a plurality of the above methods may be combined to generate a plurality of converted images from one input image.

[0026] The data augmentation unit 112, the first processing unit 113, the second processing unit 114, and the update unit 115 perform learning of a neural network by representation learning. As the representation learning, for example, contrastive learning can be used. Hereinafter, the case of performing contrastive learning using two feature extractors will be described as an example. The data augmentation unit 112, the first processing unit 113, the second processing unit 114, and the update unit 115 optimize at least one of the two feature extractors (the first feature extractor and the second feature extractor) by performing contrastive learning using two converted images generated from one converted image for all the converted images. That is, at least one of the first feature extractor and the second feature extractor is used as a learning target and is output as a learned machine learning model when the learning is completed.

[0027] The first feature extractor and the second feature extractor receive an input of an image and output a feature vector as a feature amount of the image. As the first feature extractor and the second feature extractor, for example, a general neural network used for representation learning can be used. Also, the parameters of the first feature extractor and the parameters of the second feature extractor may be the same or different.

[0028] The data augmentation unit 112 generates two transformed images for contrastive learning using each transformed image. At this time, the data augmentation unit 112 generates a transformed image using a transformation method different from the transformation method by the transformation unit 111. Hereinafter, the two generated transformed images are referred to as a first augmented image and a second augmented image. In representation learning such as contrastive learning, when generating a plurality of transformed images based on one image, a transformation method in which the local structure changes is used. The data augmentation unit 112 performs transformation by randomly combining processes such as color transformation, rigid body transformation, filtering, masking, and image cropping.

[0029] The first processing unit 113 calculates a first feature amount that is a feature amount of the first augmented image based on the first augmented image. At this time, the first processing unit 113 inputs the first augmented image to the first feature extractor. The first feature extractor receives an input of the first augmented image, calculates a first feature amount from the first augmented image based on one or more parameters, and outputs the calculated first feature amount.

[0030] The second processing unit 114 calculates a second feature amount that is a feature amount of the second augmented image based on the second augmented image. At this time, the second processing unit 114 inputs the second augmented image to the second feature extractor. The second feature extractor receives an input of the second augmented image, calculates a second feature amount from the second augmented image based on one or more parameters, and outputs the calculated second feature amount.

[0031] The update unit 115 updates at least one parameter of the first feature extractor and the second feature extractor based on the first feature amount and the second feature amount. For example, the update unit 115 updates at least one parameter of the first feature extractor and the second feature extractor so that two feature amounts (the first feature amount and the second feature amount) calculated based on the same converted image approach each other, thereby optimizing the feature extractor to be learned.

[0032] After the data augmentation unit 112, the first processing unit 113, the second processing unit 114, and the update unit 115 execute representation learning using all the converted images, they output at least one of the optimized first feature extractor and the second feature extractor as a learned neural network. The learned neural network is output to a device that executes a task targeting an image as a neural network that receives an input of an image and outputs a feature amount of the image.

[0033] Note that both the first feature extractor and the second feature extractor may be targets of learning, and a neural network including both feature extractors may be output as a learned neural network.

[0034] (Pre-training process) Next, the operation of the pre-training process executed by the pre-training device 100 will be described. The pre-training device 100 starts the pre-training process based on having acquired a small number of input images as labeled learning data. FIG. 2 is a flowchart showing an example of the procedure of the pre-training process. Further, FIG. 3 is a diagram showing an example of the data flow in the pre-training process. Here, the case where the first feature extractor is the target of learning will be described as an example. Note that the processing procedures in each process described below are merely examples, and each process can be changed as appropriate as much as possible. Also, regarding the processing procedures described below, steps can be omitted, replaced, and added as appropriate according to the embodiment.

[0035] (Step S101) In the pre-training process, first, the conversion unit 111 converts one input image to generate a converted image. FIG. 4 is a diagram showing the flow of the process of converting the input image. In one process, the conversion unit 111 converts one input image and generates a plurality of converted images in which the local structure is maintained and the global structure is deformed. For example, the conversion unit 111 generates N converted images B11 - B1N from the input image A1. When converting the input image, methods such as perturbing and replacing the above-mentioned partial images, using an affine transformation method, and rearranging the partial images can be used.

[0036] (Step S102) Next, the data augmentation unit 112 converts the converted image to generate a first augmented image and a second augmented image. FIG. 5 is a diagram showing the flow of the process related to the representation learning after step S102. For each converted image generated in the process of step S101, the conversion unit 111 generates one first augmented image and one second augmented image. For example, the conversion unit 111 generates the first augmented images C111 - C1N1 and the second augmented images C112 - C1N2 from the converted images B11 - B1N. When generating the first augmented image and the second augmented image, for example, a method combining color conversion, rigid body transformation, filtering process, masking process, and image cropping process can be used.

[0037] (Step S103) Next, the first processing unit 113 inputs the first augmented image to the first feature extractor and acquires the feature vector output from the first feature extractor as the first feature quantity.

[0038] (Step S104) Similarly, the second processing unit 114 inputs the second augmented image to the second feature extractor and acquires the feature vector output from the second feature extractor as the second feature quantity.

[0039] (Step S105) Next, the update unit 115 adjusts and updates the parameters of the first feature extractor so that the first feature quantity and the second feature quantity approach each other.

[0040] (Step S106) Next, the update unit 115 determines whether to finish adjusting the parameters of the first feature extractor. If the adjustment of the feature extractor is not finished (Step S106 - No), the process returns to Step S101. After that, for the input images not used in learning, the processes of Steps S101 - S105 are repeated to adjust the first feature extractor using all the input images.

[0041] When the adjustment of the first feature extractor is finished for all the input images (Step S106 - Yes), the pre - learning device 100 finishes the pre - learning process and outputs the first feature extractor with adjusted parameters as a learned neural network to an external device or the like.

[0042] Note that in the process of Step S105, instead of the first feature extractor, the second feature extractor may be optimized, or both the first feature extractor and the second feature extractor may be optimized.

[0043] Hereinafter, the effects of the pre - learning device 100 according to this embodiment will be described.

[0044] The pre - learning device 100 according to this embodiment includes a conversion unit 111, a data augmentation unit 112, a first processing unit 113, a second processing unit 114, and an update unit 115. The conversion unit 111 converts an input image to generate a converted image. At this time, the conversion unit 111 generates, as the converted image, an image in which the local structure of the input image is maintained and the global structure of the input image is deformed. As a method for generating the converted image, for example, a method of adding perturbations to and replacing sub - images, a method using affine transformation, a method of rearranging sub - images, etc. can be used.

[0045] The data augmentation unit 112 generates a first augmented image and a second augmented image from the transformed image based on a method different from the method for generating the transformed image. For example, the data augmentation unit 112 executes one or more of color conversion, rigid body conversion, filtering, image masking, and image cropping to generate the first augmented image and the second augmented image. The first processing unit 113 inputs the first augmented image to the first feature extractor to calculate a first feature amount, and the second processing unit 114 inputs the second augmented image to the second feature extractor to calculate a second feature amount. The update unit 115 updates at least one parameter of the first feature extractor and the second feature extractor based on the first feature amount and the second feature amount. At this time, the update unit 115 updates at least one parameter of the first feature extractor and the second feature extractor so that the first feature amount and the second feature amount output based on the same transformed image approach each other.

[0046] With the above configuration, according to the pre-training device 100 according to the present embodiment, when performing pre-training of a neural network that executes a task for an image by representation learning, even when only a small number of labeled data are available, by converting the labeled data by a method different from the conversion method during representation learning, a large amount of unlabeled data required for pre-training can be prepared. Thereby, high-precision pre-training becomes possible.

[0047] In addition, the pre-training device 100 according to the present embodiment can convert learning data with the local structure maintained by converting the input image using a method in which the local structure of the input image is maintained and the global structure of the input image is deformed. An image with an unchanged local structure can be used as learning data for learning the features of the target image regardless of the type of task. For example, even if an image that does not actually exist as a converted image is generated, since the local structure that contributes to the task targeting the image is maintained, it can be used as unlabeled data. In this way, by performing data augmentation using a conversion method that maintains the local structure, a large amount of unlabeled data can be prepared without designing or selecting data augmentation according to the type of task. Therefore, even when there is little learning data, the learning data that can be used for pre-training can be increased, and the accuracy of pre-training can be improved.

[0048] In addition, by training a neural network using the above pre-training method, a trained neural network capable of accurately performing a task targeting an image can be generated.

[0049] (Second Embodiment) The second embodiment will be described. This embodiment is a modification of the configuration of the first embodiment as follows. Descriptions of the same configurations, operations, and effects as those of the first embodiment will be omitted. In this embodiment, expression learning is performed by excluding converted images that are not suitable as learning data for pre-training.

[0050] FIG. 6 is a diagram showing a configuration example of the pre-training device 100 according to the present embodiment. As shown in FIG. 13, the processing circuit 11 further includes a determination unit 116. The processing circuit 11 further realizes a determination function by executing a pre-training program.

[0051] The determination unit 116 selects a conversion image suitable for learning by determining whether to use the generated conversion image for learning. For example, the determination unit 116 determines whether to use each conversion image for learning by using the error between the conversion image and the input image, the attributes of the conversion image and the input image, or the statistical value of the conversion image. In the present embodiment, the determination unit 116 calculates the error between the input image and the conversion image, determines that a conversion image with a large error is not suitable for learning, and selects a conversion image with a small error as an image suitable for learning.

[0052] (Pre-training process) Next, the operation of the pre-training process executed by the pre-training device 100 according to the present embodiment will be described. FIG. 7 is a flowchart showing an example of the procedure of the pre-training process. FIG. 8 is a diagram showing an example of the data flow in the pre-training process. Since the processes in steps S201, S204 - S208 in FIG. 7 are the same as the processes in steps S101, S102 - S106 in FIG. 2, detailed descriptions thereof will be omitted.

[0053] (Step S202) When a conversion image is generated based on one input image in the process of step S201, the determination unit 116 calculates the error between the generated conversion image and the original input image. As the error, for example, the mean squared error (MSE) or the mean absolute error (MAE) can be used. Alternatively, as the error, the difference in statistical values or the difference in feature amounts calculated using a feature extractor may be used. As the statistical value, for example, the average value, variance, standard deviation, median, mode, etc. of each pixel value can be used. The error may be called a difference value.

[0054] (Step S203) Next, the determination unit 116 determines whether to use the conversion image for subsequent learning processing based on the error from the original input image. At this time, the determination unit 116 determines whether the error between the input image and the target conversion image is less than or equal to a threshold value.

[0055] When the error from the original input image is equal to or less than the threshold (step S203 - Yes), the determination unit 116 determines that the converted image is suitable for learning, and proceeds to steps S204 - S207 to perform contrast learning using the converted image. On the other hand, when the error from the original input image is greater than the threshold (step S203 - No), the determination unit 116 determines that the converted image is not suitable for learning because it is too different from the original input image. In this case, the contrast learning using the converted image is not executed, and the process returns to step S201 to generate a converted image of the next input image.

[0056] Note that a converted image with too small an error from the input image may also be determined as an image not suitable for learning. In this case, the determination unit 116 determines a converted image with an error within a predetermined range from the original input image as an image suitable for learning.

[0057] Also, when generating a plurality of converted images for one input image, the processes of steps S202 - S207 are repeatedly executed for each converted image generated from the same input image.

[0058] The pre - learning device 100 repeatedly executes the processes of steps S201 - S207 until the adjustment of the parameters by contrast learning using all the input images is completed (step S208 - No), and performs contrast learning using only the converted images determined to be suitable for learning. When the processes of steps S201 - S207 for all the input images are completed (step S208 - Yes), the pre - learning device 100 determines that the adjustment of the parameters is completed and ends the pre - learning process.

[0059] The pre-training device 100 according to this embodiment further includes a determination unit 116 that determines whether to use the converted image for learning. The determination unit 116 can calculate the error between the input image and the converted image, and determine whether to use the converted image for learning based on the error. With the above configuration, a converted image suitable for learning can be selected based on the error from the input image, and pre-training can be further improved by performing pre-training only using the selected converted image. For example, by excluding converted images with too large a difference from the input image and performing pre-training, the learning accuracy can be improved.

[0060] (First Modification Example) Instead of the error between the input image and the converted image, the comparison result of the attributes of the input image and the converted image may be used. In this case, the determination unit 116 estimates the respective attributes of the input image and the converted image, and determines whether to use the converted image by determining whether the estimated attributes match.

[0061] FIG. 9 is a flowchart showing an example of the procedure of the pre-training process in this modification. The processes of step S301, step S304 - step S308 in FIG. 9 are the same as the processes of S201, step S204 - step S208 in FIG. 7, and thus detailed description is omitted.

[0062] (Step S302) In this modification, the determination unit 116 estimates the attributes for each of the input image and the converted image. As the attributes, for example, labels such as class information and type information used in the downstream task can be used. In this case, the determination unit 116 estimates the label of the converted image using a specified feature extractor for estimating the label. At this time, the label preset for the input image may be used as the attribute of the input image. Also, as the attribute, cluster information such as the classification result by unsupervised classification processing may be used. In this case, the determination unit 116 performs unsupervised classification processing on each of the input image and the converted image.

[0063] (Step S303) Next, the determination unit 116 compares the attributes of the original input image and the converted image, and determines whether to use the converted image in subsequent learning processing. At this time, the determination unit 116 determines whether the attributes of the converted image match those of the input image.

[0064] When the attributes of the converted image match those of the original input image (step S303 - Yes), the determination unit 116 determines that the converted image is suitable for learning, and proceeds to steps S304 - S307 to perform contrast learning using the converted image. On the other hand, when the attributes of the converted image do not match those of the original input image (step S303 - No), the determination unit 116 determines that the converted image is too different from the original input image and is not suitable for learning. In this case, the contrast learning using the converted image is not executed, and the process returns to step S301 to generate a converted image of the next input image.

[0065] Also in this modified example, by selecting a converted image suitable for learning and performing learning using only the selected converted image, the accuracy of pre - learning can be further improved.

[0066] (Second Modified Example) A converted image suitable for learning may be selected using the statistical value of the converted image. In this case, the determination unit 116 uses the statistical value of the pixel values of the converted image to determine whether to use the converted image for representation learning. As the statistical value, for example, the average value, variance, standard deviation, median, mode, etc. of each pixel value can be used. The determination unit 116 calculates the statistical value of the converted image and selects a converted image suitable for learning by determining whether the statistical value satisfies a predetermined condition. Then, by performing learning using only the selected converted image, the accuracy of pre - learning can be further improved. For example, the determination unit 116 selects a converted image whose error between the statistical value of the input image and the statistical value of the converted image is below a threshold as a converted image suitable for learning.

[0067] Thus, according to any of the above - described embodiments, it is possible to provide a pre - learning device, a pre - learning method, and a pre - learning program that can improve the accuracy of pre - learning regardless of the type of task.

[0068] Note that the present invention is not limited to the above-described embodiments as they are, and at the implementation stage, the components can be modified and embodied without departing from the gist thereof. Further, various inventions can be formed by appropriately combining a plurality of components disclosed in the above-described embodiments. For example, some components may be deleted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

Description of Reference Numerals

[0069] 100... pre-training device, 11... processing circuit, 12... storage device, 13... input device, 14... communication device, 15... display device, 111... conversion unit, 112... data augmentation unit, 113... first processing unit, 114... second processing unit, 115... update unit, 116... determination unit.

Claims

1. A conversion unit that converts an input image to generate a converted image, A data augmentation unit that generates a first augmented image and a second augmented image from the converted image based on a method different from the method of generating the converted image, A first processing unit that inputs the first augmented image to a first feature extractor to calculate a first feature amount, A second processing unit that inputs the second augmented image to a second feature extractor to calculate a second feature amount, An update unit that updates at least one parameter of the first feature extractor and the second feature extractor based on the first feature amount and the second feature amount, A pre-training device comprising:

2. The conversion unit generates the converted image as an image in which the local structure of the input image is maintained and the global structure of the input image is deformed. The pre-training device according to claim 1.

3. The conversion unit converts the input image into images of two or more resolutions, generates a perturbed partial image obtained by adding a perturbation to a partial image cut out from the input image of each resolution, generates an added partial image by weighted addition of the perturbed partial images having different resolutions, and generates the converted image by replacing the partial image at the corresponding position of the input image with the added partial image. The pre-training device according to claim 2.

4. The conversion unit generates the converted image by performing an affine transformation on the input image. The pre-training device according to claim 2.

5. The conversion unit generates the converted image by performing a conversion based on adjusted parameters on the input image. The pre-training device according to claim 2.

6. The data augmentation unit generates the first augmented image and the second augmented image by performing one or more of color conversion, rigid body conversion, filtering, image masking, and image cropping processing. The pre-training device according to claim 1.

7. The pre-training device further includes a determination unit that determines whether to use the converted image for learning. The pre-training device according to claim 1.

8. The determination unit calculates an error between the input image and the converted image, and determines whether to use the converted image for learning based on the error. The pre-training device according to claim 7.

9. The determination unit estimates attributes of the input image and the converted image, and determines whether to use the converted image by determining whether the attributes match. The pre-training device according to claim 7.

10. The determination unit determines whether to use the converted image for learning based on a statistical value of the converted image. The pre-training device according to claim 7.

11. The pre-training device is a learning device that learns a neural network including at least one of the first feature extractor and the second feature extractor. The pre-training device according to claim 1.

12. The data augmentation unit generates an image in which a local structure of the converted image is changed as the first augmented image and the second augmented image. The pre-training device according to claim 2.

13. The conversion unit converts an input image to generate a converted image, The data augmentation unit generates a first augmented image and a second augmented image based on a method different from the method for generating the converted image, The first processing unit inputs the first enlarged image into the first feature extractor to calculate a first feature amount, The second processing unit inputs the second enlarged image into the second feature extractor to calculate a second feature amount, The update unit updates at least one parameter of the first feature extractor and the second feature extractor based on the first feature amount and the second feature amount, A method comprising the above.

14. On a computer, A function of converting an input image to generate a converted image, A function of generating a first enlarged image and a second enlarged image based on a method different from the method for generating the converted image, A function of inputting the first enlarged image into the first feature extractor to calculate a first feature amount, A function of inputting the second enlarged image into the second feature extractor to calculate a second feature amount, A function of updating at least one parameter of the first feature extractor and the second feature extractor based on the first feature amount and the second feature amount, A program for realizing the above.

Citation Information

Patent Citations

  • Supervised contrastive learning using multiple positive examples

    JP2023523726A