Model learning method and program, and model learning device
By selectively freezing decoder layers during neural network training and using separate data sets with varying objectives, the method enhances training efficiency and accuracy, addressing inefficiencies in existing methods.
Patent Information
- Application Number
- PCT/JP2025/008692
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-14
- Filing Date
- 2025-03-10
- Publication Date
- 2025-09-18
AI Technical Summary
Existing neural network training methods are inefficient when using data sets with different purposes, leading to incomplete training of encoder and decoder units and increased learning costs and time.
A method involving a first training step to update weights of both encoder and decoder layers using a main data set, followed by a second step where only the encoder weights are updated while freezing the decoder, using a separate data set with different objectives, and optionally a third step to further refine the encoder.
This approach reduces learning costs and time, improves accuracy, and minimizes overfitting by selectively freezing decoder layers during training, allowing efficient adaptation to different data sets with varying objectives.
Smart Images

Figure JP2025008692_18092025_PF_FP_ABST
Abstract
Description
Model learning method and program, and model learning device
[0001] The present invention relates to a model learning method and program, and a model learning device, and more particularly to a technique for learning a model configured by a neural network.
[0002] Two-stage AI (Artificial Intelligence) is known, in which the output of a previous model is input to a subsequent model.
[0003] Patent Document 1 describes a technique in which, in a neural network that supplies the output of a first neural network to a second neural network, the weights of the first neural network are frozen and the second neural network is trained.
[0004] Patent Document 2 describes a technology in which an RGB (Red Green Blue) image converted by an image transformation model is passed to an LR (Lip Reading) model, and the LR model is used to infer the content of the speech shown in the RGB image, in which the image transformation model is trained while the LR model is fixed.
[0005] Patent Document 3 describes a technology in which, when training an RPN (Region Proposal Network), a convolutional layer that extracts features, and a classifier, the training results of the convolutional layer and the classifier are fixed and only the RPN is retrained.
[0006] JP-T-2023-508235 A JP-A-2020-030458 A JP-A-2018-128897 A
[0007] Mingxing Tan, Quoc V. Le,"EfficientNetV2: Smaller Models and Faster Training"<URL:https: / / arxiv.org / pdf / 2104.00298.pdf>
[0008] One embodiment of the technology of the present disclosure provides a model learning method, a program, and a model learning device for efficiently learning a model configured by a neural network.
[0009] In order to achieve the above object, a model training method according to a first aspect of the present disclosure is a model training method for a model configured by a neural network consisting of multiple layers, using a first training data set including multiple pairs of first input data and first correct answer data, and a second training data set including multiple pairs of second input data and second correct answer data, wherein the multiple layers of the neural network are configured as an earlier layer and a later layer that receives output from the earlier layer as input and outputs a final result, the model training method including: a first training step for updating weights of the earlier layer and weights of the later layer of the neural network using the first training data set so that the model outputs first correct answer data when it receives the first input data; and a second training step for updating only the weights of the earlier layer of the neural network without updating the weights of the later layer, after the first training step, using a second training data set so that the model outputs second correct answer data when it receives second input data.
[0010] The model training method according to the second aspect of the present disclosure is preferably the model training method according to the first aspect, in which the first training step and the second training step are repeated.
[0011] The model training method according to the third aspect of the present disclosure is preferably the model training method according to the first or second aspect, and further includes, after the second training step, a third training data set including a plurality of sets of third input data and third supervised data, updating weights in a previous layer and weights in a subsequent layer of the neural network so that the model outputs third supervised data when the model receives the third input data.
[0012] In the model training method according to the fourth aspect of the present disclosure, in the model training method according to the second or third aspect, it is preferable that the first training step varies the set of first input data and first supervised data for at least a portion of the first training data set for each repetition.
[0013] In the model training method according to the fifth aspect of the present disclosure, in the model training method according to any one of the second to fourth aspects, it is preferable that the second training step varies the set of second input data and second correct answer data for at least a portion of the second training data set for each repetition.
[0014] In the model training method according to the sixth aspect of the present disclosure, in the model training method according to any one of the second to fifth aspects, it is preferable that the first training step relatively increases the difficulty of the first training data set with each iteration.
[0015] In the model training method according to the seventh aspect of the present disclosure, in the model training method according to any one of the second to sixth aspects, it is preferable that the second training step relatively increases the difficulty of the second training data set with each iteration.
[0016] The model training method according to the eighth aspect of the present disclosure is preferably a model training method according to any one of the second to seventh aspects, which includes a modification step of modifying the position between an earlier layer and a later layer of multiple layers of the neural network according to the repetition.
[0017] In a model training method according to a ninth aspect of the present disclosure, in the model training method according to the eighth aspect, it is preferable that the changing step changes a position between an earlier layer and a later layer of the multiple layers of the neural network to a position where the number of later layers increases.
[0018] In a model training method according to a tenth aspect of the present disclosure, in the model training method according to the eighth aspect, it is preferable that the changing step changes a position between an earlier layer and a later layer of multiple layers of the neural network to a position where the number of later layers is reduced.
[0019] The model training method according to an eleventh aspect of the present disclosure is preferably the model training method according to any one of the second to tenth aspects, and includes an evaluation step of evaluating the output accuracy of the model according to each iteration, and evaluating the first training data set and the second training data set based on the output accuracy.
[0020] A model training method according to a twelfth aspect of the present disclosure is preferably a model training method according to any one of the first to eleventh aspects, in which the preceding layer constitutes an encoder and the subsequent layer constitutes a decoder.
[0021] In the model training method relating to the thirteenth aspect of the present disclosure, in the model training method relating to any one of the first to twelfth aspects, it is preferable that the first correct answer data and the second correct answer data are different types of data.
[0022] A model training method according to a fourteenth aspect of the present disclosure is preferably a model training method according to any one of the first to thirteenth aspects, wherein the first correct answer data is a distance image and the second correct answer data is a segmentation image.
[0023] In the model training method according to the fifteenth aspect of the present disclosure, in the model training method according to any one of the third to fourteenth aspects, it is preferable that the first correct answer data and the third correct answer data are each the same type of data.
[0024] In order to achieve the above object, a program according to a sixteenth aspect of the present disclosure is a program that causes a computer to execute the model learning method according to any one of aspects 1 to 15. The present disclosure also includes a non-transitory computer-readable storage medium that stores the program according to the sixteenth aspect.
[0025] In order to achieve the above object, a model learning device according to a seventeenth aspect of the present disclosure is a model learning device configured with a neural network consisting of multiple layers, using a first learning data set including multiple pairs of first input data and first correct answer data, and a second learning data set including multiple pairs of second input data and second correct answer data, wherein the multiple layers of the neural network are configured as an earlier layer and a later layer that receives output from the earlier layer as input and outputs a final result, and the model learning device is equipped with one or more processors and one or more memories that store programs to be executed by the one or more processors, wherein the one or more processors execute instructions of the program to perform a first learning process using the first learning data set to update weights in the earlier layer and weights in the later layer of the neural network so that the model outputs first correct answer data when it inputs the first input data, and after performing the first learning, the model learning device performs a second learning process using a second learning data set to update only the weights in the earlier layer of the neural network without updating the weights in the later layer of the neural network so that the model outputs second correct answer data when it inputs second input data.
[0026] FIG. 1 is a diagram for explaining general transfer learning. FIG. 2 is a diagram showing an example of an input image and an output image of a monocular distance map AI. FIG. 3 is a block diagram showing the configuration of a learning device for learning a monocular distance map AI. FIG. 4 is a diagram showing an example of the configuration of a learning model according to a first embodiment. FIG. 5 is a flowchart showing the processing of a learning method for a learning model by a learning device. FIG. 6 is a diagram for explaining AI parameters optimized by the learning method for a learning model. FIG. 7 is a diagram showing the learning time required to train a learning model. FIG. 8 is a table showing an example of a learning combination of a first learning data set and a second learning data set for different purposes on the same network. FIG. 9 is a flowchart showing the processing of a learning method for a learning model according to a third embodiment. FIG. 10 shows the progress of performance of a trained model through iterative learning. FIG. 11 is a flowchart showing the processing of a learning method for a learning model according to a fourth embodiment. FIG. 12 is a diagram showing images of different difficulty levels. FIG. 13 is a diagram showing an example of the configuration of a learning model according to a seventh embodiment. FIG. 14 is a diagram for explaining the relationship between nodes and edges. FIG. 15 is a diagram showing the learning range and freeze range of a learning model.
[0027] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0028] <General Transfer Learning> Fig. 1 is a diagram for explaining general transfer learning. Here, an example is described in which a model for generating a segmentation image is generated by transfer learning. A learning model 200, which is an AI network shown in Fig. 1, includes an encoder unit 202, a decoder unit 204, and a decoder unit 206.
[0029] The encoder unit 202 is a neural network including a plurality of intermediate layers 202-1, 202-2, .... Each of the intermediate layers 202-1, 202-2, ... has parameters such as weights. The encoder unit 202 encodes an input image into features.
[0030] The decoder unit 204 receives the output of the encoder unit 202. The decoder unit 204 is a neural network including multiple intermediate layers 204-1, 204-2, .... Each of the intermediate layers 204-1, 204-2, ... has parameters such as weights. The decoder unit 204 decodes the feature values output from the encoder unit 202 and classifies the object class of the input image.
[0031] The decoder unit 206 receives the output of the encoder unit 202. The decoder unit 206 is a neural network including multiple intermediate layers 206-1, 206-2, .... Each of the intermediate layers 206-1, 206-2, ... has parameters such as weights. The decoder unit 206 decodes the feature values output from the encoder unit 202 and generates a segmentation image of the input image.
[0032] In the learning model 200, there is a difference in the output shapes of the decoder unit 204 and the decoder unit 206.
[0033] When training the learning model 200, the first step is to pre-train the encoder unit 202 and the decoder unit 204 using a large-scale data set. The next step is to perform transfer training on the decoder unit 206 using a relatively small number of data sets that are training data for the intended input-output relationship. Note that when training the decoder unit 206, the parameters of the encoder unit 202 are fixed (frozen).
[0034] In this way, the training data set used to train the decoder unit 204 is often a data set for a purpose other than the original purpose, such as class classification, etc. For this reason, it only contributes to the training of the encoder unit 202, which is the feature extraction unit in the first half of the training model 200, and the decoder unit 206, which is the second half of the training model 200, is completely new, resulting in inefficient training.
[0035] Therefore, we propose a learning method, program, and device for a learning model that efficiently learns the desired learning model by designing the output shapes to be the same for a first learning data set and a second learning data set, which have different purposes, without changing the structure of the AI network, and selectively freezing the learning points.
[0036] First Embodiment In a first embodiment, the AI function that is ultimately desired to be generated is a monocular distance map AI. The monocular distance map AI is an AI whose input is a real-life image and whose output is a distance map (an example of a “distance image”).
[0037] Fig. 2 shows an example of an input image and an output image of the monocular distance map AI. Image IM1 shown in Fig. 2 is a real-life image that is an input image of the monocular distance map AI. Image IM1 may be a color image. Image IM2 shown in Fig. 2 is a distance map corresponding to image IM1. The distance map is a two-dimensional image that represents subject distances using brightness values, with subjects that are relatively closer being represented in brighter colors.
[0038] [Learning Device] Fig. 3 is a block diagram showing the configuration of a learning device 10 for learning a monocular distance map AI. The learning device 10 is realized by at least one computer. As shown in Fig. 3, the learning device 10 includes a processor 12, a memory 14, an input device 16, and an output device 18.
[0039] The processor 12 executes instructions stored in the memory 14. The hardware structure of the processor 12 is various processors as shown below. The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various functional units, a GPU (Graphics Processing Unit), which is a processor specialized for image processing, a PLD (Programmable Logic Device), which is a processor whose circuit configuration can be changed after manufacture such as an FPGA (Field Programmable Gate Array), and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor having a circuit configuration designed specifically for executing specific processing.
[0040] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types (e.g., multiple FPGAs, a combination of a CPU and an FPGA, or a combination of a CPU and a GPU). Multiple functional units may also be configured with a single processor. Examples of multiple functional units configured with a single processor include: a first configuration, as typified by a client or server computer, in which a single processor is configured with a combination of one or more CPUs and software, and this processor operates as multiple functional units; and a second configuration, as typified by a SoC (System on Chip), in which a processor is used to realize the functions of an entire system including multiple functional units on a single IC (Integrated Circuit) chip. In this way, the various functional units are configured with one or more of the above-mentioned various processors as a hardware structure.
[0041] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.
[0042] The memory 14 stores instructions to be executed by the processor 12. The memory 14 also stores parameters of the neural network. The memory 14 includes a random access memory (RAM) and a read-only memory (ROM), not shown. The processor 12 uses the RAM as a working area, executes software using various programs and parameters stored in the ROM, and performs various processes of the learning device 10 by using parameters stored in the ROM, etc.
[0043] The input device 16 may be, for example, a keyboard, a mouse, a touch panel, or other pointing device, or a voice input device, or an appropriate combination thereof. A user can use the input device 16 to input various instructions and a learning data set to the learning device 10.
[0044] The output device 18 includes a display device. The display device may be, for example, a liquid crystal display, an organic electroluminescence (OEL) display, a projector, or an appropriate combination of these. Note that the input device 16 and the display device of the output device 18 may be integrated into one device, such as a touch panel.
[0045] The learning device 10 may include a large-capacity storage device for storing a learning dataset. The learning device 10 may also include a reading device for reading data recorded on a computer-readable, non-transitory recording medium, such as a removable storage medium such as a DVD or CD-ROM. The learning device 10 may also include a communication interface for transmitting and receiving data to and from other computers. The learning device 10 may also be used as an image processing device to which monocular distance map AI is applied.
[0046] [Configuration of Learning Model] Fig. 4 is a diagram showing an example of the configuration of a learning model 100 (an example of a "model") according to the first embodiment. The learning model 100 is trained by the learning device 10 to become a monocular distance map AI, which is a trained model. The learning model 100 includes an encoder unit 102 and a decoder unit 104.
[0047] The encoder unit 102 is a neural network that includes a plurality of intermediate layers 102-1, 102-2, ... (examples of "previous layers") as a plurality of layers. Each of the intermediate layers 102-1, 102-2, ... has parameters such as weights. The encoder unit 102 encodes an input image into features.
[0048] The decoder unit 104 receives the output of the encoder unit 102. The decoder unit 104 is a neural network that includes a plurality of intermediate layers 104-1, 104-2, ... (examples of "later layers") as a plurality of layers. Each of the intermediate layers 104-1, 104-2, ... has parameters such as weights. The decoder unit 104 decodes the feature values output from the encoder unit 102 and outputs a segmentation image or a distance map as a final result. The segmentation image and the distance map are images that have different purposes but a common output shape.
[0049] The learning model 100 can be applied to, for example, U-Net, but is not limited to U-Net and can be applied to other networks.
[0050] [Learning Method] Figure 5 is a flowchart showing the processing of the learning method for the learning model by the learning device 10. The learning method for the learning model is realized by the processor 12 executing a learning program stored in the memory 14. The learning program may be provided by a computer-readable non-transitory recording medium. In this case, the learning device 10 may read the learning program from the non-transitory storage medium and store it in the memory 14.
[0051] In step S1 (an example of a "first learning step"), the processor 12 performs a first learning of the learning model 100. The first learning is "main learning," which is the learning of the AI function to be ultimately generated, and in this case, it is the learning of a monocular distance map AI. The first learning is performed using a first learning dataset (an example of a "first learning dataset"), which is the "main learning dataset" used for the main learning. The first learning dataset includes multiple sets of first input data and first correct answer data. For example, the first input data is a real-life image, and the first correct answer data is a distance map.
[0052] In the first learning, the encoder unit 102 and the decoder unit 104 of the learning model 100 are trained. That is, the processor 12 updates the parameters of the intermediate layers 102-1, 102-2, ... of the encoder unit 102 and the parameters of the intermediate layers 104-1, 104-2, ... of the decoder unit 104 so that when first input data that is a real-life image is input, first ground truth data that is a distance map is output.
[0053] In step S2 (an example of a "second learning step"), the processor 12 performs second learning of the learning model 100. The second learning is learning of a target AI function different from the AI function to be ultimately generated, in this case, learning of a segmentation image AI. The second learning is performed using a second learning dataset (an example of a "second learning dataset"), which is a "separate learning dataset." The second learning dataset includes multiple sets of second input data and second correct answer data. For example, the second input data is a CG image, and the second correct answer data is a segmentation image. In this way, the first correct answer data and the second correct answer data are different types of data. Note that here, input data and correct answer data are created and used for learning based on CG images to avoid the influence of erroneous input and misclassification by the user when generating the correct answer data. However, the second input data may be a real-life image, like the first input data.
[0054] The second learning is "freeze learning" in which the decoder unit 104 of the learning model 100 is frozen and the encoder unit 102 is trained. That is, the processor 12 updates only the parameters of the intermediate layers 102-1, 102-2, ..., of the encoder unit 102 without updating the parameters of the intermediate layers 104-1, 104-2, ..., of the decoder unit 104 so that when second input data, which is a real-life image, is input, second ground truth data, which is a segmentation image, is output.
[0055] In step S3 (an example of a "third learning step"), the processor 12 performs third learning of the learning model 100. The third learning is main learning, and is performed using a third learning dataset (an example of a "third learning data set"), which is the main learning dataset. The third learning dataset includes multiple sets of third input data and third correct answer data. The third input data is a real image, and the third correct answer data is a distance map. In this way, the first correct answer data and the third correct answer data are the same type of data.
[0056] In the third learning, the encoder unit 102 and the decoder unit 104 of the learning model 100 are trained. That is, the processor 12 updates the parameters of the intermediate layers 102-1, 102-2, ... of the encoder unit 102 and the parameters of the intermediate layers 104-1, 104-2, ... of the decoder unit 104 so that when third input data that is a real-life image is input, third ground truth data that is a distance map is output.
[0057] This completes the learning method for the learning model, and a monocular distance map AI is generated, which is the learning model 100 in which the parameters of the intermediate layers 102-1, 102-2, ... of the encoder unit 102 and the parameters of the intermediate layers 104-1, 104-2, ... of the decoder unit 104 are optimized. The learning method for the learning model corresponds to a method for producing a learning model. The learning method for the learning model may involve only the first learning and the second learning, without the third learning.
[0058] As described above, according to the training method for the learning model according to the first embodiment, the first training data set is trained first, and then the second training data set for strengthening the encoder unit 102 is trained, during which the decoder unit 104 is frozen.
[0059] 6 is a diagram for explaining AI parameters, which are parameters of the intermediate layers 102-1, 102-2, ... of the encoder unit 102 and parameters of the intermediate layers 104-1, 104-2, ... of the decoder unit 104, optimized by the learning method of the learning model. F6A in FIG. 6 shows the case of the general transfer learning shown in FIG. 1. As shown in F6A, the AI parameters are learned to evenly cover all data in the large-scale learning dataset.
[0060] F6B in Fig. 6 shows the case of the first embodiment. In the learning method of the learning model according to the first embodiment, the decoder unit 104 is frozen during learning of the second learning data set, and therefore, as shown in F6B, parameters that will produce good results even when passed through the frozen decoder unit 104 are selectively learned. This has the effect of making overlearning less likely to occur.
[0061] In general, transfer learning requires a dataset of approximately 10 to 14 million elements for learning. In contrast, in the learning method for a learning model according to the first embodiment, the dataset required for the first learning is approximately 100,000 elements. In addition, the second learning dataset is generally approximately 10 to 100 times the size of the first learning dataset, requiring, for example, approximately 1 million elements.
[0062] FIG. 7 is a diagram showing the learning time required to train a learning model. The horizontal axis of FIG. 7 represents the learning time. F7A in FIG. 7 shows the case of general transfer learning, and F7B shows the case of the first embodiment. The learning time for general transfer learning is, for example, approximately 5,000 hours for pre-learning. On the other hand, in the first embodiment, the number of datasets required for learning is relatively small, so learning is completed in a relatively short time. Furthermore, the learning time for the second learning using the second learning dataset is also shorter than that for general transfer learning.
[0063] In general transfer learning, only the encoder unit 202 (see Figure 1) is used as a common learning area. In this case, even if the output shapes can be matched, such as between a segmentation image and a distance map, if learning is also performed up to the decoder unit 204, the intermediate layers closer to the output of the decoder unit 204 are more susceptible to influence from the correct data due to differences in objectives. For this reason, when two data sets with different objectives are learned, only the parameters of the decoder unit 206, which is closer to the end, are learned, making it difficult for learning of the encoder unit 202, which is the first half, to progress.
[0064] In the first embodiment, the performance of the learning model 100 can be improved by intentionally freezing the decoder unit 104, which is the latter half of the model, and actively training the encoder unit 102.
[0065] According to the first embodiment, even if two data sets with different purposes are used, by using a learning method that can learn the same network structure, it is possible to eliminate the learning costs of hardware equipment and time required to switch between the decoder units 204 and 206 of the learning model 200 in the case of general transfer learning.
[0066] Furthermore, because the decoder unit 104 is frozen after training with the first training data set, training proceeds only for the encoder unit 102 that has characteristics in common with the decoder unit 104 trained in the first training except during the first training. Even in this case, training of the encoder unit 102 that is unrelated to the target AI function, as in the pre-training of general transfer learning, is not required, resulting in less learning loss. This is also effective when there is incomplete overlap between the pre-training data set, which is the data set used to pre-train the learning model 100, and the first training data set.
[0067] Furthermore, since the learning of the encoder unit 102 suitable for the decoder unit 104 of the target AI function is advanced, it is possible to achieve performance in a shorter time than with general transfer learning.
[0068] <Second embodiment> In the first embodiment, an example was described in which the first correct answer data of the first training data set is a monocular distance map and the second correct answer data of the second training data set is a segmentation image, but the training data sets are not limited to this example.
[0069] 8 is a table showing an example of a combination of training data sets with different purposes, such as the training model 100, in the same network. For example, the first training data set may be a data set for boundary segmentation (partial masking), and the second training data set may be a data set for segmented images. Alternatively, the first training data set may be a data set for image noise reduction, and the second training data set may be a data set for segmented images. In this way, the training method for the training model is suitable when the images in the first training data set and the images in the second training data set correspond in size and number of channels.
[0070] In this way, even if distance maps and segmentation images are not relatively close to each other in terms of the segmentation that is the purpose of large-scale datasets, it is possible to train them using the same network by sharing the output shape, which can speed up learning and improve accuracy.
[0071] It should be noted that there are some AI functions, such as image noise reduction, whose purpose is significantly different from segmentation, but even in conventional general transfer learning, there are examples in which the encoder unit 202 is trained by class classification and then style conversion is trained, so even with different combinations of data sets, it is basically effective.
[0072] Third Embodiment A learning method for a learning model according to a third embodiment aims to improve accuracy by iterative learning. Fig. 9 is a flowchart showing the process of the learning method for a learning model according to the third embodiment.
[0073] The processes in steps S1 and S2 are the same as those in the learning method for the learning model according to the first embodiment shown in FIG.
[0074] In the following step S11, the processor 12 determines whether or not to repeat the first learning in step S1 and the second learning in step S2. Step S3 may be determined by the user. If the first learning in step S1 and the second learning in step S2 are to be repeated, the process proceeds to step S1, and the processor 12 performs the same processing. If the second learning is not to be repeated, the process proceeds to step S3.
[0075] In step S3, similar to the learning method of the learning model according to the first embodiment shown in FIG. 5, the processor 12 performs third learning of the learning model 100, and ends the processing of this flowchart.
[0076] In freeze learning, the decoder unit 104 is frozen, and the encoder unit 102 learns the parts that match the decoder unit 104. However, if main learning is performed again after freeze learning, the decoder unit 104 also changes further.
[0077] Here, a difference occurs between the encoder unit 102 that can be trained by the decoder unit 104 in the first main learning and the encoder unit 102 that can be trained by the decoder unit 104 in the second main learning, and the second time is more improved.
[0078] FIG. 10 shows the change in performance of the learning model 100 through iterative learning. The horizontal axis of FIG. 10 represents learning time, with longer learning times indicated further to the right. The vertical axis of FIG. 10 represents performance of the learning model 100, with higher performance indicated further up. The performance of the learning model 100 is obtained by evaluating the output accuracy of the monocular distance map output for the input image. FIG. 10 shows the change in performance when the first learning and the second learning are repeated.
[0079] As shown in Figure 10, the second learning, which is the frozen learning, has lower performance than the first learning, which is the main learning. However, by repeating the first learning and the second learning, the results of both types of learning gradually improve. Accuracy continued to improve even when the first learning and the second learning were repeated more than 10 times. In this way, overfitting is effectively suppressed, and sufficient performance can be achieved even with a relatively small AI model.
[0080] In general transfer learning, large-scale data sets are completely learned in pre-learning, and the model structure for decoding the latter part is different, so it is impossible to repeatedly learn with multiple data sets. Also, if repeated learning is performed using only the same data, accuracy improvement will plateau after about three times. For this reason, the key to improving accuracy is to repeatedly learn using multiple data sets, as in the third embodiment.
[0081] According to the learning method of the learning model of the third embodiment, by repeating main learning and freeze learning using two data sets with different purposes, it is possible to adaptively progress the learning of the encoder section, thereby making it possible to progress the learning at high speed while suppressing overlearning.
[0082] Fourth Embodiment In the third embodiment, an example was shown in which performance improved through iterative learning, but performance may also deteriorate. For example, when iterative learning was attempted using a combination of a color distance map as the first learning data set and a monochrome image segmentation as the second learning data set, the results of both gradually deteriorated. If performance deteriorates during iterative learning in this way, it can be determined that the first learning data set and the second learning data set are an inappropriate combination. A learning method for a learning model according to the fourth embodiment evaluates the performance of the model to determine whether the combination of two learning data sets with different purposes is appropriate.
[0083] FIG. 11 is a flowchart showing the processing of the learning method for the learning model according to the fourth embodiment.
[0084] The processes of steps S1, S2, and S3 are the same as those of the learning method of the learning model according to the first embodiment shown in Fig. 6. The processor 12 performs the first learning, the second learning, and the third learning for the first time.
[0085] In the subsequent step S21 (an example of an "evaluation step"), the processor 12 evaluates the first training data set and the second training data set by evaluating the performance of the learning model 100 that has undergone the first round of third learning.
[0086] In step S22, the performance evaluated in step S21 is compared with the performance stored in memory 14 to determine whether the performance of learning model 100 has deteriorated. Here, since past evaluation results have not yet been stored in memory 14, the process proceeds to step S11.
[0087] In step S11, the processor 12 stores the evaluation result from step S21 in the memory 14, and determines whether or not to repeat the first learning in step S1 and the second learning in step S2.
[0088] If it is determined in step S11 that the process is not to be repeated, the processor 12 ends the process of this flowchart. The learning model 100 at this point becomes the final trained model.
[0089] If it is determined in step S11 that the process is to be repeated, the processor 12 performs the second first learning in step S1. The second first learning is performed on the learning model 100 that has undergone the first second learning, but before the first third learning is performed.
[0090] In steps S2 and S3, the processor 12 performs second and third learning rounds on the learning model 100 that has undergone the second first learning round.
[0091] In the following step S21, the processor 12 evaluates the performance of the learning model 100 that has undergone the second third learning.
[0092] In step S22, the processor 12 compares the performance evaluated in step S21 with the performance stored in the memory 14 to determine whether or not the performance of the learning model 100 has deteriorated. Here, the processor 12 compares the performance of the learning model 100 that has undergone the second round of third learning evaluated in step S21 with the performance of the learning model 100 that has undergone the first round of third learning stored in the memory 14.
[0093] In step S22, if the performance of the learning model 100 has not deteriorated, it is determined that there is no problem with the first learning data set and the second learning data set, and the process proceeds to step S11. The process from step S11 onwards is the same as described above. On the other hand, in step S22, if the performance of the learning model 100 has deteriorated, it is determined that there is a problem with the first learning data set and the second learning data set, and the process proceeds to step S23.
[0094] In step S23, the processor 12 presents a warning to the user and ends the processing of this flowchart. For example, the processor 12 may cause the display device of the output device 18 to display a message indicating that the combination of each of the two training data sets is inappropriate.
[0095] If it is determined in step S22 that the performance of the learning model 100 has not deteriorated and it is determined in step S11 that the process should be repeated, the process returns to step S1. Here, the processor 12 performs the third first learning on the learning model 100 that has undergone the second second learning, but before the second third learning. The subsequent processes are similar.
[0096] As described above, in the fourth embodiment, it is determined that the combination of the training data sets is not good when the performance of the training model 100 deteriorates. The meaning of the combination being not good is that the purposes of the first training data set and the second training data set do not match in the first place.
[0097] In contrast, general transfer learning does not allow for iterative learning, so it is not possible to determine whether a combination of data sets is good or bad.
[0098] According to the fourth embodiment, by using iterative learning to determine the degree of performance improvement in two training data sets with different purposes, it is possible to determine whether the combination of the two training data sets is appropriate.
[0099] In addition, the determination of whether the performance of the learning model 100 has deteriorated in step S22 is not limited to a comparison with the previous evaluation result, and the processor 12 may determine that the performance of the learning model 100 has deteriorated if it determines that the performance has gradually deteriorated based on multiple past evaluation results.
[0100] Fifth Embodiment In the fifth embodiment, images of input data in a training data set are made relatively more difficult as the training progresses.
[0101] FIG. 12 shows images with different levels of difficulty. Image IM11 shown in FIG. 12 is a normal image with a relatively low level of difficulty. Image IM12 shown in FIG. 12 is an image obtained by rotating image IM11 by a predetermined angle (here, 45 degrees counterclockwise) around an axis perpendicular to the image surface. The difficulty level of image IM12 is relatively higher than that of image IM11. Image IM13 shown in FIG. 12 is an image obtained by enlarging image IM11 horizontally by a predetermined magnification factor (e.g., 1.2 times) and rotating it by a predetermined angle (here, 30 degrees counterclockwise). The difficulty level of image IM13 is relatively higher than that of image IM12. In this way, by rotating and transforming an image to add complexity, the difficulty level of the image can be relatively increased.
[0102] The difficulty level of an image may be relatively increased by increasing image noise. Furthermore, in the case of learning using a learning model 100 that outputs a distance map, the difficulty level of an image may be relatively increased by increasing background blur. Furthermore, an image with only two objects, one close and one far away, may be used as an image with a relatively low difficulty level, and an image with a mixture of objects from close to far away may be used as an image with a relatively high difficulty level. An image with a gradual change in distance from close to far, an image with a sudden change in distance between close and far objects, or an image that is completely blurred may be used as an image with a relatively high difficulty level.
[0103] In the first learning, the processor 12 relatively increases the difficulty level of the first learning data set as the iterative learning is repeated, and in the second learning, the processor 12 relatively increases the difficulty level of the second learning data set as the iterative learning is repeated.
[0104] "Relatively increasing the difficulty level of the first training dataset with each iteration of iterative learning" does not necessarily mean relatively increasing the difficulty level of all images in the first training dataset with each iteration of iterative learning, but also includes increasing the proportion of images with relatively high difficulty levels among the multiple images in the first training dataset. In other words, it is sufficient that the difficulty level of the first training dataset increases overall when comparing the difficulty level of the previous first training dataset with the difficulty level of the current first training dataset in the iterative learning iteration. The same applies to the difficulty level of the second training dataset.
[0105] Furthermore, "relatively increasing the difficulty level of the first training data set as the iterative training is repeated" is not limited to the case where the processor 12 increases the difficulty level of the images in the first training data set as the iterative training is repeated, but also includes the case where the processor 12 acquires a first training data set in which the difficulty level of the images has increased as the iterative training is repeated and uses the first training data set for training. The same applies to the difficulty level of the second training data set.
[0106] To prevent overfitting and speed up learning, there is a technique called adaptive regularization, which applies small regularization in early epochs and increases regularization as learning progresses. Non-Patent Document 1 reports performance improvement using adaptive regularization. Note that regularization refers to processing training data such as image rotation and CutMix to prevent overfitting.
[0107] If the performance of the learning model 100 improves through iterative learning, the improvement effect can be obtained more efficiently by learning in stages using learning data sets of different difficulty levels to balance the regularization effect and faster learning.
[0108] Furthermore, even if difficult scenes are found after the training, if it is not easy to increase the number of samples in the training dataset, it is possible to improve the performance by increasing the number of training datasets other than the training dataset. If a tendency to be sensitive to blur or noise is found, one possible method is to use regularization to increase the number of images similar to the difficult scenes, but if the number of images of difficult scenes is suddenly increased, training becomes too difficult and overfitting occurs. Therefore, it is effective to combine gradual enhancement of difficult scenes with iterative training.
[0109] According to the fifth embodiment, by performing iterative learning using data sets with different levels of difficulty or data sets with gradually changing degrees of regularization, it is possible to achieve a better balance between faster learning speed and performance than by performing iterative learning using the same data set.
[0110] Sixth Embodiment The following method is available for combining iterative learning and a plurality of learning data sets.
[0111] For example, a first learning data set AD1, which is the main learning data set, and second learning data sets BD1, BD2, BD3, ..., which are separate learning data sets with increasing difficulty in this order, may be prepared, and learning may be repeated in the following order: first learning using AD1 → second learning using BD1 → first learning using AD1 → second learning using BD2 → first learning using AD1 → second learning using BD3 → ...
[0112] Furthermore, first learning data sets AD1, AD2, AD3, ... of increasing difficulty may be prepared, and learning may be repeated as follows: first learning using AD1 → second learning using BD1 → first learning using AD2 → second learning using BD2 → first learning using AD2 → second learning using BD3 → ...
[0113] Furthermore, the training data sets may be used cyclically in the iterative training. For example, the iterative training may be performed in the following order: first training using AD1 → second training using BD1 → second training using BD2 → first training using AD1 → second training using BD1 → second training using BD2 → .... Alternatively, the iterative training may be performed in the following order: first training using AD1 → second training using BD1 → first training using AD2 → second training using BD1 → first training using AD3 → second training using BD1 → ....
[0114] In the first learning, it is preferable that at least some of the sets of first input data and first correct answer data in the first learning dataset are different for each iteration of the iterative learning. For example, if there are 10,000 sets of data usable for the first learning dataset with set numbers 00001 to 10,000, it is preferable to assign set numbers 00001 to 01,000 to the first learning dataset AD11, set numbers 01,001 to 02,000 to the first learning dataset AD12, and set numbers 02,001 to 03,000 to the first learning dataset AD13, and perform iterative learning in the following order: first learning using AD11 → second learning → first learning using AD12 → second learning → first learning using AD13 → second learning → .... Note that the difficulty levels of the first learning datasets AD11, AD12, AD13, ... may be the same.
[0115] Alternatively, the set numbers 00001 to 01000 may be assigned to the first training data set AD21, the set numbers 00501 to 01500 to the first training data set AD22, and the set numbers 01001 to 02000 to the first training data set AD23, with overlapping assignments, and repeated learning may be performed in the order of first learning using AD21 → second learning → first learning using AD22 → second learning → first learning using AD23 → second learning → .... In this case as well, the difficulty levels of the first training data sets AD21, AD22, AD23, ... may be the same.
[0116] In the second learning, it is preferable that at least some of the sets of second input data and second supervised data in the second learning dataset are different for each iteration of the iterative learning. For example, if there are 100,000 sets of data usable for the second learning dataset, with set numbers 000001 to 100,000, it is preferable to assign set numbers 000001 to 010,000 to the second learning dataset BD11, set numbers 020001 to 030000 to the second learning dataset BD12, and set numbers 030001 to 040000 to the second learning dataset BD13, and perform iterative learning in the following order: first learning → second learning using BD11 → first learning → second learning using BD12 → first learning → second learning using BD13 → .... Note that the difficulty levels of the second learning datasets BD11, BD12, BD13, ... may be the same.
[0117] Alternatively, the set numbers may be assigned with overlap, such as set numbers 000001 to 010000 to the second training data set BD21, set numbers 005001 to 015000 to the second training data set BD22, and set numbers 010001 to 020000 to the second training data set BD23, and training may be performed iteratively in the order of first training → second training using BD21 → first training → second training using BD22 → first training → second training using BD23 → .... In this case, the difficulty levels of the second training data sets BD21, BD22, BD23, ... may be the same.
[0118] In the above case, it is preferable that the first learning use all 10,000 sets of data with set numbers 00001 to 10,000 as the first learning data set. The final performance of the learning model 100 will cover about 80% of the accuracy rate of features common to all data in the first learning data set. On the other hand, the second learning data set used in the second learning is used to improve the encoder unit 102 using features not included in the first learning data set, and therefore the number and type of the second learning data set may be determined appropriately.
[0119] Furthermore, at least one of the learning parameters, such as the batch size, the L1 regularization, and the L2 regularization, may be changed stepwise for each iteration of the iterative learning. The learning parameters may be changed between the first learning and the second learning.
[0120] 13 is a diagram showing an example of the configuration of a learning model 110 according to a seventh embodiment. The learning model 110 is trained by the learning device 10 to become a monocular distance map AI, which is a trained model. The learning model 110 is a neural network including an input layer 112, an intermediate layer 114, and an output layer 116.
[0121] Input data, which is a two-dimensional input image, is input to the input layer 112. The intermediate layer 114 includes multiple layers and converts the input data. The intermediate layer 114 corresponds to the encoder unit 102 and decoder unit 104 of the learning model 100 shown in FIG. 4, and the multiple layers of the intermediate layer 114 correspond to the intermediate layers 102-1, 102-2, ..., intermediate layers 104-1, 104-2, ... of the learning model 100. The output layer 116 outputs an output image. Each layer has a structure in which multiple nodes are connected by edges. In FIG. 13, white circles represent nodes, and the straight lines connecting each white circle represent edges. The number of nodes in the input layer 112 is the same as the number of nodes in the output layer 116.
[0122] Fig. 14 is a diagram for explaining the relationship between nodes and edges. Fig. 14 shows a part of the hidden layer 114 in Fig. 13, and the direction from left to right is the forward propagation direction from the input layer 112 to the output layer 116. Although some parts are omitted in Fig. 14, there are n nodes in layer 1, which are designated X1, X2, ..., x n Similarly, there are n nodes in layer 2, which are g1(z), g2(z), ..., g n The values written at the positions of each edge are the weights (coefficients) of each layer.
[0123] Here, the value g1(z) of the top node in layer 2 and the value g2(z) of the second-highest node can be expressed as the following equations 1 and 2, respectively: g1(z) = X1 × W 11 +X2×W 21 +...+X n ×W n1 ...(Formula 1) g2(z)=X1×W 12 +X2×W 22 +...+X n ×W n2 ...(Formula 2)
[0124] The processor 12 calculates the weights W of each layer so that correct data can be output for the input data. 11 , ..., W n1 , W 12 , ..., W n2 , W 13 , ..., W n3 , ..., W 1n , ..., W nn The weights W of each layer that can output the correct answer data for the input data are updated. 11 , ..., W n1 , W 12 , ..., W n2 , W 13 , ..., W n3 , ..., W 1n , ..., W nn The learning model 110 is a trained model.
[0125] [Learning Method] In the learning method of the learning model according to the seventh embodiment, the freeze range is changed in stages. The processing of the learning method of the learning model will be described with reference to the flowchart shown in FIG.
[0126] In step S1, the processor 12 performs first learning of the learning model 110. As in the previous steps, the first learning is performed using a first learning data set including a plurality of pairs of first input data and first correct answer data. For example, a relatively small first learning data set including 10,000 pairs of real-life images and distance maps is used.
[0127] FIG. 15 is a diagram showing the learning range and the frozen range of the learning model 110. F15A in FIG. 15 shows the learning range 120 in the first learning. In the example shown in F15A, the entire intermediate layer 114 is the learning range 120. The learning range 120 is the range in which the processor 12 updates the layer weights, and is also the non-frozen range. As shown in F15A, in the first learning, all the layer weights in the intermediate layer 114 are updated, and no layer is frozen. That is, the processor 12 updates the weights W of each layer. 11 , ..., W n1 , W 12 , ..., W n2 , W 13 , ..., W n3 , ..., W 1n , ..., W nn The weights of each layer are calculated so that the corresponding distance maps can be output for all input images.
[0128] In step S2, the processor 12 performs second learning of the learning model 110 after the first learning. As in the previous steps, the second learning is performed using a second learning data set including a plurality of sets of second input data and second correct answer data. For example, a relatively large second learning data set including 100,000 sets of CG images and segmentation images is used.
[0129] F15B in FIG. 15 shows the learning range 120 and the frozen range 122 in the second learning. In the example shown in F15B, the learning range 120 is the previous layer of the intermediate layer 114, and the frozen range 122 is the subsequent layer that receives the output from the previous layer as input and outputs the final result. The frozen range 122 is a range in which the processor 12 does not update the layer weights. As shown in F15B, in the second learning, the processor 12 does not update the weights of the subsequent layer of the intermediate layer 114 that is included in the frozen range 122, but only updates the weights of the previous layer of the intermediate layer 114 that is included in the learning range 120.
[0130] In step S11, the processor 12 determines whether or not to repeat the first learning and the second learning. If so, the process proceeds to step S1.
[0131] In step S1 of the second iteration, the processor 12 performs a first learning of the learning model 110. For example, the first learning data set of 10,000 sets of real-life images and distance maps used in the first learning is used again. In the second learning, the processor 12 also updates the weights of all layers of the hidden layer 114, as shown in F15A.
[0132] In step S2 of the second iteration, the processor 12 performs second learning of the learning model 110. For example, the second learning data set of 100,000 sets of CG images and segmentation images used in the first second learning is used again. F15C in FIG. 15 shows the learning range 120 and the frozen range 122 in the second second learning. As shown in F15C, in the second second learning, the processor 12 updates only the weights of the layers preceding the intermediate layer 114 included in the learning range 120, without updating the weights of the layers following the intermediate layer 114 included in the frozen range 122.
[0133] Furthermore, the processor 12 changes the position between the previous layer and the subsequent layer of the intermediate layer 114 of the learning model 110 according to the repetition (an example of a "changing step"). Here, the frozen range 122 shown in F15C is wider than the frozen range 122 shown in F15B, and the learning range 120 shown in F15C is narrower than the learning range 120 shown in F15B. In other words, the processor 12 changes the position between the previous layer and the subsequent layer of the intermediate layer 114 of the learning model 110 to a position where the number of subsequent layers increases.
[0134] In step S11, the processor 12 determines whether or not to repeat the first learning and the second learning. If it is to be repeated, the processor 12 performs the first learning and the second learning for the third time.
[0135] In the third first learning, as shown in F15A, the processor 12 updates the weights of all layers in the intermediate layer 114. In the third second learning, the processor 12 further widens the freeze range 122 and narrows the learning range 120 compared to the second second learning, and updates the weights of the layers preceding the intermediate layer 114. In other words, the processor 12 changes the position between the layers preceding and following the intermediate layer 114 in the learning model 110 to a position where the number of subsequent layers increases, and performs the second learning.
[0136] If it is determined in step S11 that the process will not be repeated, the process proceeds to step S3.
[0137] In step S3, the processor 12 performs third learning of the learning model 110 after the first learning and the second learning, and then ends the processing of this flowchart. The third learning is performed using a third learning data set including a plurality of pairs of third input data and third correct answer data. For example, a relatively small third learning data set including 10,000 pairs of real-life images and distance maps is used.
[0138] 15D shows the learning range 120 in the third learning. As shown in F15D, in the third learning, the processor 12 updates the weights of all layers in the hidden layer 114.
[0139] According to the seventh embodiment, the freeze range is changed stepwise according to the repetitions, so that initially there is no freeze until near the final stage, and as learning progresses, the freeze range is expanded to near the first half, which allows the learning model 110 to learn at a deeper level and improves performance.
[0140] Here, the processor 12 changes the position between the previous layer and the next layer of the intermediate layer 114 of the learning model 110 according to the iteration so that the number of the next layer increases, but the position between the previous layer and the next layer may also be changed so that the number of the next layer decreases. Also, as in the first embodiment, the position between the previous layer and the next layer may be fixed.
[0141] <Others> The technical scope of the present invention is not limited to the scope described in the above embodiments. The configurations of the respective embodiments can be appropriately combined with each other within the scope that does not deviate from the spirit of the present invention.
[0142] 10...Learning device 12...Processor 14...Memory 16...Input device 18...Output device 100...Learning model 102...Encoder unit 102-1...Intermediate layer 102-2...Intermediate layer 104...Decoder unit 104-1...Intermediate layer 104-2...Intermediate layer 110...Learning model 112...Input layer 114...Intermediate layer 116...Output layer 120...Learning range 122...Freeze range 200...Learning model 202...Encoder unit 202-1...Intermediate layer 202-2...Intermediate layer 204...Decoder unit 204-1...Intermediate layer 204-2...Intermediate layer 206...Decoder unit 206-1...Intermediate layer 206-2...Intermediate layer IM1...Image IM2...Image IM11...Image IM12...Image IM13...Image S1 to S3, S11, S21 to S23...Steps of the learning method for the learning model
Claims
1. A method for training a model constituted by a neural network consisting of multiple layers, using a first training data set including multiple pairs of first input data and first correct answer data, and a second training data set including multiple pairs of second input data and second correct answer data, wherein the multiple layers of the neural network are constituted by an earlier layer and a later layer that receives output from the earlier layer as input and outputs a final result, the method comprising: a first training step using the first training data set to update weights of the earlier layer and weights of the later layer of the neural network so that the model outputs the first correct answer data when it receives the first input data; and a second training step after the first training step to update only the weights of the earlier layer of the neural network without updating the weights of the later layer, using the second training data set so that the model outputs the second correct answer data when it receives the second input data.
2. The model training method according to claim 1, wherein the first training step and the second training step are repeated.
3. A method for training a model as described in claim 1, comprising, after the second training step, a third training data set including a plurality of sets of third input data and third correct answer data, updating the weights of the previous layer and the weights of the subsequent layer of the neural network so that the model outputs third correct answer data when third input data is input.
4. The model training method according to claim 2, wherein the first training step varies the set of first input data and first correct answer data for at least a portion of the first training data set for each repetition.
5. The model training method according to claim 2, wherein the second training step varies the set of second input data and second correct answer data for at least a portion of the second training data set for each repetition.
6. The model training method according to claim 2, wherein the first training step relatively increases the difficulty of the first training data set according to the repetition.
7. The model training method according to claim 2, wherein the second training step relatively increases the difficulty of the second training data set according to the repetition.
8. The model training method according to claim 2, further comprising a changing step of changing the positions of the plurality of layers of the neural network between the previous layer and the subsequent layer in accordance with the repetition.
9. The model training method according to claim 8, wherein the changing step changes a position between the previous layer and the subsequent layer of the plurality of layers of the neural network to a position where the number of subsequent layers increases.
10. The model training method according to claim 8, wherein the changing step changes a position between the previous layer and the subsequent layer of the plurality of layers of the neural network to a position that reduces the number of subsequent layers.
11. The model training method according to claim 2, further comprising an evaluation step of evaluating the output accuracy of the model according to the repetitions, and evaluating the first training data set and the second training data set based on the output accuracy.
12. The model training method of claim 1, wherein the preceding layer constitutes an encoder, and the following layer constitutes a decoder.
13. The model training method according to claim 1, wherein the first correct answer data and the second correct answer data are different types of data.
14. The model training method according to claim 13, wherein the first correct answer data is a range image, and the second correct answer data is a segmentation image.
15. The model training method according to claim 3, wherein the first correct answer data and the third correct answer data are the same type of data.
16. A program that causes a computer to execute the model learning method according to any one of claims 1 to 15.
17. A non-transitory computer-readable recording medium on which the program according to claim 16 is recorded.
18. A training device for a model constituted by a neural network consisting of multiple layers, using a first training data set including multiple pairs of first input data and first correct answer data, and a second training data set including multiple pairs of second input data and second correct answer data, wherein the multiple layers of the neural network constitute a previous layer and a subsequent layer that receives output from the previous layer as input and outputs a final result, the device comprising: one or more processors; and one or more memories that store a program to be executed by the one or more processors, wherein the one or more processors execute instructions of the program to perform a first training using the first training data set, updating weights of the previous layer and weights of the subsequent layer of the neural network so that the model outputs the first correct answer data when the first input data is input; a model learning device that, after the first learning is performed, uses the second learning data set to perform second learning, in which the weights of the subsequent layers of the neural network are not updated but only the weights of the previous layers are updated so that the model outputs the second correct answer data when the model receives the second input data.
Citation Information
Patent Citations
Neural network device
JP1993274455A
Super loss: general loss for robust curriculum learning
JP2022063250A
Inspection system, image processing system, image processing method, and program
JP2023067546A
Training method, training device, and program
WO2022239689A1