Model learning device, model learning method, and model learning program
Patent Information
- Application Number
- JP2024563146
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-07-05
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-07-05
Smart Images

Figure 00000011_0000 
Figure 00000011_0001 
Figure 00000012_0000
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a model learning device, a model learning method, and a model learning program. [Background technology]
[0002] In machine learning, when there is little learning data or when a difficult problem setting is made, such as weakly supervised learning, there is a possibility that features that are undesirable for humans will be learned. Although it is possible to judge whether the learned features are appropriate (for example, to check whether undesirable features have been acquired) through visualization by XAI (Explainable AI), it is difficult to provide feedback (for example, transfer learning) so that the learning model does not learn inappropriate features. Therefore, as a method of providing feedback so that the learning model does not learn inappropriate features, a model learning method has been proposed in which attention acquired by learning is re-learned to change the loss function for each data, the learning of features to be acquired is controlled, and transfer learning is repeated until a desirable feature is acquired (for example, see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2022-79331 A Summary of the Invention [Problem to be solved by the invention]
[0004] In the conventional model learning method described above, it is necessary to repeat transfer learning until features desired by humans are acquired, but since transfer learning is newly repeated without memorizing features previously acquired, there is a possibility that features previously learned will be reacquired. For this reason, the conventional model learning method has a problem of being inefficient.
[0005] The present disclosure has been made to solve the above-mentioned conventional problems, and aims to provide a model learning device, a model learning method, and a model learning program that make it possible to improve the efficiency of model learning. [Means for solving the problem]
[0006] The model learning device of the present disclosure is a device that performs transfer learning on a learning model stored in a storage, and is characterized in having: a fixed branch selection unit that selects a fixed branch that is a branch to be excluded from a learning target from a plurality of branches included in the learning model; a computation graph modification unit that changes a computation graph to be used to either a first computation graph using the plurality of branches or a second computation graph that uses a learning target branch obtained by excluding the fixed branch from the plurality of branches; a branch distance calculation unit that calculates a branch-to-branch distance including a distance between features generated by each of the plurality of branches in a state in which the computation graph to be used is changed to the first computation graph; a loss function calculation unit that calculates a sum of losses based on a predetermined loss function and the branch distance; and a branch update that updates weight parameters in the learning target branch based on the sum of losses in a state in which the computation graph to be used is changed to the second computation graph.
[0007] The model learning method disclosed herein is a method executed by a model learning device that performs transfer learning on a learning model stored in a storage, and includes the steps of: selecting a fixed branch, which is a branch to be excluded from a learning target, from a plurality of branches included in the learning model; changing a computation graph to be used to either a first computation graph using the plurality of branches or a second computation graph using a learning target branch obtained by excluding the fixed branch from the plurality of branches; calculating a branch-to-branch distance including a distance between features generated by each of the plurality of branches in a state in which the computation graph to be used is changed to the first computation graph; calculating a sum of losses based on a predetermined loss function and the branch-to-branch distance; and updating a weight parameter in the learning target branch based on the sum of losses in a state in which the computation graph to be used is changed to the second computation graph. Effect of the Invention
[0008] According to the present disclosure, the efficiency of model learning can be improved. [Brief description of the drawings]
[0009] [Figure 1] 1 is a block diagram illustrating a schematic configuration of a model learning device according to a first embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of the model learning device according to the first embodiment. [Diagram 3] 3 is a schematic diagram showing the operation of the model learning device according to the first embodiment. FIG. [Figure 4] FIG. 11 is a schematic diagram showing the operation of a model learning device of a comparative example. [Diagram 5] FIG. 11 is an explanatory diagram showing the operation of the model learning unit during forward propagation. [Figure 6] FIG. 11 is an explanatory diagram showing the operation of the model learning unit during error backpropagation. [Figure 7] 4 is a flowchart showing the operation of the model learning device according to the first embodiment. [Figure 8] 5 is a flowchart showing an operation during model learning of the model learning device according to the first embodiment. [Figure 9] FIG. 11 is a block diagram illustrating a schematic configuration of a model learning device according to a second embodiment. [Figure 10] FIG. 11 is a diagram illustrating an example of a hardware configuration of a model learning device according to a second embodiment. [Figure 11] 10 is a flowchart showing an operation during model learning of the model learning device according to the second embodiment. [Figure 12] FIG. 11 is a block diagram illustrating a schematic configuration of a model learning device according to a third embodiment. [Figure 13] FIG. 11 is a diagram illustrating an example of a hardware configuration of a model learning device according to a third embodiment. [Figure 14] 13 is a flowchart showing an operation during model learning of the model learning device according to the third embodiment. [Figure 15] FIG. 13 is a block diagram illustrating a schematic configuration of a model learning device according to a fourth embodiment. [Figure 16] FIG. 13 is a diagram illustrating an example of a hardware configuration of a model learning device according to a fourth embodiment. [Figure 17] 13 is a flowchart showing an operation during model learning of the model learning device according to the fourth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] A model learning device, a model learning method, and a model learning program according to embodiments will be described below with reference to the drawings. The following embodiments are merely examples, and the embodiments can be appropriately combined and each embodiment can be appropriately modified.
[0011] 1. First embodiment 《1-1》Configuration FIG. 1 is a block diagram showing a schematic configuration of a model learning device 1 according to the first embodiment. The model learning device 1 is a device capable of implementing the model learning method according to the first embodiment, and is, for example, a computer that executes the model learning program according to the first embodiment. The model learning device 1 according to the first embodiment includes a model learning unit 10, a branch visualization unit 15, and a fixed branch selection unit 16. The model learning unit 10 includes an inter-branch distance calculation unit 11, a loss function calculation unit 12, a computation graph change unit 13, and a branch update unit 14. Note that one or both of the fixed branch selection unit 16 and the branch visualization unit 15 may be part of the model learning unit 10.
[0012] 2 is a diagram showing an example of a hardware configuration of the model learning device 1 according to the first embodiment. The model learning device 1 has, for example, a processor 101 such as a CPU (Central Processing Unit), a storage 102 as a storage device, and an interface 103. Each part constituting the model learning device 1 is constituted by, for example, a processing circuit. The processing circuit may be dedicated hardware, or may include a CPU that executes a program (for example, a model learning program) stored in the storage 102. The processor 101 realizes each functional block shown in FIG. 1.
[0013] The storage 102 has, for example, a semiconductor memory such as a RAM (Random Access Memory) and a non-volatile storage device such as a HDD (Hard Disk Drive). The model learning device 1 may also be a mixture of a component made up of a processing circuit and a component made up of a processor. A part or all of the model learning device 1 may be a server computer on a network. The model learning program is provided by downloading via a network, or by a storage medium such as a USB memory for storing information.
[0014] In the example of FIG. 2, the storage 102 stores a learning model and learning data used for learning. The learning model has a plurality of attention branches (also simply called "branches"). The interface 103 has an input unit 104, which is a user interface where user operations are performed, and a display unit 105, such as a liquid crystal display, that presents information. Note that the hardware configuration of FIG. 2 is an example and can be changed.
[0015] 1 and 2, the fixed branch selection unit 16 of the model learning device 1 performs transfer learning on a learning model stored in the storage 102. From a plurality of branches included in the learning model, a fixed branch that is a branch to be excluded from learning targets (i.e., a branch in which a weight parameter is fixed) is selected. Identifying information of the fixed branch is input from, for example, an input unit 104 where an operation for inputting an operation by a user is performed.
[0016] In addition, the computation graph modification unit 13 of the model learning unit 10 changes the computation graph to be used to either a first computation graph using multiple branches included in the learning model (i.e., a computation graph corresponding to the configuration during forward propagation shown in Figure 5 described below) or a second computation graph using a learning target branch obtained by excluding a fixed branch from multiple branches included in the learning model (i.e., a computation graph corresponding to the configuration during error backpropagation shown in Figure 6 described below).
[0017] The branch distance calculation unit 11 of the model learning unit 10 calculates the branch distance including the distance between the features generated by each of the multiple branches included in the learning model, in a state where the computation graph used in learning is changed to a first computation graph (i.e., a computation graph corresponding to the configuration at the time of forward propagation). The branch distance calculated by the branch distance calculation unit 11 can include, in addition to the distance between the features generated by each of the multiple branches included in the learning model, the distance between each of the features and the feature of a predetermined target branch.
[0018] The loss function calculation unit 12 of the model learning unit 10 calculates the total loss based on a predetermined loss function and the inter-branch distance.
[0019] The branch update unit 14 of the model learning unit 10 updates the weight parameters in the learning target branch based on the sum of losses obtained by the loss function calculation unit 12, with the computation graph used in learning changed to a second computation graph (i.e., a computation graph corresponding to the configuration during error backpropagation).
[0020] The branch visualization unit 15 visualizes features generated by each of a plurality of branches included in the learning model. Specifically, the branch visualization unit 15 transmits the features to the display unit 105 and causes the display unit 105 to display the features.
[0021] 3 is a schematic diagram showing the operation of the model learning device 1 according to the first embodiment. The model learning device 1 divides features that have been learned in the past into branches A1, A2, ..., A n (n is a positive integer), the acquired branches are stored, and new features are learned by learning the distance between branches during transfer learning. In this way, by learning the distance between branches during transfer learning, and repeating the learning of the distance between branches (e.g., learning the distance between the acquired branch and the target branch, learning the distance between the acquired branches) every time a new feature is acquired, appropriate features (i.e., branch A that overlaps with the target branch B0 in the feature space in Figure 3) can be learned. n ) can be obtained. In this case, the model learning device 1 stores the acquisition branches of undesirable features acquired in the past and takes into account the distance between the branches, making it possible to perform learning that makes the most of the results of past transfer learning by feeding back the results.
[0022] 4 is a schematic diagram showing the operation of the model learning device of the comparative example. During transfer learning, the model learning device of the comparative example applies learned features to branches C1, C2, ..., C n(n is a positive integer) and repeat the transfer learning to obtain the appropriate features (i.e., the branch C that overlaps with the target branch B0 in the feature space of Figure 4). n ) is obtained. In this case, the previously acquired branch for undesirable features is not memorized, so the previously acquired branch for undesirable features cannot be utilized. In this case, the branch for undesirable features may be re-learned, resulting in inefficient model learning.
[0023] FIG. 5 is an explanatory diagram showing the operation of the model learning unit 10 during forward propagation. FIG. 5 shows a case where the computation graph modification unit 13 of the model learning unit 10 sets the branch #1 and the branch #2 as learning target branches and sets the branch #3 as a fixed branch (i.e., a branch excluded from the learning target by the fixed branch selection unit 16). During forward propagation, features #1, #2, and #3 are generated from the branches #1, #2, and #3, respectively. However, since the feature #3 is an inappropriate feature that is undesirable for humans and an instruction to exclude it from the learning target is input to the fixed branch selection unit 16, the computation graph modification unit 13 inputs the features #1 and #2 to the header and does not input the feature #3 to the header. On the other hand, since the feature #1 and the feature #2 should be learned to be features distant from the feature #3, the computation graph modification unit 13 inputs all features including the features #1, #2, and #3 to the branch distance calculation unit 11.
[0024] FIG. 6 is an explanatory diagram showing the operation of the model learning unit 10 during error backpropagation. FIG. 6 shows a case where the computation graph modification unit 13 of the model learning unit 10 sets branch #1 and branch #2 as learning target branches and sets branch #3 as a fixed branch. During error backpropagation, an instruction to set branch #3 as a fixed branch is input to the fixed branch selection unit 16, so the computation graph modification unit 13 removes the input and output edges of branch #3 from the computation graph to remove branch #3 from the learning target. Therefore, the computation graph modification unit 13 inputs features #1 and #2 to the branch distance calculation unit 11, but does not input feature #3 generated by the fixed branch to the branch distance calculation unit 11.
[0025] As described above, in the embodiment 1, the computation graph during forward propagation is different from the computation graph during error back propagation. That is, when calculating the total loss based on the branch distance, the features #1 to #3 of all branches including the feature #3 of the fixed branch are output to the branch distance calculation unit 11, but when updating the branches, the features #1 to #2 of the branches excluding the feature #3 of the fixed branch are output.
[0026] 《1-2》Operation 7 is a flowchart showing the operation during model learning of the model learning device 1 according to embodiment 1. In embodiment 1, first, the model learning unit 10 learns a model using learning data (for example, learning data in the storage 102 in FIG. 2) (step S1).
[0027] Next, the branch visualization unit 15 visualizes the characteristics acquired by each branch by the XAI, and displays the visualization result on a display unit (for example, the display unit 105 in FIG. 2) so that a human can interpret it (step S2). At this time, the visualization result may be displayed by a BI (Business Intelligence) tool or a dedicated GUI (Graphical User Interface). There are local explanations (for example, explanations for each data) and global explanations (for example, explanations of the behavior of a model) as XAI. In the conventional technology, local explanations (attention) are used as XAI, but in the first embodiment, either local explanations or global explanations may be used as XAI, or local explanations and global explanations may be used together. When it is necessary to exclude a branch, a user who has seen the visualization result inputs specific information (i.e., branch ID) of the fixed branch that is to be excluded to the fixed branch selection unit 16, for example, using the input unit 104 in FIG. 2.
[0028] Based on the results of human interpretation of the features learned by each branch, if the features learned by each branch fall under a predetermined condition, i.e., the first or second case below, the fixed branch selection unit 16 fixes the weight parameter of the branch that has acquired the feature corresponding to the first or second case, and excludes the feature corresponding to the first or second case from the learning targets.
[0029] The first case is when the features generated by the branch through learning are features that are undesirable to humans. Since the features in the first case are not used for inference after learning, in order to avoid re-learning the features in the first case, the branch that learns the features in the first case is set as a fixed branch and excluded from the learning targets.
[0030] In the second case, the features generated by the branch through learning are desirable features for humans. The features in the second case are used for inference after learning, but in order to retain the features in the second case even when transfer learning is performed, the branch that learns the features in the second case is treated as a fixed branch and is excluded from the learning target.
[0031] The model learning unit 10 determines whether re-learning is necessary, and if re-learning is necessary (YES in step S3), returns the processing to step S1, and if re-learning is not necessary (NO in step S3), ends the processing.
[0032] Fig. 8 is a flowchart showing the operation of model learning device 1 according to embodiment 1 during model learning (i.e., details of step S1 in Fig. 7). First, model learning unit 10 determines whether the learning to be performed is the first learning, and if it is the first learning (YES in step S101), the process proceeds to step S106, where loss function calculation unit 12 calculates the loss function. When performing the second or subsequent model learning (NO in step S101), model learning unit 10 proceeds to step S102.
[0033] In step S102, the model learning unit 10 determines whether there is a fixed branch with fixed weight parameters as a branch to be used to acquire data features, and if there is a fixed branch (YES in step S102), the process proceeds from step S102 to step S103 to select a fixed branch, and if there is no fixed branch (NO in step S102), the process proceeds from step S102 to step S104.
[0034] In step S104, the computation graph modification unit 13 of the model learning unit 10 modifies the computation graph for forward propagation. During forward propagation, as shown in Fig. 5, the edge from the input to branch #3, which is a fixed branch, is set to valid, the edge from branch #3 to the branch distance calculation unit 11 is set to valid, and the edge from branch #3 to the header is set to invalid.
[0035] In step S105, the branch distance calculation unit 11 of the model learning unit 10 calculates the distance between the branches. In this case, the branch distance calculation unit 11 calculates the branch distance in order to learn a branch different from a branch learned in the past. The branch distance calculation unit 11 calculates the following two types of distances, a first distance and a second distance, as the branch distance.
[0036] The first distance is the distance between the feature generated by the learning branch and the feature generated by the fixed branch. In order to learn a feature different from the features learned in the past, it is desirable for the first distance to be far.
[0037] The second distance is the distance between the features generated by the training branch. In order to prevent the features simultaneously acquired by multiple training branches (i.e., newly acquired branches) from being similar, it is desirable that the second distance is large. Note that when the number of training branches is one, the second distance does not exist.
[0038] Here, the distance may be freely defined by the user. For example, in ArcFace, a deep metric learning method, the cosine similarity between features mapped onto a hypersphere is used as the distance.
[0039] In the next step S106, the loss function calculation unit 12 of the model learning unit 10 calculates the total loss using a predetermined loss function. The loss function is defined as the sum of the task-dependent loss and the distance loss dependent on the branch distance calculation unit 11 (i.e., the total loss), and is expressed by the following formula (1). (total loss) = (task-dependent loss) + (β * distance loss) (1) Since the number of terms included in the distance loss changes depending on the result of the selection by the fixed branch selection unit 16, the hyperparameter β is adjusted (or normalized) depending on the balance with the task-dependent loss.
[0040] When the number of fixed branches is a and the number of branches to be learned is b, there are the number of terms shown in the following equation (2).
[0041]
number
[0042] In equation (2), the first term represents the number of combinations between the fixed branch and the training branch, and the second term represents the number of combinations between the training branches.
[0043] In step S107, the computation graph modification unit 13 of the model learning unit 10 modifies the computation graph for error backpropagation. During error backpropagation, as shown in Fig. 6, the edge from the input to branch #3, which is a fixed branch, is set to be invalid, the edge from branch #3 to the inter-branch distance calculation unit 11 is set to be invalid, and the edge from branch #3 to the header is set to be invalid.
[0044] In step S108, branch update unit 14 of model learning unit 10 updates the weight parameters in the learning target branch.
[0045] 1-3 Effects According to the model learning device 1 of the first embodiment, since learning can be started with a small number of branches, overlearning can be suppressed and learning can be accelerated.
[0046] 2. Second embodiment FIG. 9 is a block diagram showing a schematic configuration of a model learning device 2 according to the second embodiment. In FIG. 9, components identical to or corresponding to those shown in FIG. 1 are given the same reference numerals as those shown in FIG. 1. FIG. 10 is a diagram showing an example of a hardware configuration of the model learning device 2 according to the second embodiment. In FIG. 10, components identical to or corresponding to those shown in FIG. 2 are given the same reference numerals as those shown in FIG. 2. The model learning device 2 is a device capable of implementing the model learning method according to the second embodiment, and is, for example, a computer that executes the model learning program according to the second embodiment.
[0047] The model learning device 2 according to the second embodiment differs from the model learning device 1 according to the first embodiment in that it is provided with a branch adding unit 21 and in that the computation graph modification unit 13a modifies the computation graph based on the branches to which the branches provided by the branch adding unit 21 are added.
[0048] Generally, learning performed with many branches prepared in a learning model uses many weight parameters, which is difficult, and may result in overlearning or long processing time. Therefore, in the early stages of learning, it may be desirable to perform learning using a small number of learnable branches and increase the number of branches to be learned by adding branches to the learning model later by the branch adding unit 21. As transfer learning is repeated, the number of branches to be learned used by the fixed branch selecting unit 16 decreases, so in the model learning device 2 according to the second embodiment, branches to be learned are added later as necessary by the branch adding unit 21.
[0049] Fig. 11 is a flowchart showing the operation of model learning device 2 according to embodiment 2 during model learning. In Fig. 11, steps that are the same as or correspond to steps shown in Fig. 8 are assigned the same reference numerals as those shown in Fig. 8. The operation of model learning device 2 during model learning differs from the operation of model learning device 1 according to embodiment 1 in that it further includes step S201 of determining whether or not to add a branch by branch adding unit 21, and step S202 of branch adding unit 21 adding a branch to model learning unit 20 if a branch is to be added, and in that it executes the processes of steps S104 to S107 using the learning target branch including the branch added by model learning unit 20.
[0050] According to the model learning device 2 according to the second embodiment, since learning can be started with a small number of branches, overlearning can be suppressed and learning can be accelerated.
[0051] Furthermore, according to the model learning device 2 of embodiment 2, the branch visualization unit 15 visualizes the features acquired by each branch by XAI, and the user can add branches with appropriately initialized weight parameters to the learning model via the branch addition unit 21, thereby improving the accuracy of learning.
[0052] Other than the above, the second embodiment is the same as the first embodiment.
[0053] 3. Third embodiment FIG. 12 is a block diagram showing a schematic configuration of a model learning device 3 according to embodiment 3. In FIG. 12, components identical to or corresponding to those shown in FIG. 1 are given the same reference numerals as those shown in FIG. 1. FIG. 13 is a diagram showing an example of a hardware configuration of the model learning device 3 according to embodiment 3. In FIG. 13, components identical to or corresponding to those shown in FIG. 2 are given the same reference numerals as those shown in FIG. 2. The model learning device 3 is a device capable of implementing the model learning method according to embodiment 3, and is, for example, a computer that executes the model learning program according to embodiment 3.
[0054] The model learning device 3 according to the third embodiment differs from the model learning device 1 according to the first embodiment in that it includes a branch deletion unit 31 and that the computation graph modification unit 13b modifies the computation graph based on the branch obtained by deleting the branch instructed by the branch deletion unit 31.
[0055] In general, at the maintenance operation stage after model learning is completed, branches that have learned inappropriate features remain in memory. When there are many branches, the time required for inference using the model becomes longer, and the amount of memory used by the branches increases. Therefore, the model learning device 3 according to the third embodiment includes a branch deletion unit 31, and is configured to be able to delete a branch selected by the user based on the model definition and weight parameters. Note that, since re-learning may be required, a backup of the deleted branch may be made when deleting.
[0056] Fig. 14 is a flowchart showing the operation of the model learning device 3 according to the third embodiment during model learning. In Fig. 14, steps that are the same as or correspond to steps shown in Fig. 8 are assigned the same reference numerals as those shown in Fig. 8. The operation of the model learning device 3 during model learning differs from the operation of the model learning device 1 according to the first embodiment during model learning in that it further includes step S301 of determining whether or not to perform branch deletion by the branch deleting unit 31, and step S302 of the branch deleting unit 31 deleting a branch from the model learning unit 30 if branch deletion is to be performed, and in that the model learning unit 30 executes the processes of steps S104 to S107 using the learning target branches excluding the deleted branch.
[0057] According to the model learning device 3 of embodiment 3, the branch visualization unit 15 visualizes the features acquired by each branch by XAI, and the user can delete branches through the branch deletion unit 31, thereby improving the accuracy of learning, thereby realizing a reduction in memory usage and faster inference.
[0058] Other than the above, the third embodiment is the same as the first embodiment. In addition, the branch deleting unit 31 in the third embodiment can also be applied to the model learning device 2 in the second embodiment.
[0059] 4. Fourth embodiment FIG. 15 is a block diagram showing a schematic configuration of a model learning device 4 according to embodiment 4. In FIG. 15, components identical to or corresponding to those shown in FIG. 1 are given the same reference numerals as those shown in FIG. 1. FIG. 16 is a diagram showing an example of a hardware configuration of the model learning device 4 according to embodiment 4. In FIG. 16, components identical to or corresponding to those shown in FIG. 2 are given the same reference numerals as those shown in FIG. 2. The model learning device 4 is a device capable of implementing the model learning method according to embodiment 4, and is, for example, a computer that executes the model learning program according to embodiment 4.
[0060] The model learning device 4 according to the fourth embodiment differs from the model learning device 1 according to the first embodiment in that it has a learning target branch selection unit 41, attention correct answer data 42, and attention loss calculation unit 43, that the calculation graph modification unit 13c modifies the calculation graph based on the learning target branch instructed by the learning target branch selection unit 41, and that the loss function calculation unit 12c modifies the calculation of the loss function based on the attention loss calculation unit 43.
[0061] In the fourth embodiment, the loss function is changed for each data by directly correcting the attention acquired by learning, and the learning of the features to be acquired is controlled. For example, the model learning device 4 selects a specific learning target branch, and learns to generate attention close to the attention corrected by a human for that branch. In this way, when features that are likely to be mistaken are known in advance, the number of required transfer learnings can be reduced by deliberately preparing and learning such attention data. In addition, when features that are likely to be mistaken are deliberately learned, the reliability of the learning model can be improved by inferring not to use those features.
[0062] 15 and 16, the learning branch selection unit 41 selects a learning branch to be learned using attention correct answer data. At this time, the type of attention stored in the attention correct answer data 42 and the branch selected by the learning branch selection unit 41 may have any of a one-to-one correspondence relationship, a one-to-many correspondence relationship, and a many-to-many correspondence relationship. For example, in the case of data for human detection, the attention correct answer data 42 may be heat map data in which a heat map is applied to the upper body, heat map data in which a heat map is applied to the lower body, and heat map data in which a heat map is applied to the entire body. For example, in the case of human detection, the learning branch selection unit 41 may select a branch that recognizes the head, or a branch that recognizes a part other than the head position (for example, the upper body, the lower body).
[0063] Fig. 17 is a flowchart showing the operation of model learning device 4 according to embodiment 4 during model learning. In Fig. 17, steps that are the same as or correspond to steps shown in Fig. 8 are denoted by the same reference numerals as those shown in Fig. 8.
[0064] In the second or subsequent learning, if there is attention correct answer data 42 (YES in step S401), the model learning device 4 causes the learning target branch selection unit 41 to select a learning target branch (step S402), causes the attention loss calculation unit 43 to calculate the loss for the selected learning target branch (step S403), and then proceeds to step S106. This is different from the operation during model learning of the model learning device 1 of embodiment 1.
[0065] In the fourth embodiment, since the loss due to attention is used only for learning a specific branch, it is necessary to perform the error backpropagation multiple times as described below, and the computation graph used at that time is also stored for each of the multiple error backpropagations. Also, in the fourth embodiment, the loss dependent on the task and the distance loss dependent on the inter-branch distance calculation unit 11 can be backpropagated to all learning branches, and the attention loss can also be backpropagated only to the specific branch selected.
[0066] According to the model learning device 4 of embodiment 4, the branch visualization unit 15 visualizes the features acquired by each branch by XAI, and the user can delete branches through the branch deletion unit 31, thereby improving the accuracy of learning, thereby realizing a reduction in memory usage and faster inference.
[0067] In addition, for data whose features are known to be easily confused in advance, the number of transfer learning rounds required can be reduced by preparing attention data and training the model. Also, when training features that are easily confused, the reliability of the learning model can be improved by inferring that these features are not used.
[0068] Other than the above, the fourth embodiment is the same as the first embodiment. The learning branch selection unit 41, the attention correct answer data 42, and the attention loss calculation unit 43 in the fourth embodiment can also be applied to the model learning device 2 in the second or third embodiment. [Explanation of symbols]
[0069] 1 to 4 model learning device, 10, 20, 30, 40 model learning unit, 11 branch distance calculation unit, 12, 12c loss function calculation unit, 13, 13a, 13b, 13c computation graph modification unit, 14 branch update unit, 15 branch visualization unit, 16 fixed branch selection unit, 21 branch addition unit, 31 branch deletion unit, 41 learning target branch selection unit, 42 attention correct answer data, 43 attention loss calculation unit, 101, 101a, 101b, 101c processor, 102 storage, 103 interface.
Claims
1. A model learning device for performing transfer learning on a learning model stored in a storage, comprising: a fixed branch selection unit that selects a fixed branch, which is a branch to be excluded from the learning target, from a plurality of branches included in the learning model; a computation graph change unit that changes the used computation graph to either a first computation graph using the plurality of branches or a second computation graph using learning target branches obtained by excluding the fixed branch from the plurality of branches; a branch distance calculation unit that calculates a branch distance including the distance between features generated by each of the plurality of branches in a state where the used computation graph is changed to the first computation graph; a loss function calculation unit that calculates the total loss based on a predetermined loss function and the branch distance; a branch update unit that updates the weight parameters in the learning target branches based on the total loss in a state where the used computation graph is changed to the second computation graph; A model learning device characterized by comprising the above.
2. Further comprising a branch visualization unit that visualizes the features generated by each of the plurality of branches. The model learning device according to claim 1, characterized by the above.
3. The branch distance includes, in addition to the distance between features generated by each of the plurality of branches, the distance between each of the features and the features of a predetermined target branch. The model learning device according to claim 1 or 2, characterized by the above.
4. Further comprising an input unit for performing an operation for inputting specific information of the fixed branch. The model learning device according to claim 1 or 2, characterized by the above.
5. Further comprising a branch addition unit that adds a new branch to the learning model. The model learning device according to claim 1 or 2, characterized by the above.
6. Further comprising a branch deletion unit that deletes a fixed branch from the learning model. The model learning device according to claim 1 or 2, characterized by the above.
7. Further comprising a branch selection unit that selects the learning target branches from the plurality of branches based on pre-created correct data. The model learning device according to claim 1 or 2, characterized by the above.
8. A model learning method executed by a model learning device for performing transfer learning on a learning model stored in a storage, comprising: Selecting a fixed branch that is a branch to be excluded from the learning target from among a plurality of branches included in the learning model; Changing the calculation graph to be used to either a first calculation graph using the plurality of branches or a second calculation graph using the learning target branches obtained by excluding the fixed branch from the plurality of branches; Calculating an inter-branch distance including the distances between the features generated by each of the plurality of branches with the calculation graph being changed to the first calculation graph; Calculating the total loss based on a predetermined loss function and the inter-branch distance; Updating the weight parameters in the learning target branches based on the total loss with the calculation graph being changed to the second calculation graph; A model learning method characterized by comprising the above.
9. A model learning program for causing a computer to perform transfer learning on a learning model stored in a storage, the program causing the computer to: Selecting a fixed branch that is a branch to be excluded from the learning target from among a plurality of branches included in the learning model; Changing the calculation graph to be used to either a first calculation graph using the plurality of branches or a second calculation graph using the learning target branches obtained by excluding the fixed branch from the plurality of branches; Calculating an inter-branch distance including the distances between the features generated by each of the plurality of branches with the calculation graph being changed to the first calculation graph; Calculating the total loss based on a predetermined loss function and the inter-branch distance; Updating the weight parameters in the learning target branches based on the total loss with the calculation graph being changed to the second calculation graph; A model learning program characterized by causing the above to be executed.