Machine learning program, machine learning method, and machine learning device
The machine learning program addresses DNN model degradation by identifying and updating specific parameters based on correct and incorrect data predictions, enhancing model stability and accuracy.
Patent Information
- Application Number
- JP2021201561
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2041-12-13
AI Technical Summary
Conventional methods for correcting deep neural network (DNN) models using training data fail to effectively address errors and can lead to model degradation due to the principle of CACE (Changing Anything Changes Everything), resulting in incorrect inferences after retraining.
A machine learning program that identifies correct and incorrect prediction data, selects specific data based on output differences, and updates only the relevant parameters of the DNN model by calculating forward and backward propagation influences to generate a retrained model.
This approach suppresses the occurrence of degradation in retraining DNN models, ensuring accurate and stable model performance.
Smart Images

Figure 0007743778000010 
Figure 0007743778000011 
Figure 0007743778000012
Abstract
Description
[Technical Field]
[0001] The present invention relates to a machine learning program, a machine learning method, and a machine learning device.
[0002] In systems using machine learning models such as deep neural networks (DNNs), if an output that is undesirable to the user occurs, the machine learning model may need to be modified. Hereinafter, DNN machine learning models may be referred to as DNN models.
[0003] For example, in an autonomous driving system using a camera, if a DNN model misrecognizes a road sign, the DNN model is corrected so that it can correctly recognize the road sign.
[0004] DNN models are not built according to specifications, but according to the training data they are given, so you can modify the DNN model by inputting the training data. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 9-128358 Summary of the Invention [Problem to be solved by the invention]
[0006] However, in such conventional methods for correcting DNN models using training data, data is collected to be used for correction, but the collected data does not necessarily allow errors to be corrected.
[0007] Furthermore, retraining a DNN model may result in incorrect inferences after retraining for data that was correctly inferred before retraining, which can lead to degradation. Degradation occurs due to the principle of CACE (Changing Anything Changes Everything).
[0008] In one aspect, the present invention aims to suppress the occurrence of degradation in retraining a machine learning model. [Means for solving the problem]
[0009] In one aspect, the machine learning program may cause a computer to execute the following processes. The processes may include identifying, from a first plurality of data, second plurality of data for which prediction results according to output values of a first machine learning model for each of the first plurality of data are correct and third plurality of data for which prediction results are incorrect. The processes may also include selecting fourth plurality of data from the third plurality of data and fifth plurality of data from the second plurality of data related to the fourth data, based on a difference between the output value of the first machine learning model for data included in the third plurality of data and a correct label value corresponding to data included in the third plurality of data, and a difference between the output value of the first machine learning model for each of the third plurality of data and the output value of the first machine learning model for each of the second plurality of data. Furthermore, the processing may include identifying a second plurality of parameters from among a first plurality of parameters included in the first machine learning model based on numerical values calculated during forward propagation and numerical values calculated during backward propagation when the fourth plurality of data and the fifth plurality of data are input into the first machine learning model, and generating a second machine learning model by updating only the second plurality of parameters from among the first plurality of parameters. [Effects of the Invention]
[0010] In one aspect, the present invention can suppress the occurrence of degradation in retraining a machine learning model. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram illustrating a configuration of an information processing apparatus as an example of an embodiment. [Figure 2] FIG. 2 is a diagram illustrating a process of a model training unit in an information processing device as an example of an embodiment. [Figure 3] 10A and 10B are diagrams illustrating processing by a nearby successful data search unit of an information processing device as an example of an embodiment; [Figure 4] 10A and 10B are diagrams illustrating a process of a nearby successful data search unit in the information processing device as an example of an embodiment. [Figure 5] 10A and 10B are diagrams illustrating a process of a nearby successful data search unit in the information processing device as an example of an embodiment. [Figure 6] 10A and 10B are diagrams illustrating processing by a failure data narrowing unit of an information processing apparatus according to an example of an embodiment; [Figure 7] 10A and 10B are diagrams illustrating processing by a successful data narrowing unit of an information processing apparatus according to an example of an embodiment; [Figure 8] 10A and 10B are diagrams illustrating a process performed by a weight influence measurement unit of an information processing device as an example of an embodiment. [Figure 9] 10A and 10B are diagrams for explaining processing by a weight influence measurement unit of an information processing device as one example of an embodiment; [Figure 10] 10A and 10B are diagrams illustrating a process performed by a weight influence measurement unit of an information processing device as an example of an embodiment. [Figure 11] 10A and 10B are diagrams illustrating a process performed by a weight influence measurement unit of an information processing device as an example of an embodiment. [Figure 12] 10A and 10B are diagrams illustrating a process of determining the number of selected weights by a weight-to-be-modified specifying unit of an information processing device as an example of an embodiment; [Figure 13]10A and 10B are diagrams illustrating a process of specifying a weight to be modified by a weight to be specified unit of an information processing apparatus according to an embodiment; [Figure 14] 10A and 10B are diagrams illustrating a method for selecting weights to be modified by a weight-to-be-modified specifying unit of an information processing device as an example of an embodiment; [Figure 15] 1 is a flowchart illustrating an outline of a process performed by an information processing apparatus as an example of an embodiment. [Figure 16] 10 is a flowchart illustrating a process of a test data classification unit of an information processing device as an example of an embodiment. [Figure 17] 10 is a flowchart illustrating a process of a nearby successful data search unit of an information processing device as an example of an embodiment. [Figure 18] 10 is a flowchart illustrating a process performed by a failure data narrowing unit of an information processing apparatus as an example of an embodiment. [Figure 19] 10 is a flowchart illustrating a process of a successful data narrowing unit of an information processing apparatus as an example of an embodiment. [Figure 20] 10 is a flowchart illustrating a process of a weight influence measurement unit of an information processing device as an example of an embodiment. [Figure 21] 10 is a flowchart illustrating processing by a correction target weight specifying unit of an information processing device as an example of an embodiment. [Figure 22] FIG. 10 is a diagram for explaining the relationship between modification and degradation of a DNN model. [Figure 23] 10A and 10B are diagrams illustrating weights to be modified that are determined in an information processing device as an example of an embodiment. [Figure 24] FIG. 10 is a diagram illustrating a simulation result in an information processing apparatus as an example of an embodiment. [Figure 25] FIG. 1 is a diagram illustrating a hardware configuration of an information processing apparatus according to an embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present program, method, and device will be described with reference to the drawings. However, the embodiments shown below are merely examples, and are not intended to exclude various modifications or application of techniques not explicitly stated in the embodiments. For example, the present embodiment can be implemented with various modifications within the scope of its purpose. Furthermore, each figure does not intend to include only the components shown in the figure, but may include other functions, etc.
[0013] (A) Configuration FIG. 1 is a diagram schematically illustrating a configuration of an information processing device 1 as an example of an embodiment.
[0014] The information processing device 1 realizes a model correction function that corrects a trained DNN model (machine learning model). As shown in Fig. 1, the information processing device 1 has functions as a model training unit 101, a test data classification unit 102, a nearby successful data search unit 103, a failed data narrowing down unit 104, a successful data narrowing down unit 105, a weight influence measurement unit 106, and a correction target weight identification unit 107. These blocks 101 to 107 are examples of a control unit.
[0015] The model training unit 101 performs training (retraining, machine learning) by inputting training data into a DNN model to be trained. For example, the model training unit 101 performs DNN model retraining (machine learning) by inputting training data into a trained DNN model (first machine learning model) to create a retrained DNN model (second machine learning model).
[0016] FIG. 2 is a diagram for explaining the processing of the model training unit 101 in the information processing device 1 as an example of an embodiment.
[0017] Training data and a DNN model to be trained are input to the model training unit 101. The training data may include, for example, test data (input data) and training labels (correct data, correct labels) corresponding to the test data. A plurality of training data may be referred to as a test data list.
[0018] A DNN model performs forward propagation processing (forward propagation processing) by inputting input data such as images or audio to the input layer and sequentially performing predetermined calculations in hidden layers including convolutional layers and pooling layers, thereby transmitting the information obtained by the calculations from the input side to the output side. In each layer, nodes are connected by edges. Furthermore, edges have weights.
[0019] In Figure 2, the information of the DNN model includes layers, types, weights, and the number of weights.
[0020] Layer is identification information that identifies a layer, and natural numbers are used in the example shown in Figure 2. Type is the type of layer, and Conv2D and Dense are used in the example shown in Figure 4. Weight is the weight value used in each layer, and is flattened to one dimension in the example shown in Figure 2. Also, in the example shown in Figure 2, the weights of the DNN model before training show the initial weight values. The number of weights is the number of weights provided in each layer.
[0021] After performing forward processing, the model training unit 101 performs backward processing (backpropagation processing) to determine parameters to be used in the forward processing in order to reduce the value of the error function obtained from the output data output from the output layer and the ground truth data. Then, an update processing is performed to update weights (variables) based on the results of the backpropagation processing. For example, gradient descent may be used as an algorithm to determine the update width of the weights used in the calculation of the backpropagation processing. Hereinafter, the DNN model may be simply referred to as a "model."
[0022] The model training unit 101 stores information about the trained DNN model (first machine learning model, trained model) and the retrained DNN model (second machine learning model) in a storage area such as the storage device 13 (see FIG. 25). For example, the model training unit 101 may store information about the retrained DNN model, which is created by retraining the trained DNN model, in a storage area such as the storage device 13. The information about the DNN model stored in the storage area also includes weights.
[0023] The model training unit 101 calculates the forward propagation of the trained DNN model M using test data as input, and calculates the output y of the last hidden layer.
[0024] The model training unit 101 generates a retrained DNN model (second machine learning model) by updating (retraining) only the parameters (weights to be modified) identified by the weight to be modified identification unit 107 described later out of the multiple parameters (learning parameters) of the trained DNN model M.
[0025] The test data classification unit 102 classifies test data (input data) into success data and failure data based on the output obtained by inputting test data into a DNN model (trained model) and the training label corresponding to the input data.
[0026] The successful data is test data for which the DNN model outputs the correct answer, and the output matches the training label. The successful data is an example of the second plurality of data for which the predicted result (classification class) according to the output value of the trained DNN model (first machine learning model) is correct.
[0027] The failure data is test data for which the DNN model gave an incorrect answer, i.e., the output does not match the training label. The failure data is an example of a third plurality of data for which the predicted result (classification class) according to the output value of the trained DNN model (first machine learning model) is incorrect.
[0028] The test data classification unit 102 evaluates the test data using the DNN model M and classifies the test data set X that is successful in inference. pos and a set of failing test data X neg They are classified as follows.
[0029] A set of test data X that allows successful inference pos Success Data List X pos Also, the set of test data X that fails inference may be called neg is the failure data list X neg It may also be referred to as.
[0030] The trained DNN model M may be referred to as the trained model M. The output of the trained model M may be represented by the symbol y. The correct label of the input data may be represented by the symbol y^ (or ^ on y).
[0031] The test data classification unit 102 classifies the test data into a failure data set X for the trained model M using the output y of the trained model M and the correct answer label ŷ. neg and the successful data set X pos They are classified as follows.
[0032] The test data classification unit 102 classifies the successful data list X pos and Failure Data List X neg is stored in a predetermined storage area such as the storage device 13.
[0033] The nearby successful data search unit 103 searches for successful data in the vicinity of the failed data.
[0034] FIG. 3 is a diagram for explaining the processing of the nearby successful data searching unit 103 of the information processing device 1 as an example of an embodiment.
[0035] The neighboring successful data search unit 103 searches for the failure data set X neg For each failure data, input the failure data into the DNN model and obtain the output y neg and the correct label y^ to (or from) the classification boundary, da The classification boundary may also be referred to as the decision boundary.
[0036] Distance d to the classification boundary a can be calculated by the difference between the maximum element value of the output of the DNN model for the failed data and the index value of the correct label. That is, the distance d a can be calculated using the following formula (1):
number
[0037] Distance d from the classification boundary a is an example of the difference between the output value of the DNN model for the failure data included in the plurality of failure data (third plurality of data) and the value of the correct label corresponding to the failure data.
[0038] FIG. 4 is a diagram for explaining the processing of the neighboring successful data searching unit 103 in the information processing device 1 as an example of an embodiment, and shows information generated by the neighboring successful data searching unit 103.
[0039] 4, symbol A illustrates an output and a correct label obtained by inputting failure data (test data) into a DNN model. The output of the DNN model may be the output of the last stage of the hidden layer.
[0040] Also, symbol B represents the maximum element value (y neg [argmax(y neg )]) and the correct label value (y neg Symbol C shows the distance d of each failure data (test data) to the classification boundary calculated based on the value shown in symbol B. a Here is an example:
[0041] Hereinafter, for convenience, the output obtained by inputting failure data into a DNN model may be referred to as the output of failure data, and the output obtained by inputting success data into a DNN model may be referred to as the output of success data. Also, the distance between the output of failure data and the output of success data may be simply referred to as the distance between the failure data and the success data.
[0042] The neighboring successful data search unit 103 outputs y pos For the failure data output y neg Distance d b Calculate the distance d b can be expressed as the absolute value of the difference between the vectors of the failed data and the successful data. That is, the distance d b can be calculated using the following formula (2).
number
[0043] Distance d b may be the "data distance" between the failure data and the success data, and is an example of the difference between the output value of the DNN model for each of the multiple failure data and the output value of the DNN model for each of the multiple success data.
[0044] The neighboring successful data search unit 103 searches for each failed data x i (See symbols P1 and P3 in Figure 3) b / d a The successful data (see symbols P2 and P4 in Figure 3) whose position is smaller than the neighborhood coefficient c is selected, and the set of nearby successful data X near,i The neighborhood coefficient c is a value less than 1 and can be set arbitrarily by the user. i The failure data x i and the distance d between the classification boundary a The range specified by the failure data x and the neighborhood coefficient c may be referred to as the neighborhood range. i Radius cd centered at a The area is represented by a circle.
[0045] In FIG. 3, for convenience, the failure data x indicated by the symbol P1 is i In order to distinguish it from the failure data x i The distance to the classification boundary is denoted by d a ' and the failure data x indicated by symbol P3 i Output of y neg The distance between b ', but the failure data x indicated by symbol P1 i may be treated the same.
[0046] Failure data x i The failure data x is included in the neighborhood centered on i It can be said that the failure data x i The set of successful data included in the neighborhood range centered on is the set of nearby successful data X near,i Construct a set of nearby successful data X near,i may be expressed as a neighborhood successful data list NN.
[0047] FIG. 5 is a diagram for explaining the processing of the neighboring successful data searching unit 103 in the information processing device 1 as an example of an embodiment, and shows information generated by the neighboring successful data searching unit 103.
[0048] In the following diagrams, each piece of test data (failed data, successful data) may be represented by a unique integer.
[0049] In Figure 5, symbol A illustrates the output and correct label obtained by inputting failed data (test data) into the DNN model. Symbol B illustrates the output and correct label obtained by inputting successful data (test data) into the DNN model. The outputs of the DNN model are the outputs of the last stage of the hidden layer.
[0050] Furthermore, code C represents the distance d between each of the failed data (test data) shown in code A and each of the successful data shown in code B. b Symbol D illustrates a set of successful data in the vicinity of each of the failed data shown in symbol C.
[0051] The failure data narrowing unit 104 narrows down the failure data set X neg For each failure data x i The successful data set X in the neighborhood of niar,i Number of |X niar,i In this way, the set of the top k failure data with the fewest number of successful data in the neighborhood is called the failure data set X k The failure data set X k is the filtered failure data list X k It may also be referred to as.
[0052] Below is the failure data set X neg The failure data included in the data x is sometimes written as the failure data neg using the symbol neg. i The successful data set X in the neighborhood of niar,i Number of |X niar,i | is sometimes written as the number of nearby successful data and is represented by the symbol #NN[neg].
[0053] For example, in the example shown in Figure 3, the number #NN[neg] of nearby successful data (see symbol P2) for the failed data (see symbol P1) on the left side of the figure is 2, and the number #NN[neg] of nearby successful data (see symbol P4) for the failed data (see symbol P3) on the right side of the figure is 1.
[0054] FIG. 6 is a diagram for explaining the processing of the failure data narrowing unit 104 of the information processing device 1 as an example of an embodiment, and shows information generated by the failure data narrowing unit 104.
[0055] 6, symbol A illustrates a set of successful data in the vicinity range for each failed data. The failed data narrowing unit 104 narrows down the failed data by the number of successful data in the vicinity |X niar,i The symbol B shows an example in which the failure data shown in symbol A is sorted in ascending order of the number of neighboring successful data sets.
[0056] Code C is the failure data set X created by extracting the top k results of the sorting results shown in code B. k (Narrowed failure data list X k ) is shown below.
[0057] That is, the failure data narrowing unit 104 narrows down the number of failure data pieces within a neighborhood radius (cd a The top k pieces of data are selected from the failure data with the fewest number of successful data within the range of cd as the narrowed down failure data. The plurality of failure data are an example of the third plurality of data, and are determined as the narrowed down failure data. a ) is an example of a fifth plurality of data. The narrowed-down failure data is an example of a fourth plurality of data.
[0058] In this way, the failure data narrowing unit 104 narrows down the data by the distance d a and the data distance d between the failed data and the successful data b Based on the above, narrowed-down failure data is selected from the plurality of failure data.
[0059] The successful data narrowing unit 105 narrows down the failed data list X k For each failure data neg included in near,i By obtaining (neighborhood successful data list NN), the set of successful data X near This set of successful data X near Successful data set X that has been narrowed down near Or narrowed down successful data list X near That's fine. Narrowed down successful data list X near The successful data included in the above may be referred to as narrowed successful data. The narrowed successful data is an example of the fifth plurality of data.
[0060] The successful data narrowing unit 105 narrows down the unsuccessful data set X k (Narrowed failure data list Xk ) to obtain the set of successful data X near Create a set of successful data X near may be expressed by the following equation (3):
number
[0061] FIG. 7 is a diagram for explaining the processing of the successful data narrowing unit 105 of the information processing device 1 as an example of an embodiment, and shows information generated by the successful data narrowing unit 105.
[0062] In FIG. 7, symbol A represents the failure data set X generated by the failure data narrowing unit 104. k (Narrowed failure data list X k ) is shown as an example. Also, symbol B is the failure data set X shown in symbol A. k (Narrowed failure data list X k ) based on the filtered successful data list X near Here is an example:
[0063] Narrowed failure data list X k The radius cd is the center of each failure data a The successful data within the range (neighborhood range) of radius cd may be referred to as narrowed successful data. a may be referred to as the neighborhood radius.
[0064] The successful data narrowing unit 105 narrows down the failed data list X k The union of the successful data in the vicinity of each failed data is defined as the refined successful data list X near is determined as an element of.
[0065] In this way, the successful data narrowing unit 105 narrows down the data by the distance d a and the data distance d between the failed data and the successful data b Based on the above, narrowed-down successful data is selected from the plurality of successful data.
[0066] The weight influence measurement unit 106 uses the narrowed down failure data X k and refined successful data list X near Using these, the forward influence of the failed data and the forward influence of the successful data are calculated. The forward influence represents the degree of influence on the weights for the output in the forward processing of the DNN model. The forward influence represents the degree of influence of the value (forward propagation) obtained by multiplying the output of each layer of the DNN model by the weight connected to the next layer.
[0067] The weight influence measurement unit 106 calculates the weight influence of the failure data set X k is input to the DNN model M (trained model), and the forward influence fwd is calculated for each layer of the DNN model M. neg Calculate.
[0068] The weight influence measurement unit 106 measures the i-th value o of the output in the layer l-1 of the DNN model M. i (l-1) and the weight value W in the jth row and ith column of layer l ji (l) With i (l-1) W ji (l) Forward impact fwd neg Calculate as follows.
[0069] Then, the weight influence measurement unit 106 calculates the forward influence fwd neg The larger weight W ji (l) A list of indexes in order I (l) fwd_neg Create a.
[0070] FIG. 8 is a diagram for explaining the processing by the weight influence measuring unit 106 of the information processing device 1 as one example of an embodiment, and is a diagram illustrating a calculation process of the forward influence of failure data.
[0071] In FIG. 8, the symbol A indicates the narrowed down failure data list X created by the failure data narrowing unit 104. kThe symbol B illustrates the weights for each layer as information on the DNN model (trained model) M.
[0072] The weight influence measurement unit 106 applies the narrowed down failure data list X to the DNN model M. k The failure data is input and the output is used as the forward impact fwd neg Used to calculate.
[0073] 8, symbol C indicates forward influence information for the failure data. The forward influence information illustrated in FIG. 8 includes, for each layer of the failure data, the forward influence (the influence of the weight on the output) and the number of weights for each layer.
[0074] The weight influence degree measuring unit 106 calculates the narrowed down failure data list X k Each failure data is input to the DNN model M, and for each layer, i (l-1) W ji (l) By calculating the above, the forward influence of the failure data illustrated by symbol C is obtained.
[0075] The influence of the weight on the output is calculated for each layer in the DNN model, and is calculated for each value of the weight of the same layer. In the example shown in Figure 8, for example, for the weight of layer 1 [0.054, 0.141,...,0.122], the influence of the weight on the output is calculated as [0.001, 0.002,...,-0.004]. In addition, hereinafter, the forward influence information for the failure data may be denoted by symbol 210 and referred to as forward influence information 210.
[0076] Furthermore, the weight influence measuring unit 106 generates an index list in which the weights are sorted in descending order of their influence on the output for each layer.
[0077] In FIG. 8, symbol D indicates an index list in which the indexes are sorted in order of forward influence.
[0078] Hereinafter, with regard to forward impact information on failed data, an index list in which the impact values are sorted in descending order may be referred to as a forward impact index list 211.
[0079] The forward influence order index list 211 indicated by symbol D is an index list sorted in descending order of the influence value of the weight on the output in the forward influence information 210 indicated by symbol C. This forward influence order index list 211 includes a layer, a forward influence order index (index) of the failure data, and the number of weights.
[0080] This forward influence order index list 211 is neg The larger weight W ji (l) A list of indexes in order I (l) fwd_neg This is an example.
[0081] The forward influence index sorts the influence of the weights on the output of the forward influence information 210 indicated by symbol C in descending order of value, and represents the corresponding weight index instead of the influence value.
[0082] As a result, by referring to the forward influence order index list 211, it is possible to easily know the weights for each layer that have a high influence of failure data on the output in the DNN model.
[0083] Forward impact fwd neg is an example of a numerical value calculated during forward propagation when the narrowed-down failed data (fourth plurality of data) is input to the DNN model.
[0084] Also, forward influence fwd neg is an example of a first forward influence degree calculated based on the product of the output of a layer included in a DNN model obtained by inputting failure data (third plurality of data) into the DNN model and the weight in the layer.
[0085] The weight influence measurement unit 106 calculates the success data set X near is input to the DNN model M (trained model), and the forward influence fwd is calculated for each layer of the DNN model M. pos Calculate.
[0086] The weight influence measurement unit 106 measures the i-th value o of the output in the layer l-1 of the DNN model M. i (l-1) and the weight value W in the jth row and ith column of layer l ji (l) With i (l-1) W ji (l) Forward impact fwd pos Calculate as follows.
[0087] Then, the weight influence measurement unit 106 calculates the forward influence fwd pos The larger weight W ji (l) A list of indexes in order I (l) fwd_pos Create a.
[0088] FIG. 9 is a diagram for explaining the processing by the weight influence measurement unit 106 of the information processing device 1 as one example of the embodiment, and is a diagram illustrating the calculation process of the forward influence of successful data.
[0089] In FIG. 9, the symbol A indicates the narrowed-down successful data list X created by the successful data narrowing unit 105. near The symbol B illustrates the weights for each layer as information on the DNN model (trained model) M.
[0090] The weight influence measurement unit 106 adds a narrowed-down successful data list X to the DNN model M. near The success data of the input is used as input, and the output is used as the forward influence fwd pos Used to calculate.
[0091] 9, symbol C indicates forward influence information for successful data. The forward influence information illustrated in FIG. 9 includes the forward influence (influence of weight on output) of each layer of successful data and the number of weights for each layer.
[0092] The weight influence degree measuring unit 106 calculates the narrowed down successful data list X near The successful data of each is input to the DNN model M, and for each layer, i (l-1) W ji (l) By calculating the above, the forward influence of the successful data illustrated in symbol C is obtained.
[0093] The influence of the weight on the output is calculated for each layer in the DNN model, and is calculated for each value of the weight of the same layer. In the example shown in Figure 9, for example, for the weight of layer 1 [0.054, 0.141,...,0.122], the influence of the weight on the output is calculated as [0.001, 0.002,...,-0.004]. In addition, hereinafter, the forward influence information for the successful data may be denoted by symbol 220 and referred to as forward influence information 220.
[0094] Furthermore, the weight influence measuring unit 106 generates an index list in which the weights are sorted in descending order of their influence on the output for each layer.
[0095] In FIG. 9, symbol D indicates an index list in which the indexes are sorted in order of forward influence.
[0096] Hereinafter, with regard to forward impact information about successful data, an index list in which the impact values are sorted in descending order may be referred to as a forward impact index list 221.
[0097] The forward influence order index list 211 indicated by symbol D is an index list sorted in descending order of the influence value of the weight on the output in the forward influence information 220 indicated by symbol C. This forward influence order index list 221 includes a layer, a forward influence order index (index) of successful data, and the number of weights.
[0098] This forward influence order index list 221 is pos The larger weight W ji (l) A list of indexes in order I (l) fwd_pos This is an example.
[0099] The forward influence index sorts the influence of the weights on the output of the forward influence information 220 indicated by symbol C in descending order of value, and represents the corresponding weight index instead of the influence value.
[0100] As a result, by referring to the forward influence order index list 221, it is possible to easily know the weights for each layer that have a high influence of successful data on the output in the DNN model.
[0101] In addition, the weight influence measurement unit 104 uses the narrowed down failure data X k and refined successful data list X near Using these, the backward influence of the failed data and the backward influence of the successful data are calculated, respectively. The backward influence represents the influence on the weights for the output in the backward processing of the DNN model. The backward influence represents the loss gradient (backpropagation) obtained by differentiating the loss by the weights.
[0102] The weight influence measurement unit 106 uses the narrowed down failure data X k and Narrowed successful data list X near By differentiating each output with respect to the weight, the influence of the weight on the output is quantified (measured).
[0103] Forward impact fwd pos is an example of a numerical value calculated during forward propagation when the successfully narrowed-down data (fifth plurality of data) is input into the DNN model.
[0104] Also, forward influence fwd pos is an example of a second forward influence calculated based on the product of the output of a layer included in a DNN model obtained by inputting successful data (second plurality of data) into the DNN model and the weight in that layer.
[0105] The weight influence measurement unit 106 calculates the weight influence of the failure data set X k is input to the DNN model M (trained model), and the backward influence grad neg Calculate.
[0106] The weight influence measurement unit 106 automatically differentiates the output of the DNN model using the weights to calculate the influence (influence) that each weight value has on the output. neg , the narrowed down failure data X k The output Y when input to the DNN model k Let W be the weight in layer l. (l) The backward influence grad in layer l is calculated by automatic differentiation. neg may be expressed by the following equation (4):
number
[0107] FIG. 10 is a diagram for explaining the processing by the weight influence measurement unit 106 of the information processing device 1 as one example of an embodiment, and is a diagram illustrating a calculation process of the backward influence of failure data.
[0108] In FIG. 10, the symbol A indicates the narrowed down failure data list X created by the failure data narrowing unit 104. k The symbol B illustrates the weights for each layer as information on the DNN model (trained model) M.
[0109] The weight influence measurement unit 106 applies the narrowed down failure data list X to the DNN model M. k The failure data is input and the output is the backward impact grad neg Used to calculate.
[0110] 10, symbol C indicates backward impact information about the failed data. The backward impact information shown in Fig. 10 includes the backward impact of the failed data (the impact of the weight on the output) and the number of weights for each layer.
[0111] The influence of the weight on the output is calculated for each layer in the DNN model, and is calculated for each value of the weight in the same layer. In the example shown in Figure 10, for example, for the weight of layer 1 [0.054, 0.141,...,0.122], the influence of the weight on the output is calculated as [0.001, 0.002,...,-0.004]. In addition, hereinafter, the backward influence information for the failure data may be denoted by the symbol 230 and referred to as backward influence information 230.
[0112] Furthermore, the weight influence measuring unit 106 generates an index list in which the weights are sorted in descending order of their influence on the output for each layer.
[0113] In FIG. 10, symbol D indicates an index list in which the indexes are sorted in order of backward influence.
[0114] Hereinafter, with regard to backward impact information on failed data, an index list in which the impact values are sorted in descending order may be referred to as a backward impact index list 231 .
[0115] The backward impact order index list 231 indicated by symbol D is an index list sorted in descending order of the influence value of the weight on the output in the backward impact information 230 indicated by symbol C. This backward impact order index list 231 includes a layer, a backward impact order index (index) of the failure data, and the number of weights.
[0116] This backward influence order index list 231 is the backward influence grad neg A list of the indices of the largest weights in descending order. (l) grad_neg This is an example.
[0117] The backward influence index rearranges the influence of the weights on the output of the backward influence information 230 indicated by the symbol C in descending order of value, and represents the corresponding weight index instead of the influence value.
[0118] As a result, by referring to the backward influence order index list 231, it is possible to easily know the weights for each layer that have a high influence of failure data on the output in the DNN model.
[0119] Backward influence degree grad neg is an example of a numerical value calculated during backpropagation when the narrowed-down failed data (fourth plurality of data) is input to the DNN model.
[0120] Also, backward influence grad neg is an example of a first backward influence degree obtained by inputting failure data (third plurality of data) into a DNN model and differentiating the output obtained by the input with the weight. The weight influence measurement unit 106 calculates the success data set X near is input to the DNN model M, and the backward influence grad pos Calculate.
[0121] The weight influence degree measuring unit 106 calculates the narrowed down successful data list X near The output Y when input to the DNN modelk The weights W in layer l (l) By automatically differentiating with , the backward influence grad pos The weight influence measurement unit 106 calculates the backward influence grad pos Calculate.
[0122] Then, the weight influence measuring unit 106 calculates the backward influence grad pos A list of the indices of the largest weights in descending order. (l) grad_pos Create a.
[0123] FIG. 11 is a diagram for explaining the processing by the weight influence measurement unit 106 of the information processing device 1 as one example of an embodiment, and is a diagram illustrating a calculation process of the backward influence of successful data.
[0124] In FIG. 11, the symbol A indicates the narrowed-down successful data list X created by the successful data narrowing unit 105. near The symbol B illustrates the weights for each layer as information on the DNN model (trained model) M.
[0125] The weight influence measurement unit 106 adds a narrowed-down successful data list X to the DNN model M. near The success data of the input is used to calculate the backward influence grad pos Used to calculate.
[0126] 11, symbol C indicates backward impact information for successful data. The backward impact information shown in Fig. 11 includes the backward impact of successful data (the impact of weights on output) for each layer and the number of weights for each layer.
[0127] The influence of the weight on the output is calculated for each layer in the DNN model, and is calculated for each value of the weight of the same layer. In the example shown in Figure 11, for example, for the weight of layer 1 [0.054, 0.141,...,0.122], the influence of the weight on the output is calculated as [0.001, 0.002,...,-0.004]. In addition, hereinafter, the backward influence information for successful data may be denoted by the symbol 240 and referred to as backward influence information 240.
[0128] Furthermore, the weight influence measuring unit 106 generates an index list in which the weights are sorted in descending order of their influence on the output for each layer.
[0129] In FIG. 11, symbol D indicates an index list in which the indexes are sorted in order of backward influence.
[0130] Hereinafter, with regard to the backward impact information on successful data, an index list in which the impact values are sorted in descending order may be referred to as a backward impact index list 241.
[0131] The backward influence order index list 241 indicated by symbol D is an index list sorted in descending order of the influence value of the weight on the output in the backward influence information 240 indicated by symbol C. This backward influence order index list 241 includes a layer, a backward influence order index (index) of successful data, and the number of weights.
[0132] This backward influence order index list 241 is the backward influence grad pos A list of the indices of the largest weights in descending order. (l) grad_pos This is an example.
[0133] The backward influence index rearranges the influence of the weights on the output of the backward influence information 240 indicated by the symbol C in descending order of value, and represents the corresponding weight index instead of the influence value.
[0134] This makes it possible to easily know, for each layer, the weights that have a high influence on the output in the DNN model by referring to the backward influence order index list 241.
[0135] Backward influence degree grad pos is an example of a numerical value calculated during backpropagation when the narrowed-down successful data (fifth plurality of data) is input into the DNN model.
[0136] Also, backward influence grad pos is an example of a second backward influence obtained by inputting successful data (second plurality of data) into a DNN model and differentiating the output obtained by the weight.
[0137] The correction target weight specifying unit 107 determines the number of weights to be selected for each layer of the DNN model, and determines (specifies) the weights to be corrected.
[0138] The correction target weight specifying unit 107 determines the number of weights to be selected for each layer based on the forward influence information 210, 220 and the backward influence information 230, 240.
[0139] FIG. 12 is a diagram for explaining a process of determining the number of selected weights by the correction target weight specifying unit 107 of the information processing device 1 as an example of an embodiment.
[0140] 12, symbol A indicates the forward influence information 210 for the failed data. Based on the forward influence information 210, the correction target weight identification unit 107 calculates the average value of the forward influence (the influence of the weight on the output) of the failed data for each layer.
[0141] Furthermore, the correction target weight specifying unit 107 normalizes the calculated average value of the forward influence of the failure data for each layer so that the sum of all layers is 1. The normalized average of the forward influence of the failure data is expressed as e (l) k,fwd It is expressed as:
[0142] In FIG. 12, the symbol B indicates an example of the average forward influence degree of the normalized failure data of each layer.
[0143] The correction target weight specifying unit 107 stores the calculated average of the normalized forward influence of the failure data for each layer in a predetermined storage area such as the storage device 13.
[0144] 12, symbol C indicates forward influence information 220 for successful data. Based on the forward influence information 220, the correction target weight identification unit 107 calculates the average value of the forward influence (influence of the weight on the output) of successful data for each layer.
[0145] Furthermore, the correction target weight specifying unit 107 normalizes the calculated average value of the forward influence of the successful data for each layer so that the total for all layers is 1. The normalized average of the forward influence of the successful data is expressed as e (l) near,fwd It is expressed as:
[0146] In FIG. 12, the symbol D indicates an example of the average forward influence of the normalized successful data of each layer.
[0147] The correction target weight specifying unit 107 stores the calculated average of the forward influence degrees of the normalized successful data for each layer in a predetermined storage area such as the storage device 13.
[0148] 12, the symbol E indicates the backward impact information 230 for the failed data. Based on the backward impact information 230, the correction target weight identification unit 107 calculates the average value of the backward impact (the impact of the weight on the output) of the failed data for each layer.
[0149] Furthermore, the correction target weight specifying unit 107 normalizes the calculated average value of the backward influence of the failure data for each layer so that the total for all layers is 1. The normalized average of the backward influence of the failure data is expressed as e (l) k,grad It is expressed as:
[0150] In FIG. 12, the symbol F indicates an example of the average of the normalized backward influence degrees of the failure data in each layer.
[0151] The correction target weight specifying unit 107 stores the calculated average of the normalized backward influence degrees of the failure data for each layer in a predetermined storage area such as the storage device 13.
[0152] 12, the symbol G indicates the backward influence information 240 for the successful data. Based on the backward influence information 240, the correction target weight identification unit 107 calculates the average value of the backward influence (the influence of the weight on the output) of the successful data for each layer.
[0153] Furthermore, the correction target weight specifying unit 107 normalizes the calculated average value of the backward influence of the successful data for each layer so that the total of all layers is 1. The normalized average of the backward influence of the successful data is expressed as e (l) near,grad It is expressed as:
[0154] In FIG. 12, the symbol H indicates an example of the average of the normalized backward influence degrees of the successful data of each layer.
[0155] The correction target weight specifying unit 107 stores the calculated average of the normalized backward influence degrees of the successful data for each layer in a predetermined storage area such as the storage device 13.
[0156] The correction target weight specification unit 107 determines the suspicion value susp for each layer using the following formula (5) based on the normalized average forward influence of the failure data for each layer, the average forward influence of the success data for each layer, the average backward influence of the failure data for each layer, and the average backward influence of the success data for each layer. (l) Calculate.
number
[0157] In FIG. 12, an example of the normalized suspicion value of each layer is indicated by the symbol J.
[0158] The correction target weight specifying unit 107 calculates the suspicion value susp for each layer. (l) is stored in a predetermined storage area such as the storage device 13.
[0159] The correction target weight specification unit 107 determines the suspicious value susp (l) The correction target weight specifying unit 107 determines the number of weights to be selected for each layer based on the following equation (6).
number
[0160] In addition, |W (l) n | is the number of weights in layer l, and susp (l) is the suspicion value of layer l, and ρ is the reduction rate. The reduction rate ρ may be set arbitrarily by the user.
[0161] In FIG. 12, the symbol K indicates an example of the calculated number of selected weights for each layer.
[0162] Furthermore, the correction target weight specifying unit 107 determines (specifies) the weight to be corrected for each layer.
[0163] FIG. 13 is a diagram for explaining the process of specifying the weight to be corrected by the correction weight specifying unit 107 of the information processing device 1 as one example of an embodiment.
[0164] The correction target weight identification unit 107 selects, for each layer, indexes corresponding to the number of selected weights from the top of the forward influence ordered index of the failure data based on the forward influence ordered index list 211 based on the failure data. Hereinafter, the top indexes corresponding to the number of selected weights (see symbol E in FIG. 13) selected from the forward influence ordered index of the failure data may be referred to as the forward influence top index of the failure data. The forward influence top index of the failure data is a set of weights that have a large influence on the output. Symbol A in FIG. 13 shows an example of the forward influence top index of the failure data.
[0165] The forward impact top index of the failure data is the forward impact order index list 211 (List I (l) fwd_neg ) from the top to c (l) Set I selected weights (l)’ fwd_neg is.
[0166] The correction target weight specifying unit 107 stores the forward influence upper index of the selected failure data in a predetermined storage area such as the storage device 13 .
[0167] Furthermore, the correction target weight identification unit 107 selects, for each layer, indexes equal to the number of selected weights from the top of the forward influence ordered index of the successful data based on the forward influence ordered index list 221 based on the successful data. Hereinafter, the top indexes equal to the number of selected weights (see symbol E in FIG. 13) selected from the forward influence ordered index of the successful data may be referred to as the top forward influence index of the successful data. The top forward influence index of the successful data is a set of weights that have a large influence on the output. Symbol B in FIG. 13 shows an example of the top forward influence index of the successful data.
[0168] The forward influence top index of the successful data is the forward influence order index list 221 (List I (l) fwd_pos ) from the top to c (l) Set I selected weights (l)’fwd_pos is.
[0169] The correction target weight specifying unit 107 stores the forward influence upper index of the selected successful data in a predetermined storage area such as the storage device 13.
[0170] The correction target weight identification unit 107 selects, for each layer, indexes equal to the number of selected weights from the top of the backward influence order index of the failure data based on the backward influence order index list 231 based on the failure data. Hereinafter, the top indexes equal to the number of selected weights (see symbol E in FIG. 13) selected from the backward influence order index of the failure data may be referred to as the backward influence top index of the failure data. The backward influence top index of the failure data is a set of weights that have a large influence on the output. Symbol C in FIG. 13 shows an example of the forward influence top index of the failure data.
[0171] The backward impact top index of the failure data is the backward impact order index list 231 (List I (l) grad_neg ) from the top to c (l) Set I selected weights (l)’ grad_neg is.
[0172] The correction target weight specifying unit 107 stores the backward influence upper index of the selected failure data in a predetermined storage area such as the storage device 13 .
[0173] The weight to be corrected identification unit 107 selects, for each layer, indexes equal to the number of selected weights from the top of the backward influence order index of the successful data based on the backward influence order index list 241 based on the successful data. Hereinafter, the top indexes equal to the number of selected weights (see symbol E in FIG. 13) selected from the backward influence order index of the successful data may be referred to as the top backward influence index of the successful data. The top backward influence index of the successful data is a set of weights that have a large influence on the output. Symbol D in FIG. 13 shows an example of the top backward influence index of the successful data.
[0174] The top indexes of the success data are listed in the index list 241 (List I) (l) grad_pos ) from the top to c (l) Set I selected weights (l)’ grad_pos is.
[0175] The correction target weight specifying unit 107 stores the backward influence upper index of the selected successful data in a predetermined storage area such as the storage device 13.
[0176] Then, the correction target weight specifying unit 107 determines the failure data X k Among the weights that affect the success data X near The weight to be corrected specifying unit 107 selects, in each layer constituting the DNN model M, a weight that is included in the top ranking of influence on the failure data and is not included in the top ranking of influence on the success data, as a weight to be corrected.
[0177] Specifically, in each layer constituting the DNN model M, the correction target weight specification unit 107 determines the forward influence upper index (set I (l)’ fwd_neg ) and the backward influence index of the failure data (set I (l)’ grad_neg ) and the intersection set (∩), the forward influence upper index of the successful data (set I (l)’ fwd_pos ) and the backward influence index of the successful data (set I (l)’ grad_pos ) and the subtraction set (\) of the intersection set (∩) of (l) localize ) to identify the
[0178] That is, the correction target weight identification unit 107 determines, in each layer constituting the DNN model M, a weight that satisfies the following equation (7) as a correction target weight.
number
[0179] Forward impact index of failure data (set I (l)’ fwd_neg ) and the backward influence index of the failure data (set I (l)’ grad_neg ) and the intersection (∩) of the failure data X k On the other hand, the forward influence index of the successful data (set I (l)’ fwd_pos ) and the backward influence index of the successful data (set I (l)’ grad_pos ) and the intersection (∩) of the set of successful data X near It is a set of weights that have a high influence on (highest influence on).
[0180] Therefore, in the above equation (7), the weight to be corrected (W (l) localize ) as failure data X k Among the set of weights with high influence (top), successful data X near We show that we select a set of weights that has a small effect on
[0181] Here, in layer l, susp (l) The number of weights determined according to I is defined as m. (l) fwd_neg ,I (l) grad_neg ,I (l) fwd_pos ,I (l) grad_pos Each of the top m subsets I (l)’ fwd_neg ,I (l)’ grad_neg ,I (l)’ fwd_pos ,I (l)’ grad_pos About (I (l)’ fwd_neg ∩I (l)’ grad_neg )\(I (l)’ fwd_pos ∩I (l)’grad_pos ) and use it as the index I of the weight to be modified. (l) localize Let's say.
[0182] The weights W of layer l to be corrected are (l) localize can be expressed by the following equation (7'), and the weights to be corrected for all layers W localize can be expressed by the following formula (7").
number
[0183] FIG. 14 is a diagram illustrating a method for selecting weights to be modified by the modification weight specifying unit 107 of the information processing device 1 as one example of an embodiment.
[0184] In FIG. 14, symbol A denotes the forward impact index of the failure data (set I (l)’ fwd_neg )" and symbol B is "Failure data backward impact upper index (set I (l)’ grad_neg )" is an example.
[0185] These "failure data forward impact top index (set I (l)’ fwd_neg ) and "Failure data backward impact upper index (set I (l)’ grad_neg )" has weight elements {77, 11, 425, 572}.
[0186] Also, in FIG. 14, symbol C indicates the forward influence upper index of the successful data (set I (l)’ fwd_pos )" and symbol D is "the backward influence upper index of successful data (set I (l)’ grad_pos )" is an example.
[0187] These "top indexes of forward influence of successful data (set I (l)’ fwd_pos) and "Backward influence index of successful data (Set I (l)’ grad_pos )" has weight elements {287, 11, 425}.
[0188] The difference set (\) between the set of weight elements {77, 11, 425, 572} and the set of weight elements {287, 11, 425} is the weight element {77, 572}. In this case, the correction target weight identification unit 107 selects the weight element {77, 572} as the weight to be corrected.
[0189] In this way, the correction target weight identification unit 107 identifies the correction target weights (second multiple parameters) among the multiple learning parameters (first multiple parameters) included in the DNN model based on the numerical values calculated during forward propagation and the numerical values calculated during backward propagation when the narrowed-down failed data and the narrowed-down successful data are input into the DNN model.
[0190] The correction target weight specification unit 107 determines the correction target weight W (l) localize is stored in a predetermined storage area such as the storage device 13. Then, the correction target weight specifying unit 107 determines the correction target weight W localize Output.
[0191] The W output in this way localize For example, weights may be modified using PSO (Particle swarm optimization). Note that fitness may be calculated using the following equation (8).
number
[0192] In addition, N patched indicates the number of corrected failed test data, and N intact indicates the number of successful test data that remains unchanged. PSO is a known method, and a detailed explanation thereof will be omitted.
[0193] (B) Operation Next, an outline of the processing in the information processing device 1 as an example of the embodiment configured as described above will be described with reference to the flowchart (steps A1 to A8) shown in FIG.
[0194] In step A1, test data and a trained DNN model are input to the information processing device 1.
[0195] In step A2, the test data classification unit 102 executes a classification process for the test data. Details of the classification process for the test data by the test data classification unit 102 will be described later with reference to the flowchart shown in FIG.
[0196] In step A3, the neighboring successful data searching unit 103 executes a process of searching for successful data neighboring the unsuccessful data. Details of the process of searching for successful data neighboring the unsuccessful data by the neighboring successful data searching unit 103 will be described later with reference to the flowchart shown in FIG.
[0197] In step A4, the failure data narrowing unit 104 narrows down the failure data set X k (Narrowed failure data list X k ) is created by the failure data narrowing unit 104. k (Narrowed failure data list X k The details of the process of creating the ) will be described later with reference to the flowchart shown in FIG.
[0198] In step A5, the successful data narrowing unit 105 narrows down the successful data set X near (Narrowed successful data list X near ) is created by the successful data narrowing unit 105. near (Narrowed successful data list X near The details of the process of creating the ) will be described later with reference to the flowchart shown in FIG.
[0199] In step A6, the weight influence measurement unit 106 calculates the forward influence of the failure data and the forward influence of the success data. Details of the calculation process of the weight influence measurement unit 106 for the forward influence of the failure data and the forward influence of the success data will be described later with reference to the flowchart shown in FIG.
[0200] In step A7, the correction target weight specifying unit 107 determines the number of selected weights for each layer of the DNN model and determines (specifies) the weights to be corrected. Details of the process of determining the number of selected weights and the weights to be corrected by the correction target weight specifying unit 107 will be described later using the flowchart shown in FIG.
[0201] Thereafter, in step A8, the weight to be corrected is output, and the process ends.
[0202] Next, the processing of the test data classification unit 102 of the information processing device 1 as an example of an embodiment will be described with reference to the flowchart (steps B1 to B9) shown in FIG.
[0203] In step B1, a test data list, correct labels of the test data, and a trained DNN model M are input to the information processing device 1.
[0204] In step B2, the test data classification unit 102 classifies the failure data list X neg and Success Data List X pos The storage location is secured in the memory 12 (see FIG. 25).
[0205] In step B3, a loop process is started in which the control up to step B7 is repeatedly performed for all test data (input data) present in the test data list.
[0206] In step B4, the test data classification unit 102 inputs the test data into the trained DNN model M and calculates estimated labels.
[0207] In step B5, the test data classification unit 102 checks whether the estimated label output from the DNN model M matches the correct label of the test data.
[0208] As a result of the check, if the estimated label and the correct label match (YES in step B5), the process proceeds to step B6. In step B6, the test data classification unit 102 classifies the successful data list X pos Add test data to
[0209] If the confirmation result in step B5 indicates that the estimated label and the correct label do not match (NO in step B5), the process proceeds to step B7. In step B7, the test data classification unit 102 adds the test data to the failure data list Xneg.
[0210] Then, control proceeds to step B8. In step B8, loop end processing corresponding to step B3 is performed. When processing for all test data is completed, control proceeds to step B9.
[0211] In step B9, the failure data list X neg and Success Data List X pos is output. Then, the process ends.
[0212] Next, the processing of the nearby successful data searching unit 103 of the information processing device 1 as an example of an embodiment will be described with reference to the flowchart (steps C1 to C15) shown in FIG.
[0213] In step C1, the failure data list X neg ,Success Data List X pos and a trained DNN model M is input.
[0214] In step C2, the neighboring successful data searching unit 103 reserves in the memory 12 a storage location for the neighboring successful data list NN for the failed data.
[0215] In step C3, the failure data list X neg All the failure data neg(neg∈X neg ), a loop process is started in which the process up to step C13 is repeatedly performed.
[0216] In step C4, the neighboring successful data search unit 103 inputs the failure data neg to the DNN model M and outputs the output y neg Calculate.
[0217] In step C5, the neighboring successful data search unit 103 calculates the distance of the failed data neg to the classification boundary layer using the above formula (1).
[0218] In step C6, the neighboring successful data search unit 103 creates a neighboring successful data list X near, nag A storage area for the above is secured in the memory 12.
[0219] In step C7, the successful data list X pos All successful data pos(pos∈X pos ), a loop process is started in which the process up to step C11 is repeatedly performed.
[0220] In step C8, the neighboring successful data search unit 103 inputs the successful data pos to the DNN model M and outputs the output y pos Calculate.
[0221] In step C9, the nearby successful data searching unit 103 calculates the distance from the failed data neg to the successful data pos using the above formula (2).
[0222] In step C10, the neighboring successful data search unit 103 calculates d b >c×d a That is, the nearby successful data searching unit 103 checks whether the successful data pos is within the nearby range of the failed data neg.
[0223] d b >c×d a If the above condition is met (YES in step C10), the process proceeds to step C11.
[0224] In step C11, the neighboring successful data search unit 103 searches the neighboring successful data list X near, neg Add the success data pos to.
[0225] Also, as a result of the check in step C10, d b >c×d a If the above condition is not met (NO in step C10), the process proceeds to step C12.
[0226] In step C12, the loop end process corresponding to step C7 is performed. pos When the process for all the successful data pos is completed, the process proceeds to step C13.
[0227] In step C13, the neighboring successful data search unit 103 adds the neighboring successful data list X for the failure data neg to the neighboring successful data list NN[neg] of the failure data neg. near, nag Add.
[0228] Then, in step C14, the loop end process corresponding to step C3 is performed. neg When the process for all the failure data neg included in is completed, the process proceeds to step C15.
[0229] In step C15, the neighboring successful data list NN is output, and then the process ends.
[0230] Next, the process performed by the failure data narrowing unit 104 of the information processing device 1 as an example of an embodiment will be described with reference to the flowchart (steps D1 to D5) shown in FIG.
[0231] In step D1, the failure data list X neg and the neighboring successful data list NN is input.
[0232] In step D2, the failure data narrowing unit 104 narrows down the failure data list X k A storage location for the above is secured in a storage area such as the memory 12.
[0233] In step D3, the failure data narrowing unit 104 narrows down the failure data list X neg are sorted in ascending order according to the number of neighboring successful data (NN[neg]) for each failed data neg.
[0234] In step D4, the failure data narrowing unit 104 sorts the failure data list X neg The top k failure data are obtained and the filtered failure data list X k Save to.
[0235] In step D5, the narrowed down failure data list X k is output, and then the process ends.
[0236] Next, the processing of the successful data narrowing unit 105 of the information processing device 1 as an example of an embodiment will be described with reference to the flowchart (steps E1 to E6) shown in FIG.
[0237] In step E1, the narrowed down failure data list X k and the neighboring successful data list NN is input.
[0238] In step E2, the successful data narrowing unit 105 narrows down the successful data list X near A storage location for the above is secured in a storage area such as the memory 12.
[0239] In step E3, the refined failure data list X k All the failure data neg(neg∈X k), a loop process is started in which the process of step E4 is repeatedly performed.
[0240] In step E4, the successful data narrowing unit 105 narrows down the neighboring successful data list NN[neg] of the failed data neg to the narrowed down successful data list X near Add data so that there is no duplication.
[0241] In step E5, a loop end process corresponding to step E3 is performed (the first loop ends). k All the failure data neg(neg∈X k ) is completed, the process proceeds to step E6.
[0242] In step E6, the narrowed down successful data list X near is output and the process ends.
[0243] Next, the processing of the weight influence measurement unit 106 of the information processing device 1 as an example of an embodiment will be described with reference to the flowchart (steps F1 to F4) shown in FIG.
[0244] In step F1, the narrowed down failure data list X k ,Narrowed successful data list X near and a trained DNN model M is input.
[0245] In step F2, the weight influence measurement unit 106 calculates the narrowed down failure data list X k is input to the DNN model M. The weight influence measurement unit 106 calculates the forward influence fwd of each weight. neg and the backward influence of each weight grad neg and are stored in a storage area such as the memory 12.
[0246] In step F3, the weight influence measurement unit 106 calculates the narrowed down successful data list X nearis input to the DNN model M. The weight influence measurement unit 106 calculates the forward influence fwd of each weight. pos and the backward influence of each weight grad pos and are stored in a storage area such as the memory 12.
[0247] In step F4, the extracted failure data X k The forward influence of each weight on neg ,Extracted failure data X k The backward influence of each weight on grad neg ,Extracted successful data X near The forward influence of each weight on pos and extracted successful data X near The backward influence of each weight on grad pos is output. Then, the process ends.
[0248] Next, the processing of the correction target weight identifying unit 107 of the information processing device 1 as one example of the embodiment will be described with reference to the flowchart (steps G1 to G8) shown in FIG.
[0249] In step G1, the extracted failure data X k The forward influence of each weight on neg ,Extracted failure data X k The backward influence of each weight on grad neg ,Extracted successful data X near The forward influence of each weight on pos and extracted successful data X near The backward influence of each weight on grad pos and a trained DNN model M is input.
[0250] In step G2, the correction target weight specification unit 107 determines the influence (fwd) of each weight. neg ,grad neg ,fwd pos ,grad pos ) to determine the number of weights to be selected in each layer (number of selected weights) m.
[0251] In step G3, the correction target weight specification unit 107 specifies each weight of the DNN model M as a forward influence fwd neg sorted in descending order and the top m weight set W f,n This weight set W f,n is the forward influence order index list 211 (list I (l) fwd_neg ) Selected from the top of set I (l)’ fwd_neg This is an example of (top index of forward impact of failure data).
[0252] In step G4, the correction target weight specification unit 107 assigns each weight of the DNN model M to a backward influence grad neg sorted in descending order and the top m weight set W g,n This weight set W g,n is the backward influence order index list 231 (list I (l) grad_neg ) Selected from the top of set I (l)’ grad_neg This is an example of (higher index of backward impact of failure data).
[0253] In step G5, the correction target weight specification unit 107 determines each weight of the DNN model M as a forward influence fwd pos sorted in descending order and the top m weight set W f,p This weight set W f,p is the forward influence order index list 221 (list I (l) fwd_pos ) Selected from the top set I (l)’ fwd_pos This is an example of (top index of forward impact of successful data).
[0254] In step G6, the correction target weight specification unit 107 calculates each weight of the DNN model M as a backward influence grad pos sorted in descending order and the top m weight set W g,p This weight set W g,p is the backward influence order index list 241 (list I (l) grad_pos) Selected from the top set I (l)’ grad_pos This is an example of (higher index of backward impact of successful data).
[0255] In step G7, the correction target weight specification unit 107 specifies the correction target weight W localized is calculated using the following equation (9) and stored in a storage area such as the memory 12. W localized = (W f,n ∩W g,n ) \(W f,p ∩W g,p ) (9)
[0256] In step G8, the weight W to be corrected is localized is output and the process ends.
[0257] Then, the weights to be corrected W localize The weights are modified using, for example, PSO.
[0258] (C) Effects of one embodiment In this way, according to the information processing device 1 as an example of the embodiment, the failure data narrowing unit 104 narrows down the failure data by selecting the data within the neighborhood radius (cd a The top k pieces of data are selected preferentially from the failure data with the fewest number of successful data within the range of .
[0259] Furthermore, the successful data narrowing unit 105 selects the successful data included in the vicinity range of the narrowed-down unsuccessful data as the narrowed-down successful data.
[0260] The weight influence measurement unit 106 calculates the weight influence (fwd) based on the narrowed-down failure data and narrowed-down success data. neg ,grad neg ,fwd pos ,grad posThen, the correction target weight specifying unit 107 selects, from the set of weights that have a high (high ranking) impact on the failed data, a set of weights that have a low impact on the successful data as the correction target weights.
[0261] This makes it possible to prevent degradation when updating the weights to be modified as selected as described above when modifying a DNN model, thereby improving the accuracy of the DNN model.
[0262] Figure 22 is a diagram illustrating the relationship between DNN model correction and degradation. Correcting a DNN model changes the classification boundary (decision boundary). This can result in data (successful data) that was correctly inferred by the DNN model before correction not being correctly inferred by the corrected DNN model, resulting in so-called degradation. This type of degradation often occurs in successful data that is close to the failed data to be corrected.
[0263] Points 1 and 2 below can be seen as examples of degradation that occurs when modifying a DNN model.
[0264] Point 1: Failed data with few nearby successful data is less likely to degrade due to changes in the classification boundary that accompany modifications to the DNN model.
[0265] Point 2: When correcting failure data far from the classification boundary, the change in the classification boundary is large, and degradation is likely to occur.
[0266] FIG. 23 is a diagram illustrating the weights to be modified that are determined in the information processing device 1 as an example of an embodiment.
[0267] In the information processing device 1, based on the above point 1, failure data that has no success data nearby is used preferentially for selecting weights to be corrected (see symbol P1). This makes it possible to reduce the amount of success data that is affected by the correction of the DNN model, and to prevent the occurrence of degradation.
[0268] Furthermore, based on point 2, the neighborhood radius of the neighborhood area set for the failed data is increased as the failed data is further from the classification boundary, thereby widening the neighborhood area (see symbol P2). This makes it easier for successful data to be included in the neighborhood area of failed data far from the classification boundary, and as a result, failed data far from the classification boundary is less likely to be used in selecting weights to be corrected, thereby preventing degradation.
[0269] The successful data narrowing unit 105 narrows down the number of successful data pieces by dividing the number of unsuccessful data pieces included in the narrowed down unsuccessful data pieces by a distance d a and the neighborhood coefficient c. a ) is determined as the refined successful data.
[0270] This makes it possible to easily identify successful data that may be degraded by correcting failed data.
[0271] The failure data narrowing unit 104 narrows down the failure data by a neighborhood radius (cd a ) and determine the top k pieces of data (items) selected from those with the fewest successful data within the range as the refined failed data. This allows the use of failed data that is unlikely to cause degradation of nearby successful data due to correction when selecting weights to be corrected, reducing the number of successful data affected by correction of the DNN model and preventing degradation.
[0272] FIG. 24 is a diagram illustrating a simulation result in the information processing device 1 as an example of an embodiment.
[0273] In FIG. 24, the method for determining edges to be corrected by the information processing device 1 is compared with the conventional method (Arachne) using an image of a traffic sign of the GTSRB (The German Traffic Sign Recognition Benchmark) as target data.
[0274] As shown in FIG. 24, it can be seen that the machine learning method by the information processing device 1 has succeeded in suppressing the BreakRate while maintaining a sufficient RepairRate.
[0275] (D) Hardware configuration example FIG. 25 is a diagram illustrating a hardware configuration of an information processing device 1 as an example of an embodiment.
[0276] The information processing device 1 is an example of a computer, and includes, as components, a processor 11, a memory 12, a storage device 13, a graphics processing device 14, an input interface 15, an optical drive device 16, a device connection interface 17, and a network interface 18. These components 11 to 18 are configured to be able to communicate with each other via a bus 19.
[0277] The processor 11 controls the entire information processing device 1. The processor 11 may be a multiprocessor. The processor 11 may be, for example, any one of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array). The processor 11 may also be a combination of two or more types of elements from the CPU, MPU, DSP, ASIC, PLD, and FPGA. The processor 11 may also be a GPU (Graphics Processing Unit).
[0278] Then, when the processor 11 executes a control program (machine learning program: not shown) for the information processing device 1, the functions of a control unit including a model training unit 101, a test data classification unit 102, a nearby successful data search unit 103, a failed data narrowing down unit 104, a successful data narrowing down unit 105, a weight influence measurement unit 106, and a weight to be corrected identification unit 107, as illustrated in Figure 1, are realized.
[0279] The information processing device 1 may realize the function as a control unit by executing a program (machine learning program, OS program) recorded on a computer-readable non-transitory recording medium, for example.
[0280] The program describing the processing to be executed by the information processing device 1 can be recorded on various recording media. For example, the program to be executed by the information processing device 1 can be stored in the storage device 13. The processor 11 loads at least a part of the program in the storage device 13 into the memory 12 and executes the loaded program.
[0281] The program to be executed by the information processing device 1 (processor 11) may also be recorded on a non-transitory portable recording medium such as an optical disk 16a, a memory device 17a, or a memory card 17c. The program stored on the portable recording medium becomes executable after being installed in the storage device 13, for example, under the control of the processor 11. The processor 11 may also read and execute the program directly from the portable recording medium.
[0282] The memory 12 is a storage memory including a ROM (Read Only Memory) and a RAM (Random Access Memory). The RAM of the memory 12 is used as the main storage device of the information processing device 1. The RAM temporarily stores at least a part of the program to be executed by the processor 11. The memory 12 also stores various data used in processing by the processor 11.
[0283] The storage device 13 is a storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a storage class memory (SCM), and stores various data. The storage device 13 is used as an auxiliary storage device for the information processing device 1.
[0284] An OS program, a control program, and various data are stored in the storage device 13. The control program includes a machine learning program.
[0285] The auxiliary storage device may be a semiconductor storage device such as an SCM or a flash memory. A plurality of storage devices 13 may be used to configure a RAID (Redundant Array of Inexpensive Disks).
[0286] The storage device 13 may store information about the DNN model M, and may also store test data (failure data, success data). Furthermore, the storage device 13 may store at least a portion of the neighborhood success data, the narrowed down failure data, the narrowed down success data, and each influence. The storage device 13 may also store the number of selected weights for each layer determined by the weight influence measurement unit 106.
[0287] A monitor 14a is connected to the graphics processing device 14. The graphics processing device 14 displays an image on the screen of the monitor 14a in accordance with an instruction from the processor 11. Examples of the monitor 14a include a display device using a CRT (Cathode Ray Tube) and a liquid crystal display device.
[0288] A keyboard 15a and a mouse 15b are connected to the input interface 15. The input interface 15 transmits signals sent from the keyboard 15a and the mouse 15b to the processor 11. The mouse 15b is an example of a pointing device, and other pointing devices can also be used. Examples of other pointing devices include a touch panel, a tablet, a touch pad, and a trackball.
[0289] The optical drive device 16 uses a laser beam or the like to read data recorded on an optical disc 16a. The optical disc 16a is a portable, non-transitory recording medium on which data is recorded so that it can be read by reflected light. Examples of the optical disc 16a include a DVD (Digital Versatile Disc), a DVD-RAM, a CD-ROM (Compact Disc Read Only Memory), and a CD-R (Recordable) / RW (Rewritable).
[0290] The device connection interface 17 is a communication interface for connecting peripheral devices to the information processing device 1. For example, a memory device 17a or a memory reader / writer 17b can be connected to the device connection interface 17. The memory device 17a is a non-transitory recording medium, such as a USB (Universal Serial Bus) memory, that has a function for communicating with the device connection interface 17. The memory reader / writer 17b writes data to or reads data from a memory card 17c. The memory card 17c is a card-type non-transitory recording medium.
[0291] The network interface 18 is connected to a network. The network interface 18 transmits and receives data via the network. Other information processing devices, communication devices, etc. may be connected to the network.
[0292] Regardless of the above-described embodiment, various modifications can be made without departing from the spirit of the present embodiment.
[0293] For example, in the above-described embodiment, an example is shown in which the invention is applied to a DNN, but the invention is not limited to this and may be applied to an NN.
[0294] Furthermore, the above disclosure will enable those skilled in the art to implement and manufacture the present embodiment.
[0295] (E) Supplementary Note The following additional notes are provided regarding the above-described embodiments.
[0296] (Appendix 1) Among the first plurality of data, a second plurality of data is identified as a correct prediction result according to the output value of the first machine learning model for each of the first plurality of data, and a third plurality of data is identified as an incorrect prediction result; selecting a fourth plurality of data from the third plurality of data and a fifth plurality of data related to the fourth data from the second plurality of data based on a difference between an output value of the first machine learning model for data included in the third plurality of data and a correct label value corresponding to the data included in the third plurality of data, and a difference between an output value of the first machine learning model for each of the third plurality of data and an output value of the first machine learning model for each of the second plurality of data; identifying a second plurality of parameters from among a first plurality of parameters included in the first machine learning model based on a numerical value calculated during forward propagation and a numerical value calculated during backward propagation when the fourth plurality of data and the fifth plurality of data are input to the first machine learning model; generating a second machine learning model by updating only the second plurality of parameters among the first plurality of parameters; A machine learning program that lets a computer perform processing.
[0297] (Appendix 2) the process of selecting the fifth plurality of data includes a process of determining, as the fifth plurality of data, data of the second plurality of data that is within a range of a neighborhood radius determined by a difference between output values of the first machine learning model for the fourth plurality of data and a neighborhood coefficient, with the fourth plurality of data being at the center; The machine learning program described in Appendix 1.
[0298] (Appendix 3) the process of selecting the fourth plurality of data includes a process of determining, as the fourth plurality of data, a top plurality of data selected from the third plurality of data in order of the number of the fifth plurality of data within the neighborhood radius. The machine learning program described in Appendix 2.
[0299] (Appendix 4) The process of identifying the second plurality of parameters includes a process of identifying a third parameter group as the second plurality of parameters, the third parameter group being obtained by subtracting a second parameter group being selected by giving priority to parameters having a small influence on the second plurality of data from a first parameter group being selected by giving priority to parameters having a high influence on the third plurality of data. The machine learning program according to any one of Supplementary Note 1 to Supplementary Note 3.
[0300] (Appendix 5) The degree of influence on the third plurality of data is: a first forward influence calculated based on the product of an output of a layer included in the first machine learning model obtained by inputting the third plurality of data into the first machine learning model and a weight in the layer; The machine learning program described in Appendix 4.
[0301] (Appendix 6) The degree of influence on the third plurality of data is: a second forward influence calculated based on the product of an output of a layer included in the first machine learning model obtained by inputting the second plurality of data into the first machine learning model and a weight in the layer; 1. The machine learning program according to claim 4 or 5.
[0302] (Appendix 7) The degree of influence on the third plurality of data is: a first backward influence degree obtained by differentiating an output obtained by inputting the third plurality of data into the first machine learning model with a weight; The machine learning program according to any one of Supplementary Note 4 to Supplementary Note 6.
[0303] (Appendix 8) The degree of influence on the third plurality of data is: and a second backward influence degree obtained by differentiating an output obtained by inputting the second plurality of data into the first machine learning model with a weight. The machine learning program according to any one of Supplementary Note 4 to Supplementary Note 7.
[0304] (Appendix 9) Among the first plurality of data, a second plurality of data is identified as a correct prediction result according to the output value of the first machine learning model for each of the first plurality of data, and a third plurality of data is identified as an incorrect prediction result; selecting a fourth plurality of data from the third plurality of data and a fifth plurality of data related to the fourth data from the second plurality of data based on a difference between an output value of the first machine learning model for data included in the third plurality of data and a value of a correct label corresponding to the data included in the third plurality of data, and a difference between an output value of the first machine learning model for each of the third plurality of data and an output value of the first machine learning model for each of the second plurality of data; identifying a second plurality of parameters from among a first plurality of parameters included in the first machine learning model based on a numerical value calculated during forward propagation and a numerical value calculated during backward propagation when the fourth plurality of data and the fifth plurality of data are input to the first machine learning model; generating a second machine learning model by updating only the second plurality of parameters among the first plurality of parameters; A machine learning method in which processing is performed by a computer.
[0305] (Appendix 10) the process of selecting the fifth plurality of data includes a process of determining, as the fifth plurality of data, data of the second plurality of data that is within a range of a neighborhood radius determined by a difference between output values of the first machine learning model for the fourth plurality of data and a neighborhood coefficient, with the fourth plurality of data being at the center; 10. The machine learning method described in Appendix 9.
[0306] (Appendix 11) the process of selecting the fourth plurality of data includes a process of determining, as the fourth plurality of data, a top plurality of data selected from the third plurality of data in order of the number of the fifth plurality of data within the neighborhood radius; 11. The machine learning method of claim 10.
[0307] (Appendix 12) the process of identifying the second plurality of parameters includes a process of identifying a third parameter group as the second plurality of parameters, the third parameter group being obtained by subtracting a second parameter group being selected by giving priority to parameters having a small influence on the second plurality of data from a first parameter group being selected by giving priority to parameters having a large influence on the third plurality of data; The machine learning method according to any one of Supplementary Note 9 to Supplementary Note 11.
[0308] (Appendix 13) The degree of influence on the third plurality of data is: a first forward influence calculated based on the product of an output of a layer included in the first machine learning model obtained by inputting the third plurality of data into the first machine learning model and a weight in the layer; 13. The machine learning method of claim 12.
[0309] (Appendix 14) The degree of influence on the third plurality of data is: a second forward influence calculated based on the product of an output of a layer included in the first machine learning model obtained by inputting the second plurality of data into the first machine learning model and a weight in the layer; 14. The machine learning method according to claim 12 or 13.
[0310] (Appendix 15) The degree of influence on the third plurality of data is: a first backward influence degree obtained by differentiating an output obtained by inputting the third plurality of data into the first machine learning model with a weight; The machine learning method according to any one of Supplementary Note 12 to Supplementary Note 14.
[0311] (Appendix 16) The degree of influence on the third plurality of data is: and a second backward influence degree obtained by differentiating an output obtained by inputting the second plurality of data into the first machine learning model with a weight. The machine learning method according to any one of Supplementary Note 12 to Supplementary Note 15.
[0312] (Appendix 17) Among the first plurality of data, a second plurality of data is identified as a correct prediction result according to the output value of the first machine learning model for each of the first plurality of data, and a third plurality of data is identified as an incorrect prediction result; selecting a fourth plurality of data from the third plurality of data and a fifth plurality of data related to the fourth data from the second plurality of data based on a difference between an output value of the first machine learning model for data included in the third plurality of data and a value of a correct label corresponding to the data included in the third plurality of data (a distance from a decision boundary and a difference between an output value of the first machine learning model for each of the third plurality of data and an output value of the first machine learning model for each of the second plurality of data); identifying a second plurality of parameters from among a first plurality of parameters included in the first machine learning model based on a numerical value calculated during forward propagation and a numerical value calculated during backward propagation when the fourth plurality of data and the fifth plurality of data are input to the first machine learning model; generating a second machine learning model by updating only the second plurality of parameters among the first plurality of parameters; An information processing device comprising a control unit that executes processing.
[0313] (Appendix 18) The control unit Among the second plurality of data, data within a range of a neighborhood radius centered on the fourth plurality of data and determined by a difference between output values of the first machine learning model for the data and a neighborhood coefficient is determined as the fifth plurality of data. 18. The information processing device according to claim 17.
[0314] (Appendix 19) The control unit Among the third plurality of data, a top plurality of data selected in order from data having a small number of the fifth plurality of data within the neighborhood radius is determined as a fourth plurality of data. 19. The information processing device according to claim 18.
[0315] (Appendix 20) The control unit A third parameter group obtained by subtracting a second parameter group selected by giving priority to parameters having a small influence on the second plurality of data from a first parameter group selected by giving priority to parameters having a high influence on the third plurality of data is identified as the second plurality of parameters. 19. The information processing device according to claim 17. [Explanation of symbols]
[0316] 1. Information processing equipment 11 processors 12 Memory 13 Storage device 14 Graphics Processing Unit 14a Monitor 15 Input Interface 15a keyboard 15b Mouse 16 Optical drive device 16a Optical disc 17 Device connection interface 17a Memory Device 17b Memory reader / writer 17c memory card 18 Network Interface 18a Network 19 Bus 101 Model Training Department 102 Test Data Classification Unit 103 Nearby successful data search unit 104 Failure Data Filtering Section 105 Successful Data Refinement Section 106 Weight influence measurement unit 107 Correction target weight specification unit 210,220 Forward influence information 230,240 Backward impact information 211,221 Forward Influence Index List 231,241 Backward Impact Index List
Claims
1. Among the first plurality of data, a second plurality of data is identified as a correct prediction result according to the output value of the first machine learning model for each of the first plurality of data, and a third plurality of data is identified as an incorrect prediction result; selecting a fourth plurality of data from the third plurality of data and a fifth plurality of data related to the fourth data from the second plurality of data based on a difference between an output value of the first machine learning model for data included in the third plurality of data and a value of a correct label corresponding to the data included in the third plurality of data, and a difference between an output value of the first machine learning model for each of the third plurality of data and an output value of the first machine learning model for each of the second plurality of data; identifying a second plurality of parameters from among a first plurality of parameters included in the first machine learning model based on a numerical value calculated during forward propagation and a numerical value calculated during backward propagation when the fourth plurality of data and the fifth plurality of data are input to the first machine learning model; generating a second machine learning model by updating only the second plurality of parameters among the first plurality of parameters; A machine learning program that lets a computer perform processing.
2. The process of selecting the fifth plurality of data includes a process of determining, as the fifth plurality of data, data among the second plurality of data that is within a range of a neighborhood radius determined by a difference between output values of the first machine learning model for the fourth plurality of data and a neighborhood coefficient, with the fourth plurality of data being the center. The machine learning program according to claim 1 .
3. the process of selecting the fourth plurality of data includes a process of determining, as the fourth plurality of data, a top plurality of data selected from the third plurality of data in order of the number of the fifth plurality of data within the neighborhood radius. The machine learning program according to claim 2 .
4. the process of identifying the second plurality of parameters includes a process of identifying a third parameter group as the second plurality of parameters, the third parameter group being obtained by subtracting a second parameter group, which is a selection of parameters having a small influence on the second plurality of data, from a first parameter group, which is a selection of parameters having a large influence on the third plurality of data. The machine learning program according to any one of claims 1 to 3.
5. The degree of influence on the third plurality of data is a first forward influence degree calculated based on the product of an output of a layer included in the first machine learning model obtained by inputting the third plurality of data into the first machine learning model and a weight in the layer; The machine learning program according to claim 4 .
6. The degree of influence on the third plurality of data is a second forward influence degree calculated based on the product of an output of a layer included in the first machine learning model obtained by inputting the second plurality of data into the first machine learning model and a weight in the layer; The machine learning program according to claim 4 or 5.
7. The degree of influence on the third plurality of data is a first backward influence degree obtained by differentiating an output obtained by inputting the third plurality of data into the first machine learning model with a weight; The machine learning program according to any one of claims 4 to 6.
8. The degree of influence on the third plurality of data is a second backward influence degree obtained by differentiating an output obtained by inputting the second plurality of data into the first machine learning model with a weight; The machine learning program according to any one of claims 4 to 7.
9. Among the first plurality of data, a second plurality of data is identified as a correct prediction result according to the output value of the first machine learning model for each of the first plurality of data, and a third plurality of data is identified as an incorrect prediction result; selecting a fourth plurality of data from the third plurality of data and a fifth plurality of data related to the fourth data from the second plurality of data based on a difference between an output value of the first machine learning model for data included in the third plurality of data and a value of a correct label corresponding to the data included in the third plurality of data, and a difference between an output value of the first machine learning model for each of the third plurality of data and an output value of the first machine learning model for each of the second plurality of data; identifying a second plurality of parameters from among a first plurality of parameters included in the first machine learning model based on a numerical value calculated during forward propagation and a numerical value calculated during backward propagation when the fourth plurality of data and the fifth plurality of data are input to the first machine learning model; generating a second machine learning model by updating only the second plurality of parameters among the first plurality of parameters; A machine learning method in which processing is performed by a computer.
10. Among the first plurality of data, a second plurality of data is identified as a correct prediction result according to the output value of the first machine learning model for each of the first plurality of data, and a third plurality of data is identified as an incorrect prediction result; selecting a fourth plurality of data from the third plurality of data and a fifth plurality of data related to the fourth data from the second plurality of data based on a difference between an output value of the first machine learning model for data included in the third plurality of data and a value of a correct label corresponding to the data included in the third plurality of data, and a difference between an output value of the first machine learning model for each of the third plurality of data and an output value of the first machine learning model for each of the second plurality of data; identifying a second plurality of parameters from among a first plurality of parameters included in the first machine learning model based on a numerical value calculated during forward propagation and a numerical value calculated during backward propagation when the fourth plurality of data and the fifth plurality of data are input to the first machine learning model; generating a second machine learning model by updating only the second plurality of parameters among the first plurality of parameters; An information processing device comprising a control unit that executes processing.
Citation Information
Patent Citations
Learning method for neural network fitted to parallel processing
JP1997128358A
Learning support device and learning support method
JP2019204190A
Arithmetic processing unit, control program, and control method
JP2021015420A
Model learning device, model learning method, and program
WO2019138655A1