Weight reduction of learning model

JP2023163102A5Inactive Publication Date: 2025-05-09窪田望
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022113601
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-07-15
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing weight reduction methods for learning models are not optimally tailored to specific learning data and models, leading to suboptimal performance.

Method used

A system and method that applies multiple weight reduction techniques to learning models, evaluates their performance using supervised learning, and generates a predictive model to identify the most appropriate weight reduction method based on learning accuracy and model size.

Benefits of technology

Enables more precise and effective weight reduction of learning models by selecting the most suitable method for each model, maintaining learning accuracy while reducing model size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an information processing method, a program, and an information processing apparatus that more appropriately reduce weight of a learned model using a neural network while maintaining learning accuracy.SOLUTION: A method includes: acquiring predetermined learning data S102; inputting predetermined data to a weight learning model in which models each including at least two models of weight reduction models are weighted on a predetermined learning model using a neural network to perform machine learning S104; acquiring a learning result when machine learning is performed by inputting the predetermined learning data to each of the weight learning models S106; performing supervised learning by using learning data including the weight learning models and learning results when learning is performed by using the weight learning models S108; and when inputting arbitrary learning data through the supervised learning, creating a prediction model for predicting a learning result for every set of weights S110.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing method, a program, and an information processing device for reducing the weight of a learning model. Regarding. [Background technology]

[0002] In recent years, research has been conducted into making learning models lighter. For example, Patent Document 1 below describes A technique for reducing the weight of a learning model using parameter quantization is described. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-133628 Summary of the Invention [Problem to be solved by the invention]

[0004] Here, to make the learning model lighter, pruning and quantization are used. , distillation, etc. Learn at least one of these weight reduction techniques. The system is applied to the model, and engineers adjust parameters as necessary to reduce the weight.

[0005] However, there are many methods for reducing the size of training data and learning models. Although it may seem different, the weight reduction method was determined by the engineers adjusting the parameters as needed. However, this was not necessarily the best option.

[0006] Therefore, one of the purposes of this invention is to make the lightweighting method for learning models more appropriate. The present invention provides an information processing method, a program, and an information processing device that enable the above. [Means for solving the problem]

[0007] An information processing method according to one aspect of the present invention includes: The server acquires predetermined learning data, and executes a predetermined learning model using a neural network. For the model, the first learning model after distillation, the second learning model after pruning, and Each model, including at least two models of the quantized third learning model, is weighted The machine learning is performed by inputting predetermined data into the weight learning model thus obtained. For each weight learning model in which the weights of each model have been changed, the predetermined learning data is input. and obtaining a learning result when the machine learning is performed using the weights. Each weight learning model and each learning result when learned with each weight learning model are included. By using the learning data, supervised learning is performed, and any learning Given input data, generate a predictive model that predicts the learning outcome for each set of weights To carry out the following. [Effects of the Invention]

[0008] According to the present invention, it is possible to make the lightweighting method for the learning model more appropriate. It is possible to provide an information processing method, a program, and an information processing device. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of a system configuration according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of a physical configuration of an information processing apparatus according to an embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a processing block of the information processing apparatus according to the embodiment. [Figure 4] FIG. 1 is a diagram for explaining distillation of a trained model. [Figure 5] FIG. 10 is a diagram for explaining pruning of a trained model. [Figure 6] FIG. 2 is a diagram illustrating an example of a processing block of the information processing apparatus according to the embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of relationship information according to the embodiment. [Figure 8] FIG. 10 is a diagram illustrating a display example of relationship information according to the embodiment. [Figure 9] 10 is a flowchart illustrating an example of a process for generating a prediction model according to an embodiment. [Figure 10] 10 is a flowchart illustrating an example of processing in an information processing device used by a user according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] An embodiment of the present invention will be described with reference to the accompanying drawings. Those marked with the same or similar symbols have the same or similar configurations.

[0011] [Embodiment] <System configuration> FIG. 1 is a diagram illustrating an example of a system configuration according to an embodiment. In the example illustrated in FIG. The server 10 and each of the information processing devices 20A, 20B, 20C, and 20D are connected via a network. When the information processing devices are not individually distinguished, the information processing devices are connected so that data can be transmitted and received. It is also referred to as device 20.

[0012] The server 10 is an information processing device capable of collecting and analyzing data, and includes one or more information processing devices. The information processing device 20 may be a smartphone, a personal computer, or the like. Information on which machine learning can be performed, such as computers, tablets, servers, and connected cars The information processing device 20 is an invasive or non-invasive device that senses brain waves. It is also a device that is directly or indirectly connected to the electrodes and can analyze, send and receive EEG data. good.

[0013] In the system shown in FIG. 1, the server 10, for example, Various lightweighting methods (lightweighting algorithms) are applied to the learning model. The method can apply one existing lightweighting method or any combination of lightweighting methods. At this time, the server 10 includes a predetermined data set, a predetermined learning model, and The learning results for a predetermined weight reduction method are stored in association with each other.

[0014] Next, the server 10 receives an arbitrary data set, an arbitrary weighting method, and the learning results thereof. (e.g., learning accuracy) and the learning results are used as training data to predict the appropriate lightweighting method. The model is trained and generated. The appropriateness of the training results is determined by, for example, the training accuracy and the model size. It is determined by factors such as compression rate.

[0015] This makes it possible to more appropriately reduce the weight of trained learning models. The server 10 also uses a model in which each weighting method is weighted and linearly combined to calculate the weight of each weighting method. Each weight that determines the application rate of the quantification method may be adjusted appropriately.

[0016] <Hardware configuration> 2 is a diagram showing an example of the physical configuration of the information processing device 10 according to the embodiment. The processing device 10 includes one or more CPUs (Central Processing Units) 10 corresponding to the arithmetic unit. a, a RAM (Random Access Memory) 10b corresponding to the storage unit, and a R OM (Read Only Memory) 10c, communication unit 10d, input unit 10e, and display unit 10f. These components are connected via a bus so that they can transmit and receive data to and from each other.

[0017] In the embodiment, the information processing device 10 is configured by one computer. However, the information processing device 10 is implemented by combining multiple computers or multiple processing units. 2 is an example, and the information processing device 10 may be implemented in other ways. The present invention may have some of these configurations, or may not have some of these configurations.

[0018] The CPU 10a controls the execution of programs stored in the RAM 10b or the ROM 10c. The CPU 10a is a control unit that controls the weight reduction and performs calculations and processing of data. A learning program using a learning model to investigate the law (learning program) or Learning to generate a predictive model that outputs an appropriate weighting method when inputting data The CPU 10a is a calculation unit that executes a program (prediction program) that performs the above. 10e and communication unit 10d, and displays the results of calculations on the data on display unit 10f. It is displayed and stored in RAM 10b.

[0019] The RAM 10b is a memory unit in which data can be rewritten, and is, for example, a semiconductor memory. The RAM 10b may be configured with a memory element. Lightweighting data on the lightweighting method (e.g., lightweighting algorithm), and the appropriate lightweighting method The predictive model to be measured, information about the data to be trained, and the appropriate lightweight model corresponding to this data It is also possible to store data such as relationship information indicating the correspondence between the image processing method and the image processing method. However, the RAM 10b may store data other than these. It is not necessary for some of them to be stored.

[0020] The ROM 10c is a memory from which data can be read, for example, a semiconductor memory. The ROM 10c may be configured with a memory element. For example, the ROM 10c may store a learning program or It may store data that is not

[0021] The communication unit 10d is an interface that connects the information processing device 10 to other devices. The receiving unit 10d may be connected to a communication network such as the Internet.

[0022] The input unit 10e receives data input from the user, for example, a keyboard. It may include a keyboard and a touch panel.

[0023] The display unit 10f visually displays the results of calculations performed by the CPU 10a. For example, The display unit 10f may be configured by an LCD (Liquid Crystal Display). Displaying the information can contribute to XAI (eXplainable AI). f may display, for example, the learning results or information about the learning model.

[0024] The learning program can be read by a computer such as RAM 10b or ROM 10c. The information may be stored in a storage medium that can be used for storage, or may be provided via a communication network connected by the communication unit 10d. In the information processing device 10, the CPU 10a executes the learning program. By executing the program, various operations are realized, which will be explained later with reference to FIG. These physical configurations are merely examples and do not necessarily have to be independent configurations. For example, the information processing device 10 is an LS in which a CPU 10a is integrated with a RAM 10b and a ROM 10c. The information processing device 10 may be equipped with a GP U (Graphical Processing Unit) and ASIC (Application Specific Integrated Circuit) It may also be provided with

[0025] The configuration of the information processing device 20 is the same as the configuration of the information processing device 10 shown in FIG. The information processing device 10 and the information processing device 20 are data processing devices. The input unit 1 only needs to have a CPU 10a, a RAM 10b, etc., which are basic components for performing the above operations. The input unit 10e and the display unit 10f may not be provided. The unit may be connected to the network using an interface.

[0026] <Processing configuration> 3 is a diagram showing an example of a processing block of the information processing device 10 according to the embodiment. The processing device 10 includes an acquisition unit 101, a first learning unit 102, a change unit 103, a second learning unit 104, A prediction unit 105, a determination unit 106, a setting unit 107, an association unit 108, a specification unit 109, and a display control unit The learning unit 110 includes a control unit 110, an output unit 111, and a storage unit 112. For example, the first learning unit 110 shown in FIG. 102, a change unit 103, a second learning unit 104, a prediction unit 105, a determination unit 106, and a setting unit 107 The associating unit 108, the identifying unit 109, and the display control unit 110 are controlled by, for example, a CPU 10a. The acquisition unit 101 and the output unit 111 are implemented by, for example, the communication unit 10d. The storage unit 112 is realized by the RAM 10b and / or the ROM 10c. obtain.

[0027] The acquisition unit 101 acquires predetermined learning data. For example, the acquisition unit 101 acquires predetermined learning data. As data, publicly known data sets such as image data, sequence data, and text data are collected. The acquiring unit 101 may acquire data stored in the storage unit 112. Alternatively, data transmitted by another information processing device may be acquired.

[0028] The first learning unit 102 is configured to perform a predetermined learning process using a neural network to solve a predetermined problem. The learning model 102a is divided into a first learning model that has been distilled, a second learning model that has been pruned, and a Each model includes at least two models: a training model, and a quantized third training model. The weighted learning model is weighted by inputting the given training data. conduct.

[0029] Here, as an example of a method for reducing the weight of the trained learning model 102a, distillation ), pruning, and quantization algorithms are described below. will be briefly explained below.

[0030] FIG. 4 is a diagram for explaining the distillation of a trained model. The distillation shown in FIG. The prediction results of the model M11 are used as training data to train a smaller model M12. At this time, the smaller model M12 is about the same weight as the larger model M11. It may have precision.

[0031] For example, in distillation, the trained model M11 is a Teacher model, a small model, The M12 is called the Student model. Design appropriately.

[0032] In the example shown in Figure 4, learning data is explained using a classifier as an example. Eacher models are trained using training data expressed as 0 and 1, with 1 being the correct answer. In contrast, the Student model of model M12 is The output values ​​(e.g., A=0.7, B=0.3) are used as training data for learning. For one training model M11, multiple different distilled models M12 may be prepared. stomach.

[0033] FIG. 5 is a diagram for explaining pruning of a trained model. By deleting weights and nodes from the trained model M21, the lighter model M22 is obtained. This makes it possible to reduce the number of calculations and memory usage.

[0034] The pruning method is to remove connections between nodes with low weights. For example, unlike distillation, pruning does not require a separate model to be designed, but the parameter Therefore, it is advisable to retrain to maintain the learning accuracy. Even if the weight is reduced by cutting branches (edges) with low impact, for example, branches with weights below a certain value, good.

[0035] Quantization is the process of expressing the parameters contained in the model using fewer bits. This makes it possible to make the model smaller without changing the network structure. For example, For example, in a simple network with six data points, the total number of data points is 192 for 32-bit precision. 8 bits are required, but if we limit the precision to 8 bits, it can be expressed in a total of 48 bits. This means that the weight has been reduced.

[0036] Returning to FIG. 3, for example, the first learning unit 102 performs the following on the trained learning model 102a: At least two lightweight models are selected from the first, second, and third models. The weights assigned to each model are set to default values.

[0037] The first, second, and third models are pre-trained for each category of trained models. It may be set in advance or automatically generated according to predetermined criteria for each trained model. For example, in the case of distillation, the first learning unit 102 may The model may be determined by machine learning, and in the case of pruning, branches with weights below a predetermined value are cut. In the case of quantization, the constraint of the predetermined value bit precision (quantization ) can be used. Also, for one trained model, multiple first models and multiple second models can be used. Alternatively, multiple third models may be set and weights may be assigned to each model.

[0038] The predetermined problem may be, for example, at least one of image data, sequence data, and text data. This includes problems of classification, generation, and / or optimization. Image data includes still image data and video data. Sequence data includes audio data and stock data. Includes value data.

[0039] The predetermined learning model 102a is a trained learning model including a neural network. For example, image recognition models, sequence data analysis models, robot control models, Reinforcement learning model, speech recognition model, speech generation model, image generation model, natural language processing model As a specific example, the predetermined learning model 102a includes at least one of C NN (Convolutional Neural Network), RNN (Recurrent Neural Network), DNN (Deep Neural Network), LSTM (Long Short-Term Memory), Bidirectional LSTM, D QN (Deep Q-Network), VAE (Variational AutoEncoder), GANs (Generative Adversarial Networks), flow-based generative models, etc.

[0040] The modification unit 103 modifies predetermined learning data and / or each weight of the weight learning model. For example, the change unit 103 may select a predetermined value from among a plurality of pieces of learning data input to the first learning unit 12. The learning data is changed one by one. The changing unit 103 also changes a certain weight learning model When all the predetermined learning data is input and learning is performed, the weight learning model is To use each weight, select one set from multiple sets of weights. Learning may be performed using all the prepared sets to obtain the learning results.

[0041] The first learning unit 102 inputs predetermined learning data into the weight learning model and performs appropriate learning. The hyperparameters of the weight learning model are learned so that the results can be output. When the hyperparameters are updated (adjusted), the first learning unit 102 adjusts the weight learning model. The weights assigned to each model in the rule are also adjusted by a predetermined method.

[0042] For example, the weights are adjusted sequentially from their initial values. At this time, the weights are adjusted so that they all add up to 1, and the previous Any adjustment method may be used as long as the adjustment is different from the adjustment described above. The learning unit 102 sequentially changes each weight by a predetermined value, and performs the changes for all combinations. For example, the first learning unit 102 sets the weight w k The initial value is subtracted by a predetermined value, and the overlap Miw k+1 When either weight becomes 0 or less, k Add 1 to each weight and repeat the change from the initial value. Also, all weights add up to 1. There is no need to set conditions. In this case, the weights are added together using a Softmax function or the like. It should then be finally adjusted to become 1.

[0043] This allows us to calculate the weights for any combination of the training data and the weights. For example, the change unit 103 may change predetermined learning data and predetermined weights. The predetermined training data and / or the predetermined Each set of weights in the set may be changed one by one, or a predetermined condition may be satisfied. The set of training data and / or each of the predetermined weights may be changed one by one in turn. may be set based on, for example, the learning accuracy or the compression rate of the model size.

[0044] The acquisition unit 101 or the first learning unit 102 performs weight learning in which the weights of each model are changed. For each model, the learning results are obtained when machine learning is performed by inputting the specified learning data. For example, the acquisition unit 101 or the first learning unit 102 acquires various combinations of predetermined learning data. The training results are obtained by training using the data and / or predetermined sets of weights.

[0045] Here, the weight learning model will be described using a specific example. For example, the first learning unit 102 is a line that assigns weights w1, w2, and w3 to the first, second, and third models, respectively. A weight learning model with a combined weight function M(x) can be used. , Equation (1) is given as an example only. M1(x)=w1m1(x)+w2m2(x)+w3m3(x)...Equation (1) w n : Weight (The set of weights is also denoted as W) m n (x): nth model x: training data

[0046] The change unit 103 changes each weight according to a predetermined standard so that, for example, w1+w2+w3=1. The first learning unit 102 changes the weights one by one. The training results are obtained by associating the training accuracy and the training result with each set of weights. This is the compression ratio of the model size, which indicates the effect of weight reduction. For example, the number of parameters in the trained model after weight reduction is the ratio of the number of parameters in the trained model before weight reduction. This is the ratio to the number of items.

[0047] Furthermore, when the change unit 103 changes the predetermined learning data, the first learning unit 102 changes the predetermined learning data. For the training data, we train a weight learning model for each set of weights as described above. This allows us to obtain the training results for any training data, any set of weights, and Training data containing the learning results for each case is generated.

[0048] The second learning unit 104 learns each weight learning model to which the changed weights are assigned, and each weight Supervised learning is performed using training data that includes the results of each learning process. For example, the second learning unit 104 uses any set of learning data and any set of weights. The learning results (e.g., learning performance and / or model size compression rate) when trained using the correct answer Supervised learning is performed using training data as a baseline.

[0049] Furthermore, the second learning unit 104 performs supervised learning to learn the following when inputting arbitrary learning data: Then, a prediction model 104a is generated to predict the learning result for each set of weights. 2. When any training data is input, the training unit 104 performs each weighting procedure for the training data. For each set of weights in the algorithm, a predictive model is generated that outputs the training accuracy and the model size compression ratio. Complete.

[0050] With the above configuration, various training data and various lightweight training models can be used. By performing supervised learning using the learning results from the model as training data, the set of weights is For each learning task, a predictive model can be generated to predict the learning outcome. The prediction model generated by the unit 104 is used to improve the weight reduction method. This becomes possible.

[0051] The prediction unit 105 inputs arbitrary learning data into the prediction model 104a, and each model For each set of weights, predict the learning outcome when the weight learning model is run. For example, When a data set of images is input as training data, the prediction unit 105 predicts a specific weight set. ToW n (w 1n ,w 2n ,w3n ) for each parameter, we calculate the training accuracy and model size (e.g. Compression rate).

[0052] This allows you to determine how effective each lightweighting technique is for any data (e.g., a dataset). For each set of weights applied, a learning outcome is predicted, and based on this learning outcome, This makes it possible to select more appropriate weights.

[0053] The determining unit 106 determines the learning result when any learning data is input to the predetermined learning model 102a. The result of the weight reduction and the learning result predicted by the prediction model 104a satisfy predetermined conditions for weight reduction. For example, the determination unit 106 determines whether the weighting factor satisfies the pre-weighting learning model 1. Learning accuracy A1 when learning data A is input to 02a and prediction by prediction model 104a It is determined whether or not a first difference value between the training accuracy B1 and the training accuracy B2 is within a first threshold value. The smaller the value, the more accurate the learning can be maintained even after the learning model is made lighter. Therefore, when the learning accuracy is B1, each weight is an appropriate weight reduction method.

[0054] In addition, the determination unit 106 determines the compression rate A2 (=1 ) and the compression ratio B2 predicted by the prediction model 104a is equal to or greater than a second threshold value. The larger this second difference value, the lighter the learning model becomes. This shows that it can be quantified.

[0055] The determination unit 106 determines the validity of each weight based on the determination result regarding weight reduction. For example, the determining unit 106 determines whether the compression ratio B2 is large and the learning rate is low based on the first difference value and the second difference value. The training accuracy B1 is an effective weighting method for each weight that can maintain the accuracy before weighting. As a specific example, the determining unit 106 determines whether the first difference value is equal to or smaller than the first threshold value and whether the second difference value is equal to or smaller than the first threshold value. Weights above the second threshold are treated as effective weight reduction methods, and other weights are treated as ineffective weight reduction methods. It may be determined that:

[0056] This allows us to calculate the model size (e.g., compression ratio) and training accuracy based on the model size. By referring to these predicted values, the determining unit 106 can select appropriate weights. The weights with the highest learning accuracy may be selected, or the weights with the highest compression rate above the second threshold may be selected. The weight with the highest degree may be selected.

[0057] The setting unit 107 accepts a user operation regarding a predetermined condition related to weight reduction. For example, The setting unit 107 receives the condition input from the condition input screen displayed on the display unit 10f by the user through the input unit 10e. When a predetermined condition related to weight reduction is input by operation, this input operation is accepted.

[0058] The setting unit 107 determines predetermined conditions related to weight reduction based on the received user operation. For example, the setting unit 107 may set the following as a determination condition based on an input operation by the user: Setting a first threshold for training performance and / or a second threshold for model size may be possible.

[0059] This allows the user to specify effective weight reduction methods using the desired conditions. It becomes possible.

[0060] The associating unit 108 uses the learning accuracy included in the learning result as a first variable and the model included in the learning result as a second variable. The value related to the data size (for example, the compression ratio) is set as the second variable, and the first and second variables and each weight are used. For example, the associating unit 108 may create relationship information that associates the first variable with the vertical axis. When the horizontal axis is the second variable, a matrix in which each weight W is associated with the intersection of each variable is created. The association unit 108 may generate a learning log based on the learning log acquired from each information processing device 20. Based on the training accuracy and compression rate, relational information that associates the first and second variables with each weight W is generated. (Actual measurement related information) may be generated.

[0061] By the above process, when the first variable or the second variable is changed, the corresponding weights W are quickly adjusted. Furthermore, the first and second variables may be changed as appropriate. For example, if we apply the learning accuracy as the first variable and the weights W as the second variable, the information identified will be It may also be a value related to the model size.

[0062] The acquiring unit 101 may also acquire a first value of the first variable and a second value of the second variable. For example, the acquiring unit 101 acquires a first value of a first variable and a second value of a second variable designated by a user. The first value or the second value is appropriately designated by the user.

[0063] In this case, the identifying unit 109 determines, based on the relationship information generated by the associating unit 108, Each weight W corresponding to a first value of the first variable and a second value of the second variable is identified. For example, 109 uses the relationship information to determine whether the value of the first variable or the value of the second variable is changed. Specify the weight W.

[0064] The display control unit 110 displays each weight W identified by the identification unit 109 on a display device (display unit 10 f) The display control unit 110 controls the display in such a manner that the first variable and the second variable are changeable. The matrix may be displayed using a GUI (Graphical User Interface) (for example, (See Figure 8, etc.)

[0065] By the above process, the first variable or the second variable specified by the user is identified. Each weight W can be visualized for the user. The user can then visualize the first or second variables. By changing the variables, we can identify the desired weights W and apply them to reduce the weight of the trained model. This can be done.

[0066] The output unit 111 outputs each weight W predicted by the second learning unit 104 to the other information processing device 2. For example, the output unit 111 may output the predetermined learning data to the information processing device that transmitted the predetermined learning data. The information processing device 20 that has requested the acquisition of appropriate weights W is The output unit 111 may output appropriate weights W corresponding to the data. Each weight W may be output to the storage unit 112 .

[0067] The storage unit 112 stores data related to learning. The storage unit 112 stores a predetermined data set. The data 112a, the data on the weight reduction method 112b, the above-mentioned related information 112c, the training data It stores information such as data, data during learning, and information about learning results.

[0068] 6 is a diagram showing an example of a processing block of the information processing device 20 according to the embodiment. The processing device 20 includes an acquisition unit 201, a learning unit 202, an output unit 203, and a storage unit 204. The information processing device 20 may be configured as a general-purpose computer.

[0069] The acquisition unit 201 receives a distributed learning instruction from another information processing device (for example, the server 10). Also, information about a given weight learning model and information about a given dataset can be obtained. The information about the predetermined weight learning model may include information indicating each weight and the weight learning model itself. The information about a given dataset may be the dataset itself. Alternatively, the information may be information indicating the storage location where a predetermined data set is stored.

[0070] The learning unit 202 inputs a predetermined data set to be learned into a predetermined weight learning model 202a. The learning unit 202 feeds back the learning results to the server 10 after the learning. The learning results include, for example, learning performance, and information about the model size. The learning unit 202 may further include information about the type of dataset to be learned, and / or Depending on the problem to be solved, the learning model 202a may be selected.

[0071] The predetermined weight learning model 202a is a learning model including a neural network. For example, image recognition models, sequential data analysis models, robot control models, reinforcement Learning models, speech recognition models, speech generation models, image generation models, natural language processing models, etc. Each weighting method is based on at least one of the following models. The base of the predetermined weight learning model 202a is a CNN (Convolutional Neural Network) network), RNN (Recurrent Neural Network), DNN (Deep Neural Network), L STM (Long Short-Term Memory), bidirectional LSTM, DQN (Deep Q-Network), VA E(Variational AutoEncoder), GANs(Generative Adversarial Networks), fl It may be either a power-based generative model or a power-based generative model.

[0072] The output unit 203 outputs information about the learning results of the distributed learning to other information processing devices. For example, the output unit 203 outputs information about the learning result by the learning unit 202 to the server 10. For example, information about the learning results of distributed learning includes the learning performance as described above. It may further include information regarding the delta size.

[0073] The storage unit 204 stores data related to the learning unit 202. The storage unit 204 stores predetermined data. The data set 204a, data acquired from the server 10, data during learning, and learning results Stores information related to the

[0074] As a result, the information processing device 20 receives instructions from other information processing devices (for example, the server 10). This performs distributed learning by applying a predetermined weight learning model to a given dataset. This makes it possible to feed back the learning results to the server 10.

[0075] The output unit 203 also outputs information about the predetermined data to another information processing device (for example, a server 10). The output unit 203 outputs predetermined data (for example, a data set to be learned) to Alternatively, characteristic information of the predetermined data may be output.

[0076] The acquisition unit 201 acquires each weight W corresponding to predetermined data from another information processing device. Each weight W to be acquired may be a predetermined weight W predicted by another information processing device using a prediction model. The weights are appropriate for the given data.

[0077] The learning unit 202 applies the acquired weights to the weight learning model 202a. The weight learning model 202a is a model that assigns weights to the weight learning model 22a used in the above-mentioned learning. The weight learning model 202a may be acquired from another information processing device 10. The learning model may be a model that is managed by the device itself, or may be a learning model that is managed by the device itself.

[0078] The learning unit 202 inputs predetermined data to the weight learning model 202a to which each weight is applied. The learning results are obtained by learning using weights appropriate for the given data. The learning unit 202 uses a learning model that is appropriately lightweight while maintaining learning performance. It can be used. <Data example> 7 is a diagram illustrating an example of relationship information according to the embodiment. is the first variable (e.g., P 11 ) and each secondary variable (e.g., P 21 ) corresponding to each weight (e.g., W 1) First variable P 1n is, for example, the training accuracy, and the second variable P 2n For example, the model The compression ratio of the size, and only one of the variables can be used as a variable. (P1n, P2m) is the first variable P 1n and the second variable P 2n is the weight in the case of

[0079] Regarding the relationship information shown in FIG. 7, the server 10 determines the number of distributed instances and the number of hyperparameters. The information processing device 20 performs distributed learning using a combination of parameters, or the teacher of the device itself. From the results of the training, the training accuracy (first variable) and the compression rate (second variable) are obtained. The server 10 associates each weight W with the acquired learning accuracy and compression rate. By performing this process every time the training accuracy and compression rate measured by training are acquired, the The relationship information can be generated. Based on the results, predictive relationship information for any data set may be generated.

[0080] <User interface example> 8 is a diagram showing a display example of relationship information according to the embodiment. In the example shown in FIG. The first and second variables included in the information can be changed using a slide bar. By using the slide bar to move the first or second variable, for example, First variable (P 1n ) or a second variable (P 2m ) corresponding to each weight W set W (P1n,P 2m) are displayed in association with the corresponding points.

[0081] The user can also specify a predetermined point on a two-dimensional graph of the first and second variables. The combination of training accuracy and compression rate corresponding to the specified point is displayed. Good too.

[0082] As a result, the server 10 can select appropriate combinations of the first and second variables. It is possible to display the weight W. It also shows the correspondence visually to the user, and It is possible to select the appropriate number of distributed instances and hyperparameters for any dataset on which training is performed. This makes it possible to provide a user interface that allows the user to select data.

[0083] <Operation> FIG. 9 is a flowchart illustrating an example of a process related to generation of a prediction model according to the embodiment. The process shown in FIG.

[0084] In step S102, the acquisition unit 101 of the information processing device 10 acquires predetermined learning data. The predetermined learning data may be selected from the data set 112a in the storage unit 112. It may be predetermined data received from another device via a network, or it may be data that the user Predetermined data input in response to an operation may be acquired.

[0085] In step S104, the first learning unit 102 of the information processing device 10 performs neural network For a given learning model using a network, the first learning model that has been distilled, the first learning model that has been pruned, At least two models are used: a second learning model that has been processed by quantization, and a third learning model that has been quantized. Each model, including the model, is weighted, and the given data is input to the weight learning model. Do machine learning.

[0086] In step S106, the second learning unit 104 of the information processing device 10 performs the following for each model: For each weight learning model with the changed weights, machine learning is performed by inputting the given learning data. Obtain the learning results when

[0087] In step S108, the second learning unit 104 of the information processing device 10 Each weighted learning model and each learning result when trained with each weighted learning model Supervised learning is performed using training data including the results.

[0088] In step S110, the second learning unit 104 of the information processing device 10 performs supervised learning. Therefore, when inputting arbitrary training data, it is possible to predict the training results for each combination of weights. Generate a predictive model that

[0089] By using the prediction model generated through the above process, the neural network It is now possible to more appropriately reduce the weight of trained models that use do.

[0090] FIG. 10 shows an example of processing in the information processing device 20 used by the user according to the embodiment. In step S202, the output unit 203 of the information processing device 20 The information about the predetermined learning data to be learned is transmitted to another information processing device (for example, the server 10). Output to.

[0091] In step S204, the acquisition unit 201 of the information processing device 20 acquires the information from the other information processing device ( For example, information indicating each weight corresponding to predetermined learning data is obtained from the server 10).

[0092] In step S206, the learning unit 202 of the information processing device 20 calculates the weights obtained as The predetermined weights are applied to the learning model 202a.

[0093] In step S208, the learning unit 202 of the information processing device 20 Predetermined learning data is input to the learning model 202a to obtain the learning results.

[0094] This allows even edge-side information processing devices to generate appropriate By using a lightweight learning model for learning, it is possible to maintain learning accuracy.

[0095] The above-described embodiments are intended to facilitate understanding of the present invention and are not intended to limit the present invention. The elements of the embodiment and their arrangement, materials, conditions, etc. The shape and size are not limited to those shown in the examples and can be changed as needed. In addition, the device including the first learning unit 102 and the device including the second learning unit 104 are different computers. In this case, the learning result generated by the first learning unit 102 is may be transmitted to the device including the second learning unit 104 via the network.

[0096] Furthermore, the information processing device 10 does not necessarily have to include the change unit 103. The processing device 10 acquires the learning performance of each pair of data to be learned and a set of weights. Additionally, learning may be performed by the second learning unit 104. [Explanation of symbols]

[0097] 10...information processing device, 10a...CPU, 10b...RAM, 10c...ROM, 10d...communication a signal unit, 10e... an input unit, 10f... a display unit, 101... an acquisition unit, 102... a first learning unit, 102 a...Learning model, 103...Modification unit, 104...Second learning unit, 104a...Prediction model, 105 ...prediction unit, 106...determination unit, 107...setting unit, 108...association unit, 109...identification unit, 1 10...display control unit, 111...output unit, 112...storage unit, 112a...data set, 112 b...Lightweighting method, 112c...Relationship information, 201...Acquisition unit, 202...Learning unit, 202a...Learning Learning model, 203...output unit, 204...memory unit, 204a...data set

Claims

1. One or more processors included in the information processing device Obtaining information about a predetermined weight learning model and information about a predetermined dataset, the predetermined weight learning model being a weight learning model in which each model is weighted, the weight learning model including at least two models of a first learning model that has been distilled, a second learning model that has been pruned, and a third learning model that has been quantized, for a predetermined learning model using a neural network; performing machine learning by inputting the predetermined data set into the predetermined weight learning model; Outputting the learning result of the machine learning; An information processing method.

2. An information processing device having one or more processors, The one or more processors: Obtaining information about a predetermined weight learning model and information about a predetermined dataset, the predetermined weight learning model being a weight learning model in which each model is weighted, the weight learning model including at least two models of a first learning model that has been distilled, a second learning model that has been pruned, and a third learning model that has been quantized, for a predetermined learning model using a neural network; performing machine learning by inputting the predetermined data set into the predetermined weight learning model; Outputting the learning result of the machine learning; An information processing device that executes the above.

3. One or more processors included in the information processing device Obtaining information about a predetermined weight learning model and information about a predetermined dataset, the predetermined weight learning model being a weight learning model in which each model is weighted, the weight learning model including at least two models of a first learning model that has been distilled, a second learning model that has been pruned, and a third learning model that has been quantized, for a predetermined learning model using a neural network; performing machine learning by inputting the predetermined data set into the predetermined weight learning model; Outputting the learning result of the machine learning; A program to execute.