Information processing method, and information processing system
The information processing method addresses the challenge of suppressing quantization losses by training a larger inference model and outputs a third model that maintains performance, effectively addressing the inconsistency between reference and embedded environment networks.
Patent Information
- Application Number
- JP2025029919
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-03-03
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2041-05-28
AI Technical Summary
Existing methods for finding inference models struggle to suppress losses caused by quantization, leading to deteriorated inference performance and inconsistencies between reference and embedded environment networks.
An information processing method that acquires a first inference model, calculates a second inference model with a larger model size, reduces it to generate a third inference model, trains the third model using machine learning, and outputs it if its performance meets the specified conditions.
This method effectively finds an inference model that minimizes losses due to quantization, maintaining performance close to the reference model while being suitable for embedded environments.
Smart Images

Figure 2025075101000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an information processing method executed by a computer. [Background technology]
[0002] Non-Patent Document 1 proposes a method for searching a network structure regarding a neural network model. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Yu Liu et al., “Search to Distill: Pearls are Everywhere but not the Eyes”, CVPR(Computer Vision and Pattern Recognition) 2020 Summary of the Invention [Problem to be solved by the invention]
[0004] However, with the method proposed in Non-Patent Document 1, it is difficult to find an inference model that suppresses the loss caused by quantization.
[0005] Therefore, the present disclosure provides an information processing method, etc. that can find an appropriate inference model. [Means for solving the problem]
[0006] An information processing method according to one embodiment of the present disclosure is an information processing method executed by a computer, which obtains a first inference model as a reference, determines a second inference model having a larger model size than the first inference model, reduces the weight of the determined second inference model to generate a third inference model, trains the third inference model using machine learning, determines whether the performance of the trained third inference model satisfies a condition, and if the performance satisfies the condition, outputs the trained third inference model.
[0007] In addition, these comprehensive or specific aspects may be realized by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium. Effect of the Invention
[0008] An information processing method according to one aspect of the present disclosure makes it possible to find an appropriate inference model. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 is a conceptual diagram showing deterioration of detection accuracy in the first reference example. [Diagram 2] FIG. 2 is a conceptual diagram showing a comparison between a reference network and a network for an embedded environment in the first reference example. [Diagram 3] FIG. 3 is a conceptual diagram illustrating a comparison between the reference network and the network for an embedded environment in the first embodiment. [Figure 4] FIG. 4 is a block diagram showing a configuration of an information processing system in the second reference example. [Diagram 5] FIG. 5 is a flowchart showing the operation of the information processing system in the second reference example. [Figure 6] FIG. 6 is a block diagram illustrating an example of a configuration of the information processing system according to the first embodiment. As shown in FIG. [Figure 7] FIG. 7 is a flowchart showing a first operation example of the information processing system according to the first embodiment. [Figure 8] FIG. 8 is a flowchart showing a second operation example of the information processing system according to the first embodiment. [Figure 9] FIG. 9 is a graph showing a size difference evaluation function in the second embodiment. [Figure 10] FIG. 10 is a flowchart showing a first operation example of the information processing system in the second embodiment. [Figure 11] FIG. 11 is a flowchart showing a second operation example of the information processing system according to the second embodiment. [Figure 12] FIG. 12 is a block diagram showing an example of a configuration of an information processing system according to the third embodiment. As shown in FIG. [Figure 13] FIG. 13 is a flowchart showing a first phase of an example of the operation of the information processing system according to the third embodiment. [Figure 14] FIG. 14 is a flowchart showing the second and third phases of an operation example of the information processing system according to the third embodiment. [Figure 15] FIG. 15 is a block diagram showing a basic implementation example of an information processing system according to a plurality of embodiments. [Figure 16] FIG. 16 is a flowchart showing an example of a basic operation of the information processing system according to a plurality of embodiments. [Figure 17] FIG. 17 is a block diagram showing another basic implementation example of an information processing system according to multiple embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] (Findings on which this disclosure is based) Inference functions based on deep learning may be incorporated into Internet of Things (IoT) devices. In addition, from the viewpoint of cost and privacy, inference processing may be performed by a processor on the device rather than in a cloud or GPU environment. In these cases, the network (NW) is made lighter using methods such as quantization. This allows inference processing based on deep learning to be performed by a processor with limited computing resources such as computing power and memory capacity.
[0011] In this case, the network refers to an inference model such as a neural network model for performing inference processing including an inference process.
[0012] However, for example, quantization converts a reference network (RefNW) in floating-point representation into an embedded network (IntNW) in fixed-point representation. Such quantization may cause loss of inference performance. Specifically, it may result in a loss of accuracy or a mismatch in inference results between the reference network and the embedded network.
[0013] Non-Patent Document 1 proposes a method for searching for a network structure regarding a neural network model. In the method proposed in Non-Patent Document 1, a network structure with high inference performance and fast inference speed is searched for. The network structure corresponds to the number of layers, the number of nodes in each layer, the connection form between the nodes, etc. In other words, in the method proposed in Non-Patent Document 1, a network structure having the number of layers, the number of nodes in each layer, the connection form between the nodes, etc. with high inference performance and fast inference speed is searched for.
[0014] However, the method proposed in Non-Patent Document 1 does not take into account the loss caused by quantization. Therefore, even if the method proposed in Non-Patent Document 1 is used, there is a possibility that a loss in inference performance will occur due to quantization for reducing the network weight.
[0015] Therefore, for example, an information processing method according to one embodiment of the present disclosure is an information processing method executed by a computer, which obtains a first inference model as a reference, calculates a second inference model having a larger model size than the first inference model based on the first inference model, quantizes the calculated second inference model to generate a third inference model, trains the third inference model using machine learning, determines whether the performance of the trained third inference model satisfies a condition, and if the performance satisfies the condition, outputs the trained third inference model.
[0016] As a result, a second inference model having a larger model size than the first inference model is quantized. It is expected that performance will not deteriorate easily even if the second inference model having a larger model size is quantized. In other words, it is expected that the loss caused by quantization will be relatively small in the third inference model generated by quantizing the second inference model having a larger model size than the first inference model. Therefore, it is possible to find an inference model in which the loss caused by quantization is suppressed.
[0017] Also, for example, the information processing method further includes acquiring setting information indicating quantization settings for the second inference model, and setting initial values for the calculation of the second inference model based on the setting information and the first inference model.
[0018] This starts the calculation of the second inference model based on the quantization setting information and the first inference model, making it possible to quickly find the third inference model according to the quantization and the first inference model.
[0019] Also, for example, the information processing method further acquires difficulty information indicating the inference difficulty of at least one of the first inference model, the second inference model, and the third inference model, and sets an initial value for the calculation of the second inference model based on the difficulty information and the first inference model.
[0020] This starts the calculation of the second inference model based on the inference difficulty information and the first inference model, making it possible to quickly find a third inference model that corresponds to the inference difficulty and the first inference model.
[0021] Also, for example, the calculation of the second inference model is a search of the second inference model using a loss function, and the loss function is a function whose output value becomes smaller when the difference between the inference result of the first inference model and the inference result of the third inference model becomes smaller, and whose output value becomes smaller when the model size of the second inference model becomes relatively larger with respect to the first inference model, and the search of the second inference model is performed so as to reduce the output value of the loss function.
[0022] This makes it possible to find an inference model based on the loss function such that the loss caused by quantization is suppressed.
[0023] Also, for example, the information processing method further includes acquiring setting information indicating quantization settings for the second inference model, and changing the loss function based on the setting information.
[0024] This makes it possible to find an inference model that suppresses the loss caused by quantization, based on a loss function that corresponds to the quantization settings.
[0025] Also, for example, the loss function is changed so that the output value of the loss function becomes larger as the degree of quantization in the settings indicated by the setting information increases, and a search for the second inference model is performed so that the output value of the loss function is equal to or less than a threshold value.
[0026] As a result, the greater the degree of quantization, the greater the output value of the loss function, but the second inference model is searched for so that the output value of the loss function is equal to or less than the threshold. In other words, even if the loss is large due to the large degree of quantization, the second inference model is searched for so that certain conditions for suppressing the loss are satisfied. Therefore, it is possible to find an inference model that suppresses the loss at a certain level.
[0027] Also, for example, the information processing method further acquires difficulty information indicating the inference difficulty of at least one of the first inference model, the second inference model, and the third inference model, and changes the loss function based on the difficulty information.
[0028] This makes it possible to find an inference model that suppresses the loss caused by quantization based on a loss function that corresponds to the difficulty of inference.
[0029] Also, for example, the loss function is changed so that the output value of the loss function becomes larger as the inference difficulty indicated by the difficulty information becomes higher, and a search for the second inference model is performed so that the output value of the loss function is below a threshold value.
[0030] As a result, the higher the inference difficulty, the larger the output value of the loss function, but the second inference model is searched for so that the output value of the loss function is below the threshold. In other words, even if the loss is large because of the high inference difficulty, the second inference model is searched for so that certain conditions for suppressing the loss are satisfied. Therefore, it is possible to find an inference model that suppresses the loss at a certain level.
[0031] Also, for example, the information processing method further includes changing quantization settings for the second inference model if the performance does not satisfy the condition.
[0032] This allows the quantization settings to be changed so that the performance criteria are met, and it is then possible to find an inference model that satisfies the performance criteria.
[0033] Also, for example, the conditions include the inference result of the first inference model or the precision or accuracy of the inference of the third inference model against reference data, and the change in the settings reduces the degree of quantization when the precision or accuracy of the inference of the third inference model is below a threshold value.
[0034] As a result, when the inference precision or accuracy of the third inference model is equal to or lower than the threshold, the quantization for the second inference model is narrowed so that the inference precision or accuracy of the third inference model is increased, and thus it becomes possible to find an inference model such that the condition of the inference precision or accuracy is satisfied.
[0035] Also, for example, the conditions include the speed of the inference processing of the third inference model, and the change in the settings involves increasing the degree of quantization when the speed of the inference processing is below a threshold value.
[0036] As a result, when the inference processing speed of the third inference model is equal to or lower than the threshold, the degree of quantization for the second inference model is increased so that the inference processing speed of the third inference model is increased, and thus it becomes possible to find an inference model that satisfies the condition of the inference processing speed.
[0037] Also, for example, the information processing method further includes inputting data into the first inference model to obtain an inference result of the first inference model, inputting the data into the second inference model to obtain an inference result of the second inference model, and training the first inference model based on a difference between the inference result of the first inference model and the inference result of the second inference model.
[0038] As a result, the first inference model and the third inference model are constructed based on the same second inference model, making it possible to reduce the difference between the inference result of the first inference model and the inference result of the third inference model.
[0039] Also, for example, an information processing system according to one embodiment of the present disclosure includes at least one processor and at least one memory, and the at least one processor uses the at least one memory to obtain a first inference model as a reference, calculates a second inference model having a model size larger than the first inference model based on the first inference model, quantizes the calculated second inference model to generate a third inference model, trains the third inference model using machine learning, determines whether the performance of the trained third inference model satisfies a condition, and if the performance satisfies the condition, outputs the trained third inference model.
[0040] As a result, a second inference model having a larger model size than the first inference model is quantized. It is expected that performance will not deteriorate easily even if the second inference model having a larger model size is quantized. In other words, it is expected that the loss caused by quantization will be relatively small in the third inference model generated by quantizing the second inference model having a larger model size than the first inference model. Therefore, it is possible to find an inference model in which the loss caused by quantization is suppressed.
[0041] Furthermore, these comprehensive or specific aspects may be realized in a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized in any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
[0042] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component arrangement and connection forms, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Furthermore, among the components in the following embodiments, components that are not described in the independent claims will be described as optional components.
[0043] In addition, in the present disclosure, ordinal numbers such as first, second, and third may be attached to elements. These ordinal numbers are attached to elements to identify the elements and do not necessarily correspond to a meaningful order. These ordinal numbers may be rearranged, newly added, or removed as appropriate.
[0044] Additionally, for purposes of this disclosure, inference can include detecting, recognizing, identifying, etc. Additionally, for purposes of this disclosure, computing can include determining, searching, obtaining, deriving, and extracting, etc.
[0045] In addition, in the present disclosure, distilling the network NW1 to train the network NW2 means, for example, training the network NW2 using the network NW1 as a teacher network. In addition, in the present disclosure, training the network means, for example, adjusting the parameters of the network. Training the network may be interpreted as learning the network. In addition, the network may be interpreted as an inference model.
[0046] (Embodiment 1) 1 is a conceptual diagram showing deterioration of detection accuracy in the first reference example. Specifically, the difference between the detection result in the reference network and the detection result in the network for embedded environments is shown.
[0047] For example, the reference network is expected to be used in a cloud environment or a GPU environment and is constructed using a floating-point representation. On the other hand, the network for an embedded environment is expected to be installed in an IoT device or the like and is constructed using a fixed-point representation. Basically, the reference network is first constructed by training in a cloud environment or a GPU environment. After that, the reference network is converted into a network for an embedded environment.
[0048] Due to limited resources in embedded environments, when a reference network is converted to a network for embedded environments, the network is lightweighted, which includes quantization to convert floating-point representation to fixed-point representation. This lightweighting leads to degradation of detection accuracy.
[0049] Specifically, as shown in Figure 1, the reference network detects a dog, a person, and a horse from the image. The network for embedded environments detects a dog and a person from the image, but does not detect a horse. That is, the network for embedded environments has a deterioration in detection accuracy compared to the reference network. That is, the network for embedded environments has more false negatives (FN) and false positives (FP) than the reference network.
[0050] FIG. 2 is a conceptual diagram showing a comparison between a reference network and a network for an embedded environment in the first reference example. The network for an embedded environment is constructed by reducing the weight of the reference network. This reduction in weight causes a deterioration in inference accuracy. In other words, the inference accuracy of the network for an embedded environment is lower than that of the reference network.
[0051] Therefore, the network for embedded environments may not achieve the desired performance. Also, there may be a discrepancy in inference results between the reference network and the network for embedded environments, which may increase the amount of work required to evaluate and verify the network for embedded environments.
[0052] 3 is a conceptual diagram showing a comparison between a reference network and a network for an embedded environment in this embodiment. In this embodiment, a network for an embedded environment having an inference accuracy close to that of the reference network is constructed.
[0053] This makes it possible to obtain the desired performance in the network for an embedded environment. In addition, discrepancies in inference results between the reference network and the network for an embedded environment are reduced, and the amount of work required for evaluating and verifying the network for an embedded environment is reduced.
[0054] 4 is a block diagram showing the configuration of an information processing system in the second reference example. Specifically, the configuration of an information processing system 100 is shown, which is assumed from Non-Patent Document 1. The information processing system 100 includes a network searching unit 101, an evaluation value calculation unit 103, and a learning processing unit 104.
[0055] The network search unit 101 searches for a second network 112 that is likely to have improved inference accuracy and inference speed by distilling the first network 111, which is a reference network. This process corresponds to obtaining an evaluation value from the evaluation value calculation unit 103 and searching for a second network 112 that will improve the evaluation value.
[0056] Each of the first network 111 and the second network 112 is an inference model such as a neural network model for performing inference processing. The second network 112 is expected to have a faster inference speed than the first network 111. Therefore, the size of the second network 112 to be searched is assumed to be smaller than the size of the first network 111.
[0057] The size of a network corresponds to the number of nodes, layers, parameters, connections between nodes, etc. included in the network. The larger the number of nodes, layers, parameters, connections between nodes, etc. included in the network, the larger the size of the network. The size of a network may correspond to any one of the number of nodes, layers, parameters, and connections between nodes included in the network.
[0058] The evaluation value calculation unit 103 calculates an evaluation value. Specifically, the evaluation value calculation unit 103 acquires the inference result of the first network 111 and the inference result of the second network 112, and calculates the output value of a loss function related to inference accuracy and inference speed as the evaluation value. With this loss function, the output value increases as the difference between the inference result of the first network 111 and the inference result of the second network 112 increases, and the output value increases as the processing delay of the second network 112 increases. The smaller the evaluation value corresponding to the output value of the loss function, the better the value.
[0059] Note that the evaluation value calculation unit 103 may input the same data to the first network 111 and the second network 112 in order to obtain inference results from the first network 111 and the second network 112. Alternatively, the data may be input by the learning processing unit 104, or by an input unit, a network control unit, or the like (not shown).
[0060] The learning processing unit 104 updates the second network 112 so as to improve the inference accuracy and inference speed. Specifically, the learning processing unit 104 obtains an evaluation value from the evaluation value calculation unit 103, and updates the second network 112 so as to improve the evaluation value.
[0061] In particular, the learning processing unit 104 updates the second network 112 to improve the evaluation value, thereby reducing the difference between the inference result of the first network 111 and the inference result of the second network 112. In other words, the learning processing unit 104 distills the first network 111 to train the second network 112.
[0062] The network search unit 101, the evaluation value calculation unit 103, and the learning processing unit 104 repeat the above-mentioned processes to obtain the second network 112 with good inference accuracy and inference speed. Furthermore, the search of the second network 112 in the information processing system 100 is a search that is automatically performed based on a loss function, and is a search of a network structure. This search is also called an automatic search or NAS (Neural Architecture Search).
[0063] 5 is a flowchart showing the operation of the information processing system 100 in the second reference example. First, the network searching unit 101 sets the first network 111 as the initial value for searching the second network 112 (S101). In other words, the network searching unit 101 sets the first network 111 as the starting position for searching the second network 112.
[0064] Next, the learning processing unit 104 trains the second network 112 (S102). Specifically, the learning processing unit 104 trains the second network 112 so as to improve the evaluation value obtained from the evaluation value calculation unit 103. This process corresponds to distilling the first network 111 and training the second network 112.
[0065] Next, the network searching unit 101 searches for the second network 112 that will improve the evaluation value obtained based on the training results (S103).
[0066] Then, if the performance of the second network 112 satisfies the requirements (Yes in S104), the process ends. For example, if the inference accuracy and inference speed of the second network 112 satisfy the requirements, the process ends. On the other hand, if the performance of the second network 112 does not satisfy the requirements (No in S104), the training of the second network 112 (S102) and the search of the second network 112 (S103) are repeated.
[0067] In the second reference example, in searching the second network 112, a structure of the second network 112 is searched for that has an inference accuracy close to that of the first network 111 and an inference speed faster than that of the first network 111. Since the second network 112 has an inference speed faster than that of the first network 111, the size of the second network 112 is assumed to be smaller than the size of the first network 111.
[0068] However, in the second reference example, weight reduction for networks for embedded environments is not taken into consideration. In particular, in the second reference example, quantization is not taken into consideration. Therefore, the second network 112 may not be applicable to networks for embedded environments. In addition, weight reduction including quantization may cause degradation compared to the reference network.
[0069] In this embodiment, an information processing method and the like capable of finding an inference model that suppresses the above-mentioned loss will be described. Note that suppressing the loss may be interpreted as compensating for the loss.
[0070] 6 is a block diagram showing an example of the configuration of an information processing system according to the present embodiment. Specifically, the configuration of an information processing system 200 is shown. The information processing system 200 includes a network searching unit 201, a weight reducing unit 202, an evaluation value calculation unit 203, a learning processing unit 204, and the like.
[0071] The network search unit 201 is an information processing unit that performs information processing, and sets the first network 211 as the initial value of the search, and searches for the second network 212 that is likely to suppress losses caused by weight reduction.
[0072] For example, the network searching unit 201 acquires the first network 211 as a reference network from an external device of the information processing system 200. Then, the network searching unit 201 sets the first network 211 as the starting position for searching the second network 212. After that, the network searching unit 201 acquires an evaluation value from the evaluation value calculation unit 203, and searches for the second network 212 that will improve the evaluation value.
[0073] Also, the number of nodes, the type of kernel, etc. may be set as initial values for the search. Also, the network searching unit 201 may acquire setting information indicating a lightening setting and difficulty information indicating the difficulty of inference, and determine the initial values for the search based on the lightening setting and the difficulty of inference.
[0074] That is, the network searching unit 201 may adjust the initial value of the search based on the weight reduction setting and the difficulty of inference. Specifically, the network searching unit 201 may increase the number of nodes when the weight reduction is stronger than a reference. Also, the network searching unit 201 may increase the number of nodes when the difficulty of inference is higher than a reference. Also, a table for determining the initial value of the search may be used for the weight reduction setting, the difficulty of inference, or a combination thereof.
[0075] Each of the first network 211 and the second network 212 is an inference model such as a neural network model for performing inference processing. In particular, the network size of the second network 212 is larger than the network size of the first network 211. This is expected to suppress losses caused by weight reduction.
[0076] Moreover, the first network 211 is a base network and a reference network. For example, the first network 211 uses a floating-point representation. The second network 212 is an intermediate network different from the reference network and the network for an embedded environment. For example, the second network 212 also uses a floating-point representation.
[0077] The weight reduction unit 202 is an information processing unit that performs information processing, and reduces the weight of the second network 212 to generate the third network 213. The weight reduction is performed for the following purposes.
[0078] One of the goals is to reduce the network execution latency, which is expressed in units of milliseconds (ms) or microseconds (μs).
[0079] Another objective is to reduce the amount of calculation required, which is expressed in terms of the number of operations (Ops).
[0080] Another objective is to reduce the amount of memory required for storing weights, intermediate features, etc. The amount of memory required is expressed in units of the number of bits.
[0081] Another objective is to reduce the amount of memory transfer required. The amount of memory transfer required mainly represents the amount of data transferred between the processor of the device performing the inference process and the external DRAM, and is expressed in units of bits per second. However, the memory of the processor's counterpart is not limited to the external DRAM.
[0082] Another objective is to reduce power consumption and the amount of power consumed. Power consumption is expressed in watts (W) or milliwatts (mW), and the amount of power consumed is expressed in watt-hours (Wh). Power consumption and the amount of power consumed are determined by a combination of factors such as the hardware on which the inference process is implemented, the amount of required calculations, and the amount of required memory transfer.
[0083] Another objective is to miniaturize the devices in which the neural network model, deep learning model, or machine learning model is installed. The indicator of device size for miniaturization is cubic centimeters (cm 3 ) or cubic millimeters (mm 3 The size of a device is determined by a combination of factors such as the power consumption of the device, the heat capacity of the device, the network execution latency required for the device, and the size of the components of the device.
[0084] It is not necessary that all of the above is for the purpose of weight reduction, and only some of the above may be for the purpose of weight reduction.
[0085] The weight reduction performed for the above purpose basically includes quantization. For example, the quantization may be quantization that changes a floating-point representation to a fixed-point representation. In addition, the quantization is not limited to quantization that changes a floating-point representation to a fixed-point representation, and may be quantization that changes the representation format to a representation format with a smaller number of bits.
[0086] For example, quantization of a network changes the representation formats of the network parameters, input data values, intermediate data values, output data values, etc., to representation formats with fewer bits. The representation formats of some, but not all, of the network parameters, input data values, intermediate data values, output data values, etc. may be changed to representation formats with fewer bits.
[0087] The lightweighting may also include reducing the network size, such as reducing the number of layers, reducing the number of nodes, reducing the number of connections between nodes, etc. The lightweighting may also include only quantization, i.e., the lightweighting may be quantization. Alternatively, lightweighting without quantization may be used.
[0088] Also, for example, the weight reduction unit 202 acquires weight reduction setting information and reduces the weight of the second network 212 based on the setting information. The setting information may include the number of quantization bits (in other words, the degree of quantization), the reduction amount of the number of layers, the reduction amount of the number of nodes, and the reduction amount of the number of connections between nodes. Here, a small number of quantization bits (in other words, a large degree of quantization) corresponds to a wide quantization width, and a large number of quantization bits (in other words, a small degree of quantization) corresponds to a narrow quantization width. Also, the weight reduction setting information may be stored in a memory or the like as weight reduction setting 231.
[0089] The third network 213 is an inference model such as a neural network model for performing inference processing. Basically, the network size of the third network 213 is larger than the network size of the first network 211. However, the present invention is not limited to this configuration, and the network size of the third network 213 may be smaller than the network size of the first network 211.
[0090] The third network 213 is a network for an embedded environment. Specifically, the third network 213 is a network that achieves the purpose of weight reduction. For example, the third network 213 uses a fixed-point representation.
[0091] The evaluation value calculation unit 203 is an information processing unit that performs information processing and calculates an evaluation value. Specifically, the evaluation value calculation unit 203 calculates the inference accuracy of the third network 213 as the evaluation value.
[0092] For example, the evaluation value calculation unit 203 obtains an inference result of the first network 211 and an inference result of the third network 213 for the same input data. Then, the evaluation value calculation unit 203 may calculate the difference between the inference result of the first network 211 and the inference result of the third network 213 as the evaluation value. Alternatively, the evaluation value calculation unit 203 may calculate the difference between the correct answer data 232 stored in a memory or the like and the inference result of the third network 213 as the evaluation value. The smaller such evaluation value is, the better it is.
[0093] Note that the evaluation value calculation unit 203 may input the same data to the first network 211 and the second network 212 in order to obtain inference results from the first network 211 and the second network 212. Alternatively, the data may be input by the learning processing unit 204, or by an input unit, a network control unit, or the like (not shown).
[0094] Furthermore, for example, the evaluation value calculated by the evaluation value calculation unit 203 is used in the network search unit 201 and the learning processing unit 204. The evaluation value calculation unit 203 may calculate the evaluation value used in the network search unit 201 and the evaluation value used in the learning processing unit 204 based on different criteria.
[0095] Specifically, for example, the evaluation value calculation unit 203 may calculate the difference between the inference result of the first network 211 and the inference result of the third network 213 as the evaluation value used in the network searching unit 201. Then, the evaluation value calculation unit 203 may calculate the difference between the correct answer data 232 and the inference result of the third network 213 as the evaluation value used in the learning processing unit 204.
[0096] The learning processing unit 204 is an information processing unit that performs information processing, and updates the third network 213 so as to improve the inference accuracy of the third network 213. Specifically, the learning processing unit 204 obtains an evaluation value from the evaluation value calculation unit 203, and updates the third network 213 so as to improve the evaluation value.
[0097] For example, the learning processing unit 204 may update the inference results of the first network 211 and the third network 213 so as to reduce the difference between them, according to an evaluation value corresponding to the difference between the inference results of the first network 211 and the inference results of the third network 213. In other words, the learning processing unit 204 may distill the first network 211 to train the third network 213.
[0098] Alternatively, the learning processing unit 204 may perform adversarial learning on the third network 213. The adversarial learning may be performed based on a comparison between an inference result of the first network 211 and an inference result of the third network 213, or may be performed based on a comparison between the correct answer data 232 and an inference result of the third network 213. Alternatively, the learning processing unit 204 may perform distance learning on the third network 213.
[0099] Furthermore, the learning processing unit 204 may change the weight reduction setting 231 according to the evaluation value or the like. That is, the learning processing unit 204 may change the weight reduction setting 231 according to the inference accuracy or the like of the third network 213. For example, the learning processing unit 204 may weaken the weight reduction as the inference accuracy of the third network 213 becomes lower.
[0100] The network searching unit 201, the weight reduction unit 202, the evaluation value calculation unit 203, and the learning processing unit 204 repeat the above-mentioned processes to obtain the third network 213 with good inference accuracy as a network for an embedded environment. In addition, since the third network 213 has been weighted down, the inference speed of the third network 213 is high. The learning processing unit 204 or other components may output the final third network 213 as a network for an embedded environment.
[0101] The information processing system 200 may also include a difficulty level calculation unit 205. The difficulty level calculation unit 205 is an information processing unit that performs information processing, and calculates the difficulty level of inference based on a data set 233. The difficulty level may also be stored as a task difficulty level 234 in a memory or the like.
[0102] For example, the difficulty level calculation unit 205 may calculate the difficulty level of inference based on the amount and number of types of data in the dataset 233 for inference. The difficulty level calculation unit 205 may calculate the difficulty level according to the type of inference, regardless of the dataset 233. For example, the difficulty level of the detection process may be higher than the difficulty level of the identification process.
[0103] Furthermore, the information processing system 200 may include a learning processing unit 206. This learning processing unit 206 updates the second network 212 so as to improve the inference accuracy of the second network 212. Specifically, the learning processing unit 206 may train the second network 212 using the correct answer data 232 as teacher data. Alternatively, the learning processing unit 206 may train the second network 212 by distilling the first network 211.
[0104] Alternatively, the learning processing unit 206 may perform adversarial learning on the second network 212. The adversarial learning may be performed based on a comparison between an inference result of the first network 211 and an inference result of the second network 212, or may be performed based on a comparison between the correct answer data 232 and an inference result of the second network 212. Alternatively, the learning processing unit 206 may perform distance learning on the second network 212.
[0105] FIG. 7 is a flowchart showing a first operation example of the information processing system 200 in this embodiment.
[0106] First, the network searching unit 201 acquires the first network 211 and sets the first network 211 as an initial value for searching the second network 212 (S201). Next, the network searching unit 201 acquires setting information indicating a lightening setting and difficulty information indicating the difficulty of inference (S202). Then, the network searching unit 201 determines an initial value for searching the second network 212 based on the lightening setting and the difficulty of inference (S203).
[0107] For example, the network searching unit 201 determines the first network 211 as the second network 212. Then, the network searching unit 201 adjusts the second network 212 based on the weight reduction setting and the difficulty of inference, and determines the second network 212. In this way, the second network 212 in the initial stage of the search is determined.
[0108] Next, the learning processing unit 206 trains the second network 212 (S204). Note that this process may be omitted.
[0109] Next, the weight reduction unit 202 reduces the weight of the second network 212 based on the weight reduction setting, and generates the third network 213 (S205).
[0110] Next, the learning processing unit 204 trains the third network 213 (S206). Specifically, the learning processing unit 204 trains the third network 213 so as to improve the evaluation value obtained from the evaluation value calculation unit 203. This process may correspond to training the third network 213 by distilling the first network 211, or may correspond to other training.
[0111] Then, if the performance of the third network 213 satisfies the requirements (Yes in S207), the process ends. For example, if the inference accuracy and inference speed of the third network 213 satisfy the requirements, the process ends. On the other hand, if the performance of the third network 213 does not satisfy the requirements (No in S207), the network searching unit 201 changes the number of nodes in each layer or a specific layer (S208). Then, the process is repeated from the training of the second network 212 (S204).
[0112] The inference accuracy may include precision and accuracy. Furthermore, as the accuracy, a correct answer matching rate which is a matching rate between the correct answer data 232 and the inference result of the third network 213 may be used, or a reference matching rate which is a matching rate between the inference result of the first network 211 and the inference result of the third network 213 may be used. Furthermore, the inference speed may correspond to the processing time of the inference.
[0113] For example, the above requirement may be that the inference accuracy of the third network 213 is equal to or higher than a standard, or that the inference speed of the third network 213 is equal to or higher than a standard. Specifically, the above requirement may be that the correct answer matching rate is 90% or higher. Furthermore, the above requirement may be that the reference matching rate is 98% or higher. Furthermore, the above requirement may be that the processing time is 20 ms or less. Furthermore, the above requirement may be a combination of these.
[0114] The network searching unit 201 may determine whether the performance of the third network 213 satisfies the requirements, according to the evaluation value obtained from the evaluation value calculation unit 203. Furthermore, the network searching unit 201 may change the number of nodes in each layer or a specific layer so that the evaluation value obtained from the evaluation value calculation unit 203 improves.
[0115] For example, the evaluation value obtained from the evaluation value calculation unit 203 may indicate the inference accuracy of the third network 213. The inference accuracy of the third network 213 may indicate the difference between the inference result of the third network 213 and the inference result of the first network 211, or may indicate the difference between the inference result of the third network 213 and the correct answer data 232. When the inference accuracy of the third network 213 is poor, the network searching unit 201 may increase the number of nodes in each layer or a specific layer.
[0116] Then, the training of the second network 212 (S204), the weight reduction of the second network 212 (S205), the training of the third network 213 (S206), and the change in the number of nodes (S208) are repeated until the performance of the third network 213 satisfies the requirements. In this way, the third network 213 whose performance satisfies the requirements is obtained. The learning processing unit 204 or other components may output the final third network 213.
[0117] FIG. 8 is a flowchart showing a second operation example of the information processing system 200 in this embodiment.
[0118] First, the network searching unit 201 acquires the first network 211 and sets the first network 211 as an initial value for searching the second network 212 (S301). Next, the network searching unit 201 acquires setting information indicating a lightening setting and difficulty information indicating the difficulty of inference (S302). Then, the network searching unit 201 determines an initial value for searching the second network 212 based on the lightening setting and the difficulty of inference (S303).
[0119] Next, the learning processing unit 206 trains the second network 212 as an optional operation (S304). Next, the weight reduction unit 202 reduces the weight of the second network 212 based on the weight reduction setting to generate the third network 213 (S305). Next, the learning processing unit 204 trains the third network 213 (S306). The processing up to this point is the same as the processing in the first operation example in this embodiment.
[0120] Next, if the performance of the third network 213 does not satisfy the first requirement (No in S307), the network searching unit 201 changes the number of layers (S308). Then, the process is repeated from the training of the second network 212 (S304).
[0121] Network searching unit 201 may determine whether or not the performance of third network 213 satisfies the first requirement, according to the evaluation value obtained from evaluation value calculation unit 203. Furthermore, network searching unit 201 may change the number of layers so that the evaluation value obtained from evaluation value calculation unit 203 improves.
[0122] For example, the evaluation value obtained from the evaluation value calculation unit 203 may indicate the inference accuracy of the third network 213. The inference accuracy of the third network 213 may indicate the difference between the inference result of the third network 213 and the inference result of the first network 211, or may indicate the difference between the inference result of the third network 213 and the correct answer data 232. When the inference accuracy of the third network 213 is poor, the network searching unit 201 may increase the number of layers.
[0123] Furthermore, if the performance of the third network 213 satisfies the first requirement but does not satisfy the second requirement (Yes in S307 and No in S309), the network searching unit 201 changes the number of nodes in each layer or a specific layer (S310). This process is the same as the process in the first operation example in this embodiment. Then, the process is repeated from the training of the second network 212 (S304).
[0124] Then, until the first and second requirements are satisfied, training of the second network 212 (S304), weight reduction of the second network 212 (S305), training of the third network 213 (S306), change in the number of layers (S308), and change in the number of nodes (S310) are repeated. Then, when the performance of the third network 213 satisfies the first and second requirements (Yes in S307 and Yes in S309), the processing ends. As a result, the third network 213 whose performance satisfies the first and second requirements is obtained.
[0125] For example, the second requirement is higher than the first requirement. In other words, the second requirement is stricter than the first requirement. It is assumed that a change in the number of layers has a larger effect on performance than a change in the number of nodes. Therefore, after the number of layers is changed so that the loose first requirement is satisfied, the number of nodes is changed so that the strict second requirement is satisfied. It is assumed that this allows the third network 213 having the performance to satisfy the strict second requirement to be found early.
[0126] As described above, the information processing system 200 in this embodiment searches for the second network 212 that is larger than the first network 211. Then, the information processing system 200 generates the third network 213 by reducing the weight of the second network 212, and trains the third network 213.
[0127] This may lead to finding a third network 213 that suppresses the loss caused by quantization included in the weight reduction.
[0128] The information processing system 200 may include some of the components described in this embodiment, and may perform some of the processes described in this embodiment. Also, at least some of the components and processes described in this embodiment may be combined with at least some of the components and processes described in other embodiments.
[0129] (Embodiment 2) The configuration example in this embodiment is the same as the configuration example shown in Fig. 6. However, in this embodiment, an automatic network search is performed using a loss function. This loss function is a loss function for finding an intermediate network (i.e., the second network 212) and an embedded network (i.e., the third network 213) that suppresses the loss caused by the quantization included in the weight reduction. This allows the embedded network to be found efficiently. Specifically, the following L(x) is used as the loss function.
[0130] L(x) = CE(x, TargetNetQuant) ·α·log(LAT(TargetNetQuant)) β ·γ·Diff(RefNet(x), TargetNetQuant(x)) δ +λ·R_size(SizeDiff(RefNet, TargetNet))
[0131] where x represents the input to the network; RefNet represents the reference network (i.e., the first network 211); TargetNet represents the intermediate network (i.e., the second network 212); TargetNetQuant represents the embedding-oriented network (i.e., the third network 213).
[0132] CE(x, TargetNetQuant) is a cross-entropy term, which indicates the difference between the inference result in the embedded network and the correct answer data 232. CE is a cross-entropy function. This term has the same properties as normal training.
[0133] α log(LAT(TargetNetQuant)) β is a latency term, which is a term related to the execution speed of the embedded network. LAT is a function that represents the amount of delay. LAT(TargetNetQuant) represents the amount of delay in the embedded network. α and β are coefficients for adjusting the output value of the latency term.
[0134] γ Diff(RefNet(x), TargetNetQuant(x)) δ is a reference embedding equivalence term, which represents the difference between the reference network and the embedding network. Diff is a function that represents the difference. Diff(RefNet(x), TargetNetQuant(x)) represents the difference between the inference result of the reference network and the inference result of the embedding network.
[0135] The difference may be absolute difference, squared absolute difference, Euclidean distance, cosine distance, etc. γ and δ are coefficients for adjusting the output value of the reference embedded equality term.
[0136] λ·R_size(SizeDiff(RefNet, TargetNet)) is the size difference constraint term. SizeDiff is a function that represents the size difference. SizeDiff(RefNet, TargetNet) represents the size difference between the reference network and the intermediate network.
[0137] The size difference may be a difference in the number of parameters, a difference in the number of channels, a difference in the number of layers, etc. Alternatively, the size difference may be a difference in the number of parameters weighted for each layer, for example by reducing the weight for an important layer.
[0138] R_size is a size difference evaluation function for evaluating the size difference. R_size(SizeDiff(RefNet, TargetNet)) is smaller as the size difference between the reference network and the intermediate network is larger, and is larger as the size difference between the reference network and the intermediate network is smaller.
[0139] The intermediate network is set to be larger than the reference network. Therefore, R_size(SizeDiff(RefNet, TargetNet)) is smaller as the intermediate network is relatively larger than the reference network. Also, λ is a coefficient for adjusting the output value of the size difference constraint term.
[0140] Fig. 9 is a graph showing a size difference evaluation function in this embodiment. The larger the input value (x) of the size difference evaluation function, the smaller the output value (R_size(x)) of the size difference evaluation function. Furthermore, ε and θ included in the size difference evaluation function in Fig. 9 are coefficients for adjusting the output value of the size difference evaluation function. Fig. 9 shows an example where ε=1 and θ=10.
[0141] The evaluation value calculation unit 203 performs the calculation of the loss function described above. Then, the evaluation value calculation unit 203 outputs the output value of the loss function as the evaluation value. The network search unit 201 obtains the output value of the loss function from the evaluation value calculation unit 203 as the evaluation value, and searches the second network 212 so as to reduce the output value of the loss function.
[0142] Furthermore, the learning processing unit 204 may obtain the output value of the loss function from the evaluation value calculation unit 203 as the evaluation value, and train the third network 213 so that the output value of the loss function becomes small. Alternatively, the evaluation value calculation unit 203 may calculate an evaluation value used in the learning processing unit 204 separately from the evaluation value used in the network searching unit 201. Then, the learning processing unit 204 may train the third network 213 based on an evaluation value different from the output value of the loss function.
[0143] Furthermore, the evaluation value calculation unit 203 may set γ, δ, λ, ε, and θ based on the weight reduction setting and the difficulty level of inference. For example, these coefficients may be adjusted depending on whether the weight reduction setting is aggressive or aggressive.
[0144] With respect to pruning to remove the number of nodes, aggressive weight reduction means a high reduction rate in the number of nodes, and passive weight reduction means a low reduction rate in the number of nodes. With respect to quantization, aggressive weight reduction means a small number of quantization bits and a wide quantization width, and passive weight reduction means a large number of quantization bits and a narrow quantization width. Aggressive weight reduction may also mean a high reduction rate in the number of layers, and passive weight reduction may also mean a low reduction rate in the number of layers.
[0145] For example, when weight reduction is aggressive, it is expected that the inference accuracy will decrease. Therefore, when weight reduction is aggressive, the evaluation value calculation unit 203 changes γ, δ, λ, ε, and θ of the loss function to increase the output value of the loss function. Then, the second network 212 is searched so that the output value of the loss function becomes equal to or less than the threshold value. This can suppress the decrease in inference accuracy.
[0146] Alternatively, when weight reduction is aggressive, γ and δ are set large so that the weight of the difference between the inference result of the reference network and the inference result of the embedded network is increased. In this case, λ and θ are set large and ε is set small so that the weight of the size difference is increased.
[0147] As a result, the output value of the loss function becomes smaller as the difference between the inference results of the reference network and the inference results of the embedded network becomes smaller and as the size of the intermediate network becomes larger.
[0148] Then, the network search unit 201 searches for an intermediate network so as to reduce the output value of the loss function. In other words, the network search unit 201 searches for an intermediate network so as to reduce the difference between the inference result of the reference network and the inference result of the embedded network and to increase the size of the intermediate network. This prevents a decrease in inference accuracy. Also, even if the size of the intermediate network is large, a decrease in inference speed is prevented by actively reducing the weight.
[0149] When the weight reduction is negative, it is assumed that the inference accuracy will not decrease much. Therefore, the opposite settings to the above are performed. For example, when the weight reduction is negative, the evaluation value calculation unit 203 changes γ, δ, λ, ε, and θ of the loss function to reduce the output value of the loss function. Alternatively, in this case, γ and δ are set small so that the weight of the difference between the inference result of the reference network and the inference result of the embedded network is reduced. Also, in this case, λ and θ are set small and ε is set large so that the weight of the size difference is reduced.
[0150] Furthermore, when the difficulty of inference is high, it is expected that the inference accuracy will decrease. Therefore, when the difficulty of inference is high, the evaluation value calculation unit 203 changes γ, δ, λ, ε, and θ of the loss function to increase the output value of the loss function. Then, the second network 212 is searched so that the output value of the loss function becomes equal to or less than the threshold value. This can suppress the decrease in inference accuracy.
[0151] Alternatively, when the inference difficulty is high, γ and δ are set large so that the weight of the difference between the inference result of the reference network and the inference result of the embedded network is increased. In this case, λ and θ are set large and ε is set small so that the weight of the size difference is increased. This prevents the inference accuracy from decreasing.
[0152] When the difficulty of inference is low, it is assumed that the inference accuracy will not decrease much. Therefore, the opposite settings to those described above are performed. For example, when the difficulty of inference is low, the evaluation value calculation unit 203 changes γ, δ, λ, ε, and θ of the loss function to reduce the output value of the loss function. Alternatively, in this case, γ and δ are set small so that the weight of the difference between the inference result of the reference network and the inference result of the embedded network is reduced. Also, in this case, λ and θ are set small and ε is set large so that the weight of the size difference is reduced.
[0153] The evaluation value calculation unit 203 may adjust one or more of the coefficients γ, δ, λ, ε, and θ based on the weight reduction setting and the difficulty of inference, and may maintain the other coefficients at their initial values. For example, the evaluation value calculation unit 203 may adjust ε among ε and θ included in the R_size function, and maintain θ. Also, for example, the evaluation value calculation unit 203 may adjust only λ, θ, and ε related to the size difference among γ, δ, λ, ε, and θ.
[0154] FIG. 10 is a flowchart showing a first operation example of the information processing system 200 in this embodiment.
[0155] First, the network searching unit 201 acquires the first network 211 and sets the first network 211 as an initial value for searching the second network 212 (S401). Next, the network searching unit 201 acquires setting information indicating a lightening setting and difficulty information indicating the difficulty of inference (S402). Then, the network searching unit 201 determines an initial value for searching the second network 212 based on the lightening setting and the difficulty of inference (S403).
[0156] The processing up to this point is the same as the processing in the first and second operation examples in the first embodiment.
[0157] Next, the evaluation value calculation unit 203 sets the coefficient of the loss function (S404). Specifically, the evaluation value calculation unit 203 may acquire setting information indicating the lightening setting and difficulty information indicating the difficulty of inference, and set the coefficient of the loss function based on the lightening setting and the difficulty of inference.
[0158] Next, the learning processing unit 206 trains the second network 212 as an optional operation (S405). Next, the weight reduction unit 202 reduces the weight of the second network 212 based on the weight reduction setting to generate the third network 213 (S406). Next, the learning processing unit 204 trains the third network 213 (S407). These processes are the same as those in the first and second operation examples in the first embodiment.
[0159] Next, the network searching unit 201 searches for a second network 212 that improves the evaluation value obtained based on the training result (S408). Specifically, the network searching unit 201 searches for a new second network 212 that reduces the output value of the loss function described above. That is, the network searching unit 201 adjusts the number of layers and the number of nodes of the second network 212, etc., so that the output value of the loss function described above is reduced.
[0160] The processes from training the second network 212 (S405) to searching the second network 212 (S408) may be repeated before determining whether the performance of the third network 213 satisfies the requirements (S409).
[0161] Then, if the performance of the third network 213 satisfies the requirements (Yes in S409), the processing ends. For example, if the inference accuracy and inference speed of the third network 213 satisfy the requirements, the processing ends. On the other hand, if the performance of the third network 213 does not satisfy the requirements (No in S409), the processing from training the second network 212 (S405) to searching the second network 212 (S408) is repeated. As a result, the third network 213 whose performance satisfies the requirements is obtained.
[0162] The above requirements may be the requirements shown in the first embodiment. For example, the above requirements may be that the correct match rate is 90% or more, the reference match rate is 98% or more, and the processing time is 20 ms or less. Alternatively, the output value of the loss function may be used as an index of the performance of the third network 213. And, the above requirement may be that the output value of the loss function is equal to or less than a threshold value.
[0163] FIG. 11 is a flowchart showing a second operation example of the information processing system 200 in this embodiment.
[0164] First, the network searching unit 201 acquires the first network 211 and sets the first network 211 as an initial value for searching the second network 212 (S501). Next, the network searching unit 201 acquires setting information indicating a lightening setting and difficulty information indicating the difficulty of inference (S502). Then, the network searching unit 201 determines an initial value for searching the second network 212 based on the lightening setting and the difficulty of inference (S503).
[0165] Next, the evaluation value calculation unit 203 sets the coefficients of the loss function (S504). Next, the learning processing unit 206 trains the second network 212 as an optional operation (S505). Next, the weight reduction unit 202 reduces the weight of the second network 212 based on the weight reduction setting to generate the third network 213 (S506). Next, the learning processing unit 204 trains the third network 213 (S507). Next, the network searching unit 201 searches the second network 212 (S508).
[0166] The processing up to this point is the same as the processing in the first operation example in embodiment 1. As in the first operation example, the processing from training the second network 212 (S505) to searching the second network 212 (S508) may be repeatedly performed before determining whether or not the performance of the third network 213 satisfies the first requirement (S509).
[0167] Next, if the performance of the third network 213 does not satisfy the first requirement (No in S509), the learning processing unit 204 makes a more passive change to the weight reduction setting 231 (S510). That is, the learning processing unit 204 changes the weight reduction for the second network 212 to a more passive weight reduction. Specifically, the learning processing unit 204 may increase the number of quantization bits, or may lower the reduction rate of the number of layers and the number of nodes.
[0168] Then, the process is repeated from the step of setting the coefficient of the loss function (S504). In the step of setting the coefficient of the loss function (S504), the coefficient of the loss function is set based on the changed lightening setting 231.
[0169] Furthermore, if the performance of the third network 213 satisfies the first requirement but does not satisfy the second requirement (Yes in S509 and No in S511), the process is repeated from setting the coefficient of the loss function (S504). In this case, the light setting 231 is not changed. Therefore, the process may be repeated from training the second network 212 (S505).
[0170] Then, if the performance of the third network 213 satisfies the first requirement, satisfies the second requirement, and does not satisfy the third requirement (Yes in S509, Yes in S511, and No in S512), the learning processing unit 204 actively changes the weight reduction setting 231 (S513). That is, the learning processing unit 204 changes the weight reduction for the second network 212 to a more aggressive weight reduction. Specifically, the learning processing unit 204 may reduce the number of quantization bits, or may increase the reduction rate of the number of layers and the number of nodes.
[0171] Then, the process is repeated from the step of setting the coefficient of the loss function (S504). In the step of setting the coefficient of the loss function (S504), the coefficient of the loss function is set based on the changed lightening setting 231.
[0172] Then, the process from setting the coefficients of the loss function (S504) to searching the second network 212 (S508) is repeated until the performance of the third network 213 satisfies the first, second and third requirements. Then, when the performance of the third network 213 satisfies the first, second and third requirements (Yes in S509, Yes in S511, and Yes in S512), the process ends. This results in a third network 213 whose performance satisfies the first, second and third requirements.
[0173] For example, the second requirement is higher than the first requirement. In other words, the second requirement is stricter than the first requirement. After the lightening setting 231 is changed so that the loose first requirement is satisfied, the second network 212 is searched for so that the strict second requirement is satisfied. It is assumed that this allows the third network 213 having the performance to satisfy the strict second requirement to be found early.
[0174] Furthermore, the first and second requirements may be requirements on the inference accuracy of the third network 213, and the third requirement may be a requirement on the inference speed of the third network 213.
[0175] In this case, first, the second network 212 is searched for the inference accuracy of the third network 213. Based on the search result regarding the inference accuracy, the second network 212 is searched for the inference speed of the third network 213 together with active change of the light weight setting 231. This may enable the third network 213 that satisfies multiple requirements regarding the inference accuracy and the inference speed to be efficiently found.
[0176] Specifically, for example, the first requirement may be that the correct match rate is 70% or more and the reference match rate is 90% or more. The second requirement may be that the correct match rate is 80% or more and the reference match rate is 95% or more. The third requirement may be that the processing time is 20 ms or less.
[0177] In the above operation, the processes from setting the coefficients of the loss function (S504) to searching the second network 212 (S508) are repeated. However, the processes from determining the initial values for the search (S503) to searching the second network 212 (S508) may be repeated. Also, the processes from training the second network 212 (S505) to searching the second network 212 (S508) may be repeated.
[0178] As described above, the information processing system 200 in this embodiment searches for the second network 212 by using a loss function. This loss function is a loss function for finding the second network 212 and the third network 213 in which the loss caused by the quantization included in the weight reduction is suppressed. This allows an embedded network to be found efficiently. There is a possibility that the third network 213 will be found.
[0179] The information processing system 200 may include some of the components described in this embodiment, and may perform some of the processes described in this embodiment. Also, at least some of the components and processes described in this embodiment may be combined with at least some of the components and processes described in other embodiments.
[0180] (Embodiment 3) In the first and second embodiments, for example, the first network 211 is provided by a third party, and therefore the parameters included in the first network 211 are fixed and cannot be changed.
[0181] In this embodiment, without being limited to the above, parameters included in the first network 211 are allowed to be changed. Then, an operation is performed to bring the inference result of the first network 211 closer to the inference result of the third network 213. As a result, it is expected that similar inference results will be obtained whether the device uses floating-point representation or fixed-point representation.
[0182] Specifically, the information processing system 200 in this embodiment generates a third network 213 that obtains inference results that have a high degree of consistency with the inference results of the first network 211 by changing the parameters of the first network 211 without fixing the parameters of the first network 211.
[0183] More specifically, the information processing system 200 trains the first network 211 by distilling the second network 212. In addition, the information processing system 200 generates the third network 213 by reducing the weight of the second network 212.
[0184] After that, the information processing system 200 trains the first network 211 and the third network 213, which have the second network 212 as the same parent, so that their inference results approach each other. For training the first network 211 and the third network 213, distillation learning, adversarial learning, or distance learning may be used.
[0185] In this embodiment, since the first network 211 and the third network 213 have the same parent, it is expected that the reference matching rate will be improved compared to the first and second embodiments.
[0186] 12 is a block diagram showing a configuration example of an information processing system 200 in this embodiment. The information processing system 200 in this embodiment includes the same components as those of the information processing system 200 in the first and second embodiments. However, in this embodiment, the inference result of the second network 212 is particularly referred to in the evaluation value calculation unit 203. In addition, the learning processing unit 204 updates the first network 211.
[0187] The operation of the information processing system 200 in this embodiment is divided into three phases: a first phase, a second phase, and a third phase.
[0188] In the first phase, the second network 212, which is larger in size, is trained. In the second phase, the first network 211 is trained by distilling the second network 212. In the third phase, the third network 213 is generated by reducing the weight of the second network 212, and the third network 213 is trained by distilling the first network 211. In the third phase, the third network 213 may be trained by distilling the second network 212.
[0189] Specifically, in the first phase, the learning processing unit 206 trains the second network 212. Note that the first phase may be omitted.
[0190] In the second phase, the evaluation value calculation unit 203 acquires the inference result of the first network 211 and the inference result of the second network 212, and calculates an evaluation value indicating the difference between the inference result of the first network 211 and the inference result of the second network 212.
[0191] Then, the learning processing unit 204 trains the first network 211 based on an evaluation value indicating the difference between the inference result of the first network 211 and the inference result of the second network 212. Specifically, the learning processing unit 204 trains the first network 211 so that the difference between the inference result of the first network 211 and the inference result of the second network 212 becomes smaller.
[0192] In the third phase, the weight reduction unit 202 reduces the weight of the second network 212 to generate the third network 213. Then, the evaluation value calculation unit 203 acquires the inference result of the first network 211 and the inference result of the third network 213, and calculates an evaluation value indicating the difference between the inference result of the first network 211 and the inference result of the third network 213.
[0193] Then, the learning processing unit 204 trains the third network 213 based on an evaluation value indicating the difference between the inference result of the first network 211 and the inference result of the third network 213. Specifically, the learning processing unit 204 trains the third network 213 so that the difference between the inference result of the first network 211 and the inference result of the third network 213 becomes smaller.
[0194] In addition, in the third phase, the evaluation value calculation unit 203 may acquire the inference result of the second network 212 and the inference result of the third network 213, and calculate an evaluation value indicating the difference between the inference result of the second network 212 and the inference result of the third network 213.
[0195] Then, the learning processing unit 204 may train the third network 213 based on an evaluation value indicating the difference between the inference result of the second network 212 and the inference result of the third network 213. Specifically, the learning processing unit 204 may train the third network 213 so that the difference between the inference result of the second network 212 and the inference result of the third network 213 becomes smaller.
[0196] The second network 212 has high expressiveness. The first network 211 and the third network 213 reflect information of the second network 212, which has high expressiveness. Therefore, the first network 211 and the third network 213 are expected to have similar inference accuracy.
[0197] The information processing system 200 may include three learning processing units corresponding to the three networks, instead of the learning processing units 204 and 206. That is, the information processing system 200 may include a first network learning processing unit that trains the first network 211, a second network learning processing unit that trains the second network 212, and a third network learning processing unit that trains the third network 213. The information processing system 200 may also include three evaluation value calculation units corresponding thereto.
[0198] FIG. 13 is a flowchart showing the first phase of an example of the operation of the information processing system 200 in this embodiment.
[0199] First, the network searching unit 201 acquires the first network 211, and sets the first network 211 as an initial value for searching the second network 212 (S601).
[0200] Note that the network searching unit 201 may acquire the first network 211 by generating the first network 211. Specifically, the network searching unit 201 may determine a planned size of the first network 211 based on design requirements and the like, and may determine a configuration of the first network 211. Then, the network searching unit 201 may generate the first network 211 based on the determined configuration.
[0201] Next, the network searching unit 201 acquires setting information indicating the lightening setting and difficulty information indicating the difficulty of inference (S602). Then, the network searching unit 201 determines the initial value of the search of the second network 212 based on the lightening setting and the difficulty of inference (S603). Next, the learning processing unit 206 trains the second network 212 (S604). These processes are the same as those in the first operation example of this embodiment.
[0202] If the performance of the second network 212 does not satisfy the first requirement (No in S605), the network searching unit 201 changes the number of layers (S606). Then, the process is repeated from training the second network 212 (S604).
[0203] If the performance of the second network 212 satisfies the first requirement but does not satisfy the second requirement (Yes in S605 and No in S607), the network searching unit 201 changes the number of nodes in each layer or a specific layer (S608). Then, the process is repeated from the training of the second network 212 (S604).
[0204] Then, the training (S604), the change in the number of layers (S606), and the change in the number of nodes (S608) of the second network 212 are repeated until the performance of the second network 212 satisfies the first and second requirements, thereby obtaining the second network 212 whose performance satisfies the first and second requirements.
[0205] These processes (S605, S606, S607, and S608) correspond to the processes (S307, S308, S309, and S310) of the second operation example of the first embodiment. However, while in the process of the second operation example of the first embodiment, a determination is made regarding the performance of the third network 213, in the process of the first phase of this operation example, a determination is made regarding the performance of the second network 212. Then, when the performance of the second network 212 satisfies the first requirement and the second requirement (Yes in S605 and Yes in S607), the process of the first phase ends.
[0206] FIG. 14 is a flowchart showing the second and third phases of an example of the operation of the information processing system 200 in this embodiment.
[0207] In the second phase, the learning processing unit 204 distills the second network 212 to train the first network 211 (S609).
[0208] Specifically, the evaluation value calculation unit 203 acquires the inference result of the first network 211 and the inference result of the second network 212, and calculates an evaluation value indicating the difference between the inference result of the first network 211 and the inference result of the second network 212. The learning processing unit 204 refers to the calculated evaluation value and trains the first network 211 so as to reduce the difference between the inference result of the first network 211 and the inference result of the second network 212.
[0209] Thereafter, if the performance of the first network 211 does not satisfy the third requirement (No in S610), the learning processing unit 204 changes the number of nodes in the first network 211 (S611). Specifically, if the inference accuracy of the first network 211 is not equal to or higher than the standard, the learning processing unit 204 increases the number of nodes in the first network 211. This is expected to improve the inference accuracy of the first network 211. Then, the processing is repeated from obtaining the lightweight setting information and inference difficulty information in the first phase (S602).
[0210] If the performance of the first network 211 satisfies the third requirement (Yes in S610), the weight reduction unit 202 reduces the weight of the second network 212 based on the weight reduction setting to generate the third network 213 (S612).
[0211] Next, the learning processing unit 204 distills the second network 212 to train the third network 213 (S613). Note that this process may be omitted.
[0212] Specifically, the evaluation value calculation unit 203 acquires the inference result of the second network 212 and the inference result of the third network 213, and calculates an evaluation value indicating the difference between the inference result of the second network 212 and the inference result of the third network 213. The learning processing unit 204 refers to the calculated evaluation value and trains the third network 213 so as to reduce the difference between the inference result of the second network 212 and the inference result of the third network 213.
[0213] Next, the learning processing unit 204 distills the first network 211 to train the third network 213 (S614).
[0214] Specifically, the evaluation value calculation unit 203 acquires the inference result of the first network 211 and the inference result of the third network 213, and calculates an evaluation value indicating the difference between the inference result of the first network 211 and the inference result of the third network 213. The learning processing unit 204 refers to the calculated evaluation value and trains the third network 213 so as to reduce the difference between the inference result of the first network 211 and the inference result of the third network 213.
[0215] Thereafter, if the performance of the third network 213 does not satisfy the fourth requirement (No in S615), the process is repeated from the beginning of the first phase (S601). If the performance of the third network 213 satisfies the fourth requirement (Yes in S615), the process ends. This results in a third network 213 whose performance satisfies the fourth requirement.
[0216] Each of the above performances may be inference accuracy, inference speed, or a combination of these, as in the first and second embodiments. Moreover, each requirement is a requirement for these performances.
[0217] As described above, the first network 211 and the third network 213 in this embodiment are based on the same parent. In particular, the information processing system 200 in this embodiment distills the second network 212 to train the first network 211. Therefore, it is expected that the reference matching rate will be improved compared to the first and second embodiments.
[0218] The information processing system 200 may include some of the components described in this embodiment, and may perform some of the processes described in this embodiment. Also, at least some of the components and processes described in this embodiment may be combined with at least some of the components and processes described in other embodiments.
[0219] (Basic implementation example and basic operation example) Basic implementation examples and basic operation examples relating to the first, second and third embodiments are shown below.
[0220] 15 is a block diagram showing a basic implementation example of an information processing system 200 in the above-mentioned embodiments. The information processing system 200 includes, for example, at least one processor 301 and at least one memory 302.
[0221] The processor 301 is an information processing circuit that performs information processing. The processor 301 may perform the roles of the network searching unit 201, the weight reducing unit 202, the evaluation value calculation unit 203, the learning processing unit 204, the difficulty level calculation unit 205, and the learning processing unit 206. The processor 301 may perform these roles by reading a program from the memory 302 and executing the program. The processor 301 may also control the inference processing of the first network 211, the second network 212, and the third network 213.
[0222] The memory 302 is a storage device for storing information, and may also be expressed as a recording medium. The memory 302 may store information such as the lightening setting 231, the correct answer data 232, the data set 233, and the task difficulty level 234. The memory 302 may also store a program for the processor 301 to execute information processing. The memory 302 may also store information on the first network 211, the second network 212, and the third network 213.
[0223] The information processing system 200 is, for example, a computer. The information processing system 200 may be one information processing device, or may be configured with multiple information processing devices.
[0224] Fig. 16 is a flowchart showing an example of a basic operation of the information processing system 200 in a number of embodiments. For example, at least one processor 301 shown in Fig. 15 performs the operation shown in Fig. 16 using at least one memory 302. In the following description, the first inference model, the second inference model, and the third inference model correspond to the first network 211, the second network 212, and the third network 213, respectively.
[0225] First, the processor 301 acquires a first inference model to be used as a reference (S701). Next, the processor 301 calculates a second inference model having a model size larger than that of the first inference model based on the first inference model (S702). Next, the processor 301 quantizes the calculated second inference model to generate a third inference model (S703).
[0226] Next, the processor 301 trains a third inference model using machine learning (S704). Next, the processor 301 determines whether the performance of the trained third inference model satisfies a condition (S705). Next, the processor 301 outputs the trained third inference model if the performance satisfies the condition (S706).
[0227] As a result, a second inference model having a larger model size than the first inference model is quantized. It is expected that performance will not deteriorate easily even if the second inference model having a larger model size is quantized. In other words, it is expected that the loss caused by quantization will be relatively small in the third inference model generated by quantizing the second inference model having a larger model size than the first inference model. Therefore, it is possible to find an inference model in which the loss caused by quantization is suppressed.
[0228] Also, for example, the processor 301 may obtain setting information indicating quantization settings for the second inference model. Then, the processor 301 may set an initial value in the calculation of the second inference model based on the setting information and the first inference model.
[0229] This starts the calculation of the second inference model based on the quantization setting information and the first inference model, making it possible to quickly find the third inference model according to the quantization and the first inference model.
[0230] Also, for example, the processor 301 may obtain difficulty information indicating the inference difficulty of at least one of the first inference model, the second inference model, and the third inference model, and may set an initial value for the calculation of the second inference model based on the difficulty information and the first inference model.
[0231] This starts the calculation of the second inference model based on the inference difficulty information and the first inference model, making it possible to quickly find a third inference model that corresponds to the inference difficulty and the first inference model.
[0232] Also, for example, the calculation of the second inference model may be a search for the second inference model using a loss function. The loss function may be a function whose output value becomes smaller when the difference between the inference result of the first inference model and the inference result of the third inference model becomes smaller, and whose output value becomes smaller when the model size of the second inference model becomes larger relative to the first inference model. The search for the second inference model may be performed so that the output value of the loss function becomes smaller.
[0233] This makes it possible to find an inference model based on the loss function such that the loss caused by quantization is suppressed.
[0234] Also, for example, the processor 301 may obtain setting information indicating a quantization setting for the second inference model. Then, the processor 301 may change the loss function based on the setting information.
[0235] This makes it possible to find an inference model that suppresses the loss caused by quantization, based on a loss function that corresponds to the quantization settings.
[0236] Also, for example, the loss function may be changed so that the output value of the loss function increases as the degree of quantization increases in the settings indicated by the setting information. Then, the search for the second inference model may be performed so that the output value of the loss function becomes equal to or less than a threshold value.
[0237] As a result, the greater the degree of quantization, the greater the output value of the loss function, but the second inference model is searched for so that the output value of the loss function is equal to or less than the threshold. In other words, even if the loss is large due to the large degree of quantization, the second inference model is searched for so that certain conditions for suppressing the loss are satisfied. Therefore, it is possible to find an inference model that suppresses the loss at a certain level.
[0238] Also, for example, the processor 301 may obtain difficulty information indicating the inference difficulty of at least one of the first inference model, the second inference model, and the third inference model, and may change the loss function based on the difficulty information.
[0239] This makes it possible to find an inference model that suppresses the loss caused by quantization based on a loss function that corresponds to the difficulty of inference.
[0240] Also, for example, the loss function may be changed so that the output value of the loss function increases as the inference difficulty indicated by the difficulty information increases. Then, the search for the second inference model may be performed so that the output value of the loss function becomes equal to or less than a threshold.
[0241] As a result, the higher the inference difficulty, the larger the output value of the loss function, but the second inference model is searched for so that the output value of the loss function is below the threshold. In other words, even if the loss is large because of the high inference difficulty, the second inference model is searched for so that certain conditions for suppressing the loss are satisfied. Therefore, it is possible to find an inference model that suppresses the loss at a certain level.
[0242] Also, for example, the processor 301 may change the quantization settings for the second inference model if the performance does not meet the requirements.
[0243] This allows the quantization settings to be changed so that the performance criteria are met, and it is then possible to find an inference model that satisfies the performance criteria.
[0244] Also, for example, the conditions may include the precision or accuracy of the inference result of the first inference model or the inference of the third inference model with respect to the reference data, and the processor 301 may reduce the degree of quantization when the precision or accuracy of the inference of the third inference model is equal to or lower than a threshold.
[0245] As a result, when the inference precision or accuracy of the third inference model is equal to or lower than the threshold, the quantization for the second inference model is narrowed so that the inference precision or accuracy of the third inference model is increased, and thus it becomes possible to find an inference model such that the condition of the inference precision or accuracy is satisfied.
[0246] Also, for example, the conditions may include the speed of the inference process of the third inference model, and the processor 301 may increase the degree of quantization if the speed of the inference process is equal to or lower than a threshold.
[0247] As a result, when the inference processing speed of the third inference model is equal to or lower than the threshold, the degree of quantization for the second inference model is increased so that the inference processing speed of the third inference model is increased, and thus it becomes possible to find an inference model that satisfies the condition of the inference processing speed.
[0248] Also, for example, the processor 301 may input data to a first inference model to obtain an inference result of the first inference model. Also, the processor 301 may input data to a second inference model to obtain an inference result of the second inference model. Then, the processor 301 may train the first inference model based on the difference between the inference result of the first inference model and the inference result of the second inference model.
[0249] As a result, the first inference model and the third inference model are constructed based on the same second inference model, making it possible to reduce the difference between the inference result of the first inference model and the inference result of the third inference model.
[0250] Also, for example, the processor 301 may further perform the processes shown in any of the above embodiments.
[0251] 17 is a block diagram showing another basic implementation example of the information processing system 200 in the above-mentioned embodiments. The information processing system 200 includes, for example, a calculation unit 401, a generation unit 402, a training unit 403, a determination unit 404, and an output unit 405. Furthermore, the information processing system 200 may include at least one of an initial value setting unit 406, a loss function changing unit 407, and a quantization setting changing unit 408.
[0252] Each of these components is an information processing circuit that performs information processing. These components may be realized by the processor 301 shown in FIG.
[0253] The calculation unit 401 is a component corresponding to the network searching unit 201, etc. The calculation unit 401 performs processing related to the calculation of the second inference model. Specifically, the calculation unit 401 obtains the first inference model (S701) and calculates the second inference model (S702) in FIG. 16.
[0254] The generation unit 402 is a component corresponding to the weight reduction unit 202, etc. The generation unit 402 performs processing related to quantization of the second inference model and generation of the third inference model. Specifically, the generation unit 402 generates the third inference model in FIG. 16 (S703).
[0255] The training unit 403 is a component corresponding to the learning processing unit 204 and the like. The training unit 403 performs processing related to training of the third inference model. Specifically, the training unit 403 performs training (S704) of the third inference model in FIG. 16. The training unit 403 may also perform processing related to training of the first inference model, or may perform processing related to training of the second inference model. The information processing system 200 may include a training unit 403 for each of the first inference model, the second inference model, and the third inference model.
[0256] The determination unit 404 is a component corresponding to the evaluation value calculation unit 203, etc. The determination unit 404 performs processing related to determining whether or not the performance of the third inference model satisfies a condition. Specifically, the determination unit 404 performs the determination (S705) in FIG. 16.
[0257] The output unit 405 is a component corresponding to the evaluation value calculation unit 203, etc. The output unit 405 performs processing related to the output of the third inference model. Specifically, the output unit 405 outputs the third inference model in FIG. 16 (S706).
[0258] The initial value setting unit 406 is a component corresponding to the network searching unit 201, etc. The initial value setting unit 406 performs processing related to setting initial values in the calculation of the second inference model. The loss function changing unit 407 is a component corresponding to the evaluation value calculation unit 203, etc. The loss function changing unit 407 performs processing related to changing the loss function. The quantization setting changing unit 408 is a component corresponding to the learning processing unit 204, etc. The quantization setting changing unit 408 performs processing related to changing the quantization settings.
[0259] Conversely, the network searching unit 201 corresponds to examples of the calculation unit 401 and the initial value setting unit 406, etc. Furthermore, the weight reduction unit 202 corresponds to examples of the generation unit 402, etc. Furthermore, the evaluation value calculation unit 203 corresponds to examples of the determination unit 404, the output unit 405, and the loss function changing unit 407, etc. Furthermore, the learning processing unit 204 and the learning processing unit 206 correspond to examples of the training unit 403 and the quantization setting changing unit 408, etc.
[0260] Note that the configuration shown in Fig. 17 is an example, and the configuration of the information processing system 200 is not limited to the example shown in Fig. 17. For example, the configuration shown in Fig. 17 may be combined with any of the configurations shown in any of the above-described embodiments.
[0261] Although aspects of the information processing system have been described above based on the embodiments, the aspects of the information processing system are not limited to the embodiments. Modifications conceived by a person skilled in the art may be applied to the embodiments, and multiple components in the embodiments may be combined in any manner. For example, a process performed by a specific component in the embodiments may be performed by another component instead of the specific component. In addition, the order of multiple processes may be changed, and multiple processes may be performed in parallel.
[0262] Furthermore, the inference model is, for example, a mathematical model for performing inference processing, and may be a machine learning model, a neural network model, or a deep learning model.
[0263] Furthermore, an information processing method including steps performed by each component of the information processing system may be executed by any device or system. In other words, this information processing method may be executed by the information processing system, or may be executed by another device or system.
[0264] For example, the above-mentioned information processing method may be executed by a computer including a processor, a memory, an input / output circuit, etc. In this case, the information processing method may be executed by the computer executing a program for causing the computer to execute the information processing method. In addition, the program may be recorded on a non-transitory computer-readable recording medium.
[0265] For example, the above program causes a computer to execute an information processing method of obtaining a first inference model as a reference, calculating a second inference model having a larger model size than the first inference model based on the first inference model, quantizing the calculated second inference model to generate a third inference model, training the third inference model using machine learning, determining whether the performance of the trained third inference model satisfies a condition, and outputting the trained third inference model if the performance satisfies the condition.
[0266] Furthermore, the multiple components of the information processing system may be configured with dedicated hardware, may be configured with general-purpose hardware that executes the above-mentioned programs, etc., or may be configured with a combination of these. Furthermore, the general-purpose hardware may be configured with a memory in which the programs are stored, and a general-purpose processor that reads the programs from the memory and executes them, etc. Here, the memory may be a semiconductor memory or a hard disk, etc., and the general-purpose processor may be a CPU, etc.
[0267] Furthermore, the dedicated hardware may be configured with a memory and a dedicated processor, etc. For example, the dedicated processor may refer to the memory and execute the above-described information processing method.
[0268] Furthermore, each component of the information processing system may be an electric circuit. These electric circuits may form a single electric circuit as a whole, or each may be a separate electric circuit. These electric circuits may correspond to dedicated hardware, or may correspond to general-purpose hardware that executes the above programs, etc. [Industrial Applicability]
[0269] The present disclosure can be used in an information processing system for finding an inference model that suppresses losses caused by quantization, and is applicable to machine learning model construction systems, neural network model construction systems, deep learning model construction systems, etc. [Explanation of symbols]
[0270] 100, 200 Information Processing System 101, 201 Network Search Department 103, 203 Evaluation value calculation unit 104, 204, 206 Learning processing unit 111, 211 First Network 112, 212 Second Network 202 Lightweight section 205 Difficulty Calculation Unit 213 3rd Network 231 Lightweight Settings 232 Correct Data 233 Dataset 234 Task Difficulty 301 Processor 302 Memory 401 Calculation Department 402 Generator 403 Training Department 404 Judgment section 405 Output section 406 Initial value setting section 407 Loss function change part 408 Quantization setting change section
Claims
1. An information processing method executed by a computer, comprising: Obtain a first inference model as a reference; determining a second inference model having a model size larger than the first inference model; A third inference model is generated by reducing the determined second inference model; training the third inference model using machine learning; Determine whether the performance of the trained third inference model satisfies a condition; If the performance satisfies the condition, output the trained third inference model. Information processing methods.
2. An information processing method executed by a computer, comprising: A first network is trained by distilling a second network having a size larger than the first network; a third network is generated by reducing the weight of the second network; The third network is trained by distilling the first network, or the third network is trained by distilling the second network. Information processing methods.
3. If the performance does not satisfy the condition, the weight reduction setting for the second inference model is changed. The information processing method according to claim 1 .
4. The conditions include the accuracy or precision of the inference result of the first inference model or the inference of the third inference model with respect to the reference data, In the change of the setting, when the inference precision or accuracy of the third inference model is equal to or lower than a threshold, the degree of weight reduction is reduced. The information processing method according to claim 3.
5. The condition includes a speed of an inference process of the third inference model; In the change of the setting, when the speed of the inference process is equal to or lower than a threshold, the degree of weight reduction is increased. The information processing method according to claim 3.
6. moreover, inputting data into the first inference model to obtain an inference result of the first inference model; inputting the data into the second inference model to obtain an inference result of the second inference model; training the first inference model based on a difference between an inference result of the first inference model and an inference result of the second inference model The information processing method according to claim 3.
7. a determination unit that obtains a first inference model to be used as a reference and determines a second inference model having a model size larger than that of the first inference model; a generation unit that generates a third inference model by reducing the determined second inference model; a training unit that trains the third inference model using machine learning; A determination unit that determines whether the performance of the trained third inference model satisfies a condition; and an output unit that outputs the trained third inference model if the performance satisfies the condition. Information processing system.
8. a training unit for distilling a second network having a size larger than that of the first network and training the first network; A generation unit that generates a third network by reducing the weight of the second network; Equipped with The training unit distills the first network to train the third network, or distills the second network to train the third network. Information processing system.