Information Processing Method and Information Processing System
By calculating the difference between the outputs of the first learning model and the second learning model and the positive solution data, and using these differential data for re-learning of the first learning model, the output difference problem caused by unclear transformation processing content of the transformation tool is solved, and the effect of reducing the output difference is achieved.
Patent Information
- Application Number
- CN201910671225.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-15
- Filing Date
- 2019-07-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-07-24
AI Technical Summary
When the transformation processing content of the transformation tool transformed from the first learning model to the second learning model is unclear, it is difficult to reduce the output data difference between the first learning model and the second learning model when the same data is input.
By calculating the difference between the output of the same input data and the positive solution data of the first learning model and the second learning model, the first difference data and the second difference data are obtained respectively, and the re-learning of the first learning model is used to reduce the output difference.
Even if the transformation processing content of the transformation tool is unclear, the difference in output data between the first learning model and the second learning model when inputting the same data can be effectively reduced.
Smart Images

Figure CN110826721B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing method and an information processing system for obtaining a learning model through machine learning. Background Art
[0002] Conventionally, the following technique has been known: using a transformation tool, a second learning model suitable for a second computer environment is generated from a first learning model learned in a first arithmetic processing environment, and the generated second learning model is used in the second arithmetic processing environment.
[0003] For example, Non-Patent Document 1 describes the following technique: a technique for reducing the difference between the output data of the first learning model and the output data of the second learning model that occurs when the same data is input to the first learning model and the second learning model obtained by transforming the first learning model using a transformation tool.
[0004] Non-Patent Document 1: Quantization and Training of Neural Networks forEfficient Integer-Arithmetic-Only Inference.https: / / arxiv.org / abs / 1712.05877 Summary of the Invention
[0005] However, in the case where the transformation processing content of the transformation tool for transforming from the first learning model to the second learning model is unclear (i.e., the transformation tool is a black box), the above conventional technique cannot be used.
[0006] Therefore, an object of the present invention is to provide an information processing method and an information processing system that can reduce the difference between the output data of the first learning model and the output data of the second learning model that occurs when the same data is input to the first learning model and the second learning model, even when the transformation processing content of the transformation tool for transforming from the first learning model to the second learning model is unclear.
[0007] An information processing method according to one aspect of the present invention performs the following processing using a computer: obtaining first output data of a first learning model for input data, correct data for the input data, and second output data of a second learning model for the input data obtained by transforming the first learning model; calculating first difference data corresponding to the difference between the first output data and the correct data, and second difference data corresponding to the difference between the second output data and the correct data; and using the first difference data and the second difference data to perform learning of the first learning model.
[0008] An information processing system according to one aspect of the present invention includes: an acquisition unit that acquires first output data of a first learning model for input data, correct data for the input data, and second output data of a second learning model for the input data obtained by transformation of the first learning model; a calculation unit that calculates first difference data corresponding to a difference between the first output data and the correct data, and second difference data corresponding to a difference between the second output data and the correct data; and a learning unit that learns the first learning model using the first difference data and the second difference data.
[0009] Advantageous Effects of the Invention
[0010] According to an information processing method and an information processing system according to one aspect of the present invention, even when the transformation processing content of a transformation tool for transforming from a first learning model to a second learning model is unclear, it is possible to reduce the difference between the output data of the first learning model and the output data of the second learning model that occurs when the same data is input to the first learning model and the second learning model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a block diagram showing the configuration of an information processing system according to the first embodiment.
[0012] Figure 2 is a schematic diagram showing an example of a state in which a transformation unit according to the first embodiment transforms a first learning model into a second learning model.
[0013] Figure 3 is a schematic diagram showing an example of a state in which a learning unit according to the first embodiment re-learns the first learning model.
[0014] Figure 4 is a flowchart of a first update process of a learning model according to the first embodiment.
[0015] Figure 5 is a block diagram showing the configuration of an information processing system according to the second embodiment.
[0016] Figure 6 is a schematic diagram showing an example of generation of data for re-learning a first learning model in the information processing system according to the second embodiment.
[0017] Figure 7 is a flowchart of a second update process of a learning model according to the second embodiment.
[0018] REFERENCE SIGNS LIST
[0019] 1, 1A Information processing system
[0020] 10 Acquisition Unit
[0021] 20 Calculation Unit
[0022] 30 Learning Unit
[0023] 40 Transformation Unit
[0024] 50 First Learning Model
[0025] 60 Second Learning Model Detailed Implementation Manner
[0026] (Reason for Obtaining a Technical Solution of the Present Invention)
[0027] In recent years, in in-vehicle installed systems such as ADAS (Advanced Driver - Assistance System) and autonomous driving systems, for recognition systems using machine learning, it is required to perform inference using a learning model.
[0028] Generally, a learning model applied to an in-vehicle installed system is generated in the following manner: a transformation tool is applied to a first learning model obtained through learning in a computer system having a higher performance than the in-vehicle installed system, and it is transformed into a second learning model suitable for the in-vehicle installed system.
[0029] For example, by transforming a first learning model that learns through floating-point arithmetic processing and performs inference through floating-point arithmetic in a personal computer into a second learning model that performs integer arithmetic processing in an in-vehicle installed system, a learning model applied to the in-vehicle installed system is generated.
[0030] The processing of the first learning model and the processing of the second learning model are not necessarily exactly the same. Therefore, even when the same data is input to the first learning model and the second learning model, there may be a difference between the output of the first learning model and the output of the second learning model.
[0031] In the case where the transformation processing content of the transformation tool for transforming from the first learning model to the second learning model is disclosed, for example, the above difference can be reduced by using the technology described in Non-Patent Document 1. However, in the case where the transformation processing content of the transformation tool for transforming from the first learning model to the second learning model is unclear, the technology described in Non-Patent Document 1 cannot be used.
[0032] In view of such problems, the inventor came up with the following information processing method and information processing system.
[0033] An information processing method for a technical solution of the present invention performs the following processing using a computer: obtaining first output data of a first learning model for input data, correct solution data for the input data, and second output data of a second learning model for the input data obtained by transformation of the first learning model; calculating first difference data corresponding to the difference between the first output data and the correct solution data, and second difference data corresponding to the difference between the second output data and the correct solution data; and using the first difference data and the second difference data to perform learning of the first learning model.
[0034] According to the above information processing method, the first learning model uses the second difference data in addition to the first difference data for learning. In addition, in the learning of the first learning model, it is not necessary to reflect the transformation processing content of the transformation tool that transforms from the first learning model to the second learning model. Therefore, according to the above information processing method, even if the transformation processing content of the transformation tool that transforms from the first learning model to the second learning model is unclear, it is possible to reduce the difference between the output data of the first learning model and the output data of the second learning model that occurs when the same data is input to the first learning model and the second learning model.
[0035] In addition, it may be that in the above learning, weights are assigned to the first difference data and the second difference data. Thereby, in the learning of the first learning model, it is possible to perform learning by giving a difference to the degree of emphasizing the output of the first learning model and the degree of emphasizing the output of the second learning model.
[0036] In addition, it may be that in the above weighting, the weight of the first difference data is made larger than the weight of the second difference data. Thereby, in the learning of the first learning model, it is possible to perform learning while emphasizing the output of the first learning model compared to the output of the second learning model. In other words, it is possible to suppress the characteristics (or performance) of the first learning model from being too close to the characteristics (or performance) of the second learning model.
[0037] In addition, it may be that in the above learning, the difference between the first difference data and the second difference data is also used. Thereby, in the learning of the first learning model, it is possible to perform learning while considering the difference between the output of the first learning model and the output of the second learning model. The smaller the difference between these two difference data, the closer the characteristics (or performance) are between the first learning model and the second learning model. Therefore, it is possible to efficiently perform learning to reduce the difference between the output data of the first learning model and the output data of the second learning model.
[0038] In addition, it may also be that, in the above learning, weights are assigned to the above first difference data, the above second difference data, and the difference between the above first difference data and the above second difference data. Thereby, in the learning of the first learning model, it is possible to perform learning by differentiating the degree of emphasizing the output of the first learning model, the degree of emphasizing the output of the second learning model, and the degree of emphasizing the difference between the output of the first learning model and the output of the second learning model.
[0039] In addition, it may also be that the above first learning model and the above second learning model are neural network type learning models. Thereby, the first learning model and the second learning model can be implemented using a relatively well-known mathematical model.
[0040] An information processing system according to an aspect of the present invention includes: an acquisition unit that acquires first output data of a first learning model for input data, correct data for the input data, and second output data of a second learning model for the input data obtained by transformation of the first learning model; a calculation unit that calculates first difference data corresponding to the difference between the first output data and the correct data, and second difference data corresponding to the difference between the second output data and the correct data; and a learning unit that uses the first difference data and the second difference data to perform learning of the first learning model.
[0041] According to the above information processing system, the first learning model performs learning using the second difference data in addition to the first difference data. In addition, in the learning of the first learning model, it is not necessary to reflect the transformation processing content of the transformation tool for transforming from the first learning model to the second learning model. Therefore, according to the above information processing system, even if the transformation processing content of the transformation tool for transforming from the first learning model to the second learning model is unclear, it is possible to reduce the difference between the output data of the first learning model and the output data of the second learning model that occurs when the same data is input to the first learning model and the second learning model.
[0042] Hereinafter, specific examples of an information processing method and an information processing system according to an aspect of the present invention will be described with reference to the drawings. The embodiments shown here are all specific examples of the present invention. Therefore, the numerical values, shapes, constituent elements, arrangements and connection forms of the constituent elements, and steps (processes) and the order of steps shown in the following embodiments are examples and do not limit the present invention. Constituent elements in the following embodiments that are not described in the independent claims are constituent elements that can be arbitrarily added. In addition, each drawing is a schematic diagram and is not necessarily drawn precisely.
[0043] In addition, these inclusive or specific technical solutions can also be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0044] (First Embodiment)
[0045] First, an information processing system related to the first embodiment will be described. This information processing system is a system that transforms a first learning model that performs floating-point arithmetic processing into a second learning model that performs integer arithmetic processing, and is a system that retrains the first learning model to reduce the difference between the output data of the first learning model and the output data of the second learning model that occurs when the same data is input to the first learning model and the second learning model.
[0046] [1-1. Structure of the Information Processing System]
[0047] Figure 1 It is a block diagram showing the structure of the information processing system 1 related to the first embodiment.
[0048] As Figure 1 shown, the information processing system 1 includes an acquisition unit 10, a calculation unit 20, a learning unit 30, a transformation unit 40, a first learning model 50, and a second learning model 60.
[0049] The information processing system 1 can also be implemented, for example, by a personal computer including a processor and a memory. In this case, each component of the information processing system 1 can be implemented, for example, by a processor executing one or more programs stored in the memory. In addition, the information processing system 1 can be implemented, for example, by multiple computer devices that can communicate with each other and each include a processor and a memory coordinating their actions. In this case, each component of the information processing system 1 can be implemented, for example, by one or more processors executing one or more programs stored in one or more memories.
[0050] The first learning model 50 is a machine learning model that uses floating-point type variables. Here, it is assumed that the first learning model 50 is a neural network type learning model, and it is described as a person recognizer that has been trained to recognize a person included as a subject in an image from the image. For the first learning model 50, for example, if an image is input as input data, coordinates indicating the position of the recognized person and the reliability of the person are output as output data.
[0051] The second learning model 60 is a machine learning model obtained by transforming the first learning model 50 through a transformation unit 40 described later so as to perform processing using integer-type variables. Here, similar to the first learning model 50, the second learning model 60 is assumed to be a neural network-type learning model, and will be described as a person identifier that has been learned to recognize a person included as a subject in an image from the image. For example, similar to the first learning model 50, if an image is input as input data, the second learning model 60 outputs coordinates indicating the position of the recognized person and the reliability of the person as output data.
[0052] The second learning model 60 performs processing with lower numerical operation precision than the first learning model 50. On the other hand, even a system that cannot process floating-point variables, that is, a system that cannot utilize the first learning model 50, can utilize the second learning model 60.
[0053] For example, in an in-vehicle mounting system that is relatively lacking in computing resources and cannot process floating-point variables but can process integer-type variables, although the first learning model 50 cannot be utilized, the second learning model 60 can be utilized.
[0054] In addition, the second learning model 60 is suitable for utilization in a system that places more importance on reducing power consumption associated with operations than on the precision of operations, for example.
[0055] The transformation unit 40 transforms the first learning model 50 that performs processing using floating-point variables into the second learning model 60 that performs processing using integer-type variables.
[0056] Figure 2 FIG. is a schematic diagram showing an example of the state in which the transformation unit 40 transforms the first learning model 50 into the second learning model 60.
[0057] As Figure 2 shown, when the first learning model 50 is composed of a plurality of weights (here, for example, the first weight 51, the second weight 52, and the third weight 53) that perform processing using floating-point variables and are hierarchically structured, the transformation unit 40 transforms the plurality of weights that perform processing using floating-point variables into a plurality of weights that perform processing using integer-type variables (here, for example, the first weight 61, the second weight 62, and the third weight 63).
[0058] The first learning model 50 is a learning model that processes using floating-point variables. In contrast, the second learning model 60 is a learning model that processes using integer variables. Therefore, even if the same image A is input to the first learning model 50 and the second learning model 60, the output data A1 output from the first learning model 50 and the output data A2 output from the second learning model 60 are not necessarily the same. That is, when the correct data in the case where the input data is image A is set as the correct data A, a difference occurs between the first difference data (described later) corresponding to the difference between the output data A1 and the correct data A and the second difference data (described later) corresponding to the difference between the output data A2 and the correct data A.
[0059] Returning again to Figure 1 , the structure of the information processing system 1 will continue to be described.
[0060] The acquisition unit 10 acquires the first output data of the first learning model 50 for the input data, the second output data of the second learning model 60 for the input data, and the correct data for the input data.
[0061] The calculation unit 20 calculates the first difference data corresponding to the difference between the first output data and the correct data (hereinafter, in equations and the like, there are cases where the first difference data is referred to as "Loss1"), and the second difference data corresponding to the difference between the second output data and the correct data (hereinafter, in equations and the like, there are cases where the second difference data is referred to as "Loss2") based on the first output data, the second output data, and the correct data acquired by the acquisition unit 10.
[0062] Here, as an example that does not need to be limited, the first difference data (Loss1) is assumed to be the L2 norm of the correct data and the first output data calculated according to the following (Equation 1).
[0063] Loss1 = ||correct data - first output data|| 2 (Equation 1)
[0064] In addition, as an example that does not need to be limited, the second difference data (Loss2) is assumed to be the L2 norm of the correct data and the second output data calculated according to the following (Equation 2).
[0065] Loss2 = ||correct data - second output data|| 2 (Equation 2)
[0066] The learning unit 30 retrains the first learning model 50 using the first difference data and the second difference data.
[0067] Figure 3 is a schematic diagram showing an example of the state where the learning unit 30 retrains the first learning model 50.
[0068] As Figure 3 shown, the learning unit 30 calculates the difference data represented by (Equation 3) based on the first difference data and the second difference data (hereinafter, the difference data may also be referred to as "LOSS" in equations and the like). In addition, the correct answer data, the first output data, and the second output data used to calculate the first difference data and the second difference data can also be normalized by the number of output data.
[0069] LOSS = λ1 * Loss1 + λ2 * Loss2 + λ3 * ||Loss1 - Loss2|| (Equation 3)
[0070] Here, λ1, λ2, and λ3 are values that weight the first difference data, the second difference data, and the difference between the first difference data and the second difference data in the calculation of the difference data, and satisfy the relationships of the following (Equation 4) to (Equation 7).
[0071] λ1 + λ2 + λ3 = 1 (Equation 4)
[0072] 1 > λ1 > 0 (Equation 5)
[0073] 1 > λ2 > 0 (Equation 6)
[0074] 1 > λ3 ≥ 0 (Equation 7)
[0075] If the learning unit 30 calculates the difference data, as Figure 3 shown, the error backpropagation method using the calculated difference data as an error is used to update the weights, so that the first learning model 50 relearns.
[0076] Regarding the relearning of the first learning model 50 by the learning unit 30, the inventor repeatedly conducted experiments by changing the combination of the values of λ1, λ2, and λ3 in (Equation 3) for calculating the difference data. As a result, the inventor obtained the following insight: In order to reduce the difference between the output data of the first learning model and the output data of the second learning model, it is preferable that λ1 is larger than λ2, that is, in the weighting of the first difference data, the second difference data, and the difference between the first difference data and the second difference data in the calculation of the difference data, it is preferable that the weight of the first difference data is larger than the weight of the second difference data. It is speculated that this is because by making the first learning model 50 relearn by attaching more importance to the output of the first learning model 50, which performs higher-precision numerical operations, than the output of the second learning model 60, which performs lower-precision numerical operations, the difference between the output data of the first learning model and the output data of the second learning model can be reduced with better precision.
[0077] [1 - 2. Operation of the information processing system]
[0078] Hereinafter, the processing performed by the information processing system 1 with the above structure will be described.
[0079] The information processing system 1 performs a first update process of a learning model that updates the first learning model 50 and the second learning model 60 using the first difference data and the second difference data.
[0080] Figure 4 It is a flowchart of the first update process of the learning model.
[0081] The first update process of the learning model starts, for example, when for one input data, the first learning model 50 outputs the first output data, the second learning model 60 outputs the second output data, and then the user using the information processing system 1 performs an operation to execute the first update process of the learning model on the information processing system 1.
[0082] When the first update process of the learning model is started and when the process of step S80 described later ends, the acquisition unit 10 acquires the first output data for one input data, the second output data for one input data, and the correct answer data for one input data (step S10).
[0083] If the acquisition unit 10 acquires the first output data, the second output data, and the correct answer data, the calculation unit 20 calculates the first difference data corresponding to the difference between the first output data and the correct answer data using (Equation 1) and calculates the second difference data corresponding to the difference between the second output data and the correct answer data using (Equation 2) based on the acquired first output data, second output data, and correct answer data (step S20).
[0084] If the first difference data and the second difference data are calculated, the learning unit 30 calculates difference data using (Equation 3) based on the first difference data and the second difference data (step S30). Then, the learning unit 30 checks whether the calculated difference data is larger than a preset specified threshold (step S40).
[0085] In the process of step S40, when the calculated difference data is larger than the preset specified threshold (step S40: Yes), the learning unit 30 updates the weights using the error backpropagation method with the calculated difference data as the error, so that the first learning model 50 relearns (step S50). Then, the first learning model 50 after relearning updates the first output data for one input data (step S60).
[0086] If the first output data is updated, the transformation unit 40 transforms the first learning model 50 after relearning into the second learning model 60 (step S70). Then, the transformed second learning model 60 updates the second output data for one input data (step S80).
[0087] If the process of step S80 ends, the information processing system 1 advances to the process of step S10 again and repeats the processes after step S10.
[0088] In the process of step S40, when the calculated difference data is not greater than a predetermined threshold value (step S40: No), the information processing system 1 ends the first update process of the learning model.
[0089] [1-3. Consideration]
[0090] As described above, according to the information processing system 1, the first learning model 50 uses, in addition to the first difference data, the second difference data based on the second learning model 60 for relearning. Further, in the relearning of the first learning model 50, it is not necessary to reflect the content of the transformation process from the first learning model 50 to the second learning model 60. Therefore, according to the information processing system 1, even if the content of the transformation process from the first learning model 50 to the second learning model 60 is unclear, it is possible to reduce the difference between the output data of the first learning model 50 and the output data of the second learning model 60 that occurs when the same data is input to the first learning model 50 and the second learning model 60.
[0091] (Second Embodiment)
[0092] Next, the information processing system related to the second embodiment will be described. In addition, the description of the same structure as that of the first embodiment will be omitted.
[0093] [2-1. Structure of Information Processing System]
[0094] Figure 5 It is a block diagram showing the structure of the information processing system 1A related to the second embodiment.
[0095] As Figure 5 shown, the information processing system 1A further includes a determination unit 70 in addition to the acquisition unit 10, the calculation unit 20, the learning unit 30, the transformation unit 40, the first learning model 50, and the second learning model 60.
[0096] The determination unit 70 as Figure 6As shown, the third differential data is generated using the first output data and the second output data. Specifically, the determination unit 70 determines whether the first output data and the second output data are true data respectively. And the determination unit 70 generates the third differential data based on the determination results. For example, the determination unit 70 is the Discriminator in a GAN (Generative Adversarial Network). The determination unit 70 generates the first probability that the first output data is true data (or the probability that it is false data) and the second probability that the second output data is true data (or the probability that it is false data) as the determination results. And the determination unit 70 generates the third differential data using the first probability and the second probability. For example, the third differential data is calculated according to the following formula (Formula 8).
[0097] Loss3 = log(D(the first output data)) + log(1 - D(the second output data))... (Formula 8)
[0098] Here, D represents the Discriminator. In the above formula, the determination unit 70 (i.e., the Discriminator) generates the probabilities that the first output data and the second output data are true data.
[0099] The learning unit 30 uses the first differential data and the third differential data to relearn the first learning model 50.
[0100] The learning unit 30 calculates the differential data (i.e., LOSS) represented by the following (Formula 9) according to the first differential data and the third differential data.
[0101] LOSS = λ4 * Loss1 + λ5 * Loss3... (Formula 9)
[0102] Here, λ4 and λ5 are values for weighting the first differential data and the third differential data in the calculation of the differential data.
[0103] The learning unit 30 updates the weights by using the error backpropagation method with the calculated differential data as the error, and relearns the first learning model 50.
[0104] [2 - 2. Operations of the Information Processing System]
[0105] Hereinafter, the processing performed by the information processing system 1A with the above structure will be described. Figure 7 It is a flowchart of the second update process of the learning model.
[0106] First, the acquisition unit 10 acquires the first output data for one input data, the second output data for one input data, and the correct answer data for one input data (step S10).
[0107] If the acquisition unit 10 acquires the first output data and the second output data, the determination unit 70 determines the authenticity of the acquired first output data and second output data (step S110). For example, the determination unit 70 calculates the probability that the first output data is true data and the probability that the second output data is true data.
[0108] The determination unit 70 calculates the third difference data based on the determination result (step S120). For example, the determination unit 70 calculates the third difference data using the above (Equation 8).
[0109] The calculation unit 20 calculates the first difference data based on the acquired first output data and the correct answer data (step S130).
[0110] The learning unit 30 calculates the difference data based on the calculated first difference data and third difference data (step S140). For example, the learning unit 30 calculates the difference data using the above (Equation 9).
[0111] The subsequent processing is substantially the same as the processing of the first embodiment, so the description is omitted.
[0112] [2 - 3. Consideration]
[0113] In this way, according to the information processing system 1A of the second embodiment, the first learning model 50 performs relearning using, in addition to the first difference data, the third difference data for making the first output data and the second output data close. By performing the learning of the first learning model 50 in such a way that the second output data is close to the first output data, the recognition performance of the second learning model 60 can be made close to that of the first learning model 50. Therefore, even if the content of the transformation process from the first learning model 50 to the second learning model 60 is unclear, it is possible to reduce the difference between the output data of the first learning model 50 and the output data of the second learning model 60 that occurs when the same data is input to the first learning model 50 and the second learning model 60.
[0114] Furthermore, in the relearning of the first learning model 50, by also using the first difference data, it is possible to suppress the performance degradation of the first learning model 50 (i.e., the performance degradation of the second learning model 60) while making the recognition performance of the second learning model 60 close to the recognition performance of the first learning model 60.
[0115] (Other Embodiments)
[0116] As described above, the information processing system for one or more technical solutions related to the present invention has been described based on the first embodiment and the second embodiment. However, the present invention is not limited to these embodiments. As long as it does not depart from the gist of the present invention, various modified forms conceived by those skilled in the art applied to the present embodiment, or forms constructed by combining the constituent elements of different embodiments can also be included within the scope of one or more technical solutions of the present invention.
[0117] (1) In the first embodiment, it has been described that the first learning model 50 is a learning model that processes using floating-point type variables and the second learning model 60 is a learning model that processes using integer type variables. However, if the second learning model 60 is a learning model obtained by transformation of the first learning model 50, it is not necessarily limited to the example where the first learning model 50 is a learning model that processes using floating-point type variables and the second learning model 60 is a learning model that processes using integer type variables.
[0118] As an example, it may also be that the first learning model 50 is a learning model that processes the pixel values of each pixel in the image to be processed as quantized 8-bit RGB data, and the second learning model 60 is a learning model that processes the pixel values of each pixel in the image to be processed as quantized 4-bit RGB data. In this case, even in a system that cannot process an image whose pixel values are composed of 8-bit RGB data due to, for example, restrictions on the data transfer rate of the data to be processed or restrictions on the storage capacity for storing the data to be processed, but can process an image whose pixel values are composed of 4-bit RGB data, the second learning model 60 can be utilized. In addition, in this case, for example, in a system that places more importance on reducing power consumption associated with operations than on the accuracy of operations, there are cases where it is more appropriate to utilize the second learning model 60 than the first learning model 50.
[0119] In addition, as another example, it may also be that the first learning model 50 is a learning model that processes using 32-bit floating-point type variables and the second learning model 60 is a learning model that processes using 16-bit floating-point type variables. In this case, even in a system that cannot process 32-bit floating-point type variables but can process 16-bit floating-point type variables, the second learning model 60 can be utilized. In addition, in this case, for example, in a system that places more importance on reducing power consumption associated with operations than on the accuracy of operations, there are cases where it is more appropriate to utilize the second learning model 60 than the first learning model 50.
[0120] In addition, as another example, the first learning model 50 may be a learning model that processes the pixel values of each pixel in the image to be processed as data in the RGB color space, and the second learning model 60 may be a learning model that processes the pixel values of each pixel in the image to be processed as data in the YCbCr color space. In this case, even in a system where, for example, the pixel values of each pixel in the image to be processed cannot be processed as data in the RGB color space but can be processed as data in the YCbCr color space, the second learning model 60 can be utilized.
[0121] (2) Part or all of the components included in the information processing system 1 may also be configured by one system LSI (Large Scale Integration). A system LSI is a super-multi-functional LSI manufactured by integrating multiple structural parts onto one chip. Specifically, it is a computer system including a microprocessor, ROM (Read Only Memory), RAM (Random Access Memory), etc. Computer programs are stored in the ROM. By operating the microprocessor according to the computer programs, the system LSI realizes its functions.
[0122] Here, it is assumed to be a system LSI, but depending on the difference in integration level, there are also cases called IC, LSI, super LSI, and ultra-large scale LSI. In addition, the method of integrating circuits is not limited to LSI, and it can also be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after manufacturing the LSI, or a reconfigurable processor that can reconstruct the connection or setting of circuit units inside the LSI can also be used.
[0123] Furthermore, if an integrated circuit technology that replaces the LSI appears due to the progress of semiconductor technology or other derived technologies, it is of course possible to use such technology for the integration of functional blocks. It may be the application of biotechnology, etc.
[0124] (3) One aspect of the present invention is not only such an information processing system, but may also be an information processing method that takes the characteristic structural parts included in the information processing system as steps. In addition, one aspect of the present invention may also be a computer program that causes a computer to execute each of the characteristic steps included in the information processing method. In addition, one aspect of the present invention may also be a computer-readable non-transitory recording medium that records such a computer program.
[0125] The present invention can be widely used in systems for information processing that enable learning models to learn.
Claims
1. An information processing method, characterized in that, Perform the following processing using a computer: Obtain a second learning model by transforming a first learning model, where the first learning model takes an image as input data and outputs coordinates representing the position of a subject included in the image and the reliability of the subject; Obtain the first output data, which is the output data of the first learning model for the image, the correct answer data for the image, and the second output data, which is the output data of the second learning model for the image; Calculate first difference data corresponding to the difference between the first output data and the correct answer data, and second difference data corresponding to the difference between the second output data and the correct answer data; Use the first difference data and the second difference data to perform only relearning of the first learning model; Update the second learning model by transforming the re-learned first learning model; 2. The information processing method according to claim 1, characterized in that, In the above learning, weight the first difference data and the second difference data; 3. The information processing method according to claim 2, characterized in that, In the above weighting, make the weight of the first difference data larger than the weight of the second difference data; 4. The information processing method according to claim 1, characterized in that, In the above learning, also use the difference between the first difference data and the second difference data; 5. The information processing method according to claim 4, characterized in that, In the above learning, weight the first difference data, the second difference data, and the difference between the first difference data and the second difference data; 6. The information processing method according to any one of claims 1 to 5, characterized in that, The first learning model and the second learning model are neural network type learning models; 7. An information processing system, characterized in that, Comprising: A transformation unit that obtains a second learning model by transforming a first learning model, where the first learning model takes an image as input data and outputs coordinates representing the position of a subject included in the image and the reliability of the subject; An acquisition unit that acquires the first output data, which is the output data of the first learning model for the image, the correct answer data for the image, and the second output data, which is the output data of the second learning model for the image; A calculation unit that calculates first difference data corresponding to the difference between the first output data and the correct answer data, and second difference data corresponding to the difference between the second output data and the correct answer data; And A learning unit that uses the first difference data and the second difference data to perform only relearning of the first learning model, where the transformation unit updates the second learning model by transforming the re-learned first learning model.
Citation Information
Patent Citations
Neural network model compression method and apparatus, storage medium and electronic device
CN108229646A
Joint model training
US20170132528A1
Video prediction using spatially displaced convolution
US20190297326A1
Application Development Platform and Software Development Kits that Provide Comprehensive Machine Learning Services
US20220091837A1