Information processing device, information processing method, and information processing program
By using the core external degree loss function and gradually reducing the kernel size in the information processing equipment, the problem of outliers in the training data affecting the accuracy of neural networks is solved, and more efficient neural network learning is achieved.
Patent Information
- Application Number
- JP2021168294
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-10-13
AI Technical Summary
When using core externality as a loss function for neural network training, outliers exist in the training data, resulting in reduced neural network learning accuracy.
An information processing device is designed, including a acquisition unit, a learning unit and an update unit. The learning unit optimizes the weight coefficient of the neural network through the core external degree loss function and reduces the impact of abnormal data by gradually reducing the kernel size.
It effectively improves the accuracy of neural network learning and reduces the impact of abnormal data on training results.
Smart Images

Figure 0007678503000014 
Figure 0007678503000015 
Figure 0007678503000016
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]
[0002] In training a neural network, parameters such as weighting coefficients can be optimized using a predetermined loss function.
[0003] For example, Non-Patent Document 1 discloses a method of optimizing parameters such as weighting coefficients by using correntropy as a loss function and maximizing the correntropy.
[0004] In the method disclosed in Non-Patent Document 1, the kernel size, which is one of the parameters that define the correntropy, is selected as the maximum value among the errors between the output of the neural network for the input data of the training data and the ground truth data linked to the input data. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Leandro LS Linhares, et al. “Fuzzy Wavelet Neural Network Using a Correntropy Criterion for Nonlinear System Identification”, Proc. of Mathematical Problems in Engineering,pp. 1-12 (2015) Summary of the Invention [Problem to be solved by the invention]
[0006] Incidentally, the training data used for training a neural network may contain data including abnormal values (abnormal data).
[0007] In the method disclosed in Non-Patent Document 1, since the kernel size is large, even if the training data includes abnormal data, the contribution of the abnormal data to the correntropy cannot be ignored.
[0008] A neural network obtained by using such correntropy as a loss function may have reduced input / output accuracy due to the contribution of abnormal data.
[0009] An object of the present invention is to provide an information processing device capable of executing neural network learning with high accuracy. [Means for solving the problem]
[0010] In order to achieve the above object, an information processing device includes an acquisition unit that acquires training data including a plurality of data including input data and correct answer data linked to the input data, and a learning unit that executes training of the neural network so that a loss function based on a correntropy using an output of the neural network for the input data, an error between the correct answer data linked to the input data, and a kernel size satisfies a predetermined condition, the learning unit having a first update unit that updates the kernel size so that the kernel size gradually decreases in the training of the neural network, and a second update unit that updates a weight coefficient of the neural network so that the value of the loss function increases after the first update unit updates the kernel size. Other features of the present invention will be made clear by the description of this specification. Effect of the Invention
[0011] According to the present invention, it is possible to provide an information processing device capable of executing neural network learning with high accuracy. [Brief description of the drawings]
[0012] [Figure 1] FIG. 2 illustrates an example of a hardware configuration of an information processing device. [Diagram 2] FIG. 13 is a diagram illustrating an example of learning data. [Diagram 3] FIG. 2 is a diagram illustrating an example of a configuration of a neural network model. [Figure 4] FIG. 1 is a diagram for explaining the transition of R(t) during neural network learning. [Diagram 5] FIG. 13 is a diagram illustrating an example of an error distribution in the initial stage of learning of a neural network. [Figure 6] FIG. 13 is a diagram illustrating an example of error distribution at a stage where learning of a neural network has progressed. [Figure 7] FIG. 2 is a diagram illustrating an example of functional blocks realized in the information processing device. [Figure 8] 11 is a flowchart outlining a process up to outputting a weighting coefficient. [Figure 9] 11 is a flowchart illustrating a process in which a learning unit executes learning of a neural network. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] == Implementation form == <<Information processing devices>> The information processing device 1 is a device that executes neural network learning based on predetermined learning data. Below, the hardware configuration of the information processing device 1, learning data, neural network, loss function, neural network learning, an example of neural network learning, and functional blocks of the information processing device 1 will be described in this order.
[0014] <Hardware configuration> 1 is a diagram showing an example of hardware of an information processing device 1. The illustrated information processing device 1 includes a processor 100, a main memory device 101, an auxiliary memory device 102, an input device 103, an output device 104, and a communication device 105. Note that the illustrated information processing device 1 may be realized, in whole or in part, by using virtual information processing resources provided by using a virtualization technique, a process space separation technique, or the like, such as a virtual server provided by a cloud system. In addition, all or in part of the functions provided by the information processing device 1 may be realized by a service provided by a cloud system via an API (Application Programming Interface), for example.
[0015] [Processor 100] The processor 100 is configured using, for example, a CPU (Central Processing Unit) and the like.
[0016] [Main memory 101] The main memory device 101 is a device for storing programs and data, and is, for example, a Read Only Memory (ROM) or a Random Access Memory (RAM).
[0017] [Auxiliary storage device 102] The auxiliary storage device 102 is, for example, a hard disk drive, a storage area of a cloud server, etc. Programs and data can be read into the auxiliary storage device 102 via a recording medium reading device or a communication device 105. The programs and data stored (memorized) in the auxiliary storage device 102 are read into the main storage device 101 as needed.
[0018] [Input device] The input device 103 is an interface that accepts input from the outside, and is, for example, a keyboard, a mouse, a touch panel, or the like.
[0019] [Output device] The output device 104 is an interface that outputs various information such as the process progress, the process results, etc. The output device 104 is, for example, a display device (such as a liquid crystal monitor) that visualizes the various information described above, a device that converts the various information described above into voice (a voice output device (such as a speaker)), or a device that converts the various information described above into text (such as a printer).
[0020] The input device 103 and the output device 104 constitute a user interface that receives information from the user and presents information.
[0021] [Communication Device] The communication device 105 is a device that realizes communication with other devices. The communication device 105 is a wired or wireless communication interface that realizes communication with other devices via a communication network NW such as the Internet.
[0022] In the information processing device 1, for example, an operating system, a file system, and the like may be installed.
[0023] Each function of the information processing device 1 is realized by a processor 100 of the information processing device 1 reading and executing a program stored in a main memory device 101, or by hardware constituting the information processing device 1. The information processing device 1 stores various information (data), for example, as a table of a database or a file managed by a file system.
[0024] <Learning data> Fig. 2 is a diagram showing an example of training data used for training a neural network. The training data 2 is stored in the auxiliary storage device 102. The training data 2 includes a plurality of data p, and in this example, includes 10,000 pieces of data p. One piece of data p is arranged in one row of Fig. 2.
[0025] The learning data 2 in this embodiment is data in which the input data 1 to 3 are associated with the correct answer data. Each of the multiple data p has the input data 1 to 3 and the correct answer data linked to the input data 1 to 3.
[0026] Data for learning the behavior of a showcase can be given as a specific example of the learning data 2. In this case, for example, the input data 1, 2, and 3 can be the frequency of a refrigerator (not shown), the opening degree of an expansion valve, and the indoor air temperature, respectively, and the correct answer data can be the temperature inside the showcase.
[0027] The learning data is not limited to showcases, and other examples include data for learning about demand for electricity, abnormality detection in factory equipment, weather, and the like.
[0028] <Neural Network> 3 is a diagram illustrating the configuration of a model of a neural network NN according to this embodiment. The neural network NN has an input layer L1, an intermediate layer L2, and an output layer L3.
[0029] The input layer L1 has three elements N, the intermediate layer L2 has many elements N, and the output layer L3 has one element N. In addition, between the two elements N, there is a weighting factor W ij are assigned to the element N (the subscripts i and j are numbers assigned to the element N, but in the following they may be omitted as appropriate).
[0030] The input layer L1 accepts input data and passes it to the hidden layer L2. The hidden layer L2 receives the input data and calculates the output based on a predetermined activation function for each element N. The hidden layer L2 extracts features from the input data and passes them to the output layer L3. The output layer L3 outputs output data based on the features output from the hidden layer L2.
[0031] The neural network NN of this embodiment is merely an example, and the number of elements in the input layer L1, the number of elements in the intermediate layer L2, the number of elements in the intermediate layer L2, the number of elements in the output layer L3, etc. are not particularly limited.
[0032] <Loss function> In the learning of the neural network of this embodiment, correntropy is used as the loss function. C is defined by the following formulas 1 to 3.
number
number
number
[0033] Equation 1 is the correntropy E C In Equation 1, “p” is a number assigned to each of the multiple data p included in the training data 2. “P” is the total number of multiple data p included in the training data 2.
[0034] Equation 2 is an equation that defines E(p) on the right side of Equation 1. In Equation 2, “σ” is the kernel size. Although details will be described later, the kernel size σ is variable during the learning process of the neural network NN.
[0035] Formula 3 is the E on the right side of Formula 2 r This is an equation that defines (p). In Equation 2, "x(p)" on the right side is the output value of the neural network NN for the input data of data p. Also, "t(p)" on the right side is the value of the correct answer data linked to the input data of data p.
[0036] That is, E defined in Equation 3 r(p) is the error between the output of the neural network NN for the input data and the correct data associated with the input data. In the following, E r (p) is sometimes simply referred to as "error."
[0037] <Neural network training> Neural network training is based on the correntropy E C This is done by searching for the weight coefficients of the neural network NN that maximizes
[0038] The training of the neural network of this embodiment includes a step of updating the weight coefficients and a step of updating the kernel size σ. These steps are then alternately repeated multiple times to obtain the correntropy E C This will be explained in detail below.
[0039] Before the neural network learning is performed, the initial values of the weight coefficients of the neural network NN are ij (1)(w ij is element N i and element N j (weighting coefficient between the first and second inputs) and the initial value σ(1) of the kernel size σ are set to predetermined values.
[0040] (1) Update of weighting coefficients Below, in (1-1) to (1-3), a method for updating the weighting coefficient based on one piece of data p will be described.
[0041] (1-1) Calculate the output value for data p When updating the weighting coefficients, first, the output value of the neural network NN for the data p (x(p) shown on the right side of Equation 3) is calculated.
[0042] At this time, it is assumed that the weight coefficients based on p-1 pieces of data (data 1 to p-1) among the multiple data p (p=1 to P) have already been updated. The weight coefficients of the neural network NN at this time are expressed as w(p)ij Let us assume that.
[0043] In other words, the output x(p) calculated at this stage is the weighting coefficient w ij (p) is the output of the neural network NN.
[0044] (1-2) Calculate the error for data p Once the output value for data p is calculated, the error for data p is then calculated. The error for data p is E r (p). In other words, the error E for data p r (p) is the difference between the correct data t(p) for data p and x(p) calculated in (1-1) above.
[0045] (1-3) Update the weighting coefficient based on data p When the error for data p is calculated, the weighting coefficients are updated based on the data p. Here, the weighting coefficients before the update are expressed as w ij (p), and the updated weight coefficient is w ij Let it be (p+1).
[0046] Updated weight coefficient w ij (p+1) is expressed by the following formula 4.
number
[0047] Here, the second term on the right hand side is the correntropy E C The weight coefficient before updating w ij This is a value obtained by multiplying the value differentiated by (p) by the learning rate η. The learning rate η is set to a predetermined positive value in advance.
[0048] In the above procedure, the updated weight coefficients w for all subscripts i and j are ij (p+1) is obtained.
[0049] Here, the correntropy E Cdepends on the output x(p) on the right side of Equation 3. The output x(p) is determined by the weighting coefficient w ij (p). Therefore, the correntropy E C is the weighting factor w via the output x(p). ij Depends on (p).
[0050] From this, the weighting factor w ij (p) to w ij When the weight coefficient is updated to (p+1), the correntropy after the weight coefficient is updated is E C (w ij (p+1)), and the correntropy before the weight coefficient is updated is E C (w ij (p)).
[0051] Therefore, the correntropy E after the weight coefficients are updated C (w ij (p+1)) is the fluctuation amount of the weighting coefficient w ij (p+1)-w ij This is of first order in (p) and is formally expressed as Equation 5 below.
number
[0052] In addition, in the left side and the first term on the right side of the formula 5, the subscripts of the weighting coefficients are omitted. Here, by using the formula 2, the second term on the right side of the formula 5 is transformed to obtain the following formula 6.
number
[0053] Formula 6 plus w in formula 4 ij Substituting (p+1), we obtain the following equation 7.
number
[0054] Since the second term on the right side of Equation 7 is a positive value, the correntropy EC (w ij (p+1)) is the correntropy E before the weight coefficients are updated. C (w ij (p) is greater than
[0055] Therefore, the weight coefficient w updated by the above method ij According to (p+1), the correntropy E C increases as a first order factor of η.
[0056] In the above, in (1-1) to (1-3), the method for updating the weight coefficient based on one piece of data p has been explained. In the same manner as above, when the weight coefficient is updated based on each of all pieces of data p (p=1 to P), the final weight coefficient w(P+1) is obtained.
[0057] Hereinafter, the process in which the weight coefficients are updated based on each of all data p (p=1 to P) and the weight coefficient w(P+1) is obtained is referred to as the "coefficient update process." In this embodiment, when the coefficient update process is performed, the kernel size σ is updated. Hereinafter, the number of times the kernel size σ is updated is referred to as the "number of learning times." Also, an increase in the number of learning times may be referred to as "learning progresses."
[0058] (2) Update the kernel size σ The coefficient update process (1) is executed, and the kernel size σ is updated while the weight coefficient w(P+1) obtained in (1) is fixed. Below, the method of updating the kernel size σ will be described in (2-1) to (2-3).
[0059] (2-1) Error calculation First, the error for each of all data p (p=1 to P) is calculated. The error for data p is E r (p). The neural network NN used in this case is a neural network NN having a weight coefficient w(P+1).
[0060] (2-2) Setting R(t) R(t) is a function of the number of times of learning t, and is a function indicating a ratio serving as an index for updating the kernel size σ (R(t) corresponds to a "predetermined ratio"). The use of the ratio R(t) will be described in detail later, but in this embodiment, the ratio R(t) is defined by the following formula 8.
number
[0061] Here, T max is the maximum number of learning times, and a predetermined value is set. max There is no particular restriction on the value of as long as it is an integer of 1 or more.
[0062] Also, R max The number of learning times t is the maximum number of learning times T max This is the value of the ratio R(t) when the value reaches a certain value, and is set to a predetermined value. max The value has a dimension of a percentage and is selected from the range of values greater than 0 and less than 1.0 (or greater than 0% and less than 100%).
[0063] As an example, T max is 1000, R max is set to 0.1, the ratio R(1) when the learning count t is 1 is 0.0001, the ratio R(500) when the learning count t is 500 is 0.05, and the ratio R(1000) when the learning count t is 1000 is 0.1.
[0064] FIG. 4 is a diagram for explaining the transition of the ratio R(t) in the learning of a neural network. The horizontal axis is the number of learning times t, and the vertical axis is the ratio R(t). Here, the positions corresponding to the initial stage of learning (t=t1), the advanced stage of learning (t=t2), and the further advanced stage of learning (t=t3) are indicated by black circles (t1 <t2<t3)。
[0065] (2-3) Update kernel size σ First, the reason for updating the kernel size σ will be explained with reference to FIG. 5. FIG. 5 is a diagram for explaining an example of distribution of errors in the initial stage of learning of a neural network (for example, t=t1 in FIG. 4). The horizontal axis represents the error value E r The vertical axis is E in Equation 2. r (p) as a variable E r In other words, the vertical axis is the correntropy E of each of the multiple errors shown in the figure. C This means the contribution to (Formula 1). In addition, in Fig. 5, the boundary of the range in which the absolute value of the error is 2√2×σ or less is shown by a dashed line.
[0066] Among the errors shown in Fig. 5, errors for normal data are indicated by circles, and errors for abnormal data are indicated by stars. Here, "abnormal data" refers to data that includes abnormal values, and "normal data" refers to data other than abnormal data.
[0067] "Data containing abnormal values" is data in which the value of the correct answer data significantly deviates from the expected value, for example, when taking into account the value of the input data and the behavior of the learning target device at the time when data p is acquired.
[0068] In this embodiment, the number of pieces of data p included in the learning data 2 is 10000, but for convenience of explanation, only a part of them is shown. The following describes an initial stage of learning and a stage in which learning has progressed.
[0069] In the initial stage of learning, as can be seen from FIG. 5, the error for the abnormal data pa is smaller than the error for the normal data pb. This is because, as can be seen from Equations 1 to 3, the correntropy E C The contribution to the normal data pb is the correntropy E C This means that the contribution to
[0070] When the weighting coefficients are updated in the state shown in FIG. 5 in (1), it is considered that the abnormal data pa contributes more to the learning of the neural network than the normal data pb.
[0071] Alternatively, in the early stages of learning, if the correntropy E C When the kernel size σ is set small enough that the contribution to pb can be ignored, the correntropy E C The contribution to is negligible. In such a case, the accuracy of the neural network NN may decrease.
[0072] Therefore, it is preferable that the kernel size σ has a certain size or more in the early stage of learning, and is updated to be gradually smaller as learning progresses. In this embodiment, the kernel size σ is updated based on the error for each of all the above-mentioned data p (p=1 to P) and the ratio R(t).
[0073] Specifically, the kernel size σ is updated so that it is smaller than any error that falls within the range of the ratio R(t) from the largest error among errors for each of all data p (p=1 to P). When the number of learning times is t, the updated kernel size σ(t) is defined by the following formulas 9 and 10.
[0074]
number
number
number
[0075] Equation 9 shows the updated kernel size σ(t). ro and select(t) are defined in Equations 10 and 11.
[0076] Formula 10 is an equation that defines select(t) on the right side of Formula 9. The roundup on the right side means that the value in the parentheses is rounded up to the nearest integer. In other words, select(t) is an integer.
[0077] Moreover, from the definition of the index R(t) in Equation 5, it can be understood that when the number of learning times t increases by one, select(t) either maintains that value or increases from that value.
[0078] Equation 11 is the error E calculated in (2-1). r This is the inequality when (p)(p=1~P) are rearranged in descending order. In other words, E ro (1) is the error E r (p) (p=1 to P) is the maximum value, and E ro (P) is the error E r (p) is the minimum value among (p=1 to P).
[0079] Here, the correntropy E C From the definition of (Equation 2) and the definition of the updated kernel size σ(t) (Equation 9), the error E ro (1)~E ro The data p corresponding to (select(t)) is the error E ro (select(t)+1)~E ro Compared to the data p corresponding to (P), the correntropy E C The contribution to is small.
[0080] The smaller the contribution of data p to the correntropy, the smaller the contribution to learning. Here, "contribution to learning" refers to the fluctuation of the updated weighting coefficient w(p+1) relative to the pre-update weighting coefficient w(p) in the update of the weighting coefficient based on data p in (1-3) above. This fluctuation corresponds to the absolute value of the second term on the right-hand side of Equation 4. The smaller the contribution of data p to learning, the smaller the fluctuation of the updated weighting coefficient w(p+1).
[0081] Error E ro (1)~E roThe data p corresponding to (select(t)) may be referred to as "rejected data" below. In addition, data p other than the rejected data p (error E ro (select(t)+1)~E ro The data p) corresponding to (P) may be referred to as the "data to be considered" or the like.
[0082] By updating the kernel size σ as in Equations 9 to 11, the kernel size σ gradually decreases as the number of learning times t increases. Regarding the kernel size σ, "gradually decreasing" includes not only monotonically decreasing, but also monotonically increasing the moving average of a predetermined number of learning times t. Therefore, "gradually decreasing" does not necessarily exclude the case where the kernel size σ increases as the number of learning times t increases by one.
[0083] In particular, when the number of learning iterations is 1 and when the number of learning iterations is the maximum number of learning iterations T max The updated kernel size σ in this case is shown in the following Equations 12 and 13, respectively.
[0084]
number
number
[0085] It is sufficient that the kernel size σ gradually decreases with an increase in the number of learning times t. Therefore, the method of updating the kernel size σ is not limited to the above example.
[0086] For example, the index R(t) in the formula 8 set in (2-2) is just an example and is not limited to this function. In this example, the index R(t) is a linear function, but it may be another function that monotonically increases with the increase in the number of learning times t.
[0087] Alternatively, the kernel size σ does not have to be updated using the index R(t). For example, a kernel function σ(t) that depends on the number of times of learning t may be set in advance.
[0088] <Example of neural network learning> 6 is a diagram for explaining an example of error distribution at a stage where neural network learning has progressed, for example, at the stage of learning iteration t3 in FIG. 4. The horizontal and vertical axes are the same as in FIG.
[0089] The initial stage of learning (t=t1) and the advanced stage of learning (t=t3) will be described below with reference to Figures 4 to 6. In Figures 5 and 6, the dashed lines indicating the range of errors whose absolute value is 2√2×σ or less mean that data p shown within the range is classified as data to be considered, and data p shown outside the range is classified as data to be considered.
[0090] [Early learning stage] First, when the number of learning iterations t is 0, there is no correlation between the magnitude of the error for data p and whether data p is normal or abnormal. Therefore, in the early stages of learning, this correlation is thought to be almost nonexistent.
[0091] In this example, as can be seen from FIG. 5, the error for the abnormal data pa is smaller than the error for the normal data pb.
[0092] As shown in Fig. 4, in the early stage of learning, the index R(t) is close to zero. Accordingly, in the early stage of learning, most of the data p is classified as data to be considered. In particular, both the abnormal data pa and the normal data pb are classified as data to be considered.
[0093] The ratio of the number of normal data in the considered data to the number of considered data will be referred to as the "purity of the considered data" or simply as "purity" hereinafter.
[0094] [Advanced learning stage] In this example, as shown in Fig. 6, the error for the abnormal data pa deviates from zero compared to the initial stage of learning, and the error for the normal data pb approaches zero compared to the initial stage of learning.
[0095] As shown in Figure 4, the index R(t) increases as the learning progresses. Therefore, the kernel size σ becomes smaller than in the early stages of learning (Figure 6). And, the number of data p classified as data to be eliminated gradually increases. In particular, the abnormal data pa is classified as data to be eliminated.
[0096] As described above, in the learning of the neural network of this embodiment, as the learning progresses, the error for normal data tends to become smaller, and the error for abnormal data tends to become larger.
[0097] One of the reasons for this is that the proportion of normal data in the learning data 2 is generally greater than the proportion of abnormal data (Reason 1). In this case, the total contribution of normal data to the correntropy is greater than the total contribution of abnormal data to the correntropy. As a result, the error for abnormal data is less likely to be reduced by updating the weighting coefficient w(p), and may even increase.
[0098] Another reason is that the kernel size σ gradually becomes smaller as learning progresses (reason 2). Abnormal data that has progressed to have a larger error due to reason 1 above is more likely to be classified as data to be eliminated, and its contribution to learning becomes smaller. If the contribution of abnormal data to learning is small, the error for abnormal data is likely to become even larger due to the update of the weighting coefficient w(p).
[0099] [summary] According to this method, in the early stages of learning, most of the data p are classified as data to be considered, regardless of whether they are normal data or abnormal data. Then, as learning progresses, the number of data p classified as data to be excluded gradually increases.
[0100] The data to be eliminated is expected to be dominated by abnormal data. In other words, it is expected that the purity of the data taken into account will increase as learning progresses. Therefore, it is expected that learning will be performed based on data taken into account with an increased purity as learning progresses.
[0101] In addition, in the method of this embodiment, the initial value σ(1) of the kernel size σ is set so that all data p are classified as data to be considered in the early stage of learning. According to such a method, it is possible to prevent a part of normal data included in the learning data 2 from being classified as data to be excluded in the early stage of learning.
[0102] In the method of this embodiment, the number of data p to be removed is determined based on the index R(t) (Equations 9 to 11). In this example, the index R(t) has an upper limit R max The upper limit R max is a value that can be preset by the user.
[0103] For example, if the user has empirical knowledge about the percentage of abnormal data included in training data 2, the user can use that knowledge to max For example, if the user has empirical knowledge that learning data 2 contains about 10 percent abnormal data, the user can set R max Just set it to 0.1.
[0104] As a result, in the final stage of learning, about 10 percent of the data p in the learning data 2 is classified as data to be removed, and the remaining about 90 percent of the data p is classified as data to be considered. At this time, it is expected that about 10 percent of the abnormal data will be classified as data to be removed. In other words, it is expected that the data to be considered will contain almost no abnormal data.
[0105] <Function block> 7 is a diagram showing functional blocks of the information processing device 1 of this embodiment. In the information processing device 1 of this embodiment, an acquisition unit 110, a setting unit 111, a learning unit 112, and an output unit 113 are realized by the processor 100 executing a predetermined program.
[0106] [Acquisition unit 110] The acquiring unit 110 acquires learning data. Learning data 2 in FIG.
[0107] [Settings section 111] The setting unit 111 sets an initial value w(1) of the weighting coefficient and an initial value σ(1) of the kernel size. The initial value w(1) of the weighting coefficient is set using, for example, a random number.
[0108] The initial value σ(1) of the kernel size may be the kernel size σ(t) of the above-mentioned formula 9 calculated based on the neural network NN to which the initial value w(1) of the weighting coefficient is applied, with select(t) set to 1. In other words, the initial value σ(1) of the kernel size may be set so that all data p (p=1 to P) are classified as data to be considered.
[0109] [Learning Section 112] The learning unit 112 executes learning of the neural network. Specifically, the learning unit 112 calculates a correntropy E using an error between the output of the neural network NN for the input data and the correct answer data associated with the input data, and a kernel size σ. CThe neural network is trained so that the loss function obtained by the above equation satisfies a predetermined condition.
[0110] The above-mentioned specified conditions are the correntropy E C This is the condition for maximizing. The details of the neural network training are as explained in "Training of Neural Networks" above.
[0111] Depending on the definition of the loss function, the above-mentioned predetermined condition may be a condition that minimizes the loss function.
[0112] The learning unit 112 includes a calculation unit 112a, a calculation unit 112b, an update unit 112c (corresponding to a second update unit), a judgment unit 112d, a setting unit 112e, an update unit 112f (corresponding to a first update unit), and a judgment unit 112g. Each of these will be described below.
[0113] (Calculation section 112a) The calculation unit 112a calculates the output value of the neural network NN for each of the multiple data p ((1-1) Acquisition of output for data p).
[0114] (Calculation section 112b) The calculation unit 112b calculates a plurality of errors for each of the plurality of data p ((1-2) Calculation of error for data p, (2-1) Calculation of error).
[0115] (Updated part 112c) The update unit 112c updates the correntropy E C The weighting coefficient w(p) of the neural network NN is updated so as to increase the value of the loss function by ((1-3) Updating the weighting coefficient based on the data p). This process is executed after the setting unit 111 sets the initial value σ(1) of the kernel size and after the updating unit 112f, which will be described later, updates the kernel size σ.
[0116] (Judgment unit 112d) The determining unit 112d determines whether or not updating of the weighting factor w(p) based on all data p (p=1 to P) has been completed in one learning session.
[0117] (Setting section 112e) The setting unit 112e sets the index R(t) ((2-2) Setting of index R(t)). At this time, the setting unit 112e sets the index R(t) so that the index R(t) increases with respect to the number of times the update unit 112f updates the kernel size σ. Note that the number of times the update unit 112f updates the kernel size σ is equal to the number of times of learning t.
[0118] In this embodiment, the setting unit 112e sets the index R(t) such that the index R(t) and the number of times the updating unit 112f has updated the kernel size σ satisfy a linear function relationship (Equation 8).
[0119] In this embodiment, the setting unit 112e determines whether the index R(t) is greater than or equal to a predetermined upper limit R max An index R(t) is set in the range (Equation 8).
[0120] (Updated part 112f) The update unit 112f updates the kernel size σ so that the kernel size σ gradually decreases during neural network learning ((2-3) Update of kernel size σ).
[0121] In the present embodiment, the update unit 112f updates the kernel size σ based on a plurality of errors. Specifically, the update unit 112f updates the kernel size σ so that the kernel size σ is smaller than any of the errors that are within a range of a predetermined percentage from the largest error among the plurality of errors.
[0122] (Judgment part 112g) The determining unit 112g determines whether or not the number of times of learning t has reached the maximum number of times of learning after the updating unit 112f updates the kernel size σ.
[0123] [Output section 113] When the neural network learning process is completed, the output unit 113 outputs the final weighting coefficient w(P+1).
[0124] The information processing device 1 has the above-mentioned configuration and is therefore capable of executing the above-mentioned neural network learning.
[0125] <<Processing up to outputting weighting coefficient w(P+1)>> 8 is a flowchart illustrating an outline of the processing up to when the information processing device 1 outputs the weighting coefficient w(P+1). The processing up to when the weighting coefficient w(P+1) is output includes steps S101 to S104.
[0126] First, in step S101, the acquiring unit 110 acquires learning data. Learning data 2 in FIG.
[0127] Next, in step S102, the setting unit 111 sets an initial value w(1) of the weighting coefficient and an initial value σ(1) of the kernel size. Also in step S102, the determination unit 112g sets the number of times learning is performed t to 1.
[0128] Next, in step S103, the learning unit 112 executes learning of the neural network. The details of the flow of this process will be described later.
[0129] Next, in step S104, the output unit 113 outputs the weighting coefficient w(p) obtained in step S103.
[0130] <<Processing to execute neural network training>> Fig. 9 is a flowchart illustrating the process of learning the neural network by learning unit 112, and is a flowchart illustrating the details of step S103 in Fig. 8. The process of learning the neural network includes steps S201 to S208.
[0131] First, in step S201, the calculation unit 112a calculates the output of the neural network NN for one piece of data p.
[0132] Next, in step S202, the calculation unit 112b calculates an error for the data p based on the output calculated in step S201 ((1-2) Calculation of error for data p).
[0133] Next, in step S203, the update unit 112c updates the weighting coefficient w(p) based on the data p ((1-3) Updating the weighting coefficient based on the data p).
[0134] Next, in step S204, the determination unit 112d determines whether or not updating of the weighting coefficient w(p) based on all data p (p=1 to P) has been completed. If the determination unit 112d determines that updating has been completed (S204: Y), the information processing device 1 proceeds to step S205. If the determination unit 112d determines that updating has not been completed (S204: N), the information processing device 1 returns to step S201.
[0135] Next, in step S205, the calculation unit 112b calculates the error for each of all data p (p=1 to P) ((2-1) Calculation of Error).
[0136] Next, in step S206, the setting unit 112e sets the index R(t) ((2-2) Setting of index R(t)).
[0137] Next, in step S207, the update unit 112f updates the kernel size σ ((2-3) Update of the kernel size σ).
[0138] In step S208, the determination unit 112g determines whether the learning count t has reached the maximum learning count. If the determination unit 112g determines that the learning count has been reached (S208: Y), the information processing device 1 ends the process. If the determination unit 112g determines that the learning count has not been reached (S208: N), the information processing device 1 adds 1 to the learning count t, and returns to step S201.
[0139] By carrying out the above-mentioned procedure, it becomes possible to carry out the learning of the above-mentioned neural network.
[0140] ==Summary== As described above, the information processing device according to the embodiment includes an acquisition unit 110 that acquires learning data including a plurality of pieces of data including input data and supervised data linked to the input data, and a correntropy E using an error between the output of a neural network NN for the input data and the supervised data linked to the input data, and a kernel size σ. C The neural network learning system includes a learning unit 112 that executes learning of the neural network so that the loss function obtained by the learning unit 112 satisfies a predetermined condition, and an updating unit 112f that updates the kernel size σ so that the kernel size σ gradually decreases in the learning of the neural network, and an updating unit 112c that updates a weighting coefficient w(p) of the neural network NN so that the value of the loss function increases after the updating unit 112f updates the kernel size σ.
[0141] According to this configuration, the purity of the data taken into consideration increases as the learning progresses. Therefore, it is expected that the learning will be performed based on the data taken into consideration whose purity has increased as the learning progresses. This makes it possible to perform the learning of the neural network to obtain a neural network NN with improved input / output accuracy.
[0142] In the information processing device, the learning unit 112 includes a calculation unit 112b that calculates a plurality of errors for each of a plurality of data, and an update unit 112f updates the kernel size σ based on the plurality of errors. With this configuration, abnormal data among the learning data is easily classified as data to be excluded. This increases the purity of the data to be considered, making it possible to execute neural network learning to obtain a neural network NN with further improved input / output accuracy.
[0143] In the information processing device, the update unit 112f updates the kernel size σ so that the kernel size σ is smaller than any of the errors that are within a predetermined range from the largest error among the multiple errors. With this configuration, abnormal data among the learning data is more likely to be classified as data to be excluded. This increases the purity of the data to be considered, making it possible to execute learning of the neural network to obtain a neural network NN with further improved input / output accuracy.
[0144] In the information processing device, the learning unit 112 includes a setting unit 112e that sets a predetermined ratio so that the predetermined ratio increases with respect to the number of times the update unit 112f updates the kernel size σ. With this configuration, the number of abnormal data classified as data to be excluded from the learning data increases steadily and gradually. This further increases the purity of the data taken into consideration, making it possible to execute neural network learning to obtain a neural network NN with further improved input / output accuracy.
[0145] In the information processing device, the setting unit 112e sets the index R(t) so that the index R(t) and the number of times the update unit 112f updates the kernel size σ satisfy a linear function relationship. With this configuration, the number of abnormal data classified as data to be excluded from the learning data increases so as to approximately satisfy a linear function relationship with the number of times the kernel size σ is updated. This makes it possible to control the purity of the data to be considered according to the number of learning times t, and therefore the learning speed can be controlled.
[0146] In the information processing device, the setting unit 112e sets the index R(t) within a range where the index R(t) is up to a predetermined upper limit value. With this configuration, an upper limit value can be set for the number of data p classified as excluded data. This makes it possible to prevent normal data from being classified as excluded data, and to prevent a decrease in the purity of the data to be considered.
[0147] The information processing method of the embodiment includes a step of acquiring learning data having a plurality of pieces of data including input data and correct answer data linked to the input data, and calculating a correntropy E using an error between an output of a neural network NN for the input data and the correct answer data linked to the input data, and a variable kernel size σ. C and executing training of the neural network so that the loss function by satisfies a predetermined condition, wherein the step of executing training of the neural network includes a step of updating a kernel size σ so that the kernel size σ gradually becomes smaller in training of the neural network, and a step of updating a weighting coefficient w(p) of the neural network NN so that the value of the loss function increases after the step of updating the kernel size σ.
[0148] According to this method, the purity of the data taken into consideration increases as the learning progresses. Therefore, it is expected that the learning will be performed based on the data taken into consideration whose purity has increased as the learning progresses. This makes it possible to perform the learning of the neural network to obtain a neural network NN with improved input / output accuracy.
[0149] The information processing program according to the embodiment includes a computer, an acquisition unit 110 that acquires learning data having a plurality of pieces of data including input data and supervised data linked to the input data, and a correntropy E using an error between an output of a neural network NN for the input data and the supervised data linked to the input data, and a variable kernel size σ. C The present invention provides a learning unit 112 that executes learning of the neural network so that the loss function obtained by the above equation satisfies a predetermined condition, and the learning unit 112 is provided with an updating unit 112f that updates the kernel size σ so that the kernel size σ gradually decreases in the learning of the neural network, and an updating unit 112c that updates the weighting coefficient w(p) of the neural network NN so that the value of the loss function increases after the updating unit 112f updates the kernel size σ.
[0150] According to such a program, the purity of the data taken into consideration increases as the learning progresses. Therefore, it is expected that the learning will be performed based on the data taken into consideration whose purity has increased as the learning progresses. This makes it possible to perform the learning of the neural network to obtain a neural network NN with improved input / output accuracy.
[0151] The above-mentioned embodiment is for the purpose of facilitating understanding of the present invention, and is not intended to limit the present invention. Furthermore, the present invention can be modified or improved without departing from the spirit of the present invention, and it goes without saying that the present invention includes equivalents thereof. [Explanation of symbols]
[0152] 1: Information processing device 100: Processor 101: Main memory 102:Auxiliary storage device 103: Input device 104: Output device 105: Communication equipment 110: Acquisition Department 111: Setting section 112: Learning Department 112a: Calculation section 112b: Calculation section 112c: Update section 112d: Judgment section 112e: Setting section 112f: Update section 112g: Judgment section 113: Output section 2: Training data
Claims
1. An acquisition unit that acquires learning data including a plurality of pieces of data including input data and correct answer data linked to the input data; A learning unit that executes learning of the neural network so that a loss function based on a correntropy using an error between an output of the neural network for the input data, the correct answer data associated with the input data, and a kernel size satisfies a predetermined condition; Equipped with The learning unit is a first update unit that updates the kernel size so that the kernel size gradually decreases during training of the neural network; a second update unit that updates a weight coefficient of the neural network so that a value of the loss function increases after the first update unit updates the kernel size; having Information processing device.
2. 2. The information processing device according to claim 1, The learning unit is a calculation unit that calculates a plurality of the errors for each of the plurality of data, The first update unit updates the kernel size based on a plurality of the errors. Information processing device.
3. 3. The information processing device according to claim 1, the first update unit updates the kernel size so that the kernel size is smaller than any of the errors that are within a range of a predetermined percentage from the largest error among the plurality of errors. Information processing device.
4. 4. The information processing device according to claim 3, The learning unit is a setting unit that sets the predetermined ratio so that the predetermined ratio increases with respect to the number of times the first update unit updates the kernel size; Information processing device.
5. 5. The information processing device according to claim 4, the setting unit sets the predetermined ratio such that the predetermined ratio and the number of times the first update unit has updated the kernel size satisfy a linear function relationship. Information processing device.
6. 6. The information processing device according to claim 4, The setting unit sets the predetermined ratio within a range up to a predetermined upper limit value. Information processing device.
7. acquiring learning data including a plurality of pieces of data including input data and correct answer data linked to the input data; A step of executing learning of the neural network so that a loss function based on a correntropy using an error between an output of the neural network for the input data and the ground truth data associated with the input data and a variable kernel size satisfies a predetermined condition; Including, The step of performing training of the neural network includes: updating the kernel size so that the kernel size gradually decreases during training of the neural network; After updating the kernel size, updating weight coefficients of the neural network so that the value of the loss function increases; Including, Information processing methods.
8. On the computer, An acquisition unit that acquires learning data having a plurality of pieces of data including input data and correct answer data linked to the input data; A learning unit that executes learning of the neural network so that a loss function based on a correntropy using an error between an output of the neural network for the input data and the ground truth data linked to the input data and a variable kernel size satisfies a predetermined condition; Realize this, The learning unit, a first update unit that updates the kernel size so that the kernel size gradually decreases during training of the neural network; a second update unit that updates a weight coefficient of the neural network so that a value of the loss function increases after the first update unit updates the kernel size; To achieve this, Information processing program.